Research papers, master and doctoral theses about التدريب عبر اللغات

Cross-Lingual Training of Dense Retrievers for Document Retrieval

203 - Association for Computation Linguistics 2021 مقالة

Dense retrieval has shown great success for passage ranking in English. However, its effectiveness for non-English languages remains unexplored due to limitation in training resources. In this work, we explore different transfer techniques for docume nt ranking from English annotations to non-English languages. Our experiments reveal that zero-shot model-based transfer using mBERT improves search quality. We find that weakly-supervised target language transfer is competitive compared to generation-based target language transfer, which requires translation models.

dense retrievers retrievers for document cross-lingual training استرداد كثيف المستردون للوثيقة التدريب عبر اللغات صناعة حمض الفوسفور المزيد..

Multiple Captions Embellished Multilingual Multi-Modal Neural Machine Translation

339 - Association for Computation Linguistics 2021 مقالة

Neural machine translation based on bilingual text with limited training data suffers from lexical diversity, which lowers the rare word translation accuracy and reduces the generalizability of the translation system. In this work, we utilise the mul tiple captions from the Multi-30K dataset to increase the lexical diversity aided with the cross-lingual transfer of information among the languages in a multilingual setup. In this multilingual and multimodal setting, the inclusion of the visual features boosts the translation quality by a significant margin. Empirical study affirms that our proposed multimodal approach achieves substantial gain in terms of the automatic score and shows robustness in handling the rare word translation in the pretext of English to/from Hindi and Telugu translation tasks.

التدريب عبر اللغات embellished multilingual multi-modal multi-modal neural machine منمق متعدد اللغات متعددة الوسائط متعددة مشروط آلة العصبية صناعة حمض الفوسفور

SkoltechNLP at SemEval-2021 Task 2: Generating Cross-Lingual Training Data for the Word-in-Context Task

134 - Association for Computation Linguistics 2021 مقالة

In this paper, we present a system for the solution of the cross-lingual and multilingual word-in-context disambiguation task. Task organizers provided monolingual data in several languages, but no cross-lingual training data were available. To addre ss the lack of the officially provided cross-lingual training data, we decided to generate such data ourselves. We describe a simple yet effective approach based on machine translation and back translation of the lexical units to the original language used in the context of this shared task. In our experiments, we used a neural system based on the XLM-R, a pre-trained transformer-based masked language model, as a baseline. We show the effectiveness of the proposed approach as it allows to substantially improve the performance of this strong neural baseline model. In addition, in this study, we present multiple types of the XLM-R based classifier, experimenting with various ways of mixing information from the first and second occurrences of the target word in two samples.

generating cross-lingual training cross-lingual training data generating cross-lingual توليد التدريب عبر اللغات بيانات التدريب عبر اللغات توليد اللغة اللغوية صناعة حمض الفوسفور المزيد..

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد