Advanced search powered by artificial intelligence

New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Crosslingual Embeddings are Essential in UNMT for distant languages: An English to IndoAryan Case Study

تعد Erblingual Embeddings ضرورية في UNMT للغات البعيدة: دراسة حالة الإنجليزية إلى Indooaryan

552 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

indoaryan case study دراسة حالة الهند صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

تقدم التطورات الحديثة في الترجمة الآلية العصبية غير المدعومة (IPNMT) من الفجوة بين أداء ترجمة الآلات الخاضعة للإشراف وغير المعروضة لأزواج اللغة ذات الصلة عن كثب. ومع ذلك، والوضع مختلف جدا على أزواج اللغة البعيدة. يؤدي نقص التداخل في المعجم وانخفاض التشابه النحوي، مثل اللغة الإنجليزية واللغات الهندية إلى ضعف جودة الترجمة في أنظمة IPS الحالية. في هذه الورقة، نعرض أن تهيئة طبقة التضمين من طرازات التضمين من طرازات بروتوكول الكثال الكثال الكربون البرمجية مع ادبات متبلة يؤدي إلى تحسينات نقاط بلو كبيرة على نماذج IPS الحالية حيث تتم تهيئة أوزان طبقة تضمينها بشكل عشوائي. مما يؤدي وتجميد الأوزان طبقة التضمين إلى تحسين مكاسب أفضل مقارنة بتحديث أوزان طبقة تضمينها أثناء التدريب. لقد جربنا باستخدام تسلسل ملثمين للتسلسل (الكتلة) وتدينك مناهج AUTONCONDER (DAE) لنهج البث لمدة ثلاث أزواج لغة بعيدة. تهيئة تضمين التضمين المتبادلة المقترحة تحسن نتيجة بلو ما يصل إلى عشر مرات فوق خط الأساس للإنجليزية-الهندية والإنجليزية-البنغالية والإنجليزية-الغوجاراتية. يوضح تحليلنا أن تهيئة طبقة التضمين مع رسم خرائط تضمين التضمين الساكنة ضرورية لتدريب نماذج بعثة الأمم المتحدة في غول الرصاص على أزواج اللغة البعيدة.

Recent advances in Unsupervised Neural Machine Translation (UNMT) has minimized the gap between supervised and unsupervised machine translation performance for closely related language-pairs. However and the situation is very different for distant language pairs. Lack of overlap in lexicon and low syntactic similarity such as between English and IndoAryan languages leads to poor translation quality in existing UNMT systems. In this paper and we show that initialising the embedding layer of UNMT models with cross-lingual embeddings leads to significant BLEU score improvements over existing UNMT models where the embedding layer weights are randomly initialized. Further and freezing the embedding layer weights leads to better gains compared to updating the embedding layer weights during training. We experimented using Masked Sequence to Sequence (MASS) and Denoising Autoencoder (DAE) UNMT approaches for three distant language pairs. The proposed cross-lingual embedding initialization yields BLEU score improvement of as much as ten times over the baseline for English-Hindi and English-Bengali and English-Gujarati. Our analysis shows that initialising embedding layer with static cross-lingual embedding mapping is essential for training of UNMT models for distant language-pairs.

References used

https://aclanthology.org/

rate research

Machine Translation in the Covid domain: an English-Irish case study for LoResMT 2021

1030 - Association for Computation Linguistics 2021 مقالة

Translation models for the specific domain of translating Covid data from English to Irish were developed for the LoResMT 2021 shared task. Domain adaptation techniques, using a Covid-adapted generic 55k corpus from the Directorate General of Transla tion, were applied. Fine-tuning, mixed fine-tuning and combined dataset approaches were compared with models trained on an extended in-domain dataset. As part of this study, an English-Irish dataset of Covid related data, from the Health and Education domains, was developed. The highestperforming model used a Transformer architecture trained with an extended in-domain Covid dataset. In the context of this study, we have demonstrated that extending an 8k in-domain baseline dataset by just 5k lines improved the BLEU score by 27 points.

اليقظ قليلا ضبط english to irish covid الإنجليزية إلى الأيرلندية مرض فيروس كورونا صناعة حمض الفوسفور

Neural Machine Translation in Low-Resource Setting: a Case Study in English-Marathi Pair

710 - Association for Computation Linguistics 2021 مقالة

In this paper and we explore different techniques of overcoming the challenges of low-resource in Neural Machine Translation (NMT) and specifically focusing on the case of English-Marathi NMT. NMT systems require a large amount of parallel corpora to obtain good quality translations. We try to mitigate the low-resource problem by augmenting parallel corpora or by using transfer learning. Techniques such as Phrase Table Injection (PTI) and back-translation and mixing of language corpora are used for enhancing the parallel data; whereas pivoting and multilingual embeddings are used to leverage transfer learning. For pivoting and Hindi comes in as assisting language for English-Marathi translation. Compared to baseline transformer model and a significant improvement trend in BLEU score is observed across various techniques. We have done extensive manual and automatic and qualitative evaluation of our systems. Since the trend in Machine Translation (MT) today is post-editing and measuring of Human Effort Reduction (HER) and we have given our preliminary observations on Translation Edit Rate (TER) vs. BLEU score study and where TER is regarded as a measure of HER.

دراسة حالة الهند صناعة حمض الفوسفور

To What Extent Does Lexical Normalization Help English-as-a-Second Language Learners to Read Noisy English Texts?

608 - Association for Computation Linguistics 2021 مقالة

How difficult is it for English-as-a-second language (ESL) learners to read noisy English texts? Do ESL learners need lexical normalization to read noisy English texts? These questions may also affect community formation on social networking sites wh ere differences can be attributed to ESL learners and native English speakers. However, few studies have addressed these questions. To this end, we built highly accurate readability assessors to evaluate the readability of texts for ESL learners. We then applied these assessors to noisy English texts to further assess the readability of the texts. The experimental results showed that although intermediate-level ESL learners can read most noisy English texts in the first place, lexical normalization significantly improves the readability of noisy English texts for ESL learners.

noisy english texts read noisy english noisy english النصوص الإنجليزية صاخبة قراءة nooisy الإنجليزية الإنجليزية صاخبة صناعة حمض الفوسفور المزيد..

Linguistic Evaluation for the 2021 State-of-the-art Machine Translation Systems for German to English and English to German

881 - Association for Computation Linguistics 2021 مقالة

We are using a semi-automated test suite in order to provide a fine-grained linguistic evaluation for state-of-the-art machine translation systems. The evaluation includes 18 German to English and 18 English to German systems, submitted to the Transl ation Shared Task of the 2021 Conference on Machine Translation. Our submission adds up to the submissions of the previous years by creating and applying a wide-range test suite for English to German as a new language pair. The fine-grained evaluation allows spotting significant differences between systems that cannot be distinguished by the direct assessment of the human evaluation campaign. We find that most of the systems achieve good accuracies in the majority of linguistic phenomena but there are few phenomena with lower accuracy, such as the idioms, the modal pluperfect and the German resultative predicates. Two systems have significantly better test suite accuracy in macro-average in every language direction, Online-W and Facebook-AI for German to English and VolcTrans and Online-W for English to German. The systems show a steady improvement as compared to previous years.

تحسين بقوة صناعة حمض الفوسفور

Scrambled Translation Problem: A Problem of Denoising UNMT

979 - Association for Computation Linguistics 2021 مقالة

In this paper and we identify an interesting kind of error in the output of Unsupervised Neural Machine Translation (UNMT) systems like Undreamt1. We refer to this error type as Scrambled Translation problem. We observe that UNMT models which use wor d shuffle noise (as in case of Undreamt) can generate correct words and but fail to stitch them together to form phrases. As a result and words of the translated sentence look scrambled and resulting in decreased BLEU. We hypothesise that the reason behind scrambled translation problem is 'shuffling noise' which is introduced in every input sentence as a denoising strategy. To test our hypothesis and we experiment by retraining UNMT models with a simple retraining strategy. We stop the training of the Denoising UNMT model after a pre-decided number of iterations and resume the training for the remaining iterations- which number is also pre-decided- using original sentence as input without adding any noise. Our proposed solution achieves significant performance improvement UNMT models that train conventionally. We demonstrate these performance gains on four language pairs and viz. and English-French and English-German and English-Spanish and Hindi-Punjabi. Our qualitative and quantitative analysis shows that the retraining strategy helps achieve better alignment as observed by attention heatmap and better phrasal translation and leading to statistically significant improvement in BLEU scores.

scrambled translation problem مشكلة الترجمة المخفوقة صناعة حمض الفوسفور

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Crosslingual Embeddings are Essential in UNMT for distant languages: An English to IndoAryan Case Study

تعد Erblingual Embeddings ضرورية في UNMT للغات البعيدة: دراسة حالة الإنجليزية إلى Indooaryan

Ask ChatGPT about the research

Read More

suggested questions