Advanced search powered by artificial intelligence

New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Surprise Language Challenge: Developing a Neural Machine Translation System between Pashto and English in Two Months

تحدي لغة مفاجأة: تطوير نظام ترجمة آلية عصبية بين البشتونية والإنجليزية في شهرين

587 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

في صناعة وسائل الإعلام وتركيز التقارير العالمية قد تحول بين عشية وضحاها. هناك حاجة ملحة إلى أن تكون قادرة على تطوير أنظمة ترجمة آلية جديدة في فترة زمنية قصيرة وللغطي بشكل أكثر كفاءة تطوير القصص بسرعة أكبر. كجزء من مشروع EU Gourmet ورفع التركيز على الترجمة ذات الآلات المنخفضة وشركاؤنا الإعلامي لغة مفاجئة يجب أن يتم بناء نظام الترجمة الآلي وتقييمه خلال شهرين (فبراير وآذار / مارس 2021). كانت اللغة المختارة كانت الباشتونية ولغة هندية إيرانية تحدثت في أفغانستان وباكستان والهند. في هذه الفترة، أكملنا خط الأنابيب الكامل لتنمية نظام الترجمة الآلية العصبية: الزحف البيانات والتنظيف ومحاذاة وإنشاء مجموعات الاختبار وتطوير ونماذج الاختبار وتقديمها إلى شركاء المستخدمين. في هذه الورق، نطبق إنشاء البيانات والتجارب السريعة مع التعلم والنقل الاحتياطي لهذا زوج لغة الموارد المنخفضة. نجد أن بدءا من نموذج كبير موجود مدرب مسبقا على 50 لغة يؤدي إلى نتائج بلو أفضل بكثير من الاحيلية على زوج لغة موارد عالية مع نموذج أصغر. نقدم أيضا تقييم بشري لأنظمنا والتي تشير إلى أن النظم الناتجة أداء أفضل من النظام التجاري المتاح بحرية عند ترجمة من اللغة الإنجليزية إلى اتجاه البشتونية وبالمثل عند ترجمة من البشتو إلى الإنجليزية.

In the media industry and the focus of global reporting can shift overnight. There is a compelling need to be able to develop new machine translation systems in a short period of time and in order to more efficiently cover quickly developing stories. As part of the EU project GoURMET and which focusses on low-resource machine translation and our media partners selected a surprise language for which a machine translation system had to be built and evaluated in two months(February and March 2021). The language selected was Pashto and an Indo-Iranian language spoken in Afghanistan and Pakistan and India. In this period we completed the full pipeline of development of a neural machine translation system: data crawling and cleaning and aligning and creating test sets and developing and testing models and and delivering them to the user partners. In this paperwe describe rapid data creation and experiments with transfer learning and pretraining for this low-resource language pair. We find that starting from an existing large model pre-trained on 50languages leads to far better BLEU scores than pretraining on one high-resource language pair with a smaller model. We also present human evaluation of our systems and which indicates that the resulting systems perform better than a freely available commercial system when translating from English into Pashto direction and and similarly when translating from Pashto into English.

References used

https://aclanthology.org/

rate research

Approaching Sign Language Gloss Translation as a Low-Resource Machine Translation Task

614 - Association for Computation Linguistics 2021 مقالة

A cascaded Sign Language Translation system first maps sign videos to gloss annotations and then translates glosses into a spoken languages. This work focuses on the second-stage gloss translation component, which is challenging due to the scarcity o f publicly available parallel data. We approach gloss translation as a low-resource machine translation task and investigate two popular methods for improving translation quality: hyperparameter search and backtranslation. We discuss the potentials and pitfalls of these methods based on experiments on the RWTH-PHOENIX-Weather 2014T dataset.

approaching sign language اقترب لغة الإشارة صناعة حمض الفوسفور

AVASAG: A German Sign Language Translation System for Public Services (short paper)

659 - Association for Computation Linguistics 2021 مقالة

This paper presents an overview of AVASAG; an ongoing applied-research project developing a text-to-sign-language translation system for public services. We describe the scientific innovation points (geometry-based SL-description, 3D animation and video corpus, simplified annotation scheme, motion capture strategy) and the overall translation pipeline.

أسطورة التوقيع language translation system نظام ترجمة اللغة صناعة حمض الفوسفور

Examining Covert Gender Bias: A Case Study in Turkish and English Machine Translation Models

762 - Association for Computation Linguistics 2021 مقالة

As Machine Translation (MT) has become increasingly more powerful, accessible, and widespread, the potential for the perpetuation of bias has grown alongside its advances. While overt indicators of bias have been studied in machine translation, we ar gue that covert biases expose a problem that is further entrenched. Through the use of the gender-neutral language Turkish and the gendered language English, we examine cases of both overt and covert gender bias in MT models. Specifically, we introduce a method to investigate asymmetrical gender markings. We also assess bias in the attribution of personhood and examine occupational and personality stereotypes through overt bias indicators in MT models. Our work explores a deeper layer of bias in MT models and demonstrates the continued need for language-specific, interdisciplinary methodology in MT model development.

english machine translation ترجمة آلة اللغة الإنجليزية صناعة حمض الفوسفور

Sign Language Translation in a Healthcare Setting

952 - Association for Computation Linguistics 2021 مقالة

Communication between healthcare professionals and deaf patients is challenging, and the current COVID-19 pandemic makes this issue even more acute. Sign language interpreters can often not enter hospitals and face masks make lipreading impossible. T o address this urgent problem, we developed a system which allows healthcare professionals to translate sentences that are frequently used in the diagnosis and treatment of COVID-19 into Sign Language of the Netherlands (NGT). Translations are displayed by means of videos and avatar animations. The architecture of the system is such that it could be extended to other applications and other sign languages in a relatively straightforward way.

healthcare setting sign language sign language translation إعداد الرعاية الصحية لغة الإشارة ترجمة لغة الإشارة صناعة حمض الفوسفور المزيد..

Machine Translation in the Covid domain: an English-Irish case study for LoResMT 2021

1164 - Association for Computation Linguistics 2021 مقالة

Translation models for the specific domain of translating Covid data from English to Irish were developed for the LoResMT 2021 shared task. Domain adaptation techniques, using a Covid-adapted generic 55k corpus from the Directorate General of Transla tion, were applied. Fine-tuning, mixed fine-tuning and combined dataset approaches were compared with models trained on an extended in-domain dataset. As part of this study, an English-Irish dataset of Covid related data, from the Health and Education domains, was developed. The highestperforming model used a Transformer architecture trained with an extended in-domain Covid dataset. In the context of this study, we have demonstrated that extending an 8k in-domain baseline dataset by just 5k lines improved the BLEU score by 27 points.

اليقظ قليلا ضبط english to irish covid الإنجليزية إلى الأيرلندية مرض فيروس كورونا صناعة حمض الفوسفور

Surprise Language Challenge: Developing a Neural Machine Translation System between Pashto and English in Two Months

تحدي لغة مفاجأة: تطوير نظام ترجمة آلية عصبية بين البشتونية والإنجليزية في شهرين

Ask ChatGPT about the research

Read More

suggested questions