Research papers, master and doctoral theses about بلغ ملغات الموارد المنخفضة

Morphologically-Guided Segmentation For Translation of Agglutinative Low-Resource Languages

388 - Association for Computation Linguistics 2021 مقالة

Neural Machine Translation (NMT) for Low Resource Languages (LRL) is often limited by the lack of available training data, making it necessary to explore additional techniques to improve translation quality. We propose the use of the Prefix-Root-Post fix-Encoding (PRPE) subword segmentation algorithm to improve translation quality for LRLs, using two agglutinative languages as case studies: Quechua and Indonesian. During the course of our experiments, we reintroduce a parallel corpus for Quechua-Spanish translation that was previously unavailable for NMT. Our experiments show the importance of appropriate subword segmentation, which can go as far as improving translation quality over systems trained on much larger quantities of data. We show this by achieving state-of-the-art results for both languages, obtaining higher BLEU scores than large pre-trained models with much smaller amounts of data.

وسائل التواصل الاجتماعي التعليقات agglutinative low-resource languages improve translation quality بلغ ملغات الموارد المنخفضة تحسين جودة الترجمة صناعة حمض الفوسفور

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد