New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Scheduled Sampling Based on Decoding Steps for Neural Machine Translation

أخذ العينات المجدولة بناء على خطوات فك التشفير للترجمة الآلية العصبية

283 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

متعدد رئيس الانتباه مؤخرا صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

يتم استخدام أخذ العينات المجدولة على نطاق واسع للتخفيف من مشكلة تحيز التعرض الترجمة الآلية العصبية. الدافع الأساسي هو محاكاة مشهد الاستدلال أثناء التدريب من خلال استبدال الرموز الأرضية مع الرموز الرائعة المتوقعة، وبالتالي سد الفجوة بين التدريب والاستدلال. ومع ذلك، فإن أخذ العينات المقررة للفانيليا تعتمد فقط على خطوات التدريب وعادل على قدم المساواة جميع خطوات فك التشفير. وهي تحاكي مشهد الاستدلال بمعدلات خطأ موحدة، والتي تفحص مشهد الاستدلال الحقيقي، حيث توجد خطوات فك التشفير الكبيرة عادة معدلات خطأ أعلى بسبب تراكم الخطأ. لتخفيف التناقض أعلاه، نقترح أساليب أخذ العينات المجدولة بناء على خطوات فك التشفير، مما يزيد من فرصة اختيار الرموز المتوقعة مع نمو خطوات فك التشفير. وبالتالي، يمكننا أن نحاكي أكثر واقعية المشهد الاستدلال أثناء التدريب، وبالتالي سد الفجوة بشكل أفضل بين التدريب والاستدلال. علاوة على ذلك، نحقق في أخذ العينات المجدولة بناء على كل من خطوات التدريب وفك تشفير الخطوات لمزيد من التحسينات. تجريبيا، فإن نهجنا تتفوق بشكل كبير على خط الأساس المحول وأخذ عينات من الفانيليا المجدولة على ثلاث مهام WMT واسعة النطاق. بالإضافة إلى ذلك، تعميم نهجنا أيضا بشكل جيد لمهمة تلخيص النص على معايير شعبية.

Scheduled sampling is widely used to mitigate the exposure bias problem for neural machine translation. Its core motivation is to simulate the inference scene during training by replacing ground-truth tokens with predicted tokens, thus bridging the gap between training and inference. However, vanilla scheduled sampling is merely based on training steps and equally treats all decoding steps. Namely, it simulates an inference scene with uniform error rates, which disobeys the real inference scene, where larger decoding steps usually have higher error rates due to error accumulations. To alleviate the above discrepancy, we propose scheduled sampling methods based on decoding steps, increasing the selection chance of predicted tokens with the growth of decoding steps. Consequently, we can more realistically simulate the inference scene during training, thus better bridging the gap between training and inference. Moreover, we investigate scheduled sampling based on both training steps and decoding steps for further improvements. Experimentally, our approaches significantly outperform the Transformer baseline and vanilla scheduled sampling on three large-scale WMT tasks. Additionally, our approaches also generalize well to the text summarization task on two popular benchmarks.

References used

https://aclanthology.org/

rate research

Encouraging Lexical Translation Consistency for Document-Level Neural Machine Translation

378 - Association for Computation Linguistics 2021 مقالة

Recently a number of approaches have been proposed to improve translation performance for document-level neural machine translation (NMT). However, few are focusing on the subject of lexical translation consistency. In this paper we apply one transla tion per discourse'' in NMT, and aim to encourage lexical translation consistency for document-level NMT. This is done by first obtaining a word link for each source word in a document, which tells the positions where the source word appears. Then we encourage the translation of those words within a link to be consistent in two ways. On the one hand, when encoding sentences within a document we properly share context information of those words. On the other hand, we propose an auxiliary loss function to better constrain that their translation should be consistent. Experimental results on Chinese↔English and English→French translation tasks show that our approach not only achieves state-of-the-art performance in BLEU scores, but also greatly improves lexical consistency in translation.

متعدد رئيس الانتباه مؤخرا صناعة حمض الفوسفور

Learning to Rewrite for Non-Autoregressive Neural Machine Translation

321 - Association for Computation Linguistics 2021 مقالة

Non-autoregressive neural machine translation, which decomposes the dependence on previous target tokens from the inputs of the decoder, has achieved impressive inference speedup but at the cost of inferior accuracy. Previous works employ iterative d ecoding to improve the translation by applying multiple refinement iterations. However, a serious drawback is that these approaches expose the serious weakness in recognizing the erroneous translation pieces. In this paper, we propose an architecture named RewriteNAT to explicitly learn to rewrite the erroneous translation pieces. Specifically, RewriteNAT utilizes a locator module to locate the erroneous ones, which are then revised into the correct ones by a revisor module. Towards keeping the consistency of data distribution with iterative decoding, an iterative training strategy is employed to further improve the capacity of rewriting. Extensive experiments conducted on several widely-used benchmarks show that RewriteNAT can achieve better performance while significantly reducing decoding time, compared with previous iterative decoding strategies. In particular, RewriteNAT can obtain competitive results with autoregressive translation on WMT14 En-De, En-Fr and WMT16 Ro-En translation benchmarks.

متعدد رئيس الانتباه مؤخرا صناعة حمض الفوسفور

Unsupervised Neural Machine Translation with Universal Grammar

457 - Association for Computation Linguistics 2021 مقالة

Machine translation usually relies on parallel corpora to provide parallel signals for training. The advent of unsupervised machine translation has brought machine translation away from this reliance, though performance still lags behind traditional supervised machine translation. In unsupervised machine translation, the model seeks symmetric language similarities as a source of weak parallel signal to achieve translation. Chomsky's Universal Grammar theory postulates that grammar is an innate form of knowledge to humans and is governed by universal principles and constraints. Therefore, in this paper, we seek to leverage such shared grammar clues to provide more explicit language parallel signals to enhance the training of unsupervised machine translation models. Through experiments on multiple typical language pairs, we demonstrate the effectiveness of our proposed approaches.

متعدد رئيس الانتباه مؤخرا صناعة حمض الفوسفور

Smart-Start Decoding for Neural Machine Translation

384 - Association for Computation Linguistics 2021 مقالة

Most current neural machine translation models adopt a monotonic decoding order of either left-to-right or right-to-left. In this work, we propose a novel method that breaks up the limitation of these decoding orders, called Smart-Start decoding. Mor e specifically, our method first predicts a median word. It starts to decode the words on the right side of the median word and then generates words on the left. We evaluate the proposed Smart-Start decoding method on three datasets. Experimental results show that the proposed method can significantly outperform strong baseline models.

آلة ذات مستوى المستند صناعة حمض الفوسفور

Enlivening Redundant Heads in Multi-head Self-attention for Machine Translation

401 - Association for Computation Linguistics 2021 مقالة

Multi-head self-attention recently attracts enormous interest owing to its specialized functions, significant parallelizable computation, and flexible extensibility. However, very recent empirical studies show that some self-attention heads make litt le contribution and can be pruned as redundant heads. This work takes a novel perspective of identifying and then vitalizing redundant heads. We propose a redundant head enlivening (RHE) method to precisely identify redundant heads, and then vitalize their potential by learning syntactic relations and prior knowledge in the text without sacrificing the roles of important heads. Two novel syntax-enhanced attention (SEA) mechanisms: a dependency mask bias and a relative local-phrasal position bias, are introduced to revise self-attention distributions for syntactic enhancement in machine translation. The importance of individual heads is dynamically evaluated during the redundant heads identification, on which we apply SEA to vitalize redundant heads while maintaining the strength of important heads. Experimental results on widely adopted WMT14 and WMT16 English to German and English to Czech language machine translation validate the RHE effectiveness.

redundant heads multi-head self-attention multi-head self-attention recently رؤساء الزائدة متعدد رئيس الانتباه متعدد رئيس الانتباه مؤخرا صناعة حمض الفوسفور المزيد..

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Scheduled Sampling Based on Decoding Steps for Neural Machine Translation

أخذ العينات المجدولة بناء على خطوات فك التشفير للترجمة الآلية العصبية

Ask ChatGPT about the research

Read More

suggested questions