New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Attainable Text-to-Text Machine Translation vs. Translation: Issues Beyond Linguistic Processing

الترجمة من Text-to-to-to-to-to نص مقابل ترجمة: قضايا تتجاوز المعالجة اللغوية

724 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

linguistic processing المعالجة اللغوية صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

يترجم الأساليب الموجودة للترجمة الآلية (MT) في الغالب نص معين في لغة المصدر في اللغة المستهدفة وبدون تشير صراحة إلى المعلومات التي لا غنى عنها لإنتاج ترجمة مناسبة. لا يشمل ذلك فقط المعلومات في العناصر والطرائق النصية الأخرى من النصوص الموجودة في نفس المستند، بل أيضا معلومات إضافية وثلاثة وثيقة وغير لغوية مثل المعايير والسكوب. لتصميم تدفقات عمل الترجمة أفضل ونحن بحاجة إلى التمييز بين مشكلات الترجمة التي يمكن حلها من خلال أساليب النص إلى النص الموجودة وغيرها. تحقيقا لهذه الغاية، أجرينا تقييم تحليلي لنواتج MT وأخذ مهمة ترجمة من الأخبار الإنجليزية إلى اليابانية كدراسة حالة. أولا وأمثلة على مشكلات الترجمة وتنقيحاتها تم جمعها بواسطة طريقة ما بعد التحرير على مرحلتين (PE): أداء الحد الأدنى من PE للحصول على الترجمة التي يمكن تحقيقها بناء على المعلومات النصية المعينة وإجراء المزيد من الأداء الكامل للحصول على ترجمة مقبولة حقا تشير إلى أي المعلومات إذا لزم الأمر. ثم تم تحليل أمثلة المراجعة التي تم جمعها يدويا. كشفنا عن القضايا والمعلومات المهيمنة التي لا غنى عنها لحلها وكائن مثل مواصفات النمط المحبوسين والمعدات المصطلحات والمعرفة الخاصة بالمجال والمستندات المرجعية الخاصة بالمجال وتحديد تمييز واضح بين الترجمة وما يمكن أن يحقق MT النص إلى النص في النهاية.

Existing approaches for machine translation (MT) mostly translate given text in the source language into the target language and without explicitly referring to information indispensable for producing proper translation. This includes not only information in other textual elements and modalities than texts in the same document and but also extra-document and non-linguistic information and such as norms and skopos. To design better translation production work-flows and we need to distinguish translation issues that could be resolved by the existing text-to-text approaches and those beyond them. To this end and we conducted an analytic assessment of MT outputs and taking an English-to-Japanese news translation task as a case study. First and examples of translation issues and their revisions were collected by a two-stage post-editing (PE) method: performing minimal PE to obtain translation attainable based on the given textual information and further performing full PE to obtain truly acceptable translation referring to any information if necessary. Then and the collected revision examples were manually analyzed. We revealed dominant issues and information indispensable for resolving them and such as fine-grained style specifications and terminology and domain-specific knowledge and and reference documents and delineating a clear distinction between translation and what text-to-text MT can ultimately attain.

References used

https://aclanthology.org/

rate research

Zero-Shot Information Extraction as a Unified Text-to-Triple Translation

818 - Association for Computation Linguistics 2021 مقالة

We cast a suite of information extraction tasks into a text-to-triple translation framework. Instead of solving each task relying on task-specific datasets and models, we formalize the task as a translation between task-specific input text and output triples. By taking the task-specific input, we enable a task-agnostic translation by leveraging the latent knowledge that a pre-trained language model has about the task. We further demonstrate that a simple pre-training task of predicting which relational information corresponds to which input text is an effective way to produce task-specific outputs. This enables the zero-shot transfer of our framework to downstream tasks. We study the zero-shot performance of this framework on open information extraction (OIE2016, NYT, WEB, PENN), relation classification (FewRel and TACRED), and factual probe (Google-RE and T-REx). The model transfers non-trivially to most tasks and is often competitive with a fully supervised method without the need for any task-specific training. For instance, we significantly outperform the F1 score of the supervised open information extraction without needing to use its training set.

تعديل الرسم البياني الشرطي open information extraction unified استخراج المعلومات المفتوح موحد صناعة حمض الفوسفور

DuoRAT: Towards Simpler Text-to-SQL Models

488 - Association for Computation Linguistics 2021 مقالة

Recent neural text-to-SQL models can effectively translate natural language questions to corresponding SQL queries on unseen databases. Working mostly on the Spider dataset, researchers have proposed increasingly sophisticated solutions to the proble m. Contrary to this trend, in this paper we focus on simplifications. We begin by building DuoRAT, a re-implementation of the state-of-the-art RAT-SQL model that unlike RAT-SQL is using only relation-aware or vanilla transformers as the building blocks. We perform several ablation experiments using DuoRAT as the baseline model. Our experiments confirm the usefulness of some techniques and point out the redundancy of others, including structural SQL features and features that link the question with the schema.

simpler effectively translate natural translate natural language أبسط ترجمة فعالة الطبيعية ترجمة اللغة الطبيعية صناعة حمض الفوسفور المزيد..

Text-to-SQL in the Wild: A Naturally-Occurring Dataset Based on Stack Exchange Data

337 - Association for Computation Linguistics 2021 مقالة

Most available semantic parsing datasets, comprising of pairs of natural utterances and logical forms, were collected solely for the purpose of training and evaluation of natural language understanding systems. As a result, they do not contain any of the richness and variety of natural-occurring utterances, where humans ask about data they need or are curious about. In this work, we release SEDE, a dataset with 12,023 pairs of utterances and SQL queries collected from real usage on the Stack Exchange website. We show that these pairs contain a variety of real-world challenges which were rarely reflected so far in any other semantic parsing dataset, propose an evaluation metric based on comparison of partial query clauses that is more suitable for real-world queries, and conduct experiments with strong baselines, showing a large gap between the performance on SEDE compared to other common datasets.

stack exchange data naturally-occurring dataset based stack exchange بيانات التبادل المكدس لحالات البيانات التي تحدث بشكل طبيعي كومة البورصة صناعة حمض الفوسفور المزيد..

Itihasa: A large-scale corpus for Sanskrit to English translation

337 - Association for Computation Linguistics 2021 مقالة

This work introduces Itihasa, a large-scale translation dataset containing 93,000 pairs of Sanskrit shlokas and their English translations. The shlokas are extracted from two Indian epics viz., The Ramayana and The Mahabharata. We first describe the motivation behind the curation of such a dataset and follow up with empirical analysis to bring out its nuances. We then benchmark the performance of standard translation models on this corpus and show that even state-of-the-art transformer architectures perform poorly, emphasizing the complexity of the dataset.

work introduces itihasa english translation large-scale translation dataset العمل يقدم Itihasa. الترجمة إلى الإنجليزية مجموعة بيانات الترجمة على نطاق واسع صناعة حمض الفوسفور المزيد..

Attention Is Indeed All You Need: Semantically Attention-Guided Decoding for Data-to-Text NLG

388 - Association for Computation Linguistics 2021 مقالة

Ever since neural models were adopted in data-to-text language generation, they have invariably been reliant on extrinsic components to improve their semantic accuracy, because the models normally do not exhibit the ability to generate text that reli ably mentions all of the information provided in the input. In this paper, we propose a novel decoding method that extracts interpretable information from encoder-decoder models' cross-attention, and uses it to infer which attributes are mentioned in the generated text, which is subsequently used to rescore beam hypotheses. Using this decoding method with T5 and BART, we show on three datasets its ability to dramatically reduce semantic errors in the generated outputs, while maintaining their state-of-the-art quality.

semantically attention-guided decoding semantically attention-guided فك التشفير الدلوي توجيه الانتباه صناعة حمض الفوسفور

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Attainable Text-to-Text Machine Translation vs. Translation: Issues Beyond Linguistic Processing

الترجمة من Text-to-to-to-to-to نص مقابل ترجمة: قضايا تتجاوز المعالجة اللغوية

Ask ChatGPT about the research

Read More

suggested questions