Advanced search powered by artificial intelligence

New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Cultural and Geographical Influences on Image Translatability of Words across Languages

التأثيرات الثقافية والجغرافية على صورة ترجمة الكلمات عبر اللغات

609 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

geographical influences cultural and geographical التأثيرات الجغرافية الثقافية والجغرافية كلمات صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

لوحظت نماذج الترجمة الآلية العصبية (NMT) لإنتاج ترجمات سيئة عندما يكون هناك عدد قليل من الجمل / لا توجد جمل متوازية لتدريب النماذج. في حالة عدم وجود بيانات متوازية، تحولت عدة طرق إلى استخدام الصور لتعلم الترجمات. نظرا لأن صور الكلمات، على سبيل المثال، قد لا تتغير الحصان عبر اللغات، يمكن تحديد الترجمات عبر الصور المرتبطة بالكلمات بلغات مختلفة تحتوي على درجة عالية من التشابه البصري. ومع ذلك، تم عرض ترجمة عبر الصور تتحسن عند نماذج النص فقط بشكل هامشي. لفهم أفضل عندما تكون الصور مفيدة للترجمة، ندرس صورة ترجمتها للكلمات، والتي نحددها كترجمة الكلمات عبر الصور، من خلال قياس أوجه التشابه بين المعلومات بين التصنيفات للكلمات التي ترجمات من بعضها البعض. نجد أن صور الكلمات ليست دائما ثابتة عبر اللغات، وأن أزواج اللغة ذات الثقافة المشتركة، والتي تعني إما عائلة لغة مشتركة أو عرقية أو دين، قد تحسنت إمكانية تحسن الصور (أي صور مشابهة للكلمات المماثلة) يحادثون، بغض النظر عن قربهم الجغرافي. بالإضافة إلى ذلك، تمشيا مع الأعمال السابقة التي تظهر الصور تساعد المزيد في ترجمة الكلمات الملموسة، وجدنا أن الكلمات الملموسة قد تحسنت إمكانية الحصول على صورة حسب الاقتضاء.

Neural Machine Translation (NMT) models have been observed to produce poor translations when there are few/no parallel sentences to train the models. In the absence of parallel data, several approaches have turned to the use of images to learn translations. Since images of words, e.g., horse may be unchanged across languages, translations can be identified via images associated with words in different languages that have a high degree of visual similarity. However, translating via images has been shown to improve upon text-only models only marginally. To better understand when images are useful for translation, we study image translatability of words, which we define as the translatability of words via images, by measuring intra- and inter-cluster similarities of image representations of words that are translations of each other. We find that images of words are not always invariant across languages, and that language pairs with shared culture, meaning having either a common language family, ethnicity or religion, have improved image translatability (i.e., have more similar images for similar words) compared to its converse, regardless of their geographic proximity. In addition, in line with previous works that show images help more in translating concrete words, we found that concrete words have improved image translatability compared to abstract ones.

References used

https://aclanthology.org/

rate research

Adapting Entities across Languages and Cultures

1043 - Association for Computation Linguistics 2021 مقالة

How would you explain Bill Gates to a German? He is associated with founding a company in the United States, so perhaps the German founder Carl Benz could stand in for Gates in those contexts. This type of translation is called adaptation in the tran slation community. Until now, this task has not been done computationally. Automatic adaptation could be used in natural language processing for machine translation and indirectly for generating new question answering datasets and education. We propose two automatic methods and compare them to human results for this novel NLP task. First, a structured knowledge base adapts named entities using their shared properties. Second, vector-arithmetic and orthogonal embedding mappings methods identify better candidates, but at the expense of interpretable features. We evaluate our methods through a new dataset of human adaptations.

explain bill gates languages and cultures cultures اشرح بيل غيتس اللغات والثقافات الثقافات صناعة حمض الفوسفور المزيد..

Subword Mapping and Anchoring across Languages

750 - Association for Computation Linguistics 2021 مقالة

State-of-the-art multilingual systems rely on shared vocabularies that sufficiently cover all considered languages. To this end, a simple and frequently used approach makes use of subword vocabularies constructed jointly over several languages. We hy pothesize that such vocabularies are suboptimal due to false positives (identical subwords with different meanings across languages) and false negatives (different subwords with similar meanings). To address these issues, we propose Subword Mapping and Anchoring across Languages (SMALA), a method to construct bilingual subword vocabularies. SMALA extracts subword alignments using an unsupervised state-of-the-art mapping technique and uses them to create cross-lingual anchors based on subword similarities. We demonstrate the benefits of SMALA for cross-lingual natural language inference (XNLI), where it improves zero-shot transfer to an unseen language without task-specific data, but only by sharing subword embeddings. Moreover, in neural machine translation, we show that joint subword vocabularies obtained with SMALA lead to higher BLEU scores on sentences that contain many false positives and false negatives.

subword mapping and anchoring كلمة فرعية رسم الخرائط والرسوم صناعة حمض الفوسفور

Automatic Discrimination between Inherited and Borrowed Latin Words in Romance Languages

546 - Association for Computation Linguistics 2021 مقالة

In this paper, we address the problem of automatically discriminating between inherited and borrowed Latin words. We introduce a new dataset and investigate the case of Romance languages (Romanian, Italian, French, Spanish, Portuguese and Catalan), w here words directly inherited from Latin coexist with words borrowed from Latin, and explore whether automatic discrimination between them is possible. Having entered the language at a later stage, borrowed words are no longer subject to historical sound shift rules, hence they are presumably less eroded, which is why we expect them to have a different intrinsic structure distinguishable by computational means. We employ several machine learning models to automatically discriminate between inherited and borrowed words and compare their performance with various feature sets. We analyze the models' predictive power on two versions of the datasets, orthographic and phonetic. We also investigate whether prior knowledge of the etymon provides better results, employing n-gram character features extracted from the word-etymon pairs and from their alignment.

borrowed latin words borrowed latin latin words الكلمات اللاتينية المقترضة استعارة اللاتينية كلمات اللاتينية صناعة حمض الفوسفور المزيد..

Wine is not v i n. On the Compatibility of Tokenizations across Languages

658 - Association for Computation Linguistics 2021 مقالة

The size of the vocabulary is a central design choice in large pretrained language models, with respect to both performance and memory requirements. Typically, subword tokenization algorithms such as byte pair encoding and WordPiece are used. In this work, we investigate the compatibility of tokenizations for multilingual static and contextualized embedding spaces and propose a measure that reflects the compatibility of tokenizations across languages. Our goal is to prevent incompatible tokenizations, e.g., wine'' (word-level) in English vs. v i n'' (character-level) in French, which make it hard to learn good multilingual semantic representations. We show that our compatibility measure allows the system designer to create vocabularies across languages that are compatible -- a desideratum that so far has been neglected in multilingual models.

القراءة الصينية الفهم tokenizations compatibility of tokenizations التوصيلات توافق التوصيلات صناعة حمض الفوسفور

On the cross-lingual transferability of multilingual prototypical models across NLU tasks

1097 - Association for Computation Linguistics 2021 مقالة

Supervised deep learning-based approaches have been applied to task-oriented dialog and have proven to be effective for limited domain and language applications when a sufficient number of training examples are available. In practice, these approache s suffer from the drawbacks of domain-driven design and under-resourced languages. Domain and language models are supposed to grow and change as the problem space evolves. On one hand, research on transfer learning has demonstrated the cross-lingual ability of multilingual Transformers-based models to learn semantically rich representations. On the other, in addition to the above approaches, meta-learning have enabled the development of task and language learning algorithms capable of far generalization. Through this context, this article proposes to investigate the cross-lingual transferability of using synergistically few-shot learning with prototypical neural networks and multilingual Transformers-based models. Experiments in natural language understanding tasks on MultiATIS++ corpus shows that our approach substantially improves the observed transfer learning performances between the low and the high resource languages. More generally our approach confirms that the meaningful latent space learned in a given language can be can be generalized to unseen and under-resourced ones using meta-learning.

nlu tasks multilingual transformers-based models مهام NLU النماذج القائمة على المحولات متعددة اللغات صناعة حمض الفوسفور

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Cultural and Geographical Influences on Image Translatability of Words across Languages

التأثيرات الثقافية والجغرافية على صورة ترجمة الكلمات عبر اللغات

Ask ChatGPT about the research

Read More

suggested questions