New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Subword Mapping and Anchoring across Languages

رسم الخرائط الفرعية ومثبتة عبر اللغات

344 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

تعتمد أنظمة متعددة اللغات متعددة اللغات على المفردات المشتركة التي تغطي جميع اللغات التي تغطي بما فيه الكفاية. تحقيقا لهذه الغاية، فإن النهج البسيط والمستعمل بشكل متكرر يستفيد من مفهليات الكلمات الفرعية التي تم إنشاؤها بشكل مشترك على عدة لغات. نحن نفترض أن مثل هذه المفردات هي فرعية نفسها بسبب الإيجابيات الخاطئة (الكلمات الفرعية المماثلة مع معاني مختلفة عبر اللغات) والسلبيات الخاطئة (كلمات فرعية مختلفة مع معاني مماثلة). لمعالجة هذه المشكلات، نقترح رسم الخرائط عن طريق الكلمات الفرعية ومثبتة عبر اللغات (SMALA)، وهي طريقة لبناء مخصصات الكلمات الفرعية ثنائية اللغة. تقوم SMALA باستخراج محاذاة الكلمات الفرعية باستخدام تقنية رسم الخرائط غير المزودة بعملية رسم الخرائط واستخدامها لإنشاء مراسي عبر اللغات بناء على أوجه تشابه الكلمات الفرعية. نوضح فوائد SMALA للاستدلال اللغوي للغة الطبيعية المتبادلة (XNLI)، حيث يحسن تحويل صفرية إلى لغة غير مرئية دون بيانات مهمة، ولكن فقط من خلال تقاسم تضييق الكلمات الفرعية. علاوة على ذلك، في الترجمة الآلية العصبية، نوضح أن مفردات الكلمة الفرعية المشتركة التي تم الحصول عليها مع Smala تؤدي إلى أعلى درجات بلو على أحكام تحتوي على العديد من الإيجابيات الخاطئة والسلبيات الخاطئة.

State-of-the-art multilingual systems rely on shared vocabularies that sufficiently cover all considered languages. To this end, a simple and frequently used approach makes use of subword vocabularies constructed jointly over several languages. We hypothesize that such vocabularies are suboptimal due to false positives (identical subwords with different meanings across languages) and false negatives (different subwords with similar meanings). To address these issues, we propose Subword Mapping and Anchoring across Languages (SMALA), a method to construct bilingual subword vocabularies. SMALA extracts subword alignments using an unsupervised state-of-the-art mapping technique and uses them to create cross-lingual anchors based on subword similarities. We demonstrate the benefits of SMALA for cross-lingual natural language inference (XNLI), where it improves zero-shot transfer to an unseen language without task-specific data, but only by sharing subword embeddings. Moreover, in neural machine translation, we show that joint subword vocabularies obtained with SMALA lead to higher BLEU scores on sentences that contain many false positives and false negatives.

References used

https://aclanthology.org/

rate research

Adapting Entities across Languages and Cultures

565 - Association for Computation Linguistics 2021 مقالة

How would you explain Bill Gates to a German? He is associated with founding a company in the United States, so perhaps the German founder Carl Benz could stand in for Gates in those contexts. This type of translation is called adaptation in the tran slation community. Until now, this task has not been done computationally. Automatic adaptation could be used in natural language processing for machine translation and indirectly for generating new question answering datasets and education. We propose two automatic methods and compare them to human results for this novel NLP task. First, a structured knowledge base adapts named entities using their shared properties. Second, vector-arithmetic and orthogonal embedding mappings methods identify better candidates, but at the expense of interpretable features. We evaluate our methods through a new dataset of human adaptations.

explain bill gates languages and cultures cultures اشرح بيل غيتس اللغات والثقافات الثقافات صناعة حمض الفوسفور المزيد..

A (Non)-Perfect Match: Mapping plWordNet onto PrincetonWordNet

555 - Association for Computation Linguistics 2021 مقالة

The paper reports on the methodology and final results of a large-scale synset mapping between plWordNet and Princeton WordNet. Dedicated manual and semi-automatic mapping procedures as well as interlingual relation types for nouns, verbs, adjectives and adverbs are described. The statistics of all types of interlingual relations are also provided.

perfect match perfect match تطابق مثالي في احسن الاحوال تطابق صناعة حمض الفوسفور المزيد..

Cultural and Geographical Influences on Image Translatability of Words across Languages

269 - Association for Computation Linguistics 2021 مقالة

Neural Machine Translation (NMT) models have been observed to produce poor translations when there are few/no parallel sentences to train the models. In the absence of parallel data, several approaches have turned to the use of images to learn transl ations. Since images of words, e.g., horse may be unchanged across languages, translations can be identified via images associated with words in different languages that have a high degree of visual similarity. However, translating via images has been shown to improve upon text-only models only marginally. To better understand when images are useful for translation, we study image translatability of words, which we define as the translatability of words via images, by measuring intra- and inter-cluster similarities of image representations of words that are translations of each other. We find that images of words are not always invariant across languages, and that language pairs with shared culture, meaning having either a common language family, ethnicity or religion, have improved image translatability (i.e., have more similar images for similar words) compared to its converse, regardless of their geographic proximity. In addition, in line with previous works that show images help more in translating concrete words, we found that concrete words have improved image translatability compared to abstract ones.

geographical influences cultural and geographical التأثيرات الجغرافية الثقافية والجغرافية كلمات صناعة حمض الفوسفور

Universal Joy A Data Set and Results for Classifying Emotions Across Languages

192 - Association for Computation Linguistics 2021 مقالة

While emotions are universal aspects of human psychology, they are expressed differently across different languages and cultures. We introduce a new data set of over 530k anonymized public Facebook posts across 18 languages, labeled with five differe nt emotions. Using multilingual BERT embeddings, we show that emotions can be reliably inferred both within and across languages. Zero-shot learning produces promising results for low-resource languages. Following established theories of basic emotions, we provide a detailed analysis of the possibilities and limits of cross-lingual emotion classification. We find that structural and typological similarity between languages facilitates cross-lingual learning, as well as linguistic diversity of training data. Our results suggest that there are commonalities underlying the expression of emotion in different languages. We publicly release the anonymized data for future research.

universal joy classifying emotions results for classifying الفرح العالمي تصنيف العواطف نتائج لتصنيف صناعة حمض الفوسفور المزيد..

Gender Bias in Natural Language Processing Across Human Languages

383 - Association for Computation Linguistics 2021 مقالة

Natural Language Processing (NLP) systems are at the heart of many critical automated decision-making systems making crucial recommendations about our future world. Gender bias in NLP has been well studied in English, but has been less studied in oth er languages. In this paper, a team including speakers of 9 languages - Chinese, Spanish, English, Arabic, German, French, Farsi, Urdu, and Wolof - reports and analyzes measurements of gender bias in the Wikipedia corpora for these 9 languages. We develop extensions to profession-level and corpus-level gender bias metric calculations originally designed for English and apply them to 8 other languages, including languages that have grammatically gendered nouns including different feminine, masculine, and neuter profession words. We discuss future work that would benefit immensely from a computational linguistics perspective.

مشكلة تقسيم زمرة language processing human languages معالجة اللغة لغات بشرية صناعة حمض الفوسفور

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Subword Mapping and Anchoring across Languages

رسم الخرائط الفرعية ومثبتة عبر اللغات

Ask ChatGPT about the research

Read More

suggested questions