New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Orthographic vs. Semantic Representations for Unsupervised Morphological Paradigm Clustering

عروض إكسبروغرافية مقابل التموين الدلالي لإلقاء المجموعة المورفولوجية غير الخاضعة للكشف

211 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

تجميع النموذج morphological paradigm نموذج مورفولوجي صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

تعرض هذه الورقة أنظمة مختلفة لمجموعة مختلفة من النماذج المورفولوجية، في سياق المهمة المشتركة Sigmorphon 2021 2. الهدف من هذه المهمة هو تصحيح الكلمات العنقودية بشكل صحيح بلغة معينة من قبل نموذج اندلاطها، دون أي معرفة سابقة باللغة وبدون إشراف من البيانات المسمى لأي فرز. تعد الكلمات الموجودة في النموذج المورفولوجي الواحد بمتغيرات انتشار مختلفة من ليمما الأساسي، مما يعني أن الكلمات تشترك في معنى أساسي مشترك. كما أنها - عادة - تظهر درجة عالية من التشابه الجبادي. بعد حدس هذه الحدس، نحقق في تجميع كماينز باستخدام نوعين مختلفين من تمثيلات الكلمات: يركز المرء على التشابه الجبائي والتركيز الآخر على التشابه الدلالي. يتم تحديد الأدوار الوسطى المحددة مسبقا بناء على وجود خوارزمية فرعية مشتركة عادية أو طريقة رسم بيانية متصلة مبنية بأطول فرعية شائعة. بالنسبة لجميع لغات التطوير، فإن المدينات القائمة على الطابع تؤدي بالمثل إلى خط الأساس، وتشير المبدأ الدوالي أداء أقل بكثير من خط الأساس إلى أن أخطاء النظم تشير إلى أن التجميع القائم على تمثيلات إلكترونية مناسبة لمجموعة واسعة من الآليات المورفولوجية، لا سيما كجزء من نظام أكبر.

This paper presents two different systems for unsupervised clustering of morphological paradigms, in the context of the SIGMORPHON 2021 Shared Task 2. The goal of this task is to correctly cluster words in a given language by their inflectional paradigm, without any previous knowledge of the language and without supervision from labeled data of any sort. The words in a single morphological paradigm are different inflectional variants of an underlying lemma, meaning that the words share a common core meaning. They also - usually - show a high degree of orthographical similarity. Following these intuitions, we investigate KMeans clustering using two different types of word representations: one focusing on orthographical similarity and the other focusing on semantic similarity.Additionally, we discuss the merits of randomly initialized centroids versus pre-defined centroids for clustering. Pre-defined centroids are identified based on either a standard longest common substring algorithm or a connected graph method built off of longest common substring. For all development languages, the character-based embeddings perform similarly to the baseline, and the semantic embeddings perform well below the baseline.Analysis of the systems' errors suggests that clustering based on orthographic representations is suitable for a wide range of morphological mechanisms, particularly as part of a larger system.

References used

https://aclanthology.org/

rate research

Unsupervised Paradigm Clustering Using Transformation Rules

238 - Association for Computation Linguistics 2021 مقالة

This paper describes the submission of the CU-UBC team for the SIGMORPHON 2021 Shared Task 2: Unsupervised morphological paradigm clustering. Our system generates paradigms using morphological transformation rules which are discovered from raw data. We experiment with two methods for discovering rules. Our first approach generates prefix and suffix transformations between similar strings. Secondly, we experiment with more general rules which can apply transformations inside the input strings in addition to prefix and suffix transformations. We find that the best overall performance is delivered by prefix and suffix rules but more general transformation rules perform better for languages with templatic morphology and very high morpheme-to-word ratios.

نموذج مورفولوجي صناعة حمض الفوسفور

Adaptor Grammars for Unsupervised Paradigm Clustering

218 - Association for Computation Linguistics 2021 مقالة

This work describes the Edinburgh submission to the SIGMORPHON 2021 Shared Task 2 on unsupervised morphological paradigm clustering. Given raw text input, the task was to assign each token to a cluster with other tokens from the same paradigm. We use Adaptor Grammar segmentations combined with frequency-based heuristics to predict paradigm clusters. Our system achieved the highest average F1 score across 9 test languages, placing first out of 15 submissions.

unsupervised paradigm clustering paradigm clustering تجميع النموذج غير المنضح تجميع النموذج صناعة حمض الفوسفور

Findings of the SIGMORPHON 2021 Shared Task on Unsupervised Morphological Paradigm Clustering

328 - Association for Computation Linguistics 2021 مقالة

We describe the second SIGMORPHON shared task on unsupervised morphology: the goal of the SIGMORPHON 2021 Shared Task on Unsupervised Morphological Paradigm Clustering is to cluster word types from a raw text corpus into paradigms. To this end, we re lease corpora for 5 development and 9 test languages, as well as gold partial paradigms for evaluation. We receive 14 submissions from 4 teams that follow different strategies, and the best performing system is based on adaptor grammars. Results vary significantly across languages. However, all systems are outperformed by a supervised lemmatizer, implying that there is still room for improvement.

unsupervised morphological paradigm morphological paradigm clustering sigmorphon shared task النموذج المورفولوجي غير المدخري تجميع النماذج المورفولوجية Sigmorphon المهمة المشتركة صناعة حمض الفوسفور المزيد..

Text Document Clustering: Wordnet vs. TF-IDF vs. Word Embeddings

359 - Association for Computation Linguistics 2021 مقالة

In the paper, we deal with the problem of unsupervised text document clustering for the Polish language. Our goal is to compare the modern approaches based on language modeling (doc2vec and BERT) with the classical ones, i.e., TF-IDF and wordnet-base d. The experiments are conducted on three datasets containing qualification descriptions. The experiments' results showed that wordnet-based similarity measures could compete and even outperform modern embedding-based approaches.

text document clustering document clustering تجميع مستند النص تجميع المستندات صناعة حمض الفوسفور

Unsupervised Paraphrasing Consistency Training for Low Resource Named Entity Recognition

298 - Association for Computation Linguistics 2021 مقالة

Unsupervised consistency training is a way of semi-supervised learning that encourages consistency in model predictions between the original and augmented data. For Named Entity Recognition (NER), existing approaches augment the input sequence with t oken replacement, assuming annotations on the replaced positions unchanged. In this paper, we explore the use of paraphrasing as a more principled data augmentation scheme for NER unsupervised consistency training. Specifically, we convert Conditional Random Field (CRF) into a multi-label classification module and encourage consistency on the entity appearance between the original and paraphrased sequences. Experiments show that our method is especially effective when annotations are limited.

low resource named resource named entity الموارد المنخفضة اسمه الكيان المسمى الموارد صناعة حمض الفوسفور

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Orthographic vs. Semantic Representations for Unsupervised Morphological Paradigm Clustering

عروض إكسبروغرافية مقابل التموين الدلالي لإلقاء المجموعة المورفولوجية غير الخاضعة للكشف

Ask ChatGPT about the research

Read More

suggested questions