New community

Subscribe to the gold package and get unlimited access to Shamra Academy

WMDecompose: A Framework for Leveraging the Interpretable Properties of Word Mover's Distance in Sociocultural Analysis

WMDECompose: إطارا للاستفادة من الخصائص القابلة للتفسير لمسافة Word Mover في التحليل الاجتماعي الثقافي

181 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

framework for leveraging word mover distance sociocultural analysis إطار للاستفادة كلمة المحرك المسافة التحليل الاجتماعي الثقافي صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

على الرغم من أن شعبية متزايدة من NLP في العلوم الإنسانية والعلوم الاجتماعية، فقد ترافق التقدم في الأداء النموذجي وتعقيد مخاوف بشأن التفسير والسلطة التوضيحية للتحليل الاجتماعي الثقافي. نموذج شعبي واحد يأخذ طريقا وسط مسافة كلمة المحرك (WMD). يتم تكييفها ظاهريا لتفسيرها، ومع ذلك تم استخدام WMD وتم تطويره بشكل أكبر بطرق تجاهل الجانب الأكثر تفسيرا في كثير من الأحيان: أي مسافات مستوى الكلمات المطلوبة لترجمة مجموعة من الكلمات إلى مجموعة أخرى من الكلمات. لمعالجة هذه الفجوة الواضحة، نقدم WMDECOMPOOPE: مكتبة نموذجية ومكتبة بيثون 1) تتحلل مسافات مستوى المستند في المسافات في مستوياتها المكونة على مستوى الكلمات، و 2) مجموعات في وقت لاحق من تحفيز العناصر المواضيعية، بحيث يتم الاحتفاظ بالمعلومات المعجمية المفيدة تلخيص للتحليل. لتوضيح إمكاناتها في سياق علمي اجتماعي، نطبقها على جثة وسائل التواصل الاجتماعي الطولية لاستكشاف العلاقة المتبادلة بين نظريات المؤامرة والأحرفات الأمريكية المحافظة. أخيرا، نظرا لتعقيد الوقت الكامل في الوقت الحالي، فإننا نقترح بالإضافة إلى طريقة لأخذ عينات من مجموعات البيانات الكبيرة بطريقة استنساخ، مع حدود ضيقة تمنع استقراء النتائج غير الموثوقة بسبب سوء أخذ العينات الممارسات.

Despite the increasing popularity of NLP in the humanities and social sciences, advances in model performance and complexity have been accompanied by concerns about interpretability and explanatory power for sociocultural analysis. One popular model that takes a middle road is Word Mover's Distance (WMD). Ostensibly adapted for its interpretability, WMD has nonetheless been used and further developed in ways which frequently discard its most interpretable aspect: namely, the word-level distances required for translating a set of words into another set of words. To address this apparent gap, we introduce WMDecompose: a model and Python library that 1) decomposes document-level distances into their constituent word-level distances, and 2) subsequently clusters words to induce thematic elements, such that useful lexical information is retained and summarized for analysis. To illustrate its potential in a social scientific context, we apply it to a longitudinal social media corpus to explore the interrelationship between conspiracy theories and conservative American discourses. Finally, because of the full WMD model's high time-complexity, we additionally suggest a method of sampling document pairs from large datasets in a reproducible way, with tight bounds that prevent extrapolation of unreliable results due to poor sampling practices.

References used

https://aclanthology.org/

rate research

Interpretable Propaganda Detection in News Articles

266 - Association for Computation Linguistics 2021 مقالة

Online users today are exposed to misleading and propagandistic news articles and media posts on a daily basis. To counter thus, a number of approaches have been designed aiming to achieve a healthier and safer online news and media consumption. Auto matic systems are able to support humans in detecting such content; yet, a major impediment to their broad adoption is that besides being accurate, the decisions of such systems need also to be interpretable in order to be trusted and widely adopted by users. Since misleading and propagandistic content influences readers through the use of a number of deception techniques, we propose to detect and to show the use of such techniques as a way to offer interpretability. In particular, we define qualitatively descriptive features and we analyze their suitability for detecting deception techniques. We further show that our interpretable features can be easily combined with pre-trained language models, yielding state-of-the-art results.

interpretable propaganda detection interpretable propaganda كشف الدعاية القابلة للتفسير الدعاية القابلة للتفسير صناعة حمض الفوسفور

Monitoring geometrical properties of word embeddings for detecting the emergence of new topics.

351 - Association for Computation Linguistics 2021 مقالة

Slow emerging topic detection is a task between event detection, where we aggregate behaviors of different words on short period of time, and language evolution, where we monitor their long term evolution. In this work, we tackle the problem of early detection of slowly emerging new topics. To this end, we gather evidence of weak signals at the word level. We propose to monitor the behavior of words representation in an embedding space and use one of its geometrical properties to characterize the emergence of topics. As evaluation is typically hard for this kind of task, we present a framework for quantitative evaluation and show positive results that outperform state-of-the-art methods. Our method is evaluated on two public datasets of press and scientific articles.

monitoring geometrical properties emerging topic detection slow emerging topic رصد خصائص هندسية اكتشاف موضوع الناشئة. موضوع الناشئة البطيء صناعة حمض الفوسفور المزيد..

Variation in framing as a function of temporal reporting distance

225 - Association for Computation Linguistics 2021 مقالة

In this paper, we measure variation in framing as a function of foregrounding and backgrounding in a co-referential corpus with a range of temporal distance. In one type of experiment, frame-annotated corpora grouped under event types were contrasted , resulting in a ranking of frames with typicality rates. In contrasting between publication dates, a different ranking of frames emerged for documents that are close to or far from the event instance. In the second type of analysis, we trained a diagnostic classifier with frame occurrences in order to let it differentiate documents based on their temporal distance class (close to or far from the event instance). The classifier performs above chance and outperforms models with words.

temporal reporting distance variation in framing temporal distance المسافة التقارير الزمنية الاختلاف في تأييد المسافة الزمنية صناعة حمض الفوسفور المزيد..

Field Embedding: A Unified Grain-Based Framework for Word Representation

605 - Association for Computation Linguistics 2021 مقالة

Word representations empowered with additional linguistic information have been widely studied and proved to outperform traditional embeddings. Current methods mainly focus on learning embeddings for words while embeddings of linguistic information ( referred to as grain embeddings) are discarded after the learning. This work proposes a framework field embedding to jointly learn both word and grain embeddings by incorporating morphological, phonetic, and syntactical linguistic fields. The framework leverages an innovative fine-grained pipeline that integrates multiple linguistic fields and produces high-quality grain sequences for learning supreme word representations. A novel algorithm is also designed to learn embeddings for words and grains by capturing information that is contained within each field and that is shared across them. Experimental results of lexical tasks and downstream natural language processing tasks illustrate that our framework can learn better word embeddings and grain embeddings. Qualitative evaluations show grain embeddings effectively capture the semantic information.

unified grain-based framework unified grain-based word representations الإطار الموحد القائمة على الحبوب القائم على الحبوب الموحدة تمثيلات كلمة صناعة حمض الفوسفور المزيد..

Neural Natural Logic Inference for Interpretable Question Answering

309 - Association for Computation Linguistics 2021 مقالة

Many open-domain question answering problems can be cast as a textual entailment task, where a question and candidate answers are concatenated to form hypotheses. A QA system then determines if the supporting knowledge bases, regarded as potential pr emises, entail the hypotheses. In this paper, we investigate a neural-symbolic QA approach that integrates natural logic reasoning within deep learning architectures, towards developing effective and yet explainable question answering models. The proposed model gradually bridges a hypothesis and candidate premises following natural logic inference steps to build proof paths. Entailment scores between the acquired intermediate hypotheses and candidate premises are measured to determine if a premise entails the hypothesis. As the natural logic reasoning process forms a tree-like, hierarchical structure, we embed hypotheses and premises in a Hyperbolic space rather than Euclidean space to acquire more precise representations. Empirically, our method outperforms prior work on answering multiple-choice science questions, achieving the best results on two publicly available datasets. The natural logic inference process inherently provides evidence to help explain the prediction process.

interpretable question answering natural logic inference natural logic استجابة سؤال مفسر الاستدلال المنطقي الطبيعي المنطق الطبيعي صناعة حمض الفوسفور المزيد..

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

WMDecompose: A Framework for Leveraging the Interpretable Properties of Word Mover's Distance in Sociocultural Analysis

WMDECompose: إطارا للاستفادة من الخصائص القابلة للتفسير لمسافة Word Mover في التحليل الاجتماعي الثقافي

Ask ChatGPT about the research

Read More

suggested questions