New community

Subscribe to the gold package and get unlimited access to Shamra Academy

TITA: A Two-stage Interaction and Topic-Aware Text Matching Model

تيتا: نموذج تفاعل ذو مرحلتين ومطابقة النص

228 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

topic-aware text matching two-stage interaction text matching model وضع موضوع إعلام النص التفاعل مرحلتين نموذج مطابقة النص صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

في هذه الورقة، نركز على مشكلة الكلمات الرئيسية ومطابقة المستندات من خلال النظر في مستويات ذات صلة مختلفة. في نظام توصيتنا، يتبع أشخاص مختلفون الكلمات الرئيسية الساخنة المختلفة باهتمام. نحتاج إلى إرفاق المستندات إلى كل كلمة رئيسية ثم توزيع المستندات على الأشخاص الذين يتبعون هذه الكلمات الرئيسية. يجب أن تحتوي المستندات المثالية على نفس الموضوع مع الكلمة الأساسية، والتي نسميها ذات أهمية تدرك الموضوع. بمعنى آخر، وثائق الأهمية ذات الصلة بالموضوع أفضل من تلك الأهمية جزئيا في هذا التطبيق. ومع ذلك، فإن المهام السابقة لا تحدد أبدا أهمية علم الموضوع بوضوح. لمعالجة هذه المشكلة، نحدد صلة ثلاثية المستوى بمهمة مطابقة الوثيقة للكلمة الرئيسية: الأهمية ذات الصلة بالموضوع، والأهمية جزئيا والأهمية. لالتقاط الأهمية بين الكلمة الرئيسية القصيرة والوثيقة في المستويات الثلاثة المذكورة أعلاه، لا ينبغي لنا الجمع بين الموضوع الكامن فقط من الوثيقة بتمثيلها العصبي العميق، ولكن أيضا التفاعلات المعقدة النموذجية بين الكلمة الرئيسية والوثيقة. تحقيقا لهذه الغاية، نقترح نموذجا متطابقا على تفاعل ثنائي مرحلتين ومطابقة النص (TITA). من حيث الموضوع - أدرك "، نقدم نموذج موضوع عصبي لتحليل موضوع المستند ثم استخدامه لمزيد من تشفير المستند. من حيث التفاعل من مرحلتين "، نقترح مراحل متتالية لنموذج التفاعلات المعقدة بين الكلمة الرئيسية والوثيقة. تكشف التجارب الواسعة أن تيتا تفوقت على خطوط الأساس الأخرى المصممة بشكل جيد وتظهر أداء ممتاز في نظام توصيتنا.

In this paper, we focus on the problem of keyword and document matching by considering different relevance levels. In our recommendation system, different people follow different hot keywords with interest. We need to attach documents to each keyword and then distribute the documents to people who follow these keywords. The ideal documents should have the same topic with the keyword, which we call topic-aware relevance. In other words, topic-aware relevance documents are better than partially-relevance ones in this application. However, previous tasks never define topic-aware relevance clearly. To tackle this problem, we define a three-level relevance in keyword-document matching task: topic-aware relevance, partially-relevance and irrelevance. To capture the relevance between the short keyword and the document at above-mentioned three levels, we should not only combine the latent topic of the document with its deep neural representation, but also model complex interactions between the keyword and the document. To this end, we propose a Two-stage Interaction and Topic-Aware text matching model (TITA). In terms of topic-aware'', we introduce neural topic model to analyze the topic of the document and then use it to further encode the document. In terms of two-stage interaction'', we propose two successive stages to model complex interactions between the keyword and the document. Extensive experiments reveal that TITA outperforms other well-designed baselines and shows excellent performance in our recommendation system.

References used

https://aclanthology.org/

rate research

Neural Attention-Aware Hierarchical Topic Model

297 - Association for Computation Linguistics 2021 مقالة

Neural topic models (NTMs) apply deep neural networks to topic modelling. Despite their success, NTMs generally ignore two important aspects: (1) only document-level word count information is utilized for the training, while more fine-grained sentenc e-level information is ignored, and (2) external semantic knowledge regarding documents, sentences and words are not exploited for the training. To address these issues, we propose a variational autoencoder (VAE) NTM model that jointly reconstructs the sentence and document word counts using combinations of bag-of-words (BoW) topical embeddings and pre-trained semantic embeddings. The pre-trained embeddings are first transformed into a common latent topical space to align their semantics with the BoW embeddings. Our model also features hierarchical KL divergence to leverage embeddings of each document to regularize those of their sentences, paying more attention to semantically relevant sentences. Both quantitative and qualitative experiments have shown the efficacy of our model in 1) lowering the reconstruction errors at both the sentence and document levels, and 2) discovering more coherent topics from real-world datasets.

attention-aware hierarchical topic neural attention-aware hierarchical الانتباه تدرك موضوع هرمي الاهتمام العصبي يدرك التسلسل الهرمي صناعة حمض الفوسفور

Context-Aware Interaction Network for Question Matching

409 - Association for Computation Linguistics 2021 مقالة

Impressive milestones have been achieved in text matching by adopting a cross-attention mechanism to capture pertinent semantic connections between two sentence representations. However, regular cross-attention focuses on word-level links between the two input sequences, neglecting the importance of contextual information. We propose a context-aware interaction network (COIN) to properly align two sequences and infer their semantic relationship. Specifically, each interaction block includes (1) a context-aware cross-attention mechanism to effectively integrate contextual information when aligning two sequences, and (2) a gate fusion layer to flexibly interpolate aligned representations. We apply multiple stacked interaction blocks to produce alignments at different levels and gradually refine the attention results. Experiments on two question matching datasets and detailed analyses demonstrate the effectiveness of our model.

context-aware interaction network question matching interaction network شبكة التفاعل السياق مطابقة السؤال شبكة التفاعل صناعة حمض الفوسفور المزيد..

LAMAD: A Linguistic Attentional Model for Arabic Text Diacritization

702 - Association for Computation Linguistics 2021 مقالة

In Arabic Language, diacritics are used to specify meanings as well as pronunciations. However, diacritics are often omitted from written texts, which increases the number of possible meanings and pronunciations. This leads to an ambiguous text and m akes the computational process on undiacritized text more difficult. In this paper, we propose a Linguistic Attentional Model for Arabic text Diacritization (LAMAD). In LAMAD, a new linguistic feature representation is presented, which utilizes both word and character contextual features. Then, a linguistic attention mechanism is proposed to capture the important linguistic features. In addition, we explore the impact of the linguistic features extracted from the text on Arabic text diacritization (ATD) by introducing them to the linguistic attention mechanism. The extensive experimental results on three datasets with different sizes illustrate that LAMAD outperforms the existing state-of-the-art models.

arabic text diacritization linguistic attentional model تشكيل النص العربي نموذج الاهتمام اللغوي صناعة حمض الفوسفور

Not All Comments Are Equal: Insights into Comment Moderation from a Topic-Aware Model

355 - Association for Computation Linguistics 2021 مقالة

Moderation of reader comments is a significant problem for online news platforms. Here, we experiment with models for automatic moderation, using a dataset of comments from a popular Croatian newspaper. Our analysis shows that while comments that vio late the moderation rules mostly share common linguistic and thematic features, their content varies across the different sections of the newspaper. We therefore make our models topic-aware, incorporating semantic features from a topic model into the classification decision. Our results show that topic information improves the performance of the model, increases its confidence in correct outputs, and helps us understand the model's outputs.

equal comments are equal متساوي التعليقات مساوية صناعة حمض الفوسفور

HTCInfoMax: A Global Model for Hierarchical Text Classification via Information Maximization

339 - Association for Computation Linguistics 2021 مقالة

The current state-of-the-art model HiAGM for hierarchical text classification has two limitations. First, it correlates each text sample with all labels in the dataset which contains irrelevant information. Second, it does not consider any statistica l constraint on the label representations learned by the structure encoder, while constraints for representation learning are proved to be helpful in previous work. In this paper, we propose HTCInfoMax to address these issues by introducing information maximization which includes two modules: text-label mutual information maximization and label prior matching. The first module can model the interaction between each text sample and its ground truth labels explicitly which filters out irrelevant information. The second one encourages the structure encoder to learn better representations with desired characteristics for all labels which can better handle label imbalance in hierarchical text classification. Experimental results on two benchmark datasets demonstrate the effectiveness of the proposed HTCInfoMax.

hierarchical text classification hierarchical text تصنيف النص الهرمي النص الهرمي صناعة حمض الفوسفور

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

TITA: A Two-stage Interaction and Topic-Aware Text Matching Model

تيتا: نموذج تفاعل ذو مرحلتين ومطابقة النص

Ask ChatGPT about the research

Read More

suggested questions