New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Incremental Few-shot Text Classification with Multi-round New Classes: Formulation, Dataset and System

تصنيف نص قليل قصير بالرصاص مع فصول جديدة متعددة جولات: صياغة ومجموعات البيانات والنظام

237 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

عادة ما تتم دراسة تصنيف النص عن طريق وضع علامات نصوص اللغة الطبيعية مع الفئات ذات الصلة من مجموعة محددة مسبقا. في العالم الحقيقي، قد تستمر فصول جديدة في تحدي النظام الحالي مع بيانات محدودة المسمى. يجب أن يكون النظام ذكي بما يكفي للتعرف على الطبقات الجديدة القادمة مع بعض الأمثلة. في هذا العمل، نحدد مهمة جديدة في مجال NLP، تصنيف النص قليل الطوابق الإضافي، حيث يتعامل النظام تدريجيا جولات متعددة من الفصول الجديدة. لكل جولة، هناك مجموعة من الطبقات الجديدة مع بعض الأمثلة المسمى لكل فصل. يوجد تحديان رئيسيان في هذه المهمة الجديدة: (1) لعملية التعلم، يجب أن يتعلم النظام تدريجيا على جولة فصول جديدة جولة من الجولة دون إعادة التدريب على الأمثلة على الطبقات السابقة؛ (2) بالنسبة للأداء، يجب أن يؤدي النظام بشكل جيد على فئات جديدة دون فقدان الكثير في الفصول السابقة. بالإضافة إلى صياغة المهمة الجديدة، نقوم أيضا بإصدار مجموعة بيانات قياسية في الإعداد القليل من الرصاص الإضافي: تصنيف النوايا وتصنيف العلاقات. علاوة على ذلك، نقترح اثنين مناهج استقصاء وتتبعها والجاذبية، والتي تظهر الوعد بحل هذه المشكلة الرواية.

Text classification is usually studied by labeling natural language texts with relevant categories from a predefined set. In the real world, new classes might keep challenging the existing system with limited labeled data. The system should be intelligent enough to recognize upcoming new classes with a few examples. In this work, we define a new task in the NLP domain, incremental few-shot text classification, where the system incrementally handles multiple rounds of new classes. For each round, there is a batch of new classes with a few labeled examples per class. Two major challenges exist in this new task: (i) For the learning process, the system should incrementally learn new classes round by round without re-training on the examples of preceding classes; (ii) For the performance, the system should perform well on new classes without much loss on preceding classes. In addition to formulating the new task, we also release two benchmark datasets in the incremental few-shot setting: intent classification and relation classification. Moreover, we propose two entailment approaches, ENTAILMENT and HYBRID, which show promise for solving this novel problem.

References used

https://aclanthology.org/

rate research

TransPrompt: Towards an Automatic Transferable Prompting Framework for Few-shot Text Classification

362 - Association for Computation Linguistics 2021 مقالة

Recent studies have shown that prompts improve the performance of large pre-trained language models for few-shot text classification. Yet, it is unclear how the prompting knowledge can be transferred across similar NLP tasks for the purpose of mutual reinforcement. Based on continuous prompt embeddings, we propose TransPrompt, a transferable prompting framework for few-shot learning across similar tasks. In TransPrompt, we employ a multi-task meta-knowledge acquisition procedure to train a meta-learner that captures cross-task transferable knowledge. Two de-biasing techniques are further designed to make it more task-agnostic and unbiased towards any tasks. After that, the meta-learner can be adapted to target tasks with high accuracy. Extensive experiments show that TransPrompt outperforms single-task and cross-task strong baselines over multiple NLP tasks and datasets. We further show that the meta-learner can effectively improve the performance on previously unseen tasks; and TransPrompt also outperforms strong fine-tuning baselines when learning with full training sets.

شبكة مزدوجة متزامن automatic transferable prompting التحويل التلقائي المطالبة صناعة حمض الفوسفور

Few-Shot Text Generation with Natural Language Instructions

294 - Association for Computation Linguistics 2021 مقالة

Providing pretrained language models with simple task descriptions in natural language enables them to solve some tasks in a fully unsupervised fashion. Moreover, when combined with regular learning from examples, this idea yields impressive few-shot results for a wide range of text classification tasks. It is also a promising direction to improve data efficiency in generative settings, but there are several challenges to using a combination of task descriptions and example-based learning for text generation. In particular, it is crucial to find task descriptions that are easy to understand for the pretrained model and to ensure that it actually makes good use of them; furthermore, effective measures against overfitting have to be implemented. In this paper, we show how these challenges can be tackled: We introduce GenPET, a method for text generation that is based on pattern-exploiting training, a recent approach for combining textual instructions with supervised learning that only works for classification tasks. On several summarization and headline generation datasets, GenPET gives consistent improvements over strong baselines in few-shot settings.

ضغط نموذج اللغة natural language enables اللغة الطبيعية تمكن صناعة حمض الفوسفور

Few-shot Intent Classification and Slot Filling with Retrieved Examples

320 - Association for Computation Linguistics 2021 مقالة

Few-shot learning arises in important practical scenarios, such as when a natural language understanding system needs to learn new semantic labels for an emerging, resource-scarce domain. In this paper, we explore retrieval-based methods for intent c lassification and slot filling tasks in few-shot settings. Retrieval-based methods make predictions based on labeled examples in the retrieval index that are similar to the input, and thus can adapt to new domains simply by changing the index without having to retrain the model. However, it is non-trivial to apply such methods on tasks with a complex label space like slot filling. To this end, we propose a span-level retrieval method that learns similar contextualized representations for spans with the same label via a novel batch-softmax objective. At inference time, we use the labels of the retrieved spans to construct the final structure with the highest aggregated score. Our method outperforms previous systems in various few-shot settings on the CLINC and SNIPS benchmarks.

few-shot intent classification تصنيف القليل من الطلقات صناعة حمض الفوسفور

Continual Few-Shot Learning for Text Classification

557 - Association for Computation Linguistics 2021 مقالة

Natural Language Processing (NLP) is increasingly relying on general end-to-end systems that need to handle many different linguistic phenomena and nuances. For example, a Natural Language Inference (NLI) system has to recognize sentiment, handle num bers, perform coreference, etc. Our solutions to complex problems are still far from perfect, so it is important to create systems that can learn to correct mistakes quickly, incrementally, and with little training data. In this work, we propose a continual few-shot learning (CFL) task, in which a system is challenged with a difficult phenomenon and asked to learn to correct mistakes with only a few (10 to 15) training examples. To this end, we first create benchmarks based on previously annotated data: two NLI (ANLI and SNLI) and one sentiment analysis (IMDB) datasets. Next, we present various baselines from diverse paradigms (e.g., memory-aware synapses and Prototypical networks) and compare them on few-shot learning and continual few-shot learning setups. Our contributions are in creating a benchmark suite and evaluation protocol for continual few-shot learning on the text classification tasks, and making several interesting observations on the behavior of similarity-based methods. We hope that our work serves as a useful starting point for future work on this important topic.

نقل متبرع عبر اللغات continual few-shot learning عدد قليل من التعلم قليلا صناعة حمض الفوسفور

MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer

382 - Association for Computation Linguistics 2021 مقالة

We introduce MULTI-EURLEX, a new multilingual dataset for topic classification of legal documents. The dataset comprises 65k European Union (EU) laws, officially translated in 23 languages, annotated with multiple labels from the EUROVOC taxonomy. We highlight the effect of temporal concept drift and the importance of chronological, instead of random splits. We use the dataset as a testbed for zero-shot cross-lingual transfer, where we exploit annotated training documents in one language (source) to classify documents in another language (target). We find that fine-tuning a multilingually pretrained model (XLM-ROBERTA, MT5) in a single source language leads to catastrophic forgetting of multilingual knowledge and, consequently, poor zero-shot transfer to other languages. Adaptation strategies, namely partial fine-tuning, adapters, BITFIT, LNFIT, originally proposed to accelerate fine-tuning for new end-tasks, help retain multilingual knowledge from pretraining, substantially improving zero-shot cross-lingual transfer, but their impact also depends on the pretrained model used and the size of the label set.

multi-label legal document multi-lingual and multi-label وثيقة قانونية متعددة العلامات متعدد اللغات ومتعددة التسمية صناعة حمض الفوسفور

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Incremental Few-shot Text Classification with Multi-round New Classes: Formulation, Dataset and System

تصنيف نص قليل قصير بالرصاص مع فصول جديدة متعددة جولات: صياغة ومجموعات البيانات والنظام

Ask ChatGPT about the research

Read More

suggested questions