New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Benchmarking Commercial Intent Detection Services with Practice-Driven Evaluations

معيار خدمات الكشف عن النية التجارية مع تقييمات مدفوعة بالممارسة

375 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

intent detection services intent detection practice-driven evaluations خدمات الكشف عن النية الكشف عن النية تقييمات مدفوعة بالممارسة صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

يعد الكشف عن النية مكونا رئيسيا في أنظمة الحوار الحديثة الموجهة نحو الأهداف التي تنجز مهمة مستخدم من خلال التنبؤ بمثابة إيداع نص المستخدمين. هناك ثلاثة تحديات أساسية في تصميم نماذج الكشف عن النية قوية ودقيقة. أولا، تتطلب نماذج الكشف عن النية النموذجية كمية كبيرة من البيانات المسمى لتحقيق دقة عالية. لسوء الحظ، في السيناريوهات العملية هو أكثر شيوعا للعثور على مجموعات بيانات صغيرة وغير متوازنة وصاخبة. ثانيا، حتى مع بيانات تدريب كبيرة، يمكن أن ترى نماذج الكشف عن النية توزيعا مختلفا لبيانات الاختبار عند نشرها في العالم الحقيقي، مما يؤدي إلى دقة سيئة. أخيرا، يجب أن يكون نموذج اكتشاف نوايا عمليا فعاليا في كل من التدريب واستنتاج الاستعلام الفردي بحيث يمكن استخدامه بشكل مستمر وإعادة تدريبه بشكل متكرر. نحن نؤيد أساليب الكشف عن النية في مجموعة متنوعة من مجموعات البيانات. تظهر نتائجنا أن نموذج الكشف عن نية مساعد Watson يفوق الحلول التجارية الأخرى ومقارنة مع نماذج اللغة المحددة مسبقا كبيرة مع حدوث جزء صغير فقط من الموارد الحسابية وبيانات التدريب. يدل مساعد واتسون درجة أعلى من المتانة عند تختلف توزيعات التدريب والاختبار.

Intent detection is a key component of modern goal-oriented dialog systems that accomplish a user task by predicting the intent of users' text input. There are three primary challenges in designing robust and accurate intent detection models. First, typical intent detection models require a large amount of labeled data to achieve high accuracy. Unfortunately, in practical scenarios it is more common to find small, unbalanced, and noisy datasets. Secondly, even with large training data, the intent detection models can see a different distribution of test data when being deployed in the real world, leading to poor accuracy. Finally, a practical intent detection model must be computationally efficient in both training and single query inference so that it can be used continuously and re-trained frequently. We benchmark intent detection methods on a variety of datasets. Our results show that Watson Assistant's intent detection model outperforms other commercial solutions and is comparable to large pretrained language models while requiring only a fraction of computational resources and training data. Watson Assistant demonstrates a higher degree of robustness when the training and test distributions differ.

References used

https://aclanthology.org/

rate research

Example-Driven Intent Prediction with Observers

271 - Association for Computation Linguistics 2021 مقالة

A key challenge of dialog systems research is to effectively and efficiently adapt to new domains. A scalable paradigm for adaptation necessitates the development of generalizable models that perform well in few-shot settings. In this paper, we focus on the intent classification problem which aims to identify user intents given utterances addressed to the dialog system. We propose two approaches for improving the generalizability of utterance classification models: (1) observers and (2) example-driven training. Prior work has shown that BERT-like models tend to attribute a significant amount of attention to the [CLS] token, which we hypothesize results in diluted representations. Observers are tokens that are not attended to, and are an alternative to the [CLS] token as a semantic representation of utterances. Example-driven training learns to classify utterances by comparing to examples, thereby using the underlying encoder as a sentence similarity model. These methods are complementary; improving the representation through observers allows the example-driven model to better measure sentence similarities. When combined, the proposed methods attain state-of-the-art results on three intent prediction datasets (banking77, clinc150, hwu64) in both the full data and few-shot (10 examples per intent) settings. Furthermore, we demonstrate that the proposed approach can transfer to new intents and across datasets without any additional training.

intent prediction observers تنبؤ النية المراقبين صناعة حمض الفوسفور

Multilingual and Cross-Lingual Intent Detection from Spoken Data

748 - Association for Computation Linguistics 2021 مقالة

We present a systematic study on multilingual and cross-lingual intent detection (ID) from spoken data. The study leverages a new resource put forth in this work, termed MInDS-14, a first training and evaluation resource for the ID task with spoken d ata. It covers 14 intents extracted from a commercial system in the e-banking domain, associated with spoken examples in 14 diverse language varieties. Our key results indicate that combining machine translation models with state-of-the-art multilingual sentence encoders (e.g., LaBSE) yield strong intent detectors in the majority of target languages covered in MInDS-14, and offer comparative analyses across different axes: e.g., translation direction, impact of speech recognition, data augmentation from a related domain. We see this work as an important step towards more inclusive development and evaluation of multilingual ID from spoken data, hopefully in a much wider spectrum of languages compared to prior work.

cross-lingual intent detection spoken data الكشف عن النية عبر اللغات البيانات المنطوقة صناعة حمض الفوسفور

Dialogue State Tracking with a Language Model using Schema-Driven Prompting

323 - Association for Computation Linguistics 2021 مقالة

Task-oriented conversational systems often use dialogue state tracking to represent the user's intentions, which involves filling in values of pre-defined slots. Many approaches have been proposed, often using task-specific architectures with special -purpose classifiers. Recently, good results have been obtained using more general architectures based on pretrained language models. Here, we introduce a new variation of the language modeling approach that uses schema-driven prompting to provide task-aware history encoding that is used for both categorical and non-categorical slots. We further improve performance by augmenting the prompting with schema descriptions, a naturally occurring source of in-domain knowledge. Our purely generative system achieves state-of-the-art performance on MultiWOZ 2.2 and achieves competitive performance on two other benchmarks: MultiWOZ 2.1 and M2M. The data and code will be available at https://github.com/chiahsuan156/DST-as-Prompting.

الحقول العشوائية صناعة حمض الفوسفور

Benchmarking Transformer-based Language Models for Arabic Sentiment and Sarcasm Detection

322 - Association for Computation Linguistics 2021 مقالة

The introduction of transformer-based language models has been a revolutionary step for natural language processing (NLP) research. These models, such as BERT, GPT and ELECTRA, led to state-of-the-art performance in many NLP tasks. Most of these mode ls were initially developed for English and other languages followed later. Recently, several Arabic-specific models started emerging. However, there are limited direct comparisons between these models. In this paper, we evaluate the performance of 24 of these models on Arabic sentiment and sarcasm detection. Our results show that the models achieving the best performance are those that are trained on only Arabic data, including dialectal Arabic, and use a larger number of parameters, such as the recently released MARBERT. However, we noticed that AraELECTRA is one of the top performing models while being much more efficient in its computational cost. Finally, the experiments on AraGPT2 variants showed low performance compared to BERT models, which indicates that it might not be suitable for classification tasks.

benchmarking transformer-based language transformer-based language models transformer-based language معايير اللغة القائمة على المحولات نماذج اللغة القائمة على المحولات اللغة القائمة على المحولات صناعة حمض الفوسفور المزيد..

Few-Shot Intent Detection via Contrastive Pre-Training and Fine-Tuning

353 - Association for Computation Linguistics 2021 مقالة

In this work, we focus on a more challenging few-shot intent detection scenario where many intents are fine-grained and semantically similar. We present a simple yet effective few-shot intent detection schema via contrastive pre-training and fine-tun ing. Specifically, we first conduct self-supervised contrastive pre-training on collected intent datasets, which implicitly learns to discriminate semantically similar utterances without using any labels. We then perform few-shot intent detection together with supervised contrastive learning, which explicitly pulls utterances from the same intent closer and pushes utterances across different intents farther. Experimental results show that our proposed method achieves state-of-the-art performance on three challenging intent detection datasets under 5-shot and 10-shot settings.

few-shot intent detection الكشف عن القلة الطلقات صناعة حمض الفوسفور

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Benchmarking Commercial Intent Detection Services with Practice-Driven Evaluations

معيار خدمات الكشف عن النية التجارية مع تقييمات مدفوعة بالممارسة

Ask ChatGPT about the research

Read More

suggested questions