Subscribe to the gold package and get unlimited access to Shamra Academy

BERT-based Multi-Task Model for Country and Province Level MSA and Dialectal Arabic Identification

نموذج متعدد المهام مقرها بيرت لمقاطعة MSA والحمولية الهوية العربية

832 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

province level msa dialectal arabic identification dialectal arabic مستوى المحافظة MSA تحديد الهوية العربية الجدلي منطقيا عربي صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

Dialect and standard language identification are crucial tasks for many Arabic natural language processing applications. In this paper, we present our deep learning-based system, submitted to the second NADI shared task for country-level and province-level identification of Modern Standard Arabic (MSA) and Dialectal Arabic (DA). The system is based on an end-to-end deep Multi-Task Learning (MTL) model to tackle both country-level and province-level MSA/DA identification. The latter MTL model consists of a shared Bidirectional Encoder Representation Transformers (BERT) encoder, two task-specific attention layers, and two classifiers. Our key idea is to leverage both the task-discriminative and the inter-task shared features for country and province MSA/DA identification. The obtained results show that our MTL model outperforms single-task models on most subtasks.

References used

https://aclanthology.org/

rate research

Deep Multi-Task Model for Sarcasm Detection and Sentiment Analysis in Arabic Language

1231 - Association for Computation Linguistics 2021 مقالة

The prominence of figurative language devices, such as sarcasm and irony, poses serious challenges for Arabic Sentiment Analysis (SA). While previous research works tackle SA and sarcasm detection separately, this paper introduces an end-to-end deep Multi-Task Learning (MTL) model, allowing knowledge interaction between the two tasks. Our MTL model's architecture consists of a Bidirectional Encoder Representation from Transformers (BERT) model, a multi-task attention interaction module, and two task classifiers. The overall obtained results show that our proposed model outperforms its single-task and MTL counterparts on both sarcasm and sentiment detection subtasks.

arabic sentiment analysis arabic sentiment تحليل المشاعر العربية الجنس العربي صناعة حمض الفوسفور

Country-level Arabic Dialect Identification using RNNs with and without Linguistic Features

1078 - Association for Computation Linguistics 2021 مقالة

This work investigates the value of augmenting recurrent neural networks with feature engineering for the Second Nuanced Arabic Dialect Identification (NADI) Subtask 1.2: Country-level DA identification. We compare the performance of a simple word-le vel LSTM using pretrained embeddings with one enhanced using feature embeddings for engineered linguistic features. Our results show that the addition of explicit features to the LSTM is detrimental to performance. We attribute this performance loss to the bivalency of some linguistic items in some text, ubiquity of topics, and participant mobility.

منطقيا عربي country-level arabic dialect لهجة عربية على مستوى البلد صناعة حمض الفوسفور

abcbpc at SemEval-2021 Task 7: ERNIE-based Multi-task Model for Detecting and Rating Humor and Offense

724 - Association for Computation Linguistics 2021 مقالة

This paper describes our system participated in Task 7 of SemEval-2021: Detecting and Rating Humor and Offense. The task is designed to detect and score humor and offense which are influenced by subjective factors. In order to obtain semantic informa tion from a large amount of unlabeled data, we applied unsupervised pre-trained language models. By conducting research and experiments, we found that the ERNIE 2.0 and DeBERTa pre-trained models achieved impressive performance in various subtasks. Therefore, we applied the above pre-trained models to fine-tune the downstream neural network. In the process of fine-tuning the model, we adopted multi-task training strategy and ensemble learning method. Based on the above strategy and method, we achieved RMSE of 0.4959 for subtask 1b, and finally won the first place.

detecting and rating rating humor humor and offense الكشف والتقييم تصنيف الفكاهة الفكاهة والجريمة صناعة حمض الفوسفور المزيد..

Multi Task Learning based Framework for Multimodal Classification

776 - Association for Computation Linguistics 2021 مقالة

Large-scale multi-modal classification aim to distinguish between different multi-modal data, and it has drawn dramatically attentions since last decade. In this paper, we propose a multi-task learning-based framework for the multimodal classificatio n task, which consists of two branches: multi-modal autoencoder branch and attention-based multi-modal modeling branch. Multi-modal autoencoder can receive multi-modal features and obtain the interactive information which called multi-modal encoder feature, and use this feature to reconstitute all the input data. Besides, multi-modal encoder feature can be used to enrich the raw dataset, and improve the performance of downstream tasks (such as classification task). As for attention-based multimodal modeling branch, we first employ attention mechanism to make the model focused on important features, then we use the multi-modal encoder feature to enrich the input information, achieve a better performance. We conduct extensive experiments on different dataset, the results demonstrate the effectiveness of proposed framework.

multi task learning learning based framework task learning based التعلم المهمة المتعددة التعلم القائمة على الإطار التعلم المهمة القائمة صناعة حمض الفوسفور المزيد..

NADI 2021: The Second Nuanced Arabic Dialect Identification Shared Task

768 - Association for Computation Linguistics 2021 مقالة

We present the findings and results of theSecond Nuanced Arabic Dialect IdentificationShared Task (NADI 2021). This Shared Taskincludes four subtasks: country-level ModernStandard Arabic (MSA) identification (Subtask1.1), country-level dialect identi fication (Subtask1.2), province-level MSA identification (Subtask2.1), and province-level sub-dialect identifica-tion (Subtask 2.2). The shared task dataset cov-ers a total of 100 provinces from 21 Arab coun-tries, collected from the Twitter domain. A totalof 53 teams from 23 countries registered to par-ticipate in the tasks, thus reflecting the interestof the community in this area. We received 16submissions for Subtask 1.1 from five teams, 27submissions for Subtask 1.2 from eight teams,12 submissions for Subtask 2.1 from four teams,and 13 Submissions for subtask 2.2 from fourteams.

nuanced arabic dialect nuanced arabic arabic dialect identification لهجة عربية دقيقة nuanced العربية الهوية العربية الهوية صناعة حمض الفوسفور المزيد..

BERT-based Multi-Task Model for Country and Province Level MSA and Dialectal Arabic Identification

نموذج متعدد المهام مقرها بيرت لمقاطعة MSA والحمولية الهوية العربية

Ask ChatGPT about the research

Read More

suggested questions