New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Generic resources are what you need: Style transfer tasks without task-specific parallel training data

الموارد العامة هي ما تحتاجه: مهام نقل النمط دون بيانات تدريب موازية محددة

513 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

تهدف نقل النمط إلى إعادة كتابة نص مصدر بأسلوب مستهدف مختلف مع الحفاظ على محتواها. نقترح نهجا جديدا لهذه المهمة التي تنفد على الموارد العامة، ودون استخدام أي بيانات متوازية (الهدف - المستهدفة (المصدر) تفوقت على النهج الموجودة غير المنشورة على مهام نقل النمط الأكثر شعبية: نقل الشكليات ومبادلة القطبية. في الممارسة العملية، نعتمد إجراء متعدد الخطوات الذي يبني على نموذج تسلسل تسلسل مسبقا عام (BART). أولا، نقوم بتعزيز قدرة النموذج على إعادة الكتابة عن طريق مزيد من الردف ما قبل التدريب على كل من مجموعة موجودة من الصيارات العامة، وكذلك على أزواج الاصطناعية التي تم إنشاؤها باستخدام مورد مجمع للأغراض العامة. ثانيا، من خلال نهج الترجمة مرة أخرى تكرارية، نقوم بتدريب نماذجين، كل منها في اتجاه نقل، حتى يتمكنوا من توفير بعضهم البعض مع أزواج توليد مزخرف، ديناميكيا في عملية التدريب. أخيرا، ندعنا نطاطنا الناتج لدينا تولد أزواجا صناعية ثابتة لاستخدامها في نظام تدريبي مشترك. إلى جانب المنهجية والنتائج الحديثة، فإن المساهمة الأساسية لهذا العمل هي انعكاس على طبيعة المهامتين التي نتعامل معها، وكيف يتم تمييز اختلافاتهم عن طريق ردهم على نهجنا.

Style transfer aims to rewrite a source text in a different target style while preserving its content. We propose a novel approach to this task that leverages generic resources, and without using any task-specific parallel (source--target) data outperforms existing unsupervised approaches on the two most popular style transfer tasks: formality transfer and polarity swap. In practice, we adopt a multi-step procedure which builds on a generic pre-trained sequence-to-sequence model (BART). First, we strengthen the model's ability to rewrite by further pre-training BART on both an existing collection of generic paraphrases, as well as on synthetic pairs created using a general-purpose lexical resource. Second, through an iterative back-translation approach, we train two models, each in a transfer direction, so that they can provide each other with synthetically generated pairs, dynamically in the training process. Lastly, we let our best resulting model generate static synthetic pairs to be used in a supervised training regime. Besides methodology and state-of-the-art results, a core contribution of this work is a reflection on the nature of the two tasks we address, and how their differences are highlighted by their response to our approach.

References used

https://aclanthology.org/

rate research

Evaluating a Joint Training Approach for Learning Cross-lingual Embeddings with Sub-word Information without Parallel Corpora on Lower-resource Languages

277 - Association for Computation Linguistics 2021 مقالة

Cross-lingual word embeddings provide a way for information to be transferred between languages. In this paper we evaluate an extension of a joint training approach to learning cross-lingual embeddings that incorporates sub-word information during tr aining. This method could be particularly well-suited to lower-resource and morphologically-rich languages because it can be trained on modest size monolingual corpora, and is able to represent out-of-vocabulary words (OOVs). We consider bilingual lexicon induction, including an evaluation focused on OOVs. We find that this method achieves improvements over previous approaches, particularly for OOVs.

joint training approach learning cross-lingual embeddings parallel corpora نهج التدريب المشترك تعلم المضبوطات عبر اللغات فورانيا الموازية صناعة حمض الفوسفور المزيد..

Stepmothers are mean and academics are pretentious: What do pretrained language models learn about you?

262 - Association for Computation Linguistics 2021 مقالة

In this paper, we investigate what types of stereotypical information are captured by pretrained language models. We present the first dataset comprising stereotypical attributes of a range of social groups and propose a method to elicit stereotypes encoded by pretrained language models in an unsupervised fashion. Moreover, we link the emergent stereotypes to their manifestation as basic emotions as a means to study their emotional effects in a more generalized manner. To demonstrate how our methods can be used to analyze emotion and stereotype shifts due to linguistic experience, we use fine-tuning on news sources as a case study. Our experiments expose how attitudes towards different social groups vary across models and how quickly emotions and stereotypes can shift at the fine-tuning stage.

تحويل ملثمين language models learn نماذج اللغة تعلم صناعة حمض الفوسفور

Adversities are all you need: Classification of self-reported breast cancer posts on Twitter using Adversarial Fine-tuning

185 - Association for Computation Linguistics 2021 مقالة

In this paper, we describe our system entry for Shared Task 8 at SMM4H-2021, which is on automatic classification of self-reported breast cancer posts on Twitter. In our system, we use a transformer-based language model fine-tuning approach to automa tically identify tweets in the self-reports category. Furthermore, we involve a Gradient-based Adversarial fine-tuning to improve the overall model's robustness. Our system achieved an F1-score of 0.8625 on the Development set and 0.8501 on the Test set in Shared Task-8 of SMM4H-2021.

self-reported breast cancer breast cancer posts posts on twitter سرطان الثدي تم الإبلاغ عنها ذاتيا المشاركات سرطان الثدي المشاركات على تويتر صناعة حمض الفوسفور المزيد..

Direction is what you need: Improving Word Embedding Compression in Large Language Models

679 - Association for Computation Linguistics 2021 مقالة

The adoption of Transformer-based models in natural language processing (NLP) has led to great success using a massive number of parameters. However, due to deployment constraints in edge devices, there has been a rising interest in the compression o f these models to improve their inference time and memory footprint. This paper presents a novel loss objective to compress token embeddings in the Transformer-based models by leveraging an AutoEncoder architecture. More specifically, we emphasize the importance of the direction of compressed embeddings with respect to original uncompressed embeddings. The proposed method is task-agnostic and does not require further language modeling pre-training. Our method significantly outperforms the commonly used SVD-based matrix-factorization approach in terms of initial language model Perplexity. Moreover, we evaluate our proposed approach over SQuAD v1.1 dataset and several downstream tasks from the GLUE benchmark, where we also outperform the baseline in most scenarios. Our code is public.

improving word embedding improving word word embedding compression تحسين كلمة التضمين تحسين كلمة كلمة تضمين ضغط صناعة حمض الفوسفور المزيد..

GHOST at SemEval-2021 Task 5: Is explanation all you need?

306 - Association for Computation Linguistics 2021 مقالة

This paper discusses different approaches to the Toxic Spans Detection task. The problem posed by the task was to determine which words contribute mostly to recognising a document as toxic. As opposed to binary classification of entire texts, word-le vel assessment could be of great use during comment moderation, also allowing for a more in-depth comprehension of the model's predictions. As the main goal was to ensure transparency and understanding, this paper focuses on the current state-of-the-art approaches based on the explainable AI concepts and compares them to a supervised learning solution with word-level labels. The work consists of two xAI approaches that automatically provide the explanation for models trained for binary classification of toxic documents: an LSTM model with attention as a model-specific approach and the Shapley values for interpreting BERT predictions as a model-agnostic method. The competing approach considers this problem as supervised token classification, where models like BERT and its modifications were tested. The paper aims to explore, compare and assess the quality of predictions for different methods on the task. The advantages of each approach and further research direction are also discussed.

سامة صناعة حمض الفوسفور

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Generic resources are what you need: Style transfer tasks without task-specific parallel training data

الموارد العامة هي ما تحتاجه: مهام نقل النمط دون بيانات تدريب موازية محددة

Ask ChatGPT about the research

Read More

suggested questions