New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Sustainable Modular Debiasing of Language Models

ديوان وحدات مستدامة للنماذج اللغوية

216 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

كفاءة اللغوية adele أديل صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

التحيزات النمطية غير العادلة (على سبيل المثال، التحيزات الجنسانية أو العنصرية أو الدينية) ترميز نماذج اللغة الحديثة المحددة مسبقا (PLMS) لها آثار أخلاقية سلبية على الاعتماد الواسع النطاق لتكنولوجيا اللغات الحديثة. لعلاج ذلك، تم تقديم مجموعة واسعة من تقنيات المساواة مؤخرا لإزالة هذه التحيزات النمطية من PLMS. ومع ذلك، فإن طرق الدخل الحالية، ومع ذلك، قم بتعديل جميع معلمات PLMS مباشرة، والتي - إلى جانب كونها باهظة الثمن - مع خطر الكامنة من (كارثي) نسيان المعرفة اللغوية المفيدة المكتسبة في الاحتجاج. في هذا العمل، نقترح نهجا أكثر استدامة للدوائر على أساس محولات Deviasing المخصصة، التي دبلها أديل. بشكل ملموس، نحن (1) وحدات محول حقن في طبقات PLM الأصلية و (2) تحديث المحولات فقط (أي ونحن نعرض أديل، في الدخل الجنساني من BERT: تقييمنا الواسع، يشمل ثلاثة تدابير محلية خارجية ومثيرة للخدمة الخارجية، مما يجعل أديل، فعالة للغاية في تخفيف التحيز. نوضح كذلك - نظرا لطبيعتها المعيارية - أديل، إلى جانب محولات المهام، تحتفظ بالإنصاف حتى بعد التدريب على النمو النطاق واسع النطاق. وأخيرا، عن طريق بيرت متعددة اللغات، نجحنا في نقل أديل بنجاح إلى ست لغات مستهدفة.

Unfair stereotypical biases (e.g., gender, racial, or religious biases) encoded in modern pretrained language models (PLMs) have negative ethical implications for widespread adoption of state-of-the-art language technology. To remedy for this, a wide range of debiasing techniques have recently been introduced to remove such stereotypical biases from PLMs. Existing debiasing methods, however, directly modify all of the PLMs parameters, which -- besides being computationally expensive -- comes with the inherent risk of (catastrophic) forgetting of useful language knowledge acquired in pretraining. In this work, we propose a more sustainable modular debiasing approach based on dedicated debiasing adapters, dubbed ADELE. Concretely, we (1) inject adapter modules into the original PLM layers and (2) update only the adapters (i.e., we keep the original PLM parameters frozen) via language modeling training on a counterfactually augmented corpus. We showcase ADELE, in gender debiasing of BERT: our extensive evaluation, encompassing three intrinsic and two extrinsic bias measures, renders ADELE, very effective in bias mitigation. We further show that -- due to its modular nature -- ADELE, coupled with task adapters, retains fairness even after large-scale downstream training. Finally, by means of multilingual BERT, we successfully transfer ADELE, to six target languages.

References used

https://aclanthology.org/

rate research

Modular Networks for Compositional Instruction Following

420 - Association for Computation Linguistics 2021 مقالة

Standard architectures used in instruction following often struggle on novel compositions of subgoals (e.g. navigating to landmarks or picking up objects) observed during training. We propose a modular architecture for following natural language inst ructions that describe sequences of diverse subgoals. In our approach, subgoal modules each carry out natural language instructions for a specific subgoal type. A sequence of modules to execute is chosen by learning to segment the instructions and predicting a subgoal type for each segment. When compared to standard, non-modular sequence-to-sequence approaches on ALFRED, a challenging instruction following benchmark, we find that modularization improves generalization to novel subgoal compositions, as well as to environments unseen in training.

networks for compositional compositional instruction modular networks شبكات للتركيب التعليمات التركيبية شبكات وحدات صناعة حمض الفوسفور المزيد..

Knodle: Modular Weakly Supervised Learning with PyTorch

216 - Association for Computation Linguistics 2021 مقالة

Strategies for improving the training and prediction quality of weakly supervised machine learning models vary in how much they are tailored to a specific task or integrated with a specific model architecture. In this work, we introduce Knodle, a sof tware framework that treats weak data annotations, deep learning models, and methods for improving weakly supervised training as separate, modular components. This modularization gives the training process access to fine-grained information such as data set characteristics, matches of heuristic rules, or elements of the deep learning model ultimately used for prediction. Hence, our framework can encompass a wide range of training methods for improving weak supervision, ranging from methods that only look at correlations of rules and output classes (independently of the machine learning model trained with the resulting labels), to those that harness the interplay of neural networks and weakly labeled data. We illustrate the benchmarking potential of the framework with a performance comparison of several reference implementations on a selection of datasets that are already available in Knodle.

modular weakly supervised weakly supervised learning weakly supervised وحدات تحت إشراف ضعيف التعلم الإشرافي ضعيف تحت إشراف ضعيف صناعة حمض الفوسفور المزيد..

Numeracy enhances the Literacy of Language Models

179 - Association for Computation Linguistics 2021 مقالة

Specialized number representations in NLP have shown improvements on numerical reasoning tasks like arithmetic word problems and masked number prediction. But humans also use numeracy to make better sense of world concepts, e.g., you can seat 5 peopl e in your room' but not 500. Does a better grasp of numbers improve a model's understanding of other concepts and words? This paper studies the effect of using six different number encoders on the task of masked word prediction (MWP), as a proxy for evaluating literacy. To support this investigation, we develop Wiki-Convert, a 900,000 sentence dataset annotated with numbers and units, to avoid conflating nominal and ordinal number occurrences. We find a significant improvement in MWP for sentences containing numbers, that exponent embeddings are the best number encoders, yielding over 2 points jump in prediction accuracy over a BERT baseline, and that these enhanced literacy skills also generalize to contexts without annotated numbers. We release all code at https://git.io/JuZXn.

الاستدلال المعجمي literacy of language معرفة القراءة والكتابة اللغة صناعة حمض الفوسفور

Modular Self-Supervision for Document-Level Relation Extraction

445 - Association for Computation Linguistics 2021 مقالة

Extracting relations across large text spans has been relatively underexplored in NLP, but it is particularly important for high-value domains such as biomedicine, where obtaining high recall of the latest findings is crucial for practical applicatio ns. Compared to conventional information extraction confined to short text spans, document-level relation extraction faces additional challenges in both inference and learning. Given longer text spans, state-of-the-art neural architectures are less effective and task-specific self-supervision such as distant supervision becomes very noisy. In this paper, we propose decomposing document-level relation extraction into relation detection and argument resolution, taking inspiration from Davidsonian semantics. This enables us to incorporate explicit discourse modeling and leverage modular self-supervision for each sub-problem, which is less noise-prone and can be further refined end-to-end via variational EM. We conduct a thorough evaluation in biomedical machine reading for precision oncology, where cross-paragraph relation mentions are prevalent. Our method outperforms prior state of the art, such as multi-scale learning and graph neural networks, by over 20 absolute F1 points. The gain is particularly pronounced among the most challenging relation instances whose arguments never co-occur in a paragraph.

تمثيلات النموذج الأولي text spans يمتد النص صناعة حمض الفوسفور

Fast and Scalable Dialogue State Tracking with Explicit Modular Decomposition

363 - Association for Computation Linguistics 2021 مقالة

We present a fast and scalable architecture called Explicit Modular Decomposition (EMD), in which we incorporate both classification-based and extraction-based methods and design four modules (for clas- sification and sequence labelling) to jointly e xtract dialogue states. Experimental results based on the MultiWoz 2.0 dataset validates the superiority of our proposed model in terms of both complexity and scalability when compared to the state-of-the-art methods, especially in the scenario of multi-domain dialogues entangled with many turns of utterances.

explicit modular decomposition modular decomposition التحلل المعياري الصريح التحلل وحدات صناعة حمض الفوسفور

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Sustainable Modular Debiasing of Language Models

ديوان وحدات مستدامة للنماذج اللغوية

Ask ChatGPT about the research

Read More

suggested questions