Subscribe to the gold package and get unlimited access to Shamra Academy

Towards Benchmarking the Utility of Explanations for Model Debugging

نحو مراجعة فائدة التفسيرات لتصحيح التصحيح النموذجي

914 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

benchmarking the utility post-hoc explanation methods trained model decision معيار الفائدة طرق تفسير ما بعد الهوك قرار النموذج المدربين صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

Post-hoc explanation methods are an important class of approaches that help understand the rationale underlying a trained model's decision. But how useful are they for an end-user towards accomplishing a given task? In this vision paper, we argue the need for a benchmark to facilitate evaluations of the utility of post-hoc explanation methods. As a first step to this end, we enumerate desirable properties that such a benchmark should possess for the task of debugging text classifiers. Additionally, we highlight that such a benchmark facilitates not only assessing the effectiveness of explanations but also their efficiency.

References used

https://aclanthology.org/

rate research

Towards Task-Agnostic Privacy- and Utility-Preserving Models

640 - Association for Computation Linguistics 2021 مقالة

Modern deep learning models for natural language processing rely heavily on large amounts of annotated texts. However, obtaining such texts may be difficult when they contain personal or confidential information, for example, in health or legal domai ns. In this work, we propose a method of de-identifying free-form text documents by carefully redacting sensitive data in them. We show that our method preserves data utility for text classification, sequence labeling and question answering tasks.

task-agnostic privacy utility-preserving models الخصوصية المرجعية المهمة نماذج الحفاظ على المرافق خصوصية صناعة حمض الفوسفور

Towards a Language Model for Temporal Commonsense Reasoning

1051 - Association for Computation Linguistics 2021 مقالة

Temporal commonsense reasoning is a challenging task as it requires temporal knowledge usually not explicit in text. In this work, we propose an ensemble model for temporal commonsense reasoning. Our model relies on pre-trained contextual representat ions from transformer-based language models (i.e., BERT), and on a variety of training methods for enhancing model generalization: 1) multi-step fine-tuning using carefully selected auxiliary tasks and datasets, and 2) a specifically designed temporal masked language model task aimed to capture temporal commonsense knowledge. Our model greatly outperforms the standard fine-tuning approach and strong baselines on the MC-TACO dataset.

temporal commonsense reasoning commonsense reasoning temporal commonsense المنطق الزمني المنطقي المنطق المنطقي العمولة الزمنية صناعة حمض الفوسفور المزيد..

Contrastive Explanations for Model Interpretability

474 - Association for Computation Linguistics 2021 مقالة

Contrastive explanations clarify why an event occurred in contrast to another. They are inherently intuitive to humans to both produce and comprehend. We propose a method to produce contrastive explanations in the latent space, via a projection of th e input representation, such that only the features that differentiate two potential decisions are captured. Our modification allows model behavior to consider only contrastive reasoning, and uncover which aspects of the input are useful for and against particular decisions. Our contrastive explanations can additionally answer for which label, and against which alternative label, is a given input feature useful. We produce contrastive explanations via both high-level abstract concept attribution and low-level input token/span attribution for two NLP classification benchmarks. Our findings demonstrate the ability of label-contrastive explanations to provide fine-grained interpretability of model decisions.

contrastive explanations explanations تفسيرات تناقض التفسيرات صناعة حمض الفوسفور

Bayesian Model-Agnostic Meta-Learning with Matrix-Valued Kernels for Quality Estimation

690 - Association for Computation Linguistics 2021 مقالة

Most current quality estimation (QE) models for machine translation are trained and evaluated in a fully supervised setting requiring significant quantities of labelled training data. However, obtaining labelled data can be both expensive and time-co nsuming. In addition, the test data that a deployed QE model would be exposed to may differ from its training data in significant ways. In particular, training samples are often labelled by one or a small set of annotators, whose perceptions of translation quality and needs may differ substantially from those of end-users, who will employ predictions in practice. Thus, it is desirable to be able to adapt QE models efficiently to new user data with limited supervision data. To address these challenges, we propose a Bayesian meta-learning approach for adapting QE models to the needs and preferences of each user with limited supervision. To enhance performance, we further propose an extension to a state-of-the-art Bayesian meta-learning approach which utilizes a matrix-valued kernel for Bayesian meta-learning of quality estimation. Experiments on data with varying number of users and language characteristics demonstrates that the proposed Bayesian meta-learning approach delivers improved predictive performance in both limited and full supervision settings.

التفضيلات الاشتراكية current quality estimation bayesian meta-learning تقدير الجودة الحالية التعلم البيئي التعلم صناعة حمض الفوسفور

ASAP: A Chinese Review Dataset Towards Aspect Category Sentiment Analysis and Rating Prediction

764 - Association for Computation Linguistics 2021 مقالة

Sentiment analysis has attracted increasing attention in e-commerce. The sentiment polarities underlying user reviews are of great value for business intelligence. Aspect category sentiment analysis (ACSA) and review rating prediction (RP) are two es sential tasks to detect the fine-to-coarse sentiment polarities. ACSA and RP are highly correlated and usually employed jointly in real-world e-commerce scenarios. While most public datasets are constructed for ACSA and RP separately, which may limit the further exploitation of both tasks. To address the problem and advance related researches, we present a large-scale Chinese restaurant review dataset ASAP including 46, 730 genuine reviews from a leading online-to-offline (O2O) e-commerce platform in China. Besides a 5-star scale rating, each review is manually annotated according to its sentiment polarities towards 18 pre-defined aspect categories. We hope the release of the dataset could shed some light on the field of sentiment analysis. Moreover, we propose an intuitive yet effective joint model for ACSA and RP. Experimental results demonstrate that the joint model outperforms state-of-the-art baselines on both tasks.

بيرت و GPT. صناعة حمض الفوسفور

Towards Benchmarking the Utility of Explanations for Model Debugging

نحو مراجعة فائدة التفسيرات لتصحيح التصحيح النموذجي

Ask ChatGPT about the research

Read More

suggested questions