New community

Subscribe to the gold package and get unlimited access to Shamra Academy

Using Noisy Self-Reports to Predict Twitter User Demographics

باستخدام تقارير ذاتية صاخبة للتنبؤ برطب مستخدم Twitter

167 0 0 0.0 ( 0 )

Download Cite

Added by Association for Computation Linguistics مقالة

Publication date 2021

fields Artificial Intelligence

and research's language is English

Created by Shamra Editor

predict twitter user twitter user demographics predict twitter توقع مستخدم Twitter Twitter المستخدم التركيبة السكانية توقع Twitter. صناعة حمض الفوسفور

visit our facebook page

‎Shamra Academia - شمرا أكاديميا‎

Ask ChatGPT about the research

Abstract in Arabic Abstract in English

Computational social science studies often contextualize content analysis within standard demographics. Since demographics are unavailable on many social media platforms (e.g. Twitter), numerous studies have inferred demographics automatically. Despite many studies presenting proof-of-concept inference of race and ethnicity, training of practical systems remains elusive since there are few annotated datasets. Existing datasets are small, inaccurate, or fail to cover the four most common racial and ethnic groups in the United States. We present a method to identify self-reports of race and ethnicity from Twitter profile descriptions. Despite the noise of automated supervision, our self-report datasets enable improvements in classification performance on gold standard self-report survey data. The result is a reproducible method for creating large-scale training resources for race and ethnicity.

References used

https://aclanthology.org/

rate research

UL2C: Mapping User Locations to Countries on Arabic Twitter

214 - Association for Computation Linguistics 2021 مقالة

Mapping user locations to countries can be useful for many applications such as dialect identification, author profiling, recommendation system, etc. Twitter allows users to declare their locations as free text, and these user-declared locations are often noisy and hard to decipher automatically. In this paper, we present the largest manually labeled dataset for mapping user locations on Arabic Twitter to their corresponding countries. We build effective machine learning models that can automate this mapping with significantly better efficiency compared to libraries such as geopy. We also show that our dataset is more effective than data extracted from GeoNames geographical database in this task as the latter covers only locations written in formal ways.

mapping user locations mapping user تعيين مواقع المستخدم رسم الخرائط المستخدم صناعة حمض الفوسفور

Leveraging knowledge sources for detecting self-reports of particular health issues on social media

355 - Association for Computation Linguistics 2021 مقالة

This paper investigates incorporating quality knowledge sources developed by experts for the medical domain as well as syntactic information for classification of tweets into four different health oriented categories. We claim that resources such as the MeSH hierarchy and currently available parse information are effective extensions of moderately sized training datasets for various fine-grained tweet classification tasks of self-reported health issues.

leveraging knowledge sources detecting self-reports الاستفادة من مصادر المعرفة الكشف عن التقارير الذاتية صناعة حمض الفوسفور

Noisy Self-Knowledge Distillation for Text Summarization

327 - Association for Computation Linguistics 2021 مقالة

In this paper we apply self-knowledge distillation to text summarization which we argue can alleviate problems with maximum-likelihood training on single reference and noisy datasets. Instead of relying on one-hot annotation labels, our student summa rization model is trained with guidance from a teacher which generates smoothed labels to help regularize training. Furthermore, to better model uncertainty during training, we introduce multiple noise signals for both teacher and student models. We demonstrate experimentally on three benchmarks that our framework boosts the performance of both pretrained and non-pretrained summarizers achieving state-of-the-art results.

noisy self-knowledge distillation self-knowledge distillation تقطير المعرفة الذاتية صاخبة تقطير المعرفة الذاتية تلخيص النص صناعة حمض الفوسفور

The Korean Morphologically Tight-Fitting Tokenizer for Noisy User-Generated Texts

414 - Association for Computation Linguistics 2021 مقالة

User-generated texts include various types of stylistic properties, or noises. Such texts are not properly processed by existing morpheme analyzers or language models based on formal texts such as encyclopedias or news articles. In this paper, we pro pose a simple morphologically tight-fitting tokenizer (K-MT) that can better process proper nouns, coinages, and internet slang among other types of noise in Korean user-generated texts. We tested our tokenizer by performing classification tasks on Korean user-generated movie reviews and hate speech datasets, and the Korean Named Entity Recognition dataset. Through our tests, we found that K-MT is better fit to process internet slangs, proper nouns, and coinages, compared to a morpheme analyzer and a character-level WordPiece tokenizer.

noisy user-generated texts noisy user-generated morphologically tight-fitting tokenizer النصوص التي أنشأها المستخدم صاخبة صاخبة المستخدم مظلمة ضيقة مورفولوجية صناعة حمض الفوسفور المزيد..

Adversities are all you need: Classification of self-reported breast cancer posts on Twitter using Adversarial Fine-tuning

186 - Association for Computation Linguistics 2021 مقالة

In this paper, we describe our system entry for Shared Task 8 at SMM4H-2021, which is on automatic classification of self-reported breast cancer posts on Twitter. In our system, we use a transformer-based language model fine-tuning approach to automa tically identify tweets in the self-reports category. Furthermore, we involve a Gradient-based Adversarial fine-tuning to improve the overall model's robustness. Our system achieved an F1-score of 0.8625 on the Development set and 0.8501 on the Test set in Shared Task-8 of SMM4H-2021.

self-reported breast cancer breast cancer posts posts on twitter سرطان الثدي تم الإبلاغ عنها ذاتيا المشاركات سرطان الثدي المشاركات على تويتر صناعة حمض الفوسفور المزيد..

يمكنك البدء بجني المال وتحقيق ربح مادي من أبحاثك العلمية، المزيد

Using Noisy Self-Reports to Predict Twitter User Demographics

باستخدام تقارير ذاتية صاخبة للتنبؤ برطب مستخدم Twitter

Ask ChatGPT about the research

Read More

suggested questions