Flu Detector: Estimating influenza-like illness rates from online user-generated content

77 0 0.0 ( 0 )

تحميل البحث استخدام كمرجع

نشر من قبل Vasileios Lampos

تاريخ النشر 2016

مجال البحث الهندسة المعلوماتية

والبحث باللغة English

تأليف Vasileios Lampos

الذكاء الاصطناعي الحساب واللغة الشبكات الاجتماعية والمعلومات

قم بزيارة صفحتنا على فيسبوك

‎Shamra Academia - شمرا أكاديميا‎

اسأل ChatGPT حول البحث

الملخص بالعربية الملخص بالإنكليزية

We provide a brief technical description of an online platform for disease monitoring, titled as the Flu Detector (fludetector.cs.ucl.ac.uk). Flu Detector, in its current version (v.0.5), uses either Twitter or Google search data in conjunction with statistical Natural Language Processing models to estimate the rate of influenza-like illness in the population of England. Its back-end is a live service that collects online data, utilises modern technologies for large-scale text processing, and finally applies statistical inference models that are trained offline. The front-end visualises the various disease rate estimates. Notably, the models based on Google data achieve a high level of accuracy with respect to the most recent four flu seasons in England (2012/13 to 2015/16). This highlighted Flu Detector as having a great potential of becoming a complementary source to the domestic traditional flu surveillance schemes.

قيم البحث

242 - Konstantin Avrachenkov , Kishor Patil , Gugan Thoppe 2020

For providing quick and accurate search results, a search engine maintains a local snapshot of the entire web. And, to keep this local cache fresh, it employs a crawler for tracking changes across various web pages. It would have been ideal if the cr awler managed to update the local snapshot as soon as a page changed on the web. However, finite bandwidth availability and server restrictions mean that there is a bound on how frequently the different pages can be crawled. This then brings forth the following optimisation problem: maximise the freshness of the local cache subject to the crawling frequency being within the prescribed bounds. Recently, tractable algorithms have been proposed to solve this optimisation problem under different cost criteria. However, these assume the knowledge of exact page change rates, which is unrealistic in practice. We address this issue here. Specifically, we provide three novel schemes for online estimation of page change rates. All these schemes only need partial information about the page change process, i.e., they only need to know if the page has changed or not since the last crawl instance. Our first scheme is based on the law of large numbers, the second on the theory of stochastic approximation, while the third is an extension of the second and involves an additional momentum term. For all of these schemes, we prove convergence and, also, provide their convergence rates. As far as we know, the results concerning the third estimator is quite novel. Specifically, this is the first convergence type result for a stochastic approximation algorithm with momentum. Finally, we provide some numerical experiments (on real as well as synthetic data) to compare the performance of our proposed estimators with the existing ones (e.g., MLE).

استرجاع المعلومات التعلم الآلي الشبكات الاجتماعية والمعلومات

Seasonal Web Search Query Selection for Influenza-Like Illness (ILI) Estimation

278 - Niels Dalum Hansen , K{aa}re M{o}lbak , Ingemar J. Cox andn Christina Lioma 2018

Influenza-like illness (ILI) estimation from web search data is an important web analytics task. The basic idea is to use the frequencies of queries in web search logs that are correlated with past ILI activity as features when estimating current ILI activity. It has been noted that since influenza is seasonal, this approach can lead to spurious correlations with features/queries that also exhibit seasonality, but have no relationship with ILI. Spurious correlations can, in turn, degrade performance. To address this issue, we propose modeling the seasonal variation in ILI activity and selecting queries that are correlated with the residual of the seasonal model and the observed ILI signal. Experimental results show that re-ranking queries obtained by Google Correlate based on their correlation with the residual strongly favours ILI-related queries.

استرجاع المعلومات

Predicting User Emotional Tone in Mental Disorder Online Communities

45 - Barbara Silveira , Henrique S. Silva , Fabricio Murai 2020

In recent years, Online Social Networks have become an important medium for people who suffer from mental disorders to share moments of hardship, and receive emotional and informational support. In this work, we analyze how discussions in Reddit comm unities related to mental disorders can help improve the health conditions of their users. Using the emotional tone of users writing as a proxy for emotional state, we uncover relationships between user interactions and state changes. First, we observe that authors of negative posts often write rosier comments after engaging in discussions, indicating that users emotional state can improve due to social support. Second, we build models based on SOTA text embedding techniques and RNNs to predict shifts in emotional tone. This differs from most of related work, which focuses primarily on detecting mental disorders from user activity. We demonstrate the feasibility of accurately predicting the users reactions to the interactions experienced in these platforms, and present some examples which illustrate that the models are correctly capturing the effects of comments on the authors emotional tone. Our models hold promising implications for interventions to provide support for people struggling with mental illnesses.

التعلم الآلي الحساب واللغة الشبكات الاجتماعية والمعلومات

An Estimation of Online Video User Engagement from Features of Continuous Emotions

59 - Lukas Stappen , Alice Baird , Michelle Lienhart 2021

Portraying emotion and trustworthiness is known to increase the appeal of video content. However, the causal relationship between these signals and online user engagement is not well understood. This limited understanding is partly due to a scarcity in emotionally annotated data and the varied modalities which express user engagement online. In this contribution, we utilise a large dataset of YouTube review videos which includes ca. 600 hours of dimensional arousal, valence and trustworthiness annotations. We investigate features extracted from these signals against various user engagement indicators including views, like/dislike ratio, as well as the sentiment of comments. In doing so, we identify the positive and negative influences which single features have, as well as interpretable patterns in each dimension which relate to user engagement. Our results demonstrate that smaller boundary ranges and fluctuations for arousal lead to an increase in user engagement. Furthermore, the extracted time-series features reveal significant (p<0.05) correlations for each dimension, such as, count below signal mean (arousal), number of peaks (valence), and absolute energy (trustworthiness). From this, an effective combination of features is outlined for approaches aiming to automatically predict several user engagement indicators. In a user engagement prediction paradigm we compare all features against semi-automatic (cross-task), and automatic (task-specific) feature selection methods. These selected feature sets appear to outperform the usage of all features, e.g., using all features achieves 1.55 likes per day (Lp/d) mean absolute error from valence; this improves through semi-automatic and automatic selection to 1.33 and 1.23 Lp/d, respectively (data mean 9.72 Lp/d with a std. 28.75 Lp/d).

الوسائط المتعددة الحساب واللغة

Protocol for an Observational Study on the Effects of Social Distancing on Influenza-Like Illness and COVID-19

77 - Bo Zhang , Ting Ye , Siyu Heng 2020

The novel coronavirus disease (COVID-19) is a highly contagious respiratory disease that was first detected in Wuhan, China in December 2019, and has since spread around the globe, claiming more than 69,000 lives by the time this protocol is written. It has been widely acknowledged that the most effective public policy to mitigate the pandemic is emph{social and physical distancing}: keeping at least six feet away from people, working from home, closing non-essential businesses, etc. There have been a lot of anecdotal evidences suggesting that social distancing has a causal effect on disease mitigation; however, few studies have investigated the effect of social distancing on disease mitigation in a transparent and statistically-sound manner. We propose to perform an optimal non-bipartite matching to pair counties with similar observed covariates but vastly different average social distancing scores during the first week (March 16th through Match 22nd) of Presidents emph{15 Days to Slow the Spread} campaign. We have produced a total of $302$ pairs of two U.S. counties with good covariate balance on a total of $16$ important variables. Our primary outcome will be the average observed illness collected by Kinsa Inc. two weeks after the intervention period. Although the observed illness does not directly measure COVID-19, it reflects a real-time aspect of the pandemic, and unlike confirmed cases, it is much less confounded by counties testing capabilities. We also consider observed illness three weeks after the intervention period as a secondary outcome. We will test a proportional treatment effect using a randomization-based test with covariance adjustment and conduct a sensitivity analysis.

السكان والتطور