Abstract:Around seven-in-ten Americans use social media (SM) to connect and engage, making these platforms excellent sources of information to understand human behavior and other problems relevant to social sciences. While the presence of a behavior can be detected, it is unclear who or under what circumstances the behavior was generated. Despite the large sample sizes of SM datasets, they almost always come with significant biases, some of which have been studied before. Here, we hypothesize the presence of a largely unrecognized form of bias on SM platforms, called participation bias , that is distinct from selection bias. It is defined as the skew in the demographics of the participants who opt-in to discussions of the topic, compared to the demographics of the underlying SM platform. To infer the participant's demographics, we propose a novel generative probabilistic framework that links surveys and SM data at the granularity of demographic subgroups (and not individuals). Our method is distinct from existing approaches that elicit such information at the individual level using their profile name, images, and other metadata, thus infringing upon their privacy. We design a statistical simulation to simulate multiple SM platforms and a diverse range of topics to validate the model's estimates in different scenarios. We use Twitter data as a case study to demonstrate participation bias on the topic of gun violence delineated by political party affiliation and gender. Although Twitter's user population leans Democratic and has an equal number of men and women according to Pew, our model's estimates point to the presence of participation bias on the topic of gun control in the opposite direction, with slightly more Republicans than Democrats, and more men compared to women. Our study cautions that in the rush to use digital data for decision-making and understanding public opinions, we must account for the biases inherent in how SM data are produced, lest we may also arrive at biased inferences about the public.

When is it Biased? Assessing the Representativeness of Twitter's Streaming API

This Sample seems to be good enough! Assessing Coverage and Temporal Reliability of Twitter's Academic API

Quantifying participation biases on social media

Data and Model Biases in Social Media Analyses: A Case Study of COVID-19 Tweets.

Who Makes Trends? Understanding Demographic Biases in Crowdsourced Recommendations

Towards a Standard Sampling Methodology on Online Social Networks: Collecting Global Trends on Twitter

Quantifying Biases in Online Information Exposure

Identifying Data Noises, User Biases, and System Errors in Geo-tagged Twitter Messages (Tweets)

Intertwined Biases Across Social Media Spheres: Unpacking Correlations in Media Bias Dimensions

Measuring Online Social Bubbles

On application-unbiased benchmarking of web videos from a social network perspective

Quantifying Algorithmic Biases over Time

Big Questions for Social Media Big Data: Representativeness, Validity and Other Methodological Pitfalls

Quantifying Biases in Social Media Analysis of Recreation in Urban Parks

Sampled Datasets Risk Substantial Bias in the Identification of Political Polarization on Social Media

A survey on learning from imbalanced data streams: taxonomy, challenges, empirical study, and reproducible experimental framework

Biases in Using Social Media Data for Public Health Surveillance: A Scoping Review.

Auditing for Bias in Ad Delivery Using Inferred Demographic Attributes

Reducing Population-level Inequality Can Improve Demographic Group Fairness: a Twitter Case Study

Real Time Sentiment Change Detection of Twitter Data Streams

A Biased Review of Biases in Twitter Studies on Political Collective Action