Abstract:Background Social media are important for monitoring perceptions of public health issues and for educating target audiences about health; however, limited information about the demographics of social media users makes it challenging to identify conversations among target audiences and limits how well social media can be used for public health surveillance and education outreach efforts. Certain social media platforms provide demographic information on followers of a user account, if given, but they are not always disclosed, and researchers have developed machine learning algorithms to predict social media users’ demographic characteristics, mainly for Twitter. To date, there has been limited research on predicting the demographic characteristics of Reddit users. Objective We aimed to develop a machine learning algorithm that predicts the age segment of Reddit users, as either adolescents or adults, based on publicly available data. Methods This study was conducted between January and September 2020 using publicly available Reddit posts as input data. We manually labeled Reddit users’ age by identifying and reviewing public posts in which Reddit users self-reported their age. We then collected sample posts, comments, and metadata for the labeled user accounts and created variables to capture linguistic patterns, posting behavior, and account details that would distinguish the adolescent age group (aged 13 to 20 years) from the adult age group (aged 21 to 54 years). We split the data into training (n=1660) and test sets (n=415) and performed 5-fold cross validation on the training set to select hyperparameters and perform feature selection. We ran multiple classification algorithms and tested the performance of the models (precision, recall, F1 score) in predicting the age segments of the users in the labeled data. To evaluate associations between each feature and the outcome, we calculated means and confidence intervals and compared the two age groups, with 2-sample t tests, for each transformed model feature. Results The gradient boosted trees classifier performed the best, with an F1 score of 0.78. The test set precision and recall scores were 0.79 and 0.89, respectively, for the adolescent group (n=254) and 0.78 and 0.63, respectively, for the adult group (n=161). The most important feature in the model was the number of sentences per comment (permutation score: mean 0.100, SD 0.004). Members of the adolescent age group tended to have created accounts more recently, have higher proportions of submissions and comments in the r/teenagers subreddit, and post more in subreddits with higher subscriber counts than those in the adult group. Conclusions We created a Reddit age prediction algorithm with competitive accuracy using publicly available data, suggesting machine learning methods can help public health agencies identify age-related target audiences on Reddit. Our results also suggest that there are characteristics of Reddit users’ posting behavior, linguistic patterns, and account features that distinguish adolescents from adults.

Textual Pre-Trained Models for Age Screening Across Community Question-Answering

Toward Personalized Activity Level Prediction in Community Question Answering Websites

Speech-based Age and Gender Prediction with Transformers

An Improved Selective Facial Extraction Model for Age Estimation

Predicting Best Responder in Community Question Answering Using Topic Model Method.

Author Identity Unveiled: Gender and Age Prediction from Textual Patterns using BERT

A Hybrid Transformer-Sequencer approach for Age and Gender classification from in-wild facial images

Can Demographic Factors Improve Text Classification? Revisiting Demographic Adaptation in the Age of Transformers

TextAge: A Curated and Diverse Text Dataset for Age Classification

Leveraging Textual Features for Best Answer Prediction in Community-based Question Answering

Harnessing the Power of Metadata for Enhanced Question Retrieval in Community Question Answering

Optimized Biomedical Question-Answering Services with LLM and Multi-BERT Integration

Accurate estimation of biological age and its application in disease prediction using a multimodal image Transformer system

Topic-Level Expert Modeling in Community Question Answering

Predicting semantic category of answers for question answering systems using transformers: a transfer learning approach

Age and Gender Prediction using Caffe Model and OpenCV

Age and Gender Recognition Using a Convolutional Neural Network with a Specially Designed Multi-Attention Module through Speech Spectrograms

Predicting Age Groups of Reddit Users Based on Posting Behavior and Metadata: Classification Model Development and Validation

Statistical Machine Translation Improves Question Retrieval in Community Question Answering Via Matrix Factorization.

BERT-Based Arabic Social Media Author Profiling

Residualized Factor Adaptation for Community Social Media Prediction Tasks