Integrating Big Data and Survey Data for Efficient Estimation of the Median

Ryan Covey
DOI: https://doi.org/10.48550/arXiv.2306.16089
2023-06-28
Abstract:An ever-increasing deluge of big data is becoming available to national statistical offices globally, but it is well documented that statistics produced by big data alone often suffer from selection bias and are not usually representative of the population at large. In this paper, we construct a new design-based estimator of the median by integrating big data and survey data. Our estimator is asymptotically unbiased and has a smaller variance than a median estimator produced using survey data alone.
Methodology
What problem does this paper attempt to address?