Meta-Analysis with Untrusted Data

Shiva Kaul,Geoffrey J. Gordon
2024-07-13
Abstract:[See paper for full abstract] Meta-analysis is a crucial tool for answering scientific questions. It is usually conducted on a relatively small amount of ``trusted'' data -- ideally from randomized, controlled trials -- which allow causal effects to be reliably estimated with minimal assumptions. We show how to answer causal questions much more precisely by making two changes. First, we incorporate untrusted data drawn from large observational databases, related scientific literature and practical experience -- without sacrificing rigor or introducing strong assumptions. Second, we train richer models capable of handling heterogeneous trials, addressing a long-standing challenge in meta-analysis. Our approach is based on conformal prediction, which fundamentally produces rigorous prediction intervals, but doesn't handle indirect observations: in meta-analysis, we observe only noisy effects due to the limited number of participants in each trial. To handle noise, we develop a simple, efficient version of fully-conformal kernel ridge regression, based on a novel condition called idiocentricity. We introduce noise-correcting terms in the residuals and analyze their interaction with a ``variance shaving'' technique. In multiple experiments on healthcare datasets, our algorithms deliver tighter, sounder intervals than traditional ones. This paper charts a new course for meta-analysis and evidence-based medicine, where heterogeneity and untrusted data are embraced for more nuanced and precise predictions.
Machine Learning,Methodology
What problem does this paper attempt to address?
The paper attempts to address the issue of how to utilize unreliable data (such as observational studies, related literature, and practical experience) in meta-analysis while maintaining the rigor and unbiasedness of predictions. Traditional meta-analysis methods rely solely on "reliable" data (usually data from randomized controlled trials), which are limited in quantity and cannot effectively model the heterogeneity of treatment effects. Therefore, traditional methods perform poorly in predicting outcomes when faced with highly heterogeneous data. This paper proposes a new algorithm based on conformal prediction, which can introduce a large amount of unreliable data while still maintaining rigorous prediction intervals for causal effects. Through this method, the paper aims to improve existing meta-analysis techniques to address heterogeneity issues and enhance prediction accuracy and reliability. Specifically, the paper proposes two new algorithms to handle noise issues: one suitable for large-scale trials and the other for small-scale trials. Both algorithms are based on fully-conformal prediction and can incorporate unreliable data to improve prediction results. Experimental evidence shows that this new meta-analysis method can significantly reduce the width of prediction intervals across multiple biomedical datasets, thereby providing more precise predictions.