Abstract:We consider quantile estimation in a semi-supervised setting, characterized by two available data sets: (i) a small or moderate sized labeled data set containing observations for a response and a set of possibly high dimensional covariates, and (ii) a much larger unlabeled data set where only the covariates are observed. We propose a family of semi-supervised estimators for the response quantile(s) based on the two data sets, to improve the estimation accuracy compared to the supervised estimator, i.e., the sample quantile from the labeled data. These estimators use a flexible imputation strategy applied to the estimating equation along with a debiasing step that allows for full robustness against misspecification of the imputation model. Further, a one-step update strategy is adopted to enable easy implementation of our method and handle the complexity from the non-linear nature of the quantile estimating equation. Under mild assumptions, our estimators are fully robust to the choice of the nuisance imputation model, in the sense of always maintaining root-n consistency and asymptotic normality, while having improved efficiency relative to the supervised estimator. They also attain semi-parametric optimality if the relation between the response and the covariates is correctly specified via the imputation model. As an illustration of estimating the nuisance imputation function, we consider kernel smoothing type estimators on lower dimensional and possibly estimated transformations of the high dimensional covariates, and we establish novel results on their uniform convergence rates in high dimensions, involving responses indexed by a function class and usage of dimension reduction techniques. These results may be of independent interest. Numerical results on both simulated and real data confirm our semi-supervised approach's improved performance, in terms of both estimation and inference.

Optimal Semi-supervised Estimation and Inference for High-dimensional Linear Regression

Optimal and Safe Estimation for High-Dimensional Semi-Supervised Learning

Enhancing efficiency and robustness in high-dimensional linear regression with additional unlabeled data

Semisupervised inference for explained variance in high dimensional linear regression and its applications

Semi-supervised Inference for Explained Variance in High-dimensional Linear Regression and Its Applications

Efficient semi-supervised inference for logistic regression under case-control studies

Semi-Supervised Quantile Estimation: Robust and Efficient Inference in High Dimensional Settings

Semi-supervised inference for nonparametric logistic regression

Semi-supervised estimation for the varying coefficient regression model

Semi-supervised Regression Analysis with Model Misspecification and High-dimensional Data

Efficient Inference on High-Dimensional Linear Models with Missing Outcomes

Inference for High-Dimensional Linear Expectile Regression with De-Biasing Method

Sparse Semiparametric Efficient Estimation in High-Dimensional Linear Regression Models

Semiparametric efficient estimation in high‐dimensional partial linear regression models

Adaptive Semi-Supervised Learning with Discriminative Least Squares Regression

Imputations for High Missing Rate Data in Covariates Via Semi-supervised Learning Approach

Distributed Semi-Supervised Sparse Statistical Inference

Smoothed Estimation on Optimal Treatment Regime under Semisupervised Setting in Randomized Trials

Adaptive Semi-Supervised Inference for Optimal Treatment Decisions with Electronic Medical Record Data

Model-Assisted Uniformly Honest Inference for Optimal Treatment Regimes in High Dimension

Improved Inference for Imputation-Based Semisupervised Learning Under Misspecified Setting