StratLearn-z: Improved photo-$z$ estimation from spectroscopic data subject to selection effects

Chiara Moretti,Maximilian Autenrieth,Riccardo Serra,Roberto Trotta,David A. van Dyk,Andrei Mesinger
2024-09-30
Abstract:A precise measurement of photometric redshifts (photo-z) is key for the success of modern photometric galaxy surveys. Machine learning (ML) methods show great promise in this context, but suffer from covariate shift (CS) in training sets due to selection bias where interesting sources are underrepresented, and the corresponding ML models show poor generalisation properties. We present an application of the StratLearn method to the estimation of photo-z, validating against simulations where we enforce the presence of CS to different degrees. StratLearn is a statistically principled approach that relies on splitting the source and target datasets into strata based on estimated propensity scores (i.e. the probability for an object to be in the source set given its observed covariates). After stratification, two conditional density estimators are fit separately to each stratum, then combined via a weighted average. We benchmark our results against the GPz algorithm, quantifying the performance of the two codes with a set of metrics. Our results show that the StratLearn-z metrics are only marginally affected by the presence of CS, while GPz shows a significant degradation of performance in the photo-z prediction for fainter objects. For the strongest CS scenario, StratLearn-z yields a reduced fraction of catastrophic errors, a factor of 2 improvement for the RMSE and one order of magnitude improvement on the bias. We also assess the quality of the conditional redshift estimates with the probability integral transform (PIT). The PIT distribution obtained from StratLearn-z features fat fewer outliers and is symmetric, i.e. the predictions appear to be centered around the true redshift value, despite showing a conservative estimation of the spread of the conditional redshift distributions. Our julia implementation of the method is available at \url{<a class="link-external link-https" href="https://github.com/chiaramoretti/StratLearn-z" rel="external noopener nofollow">this https URL</a>}.
Cosmology and Nongalactic Astrophysics
What problem does this paper attempt to address?