Abstract:As neural networks become more popular, the need for accompanying uncertainty estimates increases. There are currently two main approaches to test the quality of these estimates. Most methods output a density. They can be compared by evaluating their loglikelihood on a test set. Other methods output a prediction interval directly. These methods are often tested by examining the fraction of test points that fall inside the corresponding prediction intervals. Intuitively, both approaches seem logical. However, we demonstrate through both theoretical arguments and simulations that both ways of evaluating the quality of uncertainty estimates have serious flaws. Firstly, both approaches cannot disentangle the separate components that jointly create the predictive uncertainty, making it difficult to evaluate the quality of the estimates of these components. Specifically, the quality of a confidence interval cannot reliably be tested by estimating the performance of a prediction interval. Secondly, the loglikelihood does not allow a comparison between methods that output a prediction interval directly and methods that output a density. A better loglikelihood also does not necessarily guarantee better prediction intervals, which is what the methods are often used for in practice. Moreover, the current approach to test prediction intervals directly has additional flaws. We show why testing a prediction or confidence interval on a single test set is fundamentally flawed. At best, marginal coverage is measured, implicitly averaging out overconfident and underconfident predictions. A much more desirable property is pointwise coverage, requiring the correct coverage for each prediction. We demonstrate through practical examples that these effects can result in favouring a method, based on the predictive uncertainty, that has undesirable behaviour of the confidence or prediction intervals. Finally, we propose a simulation-based testing approach that addresses these problems while still allowing easy comparison between different methods. This approach can be used for the development of new uncertainty quantification methods.

Analytical results for uncertainty propagation through trained machine learning regression models

Analytical Uncertainty Propagation in Neural Networks

Uncertainty Quantification and Propagation in Atomistic Machine Learning

Uncertainty Prediction for Machine Learning Models of Material Properties

Quantifying the Prediction Uncertainty of Machine Learning Models for Individual Data

Quantification of Deep Neural Network Prediction Uncertainties for VVUQ of Machine Learning Models

How to evaluate uncertainty estimates in machine learning for regression?

Uncertainty Quantification Metrics for Deep Regression

Addressing Uncertainty on Machine Learning Models for Long-Period Fiber Grating Signal Conditioning Using Monte Carlo Method

Uncertainty Quantification For Turbulent Flows with Machine Learning

Characterizing Uncertainty in Machine Learning for Chemistry

A comparative study of conformal prediction methods for valid uncertainty quantification in machine learning

Machine learning-based moment closure model for the semiconductor Boltzmann equation with uncertainties

Propagation of Linear Uncertainties through Multiline Thru-Reflect-Line Calibration

Physics Based & Machine Learning Methods For Uncertainty Estimation In Turbulence Modeling

Quantifying Uncertainty with Probabilistic Machine Learning Modeling in Wireless Sensing

Quantitative assessment of machine learning reliability and resilience

Negative impact of heavy-tailed uncertainty and error distributions on the reliability of calibration statistics for machine learning regression tasks

Predicting uncertainty of machine learning models for modelling nitrate pollution of groundwater using quantile regression and UNEEC methods

A review of predictive uncertainty estimation with machine learning

Uncertainty Wrapper in the medical domain: Establishing transparent uncertainty quantification for opaque machine learning models in practice