Improved prediction rule ensembling through model-based data generation

Benny Markovitch,Marjolein Fokkema
DOI: https://doi.org/10.48550/arXiv.2109.13672
2021-09-28
Abstract:Prediction rule ensembles (PRE) provide interpretable prediction models with relatively high <a class="link-external link-http" href="http://accuracy.PRE" rel="external noopener nofollow">this http URL</a> obtain a large set of decision rules from a (boosted) decision tree ensemble, and achieves sparsitythrough application of Lasso-penalized regression. This article examines the use of surrogate modelsto improve performance of PRE, wherein the Lasso regression is trained with the help of a massivedataset generated by the (boosted) decision tree ensemble. This use of model-based data generationmay improve the stability and consistency of the Lasso step, thus leading to improved overallperformance. We propose two surrogacy approaches, and evaluate them on simulated and existingdatasets, in terms of sparsity and predictive accuracy. The results indicate that the use of surrogacymodels can substantially improve the sparsity of PRE, while retaining predictive accuracy, especiallythrough the use of a nested surrogacy approach.
Machine Learning
What problem does this paper attempt to address?