Abstract:When machine learning supports decision-making in safety-critical systems, it is important to verify and understand the reasons why a particular output is produced. Although feature importance calculation approaches assist in interpretation, there is a lack of consensus regarding how features' importance is quantified, which makes the explanations offered for the outcomes mostly unreliable. A possible solution to address the lack of agreement is to combine the results from multiple feature importance quantifiers to reduce the variance of estimates. Our hypothesis is that this will lead to more robust and trustworthy interpretations of the contribution of each feature to machine learning predictions. To assist test this hypothesis, we propose an extensible Framework divided in four main parts: (i) traditional data pre-processing and preparation for predictive machine learning models; (ii) predictive machine learning; (iii) feature importance quantification and (iv) feature importance decision fusion using an ensemble strategy. We also introduce a novel fusion metric and compare it to the state-of-the-art. Our approach is tested on synthetic data, where the ground truth is known. We compare different fusion approaches and their results for both training and test sets. We also investigate how different characteristics within the datasets affect the feature importance ensembles studied. Results show that our feature importance ensemble Framework overall produces 15% less feature importance error compared to existing methods. Additionally, results reveal that different levels of noise in the datasets do not affect the feature importance ensembles' ability to accurately quantify feature importance, whereas the feature importance quantification error increases with the number of features and number of orthogonal informative features.

Applying Machine Learning Diversity Metrics To Data Fusion In Information Retrieval

A survey on machine learning for data fusion

Exploring Fusion Techniques in Multimodal AI-Based Recruitment: Insights from FairCVdb

Structural Learning of Diverse Ranking.

Diversity in Machine Learning

Combining Multiple Retrieval Systems Using Combinatorial Fusion Analysis and Rank-Score Characteristic Function

Multimodal data fusion using signal/image processing methods for multi-class machine learning

Learning Fair Classifiers via Min-Max F-divergence Regularization

Result Diversification in Search and Recommendation: A Survey

Combination of Multiple Retrieval Systems Using Rank-Score Function and Cognitive Diversity

Learning to Diversify via Weighted Kernels for Classifier Ensemble

Enhancing Criminal Case Matching through Diverse Legal Factors

Unsupervised metric fusion by cross diffusion

Towards a More Reliable Interpretation of Machine Learning Outputs for Safety-Critical Systems using Feature Importance Fusion

Improved Combination of Multiple Retrieval Systems Using a Dynamic Combinatorial Fusion Algorithm.

Advancing Academic Knowledge Retrieval via LLM-enhanced Representation Similarity Fusion

On Diversity in Discriminative Neural Networks

An Optimization Framework for Merging Multiple Result Lists

On the Relationships among Various Diversity Measures in Multiple Classifier Systems

Bias and Unfairness in Information Retrieval Systems: New Challenges in the LLM Era

Directly Optimize Diversity Evaluation Measures: A New Approach to Search Result Diversification.