Abstract:The Rashômon Effect, applied in Explainable Machine Learning, refers to the disagreement between the explanations provided by various attribution explainers and to the dissimilarity across multiple explanations generated by a particular explainer for a single instance from the dataset (differences between feature importances and their associated signs and ranks), an undesirable outcome especially in sensitive domains such as healthcare or finance. We propose a method inspired from textual-case based reasoning for aligning explanations from various explainers in order to resolve the disagreement and dissimilarity problems. We iteratively generated a number of 100 explanations for each instance from six popular datasets, using three prevalent feature attribution explainers: LIME, Anchors and SHAP (with the variations Tree SHAP and Kernel SHAP) and consequently applied a global cluster-based aggregation strategy that quantifies alignment and reveals similarities and associations between explanations. We evaluated our method by weighting the -NN algorithm with agreed feature overlap explanation weights and compared it to a non-weighted -NN predictor, having as task binary classification. Also, we compared the results of the weighted -NN algorithm using aggregated feature overlap explanation weights to the weighted -NN algorithm using weights produced by a single explanation method (either LIME, SHAP or Anchors). Our global alignment method benefited the most from a hybridization with feature importance scores (information gain), that was essential for acquiring a more accurate estimate of disagreement, for enabling explainers to reach a consensus across multiple explanations and for supporting effective model learning through improved classification performance.

Axiomatic Aggregations of Abductive Explanations

Evaluating and Aggregating Feature-based Model Explanations

Succint Interaction-Aware Explanations

Provably Better Explanations with Optimized Aggregation of Feature Attributions

Local Interpretable Model Agnostic Shap Explanations for machine learning models

Unified Explanations in Machine Learning Models: A Perturbation Approach

Model Agnostic Multilevel Explanations

Clarity in complexity: how aggregating explanations resolves the disagreement problem

An Axiomatic Approach to Model-Agnostic Concept Explanations

A Comparative Analysis of Model Agnostic Techniques for Explainable Artificial Intelligence

A Perspective on Explainable Artificial Intelligence Methods: SHAP and LIME

On the Connection between Game-Theoretic Feature Attributions and Counterfactual Explanations

Sufficient and Necessary Explanations (and What Lies in Between)

Locally-Minimal Probabilistic Explanations

Axiomatic Characterisations of Sample-based Explainers

How Well Do Feature-Additive Explainers Explain Feature-Additive Predictors?

Explaining the Model and Feature Dependencies by Decomposition of the Shapley Value

Selective Explanations

Causality-Aware Local Interpretable Model-Agnostic Explanations

Explaining Aggregates for Exploratory Analytics

LaPLACE: Probabilistic Local Model-Agnostic Causal Explanations