Abstract:With the increased deployment of machine learning models in various real-world applications, researchers and practitioners alike have emphasized the need for explanations of model behaviour. To this end, two broad strategies have been outlined in prior literature to explain models. Post hoc explanation methods explain the behaviour of complex black-box models by identifying features critical to model predictions; however, prior work has shown that these explanations may not be faithful, in that they incorrectly attribute high importance to features that are unimportant or non-discriminative for the underlying task. Inherently interpretable models, on the other hand, circumvent these issues by explicitly encoding explanations into model architecture, meaning their explanations are naturally faithful, but they often exhibit poor predictive performance due to their limited expressive power. In this work, we identify a key reason for the lack of faithfulness of feature attributions: the lack of robustness of the underlying black-box models, especially to the erasure of unimportant distractor features in the input. To address this issue, we propose Distractor Erasure Tuning (DiET), a method that adapts black-box models to be robust to distractor erasure, thus providing discriminative and faithful attributions. This strategy naturally combines the ease of use of post hoc explanations with the faithfulness of inherently interpretable models. We perform extensive experiments on semi-synthetic and real-world datasets and show that DiET produces models that (1) closely approximate the original black-box models they are intended to explain, and (2) yield explanations that match approximate ground truths available by construction. Our code is made public at <a class="link-external link-https" href="https://github.com/AI4LIFE-GROUP/DiET" rel="external noopener nofollow">this https URL</a>.

AIDE: Antithetical, Intent-based, and Diverse Example-Based Explanations

CLIMAX: An exploration of Classifier-Based Contrastive Explanations

T-Explainer: A Model-Agnostic Explainability Framework Based on Gradients

Advancing Post Hoc Case Based Explanation with Feature Highlighting

One Explanation Does Not Fit All: A Toolkit and Taxonomy of AI Explainability Techniques

A Comparative Analysis of Model Agnostic Techniques for Explainable Artificial Intelligence

I-CEE: Tailoring Explanations of Image Classification Models to User Expertise

Explanation as a process: user-centric construction of multi-level and multi-modal explanations

Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability

Selective Explanations: Leveraging Human Input to Align Explainable AI

NoMatterXAI: Generating "No Matter What" Alterfactual Examples for Explaining Black-Box Text Classification Models

Gradient-free Post-hoc Explainability Using Distillation Aided Learnable Approach

How Well Do Feature-Additive Explainers Explain Feature-Additive Predictors?

A Unified Framework for Input Feature Attribution Analysis

Post-hoc Explanation Options for XAI in Deep Learning: The Insight Centre for Data Analytics Perspective

Explain, Edit, and Understand: Rethinking User Study Design for Evaluating Model Explanations

Regularized adversarial examples for model interpretability

Explaining Machine Learning Classifiers through Diverse Counterfactual Explanations

Altruist: Argumentative Explanations through Local Interpretations of Predictive Models

Interpretable Data-Based Explanations for Fairness Debugging

Selective Explanations