CauseRuDi: Explaining Behavior Sequence Models by Causal Statistics Generation and Rule Distillation

Yao Zhang,Yun Xiong,Yiheng Sun,Yucheng Jin,Caihua Shan,Tian Lu,Hui Song,Shengli Sun
DOI: https://doi.org/10.1109/tkde.2024.3487625
IF: 9.235
2024-01-01
IEEE Transactions on Knowledge and Data Engineering
Abstract:Risk scoring systems have been widely deployed in many applications, which assign risk scores to users according to their behavior sequences. Though many deep learning methods with sophisticated designs have achieved promising results, the black-box nature hinders their applications due to fairness, explainability, and compliance consideration. Rule-based systems are considered reliable in these sensitive scenarios. However, building a rule system is labor-intensive. Experts need to find informative statistics from user behavior sequences, design rules based on statistics and assign weights to each rule. In this paper, we bridge the gap between effective but black-box models and transparent rule models. We propose a two-stage framework, CauseRuDi, that distills the knowledge of black-box teacher models into rule-based student models. We design a Monte Carlo tree search-based statistics generation method that maximizes the correlation or dependence between the generated statistics and the teacher model's outputs. We formulate a sequential move game and a simultaneous move coalitional game to generate multiple statistics. Then statistics are composed into logical rules with our proposed neural logical networks by mimicking the outputs of teacher models. We evaluate CauseRuDi on three real-world public datasets and an industrial dataset to demonstrate its effectiveness.
What problem does this paper attempt to address?