Multi-view GCN for loan default risk prediction

Zihao Li,Yakun Chen,Xianzhi Wang,Lina Yao,Guandong Xu
DOI: https://doi.org/10.1007/s00521-024-09695-x
2024-04-19
Neural Computing and Applications
Abstract:Abstract As a significant application of machine learning in financial scenarios, loan default risk prediction aims to evaluate the client’s default probability. However, most existing deep learning solutions treat each application as an independent individual, neglecting the explicit connections among different application records. Besides, these attempts suffer from the problem of missing data and imbalanced distribution (i.e., the default records are small samples against all the applications). We believe similar records could provide some auxiliary signals, which are of critical importance to alleviate the data missing issue and facilitate data argumentation. To this end, we propose multi-view loan application graphs, dubbed MLAGs. By evaluating the similarity between the records, a loan application graph can be constructed. Furthermore, we arrange different similarity thresholds to organize various graph structures for multi-graph constructions; thus, a variety of representations can be generated via information propagation and aggregation for small sample argumentation. Consequently, the imbalanced data distribution and missing values issues can be alleviated effectively. We conduct experiments on three public datasets from real-world home credit and P2P lending platforms, which show that MGCN outperforms both conventional and deep learning models. Ablation studies also illustrated the validity of each module design.
computer science, artificial intelligence
What problem does this paper attempt to address?
This paper proposes a solution to the problem of predicting loan default risk. Existing deep learning methods typically treat each loan application as an independent entity when processing loan applications, ignoring explicit connections between different application records. In addition, these methods also face the issues of data missing and imbalanced distributions (i.e., default records are small samples). The paper argues that similar records can provide auxiliary signals that help alleviate data missing problems and enhance data argumentation. To this end, the paper introduces Multi-View Loan Application Graphs (MLAGs), which construct graph structures by computing similarities between records and organize multiple graphs with different similarity thresholds to generate diversified representations for enhancing small-sample argumentation. This approach effectively mitigates data imbalance and missing values issues. Experiments conducted on three real-world public datasets from consumer credit and P2P lending platforms demonstrate that MGCN outperforms traditional and deep learning models, and the effectiveness of the module design is validated through ablation studies.