Abstract:Pattern discovery is a machine learning technique that aims to find sets of items, subsequences, or substructures that are present in a dataset with a higher frequency value than a manually set threshold. This process helps to identify recurring patterns or relationships within the data, allowing for valuable insights and knowledge extraction. In this work, we propose Information Gained Subgroup Discovery (IGSD), a new SD algorithm for pattern discovery that combines Information Gain (IG) and Odds Ratio (OR) as a multi-criteria for pattern selection. The algorithm tries to tackle some limitations of state-of-the-art SD algorithms like the need for fine-tuning of key parameters for each dataset, usage of a single pattern search criteria set by hand, usage of non-overlapping data structures for subgroup space exploration, and the impossibility to search for patterns by fixing some relevant dataset variables. Thus, we compare the performance of IGSD with two state-of-the-art SD algorithms: FSSD and SSD++. Eleven datasets are assessed using these algorithms. For the performance evaluation, we also propose to complement standard SD measures with IG, OR, and p-value. Obtained results show that FSSD and SSD++ algorithms provide less reliable patterns and reduced sets of patterns than IGSD algorithm for all datasets considered. Additionally, IGSD provides better OR values than FSSD and SSD++, stating a higher dependence between patterns and targets. Moreover, patterns obtained for one of the datasets used, have been validated by a group of domain experts. Thus, patterns provided by IGSD show better agreement with experts than patterns obtained by FSSD and SSD++ algorithms. These results demonstrate the suitability of the IGSD as a method for pattern discovery and suggest that the inclusion of non-standard SD metrics allows to better evaluate discovered patterns.

Discovering Statistically Non-Redundant Subgroups

Efficient redundancy reduced subgroup discovery via quadratic programming

Subgroup identification in clinical trials: an overview of available methods and their implementations with R

Group Detection in Real-World Social Networks

Robust subgroup discovery

A Shrinkage Likelihood Ratio Test for High-Dimensional Subgroup Analysis with a Logistic-Normal Mixture Model

Data-Driven Subgroup Identification for Linear Regression

Constrained Latent Dirichlet Allocation For Subgroup Discovery With Topic Rules

ROBUST SUBGROUP IDENTIFICATION

Using Constraints to Discover Sparse and Alternative Subgroup Descriptions

A Two-Step Non-redundant Subspace Clustering Approach.

Statistical Performance Guarantee for Subgroup Identification with Generic Machine Learning

A Subgroup Discovery Algorithm Based on Genetic Fuzzy Systems

Subgroup Identification Based on Quantitative Objectives

Subgroup Identification Using the personalized Package

Expert-Guided Subgroup Discovery: Methodology and Application

Multiply Robust Subgroup Identification for Longitudinal Data with Dropouts Via Median Regression

Optimal False Discovery Rate Control for Large Scale Multiple Testing with Auxiliary Information

Intersectional fair ranking via subgroup divergence

A new algorithm for Subgroup Set Discovery based on Information Gain

Subgroup Analysis of Linear Models with Measurement Error