Canonical Correlation Analysis for Analyzing Sequences of Medical Billing Codes

Corinne L. Jones,Sham M. Kakade,Lucas W. Thornblade,David R. Flum,Abraham D. Flaxman
DOI: https://doi.org/10.48550/arXiv.1612.00516
IF: 5.414
2016-12-01
Machine Learning
Abstract:We propose using canonical correlation analysis (CCA) to generate features from sequences of medical billing codes. Applying this novel use of CCA to a database of medical billing codes for patients with diverticulitis, we first demonstrate that the CCA embeddings capture meaningful relationships among the codes. We then generate features from these embeddings and establish their usefulness in predicting future elective surgery for diverticulitis, an important marker in efforts for reducing costs in healthcare.
What problem does this paper attempt to address?