Comparative Evaluation of Machine Learning Classifiers for Predicting Pathological Complete Response in HER2-Heterogeneous Breast Cancer Treated with T-DM1 and Pertuzumab
Open Access DepositedResponse to HER2-targeted therapy in HER2-positive breast cancer (BC) is variable, particularly in tumors with intratumoral heterogeneity, where some patients achieve pathological complete response (pCR) while others exhibit residual disease. This study evaluates whether transcriptomic profiles can be used to classify pCR status using machine learning (ML). We hypothesize that gene expression patterns derived from RNA-seq data encode predictive signals associated with treatment response. Supervised ML models—including support vector machines (SVM), random forests (RF), and k-nearest neighbors (KNN)—were implemented using scikit-learn and evaluated using accuracy, macro F1 score, and precision–recall area under the curve (PR-AUC). Across models, performance was moderate
SVMs with linear and polynomial kernels consistently outperformed other classifiers, achieving the highest F1 and PR-AUC scores. Optimal models utilized large feature subsets (~10,000 genes), suggesting a distributed transcriptomic signal underlying treatment response. Furthermore, a permutation importance of the final model found genes predicted to significantly affect model learning (IGHG1). These results demonstrate that ML models can capture biologically relevant patterns associated with pCR in HER2-positive BC, including heterogeneous tumors. This approach may support prediction of treatment response from RNA-seq data, potentially reducing reliance on extensive pathological assessment. Future work will focus on pathway-level feature engineering, model calibration, and external validation to improve clinical applicability.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.