Enhancing Data Integrity in Pharmaceutical Equipment Root Cause Analysis Through a Python-Based Data-Mining Framework
Open Access DepositedMisclassification of equipment root causes in pharmaceutical deviation records can weaken data integrity and affect trending and downstream quality decisions. This praxis evaluates whether a Python-based data-mining screening framework could help to identify equipment-labeled deviation records that were misclassified.A manually adjudicated gold-standard dataset of 478 equipment-labeled deviation records was drawn from a total dataset of 2,232 records. The gold-standard dataset was used to examine misclassification patterns across structured categorical predictors and to train supervised models that combined structured fields with Term Frequency-Inverse Document Frequency (TF-IDF) features from derived narratives. Logistic Regression (LR), Random Forest (RF), and linear Support Vector Machine (SVM) were evaluated using stratified 5-fold cross-validation and a held-out test set at selected operating thresholds. All three models showed high recall for the misclassified class on the held-out test set (0.98 for LR and RF, 1.00 for SVM) but precision, F1, and macro-F1 differed across models. Precision for the misclassified class was 0.87 for LR, 0.57 for RF, and 0.95 for linear SVM. The corresponding F1 scores were 0.92, 0.73, and 0.97. Macro-F1 across both classes was 0.90 for LR, 0.39 for RF, and 0.97 for linear SVM. Of the three models, linear SVM showed the strongest overall performance on the held-out test set. Overall, the results show that supervised screening can help to identify a smaller set of equipment-labeled records for secondary quality and engineering review while leaving final judgment to human reviewers.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.