Random Forest and LightGBM Comparison for Acute Pain Diagnosis Using SMOTE on an Expert-Labeled Dataset
DOI:
https://doi.org/10.35314/x2getn56Keywords:
acute pain diagnosis, expert-based dataset, LightGBM, Random Forest, SMOTEAbstract
Limited healthcare personnel may delay early pain assessment and encourage self-medication, increasing medication-error risk. However, evidence remains limited regarding whether bagging or boosting is more suitable for multiclass acute pain classification using imbalanced, expert-system-derived symptom data and whether SMOTE improves performance. This study compared Random Forest as a bagging approach and LightGBM as a boosting approach for classifying nine acute pain diagnostic classes without SMOTE and with SMOTE using k_neighbors=1 and 5. The dataset comprised 2,722 records and 36 discrete symptom features. Of 125 representative symptom combinations reviewed by a medical expert, 115 were considered appropriate; the remaining records were synthetically generated using the same expert-system knowledge base and inference mechanism. Data were divided using stratified 80:20 sampling, while model configuration was evaluated using five-fold cross-validation. SMOTE was applied only to training data within each fold. LightGBM without SMOTE achieved the best performance, with 83.49% accuracy, a macro F1-score of 0.81, and a weighted F1-score of 0.83, compared with 80.18%, 0.77, and 0.80 for Random Forest. With SMOTE, Random Forest achieved 78.35% and 77.61% accuracy, while LightGBM achieved 81.10% and 82.39%. Thus, LightGBM without SMOTE performed best for this dataset. Validation using real clinical data and multiple experts is required
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Innovation and Technology Polbeng Series on Informatics (INOVTEK Polbeng - Seri Informatika)

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.









