Comparison of classical machine learning methods in the task of obesity level classification

Main Article Content

Paweł Biesaga

s97568@pollub.edu.pl

https://orcid.org/0009-0004-8146-6036
Paweł Powroźnik

p.powroznik@pollub.pl

Abstract

Obesity is a critical public-health problem amenable to data-driven risk assessment. We compare eight classical classifiers on the UCI ObesityDataSet (2087 samples, 7 classes) under a leakage-free protocol: stratified 10-fold cross-validation, randomised-search tuning, pairwise McNemar and Wilcoxon tests with Holm–Bonferroni correction, and multi-method interpretability (split gain, SHAP, permutation, LIME). XGBoost reached 96.17 % test accuracy (F1 = 0.962, κ = 0.955); paired tests show no significant edge over tuned LightGBM, SVM-RBF, Random Forest or Logistic Regression. All four importance measures rank body weight and height first, and LIME recovers BMI-based thresholds, supporting gradient boosting for clinical decision support.

Keywords:

Machine Learning, obesity classification, gradient boosting, statistical significance, shap, lime

Sustainable Development Goal (SDG)

  • Good health and well-being

References

Article Details

Biesaga, P., & Powroźnik, P. (2026). Comparison of classical machine learning methods in the task of obesity level classification. Journal of Computer Sciences Institute, 40, 304-312. https://doi.org/10.35784/jcsi.9922