Comparison of classical machine learning methods in the task of obesity level classification
Article Sidebar
Issue Vol. 40 (2026)
-
Analysis of the capabilities of predictive artificial intelligence models in corporate risk management
Kacper Ziemski188-192
-
Usability and availability of selected e-commerce services
Marcin Kozicki, Maria Skublewska-Paszkowska193-200
-
Comparison of C++ and Python performance based on selected algorithms
Szymon Bogucki, Kacper Burda201-205
-
Security analysis of selected web applications using vulnerability scanners
Mariusz Choroś, Marta Dziuba-Kozieł206-212
-
Comparison of Java and .NET reflection mechanisms for dynamic module loading: a performance benchmark study
Michał Mazur, Sebastian Maruszak, Marek Miłosz213-217
-
Comparative analysis of network vulnerability detection tools
Mateusz Zdunek218-225
-
Evaluation of mobile applications for personal finance management using the MARS scale
Łukasz Nikiel, Artsiom Patskevich, Marek Miłosz226-231
-
Comparison of the effectiveness of roulette betting strategies using Monte Carlo simulation
Marek Sarnecki232-238
-
Analysis of optimization capabilities of selected database management systems
Paweł Tarkiewicz, Małgorzata Plechawska-Wójcik239-246
-
Comparative analysis of Espresso and Appium frameworks for automated UI testing of Android mobile applications
Jakub Derkacz247-254
-
Comparative analysis of the applicability of artificial intelligence models for code generation
Patryk Warchoł, Małgorzata Plechawska-Wójcik255-262
-
Comparison of the effectiveness of selected tools for detecting texts generated by artificial intelligence
Marcin Brodacki, Małgorzata Plechawska-Wójcik263-269
-
Comparative analysis of selected containerization tools in terms of MCP
Paweł Jan Tłusty, Maciej Pańczyk270-276
-
SpikeCliff effect: empirical analysis of deterministic timing discontinuities in sponge-based XOF functions
Łukasz Wójcik, Stanisław Lota277-282
-
Comparison of AI agents for creating SQL queries
Julia Sierpień, Maria Skublewska-Paszkowska283-288
-
Comparative analysis of the performance of PostgreSQL and Neo4j databases in the context of genealogical queries
Michał Muzyka, Mateusz Niedźwiedź, Marek Miłosz289-296
-
Evaluation of the effectiveness of static and dynamic methods in malware analysis
Dominik Tracz, Daniel Sawicki, Konrad Gromaszek297-303
-
Comparison of classical machine learning methods in the task of obesity level classification
Paweł Biesaga, Paweł Powroźnik304-312
Main Article Content
Authors
Abstract
Obesity is a critical public-health problem amenable to data-driven risk assessment. We compare eight classical classifiers on the UCI ObesityDataSet (2087 samples, 7 classes) under a leakage-free protocol: stratified 10-fold cross-validation, randomised-search tuning, pairwise McNemar and Wilcoxon tests with Holm–Bonferroni correction, and multi-method interpretability (split gain, SHAP, permutation, LIME). XGBoost reached 96.17 % test accuracy (F1 = 0.962, κ = 0.955); paired tests show no significant edge over tuned LightGBM, SVM-RBF, Random Forest or Logistic Regression. All four importance measures rank body weight and height first, and LIME recovers BMI-based thresholds, supporting gradient boosting for clinical decision support.
Keywords:
Sustainable Development Goal (SDG)
- Good health and well-being
References
[1] World Health Organization, Obesity and overweight – Fact sheet, https://www.who.int/news-room/fact-sheets/detail/obesity-and-overweight, [2022].
[2] F. M. Palechor, A. de la Hoz Manotas, Dataset for estimation of obesity levels based on eating habits and physical condition in individuals from Colombia, Peru and Mexico, Data in Brief 25 (2019) 104344, https://doi.org/10.1016/j.dib.2019.104344.
[3] L. Grinsztajn, E. Oyallon, G. Varoquaux, Why do tree-based models still outperform deep learning on typical tabular data?, Advances in Neural Information Processing Systems 35 (2022) 507–520, https://arxiv.org/abs/2207.08815.
[4] R. Shwartz-Ziv, A. Armon, Tabular Data: Deep Learning is Not All You Need, Information Fusion 81 (2022) 84–90, https://doi.org/10.1016/j.inffus.2021.11.011.
[5] M. T. Ribeiro, S. Singh, C. Guestrin, “Why Should I Trust You?”: Explaining the Predictions of Any Classifier, Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (2016) 1135–1144, https://doi.org/10.1145/2939672.2939778.
[6] D. Chicco, G. Jurman, Machine learning can predict survival of patients with heart failure from serum creatinine and ejection fraction alone, BMC Medical Informatics and Decision Making 20(1) (2020) 16, https://doi.org/10.1186/s12911-020-1023-5.
[7] M. Fernández-Delgado, E. Cernadas, S. Barro, D. Amorim, Do we need hundreds of classifiers to solve real world classification problems?, Journal of Machine Learning Research 15 (2014) 3133–3181, https://jmlr.org/papers/v15/delgado14a.html.
[8] F. Ferdowsy, K. S. A. Rahi, M. Jabiullah, M. T. Habib, A machine learning approach for obesity risk prediction, Current Research in Behavioral Sciences 2 (2021) 100053, https://doi.org/10.1016/j.crbeha.2021.100053.
[9] T. Chen, C. Guestrin, XGBoost: A Scalable Tree Boosting System, Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (2016) 785–794, https://doi.org/10.1145/2939672.2939785.
[10] G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, T.-Y. Liu, LightGBM: A Highly Efficient Gradient Boosting Decision Tree, Advances in Neural Information Processing Systems 30 (2017) 3146–3154, https://papers.nips.cc/paper/2017/hash/6449f44a102fde848669bdd9eb6b76fa-Abstract.html.
[11] L. Rokach, Ensemble-based classifiers, Artificial Intelligence Review 33(1–2) (2010) 1–39, https://doi.org/10.1007/s10462-009-9124-7.
[12] I. D. Mienye, Y. Sun, A Survey of Ensemble Learning: Concepts, Algorithms, Applications, and Prospects, IEEE Access 10 (2022) 99129–99149, https://doi.org/10.1109/ACCESS.2022.3207287.
[13] J. Cervantes, F. Garcia-Lamont, L. Rodríguez-Mazahua, A. Lopez, A comprehensive survey on support vector machine classification: Applications, challenges and trends, Neurocomputing 408 (2020) 189–215, https://doi.org/10.1016/j.neucom.2019.10.118.
[14] D. W. Hosmer, S. Lemeshow, R. X. Sturdivant, Applied Logistic Regression, 3rd ed., John Wiley & Sons, Hoboken, NJ, 2013, https://doi.org/10.1002/9781118548387.
[15] I. Rish, An empirical study of the naive Bayes classifier, In IJCAI 2001 Workshop on Empirical Methods in Artificial Intelligence (2001) 41–46.
[16] T. Cover, P. Hart, Nearest neighbor pattern classification, IEEE Transactions on Information Theory 13(1) (1967) 21–27, https://doi.org/10.1109/TIT.1967.1053964.
[17] C. Cortes, V. Vapnik, Support-vector networks, Machine Learning 20(3) (1995) 273–297, https://doi.org/10.1007/BF00994018.
[18] L. Breiman, J. H. Friedman, R. A. Olshen, C. J. Stone, Classification and Regression Trees, Wadsworth, Belmont, CA, 1984.
[19] L. Breiman, Random Forests, Machine Learning 45(1) (2001) 5–32, https://doi.org/10.1023/A:1010933404324.
[20] T. Fawcett, An introduction to ROC analysis, Pattern Recognition Letters 27(8) (2006) 861–874, https://doi.org/10.1016/j.patrec.2005.10.010.
[21] T. G. Dietterich, Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms, Neural Computation 10(7) (1998) 1895–1923, https://doi.org/10.1162/089976698300017197.
[22] J. Demšar, Statistical Comparisons of Classifiers over Multiple Data Sets, Journal of Machine Learning Research 7 (2006) 1–30, https://jmlr.org/papers/v7/demsar06a.html.
[23] S. Holm, A Simple Sequentially Rejective Multiple Test Procedure, Scandinavian Journal of Statistics 6(2) (1979) 65–70, https://www.jstor.org/stable/4615733.
[24] S. M. Lundberg, S.-I. Lee, A Unified Approach to Interpreting Model Predictions, Advances in Neural Information Processing Systems 30 (2017) 4765–4774, https://papers.nips.cc/paper/2017/hash/8a20a8621978632d76c43dfd28b67767-Abstract.html.
[25] N. V. Chawla, K. W. Bowyer, L. O. Hall, W. P. Kegelmeyer, SMOTE: Synthetic Minority Over-sampling Technique, Journal of Artificial Intelligence Research 16 (2002) 321–357, https://doi.org/10.1613/jair.953.
[26] J. Snoek, H. Larochelle, R. P. Adams, Practical Bayesian Optimization of Machine Learning Algorithms, Advances in Neural Information Processing Systems 25 (2012) 2951–2959, https://papers.nips.cc/paper/2012/hash/05311655a15b75fab86956663e1819cd-Abstract.html.
[27] A. Szymczyk, M. Skublewska-Paszkowska, P. Powroźnik, CNN-Based Ensemble Architectures with Explainable AI for Cutaneous Melanoma Identification, Acta Mechanica et Automatica 20(2) (2026) 317–330, https://doi.org/10.65731/ama/2026-0033.
Article Details
Abstract views: 1

