Parallelogram-mode: A novel clustering method for categorical data
Article Sidebar
Issue Vol. 22 No. 3 (2026)
-
Enforcing label consistency and lowering labelling time in infants’ pose data via semi-automatic annotation
Greta DI MARINO, Emanuele CARDINALE, Alessio CORREANI, Lucia MIGLIORELLI, Sara MOCCIA1-14
-
A data-driven framework for AI adoption efficiency assessment using hybrid DEA and machine learning methods
Ewa CHODAKOWSKA15-29
-
ARECA-Lite: A lightweight modified ArecaNet with reduced complexity for real-time robust facial emotion recognition
Mustapha Abdelkader LAOUMIR, Amina KINANE DAOUADJI, Fatima BENDELLA30-51
-
Designing vehicle structures using materials with a low carbon footprint
Bartosz ŁATA, Jacek CZARNIGOWSKI, Wiktor ISKRA, Miłosz CHWIEJCZAK, Iga KOPEĆ52-61
-
Parallelogram-mode: A novel clustering method for categorical data
Ashuza KUDERHA, Olamma IHEANETU62-81
-
Digitalisation of relay protection and implementation of main protections in digital form
Dmytro DANYLCHENKO, Vladyslav TSIUPA, Oleksandr MIROSHNYK, Taras SHCHUR, Katarzyna PIOTROWSKA82-96
-
Modified snake optimizer algorithm for solving the permutation flow shop scheduling problem
Hassan ALMAZINI, Salah MORTADA, Hussein Fouad ALMAZINI97-107
-
Computer-based data processing approaches to production scrap management
Łukasz WÓJCIK, Arkadiusz GOLA, Jakub PIZOŃ108-120
-
A hybrid parameter-tuning for adaptive variable-length particle swarm optimisation in cancer feature selection
Shir Li WANG, Siti RAMADHANI, Muhammad FIKRY, Haldi BUDIMAN, Theam Foo NG, Sumayyah DZULKIFLY, Roziana ARIFFIN121-147
-
Beyond classical optimisation: Toward feasibility-aware computational architectures for synchronised systems
Grzegorz BOCEWICZ, Czesław SMUTNICKI, Zbigniew BANASZAK148-167
-
A composite latency model for evaluating hybrid OLTP/OLAP information systems
Volodymyr SOLOHUB, Volodymyr PASHKEVYCH168-180
-
Automatic methods for 3D motion trajectories gap filling: Custom-based Kalman vs. BiLSTM
Kamil ŻELAZOWSKI, Wojciech WOJCIECHEWICZ, Maria SKUBLEWSKA-PASZKOWSKA, Paweł POWROŹNIK181-195
-
Implementation of an IEC 61215-oriented photovoltaic module test emulator with integrated predictive maintenance capabilities
Aristide TOLOK NELEM, Yannick Antoine ABANDA, Steyve Samson NYATTE, Mathieu Jean Pierre PESDJOCK, Achille MELINGUI, Pierre ELE196-218
-
Modelling the predictive reliability of rotating machines using Artificial Intelligence.
Fernand Joseph TOUKAP NONO, Tokoue Ngatcha DIANORRÉ, Offole FLORENC, Mouzong Pemi MARCELIN219-243
-
Anomaly detection in vibroarthrographic signals using handcrafted signal features and one-class methods
Robert KARPIŃSKI, Arkadiusz SYTA244–261
Archives
-
Vol. 22 No. 3
2026-09-30 15
-
Vol. 22 No. 2
2026-06-30 15
-
Vol. 22 No. 1
2026-03-31 15
-
Vol. 21 No. 4
2025-12-31 12
-
Vol. 21 No. 3
2025-09-30 12
-
Vol. 21 No. 2
2025-06-30 12
-
Vol. 21 No. 1
2025-03-31 12
-
Vol. 20 No. 4
2024-12-31 12
-
Vol. 20 No. 3
2024-09-30 12
-
Vol. 20 No. 2
2024-06-30 12
-
Vol. 20 No. 1
2024-03-30 12
-
Vol. 19 No. 4
2023-12-31 10
-
Vol. 19 No. 3
2023-09-30 10
-
Vol. 19 No. 2
2023-06-30 10
-
Vol. 19 No. 1
2023-03-31 10
-
Vol. 18 No. 4
2022-12-30 8
-
Vol. 18 No. 3
2022-09-30 8
-
Vol. 18 No. 2
2022-06-30 8
-
Vol. 18 No. 1
2022-03-31 8
Main Article Content
Authors
kuderha.ashuzapgs@stu.cu.edu.ng
olamma.iheanetu@covenantuniversity.edu.ng
Abstract
The need for clustering methods that treat the interpretability of clustering results as valuable as cluster correctness is particularly important in high-risk domains. The heterogeneous nature of categorical data makes its clustering results less interpretable than those for numerical data. This challenge is amplified by the scarcity of works investigating the clustering of categorical data. Existing interpretable categorical data clustering algorithms do not optimise for scalability. In this paper, we address the scalability limitation by employing a vectorial approach to interpretable categorical data clustering that takes full advantage of the computational acceleration offered by modern computing hardware, such as GPUs. We formulate categorical data clustering as a rule-set search task that manipulates high-dimensional vector representations of entities in the input dataset. We propose an efficient and effective algorithm called “Coupling-Sorting-Branching” to solve the search task. This approach provides a scalable and fully interpretable method for clustering categorical data. Extensive experimental results across 14 real-world datasets demonstrate that our algorithm outperforms existing methods in cluster quality and explainability by producing IF-THEN rules.
Keywords:
Sustainable Development Goal (SDG)
- Industry, Innovation, Technology and Infrastructure
References
Aha, D. (1991). Tic-Tac-Toe endgame [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5688J
Andritsos, P., Tsaparas, P., Miller, R. J., & Sevcik, K. C. (2004). LIMBO: Scalable clustering of categorical data. In Lecture Notes in Computer Science (Vol. 2992, pp. 123–146). Springer. https://doi.org/10.1007/978-3-540-24741-8_9 DOI: https://doi.org/10.1007/978-3-540-24741-8_9
Bai, L., & Liang, J. (2022). A categorical data clustering framework on graph representation. Pattern Recognition, 128, Article 108694. https://doi.org/10.1016/j.patcog.2022.108694 DOI: https://doi.org/10.1016/j.patcog.2022.108694
Ben Salem, S., Naouali, S., & Chtourou, Z. (2018). A fast and effective partitional clustering algorithm for large categorical datasets using a k-means based approach. Computers & Electrical Engineering, 68, 463–483. https://doi.org/10.1016/j.compeleceng.2018.04.023 DOI: https://doi.org/10.1016/j.compeleceng.2018.04.023
Bhattacharjee, P., & Mitra, P. (2021). A survey of density based clustering algorithms. Frontiers of Computer Science, 15(1), Article 151308. https://doi.org/10.1007/s11704-019-9059-3 DOI: https://doi.org/10.1007/s11704-019-9059-3
Bohanec, M. (1988). Car evaluation [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5JP48
Cao, F., Liang, J., Li, D., & Zhao, X. (2013). A weighting k-modes algorithm for subspace clustering of categorical data. Neurocomputing, 108, 23–30. https://doi.org/10.1016/j.neucom.2012.11.009 DOI: https://doi.org/10.1016/j.neucom.2012.11.009
Carrizosa, E., Kurishchenko, K., Marín, A., & Romero Morales, D. (2023). On clustering and interpreting with rules by means of mathematical optimization. Computers & Operations Research, 154, Article 106180. https://doi.org/10.1016/j.cor.2023.106180 DOI: https://doi.org/10.1016/j.cor.2023.106180
Congressional voting records. (1987). [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5C01P
Dinh, D.-T., & Huynh, V.-N. (2020). k-PbC: An improved cluster center initialization for categorical data clustering. Applied Intelligence, 50(8), 2610–2632. https://doi.org/10.1007/s10489-020-01677-5 DOI: https://doi.org/10.1007/s10489-020-01677-5
Dinh, T., Wong, H., Fournier-Viger, P., Lisik, D., Ha, M. Q., Dam, H. C., & Huynh, V. N. (2025). Categorical data clustering: 25 years beyond K-modes. Expert Systems with Applications, 272, Article 126608. https://doi.org/10.1016/j.eswa.2025.126608 DOI: https://doi.org/10.1016/j.eswa.2025.126608
Downey, A. (2025). Think stats: Exploratory data analysis. O’Reilly Media.
Fisher, D. H. (1987). Knowledge acquisition via incremental conceptual clustering. Machine Learning, 2(2), 139–172. https://doi.org/10.1023/A:1022852608280 DOI: https://doi.org/10.1023/A:1022852608280
Forsyth, R. (1990). Zoo [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5R59V
Fraiman, R., Ghattas, B., & Svarc, M. (2013). Interpretable clustering using unsupervised binary trees. Advances in Data Analysis and Classification, 7(2), 125–145. https://doi.org/10.1007/s11634-013-0129-3 DOI: https://doi.org/10.1007/s11634-013-0129-3
Gormley, I. C., Murphy, T. B., & Raftery, A. E. (2023). Model-based clustering. Annual Review of Statistics and Its Application, 10(1), 573–595. https://doi.org/10.1146/annurev-statistics-033121-115326 DOI: https://doi.org/10.1146/annurev-statistics-033121-115326
Guha, S., Rastogi, R., & Shim, K. (2000). Rock: A robust clustering algorithm for categorical attributes. Information Systems, 25(5), 345–366. https://doi.org/10.1016/S0306-4379(00)00022-3 DOI: https://doi.org/10.1016/S0306-4379(00)00022-3
Hayes-Roth, B., & Hayes-Roth, F. (1977). Hayes-Roth [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5501T
He, Z., Xu, X., & Deng, S. (2002). Squeezer: An efficient algorithm for clustering categorical data. Journal of Computer Science and Technology, 17(5), 611–624. https://doi.org/10.1007/BF02948829 DOI: https://doi.org/10.1007/BF02948829
Hu, L., Jiang, M., Liu, X., & He, Z. (2025). Significance-based decision tree for interpretable categorical data clustering. Information Sciences, 690, Article 121588. https://doi.org/10.1016/j.ins.2024.121588 DOI: https://doi.org/10.1016/j.ins.2024.121588
Huang, Z. (1997). A fast clustering algorithm to cluster very large categorical data sets in data mining. Research Issues on Data Mining and Knowledge Discovery, 1–8. DOI: https://doi.org/10.1023/A:1009783824328
Huang, Z., & Ng, M. K. (1999). A fuzzy k-modes algorithm for clustering categorical data. IEEE Transactions on Fuzzy Systems, 7(4), 446–452. https://doi.org/10.1109/91.784206 DOI: https://doi.org/10.1109/91.784206
Iliopoulos, V. (2015). The Quicksort algorithm and related topics. arXiv. https://doi.org/10.48550/arXiv.1503.02504
Janosi, A., Steinbrunn, W., Pfisterer, M., & Detrano, R. (1989). Heart disease [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C52P4X
Katakam, A. (2025). Mann-Whitney test, and Wilcoxon. Translational Gastroenterology, 135. DOI: https://doi.org/10.1016/B978-0-12-821426-8.00022-4
Khoei, T. T., & Singh, A. (2025). Data reduction in big data: A survey of methods, challenges and future directions. International Journal of Data Science and Analytics, 20(3), 1643–1682. https://doi.org/10.1007/s41060-024-00603-z DOI: https://doi.org/10.1007/s41060-024-00603-z
Kita, F. J., Rao, G. S., & Kirigiti, P. J. (2025). Optimizing clustering of electronic health tabular data: Generative adversarial networks and Dirichlet process mixture models for advance healthcare analytics. IISE Transactions on Healthcare Systems Engineering, 15(3), 224–240. https://doi.org/10.1080/24725579.2025.2510966 DOI: https://doi.org/10.1080/24725579.2025.2510966
Kumar, K. R. N., Lahza, H., Sreenivasa, B. R., Shawly, T., Alsheikhy, A. A., Arunkumar, H., & Nirmala, C. R. (2023). A novel cluster analysis-based crop dataset recommendation method in precision farming. Computer Systems Science and Engineering, 46(3), 3239–3260. https://doi.org/10.32604/csse.2023.036629 DOI: https://doi.org/10.32604/csse.2023.036629
Kutbay, U. (2018). Partitional clustering. In Recent Applications in Data Clustering. IntechOpen. https://doi.org/10.5772/intechopen.75836 DOI: https://doi.org/10.5772/intechopen.75836
Lahitani, A. R., Permanasari, A. E., & Setiawan, N. A. (2016). Cosine similarity to determine similarity measure: Study case in online essay assessment. In Proceedings of 2016 4th International Conference on Cyber and IT Service Management (CITSM 2016) (pp. 1–6). IEEE. https://doi.org/10.1109/CITSM.2016.7577578 DOI: https://doi.org/10.1109/CITSM.2016.7577578
Lange, F. J. D., Wilcke, J. C., Hoffmann, S., Herrmann, M., & Boulesteix, A. L. (2025). On “confirmatory” methodological research in statistics and related fields. Statistics in Medicine, 44(25–27). https://doi.org/10.1002/sim.70303 DOI: https://doi.org/10.1002/sim.70303
Lenses. (1990). [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5K88Z
Leydesdorff, L., & Vaughan, L. (2006). Co-occurrence matrices and their applications in information science: Extending ACA to the web environment. Journal of the American Society for Information Science and Technology, 57(12), 1616–1628. https://doi.org/10.1002/asi.20335 DOI: https://doi.org/10.1002/asi.20335
Li, J., Zhang, C., Zhang, J., Qin, X., & Hu, L. (2022). MiCS-P: Parallel mutual-information computation of big categorical data on Spark. Journal of Parallel and Distributed Computing, 161, 118–129. https://doi.org/10.1016/j.jpdc.2021.12.002 DOI: https://doi.org/10.1016/j.jpdc.2021.12.002
Mar San, O., Huynh, V.-N., & Nakamori, Y. (2004). An alternative extension of the k-means algorithm for clustering categorical data. International Journal of Applied Mathematics and Computer Science, 14(2), 241–247.
Mau, T. N., & Huynh, V. N. (2021). An LSH-based k-representatives clustering method for large categorical data. Neurocomputing, 463, 29–44. https://doi.org/10.1016/j.neucom.2021.08.050 DOI: https://doi.org/10.1016/j.neucom.2021.08.050
Michalski, R. (1980). Soybean (small) [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5DS3P
Mukhopadhyay, A., Maulik, U., & Bandyopadhyay, S. (2009). Multiobjective genetic algorithm-based fuzzy clustering of categorical attributes. IEEE Transactions on Evolutionary Computation, 13(5), 991–1005. https://doi.org/10.1109/TEVC.2009.2012163 DOI: https://doi.org/10.1109/TEVC.2009.2012163
Mushroom. (1981). [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5959T
Narasimhan, M., Balasubramanian, B., Kumar, S. D., & Patil, N. (2018). EGA-FMC: Enhanced genetic algorithm-based fuzzy k-modes clustering for categorical data. International Journal of Bio-Inspired Computation, 11(4), 219–228. https://doi.org/10.1504/IJBIC.2018.092801 DOI: https://doi.org/10.1504/IJBIC.2018.092801
Nguyen, T. H. T., Dinh, D. T., Sriboonchitta, S., & Huynh, V. N. (2023). A method for k-means-like clustering of categorical data. Journal of Ambient Intelligence and Humanized Computing, 14(11), 15011–15021. https://doi.org/10.1007/s12652-019-01445-5 DOI: https://doi.org/10.1007/s12652-019-01445-5
Oliveira, G. S. de, Silva, F. A., & Ferreira, R. V. (2025). Explainable clustering: A solution to interpret and describe clusters. Journal of Information and Data Management, 16(1), 170–180. https://doi.org/10.5753/jidm.2025.4663
Pang, N., Zhang, J., Zhang, C., Qin, X., & Cai, J. (2019). PUMA: Parallel subspace clustering of categorical data using multi-attribute weights. Expert Systems with Applications, 126, 233–245. https://doi.org/10.1016/j.eswa.2019.02.030 DOI: https://doi.org/10.1016/j.eswa.2019.02.030
Ran, X., Xi, Y., Lu, Y., Wang, X., & Lu, Z. (2023). Comprehensive survey on hierarchical clustering algorithms and the recent developments. Artificial Intelligence Review, 56(8), 8219–8264. https://doi.org/10.1007/s10462-022-10366-3 DOI: https://doi.org/10.1007/s10462-022-10366-3
Rao, R. K., & Mandhala, V. N. (2025). Comparative analysis of clustering algorithms for financial fraud detection. International Journal of Safety and Security Engineering, 15(4), 777–785. https://doi.org/10.18280/ijsse.150414 DOI: https://doi.org/10.18280/ijsse.150414
Sankhe, P., Hall, S. F., Sage, M., Rodriguez, M. Y., Chandola, V., & Joseph, K. (2022). Mutual information scoring: Increasing interpretability in categorical clustering tasks with applications to child welfare data. In Lecture Notes in Computer Science (Vol. 13558, pp. 165–175). Springer. https://doi.org/10.1007/978-3-031-17114-7_16 DOI: https://doi.org/10.1007/978-3-031-17114-7_16
Sedighi, M. (2016). Application of word co-occurrence analysis method in mapping of the scientific fields (case study: The field of Informetrics). Library Review, 65(1–2), 52–64. https://doi.org/10.1108/LR-07-2015-0075 DOI: https://doi.org/10.1108/LR-07-2015-0075
Siegler, R. (1976). Balance scale [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5488X
Towell, G. G., Shavlik, J. W., & Noordewier, M. O. (1993). Molecular biology (promoter gene sequences) [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5S01D
Wolberg, W. (1990). Breast cancer Wisconsin (original) [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5HP4Z
Yang, Y., Guan, X., & You, J. (2002). CLOPE: A fast clustering algorithm for transactional data. In Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 682–687). ACM. https://doi.org/10.1145/775047.775149 DOI: https://doi.org/10.1145/775047.775149
Zhang, Z., & Chen, S.-M. (2021). Optimization-based group decision making using interval-valued intuitionistic fuzzy preference relations. Information Sciences, 561, 352–370. https://doi.org/10.1016/j.ins.2020.12.047 DOI: https://doi.org/10.1016/j.ins.2020.12.047
Zwitter, M., & Soklic, M. (1987). Primary tumor [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5WK5Q
Zwitter, M., & Soklic, M. (1988). Lymphography [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C54598
Article Details
License

This work is licensed under a Creative Commons Attribution 4.0 International License.
All articles published in Applied Computer Science are open-access and distributed under the terms of the Creative Commons Attribution 4.0 International License.
