Enforcing label consistency and lowering labelling time in infants’ pose data via semi-automatic annotation
Article Sidebar
Issue Vol. 22 No. 3 (2026)
-
Enforcing label consistency and lowering labelling time in infants’ pose data via semi-automatic annotation
Greta DI MARINO, Emanuele CARDINALE, Alessio CORREANI, Lucia MIGLIORELLI, Sara MOCCIA1-14
-
A data-driven framework for AI adoption efficiency assessment using hybrid DEA and machine learning methods
Ewa CHODAKOWSKA15-29
-
ARECA-Lite: A lightweight modified ArecaNet with reduced complexity for real-time robust facial emotion recognition
Mustapha Abdelkader LAOUMIR, Amina KINANE DAOUADJI, Fatima BENDELLA30-51
-
Designing vehicle structures using materials with a low carbon footprint
Bartosz ŁATA, Jacek CZARNIGOWSKI, Wiktor ISKRA, Miłosz CHWIEJCZAK, Iga KOPEĆ52-61
-
Parallelogram-mode: A novel clustering method for categorical data
Ashuza KUDERHA, Olamma IHEANETU62-81
-
Digitalisation of relay protection and implementation of main protections in digital form
Dmytro DANYLCHENKO, Vladyslav TSIUPA, Oleksandr MIROSHNYK, Taras SHCHUR, Katarzyna PIOTROWSKA82-96
-
Modified snake optimizer algorithm for solving the permutation flow shop scheduling problem
Hassan ALMAZINI, Salah MORTADA, Hussein Fouad ALMAZINI97-107
-
Computer-based data processing approaches to production scrap management
Łukasz WÓJCIK, Arkadiusz GOLA, Jakub PIZOŃ108-120
-
A hybrid parameter-tuning for adaptive variable-length particle swarm optimisation in cancer feature selection
Shir Li WANG, Siti RAMADHANI, Muhammad FIKRY, Haldi BUDIMAN, Theam Foo NG, Sumayyah DZULKIFLY, Roziana ARIFFIN121-147
-
Beyond classical optimisation: Toward feasibility-aware computational architectures for synchronised systems
Grzegorz BOCEWICZ, Czesław SMUTNICKI, Zbigniew BANASZAK148-167
-
A composite latency model for evaluating hybrid OLTP/OLAP information systems
Volodymyr SOLOHUB, Volodymyr PASHKEVYCH168-180
-
Automatic methods for 3D motion trajectories gap filling: Custom-based Kalman vs. BiLSTM
Kamil ŻELAZOWSKI, Wojciech WOJCIECHEWICZ, Maria SKUBLEWSKA-PASZKOWSKA, Paweł POWROŹNIK181-195
-
Implementation of an IEC 61215-oriented photovoltaic module test emulator with integrated predictive maintenance capabilities
Aristide TOLOK NELEM, Yannick Antoine ABANDA, Steyve Samson NYATTE, Mathieu Jean Pierre PESDJOCK, Achille MELINGUI, Pierre ELE196-218
-
Modelling the predictive reliability of rotating machines using Artificial Intelligence.
Fernand Joseph TOUKAP NONO, Tokoue Ngatcha DIANORRÉ, Offole FLORENC, Mouzong Pemi MARCELIN219-243
-
Anomaly detection in vibroarthrographic signals using handcrafted signal features and one-class methods
Robert KARPIŃSKI, Arkadiusz SYTA244–261
Archives
-
Vol. 22 No. 3
2026-09-30 15
-
Vol. 22 No. 2
2026-06-30 15
-
Vol. 22 No. 1
2026-03-31 15
-
Vol. 21 No. 4
2025-12-31 12
-
Vol. 21 No. 3
2025-09-30 12
-
Vol. 21 No. 2
2025-06-30 12
-
Vol. 21 No. 1
2025-03-31 12
-
Vol. 20 No. 4
2024-12-31 12
-
Vol. 20 No. 3
2024-09-30 12
-
Vol. 20 No. 2
2024-06-30 12
-
Vol. 20 No. 1
2024-03-30 12
-
Vol. 19 No. 4
2023-12-31 10
-
Vol. 19 No. 3
2023-09-30 10
-
Vol. 19 No. 2
2023-06-30 10
-
Vol. 19 No. 1
2023-03-31 10
-
Vol. 18 No. 4
2022-12-30 8
-
Vol. 18 No. 3
2022-09-30 8
-
Vol. 18 No. 2
2022-06-30 8
-
Vol. 18 No. 1
2022-03-31 8
Main Article Content
Authors
emanuele.cardinale@phd.unich.it
Abstract
Infants’ spontaneous movements provide clinically relevant information about neurodevelopment, but their assessment still relies on qualitative visual inspection, which is prone to variability. Video-based systems with automatic pose estimation algorithms have been proposed to address this, yet annotating data to train these algorithms is time-consuming and prone to intra- and inter-annotator variability. To improve the efficiency and consistency of annotation, semi-automatic annotation has been proposed in the broader human pose estimation literature. To assess the extent to which semi-automatic annotation may also be beneficial for infants’ pose estimation, we benchmarked six state-of-the-art adult-trained models on a new dataset comprising 46 videos of preterm infants recorded in a neonatal unit. The predicted joints from the best-performing model were presented to human annotators for optional refinement, who were also asked to label the same joints manually for comparison. ViTPose and Sapiens achieved the highest joint localisation accuracy, with ViTPose requiring markedly lower computational cost. Semi-automatic annotation reduced inter-annotator variability by 7.82% and decreased the time to label a frame by 35.9% compared to full manual labelling. These findings support the use of semi-automatic annotation as an effective strategy to enhance the labelling process of infant pose estimation data from real clinical settings.
Keywords:
Sustainable Development Goal (SDG)
- Good health and well-being
References
Andriluka, M., Iqbal, U., Insafutdinov, E., Pishchulin, L., Milan, A., Gall, J., & Schiele, B. (2018). PoseTrack: A benchmark for human pose estimation and tracking. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 5167–5176). https://doi.org/10.1109/CVPR.2018.00542 DOI: https://doi.org/10.1109/CVPR.2018.00542
Bishop, C. E., Machut, K. Z., Dammann, C. E., Cuevas Guaman, M., Miller, E. R., & Lakshminrusimha, S. (2023). Academic neonatologist—a species at the brink of extinction? Journal of Perinatology, 43(12), 1526–1529. https://doi.org/10.1038/s41372-023-01803-4 DOI: https://doi.org/10.1038/s41372-023-01803-4
Cacciatore, A., Berardini, D., Scaraggi, V., Mancini, A., Moccia, S., & Migliorelli, L. (2025). Online knowledge distillation and deep supervision in HRNet: Green AI for preterm infants’ pose estimation. ACM Transactions on Computing for Healthcare, 6(4), 1–19. https://doi.org/10.1145/3757067 DOI: https://doi.org/10.1145/3757067
Doi, H., Furui, A., Ueda, R., Shimatani, K., Yamamoto, M., Sakurai, K., Mori, C., & Tsuji, T. (2023). Spatiotemporal patterns of spontaneous movement in neonates are significantly linked to risk of autism spectrum disorders at 18 months old. Scientific Reports, 13(1), Article 13869. https://doi.org/10.1038/s41598-023-40368-2 DOI: https://doi.org/10.1038/s41598-023-40368-2
Einspieler, C., Sigafoos, J., Bartl-Pokorny, K. D., Landa, R., Marschik, P. B., & Bölte, S. (2014). Highlighting the first 5 months of life: General movements in infants later diagnosed with autism spectrum disorder or Rett syndrome. Research in Autism Spectrum Disorders, 8(3), 286–291. https://doi.org/10.1016/j.rasd.2013.12.013 DOI: https://doi.org/10.1016/j.rasd.2013.12.013
Fang, H.-S., Li, J., Tang, H., Xu, C., Zhu, H., Xiu, Y., Li, Y.-L., & Lu, C. (2022). Alphapose: Whole-body regional multi-person pose estimation and tracking in real-time. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(6), 7157–7173. https://doi.org/10.1109/TPAMI.2022.3222784 DOI: https://doi.org/10.1109/TPAMI.2022.3222784
Feng, R., Richter, F., Mari, E., Gleason, A., Le, C., Kellner, C., Shrivastava, R., Fields, M., Rapoport, B., Bederson, J., Schadt, E. E., Glicksberg, B. S., Richter, F., & Dangayach, N. (2025). Pose AI prediction of neurological status in the Neuroscience Intensive Care Unit. MedRxiv. DOI: https://doi.org/10.1101/2025.03.12.25323510
Gama, F., Mísař, M., Navara, L., Popescu, S. T., & Hoffmann, M. (2025). Automatic infant 2D pose estimation from videos: Comparing seven deep neural network methods. Behavior Research Methods, 57(10), Article 280. https://doi.org/10.3758/s13428-025-02816-x DOI: https://doi.org/10.3758/s13428-025-02816-x
Grafton, A., Warnecke, J. M., Li, M., He, E., Thomson, L., Beardsall, K., & Lasenby, J. (2025). Neonatal pose estimation in the unaltered clinical environment with fusion of RGB, depth and IR images. npj Digital Medicine, 8(1), Article 539. https://doi.org/10.1038/s41746-025-01929-z DOI: https://doi.org/10.1038/s41746-025-01929-z
He, S., Wei, M., Meng, D., Lv, Z., Guo, H., Yang, G., & Wang, Z. (2024). Adversarially trained RTMpose: A high-performance, non-contact method for detecting Genu valgum in adolescents. Computers in Biology and Medicine, 183, Article 109214. https://doi.org/10.1016/j.compbiomed.2024.109214 DOI: https://doi.org/10.1016/j.compbiomed.2024.109214
Huang, X., Fu, N., Liu, S., & Ostadabbas, S. (2021). Invariant representation learning for infant pose estimation with small data. In IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021) (pp. 1–8). https://doi.org/10.1109/FG52635.2021.9667056 DOI: https://doi.org/10.1109/FG52635.2021.9666956
HumanSignal. (n.d.). Label Studio: Open source data labeling and AI evaluation. https://labelstud.io/
Jahn, L., Flügge, S., Zhang, D., Poustka, L., Bölte, S., Wörgötter, F., Marschik, P. B., & Kulvicius, T. (2025). Comparison of marker-less 2D image-based methods for infant pose estimation. Scientific Reports, 15(1), Article 12148. https://doi.org/10.1038/s41598-025-96206-0 DOI: https://doi.org/10.1038/s41598-025-96206-0
Jiang, T., Lu, P., Zhang, L., Ma, N., Han, R., Lyu, C., Li, Y., & Chen, K. (2023). RTMPose: Real-time multi-person pose estimation based on MMPose. ArXiv, abs/2303.07399. https://doi.org/10.48550/arXiv.2303.07399
Kali, G. T., Du Preez, J. C., Van Zyl, J. I., Burger, M., Katsabola, H., & Pepper, M. S. (2025). Predicting cerebral palsy and 18 month neurodevelopmental outcome in infants with presumed Hypoxic Ischaemic Encephalopathy: Role of General Movements Assessment and early neurological examination. Frontiers in Pediatrics, 13, Article 1638584. https://doi.org/10.3389/fped.2025.1638584 DOI: https://doi.org/10.3389/fped.2025.1638584
Khirodkar, R., Bagautdinov, T., Martinez, J., Zhaoen, S., James, A., Selednik, P., Anderson, S., & Saito, S. (2024). Sapiens: Foundation for human vision models. ArXiv, abs/2408.12569. https://doi.org/10.48550/arXiv.2408.12569 DOI: https://doi.org/10.1007/978-3-031-73235-5_12
Kyrollos, D. G., Fuller, A., Greenwood, K., Harrold, J., & Green, J. R. (2023). Under the cover infant pose estimation using multimodal data. IEEE Transactions on Instrumentation and Measurement, 72, 1–12. https://doi.org/10.1109/TIM.2023.3244220 DOI: https://doi.org/10.1109/TIM.2023.3244220
Letzkus, L., Pulido, J. V., Adeyemo, A., Baek, S., & Zanelli, S. (2024). Machine learning approaches to evaluate infants’ general movements in the writhing stage—a pilot study. Scientific Reports, 14(1), Article 4522. https://doi.org/10.1038/s41598-024-54297-1 DOI: https://doi.org/10.1038/s41598-024-54297-1
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., & Zitnick, C. L. (2014). Microsoft COCO: Common objects in context. In European Conference on Computer Vision (pp. 740–755). Springer. https://doi.org/10.1007/978-3-319-10602-1_48 DOI: https://doi.org/10.1007/978-3-319-10602-1_48
Liu, L., Sun, Y., Li, Y., & Liu, Y. (2025). A hybrid human fall detection method based on modified YOLOv8s and AlphaPose. Scientific Reports, 15(1), Article 2636. https://doi.org/10.1038/s41598-025-86429-6 DOI: https://doi.org/10.1038/s41598-025-86429-6
Maji, D., Nagori, S., Mathew, M., & Poddar, D. (2022). YOLO-Pose: Enhancing YOLO for multi person pose estimation using object keypoint similarity loss. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 2636-2645). https://doi.org/10.1109/CVPRW56347.2022.00297 DOI: https://doi.org/10.1109/CVPRW56347.2022.00297
McCay, K. D., Ho, E. S., Shum, H. P., Fehringer, G., Marcroft, C., & Embleton, N. D. (2020). Abnormal infant movements classification with deep learning on pose-based features. IEEE Access, 8, 51582–51592. https://doi.org/10.1109/ACCESS.2020.2980269 DOI: https://doi.org/10.1109/ACCESS.2020.2980269
Migliorelli, L., Cacciatore, A., Ottaviani, V., Berardini, D., Dellacà, R. L., Frontoni, E., & Moccia, S. (2023a). TwinEDA: A sustainable deep-learning approach for limb-position estimation in preterm infants’ depth images. Medical & Biological Engineering & Computing, 61(2), 387–397. https://doi.org/10.1007/s11517-022-02724-z DOI: https://doi.org/10.1007/s11517-022-02696-9
Migliorelli, L., Tiribelli, S., Cacciatore, A., Giovanola, B., Frontoni, E., & Moccia, S. (2023b). Accountable deep-learning-based vision systems for preterm infant monitoring. Computer, 56(8), 84–93. https://doi.org/10.1109/MC.2023.3235987 DOI: https://doi.org/10.1109/MC.2023.3235987
Moccia, S., Migliorelli, L., Carnielli, V., & Frontoni, E. (2019). Preterm infants’ pose estimation with spatio-temporal features. IEEE Transactions on Biomedical Engineering, 67(8), 2370–2380. https://doi.org/10.1109/tbme.2019.2961448 DOI: https://doi.org/10.1109/TBME.2019.2961448
Newell, A., Yang, K., & Deng, J. (2016). Stacked hourglass networks for human pose estimation. In European Conference on Computer Vision (pp. 483–499). Springer. https://doi.org/10.1007/978-3-319-46484-8_29 DOI: https://doi.org/10.1007/978-3-319-46484-8_29
OpenMMLab. (n.d.). Welcome to MMPose’s documentation (wersja 1.3.2). https://mmpose.readthedocs.io/en/latest/mmpose.readthedocs
Sagonas, C., Tzimiropoulos, G., Zafeiriou, S., & Pantic, M. (2013). A semi-automatic methodology for facial landmark annotation. In IEEE Conference on Computer Vision and Pattern Recognition (pp. 896–903). https://doi.org/10.1109/CVPRW.2013.132 DOI: https://doi.org/10.1109/CVPRW.2013.132
Sato, K., Nagashima, Y., Mano, T., Iwata, A., & Toda, T. (2019). Quantifying normal and parkinsonian gait features from home movies: Practical application of a deep learning-based 2D pose estimator. PLoS ONE, 14(11), Article e0223549. https://doi.org/10.1371/journal.pone.0223549 DOI: https://doi.org/10.1371/journal.pone.0223549
Shamsipour, G., Fekri-Ershad, S., Sharifi, M., & Alaei, A. (2024). Improve the efficiency of handcrafted features in image retrieval by adding selected feature generating layers of deep convolutional neural networks. Signal, Image and Video Processing, 18(3), 2607–2620. https://doi.org/10.1007/s11760-023-02934-z DOI: https://doi.org/10.1007/s11760-023-02934-z
Shin, H. I., Shin, H.-I., Bang, M. S., Kim, D.-K., Shin, S. H., Kim, E.-K., Kim, Y.-J., Lee, E. S., Park, S. G., Ji, H. M., & Lee, W. H. (2022). Deep learning-based quantitative analyses of spontaneous movements and their association with early neurological development in preterm infants. Scientific Reports, 12(1), Article 3138. https://doi.org/10.1038/s41598-022-07139-x DOI: https://doi.org/10.1038/s41598-022-07139-x
Sun, K., Xiao, B., Liu, D., & Wang, J. (2019). Deep high-resolution representation learning for human pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 5693–5703). https://doi.org/10.1109/CVPR.2019.00584 DOI: https://doi.org/10.1109/CVPR.2019.00584
Turner, A., & Sharkey, D. (2024). Enhanced infant movement analysis using transformer-based fusion of diverse video features for neurodevelopmental monitoring. Sensors, 24(20), Article 6619. https://doi.org/10.3390/s24206619 DOI: https://doi.org/10.3390/s24206619
Ultralytics. (2024). Ultralytics YOLO11. https://docs.ultralytics.com/models/yolo11/
Xu, Y., Zhang, J., Zhang, Q., & Tao, D. (2022). ViTPose: Simple vision transformer baselines for human pose estimation. Advances in Neural Information Processing Systems, 35, 38571–38584. https://doi.org/10.52202/068431-2795 DOI: https://doi.org/10.52202/068431-2795
Article Details
License

This work is licensed under a Creative Commons Attribution 4.0 International License.
All articles published in Applied Computer Science are open-access and distributed under the terms of the Creative Commons Attribution 4.0 International License.
