Efficiency of Temporal Convolutional Networks in Karate Kata Evaluation: A Comparative Study

Authors

DOI:

https://doi.org/10.30872/jim.v21i1.28888

Keywords:

Temporal Convolutional Network, Karate Kata, Deep Learning Efficiency, Pose Estimation, MoveNet,

Abstract

The objective quantification of complex martial-arts movement is a long-standing challenge in computer vision. This study reports a feasibility prototype that pairs MoveNet Lightning pose estimation with a lightweight Temporal Convolutional Network for two tasks on a small custom dataset of four foundational Karate Kata (Heian Shodan through Heian Yondan): (1) multi-class kata recognition and (2) binary correctness evaluation. A 132-dimensional kinematic feature vector is constructed per frame from 17 MoveNet keypoints, combining hip-centered normalized coordinates, confidence scores, joint angles, bone-length ratios, and end-effector velocities. The full dataset comprises 100 lateral-view videos (80 train/10 validation/10 test) collected from a single Indonesian dojo. A multi-task model with two causal-dilated 1D-convolution layers (32 filters, dilations 1 and 2) is trained for up to 100 epochs on Apple M1 hardware over five random seeds. The kata head reaches 100% validation accuracy with zero variance across all five seeds; however, this result is interpreted with strong reservations because the validation set contains only 10 samples. The correctness head consistently collapses to majority-class prediction (60.0% ± 0.0%, identical to the all-positive baseline of 6/10). The contribution of this paper is therefore a fully reproducible end-to-end pipeline (MoveNet → 132-D features → multi-task TCN → real-time webcam demo) together with a candid characterization of the dataset-driven limits that block the correctness task at this scale.

Author Biographies

  • Viona Zatil Aqmar Kaleb, President University
    Master of Informatics Study Program, President University
  • Wiranto Herry Utomo, President University
    Master of Informatics Study Program, President University

References

Abdulrahman, N., Zainuddin, Z., & Nurtanio, I. (2026). A deep learning-based framework using CNN+LSTM for karate kata classification and correctness evaluation. Engineering, Technology & Applied Science Research, 16(1), 32619–32624. https://doi.org/10.48084/etasr.12147

Armstrong, K., Rodrigues, A., Willmott, A. P., Zhang, L., & Ye, X. (2025). Validation of human pose estimation and human mesh recovery for extracting clinically relevant motion data from videos. arXiv. https://doi.org/10.48550/arXiv.2503.14760

Bai, S., Kolter, J. Z., & Koltun, V. (2018). An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv. https://doi.org/10.48550/arXiv.1803.01271

Chen, J., Wang, J., Yuan, Q., & Yang, Z. (2023). CNN-LSTM model for recognizing video-recorded actions performed in a traditional Chinese exercise. IEEE Journal of Translational Engineering in Health and Medicine, 11, 351–359. https://doi.org/10.1109/JTEHM.2023.3282245

Cornman, H. L., Stenum, J., & Roemmich, R. T. (2021). Video-based quantification of human movement frequency using pose estimation: A pilot study. PLoS ONE, 16(12), e0261450. https://doi.org/10.1371/journal.pone.0261450

Ding, X., Li, Z., Yu, J., Xie, W., Li, X., & Jiang, T. (2023). A novel lightweight human activity recognition method via L-CTCN. Sensors, 23(24), 9681. https://doi.org/10.3390/s23249681

Echeverria, J., & Santos, O. C. (2021). Toward modeling psychomotor performance in karate combats using computer vision pose estimation. Sensors, 21(24), 8378. https://doi.org/10.3390/s21248378

Ercolano, G., & Rossi, S. (2021). Combining CNN and LSTM for activity of daily living recognition with a 3D matrix skeleton representation. Intelligent Service Robotics, 14(2), 175–185. https://doi.org/10.1007/s11370-021-00358-7

Fukushima, T., Blauberger, P., Guedes Russomanno, T., & Lames, M. (2025). Comparison of different pose estimation models for lower-body kinematics: A validation study. Scientific Journal of Sport and Performance, 5(2), 253–268. https://doi.org/10.55860/XWJL7156

Gupta, A., Gupta, K., Gupta, K., & Gupta, K. (2021). Human activity recognition using pose estimation and machine learning algorithm. Proceedings of the 2nd International Semantic Intelligence Conference (ISIC 2021), 2786, 323–330. https://ceur-ws.org/Vol-2786/Paper40.pdf

He, J.-L., Wang, J.-H., Lo, C.-M., & Jiang, Z. (2025). Human activity recognition via attention-augmented TCN-BiGRU fusion. Sensors, 25(18), 5765. https://doi.org/10.3390/s25185765

Islam, M. M., Nooruddin, S., Karray, F., & Muhammad, G. (2022). Human activity recognition using tools of convolutional neural networks: A state of the art review, data sets, challenges, and future prospects. Computers in Biology and Medicine, 149, 106060. https://doi.org/10.1016/j.compbiomed.2022.106060

Kim, J.-W., Choi, J.-Y., Ha, E.-J., & Choi, J.-H. (2023). Human pose estimation using MediaPipe Pose and optimization method based on a humanoid model. Applied Sciences, 13(4), 2700. https://doi.org/10.3390/app13042700

Ma, N., Zhang, X., Zheng, H.-T., & Sun, J. (2018). ShuffleNet V2: Practical guidelines for efficient CNN architecture design. In V. Ferrari, M. Hebert, C. Sminchisescu, & Y. Weiss (Eds.), Computer Vision – ECCV 2018 (pp. 122–138). Springer. https://doi.org/10.1007/978-3-030-01264-9_8

Malik, N. U. R., Abu-Bakar, S. A. R., Sheikh, U. U., Channa, A., & Popescu, N. (2023). Cascading pose features with CNN-LSTM for multiview human action recognition. Signals, 4(1), 40–55. https://doi.org/10.3390/signals4010002

Mortezapour Shiri, F., Yamaguchi, S., & Ahmadon, M. A. B. (2025). A deep learning model based on bidirectional temporal convolutional network (Bi-TCN) for predicting employee attrition. Applied Sciences, 15(6), 2984. https://doi.org/10.3390/app15062984

Priya, B. G., & Arulselvi, D. M. (2018). Action classification for karate dataset using deep learning. 2018 International Conference on Intelligent Computing and Communication for Smart World (I2C2SW), 1–4. https://doi.org/10.1109/I2C2SW45816.2018.8997335

Saeed, S. M., Akbar, H., Nawaz, T., Elahi, H., & Khan, U. S. (2023). Body-pose-guided action recognition with convolutional long short-term memory (LSTM) in aerial videos. Applied Sciences, 13(16), 9384. https://doi.org/10.3390/app13169384

Sekaran, S. R., Pang, Y. H., Ling, G. F., & Yin, O. S. (2022). MSTCN: A multiscale temporal convolutional network for user independent human activity recognition. F1000Research, 11, 1227. https://doi.org/10.12688/f1000research.126090.2

Shuvo, M. R., Mekala, M. S., & Elyan, E. (2025). Deep learning and attention-based methods for human activity recognition and anticipation: A comprehensive review. Cognitive Computation, 17(6), 158. https://doi.org/10.1007/s12559-025-10513-2

Washabaugh, E. P., Shanmugam, T. A., Ranganathan, R., & Krishnan, C. (2022). Comparing the accuracy of open-source pose estimation methods for measuring gait kinematics. Gait & Posture, 97, 188–195. https://doi.org/10.1016/j.gaitpost.2022.08.008

Yao, S., Ping, Y., Yue, X., & Chen, H. (2025). Graph convolutional networks for multi-modal robotic martial arts leg pose recognition. Frontiers in Neurorobotics, 18, 1520983. https://doi.org/10.3389/fnbot.2024.1520983

Ye, R. Z., Subramanian, A., Diedrich, D., Lindroth, H., Pickering, B., & Herasevich, V. (2022). Effects of image quality on the accuracy human pose estimation and detection of eye lid opening/closing using OpenPose and DLib. Journal of Imaging, 8(12), 330. https://doi.org/10.3390/jimaging8120330

Downloads

Published

2026-06-29