Efficiency of Temporal Convolutional Networks in Karate Kata Evaluation: A Comparative Study
DOI:
https://doi.org/10.30872/jim.v21i1.28888Keywords:
Temporal Convolutional Network, Karate Kata, Deep Learning Efficiency, Pose Estimation, MoveNet,Abstract
The objective quantification of complex martial-arts movement is a long-standing challenge in computer vision. This study reports a feasibility prototype that pairs MoveNet Lightning pose estimation with a lightweight Temporal Convolutional Network for two tasks on a small custom dataset of four foundational Karate Kata (Heian Shodan through Heian Yondan): (1) multi-class kata recognition and (2) binary correctness evaluation. A 132-dimensional kinematic feature vector is constructed per frame from 17 MoveNet keypoints, combining hip-centered normalized coordinates, confidence scores, joint angles, bone-length ratios, and end-effector velocities. The full dataset comprises 100 lateral-view videos (80 train/10 validation/10 test) collected from a single Indonesian dojo. A multi-task model with two causal-dilated 1D-convolution layers (32 filters, dilations 1 and 2) is trained for up to 100 epochs on Apple M1 hardware over five random seeds. The kata head reaches 100% validation accuracy with zero variance across all five seeds; however, this result is interpreted with strong reservations because the validation set contains only 10 samples. The correctness head consistently collapses to majority-class prediction (60.0% ± 0.0%, identical to the all-positive baseline of 6/10). The contribution of this paper is therefore a fully reproducible end-to-end pipeline (MoveNet → 132-D features → multi-task TCN → real-time webcam demo) together with a candid characterization of the dataset-driven limits that block the correctness task at this scale.References
Abdulrahman, N., Zainuddin, Z., & Nurtanio, I. (2026). A deep learning-based framework using CNN+LSTM for karate kata classification and correctness evaluation. Engineering, Technology & Applied Science Research, 16(1), 32619–32624. https://doi.org/10.48084/etasr.12147
Armstrong, K., Rodrigues, A., Willmott, A. P., Zhang, L., & Ye, X. (2025). Validation of human pose estimation and human mesh recovery for extracting clinically relevant motion data from videos. arXiv. https://doi.org/10.48550/arXiv.2503.14760
Bai, S., Kolter, J. Z., & Koltun, V. (2018). An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv. https://doi.org/10.48550/arXiv.1803.01271
Chen, J., Wang, J., Yuan, Q., & Yang, Z. (2023). CNN-LSTM model for recognizing video-recorded actions performed in a traditional Chinese exercise. IEEE Journal of Translational Engineering in Health and Medicine, 11, 351–359. https://doi.org/10.1109/JTEHM.2023.3282245
Cornman, H. L., Stenum, J., & Roemmich, R. T. (2021). Video-based quantification of human movement frequency using pose estimation: A pilot study. PLoS ONE, 16(12), e0261450. https://doi.org/10.1371/journal.pone.0261450
Ding, X., Li, Z., Yu, J., Xie, W., Li, X., & Jiang, T. (2023). A novel lightweight human activity recognition method via L-CTCN. Sensors, 23(24), 9681. https://doi.org/10.3390/s23249681
Echeverria, J., & Santos, O. C. (2021). Toward modeling psychomotor performance in karate combats using computer vision pose estimation. Sensors, 21(24), 8378. https://doi.org/10.3390/s21248378
Ercolano, G., & Rossi, S. (2021). Combining CNN and LSTM for activity of daily living recognition with a 3D matrix skeleton representation. Intelligent Service Robotics, 14(2), 175–185. https://doi.org/10.1007/s11370-021-00358-7
Fukushima, T., Blauberger, P., Guedes Russomanno, T., & Lames, M. (2025). Comparison of different pose estimation models for lower-body kinematics: A validation study. Scientific Journal of Sport and Performance, 5(2), 253–268. https://doi.org/10.55860/XWJL7156
Gupta, A., Gupta, K., Gupta, K., & Gupta, K. (2021). Human activity recognition using pose estimation and machine learning algorithm. Proceedings of the 2nd International Semantic Intelligence Conference (ISIC 2021), 2786, 323–330. https://ceur-ws.org/Vol-2786/Paper40.pdf
He, J.-L., Wang, J.-H., Lo, C.-M., & Jiang, Z. (2025). Human activity recognition via attention-augmented TCN-BiGRU fusion. Sensors, 25(18), 5765. https://doi.org/10.3390/s25185765
Islam, M. M., Nooruddin, S., Karray, F., & Muhammad, G. (2022). Human activity recognition using tools of convolutional neural networks: A state of the art review, data sets, challenges, and future prospects. Computers in Biology and Medicine, 149, 106060. https://doi.org/10.1016/j.compbiomed.2022.106060
Kim, J.-W., Choi, J.-Y., Ha, E.-J., & Choi, J.-H. (2023). Human pose estimation using MediaPipe Pose and optimization method based on a humanoid model. Applied Sciences, 13(4), 2700. https://doi.org/10.3390/app13042700
Ma, N., Zhang, X., Zheng, H.-T., & Sun, J. (2018). ShuffleNet V2: Practical guidelines for efficient CNN architecture design. In V. Ferrari, M. Hebert, C. Sminchisescu, & Y. Weiss (Eds.), Computer Vision – ECCV 2018 (pp. 122–138). Springer. https://doi.org/10.1007/978-3-030-01264-9_8
Malik, N. U. R., Abu-Bakar, S. A. R., Sheikh, U. U., Channa, A., & Popescu, N. (2023). Cascading pose features with CNN-LSTM for multiview human action recognition. Signals, 4(1), 40–55. https://doi.org/10.3390/signals4010002
Mortezapour Shiri, F., Yamaguchi, S., & Ahmadon, M. A. B. (2025). A deep learning model based on bidirectional temporal convolutional network (Bi-TCN) for predicting employee attrition. Applied Sciences, 15(6), 2984. https://doi.org/10.3390/app15062984
Priya, B. G., & Arulselvi, D. M. (2018). Action classification for karate dataset using deep learning. 2018 International Conference on Intelligent Computing and Communication for Smart World (I2C2SW), 1–4. https://doi.org/10.1109/I2C2SW45816.2018.8997335
Saeed, S. M., Akbar, H., Nawaz, T., Elahi, H., & Khan, U. S. (2023). Body-pose-guided action recognition with convolutional long short-term memory (LSTM) in aerial videos. Applied Sciences, 13(16), 9384. https://doi.org/10.3390/app13169384
Sekaran, S. R., Pang, Y. H., Ling, G. F., & Yin, O. S. (2022). MSTCN: A multiscale temporal convolutional network for user independent human activity recognition. F1000Research, 11, 1227. https://doi.org/10.12688/f1000research.126090.2
Shuvo, M. R., Mekala, M. S., & Elyan, E. (2025). Deep learning and attention-based methods for human activity recognition and anticipation: A comprehensive review. Cognitive Computation, 17(6), 158. https://doi.org/10.1007/s12559-025-10513-2
Washabaugh, E. P., Shanmugam, T. A., Ranganathan, R., & Krishnan, C. (2022). Comparing the accuracy of open-source pose estimation methods for measuring gait kinematics. Gait & Posture, 97, 188–195. https://doi.org/10.1016/j.gaitpost.2022.08.008
Yao, S., Ping, Y., Yue, X., & Chen, H. (2025). Graph convolutional networks for multi-modal robotic martial arts leg pose recognition. Frontiers in Neurorobotics, 18, 1520983. https://doi.org/10.3389/fnbot.2024.1520983
Ye, R. Z., Subramanian, A., Diedrich, D., Lindroth, H., Pickering, B., & Herasevich, V. (2022). Effects of image quality on the accuracy human pose estimation and detection of eye lid opening/closing using OpenPose and DLib. Journal of Imaging, 8(12), 330. https://doi.org/10.3390/jimaging8120330
Downloads
Published
Issue
Section
License
Copyright Transfer StatementThe copyright of this article is transferred to Informatika Mulawarman : Jurnal Ilmiah Ilmu Komputer and when the article is accepted for publication. the authors transfer all and all rights into and to paper including but not limited to all copyrights in the Informatika Mulawarman. The author represents and warrants that the original is the original and that he/she is the author of this paper unless the material is clearly identified as the original source, with notification of the permission of the copyright owner if necessary. The author states that he has the authority and authority to make and carry out this task.
The author states that:
- This paper has not been published in the same form elsewhere.
- This will not be submitted elsewhere for publication prior to acceptance/rejection by this Journal.
A Copyright permission is obtained for material published elsewhere and who require permission for this reproduction. Furthermore, I / We hereby transfer the unlimited publication rights of the above paper to Informatika Mulawarman : Jurnal Ilmiah Ilmu Komputer. Copyright transfer includes exclusive rights to reproduce and distribute articles, including reprints, translations, photographic reproductions, microforms, electronic forms (offline, online), or other similar reproductions.
The author's mark is appropriate for and accepts responsibility for releasing this material on behalf of any and all coauthor. This Agreement shall be signed by at least one author who has obtained the consent of the co-author (s) if applicable. After the submission of this agreement is signed by the author concerned, the amendment of the author or in the order of the author listed shall not be accepted.
Rights / Terms and Conditions Saved
- The author keeps all proprietary rights in every process, procedure, or article creation described in Work.
- The author may reproduce or permit others to reproduce the work or derivative works for the author's personal use or for the use of the company, provided that the source and the Informatika Mulawarman copyright notice are indicated, the copy is not used in any way implying the Journal of Informatika Mulawarman (JIM) approval of the product or service from any company, and the copy itself is not offered for sale.
- Although authors are permitted to reuse all or part of the Works in other works, this does not include granting third-party requests to reprint, republish, or other types of reuse.

Informatika Mulawarman by http://e-journals.unmul.ac.id/index.php/JIM/index is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
under the CC BY-SA license, authors and other users are able to reprint, distribute or use the material for commercial purposes so long as they give attribution to the journal Informatika Mulawarman and license the republished material under the same license.