COMPARATIVE EVALUATION OF AI-GENERATED AND COACH-DESIGNED TRAINING PROGRAMS IN YOUTH KICKBOXING
DOI:
https://doi.org/10.67034/2786-8354.2026.20.2.21Keywords:
artificial intelligence, kickboxing, combat sports, coaching expertise, training programme, expert evaluationAbstract
Introduction. The aim of this study was to conduct an expert evaluation and comparative analysis of kickboxing training programmes developed by coaches of different qualification levels and by an artificial intelligence language model (ChatGPT).
Materials and Methods. A blinded repeated-measures design was applied. Sixteen independent experts in striking combat sports evaluated four 8-week training programmes for beginner kickboxers aged 10–11 years: three coach-designed programmes (TP-L1 – expert, TP-L2 – intermediate, TP-L3 – novice) and one AI-generated programme (TP-AI). Programme quality was assessed using a 12-item questionnaire covering three domains: structure, combat sports specificity, and practicality. All items were rated on a 5-point Likert scale. Data were analysed using non-parametric statistics (Friedman test, Wilcoxon signed-rank test with Bonferroni correction, Kendall’s W).
Results. Significant differences were found between programmes (χ²(3) = 32.30, p < 0.01) with a strong level of inter-expert agreement (W = 0.67). The highest overall ratings were observed for TP-L1 (4.08 ± 0.32), followed by TP-L2 (3.50 ± 0.42) and TP-AI (3.32 ± 0.48), while TP-L3 received the lowest scores (3.01 ± 0.48). Pairwise comparisons revealed significant differences between TP-L1 and all other programmes (p < 0.05–0.01), as well as between TP-L2 and TP-L3 (p < 0.05). No significant differences were found between TP-L2 and TP-AI (p = 0.18) or TP-L3 and TP-AI (p = 0.63). Effect sizes were large for all significant comparisons (r = 0.67–0.99). Domain-specific analysis showed that TP-L1 achieved the highest scores across all domains, while TP-AI demonstrated relatively high structural quality (3.90 ± 0.57) but lower combat sports specificity (2.65 ± 0.64). Item-level analysis confirmed that the AI-generated programme performed well in general structure (Q1–Q4) but showed weaker results in sport-specific components (Q6–Q8).
Conclusions. AI-generated training programmes demonstrate acceptable methodological quality and may reach a level comparable to those developed by less experienced coaches. However, expert-designed programmes remain superior, particularly in sport-specific integration. Artificial intelligence can be considered a supportive tool in training design, but professional coaching expertise remains essential in combat sports.
References
1. Ambroży, T., Bąk, R., Niewczas, M., & Rydzik, Ł. (2022). Physical and physiological characteristics of kickboxers: a systematic review. Science of Martial Arts, 18, 111-120.
2. Ayers, J. W., Poliak, A., Dredze, M., Leas, E. C., Zhu, Z., Kelley, J. B., ... & Smith, D. M. (2023). Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Internal Medicine, 183(6), 589–596. https://doi.org/10.1001/jamainternmed.2023.1838
3. Bommasani, R., Klyman, K., Longpre, S., Kapoor, S., Maslej, N., Xiong, B., ... & Liang, P. (2023). The foundation model transparency index. arXiv preprint arXiv:2310.12941. https://doi.org/10.48550/arXiv.2310.12941
4. Bompa, T. O., & Buzzichelli, C. (2019). Periodization: Theory and methodology of training (6th ed.). Human Kinetics.
5. Bujak, Z., Gierczuk, D., & Litwiniuk, S. (2013). Professional activities of a coach of martial arts and combat sports. Journal of Combat Sports and Martial Arts, 4(2), 191-195. https://doi.org/10.2478/pjst-2013-0004
6. Cascella, M., Montomoli, J., Bellini, V., & Bignami, E. (2023). Evaluating the feasibility of ChatGPT in healthcare: an analysis of multiple clinical and research scenarios. Journal of medical systems, 47(1), 33. https://doi.org/10.1007/s10916-023-01925-4
7. Cavazzotto, T. G., Dantas, D. B., & Queiroga, M. R. (2024). ChatGPT and exercise prescription: Human vs. machine or human plus machine?. Journal of Sport and Health Science, 13(5), 661-662. https://doi.org/10.1016/j.jshs.2023.10.008
8. Côté, J., & Gilbert, W. (2009). An integrative definition of coaching effectiveness and expertise. International Journal of Sports Science & Coaching, 4(3), 307–323. https://doi.org/10.1260/174795409789623892
9. Curby, D., Dokmanac, M., Kerimov, F., Tropin, Y., Latyshev, M., Bezkorovainyi, D., & Korobeynikov, G. (2023). Performance of wrestlers at the Olympic Games: Gender aspect. Pedagogy of Physical Culture and Sports, 27(6), 487–493. https://doi.org/10.15561/26649837.2023.0607
10. Dergaa, I., Saad, H. B., El Omri, A., Glenn, J., Clark, C., Washif, J., ... & Chamari, K. (2024). Using artificial intelligence for exercise prescription in personalised health promotion: A critical evaluation of OpenAI’s GPT-4 model. Biology of Sport, 41(2), 221-241. https://doi.org/10.5114/biolsport.2024.133661
11. Düking, P., Sperlich, B., Voigt, L., Van Hooren, B., Zanini, M., & Zinner, C. (2024). ChatGPT generated training plans for runners are not rated optimal by coaching experts, but increase in quality with additional input information. Journal of Sports Science & Medicine, 23(1), 56–64. https://doi.org/10.52082/jssm.2024.56
12. Erickson, K., Côté, J., & Fraser-Thomas, J. (2007). Sport experiences, milestones, and educational activities associated with high-performance coaches’ development. The Sport Psychologist, 21(3), 302–316.
13. Genç, A., Aydın, G. R., Kasap, M., Özkan, A., Rakhymzhanov, A., Gürcan, H. H., ... & Demirhan, B. (2026). Comparative evaluation of ChatGPT versions in training program design: scientific approach, accuracy, and practical applicability. BMC Sports Science, Medicine and Rehabilitation, 18(1), 19. https://doi.org/10.1186/s13102-025-01409-7
14. Guatam, D., Saini, P., Haq, A. U., Yadav, S. K., Singh, A., Kumar, R., ... & Reddy, T. O. (2026). Artificial intelligence in sport sciences: A systematic review of models and methods. Sports Orthopaedics and Traumatology. https://doi.org/10.1016/j.orthtr.2025.08.001
15. Havers, T., Masur, L., Isenmann, E., Geisler, S., Zinner, C., Sperlich, B., & Düking, P. (2025). Reproducibility and quality of hypertrophy-related training plans generated by GPT-4 and Google Gemini as evaluated by coaching experts. Biology of Sport, 42(2), 289–329. https://doi.org/10.5114/biolsport.2025.145911
16. Kuliś, S., Zając, A., & Gołaś, A. (2024). The role of artificial intelligence in sports analytics: A systematic review and meta-analysis of performance trends. Applied Sciences, 15(13), 7254.
17. LaPlaca, D. A., & Schempp, P. G. (2020). The characteristics differentiating expert and competent strength and conditioning coaches. Research Quarterly for Exercise and Sport, 91(3), 488-499. https://doi.org/10.1080/02701367.2019.1686451
18. Latyshev, M., Lopatenko, G., Shandryhos, V., Yarmoliuk, O., Pryimak, M., & Kvasnytsia, I. (2024). Computer vision technologies for human pose estimation in exercise: Accuracy and practicality. In Society. Integration. Education. Proceedings of the International Scientific Conference. Vol. 2, 626–636. https://doi.org/10.17770/sie2024vol2.7842
19. Li, L., Olson, H. O., Tereschenko, I., Wang, A., & McCleery, J. (2025). Impact of coach education on coaching effectiveness in youth sport: A systematic review and meta-analysis. International Journal of Sports Science & Coaching, 20(1), 340-356. https://doi.org/10.1177/17479541241283442
20. Lukac, P. J., Turner, W., Vangala, S., Chin, A. T., Khalili, J., Shih, Y. C. T., ... & Mafi, J. N. (2025). A randomized-clinical trial of two ambient artificial intelligence scribes: Measuring documentation efficiency and physician burnout. medRxiv. Preprint. https://doi.org/10.1101/2025.07.10.25331333
21. Nash, C., Ashford, M., & Collins, L. (2023). Expertise in coach development: The need for clarity. Behavioral Sciences, 13(11), 924. https://doi.org/10.3390/bs13110924
22. Ouergui, I., Hssin, N., Haddad, M., Padulo, J., Franchini, E., Gmada, N., & Bouhlel, E. (2014). The effects of five weeks of kickboxing training on physical fitness. Muscles, ligaments and tendons journal, 4(2), 106.
23. Philuek, P., Kusump, S., Sathianpoonsook, T., Jansupom, C., Sawanyawisuth, P., Sawanyawisuth, K., & Chainarong, A. (2025). The effects of chat GPT generated exercise program in healthy overweight young adults: A pilot study. Journal of Human Sport and Exercise, 20(1), 169-179. https://doi.org/10.55860/1epqgp77
24. Puce, L., Bragazzi, N. L., Curra, A., & Trompetto, C. (2025). Harnessing Generative Artificial Intelligence for Exercise and Training Prescription: Applications and Implications in Sports and Physical Activity – A Systematic Literature Review. Applied Sciences (2076-3417), 15(7). https://doi.org/10.3390/app15073497
25. Shin, D., Hsieh, G., & Kim, Y. H. (2025, July). PlanFitting: Personalized Exercise Planning with Large Language Model-driven Conversational Agent. In Proceedings of the 7th ACM Conference on Conversational User Interfaces. 1-19. https://doi.org/10.1145/3719160.3736607
26. Sperlich, B., Düking, P., Leppich, R., & Holmberg, H. C. (2023). Strengths, weaknesses, opportunities, and threats associated with the application of artificial intelligence in connection with sport research, coaching, and optimization of athletic performance: a brief SWOT analysis. Frontiers in Sports and Active Living, 5, 1258562. https://doi.org/10.3389/fspor.2023.1258562
27. Tomczak, M., & Tomczak, E. (2014). The need to report effect size estimates revisited. An overview of some recommended measures of effect size. TRENDS in Sport Sciences. 1(21): 19-25.
28. Toros, T. (2011). Training exercise performance questionnaire (TEPQ) – development study: A study on sportsmen from branches of judo, taekwondo, karate. Archives of Budo, 7(2), 81–86.
29. Wachholz, F., Manno, S., Schlachter, D., Gamper, N., & Schnitzer, M. (2025). Acceptance and trust in AI-generated exercise plans among recreational athletes and quality evaluation by experienced coaches: A pilot study. BMC Research Notes, 18(1), 112. https://doi.org/10.1186/s13104-025-07172-9
30. Yang, J., Qin, S., & Ren, D. (2025). Artificial intelligence coaches on the sidelines: evaluating readability and quality of soccer training plans from six generative models. International Journal of Sports Science & Coaching, 17479541251369593. https://doi.org/10.1177/17479541251369593
31. Zhou, D., Keogh, J. W. L., Ma, Y., Tong, R. K. Y., Khan, A. R., & Jennings, N. R. (2025). Artificial intelligence in sport: A narrative review of applications, challenges and future trends. Journal of Sports Sciences, https://doi.org/10.1080/02640414.2025.2518694
