- AUTHOR :
- Hayang You
- INFORMATION:
- page. 231~255 / 2026 Vol.55 No.1
| e-ISSN |
2713-3788 |
| p-ISSN |
1229-4179 |
ABSTRACT
This study investigates the potential of generative artificial intelligence (AI), particularly ChatGPT, as a supportive tool for item development and scoring of restricted and extended essay assessments in music education. For this purpose, assessment items appropriate for first-grade middle school music classes were developed using ChatGPT, then implemented and scored in an actual classroom setting. The findings indicate that during the item development phase, the generated items varied in content and scope despite using identical prompts. More advanced versions tended to offer clearer explanations of musical concepts and presented them in a more systematically organized manner. In the scoring phase, using customized GPT improved both efficiency and consistency compared to prompt-based approaches, which required the presentation of multiple materials simultaneously. Notably, incorporating teacher scoring results into the process significantly increased the level of agreement between ChatGPT and teacher scores. These findings highlight both the potential and limitations of generative AI as a tool to support music teachers’ assessment practices. They also suggest that teachers’ critical review and professional judgment are essential for the effective and responsible integration of generative AI in educational assessment.
Keyword :
REFERENCES
- Baek, M., & Park, C. (2025). A study on preservice home economics teachers' experiences in developing essay-type assessment items using generative AI. Comprehensive Research in Education, 23(1), 153-168. https://doi.org/10.31352/JER.23.1.153
[Crossref]
- Baek, S. (2002). Performance assessment: Theoretical perspectives. Kyoyook Kwahaksa.
- Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarw al, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., ... Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877-1901. https://doi.org/10.48550/arxiv.2005.14165
[Crossref]
- Choi, S., & Park, J. (2024). Analysis of in-service Korean-language teachers' development of essay-type assessment items using generative AI. Journal of CheongRam Korean Language Education, 97, 243-270. https://doi.org/10.26589/jockle..97.202401.243
[Crossref]
- Downing, S. M. (2006). Selected-response item formats in test development. In S. M. Downing & T. M. Haladyna (Eds.), Handbook of test development (pp. 287-301). Lawrence Erlbaum Associates.
- García-Varela, F., Nussbaum, M., Mendoza, M., Martínez-Troncoso, C., & Bekerman, Z. (2025). ChatGPT as a stable and fair tool for automated essay scoring. Education Sciences, 15(8), 946. https://doi.org/10.3390/educsci15080946
[Crossref]
- Gyeonggido Office of Education (2025). Guidebook for essay-type assessment in secondary education. Gyeonggido Office of Education.
- Ham, E., Park, S., Lee, B., Kim, G., & Lee, D. (2024). Protocol for developing constructedresponse items using ChatGPT and quality evaluation: Focusing on Korean language arts. Korean Journal of Educational Research, 62(8), 63-93. https://doi.org/10.30916/KERA.62.8.63
[Crossref]
- Han, S., & Han, K. (2025). Effects of parameter-based AI composition on self-efficacy and interest in elementary music education. Korean Journal of Research in Music Education, 54(3), 283-304. https://doi.org/10.30775/KMES.54.3.283
[Crossref]
- Holster, J. (2024). Augmenting music education through AI: Practical applications of ChatGPT. Music Educators Journal, 110(4), 36-42. https://doi.org/10.1177/00274321241255938
[Crossref]
- Im, J., & Lim, J. (2024). Action research on the development program of ChatGPT for automated scoring of essay type items. Journal of Educational Innovation Research, 34(1), 349-370. https://doi.org/10.21024/pnuedi.34.1.202403.349
[Crossref]
- Jeong, Y., & Baek, S. (2025). Comparison of the usefulness of prompt-engineering techniques for automatic scoring of social-studies essay-type items. Korean Journal of Educational Research, 63(4), 95-124. https://doi.org/10.30916/KERA.63.4.95
[Crossref]
- Ji, M., Oh, H., Noh, M., Kwon, U., Kim, Y., Lee, J., Jang, H., So, M., & Cho, B. (2024). Evaluating the 2022 revised curriculum: Adding wings with AI. AnswerBook.
- Kim, H. (2024). Educational significance and application directions of descriptive and essay-type assessment in the 2022 revised art curriculum. Art Education Review, 92, 27-44. https://doi.org/10.25297/AER.2024.92.27
[Crossref]
- Kim, K. (2020). Exploring the evaluative meaning of essay-type assessment. Journal of Educational Evaluation Research, 33(4), 839-862. https://doi.org/10.31158/JEEV.2020.33.4.839
[Crossref]
- Kim, S., & Hong, J. (2024). Popular music education using generative artificial intelligence: Focusing on ChatGPT. Korean Journal of Popular Music, 33, 35-68. https://doi.org/10.36775/kjpm.2024.33.35
[Crossref]
- Kwon, T. (2024). A study on writing scoring and feedback practices using ChatGPT: Focusing on prompting strategies. The Academy for Korean Language Education, 141, 7-42. https://doi.org/10.15734/koed..141.202412.7
[Crossref]
- Lee, B. (2015). Development and application of essay-type assessment items in middle school music education. Master's thesis, Korea National University of Education.
- Lee, J., Lee, E., & Kim, K. (2025). Enhancing alignment between large language models and teacher in open-ended assessment through in-context Learning. Brain, Digital, & Learning, 15(3), 375-401. https://doi.org/10.31216/BDL.2025.15.3.4
[Crossref]
- McMillan, J. H. (2013). Classroom assessment: Principles and practices for effective standardsbased instruction (S. Son, J. Park, S. Kang, C. Park, & K. Kim, Trans.). Kyoyook Kwahaksa.
- Ministry of Education (2022) Music curriculum. No. 2022-33. Ministry of Education.
- Ministry of Education (2023, October 10). Draft plan for the 2028 university admission system reform. Retrieved June 10, 2025, from https://moe.go.kr
- Moore, S., Nguyen, H. A., Bier, N., Domadia, T., & Stamper, J. (2022). Assessing the quality of student-generated short answer questions using GPT-3. Paper presented at the 17th European Conference on Technology Enhanced Learning (pp. 243-257), Toulouse. https://doi.org/10.1007/978-3-031-16290-9_18
[Crossref]
- OpenAI (2025). ChatGPT-5. Retrieved August 10, 2025, from https://platform.openai.com/docs/models
- Park, G., & Choi, S. (2025). A study on automatic item generation for multiple-choice assessment in the reading domain of Korean language education. Journal of Curriculum and Evaluation, 28(1), 215-246. https://doi.org/10.29221/jce.2025.28.1.215
[Crossref]
- Park, H., Kim, S., Kim, K., Lee, M., & Kim, K. (2019). Plans for enhancing the quality of essay-type assessment in individual schools. Research Report ORM 2019-54-10. Korea Institute for Curriculum and Evaluation.
- Park, J. (2024). An exploration of the current introduction and implementation issues of essay-type assessment in Korean language education. Journal of Cheong Ram Korean Language Education, 101, 273-307. https://doi.org/10.26589/jockle..101.202409.273
[Crossref]
- Park, S., Hong, Y., & Lee, B. (2024). A study on exploring the potential of ChatGPT in writing skills assessment: Focusing on essay writing. Korean Journal of Educational Research, 62(5), 219-248. https://doi.org/10.30916/KERA.62.5.219
[Crossref]
- Park, S., Lee, B., Ham, E., Lee, Y., & Lee, S. (2023). Exploring the possibility of science-inquiry competence assessment by ChatGPT-4: Comparisons with human evaluators. Korean Journal of Educational Research, 61(4), 299-332. https://doi.org/10.30916/KERA.61.4.299
[Crossref]
- Sengar, S. S., Hasan, A. B., Kumar, S., & Carroll, F. (2024). Generative artificial intelligence: A systematic review and applications. arXiv. https://doi.org/10.48550/arxiv.2405.11029 https://doi.org/10.1007/s11042-024-20016-1
[Crossref]
- Seong, J., & Shin, B. (2023). Exploring the feasibility of automatic scoring of written tests using ChatGPT: Focusing on the world geography written test. Journal of the Association of Korean Geographers, 12(3), 415-432. https://doi.org/10.25202/JAKG.12.3.3
[Crossref]
- Seong, T., Lim, H., Jeon, K., & Choi, Y. (2025). Foundations of educational evaluation (4th ed.). Hakjisa.
- Shaw, B. P. (2024). Artificial intelligence and assessment: Three implications for music educators. Music Educators Journal, 111(2), 19-25. https://doi.org/10.1177/00274321241296118
[Crossref]
- Shin, B., Lee, J., & Yoo, Y. (2024). Exploring automatic scoring of mathematical descriptive assessment using prompt engineering with the GPT-4 model: Focused on permutations and combinations. The Mathematical Education, 63(2), 187-207. https://doi.org/10.7468/mathedu.2024.63.2.187 https://doi.org/10.63311/mathedu.2024.63.2.187
[Crossref]
- Song, J. (2022). Improvement plans for essay-based assessment and university admissions in the transition to a future-oriented education system. Daegu Education 2022-102. Daegu Metropolitan Office of Education.
- Song, M., Kim, D., Kim, J., Park, S., Park, J., & Jeong, S. (2024). A study on the application of AI models for automated grading of subject-specific essay-type assessments. RRE 2024-9. Korea Institute for Curriculum and Evaluation.
- Swiecki, Z., Khosravi, H., Chen, G., Martinez-Maldonado, R., Lodge, J. M., Milligan, S., Selwyn, N., & Gašević, D. (2022). Assessment in the age of artificial intelligence. Computers and Education. Artificial Intelligence, 3, 100075. https://doi.org/10.1016/j.caeai.2022.100075
[Crossref]
- Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., & Zhou, D. (2023). Chain-of-thought prompting elicits reasoning in large language models. arXiv. https://doi.org/10.48550/arxiv.2201.11903 https://doi.org/10.52202/068431-1800
[Crossref]
- Yoon, K. (2024). A study on teaching plans for music creation applying generative AI based on 'Text to Music' to group investigation model. Korean Journal of Research in Music Education, 53(1), 143-164. https://doi.org/10.30775/KMES.53.1.143
[Crossref]