Articles
| e-ISSN | 2713-3788 |
| p-ISSN | 1229-4179 |
This study developed and evaluated an open-source, locally deployed LLM-based assessment tool for constructed responses in music education, designed to minimize external data transfer and support school-level privacy constraints. The system implements an end-to-end workflow—item input, automatic rubric generation and teacher editing, Excel-based response upload, rubric-based three-level classification, evidence citation, personalized improvement feedback, and Excel export. Using 200 labeled constructed responses in a Western Music History unit, we examined (1) classification performance and (2) scoring consistency across repeated runs after random shuffling of response order under fixed model and inference settings. To mitigate known issues in LLM scoring, the tool restricts the LLM to criterion-level judgments and evidence extraction while computing final scores via deterministic Python logic, combining rule-based scoring signals (RBSS) with model metadata in a hybrid architecture. Results indicated that the overall score and rank structure remained stable across runs, and the observed relationship between confidence and RBSS suggests that confidence can be interpreted as a decision-certainty signal rather than a direct proxy for correctness.
Keyword :
Review Fee: $100
Publication Fee: $145(~$20, when exceeding 20 pages)
Bank Account: https://www.paypal.me/kmes727