Evaluating large language models as raters in large-scale writing assessments: A psychometric framework for reliability and validity
Abstract
In large-scale international writing assessments, human raters often exhibit inconsistency, undermining reliability and validity. Large language models (LLMs) offer a potential solution, but their assessment reliability remains underexplored. This study employed generalizability theory and many-facet Rasch modeling to compare human and LLM raters across three essay genres (4315 samples). Findings reveal that human-LLM discrepancies stem from fundamental evaluation differences, with minimal divergence in key-point scoring. Humans excel in holistic scoring scenarios but struggle with complex analytical rubrics where LLMs demonstrate advantages. While LLMs perform adequately for relative ranking tasks, they remain less reliable for absolute standard judgments. Claude models exhibited superior scoring stability compared to GPT models, approaching perfect reliability in key-point scoring. Detailed hierarchical rubrics enabled LLMs to achieve human-comparable consistency even on subjective dimensions. Both human and LLM raters demonstrated random scoring behaviors with different patterns. LLMs rely on surface similarities rather than deep semantic understanding, while humans struggle with lengthy, complex rubrics. All scoring systems suffered from restriction-of-range effects, with model scores clustering around specific rating levels (particularly scores 2-4). Additionally, GPT models and human raters both exhibited halo effects, where overall scores were heavily influenced by single dominant dimensions. Information function analysis indicated humans better suit broad-spectrum assessment, while LLMs excel at fine-grained evaluation within narrow intervals. Regarding severity, humans typically assigned higher scores than LLMs, with GPT models being most stringent and Claude positioned intermediately. These findings contribute significantly to educational assessment by establishing a systematic framework for evaluating automated scoring systems.
Identifier Metadata
| Identifier | 110.0821/CON.2026.00792 |
| Canonical | mdoi:110.0821/CON.2026.00792 |
| Resolver URL | https://mdoi.org/110.0821/CON.2026.00792 |
| Resource URL | Open resource |
| Document URL | Open document |
| Content Type | Article |
| Authors | Yuehan Wang, Jinyan Huang, Lun Du, Yuxin Guo, Ying Liu, Rong Wang |
| Year | 2025 |
| Depositor | Convergence Chronicles Organisation |
| Prefix | 110.0821 |
| Registered | July 28, 2026 |
| Updated | July 28, 2026 |
| Status | Active |
| Visibility | Public |
Cite This Identifier
APA 7th Edition
Click to copy
MLA 9th Edition
Click to copy
Chicago 17th Edition
Click to copy
BibTeX
Click to copy
Persistent Identifier
mdoi:110.0821/CON.2026.00792Click to copy