MDOI Convergence Chronicles 110.0821/CON.2026.00792
110.0821/CON.2026.00792
Article

Evaluating large language models as raters in large-scale writing assessments: A psychometric framework for reliability and validity

Yuehan Wang, Jinyan Huang, Lun Du, Yuxin Guo, Ying Liu, Rong Wang 2025 Convergence Chronicles

Abstract

In large-scale international writing assessments, human raters often exhibit inconsistency, undermining reliability and validity. Large language models (LLMs) offer a potential solution, but their assessment reliability remains underexplored. This study employed generalizability theory and many-facet Rasch modeling to compare human and LLM raters across three essay genres (4315 samples). Findings reveal that human-LLM discrepancies stem from fundamental evaluation differences, with minimal divergence in key-point scoring. Humans excel in holistic scoring scenarios but struggle with complex analytical rubrics where LLMs demonstrate advantages. While LLMs perform adequately for relative ranking tasks, they remain less reliable for absolute standard judgments. Claude models exhibited superior scoring stability compared to GPT models, approaching perfect reliability in key-point scoring. Detailed hierarchical rubrics enabled LLMs to achieve human-comparable consistency even on subjective dimensions. Both human and LLM raters demonstrated random scoring behaviors with different patterns. LLMs rely on surface similarities rather than deep semantic understanding, while humans struggle with lengthy, complex rubrics. All scoring systems suffered from restriction-of-range effects, with model scores clustering around specific rating levels (particularly scores 2-4). Additionally, GPT models and human raters both exhibited halo effects, where overall scores were heavily influenced by single dominant dimensions. Information function analysis indicated humans better suit broad-spectrum assessment, while LLMs excel at fine-grained evaluation within narrow intervals. Regarding severity, humans typically assigned higher scores than LLMs, with GPT models being most stringent and Claude positioned intermediately. These findings contribute significantly to educational assessment by establishing a systematic framework for evaluating automated scoring systems.

Identifier Metadata

Identifier 110.0821/CON.2026.00792
Canonical mdoi:110.0821/CON.2026.00792
Resolver URL https://mdoi.org/110.0821/CON.2026.00792
Resource URL Open resource
Document URL Open document
Content Type Article
Authors Yuehan Wang, Jinyan Huang, Lun Du, Yuxin Guo, Ying Liu, Rong Wang
Year 2025
Depositor Convergence Chronicles Organisation
Prefix 110.0821
Registered July 28, 2026
Updated July 28, 2026
Status Active
Visibility Public

Cite This Identifier

APA 7th Edition

Click to copy

MLA 9th Edition

Click to copy

Chicago 17th Edition

Click to copy

BibTeX

Click to copy

Persistent Identifier

mdoi:110.0821/CON.2026.00792

Click to copy

About MDOI

MDOI identifiers are permanent and unique identifiers assigned to digital objects to ensure long-term access, tracking, and referencing.

  • MDOI provides a permanent identity for digital objects.
  • Each MDOI is unique and points to one specific resource.
  • The prefix, such as 110.XXXX, identifies the registrant.
  • The suffix identifies the exact digital object.
  • MDOI remains stable even when a website URL changes.
  • It helps prevent broken links in digital publishing.
  • It makes academic and digital resources easier to find and cite.
  • MDOI supports proper tracking and management of digital content.
  • It improves the credibility and visibility of published resources.
  • MDOI ensures digital objects remain accessible, traceable, and reliable over time.
CO
Registered by Convergence Chronicles