Abstract
Researchers have sought for decades to automate holistic essay scoring. Over the years, these programs have improved significantly. However, accuracy requires significant amounts of training on human-scored texts—reducing the expediency and usefulness of such programs for routine uses by teachers across the nation on non-standardized prompts. This study analyzes the output of multiple versions of ChatGPT scoring of secondary student essays from three extant corpora and compares it to quality human ratings. We find that the current iteration of ChatGPT scoring is not statistically significantly different from human scoring; substantial agreement with humans is achievable and may be sufficient for low-stakes, formative assessment purposes. However, as large language models evolve additional research will be needed to continue to assess their aptitude for this task as well as determine whether their proximity to human scoring can be improved through prompting or training.
Identifier Metadata
| Identifier | 110.1060/CON.2026.01031 |
| Canonical | mdoi:110.1060/CON.2026.01031 |
| Resolver URL | https://mdoi.org/110.1060/CON.2026.01031 |
| Resource URL | Open resource |
| Document URL | Open document |
| Content Type | Article |
| Authors | Tamara P. Tate, Jacob Steiss, Drew Bailey, Steve Graham, Youngsun Moon, Daniel Ritchie, Waverly Tseng, Mark Warschauer |
| Year | 2024 |
| Depositor | Convergence Chronicles Organisation |
| Prefix | 110.1060 |
| Registered | Aug. 12, 2026 |
| Updated | Aug. 12, 2026 |
| Status | Active |
| Visibility | Public |
Cite This Identifier
APA 7th Edition
Click to copy
MLA 9th Edition
Click to copy
Chicago 17th Edition
Click to copy
BibTeX
Click to copy
Persistent Identifier
mdoi:110.1060/CON.2026.01031Click to copy