MDOI Convergence Chronicles 110.0931/CON.2026.00902
110.0931/CON.2026.00902
Article

Validating AI-generated classroom observations: Reliability, accuracy, and limits of LLM-based pedagogical judgment

Carolina Melo, Javiera de la Maza, Matías Recabarren 2025 Convergence Chronicles

Abstract

This study examines the reliability and accuracy of large language models (LLMs) for automated classroom observation using the World Bank's TEACH Primary framework. As education systems increasingly explore AI-based tools to scale teacher feedback and professional development, empirical validation of these systems is critical. Using a corpus of 12 primary classroom videos, we compared 8618 AI-generated evaluations from eight LLM endpoints against consensus-based ratings from certified TEACH experts. To account for model stochasticity, each model produced 10 independent evaluations per video–element pair. Reliability was assessed using variability and inter-rater consistency indicators, while accuracy was evaluated using error-based and concordance-based agreement measures. Results show substantial stochastic variability across repeated evaluations, with no model achieving uniformly high reliability across instructional elements. Agreement with expert ratings remained moderate at best. Importantly, reliability and accuracy did not co-vary systematically: models producing more stable scores did not necessarily align better with expert judgments, and models with stronger expert agreement often exhibited higher internal variability. In an exploratory analysis of model justifications, patterns suggest that LLMs tend to prioritize explicit verbal cues over contextual or implicit pedagogical evidence when generating high-inference judgments. These findings highlight structural limitations of current text-based AI observation pipelines and demonstrate that automated classroom observation cannot be treated as a uniform capability. The study provides empirical evidence to inform the design and validation of AI-assisted observation systems that integrate pedagogical expertise, measurement constraints, and complementary human judgment.

Identifier Metadata

Identifier 110.0931/CON.2026.00902
Canonical mdoi:110.0931/CON.2026.00902
Resolver URL https://mdoi.org/110.0931/CON.2026.00902
Resource URL Open resource
Document URL Open document
Content Type Article
Authors Carolina Melo, Javiera de la Maza, Matías Recabarren
Year 2025
Depositor Convergence Chronicles Organisation
Prefix 110.0931
Registered Aug. 1, 2026
Updated Aug. 1, 2026
Status Active
Visibility Public

Cite This Identifier

APA 7th Edition

Click to copy

MLA 9th Edition

Click to copy

Chicago 17th Edition

Click to copy

BibTeX

Click to copy

Persistent Identifier

mdoi:110.0931/CON.2026.00902

Click to copy

About MDOI

MDOI identifiers are permanent and unique identifiers assigned to digital objects to ensure long-term access, tracking, and referencing.

  • MDOI provides a permanent identity for digital objects.
  • Each MDOI is unique and points to one specific resource.
  • The prefix, such as 110.XXXX, identifies the registrant.
  • The suffix identifies the exact digital object.
  • MDOI remains stable even when a website URL changes.
  • It helps prevent broken links in digital publishing.
  • It makes academic and digital resources easier to find and cite.
  • MDOI supports proper tracking and management of digital content.
  • It improves the credibility and visibility of published resources.
  • MDOI ensures digital objects remain accessible, traceable, and reliable over time.
CO
Registered by Convergence Chronicles