MDOI Convergence Chronicles 110.0978/CON.2026.00949
110.0978/CON.2026.00949
Article

Comparing human and LLM ordered coding of qualitative data: How coding differences cascade through temporal analysis

Kamila Misiejuk, Sonsoles Lopez-Pernas, Eduardo Araujo Oliveira, Brendan Eagan, Mohammed Saq 2026 Convergence Chronicles

Abstract

Automating the process of qualitatively coding text data from learners has been a long-standing ambition of learning analytics researchers since it represents an essential step toward delivering timely and scalable feedback. Automating this process is especially challenging in the case of ordered coding schemes —necessary for temporal analytical methods— where one text utterance can be assigned more than one qualitative code and the assignment order matters. This problem goes beyond multi-class and multi-label classification and, therefore, cannot be easily tackled using classic language models such as BERT. Recent advances in generative artificial intelligence, especially with the advent of large language models, have —allegedly— created a substantial step forward in making the goal of automatically coding complex temporal data attainable. However, little is yet known about how to implement this process in a way that most closely resembles human coding, i.e., taking into account the context in which the textual data appears for accurate interpretation. Moreover, due to the complexity of the data and its shape, the accuracy of the results cannot be computed using classic accuracy metrics. This study makes two main contributions: first, it presents two evaluation approaches for assessing the quality of ordered data coding and the usability of LLM in automatically coding ordered processes; and second, it demonstrates a method of LLM prompting that leverages a consistent context window. Our results reveal systematic and statistically significant differences between LLM and human coding across structural, transitional, and code-level metrics for binary and ordered tasks. As classification errors can propagate through automated feedback systems, relying on LLM outputs risks amplifying inaccuracies and producing misleading interpretations of learning processes.

Identifier Metadata

Identifier 110.0978/CON.2026.00949
Canonical mdoi:110.0978/CON.2026.00949
Resolver URL https://mdoi.org/110.0978/CON.2026.00949
Resource URL Open resource
Document URL Open document
Content Type Article
Authors Kamila Misiejuk, Sonsoles Lopez-Pernas, Eduardo Araujo Oliveira, Brendan Eagan, Mohammed Saq
Year 2026
Depositor Convergence Chronicles Organisation
Prefix 110.0978
Registered Aug. 3, 2026
Updated Aug. 3, 2026
Status Active
Visibility Public

Cite This Identifier

APA 7th Edition

Click to copy

MLA 9th Edition

Click to copy

Chicago 17th Edition

Click to copy

BibTeX

Click to copy

Persistent Identifier

mdoi:110.0978/CON.2026.00949

Click to copy

About MDOI

MDOI identifiers are permanent and unique identifiers assigned to digital objects to ensure long-term access, tracking, and referencing.

  • MDOI provides a permanent identity for digital objects.
  • Each MDOI is unique and points to one specific resource.
  • The prefix, such as 110.XXXX, identifies the registrant.
  • The suffix identifies the exact digital object.
  • MDOI remains stable even when a website URL changes.
  • It helps prevent broken links in digital publishing.
  • It makes academic and digital resources easier to find and cite.
  • MDOI supports proper tracking and management of digital content.
  • It improves the credibility and visibility of published resources.
  • MDOI ensures digital objects remain accessible, traceable, and reliable over time.
CO
Registered by Convergence Chronicles