MDOI Convergence Chronicles 110.0886/CON.2026.00857
110.0886/CON.2026.00857
Article

EvalYaks: Instruction tuning datasets and LoRA fine-tuned models for automated scoring of CEFR B2 speaking assessment transcripts

Nicy Scaria, Silvester John Joseph Kennedy, Thomas Latinovich, Deepak Subramani 2026 Convergence Chronicles

Abstract

Relying on human experts to evaluate the Common European Framework of Reference for Languages (CEFR) speaking assessments in an e-learning environment creates scalability challenges, as it limits how quickly and widely assessments can be conducted. We aim to automate the evaluation of CEFR B2 English speaking assessments in e-learning environments from conversation transcripts. First, we evaluate the capability of leading open source and commercial Large Language Models (LLMs) to score a candidate’s performance across various criteria in the CEFR B2 speaking exam in both global and India-specific contexts. Next, we create a new expert-validated, CEFR-aligned synthetic conversational dataset with transcripts that are rated at different assessment scores. In addition, new instruction-tuned datasets are developed from the English Vocabulary Profile (up to CEFR B2 level) and the CEFR-SP WikiAuto datasets. Finally, using these new datasets, we perform parameter efficient instruction tuning of Mistral Instruct 7B v0.2 to develop a family of models called EvalYaks. Four models in this family are for assessing the four sections of the CEFR B2 speaking exam, one for identifying the CEFR level of vocabulary and generating level-specific vocabulary, and another for detecting the CEFR level of text and generating level-specific text. EvalYaks achieved an average acceptable accuracy of 96 %, a degree of variation of 0.35 levels, achieving performance competitive with state-of-the-art frontier models like GPT-4o and Gemini Flash 2.5. Furthermore, a pilot validation on real-world learner transcripts verified the model’s transferability to real-world assessment contexts. This demonstrates that a 7B parameter LLM instruction tuned with high-quality CEFR-aligned assessment data can effectively evaluate and score CEFR B2 English speaking assessments, offering a promising solution for scalable, automated language proficiency evaluation. The methodology is adaptable to other regional contexts and CEFR levels through appropriate data generation and validation protocols.

Identifier Metadata

Identifier 110.0886/CON.2026.00857
Canonical mdoi:110.0886/CON.2026.00857
Resolver URL https://mdoi.org/110.0886/CON.2026.00857
Resource URL Open resource
Document URL Open document
Content Type Article
Authors Nicy Scaria, Silvester John Joseph Kennedy, Thomas Latinovich, Deepak Subramani
Year 2026
Depositor Convergence Chronicles Organisation
Prefix 110.0886
Registered July 29, 2026
Updated July 29, 2026
Status Active
Visibility Public

Cite This Identifier

APA 7th Edition

Click to copy

MLA 9th Edition

Click to copy

Chicago 17th Edition

Click to copy

BibTeX

Click to copy

Persistent Identifier

mdoi:110.0886/CON.2026.00857

Click to copy

About MDOI

MDOI identifiers are permanent and unique identifiers assigned to digital objects to ensure long-term access, tracking, and referencing.

  • MDOI provides a permanent identity for digital objects.
  • Each MDOI is unique and points to one specific resource.
  • The prefix, such as 110.XXXX, identifies the registrant.
  • The suffix identifies the exact digital object.
  • MDOI remains stable even when a website URL changes.
  • It helps prevent broken links in digital publishing.
  • It makes academic and digital resources easier to find and cite.
  • MDOI supports proper tracking and management of digital content.
  • It improves the credibility and visibility of published resources.
  • MDOI ensures digital objects remain accessible, traceable, and reliable over time.
CO
Registered by Convergence Chronicles