MDOI Convergence Chronicles 110.0912/CON.2026.00883
110.0912/CON.2026.00883
Article

FermBench: A new benchmark for measuring the capabilities of LLMs on fermentation knowledge

Fiammetta Caccavale, Adem R.N. Aouichaoui, Ulrich Krühne, Krist V. Gernaey, Carina L. Gargalo 2026 Convergence Chronicles

Abstract

Generative Artificial Intelligence (GenAI) chatbots continue to amaze users worldwide with their rapid improvements. These tools possess vast general knowledge and can thus be used in various fields, including education. However, before rolling out these models in pedagogical applications, it is fundamental to understand whether the information provided is reliable and if any of the currently available chatbots are best suited for domain-specific tasks. The objective of this study is to thoroughly investigate these aspects in a specific domain, fermentation, with the overarching goal of providing guidelines to students and teachers to select the best GenAI assistant. To achieve this goal, we introduce FermBench, a dataset specifically designed for fermentation processes. We use the collected data to benchmark five large language models (LLMs) powering commercially available GenAI chatbots, including ChatGPT, Gemini, DeepSeek, Claude and le Chat. To evaluate the responses of these models, we propose a robust experimental framework that includes automated metrics, human annotations, and the LLM-as-a-Judge approach. The obtained results suggest that, given the high baseline and the fact that the judges were unable to agree on an overall best model, the current knowledge embedded within these models is adequate and the standalone results cannot provide pedagogical guidelines regarding which chatbot should be used in education. These results suggest that the choice of which GenAI chatbot should be supported by institutional or government guidance, as well as individual preferences, perhaps informed by the parameters identified in our analysis. An interesting finding of the study is that curated answers are not necessarily better than generated ones.

Identifier Metadata

Identifier 110.0912/CON.2026.00883
Canonical mdoi:110.0912/CON.2026.00883
Resolver URL https://mdoi.org/110.0912/CON.2026.00883
Resource URL Open resource
Document URL Open document
Content Type Article
Authors Fiammetta Caccavale, Adem R.N. Aouichaoui, Ulrich Krühne, Krist V. Gernaey, Carina L. Gargalo
Year 2026
Depositor Convergence Chronicles Organisation
Prefix 110.0912
Registered July 31, 2026
Updated July 31, 2026
Status Active
Visibility Public

Cite This Identifier

APA 7th Edition

Click to copy

MLA 9th Edition

Click to copy

Chicago 17th Edition

Click to copy

BibTeX

Click to copy

Persistent Identifier

mdoi:110.0912/CON.2026.00883

Click to copy

About MDOI

MDOI identifiers are permanent and unique identifiers assigned to digital objects to ensure long-term access, tracking, and referencing.

  • MDOI provides a permanent identity for digital objects.
  • Each MDOI is unique and points to one specific resource.
  • The prefix, such as 110.XXXX, identifies the registrant.
  • The suffix identifies the exact digital object.
  • MDOI remains stable even when a website URL changes.
  • It helps prevent broken links in digital publishing.
  • It makes academic and digital resources easier to find and cite.
  • MDOI supports proper tracking and management of digital content.
  • It improves the credibility and visibility of published resources.
  • MDOI ensures digital objects remain accessible, traceable, and reliable over time.
CO
Registered by Convergence Chronicles