Evaluation of early student performance prediction given concept drift
Abstract
Forecasting student performance can help to identify students at risk and aids in recommending actions to improve their learning outcomes. That often involves elaborate machine learning pipelines. These tend to use large feature sets including behavioral data from learning management systems or demographic information. However, this complexity can lead to inaccurate predictions when concept drift occurs, or when a large number of features are used with a limited sample size. We investigate the performance of different machine learning pipelines on a data set with change in study behavior during the Covid-19 period. We demonstrate that (i) LASSO, a shrinkage estimator that reduces complexity and overfitting, outperforms several machine learning models under these circumstances, (ii) a linear regression relying on only two handcrafted features achieves higher accuracy and substantially less predictive bias than commonly used, more complex models with large feature sets. Due to their simplicity, these models can serve as a benchmark for future studies and a fallback model when substantial concept or covariate drift is encountered.
Identifier Metadata
| Identifier | 110.0716/CON.2026.00687 |
| Canonical | mdoi:110.0716/CON.2026.00687 |
| Resolver URL | https://mdoi.org/110.0716/CON.2026.00687 |
| Resource URL | Open resource |
| Document URL | Open document |
| Content Type | Article |
| Authors | Benedikt Sonnleitner, Tom Madou, Matthias Deceuninck, Filotas Theodosiou, Yves R. Sagaert |
| Year | 2025 |
| Depositor | Convergence Chronicles Organisation |
| Prefix | 110.0716 |
| Registered | July 21, 2026 |
| Updated | July 21, 2026 |
| Status | Active |
| Visibility | Public |
Cite This Identifier
APA 7th Edition
Click to copy
MLA 9th Edition
Click to copy
Chicago 17th Edition
Click to copy
BibTeX
Click to copy
Persistent Identifier
mdoi:110.0716/CON.2026.00687Click to copy