How Classification Models Degrade as Data Diminish: A Multi-Criteria Analysis of Tabular Classifiers Under Controlled Data Scarcity

machine learning, data scarcity, model selection, classification, probability calibration, bias–variance trade-off,tabular data

Authors

August 21, 2026

Downloads

Machine learning research is frequently conducted under an implicit assumption of data abundance, yet applied software development often proceeds under severe data constraints. This study evaluated how five widely used classification algorithms—logistic regression, support vector machines, Gaussian naive Bayes, decision trees, and random forests—degrade when training data are systematically reduced. Eight tabular datasets drawn from the UCI Machine Learning Repository, spanning small, medium, and large volume tiers, were analysed within a quantitative comparative experimental design. Training folds were reduced to 100%, 75%, 50%, and 25% of their original size through stratified fractional sampling inside a stratified five-fold cross-validation loop, while validation folds were retained at full volume. Ten criteria were recorded: accuracy, precision, recall, F1 score, AUC-ROC, AUC-PR, mean absolute error and mean squared error of predicted probabilities, training time, and interpretability. One-way analyses of variance and Tukey honestly significant difference tests were applied at an alpha level of .05. Between-model differences were statistically significant in 31 of 32 dataset-by-volume configurations. Tree-based models exhibited the largest reductions in accuracy and the steepest increases in calibration error under scarcity, whereas logistic regression preserved threshold balance and probability calibration at negligible computational cost. Random forests attained the highest scores only when data were abundant. Pairwise comparisons aggregated across datasets were not significant, indicating that dataset context, rather than architectural complexity alone, governs absolute performance.