How Classification Models Degrade as Data Diminish: A Multi-Criteria Analysis of Tabular Classifiers Under Controlled Data Scarcity
Downloads
Machine learning research is frequently conducted under an implicit assumption of data abundance, yet applied software development often proceeds under severe data constraints. This study evaluated how five widely used classification algorithms—logistic regression, support vector machines, Gaussian naive Bayes, decision trees, and random forests—degrade when training data are systematically reduced. Eight tabular datasets drawn from the UCI Machine Learning Repository, spanning small, medium, and large volume tiers, were analysed within a quantitative comparative experimental design. Training folds were reduced to 100%, 75%, 50%, and 25% of their original size through stratified fractional sampling inside a stratified five-fold cross-validation loop, while validation folds were retained at full volume. Ten criteria were recorded: accuracy, precision, recall, F1 score, AUC-ROC, AUC-PR, mean absolute error and mean squared error of predicted probabilities, training time, and interpretability. One-way analyses of variance and Tukey honestly significant difference tests were applied at an alpha level of .05. Between-model differences were statistically significant in 31 of 32 dataset-by-volume configurations. Tree-based models exhibited the largest reductions in accuracy and the steepest increases in calibration error under scarcity, whereas logistic regression preserved threshold balance and probability calibration at negligible computational cost. Random forests attained the highest scores only when data were abundant. Pairwise comparisons aggregated across datasets were not significant, indicating that dataset context, rather than architectural complexity alone, governs absolute performance.
Aeberhard, S., & Forina, M. (1992). *Wine* [Data set]. UCI Machine Learning Repository. https://doi.org/10.24432/C5PC7J
Al-Bahadili, H., Abu-Dalo, A., & Khreisat, L. (2021). Impact of dataset size on classification performance: An empirical evaluation in the medical domain. *Applied Sciences, 11*(2), 796. https://doi.org/10.3390/app11020796
Banko, M., & Brill, E. (2001). Scaling to very large corpora for natural language disambiguation. In *Proceedings of the 39th Annual Meeting of the Association for Computational Linguistics* (pp. 26–33).
Association for Computational Linguistics. https://aclanthology.org/P01-1005/
Becker, B., & Kohavi, R. (1996). *Adult* [Data set]. UCI Machine Learning Repository. https://doi.org/10.24432/C5XW20
Bock, R. (2004). *MAGIC gamma telescope* [Data set]. UCI Machine Learning Repository.
https://doi.org/10.24432/C52C8B
Bohanec, M. (1988). *Car evaluation* [Data set]. UCI Machine Learning Repository.
https://doi.org/10.24432/C5JP48
Breiman, L. (2001). Random forests. *Machine Learning, 45*(1), 5–32.
https://doi.org/10.1023/A:1010933404324
Breiman, L., Friedman, J. H., Olshen, R. A., & Stone, C. J. (1984). *Classification and regression trees*. Wadsworth.
Choi, Y., & Boo, Y. (2020). Comparing logistic regression models with alternative machine learning methods to predict the risk of drug intoxication mortality. *International Journal of Environmental Research and Public Health, 17*(3), 897.
https://doi.org/10.3390/ijerph17030897
Corea, P. M., Liu, Y., Wang, J., Niu, S., & Song, H. (2024). *Explainable AI for comparative analysis of intrusion detection models*. arXiv. https://arxiv.org/abs/2406.09684
Cortes, C., & Vapnik, V. (1995). Support-vector networks. *Machine Learning, 20*(3), 273–297. https://doi.org/10.1007/BF00994018
Domingos, P., & Pazzani, M. (1997). On the optimality of the simple Bayesian classifier under zero-one loss. *Machine Learning, 29*(2–3), 103–130. https://doi.org/10.1023/A:1007413511361
Feldman, K. (2025, January 23). A comprehensive guide to input-process-output models. *iSixSigma*.
https://www.isixsigma.com/dictionary/input-process-output-i-p-o/
Fernández-Delgado, M., Cernadas, E., Barro, S., & Amorim, D. (2014). Do we need hundreds of classifiers to solve real world classification problems? *Journal of Machine Learning Research, 15*(1), 3133–3181. https://jmlr.org/papers/v15/delgado14a.html
Gorman, R. P., & Sejnowski, T. J. (1988). *Connectionist bench (sonar, mines vs. rocks)* [Data set]. UCI Machine Learning Repository. https://doi.org/10.24432/C5T01Q
Halevy, A., Norvig, P., & Pereira, F. (2009). The unreasonable effectiveness of data. *IEEE Intelligent Systems, 24*(2), 8–12.
https://doi.org/10.1109/MIS.2009.36
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., … Oliphant, T. E. (2020). Array programming with NumPy. *Nature, 585*(7825), 357–362.
https://doi.org/10.1038/s41586-020-2649-2
Hopkins, M., Reeber, E., Forman, G., & Suermondt, J. (1999). *Spambase* [Data set]. UCI Machine Learning Repository.
https://doi.org/10.24432/C53G6X
Kelly, M., Longjohn, R., & Nottingham, K. (n.d.). *UCI Machine Learning Repository*. https://archive.ics.uci.edu/
Kohavi, R. (1995). A study of cross-validation and bootstrap for accuracy estimation and model selection. In *Proceedings of the 14th International Joint Conference on Artificial Intelligence* (pp. 1137–1143). Morgan Kaufmann. https://www.ijcai.org/Proceedings/95-2/Papers/016.pdf
Lohweg, V. (2012). *Banknote authentication* [Data set]. UCI Machine Learning Repository. https://doi.org/10.24432/C55P57
Lynam, A. L., Dennis, J. M., Owen, K. R., Oram, R. A., Jones, A. G., Shields, B. M., & Ferrat, L. A. (2020). Logistic regression has similar performance to optimised machine learning algorithms in a clinical setting: Application to the discrimination between type 1 and type 2 diabetes in young adults. *Diagnostic and Prognostic Research, 4*(1), 6.
https://doi.org/10.1186/s41512-020-00075-2
McCallum, A., & Nigam, K. (1998). A comparison of event models for naive Bayes text classification. In *AAAI-98 Workshop on Learning for Text Categorization* (pp. 41–48). AAAI Press.
https://cdn.aaai.org/Workshops/1998/WS-98-05/WS98-05-007.pdf
McKinney, W. (2010). Data structures for statistical computing in Python. In *Proceedings of the 9th Python in Science Conference* (pp. 56–61).
https://doi.org/10.25080/Majora-92bf1922-00a
Naser, M. Z. (2026). A review of machine learning with small and limited data. *Journal of Big Data, 13*(1), 18.
https://doi.org/10.1186/s40537-025-01346-9
Park, H.-A. (2013). An introduction to logistic regression: From basic concepts to interpretation with particular attention to nursing domain. *Journal of Korean Academy of Nursing, 43*(2),
–164. https://doi.org/10.4040/jkan.2013.43.2.154
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, R., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. *Journal of Machine Learning Research, 12*,
–2830. https://jmlr.org/papers/v12/pedregosa11a.html
Powers, D. M. W. (2011). Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation. *Journal of Machine Learning Technologies, 2*(1), 37–63.
Quinlan, J. R. (1986). Induction of decision trees. *Machine Learning, 1*(1), 81–106.
https://doi.org/10.1023/A:1022643204877
Refaeilzadeh, P., Tang, L., & Liu, H. (2009). Cross-validation. In L. Liu & M. T. Özsu (Eds.), *Encyclopedia of database systems* (pp. 532–538). Springer.
https://doi.org/10.1007/978-0-387-39940-9_565
Sokolova, M., & Lapalme, G. (2009). A systematic analysis of performance measures for classification tasks. *Information Processing & Management, 45*(4), 427–437.
https://doi.org/10.1016/j.ipm.2009.03.002
Van Smeden, M., Groenwold, R. H. H., Moons, K. G. M., Collins, G. S., & Riley, R. D. (2024). Sample size requirements for popular classification algorithms in tabular clinical data: Empirical study. *Journal of Medical Internet Research, 26*, e60231.
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, P., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., … SciPy 1.0 Contributors. (2020). SciPy 1.0: Fundamental algorithms for scientific computing in Python. *Nature Methods, 17*(3), 261–272. https://doi.org/10.1038/s41592-019-0686-2
Wolberg, W., Mangasarian, O., Street, N., & Street, W. (1993). *Breast cancer Wisconsin (diagnostic)* [Data set]. UCI Machine Learning Repository. https://doi.org/10.24432/C5DW2B
Zantvoort, K., Nacke, B., Görlich, D., Hornstein, S., Jacobi, C., & Funk, B. (2024). Estimation of minimal data set sizes for machine learning predictions in digital mental health interventions. *npj Digital Medicine, 7*, Article 361. https://doi.org/10.1038/s41746-024-01360-w
