Predicting Missing Values in Randomly Incomplete Data Using Artificial Intelligence Algorithms: An Experimental Study with More Than 95% Normalized Accuracy

missing values; artificial intelligence; machine learning; data imputation; MICE; Bayesian Ridge; KNN; random missing data.

Authors

  • Ahmed K Alhebshi Department of Engineering and Computing, University of Science and Technology Hadramout, Yemen
June 18, 2026
June 22, 2026

Downloads

Missing values often appear in operational datasets long before any statistical analysis or machine learning model is built. When the incomplete entries are treated with overly simple rules, such as row deletion or column means, the resulting dataset may no longer reflect the original relationships among variables. This paper presents an experimental framework for estimating randomly missing numerical values using artificial intelligence and machine learning techniques. The Wisconsin Diagnostic Breast Cancer dataset was used as a reproducible benchmark. Three random missingness levels were simulated: 10%, 20%, and 30%. Mean imputation, K-Nearest Neighbors imputation, and Bayesian Ridge-based Multivariate Imputation by Chained Equations were evaluated using Mean Absolute Error, Root Mean Squared Error, and Normalized MAE Accuracy. The Bayesian Ridge MICE approach produced the strongest results, achieving 97.51%, 96.93%, and 96.60% Normalized MAE Accuracy at 10%, 20%, and 30% missingness, respectively. These findings show that a multivariate AI-based imputation workflow can recover randomly missing values with more than 95% normalized accuracy in this experimental setting and can outperform the tested baseline methods.