Comparative evaluation of bayesian and neural network models for cardiovascular disease prediction considering calibration and uncertainty
Main Article Content
Abstract
This paper investigates the problem of reliability in probabilistic forecasting of cardiovascular diseases within decision support systems. Unlike the conventional approach, where model quality is primarily evaluated using a single classification metric, this study considers a comprehensive set of interrelated properties of the predictive system: discriminative ability, alignment of predicted probabilities with actual frequencies, uncertainty estimation, error detection capability, and the possibility of rejecting automated decisions for complex cases. The experiments were conducted on the open-source Cardiovascular Disease Dataset. Following conservative data cleaning, over sixty-eight thousand observations were analyzed; twenty-one input variables were generated from the initial features after encoding and scaling. A unified experimental framework was used to compare a Naive Bayes classifier, a Bayesian Network, Logistic Regression, a Multilayer Perceptron, Monte Carlo Dropout, Last-Layer Laplace Approximation (Neural-Laplace), a Hamiltonian Monte Carlo Bayesian Last Layer, and a Deep Ensemble. The results demonstrate that on the full test set, the differences among models in their ability to rank observations are small: the Deep Ensemble achieved the highest area under the receiver operating characteristic curve, whereas the Neural-Laplace model exhibited the lowest expected calibration error among the studied models. For models yielding predictive distributions, uncertainty estimation detected erroneous predictions better than random ranking; the highest area under the receiver operating characteristic curve was obtained by the Neural-Laplace model. Selective prediction revealed that excluding cases with the highest estimated uncertainty from automated processing reduces the risk on the remaining subset. The experiment involving the reduction of the training sample size showed that when using only ten percent of the training data, the discrepancy in calibration metrics between the Neural-Laplace and the Multilayer Perceptron models was most pronounced: the expected calibration error of the Neural-Laplace model was over seven times lower. The scientific result consists in establishing, within a unified experimental framework, that the Bayesian component does not automatically increase discriminative ability but can alter the predictive system's quality profile regarding calibration and uncertainty: the Neural-Laplace model demonstrated the lowest expected calibration error and the best ability to detect erroneous predictions based on uncertainty estimation. The practical value lies in utilizing such a performance profile to develop selective prediction mechanisms in decision support systems.

