Now showing 1 - 6 of 6
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    425-P: Machine Learning Identifies Glypican 4 as Key Predictor of Five-Year Mortality in Heart Failure Patients with Prediabetes or Diabetes
    (2025-06-20)
    Leiherer, Andreas
    ;
    Muendlein, Axel
    ;
    Schnetzer, Laura
    ;
    Mink, Sylvia
    ;
    Heinzle, Christine
    Introduction and Objective: The rise of Big Data necessitates artificial intelligence-driven analyses to extract valuable insights, particularly for risk prediction in high-risk patient populations. This observational study applied machine learning (ML) algorithms to predict 5-year overall mortality in heart failure patients with type 2 diabetes mellitus (T2DM) or prediabetes Methods: A cohort of 290 heart failure patients with T2DM or prediabetes was followed for 5 years, during which 54% of participants died. The dataset comprised 470 variables, e.g. anthropometric, clinical, social, family history, and lifestyle factors. After preprocessing, the data were analyzed using ML techniques implemented in R’s caret package. The dataset was split into training (75%) and test (25%) subsets. Results: Among the ML models tested, the Random Forest algorithm demonstrated the best predictive performance, with a sensitivity of 82%, specificity of 89%, and overall accuracy of 85%. Clinical parameters were the most significant predictors, with the multimorbidity marker Glypican-4, hemoglobin, and glomerular filtration rate identified as the top three contributors. Conclusion: In conclusion, ML-based Big Data analysis holds great potential for predicting mortality risk in pre-/diabetic heart failure patients, paving the way for personalized and timely interventions.
    Type:
    Journal:
    Volume:
    Issue:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    1386-P: Machine Learning Predicts T2DM Incidence Using Basic Clinical and Laboratory Parameters
    (2025-06-20)
    Leiherer, Andreas
    ;
    Schnetzer, Laura
    ;
    Mink, Sylvia
    ;
    Mader, Arthur
    ;
    Muendlein, Axel
    Introduction and Objective: Artificial Intelligence (AI) and Machine Learning (ML) have the potential to improve risk prediction by identifying complex patterns in clinical and laboratory data, surpassing traditional approaches. ML has already shown success in detecting metabolic diseases, including Type 2 Diabetes Mellitus (T2DM). However, the ability to accurately forecast T2DM incidence is even more beneficial, enabling earlier interventions and treatment. Methods: This observational study aimed to leverage ML to predict the 4-year risk of developing T2DM. A cohort of 904 cardiovascular risk patients, initially free of T2DM, was analyzed at baseline, with data including anthropometric measurements, clinical and laboratory parameters, and recent metabolic biomarkers. Over four years of follow-up, 10.2% of the patients developed T2DM. Results: The ML approach, utilizing the Caret package in R, applied 50 variables. Patients were randomly split into training and test cohorts (75:25), with oversampling used to address class imbalance in T2DM incidence. Recursive feature elimination (RFE) was employed to identify the most relevant variables. A Support Vector Machine (SVM) model with a linear kernel demonstrated the most promising predictive performance, achieving a balanced accuracy of 73%, a sensitivity of 74%, a specificity of 71%, and an AUC of 0.727. The top-ranked predictors for T2DM were glucose measurements (2-hour OGTT glucose, fasting glucose, and HbA1c), HDL-cholesterol, and the triglyceride-glucose (TyG) index. Conclusion: In conclusion, ML proves to be a valuable tool for identifying individuals at risk of T2DM, paving the way for personalized medicine through earlier diagnosis and tailored interventions.
    Type:
    Journal:
    Volume:
    Issue:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    426-P: Predicting Coronary Stenoses Using Machine Learning to Reduce Unnecessary Angiographies
    (2025-06-20)
    Leiherer, Andreas
    ;
    Schnetzer, Laura
    ;
    Mink, Sylvia
    ;
    Muendlein, Axel
    ;
    Introduction and Objective: Coronary angiography is the gold standard for diagnosing coronary artery stenoses, but is invasive and confers potential risks. This study aimed to develop a Machine Learning (ML) model to predict significant stenoses while minimizing false negatives, ensuring accurate risk stratification and better patient selection. Methods: Data from 2,310 patients undergoing coronary angiography were analyzed, with outcomes classified as no stenoses (X0), non-significant stenoses (X1), or significant stenoses (X2). Results: XGBoost, optimized through grid search and 5-fold cross-validation, emerged as the top-performing ML algorithm. Of 114 clinical and laboratory variables, fibrinogen, HbA1c, BMI, waist-hip ratio, TyG index, FGF23, ceramides, and vitamin D - all associated with insulin resistance and diabetes - were identified as key contributors to the ML model (figure). Overall, the model achieved 62.9% accuracy (95% CI: 57.8-67.9). Sensitivity, precision, and F1 score for X2 were 94.7\%, 62.7\%, and 74.5\%. For X0, sensitivity was 37.3%, with precision and F1 scores of 67.6% and 48.1%. Conclusion: This ML-based approach has the potential to reduce unnecessary angiographies and optimize patient selection in clinical practice. Highlighting the relevance of diabetes-linked variables, the study also underscores the potential of metabolic profiling in coronary risk stratification.
    Type:
    Journal:
    Volume:
    Issue:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Subset Pretraining for Enhancing Neural Network Training Efficiency
    We propose a novel alternative to traditional randomly sampled mini-batches for gradient computation: using a fixed subset for complete pretraining of a neural network model. This approach enables deterministic convergence instead of a merely probabilistic one, as proven by the stochastic approximation theory, whose prerequisites are frequently violated by popular optimization algorithms. The approach is justified by the hypothesis that the loss minimum of the training set can be expected to be well-approximated by the minima of its subsets. Such subset minima can be computed in a fraction of the time necessary for optimizing with the whole training set. They are also compatible with efficient second-order optimization methods, such as the conjugate gradient optimizer. These methods are particularly efficient in the convex environment of the loss minimum. The image classification datasets MNIST, CIFAR-10, and CIFAR-100, (optionally extended by augmentation of training data) test this hypothesis. The experiments confirm that the models achieve performance equivalent to that when trained with the conventional training scheme. In conclusion, if the overdetermination ratio for the given model and dataset sufficiently exceed unity, even small subsets are representative. This results in a possible reduction of the computing expense to a tenth or less. This paper is an extended version of Spörer et al. [13].
    Type:
    Journal:
    Volume:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Training Neural Networks in Single vs Double Precision
    The commitment to single-precision floating-point arithmetic is widespread in the deep learning community. To evaluate whether this commitment is justified, the influence of computing precision (single and double precision) on the optimization performance of the Conjugate Gradient (CG) method (a second-order optimization algorithm) and RMSprop (a first-order algorithm) has been investigated. Tests of neural networks with one to five fully connected hidden layers and moderate or strong nonlinearity with up to 4 million network parameters have been optimized for Mean Square Error (MSE). The training tasks have been set up so that their MSE minimum was known to be zero. Computing experiments have disclosed that single-precision can keep up (with superlinear convergence) with double-precision as long as line search finds an improvement. First-order methods such as RMSprop do not benefit from double precision. However, for moderately nonlinear tasks, CG is clearly superior. For strongly nonlinear tasks, both algorithm classes find only solutions fairly poor in terms of mean square error as related to the output variance. CG with double floating-point precision is superior whenever the solutions have the potential to be useful for the application goal.
    Type:
    Journal:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Mathematical Foundations of Data Science
    This textbook aims to point out the most important principles of data analysis from the mathematical point of view. Specifically, it selected these questions for exploring: Which are the principles necessary to understand the implications of an application, and which are necessary to understand the conditions for the success of methods used? Theory is presented only to the degree necessary to apply it properly, striving for the balance between excessive complexity and oversimplification. Its primary focus is on principles crucial for application success. Topics and features: Focuses on approaches supported by mathematical arguments, rather than sole computing experiences Investigates conditions under which numerical algorithms used in data science operate, and what performance can be expected from them Considers key data science problems: problem formulation including optimality measure; learning and generalization in relationships to training set size and number of free parameters; and convergence of numerical algorithms Examines original mathematical disciplines (statistics, numerical mathematics, system theory) as they are specifically relevant to a given problem Addresses the trade-off between model size and volume of data available for its identification and its consequences for model parametrization Investigates the mathematical principles involves with natural language processing and computer vision Keeps subject coverage intentionally compact, focusing on key issues of each topic to encourage full comprehension of the entire book Although this core textbook aims directly at students of computer science and/or data science, it will be of real appeal, too, to researchers in the field who want to gain a proper understanding of the mathematical foundations “beyond” the sole computing experience.
    Type:
    Volume:
    Scopus© Citations 8