Repository logo
Research Outputs
Projects
People
Statistics
  1. Home
  2. HSG CRIS
  3. HSG Publications
  4. A Convexity-dependent Two-Phase Training Algorithm for Deep Neural Networks
Details

A Convexity-dependent Two-Phase Training Algorithm for Deep Neural Networks

Type
forthcoming
Date Issued
2025-10-29
Author(s)
Tomas Hrycej  
;
Bernhard Bermeitinger  
;
Massimo Pavone  
;
Götz-Henrik Wiegand  
;
Siegfried Handschuh  
DOI
10.48550/arXiv.2510.25366
Abstract
The key task of machine learning is to minimize the loss function that measures the model fit to the training data. The numerical methods to do this efficiently depend on the properties of the loss function. The most decisive among these properties is the convexity or non-convexity of the loss function. The fact that the loss function can have, and frequently has, non-convex regions has led to a widespread commitment to non-convex methods such as Adam. However, a local minimum implies that, in some environment around it, the function is convex. In this environment, second-order minimizing methods such as the Conjugate Gradient (CG) give a guaranteed superlinear convergence. We propose a novel framework grounded in the hypothesis that loss functions in real-world tasks swap from initial non-convexity to convexity towards the optimum. This is a property we leverage to design an innovative two-phase optimization algorithm. The presented algorithm detects the swap point by observing the gradient norm dependence on the loss. In these regions, non-convex (Adam) and convex (CG) algorithms are used, respectively. Computing experiments confirm the hypothesis that this simple convexity structure is frequent enough to be practically exploited to substantially improve convergence and accuracy.
Abstract (De)
Die Hauptaufgabe des maschinellen Lernens liegt in der Minimierung der Verlustfunktion, welche die Modellanpassung an die Trainingsdaten quantifiziert. Die effizienten numerischen Verfahren hierzu sind von den Eigenschaften der Verlustfunktion abhängig. Die ausschlaggebendste Eigenschaft ist dabei die Konvexität beziehungsweise Nicht-Konvexität der Verlustfunktion. Der Umstand, dass Verlustfunktionen nicht-konvexe Regionen aufweisen können und dies häufig tun, hat zu einer weitreichenden Präferenz für nicht-konvexe Methoden wie Adam geführt. Allerdings impliziert ein lokales Minimum, dass die Funktion in dessen Umgebung konvex ist. In dieser Umgebung garantieren Minimierungsverfahren zweiter Ordnung wie das Konjugierte Gradientenverfahren (CG) superlineare Konvergenz. Wir präsentieren ein neuartiges Framework, das auf der Hypothese basiert, dass Verlustfunktionen bei realen Aufgabenstellungen von initialer Nicht-Konvexität zur Konvexität in Richtung Optimum übergehen. Diese Eigenschaft nutzen wir zur Konstruktion eines innovativen zweiphasigen Optimierungsalgorithmus. Der vorgestellte Algorithmus detektiert den Übergangspunkt mittels Beobachtung der Abhängigkeit der Gradientennorm vom Verlust. In den jeweiligen Regionen kommen nicht-konvexe (Adam) respektive konvexe (CG) Algorithmen zum Einsatz. Rechenexperimente bestätigen die Hypothese, dass diese simple Konvexitätsstruktur hinreichend frequent auftritt, um sie praktisch zur substanziellen Verbesserung von Konvergenz und Genauigkeit zu nutzen.
Language
English
HSG Classification
contribution to scientific community
Refereed
Yes
Pages
9
Official URL
https://arxiv.org/abs/2510.25366
URL
https://www.alexandria.unisg.ch/handle/20.500.14171/124178
Subject(s)

computer science

Division(s)

ICS - Institute of Co...

ICV - Institute for C...

Contact Email Address
bernhard.bermeitinger@unisg.ch
Additional Information
Appeared on KDIR IC3K Conference 2025 (Best Paper Award)
File(s)
Thumbnail Image
Name

2510.25366v1.pdf

Size

630.98 KB

Format

Adobe PDF

Checksum (MD5)

d424bf25648322e016588c386b153bd2

Support
HSG researchers can find instructions here for adding or importing publications (DOI, ORCID). Please send questions to alexandria@unisg.ch

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Accessibility settings
  • Privacy policy
  • End User Agreement
  • Send Feedback
Repository logo COAR Notify