Repository logo
Research Outputs
Projects
People
Statistics
  1. Home
  2. HSG CRIS
  3. HSG Publications
  4. Efficient Neural Network Training via Subset Pretraining
Details

Efficient Neural Network Training via Subset Pretraining

Journal
Proceedings of the 16th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management
ISSN
2184-3228
Type
conference paper
Date Issued
2024-11
Author(s)
Jan Spörer
;
Bernhard Bermeitinger  
;
Tomas Hrycej  
;
Niklas Limacher
;
Siegfried Handschuh  
DOI
10.5220/0012893600003838
Abstract
In training neural networks, it is common practice to use partial gradients computed over batches, mostly very small subsets of the training set. This approach is motivated by the argument that such a partial gradient is close to the true one, with precision growing only with the square root of the batch size. A theoretical justification is with the help of stochastic approximation theory. However, the conditions for the validity of this theory are not satisfied in the usual learning rate schedules. Batch processing is also difficult to combine with efficient second-order optimization methods. This proposal is based on another hypothesis: the loss minimum of the training set can be expected to be well-approximated by the minima of its subsets. Such subset minima can be computed in a fraction of the time necessary for optimizing over the whole training set. This hypothesis has been tested with the help of the MNIST, CIFAR-10, and CIFAR-100 image classification benchmarks, optionally extended by training data augmentation. The experiments have confirmed that results equivalent to conventional training can be reached. In summary, even small subsets are representative if the overdetermination ratio for the given model parameter set sufficiently exceeds unity. The computing expense can be reduced to a tenth or less.
Language
English
Keywords
Deep Neural Network
Convolutional Network
Computer Vision
Efficient Training
Resource Optimization
Training Strategies
Determination Ratio
Stochastic Approximation Theory
HSG Classification
contribution to scientific community
Refereed
Yes
Publisher
SciTePress
Start page
242
End page
249
Pages
8
Event Title
16th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management
Event Location
Porto, Portugal
Event Date
17-19 November 2024
Official URL
https://www.scitepress.org/Link.aspx?doi=10.5220%2f0012893600003838
URL
https://www.alexandria.unisg.ch/handle/20.500.14171/123208
Contact Email Address
jan.spoerer@unisg.ch
File(s)
Thumbnail Image

restricted

Name

128936.pdf

Size

1 MB

Format

Adobe PDF

Checksum (MD5)

249f96220ed6fd658c1c7679f65ed1b3

Support
HSG researchers can find instructions here for adding or importing publications (DOI, ORCID). Please send questions to alexandria@unisg.ch

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Accessibility settings
  • Privacy policy
  • End User Agreement
  • Send Feedback
Repository logo COAR Notify