Repository logo
Research Outputs
Projects
People
Statistics
  1. Home
  2. HSG CRIS
  3. HSG Publications
  4. Make Deep Networks Shallow Again
Details

Make Deep Networks Shallow Again

Journal
Proceedings of the 15th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management
ISSN
2184-3228
ISBN
978-989-758-671-2
Type
conference paper
Date Issued
2023-11-15
Author(s)
Bernhard Bermeitinger  
;
Tomas Hrycej  
;
Siegfried Handschuh  
DOI
10.5220/0012203800003598
Abstract
Deep neural networks have a good success record and are thus viewed as the best architecture choice for complex applications. Their main shortcoming has been, for a long time, the vanishing gradient which prevented the numerical optimization algorithms from acceptable convergence. An important special case of network architecture, frequently used in computer vision applications, consists of using a stack of layers of the same dimension. For this architecture, a breakthrough has been achieved by the concept of residual connections-an identity mapping parallel to a conventional layer. This concept substantially alleviates the vanishing gradient problem and is thus widely used. The focus of this paper is to show the possibility of substituting the deep stack of residual layers with a shallow architecture with comparable expressive power and similarly good convergence properties. A stack of residual layers can be expressed as an expansion of terms similar to the Taylor expansion. This expansion suggests the possibility of truncating the higher-order terms and receiving an architecture consisting of a single broad layer composed of all initially stacked layers in parallel. In other words, a sequential deep architecture is substituted by a parallel shallow one. Prompted by this theory, we investigated the performance capabilities of the parallel architecture in comparison to the sequential one. The computer vision datasets MNIST and CIFAR10 were used to train both architectures for a total of 6,912 combinations of varying numbers of convolutional layers, numbers of filters, kernel sizes, and other meta parameters. Our findings demonstrate a surprising equivalence between the deep (sequential) and shallow (parallel) architectures. Both layouts produced similar results in terms of training and validation set loss. This discovery implies that a wide, shallow architecture can potentially replace a deep network without sacrificing performance. Such substitution has the potential to simplify network architectures, improve optimization efficiency, and accelerate the training process.
Language
English (United States)
Keywords
Residual Connection
Deep Neural Network
Shallow Neural Network
Computer Vision
Image Classification
Convolutional Networks
Book title
Proceedings of the 15th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management - KDIR
Publisher
SciTePress
Start page
339
End page
346
Pages
8
Event Location
Rome, Italy
Event Date
13-15 November 2023
Official URL
https://www.scitepress.org/PublicationsDetail.aspx?ID=njj8erFelyk%3d&t=1
URL
https://www.alexandria.unisg.ch/handle/20.500.14171/118680
Contact Email Address
bernhard.bermeitinger@unisg.ch
File(s)
Thumbnail Image

open.access

Name

122038.pdf

Type

Main Article

Size

269.06 KB

Format

Adobe PDF

Checksum (MD5)

85b15080a69c40b03016859fc50f37be

Thumbnail Image

open.access

Name

Make Deep Networks Shallow Again (KDIR2023, Rome).pdf

Description
Slides for the live presentation
Size

2.24 MB

Format

Adobe PDF

Checksum (MD5)

f72212765ad81e31db53d66348034c9a

Support
HSG researchers can find instructions here for adding or importing publications (DOI, ORCID). Please send questions to alexandria@unisg.ch

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Accessibility settings
  • Privacy policy
  • End User Agreement
  • Send Feedback
Repository logo COAR Notify