Jan Spörer
Last Name
Spörer
First name
Jan
Email
jan.spoerer@unisg.ch
3 results
Now showing 1 - 3 of 3
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, Subset Pretraining for Enhancing Neural Network Training EfficiencyWe propose a novel alternative to traditional randomly sampled mini-batches for gradient computation: using a fixed subset for complete pretraining of a neural network model. This approach enables deterministic convergence instead of a merely probabilistic one, as proven by the stochastic approximation theory, whose prerequisites are frequently violated by popular optimization algorithms. The approach is justified by the hypothesis that the loss minimum of the training set can be expected to be well-approximated by the minima of its subsets. Such subset minima can be computed in a fraction of the time necessary for optimizing with the whole training set. They are also compatible with efficient second-order optimization methods, such as the conjugate gradient optimizer. These methods are particularly efficient in the convex environment of the loss minimum. The image classification datasets MNIST, CIFAR-10, and CIFAR-100, (optionally extended by augmentation of training data) test this hypothesis. The experiments confirm that the models achieve performance equivalent to that when trained with the conventional training scheme. In conclusion, if the overdetermination ratio for the given model and dataset sufficiently exceed unity, even small subsets are representative. This results in a possible reduction of the computing expense to a tenth or less. This paper is an extended version of Spörer et al. [13].Type:conference paperJournal:Communications in Computer and Information ScienceVolume:2703 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Efficient Neural Network Training via Subset Pretraining(SCITEPRESS - Science and Technology Publications, 2024-11-22); ; ; ;Limacher, NiklasIn training neural networks, it is common practice to use partial gradients computed over batches, mostly very small subsets of the training set. This approach is motivated by the argument that such a partial gradient is close to the true one, with precision growing only with the square root of the batch size. A theoretical justification is with the help of stochastic approximation theory. However, the conditions for the validity of this theory are not satisfied in the usual learning rate schedules. Batch processing is also difficult to combine with efficient second-order optimization methods. This proposal is based on another hypothesis: the loss minimum of the training set can be expected to be well-approximated by the minima of its subsets. Such subset minima can be computed in a fraction of the time necessary for optimizing over the whole training set. This hypothesis has been tested with the help of the MNIST, CIFAR-10, and CIFAR-100 image classification benchmarks, optionally extended by training data augmentation. The experiments have confirmed that results equivalent to conventional training can be reached. In summary, even small subsets are representative if the overdetermination ratio for the given model parameter set sufficiently exceeds unity. The computing expense can be reduced to a tenth or less.Type:conference paperJournal:Proceedings of the 16th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge ManagementScopus© Citations 3 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Is Small Really Beautiful for Central Bank Communication? Evaluating Language Models for Finance: Llama-3-70B, GPT-4, FinBERT-FOMC, FinBERT, and VADERThis study compares the sentiment detection capabilities of language models for the domain of central bank communication, particularly the official statements released by the U.S. Federal Open Market Committee (FOMC). While previous studies have explored FOMC communication, this work is one of the few studies that use a natural language processing-based approach. The analysis employs VADER, FinBERT, a fine-tuned FinBERT model (FinBERT-FOMC), GPT-4, and Llama-3-70B. Within the scope of our labeled dataset on FOMC minutes, Llama 3 is the most accurate model, followed by GPT-4, FinBERT-FOMC, FinBERT, and VADER. The FinBERT-FOMC model, which was fine-tuned on central bank communication and utilizes a text simplification pipeline, performs better than the original FinBERT model. Llama 3 and GPT-4 outperform at the expense of large model sizes. Unlike GPT-4, FinBERT and FinBERT-FOMC are open-source and can be deployed on consumer-grade hardware. Llama 3 requires substantial hardware investments to deploy. The work thus finds that there is still a trade-off between model size and performance, and that the notion that “small is beautiful” can still hold for use cases where maximum accuracy is a lesser concern than inference speed and cost. Human performance is still significantly above all models, indicating that further improvements in language models and FOMC-specific prompting are possible. The labeled dataset for central bank communication we present in this paper is thus a challenging benchmark for future research.Type:conference paperJournal:Proceedings of the 5th ACM International Conference on AI in FinanceScopus© Citations 8