Julius Schulte
Now showing 1 - 2 of 2
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, sparsesurv: a Python package for fitting sparse survival models via knowledge distillation(2024-09-18) ;Wissel, David ;Janakarajan, Nikita; ;Rowson, DanielYuan, XintianMotivation: Sparse survival models are statistical models that select a subset of predictor variables while modeling the time until an event occurs, which can subsequently help interpretability and transportability. The subset of important features is often obtained with regularized models, such as the Cox Proportional Hazards model with Lasso regularization, which limit the number of non-zero coefficients. However, such models can be sensitive to the choice of regularization hyperparameter. Results: In this work, we develop a software package and demonstrate how knowledge distillation, a powerful technique in machine learning that aims to transfer knowledge from a complex teacher model to a simpler student model, can be leveraged to learn sparse survival models while mitigating this challenge. For this purpose, we present sparsesurv, a Python package that contains a set of teacher-student model pairs, including the semi-parametric accelerated failure time and the extended hazards models as teachers, which currently do not have Python implementations. It also contains in-house survival function estimators, removing the need for external packages. Sparsesurv is validated against R-based Elastic Net regularized linear Cox proportional hazards models as implemented in the commonly used glmnet package. Our results reveal that knowledge distillation-based approaches achieve competitive discriminative performance relative to glmnet across the regularization path while making the choice of the regularization hyperparameter significantly easier. All of these features, combined with a sklearn-like API, make sparsesurv an easy-to-use Python package that enables survival analysis for high-dimensional datasets through fitting sparse survival models via knowledge distillation.Type:case review (law)Scopus© Citations 6 - Some of the metrics are blocked by yourconsent settings
Item type:Research Project, Monetization Risks in AI Systems: A Bing Copilot case study Submission to the European Commission(2024-04) ;Strauss, Ilan ;O'reilly, Tim ;Mazzucato, MarianaWe conduct a descriptive analysis and simple experiments on Microsoft's AI Copilot LLM-a service that provides answers in a comparable manner to a Search results engine. We interrogate the nature of Copilot's generated results, characterized by notable use of advertising to monetize user attention already, and compare it to Microsoft's traditional Bing search engine results on a series of e-commerce product queries. We find that the underlying sources from which the generated results in Copilot ostensibly come from (the footnote links) differs substantially from the websites listed in Bing's traditional search results, but for no apparent reason. A lack of algorithmic transparency and explainability in Bing's Copilot opens up room for providing suboptimal text results in return for paid placement (advertising). We highlight that use of retrieval-augmented generation (RAG) by LLMs, to draw on a wider body of relevant material, creates significant monetization opportunities for LLMs. To avoid past mistakes of digital platforms, advertising output needs to be clearly demarcated, limited in scope, and regulated for quality. We also recommend that LLM generated results should