Thomas Huber
Last Name
Huber
First name
Thomas
Email
thomas.huber@unisg.ch
Phone
+41 71 224 79 20
6 results
Now showing 1 - 6 of 6
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, ARTIST: A Learning Support System for Fostering Students’ Argumentative Writing Skills(2025-10); We present ARTIST, a learning support system that can help students assess their argumentative writing and provide automated, individual feedback, thus improving their writing performance. It analyzes student-written argumentative texts by identifying argument components and their relationships. The resulting argumentative discourse structure is displayed in an interactive interface. In that way, the ARTIST tool provides immediate and personalized visual feedback on the quality of students’ texts, supporting self-monitoring and reflection on how to improve their texts.Type:conference paper - Some of the metrics are blocked by yourconsent settings
Item type:Publication, CLEAR: A Comprehensive Linguistic Evaluation of Argument Rewriting by Large Language ModelsWhile LLMs have been extensively studied on general text generation tasks, there is less research on text rewriting, a task related to general text generation, and particularly on the behavior of models on this task. In this paper we analyze what changes LLMs make in a text rewriting setting. We focus specifically on argumentative texts and their improvement, a task named Argument Improvement (ArgImp). We present CLEAR: an evaluation pipeline consisting of 57 metrics mapped to four linguistic levels: lexical, syntactic, semantic and pragmatic. This pipeline is used to examine the qualities of LLM-rewritten arguments on a broad set of argumentation corpora and compare the behavior of different LLMs on this task and analyze the behavior of different LLMs on this task in terms of linguistic levels. By taking all four linguistic levels into consideration, we find that the models perform ArgImp by shortening the texts while simultaneously increasing average word length and merging sentences. Overall we note an increase in the persuasion and coherence dimensions.Type:conference paperJournal:Findings of the Association for Computational Linguistics: EMNLP 2025 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, LLMs meet Bloom’s Taxonomy: A Cognitive View on Large Language Model EvaluationsCurrent evaluation approaches for Large Language Models (LLMs) lack a structured approach that reflects the underlying cognitive abilities required for solving the tasks. This hinders a thorough understanding of the current level of LLM capabilities. For instance, it is widely accepted that LLMs perform well in terms of grammar, but it is unclear in what specific cognitive areas they excel or struggle in. This paper introduces a novel perspective on the evaluation of LLMs that leverages a hierarchical classification of tasks. Specifically, we explore the most widely used benchmarks for LLMs to systematically identify how well these existing evaluation methods cover the levels of Bloom’s Taxonomy, a hierarchical framework for categorizing cognitive skills. This comprehensive analysis allows us to identify strengths and weaknesses in current LLM assessment strategies in terms of cognitive abilities and suggest directions for both future benchmark development as well as highlight potential avenues for LLM research. Our findings reveal that LLMs generally perform better on the lower end of Bloom’s Taxonomy. Additionally, we find that there are significant gaps in the coverage of cognitive skills in the most commonly used benchmarks.Type:conference paper - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Let's discuss! Quality Dimensions and Annotated Datasets for Computational Argument Quality Assessment(Empirical Methods in Natural Language Processing, 2024-11-15); ; Research in the computational assessment of Argumentation Quality has gained popularity over the last ten years. Various quality dimensions have been explored through the creation of domain-specific datasets and assessment methods. We survey the related literature (211 publications and 32 datasets), while addressing potential overlaps and blurry boundaries to related domains. This paper provides a representative overview of the state of the art in Computational Argument Quality Assessment with a focus on annotated datasets. The aim of the survey is to identify research gaps and to aid future discussions and work in the domain.Type:conference paperScopus© Citations 5 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Enhancing Educational Dialogues: A Reinforcement Learning Approach for Generating AI Teacher Responses(Association for Computational Linguistics, 2023-07-13); ; Reinforcement Learning remains an underutilized method of training and fine-tuning Language Models (LMs) despite recent successes. This paper presents a simple approach of finetuning a language model with Reinforcement Learning to achieve competitive performance on the BEA 2023 Shared Task whose goal is to automatically generate teacher responses in educational dialogues. We utilized the novel NLPO algorithm that masks out tokens during generation to direct the model towards generations that maximize a reward function. We show results for both the t5-base model with 220 million parameters from the HuggingFace repository submitted to the leaderboard that, despite its comparatively small size, has achieved a good performance on both test and dev set, as well as GPT-2 with 124 million parameters. The presented results show that despite maximizing only one of the metrics used in the evaluation as a reward function our model scores highly in the other metrics as well.Type:conference paperJournal:Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023)Volume:Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023)Scopus© Citations 5 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Mine, Assess, Enhance: Dialogue-based Adaptive Argumentative Writing Support(2023); ; Ongoing SNF (Schweizerischer Nationalfonds) project. Goals: (1) automatically judge the quality of arguments, (2) develop an argumentative writing support system for students (3) that can debate with students about controversial topics and (4) give interactive feedback about the quality of students' arguments.Type:conference poster