Nick von Felten
Last Name
von Felten
First name
Nick
Email
nick.vonfelten@unisg.ch
Phone
071 224 7489
9 results
Now showing 1 - 9 of 9
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, Does Carbon Label Design Affect Flight Choices and Trustworthiness?Carbon emissions labels shown during flight booking may shift flight choices but also affect perceived trustworthiness of booking‐page information. This study tests how carbon label design influences flight choices, booking propensity and trustworthiness. In a preregistered online experiment (N = 939), participants completed a simulated flight booking task and were randomly assigned to one of five carbon label conditions: no label, numeric emissions only, traffic‐light colour coding, colour coding plus a contextual benchmark or colour coding plus the European Union Flight Emissions Label. Flight choice was analysed with multinomial logistic regression, booking propensity with binary logistic regression and trustworthiness with multigroup confirmatory factor analysis. Traffic‐light and enhanced carbon label formats increased selection of the lowest‐emission option relative to no label but did not robustly change booking propensity. These formats also reduced perceived trustworthiness of booking‐page information, even as they increased selection of the lowest‐emission option.Type:journal articleJournal:Business Strategy and the Environment - Some of the metrics are blocked by yourconsent settings
Item type:Publication, CHI’s Greatest Hits: Analyzing the 100 Most-Cited Papers in 43 Years of Research at ACM CHI(Association for Computing Machinery (ACM), 2025-01); ; ; Type:journal articleJournal:InteractionsVolume:32Issue:1DOI:10.1145/3704804Scopus© Citations 5 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Development and Psychometric Validation of a Positively Worded German Version of the System Usability Scale (SUS)(Informa UK Limited, 2024-12-11) ;Perrig Sebastian Andrea Caesar; ;Vollenwyder, BeatOpwis, KlausDespite its popularity, knowledge of the System Usability Scale's (SUS) psychometric quality in German is limited. This article developed a positively worded German SUS version and investigated two existing versions. In a preregistered first study (N ¼ 250), positive alternatives for negative items were evaluated using item analyses, exploratory factor analyses, and language expert suggestions. The two existing versions were also compared. A preregistered second study (N ¼ 877) involved participants interacting with 12 websites and completing the different SUS versions to validate the new positive-only and the original versions. Analyses included item analyses, confirmatory factor analyses, correlations with related scales, and error analyses. Findings indicated that the SUS-DE-Pos, a positive version based on Rummel, performed best. Researchers should use this version or Rummel's version if negative items are needed. Overall, this work provides validated German SUS versions, with and without negative items, addressing a critical gap for German-speaking HCI researchers and practitioners.Type:journal articleJournal:International Journal of Human–Computer InteractionScopus© Citations 7 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Using Large Language Models to Detect Insufficient Effort Responding in Open-Ended Survey Questions(2026-03-01)Careless responses pose a challenge for data quality in online survey research, a core method in human–computer interaction (HCI). Open-ended answers can reveal such insufficient effort responding (IER), but are costly to evaluate manually. I explore the use of two large language model (LLM) pipelines to automate IER detection in a dataset of 1,551 open-text responses: using open-source embedding models with standard classifiers, and using text-generation labelling with GPT-4o-mini. Embedding-based models achieved higher precision, but missed inattentive responses, whereas text generation showed better accuracy yet tended to overpredict IER. These patterns were explained by severe class imbalance, which was identified as a typical feature of high-quality crowdsourced samples and thus a central challenge for automated IER detection. I discuss how such pipelines could be integrated into human-in-the-loop workflows and emphasize the need for curated, openly available datasets and improved model engineering to advance reliable IER detection.Type:conference paperJournal:1st International Conference on Human-Computer Interaction in the Alps (AlpCHI 2026)Scopus© Citations 1 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, To Trust or Distrust AI: A Questionnaire Validation Study(ACM, 2025-06-23) ;Scharowski, Nicolas ;Perrig Sebastian Andrea Caesar; ;Wintersberger, PhilippOpwis, KlausDespite the importance of trust in human-AI interactions, researchers must often rely on questionnaires adapted from other fields, which lack validation in the AI context. Motivated by the need for reliable and valid measures, we investigated the psychometric quality of the most commonly used trust questionnaire in the context of AI by Jian, Bisantz, and Drury (2000). In a pre-registered online experiment (= 1485), participants observed interactions with both trustworthy and untrustworthy AI and rated their trust. Our results did not support the originally proposed single-factor structure for the questionnaire, but instead suggested a two-factor solution that distinguishes between trust and distrust. Based on our findings, we provide recommendations for future studies on how to use the questionnaire. Finally, we present arguments for considering trust and distrust as two distinct constructs, emphasizing the opportunities of considering and measuring both in human-AI interactions.Type:conference paperJournal:Proceedings of the 2025 ACM Conference on Fairness, Accountability, and TransparencyScopus© Citations 12 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, AI Washing Inflates Expected Performance but Not Interaction Outcomes: An AI Placebo Study Using Fitts' Law(Association of computing macheiner, 2026-06-27); ;Müller Luisa EllaExpectations about the support of artificial intelligence (AI) may influence interaction outcomes similar to placebos. Such expectations may result from AI washing, a practice of overstating a system's AI capabilities when actual functionality is limited. For example, some computer mice are marketed as "AI-assisted" despite lacking AI in core functions. In a withinsubjects study, 28 participants completed Fitts' Law tasks with a computer mouse under three conditions: no support, supposed predictive AI support, and supposed biosignal-enhanced AI support. Objective Fitts' Law performance indicators and subjective performance expectations, perceived workload, and perceived usability were measured. Compared to baseline, participants expected significantly improved performance in placebo conditions. However, these expectations did not translate into differences in objective or subjective assessments. This paper contributes evidence that AI washing inflates user expectations without altering actual interaction outcomes, highlighting a critical transparency issue. By exposing how deceptive AI marketing can shape user expectations, we underscore the need for accountability in AI product claims. Further, we establish Fitts' Law as a rigorous methodological lens for auditing AI-labelled input devices.Type:conference contributionJournal:ACM Conference on Fairness, Accountability, and Transparency (FAccT '26) - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Carbon Labels and Flight Choices: Insights from Experimental Research(2025-07-03); Type:conference contribution - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Beyond Isolation: Towards an Interactionist Perspective on Human Cognitive Bias and AI BiasIsolated perspectives have often paved the way for great scientific discoveries. However, many breakthroughs only emerged when moving away from singular views towards interactions. Discussions on Artificial Intelligence (AI) typically treat human and AI bias as distinct challenges, leaving their dynamic interplay and compounding potential largely unexplored. Recent research suggests that biased AI can amplify human cognitive biases, while well-calibrated systems might help mitigate them. In this position paper, I advocate for transcending beyond separate treatment of human and AI biases and instead focus on their interaction effects. I argue that a comprehensive framework, one that maps (compound human-AI) biases to mitigation strategies, is essential for understanding and protecting human cognition, and I outline concrete steps for its development. CCS CONCEPTS • Human-centered computing → HCI theory, concepts and models.Type:conference contributionJournal:CHI 2025: Tools for Thought: Research and Design for Understanding, Protecting, and Augmenting Human Cognition with Generative AI - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Bewertung der Datenqualität von Online-Umfragen: Eine explorative Untersuchung der Potenziale von Large-Language Models zur Detektion von unzureichendem Antwortaufwand(2025); Online-Umfragen sind in der empirischen Forschung als Instrument zur Datenerhebung weit verbreitet. Um Teilnehmer für Online-Umfragen zu erreichen, gewinnen Crowdsourcing-Plattformen wie Amazon Mechanical Turk oder Prolific zunehmend an Bedeutung. Solche Plattformen ermöglichen eine schnelle Datenerhebung und den Zugang zu oft diverseren Stichproben als herkömmliche Rekrutierungsmethoden. Es besteht jedoch die Gefahr, dass Teilnehmende hauptsächlich aus äusseren Anreizen wie beispielsweise der Vergütung teil-nehmen und die Fragen daher unachtsam beantworten. Dieses als „Insufficient Effort Res-ponding“ (IER) bekannte Phänomen kann die Untersuchungsergebnisse erheblich verzerren. Zur Erkennung und Kontrolle von IER wurden verschiedene Ansätze entwickelt. Beispiels-weise testen instruierte Antworten, ob Teilnehmende Anweisungen befolgen, während sta-tistische Verfahren Unstimmigkeiten in Antwortmustern aufdecken. Letztere erfordern je-doch häufig komplexe Analysen oder spezifische Datenvoraussetzungen. Besonders heraus-fordernd ist die Überprüfung offener Antworten, da diese meist manuell beurteilt werden müssen und stark vom jeweiligen Kontext abhängen, was besonders bei grösseren Stichpro-ben erheblichen Aufwand bedeutet. Large Language Models (LLMs) können dabei unter-stützen, doch ihre Zuverlässigkeit bei der Erkennung von IER ist noch unklar. In diesem Workshop werden wir gemeinsam mit den Teilnehmenden explorieren, inwieweit sich LLMs zur Identifikation von IER eignen. Anhand öffentlich zugänglicher Datensätze sowie grösseren Datensätze aus unserer eigenen Forschung wollen wir gemeinsam erarbei-ten, wie gut LLMs Vergleich zu menschlichen Einschätzungen und statistischen Kennwerten abschneiden. Der Workshop beginnt mit einer kurzen Einführung in die Thematik der Datenqualität in Online-Untersuchungen und den Herausforderungen von IER. Wir stellen gängige Techni-ken vor, beleuchten deren Vor- und Nachteile und entwickeln gemeinsam erste Hypothesen zu den Potenzialen und Schwächen von LLMs in diesen Aufgaben. Anschliessend erhalten die Teilnehmenden einen Datensatz mit offenen Antworten aus einer früheren Studie, die bereits manuell bewertet wurden. In Gruppenarbeit erfolgt zunächst eine eigene Bewertung, die als zusätzliche Baseline dient. Im Anschluss vergleichen wir die aktuellen Einschätzungen mit früheren Ergebnissen, den Bewertungen eines LLMs und statistischen Kennzahlen. Ziel ist es, systematische Übereinstimmungen oder Abweichungen herauszuarbeiten. Abschliessend folgt eine gemeinsame Reflexion darüber, unter welchen Bedingungen LLMs eine sinnvolle Ergänzung oder gar Alternative zu etablierten Verfahren darstellen könnten, wie sich mögliche Fehler systematisch erklären lassen und welche Prompt-Strategien zu bes-seren Ergebnissen führen könnten. Ziel des Workshops ist es, praxisnah zu ermitteln, ob und unter welchen Bedingungen LLMs die Identifikation von IER sinnvoll unterstützen können. Dabei werden systematisch Stärken, Schwächen und Optimierungsmöglichkeiten herausgearbeitet.Type:presentation