Using Large Language Models to Detect Insufficient Effort Responding in Open-Ended Survey Questions
Journal
1st International Conference on Human-Computer Interaction in the Alps (AlpCHI 2026)
Type
conference paper
Date Issued
2026-03-01
Author(s)
Abstract
Careless responses pose a challenge for data quality in online survey research, a core method in human–computer interaction (HCI).
Open-ended answers can reveal such insufficient effort responding
(IER), but are costly to evaluate manually. I explore the use of two
large language model (LLM) pipelines to automate IER detection in
a dataset of 1,551 open-text responses: using open-source embedding models with standard classifiers, and using text-generation
labelling with GPT-4o-mini. Embedding-based models achieved
higher precision, but missed inattentive responses, whereas text
generation showed better accuracy yet tended to overpredict IER.
These patterns were explained by severe class imbalance, which was
identified as a typical feature of high-quality crowdsourced samples
and thus a central challenge for automated IER detection. I discuss
how such pipelines could be integrated into human-in-the-loop
workflows and emphasize the need for curated, openly available
datasets and improved model engineering to advance reliable IER
detection.
Open-ended answers can reveal such insufficient effort responding
(IER), but are costly to evaluate manually. I explore the use of two
large language model (LLM) pipelines to automate IER detection in
a dataset of 1,551 open-text responses: using open-source embedding models with standard classifiers, and using text-generation
labelling with GPT-4o-mini. Embedding-based models achieved
higher precision, but missed inattentive responses, whereas text
generation showed better accuracy yet tended to overpredict IER.
These patterns were explained by severe class imbalance, which was
identified as a typical feature of high-quality crowdsourced samples
and thus a central challenge for automated IER detection. I discuss
how such pipelines could be integrated into human-in-the-loop
workflows and emphasize the need for curated, openly available
datasets and improved model engineering to advance reliable IER
detection.
Language
German
Keywords
Sentence Embeddings
Large Language Models
Open-ended Responses
Insufficient Effort Responding
Data Quality
AI-assisted research methods, Crowdsourced Survey Research
Human-Centered AI
HSG Classification
contribution to scientific community
Refereed
Yes
Event Title
1st International Conference on Human-Computer Interaction in the Alps (AlpCHI 2026),
Event Location
Ascona, Switzerland
Event Date
01.03.2026-05.03.2026
Subject(s)
Division(s)
File(s)![Thumbnail Image]()
open.access
Name
alpchi2026-14.pdf
Size
1.12 MB
Format
Adobe PDF
Checksum (MD5)
4185b9159b95f701528c61b2da0e09ba