PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 15, 2026ACM Transactions on Interactive Intelligent Systems0 citations

AI Can See What You Can't See: How LLM-Agents Complement Human-Based Gender-Inclusive Usability Testing

View Full Paper
JGJoy GeuenichChemnitz University of TechnologyFJFrank JoublinHonda (Japan)ACAntonello CeravolaHonda (Japan)

Key Points

  • This research aims to assess the effectiveness of LLM-agent systems in complementing human-based gender-inclusive usability testing.
  • Evaluation of LLM-agent system in GenderMag persona-based usability testing
  • Comparison of agent performance against traditional human-led evaluations
  • Analysis of usability issues from four web interfaces, including three generic and one intentionally flawed interface
  • LLM-agents and humans detect overlapping usability issues in generic interfaces
  • Agents assign higher severity and relevance ratings compared to human evaluations
  • Overlap in detecting gender-related usability issues is low between LLM-agents and humans

Abstract

Inclusive usability testing, such as the GenderMag method, wants to identify gender-related usability problems in digital interfaces. Large Language Models (LLMs) have been used by usability engineers in usability evaluations but their contribution is still underexplored, especially regarding inclusive usability testing. Research has shown that GenderMag workshops can produce valuable insights but are resource intense and might show effects from the evaluators’ ability to embody personas with a different cognitive style. Therefore, we need to assess if LLM-agent based testing can aid human-based evaluations. This study evaluates an LLM-agent system for GenderMag persona-based usability testing, and compares its performance to traditional human-led evaluations. The agent system integrates GenderMag persona facets into three LLM-agents, which analyze usability issues of four web interfaces, three generic and one intentionally flawed interface with gender-related usability issues. We quantitatively and qualitatively compare the types, severity, and relevance of usability problems LLM-agents identified to those produced in three GenderMag workshops involving nine participants. Findings show a broad overlap in detecting usability issues of humans and LLM-agents that are not gender specific via the generic interfaces. Agents thereby consistently assign significantly higher severity and relevance ratings. For the intentionally flawed interface, humans and LLM-agents assign similar ratings but the overlap between human and agent gender-related usability issues was low, with each missing issues the other caught. The agent system is an efficient tool for gender-related usability evaluation that humans may overlook, thereby expanding the coverage of evaluation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Geuenich et al. (2026) studied this question.

synapsesocial.com/papers/69b5ff6e83145bc643d1bf9bhttps://doi.org/10.1145/3801978
Ask AI
Helpful
Bookmark
Share
View Full Paper