PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 8, 20260 citationsOpen Access

Does GenAI Make Usability Testing Obsolete?

APAli Ebrahimi PourasadWMWalid Maalej

Key Points

  • The research aims to evaluate the effectiveness of UX-LLM in identifying usability issues in iOS apps.
  • Presented at the ICSE 2025 conference and awarded ACM SIGSOFT Distinguished Paper Award.
  • Predicted usability issues in two open-source iOS apps using UX-LLM.
  • Conducted traditional usability testing and expert review for comparison.
  • Engaged a focus group from an app development team to assess UX-LLM's effectiveness.
  • UX-LLM showed precision between 0.61 and 0.66 and recall between 0.35 and 0.38.
  • The tool helped identify previously unknown usability issues in the development team's app.
  • Concerns were raised regarding UX-LLM's integration into existing workflows.

Abstract

This paper was presented at the 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE) and received the ACM SIGSOFT Distinguished Paper Award. Ensuring usability is crucial for the success of mobile apps. Usability issues can compromise user experience and negatively impact the perceived app quality. This paper presents UX-LLM, a novel tool powered by a Large Vision-Language Model that predicts usability issues in iOS apps. To evaluate the performance of UX-LLM, we predicted usability issues in two open-source apps of a medium complexity and asked two usability experts to assess the predictions. We also performed traditional usability testing and expert review for both apps and compared the results to those of UX-LLM. UX-LLM demonstrated precision ranging from 0.61 and 0.66 and recall between 0.35 and 0.38, indicating its ability to identify valid usability issues, yet failing to capture the majority of issues. Finally, we conducted a focus group with an app development team of a capstone project developing a transit app for visually impaired persons. The focus group expressed positive perceptions of UX-LLM as it identified unknown usability issues in their app. However, they also raised concerns about its integration into the development workflow, suggesting potential improvements. Our results show that UX-LLM cannot fully replace traditional usability evaluation methods but serves as a valuable supplement particularly for small teams with limited resources, to identify issues in less common user paths, due to its ability to inspect the source code.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Pourasad et al. (2026) studied this question.

synapsesocial.com/papers/698827570fc35cd7a88460fahttps://doi.org/10.18420/se2026_12
Ask AI
Helpful
Bookmark
Share
View Full Paper