Generative AI tools such as chatbots powered by large language models (LLMs) present a new modality for UXresearchers when evaluating system interfaces with users. While conversational UIs help users understandthese tools more quickly, they can present new challenges for researchers. These challenges include the way user prompting behaviors vary, how generative content changes between instances, and how to contextualize AI-generated content alongside other parts of the interface. How do you evaluate UI elements of generative AI chatbots? How does content generated in real time andconversational UIs change your approach and evaluation criteria? This presentation will cover a case studywhere Northwestern University Libraries evaluated an LLM-based research tool and compare and contrast theapproach with non-generative AI experiences. The tool evaluated — the homegrown generative AI research toolfor Northwestern’s Digital Collections (supported by an Institute of Museum and Library Services grant) —utilizes a conversational UI to help users learn and discover resources from a large catalog of materials. Attendees will learn approaches for researching and evaluating generative AI experiences, best practices fordeveloping a test plan, and how to translate findings and recommendations into actionable changes to theexperience.
Frank Sweis (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: