This lightning talk shares a case study from the University at Buffalo Libraries about using generative AI to create alternative text and long descriptions for digital collections. To assess AI-generated responses, librarians developed a scoring rubric with three criteria: factual accuracy and correctness, relevance and task completion, and clarity and communication quality. This approach allowed an objective review of three AI tools which were tested on 45 images. The rubric showed problems with the responses, including hallucinated details, omissions of key visuals from the photographs, and cultural insensitivity. The rubric also showed the importance of incorporating iterative changes to prompts and workflows. Lessons learned include the importance of metrics, human review, and collaboration. This talk offers suggestions for libraries and repositories seeking scalable approaches to accessibility compliance and a rubric for evaluating AI tools.
Cogley et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: