PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 24, 2026Reference Services Review2 citations

Deploying and evaluating a conversational agent using LLMs for academic library reference

View Full Paper
MFMegan FitzgibbonsFBFrancisco BerrizbeitiaJCJoshua Chalifour

Key Points

  • To implement and evaluate a RAG-based chatbot for answering reference questions in libraries using different LLMs.
  • Developed a RAG-based GenAI system for reference queries.
  • Implemented a two-step approach: document retrieval and context-aware response generation.
  • Tested the chatbot with fourteen common reference questions.
  • Created and piloted an evaluation rubric assessing accuracy, groundedness, and completeness.
  • Chatbot effectively restricted responses to the relevant knowledge base.
  • Evaluation rubric successfully highlighted strengths and weaknesses of different LLMs.
  • Evaluators produced similar scores, despite subjectivity, with varied results in elicitation.

Abstract

Purpose This study has two aims. First, we sought to implement a RAG-based GenAI system capable of answering reference questions. Second, we aimed to develop an evaluation protocol to assess the chatbot by means of comparing implementations that use three different LLMs. An evaluation rubric was piloted to gauge its viability as an assessment tool. Design/methodology/approach The RAG-based chatbot uses a two-step approach. First, in response to a query, the system retrieves relevant documents from a knowledge base. Each document is vectorized and matched by relevance. Second, retrieved data is combined with an LLM's generative capabilities to produce a context-aware response. Fourteen common questions representing different areas of the knowledge base were tested with the chatbot versions. The research team developed and then used an evaluation rubric to score the chatbots' responses according to: accuracy, groundedness, elicitation, completeness and further assistance. The rubric was also evaluated by calculating the standard deviation among reviewers' scores. Findings The RAG implementations were largely successful in restricting the chatbot's responses to the knowledge base. The evaluation rubric was effective for assessing the models, highlighting each's strengths and weaknesses. Despite the evaluation being subjective, the evaluators gave similar scores, with the greatest variation in the elicitation dimension. Originality/value This study offers a technical description of a practical way to implement a RAG-based chatbot in a library setting as well as a protocol for evaluating such chatbots in multiple dimensions that hasn't been discussed in previous literature.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fitzgibbons et al. (2026) studied this question.

synapsesocial.com/papers/697460acbb9d90c67120a851https://doi.org/10.1108/rsr-05-2025-0030
Ask AI
Helpful
Bookmark
Share
View Full Paper