PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 1, 20260 citationsOpen Access

NKUST at the NTCIR-16 DialEval-2 Task

View Full Paper
TCTao-Hsing ChangJCJian-He Chen

Key Points

  • The study aims to assess the efficiency of different models in evaluating dialogue quality generated by chatbots.
  • Developed three models for dialogue evaluation: Pegasus for summarization, Bi-LSTM for structural adjustments, and a multi-agent model for diverse evaluations.
  • Conducted nugget detection as part of the evaluation process.
  • Performed experimental tests comparing effectiveness of different models.
  • Identified that some methods require refined experimental designs for effective application.
  • Highlighted the necessity of tuning model parameters for better dialogue evaluation.

Abstract

It is important to evaluate the quality of dialogues generated by chatbots. Most previous automatic evaluation methods have been based on models (e.g., LSTM ) that are capable of processing time series. This study presents three models for dialogue quality and two nugget detection subtasks, respectively. Specifically, the first model uses a Pegasus model that can transform dialogues into short summaries; the second model uses a Bi-LSTM that merely adjusts the internal model structure; and the third model is a multi-agent model simulating situations in which multiple annotators generate different evaluation results for the same text. The experimental results show that certain opinions may need to be corroborated by more refined experimental design and the testing of more model parameters before they are applicable to this issue.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chang et al. (2022) studied this question.

synapsesocial.com/papers/69cd7e935652765b073a996fhttps://doi.org/10.20736/0002002262
Ask AI
Helpful
Bookmark
Share
View Full Paper