It is important to evaluate the quality of dialogues generated by chatbots. Most previous automatic evaluation methods have been based on models (e.g., LSTM ) that are capable of processing time series. This study presents three models for dialogue quality and two nugget detection subtasks, respectively. Specifically, the first model uses a Pegasus model that can transform dialogues into short summaries; the second model uses a Bi-LSTM that merely adjusts the internal model structure; and the third model is a multi-agent model simulating situations in which multiple annotators generate different evaluation results for the same text. The experimental results show that certain opinions may need to be corroborated by more refined experimental design and the testing of more model parameters before they are applicable to this issue.
Chang et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: