This paper provides an overview of the NTCIR-16 Dialogue Evaluation (DialEval-2) task. DialEval-2 is the successor of The NTCIR-15 DialEval-1 task and the NTCIR-14 Short Text Conversation STC- 3 task. DialEval-2 consists of two subtasks: the Dialogue Quality (DQ) subtask and the Nugget Detection (ND) subtask. Both of the subtasks are designed to aim automatical evaluation of customerhelpdesk dialogues. The DQ subtask requires our participants to estimate three kinds of quality score for each dialogue: task accomplishment, customer satisfaction, and dialogue effectiveness. The ND subtask is set as a classification task, where participants are asked to classify every turn of a dialogue to detect nugget turns. A nugget stands for a turn being helpful for problem solving in the dialogue. In this paper, we introduce the task definition, data collection, evaluation measures, and the official evaluation results on the runs from the participant teams.
Tao et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: