In recent years, there has been a surge in interest in evaluating the quality of chatbot conversation. We participated in the Dialogue Quality (DQ) and Nugget Detection (ND) subtasks in both Chinese and English. However, the majority of existing conventional approaches are based on the long short-term memory (LSTM) model. The paper suggests a method for assisting customers in resolving problems. The goal of this subtask is to automatically determine the status of dialogue sentences in a dialogue system's logs. On conversation tasks, we developed fine-tuning methodologies for the Transformer model. To evaluate and show the concept, we created a wide framework for testing and displaying the XLM-RoBERTa model's performance on conversational texts. Finally, the experimental findings of the two subtasks demonstrate the efficacy of our strategy. The experimental findings for the DialEval-2 task show that the suggested method's performance is reasonably equal to that of an LSTM-based baseline model. The main contribution of our study is that we suggested two crucial elements for increasing conversation quality and nugget identification subtasks in dialogue assessment, namely tokenization methods and finetuning procedures.
Hsiao et al. (Tue,) studied this question.