The results indicate that experts largely agree on metrics and that the consensus set is broad. This implies that health chatbot evaluation must be multifaceted to ensure acceptability.
No takes yet. Share an insight, caveat, or question.
Denecke et al. (2021) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: