Computational study demonstrates adversarial dialogue training reduces toxic responses in conversational agents, indicating improved safety during interaction.
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, Emily Dinan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
No takes yet. Share an insight, caveat, or question.
Xu et al. (2021) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: