Randomized trial investigates how agent-generated feedback improves report quality and tester performance in crowdsourced testing.
Agentic AI is increasingly being integrated into software engineering workflows, and crowdsourced testing is no exception. In this setting, the large volume and uneven quality of submitted test reports still create a substantial review burden for developers. To help address this burden, we previously developed and validated a multi-agent assessment backbone based on the LLM-as-a-Judge paradigm. The backbone, which assesses reports along the dimensions of textuality, adequacy, and competitiveness, was shown to align well with human consensus while substantially reducing assessment effort. Yet reliable automated judging does not by itself show whether agent outputs can improve human work when embedded into the workflow. This paper studies that missing question in the context of crowdsourced testing. We investigate whether assessment-derived, actionable feedback can improve how testers revise reports, perform on later tasks, and transfer reporting practices across applications. To do so, we conducted a controlled four-stage human-subject study with 20 testers across three real-world applications. The results show that agent-generated feedback supports immediate improvements in revised reports, better first submissions on a new task after prior feedback exposure, and evidence of partial but meaningful transfer to a later application. A post-task questionnaire completed by 17 participants complements these artifact-based findings by suggesting that the feedback was generally understandable, acted upon in revision, and carried into later tasks, while also revealing remaining friction in specificity and execution. Overall, the study provides empirical evidence that, in the studied crowdsourced testing setting, assessment agents can serve not only as post-hoc judges but also as workflow-integrated feedback providers that support upstream report-quality improvement.
No takes yet. Share an insight, caveat, or question.
Wang et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: