Abstract The rapid integration of large language models (LLMs) into scientific research raises a fundamental question for research evaluation: what is being assessed when core research activities become partially automated? Although prior studies consistently report efficiency gains, existing evaluation practices do not fully capture changes in verification behavior, epistemic reliability, and dependency risk. This paper draws on a systematic review of 90 peer-reviewed articles to examine how LLM-assisted research activities are evaluated across four workflow stages: search, screening, summarizing, and drafting. The results indicate a systematic asymmetry in evaluation practices. Measures of time savings and productivity are widely reported and typically positive, whereas verification practices, trust calibration, and reliability concerns are rarely specified or directly measured, particularly in high-transformation stages such as summarizing and drafting. These patterns suggest that prevailing evaluation approaches tend to emphasize efficiency while leaving epistemic safeguards insufficiently articulated. In response, this study develops a stage-sensitive evaluation framework that distinguishes between low- and high-transformation research activities and specifies proportionate verification checkpoints. By framing LLM-assisted research as a problem of evaluation design rather than merely productivity enhancement, the study contributes to research evaluation theory and practice.
Kim et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: