Comparison shows Elicit AI has low sensitivity but high precision in literature searches, suggesting it may assist researchers.
Background Elicit AI aims to simplify and accelerate the systematic review process without compromising accuracy. However, research on Elicit's performance is limited. Objectives To determine whether Elicit AI is a viable tool for systematic literature searches and title/abstract screening stages. Methods We compared the included studies in four evidence syntheses to those identified using the subscription‐based version of Elicit Pro in Review mode. We calculated sensitivity, precision and observed patterns in the performance of Elicit. Results The sensitivity of Elicit was poor, averaging 39.5% (25.5–69.2%) compared to 94.5% (91.1–98.0%) in the original reviews. However, Elicit identified some included studies not identified by the original searches and had an average of 41.8% precision (35.6–46.2%) which was higher than the 7.55% average of the original reviews (0.65–14.7%). Discussion At the time of this evaluation, Elicit did not search with high enough sensitivity to replace traditional literature searching. However, the high precision of searching in Elicit could prove useful for preliminary searches, and the unique studies identified mean that Elicit can be used by researchers as a useful adjunct. Conclusion Whilst Elicit searches are currently not sensitive enough to replace traditional searching, Elicit is continually improving, and further evaluations should be undertaken as new developments take place.
No takes yet. Share an insight, caveat, or question.
Lau et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: