Optimizing scholarly literature search: a comparative study of relevance and evidence quality across AI-powered, semantic, and traditional search engines
Comparative study evaluates AI-powered and traditional search engines for literature retrieval, suggesting AI tools enhance relevance and precision.
Key Points
This study evaluates the performance of AI-powered search engines and traditional platforms in retrieving relevant literature in tissue engineering.
Used PICO queries for Google Scholar, natural language for Semantic Scholar, and targeted prompting for Consensus.
Assessed the top 50 results for precision, false drop rates, study designs, and citation metrics.
Judged relevance by experts using fuzzy logic and evaluated evidence quality against medical hierarchies.
Consensus achieved 86.0% precision with no false drops, outperforming Semantic Scholar (75.5% precision, 10.0% false drop rate) and Google Scholar (72.0% precision, 10.0% false drop rate).
Consensus retrieved articles with an average of 341 citations, showing a moderate positive correlation (rs = 0.65) between relevance and citation frequency.
Minimal duplicate retrievals were observed (8%) with Semantic Scholar yielding newer publications but no significant differences in study types.