PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 5, 2026Cancer Research0 citations

Abstract 2748: Arkangel AI, OpenEvidence, ChatGPT, Medisearch: Are they objectively up to medical standards? A real-life assessment of LLMs in healthcare.

View Full Paper
N-Natalia Castano -VillegasMVMaria Camila VillaKMKatherine Monsalve

Key Result

ArkangelAI-Deep achieved the highest clinician satisfaction (92.9%) in answering clinical vignettes, followed by OpenEvidence (83.6%), ChatGPT-Deep (80.5%), and Medisearch (71.1%).

Key Points

  • To evaluate the validity and safety of multiple large language models in real healthcare scenarios.
  • Developed four fictitious clinical vignettes tested in four conversational agents.
  • Responses evaluated by independent clinicians using an eight-criterion Likert scale.
  • Criteria included correctness, consensus, bias, and patient safety.
  • Measured response times to assess usability.
  • ArkangelAI-Deep scored 92.9% satisfaction, outperforming the others.
  • Dissatisfaction was highest for the real-source-of-references criteria in GPT models.
  • All models performed well in correctness and consensus agreement.
  • Medisearch had the fastest response time at 18 seconds, while GPT-Deep took 13 minutes.

Structured PICO

Do Arkangel AI, OpenEvidence, ChatGPT, and Medisearch provide satisfactory and safe responses to clinical vignettes?

P
Population
Four fictitious clinical vignettes developed by independent specialists, each including four questions
I
Intervention
Four conversational agents: ArkangelAI, OpenEvidence, ChatGPT, and Medisearch
C
Comparator
Compared against each other
O
Outcome
Clinician satisfaction evaluated using an eight-criterion Likert scale (correctness, consensus, bias, standard of care, updated information, patient safety, real sources in references, and context-awareness)

ArkangelAI-Deep and OpenEvidence provided the highest clinician satisfaction for answering clinical vignettes, though trade-offs exist between depth and response time.

Abstract

Abstract Background: Large language models (LLMs) are increasingly used in healthcare, but standardized benchmarks fail to capture their validity and safety in real-world scenarios. Evaluating their quality is critical for safe integration into practice. Methods: Four fictitious clinical vignettes were developed by independent specialists and tested in four conversational agents: ArkangelAI, OpenEvidence, ChatGPT, and Medisearch. Each vignette included four questions. Responses were evaluated by four external clinicians using an eight-criterion Likert scale: 1-2 = dissatisfaction, 3 = neutral, 4-5 = satisfaction, 6 = not applicable. The criteria considered correctness, consensus, bias, standard of care, updated information, patient safety, real sources in references, and context-awareness. Response times were measured with medians/interquartile ranges (IQR). Results were reported as frequencies. Hypothesis tests were applied (α= 0.05). Results: There were 128 Question-answer pairs. ArkangelAI-Deep had the highest satisfaction (92.9%), followed by OpenEvidence (83.6%), ChatGPT-Deep (80.5%), and Medisearch (71.1%). Most dissatisfaction was for the real-source-of-references criteria: GPT-Personalized 75%, GPT-Regular 97%. Conversely, ArkangelAI-Deep, ChatGPT-Deep, and OpenEvidence obtained 100% satisfaction. All performed well in correctness and agreement with the consensus. ChatGPT was the lowest-scoring in non-biased answers. The safest for patients was GPT-Personalized, followed by Arkagel AI-Deep. Medisearch had the fastest response time (18 s), while GPT-Deep (13 min) and ArkangelAI-Deep (7.4 min) were slowest, showing a trade-off between depth and usability. Conclusions: ArkangelAI-Deep and OpenEvidence consistently outperformed others, while Medisearch and GPT-Regular had significant limitations. These results underscore the need for standardized frameworks to ensure safe use of LLMs in healthcare. Citation Format: Natalia Castano -Villegas, Maria Camila Villa, Katherine Monsalve, Isabella Llano, Laura Velásquez, Jose Zea. Arkangel AI, OpenEvidence, ChatGPT, Medisearch: Are they objectively up to medical standards? A real-life assessment of LLMs in healthcare abstract. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 2748.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

-Villegas et al. (2026) studied this question. Large language models (ArkangelAI, OpenEvidence, ChatGPT, Medisearch) was evaluated on Clinician satisfaction evaluated using an eight-criterion Likert scale. ArkangelAI-Deep achieved the highest clinician satisfaction (92.9%) in answering clinical vignettes, followed by OpenEvidence (83.6%), ChatGPT-Deep (80.5%), and Medisearch (71.1%).

synapsesocial.com/papers/69d1fdf7a79560c99a0a4594https://doi.org/10.1158/1538-7445.am2026-2748
Ask AI
Helpful
Bookmark
Share
View Full Paper