PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 4, 2026Work0 citations

Assessing the efficacy of LLMs in neuroscience: A performance analysis of ChatGPT and gemini on synaptic plasticity

View Full Paper
MAMelek AltunkayaEBErcan BabürECEmine Cihan

Key Points

  • The study evaluates the accuracy, quality, and readability of responses generated by LLMs regarding synaptic plasticity.
  • Selected LLMs ChatGPT-4 and Gemini 2.5 for evaluation.
  • Posed ten questions to each model and evaluated responses qualitatively using a 4-point Likert scale.
  • Analyzed readability of answers using the Flesch-Kincaid Grade Level test.
  • Both models provided generally accurate and acceptable information.
  • Gemini received higher median scores in some instances, but no statistically significant difference was observed overall.
  • Gemini's responses were longer and had a higher Flesch-Kincaid Grade Level, indicating a more technical structure.

Abstract

BackgroundSynaptic plasticity, which plays a critical role in fundamental neurological processes, is a complex subject to master. Therefore, large language models (LLMs) are increasingly being used to facilitate the learning of such complex topics. However, these models have limitations, including producing inaccurate information and failing to capture the nuances of scientific terminology.ObjectivesThis study aimed to evaluate the accuracy, quality and readability of LLM responses to questions on synaptic plasticity.MethodsThe widely used LLMs ChatGPT-4 and Gemini 2.5 were selected in the study. Ten questions were posed to each LLM, and the initial responses were recorded. Five neurophysiologists evaluated the responses qualitatively using a 4-point Likert scale. Readability level of the answers was analyzed using Flesch-Kincaid Grade Level test.ResultsIn the qualitative assessment, both models generally provided accurate and acceptable information. Within the limited scope of the questions analyzed, Gemini received higher median scores in certain instances; however, no statistically significant difference was observed between the two models across most of the question set. Linguistic analysis showed that Gemini's responses were longer and featured a higher Flesch-Kincaid Grade Level, suggesting a structure more aligned with academic or technical discourse.ConclusionFor the specific neuroscientific inquiries examined in this study, both LLMs demonstrated a high capacity for generating accurate content. While Gemini's responses exhibited a more technical linguistic profile, the findings are context-specific and further research is needed to determine if these trends persist across broader scientific domains and larger datasets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Altunkaya et al. (2026) studied this question.

synapsesocial.com/papers/6a48a36b89561a0c2d78d64fhttps://doi.org/10.1177/10519815261462135
Ask AI
Helpful
Bookmark
Share
View Full Paper