PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 6, 2026Circulation0 citations

Abstract WE488: A Comparison of English and Spanish AI Chatbot Responses to Online Cardiovascular Questions: Differences in Response Length, Readability, and Guidance

View Full Paper
YBYoel BenarrochJGJason GusdorfBeth Israel Deaconess Medical CenterNINicolás IsazaBeth Israel Deaconess Medical Center

Key Result

ChatGPT offered further medical guidance in 10/12 English and 11/12 Spanish responses compared to 0/12 for Google, while readability for both exceeded the recommended 6th grade level.

Key Points

  • The study aims to compare AI chatbot responses to cardiovascular questions in English and Spanish, focusing on response length, readability, and medical guidance.
  • Analyzed 48 responses to twelve cardiovascular questions using ChatGPT and Google's AI in English and Spanish.
  • Descriptive statistics and paired t-tests assessed differences in response length, sentence length, and readability.
  • Readability measured using Flesch-Kincaid Grade Level for English and Fernández-Huerta index for Spanish.
  • ChatGPT generated longer answers than Google in both languages, with significant differences in English.
  • Readability exceeded the recommended 6th grade level across both platforms, indicating challenges for general understanding.
  • ChatGPT provided medical guidance in nearly all responses, while Google offered none, with disclaimers present in all Google responses.

Study Design

Type

Cross-Sectional (n=48)

Structured PICO

How do English and Spanish AI chatbot responses to online cardiovascular questions differ in response length, readability, and guidance?

P
Population
12 common cardiovascular questions queried in English and Spanish
I
Intervention
ChatGPT (4o) and Google's AI summaries
C
Comparator
Comparison between English and Spanish languages, and between ChatGPT and Google platforms
O
Outcome
Response length, sentence length, disclaimers, clinical guidance elements, and readability (Flesch-Kincaid Grade Level for English and Fernández-Huerta index for Spanish)

AI chatbot responses to cardiovascular questions vary by language and platform, with readability consistently exceeding the recommended 6th-grade level, highlighting the need for linguistic tailoring to improve health equity.

Limitations

  • small sample size
  • evaluation did not analyze quality or meaning

Abstract

Introduction: Large language models (LLMs) such as ChatGPT (OpenAI) and Gemini (Google) have become a dominant source of information. These tools contain nuanced and personalized information; they are also prone to hallucinations, societal biases, and sycophancy. While studies suggest LLMs are capable of providing high quality health information, less is known how this information varies by language and reading level. Because cardiovascular disease remains a leading cause of morbidity and mortality, and access to understandable health information is critical to promoting health equity, we studied the abilities of these models in providing approachable health information. Methods: Twelve common cardiovascular questions were queried in English and Spanish using two popular sources for patient LLM usage, the ChatGPT (4o) and Google’s AI summaries (Table 1). Response length, sentence length, disclaimers, and clinical guidance elements were recorded and analyzed using descriptive statistics and paired t-tests were performed to assess between-language differences. Readability was assessed using validated tools - Flesch-Kincaid Grade Level for English and Fernández-Huerta index for Spanish. Results: Across 48 responses, ChatGPT produced longer answers than Google in both languages. Mean sentence length was significantly longer in ChatGPT English compared with ChatGPT Spanish. Conversely, Google responses were slightly longer in Spanish. ChatGPT offered further medical guidance in 10 of 12 English and 11 of 12 Spanish responses, whereas Google provided none. Medical disclaimers were present in all Google responses and in approximately half of ChatGPT responses. Readability was above the recommended 6th grade level across languages and platforms. Conclusion: In conclusion, AI response length, readability, and guidance to cardiovascular questions varied by language and platform. AI companies should target health-related outputs to reading levels that are appropriate to users. Our finding of decreased medical disclaimers is consistent with other work in this field. Further research should evaluate the quality of responses in different languages with validated metrics. There are major limitations to our study, most notably that it had a small sample size, and our evaluation did not analyze quality or meaning. Greater attention to readability and linguistic tailoring of AI outputs may enhance equitable access to understandable online health information.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Benarroch et al. (2026) conducted a cross-sectional in Cardiovascular questions (n=48). ChatGPT (4o) vs. Google's AI summaries was evaluated on Response length, sentence length, disclaimers, clinical guidance elements, and readability. ChatGPT offered further medical guidance in 10/12 English and 11/12 Spanish responses compared to 0/12 for Google, while readability for both exceeded the recommended 6th grade level.

synapsesocial.com/papers/69fadad703f892aec9b1e850https://doi.org/10.1161/cir.153.suppl_1.we488
Ask AI
Helpful
Bookmark
Share
View Full Paper