PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 16, 2025Scientometrics21 citationsOpen Access

Research evaluation with ChatGPT: is it age, country, length, or field biased?

View Full Paper
MTMike ThelwallZKZeyneb Kurt

Key Points

  • ChatGPT's quality evaluation of research shows a trend of increasing scores over time, not influenced by author nationality or title length.
  • The investigation involved a balanced dataset of 117,650 articles across 26 fields, published from 2003 to 2023, highlighting field variability.
  • Longer abstracts typically lead to higher quality scores, largely due to these articles being more substantial and published in top-tier journals.
  • Normalizing ChatGPT scores for field and year is essential for reliable research quality evaluations, addressing potential anomalies.

Abstract

Abstract Some research now suggests that ChatGPT can estimate the quality of journal articles from their titles and abstracts. This has created the possibility to use ChatGPT quality scores, perhaps alongside citation-based formulae, to support peer review for research evaluation. Nevertheless, ChatGPT’s internal processes are effectively opaque, despite it writing a report to support its scores, and its biases are unknown. This article investigates whether publication date and field are biasing factors. Based on submitting a monodisciplinary journal-balanced set of 117,650 articles from 26 fields published in the years 2003, 2008, 2013, 2018 and 2023 to ChatGPT 4o-mini, the results show that average scores increased over time, and this was not due to author nationality or title and abstract length changes. The results also varied substantially between fields, and first author countries. In addition, articles with longer abstracts tended to receive higher scores, mostly due to such articles tending to be better (e.g., more likely to be in higher impact journals) but also partly due to ChatGPT analysing more text. For the most accurate research quality evaluation results from ChatGPT, it is important to normalise ChatGPT scores for field and year and check for anomalies caused by sets of articles with short abstracts.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Thelwall et al. (2025) studied this question.

synapsesocial.com/papers/68a366a80a429f797332c9aahttps://doi.org/10.1007/s11192-025-05393-0
Ask AI
Helpful
Bookmark
Share
View Full Paper