PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 18, 2025Systems6 citationsOpen Access

The AI Annotator: Large Language Models’ Potential in Scoring Sustainability Reports

View Full Paper
YWYue WuPHPeng HuDWDerek Wang

Key Points

  • GPT-4o achieved an average accuracy of 58%, highlighting potential but not yet ready to fully replace human experts.
  • The study noted hallucination rates in models; specific data processing methods like CoVe help improve performance.
  • Three LLMs were benchmarked against human expert scores, assessing key metrics such as accuracy and mean absolute error.
  • Effective verification strategies are essential to address risks associated with LLMs in sustainability reporting applications.

Abstract

To explore the potential of Large Language Models (LLMs) as AI Annotators in the domain of sustainability reporting, this study establishes a systematic evaluation methodology. We use the specific case of European football clubs, quantifying their sustainability reports based on the sport Positive matrix as a benchmark to compare the performance of three state-of-the-art models (i.e., GPT-4o, Qwen-2-72b-instruct, and Llama-3-70b-instruct) against human expert scores. The evaluation is benchmarked on dimensions including accuracy, mean absolute error (MAE), and hallucination rates. The results indicate that GPT-4o is the top performer, yet its average accuracy of approximately 56% shows it cannot fully replace human experts at present. The study also reveals significant issues with overconfidence and factual hallucinations in models like Qwen-2-72b-instructon. Critically, we find that by implementing further data processing, specifically a Chain-of-Verification (CoVe) self-correction method, GPT-4o’s initial hallucination rate is successfully reduced from 16% to 10%, while accuracy improved to 58%. In conclusion, while LLMs demonstrate immense potential to streamline and democratize sustainability ratings, inherent risks like hallucinations remain a primary obstacle. Adopting verification strategies such as CoVe is a crucial pathway to enhancing model reliability and advancing their effective application in this field.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wu et al. (2025) studied this question.

synapsesocial.com/papers/68f35bfc73f0a7d050f47d86https://doi.org/10.3390/systems13100899
Ask AI
Helpful
Bookmark
Share
View Full Paper