PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 18, 2026Information1 citationsOpen Access

Machines Prefer Humans as Literary Authors: Evaluating Authorship Bias in Large Language Models

View Full Paper
MRMarco RospocherUniversity of VeronaMSMassimo SalgaroUniversity of VeronaSRSimone ReboraUniversity of Verona

Key Points

  • The aim is to investigate how large language models assess literary quality when authorship is framed as human, AI, or a collaboration.
  • Used an experimental design with a questionnaire to prompt four instruction-tuned LLMs.
  • Evaluated three short stories generated by ChatGPT 4 in the style of Roald Dahl.
  • Collected 3600 responses across different authorship framings to assess model judgments.
  • Identical stories received higher ratings when framed as human-authored or co-authored compared to AI-authored.
  • A robust negative bias against AI authorship was observed in model evaluations.
  • Distinct evaluative profiles were identified among different LLMs.

Abstract

Automata and artificial intelligence (AI) have long occupied a central place in cultural and artistic imagination, and the recent proliferation of AI-generated artworks has intensified debates about authorship, creativity, and human agency. Empirical studies show that audiences often perceive AI-generated works as less authentic or emotionally resonant than human creations, with authorship attribution strongly shaping esthetic judgments. Yet little attention has been paid to how AI systems themselves evaluate creative authorship. This study investigates how large language models (LLMs) evaluate literary quality under different framings of authorship—Human, AI, or Human+AI collaboration. Using a questionnaire-based experimental design, we prompted four instruction-tuned LLMs (ChatGPT 4, Gemini 2, Gemma 3, and LLaMA 3) to read and assess three short stories in Italian, originally generated by ChatGPT 4 in the narrative style of Roald Dahl. For each story × authorship condition × model combination, we collected 100 questionnaire completions, yielding 3600 responses in total. Across esthetic, literary, and inclusiveness dimensions, the stated authorship systematically conditioned model judgments: identical stories were consistently rated more favorably when framed as human-authored or human–AI co-authored than when labeled as AI-authored, revealing a robust negative bias toward AI authorship. Model-specific analyses further indicate distinctive evaluative profiles and inclusiveness thresholds across proprietary and open-source systems. Our findings extend research on attribution bias into the computational realm, showing that LLM-based evaluations reproduce human-like assumptions about creative agency and literary value. We publicly release all materials to facilitate transparency and future comparative work on AI-mediated literary evaluation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rospocher et al. (2026) studied this question.

synapsesocial.com/papers/696c79cde45ebfc9113cd45ehttps://doi.org/10.3390/info17010095
Ask AI
Helpful
Bookmark
Share
View Full Paper