PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 5, 2026Assessing Writing1 citationsOpen Access

Assessing fairness in finetuned scoring models with demographically restricted training data

View Full Paper
LHLangdon HolmesWMWesley MorrisSCScott A. Crossley

Key Points

  • The aim is to assess how demographic composition in training data affects fairness in automated essay scoring systems.
  • Analyzed the PERSUADE corpus of 26,000 student essays.
  • Developed four variants of a Longformer-based automated essay scoring system.
  • Created demographically restricted training sets from specific racial/ethnic groups.
  • Conducted analysis comparing systems trained on balanced vs. restricted data.
  • Demographic factors accounted for 12.5% of variance in human essay scores.
  • LLM-AES trained on restricted data showed systematic biases (R² = 0.043).
  • Balanced training data minimized bias in LLM-AES performance across demographic groups.

Abstract

The increasing adoption of automated essay scoring (AES) in high-stakes educational contexts necessitates careful examination of potential biases within the systems. This study investigates how the demographic composition of training data influences fairness in AES systems developed from finetuned large language models (LLMs). Using the PERSUADE corpus of 26,000 student essays, we conducted a systematic analysis using demographically restricted training sets to isolate the impact of training data demographics on LLM-AES performance. Each demographically restricted training set comprised essays written by one racial/ethnic group. Four variants of a Longformer-based AES were developed: one trained on demographically balanced data and three trained on demographically restricted datasets. An initial analysis of the human ratings indicated that demographic factors significantly predict human essay scores (marginal R² = 0.125), a pattern that is paralleled in national writing assessment data. LLM-AES systems trained on demographically restricted data exhibited small systematic biases (marginal R² = 0.043). However, the LLM trained on balanced data showed minimal demographic bias, suggesting that representative training data can effectively prevent amplification of demographic disparities beyond those present in human ratings. These results highlight both the importance and limitations of training data diversity in achieving fair assessment outcomes. • 12.5% of variance in human essay ratings was explained by demographics. • We construct demographically restricted training sets to isolate bias. • Balanced training data minimized LLM-AES bias across demographic groups. • LLM-AES trained on demographically restricted data showed more bias.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Holmes et al. (2026) studied this question.

synapsesocial.com/papers/69a91df9d6127c7a504c169dhttps://doi.org/10.1016/j.asw.2026.101032
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Discriminated by an algorithm: a systematic review of discrimination and fairness by algorithmic decision-making in the context of HR recruitment and HR development2020 · 471 citations
  2. 2Fitting Linear Mixed-Effects Models Using lme42015 · 88,543 citations
  3. 3Effectiveness of large language models in automated evaluation of argumentative essays: finetuning vs. zero-shot prompting2024 · 36 citations
  4. 4For a Greater Good: Bias Analysis in Writing Assessment2019 · 26 citations
  5. 5An Empirical Analysis of BERT Embedding for Automated Essay Scoring2020 · 59 citations