PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 23, 2026Information and Software Technology2 citationsOpen Access

Meta-Fair: AI-assisted fairness testing of large language models

View Full Paper
MRMiguel Romero-ArjonaJPJosé A. ParejoJAJuan J. Alonso

Key Points

  • The study aims to create an automated testing method for fairness in large language models to reduce bias assessment challenges.
  • Developed Meta-Fair, an automated fairness testing approach for LLMs.
  • Utilized metamorphic testing to reveal bias by altering input prompts.
  • Generated diverse test cases using LLMs for effective evaluation.
  • Achieved an average precision of 92% in uncovering bias.
  • Revealed biased behavior in 29% of tests executed.
  • Best-performing models reached F1-scores of up to 0.79.

Abstract

Fairness—the absence of unjustified bias—is a core principle in the development of Artificial Intelligence (AI) systems, yet it remains difficult to assess and enforce. Current approaches to fairness testing in large language models (LLMs) often rely on manual evaluation, fixed templates, deterministic heuristics, and curated datasets, making them resource-intensive and difficult to scale. This work aims to lay the groundwork for a novel, automated method for testing fairness in LLMs, reducing the dependence on domain-specific resources and broadening the applicability of current approaches. Our approach, Meta-Fair, is based on two key ideas. First, we adopt metamorphic testing to uncover bias by examining how model outputs vary in response to controlled modifications of input prompts, defined by metamorphic relations (MRs). Second, we propose exploiting the potential of LLMs for both test case generation and output evaluation, leveraging their capability to generate diverse inputs and classify outputs effectively. The proposal is complemented by three open-source tools supporting LLM-driven generation, execution, and evaluation of test cases. We report the findings of several experiments involving 12 pre-trained LLMs, 14 MRs, 5 bias dimensions, and 7.9K automatically generated test cases. The results show that Meta-Fair is effective in uncovering bias in LLMs, achieving an average precision of 92% and revealing biased behaviour in 29% of executions. Additionally, LLMs prove to be reliable and consistent evaluators, with the best-performing models achieving F1-scores of up to 0.79. Although non-determinism affects consistency, these effects can be mitigated through careful MR design. This work highlights the feasibility and potential of integrating metamorphic testing with LLM-driven test generation and assessment. While challenges remain to ensure broader applicability, the results indicate a promising path towards an unprecedented level of automation in LLM testing.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Romero-Arjona et al. (2026) studied this question.

synapsesocial.com/papers/699bee1c1c6c6bad5397fdc1https://doi.org/10.1016/j.infsof.2026.108075
Ask AI
Helpful
Bookmark
Share
View Full Paper