ABSTRACT Recent advances in text‐to‐image generation have enabled generative models to produce realistic visuals from textual descriptions, transforming creative workflows in domains like fashion. However, these systems may encode and reproduce societal biases, particularly in gender representation. This study proposes a systematic and interpretable methodology for analysing gender bias in text‐to‐image generation models. The framework is model‐agnostic and applicable to any generative system, combining quantitative evaluation with interpretable analysis. Our proposal is structured in two main components: (1) the creation of a controlled corpus; and (2) the evaluation of the generated outputs through manual annotations and three complementary analyses: (i) model neutrality, assessing gender balance under neutral prompts; (ii) model accuracy, measuring adherence to gendered instructions; and (iii) interpretable pattern discovery, uncovering the semantic attributes that drive gendered generations via decision tree modelling. Concretely, we focus on the fashion domain and employ Stable Diffusion as a representative state‐of‐the‐art text‐to‐image model, given the relevance of fashion and the scarcity of resources addressing bias in this field. To this end, we build a controlled corpus of 300 fashion‐related descriptions, each adapted into neutral, male and female versions. Empirically, experiments show that Stable Diffusion exhibits significant gender imbalances when generating images from neutral prompts, associating traditionally masculine outfits with male figures and traditionally feminine outfits with female figures. Theoretically, this methodology offers a reproducible approach for detecting and interpreting bias in multimodal generative models, and the resources created in this research are publicly available to scientific community, contributing to the development of fairer and more transparent AI systems.
Villalba‐Osés et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: