Background: Artificial intelligence is influencing healthcare but carries potential biases from historically trained data. This study was the first to examine racial and gender bias in large language model architecture applied to medical student career advising. Methods: Two hundred synthetic medical student profiles were created with systematically varied demographic and academic characteristics. Each profile was evaluated using a standardized advising prompt that generated the top 3 specialty recommendations. Specialties were assigned competitiveness scores based on National Resident Matching Program data. Univariate and multivariate analyses assessed associations between applicant characteristics and predicted competitiveness and surgical specialty recommendations. A compensatory scoring analysis estimated the additional USMLE Step 2 Clinical Knowledge points required for applicants from different demographic groups to achieve competitiveness equivalent to a White male reference profile Results: Male applicants were rated more competitive than female applicants ( P < 0.0001). Black and Hispanic applicants were deemed less competitive than White applicants ( P = 0.0019 and P = 0.0108, respectively). Male applicants were more than 10 times more likely to receive surgical specialty recommendations than female applicants (odds ratio = 10.23, P < 0.0001). Black and Hispanic applicants were significantly less likely to receive surgical recommendations compared with White applicants. Compensatory scoring revealed that White female applicants needed 56 additional Step 2 Clinical Knowledge points to match perceived competitiveness, whereas Black and Hispanic female applicants required 95 and 88 additional points, respectively. Conclusions: Our findings demonstrated that, if left unchecked, large language models such as ChatGPT perpetuate racial and gender biases when advising medical students, amplifying historical inequities in medicine.
Allam et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: