Peer performance evaluations are an important determinant of hiring and promotion decisions, but how objective are they? To tackle this question, we measure gender bias in performance evaluations in a large course at a flagship public university. We exploit the random assignments of both peer evaluators and blinded official graders over several essay assignments, where they are incentivized to match official blinded grades. We find that male peer graders assign higher scores to classmates without female-sounding names in content, but lower scores in writing style. Interestingly, we do not find such biases for female graders. Our findings highlight potential challenges of designing fair assessment practices. • We assign peer evaluators randomly to score essays using a rubric. • Evaluators were told that the official grading is done blindly and were incentivized to match official grades, adding a monitoring effect. • We exploit the random assignments of both peer evaluators and blinded official graders over several essay assignments. • Using student and peer evaluator fixed effects and conditioning on blinded official grades, we find that female peers grade essays written by female-sounding names similarly with essays written by students without female-sounding names. • However, male peers are more generous graders of essays written by students without female- sounding names in content scores, while this disparity reverses in favor of female-sounding names when it comes to evaluating writing and grammar. • These results suggest that biased performance evaluations could be at least partly responsible for gender gaps in hiring and promotion, particularly with tasks associated with gender stereotypes in male-dominated fields where evaluators are more likely to be men.
Saygin et al. (Wed,) studied this question.