Abstract:The use of Large Language Models (LLMs) in hiring has led to legislative actions to protect vulnerable demographic groups. This paper presents a novel framework for benchmarking hierarchical gender hiring bias in Large Language Models (LLMs) for resume scoring, revealing significant issues of reverse gender hiring bias and overdebiasing. Our contributions are fourfold: Firstly, we introduce a new construct grounded in labour economics, legal principles, and critiques of current bias benchmarks: hiring bias can be categorized into two types: Level bias (difference in the average outcomes between demographic counterfactual groups) and Spread bias (difference in the variance of outcomes between demographic counterfactual groups); Level bias can be further subdivided into statistical bias (i.e. changing with non-demographic content) and taste-based bias (i.e. consistent regardless of non-demographic content). Secondly, the framework includes rigorous statistical and computational hiring bias metrics, such as Rank After Scoring (RAS), Rank-based Impact Ratio, Permutation Test, and Fixed Effects Model. Thirdly, we analyze gender hiring biases in ten state-of-the-art LLMs. Seven out of ten LLMs show significant biases against males in at least one industry. An industry-effect regression reveals that the healthcare industry is the most biased against males. Moreover, we found that the bias performance remains invariant with resume content for eight out of ten LLMs. This indicates that the bias performance measured in this paper might apply to other resume datasets with different resume qualities. Fourthly, we provide a user-friendly demo and resume dataset to support the adoption and practical use of the framework, which can be generalized to other social traits and tasks.

Gender bias and stereotypes in Large Language Models

Gender bias and stereotypes in Large Language Models

Evaluating Gender, Racial, and Age Biases in Large Language Models: A Comparative Analysis of Occupational and Crime Scenarios

Protected group bias and stereotypes in Large Language Models

Gender Bias in Large Language Models across Multiple Languages

Assessing Gender Bias in LLMs: Comparing LLM Outputs with Human Perceptions and Official Statistics

Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts

Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias

The Unequal Opportunities of Large Language Models: Revealing Demographic Bias through Job Recommendations

Hire Me or Not? Examining Language Model's Behavior with Occupation Attributes

Locating and Mitigating Gender Bias in Large Language Models

JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models

Large Language Models Portray Socially Subordinate Groups as More Homogeneous, Consistent with a Bias Observed in Humans

Evaluation of Large Language Models: STEM education and Gender Stereotypes

Unveiling Gender Bias in Large Language Models: Using Teacher's Evaluation in Higher Education As an Example

Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models

Gender Bias of LLM in Economics: An Existentialism Perspective

Towards Auditing Large Language Models: Improving Text-based Stereotype Detection

Evaluating Gender Bias of LLMs in Making Morality Judgements

Measuring Gender and Racial Biases in Large Language Models

White Men Lead, Black Women Help? Benchmarking Language Agency Social Biases in LLMs