DOI: 10.3390/analytics5030030 ISSN: 2813-2203

Minimal but Conditional: Auditing Demographic Bias in Large Language Model Résumé Evaluation Across Commercial and Open-Weight Models

Vasileios Pavlopoulos

Large language models are increasingly used to read résumés and judge who advances in hiring, a task once reserved for people and now handed to systems whose reasoning is hard to inspect. Whether these models carry the demographic biases that have long shaped human hiring is therefore an urgent question, and the published evidence so far is mixed and difficult to interpret, partly because studies tend to test a single condition and rarely confirm that their measurement instrument can detect bias at all. This paper audits demographic bias in résumé evaluation across three current models, one of them open-weight, and it treats robustness as a central concern rather than seeking a single verdict. Each résumé is scored through a reference-anchored comparison task in which the model rates the candidate against a fixed neutral reference for the same occupation. Effect sizes are estimated as standardised mean differences under false-discovery control, the sensitivity of the instrument is tested with an embedded seniority control, and the null findings are corroborated by formal equivalence tests against a justified smallest effect size of interest and by mixed-effects models that account for the clustered structure of repeated evaluations. The audit pairs a positive control that confirms the models read genuine differences in candidate quality with a deliberate attempt to provoke bias by weakening candidates, relaxing the prompt, and adding culture-fit language of the kind used in real hiring. Across more than thirty thousand evaluations, gender and race effects prove negligible and remain so under every one of these conditions. The one systematic preference that emerges favours candidates who appear more experienced, and closer inspection shows that most of it is an artefact of how the résumés were built rather than a bias against age, leaving only a modest effect that surfaces when the prompt is casual. A separate and quieter pattern appears in the open-weight model, which reacts to a few explicit signals of minority status. The broader lesson is that fairness measured on a clean benchmark does not by itself guarantee fairness in deployment, because how a model is prompted can decide whether bias appears.

More from our Archive