Like humans, language models demonstrate face-to-character biases
Steven A Lehr, Yash Lothe, Mahzarin R BanajiAbstract
Humans routinely make unjustified character inferences, such as labeling people as trustworthy or untrustworthy, based on facial features. We conducted 13 experiments, using four models and totaling nearly 8,000 trials, to ask: Would large language models (LLMs), trained foundationally on language and not images, mirror human biases by inferring character traits from 2D facial images, or would they be free from this human error? GPT-4o reliably exhibited face-biased judgments of competence and trustworthiness (experiments 1 and 2), generalized these judgments to semantically related traits (experiments 3 and 4), and even to primate faces that were absent from its training data (experiment 5). Strikingly, the LLM's face-to-character inferences escalated to ascriptions of extreme negative behavior such as murder and human trafficking (experiment 6) as well as to positive high-impact decisions like selection for employment or venture capital funding (experiment 7). To ensure these results were not an idiosyncratic feature of a particular LLM (GPT-4o), in experiments 8–13, we showed the same patterns, usually with an even greater degree of bias, in GPT-5, Gemini 3 Flash Preview, and Claude Sonnet 4.5. These robust biases stand in contrast to LLMs’ reluctance to explicitly endorse race/gender stereotypes, and suggest that alignment efforts to date have been domain-specific, with current models lacking generalized egalitarian decision-making.