Safety metrics mask persistent gender bias across GPT generations

A new study challenges the validity of safety evaluations for large language models by demonstrating that harmful biases are being masked rather than eliminated across model generations. Researchers analyzed 450,000 gender-directed completions from GPT-2 through GPT-5 and found that while explicit discriminatory content becomes less visible in newer models, underlying representational disparities persist and shift form. The work exposes a critical gap in current evaluation methodology: surface-level harm metrics may create false confidence in safety progress while structural inequities remain embedded in model outputs. This finding has direct implications for how labs validate safety claims and how regulators assess model safety compliance.
Modelwire context
Skeptical readThe paper's real contribution isn't that bias persists (known for years) but that it *shifts form* across generations in ways current evals miss. The critical omission: the study doesn't establish whether this transformation makes the bias harder to trigger in practice, or just harder to detect in their particular test set of 450,000 completions.
This connects directly to the September safety evaluation work on coding agents ('Coding Agents with an Obstacle-Aware Harness'), which found that LLMs receive explicit safety instructions yet still fail because planning algorithms treat constraints as secondary. Both papers expose the same gap: surface-level compliance metrics (explicit discrimination disappears, safety instructions are acknowledged) mask deeper misalignment in how models actually reason about trade-offs. The difference is scope: one targets gender bias in text, the other targets physical safety in robotics, but both suggest that safety training optimizes for evaluation visibility rather than structural change.
If OpenAI or Anthropic release detailed ablations showing which training interventions actually reduce the *triggerable* rate of gender discrimination (not just its surface form), versus which merely redistribute it, that would validate or refute the 'laundering' framing. If they don't publish that level of detail within six months, the skepticism stands: we'll have no way to know if the bias is genuinely harder to activate or just harder to measure.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsOpenAI · GPT-2 · GPT-4 · GPT-5
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as “Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.