Modelwire
Subscribe

GPT-5.2 and Gemini show divergent bias patterns across English and Swahili

Researchers tested GPT-5.2 and Gemini 2.5 Flash across 4,900 English-Swahili prompt pairs to measure whether social biases persist uniformly across languages. The study reveals that bias manifests differently rather than identically: stereotype prevalence shifted by up to 12 percentage points on specific demographic axes, Gemini's neutral-sentiment outputs nearly doubled in Swahili, and GPT-5.2 exhibited stark refusal asymmetry, blocking 169 English prompts while allowing zero refusals in Swahili. This finding challenges the assumption that alignment training generalizes across languages and exposes a critical gap in multilingual safety evaluation, particularly for non-English deployments where guardrails may be substantially weaker.

Modelwire context

Explainer

The critical finding isn't that bias exists across languages, but that safety mechanisms (refusals, guardrails) actively diverge by language while bias itself shifts unpredictably. This means a model can be more permissive in one language while appearing aligned in another, creating a false sense of safety in non-English deployments.

This connects directly to the multilingual planning failures documented in early August, where low-resource languages systematically accumulated errors during conversion. But this bias study reveals a sharper problem: the safety layer itself becomes language-dependent. Unlike the TreeProbe benchmark (August 1st), which exposed gaps in training data representation, this work shows that alignment training itself fails to generalize across linguistic boundaries. The refusal asymmetry in GPT-5.2 mirrors the sycophancy vulnerabilities found in medical contexts, where context and framing determine whether guardrails activate, except here the context is simply the language of the prompt.

If OpenAI and Google release multilingual safety audits for their next model versions within the next two quarters that explicitly measure refusal rates and bias prevalence by language, that signals they're treating this as a priority. If they don't, and deployments in non-English markets proceed without language-specific safety benchmarks, that confirms the gap will persist in production systems.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsGPT-5.2 · Gemini 2.5 Flash · OpenAI · Google

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. arXiv cs.CL originally reported this story as Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili”. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

GPT-5.2 and Gemini show divergent bias patterns across English and Swahili · Modelwire