Modelwire
Subscribe

Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts

Illustration accompanying: Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts

A new research benchmark exposes a critical tension in deploying safety-aligned LLMs in high-stakes professional contexts. Swiss courts already use smaller models for legal translation and summarization, but criminal law work triggers excessive refusals when models encounter descriptions of violent or sexual offenses, even in legitimate professional summarization tasks. This paper introduces TF-RefusalBench, a multilingual evaluation framework to measure and address over-alignment in legal domains. The finding matters because it reveals how guardrails designed for consumer safety can actively degrade utility in regulated industries where content moderation must yield to professional necessity, forcing a reckoning between alignment philosophy and real-world deployment constraints.

Modelwire context

Explainer

The paper's sharpest contribution isn't the benchmark itself but the framing: over-alignment isn't a bug to patch in isolation, it's a structural conflict between how alignment is trained (optimizing against consumer harm signals) and where models are increasingly being deployed (professional contexts where graphic content is legally necessary to process). Swiss courts are already past the pilot stage, which means this isn't a hypothetical tension.

This is largely disconnected from recent activity in our archive, as we have no prior coverage to anchor it to. It belongs to a growing body of work examining alignment tax in enterprise and regulated-industry deployments, sitting alongside (but distinct from) debates about RLHF calibration and constitutional AI approaches. The legal domain is a particularly pointed test case because the content triggering refusals isn't ambiguous or adversarial, it's routine professional material that happens to describe violence or sexual offenses by definition.

Watch whether TF-RefusalBench gets adopted by any of the major model providers as an evaluation layer in their legal or enterprise fine-tuning pipelines within the next 12 months. Adoption by even one named provider would signal that the field is treating domain-specific over-alignment as a first-class problem rather than an edge case.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsSwiss Federal Supreme Court · TF-RefusalBench · LLMs

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. The full content lives on arxiv.org. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts · Modelwire