
Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code
A new attack surface has emerged in code generation workflows: the very safety mechanisms designed to enforce syntactic correctness can be weaponized to bypass LLM safeguards. Researchers demonstrate that grammar-constrained decoding, widely adopted to improve code reliability, enables a jailbreak technique called CodeSpear that forces models into generating malicious code. This finding inverts conventional wisdom about constraint-based reliability and signals that alignment researchers must rethink how guardrails interact with structured output requirements. The proposed CodeShield defense suggests the field is moving toward adversarial robustness testing of code generation infrastructure before deployment at scale.68



























