Legal RAG systems still hallucinate in up to half of responses
A systematic evaluation of eight legal RAG systems reveals that hallucination remains a critical failure mode in high-stakes domains, with error rates spanning 10% to 50% depending on the system. The study, spanning English GDPR and French civil law corpora, moves beyond generic hallucination metrics to measure severity and performance variance across question types and user expertise levels. This work signals a maturation in how the field assesses production-grade AI reliability: legal applications cannot tolerate the permissiveness of consumer chatbots, forcing vendors and researchers to confront whether current retrieval and grounding techniques are sufficient for regulated deployment.62












