
REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection
REDACT addresses a critical gap in AI safety infrastructure: systematic evaluation of personal information detection across languages and contexts. The benchmark spans 25 languages, 51 entity types, and 4,127 surface patterns, using controlled generation axes to isolate failure modes rather than aggregate performance metrics. This matters because PII detection underpins compliance (GDPR, data privacy) and model safety, yet existing benchmarks remain fragmented and ad hoc. Stratified evaluation via sensitivity tiers enables practitioners to measure real-world robustness, not just F1 scores. For teams building multilingual systems or auditing model behavior on sensitive data, REDACT establishes a reproducible standard that exposes which linguistic and formatting conditions break current detectors.62




























