
RuBench benchmarks coding agents on native Russian task specifications
Coding agents trained on English-heavy datasets now face a real-world gap: developers worldwide file maintenance requests in their native languages. RuBench closes this measurement blind spot by introducing 25 repository-level tasks specified natively in Russian, sourced from live open-source projects and validated against maintainer regression tests. The benchmark spans Python, PHP, TypeScript, and JavaScript ecosystems, forcing agent evaluators to confront multilingual task comprehension beyond translation pipelines. This matters because production coding agents must handle customer requests as written, not sanitized English proxies, making RuBench a critical stress test for deployment readiness.62

























