Skip to main content
Back to Newswire
Science AI

RuBench Tests Coding Agents on Russian-Language Tasks From Real Repositories

A new benchmark called RuBench 1.0 evaluates product-grade coding agents on 25 repository-level tasks mined from recent fix commits in five live open-source repositories (aiohttp, aiogram, Laravel, NestJS, Fastify) spanning Python, PHP, TypeScript, and JavaScript. Each task is specified natively in Russian, written from scratch in the style of an actual customer request rather than translated from English. Tasks are judged by upstream maintainer regression tests withheld from release. All 25 fix commits postdate the training-data cutoffs of every evaluated model, providing a contamination argument that holds task-by-task. The study evaluates deployed product configurations -- Claude Code with Opus 4.8, Sonnet 5, and Haiku 4.5, and Codex CLI with GPT-5.5 -- with three independent runs each, reporting pass@1 with task-level confidence intervals, paired comparisons, dollar cost, and token usage. The best configuration resolves 78.7% of tasks; at N=25 only gaps to the weakest model are statistically resolvable. Auditing full trajectories of a fifth configuration (Claude Code + Fable 5, July 2, 2026 release) caught the product silently substituting the model: on 5 of 25 tasks (20%) an official safeguard fallback re-routed routine HTTP-protocol fixes to Opus 4.8, providing direct evidence that the deployed product, not the model alone, is the unit actually measured.
Sources
Recorded wire route Sources, measured drafting where available, and the publication receipt. See concurrent Machine
Evidence entered
Admission Evidence and chronology passed Science AI
Publication receipt Entered the validated Newswire
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire