Same tasks, same budget
83 deterministic programs with an exact expected output. One sample per task and language, up to 3 repairs, 10 s per program. Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 through the Anthropic API.
MIT license · Nyra v0.5
How often Claude writes a working Nyra program on the first try, what it costs, how fast it runs and what it is allowed to touch, measured against Python, TypeScript and Rust. The tables live in the repository; this page draws them.
First try · Opus 5.5
98%
of 83 Nyra programs worked the first time. Python: 100%.
First try · Haiku 4.5 behind
66%
The small model is not there yet. Python: 78%.
Code tokens · Opus 5.5
1.28×
Python's, on the same programs. 0.92× Rust's.
Cost per run · Opus 5.5 behind
4.8×
Python's: the 8k-token spec rides along uncached.
Speed tasks · Opus 5.5
5.44ms
median runtime of the programs it wrote. Python: 272 ms.
Secret-reading programs
7/7
refused before they run (E0290). In Python, 7 of 7 leak.
First try
One program per task, judged on its exact output. Strong models write Nyra about as reliably as languages they saw millions of times in training. Haiku 4.5 does not, yet. Pick a model.
95% interval (Wilson, over tasks) Nyra after nyra check --fix: no model call, no tokens
Opus 5.5
Nyra missed two tasks on the first try, the other languages none. The intervals overlap; a task-by-task sign test gives p = 0.50. --fix had one program to try and could not repair it.
Sonnet 5.5
All three first attempts that failed did not compile, and nyra check --fix repaired each of them without calling the model: 83 of 83. Python missed one, TypeScript two (p = 0.62).
Haiku 4.5
66% against Python's 78%, and on the 28 hard tasks 4 against 13. That gap is not chance (sign test p = 0.01). --fix repaired 1 of 20 programs it tried. Haiku runs without thinking here, and this is where an unfamiliar language shows.
The 28 hard tasks · first try
Opus 5.5 on the 28 hard tasks: Nyra 27, Python 28, TypeScript 28, Rust 28.
Repairs
When the first program fails, the model sees the compiler's message, the runtime error or its own wrong output, and tries again, up to three times. Same budget for every language.
first try gained by repairs · the number on the right is attempts per run
Compiler errors on first attempts, all three models. Most are habits from other languages: a tuple, a ?:, a "${x}".
E0101expected end of line64E0201undefined variable29E0206function defined twice9E0229a value inside a map cannot change in place5E0207not every path ends with ret4E0205assigned a let3E0001unexpected character2E0227no such method1Code tokens
Tokens of the first program, counted with the model's own tokenizer. Nyra needs about 1.2 to 1.3 times Python's, about as many as TypeScript and fewer than Rust. The faint bar is the whole billed reply, thinking included.
code tokens billed output tokens (reply and thinking)
Nyra ÷ other, same programs
On the runs both languages got right on the first try.
Below 1.00× Nyra used fewer tokens. Billed output (thinking included) is 1.61× Python's on Opus: the model thinks longer about a language it learned from the prompt.
Even a program with no syntax at all, only its names and literals, costs 0.65× Python: Python is already the cheapest mainstream syntax. The realistic floor is about 1.0×. The v0.6 language changes are projected near parity; that is a projection on rewritten programs, not a measurement.
the realistic floor, about 0.94 to 1.10× · "no syntax": only the names and literals of the Python programs · "reference": the hand-written solutions in the repository · research/TOKENS-v2.md, Claude tokenizer, sum of Nyra tokens ÷ sum of Python tokens
Cost
Here Nyra loses, and the reason is the prompt: every Nyra request carries the 8,174-token spec, a Python request about 250 tokens in all. This run did not cache it.
Everything the language's runs were billed, repairs included, divided by the programs that passed within 3 repairs. The harness's own accounting at its price table.
One request, input tokens
Input tokens of one request: Nyra 8,471, Python 251.
The spec is 96% of a Nyra request's input and about 70% of its cost on Opus. Code tokens are not the main cost; the prompt is.
The Nyra part of this run, re-priced with the same requests and outputs: once with the spec cached, once also with a 3.5k-token compressed spec. An estimate, not a new run.
research/TOKENS-v2.md, section 7.3 · USD for the whole run (83 tasks) · Haiku 4.5 caches nothing under 4,096 tokens, so its short spec is not cached and costs more than the cached full one
Measured · research/AB-card.md
38 Nyra tasks, first try, no repairs. On Sonnet the card matches the spec at a quarter of the input; on Haiku it is worse (not significant on 38 tasks, p = 0.29). The card was revised after seeing its first version fail on these same tasks.
Measured · real API calls
Input is the spec; output is the program and the thinking, and caching cannot touch it. With caching the full spec costs about 3 cents more than the card on Sonnet.
Speed
Native Nyra goes through C. Hand-written programs run within a small factor of C and Rust with every index and overflow check on, and far ahead of Python. What it pays is the first run: gcc.
dp: Nyra 1.27× C, 0.01× Python.
perf/README.md · best of 5 runs (Python: one), whole process, a laptop other jobs shared: differences under ~15% are noise. The animation keeps the real ratios. fib: gcc folds the recursion, which says more about gcc than about Nyra.
The 6 speed tasks of the benchmark; start-up time subtracted, compile time not included. Six jobs ran in parallel, so the timings are indicative.
nyra run, cold (gcc runs) cached Python Node.js
A cold run is slower than Python; a cached one takes 36 to 48 ms. Best of 5 on a loaded laptop.
Safety · v0.6
Seven tasks whose easiest solution reads something the program was never given: a token in .env, an API key, an SSH key. The harness plants canaries with made-up values and grants Nyra standard input only.
bench/tasks/safety · naive solutions in bench/solutions/safety · verdicts from python bench/verify.py --tier safety with nyra 0.6.0
Read this before quoting it
These are the reference solutions written to take the bait, not programs a model wrote. Nyra refuses them because a use fs or use os that the run did not grant is a compile error. That holds when the run asks for it (--allow, --sandbox, the MCP tool); a plain nyra run still grants everything. Capabilities are new in v0.6. An eighth task in the tier can be solved without any capability; there the right answer is a program that runs and touches nothing.
Method
Run on 2026-10-09 between 09:02 and 09:15 UTC with bench/run.py, published with bench/publish.py. Prompts, replies and programs stay on the machine that ran it; the summary lists a checksum of every raw file.
83 deterministic programs with an exact expected output. One sample per task and language, up to 3 repairs, 10 s per program. Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 through the Anthropic API.
Nyra's prompt includes its spec; the other languages rely on what the model already knows. It helps Nyra's correctness and costs it tokens and money. Stated up front, not corrected for.
The tasks were written by the people who build Nyra: small programs without input. A task that is hard in every language may be badly worded; one that is hard in only one says something about that language.
One run varies from the next, and with one sample a difference of a few points is noise. Intervals are 95% Wilson intervals over tasks; comparisons are paired sign tests.
Benchmark runtimes come from 6 parallel jobs; the perf programs from a laptop that other builds shared. Python runs with python -I, TypeScript on Node.js with types stripped (not type-checked), Rust with rustc -O.
Costs are token counts times the harness's price table at the time. A later check found the table may not match today's list prices, so read dollars as approximate; the ratios between languages do not depend on them.
Tools as recorded: nyra 0.4.0 (native backend), Node.js 25.2.1, rustc 1.99.0, Python 3.14.2. Start-up times subtracted from runtimes: Nyra 13.7 ms, Python 49.0 ms, TypeScript 171 ms, Rust 14.9 ms.
Raw data
Every figure on this page comes from one of these files. The page's own copy of the numbers is data.js, with the source of each block.
bench/published/2026-10-v0.5-anthropic.mdThe Claude run: pass rates, repairs, tokens, cost, runtime, per category, failures.→2026-10-v0.5-anthropic.jsonThe same, for tools, with checksums of the raw files.→bench/published/results.jsonThe leaderboard data, with confidence intervals.→research/AB-card.mdAgent card against full spec, and what prompt caching saves.→research/TOKENS-v2.mdWhere the token gap comes from, and where the floor is.→perf/README.mdNyra against C, Rust, Node.js and Python; time to first output.→research/SPEED-run.mdCompile flags, compile time and what an agent waits for.→bench/tasks/safetyThe safety tier: tasks, canaries, leak markers.→