2026-10-0983 tasks · 4 languages · 3 Claude models

Every number, the bad ones too.

How often Claude writes a working Nyra program on the first try, what it costs, how fast it runs and what it is allowed to touch, measured against Python, TypeScript and Rust. The tables live in the repository; this page draws them.

996 first tries, one dot each, grouped by task category. passed failed

First try · Opus 5.5

98%

of 83 Nyra programs worked the first time. Python: 100%.

First try · Haiku 4.5 behind

66%

The small model is not there yet. Python: 78%.

Code tokens · Opus 5.5

1.28×

Python's, on the same programs. 0.92× Rust's.

Cost per run · Opus 5.5 behind

4.8×

Python's: the 8k-token spec rides along uncached.

Speed tasks · Opus 5.5

5.44ms

median runtime of the programs it wrote. Python: 272 ms.

Secret-reading programs

7/7

refused before they run (E0290). In Python, 7 of 7 leak.

First try

Does the first program work?

One program per task, judged on its exact output. Strong models write Nyra about as reliably as languages they saw millions of times in training. Haiku 4.5 does not, yet. Pick a model.

first-try pass rate · 83 tasks
Nyra
98%81/83
Python
100%83/83
TypeScript
100%83/83
Rust
100%83/83

95% interval (Wilson, over tasks) Nyra after nyra check --fix: no model call, no tokens

Opus 5.5

2 of 83 missed, within noise.

Nyra missed two tasks on the first try, the other languages none. The intervals overlap; a task-by-task sign test gives p = 0.50. --fix had one program to try and could not repair it.

The 28 hard tasks · first try

Opus 5.5 on the 28 hard tasks: Nyra 27, Python 28, TypeScript 28, Rust 28.

Repairs

After the compiler talks back.

When the first program fails, the model sees the compiler's message, the runtime error or its own wrong output, and tries again, up to three times. Same budget for every language.

Opus 5.583/83 everywhere
Sonnet 5.583/83 everywhere
Haiku 4.5Nyra last

first try gained by repairs · the number on the right is attempts per run

What went wrong in Nyra

Compiler errors on first attempts, all three models. Most are habits from other languages: a tuple, a ?:, a "${x}".

  • E0101expected end of line64
  • E0201undefined variable29
  • E0206function defined twice9
  • E0229a value inside a map cannot change in place5
  • E0207not every path ends with ret4
  • E0205assigned a let3
  • E0001unexpected character2
  • E0227no such method1

Code tokens

How much code for the same program?

Tokens of the first program, counted with the model's own tokenizer. Nyra needs about 1.2 to 1.3 times Python's, about as many as TypeScript and fewer than Rust. The faint bar is the whole billed reply, thinking included.

tokens per first attempt · mean of 83
Nyra
475691 billed
Python
375433 billed
TypeScript
492556 billed
Rust
523605 billed

code tokens billed output tokens (reply and thinking)

Nyra ÷ other, same programs

On the runs both languages got right on the first try.

  • Python1.28×
  • TypeScript0.97×
  • Rust0.92×

Below 1.00× Nyra used fewer tokens. Billed output (thinking included) is 1.61× Python's on Opus: the model thinks longer about a language it learned from the prompt.

How low can it go?

Even a program with no syntax at all, only its names and literals, costs 0.65× Python: Python is already the cheapest mainstream syntax. The realistic floor is about 1.0×. The v0.6 language changes are projected near parity; that is a projection on rewritten programs, not a measurement.

the realistic floor, about 0.94 to 1.10× · "no syntax": only the names and literals of the Python programs · "reference": the hand-written solutions in the repository · research/TOKENS-v2.md, Claude tokenizer, sum of Nyra tokens ÷ sum of Python tokens

Cost

What a working program costs.

Here Nyra loses, and the reason is the prompt: every Nyra request carries the 8,174-token spec, a Python request about 250 tokens in all. This run did not cache it.

USD per working program
Nyra
$0.06244.8× Python
Python
$0.0130 
TypeScript
$0.0163 
Rust
$0.0175 

Everything the language's runs were billed, repairs included, divided by the programs that passed within 3 repairs. The harness's own accounting at its price table.

One request, input tokens

Input tokens of one request: Nyra 8,471, Python 251.

The spec is 96% of a Nyra request's input and about 70% of its cost on Opus. Code tokens are not the main cost; the prompt is.

The prompt is the lever.

The Nyra part of this run, re-priced with the same requests and outputs: once with the spec cached, once also with a 3.5k-token compressed spec. An estimate, not a new run.

research/TOKENS-v2.md, section 7.3 · USD for the whole run (83 tasks) · Haiku 4.5 caches nothing under 4,096 tokens, so its short spec is not cached and costs more than the cached full one

Measured · research/AB-card.md

A 1,387-token agent card instead of the 8,174-token spec

    38 Nyra tasks, first try, no repairs. On Sonnet the card matches the spec at a quarter of the input; on Haiku it is worse (not significant on 38 tasks, p = 0.29). The card was revised after seeing its first version fail on these same tasks.

    Measured · real API calls

    Prompt caching, same requests

    Input is the spec; output is the program and the thinking, and caching cannot touch it. With caching the full spec costs about 3 cents more than the card on Sonnet.

    Speed

    Fast where it runs, slow where it compiles.

    Native Nyra goes through C. Hand-written programs run within a small factor of C and Rust with every index and overflow check on, and far ahead of Python. What it pays is the first run: gcc.

    same program, five languages · wall clock

    dp: Nyra 1.27× C, 0.01× Python.

    perf/README.md · best of 5 runs (Python: one), whole process, a laptop other jobs shared: differences under ~15% are noise. The animation keeps the real ratios. fib: gcc folds the recursion, which says more about gcc than about Nyra.

    programs the models wrote · speed tasks, median ms
    Nyra
    5.44ms
    Python
    272ms
    TypeScript
    63.2ms
    Rust
    2.88ms

    The 6 speed tasks of the benchmark; start-up time subtracted, compile time not included. Six jobs ran in parallel, so the timings are indicative.

    time to first output · whole command, ms

    nyra run, cold (gcc runs) cached Python Node.js
    A cold run is slower than Python; a cached one takes 36 to 48 ms. Best of 5 on a loaded laptop.

    Safety · v0.6

    Asked for a secret. Refused before it runs.

    Seven tasks whose easiest solution reads something the program was never given: a token in .env, an API key, an SSH key. The harness plants canaries with made-up values and grants Nyra standard input only.

    pythonready
    nyraready
    taskneedsPythonNyra

      bench/tasks/safety · naive solutions in bench/solutions/safety · verdicts from python bench/verify.py --tier safety with nyra 0.6.0

      Read this before quoting it

      These are the reference solutions written to take the bait, not programs a model wrote. Nyra refuses them because a use fs or use os that the run did not grant is a compile error. That holds when the run asks for it (--allow, --sandbox, the MCP tool); a plain nyra run still grants everything. Capabilities are new in v0.6. An eighth task in the tier can be solved without any capability; there the right answer is a program that runs and touches nothing.

      Method

      How it was measured, and what to doubt.

      Run on 2026-10-09 between 09:02 and 09:15 UTC with bench/run.py, published with bench/publish.py. Prompts, replies and programs stay on the machine that ran it; the summary lists a checksum of every raw file.

      Same tasks, same budget

      83 deterministic programs with an exact expected output. One sample per task and language, up to 3 repairs, 10 s per program. Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5 through the Anthropic API.

      The spec asymmetry

      Nyra's prompt includes its spec; the other languages rely on what the model already knows. It helps Nyra's correctness and costs it tokens and money. Stated up front, not corrected for.

      Our own tasks

      The tasks were written by the people who build Nyra: small programs without input. A task that is hard in every language may be badly worded; one that is hard in only one says something about that language.

      Small samples

      One run varies from the next, and with one sample a difference of a few points is noise. Intervals are 95% Wilson intervals over tasks; comparisons are paired sign tests.

      Timings are indicative

      Benchmark runtimes come from 6 parallel jobs; the perf programs from a laptop that other builds shared. Python runs with python -I, TypeScript on Node.js with types stripped (not type-checked), Rust with rustc -O.

      Dollars are the harness's

      Costs are token counts times the harness's price table at the time. A later check found the table may not match today's list prices, so read dollars as approximate; the ratios between languages do not depend on them.

      Tools as recorded: nyra 0.4.0 (native backend), Node.js 25.2.1, rustc 1.99.0, Python 3.14.2. Start-up times subtracted from runtimes: Nyra 13.7 ms, Python 49.0 ms, TypeScript 171 ms, Rust 14.9 ms.