Code-generation benchmark · 2026-10-02 · 8 Python tasks

Muse vs Codex vs Astra

Three models, the same eight coding tasks, one shared assert-based test harness. Muse and Codex went 8-for-8. Astra's code almost did — and the reason it didn't is the whole story.

8/8
Muse Spark
Direct generation, same tests
Every task passed on the first attempt. No generation timing was measured on this side, so speed isn't part of the claim.
8/8
Codex default gpt-5.6-sol
Via codex exec CLI · codex-cli 0.149.0
Avg ~11s per task wall clock (includes ~2–3s CLI startup). Independently nailed the tricky RFC 4180 quoting rule in the CSV task.
as-shipped
0/8
GPT-6 Astra
Via ChatGPT web app · Work mode · GPT-6 Astra selected
6/8 on logic after a purely mechanical re-indentation (no logic edits). The defect is in how the code arrived, not in the reasoning — details below.

The 8 tasks

Tap a task to expand the spec and the side-by-side code. On phones, use the tabs to switch models.

What happened with Astra

Every one of Astra's eight responses has the same structural defect. The model emits the opening lines — def / import lines — as unfenced plain text, then puts the rest of the code in a fenced code block whose content starts at column 0. The function body is detached from its def, so as a literal file nothing imports. This was verified twice via independent copy-button captures, read back character-exact, plus screenshot checks.

Below is Astra's actual binary_search response, exactly as captured — plain-text head, then the fenced block:

Sent as plain text (not in a code block)
Fenced code block — starts at column 0
What Python says
SyntaxError: 'return' outside function (astra_1.py, line 7)

Two genuine bugs underneath the formatting

Re-attaching the bodies (indent the fenced part one level, zero logic changes) gets Astra to 6/8. The two that still fail are real model bugs:

1. def init, not __init__ (LRUCache task) — confirmed in the raw message markdown, not a rendering artifact. LRUCache(2) raises TypeError: LRUCache() takes no arguments.
2. Broken regex (md_to_text task) — unescaped * quantifiers with nothing to repeat: re.error: nothing to repeat at position 41. The link/image alternative is malformed too.
The takeaway: Astra's code reasoning is in the same ballpark as the others (6/8 on logic), but its output serialization through the ChatGPT app was systematically broken — 8 out of 8 responses. Nothing it produced runs as-pasted. Whether the defect originates in the model or in how the app renders/copies Astra output couldn't be determined from the app alone.

Methodology

Codex

  • codex exec CLI, codex-cli 0.149.0
  • Model resolved to gpt-5.6-sol (Codex default; "astra" is blocked with a ChatGPT account)
  • Wall time measured externally around each CLI call: avg ~11s/task, includes ~2–3s CLI startup
  • Token counts from the CLI's own "tokens used" line (~2.3k–3.0k per task)
  • Zero failures, zero retries; the "output ONLY code" prompt was honored

Muse Spark

  • Solutions written directly by the assistant as the fair stand-in
  • Same specs, same tests, same grade.py harness
  • No generation timing measured — timing comparison across models isn't claimed

GPT-6 Astra

  • ChatGPT web app, Work mode, GPT-6 Astra explicitly selected (verified in composer before every prompt)
  • 8 prompts sent verbatim, one at a time, each awaited to completion
  • Responses collected via message/code-block Copy buttons — no timing possible through the interactive UI
  • Defect verified by two independent captures (copy-paste + read-back, character-exact)

Shared harness