The 8 tasks
Tap a task to expand the spec and the side-by-side code. On phones, use the tabs to switch models.
What happened with Astra
Every one of Astra's eight responses has the same structural defect. The model emits the opening lines — def / import lines — as unfenced plain text, then puts the rest of the code in a fenced code block whose content starts at column 0. The function body is detached from its def, so as a literal file nothing imports. This was verified twice via independent copy-button captures, read back character-exact, plus screenshot checks.
Below is Astra's actual binary_search response, exactly as captured — plain-text head, then the fenced block:
Two genuine bugs underneath the formatting
Re-attaching the bodies (indent the fenced part one level, zero logic changes) gets Astra to 6/8. The two that still fail are real model bugs:
Methodology
Codex
- codex exec CLI, codex-cli 0.149.0
- Model resolved to gpt-5.6-sol (Codex default; "astra" is blocked with a ChatGPT account)
- Wall time measured externally around each CLI call: avg ~11s/task, includes ~2–3s CLI startup
- Token counts from the CLI's own "tokens used" line (~2.3k–3.0k per task)
- Zero failures, zero retries; the "output ONLY code" prompt was honored
Muse Spark
- Solutions written directly by the assistant as the fair stand-in
- Same specs, same tests, same grade.py harness
- No generation timing measured — timing comparison across models isn't claimed
GPT-6 Astra
- ChatGPT web app, Work mode, GPT-6 Astra explicitly selected (verified in composer before every prompt)
- 8 prompts sent verbatim, one at a time, each awaited to completion
- Responses collected via message/code-block Copy buttons — no timing possible through the interactive UI
- Defect verified by two independent captures (copy-paste + read-back, character-exact)
Shared harness
- Assert-based tests (test_N.py), no dependencies, run via grade.py in a subprocess with a 30s timeout
- Spec ambiguities found during grading were fixed in the tests before scoring (e.g. the CSV mid-field-quote case was pinned to RFC 4180 and checked against Python's own csv module)
- Honest caveat: these are LeetCode easy/medium tasks — too easy to separate the models on reasoning. The differentiator this round turned out to be output formatting, not intelligence.