Claude Code vs Codex, round 2: 3 practical tasks across multiple files, each run 3 times
Both tools scored full marks in our first round, which used small, single-file tasks. This time, we built a small online store project closer to real service code (pricing, inventory, orders, and storage across 6 files). Tasks covered a bug described only through customer-reported symptoms, a feature defined by a specification, and a bug triggered only by concurrent orders.
Run on 2026-10-04
At a Glance
- Accuracy: both tools again passed every hidden test in all 9 runs each. Because the concurrency task uses random delays, each run was graded three times, with no failures.
- Speed: based on the median for each task, Codex (GPT-6.1 Sol) was approximately 1.3~1.9 times faster than Claude Code (Claude Opus 5). The gap was 24~42 seconds depending on the task.
- Tokens: based on the median for each task, Claude Code used approximately 4~6 times as many input tokens as Codex, including cache reuse, and approximately 2.5~3.7 times as many output tokens.
- Work habits: Codex added new regression test files in all 6 runs of the bug task (T4) and concurrency task (T6). In one run, it also changed the README and storage code (store.js). Claude Code changed only source files in all 9 runs and added no tests.
- Design choices: across all 6 concurrency runs, both tools used a global lock that 'puts order processing into a single queue'. The results are correct, but orders for different products must also wait their turn.
Results
Times are medians of 3 runs; parentheses show the fastest~slowest values.
| Task | Tool | Passed (all 3 runs) | Time taken | Input tokens | Output tokens | Lines changed (+/−) | |
|---|---|---|---|---|---|---|---|
| T4 Fix bugs across multiple files · Claude Code | T4 Fix bugs across multiple files | Claude Code | 3/3 | 60.6s (59.7–75.2) | 564,618 | 4,495 | +6 / −5 |
| T4 Fix bugs across multiple files · Codex CLI | Codex CLI | 3/3 | 36.3s (31.9–44.2) | 94,160 | 1,363 | +4 / −4 | |
| T5 Add a feature from a specification · Claude Code | T5 Add a feature from a specification | Claude Code | 3/3 | 90.1s (89.8–110.6) | 592,485 | 7,366 | +42 / −2 |
| T5 Add a feature from a specification · Codex CLI | Codex CLI | 3/3 | 47.8s (47–49.3) | 95,745 | 2,002 | +41 / −2 | |
| T6 Concurrency bug + API migration · Claude Code | T6 Concurrency bug + API migration | Claude Code | 3/3 | 111.9s (104.4–147.4) | 660,107 | 9,554 | +87 / −16 |
| T6 Concurrency bug + API migration · Codex CLI | Codex CLI | 3/3 | 85s (79–93.1) | 162,193 | 3,818 | +49 / −18 |
Input tokens largely include cached tokens reused when rereading earlier conversation steps. The two CLIs count tokens differently, so use these figures only to compare approximate scale.
Claude Code reported an 'equivalent cost at standard API rates' of approximately 0.78~1.25 USD per task. Subscription accounts do not pay this amount separately; usage counts toward the plan's limits. Codex did not report an equivalent figure, so we did not compare it.
3 tasks
T4 Fix bugs across multiple files
We supplied only 4 customer reports in ISSUES.md (JPY amounts have decimal places, buying exactly 10 items does not trigger the bulk discount, a 15% coupon produces a 1-cent discrepancy, and a coupon makes the total negative), leaving the tools to find the correct rules in the README. The causes are spread across money.js and cart.js.
T5 Add a feature from a specification
Implement partial refunds (refundOrder) according to the README specification. Hidden tests check reduced refunds when returns invalidate the bulk discount, coupon recalculation, inventory restoration, refunds split across multiple requests, rejection of excessive refunds, and the requirement that nothing changes on failure.
T6 Concurrency bug + API migration
Fix a bug where multiple people buying the last item simultaneously all succeed, making inventory negative, and add await support while preserving existing callback-based calls. Hidden tests check 20 concurrent buyers, deadlocks in orders involving multiple products, all-or-nothing behavior, and error propagation.
Prompt used for every run
Only the sentence after 'Task:' in the prompt below changes for each task. Unlike the first round, we instructed the tools to read 'README.md and other documentation'.
This folder is a small Node.js project (ES modules, no dependencies). Read the code, README.md and any other docs, and the tests. Task: {task} Do not modify existing test files. When you are done, make sure `node --test` passes.- T4 Fix bugs across multiple files —
Fix the bugs reported in ISSUES.md. README.md describes the correct pricing rules. - T5 Add a feature from a specification —
Implement refundOrder in src/orders.js according to the Refunds section of README.md. - T6 Concurrency bug + API migration —
Fix the problems described in the Orders section of README.md (overselling, and supporting both the Promise and the callback style).
How We Measured
- Run date: 2026-10-04 (night, South Korean time), Apple M1 Pro · macOS 26.6 · Node.js 22. The same computer, versions, and settings as the first round.
- Claude Code 2.1.250 — default model Claude Opus 5 (Claude Haiku 4.5 was also called for short supporting tasks). Codex CLI 0.159.2 — model GPT-6.1 Sol, reasoning effort medium.
- Both tools ran in non-interactive mode (claude -p, codex exec) using the operator's personal subscription accounts, with file edits and node execution automatically allowed. Runs were covered by existing subscriptions at no additional cost.
- We varied the execution order for each task (Claude Code first in run 1, Codex first in run 2, Claude Code first in run 3). No heavy work, such as builds, ran on the same computer during testing.
Grading method
- Each task included 3 tests visible to the agent and hidden tests (9~12) added only during grading. We first validated the hidden tests using reference code written by the operator, then ran the concurrency reference code 10 times to check that results were consistent.
- Before grading, we restored the test folder to its original state. Test files added by the agents were not used for grading.
- Because concurrency bugs can appear intermittently, we graded every result three times at different times and recorded the lowest score.
How the code written by the two tools differed
- T4: both tools found all 4 bugs in money.js (JPY decimal places and rounding) and cart.js (the 10-item threshold and coupon cap). Source changes were similar: 5~9 lines for Claude Code and 4~5 lines for Codex. In all three runs, Codex created a separate regression test file reproducing the reported issues.
- T5: both tools followed the specification to 'recalculate the remaining items using the original pricing rules and subtract that amount'. Added source lines were also similar, at 40~47.
- T6: across all 6 runs, both tools serialized all order processing through a single queue (global lock). In runs 2·3, Codex scoped the lock to the store so multiple order services using the same store shared the queue. In run 1, it added a transaction feature to the store and documented it in the README. Claude Code made changes only within orders.js. Neither tool used separate locks for each product.
Limitations of This Test
- The project is small: approximately 200 lines across 6 files. It is closer to practical work than the first round, but much smaller than a real service repository.
- Both tools scored full marks, so the test did not reveal an accuracy difference. Results may differ with larger repositories or tasks with ambiguous requirements.
- Time taken varies with server conditions and the network. Alternating runs on the same night reduced these effects but did not eliminate them.
- Codex used reasoning effort medium, while Claude Code used default settings. We did not measure how much of each subscription's usage allowance was consumed.
Which Tool Should You Use?
The following reflects the site operator's opinion based on these test results.
- Even with practical tasks at this scale, no difference in output accuracy emerged. It makes sense to start with the tool included in a subscription you already pay for.
- If you want faster runs, Codex was again 1.3~1.9 times faster.
- Claude Code was better at keeping changes narrow, modifying only source files. Codex added regression tests on its own, but sometimes also changed the README and storage code without being asked, so its changes need review.
- For problems requiring design judgment, such as concurrency and performance, both tools chose the simplest correct solution. If requirements such as throughput matter, state them explicitly in the prompt.
All 18 run records
| Task | Tool | Run | Seconds | Tests passed | Input tokens | Output tokens | Lines changed (+/−) | |
|---|---|---|---|---|---|---|---|---|
| T4 · Claude Code · Run 1 | T4 | Claude Code | 1 | 59.7 | 15/15 | 567,694 | 4,495 | +9 / −5 |
| T4 · Claude Code · Run 2 | T4 | Claude Code | 2 | 75.2 | 15/15 | 496,174 | 4,638 | +5 / −5 |
| T4 · Claude Code · Run 3 | T4 | Claude Code | 3 | 60.6 | 15/15 | 564,618 | 4,458 | +6 / −5 |
| T4 · Codex CLI · Run 1 | T4 | Codex CLI | 1 | 44.2 | 15/15 | 95,606 | 1,850 | +5 / −5 |
| T4 · Codex CLI · Run 2 | T4 | Codex CLI | 2 | 36.3 | 15/15 | 94,160 | 1,363 | +4 / −4 |
| T4 · Codex CLI · Run 3 | T4 | Codex CLI | 3 | 31.9 | 15/15 | 89,421 | 1,168 | +4 / −4 |
| T5 · Claude Code · Run 1 | T5 | Claude Code | 1 | 89.8 | 12/12 | 654,717 | 7,366 | +42 / −2 |
| T5 · Claude Code · Run 2 | T5 | Claude Code | 2 | 90.1 | 12/12 | 585,062 | 7,327 | +47 / −2 |
| T5 · Claude Code · Run 3 | T5 | Claude Code | 3 | 110.6 | 12/12 | 592,485 | 9,232 | +40 / −2 |
| T5 · Codex CLI · Run 1 | T5 | Codex CLI | 1 | 49.3 | 12/12 | 95,745 | 2,054 | +40 / −2 |
| T5 · Codex CLI · Run 2 | T5 | Codex CLI | 2 | 47 | 12/12 | 95,683 | 1,989 | +41 / −2 |
| T5 · Codex CLI · Run 3 | T5 | Codex CLI | 3 | 47.8 | 12/12 | 116,203 | 2,002 | +43 / −2 |
| T6 · Claude Code · Run 1 | T6 | Claude Code | 1 | 147.4 | 12/12 | 1,100,259 | 10,403 | +105 / −16 |
| T6 · Claude Code · Run 2 | T6 | Claude Code | 2 | 111.9 | 12/12 | 660,107 | 9,554 | +87 / −16 |
| T6 · Claude Code · Run 3 | T6 | Claude Code | 3 | 104.4 | 12/12 | 646,651 | 8,335 | +63 / −16 |
| T6 · Codex CLI · Run 1 | T6 | Codex CLI | 1 | 85 | 12/12 | 173,381 | 3,818 | +64 / −25 |
| T6 · Codex CLI · Run 2 | T6 | Codex CLI | 2 | 93.1 | 12/12 | 162,193 | 4,385 | +49 / −18 |
| T6 · Codex CLI · Run 3 | T6 | Codex CLI | 3 | 79 | 12/12 | 157,801 | 3,694 | +44 / −17 |
Download task and grading files
The bundle includes the original tasks (online store code), hidden tests, reference code, execution scripts, a summary of all 18 results (JSON), and diffs for each run.
Download coding-agents-round2-2026-10.zip