Claude Code vs Codex: We Ran the Same 3 Coding Tasks 3 Times Each
When choosing a coding agent, the main question is which one does the same job better and faster. We created 3 small tasks, gave Claude Code and Codex CLI identical instructions 3 times each, and scored them using hidden tests the agents could not see.
Run on 2026-10-04
At a Glance
- Accuracy: Both tools passed every hidden test in all 9 runs. Tasks of this size revealed no difference in output accuracy.
- Speed: Based on the median for each task, Codex (GPT-6.1 Sol) finished about 1.8~2 times faster than Claude Code (Claude Opus 5). The difference was 16~43 seconds per task.
- Tokens: Based on the median for each task, Claude Code processed about 4.2~5.7 times as many input tokens (including cache reuse) and about 3.5~4.3 times as many output tokens as Codex. When we disabled all MCP tool connections and ran T1 once more, the input token count was similar (about 397,000), but we could not determine the cause of the difference.
- Code: Both tools fixed the same three lines in the bug-fixing task. In the performance task, Codex preserved even a very rare behavior in the original code (when id is NaN), while Claude Code wrote shorter code with comments explaining its reasoning.
- Gemini: We also ran Gemini CLI under the same conditions, but it stopped immediately during authentication. This was because Gemini CLI service for individual, Google AI Pro, and Ultra accounts ended on June 18, 2026.
Results
Time is the median of 3 runs; parentheses show the fastest~slowest values.
| Task | Tool | Passed (all 3 runs) | Elapsed Time | Input Tokens | Output Tokens | Lines Changed (+/−) |
|---|---|---|---|---|---|---|
| T1 Bug Fixing | Claude Code | 3/3 | 36.6s (36.1–48.7) | 356,196 | 1,913 | +3 / −3 |
| Codex CLI | 3/3 | 20.3s (19.1–32.5) | 85,176 | 486 | +3 / −3 | |
| T2 Feature Implementation | Claude Code | 3/3 | 69.8s (66–73.6) | 615,019 | 5,379 | +64 / −2 |
| Codex CLI | 3/3 | 38.2s (33.2–39.1) | 107,428 | 1,265 | +53 / −2 | |
| T3 Performance Optimization | Claude Code | 3/3 | 87.2s (85.3–104.3) | 561,732 | 6,085 | +26 / −12 |
| Codex CLI | 3/3 | 44.2s (39.8–55.1) | 109,749 | 1,742 | +21 / −6 |
Input tokens largely include cached tokens reused when rereading earlier conversation steps. The two CLIs count tokens differently, so treat this only as a rough comparison of scale.
Claude Code also reports a cost calculated at standard API rates for each run (about 0.53~1.04 USD per task in this test). Subscription accounts do not pay this amount separately; usage counts against the plan's limits. Codex does not report an equivalent value, so we did not compare it.
3 Tasks
T1 Bug Fixing
The date calculation code for accommodation bookings (dates.js) has 3 bugs: months are not counted from 0, stays are counted as 1 day too long, and the final day is omitted from business-day calculations. Hidden tests check leap years, year-end dates, daylight saving time transitions, and other cases.
T2 Feature Implementation
The task is to build a CSV parser from scratch following the README. Hidden tests check commas and line breaks inside quotes, escaped double quotes (""), Windows line endings (CRLF), empty fields, and errors for unclosed quotes.
T3 Performance Optimization
Code that aggregates 200,000 orders (deduplication, customer rankings, and daily totals) is too slow. The task is to make it finish within 1 second without changing the results. Hidden tests check whether the results match the original code, tie order is preserved, and 300,000 orders also finish within 1 second.
Instructions Used for Every Run
Only the sentence after 'Task:' in the text below changes for each task.
This folder is a small Node.js project (ES modules, no dependencies). Read the code, README.md if present, and the tests. Task: {task} Do not modify existing test files. When you are done, make sure `node --test` passes.- T1 Bug Fixing —
Fix the bugs in dates.js so that every function behaves as its comment describes and the tests pass. - T2 Feature Implementation —
Implement parseCSV in csv.js according to README.md. - T3 Performance Optimization —
Make report.js fast enough as described in README.md, without changing any results.
How We Measured
- Run on 2026-10-04 (evening, Korea time), Apple M1 Pro · macOS 26.6 · Node.js 22.
- Claude Code 2.1.250 — default model Claude Opus 5 (Claude Haiku 4.5 was also called for brief supporting tasks). Codex CLI 0.159.2 — model GPT-6.1 Sol, reasoning effort medium.
- Both tools ran in non-interactive mode (claude -p, codex exec) while signed in to the site operator's personal subscription accounts. File edits and node execution were automatically allowed.
- We copied the original task into a fresh folder for every run. To reduce order effects, Claude Code ran first in rounds 1 and 3, and Codex ran first in round 2.
- Time was measured from the moment the command started until it finished. Token counts were copied directly from each CLI's run report.
Scoring Method
- Each task had 3 tests visible to the agent and hidden tests (9~14) added only during scoring. The hidden tests were first verified against reference code written by the site operator.
- Before scoring, we restored the visible test files to their originals. The agents did not modify test files in any of the 18 runs.
- Because date bugs may appear only in regions with daylight saving time, we ran the tests in both New York and Berlin time zones and recorded the lower score.
How Did the Code Differ?
- T1: Each tool ran 3 times, and all 6 runs changed the same three lines (month calculation, removal of +1 from the number of nights, and inclusion of the last day). Every run also had the same number of changed lines: 3 added and 3 deleted.
- T2: Both used parsers that read one character at a time. Claude Code added 63~67 lines, while Codex added 53~55 lines; Claude Code included slightly more explanatory comments.
- T3: Both replaced repeated searches from the start of a list with hash-based lookups (Map·Set). Codex even preserved the original behavior of not treating an id of NaN as a duplicate. Claude Code left a comment explaining why the order of first appearance is preserved for ties.
Why Was Gemini Excluded?
- We also ran Gemini CLI 0.58.0 with the same tasks and instructions, but all three tasks ended with an authentication error within 4 seconds. Error message: "This client is no longer supported for Gemini Code Assist for individuals. To continue using Gemini, please migrate to the Antigravity suite of products".
- According to Google's official announcement, Gemini CLI and the Gemini Code Assist IDE extension no longer process requests from Gemini Code Assist for individuals, Google AI Pro, and Google AI Ultra tiers as of June 18, 2026. Users are directed to Antigravity and Antigravity CLI (agy) instead.
- Running the same measurements with Antigravity CLI requires a one-time interactive login, so we did not include it in this test.
Limitations of This Test
- These were 3 small, single-file tasks, each run only 3 times. Results may differ for large repositories, ambiguous requirements, or changes across multiple files.
- Elapsed time varies with server load, network conditions, and model and CLI versions. Alternating runs on the same computer on the same day reduced these effects but did not eliminate them.
- Codex ran with reasoning effort medium (the site operator's setting), while Claude Code used default settings. Changing the reasoning effort or model will affect speed and token usage.
- We did not measure how much of the subscription plans' usage allowances was consumed.
Which Tool Should You Use?
The following reflects the site operator's opinion based on these test results.
- For clearly scoped work such as small bug fixes or function implementation, this test showed no difference in the tools' outputs. It makes sense to try the one included in the subscription you already pay for (ChatGPT or Claude) first.
- If you need to run the same work repeatedly and quickly, Codex was about 2 times faster in this test and used fewer tokens.
- Claude Code tended to verify its work through more steps and leave comments explaining its reasoning. This may be an advantage when people need to read and review the resulting code.
- Individual Gemini users need to move to Antigravity CLI because Gemini CLI no longer works for them.
All 18 Runs
| Task | Tool | Round | Seconds | Tests Passed | Input Tokens | Output Tokens | Lines Changed (+/−) |
|---|---|---|---|---|---|---|---|
| T1 | Claude Code | 1 | 48.7 | 15/15 | 356,196 | 2,540 | +3 / −3 |
| T1 | Claude Code | 2 | 36.1 | 15/15 | 291,026 | 1,913 | +3 / −3 |
| T1 | Claude Code | 3 | 36.6 | 15/15 | 537,029 | 1,781 | +3 / −3 |
| T1 | Codex CLI | 1 | 32.5 | 15/15 | 107,102 | 776 | +3 / −3 |
| T1 | Codex CLI | 2 | 19.1 | 15/15 | 85,176 | 485 | +3 / −3 |
| T1 | Codex CLI | 3 | 20.3 | 15/15 | 70,801 | 486 | +3 / −3 |
| T2 | Claude Code | 1 | 66 | 17/17 | 551,891 | 5,379 | +63 / −2 |
| T2 | Claude Code | 2 | 73.6 | 17/17 | 618,636 | 5,603 | +64 / −2 |
| T2 | Claude Code | 3 | 69.8 | 17/17 | 615,019 | 5,240 | +67 / −2 |
| T2 | Codex CLI | 1 | 38.2 | 17/17 | 107,428 | 1,228 | +53 / −2 |
| T2 | Codex CLI | 2 | 33.2 | 17/17 | 87,432 | 1,265 | +53 / −2 |
| T2 | Codex CLI | 3 | 39.1 | 17/17 | 107,461 | 1,268 | +55 / −2 |
| T3 | Claude Code | 1 | 87.2 | 12/12 | 428,501 | 5,398 | +19 / −11 |
| T3 | Claude Code | 2 | 104.3 | 12/12 | 628,309 | 6,880 | +29 / −15 |
| T3 | Claude Code | 3 | 85.3 | 12/12 | 561,732 | 6,085 | +26 / −12 |
| T3 | Codex CLI | 1 | 44.2 | 12/12 | 89,678 | 1,742 | +20 / −8 |
| T3 | Codex CLI | 2 | 55.1 | 12/12 | 112,393 | 2,561 | +21 / −6 |
| T3 | Codex CLI | 3 | 39.8 | 12/12 | 109,749 | 1,593 | +21 / −5 |
Download Task and Scoring Files
The bundle includes the original tasks, hidden tests, run scripts, a summary of all 18 runs (JSON), and the code changes from each run (diff). You can repeat the measurements using the same method.
Download coding-agents-2026-10.zip