Hands-on Test Reports
We publish only results the site operator has directly run and measured. Tasks, prompts, scoring methods, and original outputs are provided so anyone can repeat the tests using the same method.
Services we have not measured are not included here. For services we do not personally use, the service guide pages summarize only official sources.
Claude Code vs Codex: We Ran the Same 3 Coding Tasks 3 Times Each
We gave both coding agents identical bug-fixing, feature implementation, and performance optimization tasks, then scored them with hidden tests. We publish all 18 runs, including time, tokens, and lines changed.
Run on 2026-10-04
Can Image AI Render Korean (Hangul) Correctly? We Gave 4 Models the Same Prompt
We used 4 models to generate images containing Korean (Hangul) text, such as signs and menus, and scored each character. We publish the resulting images and credits used as recorded.
Run on 2026-10-04