Valumigo

OpenRouter · Customer service agents (τ-bench airline)

We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: Results from OpenRouter's evaluation of τ-bench (verified version, airline domain). It tests whether agents use tools to handle requests such as booking changes according to the rules in airline customer service scenarios.

How to read the score: Higher success rates (%) are better. You can also compare average cost per task.

Caveats: It covers scenarios in only one domain (airlines), so agent performance on other types of work may differ.

Insights from this ranking

  • The highest score is 79.9%, achieved by Claude Fable 5.
  • The gap between 1st and 5th place is 1.4%p.
  • Anthropic has the most models among the top 10, with 4.
  • Among the top 10, Step 3.7 Flash has the lowest average cost per task ($0.0201, score 77.3%).

Customer service agents (τ-bench airline) OpenRouter

Published 2026-10-04 · Retrieved 2026-10-05

AnthropicClaude Fable 5
79.9%
XiaomiMiMo-V2.6-Flash
79.3%
AnthropicClaude Fable 5.1
79.2%
AnthropicClaude Opus 5
78.6%
GoogleGemini 3.7 Flash
78.5%
ANova Micro 1.0
78.0%
SStep 3.7 Flash
77.3%
AnthropicClaude Opus 4.5
77.1%
OGPT-5.1
77.1%
AlibabaQwen3.8 27B
77.0%
ZGLM 5
77.0%
DeepSeekDeepSeek V4 Pro 0813
77.0%
AlibabaQwen3.5 397B A17B
76.9%
GoogleGemini 3.5 Flash Lite
76.9%
GoogleGemma 4 31B
76.7%
AnthropicClaude Opus 4.8
76.4%
DeepSeekDeepSeek V4 Pro 0423
76.4%
DeepSeekDeepSeek V4.1 Flash
76.2%
ZGLM 5.3
76.1%
AlibabaQwen3.5-122B-A10B
76.1%

The chart shows only the highest-scoring reasoning setting, such as high or max, for each model. The table shows rankings for all settings.

View table (top 60)
RankModelDeveloperScoreAverage cost per task
1#1AnthropicClaude Fable 5Anthropic79.9%$0.9265
2#2XiaomiMiMo-V2.6-FlashXiaomi79.3%$0.0204
3#3AnthropicClaude Fable 5.1Anthropic79.2%$0.7162
4#4AnthropicClaude Opus 5Anthropic78.6%$0.4868
5#5GoogleGemini 3.7 FlashGoogle78.5%$0.1149
6#6ANova Micro 1.0Amazon78%$1.2758
7#7SStep 3.7 FlashStepFun77.3%$0.0201
8#8AnthropicClaude Opus 4.5Anthropic77.1%$0.5319
9#9OGPT-5.1OpenAI77.1%$0.4149
10#10AlibabaQwen3.8 27BAlibaba77%$0.0811
11#11ZGLM 5Z.ai77%$0.0447
12#12DeepSeekDeepSeek V4 Pro 0813DeepSeek77%$0.1001
13#13AlibabaQwen3.5 397B A17BAlibaba76.9%$0.1697
14#14GoogleGemini 3.5 Flash LiteGoogle76.9%$0.0978
15#15GoogleGemma 4 31BGoogle76.7%$0.0282
16#16AnthropicClaude Opus 4.8Anthropic76.4%$0.4977
17#17DeepSeekDeepSeek V4 Pro 0423DeepSeek76.4%$0.0399
18#18DeepSeekDeepSeek V4.1 FlashDeepSeek76.2%$0.0315
19#19ZGLM 5.3Z.ai76.1%$0.062
20#20AlibabaQwen3.5-122B-A10BAlibaba76.1%$0.169
21#21ZGLM 5.3Z.ai76%$0.016
22#22AnthropicClaude Opus 4.7Anthropic75.9%$0.4192
23#23GoogleGemini 3.1 Pro PreviewGoogle75.8%$0.3592
24#24ZGLM 5.3 FlashZ.ai75.6%$0.0067
25#25AlibabaQwen3.8 2.4T A95BAlibaba75.5%$0.1978
26#26ZGLM 5.1Z.ai75.4%$0.0778
27#27xGrok 4.6xAI75.3%$0.3309
28#28AnthropicClaude Sonnet 5Anthropic75.3%$0.1994
29#29OGPT-5.4OpenAI75.3%$0.2916
30#30GoogleGemini 3 Flash PreviewGoogle75.3%$0.1221
31#31OGPT-5.3-CodexOpenAI75.3%$0.2997
32#32OGPT-5.5OpenAI75.1%$0.4875
33#33DeepSeekDeepSeek V4 Flash Vision ExpDeepSeek74.8%$0.0302
34#34AnthropicClaude Opus 4.6Anthropic74.7%$0.4898
35#35AnthropicClaude Sonnet 5.5Anthropic74.7%$0.1787
36#36GoogleGemini 3.6 FlashGoogle74.7%$0.2141
37#37OGPT-5.6 SolOpenAI74.7%$0.184
38#38MetaMuse Glimmer 30BMeta74.7%$0.0374
39#39AnthropicClaude Sonnet 4.6Anthropic74.6%$0.3238
40#40DeepSeekDeepSeek V4 Flash 0423DeepSeek74.4%$0.0111
41#41AnthropicClaude Opus 5.5Anthropic74.3%$0.3458
42#42AlibabaQwen3.6 27BAlibaba74.3%$0.1343
43#43Moonshot AIKimi K2.6Moonshot AI74.1%$0.0733
44#44MiniMaxMiniMax M3MiniMax74.1%$0.0235
45#45DeepSeekDeepSeek V3.2DeepSeek74%$0.0299
46#46OGPT-5.6 Sol ProOpenAI73.7%$1.288
47#47XiaomiMiMo-V2.5-ProXiaomi73.6%$0.0304
48#48DeepSeekDeepSeek V4 Flash 0731DeepSeek73.4%$0.0092
49#49OGPT-5.2OpenAI73.3%$0.2413
50#50OGPT-5.6 TerraOpenAI73.1%$0.1054
51#51GoogleGemini 3.5 FlashGoogle73.1%$0.3988
52#52OGPT-6 Astra ProOpenAI72.7%$2.9136
53#53ZGLM 5.2Z.ai72.7%$0.034
54#54Moonshot AIKimi K2.7 CodeMoonshot AI72.6%$0.0736
55#55AnthropicClaude Sonnet 4.5Anthropic72.4%$0.3328
56#56ZGLM 4.6Z.ai72.1%$0.0499
57#57ZGLM 4.7Z.ai72.1%$0.0581
58#58NVIDIANemotron 3 UltraNVIDIA72%$0.1515
59#59AlibabaQwen3.7 MaxAlibaba72%$0.4029
60#60GoogleGemini 3.1 Flash LiteGoogle72%$0.0932

Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). Licensed under CC BY 4.0.

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings). Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). Source: Design Arena (www.designarena.ai) via OpenRouter (openrouter.ai/rankings). Licensed under CC BY 4.0. Source: OpenRouter (openrouter.ai/rankings), as of 2026-10-05. Licensed under CC BY 4.0. https://openrouter.ai/rankings