Valumigo

LMArena · Image editing

We regularly fetch and display public benchmark data: LMArena user-voted rankings and Epoch AI test scores. Each leaderboard shows its publication and retrieval dates. These scores are not produced by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: A ranking based on votes comparing two models' results when given an image and editing instructions, such as changing the background or colors.

How to read the score: A relative score reflecting which edited results users preferred. It does not score instruction following and preservation of the original separately.

Caveats: To check suitability for the edits you need, compare models directly using the same image and instructions.

Insights from this ranking

  • The 95% confidence intervals for 1st-place gpt-image-2.5-sunburst and 2nd-place gpt-image-2.5-flare do not overlap. This aggregation provides relatively clear evidence that the 1st-place model leads.
  • OpenAI has the most models among the top 10, with 3.
  • The 1st-place model received 47,289 votes, and this leaderboard includes 50 models.

Image editing LMArena

Published 2026-09-30 · Retrieved 2026-10-04

13201380144015001560
Ogpt-image-2.5-sunburst
1522
Ogpt-image-2.5-flare
1478
Ogpt-image-2 (medium)
1461
xgrok-imagine-image-2.0 (low)
1427
Mmai-image-2.6
1427
Metamuse-image
1403
Mmai-image-2.5
1401
ByteDanceseedream-5.0-pro
1394
xgrok-imagine-image-quality (20260519)
1391
Googlegemini-3-pro-image-2k (nano-banana-pro)
1390
Ochatgpt-image-latest-high-fidelity (20251216)
1389
Googlegemini-3.1-flash-image (nano-banana-2) [web-search]
1387
Googlegemini-3-pro-image-preview (nano-banana-pro)
1386
Rreve-2.1
1374
Ogpt-image-1.5-high-fidelity
1370
Alibabaqwen-image-2.1
1366
Rreve-2.0
1359
Iideogram-4.5
1351
Luni-1.1-max
1334
xgrok-imagine-image
1329

Dots show scores; horizontal lines show 95% confidence intervals. Overlapping intervals suggest similar performance.

View table (top 50)
RankModelDeveloperScore95% confidence intervalVotes
1Ogpt-image-2.5-sunburstOpenAI15221517–152847,289
2Ogpt-image-2.5-flareOpenAI14781473–148444,792
3Ogpt-image-2 (medium)OpenAI14611458–1465293,819
4xgrok-imagine-image-2.0 (low)xAI14271422–143236,022
5Mmai-image-2.6Microsoft AI14271422–143241,512
6Metamuse-imageMeta14031399–1408130,653
7Mmai-image-2.5Microsoft AI14011397–1405190,356
8ByteDanceseedream-5.0-proByteDance13941391–1397270,020
9xgrok-imagine-image-quality (20260519)xAI13911386–139738,256
10Googlegemini-3-pro-image-2k (nano-banana-pro)Google13901388–1393638,610
11Ochatgpt-image-latest-high-fidelity (20251216)OpenAI13891386–1391617,941
12Googlegemini-3.1-flash-image (nano-banana-2) [web-search]Google13871384–1391216,640
13Googlegemini-3-pro-image-preview (nano-banana-pro)Google13861383–1389540,778
14Rreve-2.1Reve13741368–138020,823
15Ogpt-image-1.5-high-fidelityOpenAI13701367–1372642,272
16Alibabaqwen-image-2.1Alibaba13661360–137214,617
17Rreve-2.0Reve13591352–136519,291
18Iideogram-4.5Ideogram13511343–13605,441
19Luni-1.1-maxLuma AI13341329–133943,451
20xgrok-imagine-imagexAI13291327–1332737,207
21Luni-1.1Luma AI13151311–1318177,667
22Googlegemini-3.1-flash-lite-image (nano-banana-2-lite)Google13141309–131865,764
23Alibabaqwen-image-2.0-pro-2026-06-22Alibaba13031299–130856,781
24Thunyuan-image-3.0-instructTencent13031299–1306331,592
25Alibabawan2.7-image-proAlibaba13031299–130743,992
26ByteDanceseedream-4.5ByteDance13021300–13041,266,531
27Alibabawan2.7-imageAlibaba13011297–130544,814
28ByteDanceseedream-5.0-liteByteDance12941291–1297446,300
29Googlegemini-2.5-flash-image-preview (nano-banana)Google12931291–129511,355,727
30ByteDanceseedream-4-2kByteDance12701264–1277210,076
31Rreve-v1.1Reve12621259–1264672,675
32Bflux-2-maxBlack Forest Labs12621258–1265375,959
33Klingkling-image-o1Kling12511248–1255143,692
34Bflux-2-proBlack Forest Labs12451243–1248624,352
35Alibabaqwen-image-editAlibaba12411239–12442,022,077
36Rreve-v1Reve12351230–1239395,314
37Alibabaqwen-image-edit-2511Alibaba12341231–1237485,308
38Alibabawan2.6-imageAlibaba12311228–1233664,189
39Bflux-2-flexBlack Forest Labs12251222–1228443,953
40Bflux-2-klein-9bBlack Forest Labs12251222–1228540,833
41Bflux-2-devBlack Forest Labs12251221–1229254,450
42ByteDanceseedream-4-high-res-falByteDance12171214–12201,229,943
43?p-image-edit?12111208–1214359,965
44ByteDanceseedream-4-falByteDance12101204–1216152,114
45Rreve-v1.1-fastReve12071204–1210553,361
46Rreve-edit-fastReve11981194–1202230,289
47Bflux-2-klein-4bBlack Forest Labs11871185–1190541,422
48Bflux-1-kontext-maxBlack Forest Labs11811178–1183390,877
49Alibabawan2.5-i2i-previewAlibaba11811178–1183715,740
50Bflux-1-kontext-proBlack Forest Labs11761174–11796,404,790

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarks