数据同步时间: 2026-08-17 20:11:32
排名 | 模型 | Elo 评分 Arena Elo 评分:基于用户盲测两两对决的胜负关系估算出的模型综合能力评分 | ||||
|---|---|---|---|---|---|---|
1 | gpt-image-2 (medium)OpenAI • Proprietary | 1381±5 | ||||
2 | mai-image-2.6-previewMicrosoft AI • Proprietary | 1336±11 | ||||
3 | grok-imagine-image-2.0 (low)SpaceXAI • Proprietary | 1316±12 | ||||
4 | reve-2.1Reve • Proprietary | 1302±8 | ||||
5 | muse-imageMeta • Proprietary | 1282±7 | ||||
6 | reve-2.0Reve • Proprietary | 1270±6 | ||||
7 | gemini-3.1-flash-image (nano-banana-2) [web-search]Google • Proprietary | 1264±5 | ||||
8 | seedream-5.0-proBytedance • Proprietary | 1258±5 | ||||
9 | qwen-image-3.0-proAlibaba • Proprietary | 1257±9 | ||||
10 | mai-image-2.5Microsoft AI • Proprietary | 1256±4 | ||||
11 | gemini-3.1-flash-lite-image (nano-banana-2-lite)Google • Proprietary | 1251±6 | ||||
12 | gemini-3-pro-image-2k (nano-banana-pro)Google • Proprietary | 1246±3 | ||||
13 | gpt-image-1.5-high-fidelityOpenAI • Proprietary | 1239±3 | ||||
14 | gemini-3-pro-image-preview (nano-banana-pro)Google • Proprietary | 1232±5 | ||||
15 | grok-imagine-image-qualitySpaceXAI • Proprietary | 1228±4 | ||||
16 | ideogram-4.0-qualityIdeogram • Open | 1204±5 | ||||
17 | qwen-image-2.0-pro-2026-06-22Alibaba • Proprietary | 1191±6 | ||||
18 | uni-1.1-maxLuma AI • Proprietary | 1188±6 | ||||
19 | mai-image-2Microsoft AI • Proprietary | 1183±5 | ||||
20 | uni-1.1Luma AI • Proprietary | 1181±5 | ||||
21 | Cosmos3-Super-Text2Image (Agentic)Nvidia • Open | 1175±10 | ||||
22 | grok-imagine-imageSpaceXAI • Proprietary | 1171±3 | ||||
23 | recraft-v4.1-utility-proRecraft • Proprietary | 1169±11 | ||||
24 | flux-2-maxBlack Forest Labs • Proprietary | 1162±4 | ||||
25 | grok-imagine-image-proSpaceXAI • Proprietary | 1161±4 | ||||
显示第 1 到 25 条,共 77 条记录
每页显示:
1 / 4
排名计算方法: 系统利用在竞技场中收集到的海量双盲人类盲测对抗数据,构建 Bradley-Terry 模型以极大似然估计模型在各个单项分类下的 Elo 相对竞技值。95% 置信区间 (CI) 表明当误差区间发生重叠时,模型的真实相对实力可能不存在统计学上的显著性差异。