前沿智能 · 版本化观测

AI 模型观测站

收录模型29frontier models
Benchmark 目录727 capability families
AA IntelligenceClaude Opus 561 · Anthropic
输出速度最快Gemini 3.6 Flash217.2 output tok/s
最佳性价比DeepSeek V4 Flash 0731$0.03 / AA task
01

前沿模型排行

能力、系统表现与人工偏好分开排序,避免把不同评价对象混成一个总分。

更新于 2026 年 8 月 10 日
#模型综合能力AAArena编程组合Agent速度对比
AA 61Arena 1488综合能力 61.0
AA 60Arena 1507综合能力 60.0
AA 59Arena 1482综合能力 59.0
AA 57Arena 1485综合能力 57.0
AA 56Arena 1473综合能力 56.0
AA 55Arena 1467综合能力 55.0
AA 55Arena 1477综合能力 55.0
AA 54Arena 1468综合能力 54.0
AA 53Arena 1463综合能力 53.0
AA 51Arena 1471综合能力 51.0
当前视图综合能力

N/A 保持缺失,不参与当前排序。

02

七维能力画像

雷达图只使用有公开分数的能力轴;缺失数据保持 N/A,不做补零或估算。

推理与知识数学与科学编程与软件工程Agent 与工具专业知识工作多模态理解长上下文与记忆
数据覆盖86%
19 / 22 项核心指标 · 广泛覆盖
03

Benchmark 组合面板

按能力族查看原始分数、版本、评测对象与评分方法;最多对比三个模型。

0255075100DeepSWEv1.1Terminal-Bench2.1ProgramBench2026FrontierSWE2026-07SWE-Marathonv1.1PostTrainBenchv1.1SciCode2026Vals · LiveCodeBench2026Vals · IOI2026Vals · Code Migration2026Vals · Vibe Code Bench2026LiveBench · Code Generation2026-06-25LiveBench · Code Completion2026-06-25LiveBench · Agentic Python2026-06-25LiveBench · Agentic JavaScript2026-06-25LiveBench · Agentic TypeScript2026-06-25

表格可横向滑动查看全部指标

已隐藏 2 项:对比中的模型都还没有它的数据 · SWE-Bench Pro · SWE-EVO

模型DeepSWEv1.1Terminal-Bench2.1ProgramBench2026FrontierSWE2026-07SWE-Marathonv1.1PostTrainBenchv1.1SciCode2026Vals · LiveCodeBench2026Vals · IOI2026Vals · Code Migration2026Vals · Vibe Code Bench2026LiveBench · Code Generation2026-06-25LiveBench · Code Completion2026-06-25LiveBench · Agentic Python2026-06-25LiveBench · Agentic JavaScript2026-06-25LiveBench · Agentic TypeScript2026-06-25
GPT-5.6 Sol72.7%benchmark +589.51%independent +71.5%independent +171.3%vendor39%vendor34.6%vendor56.9%independent +682.604%independent86.667%independent52.916%independent80.495%independent83.099%benchmark84.783%benchmark55%benchmark63.636%benchmark50%benchmark
Claude Fable 569.9%benchmark +583.82%benchmark +42%independent +189%benchmark +135%vendor41.4%vendor60.2%independent +189.778%independent72.25%independent55.065%independent90.352%independent91.549%benchmark80.435%benchmark65%benchmark68.182%benchmark53.333%benchmark
Kimi K368.5%benchmark +185.02%independent +32%independent +181.2%vendor42%vendor36.6%vendor58.7%independent +287.185%independentN/AN/A84.963%independent80.282%benchmark82.609%benchmark65%benchmark68.182%benchmark53.333%benchmark
04

Token 价格

每百万 Token 的厂商标价,取自存档、可溯源;OpenRouter 实时价仅作对照,不覆盖目录数字。

GPT-5.6 SolOpenAI
输入$5
输出$30
7:2:1 混合价$9.55
Claude Fable 5Anthropic
输入$10
输出$50
7:2:1 混合价$17.1
Kimi K3Moonshot
输入$3
输出$15
7:2:1 混合价$5.13
05

评测目录

核心指标进入能力面板;观察指标先展示、不计综合排行;历史指标仅作趋势参考。

22 CORE · 46 OBSERVE
06

数据来源与可比性

21 已接入0 接入队列
读取于 2026-08-08已接入 · 444
Artificial Analysisindependent capability, speed, price and GDPval index
27 Jul 2026已接入
LM Arenalarge-scale human preference signal
读取于 2026-08-10已接入 · 415
Vals AIindependent professional-work evaluations
读取于 2026-08-09已接入 · 210
Epoch AIFrontierMath and benchmark methodology cross-checks
读取于 2026-08-07已接入 · 135
ARC Prize verifiedverified ARC-AGI fluid-intelligence results
读取于 2026-08-01已接入 · 23
Terminal-Benchbenchmark-native verified terminal-agent runs
读取于 2026-08-08已接入 · 49
DeepSWE official leaderboardbenchmark-native long-horizon coding runs
读取于 2026-07-31已接入 · 14
Scale LabsMCP-Atlas and SWE-Bench Pro leaderboards
读取于 2026-08-08已接入 · 11
OSWorld 2.0execution-based long-horizon computer use
读取于 2026-08-07已接入 · 42
Agents' Last Examverifiable real-world professional tasks, read by script
读取于 2026-07-31已接入 · 9
FrontierSWEfrontier engineering tasks, Mean@5
读取于 2026-07-31已接入 · 4
Mercor APEX-Agentsbanking, consulting and legal deliverables
读取于 2026-07-31已接入 · 10
Toolathlon-Verifiedlong-horizon cross-application tool use
读取于 2026-08-07已接入 · 1
MMMUexpert multimodal reasoning
读取于 2026-07-31已接入 · 1
SWE-benchofficial software-engineering leaderboard and archive
读取于 2026-08-08已接入 · 575
LiveBenchobjective, contamination-limited general evaluation
live feed已接入
OpenRouter Models APIprovider pricing and context-window metadata
评测于 2026-07-31已接入 · 36
Google DeepMind model cardsvendor results with harness notes
评测于 2026-07-31已接入 · 15
DeepSeek V4 model cardsvendor results with effort modes
评测于 2026-07-23已接入 · 125
Kimi K3 release tablecomparison seed only, not the global standard
读取于 2026-08-05已接入 · 13
Qwen release tablesvendor results, captured per release

“已接入”是数出来的,不是声明的:只有当存档里确实有来自该来源的观测行时才算接入,数字即行数。“接入队列”只是下一批采集目标,不参与现有分数。日期同样是数出来的:“读取于”是本项目最后一次抄录该来源的时间,“评测于”是该来源已发布的最新结果时间,两者不可互换。十六个批次中有十一个靠人工抄录、无法与上游自动比对,所以超过 30 天未读取的来源标为“待复核”——最近读过但评测日期很旧,说明的是该榜单本身没有更新。Arena 衡量人工偏好,系统类结果还同时反映 harness、工具与预算。