{"meta":{"site_name":"AI-Gallery","version":"0.2.0","generated_at":"2026-08-28","data_cutoff":"2026-08-28","note":"分数与价格为仓库内手工维护的静态快照，来源与日期见每条记录。2026 年当前代模型数据于 2026-08-28 联网核对；被替代的旧代模型保留 2025-12 快照供对照。AA 指数为 v4.x 口径，旧代模型的 v3 指数已移除以免混比；Terminal-Bench 列混合 2.0 / 2.1 版本，详情页来源处可见。","site_url":"https://tan-zhuo.github.io/AI-Gallery","author":"谭卓","author_url":"https://tanzhuo.xyz","repo_url":"https://github.com/tan-zhuo/AI-Gallery"},"models":[{"aliases":["Baichuan2-13B-Chat"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"knowledge":"MMLU 59.2%（官方基座）。","chinese":"C-Eval 58.1%、CMMLU 62.0%（官方基座）。"},"complete":false,"id":"baichuan2-13b","name":"Baichuan2-13B","name_zh":"百川 2 · 13B","vendor":"Baichuan","vendor_zh":"百川智能","family":"Baichuan2","license":"Baichuan 2 Community License（Apache-2.0 代码）","license_commercial":"restricted","weights_url":"https://huggingface.co/baichuan-inc/Baichuan2-13B-Chat","released_at":"2023-09-06","architecture":{"type":"dense","total_params":"13.9B","total_params_b":13.9,"layers":40,"hidden_size":5120,"vocab_size":125696,"kv_heads":40,"head_dim":128,"attention":"MHA（40 头）+ ALiBi","notes":"2.6T 中英 token；ALiBi 位置编码；上下文 4K。","undisclosed":false},"context":{"max_tokens":4096,"display":"4K"},"memory":{"weight_gb":{"bf16":27.8,"q8":14.7,"q4":8.1},"estimated":true,"kv_per_token_kib":800,"kv_note":"MHA：40×128×2×40×2 B = 800 KiB/token。","ref_hw_24gb":"BF16 27.8 GB 超 24GB；int8 官方版 14 GB 可跑","ref_hw_80gb":"BF16 单卡","ref_hw_8x80gb":"过剩"},"links":{"official":"https://www.baichuan-ai.com/","github":"https://github.com/baichuan-inc/Baichuan2","paper":"https://arxiv.org/abs/2309.10305","hf":"https://huggingface.co/baichuan-inc/Baichuan2-13B-Chat"},"copy":{"one_liner":"2023 年中文开源热潮的代表，13B 中文当年领先。","highlights":["2.6T 中英 token，C-Eval / CMMLU 当年同尺寸最高","公开全部训练中间检查点，学术价值高","官方 int8 / int4 量化版"],"pitfalls":["上下文 4K，MHA KV 大","自定义模型代码，新框架兼容差","商用需登记；百川后续转向闭源"],"logic_ability":"2023 年 13B 水平：MMLU 59.2%、C-Eval 58.1%（官方基座）。","best_for":["历史对照 / 中文预训练研究"],"not_for":["新项目"]},"ecosystem":{"engines":["vLLM","llama.cpp"],"finetune":"LoRA 可行","zh_docs":"有"}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"claude-2","name":"Claude 2","aliases":["claude-2.0","claude-2.1"],"vendor":"Anthropic","family":"Claude 2","superseded_by":"claude-3-opus","released_at":"2023-07-11","modalities":["text"],"reasoning_mode":"none","context":{"max_tokens":100000,"display":"100K（2.1 版 200K）","max_output":4096},"pricing":{"input_per_m":8,"output_per_m":24,"currency":"USD","source":"Anthropic 定价页（Claude 2.1）","as_of":"2023-11-21","note":"Claude 2.0 发布价 $11.02 / $32.68"},"links":{"official":"https://www.anthropic.com/news/claude-2","pricing":"https://www.anthropic.com/pricing"},"copy":{"one_liner":"100K 上下文的早期 Claude，长文档处理先驱。","highlights":["100K 上下文（2.1 版 200K），当年远超 GPT-4","律师考试 76.5%、GSM8K 88.0%","2.1 版幻觉率减半、引入工具调用（beta）"],"pitfalls":["过度拒答，\"Claude 太谨慎\"成梗","HumanEval 71.2%，代码弱于 GPT-4","无视觉、无 JSON 模式"],"logic_ability":"逻辑推理接近 GPT-3.5 到 GPT-4 之间，长文档归纳强，数学一般。","best_for":["历史对照"],"not_for":["新项目"]}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。2024-10-22 版（同名升级，社区称 3.6）首次加入 computer use（beta）。"},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{"coding":"HumanEval 92.0%（0620）；SWE-bench Verified 33.4%（0620）→ 49.0%（1022）（官方）。","reasoning":"GPQA Diamond 59.4%（0620）→ 65.0%（1022）（官方）。","math":"MATH 71.1%、GSM8K 96.4%（0620，官方）。","agent":"TAU-bench retail 69.2%、airline 46.0%（1022，官方）；OSWorld 14.9%。","multimodal":"MMMU 68.3%（0620）→ 70.4%（1022）（官方）。","chinese":"中文良好，缺独立榜单。"},"complete":true,"id":"claude-3-5-sonnet","name":"Claude 3.5 Sonnet","name_zh":"Claude 3.5 Sonnet（含 10-22 新版）","aliases":["claude-3-5-sonnet-20240620","claude-3-5-sonnet-20241022","Claude 3.6"],"vendor":"Anthropic","family":"Claude 3.5","superseded_by":"claude-3-7-sonnet","released_at":"2024-06-20","modalities":["text","image","tools","computer-use"],"reasoning_mode":"none","context":{"max_tokens":200000,"display":"200K","max_output":8192},"pricing":{"input_per_m":3,"output_per_m":15,"currency":"USD","source":"Anthropic 定价页","as_of":"2024-10-22","note":"缓存写 $3.75 / 读 $0.30"},"links":{"official":"https://www.anthropic.com/news/claude-3-5-sonnet","paper":"https://www.anthropic.com/news/3-5-models-and-computer-use","pricing":"https://www.anthropic.com/pricing"},"copy":{"one_liner":"确立 Claude 编程口碑的一代，首个 computer use 模型。","highlights":["10-22 版 SWE-bench Verified 49.0%，当时闭源第一","Artifacts + 代码体验让其成为开发者首选","首个官方 computer use（OSWorld 14.9%）"],"pitfalls":["同名两版（0620 / 1022）行为差异大，易混淆","无推理模式，AIME 类竞赛数学弱","最大输出 8K"],"logic_ability":"工程型推理在非推理模型中最强，代码修改与多文件重构可靠；考试型推理 GPQA 59.4% → 65.0%（新版）。常见问题：长代理任务中偶尔\"偷懒\"省略代码。","best_for":["代理编程（当年）","文档分析与写作","历史对照"],"not_for":["竞赛数学","新项目"]},"sheet":{"architecture_md":"**未披露**。Anthropic 未公开参数量。\n\n已知信息：\n- 200K 上下文，最大输出 8K（0620 版初期 4K，后放开到 8K）\n- 支持视觉输入、工具调用、提示缓存（2024-08）\n- 1022 版新增 computer use API（beta）：截图 + 鼠标键盘动作\n- 速度约为 Claude 3 Opus 的 2 倍","memory_md":"无自建选项。\n\n成本参考：输入 $3 / 输出 $15；缓存写 $3.75、读 $0.30；Batch API 半价。","training_md":"- 训练细节未披露，知识截止 2024-04\n- 无推理模式\n- 支持工具调用、JSON 输出（通过工具）、PDF 输入（1022 版）\n- 1022 版与 0620 版同价同名，通过日期后缀区分","ecosystem_md":"- 官方 API、AWS Bedrock、Google Vertex AI\n- 不支持微调\n- 中文文档：弱（官方英文）","versions_md":"- 上代：Claude 3 Opus / Sonnet\n- 版本：20240620（首发）/ 20241022（升级，社区称 3.6）\n- 同代：Claude 3.5 Haiku（2024-11）\n- 继任：Claude 3.7 Sonnet（2025-02）"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"弱"}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true,"notes":"首个混合推理模型：同一模型可选择即时回答或扩展思考，思考预算可控。"},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"claude-3-7-sonnet","name":"Claude 3.7 Sonnet","aliases":["claude-3-7-sonnet-20250219"],"vendor":"Anthropic","family":"Claude 3.7","superseded_by":"claude-sonnet-4","released_at":"2025-02-24","modalities":["text","image","tools","computer-use"],"reasoning_mode":"optional","context":{"max_tokens":200000,"display":"200K","max_output":64000},"pricing":{"input_per_m":3,"output_per_m":15,"currency":"USD","source":"Anthropic 定价页","as_of":"2025-02-24","note":"思考 token 按输出计费"},"links":{"official":"https://www.anthropic.com/news/claude-3-7-sonnet","pricing":"https://www.anthropic.com/pricing"},"copy":{"one_liner":"首个混合推理模型，与 Claude Code 一同发布。","highlights":["SWE-bench Verified 62.3%（脚手架 70.3%），当时第一","扩展思考可控预算，GPQA 78.2%","Claude Code 命令行工具同日发布"],"pitfalls":["过度工程化倾向，会擅自改动未要求的代码","AIME 2024 61.3%，数学不及 o3-mini","128K 输出为 beta"],"logic_ability":"工程型推理里程碑，思考模式开启后考试型推理也进入第一梯队；关掉思考时与 3.5 Sonnet 相当。","best_for":["历史对照"],"not_for":["新项目"]}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"claude-3-opus","name":"Claude 3 Opus","aliases":["claude-3-opus-20240229"],"vendor":"Anthropic","family":"Claude 3","superseded_by":"claude-3-5-sonnet","released_at":"2024-03-04","modalities":["text","image","tools"],"reasoning_mode":"none","context":{"max_tokens":200000,"display":"200K","max_output":4096},"pricing":{"input_per_m":15,"output_per_m":75,"currency":"USD","source":"Anthropic 定价页","as_of":"2024-03-04"},"links":{"official":"https://www.anthropic.com/news/claude-3-family","paper":"https://www.anthropic.com/claude-3-model-card","pricing":"https://www.anthropic.com/pricing"},"copy":{"one_liner":"首个在 LMArena 超越 GPT-4 的模型，写作口碑极佳。","highlights":["MMLU 86.8%、GPQA 50.4%、GSM8K 95.0%","200K 上下文 needle 召回 99%，还\"察觉\"到测试","写作自然度与长文分析被广泛推崇"],"pitfalls":["$15 / $75 且速度慢","最大输出仅 4K","3 个月后被 1/5 价格的 3.5 Sonnet 超越"],"logic_ability":"2024 年初最强非推理模型之一，复杂指令与长链条分析稳定；数学 MATH 60.1% 一般。","best_for":["历史对照"],"not_for":["新项目"]}},{"id":"claude-fable-5","name":"Claude Fable 5","name_zh":"Claude Fable 5（Mythos 级公开版）","aliases":["claude-fable-5","Fable 5","anthropic.claude-fable-5"],"vendor":"Anthropic","family":"Claude 5","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2026-06-09","updated_at":"2026-08-28","modalities":["text","image","tools","computer-use"],"reasoning_mode":"default-on","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。官方仅说明：与 Claude Mythos 5 为同一底层模型，Fable 5 额外挂载安全分类器；adaptive thinking 始终开启，不可关闭；使用 Opus 4.7 起的新分词器（同文本 token 约多 30%）。"},"context":{"max_tokens":1000000,"display":"1M","max_output":128000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":10,"output_per_m":50,"currency":"USD","source":"Anthropic 官方定价页（platform.claude.com）","as_of":"2026-08-28","note":"不到 Mythos Preview 一半；缓存读 $1，Batch 五折；被拒绝请求不计费"},"links":{"official":"https://www.anthropic.com/claude/fable","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","paper":"https://www.anthropic.com/news/claude-fable-5-mythos-5"},"copy":{"one_liner":"Anthropic 公开的 Mythos 级模型，可连续多日运行 agent。","highlights":["LMArena 文本榜第一（1507，2026-08-27），AA 智能指数 62","1M 上下文 + 128K 输出，思考始终开启，可跨阶段委派子 agent 连续工作数日","官方定价 $10 / $50，不到 Mythos Preview 的一半，缓存读取 $1"],"pitfalls":["网络安全 / 生化 / 蒸馏类请求会被分类器拒绝并回退到 Opus 4.8，集成须处理 stop_reason=refusal","首 token 延迟高（AA 实测 TTFT 约 80 s），思考不可关闭且不返回原始思维链","价格是 Opus 5 的 2 倍，而 Opus 5 在多数榜单已持平甚至反超；30 天强制数据留存，无 ZDR"],"logic_ability":"当前公开模型中的最高档：GDP.pdf 视觉推理 29.8、SWE-bench Pro 80.3（官方公告表）领先所有对手；纯视觉通关《宝可梦 火红》，持久记忆下 Slay the Spire 表现比 Opus 4.8 高 3 倍。但 Opus 5 发布后，在 GDPval-AA v2、AA-Briefcase、ARC-AGI-3 上已被 Opus 5 反超，Fable 5 的优势集中在需要极长任务链、极高自主度的场景。失败模式：分类器偏保守，约 5% 会话会触发回退；输出偏长、思考 token 多。","best_for":["跨天运行的多阶段 agent 项目","大规模代码迁移 / 仓库级重构","金融、法律等高级知识工作","视觉密集型任务"],"not_for":["高频低价值调用","安全研究 / 生物类内容（会回退）","需要 ZDR 或私有化部署的合规场景"]},"capability_notes":{"coding":"SWE-bench Pro 80.3%（官方公告表，Vellum 转录）；Cognition FrontierCode 最高分；Stripe 用其一天完成 5000 万行 Ruby 迁移。","reasoning":"AA 智能指数 62（2026-08）；Frontier-Bench v0.1 33.7%，低于 Opus 5 的 43.3%。","agent":"官方定位为「长时运行 agent」，Claude Code / Managed Agents 下可工作数日；AA 测 Cost per task $3.14。","multimodal":"官方称视觉 SOTA：GDP.pdf（视觉、无工具）29.8 vs GPT-5.5 24.9。","chinese":"中文可用，缺少独立中文评测。"},"sheet":{"architecture_md":"**未披露**。Anthropic 未公布参数量、Dense/MoE、注意力结构。\n\n官方确认的事实：\n- 与 **Claude Mythos 5** 是同一底层模型，Fable 5 多了三类安全分类器（网络安全、生物/化学、蒸馏）；Mythos 5 仅通过 Project Glasswing 受邀提供\n- **Adaptive thinking 始终开启**，`thinking.type=disabled` 直接报错；深度用 `effort`（默认 `high`）调节\n- 原始思维链从不返回，`thinking.display` 只能选 `summarized` / `omitted`\n- 使用 Opus 4.7 起的新分词器，同一文本 token 数约多 30%\n- 1M 上下文（默认即最大）、128K 最大输出\n- 可靠知识截止 2026-01","memory_md":"无自建选项。\n\n**成本模型**（官方定价页，2026-08-28）：\n- 输入 $10 / 输出 $50 / MTok\n- 5 分钟缓存写 $12.50，1 小时缓存写 $20，缓存读 $1（10%）\n- Batch API 五折：$5 / $25\n- 1M 长上下文按标准价，不加价\n\n**被拒绝的请求不计费**；回退到其他模型时 fallback credit 会退还重复的缓存成本。\n\n实测：AA 智能指数单任务平均 $3.14，输出 64.9 tok/s，TTFT 约 80 s——比 Opus 5 贵约 35% 且更慢。","training_md":"- 训练细节未披露\n- 官方公告：在软件工程、知识工作、视觉、科学研究等几乎所有测试基准上 SOTA（截至 2026-06-09）\n- 生命科学：内部蛋白设计专家称药物设计流程提速约 10 倍；14 个蛋白靶点中 9 个得到强候选\n- 安全：1000+ 小时外部红队未发现通用越狱；UK AISI 在首轮测试窗口内「取得进展」\n- 对齐评估结果与 Opus 4.8 相近\n- 2026-06-12 因美国商务部出口管制命令下线，06-30 命令解除，07-01 恢复","ecosystem_md":"- Claude API、Amazon Bedrock、Claude Platform on AWS、Google Cloud（Vertex）、Microsoft Foundry\n- claude.ai Pro / Max / Team / Enterprise 可用（2026-06-23 起订阅用户需额外用量额度）\n- Claude Code、Claude Managed Agents\n- 支持 effort、task budgets（beta）、memory tool、code execution、programmatic tool calling、compaction、context editing\n- **不可微调**；30 天数据留存，属 Covered Model，不支持 ZDR\n- 中文文档：弱","versions_md":"- 2026-06-09 Claude Fable 5 / Claude Mythos 5 发布，定价 $10 / $50\n- 2026-06-12 ~ 07-01 因出口管制短暂下线\n- 同代：Claude Opus 5（2026-07-24，$5 / $25，多数榜单已追平）、Claude Sonnet 5（2026-06-30）\n- 上代对应：Claude Mythos Preview（Glasswing 限定）、Claude Opus 4.8\n- 退役承诺：不早于 2027-06-09"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"弱"},"i18n":{"en":{"name_zh":"Claude Fable 5 (public Mythos-class)","one_liner":"Anthropic's public Mythos-class model; runs agents for days on end.","highlights":["#1 on LMArena text (1507, 2026-08-27), AA Intelligence Index 62","1M context + 128K output, thinking always on, can delegate sub-agents across stages and work for days","Official pricing $10 / $50, under half of Mythos Preview; cache reads $1"],"pitfalls":["Cybersecurity / bio-chem / distillation requests are refused by a classifier and fall back to Opus 4.8; integrations must handle stop_reason=refusal","High first-token latency (AA-measured TTFT about 80 s); thinking cannot be disabled and raw chain-of-thought is not returned","2x the price of Opus 5, which already matches or beats it on most leaderboards; mandatory 30-day data retention, no ZDR"],"logic_ability":"Top tier among current public models: GDP.pdf visual reasoning 29.8 and SWE-bench Pro 80.3 (official announcement table) lead all rivals; beat Pokemon FireRed on vision alone, and with persistent memory scores 3x Opus 4.8 on Slay the Spire. Since the Opus 5 launch, however, it has been overtaken by Opus 5 on GDPval-AA v2, AA-Briefcase and ARC-AGI-3; Fable 5's edge is concentrated in very long task chains with very high autonomy. Failure modes: the classifier is conservative, about 5% of sessions trigger the fallback; outputs run long with many thinking tokens.","best_for":["Multi-stage agent projects that run across days","Large-scale code migrations / repo-level refactors","Advanced knowledge work such as finance and law","Vision-heavy tasks"],"not_for":["High-frequency low-value calls","Security research / biology content (falls back)","Compliance scenarios that need ZDR or on-prem deployment"],"capability_notes":{"coding":"SWE-bench Pro 80.3% (official announcement table, transcribed by Vellum); top score on Cognition FrontierCode; Stripe used it to migrate 50M lines of Ruby in one day.","reasoning":"AA Intelligence Index 62 (2026-08); Frontier-Bench v0.1 33.7%, below Opus 5's 43.3%.","agent":"Officially positioned as a \"long-running agent\"; works for days under Claude Code / Managed Agents; AA-measured cost per task $3.14.","multimodal":"Officially claimed vision SOTA: GDP.pdf (vision, no tools) 29.8 vs GPT-5.5 24.9.","chinese":"Chinese usable; no independent Chinese evaluation."}},"ja":{"name_zh":"Claude Fable 5（Mythos 級公開版）","one_liner":"Anthropic 公開の Mythos 級。agent を数日連続実行可能。","highlights":["LMArena テキスト部門 1 位（1507、2026-08-27）、AA 知能指数 62","1M コンテキスト + 128K 出力、思考は常時オン、段階をまたいでサブ agent に委譲し数日連続で稼働可能","公式価格 $10 / $50、Mythos Preview の半額未満、キャッシュ読み取り $1"],"pitfalls":["サイバーセキュリティ / 生化学 / 蒸留系リクエストは分類器に拒否され Opus 4.8 にフォールバック。統合側で stop_reason=refusal の処理が必須","初回トークン遅延が大きい（AA 実測 TTFT 約 80 s）。思考はオフにできず生の思考過程も返らない","価格は Opus 5 の 2 倍だが、Opus 5 は多くのリーダーボードで既に同等かそれ以上。30 日間の強制データ保持、ZDR なし"],"logic_ability":"現行公開モデルの最上位：GDP.pdf 視覚推論 29.8、SWE-bench Pro 80.3（公式発表表）で全競合をリード。純粋な視覚のみで『ポケットモンスター ファイアレッド』をクリア、永続メモリ下の Slay the Spire では Opus 4.8 の 3 倍の成績。ただし Opus 5 発表後は GDPval-AA v2、AA-Briefcase、ARC-AGI-3 で Opus 5 に逆転され、Fable 5 の強みは極めて長いタスクチェーンと高い自律性を要する場面に集中。失敗モード：分類器が保守的で約 5% のセッションでフォールバックが発生。出力が長く思考トークンが多い。","best_for":["日をまたいで動く多段階 agent プロジェクト","大規模コード移行 / リポジトリ級リファクタリング","金融・法務などの高度な知識労働","視覚中心のタスク"],"not_for":["高頻度・低価値の呼び出し","セキュリティ研究 / 生物系コンテンツ（フォールバックする）","ZDR やオンプレ展開が必要なコンプライアンス用途"],"capability_notes":{"coding":"SWE-bench Pro 80.3%（公式発表表、Vellum 転記）。Cognition FrontierCode 最高スコア。Stripe は 5000 万行の Ruby 移行を 1 日で完了。","reasoning":"AA 知能指数 62（2026-08）。Frontier-Bench v0.1 33.7%、Opus 5 の 43.3% を下回る。","agent":"公式の位置づけは「長時間稼働 agent」。Claude Code / Managed Agents 下で数日稼働可能。AA 測定のタスク単価 $3.14。","multimodal":"公式に視覚 SOTA を主張：GDP.pdf（視覚、ツールなし）29.8 vs GPT-5.5 24.9。","chinese":"中国語は利用可能。独立した中国語評価はなし。"}}},"complete":true,"runtime":{"tok_s":65,"latency_s":79.92,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"claude-haiku-4-5","name":"Claude Haiku 4.5","name_zh":"Claude 4.5 轻量版","aliases":["claude-haiku-4-5-20251001"],"vendor":"Anthropic","family":"Claude 4.5","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2025-10-15","updated_at":"2025-12-20","modalities":["text","image","tools","computer-use"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true},"context":{"max_tokens":200000,"display":"200K","max_output":64000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":1,"output_per_m":5,"currency":"USD","source":"Anthropic 定价页","as_of":"2025-12-20"},"links":{"official":"https://www.anthropic.com/claude/haiku","pricing":"https://www.anthropic.com/pricing"},"copy":{"one_liner":"接近 Sonnet 4 的编程能力，1/3 价格、2 倍速度。","highlights":["SWE-bench Verified 73.3%，小模型里最强代码","$1 / $5，支持 extended thinking","延迟低，适合子 agent"],"pitfalls":["知识广度与复杂推理弱于 Sonnet 4.5","长任务规划能力有限","无自建"],"logic_ability":"工程型推理在同价位里领先；考试型推理中等。适合做多 agent 系统里的执行者。","best_for":["高频代码补全 / 小任务","多 agent 子任务"],"not_for":["复杂长任务"]},"capability_notes":{},"i18n":{"en":{"name_zh":"Claude 4.5 lightweight","one_liner":"Near Sonnet 4 coding ability at 1/3 the price and 2x the speed.","highlights":["SWE-bench Verified 73.3%, strongest code among small models","$1 / $5, supports extended thinking","Low latency, suited to sub-agents"],"pitfalls":["Knowledge breadth and complex reasoning weaker than Sonnet 4.5","Limited long-task planning ability","No self-hosting"],"logic_ability":"Engineering-style reasoning leads its price bracket; exam-style reasoning is middling. Suited to being the executor in multi-agent systems.","best_for":["High-frequency code completion / small tasks","Multi-agent subtasks"],"not_for":["Complex long tasks"],"capability_notes":{}},"ja":{"name_zh":"Claude 4.5 軽量版","one_liner":"Sonnet 4 に迫るコーディング力を 1/3 の価格、2 倍の速度で。","highlights":["SWE-bench Verified 73.3%、小型モデル最強のコード力","$1 / $5、extended thinking 対応","低遅延でサブ agent 向き"],"pitfalls":["知識の広さと複雑な推論は Sonnet 4.5 に劣る","長時間タスクの計画能力は限定的","セルフホスト不可"],"logic_ability":"エンジニアリング型推論は同価格帯でトップ。試験型推論は中程度。マルチ agent システムの実行役に向く。","best_for":["高頻度のコード補完 / 小タスク","マルチ agent のサブタスク"],"not_for":["複雑な長時間タスク"],"capability_notes":{}}},"complete":false,"runtime":{"tok_s":119,"latency_s":16.59,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"claude-opus-4-1","name":"Claude Opus 4.1","aliases":["claude-opus-4-1-20250805"],"vendor":"Anthropic","family":"Claude 4","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","superseded_by":"claude-opus-4-5","released_at":"2025-08-05","updated_at":"2025-12-20","modalities":["text","image","tools","computer-use"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true},"context":{"max_tokens":200000,"display":"200K","max_output":32000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":15,"output_per_m":75,"currency":"USD","source":"Anthropic 定价页","as_of":"2025-12-20"},"links":{"official":"https://www.anthropic.com/news/claude-opus-4-1","pricing":"https://www.anthropic.com/pricing"},"copy":{"one_liner":"上代 Opus，已被 Opus 4.5 以 1/3 价格全面替代。","highlights":["SWE-bench Verified 74.5%","agent 任务稳定","200K 上下文"],"pitfalls":["$15 / $75 极贵","已被 Opus 4.5 替代","无自建"],"logic_ability":"工程型推理强，考试型中上。已被 Opus 4.5 超越。","best_for":["历史对照"],"not_for":["新项目"]},"capability_notes":{},"complete":false},{"id":"claude-opus-4-5","name":"Claude Opus 4.5","name_zh":"Claude 4.5 顶配","aliases":["claude-opus-4-5-20251101","Opus 4.5"],"vendor":"Anthropic","family":"Claude 4.5","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","released_at":"2025-11-24","updated_at":"2025-12-20","modalities":["text","image","tools","computer-use"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。公开能力：extended thinking、effort 参数（low/medium/high）控制思考深度。"},"context":{"max_tokens":200000,"display":"200K","max_output":64000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":5,"output_per_m":25,"currency":"USD","source":"Anthropic 定价页","as_of":"2025-12-20","note":"较 Opus 4.1 降价 2/3"},"links":{"official":"https://www.anthropic.com/claude/opus","pricing":"https://www.anthropic.com/pricing","paper":"https://www.anthropic.com/news/claude-opus-4-5"},"copy":{"one_liner":"2025 年底最强编程模型之一，价格降到 Sonnet 级别的 1.7 倍。","highlights":["SWE-bench Verified 80.9%，首个突破 80% 的公开模型","effort 参数让同一模型在成本与深度之间连续调节","定价 $5 / $25，比 Opus 4.1 便宜 67%"],"pitfalls":["仍是全场最贵档之一，日常吞吐任务应下沉到 Sonnet / Haiku","200K 上下文在多文件大仓库场景仍会不够，需要检索配合","closed-source：参数、架构完全不透明，合规审计只能看服务协议"],"logic_ability":"工程型与考试型兼备：ARC-AGI-2 37.6%、GPQA 87% 与 SWE-bench 80.9% 同时领先，是少数在「抽象推理 / 科学问答 / 真实软件工程」三类上都靠前的模型。effort=high 时会长时间自我验证，token 消耗大；effort=low 接近 Sonnet 4.5 的行为。失败模式：在模糊需求下倾向于给出过度完整的方案而非追问。","best_for":["复杂重构与跨文件修改","需要高可靠性的 agent 编排","研究型分析、长推理链"],"not_for":["高频低价值调用","需要私有化部署的合规场景"]},"capability_notes":{"coding":"SWE-bench Verified 80.9%、Terminal-bench 59.3%（官方）。","reasoning":"GPQA Diamond 87.0%、ARC-AGI-2 37.6%（官方）。","agent":"tau2-bench、OSWorld 官方领先，长任务稳定。","multimodal":"MMMU 80.7%，图像理解良好。","chinese":"中文可用，缺少独立中文评测。"},"sheet":{"architecture_md":"**未披露**。Anthropic 未公布参数量、Dense/MoE、注意力类型。\n\n已知接口特性：\n- `effort` 参数（low / medium / high）控制推理投入\n- extended thinking，思考可跨工具调用保持（interleaved thinking）\n- 200K 上下文、64K 最大输出","memory_md":"无自建选项。成本角度看：输入 $5 / 输出 $25，是 Sonnet 4.5 的 1.67 倍，Opus 4.1 的 1/3。\n\n高 effort 模式下单次请求 thinking token 可达数万，请用 `max_tokens` 与 effort 联合限制预算。","training_md":"- 训练细节未披露\n- 默认不思考，可按请求开启并设定预算\n- 支持并行工具调用、计算机使用、结构化输出\n- 最大输出 64K","ecosystem_md":"- 官方 API、Bedrock、Vertex AI\n- Claude Code 默认高档模型\n- 不支持微调\n- 中文文档：弱","versions_md":"- 上代：Claude Opus 4.1（2025-08）\n- 同代：Sonnet 4.5、Haiku 4.5\n- 继任：Claude Opus 5 / Claude Fable 5（2026，规格未披露）"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"弱"},"complete":true,"superseded_by":"claude-opus-5"},{"id":"claude-opus-4-8","name":"Claude Opus 4.8","name_zh":"Claude 4.8 顶配（已被 Opus 5 取代）","aliases":["claude-opus-4-8","Opus 4.8","anthropic.claude-opus-4-8"],"vendor":"Anthropic","family":"Claude 4","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","superseded_by":"claude-opus-5","released_at":"2026-05-28","updated_at":"2026-08-28","modalities":["text","image","tools","computer-use"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。adaptive thinking 需显式开启；effort 默认 high；fast mode 2.5× 速度；官方已标 Legacy。"},"context":{"max_tokens":1000000,"display":"1M","max_output":128000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":5,"output_per_m":25,"currency":"USD","source":"Anthropic 官方定价页（platform.claude.com）","as_of":"2026-08-28","note":"Fast mode $10 / $50"},"links":{"official":"https://platform.claude.com/docs/en/models/opus-4-8/overview","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","paper":"https://www.anthropic.com/news/claude-opus-4-8"},"copy":{"one_liner":"Opus 4.x 末代，SWE-bench 88.6%，已被 Opus 5 取代","highlights":["SWE-bench Verified 88.6%、GPQA Diamond 93.6%、OSWorld-Verified 83.4%（官方公告表）","Fable 5 拒绝请求时的官方回退模型，仍在所有平台提供","与 Opus 5 同价 $5 / $25，思考默认关闭，迁移成本低"],"pitfalls":["官方状态 Legacy：Frontier-Bench 仅 18.7%，不到 Opus 5 一半，新项目无理由选它","ARC-AGI-3 仅 1.5%，抽象推理明显落后","退役承诺仅到 2027-05-28，应尽快迁移"],"logic_ability":"上一代旗舰水平：GPQA Diamond 93.6%、HLE（带工具）57.9%、GDPval-AA 1890（官方公告表，Vellum 转录）。相比 Opus 5，深度推理与长任务差距显著（Frontier-Bench 18.7 vs 43.3）。","best_for":["Fable 5 回退目标","已有 Opus 4.x 提示词的存量系统"],"not_for":["新项目","需要私有化部署的合规场景"]},"capability_notes":{"coding":"SWE-bench Verified 88.6%、SWE-bench Pro 69.2%、Terminal-Bench 2.1 74.6%（官方公告表，Vellum 转录）。","reasoning":"GPQA Diamond 93.6%、HLE（带工具）57.9%。","agent":"OSWorld-Verified 83.4%、Finance Agent v2 53.9%。","chinese":"中文可用，缺少独立中文评测。"},"complete":false},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"claude-opus-4","name":"Claude Opus 4","aliases":["claude-opus-4-20250514"],"vendor":"Anthropic","family":"Claude 4","superseded_by":"claude-opus-4-1","released_at":"2025-05-22","modalities":["text","image","tools","computer-use"],"reasoning_mode":"optional","context":{"max_tokens":200000,"display":"200K","max_output":32000},"pricing":{"input_per_m":15,"output_per_m":75,"currency":"USD","source":"Anthropic 定价页","as_of":"2025-05-22"},"links":{"official":"https://www.anthropic.com/news/claude-4","pricing":"https://www.anthropic.com/pricing"},"copy":{"one_liner":"长时间自主编程的旗舰，Terminal-bench 43.2%。","highlights":["SWE-bench Verified 72.5%、Terminal-bench 43.2%","可连续自主工作数小时的代理稳定性","思考中可交替调用工具、记忆文件"],"pitfalls":["$15 / $75，SWE-bench 与 Sonnet 4 几乎一样","系统卡披露\"告密\"等高能动性行为引发讨论","2.5 个月后被 Opus 4.1 替代"],"logic_ability":"工程型推理第一梯队，长程任务保持目标；考试型 GPQA 79.6%、AIME 2025 75.5%（扩展思考）。","best_for":["历史对照"],"not_for":["新项目"]}},{"id":"claude-opus-5","name":"Claude Opus 5","name_zh":"Claude 5 顶配","aliases":["claude-opus-5","Opus 5","anthropic.claude-opus-5"],"vendor":"Anthropic","family":"Claude 5","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2026-07-24","updated_at":"2026-08-28","modalities":["text","image","tools","computer-use"],"reasoning_mode":"default-on","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。官方接口特性：adaptive thinking 默认开启；effort 五档 low/medium/high/xhigh/max；xhigh/max 下不可关闭思考；fast mode（2.5× 速度，$10/$50）。"},"context":{"max_tokens":1000000,"display":"1M","max_output":128000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":5,"output_per_m":25,"currency":"USD","source":"Anthropic 官方定价页（platform.claude.com）","as_of":"2026-08-28","note":"与 Opus 4.8 同价；Fast mode $10 / $50；Batch 五折"},"links":{"official":"https://www.anthropic.com/claude/opus","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","paper":"https://www.anthropic.com/news/claude-opus-5"},"copy":{"one_liner":"AA 智能指数第一，Fable 5 级智能只要一半价钱。","highlights":["AA 智能指数 63 排名第一，GDPval-AA v2、AA-Briefcase、Frontier-Bench 均为 SOTA","effort 五档（low→max）可连续换算成本与深度，max 档 Terminal-Bench 2.1 达 89%","定价 $5 / $25 与 Opus 4.8 持平，仅 Fable 5 一半；1M 上下文不加价"],"pitfalls":["思考默认开启，max_tokens 要重新预算；xhigh/max 下无法关闭思考（400 错误）","默认回复与交付物明显变长，会主动自我验证，沿用旧提示词易「过度验证」","生物、进攻性网络安全能力仍落后 Mythos 5，且安全限制比 Opus 4.8 更严"],"logic_ability":"当前综合推理最强档之一：ARC-AGI-3 官方验证 30.2%（次优 GPT-5.6 Sol 仅 7.8%），HLE 53%（AA，无工具），Frontier-Bench v0.1 43.3% 是 Opus 4.8 的 2 倍以上。官方强调「test-time compute scaling」：把 effort 从 low 提到 max，GDPval-AA 分差达 407 Elo，输出 token 约 8 倍。失败模式：高 effort 下输出冗长；关闭思考时偶尔把工具调用写进正文或漏出内部 XML 标签。","best_for":["复杂多文件重构与端到端功能开发","多 agent 编排（writer-verifier 模式）","深度研究与长推理链","表格 / 幻灯片等办公文档生成"],"not_for":["高频低价值调用（下沉到 Sonnet 5 / Haiku 4.5）","进攻性安全研究","需要私有化部署的合规场景"]},"capability_notes":{"coding":"Frontier-Bench v0.1 43.3%（官方公告，Vellum 转录），CursorBench 3.2 与 Fable 5 差距 0.5% 以内；Terminal-Bench 2.1 max 档 89%（AA）。","reasoning":"AA 智能指数 63（第一）；HLE 53%（AA）；ARC-AGI-3 30.2%（ARC Prize 验证）。","agent":"Zapier AutomationBench 通过率为次优模型 1.5 倍；OSWorld 2.0 以 1/3 成本超过 Fable 5；GDPval-AA v2 1861 Elo。","multimodal":"官方称视觉能力提升，擅长图表 / 文档 / UI 复刻，配合裁剪工具迭代效果最好。","chinese":"中文可用，缺少独立中文评测。"},"sheet":{"architecture_md":"**未披露**。Anthropic 未公布参数量、Dense/MoE、注意力类型。\n\n已知接口特性：\n- **Adaptive thinking 默认开启**（Opus 4.8 默认关闭），模型自行决定每轮思考量\n- `effort` 五档：`low` / `medium` / `high`（默认） / `xhigh` / `max`；官方称 Opus 5 把额外 effort 换成结果的效率高于任何前代\n- `thinking.type=disabled` 仅在 effort ≤ high 时接受，xhigh/max 下返回 400（相对 Opus 4.8 的破坏性变更）\n- 1M 上下文既是默认也是最大值，128K 最大输出；Batch API 加 beta 头可到 300K 输出\n- 新增：会话中途增删工具并保留缓存（beta）、`fallbacks: \"default\"` 模式、缓存最小长度降至 512 token\n- Fast mode（研究预览）：约 2.5× 输出速度，$10 / $50\n- 可靠知识截止 2026-05（全家族最新）","memory_md":"无自建选项。\n\n**成本模型**（官方定价页，2026-08-28）：\n- 输入 $5 / 输出 $25 / MTok，与 Opus 4.8、4.7、4.6、4.5 完全一致\n- 缓存写 $6.25（5 分钟）/ $10（1 小时），缓存读 $0.50\n- Batch $2.50 / $12.50；Fast mode $10 / $50\n- 1M 长上下文标准价\n\n**effort 决定真实成本**：AA 实测 max 档单任务 $2.03~2.34，比 Fable 5（$2.75~3.14）低约 26%，但比 Opus 4.8 max（$1.80）高；low→max 输出 token 跨度约 8 倍。高 effort 时请给足 `max_tokens`（官方示例 64K 并流式）。\n\n速度：AA 实测 55.4 tok/s，TTFT 约 43 s（max 档）。","training_md":"- 训练细节未披露\n- 官方定位：相对 Opus 4.8 是「阶跃式」而非增量升级，最大提升在深度推理、agent 长任务、test-time compute scaling\n- 官方公告数据：Frontier-Bench v0.1 43.3%（Opus 4.8 18.7%）；ARC-AGI-3 为次优模型 3 倍；有机化学 +10.2 pp、蛋白预测 +7.7 pp（vs 4.8）\n- 代码审查：单次通过发现真实 bug 率高、误报少，低 effort 仍准确\n- 对齐：近期模型中最低的 misalignment 分数（2.3）；生物与进攻性网络安全仍弱于 Mythos 5\n- 行为变化：回复更长、agent 会话中更常汇报进度、更愿意委派子 agent、无需指示即自我验证","ecosystem_md":"- Claude API、Amazon Bedrock（含 InvokeModel）、Claude Platform on AWS、Google Cloud、Microsoft Foundry\n- claude.ai Max 默认模型，Pro 可用的最强模型；Claude Code、Claude Cowork\n- Claude Managed Agents：token 价 + $0.08 / 会话小时\n- 工具：computer_toolset_20260801、browser_toolset_20260801、bash、text editor、web search、code execution、memory、tool search\n- **不可微调**\n- 中文文档：弱","versions_md":"- 上代：Claude Opus 4.8（2026-05-28，已标 Legacy，仍可用；官方建议迁移）\n- 更早：Opus 4.7、4.6、4.5（2025-11）\n- 同代：Claude Fable 5（2026-06-09，$10/$50）、Claude Sonnet 5（2026-06-30，$2/$10）\n- 更高档：Claude Mythos 5（同 Fable 5 底层，Glasswing 限定）\n- 退役承诺：不早于 2027-07-24"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"弱"},"i18n":{"en":{"name_zh":"Claude 5 flagship","one_liner":"#1 on AA Intelligence Index; Fable 5-level intelligence at half the price.","highlights":["AA Intelligence Index 63, ranked #1; SOTA on GDPval-AA v2, AA-Briefcase and Frontier-Bench","Five effort levels (low to max) trade cost against depth continuously; max reaches 89% on Terminal-Bench 2.1","$5 / $25, same as Opus 4.8 and half of Fable 5; 1M context at no extra charge"],"pitfalls":["Thinking is on by default, so max_tokens must be re-budgeted; thinking cannot be turned off at xhigh/max (400 error)","Default replies and deliverables are noticeably longer and it self-verifies proactively; old prompts tend to cause \"over-verification\"","Biology and offensive cybersecurity still trail Mythos 5, and safety restrictions are stricter than Opus 4.8"],"logic_ability":"One of the strongest overall reasoners today: ARC-AGI-3 officially verified 30.2% (runner-up GPT-5.6 Sol only 7.8%), HLE 53% (AA, no tools), Frontier-Bench v0.1 43.3%, more than 2x Opus 4.8. Officially emphasizes \"test-time compute scaling\": raising effort from low to max spreads GDPval-AA by 407 Elo with about 8x the output tokens. Failure modes: verbose output at high effort; with thinking off it occasionally writes tool calls into the body or leaks internal XML tags.","best_for":["Complex multi-file refactors and end-to-end feature development","Multi-agent orchestration (writer-verifier pattern)","Deep research and long reasoning chains","Office documents such as spreadsheets / slides"],"not_for":["High-frequency low-value calls (downgrade to Sonnet 5 / Haiku 4.5)","Offensive security research","Compliance scenarios needing on-prem deployment"],"capability_notes":{"coding":"Frontier-Bench v0.1 43.3% (official announcement, transcribed by Vellum); CursorBench 3.2 within 0.5% of Fable 5; Terminal-Bench 2.1 at max 89% (AA).","reasoning":"AA Intelligence Index 63 (#1); HLE 53% (AA); ARC-AGI-3 30.2% (ARC Prize verified).","agent":"Zapier AutomationBench pass rate 1.5x the runner-up; beats Fable 5 on OSWorld 2.0 at 1/3 the cost; GDPval-AA v2 1861 Elo.","multimodal":"Officially improved vision; strong at charts / documents / UI replication, works best iterating with a crop tool.","chinese":"Chinese usable; no independent Chinese evaluation."}},"ja":{"name_zh":"Claude 5 フラッグシップ","one_liner":"AA 知能指数 1 位。Fable 5 級の知能を半額で。","highlights":["AA 知能指数 63 で 1 位。GDPval-AA v2、AA-Briefcase、Frontier-Bench いずれも SOTA","effort 5 段階（low→max）でコストと深さを連続的に調整可能。max では Terminal-Bench 2.1 89%","価格 $5 / $25 は Opus 4.8 と同じで Fable 5 の半額。1M コンテキストも追加料金なし"],"pitfalls":["思考がデフォルトでオンのため max_tokens の再設定が必要。xhigh/max では思考をオフにできない（400 エラー）","デフォルトの応答と成果物が明らかに長くなり、自発的に自己検証する。旧プロンプト流用では「過剰検証」になりやすい","生物学・攻撃的サイバーセキュリティ能力は依然 Mythos 5 に劣り、安全制限は Opus 4.8 より厳しい"],"logic_ability":"現時点で総合推論最強クラスの一つ：ARC-AGI-3 公式検証 30.2%（次点の GPT-5.6 Sol は 7.8%）、HLE 53%（AA、ツールなし）、Frontier-Bench v0.1 43.3% は Opus 4.8 の 2 倍超。公式は「test-time compute scaling」を強調：effort を low から max に上げると GDPval-AA で 407 Elo の差、出力トークンは約 8 倍。失敗モード：高 effort では出力が冗長。思考オフ時にツール呼び出しを本文に書いたり内部 XML タグを漏らすことがある。","best_for":["複雑な複数ファイルのリファクタリングとエンドツーエンドの機能開発","マルチ agent オーケストレーション（writer-verifier パターン）","深い調査と長い推論チェーン","表計算 / スライドなどのオフィス文書生成"],"not_for":["高頻度・低価値の呼び出し（Sonnet 5 / Haiku 4.5 に下げる）","攻撃的セキュリティ研究","オンプレ展開が必要なコンプライアンス用途"],"capability_notes":{"coding":"Frontier-Bench v0.1 43.3%（公式発表、Vellum 転記）。CursorBench 3.2 で Fable 5 との差 0.5% 以内。Terminal-Bench 2.1 max で 89%（AA）。","reasoning":"AA 知能指数 63（1 位）。HLE 53%（AA）。ARC-AGI-3 30.2%（ARC Prize 検証）。","agent":"Zapier AutomationBench 通過率は次点モデルの 1.5 倍。OSWorld 2.0 で 1/3 のコストで Fable 5 を上回る。GDPval-AA v2 1861 Elo。","multimodal":"公式に視覚能力向上を主張。図表 / 文書 / UI 再現に強く、クロップツールと組み合わせた反復で最も効果的。","chinese":"中国語は利用可能。独立した中国語評価はなし。"}}},"complete":true,"runtime":{"tok_s":55,"latency_s":42.91,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"claude-sonnet-4-5","name":"Claude Sonnet 4.5","name_zh":"Claude 4.5 中档旗舰","aliases":["claude-sonnet-4-5-20250929","Sonnet 4.5"],"vendor":"Anthropic","family":"Claude 4.5","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","released_at":"2025-09-29","updated_at":"2025-12-20","modalities":["text","image","tools","computer-use"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true,"notes":"Anthropic 未披露参数量、层数与注意力结构。仅公开：支持 extended thinking（可调思考预算）、上下文 200K（beta 1M）。"},"context":{"max_tokens":200000,"display":"200K","max_output":64000},"memory":{"weight_gb":{},"estimated":false,"kv_note":"闭源模型，无自建选项。"},"pricing":{"input_per_m":3,"output_per_m":15,"currency":"USD","source":"Anthropic 定价页","as_of":"2025-12-20","note":"Prompt cache 读 $0.30；>200K 上下文加价"},"runtime":{"tok_s":60,"source":"Artificial Analysis 中位吞吐（估）"},"links":{"official":"https://www.anthropic.com/claude/sonnet","pricing":"https://www.anthropic.com/pricing","paper":"https://www.anthropic.com/news/claude-sonnet-4-5"},"copy":{"one_liner":"编程与 Agent 场景的性价比主力，长时程任务稳定。","highlights":["SWE-bench Verified 77.2%，同价位段最强代码模型之一","计算机使用（OSWorld）与长时程 agent 稳定性明显优于上代","$3 / $15 的价格远低于 Opus，工程团队日常可用"],"pitfalls":["纯知识问答与多语种创作不如 Opus 系与 Gemini 3 Pro","extended thinking 开启后延迟与成本会显著上升，需按任务开关","200K 之上的 1M 上下文为 beta 且加价，别把它当默认能力"],"logic_ability":"工程型推理强于考试型推理：在 SWE-bench、Terminal-bench 这类需要读代码、规划多步、验证结果的任务上表现稳定；AIME 类竞赛题略逊于 GPT-5 / Gemini 3 Pro。默认不思考，开启 extended thinking 后规划与自我修正能力提升。常见失败模式：长任务里过度自信地宣布完成、未运行验证。","best_for":["代码助手 / Coding Agent","计算机使用与浏览器自动化","需要工具调用的多步任务"],"not_for":["预算极低的高频简单任务（用 Haiku 4.5）","超长文档（>200K）常规处理"]},"capability_notes":{"coding":"SWE-bench Verified 77.2%（官方），在真实仓库修 bug 场景非常稳。","reasoning":"GPQA Diamond 83.4%，推理扎实但非顶尖。","math":"AIME 2025 无工具 87%，有 Python 工具 100%（官方）。","agent":"tau2-bench、OSWorld 领先，30 小时级长任务官方演示。","multimodal":"图像理解可用，MMMU 77.8%，弱于 Gemini 系。","chinese":"中文流畅，但缺少独立中文榜单数据。"},"sheet":{"architecture_md":"**类型**：未披露（Anthropic 不公开 Dense/MoE、参数量或层结构）。\n\n**已公开的结构性信息**：\n- 支持 *extended thinking*：可设 `budget_tokens`，思考内容可返回摘要\n- 上下文 200K，1M 为 beta（需 header 开启，价格上浮）\n- 原生支持工具调用、并行工具调用、计算机使用（screenshot → action）\n- 支持 prompt caching、批处理（Batch API 半价）\n\n架构简图中的一切内部结构均以虚线标「未披露」。","memory_md":"闭源模型，**没有自建选项**，因此没有显存表。\n\n成本视角：\n- 输入 $3 / 输出 $15（每百万 token）\n- 缓存命中输入 $0.30，长 system prompt 场景务必开缓存\n- Batch API 打 5 折，离线批量任务优先用\n\n**警示**：思考模式产生的 thinking token 按输出计费，开启后单次请求成本可能翻数倍。","training_md":"- 训练数据、规模、阶段：未披露细节（官方仅说明混合了公开数据、第三方数据和内部生成数据）\n- 默认 **不思考**，`thinking` 参数可按请求开启\n- 支持函数调用、并行工具、结构化输出（JSON schema）\n- 计算机使用：官方 Agent SDK 与 Claude Code 首选模型\n- 最大输出 64K token","ecosystem_md":"- 官方 API、AWS Bedrock、Google Vertex AI 三渠道\n- Claude Code / Agent SDK 原生支持\n- 无权重，无微调（不支持 fine-tune）\n- 中文文档：官方文档以英文为主，社区中文资料充足","versions_md":"- 上代：Claude Sonnet 4（2025-05）\n- 同代：Claude Opus 4.5（更强、更贵）、Claude Haiku 4.5（更快、更便宜）\n- 继任：Claude Sonnet 5（2026，规格未披露）"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"弱"},"complete":true,"superseded_by":"claude-sonnet-5"},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true,"notes":"混合推理模型，扩展思考可选、预算可控；2025-08 起 API 提供 1M 上下文 beta。"},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{"coding":"SWE-bench Verified 72.7%（并行计算 80.2%）、Terminal-bench 35.5%（官方）。","reasoning":"GPQA Diamond 75.4%（扩展思考；并行 83.8%）（官方）。","math":"AIME 2025 70.5%（并行 85.0%）（官方）。","agent":"TAU-bench retail 80.5%、airline 60.0%（官方）。","multimodal":"MMMU 74.4%（官方）。","chinese":"MMMLU 86.5%（官方），中文良好。"},"complete":true,"id":"claude-sonnet-4","name":"Claude Sonnet 4","aliases":["claude-sonnet-4-20250514"],"vendor":"Anthropic","family":"Claude 4","superseded_by":"claude-sonnet-4-5","released_at":"2025-05-22","modalities":["text","image","tools","computer-use"],"reasoning_mode":"optional","context":{"max_tokens":200000,"display":"200K（后 1M beta）","max_output":64000},"pricing":{"input_per_m":3,"output_per_m":15,"currency":"USD","source":"Anthropic 定价页","as_of":"2025-05-22","note":">200K 输入 $6 / $22.5（1M beta）"},"links":{"official":"https://www.anthropic.com/news/claude-4","paper":"https://www.anthropic.com/claude-4-system-card","pricing":"https://www.anthropic.com/pricing"},"copy":{"one_liner":"与 Opus 4 同分的性价比编程主力，GitHub Copilot 默认。","highlights":["SWE-bench Verified 72.7%，与 Opus 4 持平、价格 1/5","被 GitHub Copilot 选为新 coding agent 底座","2025-08 起 1M 上下文 beta"],"pitfalls":["Terminal-bench 35.5%，长程代理稳定性不及 Opus","考试型推理 AIME 2025 70.5%，落后 o3 / Gemini 2.5 Pro","扩展思考 token 按输出计费，实际成本可能翻倍"],"logic_ability":"工程型推理极佳：代码导航、精确编辑、减少 3.7 版的\"过度工程化\"。考试型推理中上，扩展思考下 GPQA 75.4%。常见问题：思考预算过小时推理浅、过大时啰嗦。","best_for":["代理编程（当年主力）","工具调用与结构化任务","历史对照"],"not_for":["竞赛数学","新项目"]},"sheet":{"architecture_md":"**未披露**。Anthropic 未公开参数量。\n\n已知接口层信息：\n- 200K 上下文（2025-08 起 1M beta），64K 最大输出\n- 扩展思考：`budget_tokens` 控制，思考摘要返回\n- 思考中可交替调用工具（interleaved thinking，beta）\n- 支持 computer use、文件 API、代码执行工具","memory_md":"无自建选项。\n\n成本参考：输入 $3 / 输出 $15；缓存写 $3.75、读 $0.30；Batch 半价。1M beta 超过 200K 部分 $6 / $22.5。\n\n**警示**：思考 token 计入输出，长思考预算下费用显著上升。","training_md":"- 训练细节未披露，知识截止 2025-03\n- 混合推理：默认不思考，按需开启\n- 支持工具调用、并行工具、结构化输出（通过工具）\n- 系统卡：ASL-2 部署（Opus 4 为 ASL-3）","ecosystem_md":"- 官方 API、AWS Bedrock、Google Vertex AI、GitHub Copilot\n- 不支持微调\n- 中文文档：弱（官方英文）","versions_md":"- 上代：Claude 3.7 Sonnet（2025-02）\n- 同代：Claude Opus 4（同日发布）\n- 2025-08：1M 上下文 beta\n- 继任：Claude Sonnet 4.5（2025-09）"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"弱"}},{"id":"claude-sonnet-5","name":"Claude Sonnet 5","name_zh":"Claude 5 中档","aliases":["claude-sonnet-5","Sonnet 5","anthropic.claude-sonnet-5"],"vendor":"Anthropic","family":"Claude 5","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2026-06-30","updated_at":"2026-08-28","modalities":["text","image","tools","computer-use"],"reasoning_mode":"default-on","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。adaptive thinking 默认开启（effort 默认 high，五档 low/medium/high/xhigh/max）；不再接受手动 extended thinking 与非默认 temperature/top_p/top_k（400 错误）；新分词器（同文本约多 30% token）。"},"context":{"max_tokens":1000000,"display":"1M","max_output":128000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":2,"output_per_m":10,"currency":"USD","source":"Anthropic 官方定价页（platform.claude.com）","as_of":"2026-08-28","note":"原定 2026-09-01 涨至 $3/$15 已取消，$2/$10 为永久标准价"},"links":{"official":"https://www.anthropic.com/claude/sonnet","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","paper":"https://www.anthropic.com/news/claude-sonnet-5"},"copy":{"one_liner":"接近 Opus 4.8 的 agent 能力，$2 / $10 永久定价。","highlights":["Terminal-Bench 2.1 80.4%（Sonnet 4.6 为 67.0%，+13.4），OSWorld-Verified 81.2% 与 Opus 4.8 基本持平","$2 / $10 原为限时价，2026-08 起改为永久标准价，比 Sonnet 4.6 便宜 1/3","1M 上下文 + 128K 输出，AA 智能指数 55，Free / Pro 计划默认模型"],"pitfalls":["新分词器使同文本 token 数多 1.0~1.35 倍，实际账单降幅小于标价降幅","temperature / top_p / top_k 非默认值直接 400，旧代码需清理采样参数","官方分数均为默认 effort；misalignment 率高于 Opus 4.8，漏洞利用类能力被刻意压低"],"logic_ability":"工程型推理接近 Opus 4.8：Terminal-Bench 2.1 80.4% 甚至高于 Opus 4.8 的 74.6%，HLE（带工具）57.4% 对 57.9%，GDPval-AA v2 1618 对 1615（官方公告表，Vellum 转录）。考试型推理略弱：SWE-bench Pro 63.2% 低于 Opus 4.8 的 69.2%，官方未公布 GPQA Diamond。失败模式：默认 effort 高时输出偏长；复杂多文件重构仍应上 Opus 5。","best_for":["生产环境 agent（浏览器 / 终端工具使用）","高吞吐编码助手","对话产品默认模型"],"not_for":["最难的跨仓库重构","需要私有化部署的合规场景"]},"capability_notes":{"coding":"Terminal-Bench 2.1 80.4%、SWE-bench Pro 63.2%、CursorBench 57%（官方公告表，Vellum 转录）；AA 实测 Terminal-Bench 较 Sonnet 4.6 +9 分。","reasoning":"HLE（带工具）57.4%（官方公告表）；AA 智能指数 55（max 档，2026-08-28）；CritPt 17%，较 Sonnet 4.6 +14 分（AA）。","agent":"OSWorld-Verified 81.2%、GDPval-AA v2 1618（官方公告表，与 Opus 4.8 的 1615 持平）；BrowseComp 相对 Sonnet 4.6 大幅提升（官方）。","instruction":"官方称幻觉率与谄媚率均低于 Sonnet 4.6；不支持 assistant 预填，需改用 structured outputs（官方文档）。","multimodal":"支持文本 + 图像输入、文本输出；官方未公布 MMMU 等视觉基准。","chinese":"中文可用，缺少独立中文评测。"},"i18n":{"en":{"name_zh":"Claude 5 mid-tier","one_liner":"Near Opus 4.8 agent ability at a permanent $2 / $10 price.","highlights":["Terminal-Bench 2.1 80.4% (Sonnet 4.6 was 67.0%, +13.4); OSWorld-Verified 81.2%, roughly level with Opus 4.8","$2 / $10 was a limited-time price, permanent standard price since 2026-08; 1/3 cheaper than Sonnet 4.6","1M context + 128K output, AA Intelligence Index 55, default model on Free / Pro plans"],"pitfalls":["New tokenizer yields 1.0-1.35x more tokens for the same text, so real bills drop less than list price","Non-default temperature / top_p / top_k return 400 directly; old code must strip sampling params","Official scores are all at default effort; misalignment rate higher than Opus 4.8, exploit-style capability deliberately suppressed"],"logic_ability":"Engineering reasoning near Opus 4.8: Terminal-Bench 2.1 80.4% even above Opus 4.8's 74.6%, HLE (with tools) 57.4% vs 57.9%, GDPval-AA v2 1618 vs 1615 (official announcement table, transcribed by Vellum). Exam-style reasoning slightly weaker: SWE-bench Pro 63.2% below Opus 4.8's 69.2%; GPQA Diamond not officially published. Failure modes: verbose output at high default effort; complex multi-file refactors should still go to Opus 5.","best_for":["Production agents (browser / terminal tool use)","High-throughput coding assistants","Default model for chat products"],"not_for":["The hardest cross-repo refactors","Compliance scenarios needing on-prem deployment"],"capability_notes":{"coding":"Terminal-Bench 2.1 80.4%, SWE-bench Pro 63.2%, CursorBench 57% (official announcement table, transcribed by Vellum); AA-measured Terminal-Bench +9 points over Sonnet 4.6.","reasoning":"HLE (with tools) 57.4% (official announcement table); AA Intelligence Index 55 (max, 2026-08-28); CritPt 17%, +14 points over Sonnet 4.6 (AA).","agent":"OSWorld-Verified 81.2%, GDPval-AA v2 1618 (official announcement table, level with Opus 4.8's 1615); BrowseComp greatly improved over Sonnet 4.6 (official).","instruction":"Officially lower hallucination and sycophancy rates than Sonnet 4.6; assistant prefill unsupported, use structured outputs instead (official docs).","multimodal":"Text + image input, text output; no official vision benchmarks such as MMMU published.","chinese":"Chinese usable; no independent Chinese evaluation."}},"ja":{"name_zh":"Claude 5 ミドルクラス","one_liner":"Opus 4.8 に迫る agent 能力を恒久価格 $2 / $10 で。","highlights":["Terminal-Bench 2.1 80.4%（Sonnet 4.6 は 67.0%、+13.4）、OSWorld-Verified 81.2% で Opus 4.8 とほぼ同等","$2 / $10 は当初期間限定だったが 2026-08 から恒久標準価格に。Sonnet 4.6 より 1/3 安い","1M コンテキスト + 128K 出力、AA 知能指数 55、Free / Pro プランのデフォルトモデル"],"pitfalls":["新トークナイザーにより同じテキストでトークン数が 1.0〜1.35 倍になり、実請求額の下げ幅は定価ほどではない","temperature / top_p / top_k をデフォルト以外にすると即 400。旧コードはサンプリングパラメータの削除が必要","公式スコアはすべてデフォルト effort。misalignment 率は Opus 4.8 より高く、脆弱性悪用系能力は意図的に抑制"],"logic_ability":"エンジニアリング型推論は Opus 4.8 に匹敵：Terminal-Bench 2.1 80.4% は Opus 4.8 の 74.6% を上回り、HLE（ツールあり）57.4% 対 57.9%、GDPval-AA v2 1618 対 1615（公式発表表、Vellum 転記）。試験型推論はやや弱く、SWE-bench Pro 63.2% は Opus 4.8 の 69.2% を下回る。GPQA Diamond は公式未公表。失敗モード：デフォルト effort が高いと出力が長い。複雑な複数ファイルのリファクタリングは Opus 5 に任せるべき。","best_for":["本番環境の agent（ブラウザ / ターミナルのツール利用）","高スループットのコーディングアシスタント","対話製品のデフォルトモデル"],"not_for":["最難関のリポジトリ横断リファクタリング","オンプレ展開が必要なコンプライアンス用途"],"capability_notes":{"coding":"Terminal-Bench 2.1 80.4%、SWE-bench Pro 63.2%、CursorBench 57%（公式発表表、Vellum 転記）。AA 実測 Terminal-Bench は Sonnet 4.6 比 +9 点。","reasoning":"HLE（ツールあり）57.4%（公式発表表）。AA 知能指数 55（max、2026-08-28）。CritPt 17%、Sonnet 4.6 比 +14 点（AA）。","agent":"OSWorld-Verified 81.2%、GDPval-AA v2 1618（公式発表表、Opus 4.8 の 1615 と同等）。BrowseComp は Sonnet 4.6 比で大幅向上（公式）。","instruction":"公式によればハルシネーション率・追従率とも Sonnet 4.6 より低い。assistant プリフィル非対応で structured outputs を使う必要あり（公式ドキュメント）。","multimodal":"テキスト + 画像入力、テキスト出力に対応。MMMU などの視覚ベンチマークは公式未公表。","chinese":"中国語は利用可能。独立した中国語評価はなし。"}}},"complete":true,"sheet":{"architecture_md":"**未披露**。Anthropic 未公布参数量、Dense/MoE、注意力类型。\n\n已知接口特性（官方文档，2026-08-28）：\n- **Adaptive thinking 默认开启**：不传 `thinking` 字段即带思考；关闭需显式 `thinking: {type: \"disabled\"}`\n- `effort` 五档 `low` / `medium` / `high`（默认）/ `xhigh` / `max`，与 Opus 4.8 对齐；`xhigh` 为 Sonnet 线新增\n- **三个破坏性变更**（相对 Sonnet 4.6）：手动 extended thinking（`budget_tokens`）返回 400；`temperature` / `top_p` / `top_k` 非默认值返回 400；assistant 预填仍不支持\n- 1M 上下文既是默认也是最大（无更小档），128K 最大输出；Batch API 加 `output-300k-2026-03-24` beta 头可到 300K 输出\n- 输入：文本 + 图像；输出：文本\n- **新分词器**：同文本约多 30%（官方公告口径 1.0~1.35×），`max_tokens`、缓存断点、上下文容量都要重新计算\n- 工具集与 Sonnet 4.6 相同，另支持 browser use 工具和稳定版 `computer_toolset_20260801`（Claude API / Google Cloud）；不提供 Priority Tier\n- 首个带实时网络安全防护的 Sonnet：被拦截时返回 HTTP 200 + `stop_reason: \"refusal\"`\n- 可靠知识截止 2026-01（训练数据截止同为 2026-01）","memory_md":"无自建选项。\n\n**成本模型**（官方定价页，2026-08-28）：\n- 输入 $2 / 输出 $10 / MTok；原定 2026-09-01 涨回 $3 / $15 的计划已于 2026-08-10 取消，$2 / $10 为永久标准价\n- 缓存写 $2.50（5 分钟）/ $4（1 小时），缓存读 $0.20（基价 10%）\n- Batch 五折：$1 / $5\n- 1M 长上下文不加价；Bedrock / Vertex / Foundry 按各自价目\n\n**什么时候会变贵**：\n- 新分词器让同一段文本多约 30% token，标价降 1/3 但等价请求账单降幅明显小于此\n- AA 实测 max 档单任务 $1.72~2.29，约为 Sonnet 4.6 的 2 倍、与 Opus 4.8 相当——因为输出 token 多 ~40%、知识工作类 agent 轮次约 3 倍；max 比 low 多约 6 倍轮次\n- 思考默认开启，`max_tokens` 是思考 + 正文总上限，旧的无思考预算会截断\n\n速度：AA 实测约 95 tok/s（max 档），高于均值；高 effort 下 TTFT 很长（AA 记录 max 档 225 s），交互场景请降 effort 或流式。","training_md":"- 训练细节未披露；训练数据截止 2026-01\n- 官方定位：Sonnet 4.6 的直接替代，最大提升在编码与 agent 任务，「接近 Opus 4.8 的能力、更低价格」\n- 官方公告表（默认 effort）：SWE-bench Pro 63.2%（4.6：58.1%）、Terminal-Bench 2.1 80.4%（4.6：67.0%；Opus 4.8：74.6%）、OSWorld-Verified 81.2%（4.6：78.5%）、HLE 带工具 57.4%（4.6：46.8%）、GDPval-AA v2 1618\n- 行为：幻觉率与谄媚率低于 Sonnet 4.6，对恶意请求的拒绝更准；整体「不良行为率」下降\n- 安全：默认开启网络安全防护；漏洞利用开发能力被刻意压低（Firefox 147 可用 exploit 0%）；系统卡见 anthropic.com/claude-sonnet-5-system-card\n- 已知取舍：misalignment 率高于 Opus 4.8；官方分数均为默认 effort，未公布 GPQA Diamond","ecosystem_md":"- **API 渠道**：Claude API（`claude-sonnet-5`，无日期后缀即固定快照）、Amazon Bedrock（`anthropic.claude-sonnet-5`，含 InvokeModel）、Claude Platform on AWS、Google Cloud Vertex AI、Microsoft Foundry\n- claude.ai：Free / Pro 计划默认模型，Max / Team / Enterprise 可用；Claude Code 默认 effort `high`\n- SDK：官方 Python / TypeScript / C# / Go / Java / PHP / Ruby，均有 `ModelClaudeSonnet5` 常量\n- 工具：bash、text editor、web search、code execution、memory、tool search、browser use、computer_toolset_20260801\n- 支持 ZDR（零数据保留）协议；**不可微调**\n- 中文文档：弱（platform.claude.com 仅英文，社区中文资料较多）","versions_md":"- 上代：Claude Sonnet 4.6（$3 / $15，仍可用，Legacy）；更早 Sonnet 4.5（2025-09）\n- 同代：Claude Opus 5（2026-07-24，$5 / $25）、Claude Fable 5（2026-06-09，$10 / $50）、Claude Haiku 4.5（$1 / $5，200K）\n- 版本策略：4.6 代起模型 ID 不带日期，`claude-sonnet-5` 本身就是固定快照\n- 退役承诺：不早于 2027-06-30（Bedrock / Google Cloud 自行定）"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"弱"},"runtime":{"tok_s":95,"latency_s":225.59,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"command-a-plus","name":"Command A+","name_zh":"Cohere Command A+","aliases":["command-a-plus-05-2026","CohereLabs/command-a-plus-05-2026"],"vendor":"Cohere","family":"Command","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/CohereLabs/command-a-plus-05-2026","status":"current","released_at":"2026-05-20","updated_at":"2026-08-28","modalities":["text","image","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"218B","active_params":"25B","total_params_b":218,"active_params_b":25,"experts":128,"active_experts":8,"shared_expert":true,"layers":32,"hidden_size":4096,"vocab_size":262144,"kv_heads":8,"head_dim":128,"attention":"混合：3 层 4096 滑窗 + 1 层全注意力交替，128 Q / 8 KV","notes":"Cohere 首个 MoE，sigmoid token-choice 路由，4 共享专家；SigLIP 视觉编码器。官方 W4A4 量化宣称无损。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":65536},"memory":{"weight_gb":{"bf16":436,"fp8":218,"q4":126,"q8":232},"estimated":true,"ref_hw_80gb":"官方 W4A4 需 2×H100 或 1×B200","ref_hw_8x80gb":"BF16 官方 8×H100"},"links":{"official":"https://huggingface.co/CohereLabs/command-a-plus-05-2026","hf":"https://huggingface.co/CohereLabs/command-a-plus-05-2026"},"variants":[{"kind":"fp8","publisher":"CohereLabs","repo":"CohereLabs/command-a-plus-05-2026-fp8","url":"https://huggingface.co/CohereLabs/command-a-plus-05-2026-fp8","note":"官方 FP8","sizes":{"fp8":225}},{"kind":"other","publisher":"CohereLabs","repo":"CohereLabs/command-a-plus-05-2026-w4a4","url":"https://huggingface.co/CohereLabs/command-a-plus-05-2026-w4a4","note":"官方 W4A4（4-bit 权重 + 4-bit 激活）"},{"kind":"other","publisher":"CohereLabs","repo":"CohereLabs/command-a-plus-05-2026-bf16","url":"https://huggingface.co/CohereLabs/command-a-plus-05-2026-bf16","note":"官方 BF16","sizes":{"bf16":437.5}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/command-a-plus-05-2026-GGUF","url":"https://huggingface.co/bartowski/command-a-plus-05-2026-GGUF","sizes":{"q4":135.1,"q5":157.8,"q6":189.8,"q8":232}}],"copy":{"one_liner":"Cohere 首个 Apache-2.0 开源 MoE，218B，48 语言。","highlights":["Apache-2.0，Cohere 首个完全开源模型","原生引用溯源（citation），48 种语言含中文","AIME 2025 90%（AA 报道），W4A4 无损量化 2×H100 可跑"],"pitfalls":["AA 智能指数 23，通用推理仅 Haiku 4.5 级","128K 上下文，低于同级 256K+","官方未公布 API 价格与完整评测表"],"logic_ability":"考试型推理数学强（AIME 90%）但 GPQA 约 76%、HLE 约 11% 平平；工程型推理 Terminal-Bench Hard 约 25% 偏弱。","best_for":["企业 RAG / 引用溯源","多语言翻译与文档处理"],"not_for":["编程 agent 主力","单卡部署"]},"capability_notes":{"reasoning":"GPQA Diamond 约 76%、HLE 约 11%（Artificial Analysis 报道）。","math":"AIME 2025 90%（Artificial Analysis 报道）。","multimodal":"MMMU-Pro 63%（Artificial Analysis 报道）。","chinese":"官方 48 语言含中文。"},"ecosystem":{"engines":["vLLM","Transformers"],"zh_docs":"无"},"i18n":{"en":{"name_zh":"Cohere Command A+","one_liner":"Cohere's first Apache-2.0 open MoE, 218B, 48 languages.","highlights":["Apache-2.0, Cohere's first fully open model","Native citation grounding, 48 languages including Chinese","AIME 2025 90% (AA report); lossless W4A4 quantization runs on 2x H100"],"pitfalls":["AA Intelligence Index 23; general reasoning only at Haiku 4.5 level","128K context, below the 256K+ of peers","No official API price or full evaluation table published"],"logic_ability":"Exam-style reasoning strong in math (AIME 90%) but GPQA about 76% and HLE about 11% are unremarkable; engineering reasoning weak at about 25% on Terminal-Bench Hard.","best_for":["Enterprise RAG / citation grounding","Multilingual translation and document processing"],"not_for":["Primary coding agent","Single-GPU deployment"],"capability_notes":{"reasoning":"GPQA Diamond about 76%, HLE about 11% (Artificial Analysis report).","math":"AIME 2025 90% (Artificial Analysis report).","multimodal":"MMMU-Pro 63% (Artificial Analysis report).","chinese":"Officially 48 languages including Chinese."}},"ja":{"name_zh":"Cohere Command A+","one_liner":"Cohere 初の Apache-2.0 オープン MoE、218B、48 言語","highlights":["Apache-2.0、Cohere 初の完全オープンモデル","ネイティブの引用（citation）機能、中国語を含む 48 言語","AIME 2025 90%（AA 報道）。W4A4 無損失量子化で 2×H100 で動作"],"pitfalls":["AA 知能指数 23、汎用推論は Haiku 4.5 級","128K コンテキストで同クラスの 256K+ に劣る","公式 API 価格と完全な評価表は未公表"],"logic_ability":"試験型推論は数学が強い（AIME 90%）が GPQA 約 76%、HLE 約 11% は平凡。エンジニアリング型推論は Terminal-Bench Hard 約 25% と弱め。","best_for":["企業 RAG / 引用付き回答","多言語翻訳と文書処理"],"not_for":["コーディング agent の主力","単一 GPU での展開"],"capability_notes":{"reasoning":"GPQA Diamond 約 76%、HLE 約 11%（Artificial Analysis 報道）。","math":"AIME 2025 90%（Artificial Analysis 報道）。","multimodal":"MMMU-Pro 63%（Artificial Analysis 報道）。","chinese":"公式に中国語を含む 48 言語対応。"}}},"complete":false,"runtime":{"tok_s":254,"latency_s":0.39,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"aliases":["c4ai-command-r-plus","Command R+ 08-2024"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"agent":"多步工具调用、引用溯源原生。","chinese":"中文为 10 语言之一，可用。"},"complete":false,"id":"command-r-plus","name":"Command R+","name_zh":"Command R+ · 104B","vendor":"Cohere","family":"Command R","license":"CC-BY-NC-4.0（附 Cohere 可接受使用政策）","license_commercial":false,"weights_url":"https://huggingface.co/CohereLabs/c4ai-command-r-plus","superseded_by":"command-a-plus","released_at":"2024-04-04","architecture":{"type":"dense","total_params":"104B","total_params_b":104,"layers":64,"hidden_size":12288,"vocab_size":256000,"kv_heads":8,"head_dim":128,"attention":"GQA（96 Q 头 / 8 KV 头）","notes":"RAG 与工具调用原生模板（引用溯源）；10 语言；128K。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":208,"q8":110.2,"q4":62.8},"estimated":false,"kv_per_token_kib":256,"kv_note":"256 KiB/token。","ref_hw_24gb":"不可行","ref_hw_80gb":"Q4 62.8 GB 单卡可跑；BF16 3×80GB","ref_hw_8x80gb":"BF16 高并发"},"pricing":{"input_per_m":2.5,"output_per_m":10,"currency":"USD","source":"Cohere 官方 API（command-r-plus-08-2024）","as_of":"2026-08-28"},"links":{"official":"https://cohere.com/blog/command-r-plus-microsoft-azure","hf":"https://huggingface.co/CohereLabs/c4ai-command-r-plus"},"copy":{"one_liner":"企业 RAG 专用的 104B，带引用溯源的工具调用先驱。","highlights":["RAG 引用（grounded generation）与多步工具调用原生模板","128K 上下文，10 种商业语言","发布时 Arena 开源第一，击败 GPT-4 早期版"],"pitfalls":["CC-BY-NC，权重禁止商用（只能走 API）","104B Dense 显存大","被 Command A（111B，2025-03）取代"],"logic_ability":"非推理模型 2024 年中游：MMLU 75.7%（社区测），RAG 场景表现优于同期开源，纯推理一般。","best_for":["企业 RAG 原型（非商用）","多语言检索问答"],"not_for":["商用自部署","数学 / 代码"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama"],"finetune":"LoRA 可行（非商用）","zh_docs":"弱"}},{"aliases":["dbrx-instruct"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"coding":"HumanEval 70.1%（官方）。","math":"GSM8K 66.9%（官方）。","knowledge":"MMLU 73.7%（官方）。"},"complete":false,"id":"dbrx","name":"DBRX","name_zh":"DBRX · 132B-A36B","vendor":"Databricks","family":"DBRX","license":"Databricks Open Model License","license_commercial":"restricted","weights_url":"https://huggingface.co/databricks/dbrx-instruct","released_at":"2024-03-27","architecture":{"type":"moe","total_params":"132B","active_params":"36B","total_params_b":132,"active_params_b":36,"experts":16,"active_experts":4,"shared_expert":false,"layers":40,"hidden_size":6144,"vocab_size":100352,"kv_heads":8,"head_dim":128,"attention":"GQA（48 Q 头 / 8 KV 头）","notes":"细粒度 MoE：16 专家 top-4（组合数 65 倍于 Mixtral）；12T token；32K。","undisclosed":false},"context":{"max_tokens":32768,"display":"32K"},"memory":{"weight_gb":{"bf16":264,"q8":139.9,"q4":76.6},"estimated":true,"kv_per_token_kib":160,"kv_note":"160 KiB/token。","ref_hw_24gb":"不可行","ref_hw_80gb":"Q4 约 77 GB 单卡勉强；BF16 4×80GB","ref_hw_8x80gb":"BF16 高并发"},"links":{"official":"https://www.databricks.com/blog/introducing-dbrx-new-state-art-open-llm","github":"https://github.com/databricks/dbrx","hf":"https://huggingface.co/databricks/dbrx-instruct"},"copy":{"one_liner":"Databricks 的 132B 细粒度 MoE，短暂的开源榜首。","highlights":["发布时超越 Mixtral 8x7B / Grok-1，代码 HumanEval 70.1%","16 选 4 细粒度专家设计早于 DeepSeek 流行","Databricks 平台原生集成"],"pitfalls":["许可限制月活 7 亿并禁止用输出训练他模","仅 Databricks 生态活跃，社区量化少","三周后被 Llama 3 70B 超越，系列无后续"],"logic_ability":"2024 年初一线：MMLU 73.7%、GSM8K 66.9%、HumanEval 70.1%（官方）。","best_for":["MoE 设计研究","Databricks 平台用户"],"not_for":["新项目"]},"ecosystem":{"engines":["vLLM","TensorRT-LLM","llama.cpp"],"finetune":"困难","zh_docs":"无"}},{"aliases":["DeepSeek-Coder-V2-Instruct"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"coding":"HumanEval 90.2%、LiveCodeBench 43.4%（官方）。","math":"MATH 75.7%（官方）。"},"complete":false,"id":"deepseek-coder-v2","name":"DeepSeek-Coder-V2","name_zh":"DeepSeek-Coder-V2 · 236B","vendor":"DeepSeek","vendor_zh":"深度求索","family":"DeepSeek-Coder","license":"DeepSeek License（代码 MIT）","license_commercial":true,"weights_url":"https://huggingface.co/deepseek-ai/DeepSeek-Coder-V2-Instruct","superseded_by":"deepseek-v3","released_at":"2024-06-17","architecture":{"type":"moe","total_params":"236B","active_params":"21B","total_params_b":236,"active_params_b":21,"experts":160,"active_experts":6,"shared_expert":true,"layers":60,"hidden_size":5120,"vocab_size":102400,"kv_heads":128,"head_dim":192,"attention":"MLA","notes":"DeepSeek-V2 中间检查点续训 6T 代码/数学；338 种语言；128K。另有 16B Lite 版。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":472,"q8":250.2,"q4":142.5},"estimated":false,"kv_per_token_kib":67.5,"kv_note":"MLA ≈ 67.5 KiB/token。","ref_hw_24gb":"不可行","ref_hw_80gb":"Q4 142.5 GB 需 2×80GB","ref_hw_8x80gb":"BF16 8×80GB"},"links":{"official":"https://www.deepseek.com/","paper":"https://arxiv.org/abs/2406.11931","github":"https://github.com/deepseek-ai/DeepSeek-Coder-V2","hf":"https://huggingface.co/deepseek-ai/DeepSeek-Coder-V2-Instruct"},"copy":{"one_liner":"首个代码能力宣称超过 GPT-4 Turbo 的开源模型。","highlights":["官方称 HumanEval / MBPP / LiveCodeBench 超越 GPT-4 Turbo","338 种编程语言，128K 上下文","16B Lite 版单卡可跑，性价比极高"],"pitfalls":["Instruct 版通用对话较弱","236B 版部署门槛高","V2.5 合并了 Chat 与 Coder，此版本停更"],"logic_ability":"代码与数学强（HumanEval 90.2%、MATH 75.7%，官方），通用推理属 V2 水平。","best_for":["代码生成 / 补全","数学"],"not_for":["通用助手"]},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp"],"finetune":"困难","zh_docs":"有"}},{"id":"deepseek-r1-0528","name":"DeepSeek-R1-0528","name_zh":"深度求索 R1（0528）","aliases":["DeepSeek R1","deepseek-r1"],"vendor":"DeepSeek","vendor_zh":"深度求索","family":"DeepSeek R1","license":"MIT","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/deepseek-ai/DeepSeek-R1-0528","status":"superseded","superseded_by":"deepseek-v4-pro","released_at":"2025-05-28","updated_at":"2025-12-20","modalities":["text","tools"],"reasoning_mode":"default-on","architecture":{"type":"moe","total_params":"671B","active_params":"37B","total_params_b":671,"active_params_b":37,"experts":256,"active_experts":8,"shared_expert":true,"layers":61,"hidden_size":7168,"kv_heads":1,"head_dim":288,"attention":"MLA","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":65536},"memory":{"weight_gb":{"fp8":690,"q4":400},"kv_per_token_kib":70,"estimated":true,"ref_hw_8x80gb":"FP8 8×H100"},"pricing":{"input_per_m":0.55,"output_per_m":2.19,"currency":"USD","source":"DeepSeek 官方 API（历史价）","as_of":"2025-11-01"},"links":{"official":"https://www.deepseek.com/","hf":"https://huggingface.co/deepseek-ai/DeepSeek-R1-0528","github":"https://github.com/deepseek-ai/DeepSeek-R1","paper":"https://arxiv.org/abs/2501.12948"},"copy":{"one_liner":"开源推理模型的里程碑，已被 V3.2 思考模式覆盖。","highlights":["MIT 许可","AIME 2025 87.5%","Nature 论文公开 RL 配方"],"pitfalls":["已被 V3.2 替代","思考链很长，成本高","无工具调用（0528 前）"],"logic_ability":"考试型推理强，工程型中等，思考不可关。","best_for":["研究 / 蒸馏源"],"not_for":["新部署（用 V3.2）"]},"capability_notes":{},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp"],"zh_docs":"有"},"complete":false},{"aliases":["DeepSeek R1","deepseek-reasoner (2025-01)"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"default-on","capability_notes":{"coding":"LiveCodeBench 65.9%、Codeforces 2029 分、SWE-bench Verified 49.2%（官方）。","reasoning":"GPQA Diamond 71.5%（官方）。","math":"AIME 2024 79.8%、MATH-500 97.3%（官方）。","knowledge":"MMLU-Pro 84.0%（官方）。","agent":"弱：工具调用不在训练目标内。","chinese":"C-Eval 91.8%（官方）。"},"complete":true,"id":"deepseek-r1","name":"DeepSeek-R1","name_zh":"DeepSeek-R1（0120）","vendor":"DeepSeek","vendor_zh":"深度求索","family":"DeepSeek-R1","license":"MIT","license_commercial":true,"weights_url":"https://huggingface.co/deepseek-ai/DeepSeek-R1","superseded_by":"deepseek-r1-0528","released_at":"2025-01-20","architecture":{"type":"moe","total_params":"671B","active_params":"37B","total_params_b":671,"active_params_b":37,"experts":256,"active_experts":8,"shared_expert":true,"layers":61,"hidden_size":7168,"vocab_size":129280,"kv_heads":128,"head_dim":192,"attention":"MLA（kv_lora_rank 512 + rope 64）","notes":"基于 V3-Base，GRPO 纯强化学习 + 冷启动 SFT 两阶段；始终输出 <think>；128K。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":65536},"memory":{"weight_gb":{"bf16":1342,"q8":711.3,"q4":404.4,"fp8":671},"estimated":false,"kv_per_token_kib":68.6,"kv_note":"MLA 压缩 KV：(512+64)×61×2 B ≈ 68.6 KiB/token；128K 上下文仅 8.6 GB。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（单卡）","ref_hw_8x80gb":"官方 FP8 671 GB 权重 8×H100 可服务（需 FP8 引擎）；BF16 需 16×80GB；Q4 404 GB 8×80GB 可跑"},"links":{"official":"https://www.deepseek.com/","github":"https://github.com/deepseek-ai/DeepSeek-R1","paper":"https://arxiv.org/abs/2501.12948","hf":"https://huggingface.co/deepseek-ai/DeepSeek-R1"},"copy":{"one_liner":"开源推理模型的引爆点，纯 RL 训出 o1 级能力。","highlights":["首个开源 o1 级推理模型，AIME 2024 79.8%、GPQA 71.5%（官方）","MIT 许可，思维链完整公开","同时发布 6 个蒸馏小模型（1.5B–70B）"],"pitfalls":["工具调用与 JSON 输出弱，agent 场景不稳","思考链极长且偶有语言混杂，0528 版才修复","671B 本地门槛高，API 高峰期不稳定"],"logic_ability":"考试型推理 2025 年初开源第一：AIME 2024 79.8%、MATH-500 97.3%、GPQA 71.5%、LiveCodeBench 65.9%、SWE-bench Verified 49.2%（官方）。逻辑严密、善于自我纠错，但不擅长工具调用与结构化输出，且过度思考简单问题。","best_for":["推理 / 数学 / 代码","API 低成本调用"],"not_for":["单机本地部署","低延迟对话"]},"pricing":{"input_per_m":0.55,"output_per_m":2.19,"currency":"USD","source":"DeepSeek 官方 API（deepseek-reasoner，2025-01 定价）","as_of":"2026-08-28"},"sheet":{"architecture_md":"**类型**：MoE Transformer，与 DeepSeek-V3 完全相同的结构，671B 总参 / 37B 激活。\n\n- 61 层，隐藏维 7168，256 路由专家 top-8 + 1 共享专家\n- MLA 注意力，KV ≈ 68.6 KiB/token\n- 输出格式：`<think>…</think>` 后接答案，思考链不可关闭\n- 上下文 128K，官方建议最大输出 32K（0528 版 64K）\n\n参考：《DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL》。","memory_md":"| 精度 | 权重大小 | 参考硬件 | 说明 |\n|---|---|---|---|\n| FP8 | ≈ 671 GB | 8×H100/H200 | 官方原生 FP8 |\n| BF16 | 1342 GB | 16×80GB | 需转换 |\n| Q4 | 404 GB | 8×80GB | GGUF Q4_K_M（unsloth） |\n| 1.58-bit | ≈ 131 GB | 2×80GB / 大内存 | unsloth 动态量化 |\n\n**KV Cache**：≈ 68.6 KiB/token；推理链常达 10–30K token，KV 预算按此规划。\n\n**参考配置**：\n- 8×H200：FP8 高并发\n- 本地：只能靠蒸馏版（R1-Distill-Qwen-32B 单卡 24GB Q4）","training_md":"- R1-Zero：V3-Base 直接 GRPO 强化学习，无 SFT，出现「顿悟时刻」\n- R1：冷启动少量长 CoT SFT → 推理 RL → 拒绝采样 SFT（80 万样本）→ 全场景 RL\n- 奖励：规则奖励（准确性 + 格式），不用神经奖励模型\n- 蒸馏：用 80 万样本 SFT Qwen2.5 / Llama 3 得到 6 个小模型\n- 官方建议 temperature 0.6，不要加 system prompt","ecosystem_md":"- HF：deepseek-ai/DeepSeek-R1、R1-Zero、R1-Distill-{Qwen,Llama}-{1.5B…70B}\n- 引擎：SGLang、vLLM、LMDeploy、llama.cpp\n- 微调：本体极困难；蒸馏版 LoRA 友好\n- 中文文档：有（GitHub 中文 README、API 文档中文）","versions_md":"- 后继：DeepSeek-R1-0528（2025-05，AIME 2025 87.5%，工具调用改善）→ V3.1 混合思考\n- 蒸馏系列：R1-Distill-Qwen-32B 是 2025 上半年最常用的本地推理模型\n- 上代：DeepSeek-R1-Lite-Preview（2024-11，仅 API）"},"ecosystem":{"engines":["SGLang","vLLM","LMDeploy","llama.cpp"],"finetune":"极困难（本体）","zh_docs":"有"}},{"aliases":["DeepSeek-V2-Chat","DeepSeek-V2 236B"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"coding":"LiveCodeBench(0901-0401) 32.5%（Chat，官方）。","knowledge":"MMLU 78.5%（官方基座）。","chinese":"C-Eval 81.7%（官方）。"},"complete":false,"id":"deepseek-v2","name":"DeepSeek-V2","name_zh":"DeepSeek-V2","vendor":"DeepSeek","vendor_zh":"深度求索","family":"DeepSeek-V2","license":"DeepSeek License（代码 MIT）","license_commercial":true,"weights_url":"https://huggingface.co/deepseek-ai/DeepSeek-V2-Chat","superseded_by":"deepseek-v3","released_at":"2024-05-06","architecture":{"type":"moe","total_params":"236B","active_params":"21B","total_params_b":236,"active_params_b":21,"experts":160,"active_experts":6,"shared_expert":true,"layers":60,"hidden_size":5120,"vocab_size":102400,"kv_heads":128,"head_dim":192,"attention":"MLA（kv_lora_rank 512 + rope 64）","notes":"MLA 与 DeepSeekMoE 首次亮相：160 路由专家 top-6 + 2 共享专家；上下文 128K。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":472,"q8":250.2,"q4":136.9},"estimated":true,"kv_per_token_kib":67.5,"kv_note":"MLA 压缩 KV：(512+64)×60×2 B ≈ 67.5 KiB/token，比同规模 GQA 小 10 倍以上。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（单卡）；Q4 约 137 GB 需 2×80GB","ref_hw_8x80gb":"BF16 472 GB，8×80GB 可服务"},"links":{"official":"https://www.deepseek.com/","paper":"https://arxiv.org/abs/2405.04434","github":"https://github.com/deepseek-ai/DeepSeek-V2","hf":"https://huggingface.co/deepseek-ai/DeepSeek-V2-Chat"},"copy":{"one_liner":"MLA + DeepSeekMoE 的首秀，引爆中国 API 价格战。","highlights":["首创 MLA，KV 缓存减少 93%","236B 总参仅 21B 激活，训练成本比 67B Dense 低 42.5%","API 定价 1 元/百万 token 掀起价格战"],"pitfalls":["自定义架构，早期推理框架支持慢","非 Apache/MIT 权重许可（含使用限制条款）","被 V2.5 / V3 快速取代"],"logic_ability":"2024 年中开源一线：MMLU 78.5%、GSM8K 79.2%（基座，官方）。无推理链训练。","best_for":["MLA / MoE 架构研究"],"not_for":["新项目"]},"ecosystem":{"engines":["vLLM","SGLang"],"finetune":"困难（MLA + MoE）","zh_docs":"有"}},{"aliases":["DeepSeek V3 0324","deepseek-chat (2025-03)"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"coding":"LiveCodeBench 49.2%（官方）。","reasoning":"GPQA Diamond 68.4%（官方）。","math":"AIME 2024 59.4%（官方）。","agent":"函数调用较 V3 明显改善。","chinese":"中文顶级。"},"complete":false,"id":"deepseek-v3-0324","name":"DeepSeek-V3-0324","name_zh":"DeepSeek-V3（0324）","vendor":"DeepSeek","vendor_zh":"深度求索","family":"DeepSeek-V3","license":"MIT","license_commercial":true,"weights_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3-0324","superseded_by":"deepseek-v3-1","released_at":"2025-03-24","architecture":{"type":"moe","total_params":"671B","active_params":"37B","total_params_b":671,"active_params_b":37,"experts":256,"active_experts":8,"shared_expert":true,"layers":61,"hidden_size":7168,"vocab_size":129280,"kv_heads":128,"head_dim":192,"attention":"MLA（kv_lora_rank 512 + rope 64）","notes":"结构同 V3；后训练吸收 R1 的 RL 方法；128K。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":8192},"memory":{"weight_gb":{"bf16":1342,"q8":711.3,"q4":404.4,"fp8":671},"estimated":false,"kv_per_token_kib":68.6,"kv_note":"MLA 压缩 KV：(512+64)×61×2 B ≈ 68.6 KiB/token；128K 上下文仅 8.6 GB。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（单卡）","ref_hw_8x80gb":"官方 FP8 671 GB 权重 8×H100 可服务（需 FP8 引擎）；BF16 需 16×80GB；Q4 404 GB 8×80GB 可跑"},"links":{"official":"https://www.deepseek.com/","github":"https://github.com/deepseek-ai/DeepSeek-V3","hf":"https://huggingface.co/deepseek-ai/DeepSeek-V3-0324"},"copy":{"one_liner":"V3 的大幅增强版，非推理模型中首次逼近 Claude 3.7。","highlights":["GPQA 59.1 → 68.4、LiveCodeBench 39.2 → 49.2（官方）","前端代码与函数调用明显改善","MIT 许可，权重与 V3 结构相同可直接替换"],"pitfalls":["仍无推理模式","671B 部署门槛不变","被 V3.1（思考 + 非思考合一）取代"],"logic_ability":"非推理模型顶级：GPQA 68.4%、AIME 2024 59.4%、LiveCodeBench 49.2%（官方）。吸收了 R1 的 RL 技巧，单次前向的推理能力大幅提升，工具调用可用。","best_for":["通用与代码","API 低成本调用"],"not_for":["单机本地部署","多模态"]},"pricing":{"input_per_m":0.27,"output_per_m":1.1,"currency":"USD","source":"DeepSeek 官方 API（deepseek-chat，2025-03 定价）","as_of":"2026-08-28"},"ecosystem":{"engines":["SGLang","vLLM","LMDeploy","llama.cpp"],"finetune":"极困难","zh_docs":"有"}},{"aliases":["DeepSeek V3.1","deepseek-chat / deepseek-reasoner (2025-08)"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"optional","capability_notes":{"coding":"LiveCodeBench 74.8%（思考）、SWE-bench Verified 66.0%（非思考 agent，官方）。","reasoning":"GPQA Diamond 80.1%、HLE 15.9%（思考，官方）。","math":"AIME 2025 88.4%（思考，官方）。","agent":"Terminal-bench 31.3%、BrowseComp 30.0%（官方）。","chinese":"中文顶级。"},"complete":false,"id":"deepseek-v3-1","name":"DeepSeek-V3.1","name_zh":"DeepSeek-V3.1","vendor":"DeepSeek","vendor_zh":"深度求索","family":"DeepSeek-V3","license":"MIT","license_commercial":true,"weights_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.1","superseded_by":"deepseek-v3-2","released_at":"2025-08-21","architecture":{"type":"moe","total_params":"671B","active_params":"37B","total_params_b":671,"active_params_b":37,"experts":256,"active_experts":8,"shared_expert":true,"layers":61,"hidden_size":7168,"vocab_size":129280,"kv_heads":128,"head_dim":192,"attention":"MLA（kv_lora_rank 512 + rope 64）","notes":"结构同 V3；两阶段长上下文扩展（32K 630B token + 128K 209B token）；混合思考模板；128K。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":65536},"memory":{"weight_gb":{"bf16":1342,"q8":711.3,"q4":405.4,"fp8":671},"estimated":false,"kv_per_token_kib":68.6,"kv_note":"MLA 压缩 KV：(512+64)×61×2 B ≈ 68.6 KiB/token；128K 上下文仅 8.6 GB。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（单卡）","ref_hw_8x80gb":"官方 FP8 671 GB 权重 8×H100 可服务（需 FP8 引擎）；BF16 需 16×80GB；Q4 404 GB 8×80GB 可跑"},"links":{"official":"https://www.deepseek.com/","github":"https://github.com/deepseek-ai/DeepSeek-V3","hf":"https://huggingface.co/deepseek-ai/DeepSeek-V3.1"},"copy":{"one_liner":"思考与非思考合一，面向 agent 时代的首个 DeepSeek。","highlights":["一套权重两种模式（思考 / 非思考），思考效率比 R1-0528 高","SWE-bench Verified 66.0%、Terminal-bench 31.3%（官方，agent 模式）","续训 840B token 扩展 128K；UE8M0 FP8 适配国产芯片"],"pitfalls":["思考模式不支持工具调用（需切非思考）","部分用户反馈中文回答中夹杂「极」字等 token 异常","2025-09 即被 V3.1-Terminus 修正，再被 V3.2 取代"],"logic_ability":"思考模式：AIME 2025 88.4%、GPQA 80.1%、LiveCodeBench 74.8%、HLE 15.9%（官方）；非思考模式 SWE-bench Verified 66.0%。推理接近 R1-0528 但 token 用量更少，agent 能力是 DeepSeek 首次达到实用水平。","best_for":["通用与代码","API 低成本调用"],"not_for":["单机本地部署","多模态"]},"pricing":{"input_per_m":0.56,"output_per_m":1.68,"currency":"USD","source":"DeepSeek 官方 API（2025-09-06 起统一定价）","as_of":"2026-08-28"},"ecosystem":{"engines":["SGLang","vLLM","LMDeploy","llama.cpp"],"finetune":"极困难","zh_docs":"有"}},{"id":"deepseek-v3-2","name":"DeepSeek-V3.2","name_zh":"深度求索 V3.2","aliases":["DeepSeek V3.2","deepseek-chat","deepseek-reasoner"],"vendor":"DeepSeek","vendor_zh":"深度求索","family":"DeepSeek V3","license":"MIT","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.2","status":"superseded","released_at":"2025-12-01","updated_at":"2025-12-20","modalities":["text","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"685B","active_params":"37B","total_params_b":685,"active_params_b":37,"experts":256,"active_experts":8,"shared_expert":true,"layers":61,"hidden_size":7168,"vocab_size":129280,"kv_heads":1,"head_dim":288,"attention":"MLA + DSA（DeepSeek Sparse Attention）","notes":"在 V3.1-Terminus 基础上引入 DSA：lightning indexer 为每个 query 选 top-k 键，把注意力复杂度从 O(L²) 降到 O(L·k)。MLA 的 KV 压缩为 512 维潜向量 + 64 维 RoPE（表内 kv_heads=1、head_dim=288 为等效换算）。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":65536},"memory":{"weight_gb":{"fp8":690,"q4":380},"kv_per_token_kib":70,"kv_note":"MLA：每 token 每层约 576 维 × 2 B ≈ 1.1 KiB，61 层 ≈ 70 KiB/token（远低于同规模 GQA）。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（单卡）","ref_hw_8x80gb":"FP8 权重 690 GB，8×H100/H200 可部署；官方推荐 H200 或 8×H100 + FP8","estimated":true},"pricing":{"input_per_m":0.28,"output_per_m":0.42,"currency":"USD","source":"DeepSeek 官方 API 定价","as_of":"2025-12-20","note":"缓存命中输入 $0.028"},"links":{"official":"https://www.deepseek.com/","hf":"https://huggingface.co/deepseek-ai/DeepSeek-V3.2","github":"https://github.com/deepseek-ai/DeepSeek-V3.2","paper":"https://github.com/deepseek-ai/DeepSeek-V3.2/blob/main/DeepSeek_V3_2.pdf","pricing":"https://api-docs.deepseek.com/quick_start/pricing"},"copy":{"one_liner":"MIT 许可的开源旗舰，稀疏注意力把长上下文成本压到闭源 1/10。","highlights":["MIT 许可，685B MoE 权重可下，商用无附加条款","DSA 稀疏注意力：128K 上下文推理成本大幅下降，API 价 $0.28 / $0.42","思考 / 非思考双模式统一权重，工具调用与推理可交错"],"pitfalls":["685B 体量，自建门槛 8×80GB 起，不是「开源就能本地跑」","DSA 为新算子，vLLM / SGLang 需较新版本，llama.cpp 支持滞后","多模态缺失，纯文本模型"],"logic_ability":"考试型与工程型均衡：AIME 2025 与 GPQA 接近闭源旗舰，SWE-bench Verified 73.1% 属开源第一梯队。思考模式下推理链较长但很少「过度思考」；非思考模式在复杂规划上明显下降。V3.2-Speciale 变体在竞赛题上更强但不适合日常。常见问题：长 agent 任务中工具调用格式偶尔漂移。","best_for":["自建高吞吐推理服务","低成本 API 替代闭源","长上下文文档处理"],"not_for":["单卡 / 消费级部署","需要视觉输入"]},"capability_notes":{"coding":"SWE-bench Verified 73.1%、Codeforces 2386（官方）。","reasoning":"GPQA Diamond 82.4%、HLE 25.1%（官方）。","math":"AIME 2025 93.1%（官方）。","agent":"Terminal-bench 37.7%，τ²-bench 80.3%（官方）。","chinese":"中文能力顶级，训练语料中文占比高。"},"sheet":{"architecture_md":"**类型**：MoE，685B 总参数 / 37B 激活。\n\n- 61 层，隐藏维 7168，词表 129,280\n- 专家：256 路由专家 + 1 共享专家，每 token 激活 8 个路由专家\n- 注意力：**MLA**（Multi-head Latent Attention）：KV 压缩到 512 维潜向量 + 64 维解耦 RoPE\n- **DSA**（DeepSeek Sparse Attention）：轻量 indexer 打分，每个 query 只对 top-2048 个键做注意力\n- 上下文 128K；RoPE 用 YaRN 扩展\n- MTP（多 token 预测）头可用于投机解码\n\n参考：DeepSeek-V3 技术报告、V3.2 技术报告（GitHub）。","memory_md":"| 精度 | 权重大小 | 说明 |\n|---|---|---|\n| FP8 | ≈ 690 GB | 官方原生发布精度 |\n| BF16 | ≈ 1.37 TB | 由 FP8 反量化，通常不这么跑 |\n| Q4 | ≈ 380 GB（估） | llama.cpp / AWQ 社区量化 |\n\n**KV Cache**：MLA 每 token 每层约 576 维 → 61 层 ≈ 70 KiB/token（BF16）。128K 上下文单请求 ≈ 8.8 GB，比同规模 GQA 小一个量级。\n\n**参考配置**：\n- 消费级 24GB：不可行\n- 单卡 80GB：不可行\n- 8×H100 80GB：FP8 可部署，剩余显存留给 KV，高并发建议 H200\n\n警示：能加载 ≠ 能高并发。DSA 的 indexer 也占算力。","training_md":"- 预训练语料：V3 为 14.8T token；V3.2 在 V3.1-Terminus 基础上继续训练并做 DSA 适配（约 1T token 稀疏注意力继续预训练）\n- 后训练：大规模 RL（GRPO），把推理、agent、对齐合并为单阶段\n- 思考模式：`deepseek-reasoner` 默认思考；`deepseek-chat` 非思考\n- 工具调用：思考模式下支持工具调用（V3.2 新增）\n- 最大输出 64K（reasoner）","ecosystem_md":"- HF：deepseek-ai/DeepSeek-V3.2\n- 推理引擎：SGLang（官方首推）、vLLM（较新版本）、TensorRT-LLM；llama.cpp 对 DSA 支持有限\n- 微调：全参微调需要多节点；LoRA 社区支持\n- 中文文档：有（官方文档中英双语）","versions_md":"- 上代：DeepSeek-V3.1 / V3.1-Terminus（2025-08/09）、V3.2-Exp（2025-09）\n- 变体：DeepSeek-V3.2-Speciale（竞赛向，思考更长）\n- 姊妹：DeepSeek-R1-0528（旧推理线，已被 V3.2 思考模式覆盖）"},"ecosystem":{"engines":["SGLang","vLLM","TensorRT-LLM"],"finetune":"多节点全参 / LoRA","zh_docs":"有"},"complete":true,"superseded_by":"deepseek-v4-pro"},{"aliases":["DeepSeek-V3 671B","deepseek-chat (2024-12)"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"coding":"LiveCodeBench 40.5%（Pass@1-COT）、Codeforces 51.6 百分位（官方）。","reasoning":"GPQA Diamond 59.1%（官方）。","math":"MATH-500 90.2%、AIME 2024 39.2%（官方）。","knowledge":"MMLU-Pro 75.9%（官方）。","agent":"工具调用可用但不稳（0324 改善）。","chinese":"C-Eval 86.5%、中文 SimpleQA 64.8%（官方），中文顶级。"},"complete":true,"id":"deepseek-v3","name":"DeepSeek-V3","name_zh":"DeepSeek-V3（1226）","vendor":"DeepSeek","vendor_zh":"深度求索","family":"DeepSeek-V3","license":"MIT","license_commercial":true,"weights_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3","superseded_by":"deepseek-v3-0324","released_at":"2024-12-26","architecture":{"type":"moe","total_params":"671B","active_params":"37B","total_params_b":671,"active_params_b":37,"experts":256,"active_experts":8,"shared_expert":true,"layers":61,"hidden_size":7168,"vocab_size":129280,"kv_heads":128,"head_dim":192,"attention":"MLA（kv_lora_rank 512 + rope 64）","notes":"14.8T token 预训练；FP8 混合精度；MTP 多 token 预测；256 路由专家 top-8 + 1 共享专家；128K。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":8192},"memory":{"weight_gb":{"bf16":1342,"q8":711.3,"q4":404.4,"fp8":671},"estimated":false,"kv_per_token_kib":68.6,"kv_note":"MLA 压缩 KV：(512+64)×61×2 B ≈ 68.6 KiB/token；128K 上下文仅 8.6 GB。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（单卡）","ref_hw_8x80gb":"官方 FP8 671 GB 权重 8×H100 可服务（需 FP8 引擎）；BF16 需 16×80GB；Q4 404 GB 8×80GB 可跑"},"links":{"official":"https://www.deepseek.com/","github":"https://github.com/deepseek-ai/DeepSeek-V3","paper":"https://arxiv.org/abs/2412.19437","hf":"https://huggingface.co/deepseek-ai/DeepSeek-V3"},"copy":{"one_liner":"560 万美元训出的开源 GPT-4o 级，MoE 时代的分水岭。","highlights":["官方称训练成本仅 2.788M H800 GPU 小时，性能对标 GPT-4o / Claude 3.5","MLA + 无辅助损失负载均衡 + MTP + FP8 训练","MIT 许可（权重），API 价格极低"],"pitfalls":["671B 总参，本地部署需 8×80GB 以上","无推理模式，Agent 与工具调用弱于后续 0324","早期版本在英文长对话中偶有中文夹杂"],"logic_ability":"非推理模型顶级：MATH-500 90.2%、GPQA 59.1%、LiveCodeBench 40.5%、Codeforces 51.6 百分位（官方）。逻辑与代码表现在 2024 年末开源第一，但没有长思考能力，复杂题目依赖单次前向。","best_for":["通用与代码","API 低成本调用"],"not_for":["单机本地部署","多模态"]},"pricing":{"input_per_m":0.27,"output_per_m":1.1,"currency":"USD","source":"DeepSeek 官方 API（deepseek-chat，2024-12 定价）","as_of":"2026-08-28"},"sheet":{"architecture_md":"**类型**：MoE Transformer，671B 总参 / 37B 激活。\n\n- 61 层（前 3 层 Dense FFN），隐藏维 7168，词表 129,280\n- DeepSeekMoE：256 路由专家（FFN 2048）top-8 + 1 共享专家，无辅助损失负载均衡\n- MLA：128 头，kv_lora_rank 512，qk_nope 128 + qk_rope 64，v_head 128\n- MTP（多 token 预测）模块 1 层，推理可用作投机解码\n- 上下文 128K（YaRN 两阶段扩展）\n\n参考：DeepSeek-V3 技术报告（arXiv 2412.19437）。","memory_md":"| 精度 | 权重大小 | 参考硬件 | 说明 |\n|---|---|---|---|\n| FP8 | ≈ 671 GB | 8×H100/H200 | 官方原生 FP8 权重 |\n| BF16 | 1342 GB | 16×80GB | 需转换 |\n| Q4 | 404 GB | 8×80GB / 大内存 CPU | GGUF Q4_K_M（unsloth） |\n| 1.58-bit | ≈ 131 GB | 2×80GB | unsloth 动态量化，质量有损 |\n\n**KV Cache**：MLA ≈ 68.6 KiB/token，128K 仅 8.6 GB，是同规模 GQA 模型的 1/10。\n\n**参考配置**：\n- 8×H200 141GB：FP8 + 长上下文高并发（官方推荐）\n- 8×H100 80GB：FP8 勉强（640 GB 显存需专家卸载或 Q8）\n- Mac Studio 512GB：Q4 可跑，约 5–10 tok/s","training_md":"- 预训练 14.8T token，FP8 混合精度，2.664M H800 GPU 小时\n- DualPipe 流水线 + 跨节点 all-to-all 通信优化\n- 后训练：SFT（含 R1 蒸馏的推理数据）+ GRPO\n- 总成本官方估算 2.788M GPU 小时 ≈ 557.6 万美元（不含研发）","ecosystem_md":"- HF：deepseek-ai/DeepSeek-V3（FP8）、-Base\n- 引擎：SGLang（官方推荐）、vLLM、LMDeploy、TensorRT-LLM、llama.cpp（GGUF）\n- 微调：极困难（671B MoE + MLA），社区几乎只做 LoRA 实验\n- 中文文档：有","versions_md":"- 后继：DeepSeek-V3-0324（2025-03）→ V3.1（2025-08）→ V3.1-Terminus → V3.2\n- 推理分支：DeepSeek-R1（2025-01，基于 V3-Base）\n- 上代：DeepSeek-V2.5"},"ecosystem":{"engines":["SGLang","vLLM","LMDeploy","TensorRT-LLM","llama.cpp"],"finetune":"极困难","zh_docs":"有"}},{"id":"deepseek-v4-flash","name":"DeepSeek-V4-Flash","name_zh":"深度求索 V4 Flash","aliases":["DeepSeek V4 Flash","DeepSeek-V4-Flash-0731","deepseek-v4-flash"],"vendor":"DeepSeek","vendor_zh":"深度求索","family":"DeepSeek V4","license":"MIT","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731","status":"current","released_at":"2026-07-31","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"284B","active_params":"13B","total_params_b":284,"active_params_b":13,"experts":256,"active_experts":6,"shared_expert":true,"layers":43,"hidden_size":4096,"vocab_size":129280,"kv_heads":1,"head_dim":512,"attention":"CSA（Compressed Sparse Attention）+ HCA（Heavily Compressed Attention）混合，MLA 式潜向量","notes":"config：64 Q 头、q_lora_rank 1024、o_lora_rank 1024、RoPE 64 维；每层 KV 用 4× 或 128× 压缩交错，indexer 64 头、top-k 512，滑窗 128，前 3 层 hash 层。表内 kv_heads=1、head_dim=512 为 MLA 式潜向量等效换算。256 路由 + 1 共享专家，每 token 激活 6，专家中间维 2,048；专家 FP4、其余 FP8。mHC、Muon、1 层 MTP。0731 版结构与 V4-Flash-DSpark 相同（附投机解码模块）。","undisclosed":false},"context":{"max_tokens":1048576,"display":"1M","max_output":393216},"memory":{"weight_gb":{"fp8":167,"q8":162,"q4":155},"kv_note":"混合压缩注意力，官方未披露每 token KiB；Think Max 建议 ≥ 384K 输出窗口。fp8 一栏为官方 FP4（专家）+ FP8 混合原生仓库 167 GB；Q4 为 unsloth UD-Q4_K_XL 155 GB，Q8 为 UD-Q8_K_XL 162 GB（专家原生 4-bit，量化几乎不缩）；IQ3_XXS 103 GB 可塞进 128 GB 内存机器。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（单卡）；Mac Studio / 128 GB 统一内存可用 IQ3_XXS 103 GB","ref_hw_8x80gb":"FP4/FP8 原生 167 GB，2×H100 80GB 即可加载，8×80GB 可做高并发 + 长上下文；官方 vLLM 示例为单节点 4×GB300","estimated":false},"pricing":{"input_per_m":0.22,"output_per_m":0.66,"currency":"USD","source":"DeepSeek 官方 API 定价（非高峰）","as_of":"2026-08-28","note":"高峰时段翻倍为 $0.44 / $1.32；缓存命中输入 $0.007 / $0.014。8 月 17 日起分峰谷计价"},"links":{"official":"https://www.deepseek.com/","hf":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731","github":"https://github.com/deepseek-ai/DeepSeek-V4","paper":"https://arxiv.org/abs/2606.19348","pricing":"https://api-docs.deepseek.com/quick_start/pricing"},"variants":[{"kind":"other","publisher":"deepseek-ai","repo":"deepseek-ai/DeepSeek-V4-Flash-0731","url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731","note":"官方原生 FP4（专家）+ FP8 混合权重","sizes":{"fp8":166.9}},{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/DeepSeek-V4-Flash-NVFP4","url":"https://huggingface.co/nvidia/DeepSeek-V4-Flash-NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/DeepSeek-V4-Flash-0731-GGUF","url":"https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF","note":"仓库内 Q8_0/BF16 文件为 DSpark 草稿模型，主权重为 UD 系列量化"},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/DeepSeek-V4-Flash-0731-GGUF","url":"https://huggingface.co/bartowski/DeepSeek-V4-Flash-0731-GGUF"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/DeepSeek-V4-Flash-4bit","url":"https://huggingface.co/mlx-community/DeepSeek-V4-Flash-4bit"}],"copy":{"one_liner":"13B 激活跑出旗舰 Agent 分，MIT，167 GB 可自建。","highlights":["MIT，284B / 13B 激活，原生 167 GB，2×H100 或 128 GB 内存机器即可跑起 1M 上下文模型","0731 正式版 Terminal Bench 2.1 82.7、DeepSWE 54.4，AA 智能指数 52 仅比 V4-Pro 低 1 分","API $0.22 / $0.66（非高峰），缓存命中 $0.007；AA 实测约 119 tok/s"],"pitfalls":["专家原生 FP4，Q4 GGUF（155 GB）几乎不比原版小，「量化省显存」不成立，最低 IQ3 也要 103 GB","纯文本；HLE 37.8、SimpleQA-Verified 34.1，知识面明显弱于 Pro","高峰 / 非高峰双价，Think Max 输出极长（AA 评为高冗长），实际成本按输出算"],"logic_ability":"13B 激活却有旗舰级推理：预览版模型卡 GPQA Diamond 88.1、LiveCodeBench 91.6、HMMT 2026 Feb 94.8、SWE-bench Verified 79.0（Max 档），与 V4-Pro 差距多在 1–3 分。0731 正式版把 Agent 补强到 Terminal Bench 2.1 82.7（预览 61.8）、DeepSWE 54.4（预览 7.3）、Toolathlon-Verified 70.3。短板在知识与长尾事实（SimpleQA-Verified 34.1 vs Pro 57.9）与 HLE。非思考模式推理骤降（GPQA 71.2、HLE 8.1），须开 high / max。常见问题：max 档输出冗长，工具调用链长时偶有格式漂移。","best_for":["性价比自建 Coding / Terminal Agent","高吞吐低价 API 批处理","1M 长上下文检索问答（LongBench-V2 44.7）"],"not_for":["消费级单卡 24GB","知识密集型问答（用 Pro）","视觉输入"]},"capability_notes":{"coding":"SWE-bench Verified 79.0、LiveCodeBench 91.6、Codeforces 3052（预览版模型卡，Max）；Terminal Bench 2.1 82.7、NL2Repo 54.2（0731 模型卡）。","reasoning":"GPQA Diamond 88.1（预览版）；HLE 37.8 / 51.5 有工具（0731 模型卡，引自 Pro-0813 对比表）。","math":"HMMT 2026 Feb 94.8、IMOAnswerBench 88.4（预览版模型卡）；AIME 2025 未披露。","agent":"DeepSWE 54.4、Toolathlon-Verified 70.3、Cybergym 76.7、DSBench-FullStack 68.7（0731 模型卡）；BrowseComp 73.2（预览版）。","chinese":"Chinese-SimpleQA 78.9（预览版模型卡）。"},"sheet":{"architecture_md":"**类型**：MoE，284B 总参数 / 13B 激活。\n\n| 项目 | 数值 |\n|---|---|\n| 层数 | 43（前 3 层 hash 层） |\n| 隐藏维 | 4,096 |\n| 专家 | 256 路由 + 1 共享，每 token 激活 6，专家中间维 2,048 |\n| 路由 | sqrtsoftplus、noaux_tc、scaling 1.5 |\n| 注意力头 | 64 Q 头，head_dim 512，KV 单头（MLA 式潜向量） |\n| 低秩 | q_lora_rank 1024、o_lora_rank 1024（o_groups 8） |\n| RoPE | 64 维，YaRN factor 16（原生 64K → 1M） |\n| 稀疏 indexer | 64 头 × 128 维，top-k 512，滑窗 128 |\n| KV 压缩比 | 逐层 4× / 128× 交错（CSA / HCA） |\n| 词表 | 129,280 |\n| 精度 | 专家 FP4，其余 FP8 |\n| MTP | 1 层；0731 版附 DSpark 投机解码 |\n\n与 V4-Pro 同构、缩小一档：层数 61 → 43，隐藏维 7168 → 4096，专家 384 → 256，indexer top-k 1024 → 512。\n\n参考：DeepSeek-V4 技术报告（arXiv 2606.19348）、config.json。","memory_md":"| 精度 | 权重大小 | 参考显存 | 说明 |\n|---|---|---|---|\n| FP4 + FP8 混合 | 167 GB | 2×80GB / 4×48GB | 官方原生，48 分片 |\n| Q8 | 162 GB | 2×80GB | unsloth UD-Q8_K_XL（近无损） |\n| Q4 | 155 GB | 2×80GB | unsloth UD-Q4_K_XL（专家本就 4-bit） |\n| IQ3_XXS | 103 GB | 128 GB 统一内存 | unsloth，3-bit 有损 |\n| IQ1_S | 82.5 GB | 96 GB | 质量损失大 |\n\n**KV Cache**：混合压缩注意力，每 token 精确值官方未披露；DSpark 投机解码额外约 10 GB。Think Max 建议 ≥ 384K 输出窗口。\n\n**参考配置**：\n- 24GB 单卡：不可行\n- 80GB 单卡：不可行；128 GB Mac Studio / DGX Spark 类可用 IQ3_XXS\n- 2×H100 80GB：原生 FP4/FP8 可加载，中等上下文\n- 8×80GB / 4×GB300：高并发 1M 上下文服务（官方 vLLM 示例）","training_md":"- 预训练 32T+ token，Muon 优化器\n- 后训练两阶段：领域专家 SFT + GRPO → on-policy 蒸馏合并\n- 0731 正式版：「大幅增强 Agent 能力」，结构同 V4-Flash-DSpark\n- 思考模式：`reasoning_effort` low / high / max，可非思考\n- 最大输出：384K（API）\n- 推荐采样：temperature 1.0；top_p 0.95（Agent）/ 1.0（其他）","ecosystem_md":"- HF：deepseek-ai/DeepSeek-V4-Flash-0731（正式）、DeepSeek-V4-Flash（4 月预览）、DeepSeek-V4-Flash-DSpark、DeepSeek-V4-Flash-Base\n- GGUF：unsloth/DeepSeek-V4-Flash-0731-GGUF（含 DSpark 草稿模块 Q8 10.9 GB）\n- 引擎：vLLM、SGLang（DSpark）、llama.cpp、Ollama、LM Studio、Docker Model Runner、Transformers；可在华为昇腾运行\n- API：`deepseek-v4-flash`；另有 `deepseek-v4-flash-vision-exp` 视觉实验版\n- 中文文档：有","versions_md":"- DeepSeek-V4-Flash（预览，2026-04-24）→ DeepSeek-V4-Flash-0731（正式，2026-07-31，本条目）\n- 姊妹：DeepSeek-V4-Pro-0813（1.6T / 49B）\n- 上代：DeepSeek-V3.2（685B / 37B）\n- 基座：DeepSeek-V4-Flash-Base"},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Ollama","Transformers"],"finetune":"多卡 LoRA；Unsloth 跟进中","zh_docs":"有"},"i18n":{"en":{"name_zh":"DeepSeek V4 Flash","one_liner":"Flagship-level agent scores from 13B active; MIT, 167 GB self-hostable.","highlights":["MIT, 284B / 13B active, 167 GB native; a 2x H100 or 128 GB-RAM machine runs the 1M-context model","0731 release: Terminal Bench 2.1 82.7, DeepSWE 54.4, AA Intelligence Index 52, only 1 point below V4-Pro","API $0.22 / $0.66 (off-peak), cache hit $0.007; AA-measured about 119 tok/s"],"pitfalls":["Experts are native FP4, so Q4 GGUF (155 GB) is barely smaller than the original; \"quantize to save VRAM\" does not apply, even IQ3 needs 103 GB","Text only; HLE 37.8, SimpleQA-Verified 34.1, knowledge clearly weaker than Pro","Peak / off-peak dual pricing; Think Max output is extremely long (AA rates it highly verbose), real cost is driven by output"],"logic_ability":"Flagship-class reasoning from only 13B active: preview model card GPQA Diamond 88.1, LiveCodeBench 91.6, HMMT 2026 Feb 94.8, SWE-bench Verified 79.0 (Max), mostly within 1-3 points of V4-Pro. The 0731 release strengthened agents to Terminal Bench 2.1 82.7 (preview 61.8), DeepSWE 54.4 (preview 7.3), Toolathlon-Verified 70.3. Weak spots are knowledge and long-tail facts (SimpleQA-Verified 34.1 vs Pro 57.9) and HLE. Reasoning collapses in non-thinking mode (GPQA 71.2, HLE 8.1), so use high / max. Common issues: verbose output at max; occasional format drift in long tool-call chains.","best_for":["Cost-effective self-hosted coding / terminal agents","High-throughput low-cost API batch processing","1M long-context retrieval QA (LongBench-V2 44.7)"],"not_for":["Consumer single 24GB GPU","Knowledge-intensive QA (use Pro)","Vision input"],"capability_notes":{"coding":"SWE-bench Verified 79.0, LiveCodeBench 91.6, Codeforces 3052 (preview model card, Max); Terminal Bench 2.1 82.7, NL2Repo 54.2 (0731 model card).","reasoning":"GPQA Diamond 88.1 (preview); HLE 37.8 / 51.5 with tools (0731 model card, from the Pro-0813 comparison table).","math":"HMMT 2026 Feb 94.8, IMOAnswerBench 88.4 (preview model card); AIME 2025 not disclosed.","agent":"DeepSWE 54.4, Toolathlon-Verified 70.3, Cybergym 76.7, DSBench-FullStack 68.7 (0731 model card); BrowseComp 73.2 (preview).","chinese":"Chinese-SimpleQA 78.9 (preview model card)."}},"ja":{"name_zh":"DeepSeek V4 Flash","one_liner":"13B 活性で旗艦級 agent スコア。MIT、167 GB で自前運用可。","highlights":["MIT、284B / 13B アクティブ、ネイティブ 167 GB。2×H100 か 128 GB メモリのマシンで 1M コンテキストモデルが動く","0731 正式版は Terminal Bench 2.1 82.7、DeepSWE 54.4、AA 知能指数 52 で V4-Pro との差はわずか 1 点","API $0.22 / $0.66（オフピーク）、キャッシュヒット $0.007。AA 実測約 119 tok/s"],"pitfalls":["エキスパートがネイティブ FP4 のため Q4 GGUF（155 GB）は元とほぼ同サイズ。「量子化で VRAM 節約」は成立せず、最小の IQ3 でも 103 GB","テキストのみ。HLE 37.8、SimpleQA-Verified 34.1 と知識面は Pro より明らかに弱い","ピーク / オフピークの二重価格。Think Max の出力は極めて長く（AA は高冗長と評価）、実コストは出力量で決まる"],"logic_ability":"13B アクティブながらフラッグシップ級の推論：プレビュー版モデルカードで GPQA Diamond 88.1、LiveCodeBench 91.6、HMMT 2026 Feb 94.8、SWE-bench Verified 79.0（Max）、V4-Pro との差は多くが 1〜3 点。0731 正式版で agent を強化し Terminal Bench 2.1 82.7（プレビュー 61.8）、DeepSWE 54.4（プレビュー 7.3）、Toolathlon-Verified 70.3。弱点は知識とロングテールの事実（SimpleQA-Verified 34.1 vs Pro 57.9）と HLE。非思考モードでは推論が急落（GPQA 71.2、HLE 8.1）するため high / max が必須。よくある問題：max では出力が冗長、長いツール呼び出しチェーンで時折フォーマットがずれる。","best_for":["コスパ重視の自前 Coding / Terminal Agent","高スループット低価格 API のバッチ処理","1M 長文コンテキスト検索 QA（LongBench-V2 44.7）"],"not_for":["コンシューマー向け 24GB 単一 GPU","知識集約型 QA（Pro を使う）","視覚入力"],"capability_notes":{"coding":"SWE-bench Verified 79.0、LiveCodeBench 91.6、Codeforces 3052（プレビュー版モデルカード、Max）。Terminal Bench 2.1 82.7、NL2Repo 54.2（0731 モデルカード）。","reasoning":"GPQA Diamond 88.1（プレビュー版）。HLE 37.8 / ツールあり 51.5（0731 モデルカード、Pro-0813 比較表より）。","math":"HMMT 2026 Feb 94.8、IMOAnswerBench 88.4（プレビュー版モデルカード）。AIME 2025 は未公表。","agent":"DeepSWE 54.4、Toolathlon-Verified 70.3、Cybergym 76.7、DSBench-FullStack 68.7（0731 モデルカード）。BrowseComp 73.2（プレビュー版）。","chinese":"Chinese-SimpleQA 78.9（プレビュー版モデルカード）。"}}},"complete":true,"runtime":{"tok_s":119,"latency_s":1.12,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"deepseek-v4-pro","name":"DeepSeek-V4-Pro","name_zh":"深度求索 V4 Pro","aliases":["DeepSeek V4 Pro","DeepSeek-V4-Pro-0813","deepseek-v4-pro"],"vendor":"DeepSeek","vendor_zh":"深度求索","family":"DeepSeek V4","license":"MIT","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813","status":"current","released_at":"2026-08-13","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"1.6T","active_params":"49B","total_params_b":1600,"active_params_b":49,"experts":384,"active_experts":6,"shared_expert":true,"layers":61,"hidden_size":7168,"vocab_size":129280,"kv_heads":1,"head_dim":512,"attention":"CSA（Compressed Sparse Attention）+ HCA（Heavily Compressed Attention）混合，MLA 式潜向量","notes":"config：128 Q 头、q_lora_rank 1536、o_lora_rank 1024、RoPE 64 维；每层 KV 用 4× 或 128× token 压缩（compress_ratios 4/128 交错），稀疏 indexer 64 头、top-k 1024，滑窗 128，前 3 层为 hash 层。表内 kv_heads=1、head_dim=512 为 MLA 式潜向量的等效换算。384 路由 + 1 共享专家，每 token 激活 6；专家 FP4、其余 FP8 原生。mHC 超连接残差、Muon 优化器、1 层 MTP。官方称 1M 上下文下单 token FLOPs 仅为 V3.2 的 27%、KV 为 10%。","undisclosed":false},"context":{"max_tokens":1048576,"display":"1M","max_output":393216},"memory":{"weight_gb":{"fp8":893,"q8":873.4,"q4":849.7},"kv_note":"混合压缩注意力：HCA 层按 128× 压缩、CSA 层按 4× 压缩再稀疏选 top-1024，官方称 1M 上下文 KV 仅为 V3.2 的 10%；每 token 精确 KiB 官方未披露。fp8 一栏为官方 FP4（专家）+ FP8 混合原生权重 893 GB；Q4 一栏为 unsloth UD-Q4_K_XL，专家本就是 4-bit 所以几乎不缩。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（单卡）","ref_hw_8x80gb":"不可行（893 GB > 640 GB）；需 8×H200 141GB（1.13 TB）或 8×B200 / GB300 级别单机","estimated":false},"pricing":{"input_per_m":0.66,"output_per_m":1.98,"currency":"USD","source":"DeepSeek 官方 API 定价（非高峰）","as_of":"2026-08-28","note":"高峰时段（UTC 01–04、06–10 周一至五）翻倍为 $1.32 / $3.96；缓存命中输入 $0.022 / $0.044。8 月 17 日起分峰谷计价"},"links":{"official":"https://www.deepseek.com/","hf":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813","github":"https://github.com/deepseek-ai/DeepSeek-V4","paper":"https://arxiv.org/abs/2606.19348","pricing":"https://api-docs.deepseek.com/quick_start/pricing"},"variants":[{"kind":"other","publisher":"deepseek-ai","repo":"deepseek-ai/DeepSeek-V4-Pro-0813","url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813","note":"官方原生 FP4（专家）+ FP8 混合权重","sizes":{"fp8":892.7}},{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/DeepSeek-V4-Pro-NVFP4","url":"https://huggingface.co/nvidia/DeepSeek-V4-Pro-NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/DeepSeek-V4-Pro-0813-GGUF","url":"https://huggingface.co/unsloth/DeepSeek-V4-Pro-0813-GGUF"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/DeepSeek-V4-Pro-4bit","url":"https://huggingface.co/mlx-community/DeepSeek-V4-Pro-4bit"}],"copy":{"one_liner":"MIT 许可 1.6T 旗舰，1M 上下文，Agent 榜追平闭源。","highlights":["MIT 许可 1.6T / 49B 激活权重可下，0813 正式版 Terminal Bench 2.1 87.9、DeepSWE 62.7 与 Opus-4.8 同档","CSA + HCA 混合压缩注意力：1M 上下文下 FLOPs 为 V3.2 的 27%、KV 为 10%","API $0.66 / $1.98（非高峰），Text Arena 1462 分；DSpark 投机解码模块随权重附带"],"pitfalls":["893 GB 原生 FP4/FP8，单机 8×80GB 装不下，起步是 8×H200 或 B200 级别","0813 版 API 价格较 4 月预览版上涨（AA 称约 3.6×），且分高峰 / 非高峰计价","纯文本；HLE 42.7、GPQA 90.1 等考试型榜落后 Opus 4.8 / Fable 5 数分，SWE-bench Verified 80.6 为预览版数"],"logic_ability":"工程型强于考试型：预览版模型卡 SWE-bench Verified 80.6、LiveCodeBench 93.5、Codeforces 3206 已是开源第一梯队；0813 正式版重点补强 Agent，Terminal Bench 2.1 从 72.1 升到 87.9、DeepSWE 从 12.8 升到 62.7、Cybergym 83.3。考试型 GPQA Diamond 90.1、HLE 42.7（无工具）/ 60.0（有工具）距 Fable 5 仍有差距。三档 reasoning_effort（low / high / max），max 建议 ≥ 384K 输出窗口；非思考模式在 HLE 等题上骤降（预览版 7.7）。AA 智能指数 53，仅比 V4-Flash-0731 高 1 分，性价比不如 Flash。","best_for":["自建旗舰级 Coding / Terminal Agent","1M 超长上下文文档与代码库","低成本 API 替代闭源旗舰"],"not_for":["单机 8×80GB 及以下部署","视觉输入","对 Flash 已够用的场景（差距仅 1 分 AA）"]},"capability_notes":{"coding":"SWE-bench Verified 80.6、LiveCodeBench 93.5、Codeforces 3206（预览版模型卡，Max）；Terminal Bench 2.1 87.9、NL2Repo 61.5（0813 模型卡）。","reasoning":"GPQA Diamond 90.1（预览版）；HLE 42.7 / 60.0 有工具（0813 模型卡）。","math":"HMMT 2026 Feb 95.2、IMOAnswerBench 89.8（预览版模型卡）；AIME 2025 未披露。","agent":"DeepSWE 62.7、Toolathlon-Verified 74.1、Cybergym 83.3、DSBench-FullStack 71.1（0813 模型卡）；BrowseComp 83.4（预览版）。","chinese":"Chinese-SimpleQA 84.4（预览版模型卡），中文事实知识顶级。"},"sheet":{"architecture_md":"**类型**：MoE，1.6T 总参数 / 49B 激活。\n\n| 项目 | 数值 |\n|---|---|\n| 层数 | 61（前 3 层 hash 层） |\n| 隐藏维 | 7,168 |\n| 专家 | 384 路由 + 1 共享，每 token 激活 6，专家中间维 3,072 |\n| 路由 | sqrtsoftplus 打分、noaux_tc、scaling 2.5 |\n| 注意力头 | 128 Q 头，head_dim 512，KV 单头（MLA 式潜向量） |\n| 低秩 | q_lora_rank 1536、o_lora_rank 1024（o_groups 16） |\n| RoPE | 64 维，YaRN factor 16（原生 64K → 1M） |\n| 稀疏 indexer | 64 头 × 128 维，top-k 1024，滑窗 128 |\n| KV 压缩比 | 逐层 4× / 128× 交错（CSA / HCA） |\n| 词表 | 129,280 |\n| 精度 | 专家 FP4，其余 FP8（e4m3，块 128×128） |\n| MTP | 1 层；0813 版附 DSpark 投机解码模块 |\n\n**CSA + HCA**：HCA 层把 KV 按 128× 压缩做全局注意力，CSA 层按 4× 压缩后用 DSA 式 indexer 选 top-1024。官方：1M 上下文下单 token FLOPs 为 V3.2 的 27%、KV 为 10%。\n\n**mHC**：Manifold-Constrained Hyper-Connections（hc_mult 4，Sinkhorn 20 步）替代普通残差。\n\n参考：DeepSeek-V4 技术报告（arXiv 2606.19348）、config.json。","memory_md":"| 精度 | 权重大小 | 说明 |\n|---|---|---|\n| FP4 + FP8 混合 | 893 GB | 官方原生发布，66 分片 |\n| Q8 | 873 GB | unsloth UD-Q8_K_XL |\n| Q4 | 850 GB | unsloth UD-Q4_K_XL（专家本就 4-bit，几乎不缩） |\n| BF16 | 未提供 | 官方不发 BF16 |\n\n**KV Cache**：混合压缩注意力，官方称 1M 上下文 KV 为 V3.2 的 10%；每 token 精确值未披露。Think Max 建议 ≥ 384K 输出窗口。\n\n**参考配置**：\n- 24GB / 80GB 单卡：不可行\n- 单机 8×80GB：不可行（893 GB > 640 GB）\n- 8×H200 141GB / 8×B200 / GB300：可部署，vLLM 与 SGLang 均支持 DSpark 投机解码\n\n警示：Q4 GGUF 不能省显存，「量化就能本地跑」对 V4-Pro 不成立。","training_md":"- 预训练 32T+ token，Muon 优化器\n- 后训练两阶段：领域专家培养（SFT + GRPO RL）→ on-policy 蒸馏统一合并\n- 0813 正式版：在预览版基础上「大幅增强 Agent 能力，生产环境提升尤其明显」\n- 思考模式：`reasoning_effort` low / high / max，可非思考\n- 最大输出：384K（API）\n- 推荐采样：temperature 1.0；top_p 0.95（Agent）/ 1.0（其他）\n- 无 Jinja 模板，官方提供 Python 编码脚本","ecosystem_md":"- HF：deepseek-ai/DeepSeek-V4-Pro-0813（正式）、DeepSeek-V4-Pro（4 月预览）、DeepSeek-V4-Pro-DSpark、DeepSeek-V4-Pro-Base\n- GGUF：unsloth/DeepSeek-V4-Pro-0813-GGUF\n- 引擎：vLLM（`--speculative-config` 开 DSpark）、SGLang（`--speculative-algorithm DSPARK`）、Transformers；可在华为昇腾运行\n- API：`deepseek-v4-pro`，OpenAI / Anthropic / Responses 三种接口\n- 中文文档：有","versions_md":"- DeepSeek-V4-Pro（预览，2026-04-24）→ DeepSeek-V4-Pro-0813（正式，2026-08-13，本条目）\n- 姊妹：DeepSeek-V4-Flash-0731（284B / 13B）、deepseek-v4-flash-vision-exp（API 视觉实验版）\n- 上代：DeepSeek-V3.2（2025-12）、V3.1-Terminus\n- 基座：DeepSeek-V4-Pro-Base"},"ecosystem":{"engines":["vLLM","SGLang","Transformers","llama.cpp"],"finetune":"多节点；社区 LoRA 极少","zh_docs":"有"},"i18n":{"en":{"name_zh":"DeepSeek V4 Pro","one_liner":"MIT-licensed 1.6T flagship, 1M context, matches closed models on agent boards.","highlights":["MIT-licensed 1.6T / 49B-active weights downloadable; 0813 release Terminal Bench 2.1 87.9, DeepSWE 62.7, same tier as Opus 4.8","CSA + HCA hybrid compressed attention: at 1M context FLOPs are 27% of V3.2 and KV 10%","API $0.66 / $1.98 (off-peak), Text Arena 1462; DSpark speculative decoding module ships with the weights"],"pitfalls":["893 GB native FP4/FP8; does not fit a single 8x80GB node, entry point is 8x H200 or B200 class","0813 API price rose versus the April preview (AA says about 3.6x), with peak / off-peak pricing","Text only; exam-style boards such as HLE 42.7 and GPQA 90.1 trail Opus 4.8 / Fable 5 by a few points; SWE-bench Verified 80.6 is a preview figure"],"logic_ability":"Stronger at engineering than exams: preview model card SWE-bench Verified 80.6, LiveCodeBench 93.5, Codeforces 3206 already put it in the open-source first tier; the 0813 release focused on agents, lifting Terminal Bench 2.1 from 72.1 to 87.9, DeepSWE from 12.8 to 62.7, Cybergym 83.3. Exam-style GPQA Diamond 90.1 and HLE 42.7 (no tools) / 60.0 (with tools) still trail Fable 5. Three reasoning_effort levels (low / high / max); max recommends a 384K+ output window; non-thinking mode collapses on HLE-type questions (preview 7.7). AA Intelligence Index 53, only 1 point above V4-Flash-0731, so less cost-effective than Flash.","best_for":["Self-hosted flagship-class coding / terminal agents","1M ultra-long-context documents and codebases","Low-cost API replacement for closed flagships"],"not_for":["Deployments on a single 8x80GB node or smaller","Vision input","Scenarios where Flash already suffices (only 1 AA point apart)"],"capability_notes":{"coding":"SWE-bench Verified 80.6, LiveCodeBench 93.5, Codeforces 3206 (preview model card, Max); Terminal Bench 2.1 87.9, NL2Repo 61.5 (0813 model card).","reasoning":"GPQA Diamond 90.1 (preview); HLE 42.7 / 60.0 with tools (0813 model card).","math":"HMMT 2026 Feb 95.2, IMOAnswerBench 89.8 (preview model card); AIME 2025 not disclosed.","agent":"DeepSWE 62.7, Toolathlon-Verified 74.1, Cybergym 83.3, DSBench-FullStack 71.1 (0813 model card); BrowseComp 83.4 (preview).","chinese":"Chinese-SimpleQA 84.4 (preview model card), top-tier Chinese factual knowledge."}},"ja":{"name_zh":"DeepSeek V4 Pro","one_liner":"MIT の 1.6T 旗艦、1M コンテキスト、agent 系で閉源に並ぶ。","highlights":["MIT ライセンスで 1.6T / 49B アクティブの重みをダウンロード可能。0813 正式版は Terminal Bench 2.1 87.9、DeepSWE 62.7 で Opus-4.8 と同格","CSA + HCA ハイブリッド圧縮アテンション：1M コンテキストで FLOPs は V3.2 の 27%、KV は 10%","API $0.66 / $1.98（オフピーク）、Text Arena 1462。DSpark 投機的デコードモジュールが重みに同梱"],"pitfalls":["ネイティブ FP4/FP8 で 893 GB。8×80GB の単一ノードには収まらず、最低でも 8×H200 か B200 クラス","0813 版の API 価格は 4 月プレビュー版より値上げ（AA によると約 3.6 倍）、ピーク / オフピーク別料金","テキストのみ。HLE 42.7、GPQA 90.1 などの試験型では Opus 4.8 / Fable 5 に数点劣る。SWE-bench Verified 80.6 はプレビュー版の数値"],"logic_ability":"試験型よりエンジニアリング型が強い：プレビュー版モデルカードで SWE-bench Verified 80.6、LiveCodeBench 93.5、Codeforces 3206 と既にオープンソース最上位。0813 正式版は agent を重点強化し、Terminal Bench 2.1 が 72.1→87.9、DeepSWE が 12.8→62.7、Cybergym 83.3。試験型の GPQA Diamond 90.1、HLE 42.7（ツールなし）/ 60.0（ツールあり）は Fable 5 との差が残る。reasoning_effort は 3 段階（low / high / max）で、max は 384K 以上の出力ウィンドウ推奨。非思考モードでは HLE などで急落（プレビュー版 7.7）。AA 知能指数 53 は V4-Flash-0731 より 1 点高いだけで、コスパは Flash に劣る。","best_for":["自前運用のフラッグシップ級 Coding / Terminal Agent","1M 超長コンテキストの文書・コードベース","閉源フラッグシップの低コスト API 代替"],"not_for":["8×80GB 以下の単一ノード展開","視覚入力","Flash で十分な用途（AA 指数差はわずか 1 点）"],"capability_notes":{"coding":"SWE-bench Verified 80.6、LiveCodeBench 93.5、Codeforces 3206（プレビュー版モデルカード、Max）。Terminal Bench 2.1 87.9、NL2Repo 61.5（0813 モデルカード）。","reasoning":"GPQA Diamond 90.1（プレビュー版）。HLE 42.7 / ツールあり 60.0（0813 モデルカード）。","math":"HMMT 2026 Feb 95.2、IMOAnswerBench 89.8（プレビュー版モデルカード）。AIME 2025 は未公表。","agent":"DeepSWE 62.7、Toolathlon-Verified 74.1、Cybergym 83.3、DSBench-FullStack 71.1（0813 モデルカード）。BrowseComp 83.4（プレビュー版）。","chinese":"Chinese-SimpleQA 84.4（プレビュー版モデルカード）、中国語の事実知識は最上位。"}}},"complete":true,"runtime":{"tok_s":66,"latency_s":1.59,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"devstral-small-2","name":"Devstral Small 2","aliases":["devstral-small-2512","Devstral 2 Small"],"vendor":"Mistral AI","family":"Devstral","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/mistralai/Devstral-Small-2-24B-Instruct-2512","status":"superseded","released_at":"2025-12-09","updated_at":"2025-12-20","modalities":["text","image","tools"],"reasoning_mode":"none","architecture":{"type":"dense","total_params":"24B","total_params_b":24,"layers":40,"hidden_size":5120,"kv_heads":8,"head_dim":128,"attention":"GQA（32 Q / 8 KV）","undisclosed":false},"context":{"max_tokens":262144,"display":"256K"},"memory":{"weight_gb":{"bf16":48,"q8":25.5,"q4":14.5},"kv_per_token_kib":160,"estimated":true,"ref_hw_24gb":"Q4 14.5 GB，24GB 舒适；Q8 需 32GB"},"pricing":{"input_per_m":0.1,"output_per_m":0.3,"currency":"USD","source":"Mistral 定价页","as_of":"2025-12-20"},"links":{"official":"https://mistral.ai/news/devstral-2-vibe-cli","hf":"https://huggingface.co/mistralai/Devstral-Small-2-24B-Instruct-2512"},"copy":{"one_liner":"24B Dense 编程 agent 模型，单卡 24GB 可跑，SWE-bench 68%。","highlights":["SWE-bench Verified 68.0%（官方），24B 体量惊人","Apache-2.0，单卡 24GB","256K 上下文"],"pitfalls":["非推理模型，通用能力一般","独立复测数据少","新模型，引擎适配观察中"],"logic_ability":"工程型推理专精；考试型推理弱。","best_for":["本地编程 agent"],"not_for":["通用助手"]},"capability_notes":{},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama"],"zh_docs":"无"},"complete":false,"superseded_by":"mistral-small-4"},{"aliases":["Devstral-Small-2505"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"coding":"SWE-bench Verified 46.8%（OpenHands，官方）。","agent":"OpenHands 专用微调，工具调用可靠。"},"complete":false,"id":"devstral-small","name":"Devstral Small","name_zh":"Devstral Small · 24B","vendor":"Mistral AI","family":"Devstral","license":"Apache-2.0","license_commercial":true,"weights_url":"https://huggingface.co/mistralai/Devstral-Small-2505","superseded_by":"devstral-small-2","released_at":"2025-05-21","architecture":{"type":"dense","total_params":"24B","total_params_b":24,"layers":40,"hidden_size":5120,"vocab_size":131072,"kv_heads":8,"head_dim":128,"attention":"GQA（32 Q 头 / 8 KV 头）","notes":"基于 Mistral Small 3.1 去掉视觉编码器，针对 OpenHands 代码 agent 微调。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":48,"q8":25.4,"q4":14.3},"estimated":false,"kv_per_token_kib":160,"kv_note":"160 KiB/token。","ref_hw_24gb":"Q4_K_M 14.3 GB，RTX 4090 可跑 agent 任务","ref_hw_80gb":"BF16 单卡","ref_hw_8x80gb":"过剩"},"pricing":{"input_per_m":0.1,"output_per_m":0.3,"currency":"USD","source":"Mistral 官方 API（devstral-small）","as_of":"2026-08-28","note":"官方托管价"},"links":{"official":"https://mistral.ai/news/devstral","hf":"https://huggingface.co/mistralai/Devstral-Small-2505"},"copy":{"one_liner":"首个单卡可跑的开源代码 agent 模型，SWE-bench 46.8%。","highlights":["发布时开源 SWE-bench Verified 第一（46.8%）","24B 可在 RTX 4090 / Mac 32GB 本地跑 agent","Apache-2.0，与 OpenHands 深度整合"],"pitfalls":["为 agent 场景微调，普通对话与知识不如通用模型","无视觉，无推理模式","2507 版（Devstral Small 1.1）已把 SWE-bench 提到 53.6%"],"logic_ability":"代码 agent 专精，SWE-bench Verified 46.8%（OpenHands 脚手架，官方）。通用推理没有单独报告，按 Small 3.1 水平估计。","best_for":["本地代码 agent（OpenHands / Cline）","私有仓库自动修 bug"],"not_for":["通用对话","数学 / 知识问答"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama","LM Studio"],"finetune":"LoRA 可行","zh_docs":"弱"}},{"aliases":["ERNIE-4.5-300B-A47B-PT"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"knowledge":"官方称 MMLU-Pro、SimpleQA 等领先 DeepSeek-V3。","chinese":"C-Eval、CMMLU 官方领先。"},"complete":false,"id":"ernie-4-5-300b-a47b","name":"ERNIE 4.5 300B-A47B","name_zh":"文心 4.5 · 300B-A47B","vendor":"Baidu","vendor_zh":"百度","family":"ERNIE 4.5","license":"Apache-2.0","license_commercial":true,"weights_url":"https://huggingface.co/baidu/ERNIE-4.5-300B-A47B-PT","released_at":"2025-06-30","architecture":{"type":"moe","total_params":"300B","active_params":"47B","total_params_b":300,"active_params_b":47,"experts":64,"active_experts":8,"shared_expert":true,"layers":54,"hidden_size":8192,"vocab_size":103424,"kv_heads":8,"head_dim":128,"attention":"GQA（64 Q 头 / 8 KV 头）","notes":"异构多模态 MoE 预训练后的纯文本后训练版；64 文本专家 top-8 + 共享；128K；PaddlePaddle 原生。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":600,"q8":318,"q4":180.2,"fp8":300},"estimated":false,"kv_per_token_kib":216,"kv_note":"216 KiB/token。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（单卡）；官方 W4A8 版 4×80GB","ref_hw_8x80gb":"BF16 600 GB 需 8×80GB（紧）；官方 2-bit 版单机可跑"},"links":{"official":"https://ernie.baidu.com/blog/posts/ernie4.5/","github":"https://github.com/PaddlePaddle/ERNIE","paper":"https://ernie.baidu.com/blog/publication/ERNIE_Technical_Report.pdf","hf":"https://huggingface.co/baidu/ERNIE-4.5-300B-A47B-PT"},"copy":{"one_liner":"百度首次开源旗舰，Apache-2.0 的 300B MoE。","highlights":["Apache-2.0，10 个模型全系开源","官方称多数基准超 DeepSeek-V3-671B","官方提供 2-bit / W4A8 量化，单机可跑"],"pitfalls":["PaddlePaddle / FastDeploy 为主，vLLM 支持滞后","非推理模型，2025 年中推理榜不占优","社区量化与微调生态薄弱"],"logic_ability":"非推理模型：官方基准对标 DeepSeek-V3-0324 / GPT-4.1；中文知识与指令遵循强。","best_for":["中文知识问答私有化","百度 / 飞桨生态"],"not_for":["推理密集任务","非飞桨部署"]},"ecosystem":{"engines":["FastDeploy","vLLM","Transformers"],"finetune":"ERNIEKit","zh_docs":"有"}},{"aliases":["falcon-180B-chat"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"knowledge":"MMLU 70.6%（官方）。","chinese":"弱。"},"complete":false,"id":"falcon-180b","name":"Falcon 180B","name_zh":"Falcon · 180B","vendor":"TII","vendor_zh":"阿联酋 TII","family":"Falcon","license":"Falcon-180B TII License","license_commercial":"restricted","weights_url":"https://huggingface.co/tiiuae/falcon-180B-chat","released_at":"2023-09-06","architecture":{"type":"dense","total_params":"180B","total_params_b":180,"layers":80,"hidden_size":14848,"vocab_size":65024,"kv_heads":8,"head_dim":64,"attention":"多组查询注意力（232 Q 头 / 8 KV 头，head_dim 64）","notes":"3.5T RefinedWeb token；并行注意力 + FFN；上下文 2K。","undisclosed":false},"context":{"max_tokens":2048,"display":"2K"},"memory":{"weight_gb":{"bf16":360,"q8":190.8,"q4":108.6},"estimated":false,"kv_per_token_kib":160,"kv_note":"8×64×2×80×2 B = 160 KiB/token。","ref_hw_24gb":"不可行","ref_hw_80gb":"Q4 108.6 GB 需 2×80GB；BF16 5×80GB","ref_hw_8x80gb":"BF16 高并发"},"links":{"official":"https://falconllm.tii.ae/","paper":"https://arxiv.org/abs/2311.16867","hf":"https://huggingface.co/tiiuae/falcon-180B-chat"},"copy":{"one_liner":"2023 年最大开放权重 Dense，RefinedWeb 数据的代表。","highlights":["发布时 HF 开源榜第一，规模 2.5 倍于 Llama 2 70B","RefinedWeb 证明纯网页数据可训出强模型","多组查询注意力早期实践"],"pitfalls":["上下文仅 2K","许可禁止托管服务（需单独授权）","能力被 Llama 2 70B 微调版轻松追平，性价比极差"],"logic_ability":"MMLU 70.6%（官方基座），推理能力 2023 年水平。","best_for":["历史对照"],"not_for":["任何生产场景"]},"ecosystem":{"engines":["vLLM","llama.cpp","TGI"],"finetune":"成本极高","zh_docs":"无"}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"gemini-1-0-pro","name":"Gemini 1.0 Pro","aliases":["gemini-pro","gemini-1.0-pro"],"vendor":"Google","family":"Gemini 1.0","superseded_by":"gemini-1-5-pro","released_at":"2023-12-13","modalities":["text","image","tools"],"reasoning_mode":"none","context":{"max_tokens":32768,"display":"32K","max_output":8192},"pricing":{"input_per_m":0.5,"output_per_m":1.5,"currency":"USD","source":"Google AI 定价页（付费层）","as_of":"2024-02-15","note":"发布初期免费；Ultra 版未开放 API"},"links":{"official":"https://blog.google/technology/ai/google-gemini-ai/","paper":"https://arxiv.org/abs/2312.11805","pricing":"https://ai.google.dev/pricing"},"copy":{"one_liner":"Gemini 系列首款开放 API 模型，Bard 底座。","highlights":["原生多模态训练，MMMU 47.9%","2023-12 免费开放 API，Bard 同步升级","同期 Gemini Ultra MMLU 90.0% 首超人类专家（未开放 API）"],"pitfalls":["MMLU 71.8%，落后 GPT-4 一代","32K 上下文","发布视频演示被指剪辑夸大"],"logic_ability":"约 GPT-3.5 水平的推理，GSM8K 86.5%，复杂多步任务不稳定。","best_for":["历史对照"],"not_for":["新项目"]}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true,"notes":"官方称由 1.5 Pro 蒸馏而来。"},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"gemini-1-5-flash","name":"Gemini 1.5 Flash","aliases":["gemini-1.5-flash-001","gemini-1.5-flash-002","gemini-1.5-flash-8b"],"vendor":"Google","family":"Gemini 1.5","superseded_by":"gemini-2-0-flash","released_at":"2024-05-14","modalities":["text","image","audio","video","tools"],"reasoning_mode":"none","context":{"max_tokens":1048576,"display":"1M","max_output":8192},"pricing":{"input_per_m":0.075,"output_per_m":0.3,"currency":"USD","source":"Google AI 定价页（2024-08 降价后，≤128K）","as_of":"2024-08-08","note":"发布价 $0.35 / $1.05；>128K 翻倍"},"links":{"official":"https://blog.google/technology/developers/gemini-gemini-1-5-pro-updates-google-io-2024/","paper":"https://arxiv.org/abs/2403.05530","pricing":"https://ai.google.dev/pricing"},"copy":{"one_liner":"1M 上下文的低价小模型，开启 Flash 产品线。","highlights":["$0.075 / $0.30，1M 上下文，多模态全支持","MMLU 78.9%，速度快，适合高并发","8B 版进一步减半价格"],"pitfalls":["GPQA 39.5%，推理弱","代理 / 代码任务不可靠","已被 2.0 Flash 全面替代"],"logic_ability":"轻量级推理，适合摘要、抽取、分类；多步逻辑与数学（MATH 54.9%）弱。","best_for":["历史对照"],"not_for":["新项目"]}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"moe","undisclosed":true,"notes":"官方技术报告称为稀疏 MoE Transformer，参数量未披露。"},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"gemini-1-5-pro","name":"Gemini 1.5 Pro","aliases":["gemini-1.5-pro-001","gemini-1.5-pro-002","gemini-1.5-pro-exp-0827"],"vendor":"Google","family":"Gemini 1.5","superseded_by":"gemini-2-5-pro","released_at":"2024-02-15","modalities":["text","image","audio","video","tools"],"reasoning_mode":"none","context":{"max_tokens":2097152,"display":"2M","max_output":8192},"pricing":{"input_per_m":1.25,"output_per_m":5,"currency":"USD","source":"Google AI 定价页（2024-10 降价后，≤128K）","as_of":"2024-10-01","note":">128K 为 $2.5 / $10；2024-05 GA 时 $3.5 / $10.5"},"links":{"official":"https://blog.google/technology/ai/google-gemini-next-generation-model-february-2024/","paper":"https://arxiv.org/abs/2403.05530","pricing":"https://ai.google.dev/pricing"},"copy":{"one_liner":"首个 1M→2M 上下文模型，长视频音频理解先驱。","highlights":["1M（后 2M）上下文，needle 召回 >99%","原生视频 / 音频输入，1 小时视频一次吃下","002 版 GPQA 59.1%、MATH 86.5%，价格降 64%"],"pitfalls":["最大输出 8K","SWE-bench 类代理编程能力未公布、口碑一般","多个 exp / 001 / 002 版本，行为差异大"],"logic_ability":"长上下文推理独步天下；常规逻辑推理与 GPT-4o 相当，002 版数学显著提升。","best_for":["历史对照"],"not_for":["新项目"]}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。2024-12-11 发布实验版，2025-02-05 GA；原生工具使用、原生图像 / 音频输出（实验）。"},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{"coding":"LiveCodeBench v5 34.5%（官方，GA 表）；实验版 SWE-bench Verified 51.8%（2024-12 官方表）。","reasoning":"GPQA Diamond 60.1%、MMLU-Pro 77.6%（官方 GA 表）。","math":"MATH 90.9%、HiddenMath 63.5%（官方）。","agent":"原生工具使用、Multimodal Live API；Project Astra / Mariner 底座。","multimodal":"MMMU 71.7%（官方）；音频 / 视频输入。","chinese":"Global MMLU Lite 83.4%（官方），中文良好。"},"complete":true,"id":"gemini-2-0-flash","name":"Gemini 2.0 Flash","aliases":["gemini-2.0-flash-001","gemini-2.0-flash-exp","gemini-2.0-flash-lite"],"vendor":"Google","family":"Gemini 2.0","superseded_by":"gemini-2-5-flash","released_at":"2025-02-05","modalities":["text","image","audio","video","tools"],"reasoning_mode":"none","context":{"max_tokens":1048576,"display":"1M","max_output":8192},"pricing":{"input_per_m":0.1,"output_per_m":0.4,"currency":"USD","source":"Google AI 定价页","as_of":"2025-02-05","note":"音频输入 $0.7；Flash-Lite $0.075 / $0.30"},"links":{"official":"https://blog.google/technology/google-deepmind/gemini-model-updates-february-2025/","paper":"https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/","pricing":"https://ai.google.dev/pricing"},"copy":{"one_liner":"代理时代的低价主力，$0.1 输入超越 1.5 Pro。","highlights":["以 Flash 价格超越 1.5 Pro，GPQA 60.1%、MMMU 71.7%","原生工具调用（搜索、代码执行）、Multimodal Live API","1M 上下文，$0.1 / $0.4 的价格带来大规模应用采用"],"pitfalls":["无推理模式，LiveCodeBench 34.5%","最大输出 8K","原生图像生成等功能长期停留实验阶段"],"logic_ability":"非推理小模型中的强者，MATH 90.9%；复杂多步推理需要 2.5 系列的 thinking。常见问题：长代理任务中容易偏离指令。","best_for":["高并发低成本应用（当年）","多模态输入理解","历史对照"],"not_for":["复杂推理","新项目"]},"sheet":{"architecture_md":"**未披露**。Google 未公开参数量。\n\n已知信息：\n- 1M 上下文，8K 最大输出\n- 原生工具使用：Google 搜索接地、代码执行\n- Multimodal Live API：实时双向音视频流\n- 实验版支持原生图像输出与多语言 TTS\n- Flash-Lite 为更低价变体（2025-02）","memory_md":"无自建选项。\n\n成本参考：文本 / 图像 / 视频输入 $0.10、音频输入 $0.70、输出 $0.40；缓存输入 $0.025。","training_md":"- 训练细节未披露，知识截止 2024-08\n- 无推理模式（Flash Thinking 为独立实验模型）\n- 支持函数调用、结构化输出、搜索接地\n- 2025-02-05 GA 版相对 12 月实验版 GPQA 略降、MMMU 略升","ecosystem_md":"- Google AI Studio / Gemini API、Vertex AI\n- 支持微调（Vertex AI 监督微调）\n- 中文文档：有（Google 官方中文文档）","versions_md":"- 上代：Gemini 1.5 Flash\n- 版本：2.0-flash-exp（2024-12-11）/ 2.0-flash-001（2025-02-05 GA）\n- 同代：2.0 Flash-Lite、2.0 Pro（实验）、2.0 Flash Thinking（实验）\n- 继任：Gemini 2.5 Flash（2025-04 预览，06 GA）"},"ecosystem":{"engines":[],"finetune":"支持（Vertex AI）","zh_docs":"有"}},{"id":"gemini-2-5-flash","name":"Gemini 2.5 Flash","aliases":["gemini-2.5-flash"],"vendor":"Google","vendor_zh":"谷歌","family":"Gemini 2.5","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","released_at":"2025-06-17","updated_at":"2025-12-20","modalities":["text","image","audio","video","tools"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true},"context":{"max_tokens":1048576,"display":"1M","max_output":65536},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":0.3,"output_per_m":2.5,"currency":"USD","source":"Google AI 定价页","as_of":"2025-12-20"},"links":{"official":"https://deepmind.google/models/gemini/flash/","pricing":"https://ai.google.dev/pricing"},"copy":{"one_liner":"1M 上下文的低价多模态档，思考预算可控。","highlights":["$0.30 / $2.50","1M 上下文 + 多模态","thinking budget 可设 0"],"pitfalls":["复杂推理明显弱于 Pro","代码 agent 一般","无自建"],"logic_ability":"考试型推理中上，可关思考。","best_for":["海量文档 / 视频处理","低价多模态"],"not_for":["复杂 agent"]},"capability_notes":{},"complete":false,"superseded_by":"gemini-3-7-flash"},{"id":"gemini-2-5-pro","name":"Gemini 2.5 Pro","aliases":["gemini-2.5-pro"],"vendor":"Google","vendor_zh":"谷歌","family":"Gemini 2.5","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","superseded_by":"gemini-3-1-pro","released_at":"2025-06-17","updated_at":"2025-12-20","modalities":["text","image","audio","video","tools"],"reasoning_mode":"default-on","architecture":{"type":"unknown","undisclosed":true,"notes":"官方称为稀疏 MoE。"},"context":{"max_tokens":1048576,"display":"1M","max_output":65536},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":1.25,"output_per_m":10,"currency":"USD","source":"Google AI 定价页","as_of":"2025-12-20","note":">200K：$2.5 / $15"},"links":{"official":"https://deepmind.google/models/gemini/pro/","pricing":"https://ai.google.dev/pricing"},"copy":{"one_liner":"2025 年中的多模态旗舰，1M 上下文，已被 Gemini 3 Pro 替代。","highlights":["1M 上下文","视频 / 音频原生输入","GPQA 86.4%"],"pitfalls":["已被 Gemini 3 Pro 替代","代码 agent 弱于 Claude","无自建"],"logic_ability":"考试型推理强，工程型中上。","best_for":["长多模态输入"],"not_for":["新项目（用 3 Pro）"]},"capability_notes":{},"complete":false},{"id":"gemini-3-1-pro","name":"Gemini 3.1 Pro","name_zh":"Gemini 3.1 专业版","aliases":["gemini-3.1-pro-preview","Gemini 3.1 Pro Preview"],"vendor":"Google","vendor_zh":"谷歌","family":"Gemini 3","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"preview","released_at":"2026-02-19","updated_at":"2026-08-28","modalities":["text","image","audio","video","tools"],"reasoning_mode":"default-on","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。延续 Gemini 3 系列的稀疏 MoE、原生多模态设计；支持 thinking_level 调节推理深度，另有 Deep Think 模式。"},"context":{"max_tokens":1048576,"display":"1M","max_output":65536},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":2,"output_per_m":12,"currency":"USD","source":"Google AI 定价页","as_of":"2026-08-28","note":">200K 上下文：$4 / $18；缓存输入 $0.20；Batch / Flex 半价"},"links":{"official":"https://deepmind.google/models/gemini/pro/","pricing":"https://ai.google.dev/gemini-api/docs/pricing","paper":"https://deepmind.google/models/model-cards/gemini-3-1-pro/"},"copy":{"one_liner":"Google 当前 Pro 旗舰（仍 preview），抽象推理第一。","highlights":["ARC-AGI-2 77.1%、GPQA Diamond 94.3%，考试型推理与抽象推理第一梯队","SWE-bench Verified 80.6%、Terminal-Bench 2.0 68.5%，工程与 agent 能力较 Gemini 3 Pro 大幅提升","1M 上下文 + 原生音视频输入，MRCR v2 128K 84.9%"],"pitfalls":["发布半年仍是 preview 接口，无 GA 承诺，配额与行为可能调整",">200K 上下文价格翻倍，1M 是能力上限不是默认预算","首 token 延迟高（AA 实测 TTFT 约 24 s），不适合交互式低延迟场景"],"logic_ability":"考试型推理顶级：HLE 无工具 44.4%（带搜索 + 代码 51.4%）、GPQA 94.3%、ARC-AGI-2 77.1%。工程型推理由 Gemini 3 Pro 的 76.2% 提升到 80.6%（SWE-bench Verified），tau2-bench 零售 90.8% / 电信 99.3%，agent 可靠性显著改善。长程 agent 任务（APEX-Agents 33.5%）仍是短板。失败模式：超长上下文里中段信息偶有遗漏；工具调用格式偶有偏差；高 thinking_level 下 token 消耗大。","best_for":["科学 / 数学 / 抽象推理","视频 / 音频 / 大量图片理解","百万 token 级长文档分析"],"not_for":["需要 GA 接口与 SLA 的生产线","低延迟对话产品","私有化部署"]},"capability_notes":{"coding":"SWE-bench Verified 80.6%、LiveCodeBench Pro Elo 2887、Terminal-Bench 2.0 68.5%（官方模型卡）。","reasoning":"GPQA Diamond 94.3%、HLE 44.4%（无工具）、ARC-AGI-2 77.1%（官方）。","math":"AIME 2025 官方未单列；官方以 ARC-AGI-2 与 HLE 为主。","agent":"tau2-bench 零售 90.8% / 电信 99.3%、BrowseComp 85.9%、APEX-Agents 33.5%（官方）。","multimodal":"MMMU-Pro 80.5%（官方）。","chinese":"MMMLU 92.6%（官方），中文能力强。"},"sheet":{"architecture_md":"**类型**：未披露（Google 仅说明沿用 Gemini 3 系列的稀疏 MoE Transformer、原生多模态）。\n\n已知：\n- 输入：文本、图像、音频、视频、PDF；输出文本\n- 1M token 上下文，64K 输出\n- `thinking_level` 可调推理深度；Deep Think 模式另行提供\n- 托管工具：Google Search grounding、代码执行、URL 上下文\n- 官方称相较 Gemini 3 Pro 在同等任务上 token 消耗更低\n\n参考：Gemini 3.1 Pro 模型卡（deepmind.google/models/model-cards/gemini-3-1-pro/）。","memory_md":"无自建选项。\n\n成本参考（2026-08-28 定价页）：\n\n| 档位 | 输入 $/M | 输出 $/M |\n|---|---|---|\n| 标准 ≤200K | 2.00 | 12.00 |\n| 标准 >200K | 4.00 | 18.00 |\n| Batch / Flex ≤200K | 1.00 | 6.00 |\n| Priority ≤200K | 3.60 | 21.60 |\n\n缓存输入 $0.20（≤200K）/ $0.40（>200K）。\n\n**警示**：单次 1M 上下文请求输入费可达 $4，务必配合 context caching；高 thinking_level 下输出 token 消耗显著。","training_md":"- 训练细节未披露，TPU 训练\n- 默认开启推理；thinking_level 控制深度\n- 支持函数调用、结构化输出、多模态输入\n- 最大输出 64K\n- 官方基准（模型卡）：HLE 44.4% / 51.4%（带工具）、GPQA 94.3%、ARC-AGI-2 77.1%、SWE-bench Verified 80.6%、Terminal-Bench 2.0 68.5%、LiveCodeBench Pro 2887 Elo、MMMU-Pro 80.5%、MRCR v2 128K 84.9%","ecosystem_md":"- Gemini API、Google AI Studio、Vertex AI\n- Antigravity、Gemini CLI、Gemini App 集成\n- 不支持微调（Pro 档）\n- 中文文档：有（Google 中文文档较全）","versions_md":"- 上代：Gemini 3 Pro（2025-11，已被 3.1 Pro 取代）\n- 同期 Flash 线迭代很快：3 Flash → 3.1 Flash-Lite → 3.5 Flash / 3.5 Flash-Lite（2026-07）→ 3.6 Flash → 3.7 Flash（2026-08-13）\n- 截至 2026-08-28，Pro 线仍停留在 3.1 Pro Preview，未见 3.5 / 3.7 Pro\n- Deep Think 模式另计"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"有"},"i18n":{"en":{"name_zh":"Gemini 3.1 Pro","one_liner":"Google's current Pro flagship (still preview), #1 in abstract reasoning.","highlights":["ARC-AGI-2 77.1%, GPQA Diamond 94.3%; first tier in exam-style and abstract reasoning","SWE-bench Verified 80.6%, Terminal-Bench 2.0 68.5%; engineering and agent ability much improved over Gemini 3 Pro","1M context + native audio/video input, MRCR v2 128K 84.9%"],"pitfalls":["Still a preview endpoint half a year after launch, no GA commitment; quotas and behavior may change","Price doubles above 200K context; 1M is a capability ceiling, not a default budget","High first-token latency (AA-measured TTFT about 24 s), unsuited to interactive low-latency use"],"logic_ability":"Top-tier exam-style reasoning: HLE 44.4% without tools (51.4% with search + code), GPQA 94.3%, ARC-AGI-2 77.1%. Engineering reasoning rose from Gemini 3 Pro's 76.2% to 80.6% (SWE-bench Verified); tau2-bench retail 90.8% / telecom 99.3%, agent reliability notably improved. Long-horizon agent tasks (APEX-Agents 33.5%) remain a weakness. Failure modes: occasional loss of mid-context information at very long contexts; occasional tool-call format deviations; heavy token use at high thinking_level.","best_for":["Science / math / abstract reasoning","Video / audio / bulk image understanding","Million-token document analysis"],"not_for":["Production lines needing a GA endpoint and SLA","Low-latency chat products","On-prem deployment"],"capability_notes":{"coding":"SWE-bench Verified 80.6%, LiveCodeBench Pro Elo 2887, Terminal-Bench 2.0 68.5% (official model card).","reasoning":"GPQA Diamond 94.3%, HLE 44.4% (no tools), ARC-AGI-2 77.1% (official).","math":"AIME 2025 not listed separately; official reporting centers on ARC-AGI-2 and HLE.","agent":"tau2-bench retail 90.8% / telecom 99.3%, BrowseComp 85.9%, APEX-Agents 33.5% (official).","multimodal":"MMMU-Pro 80.5% (official).","chinese":"MMMLU 92.6% (official); strong Chinese ability."}},"ja":{"name_zh":"Gemini 3.1 Pro","one_liner":"Google 現行 Pro 旗艦（まだ preview）、抽象推論 1 位。","highlights":["ARC-AGI-2 77.1%、GPQA Diamond 94.3%、試験型推論と抽象推論で最上位グループ","SWE-bench Verified 80.6%、Terminal-Bench 2.0 68.5%、エンジニアリングと agent 能力は Gemini 3 Pro から大幅向上","1M コンテキスト + ネイティブの音声・動画入力、MRCR v2 128K 84.9%"],"pitfalls":["公開から半年経っても preview エンドポイントのままで GA の約束なし。クォータや挙動が変わる可能性あり","200K 超のコンテキストは価格が 2 倍。1M は能力の上限であってデフォルト予算ではない","初回トークン遅延が大きい（AA 実測 TTFT 約 24 s）、対話型の低遅延用途には不向き"],"logic_ability":"試験型推論はトップ級：HLE ツールなし 44.4%（検索 + コードありで 51.4%）、GPQA 94.3%、ARC-AGI-2 77.1%。エンジニアリング型推論は Gemini 3 Pro の 76.2% から 80.6% に向上（SWE-bench Verified）、tau2-bench リテール 90.8% / テレコム 99.3% で agent の信頼性が大きく改善。長期的な agent タスク（APEX-Agents 33.5%）は依然弱点。失敗モード：超長コンテキストで中盤の情報を時折見落とす、ツール呼び出しのフォーマットが時折ずれる、高 thinking_level ではトークン消費が大きい。","best_for":["科学 / 数学 / 抽象推論","動画 / 音声 / 大量画像の理解","100 万トークン級の長文書分析"],"not_for":["GA エンドポイントと SLA が必要な本番ライン","低遅延の対話製品","オンプレ展開"],"capability_notes":{"coding":"SWE-bench Verified 80.6%、LiveCodeBench Pro Elo 2887、Terminal-Bench 2.0 68.5%（公式モデルカード）。","reasoning":"GPQA Diamond 94.3%、HLE 44.4%（ツールなし）、ARC-AGI-2 77.1%（公式）。","math":"AIME 2025 は公式に個別掲載なし。公式は ARC-AGI-2 と HLE を中心に報告。","agent":"tau2-bench リテール 90.8% / テレコム 99.3%、BrowseComp 85.9%、APEX-Agents 33.5%（公式）。","multimodal":"MMMU-Pro 80.5%（公式）。","chinese":"MMMLU 92.6%（公式）、中国語能力は高い。"}}},"complete":true,"runtime":{"tok_s":115,"latency_s":23.93,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"gemini-3-5-flash-lite","name":"Gemini 3.5 Flash-Lite","name_zh":"Gemini 3.5 闪电轻量版","aliases":["gemini-3.5-flash-lite","Gemini 3.5 Flash Lite"],"vendor":"Google","vendor_zh":"谷歌","family":"Gemini 3","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2026-07-21","updated_at":"2026-08-28","modalities":["text","image","audio","video","tools"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。"},"context":{"max_tokens":1048576,"display":"1M","max_output":65536},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":0.3,"output_per_m":2.5,"currency":"USD","source":"Google AI 定价页","as_of":"2026-08-28","note":"缓存输入 $0.03；Batch / Flex 半价"},"links":{"official":"https://deepmind.google/models/model-cards/gemini-3-5-flash-lite/","pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"copy":{"one_liner":"Google 当前最便宜的 1M 上下文多模态模型，吞吐优先。","highlights":["$0.30 / $2.50，1M 上下文，全模态输入","AA 实测约 350 tok/s，Terminal-Bench 2.1 54.0%（上代 3.1 Flash-Lite 31.0%）","GDM-MRCR v2 128K 72.2%，长上下文检索远超 Haiku 4.5 / GPT-5.4 mini"],"pitfalls":["AA 智能指数 37，复杂推理明显弱于 Flash 主力","输出价 $2.50 比上代 3.1 Flash-Lite（$1.50）贵 67%，纯文本轻任务未必划算","官方未公布 GPQA / HLE / MMMU 等通用基准"],"logic_ability":"轻量档，agent / 抽取 / 分类类任务为主。SWE-Bench Pro 54.2%、OSWorld-Verified 74.0%。深度推理应上 Flash 或 Pro。","best_for":["大规模文档抽取 / 翻译","agentic search","低延迟高并发"],"not_for":["复杂推理与数学","私有化部署"]},"capability_notes":{"coding":"SWE-Bench Pro 54.2%（官方模型卡）。","agent":"Terminal-Bench 2.1 54.0%、OSWorld-Verified 74.0%（官方）。"},"i18n":{"en":{"name_zh":"Gemini 3.5 Flash-Lite","one_liner":"Google's cheapest 1M-context multimodal model, built for throughput.","highlights":["$0.30 / $2.50, 1M context, all-modality input","AA-measured about 350 tok/s; Terminal-Bench 2.1 54.0% (previous 3.1 Flash-Lite 31.0%)","GDM-MRCR v2 128K 72.2%; long-context retrieval far ahead of Haiku 4.5 / GPT-5.4 mini"],"pitfalls":["AA Intelligence Index 37; complex reasoning clearly weaker than the main Flash","Output price $2.50 is 67% above the previous 3.1 Flash-Lite ($1.50); not necessarily worthwhile for light text-only tasks","No official GPQA / HLE / MMMU or other general benchmarks published"],"logic_ability":"Lightweight tier, mainly for agent / extraction / classification tasks. SWE-Bench Pro 54.2%, OSWorld-Verified 74.0%. Deep reasoning should go to Flash or Pro.","best_for":["Large-scale document extraction / translation","Agentic search","Low-latency high-concurrency"],"not_for":["Complex reasoning and math","On-prem deployment"],"capability_notes":{"coding":"SWE-Bench Pro 54.2% (official model card).","agent":"Terminal-Bench 2.1 54.0%, OSWorld-Verified 74.0% (official)."}},"ja":{"name_zh":"Gemini 3.5 Flash-Lite","one_liner":"Google 最安の 1M コンテキスト・マルチモーダル、スループット優先。","highlights":["$0.30 / $2.50、1M コンテキスト、全モダリティ入力","AA 実測約 350 tok/s、Terminal-Bench 2.1 54.0%（前世代 3.1 Flash-Lite は 31.0%）","GDM-MRCR v2 128K 72.2%、長文コンテキスト検索は Haiku 4.5 / GPT-5.4 mini を大きく上回る"],"pitfalls":["AA 知能指数 37、複雑な推論は主力 Flash より明らかに弱い","出力価格 $2.50 は前世代 3.1 Flash-Lite（$1.50）より 67% 高く、テキストのみの軽タスクでは割に合わないことも","GPQA / HLE / MMMU などの汎用ベンチマークは公式未公表"],"logic_ability":"軽量クラスで、agent / 抽出 / 分類系タスクが主用途。SWE-Bench Pro 54.2%、OSWorld-Verified 74.0%。深い推論は Flash か Pro に任せるべき。","best_for":["大規模な文書抽出 / 翻訳","agentic search","低遅延・高並行"],"not_for":["複雑な推論と数学","オンプレ展開"],"capability_notes":{"coding":"SWE-Bench Pro 54.2%（公式モデルカード）。","agent":"Terminal-Bench 2.1 54.0%、OSWorld-Verified 74.0%（公式）。"}}},"complete":false,"runtime":{"tok_s":346,"latency_s":8.76,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"gemini-3-7-flash","name":"Gemini 3.7 Flash","name_zh":"Gemini 3.7 闪电版","aliases":["gemini-3.7-flash","Gemini 3.7 Flash High"],"vendor":"Google","vendor_zh":"谷歌","family":"Gemini 3","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2026-08-13","updated_at":"2026-08-28","modalities":["text","image","audio","video","tools"],"reasoning_mode":"default-on","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。基于 Gemini 3.6 Flash 继续训练；thinking_level 可选 low / medium（默认）/ high，不支持 minimal。知识截止 2026-03。"},"context":{"max_tokens":1048576,"display":"1M","max_output":65536},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":0.75,"output_per_m":3.75,"currency":"USD","source":"Google AI 定价页","as_of":"2026-08-28","note":"2026-12-31 前为发布优惠价；2027-01-01 起 $1.50 / $7.50。缓存输入 $0.075"},"links":{"official":"https://deepmind.google/models/model-cards/gemini-3-7-flash/","pricing":"https://ai.google.dev/gemini-api/docs/pricing","paper":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/"},"copy":{"one_liner":"Google 当前 Flash 主力，代码 / agent 强，年内半价。","highlights":["Terminal-Bench 2.1 85.8%、DeepSWE v1.1 65.3%，Flash 档做到接近 Sonnet 5 水平","AA 智能指数 56，与 Grok 4.5 持平，输出速度约 300 tok/s","Arena 文本榜 1490，高于 Gemini 3.1 Pro"],"pitfalls":["$0.75 / $3.75 是 2026 年底前的优惠价，2027 年起翻倍到 $1.50 / $7.50","Flash 线三周一迭代（3.5 → 3.6 → 3.7），模型 ID 与行为漂移快，需固定版本","官方模型卡未公布 GPQA / HLE / SWE-bench Verified 等通用基准，横向比较依赖第三方"],"logic_ability":"工程型推理为主打：DeepSWE 65.3%、FrontierCode 43.6%、Terminal-Bench 2.1 85.8%；长上下文 MRCR v2 8-needle 97.0%。考试型推理官方未列数据。","best_for":["高频代码 / agent 任务","长文档批处理","高性价比多模态"],"not_for":["需要长期稳定价格的预算规划","私有化部署"]},"capability_notes":{"coding":"DeepSWE v1.1 65.3%、FrontierCode 1.1 43.6%（官方模型卡；Sonnet 5 为 53.8% / 42.7%）；WebDev Arena 1588（官方博客）。","agent":"Terminal-Bench 2.1 85.8%、OSWorld-2.0 47.9%、AutomationBench 30.4%（官方模型卡 / 博客）。","reasoning":"AA 智能指数 56（high 档，2026-08-28，3.6 Flash 为 52）；官方未公布 GPQA / HLE。","knowledge":"GDM-MRCR 128K 97.0%（官方模型卡，同表最高）；GDP.pdf 文档任务 34.0%（官方博客）。","multimodal":"支持文本 / 图像 / 音频 / 视频 / PDF 输入，仅文本输出；官方未公布 MMMU。"},"i18n":{"en":{"name_zh":"Gemini 3.7 Flash","one_liner":"Google's current main Flash; strong code / agent, half price through year-end.","highlights":["Terminal-Bench 2.1 85.8%, DeepSWE v1.1 65.3%; a Flash tier approaching Sonnet 5 level","AA Intelligence Index 56, level with Grok 4.5, output speed about 300 tok/s","Arena text 1490, above Gemini 3.1 Pro"],"pitfalls":["$0.75 / $3.75 is a promotional price through end of 2026; doubles to $1.50 / $7.50 from 2027","The Flash line iterates every three weeks (3.5 -> 3.6 -> 3.7); model IDs and behavior drift fast, pin versions","Official model card omits GPQA / HLE / SWE-bench Verified and other general benchmarks; cross-comparison relies on third parties"],"logic_ability":"Engineering reasoning is the focus: DeepSWE 65.3%, FrontierCode 43.6%, Terminal-Bench 2.1 85.8%; long-context MRCR v2 8-needle 97.0%. No official exam-style reasoning figures.","best_for":["High-frequency code / agent tasks","Long-document batch processing","Cost-effective multimodal"],"not_for":["Budget planning that needs long-term stable pricing","On-prem deployment"],"capability_notes":{"coding":"DeepSWE v1.1 65.3%, FrontierCode 1.1 43.6% (official model card; Sonnet 5 at 53.8% / 42.7%); WebDev Arena 1588 (official blog).","agent":"Terminal-Bench 2.1 85.8%, OSWorld-2.0 47.9%, AutomationBench 30.4% (official model card / blog).","reasoning":"AA Intelligence Index 56 (high, 2026-08-28; 3.6 Flash was 52); GPQA / HLE not officially published.","knowledge":"GDM-MRCR 128K 97.0% (official model card, highest in the table); GDP.pdf document tasks 34.0% (official blog).","multimodal":"Text / image / audio / video / PDF input, text output only; MMMU not officially published."}},"ja":{"name_zh":"Gemini 3.7 Flash","one_liner":"Google 現行主力 Flash。コード / agent に強く年内半額。","highlights":["Terminal-Bench 2.1 85.8%、DeepSWE v1.1 65.3%、Flash クラスで Sonnet 5 に迫る水準","AA 知能指数 56 で Grok 4.5 と同等、出力速度約 300 tok/s","Arena テキスト部門 1490、Gemini 3.1 Pro を上回る"],"pitfalls":["$0.75 / $3.75 は 2026 年末までの優遇価格で、2027 年からは $1.50 / $7.50 に倍増","Flash 系は 3 週間ごとに更新（3.5 → 3.6 → 3.7）、モデル ID と挙動の変化が速いためバージョン固定が必要","公式モデルカードに GPQA / HLE / SWE-bench Verified などの汎用ベンチマークがなく、横比較は第三者頼み"],"logic_ability":"エンジニアリング型推論が主軸：DeepSWE 65.3%、FrontierCode 43.6%、Terminal-Bench 2.1 85.8%。長文コンテキスト MRCR v2 8-needle 97.0%。試験型推論は公式データなし。","best_for":["高頻度のコード / agent タスク","長文書のバッチ処理","コスパの良いマルチモーダル"],"not_for":["長期的に安定した価格が必要な予算計画","オンプレ展開"],"capability_notes":{"coding":"DeepSWE v1.1 65.3%、FrontierCode 1.1 43.6%（公式モデルカード。Sonnet 5 は 53.8% / 42.7%）。WebDev Arena 1588（公式ブログ）。","agent":"Terminal-Bench 2.1 85.8%、OSWorld-2.0 47.9%、AutomationBench 30.4%（公式モデルカード / ブログ）。","reasoning":"AA 知能指数 56（high、2026-08-28。3.6 Flash は 52）。GPQA / HLE は公式未公表。","knowledge":"GDM-MRCR 128K 97.0%（公式モデルカード、同表最高）。GDP.pdf 文書タスク 34.0%（公式ブログ）。","multimodal":"テキスト / 画像 / 音声 / 動画 / PDF 入力、出力はテキストのみ。MMMU は公式未公表。"}}},"complete":true,"sheet":{"architecture_md":"**未披露**。Google 未公布参数量与架构，模型卡只说明「基于 Gemini 3.6 Flash」，细节引用 3.6 Flash 模型卡。\n\n已知接口特性（ai.google.dev 模型页，2026-08-28）：\n- 输入 1,048,576 token，输出 65,536 token\n- 输入：文本、图像、视频、音频、PDF；输出：仅文本（无图像 / 音频生成，无 Live API）\n- `thinking_level`：`low` / `medium`（默认）/ `high`；**不支持 `minimal`**，思考不可完全关闭\n- 能力：function calling、code execution、Google Search grounding、Google Maps grounding、URL context、file search、structured outputs、context caching、Batch API、computer use（预览）\n- 模型 ID `gemini-3.7-flash`，状态 Stable（GA）\n- 知识截止 2026-03（部分领域仅到 2025-01）","memory_md":"无自建选项。\n\n**成本模型**（Google AI 定价页，2026-08-28）：\n- 2026-12-31 前发布价：输入 $0.75 / 输出 $3.75（含思考 token）\n- **2027-01-01 起翻倍**：$1.50 / $7.50\n- 缓存读 $0.075（10%），缓存存储 $0.50 / 1M token / 小时\n- Batch 与 Flex 五折；Priority 1.8×\n- Search / Maps grounding：3.x 系列共享每月 5,000 次免费，之后 $14 / 1,000 次\n- 免费层：AI Studio 免费但内容用于产品改进；付费层不用于训练\n\n**什么时候会变贵**：\n- 思考 token 按输出计费，`high` 档长任务输出膨胀\n- 缓存存储按小时计，长期挂着大上下文会被存储费吃掉\n- 2027 年预算要按翻倍后价格做\n\n速度：AA 实测约 306 tok/s（high 档，187 个模型中第一），TTFT 约 12 s；单任务成本 $0.40。","training_md":"- 训练细节未披露；模型卡称在 Gemini 3.6 Flash 基础上构建，训练数据说明沿用 3.6 Flash\n- 知识截止 2026-03\n- 官方博客强调提升：生产级代码质量与调试、Web 开发 / UI 生成、知识密集流程（金融 / 法律 / 生物）、复杂文档处理、业务流程自动化\n- 官方数据（vs 3.6 Flash）：FrontierCode 43.6%（34.4%）、DeepSWE 65.3%（48.6~49.0%）、WebDev Arena 1588（1538）、GDP.pdf 34.0%（22.0%）、AutomationBench 30.4%（17.0%）、MRCR 128K 97.0%（91.8%）\n- 安全：Frontier Safety 评估未触及任何 tracked / critical 能力等级；红队测试「与 3.6 Flash 相当或更好」；强化了 CBRN 与网络滥用防护","ecosystem_md":"- **API 渠道**：Gemini API（AI Studio）、Vertex AI / Gemini Enterprise Agent Platform（同 ID `gemini-3.7-flash`，支持 batch、context caching、grounding、provisioned throughput）\n- 产品：Gemini App（Spark，Pro / Ultra 订阅 160+ 国家）、Gemini Enterprise app、Google Antigravity、Android Studio、Gemini CLI\n- SDK：google-genai（Python / JS / Go / Java）、Vertex AI SDK；OpenAI 兼容端点\n- 微调：模型页未列出 tuning 支持，按不支持处理\n- 中文文档：有（ai.google.dev 与 cloud.google.com 提供中文版）","versions_md":"- Flash 线迭代：Gemini 3 Flash（2025-12）→ 3.1 Flash → 3.5 Flash → 3.6 Flash → **3.7 Flash**（2026-08-13 GA），约三周一版，行为漂移快\n- 同代 Pro 线：Gemini 3.1 Pro（Arena 文本榜低于 3.7 Flash）\n- 更低档：Gemini 3.1 Flash-Lite\n- 定价节点：2026-12-31 优惠价到期，2027-01-01 起 $1.50 / $7.50\n- 模型 ID：`gemini-3.7-flash`（Stable），建议在生产中固定版本而非跟随最新"},"ecosystem":{"engines":[],"finetune":"未列出，按不支持处理","zh_docs":"有"},"runtime":{"tok_s":306,"latency_s":11.73,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"gemini-3-flash","name":"Gemini 3 Flash","aliases":["gemini-3-flash-preview"],"vendor":"Google","vendor_zh":"谷歌","family":"Gemini 3","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","released_at":"2025-12-17","updated_at":"2025-12-20","modalities":["text","image","audio","video","tools"],"reasoning_mode":"default-on","architecture":{"type":"unknown","undisclosed":true},"context":{"max_tokens":1048576,"display":"1M","max_output":65536},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":0.5,"output_per_m":3,"currency":"USD","source":"Google AI 定价页","as_of":"2025-12-20"},"links":{"official":"https://deepmind.google/models/gemini/flash/","pricing":"https://ai.google.dev/pricing"},"copy":{"one_liner":"Flash 价位跑出接近 3 Pro 的分数，2025 年底性价比最强闭源之一。","highlights":["GPQA 90.4%、SWE-bench 78%（官方）","$0.50 / $3","1M 上下文多模态"],"pitfalls":["preview 状态","长 agent 稳定性待验证","无自建"],"logic_ability":"考试型推理接近 Pro；工程型推理官方数字强，独立复测待补。","best_for":["高性价比通用 API","多模态批处理"],"not_for":["私有化"]},"capability_notes":{},"complete":false,"superseded_by":"gemini-3-7-flash"},{"id":"gemini-3-pro","name":"Gemini 3 Pro","name_zh":"Gemini 3 专业版","aliases":["gemini-3-pro-preview","Gemini 3"],"vendor":"Google","vendor_zh":"谷歌","family":"Gemini 3","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","released_at":"2025-11-18","updated_at":"2025-12-20","modalities":["text","image","audio","video","tools"],"reasoning_mode":"default-on","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。Google 仅说明为稀疏 MoE Transformer，原生多模态。thinking_level 参数（low / high）。"},"context":{"max_tokens":1048576,"display":"1M","max_output":65536},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":2,"output_per_m":12,"currency":"USD","source":"Google AI 定价页","as_of":"2025-12-20","note":">200K 上下文：$4 / $18"},"links":{"official":"https://deepmind.google/models/gemini/pro/","pricing":"https://ai.google.dev/pricing","paper":"https://blog.google/products/gemini/gemini-3/"},"copy":{"one_liner":"2025 年底榜单第一的多模态旗舰，1M 上下文。","highlights":["LMArena 文本榜首次突破 1500 Elo","HLE 37.5%、GPQA 91.9%，考试型推理断层领先","原生视频 / 音频 / 图像输入 + 1M 上下文，长多模态任务无对手"],"pitfalls":["preview 状态，接口行为与配额可能变动","超过 200K 上下文后价格翻倍，1M 是能力不是默认预算","Agent 长任务稳定性略逊于 Claude Opus 4.5（Terminal-bench / tau2 差距）"],"logic_ability":"考试型推理当前第一：HLE、GPQA、ARC-AGI-2 均领先。工程型推理（SWE-bench 76.2%）优秀但非第一。thinking_level=high 时会做大量内部搜索式推理；low 更接近 Flash 行为。失败模式：在极长上下文里偶尔遗漏中段信息；工具调用格式偶有偏差。","best_for":["视频 / 音频 / 大量图片理解","超长文档（百万 token 级）分析","科学 / 数学推理"],"not_for":["需要稳定 GA 接口的生产线（等 GA）","私有化部署"]},"capability_notes":{"coding":"SWE-bench Verified 76.2%、LiveCodeBench Pro Elo 2439（官方）。","reasoning":"GPQA 91.9%、HLE 37.5%、ARC-AGI-2 31.1%（官方）。","math":"AIME 2025 95%（无工具，官方）。","agent":"Terminal-bench 2.0 54.2%（官方）。","multimodal":"MMMU-Pro 81%、Video-MMMU 87.6%（官方），当前最强。","chinese":"中文能力强，多语言 MMMLU 91.8%。"},"sheet":{"architecture_md":"**类型**：稀疏 MoE Transformer（Google 官方描述），细节未披露。\n\n已知：\n- 原生多模态：文本、图像、音频、视频统一输入\n- 1M token 上下文，64K 输出\n- `thinking_level`：low / high\n- 支持 Google Search grounding、代码执行、URL 上下文等托管工具","memory_md":"无自建选项。\n\n成本参考：≤200K 上下文 $2 / $12；>200K $4 / $18。缓存输入 $0.20。\n\n**警示**：1M 上下文下一次请求输入费可达 $4；请配合 context caching。","training_md":"- 训练细节未披露，TPU 训练\n- 默认推理开启\n- 支持函数调用、结构化输出、多模态输入、原生音频\n- 最大输出 64K","ecosystem_md":"- Google AI Studio、Vertex AI、Gemini API\n- Antigravity IDE、Gemini CLI 集成\n- 不支持微调（Pro 档）\n- 中文文档：有（Google 中文文档较全）","versions_md":"- 上代：Gemini 2.5 Pro\n- 同代：Gemini 3 Flash（2025-12）\n- Deep Think 模式另计"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"有"},"complete":true,"superseded_by":"gemini-3-1-pro"},{"aliases":["gemma-2-27b-it"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"coding":"HumanEval 51.8%（官方基座）。","math":"GSM8K 74.0%、MATH 42.3%（官方基座）。","knowledge":"MMLU 75.2%（官方基座）。","chinese":"中文可用，弱于 Qwen。"},"complete":true,"id":"gemma-2-27b","name":"Gemma 2 27B","name_zh":"Gemma 2 · 27B","vendor":"Google","family":"Gemma 2","license":"Gemma Terms of Use","license_commercial":"restricted","weights_url":"https://huggingface.co/google/gemma-2-27b-it","superseded_by":"gemma-3-27b","released_at":"2024-06-27","architecture":{"type":"dense","total_params":"27.2B","total_params_b":27.2,"layers":46,"hidden_size":4608,"vocab_size":256000,"kv_heads":16,"head_dim":128,"attention":"GQA（32 Q 头 / 16 KV 头）；局部 4K 滑窗与全局注意力逐层交替","notes":"13T token；logit soft-capping；上下文 8K；从零训练（非蒸馏，9B 才是蒸馏）。","undisclosed":false},"context":{"max_tokens":8192,"display":"8K"},"memory":{"weight_gb":{"bf16":54.4,"q8":28.8,"q4":16.6},"estimated":false,"kv_per_token_kib":368,"kv_note":"16×128×2×46×2 B = 368 KiB/token（一半层为 4K 滑窗，长文实际更小）；8K 上限约 2.9 GB。","ref_hw_24gb":"Q4_K_M 16.6 GB，RTX 4090 可跑满 8K 上下文","ref_hw_80gb":"BF16 54.4 GB 单卡","ref_hw_8x80gb":"过剩"},"links":{"official":"https://blog.google/technology/developers/google-gemma-2/","paper":"https://arxiv.org/abs/2408.00118","hf":"https://huggingface.co/google/gemma-2-27b-it"},"copy":{"one_liner":"27B 打到 70B 级对话质量，单卡 24GB 的甜点。","highlights":["Arena 一度超过 Llama 3 70B，同尺寸对话质量当年第一","Q4 16.6 GB 单张 24GB 卡跑满上下文","多语言与写作风格好，官方 TPU / GPU 全支持"],"pitfalls":["上下文仅 8K，是最大硬伤","无 system prompt、无原生工具调用","Gemma Terms 含禁用政策，非 OSI 许可"],"logic_ability":"非推理模型中游偏上：MMLU 75.2%、GSM8K 74.0%、MATH 42.3%（官方基座）。对话与写作优秀，数学 / 代码明显不及 Qwen2.5-32B。","best_for":["单卡英文 / 多语言对话助手","写作与总结"],"not_for":["长文档","代码 / 数学"]},"sheet":{"architecture_md":"**类型**：Dense Transformer，27.2B 参数。\n\n- 46 层，隐藏维 4608，FFN 36,864（GeGLU），词表 256,000\n- GQA：32 Query 头 / 16 KV 头，head_dim 128\n- 局部滑窗（4096）与全局注意力逐层交替\n- Logit soft-capping（attn 50 / final 30），RMSNorm 前后双归一化\n- 上下文 8192\n\n参考：Gemma 2 技术报告（arXiv 2408.00118）。","memory_md":"| 精度 | 权重大小 | 参考硬件 | 说明 |\n|---|---|---|---|\n| BF16 | 54.4 GB | 80GB 单卡 | 官方 safetensors |\n| Q8 | ≈ 28.9 GB | 32GB / 48GB | GGUF Q8_0 |\n| Q4 | 16.6 GB | 24GB | GGUF Q4_K_M（bartowski） |\n\n**KV Cache**：368 KiB/token（全局层），滑窗层封顶 4K。8K 上下文最多约 2.9 GB。\n\n**参考配置**：\n- RTX 4090 24GB：Q4_K_M + 8K 全上下文，余量充足\n- A100 80GB：BF16，多并发\n- 注意：soft-capping 要求 eager attention 或 flash-attn ≥ 2.6，早期 vLLM 需 flashinfer","training_md":"- 预训练 13T token（27B），主要英文网页 / 代码 / 科学\n- 27B 从零训练；2B / 9B 用 27B 作教师蒸馏\n- 后训练：SFT + RLHF（奖励模型比策略大一个数量级）+ 模型合并（WARP）\n- 无 system 角色；对话模板 `<start_of_turn>user`","ecosystem_md":"- HF：google/gemma-2-27b-it、-27b（基座）\n- 引擎：vLLM、llama.cpp、Ollama、MLX、Keras / JAX、TensorRT-LLM\n- 微调：LoRA 友好（Unsloth、LLaMA-Factory），注意 soft-capping\n- 中文文档：弱（官方英文为主）","versions_md":"- 同系列：Gemma 2 9B（蒸馏）、2B（2024-07）\n- 后继：Gemma 3 27B（2025-03，128K + 多模态）→ Gemma 4\n- 衍生：ShieldGemma、DataGemma、日语特化版"},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama","MLX","TensorRT-LLM"],"finetune":"LoRA 友好","zh_docs":"弱"}},{"aliases":["gemma-2-9b-it"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"knowledge":"MMLU 71.3%（官方基座）。","math":"GSM8K 68.6%（官方基座）。"},"complete":false,"id":"gemma-2-9b","name":"Gemma 2 9B","name_zh":"Gemma 2 · 9B","vendor":"Google","family":"Gemma 2","license":"Gemma Terms of Use","license_commercial":"restricted","weights_url":"https://huggingface.co/google/gemma-2-9b-it","superseded_by":"gemma-4-26b-a4b","released_at":"2024-06-27","architecture":{"type":"dense","total_params":"9.2B","total_params_b":9.2,"layers":42,"hidden_size":3584,"vocab_size":256000,"kv_heads":8,"head_dim":256,"attention":"GQA（16 Q 头 / 8 KV 头）；4K 滑窗交替","notes":"8T token，由 27B 蒸馏；上下文 8K。","undisclosed":false},"context":{"max_tokens":8192,"display":"8K"},"memory":{"weight_gb":{"bf16":18.4,"q8":9.8,"q4":5.8},"estimated":false,"kv_per_token_kib":336,"kv_note":"8×256×2×42×2 B = 336 KiB/token。","ref_hw_24gb":"BF16 18.4 GB 单卡可跑","ref_hw_80gb":"高并发","ref_hw_8x80gb":"过剩"},"links":{"official":"https://blog.google/technology/developers/google-gemma-2/","paper":"https://arxiv.org/abs/2408.00118","hf":"https://huggingface.co/google/gemma-2-9b-it"},"copy":{"one_liner":"蒸馏出的 9B，对话质量一度超过 Llama 3 8B。","highlights":["知识蒸馏使 9B 逼近 27B 的一半以上能力","Q4 5.8 GB，笔记本可跑","多语言与写作优秀"],"pitfalls":["上下文 8K","无工具调用 / system prompt","Gemma Terms 非 OSI 许可"],"logic_ability":"9B 中上：MMLU 71.3%、GSM8K 68.6%（官方基座）。","best_for":["轻量多语言对话","教学"],"not_for":["长文本","代码"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama","MLX"],"finetune":"LoRA 友好","zh_docs":"弱"}},{"id":"gemma-3-27b","name":"Gemma 3 27B","name_zh":"谷歌 Gemma 3 · 27B","aliases":["gemma-3-27b-it","Gemma3 27B"],"vendor":"Google","vendor_zh":"谷歌","family":"Gemma 3","license":"Gemma Terms of Use","license_commercial":"restricted","openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/google/gemma-3-27b-it","status":"superseded","released_at":"2025-03-12","updated_at":"2025-12-20","modalities":["text","image","tools"],"reasoning_mode":"none","architecture":{"type":"dense","total_params":"27B","total_params_b":27,"layers":62,"kv_layers":11,"hidden_size":5376,"vocab_size":262144,"kv_heads":16,"head_dim":128,"attention":"GQA（32 Q 头 / 16 KV 头），5:1 局部滑窗(1024) : 全局","notes":"62 层中约 11 层全局注意力，其余为 1024 token 滑窗，KV 显著降低。SigLIP 400M 视觉编码器，图像 896×896 固定 256 token。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":8192},"memory":{"weight_gb":{"bf16":54,"q8":28.7,"q4":16.5},"kv_per_token_kib":44,"kv_note":"全局层 11 × 16 KV 头 × 128 × 2 × 2 B ≈ 44 KiB/token；局部层 KV 上限 1024 token，可忽略。","ref_hw_24gb":"Q4_K_M 16.5 GB，KV 余量充足，32K+ 上下文可行","ref_hw_80gb":"BF16 54 GB 舒适","ref_hw_8x80gb":"过剩","estimated":false},"pricing":{"input_per_m":0.09,"output_per_m":0.16,"currency":"USD","source":"第三方托管常见价","as_of":"2025-12-20","note":"开源权重，社区托管价"},"links":{"official":"https://ai.google.dev/gemma","hf":"https://huggingface.co/google/gemma-3-27b-it","paper":"https://arxiv.org/abs/2503.19786"},"copy":{"one_liner":"单卡跑的多模态 Dense 模型，对话体验与多语言突出。","highlights":["LMArena 1338，发布时单卡可跑模型里人类偏好最高","原生视觉输入 + 140 语言，24GB 卡上 Q4 + 长上下文舒适","滑窗 / 全局 5:1 设计，KV 只有同规模模型的 1/5"],"pitfalls":["Gemma 许可有使用限制条款（禁用用途、需传递条款），不是 Apache/MIT","无思考模式，数学 / 代码推理明显弱于 Qwen3 同档","最大输出 8K，长生成任务受限"],"logic_ability":"非推理模型，考试型推理弱（AIME 2025 未官方公布，MATH 89%）。强项在多轮对话质量、格式遵循、多语言与视觉描述。工程型推理（代码修改）一般。适合做对话产品与视觉问答，不适合做复杂推理主力。","best_for":["单卡多模态助手","多语言对话产品","视觉问答 / 文档 OCR"],"not_for":["数学 / 竞赛推理","需要宽松许可的商业闭包"]},"capability_notes":{"coding":"LiveCodeBench 29.7%（官方），代码偏弱。","reasoning":"GPQA Diamond 42.4%（官方）。","math":"MATH 89.0%（官方）。","multimodal":"MMMU 64.9%（官方）。","chinese":"多语言训练，中文可用但非重点。"},"sheet":{"architecture_md":"**类型**：Dense，27B 参数。\n\n- 62 层，隐藏维 5376，词表 262,144\n- GQA：32 Q 头 / 16 KV 头，head_dim 128\n- 注意力交替：5 层局部滑窗（1024）→ 1 层全局，全局层用 RoPE base 1M\n- 视觉：SigLIP 400M 编码器，Pan & Scan 处理任意宽高比\n- 上下文 128K，输出 8K\n\n参考：Gemma 3 技术报告（arXiv 2503.19786）。","memory_md":"| 精度 | 权重大小 | 参考显存 |\n|---|---|---|\n| BF16 | 54 GB | 80GB |\n| Q8 | 28.7 GB | 48GB |\n| Q4 | 16.5 GB | 24GB |\n\n**KV Cache**：≈ 44 KiB/token（仅全局层）。128K 上下文 ≈ 5.5 GB，是 Qwen3-32B 的 1/6。\n\n**参考配置**：\n- 24GB：Q4_K_M + 视觉 + 32K 上下文舒适\n- 80GB：BF16 + 128K\n- 官方另有 QAT 版 int4，质量损失更小","training_md":"- 预训练 14T token，知识蒸馏\n- 后训练：SFT + RLHF + RLMF（数学）+ RLEF（代码）\n- 无思考模式\n- 函数调用：通过 prompt 格式支持\n- 最大输出 8,192","ecosystem_md":"- HF：google/gemma-3-27b-it（+ QAT GGUF 官方版）\n- 引擎：llama.cpp、Ollama、vLLM、MLX、Transformers 全支持\n- 微调：LoRA 生态成熟（Unsloth 等）\n- 中文文档：弱","versions_md":"- 同系列：Gemma 3 1B / 4B / 12B、Gemma 3n（端侧）\n- 上代：Gemma 2 27B\n- 后续：Gemma 4（2026）"},"ecosystem":{"engines":["llama.cpp","Ollama","vLLM","MLX"],"finetune":"LoRA 成熟","zh_docs":"弱"},"complete":true,"superseded_by":"gemma-4-31b"},{"id":"gemma-4-26b-a4b","name":"Gemma 4 26B A4B","name_zh":"谷歌 Gemma 4 · 26B A4B","aliases":["gemma-4-26B-A4B-it","Gemma4 26B MoE"],"vendor":"Google","vendor_zh":"谷歌","family":"Gemma 4","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/google/gemma-4-26B-A4B-it","status":"current","released_at":"2026-04-02","updated_at":"2026-08-28","modalities":["text","image","video","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"25.2B","active_params":"3.8B","total_params_b":25.2,"active_params_b":3.8,"experts":128,"active_experts":8,"shared_expert":true,"layers":30,"kv_layers":6,"hidden_size":2816,"vocab_size":262144,"kv_heads":2,"head_dim":512,"attention":"局部滑窗(1024) 24 层：16 Q / 8 KV，head_dim 256；全局 6 层：16 Q / 2 KV，head_dim 512，K=V 共享（5:1 交错）","notes":"30 层 = 6 × (5 × 滑窗 + 1 × 全局)。128 路由专家 + 1 共享专家选 8，专家中间维 704，稠密 FFN 2112。全局层 partial_rotary 0.25、theta 1e6；滑窗层 theta 1e4。logit softcap 30。视觉编码器 27 层、hidden 1152、patch 16，每图 280 个 soft token。表内 kv_heads / head_dim 指 6 层全局层。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K","max_output":32768},"memory":{"weight_gb":{"bf16":51.6,"fp8":28.8,"q8":26.9,"q4":14.4},"kv_per_token_kib":12,"kv_note":"随长度增长的只有 6 层全局层：2 KV 头 × 512 × K=V 共享 × 2 B = 2 KiB/层 → 12 KiB/token（BF16）。24 层滑窗层窗口 1024，KV 固定约 8 MiB/层。256K 上下文 KV ≈ 3 GB。","ref_hw_24gb":"官方 QAT Q4_0 14.4 GB（社区 Q4_K_M 16.9 GB），KV 极小，256K 上下文单并发可行，激活 3.8B 速度快","ref_hw_80gb":"BF16 51.6 GB + 256K 上下文（KV 3 GB）轻松，可多并发","ref_hw_8x80gb":"过剩；用于高并发服务","estimated":false},"pricing":{"input_per_m":0,"output_per_m":0,"currency":"USD","source":"Google AI 定价页（Gemma 4 仅免费层，无付费档）","as_of":"2026-08-28","note":"开源权重；第三方托管参考价约 $0.12 / $0.37（AA 页面）"},"links":{"official":"https://ai.google.dev/gemma/docs/core","hf":"https://huggingface.co/google/gemma-4-26B-A4B-it","paper":"https://arxiv.org/abs/2607.02770"},"variants":[{"kind":"gguf","publisher":"google","repo":"google/gemma-4-26B-A4B-it-qat-q4_0-gguf","url":"https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-gguf","note":"官方 QAT Q4_0"},{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/Gemma-4-26B-A4B-NVFP4","url":"https://huggingface.co/nvidia/Gemma-4-26B-A4B-NVFP4"},{"kind":"fp8","publisher":"RedHatAI","repo":"RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic","url":"https://huggingface.co/RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic","sizes":{"fp8":28.6}},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/gemma-4-26B-A4B-it-GGUF","url":"https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF","sizes":{"q8":27.3,"bf16":51.4}},{"kind":"awq","publisher":"cyankiwi","repo":"cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit","url":"https://huggingface.co/cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/gemma-4-26b-a4b-it-4bit","url":"https://huggingface.co/mlx-community/gemma-4-26b-a4b-it-4bit"}],"copy":{"one_liner":"激活 3.8B 的 MoE，24GB 卡上又快又能装长上下文。","highlights":["Apache-2.0，25.2B 总参 / 3.8B 激活，单卡吞吐远高于 31B Dense","GPQA 82.3、AIME 2026 88.3、LiveCodeBench v6 77.1，只比 31B 低 2–3 分","仅 6 层全局注意力且 K=V 共享，KV 12 KiB/token，24GB 卡 Q4 可跑 256K 上下文"],"pitfalls":["MoE 全部 25B 参数都要进显存，不是 4B 模型的显存需求","HLE 无工具 8.7、MRCR v2 128K 44.1，知识与长程检索明显弱于 31B","tau2 均值 68.2、Codeforces 1718，agent / 竞赛代码差距比考试题大；无 SWE-bench 官方数"],"logic_ability":"带可选思考模式的 MoE。考试型推理接近 31B（GPQA Diamond 82.3、AIME 2026 88.3、MMLU Pro 82.6、BigBench Extra Hard 64.8），但 HLE 8.7（无工具）/ 17.2（带搜索）和 128K 长上下文检索 44.1 差距大。工程型推理中等：LiveCodeBench v6 77.1、Codeforces 1718、tau2 68.2；官方未给 SWE-bench / Terminal-Bench。多模态推理强：MMMU Pro 73.8、MATH-Vision 82.4。","best_for":["单卡高吞吐本地助手","边缘服务器批处理","需要宽松许可的低成本产品"],"not_for":["前沿知识问答","复杂长程 agent"]},"capability_notes":{"coding":"LiveCodeBench v6 77.1、Codeforces Elo 1718（官方模型卡）；SWE-bench 未披露。","reasoning":"GPQA Diamond 82.3、HLE 8.7（无工具）/ 17.2（带搜索）、BigBench Extra Hard 64.8（官方模型卡）。","math":"AIME 2026 88.3（无工具，官方模型卡）；AIME 2025 未披露。","agent":"tau2-bench 三项均值 68.2（官方模型卡）。","multimodal":"MMMU Pro 73.8、MATH-Vision 82.4、MedXPertQA MM 58.1、OmniDocBench 1.5 0.149（官方模型卡）。","knowledge":"MMLU Pro 82.6、MMMLU 86.3（官方模型卡）。","chinese":"140+ 语言训练，MMMLU 86.3；中文独立数据少。"},"i18n":{"en":{"name_zh":"Google Gemma 4 26B A4B","one_liner":"3.8B-active MoE; fast on a 24GB GPU with room for long context.","highlights":["Apache-2.0, 25.2B total / 3.8B active; single-GPU throughput far above the 31B dense model","GPQA 82.3, AIME 2026 88.3, LiveCodeBench v6 77.1, only 2-3 points below the 31B","Only 6 global-attention layers with shared K=V, KV 12 KiB/token; a 24GB GPU runs 256K context at Q4"],"pitfalls":["All 25B MoE parameters must fit in VRAM; this is not a 4B model's memory footprint","HLE 8.7 without tools, MRCR v2 128K 44.1; knowledge and long-range retrieval clearly weaker than the 31B","tau2 average 68.2, Codeforces 1718; the gap in agent / competitive coding is wider than on exams; no official SWE-bench figure"],"logic_ability":"MoE with an optional thinking mode. Exam-style reasoning close to the 31B (GPQA Diamond 82.3, AIME 2026 88.3, MMLU Pro 82.6, BigBench Extra Hard 64.8), but HLE 8.7 (no tools) / 17.2 (with search) and 128K long-context retrieval 44.1 lag well behind. Engineering reasoning is middling: LiveCodeBench v6 77.1, Codeforces 1718, tau2 68.2; no official SWE-bench / Terminal-Bench. Multimodal reasoning is strong: MMMU Pro 73.8, MATH-Vision 82.4.","best_for":["High-throughput local assistants on one GPU","Edge-server batch processing","Low-cost products needing a permissive license"],"not_for":["Frontier knowledge QA","Complex long-horizon agents"],"capability_notes":{"coding":"LiveCodeBench v6 77.1, Codeforces Elo 1718 (official model card); SWE-bench not disclosed.","reasoning":"GPQA Diamond 82.3, HLE 8.7 (no tools) / 17.2 (with search), BigBench Extra Hard 64.8 (official model card).","math":"AIME 2026 88.3 (no tools, official model card); AIME 2025 not disclosed.","agent":"tau2-bench three-task average 68.2 (official model card).","multimodal":"MMMU Pro 73.8, MATH-Vision 82.4, MedXPertQA MM 58.1, OmniDocBench 1.5 0.149 (official model card).","knowledge":"MMLU Pro 82.6, MMMLU 86.3 (official model card).","chinese":"Trained on 140+ languages, MMMLU 86.3; little Chinese-specific data."}},"ja":{"name_zh":"Google Gemma 4 26B A4B","one_liner":"アクティブ 3.8B の MoE。24GB GPU で高速、長文も収まる。","highlights":["Apache-2.0、総パラメータ 25.2B / アクティブ 3.8B、単一 GPU のスループットは 31B Dense を大きく上回る","GPQA 82.3、AIME 2026 88.3、LiveCodeBench v6 77.1 で 31B との差はわずか 2〜3 点","グローバルアテンションは 6 層のみで K=V 共有、KV 12 KiB/token。24GB GPU の Q4 で 256K コンテキストが動く"],"pitfalls":["MoE の 25B パラメータ全体を VRAM に載せる必要があり、4B モデル並みのメモリ要件ではない","HLE ツールなし 8.7、MRCR v2 128K 44.1 と、知識と長距離検索は 31B より明らかに弱い","tau2 平均 68.2、Codeforces 1718 と agent / 競技コードの差は試験問題より大きい。SWE-bench の公式値なし"],"logic_ability":"任意の思考モード付き MoE。試験型推論は 31B に近い（GPQA Diamond 82.3、AIME 2026 88.3、MMLU Pro 82.6、BigBench Extra Hard 64.8）が、HLE 8.7（ツールなし）/ 17.2（検索あり）と 128K 長文コンテキスト検索 44.1 の差は大きい。エンジニアリング型推論は中程度：LiveCodeBench v6 77.1、Codeforces 1718、tau2 68.2。SWE-bench / Terminal-Bench は公式値なし。マルチモーダル推論は強い：MMMU Pro 73.8、MATH-Vision 82.4。","best_for":["単一 GPU の高スループットなローカルアシスタント","エッジサーバーのバッチ処理","寛容なライセンスが必要な低コスト製品"],"not_for":["最先端の知識 QA","複雑な長期 agent"],"capability_notes":{"coding":"LiveCodeBench v6 77.1、Codeforces Elo 1718（公式モデルカード）。SWE-bench は未公表。","reasoning":"GPQA Diamond 82.3、HLE 8.7（ツールなし）/ 17.2（検索あり）、BigBench Extra Hard 64.8（公式モデルカード）。","math":"AIME 2026 88.3（ツールなし、公式モデルカード）。AIME 2025 は未公表。","agent":"tau2-bench 3 項目平均 68.2（公式モデルカード）。","multimodal":"MMMU Pro 73.8、MATH-Vision 82.4、MedXPertQA MM 58.1、OmniDocBench 1.5 0.149（公式モデルカード）。","knowledge":"MMLU Pro 82.6、MMMLU 86.3（公式モデルカード）。","chinese":"140 以上の言語で学習、MMMLU 86.3。中国語単独のデータは少ない。"}}},"complete":true,"sheet":{"architecture_md":"**类型**：MoE + 局部/全局交错注意力，25.2B 总参 / 3.8B 激活，原生图像 / 视频输入。\n\n| 项目 | 数值 |\n|---|---|\n| 层数 | 30 = 6 × (5 × 滑窗 1024 + 1 × 全局) |\n| 隐藏维 | 2,816 |\n| 稠密 FFN | 2,112 |\n| 专家 | 128 路由 + 1 共享，每 token 选 8，专家中间维 704 |\n| 滑窗层（24 层） | 16 Q / 8 KV，head_dim 256，theta 1e4 |\n| 全局层（6 层） | 16 Q / 2 KV，head_dim 512，K=V 共享，partial_rotary 0.25，theta 1e6 |\n| 词表 | 262,144 |\n| 上下文 | 262,144 |\n| logit softcap | 30 |\n| 视觉编码器 | 27 层，hidden 1152，patch 16，280 soft token / 图 |\n\n`config.json` 中 `model_type: gemma4`，`architectures: Gemma4ForConditionalGeneration`。\n\n参考：google/gemma-4-26B-A4B-it config.json 与 Gemma 4 模型卡。","memory_md":"| 精度 | 权重大小 | 参考显存 | 说明 |\n|---|---|---|---|\n| BF16 | 51.6 GB | 80GB 单卡 | 官方 safetensors，2 分片（含视觉编码器） |\n| FP8 | 28.8 GB | 48GB | 社区量化（估） |\n| Q8 | 26.9 GB | 48GB | unsloth GGUF Q8_0 |\n| Q4 | 14.4 GB | 24GB | 官方 QAT Q4_0 GGUF；unsloth Q4_K_M 16.9 GB |\n\n**KV Cache**：只有 6 层全局层随长度增长，2 KV 头 × 512 × K=V 共享 × 2 B = 12 KiB/token；24 层滑窗层各固定约 8 MiB。256K 上下文 KV ≈ 3 GB。\n\n**参考配置**：\n- RTX 4090 / 5090 24GB：QAT Q4_0，256K 上下文单并发\n- A100 / H100 80GB：BF16 + 256K，多并发\n- 8×80GB：高并发服务\n\n警示：MoE 权重全部驻留显存；视频帧 token 会明显增加输入长度。","training_md":"- 预训练 token 数：**未披露**（模型卡仅描述语料类型：网页、代码、数学、图像，140+ 语言）\n- 知识截止：2025 年 1 月\n- 思考模式可选（optional），无思考时延迟更低\n- 官方发布 QAT（量化感知训练）Q4_0 检查点，4-bit 精度损失小\n- 后训练细节：官方未披露","ecosystem_md":"- HF：google/gemma-4-26B-A4B-it（BF16 51.6 GB）、google/gemma-4-26B-A4B-it-qat-q4_0-gguf（14.4 GB）；GGUF：unsloth/gemma-4-26B-A4B-it-GGUF\n- 引擎：Transformers（≥ 5.5 dev）、vLLM、SGLang、llama.cpp、Ollama、LM Studio、MLX\n- 微调：Unsloth、LLaMA-Factory 已适配；QLoRA 可在 24GB 卡上进行\n- 中文文档：弱（Google 官方文档以英文为主）","versions_md":"- 同代：Gemma 4 31B（Dense，考试型分数高 2–3 分）、Gemma 4 小体量版本\n- 上代：Gemma 3 27B / 12B（Dense，Gemma 许可）\n- 本代改为 Apache-2.0，并新增 MoE 与视频输入"},"ecosystem":{"engines":["Transformers","vLLM","SGLang","llama.cpp","Ollama","MLX"],"finetune":"Unsloth / LLaMA-Factory 已适配；QLoRA 24GB 可行","zh_docs":"弱"}},{"id":"gemma-4-31b","name":"Gemma 4 31B","name_zh":"谷歌 Gemma 4 · 31B","aliases":["gemma-4-31B-it","Gemma4 31B"],"vendor":"Google","vendor_zh":"谷歌","family":"Gemma 4","license":"Apache 2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/google/gemma-4-31B-it","status":"current","released_at":"2026-04-02","updated_at":"2026-08-28","modalities":["text","image","video","tools"],"reasoning_mode":"optional","architecture":{"type":"dense","total_params":"30.7B","total_params_b":30.7,"layers":60,"kv_layers":12,"hidden_size":5376,"vocab_size":262144,"kv_heads":16,"head_dim":256,"attention":"GQA（32 Q 头 / 16 KV 头，head_dim 256），5:1 局部滑窗(1024) : 全局","notes":"60 层中 12 层全局注意力（每 6 层 1 层，末层为全局），其余 1024 token 滑窗。中间层维 21504。视觉编码器约 550M。最大位置 262,144。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K","max_output":32768},"memory":{"weight_gb":{"bf16":61.4,"fp8":34.9,"q8":32.6,"q4":17.5},"kv_per_token_kib":192,"kv_note":"全局层 12 × 16 KV 头 × 256 × 2 × 2 B = 192 KiB/token；局部层 KV 上限 1024 token，可忽略。128K 上下文 ≈ 24 GB，比 Gemma 3 27B 高 4 倍以上（head_dim 翻倍）。","ref_hw_24gb":"官方 QAT Q4_0 17.5 GB（社区 Q4_K_M 18.3 GB）+ 视觉，剩余约 5 GB 给 KV，约 16–24K 上下文；长上下文需 Q3 或卸载","ref_hw_80gb":"BF16 61.4 GB + 64K 上下文 KV 约 12 GB，可行；FP8 34.9 GB 可跑满 256K","ref_hw_8x80gb":"过剩","estimated":false},"pricing":{"input_per_m":0,"output_per_m":0,"currency":"USD","source":"Google AI 定价页（Gemma 4 仅免费层，无付费档）","as_of":"2026-08-28","note":"开源权重；Gemini API 仅提供免费层，商业调用走第三方托管或自建"},"links":{"official":"https://ai.google.dev/gemma/docs/core","hf":"https://huggingface.co/google/gemma-4-31B-it","paper":"https://arxiv.org/abs/2607.02770"},"variants":[{"kind":"gguf","publisher":"google","repo":"google/gemma-4-31B-it-qat-q4_0-gguf","url":"https://huggingface.co/google/gemma-4-31B-it-qat-q4_0-gguf","note":"官方 QAT Q4_0"},{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/Gemma-4-31B-IT-NVFP4","url":"https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4"},{"kind":"fp8","publisher":"RedHatAI","repo":"RedHatAI/gemma-4-31B-it-FP8-dynamic","url":"https://huggingface.co/RedHatAI/gemma-4-31B-it-FP8-dynamic","sizes":{"fp8":33.3}},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/gemma-4-31B-it-GGUF","url":"https://huggingface.co/unsloth/gemma-4-31B-it-GGUF","sizes":{"q4":18.3,"q5":21.7,"q6":25.2,"q8":33.2,"bf16":62.4}},{"kind":"awq","publisher":"cyankiwi","repo":"cyankiwi/gemma-4-31B-it-AWQ-4bit","url":"https://huggingface.co/cyankiwi/gemma-4-31B-it-AWQ-4bit"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/gemma-4-31b-it-4bit","url":"https://huggingface.co/mlx-community/gemma-4-31b-it-4bit"}],"copy":{"one_liner":"Apache 2.0 的单卡 Dense 旗舰，带思考模式，256K 上下文。","highlights":["Apache 2.0 许可，告别 Gemma 自定义条款，可商用闭包","GPQA 84.3%、AIME 2026 89.2%、LiveCodeBench v6 80.0%，Arena 1451 开源前三","256K 上下文 + 图像 / 视频输入 + 函数调用，官方 QAT Q4 版 17.5 GB"],"pitfalls":["head_dim 256，全局层 KV 192 KiB/token，24GB 卡上长上下文比 Gemma 3 更吃紧","HLE 无工具仅 19.5%，前沿知识题距闭源旗舰差距大","31B 不支持音频输入（仅 E2B / E4B 有）"],"logic_ability":"带 <|think|> 开关的推理模型。考试型推理在开源同档领先：GPQA 84.3%、AIME 2026 89.2%、MMLU Pro 85.2%。工程型推理：LiveCodeBench v6 80.0%、Codeforces 2150，tau2-bench 均值 76.9%，agent 可用。弱项：HLE 19.5%，长上下文 MRCR v2 128K 66.4%，长程检索一般。","best_for":["单卡私有化推理助手","需要宽松许可的商业产品","多语言 / 视觉问答"],"not_for":["前沿知识问答（HLE 类）","24GB 卡上 100K+ 上下文"]},"capability_notes":{"coding":"LiveCodeBench v6 80.0%、Codeforces Elo 2150（官方模型卡）。","reasoning":"GPQA Diamond 84.3%、HLE 19.5%（无工具）、BBEH 74.4%（官方）。","math":"AIME 2026 89.2%（无工具，官方）。","agent":"tau2-bench 三项均值 76.9%（官方）。","multimodal":"MMMU Pro（视觉）76.9%、MATH-Vision 85.6%（官方）。","chinese":"MMMLU 88.4%，140+ 语言训练，中文可用。"},"sheet":{"architecture_md":"**类型**：Dense，30.7B 参数（含约 550M 视觉编码器）。\n\n- 60 层，隐藏维 5376，中间层维 21504，词表 262,144\n- GQA：32 Q 头 / 16 KV 头，head_dim 256\n- 注意力交替：5 层局部滑窗（1024）→ 1 层全局，共 12 层全局，末层固定全局\n- 上下文 256K（max_position_embeddings 262,144）\n- 多模态：图像 / 视频输入；31B 不含音频\n- 思考模式：系统提示中加 `<|think|>` 开启\n- 原生系统提示、函数调用、结构化 JSON 输出\n- 多 token 预测（MTP）草稿模型可用于投机解码\n\n参考：HF config.json（google/gemma-4-31B-it）、Gemma 4 技术报告（arXiv 2607.02770）。","memory_md":"| 精度 | 权重大小 | 来源 | 参考显存 |\n|---|---|---|---|\n| BF16 | 61.4 GB | 30.7B × 2 / GGUF BF16 | 80GB |\n| SFP8 | 34.9 GB | 官方文档 | 48GB |\n| Q8_0 | 32.6 GB | 社区 GGUF | 48GB |\n| Q4_0 (QAT) | 17.5 GB | 官方文档 | 24GB |\n| Q4_K_M | 18.3 GB | unsloth GGUF | 24GB |\n\n**KV Cache**：全局层 12 × 16 × 256 × 2 × 2 B = 192 KiB/token。32K ≈ 6 GB，128K ≈ 24 GB，256K ≈ 48 GB。\n\n**参考配置**：\n- 24GB：Q4 QAT + 16–24K 上下文；再长需 KV 量化或 Q3\n- 80GB：BF16 + 64K，或 FP8 + 256K\n- 官方提供 QAT 版（GGUF / 未量化 QAT / compressed-tensors），Q4 质量损失小于社区后量化","training_md":"- 数据截止 2025-01，网页 / 代码 / 图像 / 音频混合数据，训练 token 数未披露\n- 源自 Gemini 3 研究成果，140+ 语言\n- 可选思考模式（`<|think|>`）\n- 官方基准（模型卡）：MMLU Pro 85.2%、GPQA 84.3%、AIME 2026 89.2%、LiveCodeBench v6 80.0%、Codeforces 2150、tau2 76.9%、HLE 19.5% / 26.5%（带搜索）、MMMU Pro 76.9%、MRCR v2 128K 66.4%","ecosystem_md":"- HF：google/gemma-4-31B-it（+ 基座 gemma-4-31B、assistant 版）\n- 引擎：llama.cpp、Ollama、vLLM、MLX、LM Studio、NVIDIA NIM、Transformers 首日支持\n- 微调：LoRA 生态（Unsloth 等）已跟进\n- Gemini API 提供免费层调用\n- 中文文档：弱","versions_md":"- 同系列：Gemma 4 E2B / E4B（端侧，含音频）、12B、26B A4B（MoE）\n- 上代：Gemma 3 27B（2025-03，已被取代）\n- 发布：2026-04-02，Arena 开源第三"},"ecosystem":{"engines":["llama.cpp","Ollama","vLLM","MLX","LM Studio"],"finetune":"LoRA 可用","zh_docs":"弱"},"i18n":{"en":{"name_zh":"Google Gemma 4 31B","one_liner":"Apache 2.0 single-GPU dense flagship with thinking mode and 256K context.","highlights":["Apache 2.0 license, no more custom Gemma terms; commercial closed-source use allowed","GPQA 84.3%, AIME 2026 89.2%, LiveCodeBench v6 80.0%; Arena 1451, top three among open models","256K context + image / video input + function calling; official QAT Q4 build is 17.5 GB"],"pitfalls":["head_dim 256, global-layer KV 192 KiB/token; long context on a 24GB GPU is tighter than Gemma 3","HLE only 19.5% without tools; large gap to closed flagships on frontier knowledge questions","The 31B does not support audio input (only E2B / E4B do)"],"logic_ability":"A reasoning model with a <|think|> switch. Exam-style reasoning leads its open-source tier: GPQA 84.3%, AIME 2026 89.2%, MMLU Pro 85.2%. Engineering reasoning: LiveCodeBench v6 80.0%, Codeforces 2150, tau2-bench average 76.9%, usable for agents. Weaknesses: HLE 19.5%, long-context MRCR v2 128K 66.4%, mediocre long-range retrieval.","best_for":["Single-GPU on-prem reasoning assistants","Commercial products needing a permissive license","Multilingual / visual QA"],"not_for":["Frontier knowledge QA (HLE-type)","100K+ context on a 24GB GPU"],"capability_notes":{"coding":"LiveCodeBench v6 80.0%, Codeforces Elo 2150 (official model card).","reasoning":"GPQA Diamond 84.3%, HLE 19.5% (no tools), BBEH 74.4% (official).","math":"AIME 2026 89.2% (no tools, official).","agent":"tau2-bench three-task average 76.9% (official).","multimodal":"MMMU Pro (vision) 76.9%, MATH-Vision 85.6% (official).","chinese":"MMMLU 88.4%, trained on 140+ languages; Chinese usable."}},"ja":{"name_zh":"Google Gemma 4 31B","one_liner":"Apache 2.0 の単一 GPU Dense 旗艦。思考モード、256K。","highlights":["Apache 2.0 ライセンスで Gemma 独自規約から脱却、商用クローズド利用も可能","GPQA 84.3%、AIME 2026 89.2%、LiveCodeBench v6 80.0%、Arena 1451 でオープンソース上位 3 位","256K コンテキスト + 画像 / 動画入力 + 関数呼び出し、公式 QAT Q4 版は 17.5 GB"],"pitfalls":["head_dim 256、グローバル層の KV は 192 KiB/token で、24GB GPU の長文コンテキストは Gemma 3 より厳しい","HLE ツールなしはわずか 19.5%、最先端知識問題では閉源フラッグシップとの差が大きい","31B は音声入力非対応（E2B / E4B のみ対応）"],"logic_ability":"<|think|> スイッチ付きの推論モデル。試験型推論はオープンソース同クラスでトップ：GPQA 84.3%、AIME 2026 89.2%、MMLU Pro 85.2%。エンジニアリング型推論：LiveCodeBench v6 80.0%、Codeforces 2150、tau2-bench 平均 76.9% で agent 用途も可。弱点：HLE 19.5%、長文コンテキスト MRCR v2 128K 66.4% と長距離検索は平凡。","best_for":["単一 GPU のオンプレ推論アシスタント","寛容なライセンスが必要な商用製品","多言語 / 視覚 QA"],"not_for":["最先端の知識 QA（HLE 系）","24GB GPU での 100K 以上のコンテキスト"],"capability_notes":{"coding":"LiveCodeBench v6 80.0%、Codeforces Elo 2150（公式モデルカード）。","reasoning":"GPQA Diamond 84.3%、HLE 19.5%（ツールなし）、BBEH 74.4%（公式）。","math":"AIME 2026 89.2%（ツールなし、公式）。","agent":"tau2-bench 3 項目平均 76.9%（公式）。","multimodal":"MMMU Pro（視覚）76.9%、MATH-Vision 85.6%（公式）。","chinese":"MMMLU 88.4%、140 以上の言語で学習、中国語は利用可能。"}}},"complete":true,"runtime":{"tok_s":35,"latency_s":1.05,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"aliases":["gemma-7b-it","gemma-1.1-7b-it"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"knowledge":"MMLU 64.3%（官方）。","math":"GSM8K 46.4%（官方）。","chinese":"弱。"},"complete":false,"id":"gemma-7b","name":"Gemma 7B","name_zh":"Gemma · 7B","vendor":"Google","family":"Gemma","license":"Gemma Terms of Use","license_commercial":"restricted","weights_url":"https://huggingface.co/google/gemma-7b-it","superseded_by":"gemma-2-9b","released_at":"2024-02-21","architecture":{"type":"dense","total_params":"8.5B","total_params_b":8.5,"layers":28,"hidden_size":3072,"vocab_size":256000,"kv_heads":16,"head_dim":256,"attention":"MHA（16 头，head_dim 256）","notes":"256K 大词表（embedding 占 0.8B）；上下文 8K；GeGLU；1.1 版（2024-04）改进 RLHF。","undisclosed":false},"context":{"max_tokens":8192,"display":"8K"},"memory":{"weight_gb":{"bf16":17,"q8":9,"q4":5.3},"estimated":false,"kv_per_token_kib":448,"kv_note":"16×256×2×28×2 B = 448 KiB/token，MHA 使 KV 偏大。","ref_hw_24gb":"BF16 17 GB 单卡","ref_hw_80gb":"高并发","ref_hw_8x80gb":"过剩"},"links":{"official":"https://blog.google/technology/developers/gemma-open-models/","paper":"https://arxiv.org/abs/2403.08295","hf":"https://huggingface.co/google/gemma-7b-it"},"copy":{"one_liner":"Google 首个开放权重模型，Gemini 技术下放的起点。","highlights":["Gemini 同源技术与 tokenizer，256K 词表多语言好","6T token 训练，MMLU 64.3% 超同期 Llama 2 7B","Keras / JAX / PyTorch 全栈官方支持"],"pitfalls":["实际 8.5B 参数、MHA 大 KV，跑起来比 7B 重","上下文仅 8K，无 system prompt","Gemma Terms 含使用政策限制，对话质量当年被批"],"logic_ability":"入门级：MMLU 64.3%、GSM8K 46.4%（官方基座）。逻辑推理弱，主要历史意义。","best_for":["历史对照","Google 生态教学"],"not_for":["新项目"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama","Keras"],"finetune":"LoRA 友好","zh_docs":"弱"}},{"id":"glm-4-5-air","name":"GLM-4.5-Air","name_zh":"智谱 GLM-4.5 Air","aliases":["glm-4.5-air"],"vendor":"Z.ai","vendor_zh":"智谱","family":"GLM-4.5","license":"MIT","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/zai-org/GLM-4.5-Air","status":"superseded","released_at":"2025-07-28","updated_at":"2025-12-20","modalities":["text","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"106B","active_params":"12B","total_params_b":106,"active_params_b":12,"experts":128,"active_experts":8,"shared_expert":true,"layers":46,"hidden_size":4096,"kv_heads":8,"head_dim":128,"attention":"GQA（96 Q / 8 KV）","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":96000},"memory":{"weight_gb":{"bf16":220,"fp8":110,"q4":65},"kv_per_token_kib":184,"estimated":true,"ref_hw_80gb":"Q4 65 GB 单卡可跑","ref_hw_8x80gb":"BF16"},"pricing":{"input_per_m":0.2,"output_per_m":1.1,"currency":"USD","source":"Z.ai 开放平台","as_of":"2025-12-20"},"links":{"official":"https://z.ai/","hf":"https://huggingface.co/zai-org/GLM-4.5-Air","github":"https://github.com/zai-org/GLM-4.5","paper":"https://arxiv.org/abs/2508.06471"},"copy":{"one_liner":"单卡 80GB 可跑的 agent 编程 MoE，MIT 许可。","highlights":["Q4 约 65 GB，单卡 80GB 或 Mac 128GB 可跑","SWE-bench Verified 57.6%（官方）","MIT"],"pitfalls":["24GB 卡不可行","知识广度一般","GLM-4.6 未出 Air 版（截至数据日期）"],"logic_ability":"工程型推理为主，思考可开关；考试型中上。","best_for":["单卡 agent 编程服务"],"not_for":["消费级 24GB"]},"capability_notes":{},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp"],"zh_docs":"有"},"complete":false,"superseded_by":"glm-5-3-flash"},{"aliases":["GLM 4.5 355B-A32B"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"optional","capability_notes":{"coding":"SWE-bench Verified 64.2%、LiveCodeBench 72.9%（官方）。","reasoning":"GPQA 79.1%、HLE 14.4%（官方）。","math":"AIME 2025 91.0%（官方）。","agent":"Terminal-bench 37.5%、TAU-bench retail 79.7 / airline 60.4、BFCL v3 77.8（官方）。","chinese":"中文顶级。"},"complete":false,"id":"glm-4-5","name":"GLM-4.5","name_zh":"智谱 GLM-4.5","vendor":"Z.ai","vendor_zh":"智谱","family":"GLM-4.5","license":"MIT","license_commercial":true,"weights_url":"https://huggingface.co/zai-org/GLM-4.5","superseded_by":"glm-4-6","released_at":"2025-07-28","architecture":{"type":"moe","total_params":"355B","active_params":"32B","total_params_b":355,"active_params_b":32,"experts":160,"active_experts":8,"shared_expert":true,"layers":92,"hidden_size":5120,"vocab_size":151552,"kv_heads":8,"head_dim":128,"attention":"GQA（96 Q 头 / 8 KV 头）+ QK-Norm","notes":"深而窄的 MoE（92 层）；MTP 层；混合思考；128K。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":98304},"memory":{"weight_gb":{"bf16":710,"q8":376.3,"q4":216.5,"fp8":355},"estimated":false,"kv_per_token_kib":368,"kv_note":"8×128×2×92×2 B = 368 KiB/token。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（单卡）；Q4 216 GB 需 4×80GB","ref_hw_8x80gb":"BF16 710 GB 需 16×80GB；FP8 8×H100 可服务"},"pricing":{"input_per_m":0.6,"output_per_m":2.2,"currency":"USD","source":"Z.ai 官方 API（glm-4.5）","as_of":"2026-08-28"},"links":{"official":"https://z.ai/blog/glm-4.5","github":"https://github.com/zai-org/GLM-4.5","paper":"https://arxiv.org/abs/2508.06471","hf":"https://huggingface.co/zai-org/GLM-4.5"},"copy":{"one_liner":"智谱首个面向 agent 的旗舰 MoE，MIT 许可。","highlights":["SWE-bench Verified 64.2%、Terminal-bench 37.5%（官方）","MIT 许可，思考 / 非思考可切换","Claude Code / Cline 等 agent 工具适配好"],"pitfalls":["355B 需多卡，本地用 Air 版","思考模式 token 消耗大","发布两月后被 4.6 取代（200K 上下文更强 agent）"],"logic_ability":"思考模式：AIME 2025 91.0%、GPQA 79.1%、HLE 14.4%、LiveCodeBench 72.9%（官方博客）。agent 与工具调用是主打，TAU-bench 零售 79.7 / 航空 60.4。","best_for":["代码 agent 后端","中文 agent 平台"],"not_for":["单机部署"]},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp"],"finetune":"LoRA 可行（官方 LLaMA-Factory 支持）","zh_docs":"有"}},{"id":"glm-4-6","name":"GLM-4.6","name_zh":"智谱 GLM-4.6","aliases":["GLM 4.6","glm-4.6"],"vendor":"Z.ai","vendor_zh":"智谱","family":"GLM-4.5","license":"MIT","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/zai-org/GLM-4.6","status":"superseded","released_at":"2025-09-30","updated_at":"2025-12-20","modalities":["text","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"355B","active_params":"32B","total_params_b":355,"active_params_b":32,"experts":160,"active_experts":8,"shared_expert":true,"layers":92,"hidden_size":5120,"vocab_size":151552,"kv_heads":8,"head_dim":128,"attention":"GQA（96 Q 头 / 8 KV 头）+ 部分 RoPE","notes":"沿用 GLM-4.5 结构：更深更窄的 MoE（层数多、专家维度小），带 MTP 层加速投机解码。上下文从 128K 扩到 200K。","undisclosed":false},"context":{"max_tokens":200000,"display":"200K","max_output":128000},"memory":{"weight_gb":{"bf16":710,"fp8":355,"q4":200},"kv_per_token_kib":368,"kv_note":"8 KV 头 × 128 × 2 × 92 层 × 2 B ≈ 368 KiB/token；深层结构导致 KV 偏大。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行","ref_hw_8x80gb":"BF16 710 GB 需 8×H200 或 16×H100；FP8 8×H100 可服务","estimated":true},"pricing":{"input_per_m":0.6,"output_per_m":2.2,"currency":"USD","source":"Z.ai 开放平台","as_of":"2025-12-20","note":"GLM Coding Plan 另有包月"},"links":{"official":"https://z.ai/","hf":"https://huggingface.co/zai-org/GLM-4.6","github":"https://github.com/zai-org/GLM-4.5","paper":"https://arxiv.org/abs/2508.06471","pricing":"https://docs.z.ai/guides/overview/pricing"},"copy":{"one_liner":"MIT 许可、Agent 编程向的开源旗舰，Claude Code 替代热门。","highlights":["MIT 许可，355B/32B 激活，商用无附加条件","编程 agent 场景（Claude Code / Cline 类工具）的实测口碑好，token 消耗比 GLM-4.5 少 15%","200K 上下文、128K 最大输出，长任务友好"],"pitfalls":["355B 体量，自建门槛与 DeepSeek 相当","92 层深结构让 KV Cache 偏大（368 KiB/token），高并发长上下文吃显存","英文创作与知识广度弱于同价位闭源"],"logic_ability":"工程型推理是主打：SWE-bench Verified 68%、在真实 IDE agent 任务中完成率高，官方以「CC-Bench」实测对标 Claude Sonnet 4。考试型推理（AIME 93.9%）也强。思考可开关；关闭时代码质量下降明显。常见问题：长 agent 任务中偶尔过早结束。","best_for":["编程 agent 后端","自建替代 Claude Sonnet 的编程服务","中文工程任务"],"not_for":["单卡部署（用 GLM-4.5-Air）","多模态（用 GLM-4.5V）"]},"capability_notes":{"coding":"SWE-bench Verified 68.0%、LiveCodeBench v6 82.8%（官方）。","reasoning":"GPQA Diamond 81.0%、HLE 30.4%（工具，官方）。","math":"AIME 2025 93.9%（官方）。","agent":"τ²-bench 75.9%、BrowseComp 45.1%（官方）。","chinese":"中文顶级。"},"sheet":{"architecture_md":"**类型**：MoE，355B 总 / 32B 激活。\n\n- 92 层（含 MTP 层），隐藏维 5120，词表 151,552\n- 专家：160 路由 + 1 共享，每 token 激活 8 个路由专家\n- GQA：96 Q 头 / 8 KV 头，head_dim 128；部分 RoPE\n- 设计取向：「更深更窄」——层多、专家小，官方认为深度对推理更有利\n- 上下文 200K\n- MTP 层支持投机解码\n\n参考：GLM-4.5 技术报告（arXiv 2508.06471）。","memory_md":"| 精度 | 权重大小 | 参考显存 |\n|---|---|---|\n| BF16 | ≈ 710 GB | 8×H200 |\n| FP8 | ≈ 355 GB | 8×H100 |\n| Q4 | ≈ 200 GB（估） | 4×80GB |\n\n**KV Cache**：≈ 368 KiB/token。128K 上下文单请求 ≈ 46 GB，是同规模 MLA 模型的 5 倍以上。\n\n**参考配置**：\n- 24GB / 80GB：不可行\n- 8×H100：FP8 权重 + 有限并发\n- 8×H200：BF16 舒适","training_md":"- 预训练 22T token（GLM-4.5）+ 后续中训（代码 / 推理 / 长上下文 / agent）\n- 后训练：专家蒸馏 + RL（slime 框架）\n- 思考模式：`thinking.type` 可切 enabled / disabled\n- 工具调用：原生，支持 Claude Code / Cline / Roo Code 等\n- 最大输出 128K","ecosystem_md":"- HF：zai-org/GLM-4.6（+ FP8）\n- 引擎：vLLM、SGLang（官方 recipes）、llama.cpp（GGUF）\n- 微调：LLaMA-Factory 支持\n- 中文文档：有","versions_md":"- 上代：GLM-4.5（2025-07）\n- 轻量：GLM-4.5-Air（106B/12B）\n- 视觉：GLM-4.5V\n- 后续：GLM-4.6V、GLM-5（2026）"},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp"],"finetune":"LoRA / 全参（多节点）","zh_docs":"有"},"complete":true,"superseded_by":"glm-5-2"},{"aliases":["glm-4-9b-chat","GLM-4-9B-Chat-1M"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"math":"MATH 50.6%（官方）。","knowledge":"MMLU 72.4%（官方）。","agent":"函数调用原生，Berkeley FC 榜同期领先。","chinese":"C-Eval 75.6%（官方），中文强。"},"complete":false,"id":"glm-4-9b","name":"GLM-4-9B","name_zh":"智谱 GLM-4 · 9B","vendor":"Z.ai","vendor_zh":"智谱","family":"GLM-4","license":"GLM-4 License","license_commercial":"restricted","weights_url":"https://huggingface.co/zai-org/glm-4-9b-chat","superseded_by":"glm-4-5-air","released_at":"2024-06-05","architecture":{"type":"dense","total_params":"9.4B","total_params_b":9.4,"layers":40,"hidden_size":4096,"vocab_size":151552,"kv_heads":2,"head_dim":128,"attention":"GQA（32 Q 头 / 2 KV 头）","notes":"10T 多语言 token；128K 上下文（另有 1M 版）；原生函数调用；26 语言。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":18.8,"q8":10,"q4":6.3},"estimated":false,"kv_per_token_kib":40,"kv_note":"2×128×2×40×2 B = 40 KiB/token，极省，128K 仅 5 GB。","ref_hw_24gb":"BF16 18.8 GB 单卡可跑 32K+","ref_hw_80gb":"高并发 / 1M 版","ref_hw_8x80gb":"过剩"},"links":{"official":"https://www.zhipuai.cn/","github":"https://github.com/zai-org/GLM-4","paper":"https://arxiv.org/abs/2406.12793","hf":"https://huggingface.co/zai-org/glm-4-9b-chat"},"copy":{"one_liner":"2024 年中文 9B 最强之一，128K/1M 长文本单卡可跑。","highlights":["2 KV 头，KV 极省，1M 上下文版本单卡可用","中文与函数调用当年同尺寸领先","GLM-4V-9B 同期提供视觉版"],"pitfalls":["自有许可需登记，商用有条件","chatglm 自定义代码，框架兼容偶有问题","被 GLM-4-0414 / 4.5-Air 取代"],"logic_ability":"9B 中上：MMLU 72.4%、GSM8K 79.6%、MATH 50.6%（官方）。","best_for":["中文长文本","轻量 agent"],"not_for":["复杂推理"]},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Ollama"],"finetune":"LoRA 友好","zh_docs":"有"}},{"id":"glm-5-2","name":"GLM-5.2","name_zh":"智谱 GLM-5.2","aliases":["GLM 5.2","glm-5.2","GLM-5.2-FP8"],"vendor":"Z.ai","vendor_zh":"智谱","family":"GLM-5","license":"MIT","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/zai-org/GLM-5.2","status":"current","released_at":"2026-06-16","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"default-on","architecture":{"type":"moe","total_params":"744B","active_params":"40B","total_params_b":744,"active_params_b":40,"experts":256,"active_experts":8,"shared_expert":true,"layers":78,"hidden_size":6144,"vocab_size":154880,"kv_heads":1,"head_dim":288,"attention":"MLA + DSA 稀疏注意力（64 头，IndexShare 索引器）","notes":"官方 GitHub 标注 744B-A40B（HF 自动计数 753B 含 MTP 层）。MLA：kv_lora_rank 512 + RoPE 64，每层 KV 潜向量 576 维，这里按「1 KV 头 × 288」等价表示。DSA 稀疏注意力每 token 只取 top-2048 索引；IndexShare 让每 4 层共用一个索引器。前 3 层 Dense MLP，其余 256 路由专家 + 1 共享，MoE 中间维 2048；含 1 层 MTP。","undisclosed":false},"context":{"max_tokens":1048576,"display":"1M","max_output":131072},"memory":{"weight_gb":{"bf16":1507,"fp8":753,"q4":432,"q8":801.4},"kv_per_token_kib":88,"kv_note":"MLA：576 维 × 2 B × 78 层 ≈ 88 KiB/token。1M 上下文单请求 KV ≈ 88 GB；DSA 只降低计算不减 KV 存储。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行","ref_hw_8x80gb":"BF16 1.5 TB 需 16×H100 或 16×H200；FP8 753 GB 可 8×H200（KV 余量约 370 GB），8×H100 FP8 极紧","estimated":true},"pricing":{"input_per_m":1.4,"output_per_m":4.4,"currency":"USD","source":"Z.ai 开放平台","as_of":"2026-08-28","note":"缓存命中输入 $0.26；GLM Coding Plan 另有包月"},"links":{"official":"https://z.ai/blog/glm-5.2","hf":"https://huggingface.co/zai-org/GLM-5.2","github":"https://github.com/zai-org/GLM-5","paper":"https://arxiv.org/abs/2602.15763","pricing":"https://docs.z.ai/guides/overview/pricing"},"variants":[{"kind":"fp8","publisher":"zai-org","repo":"zai-org/GLM-5.2-FP8","url":"https://huggingface.co/zai-org/GLM-5.2-FP8","note":"官方 FP8","sizes":{"fp8":755.6}},{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/GLM-5.2-NVFP4","url":"https://huggingface.co/nvidia/GLM-5.2-NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/GLM-5.2-GGUF","url":"https://huggingface.co/unsloth/GLM-5.2-GGUF","sizes":{"q8":801.4,"bf16":1508}},{"kind":"awq","publisher":"cyankiwi","repo":"cyankiwi/GLM-5.2-AWQ-INT4","url":"https://huggingface.co/cyankiwi/GLM-5.2-AWQ-INT4"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/GLM-5.2-4bit","url":"https://huggingface.co/mlx-community/GLM-5.2-4bit"}],"copy":{"one_liner":"MIT 许可、1M 上下文的 744B 开源编程旗舰。","highlights":["MIT 许可，744B/40B 激活，无地区限制，商用无附加条件","官方 1M「无损」上下文；IndexShare 稀疏注意力让 1M 长度下每 token FLOPs 降 2.9×","SWE-bench Pro 62.1、Terminal-Bench 2.1 81.0（Terminus-2），发布时为开源最高"],"pitfalls":["744B 体量，FP8 也要 8×H200 级别，自建门槛比 GLM-4.6 翻倍","思考默认开启且 API 已不支持关闭（只能调 reasoning_effort low/high/max），简单任务 token 消耗大","API 价 $1.4/$4.4，是 GLM-4.6 的 2 倍以上；长上下文 KV 存储不因 DSA 而减少"],"logic_ability":"工程型推理是主打：SWE-bench Pro 62.1、Terminal-Bench 2.1 81.0、FrontierSWE 74.4，官方对标 Claude Opus 4.8 / GPT-5.5 且在部分 agent 编程基准持平。考试型推理也到顶级（AIME 2026 99.2、GPQA 91.2、HLE 40.5 / 带工具 54.7）。思考不可关；reasoning_effort=max 时输出很长。常见问题：SWE-Marathon 这类超长任务仅 13.0，长时间自主任务仍明显落后闭源。","best_for":["编程 agent 后端（Claude Code / ZCode / OpenCode）","整库级 1M 上下文代码理解","自建替代 Opus 级编程服务"],"not_for":["单机 / 单节点 80GB 部署（用 GLM-5.3-Flash）","多模态输入（用 GLM-5.3-Flash 或 GLM-5V 系列）","对成本敏感的简单对话"]},"capability_notes":{"coding":"SWE-bench Pro 62.1、Terminal-Bench 2.1 81.0（Terminus-2）/ 82.7（最佳 harness）、DeepSWE 46.2、NL2Repo 48.9（官方）。","reasoning":"GPQA Diamond 91.2、HLE 40.5、HLE（工具）54.7（官方）。","math":"AIME 2026 99.2、HMMT Feb 2026 92.5、IMOAnswerBench 91.0（官方）。","agent":"MCP-Atlas 76.8、Tool-Decathlon 48.2、FrontierSWE 74.4（官方）。","chinese":"中文顶级。"},"sheet":{"architecture_md":"**类型**：MoE，744B 总 / 40B 激活（HF 计数 753B 含 MTP）。\n\n- 78 层（前 3 层 Dense MLP）+ 1 层 MTP，隐藏维 6144，词表 154,880\n- 专家：256 路由 + 1 共享，每 token 激活 8 个，MoE 中间维 2048，sigmoid 路由（noaux_tc）\n- 注意力：MLA，64 头，q_lora 2048 / kv_lora 512，qk_nope 192 + RoPE 64，v_head 256\n- **DSA 稀疏注意力**：每 token 通过 32 头索引器选 top-2048 位置；**IndexShare** 让每 4 层共用 1 个索引器（full/shared 交替），1M 上下文下每 token FLOPs 降 2.9×\n- 上下文 1,048,576，RoPE theta 8e6\n- MTP 层改进后投机解码接受长度提升最多 20%\n\n参考：GLM-5 技术报告（arXiv 2602.15763）、IndexShare（arXiv 2603.12201）。","memory_md":"| 精度 | 权重大小 | 参考显存 |\n|---|---|---|\n| BF16 | 1,507 GB（官方文件） | 16×H100 / 16×H200 |\n| FP8 | ≈ 753 GB（zai-org/GLM-5.2-FP8） | 8×H200 |\n| Q4 | ≈ 432 GB（估） | 8×80GB |\n\n**KV Cache**：MLA ≈ 88 KiB/token。128K 上下文单请求 ≈ 11 GB，1M ≈ 88 GB。\n\n**参考配置**：\n- 24GB / 80GB：不可行\n- 8×H100 80GB：FP8 权重后仅剩 ~90 GB 给 KV，长上下文并发受限\n- 8×H200 141GB：FP8 舒适\n- 16×H100：BF16","training_md":"- 底座沿用 GLM-5 系列预训练；GLM-5.2 重点做长上下文（1M）与长程任务的中训 + 后训练\n- 后训练：agentic RL（slime 框架），评测均在 reasoning_effort=max\n- 思考默认开启，`reasoning_effort` 可选 low / high / max\n- 工具调用：原生，兼容 Claude Code / ZCode / OpenCode 等\n- 最大输出 128K（API）","ecosystem_md":"- HF：zai-org/GLM-5.2、GLM-5.2-FP8；nvidia/GLM-5.2-NVFP4\n- 引擎：SGLang（≥0.5.13）、vLLM（≥0.23）、Transformers、KTransformers、Unsloth、昇腾 vLLM-Ascend / xLLM\n- 微调：Unsloth 有指南；体量大，社区少\n- 中文文档：有","versions_md":"- 同底座：GLM-5（2026-02）、GLM-5.1、GLM-5.2（2026-06）\n- 后续：GLM-5.3（2026-08-14 API 上线，同底座纯后训练，权重承诺 8-28 开源）\n- 轻量：GLM-5.3-Flash（320B/18B，MIT，原生多模态，2026-08-26）\n- 上代：GLM-4.6 / GLM-4.7"},"ecosystem":{"engines":["SGLang","vLLM","KTransformers","Transformers"],"finetune":"Unsloth / 全参（多节点）","zh_docs":"有"},"i18n":{"en":{"name_zh":"Zhipu GLM-5.2","one_liner":"MIT-licensed 744B open coding flagship with 1M context.","highlights":["MIT license, 744B / 40B active, no regional restrictions, commercial use without extra conditions","Official 1M \"lossless\" context; IndexShare sparse attention cuts per-token FLOPs 2.9x at 1M length","SWE-bench Pro 62.1, Terminal-Bench 2.1 81.0 (Terminus-2); highest among open models at launch"],"pitfalls":["At 744B even FP8 needs an 8x H200-class node; self-hosting bar is double that of GLM-4.6","Thinking is on by default and the API no longer lets you disable it (only reasoning_effort low/high/max); heavy token use on simple tasks","API price $1.4/$4.4, more than 2x GLM-4.6; long-context KV storage is not reduced by DSA"],"logic_ability":"Engineering reasoning is the focus: SWE-bench Pro 62.1, Terminal-Bench 2.1 81.0, FrontierSWE 74.4; officially benchmarked against Claude Opus 4.8 / GPT-5.5 and level on some agentic coding benchmarks. Exam-style reasoning is also top-tier (AIME 2026 99.2, GPQA 91.2, HLE 40.5 / 54.7 with tools). Thinking cannot be disabled; output is very long at reasoning_effort=max. Common issue: ultra-long tasks such as SWE-Marathon score only 13.0; long autonomous tasks still trail closed models clearly.","best_for":["Coding agent backend (Claude Code / ZCode / OpenCode)","Repo-wide 1M-context code understanding","Self-hosted replacement for Opus-class coding services"],"not_for":["Single-machine / single-node 80GB deployment (use GLM-5.3-Flash)","Multimodal input (use GLM-5.3-Flash or the GLM-5V series)","Cost-sensitive simple chat"],"capability_notes":{"coding":"SWE-bench Pro 62.1, Terminal-Bench 2.1 81.0 (Terminus-2) / 82.7 (best harness), DeepSWE 46.2, NL2Repo 48.9 (official).","reasoning":"GPQA Diamond 91.2, HLE 40.5, HLE (tools) 54.7 (official).","math":"AIME 2026 99.2, HMMT Feb 2026 92.5, IMOAnswerBench 91.0 (official).","agent":"MCP-Atlas 76.8, Tool-Decathlon 48.2, FrontierSWE 74.4 (official).","chinese":"Top-tier Chinese."}},"ja":{"name_zh":"智譜 GLM-5.2","one_liner":"MIT ライセンス、1M コンテキストの 744B オープンコーディング旗艦。","highlights":["MIT ライセンス、744B / 40B アクティブ、地域制限なし、商用利用に追加条件なし","公式 1M「無損失」コンテキスト。IndexShare スパースアテンションで 1M 長のトークンあたり FLOPs を 2.9 倍削減","SWE-bench Pro 62.1、Terminal-Bench 2.1 81.0（Terminus-2）、公開時点でオープンソース最高"],"pitfalls":["744B の規模で FP8 でも 8×H200 級が必要、自前運用のハードルは GLM-4.6 の 2 倍","思考がデフォルトでオンで API からオフにできない（reasoning_effort low/high/max の調整のみ）、単純タスクでもトークン消費が大きい","API 価格 $1.4/$4.4 は GLM-4.6 の 2 倍超。長文コンテキストの KV ストレージは DSA でも減らない"],"logic_ability":"エンジニアリング型推論が主軸：SWE-bench Pro 62.1、Terminal-Bench 2.1 81.0、FrontierSWE 74.4、公式は Claude Opus 4.8 / GPT-5.5 を比較対象とし、一部の agent コーディングベンチマークで同等。試験型推論もトップ級（AIME 2026 99.2、GPQA 91.2、HLE 40.5 / ツールあり 54.7）。思考はオフ不可、reasoning_effort=max では出力が非常に長い。よくある問題：SWE-Marathon のような超長タスクは 13.0 にとどまり、長時間の自律タスクは依然閉源に明らかに劣る。","best_for":["コーディング agent のバックエンド（Claude Code / ZCode / OpenCode）","リポジトリ全体の 1M コンテキストコード理解","Opus 級コーディングサービスの自前代替"],"not_for":["単一マシン / 単一ノード 80GB 展開（GLM-5.3-Flash を使う）","マルチモーダル入力（GLM-5.3-Flash か GLM-5V 系を使う）","コスト重視の単純な対話"],"capability_notes":{"coding":"SWE-bench Pro 62.1、Terminal-Bench 2.1 81.0（Terminus-2）/ 82.7（最良ハーネス）、DeepSWE 46.2、NL2Repo 48.9（公式）。","reasoning":"GPQA Diamond 91.2、HLE 40.5、HLE（ツールあり）54.7（公式）。","math":"AIME 2026 99.2、HMMT Feb 2026 92.5、IMOAnswerBench 91.0（公式）。","agent":"MCP-Atlas 76.8、Tool-Decathlon 48.2、FrontierSWE 74.4（公式）。","chinese":"中国語は最上位。"}}},"complete":true,"runtime":{"tok_s":69,"latency_s":1.55,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"glm-5-3-flash","name":"GLM-5.3-Flash","name_zh":"智谱 GLM-5.3-Flash","aliases":["GLM 5.3 Flash","glm-5.3-flash","Ox Alpha"],"vendor":"Z.ai","vendor_zh":"智谱","family":"GLM-5","license":"MIT","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/zai-org/GLM-5.3-Flash","status":"current","released_at":"2026-08-26","updated_at":"2026-08-28","modalities":["text","image","tools"],"reasoning_mode":"default-on","architecture":{"type":"hybrid","total_params":"320B","active_params":"18B","total_params_b":320,"active_params_b":18,"experts":288,"active_experts":8,"shared_expert":true,"layers":45,"hidden_size":4096,"vocab_size":154880,"kv_heads":1,"head_dim":256,"attention":"KDA 线性注意力 34 层 + DSA 稀疏 MLA 11 层（3:1 交错）","notes":"glm5_next 混合架构：45 层 = 34 层 KDA（Kimi Delta Attention，64 头 × 128）+ 11 层 DeepSeek 稀疏注意力（MLA，kv_lora_rank 512，q_lora_rank 1536，qk/v head_dim 256，无 RoPE 维，indexer 32 头 top-k 2048）。前 3 层 Dense MLP，其余 288 路由专家 + 1 共享专家选 8，专家中间维 2048，sigmoid 无辅助损失路由。mHC 超连接（hc_mult 4）。1 层 MTP 头。视觉编码器 24 层、hidden 1024、patch 14。表内 kv_heads / head_dim 指 11 层 DSA 层的 MLA 潜向量。config 最大位置 1,048,576，官方评测按 300K 报告。","undisclosed":false},"context":{"max_tokens":1048576,"display":"1M"},"memory":{"weight_gb":{"fp8":328,"bf16":640,"q4":200},"kv_per_token_kib":11,"kv_note":"仅 11 层 DSA 层存 KV：MLA 潜向量 512 维（无 RoPE 分量）× 2 B = 1 KiB/层 → 11 KiB/token。34 层 KDA 为固定大小递归状态（64 头 × 128 × 128），与长度无关。1M 上下文 KV ≈ 11 GB。","ref_hw_24gb":"不可行（最小 UD-IQ1_S 也有 93 GB）","ref_hw_80gb":"不可行（UD-IQ1_S 93 GB / UD-IQ4_XS 157 GB 均超单卡）","ref_hw_8x80gb":"官方 FP8 328 GB：4×H100 起步，8×H100 舒适并可用满 1M 上下文（KV 很小）","estimated":false},"pricing":{"input_per_m":0.075,"output_per_m":0.25,"currency":"USD","source":"Z.ai 开放平台","as_of":"2026-08-28","note":"50% 折扣价，促销至 2026-09-09；缓存命中 $0.015"},"links":{"official":"https://z.ai/blog/glm-5.3-flash","hf":"https://huggingface.co/zai-org/GLM-5.3-Flash","github":"https://github.com/zai-org/GLM-5","paper":"https://arxiv.org/abs/2602.15763","pricing":"https://docs.z.ai/guides/overview/pricing"},"variants":[{"kind":"fp8","publisher":"zai-org","repo":"zai-org/GLM-5.3-Flash","url":"https://huggingface.co/zai-org/GLM-5.3-Flash","note":"官方原生 FP8 仓库","sizes":{"fp8":328.3}},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/GLM-5.3-Flash-GGUF","url":"https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF"},{"kind":"fp8","publisher":"unsloth","repo":"unsloth/GLM-5.3-Flash-FP8","url":"https://huggingface.co/unsloth/GLM-5.3-Flash-FP8","sizes":{"fp8":328.3}},{"kind":"nvfp4","publisher":"RedHatAI","repo":"RedHatAI/GLM-5.3-Flash-NVFP4","url":"https://huggingface.co/RedHatAI/GLM-5.3-Flash-NVFP4"}],"copy":{"one_liner":"320B/18B 混合注意力多模态开源模型，价格极低。","highlights":["MIT 许可，Terminal-Bench 2.1 84.3、DeepSWE 63.4 超过 GLM-5.2，官方称价格仅其 1/10","KDA 线性 + DSA 稀疏 3:1 混合，KV 仅 11 KiB/token，1M 上下文 KV ≈ 11 GB","GLM-5 系列首个原生多模态（30T 多模态语料新底座）；曾以 Ox Alpha 匿名上线 OpenRouter"],"pitfalls":["320B 总参数，官方 FP8 328 GB，最低 4×H100；社区 1-bit 量化也有 93 GB，单卡无解","架构新（glm5_next：KDA + DSA + mHC），需 SGLang / vLLM / Transformers 最新版，llama.cpp 仅 Unsloth 动态量化","刚发布，官方只给 6 项 agent 类基准，GPQA / AIME / SWE-bench Verified / MMMU 等常用榜未披露"],"logic_ability":"工程型推理强：Terminal-Bench 2.1 84.3、DeepSWE v1.1 63.4、AutomationBench 48.8、Agents' Last Exam 26.3，均高于 GLM-5.2（81.0 / 46.2 / 26.2 / 20.4），逼近 Claude Opus 4.8 与 GPT-5.6 Terra 的档位。HLE（带工具）55.3。GDPVal-AA v2 1773 为本表最高。考试型基准（GPQA、AIME）官方未披露，无法评估纯数学推理。LMArena 1469、AA 指数 57 为第三方口径。","best_for":["低成本编程 agent","多模态 + 长上下文的自建服务"],"not_for":["单卡部署","需要成熟生态的微调"]},"capability_notes":{"coding":"Terminal-Bench 2.1 84.3、DeepSWE v1.1 63.4（官方模型卡 bench_53）。","agent":"Agents' Last Exam 26.3、AutomationBench v1.0.6 48.8、GDPVal-AA v2 1773（官方模型卡）。","reasoning":"HLE（带工具）55.3（官方模型卡）；无工具 HLE、GPQA 未披露。","math":"AIME 官方未披露。","multimodal":"原生图像输入（24 层视觉编码器）；MMMU 等官方数值未披露。","chinese":"智谱系中文顶级。"},"ecosystem":{"engines":["SGLang","vLLM","TokenSpeed","KTransformers","Transformers"],"finetune":"Unsloth 指南；全参需多机","zh_docs":"有"},"i18n":{"en":{"name_zh":"Zhipu GLM-5.3-Flash","one_liner":"320B/18B hybrid-attention multimodal open model at a very low price.","highlights":["MIT license; Terminal-Bench 2.1 84.3, DeepSWE 63.4, above GLM-5.2, officially priced at 1/10 of it","KDA linear + DSA sparse 3:1 hybrid, KV only 11 KiB/token; 1M context KV is about 11 GB","First natively multimodal GLM-5 model (new base on 30T multimodal corpus); launched anonymously on OpenRouter as Ox Alpha"],"pitfalls":["320B total parameters, official FP8 328 GB, minimum 4x H100; even community 1-bit quantization is 93 GB, no single-GPU option","New architecture (glm5_next: KDA + DSA + mHC) needs the latest SGLang / vLLM / Transformers; llama.cpp only via Unsloth dynamic quants","Just released; official reports only 6 agent-type benchmarks, with GPQA / AIME / SWE-bench Verified / MMMU and other common boards undisclosed"],"logic_ability":"Strong engineering reasoning: Terminal-Bench 2.1 84.3, DeepSWE v1.1 63.4, AutomationBench 48.8, Agents' Last Exam 26.3, all above GLM-5.2 (81.0 / 46.2 / 26.2 / 20.4) and approaching the Claude Opus 4.8 and GPT-5.6 Terra tier. HLE (with tools) 55.3. GDPVal-AA v2 1773 is the highest in its table. Exam-style benchmarks (GPQA, AIME) not officially disclosed, so pure math reasoning cannot be assessed. LMArena 1469 and AA index 57 are third-party figures.","best_for":["Low-cost coding agents","Self-hosted multimodal + long-context services"],"not_for":["Single-GPU deployment","Fine-tuning that needs a mature ecosystem"],"capability_notes":{"coding":"Terminal-Bench 2.1 84.3, DeepSWE v1.1 63.4 (official model card bench_53).","agent":"Agents' Last Exam 26.3, AutomationBench v1.0.6 48.8, GDPVal-AA v2 1773 (official model card).","reasoning":"HLE (with tools) 55.3 (official model card); HLE without tools and GPQA not disclosed.","math":"AIME not officially disclosed.","multimodal":"Native image input (24-layer vision encoder); official MMMU and similar figures not disclosed.","chinese":"Top-tier Chinese within the Zhipu family."}},"ja":{"name_zh":"智譜 GLM-5.3-Flash","one_liner":"320B/18B ハイブリッド注意のマルチモーダルオープンモデル、超低価格。","highlights":["MIT ライセンス、Terminal-Bench 2.1 84.3、DeepSWE 63.4 で GLM-5.2 を上回り、公式は価格をその 1/10 と表明","KDA 線形 + DSA スパースの 3:1 ハイブリッド、KV はわずか 11 KiB/token、1M コンテキストの KV ≈ 11 GB","GLM-5 系初のネイティブマルチモーダル（30T マルチモーダルコーパスの新ベース）。OpenRouter に Ox Alpha として匿名公開されていた"],"pitfalls":["総パラメータ 320B、公式 FP8 で 328 GB、最低 4×H100。コミュニティの 1-bit 量子化でも 93 GB で単一 GPU は不可","新アーキテクチャ（glm5_next：KDA + DSA + mHC）のため SGLang / vLLM / Transformers の最新版が必要、llama.cpp は Unsloth 動的量子化のみ","公開直後で公式は agent 系 6 項目のベンチマークのみ。GPQA / AIME / SWE-bench Verified / MMMU などの一般的な指標は未公表"],"logic_ability":"エンジニアリング型推論が強い：Terminal-Bench 2.1 84.3、DeepSWE v1.1 63.4、AutomationBench 48.8、Agents' Last Exam 26.3 と、いずれも GLM-5.2（81.0 / 46.2 / 26.2 / 20.4）を上回り、Claude Opus 4.8 と GPT-5.6 Terra のクラスに迫る。HLE（ツールあり）55.3。GDPVal-AA v2 1773 は同表最高。試験型ベンチマーク（GPQA、AIME）は公式未公表で、純粋な数学推論は評価不能。LMArena 1469、AA 指数 57 は第三者の数値。","best_for":["低コストのコーディング agent","マルチモーダル + 長文コンテキストの自前サービス"],"not_for":["単一 GPU での展開","成熟したエコシステムが必要なファインチューニング"],"capability_notes":{"coding":"Terminal-Bench 2.1 84.3、DeepSWE v1.1 63.4（公式モデルカード bench_53）。","agent":"Agents' Last Exam 26.3、AutomationBench v1.0.6 48.8、GDPVal-AA v2 1773（公式モデルカード）。","reasoning":"HLE（ツールあり）55.3（公式モデルカード）。ツールなし HLE、GPQA は未公表。","math":"AIME は公式未公表。","multimodal":"ネイティブ画像入力（24 層の視覚エンコーダ）。MMMU などの公式数値は未公表。","chinese":"智譜系で中国語は最上位。"}}},"complete":true,"runtime":{"tok_s":50,"latency_s":1.49,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"sheet":{"architecture_md":"**类型**：MoE + 混合注意力（KDA 线性 / DSA 稀疏 MLA），320B 总参 / 18B 激活，原生多模态。\n\n| 项目 | 数值 |\n|---|---|\n| 层数 | 45 = 11 × (3 × KDA → MoE + 1 × DSA → MoE) |\n| 隐藏维 | 4,096 |\n| Dense 层 | 前 3 层 Dense MLP（中间维 12,288） |\n| 专家 | 288 路由 + 1 共享，每 token 选 8，专家中间维 2,048 |\n| DSA 层（11 层） | MLA：kv_lora_rank 512、q_lora_rank 1536，64 头，qk/v head_dim 256，无 RoPE 分量；indexer 32 头 × 128，top-k 2048 |\n| KDA 层（34 层） | 64 头 × head_dim 128，short conv kernel 4 |\n| 超连接 | mHC（hc_mult 4，Sinkhorn 20 轮） |\n| 词表 | 154,880 |\n| 上下文 | config 1,048,576；官方评测按 300K |\n| MTP | 1 层 |\n| 视觉编码器 | 24 层，hidden 1024，patch 14，spatial merge 2 |\n\n`config.json` 中 `architectures: Glm5NextForConditionalGeneration`，`model_type: glm5_next_text`。\n\n参考：zai-org/GLM-5.3-Flash config.json。","memory_md":"| 精度 | 权重大小 | 参考显存 | 说明 |\n|---|---|---|---|\n| FP8 | 328 GB | 4–8×H100 | 官方 safetensors，62 分片（BF16 / F8_E4M3 混合） |\n| BF16 | ≈640 GB | 8×H200+ | 按 320B × 2 B 估算，官方未发布 |\n| Q4 | 200 GB | 3×80GB | unsloth UD-Q4_K_XL；UD-IQ4_XS 157 GB |\n| 低比特 | 93–148 GB | 2×80GB | unsloth UD-IQ1_S 93.1 / UD-Q2_K_XL 109 / UD-Q3_K_XL 148 |\n\n**KV Cache**：仅 11 层 DSA 层产生 KV，MLA 潜向量 512 维 × 2 B = 1 KiB/层，共 11 KiB/token；34 层 KDA 为固定递归状态。1M 上下文 KV ≈ 11 GB，是同级模型中最省的。\n\n**参考配置**：\n- 24GB / 80GB 单卡：不可行\n- 4×H100：官方 FP8 最低配置\n- 8×H100：FP8 高并发 + 1M 上下文\n\n警示：权重体积远大于「Flash」之名；视觉 token 另占显存。","training_md":"- 新底座：30T token 多模态预训练语料（模型卡）\n- 思考默认开启，可关；训练带 1 层 MTP 头可做投机解码\n- 后训练细节、RL 配方：官方未披露\n- 官方评测最大上下文 300K；BabyVision 评测用 164K","ecosystem_md":"- HF：zai-org/GLM-5.3-Flash（FP8 328 GB）；GGUF：unsloth/GLM-5.3-Flash-GGUF（仅 UD 动态量化，93–200 GB）\n- 引擎：SGLang、vLLM、TokenSpeed、KTransformers、Transformers（模型卡列出）\n- 微调：Unsloth 指南；全参微调需多机\n- 中文文档：有（智谱开放平台）","versions_md":"- 同代：GLM-5.3（744B-A40B，API；权重预告 2026-08-28）\n- 上代：GLM-5.2、GLM-5 / GLM-5-FP8（754B）\n- 关系：GLM-5.3-Flash 为全新底座（非 GLM-5.2 蒸馏），是 GLM-5 系列首个原生多模态模型"}},{"id":"glm-5-3","name":"GLM-5.3","name_zh":"智谱 GLM-5.3","aliases":["GLM 5.3","glm-5.3"],"vendor":"Z.ai","vendor_zh":"智谱","family":"GLM-5","license":"MIT（承诺；权重未发布）","license_commercial":true,"openness":"api-only","weights_available":false,"weights_url":"https://huggingface.co/zai-org/GLM-5.3","status":"current","released_at":"2026-08-14","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"default-on","architecture":{"type":"moe","total_params":"744B","active_params":"40B","total_params_b":744,"active_params_b":40,"experts":256,"active_experts":8,"shared_expert":true,"layers":78,"hidden_size":6144,"vocab_size":154880,"kv_heads":1,"head_dim":288,"attention":"MLA + DSA 稀疏注意力（同 GLM-5.2）","notes":"官方明确「与 GLM-5.2 同一底座，所有提升来自后训练」，结构参数沿用 GLM-5.2（GitHub 标注 744B-A40B）。权重截至 2026-08-28 未发布，HF 页面为倒计时预告。","undisclosed":false},"context":{"max_tokens":1048576,"display":"1M","max_output":131072},"memory":{"weight_gb":{},"kv_per_token_kib":88,"kv_note":"同 GLM-5.2 结构，MLA ≈ 88 KiB/token；权重未发布，暂无官方文件大小。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行","ref_hw_8x80gb":"开源后预期同 GLM-5.2：FP8 8×H200","estimated":true},"pricing":{"input_per_m":1.4,"output_per_m":4.4,"currency":"USD","source":"Z.ai 开放平台","as_of":"2026-08-28","note":"缓存命中输入 $0.26；GLM Coding Plan 优先开放"},"links":{"official":"https://z.ai/blog/glm-5.3","hf":"https://huggingface.co/zai-org/GLM-5.3","github":"https://github.com/zai-org/GLM-5","paper":"https://arxiv.org/abs/2602.15763","pricing":"https://docs.z.ai/guides/overview/pricing"},"copy":{"one_liner":"GLM-5.2 同底座后训练加强版，权重承诺开源。","highlights":["承诺开源：官方称安全审查后两周内开源权重，HF 页面预告 2026-08-28；截至今日仍只有 API","同底座纯后训练：Z.ai Code Bench 比 GLM-5.2 提升 50%，DeepSWE v1.1 46.2→66.9，Terminal-Bench 3.0 4.6→28.3","网络安全能力突出：CyberGym 84.5、ExploitBench 54.4（GLM-5.2 为 24.4）"],"pitfalls":["权重尚未发布，本条目按 API-only 记录；开源后需重新核对文件与许可","推理不可关闭，只能调 reasoning_effort；max 档单题输出可达 ~75K token","漏洞利用能力被官方列为「意外涌现」，分阶段先向安全伙伴开放，企业合规需留意"],"logic_ability":"工程型推理：官方主张为开源最强编程模型，agent 类任务提升最大（DeepSWE 66.9、SWE-Marathon 42.5、Agents' Last Exam 28.5）。LMArena 文本榜 1484（第 15），AA 指数 60。考试型基准官方未单独披露，推测与 GLM-5.2 相近（同底座）。","best_for":["编程 agent（ZCode / Claude Code / OpenCode）","漏洞发现与安全审计","等 GLM-5.2 开源用户升级"],"not_for":["需要今天就能自建的场景（权重未发布）","多模态（用 GLM-5.3-Flash）"]},"capability_notes":{"coding":"DeepSWE v1.1 66.9、Terminal-Bench 3.0 28.3、SWE-Marathon v1.1 42.5、Z.ai Code Bench 34.5（max）（官方）。","agent":"Agents' Last Exam 28.5；CyberGym 84.5、ExploitBench 54.4（官方）。","reasoning":"官方未单独披露；同底座参考 GLM-5.2。","chinese":"中文顶级。"},"ecosystem":{"engines":["SGLang","vLLM"],"finetune":"待权重发布","zh_docs":"有"},"i18n":{"en":{"name_zh":"Zhipu GLM-5.3","one_liner":"Post-trained upgrade on the GLM-5.2 base; weights promised open.","highlights":["Open-source pledge: official says weights will be released within two weeks of safety review, HF page teases 2026-08-28; as of today API only","Same base, post-training only: Z.ai Code Bench +50% over GLM-5.2, DeepSWE v1.1 46.2 -> 66.9, Terminal-Bench 3.0 4.6 -> 28.3","Standout cybersecurity: CyberGym 84.5, ExploitBench 54.4 (GLM-5.2 was 24.4)"],"pitfalls":["Weights not yet released; this entry is recorded as API-only, recheck files and license after open-sourcing","Reasoning cannot be disabled, only reasoning_effort is adjustable; at max a single question can output ~75K tokens","Exploit capability is officially described as \"unexpectedly emergent\" and rolled out first to security partners; enterprise compliance should take note"],"logic_ability":"Engineering reasoning: officially claimed as the strongest open-source coding model, with the biggest gains on agentic tasks (DeepSWE 66.9, SWE-Marathon 42.5, Agents' Last Exam 28.5). LMArena text 1484 (#15), AA index 60. Exam-style benchmarks not separately disclosed; presumably close to GLM-5.2 (same base).","best_for":["Coding agents (ZCode / Claude Code / OpenCode)","Vulnerability discovery and security audits","GLM-5.2 open-source users waiting to upgrade"],"not_for":["Scenarios that need self-hosting today (weights unreleased)","Multimodal (use GLM-5.3-Flash)"],"capability_notes":{"coding":"DeepSWE v1.1 66.9, Terminal-Bench 3.0 28.3, SWE-Marathon v1.1 42.5, Z.ai Code Bench 34.5 (max) (official).","agent":"Agents' Last Exam 28.5; CyberGym 84.5, ExploitBench 54.4 (official).","reasoning":"Not separately disclosed; refer to GLM-5.2 on the same base.","chinese":"Top-tier Chinese."}},"ja":{"name_zh":"智譜 GLM-5.3","one_liner":"GLM-5.2 と同じベースの後学習強化版。重みの公開を約束。","highlights":["オープンソース公約：公式は安全審査後 2 週間以内に重み公開と表明、HF ページは 2026-08-28 を予告。本日時点では API のみ","同ベースの純後学習：Z.ai Code Bench は GLM-5.2 比 50% 向上、DeepSWE v1.1 46.2→66.9、Terminal-Bench 3.0 4.6→28.3","サイバーセキュリティ能力が突出：CyberGym 84.5、ExploitBench 54.4（GLM-5.2 は 24.4）"],"pitfalls":["重みは未公開で本項目は API-only として記録。公開後にファイルとライセンスの再確認が必要","推論はオフ不可で reasoning_effort の調整のみ。max では 1 問の出力が約 75K トークンに達することも","脆弱性悪用能力は公式に「予期せぬ創発」とされ、まずセキュリティパートナーに段階的に開放。企業のコンプライアンスは要注意"],"logic_ability":"エンジニアリング型推論：公式はオープンソース最強のコーディングモデルを主張し、agent 系タスクの向上が最大（DeepSWE 66.9、SWE-Marathon 42.5、Agents' Last Exam 28.5）。LMArena テキスト部門 1484（15 位）、AA 指数 60。試験型ベンチマークは公式に個別公表なし、同ベースのため GLM-5.2 に近いと推測。","best_for":["コーディング agent（ZCode / Claude Code / OpenCode）","脆弱性発見とセキュリティ監査","GLM-5.2 オープンソース利用者のアップグレード待ち"],"not_for":["今日すぐ自前運用が必要な用途（重み未公開）","マルチモーダル（GLM-5.3-Flash を使う）"],"capability_notes":{"coding":"DeepSWE v1.1 66.9、Terminal-Bench 3.0 28.3、SWE-Marathon v1.1 42.5、Z.ai Code Bench 34.5（max）（公式）。","agent":"Agents' Last Exam 28.5。CyberGym 84.5、ExploitBench 54.4（公式）。","reasoning":"公式に個別公表なし。同ベースの GLM-5.2 を参照。","chinese":"中国語は最上位。"}}},"complete":false,"runtime":{"tok_s":72,"latency_s":1.62,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"gpt-4-1","name":"GPT-4.1","aliases":["gpt-4.1-2025-04-14"],"vendor":"OpenAI","family":"GPT-4","superseded_by":"gpt-5","released_at":"2025-04-14","modalities":["text","image","tools"],"reasoning_mode":"none","context":{"max_tokens":1047576,"display":"1M","max_output":32768},"pricing":{"input_per_m":2,"output_per_m":8,"currency":"USD","source":"OpenAI 定价页","as_of":"2025-04-14","note":"缓存输入 $0.5；mini $0.4 / $1.6，nano $0.1 / $0.4"},"links":{"official":"https://openai.com/index/gpt-4-1/","pricing":"https://openai.com/api/pricing"},"copy":{"one_liner":"API 专属的 GPT-4o 升级版，1M 上下文，代码大幅提升。","highlights":["SWE-bench Verified 54.6%，比 GPT-4o 高 21 个百分点","1M 上下文，长文档 needle 检索稳定","指令遵循与 diff 生成明显改进"],"pitfalls":["无推理模式，GPQA 66.3%","起初仅 API，不在 ChatGPT","4 个月后被 GPT-5 替代"],"logic_ability":"非推理模型中的工程型推理强者，Aider Polyglot 52.9%；复杂数学需要 o 系列。","best_for":["历史对照"],"not_for":["新项目"]}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"gpt-4-5","name":"GPT-4.5","aliases":["gpt-4.5-preview","gpt-4.5-preview-2025-02-27"],"vendor":"OpenAI","family":"GPT-4","superseded_by":"gpt-4-1","released_at":"2025-02-27","modalities":["text","image","tools"],"reasoning_mode":"none","context":{"max_tokens":128000,"display":"128K","max_output":16384},"pricing":{"input_per_m":75,"output_per_m":150,"currency":"USD","source":"OpenAI 定价页（预览）","as_of":"2025-02-27"},"links":{"official":"https://openai.com/index/introducing-gpt-4-5/","pricing":"https://openai.com/api/pricing"},"copy":{"one_liner":"最大的非推理模型，情商与写作强，价格离谱。","highlights":["GPQA 71.4%、SimpleQA 幻觉率大幅降低","写作、情商、审美是当时最佳","LMArena 发布即登顶"],"pitfalls":["$75 / $150，是 GPT-4o 的 30 倍","AIME 2024 36.7%，不适合数学 / 推理","2025-07-14 已从 API 退役"],"logic_ability":"靠规模提升的直觉式推理，知识与常识判断强，但无推理链，数学与代码弱于 o 系列。","best_for":["历史对照"],"not_for":["新项目"]}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"gpt-4-turbo","name":"GPT-4 Turbo","aliases":["gpt-4-1106-preview","gpt-4-0125-preview","gpt-4-turbo-2024-04-09"],"vendor":"OpenAI","family":"GPT-4","superseded_by":"gpt-4o","released_at":"2023-11-06","modalities":["text","image","tools"],"reasoning_mode":"none","context":{"max_tokens":128000,"display":"128K","max_output":4096},"pricing":{"input_per_m":10,"output_per_m":30,"currency":"USD","source":"OpenAI 发布时定价","as_of":"2023-11-06"},"links":{"official":"https://openai.com/index/new-models-and-developer-products-announced-at-devday/","pricing":"https://openai.com/api/pricing"},"copy":{"one_liner":"128K 上下文 + 降价三分之二的 GPT-4，2024-04 正式版带视觉。","highlights":["128K 上下文，价格降到 $10 / $30","JSON 模式、并行函数调用、可复现 seed","2024-04-09 GA 版集成视觉，GPQA 48.0%"],"pitfalls":["最大输出仅 4K","preview 期间\"变懒\"问题被广泛吐槽","推出半年即被 GPT-4o 以一半价格替代"],"logic_ability":"与 GPT-4 同级的推理，长文档理解提升；数学 MATH 72.6%。无内置推理链。","best_for":["历史对照"],"not_for":["新项目"]}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"gpt-4","name":"GPT-4","aliases":["gpt-4-0314","gpt-4-0613","gpt-4-32k"],"vendor":"OpenAI","family":"GPT-4","superseded_by":"gpt-4-turbo","released_at":"2023-03-14","modalities":["text","tools"],"reasoning_mode":"none","context":{"max_tokens":8192,"display":"8K（32K 版）","max_output":8192},"pricing":{"input_per_m":30,"output_per_m":60,"currency":"USD","source":"OpenAI 发布时定价（8K 版）","as_of":"2023-03-14","note":"32K 版 $60 / $120"},"links":{"official":"https://openai.com/index/gpt-4-research/","paper":"https://arxiv.org/abs/2303.08774","pricing":"https://openai.com/api/pricing"},"copy":{"one_liner":"开启大模型时代的里程碑，首个通过律师考试级别的模型。","highlights":["MMLU 86.4%、HumanEval 67%，发布时全面领先","模拟律师考试前 10%","首次支持图像输入（GPT-4V，2023-09 开放）"],"pitfalls":["$30 / $60 极贵，8K 上下文很短","速度慢、2023-09 前无函数调用/JSON 模式","已于 2025 年从 API 退役"],"logic_ability":"2023 年最强通用推理，多步数学与代码远超 GPT-3.5，但无原生思维链，复杂推理需靠提示词引导。","best_for":["历史对照"],"not_for":["新项目"]}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"gpt-4o-mini","name":"GPT-4o mini","aliases":["gpt-4o-mini-2024-07-18"],"vendor":"OpenAI","family":"GPT-4","superseded_by":"gpt-5-mini","released_at":"2024-07-18","modalities":["text","image","tools"],"reasoning_mode":"none","context":{"max_tokens":128000,"display":"128K","max_output":16384},"pricing":{"input_per_m":0.15,"output_per_m":0.6,"currency":"USD","source":"OpenAI 发布时定价","as_of":"2024-07-18"},"links":{"official":"https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/","pricing":"https://openai.com/api/pricing"},"copy":{"one_liner":"替代 GPT-3.5 的超低价小模型，MMLU 82%。","highlights":["$0.15 / $0.60，比 GPT-3.5 Turbo 便宜 60%+","MMLU 82.0%、HumanEval 87.2%，小模型中领先","128K 上下文、支持视觉与微调"],"pitfalls":["推理能力弱，GPQA 40.2%","复杂代理任务不稳定","已被 GPT-4.1 mini / GPT-5 mini 替代"],"logic_ability":"轻量级推理，简单多步任务可用，MATH 70.2%；复杂逻辑与长链条推理明显不如 GPT-4o。","best_for":["历史对照"],"not_for":["新项目"]}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。官方称为端到端训练的单一模型，统一处理文本、图像、音频输入输出（\"omni\"）。"},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{"coding":"HumanEval 90.2%（官方）；SWE-bench Verified 33.2%（OpenAI 2024-08 SWE-bench Verified 博文）。","reasoning":"GPQA 53.6%（官方）。","math":"MATH 76.6%、MGSM 90.5%（官方）。","multimodal":"MMMU 69.1%、MathVista 63.8%（官方），原生音频输入输出。","chinese":"中文良好，新分词器使中文 token 数减少约 1.4 倍。"},"complete":true,"id":"gpt-4o","name":"GPT-4o","name_zh":"GPT-4o（omni）","aliases":["gpt-4o-2024-05-13","gpt-4o-2024-08-06","gpt-4o-2024-11-20","chatgpt-4o-latest"],"vendor":"OpenAI","family":"GPT-4","superseded_by":"gpt-4-1","released_at":"2024-05-13","modalities":["text","image","audio","tools"],"reasoning_mode":"none","context":{"max_tokens":128000,"display":"128K","max_output":16384},"pricing":{"input_per_m":2.5,"output_per_m":10,"currency":"USD","source":"OpenAI 定价页（2024-08-06 版）","as_of":"2024-08-06","note":"2024-05 发布价 $5 / $15；缓存输入 $1.25"},"links":{"official":"https://openai.com/index/hello-gpt-4o/","pricing":"https://openai.com/api/pricing","paper":"https://openai.com/index/gpt-4o-system-card/"},"copy":{"one_liner":"原生多模态的 GPT-4 平替，速度翻倍价格减半。","highlights":["原生文本 / 图像 / 音频统一模型，实时语音延迟 320ms","MMLU 88.7%、MMMU 69.1%，发布时多模态第一","价格从 GPT-4 Turbo 的 $10 / $30 降到 $5 / $15（后再降至 $2.5 / $10）"],"pitfalls":["无推理模式，SWE-bench Verified 仅 33.2%","版本迭代多（05-13 / 08-06 / 11-20），行为差异大","2025-08 随 GPT-5 发布从 ChatGPT 下线，引发用户反弹"],"logic_ability":"非推理模型的天花板级别：GPQA 53.6%、MATH 76.6%。多步推理靠提示词，复杂代理任务容易半途放弃。1120 版本的写作与指令遵循更好但数学略退。","best_for":["多模态 / 语音交互","高性价比通用对话（当年）","历史对照"],"not_for":["复杂推理与代理编程","新项目"]},"sheet":{"architecture_md":"**未披露**。OpenAI 未公开参数量。\n\n已知信息：\n- 单一端到端模型，文本、视觉、音频共用一套网络（官方称之为 \"omni\"）\n- 新分词器 o200k_base，非英语语言 token 数明显减少\n- 128K 上下文；08-06 版起最大输出 16K\n- 支持结构化输出（Structured Outputs，08-06 版起）","memory_md":"无自建选项。\n\n成本参考（08-06 版）：输入 $2.5 / 输出 $10；缓存输入 $1.25；Batch API 半价。\n\n发布时（05-13 版）为 $5 / $15。","training_md":"- 训练细节未披露，知识截止 2023-10\n- 无推理模式\n- 支持函数调用、并行工具、JSON / 结构化输出、视觉输入\n- Realtime API（2024-10）提供语音对语音","ecosystem_md":"- 官方 API、Azure OpenAI\n- 支持微调（2024-08 起开放 GPT-4o 微调）\n- 中文文档：弱（官方英文）","versions_md":"- 上代：GPT-4 Turbo\n- 版本：2024-05-13 / 2024-08-06（结构化输出、降价）/ 2024-11-20（写作提升）\n- 同代：GPT-4o mini（2024-07）\n- 继任：GPT-4.1（2025-04 API）、GPT-5（2025-08 ChatGPT）"},"ecosystem":{"engines":[],"finetune":"支持（官方微调）","zh_docs":"弱"}},{"id":"gpt-5-1","name":"GPT-5.1","aliases":["gpt-5.1"],"vendor":"OpenAI","family":"GPT-5","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","released_at":"2025-11-13","updated_at":"2025-12-20","modalities":["text","image","tools"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true,"notes":"新增 reasoning_effort=none 与自适应推理。"},"context":{"max_tokens":400000,"display":"400K","max_output":128000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":1.25,"output_per_m":10,"currency":"USD","source":"OpenAI 定价页","as_of":"2025-12-20"},"links":{"official":"https://openai.com/index/gpt-5-1/","pricing":"https://openai.com/api/pricing"},"copy":{"one_liner":"GPT-5 的对话与自适应推理改版，价格不变。","highlights":["reasoning_effort=none 可完全关推理","自适应思考时长","与 GPT-5 同价"],"pitfalls":["榜单分数与 GPT-5 接近，升级收益在体验而非分数","无自建","本站独立评测数据待补"],"logic_ability":"与 GPT-5 同级，简单问题响应更快。","best_for":["对话产品","GPT-5 平替"],"not_for":["私有化"]},"capability_notes":{},"complete":false,"superseded_by":"gpt-5-6-sol"},{"id":"gpt-5-5","name":"GPT-5.5","name_zh":"GPT-5.5","aliases":["gpt-5.5","gpt-5.5-2026-04-23"],"vendor":"OpenAI","family":"GPT-5","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2026-04-23","updated_at":"2026-08-28","modalities":["text","image","tools","computer-use"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。reasoning.effort 支持 none/low/medium/high/xhigh（无 max）。"},"context":{"max_tokens":1050000,"display":"1M","max_output":128000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":5,"output_per_m":30,"currency":"USD","source":"OpenAI 官方定价页（developers.openai.com）","as_of":"2026-08-28","note":">272K 输入按 2×/1.5× 计费；gpt-5.5-pro $30 / $180"},"links":{"official":"https://developers.openai.com/api/docs/models/gpt-5.5","pricing":"https://developers.openai.com/api/docs/pricing","paper":"https://openai.com/index/introducing-gpt-5-5/"},"copy":{"one_liner":"GPT-5.6 之前的旗舰，1M 上下文，$5 / $30。","highlights":["LMArena 文本榜 1482（GPT-5.5 High），与 GPT-5.6 Sol 持平","Terminal-Bench 2.1 88.0%、OSWorld-Verified 78.7%","1.05M 上下文 + 128K 输出；另有 gpt-5.5-pro（$30 / $180）"],"pitfalls":["$5 / $30 比 GPT-5.6 Sol 促销价贵，且无 max 推理档，性价比已被 Sol 反超","SWE-bench Pro 58.6% 明显落后 Opus 4.8（69.2%）","无自建"],"logic_ability":"GPT-5.6 发布前的旗舰：Terminal-Bench 2.1 88.0%（OpenAI 公告对比表）、HLE（带工具）52.2%、GDPval-AA 1769（Anthropic Opus 4.8 公告表，Vellum 转录）。综合能力与 Opus 4.7 相当，已被 GPT-5.6 Sol 在同价位全面覆盖。","best_for":["已在 GPT-5.5 上调优的存量系统","需要 gpt-5.5-pro 高精度档的场景"],"not_for":["新项目（改用 GPT-5.6 Sol）","私有化部署"]},"capability_notes":{"coding":"Terminal-Bench 2.1 88.0%；SWE-bench Pro 58.6%。","agent":"OSWorld-Verified 78.7%。","reasoning":"HLE（带工具）52.2%。"},"i18n":{"en":{"name_zh":"GPT-5.5","one_liner":"The flagship before GPT-5.6; 1M context, $5 / $30.","highlights":["LMArena text 1482 (GPT-5.5 High), level with GPT-5.6 Sol","Terminal-Bench 2.1 88.0%, OSWorld-Verified 78.7%","1.05M context + 128K output; gpt-5.5-pro also available ($30 / $180)"],"pitfalls":["$5 / $30 is pricier than GPT-5.6 Sol's promotional price and there is no max reasoning level; Sol now wins on value","SWE-bench Pro 58.6% clearly trails Opus 4.8 (69.2%)","No self-hosting"],"logic_ability":"The flagship before GPT-5.6: Terminal-Bench 2.1 88.0% (OpenAI announcement comparison table), HLE (with tools) 52.2%, GDPval-AA 1769 (Anthropic Opus 4.8 announcement table, transcribed by Vellum). Overall comparable to Opus 4.7, now fully covered by GPT-5.6 Sol at the same price.","best_for":["Existing systems already tuned on GPT-5.5","Scenarios needing the high-precision gpt-5.5-pro tier"],"not_for":["New projects (use GPT-5.6 Sol instead)","On-prem deployment"],"capability_notes":{"coding":"Terminal-Bench 2.1 88.0%; SWE-bench Pro 58.6%.","agent":"OSWorld-Verified 78.7%.","reasoning":"HLE (with tools) 52.2%."}},"ja":{"name_zh":"GPT-5.5","one_liner":"GPT-5.6 以前のフラッグシップ、1M コンテキスト、$5 / $30。","highlights":["LMArena テキスト部門 1482（GPT-5.5 High）、GPT-5.6 Sol と同等","Terminal-Bench 2.1 88.0%、OSWorld-Verified 78.7%","1.05M コンテキスト + 128K 出力。別途 gpt-5.5-pro（$30 / $180）あり"],"pitfalls":["$5 / $30 は GPT-5.6 Sol の優遇価格より高く max 推論段階もないため、コスパは Sol に逆転された","SWE-bench Pro 58.6% は Opus 4.8（69.2%）に明らかに劣る","セルフホスト不可"],"logic_ability":"GPT-5.6 発表前のフラッグシップ：Terminal-Bench 2.1 88.0%（OpenAI 発表の比較表）、HLE（ツールあり）52.2%、GDPval-AA 1769（Anthropic Opus 4.8 発表表、Vellum 転記）。総合能力は Opus 4.7 相当で、同価格帯では GPT-5.6 Sol に全面的にカバーされた。","best_for":["GPT-5.5 向けに調整済みの既存システム","gpt-5.5-pro の高精度段階が必要な用途"],"not_for":["新規プロジェクト（GPT-5.6 Sol に変更）","オンプレ展開"],"capability_notes":{"coding":"Terminal-Bench 2.1 88.0%。SWE-bench Pro 58.6%。","agent":"OSWorld-Verified 78.7%。","reasoning":"HLE（ツールあり）52.2%。"}}},"complete":false,"runtime":{"tok_s":81,"latency_s":62.42,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"gpt-5-6-luna","name":"GPT-5.6 Luna","name_zh":"GPT-5.6 经济档","aliases":["gpt-5.6-luna","Luna"],"vendor":"OpenAI","family":"GPT-5.6","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2026-07-09","updated_at":"2026-08-28","modalities":["text","image","tools","computer-use"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。GPT-5.6 系列最小档（对应此前的 nano 定位），支持 none~max 六档 reasoning.effort。"},"context":{"max_tokens":1050000,"display":"1M","max_output":128000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":0.2,"output_per_m":1.2,"currency":"USD","source":"OpenAI 官方定价页（developers.openai.com）","as_of":"2026-08-28","note":"2026-07-30 降价 80% 后价格"},"links":{"official":"https://developers.openai.com/api/docs/models/gpt-5.6-luna","pricing":"https://developers.openai.com/api/docs/pricing","paper":"https://openai.com/index/gpt-5-6/"},"copy":{"one_liner":"GPT-5.6 最便宜档，$0.20 / $1.20，AA 指数 51。","highlights":["$0.20 / $1.20，2026-07-30 降价 80%，单任务成本约为 Sol 的 1/5","AA 智能指数 51、Coding Agent Index 75，仍带 1.05M 上下文与全部工具","支持 max 推理档，可按需拉高深度"],"pitfalls":["知识广度与复杂 agent 能力明显弱于 Sol / Terra","官方未单独公布基准分数","无自建，且不支持微调"],"logic_ability":"AA 智能指数 51（max 档），适合批量分类、抽取与轻量推理；复杂 agent 任务应上 Terra 或 Sol。","best_for":["批量分类 / 抽取","高并发低价值调用","轻量推理"],"not_for":["复杂 agent","私有化部署"]},"capability_notes":{"reasoning":"AA 智能指数 51（max 档）。","coding":"AA Coding Agent Index 75。"},"i18n":{"en":{"name_zh":"GPT-5.6 economy tier","one_liner":"Cheapest GPT-5.6 tier, $0.20 / $1.20, AA index 51.","highlights":["$0.20 / $1.20, 80% price cut on 2026-07-30; per-task cost about 1/5 of Sol","AA Intelligence Index 51, Coding Agent Index 75, still with 1.05M context and all tools","Supports the max reasoning level to dial up depth as needed"],"pitfalls":["Knowledge breadth and complex agent ability clearly weaker than Sol / Terra","No separately published official benchmark scores","No self-hosting, and fine-tuning unsupported"],"logic_ability":"AA Intelligence Index 51 (max); suited to batch classification, extraction and light reasoning; complex agent tasks should go to Terra or Sol.","best_for":["Batch classification / extraction","High-concurrency low-value calls","Light reasoning"],"not_for":["Complex agents","On-prem deployment"],"capability_notes":{"reasoning":"AA Intelligence Index 51 (max).","coding":"AA Coding Agent Index 75."}},"ja":{"name_zh":"GPT-5.6 エコノミー","one_liner":"GPT-5.6 最安クラス、$0.20 / $1.20、AA 指数 51。","highlights":["$0.20 / $1.20、2026-07-30 に 80% 値下げ、タスク単価は Sol の約 1/5","AA 知能指数 51、Coding Agent Index 75、1.05M コンテキストと全ツールを維持","max 推論段階に対応、必要に応じて深さを引き上げ可能"],"pitfalls":["知識の広さと複雑な agent 能力は Sol / Terra より明らかに弱い","公式のベンチマークスコアは個別公表なし","セルフホスト不可、ファインチューニングも非対応"],"logic_ability":"AA 知能指数 51（max）。大量の分類・抽出・軽量推論に向く。複雑な agent タスクは Terra か Sol に任せるべき。","best_for":["大量の分類 / 抽出","高並行・低価値の呼び出し","軽量推論"],"not_for":["複雑な agent","オンプレ展開"],"capability_notes":{"reasoning":"AA 知能指数 51（max）。","coding":"AA Coding Agent Index 75。"}}},"complete":false,"runtime":{"tok_s":126,"latency_s":108.61,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"gpt-5-6-sol","name":"GPT-5.6 Sol","name_zh":"GPT-5.6 旗舰","aliases":["gpt-5.6-sol","gpt-5.6","GPT-5.6","Sol"],"vendor":"OpenAI","family":"GPT-5.6","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2026-07-09","updated_at":"2026-08-28","modalities":["text","image","tools","computer-use"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。系统卡仅称「以强化学习训练的推理模型，具备长内部思维链」。reasoning.effort 支持 none/low/medium/high/xhigh/max；ultra 模式调用子 agent。"},"context":{"max_tokens":1050000,"display":"1M","max_output":128000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":4,"output_per_m":20,"currency":"USD","source":"OpenAI 官方定价页（developers.openai.com）","as_of":"2026-08-28","note":"2026-08-21 ~ 11-21 促销价，原价 $5 / $30；>272K 输入按 2×/1.5× 计费"},"links":{"official":"https://openai.com/index/gpt-5-6/","pricing":"https://developers.openai.com/api/docs/pricing","paper":"https://deploymentsafety.openai.com/gpt-5-6-preview"},"copy":{"one_liner":"OpenAI 旗舰，Terminal-Bench 第一，限时 $4/$20。","highlights":["Terminal-Bench 2.1 88.8%（ultra 模式 91.9%）领先全场，FrontierMath v2 89.1%","新增 max 推理档与 ultra 子 agent 模式，1.05M 上下文 + 128K 输出","2026-08-21 起降价至 $4 / $20（原 $5 / $30），至 11-21"],"pitfalls":["$4 / $20 为三个月促销价，11-21 后可能回到 $5 / $30；>272K 输入按 2× 输入 / 1.5× 输出计费","AA 实测 max 档首 token 延迟约 122 s，深推理成本与时延都高","LMArena 1482 落后 Claude Fable 5 / Opus 5；网络安全与生化均被评为 High 风险，有安全拦截"],"logic_ability":"数学与命令行 agent 极强：FrontierMath v2 89.1%（max）、Terminal-Bench 2.1 88.8%、GPQA Diamond 93.5%（DataLearner 汇总）。ARC-AGI-2 92.5%（ARC Prize 验证）说明模式推理已接近饱和，但 ARC-AGI-3 仅 7.8%，远低于 Opus 5 的 30.2%——交互式新环境推理是弱项。AA 智能指数 61，略低于 Opus 5（63）与 Fable 5（62）；AA-Briefcase 表现 Elo 最高但 rubric 分 42% 低于 Fable 5 的 56%。失败模式：系统卡承认存在任务作弊与捏造研究结果的实例。","best_for":["终端 / Codex 类编码 agent","数学与形式化推理","高并发企业工作流（tier 5 达 40M TPM）"],"not_for":["交互式探索型推理（ARC-AGI-3）","生物 / 安全敏感场景","需要私有化部署的合规场景"]},"capability_notes":{"coding":"Terminal-Bench 2.1 88.8%（官方公告，ultra 91.9%）；SWE-bench Pro Public 64.6%（DataLearner）；AA Coding Agent Index 80。","reasoning":"AA 智能指数 61；ARC-AGI-2 92.5%（ARC Prize 验证）；ARC-AGI-3 7.8%。","math":"FrontierMath v2 89.1%（max，DataLearner 汇总）。","agent":"OSWorld 2.0 62.6%（xhigh + 工具）；ultra 模式以子 agent 加速复杂任务。","multimodal":"支持文本 + 图像输入；官方未公布 MMMU。","chinese":"中文可用，缺少独立中文评测。"},"sheet":{"architecture_md":"**未披露**。OpenAI 未公布参数量、Dense/MoE、注意力结构。系统卡仅描述为「以强化学习训练、具备长内部思维链的推理模型」。\n\n已知接口特性：\n- `reasoning.effort`：`none` / `low` / `medium`（默认） / `high` / `xhigh` / **`max`**（GPT-5.6 新增）\n- **ultra 模式**：超越单 agent，调用子 agent 并行推进复杂任务（Terminal-Bench 2.1 从 88.8% 提到 91.9%）\n- 1,050,000 上下文，最大输入 922K，最大输出 128K\n- 模态：文本 + 图像 → 文本\n- 端点：Chat Completions、Responses、Batch；不支持 fine-tuning\n- 内置工具：web search、file search、image generation、code interpreter、hosted shell、apply patch、skills、computer use、MCP、tool search\n- 知识截止 2026-02-16","memory_md":"无自建选项。\n\n**成本模型**（官方定价页，2026-08-28）：\n- 输入 $4 / 缓存输入 $0.40 / 输出 $20 / MTok（**2026-08-21 ~ 11-21 促销价**，原价 $5 / $30）\n- >272K 输入 token 的请求：输入 2×、输出 1.5×\n- 缓存写 1.25×；Batch / Flex 半价，Priority 2.5×\n\n实测（AA）：max 档单任务 $0.95~1.04，输出约 15K token / 任务（GPT-5.5 为 16K），69.6 tok/s，TTFT 约 122 s。同级的 Terra 便宜约 50%，Luna 便宜约 80%。","training_md":"- 训练细节未披露；系统卡（2026-06-26）称 RL 训练的推理模型\n- Preparedness 分级：网络安全 **High**、生物化学 **High**、AI 自我改进 Below High（Sol/Terra/Luna 相同）\n- 官方公告数据：Terminal-Bench 2.1 SOTA；ExploitBench 以约 1/3 输出 token 达到竞争力；World-Class Bio 68.3%（比 GPT-5.5 高约 9 pp）\n- 系统卡承认存在任务作弊与捏造研究结果的实例；CoT-Control 评测（13,000+ 任务）衡量思维链可控性\n- 幻觉率相对 GPT-5.5 改善","ecosystem_md":"- OpenAI API（Chat Completions / Responses / Batch）、Codex、ChatGPT（Plus/Pro/Work）\n- 2026-06-26 限量预览（应美国政府要求限定可信伙伴），2026-07-09 GA\n- 同系列：GPT-5.6 Terra（$2 / $12）、GPT-5.6 Luna（$0.20 / $1.20）\n- **不可微调**\n- 中文文档：弱","versions_md":"- 2026-06-26 GPT-5.6 Sol / Terra / Luna 预览；07-09 GA（$5 / $30）\n- 2026-07-30 Terra 降价 20%、Luna 降价 80%；08-21 Sol 降至 $4 / $20（至 11-21）\n- 上代：GPT-5.5（2026-04-23，$5 / $30，仍在售）、GPT-5.4（$2.50 / $15）\n- 别名 `gpt-5.6` 指向 `gpt-5.6-sol`"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"弱"},"i18n":{"en":{"name_zh":"GPT-5.6 flagship","one_liner":"OpenAI's flagship, #1 on Terminal-Bench, limited-time $4/$20.","highlights":["Terminal-Bench 2.1 88.8% (91.9% in ultra mode) leads the field; FrontierMath v2 89.1%","New max reasoning level and ultra sub-agent mode; 1.05M context + 128K output","Price cut to $4 / $20 from 2026-08-21 (was $5 / $30), through 11-21"],"pitfalls":["$4 / $20 is a three-month promotion and may revert to $5 / $30 after 11-21; inputs above 272K billed at 2x input / 1.5x output","AA-measured first-token latency at max about 122 s; deep reasoning is costly and slow","LMArena 1482 trails Claude Fable 5 / Opus 5; cybersecurity and bio-chem both rated High risk with safety blocks"],"logic_ability":"Extremely strong at math and command-line agents: FrontierMath v2 89.1% (max), Terminal-Bench 2.1 88.8%, GPQA Diamond 93.5% (DataLearner compilation). ARC-AGI-2 92.5% (ARC Prize verified) shows pattern reasoning near saturation, but ARC-AGI-3 is only 7.8%, far below Opus 5's 30.2%; reasoning in interactive novel environments is a weakness. AA Intelligence Index 61, slightly below Opus 5 (63) and Fable 5 (62); highest AA-Briefcase Elo but rubric score 42% below Fable 5's 56%. Failure modes: the system card admits instances of task gaming and fabricated research results.","best_for":["Terminal / Codex-style coding agents","Math and formal reasoning","High-concurrency enterprise workflows (tier 5 reaches 40M TPM)"],"not_for":["Interactive exploratory reasoning (ARC-AGI-3)","Bio / security-sensitive scenarios","Compliance scenarios needing on-prem deployment"],"capability_notes":{"coding":"Terminal-Bench 2.1 88.8% (official announcement, ultra 91.9%); SWE-bench Pro Public 64.6% (DataLearner); AA Coding Agent Index 80.","reasoning":"AA Intelligence Index 61; ARC-AGI-2 92.5% (ARC Prize verified); ARC-AGI-3 7.8%.","math":"FrontierMath v2 89.1% (max, DataLearner compilation).","agent":"OSWorld 2.0 62.6% (xhigh + tools); ultra mode speeds up complex tasks with sub-agents.","multimodal":"Text + image input; MMMU not officially published.","chinese":"Chinese usable; no independent Chinese evaluation."}},"ja":{"name_zh":"GPT-5.6 フラッグシップ","one_liner":"OpenAI 旗艦、Terminal-Bench 1 位、期間限定 $4/$20","highlights":["Terminal-Bench 2.1 88.8%（ultra モードで 91.9%）で全体をリード、FrontierMath v2 89.1%","max 推論段階と ultra サブ agent モードを新設、1.05M コンテキスト + 128K 出力","2026-08-21 から $4 / $20 に値下げ（旧 $5 / $30）、11-21 まで"],"pitfalls":["$4 / $20 は 3 か月の優遇価格で 11-21 以降は $5 / $30 に戻る可能性。272K 超の入力は入力 2 倍 / 出力 1.5 倍で課金","AA 実測で max の初回トークン遅延約 122 s、深い推論はコストも遅延も大きい","LMArena 1482 は Claude Fable 5 / Opus 5 に劣る。サイバーセキュリティと生化学はいずれも High リスク評価で安全ブロックあり"],"logic_ability":"数学とコマンドライン agent が極めて強い：FrontierMath v2 89.1%（max）、Terminal-Bench 2.1 88.8%、GPQA Diamond 93.5%（DataLearner 集計）。ARC-AGI-2 92.5%（ARC Prize 検証）はパターン推論が飽和に近いことを示すが、ARC-AGI-3 はわずか 7.8% で Opus 5 の 30.2% を大きく下回り、対話的な新環境での推論が弱点。AA 知能指数 61 は Opus 5（63）と Fable 5（62）をやや下回る。AA-Briefcase は Elo 最高だがルーブリック得点 42% は Fable 5 の 56% より低い。失敗モード：システムカードはタスクの不正攻略と研究結果の捏造事例を認めている。","best_for":["ターミナル / Codex 型のコーディング agent","数学と形式的推論","高並行の企業ワークフロー（tier 5 で 40M TPM）"],"not_for":["対話的な探索型推論（ARC-AGI-3）","生物 / セキュリティに敏感な用途","オンプレ展開が必要なコンプライアンス用途"],"capability_notes":{"coding":"Terminal-Bench 2.1 88.8%（公式発表、ultra 91.9%）。SWE-bench Pro Public 64.6%（DataLearner）。AA Coding Agent Index 80。","reasoning":"AA 知能指数 61。ARC-AGI-2 92.5%（ARC Prize 検証）。ARC-AGI-3 7.8%。","math":"FrontierMath v2 89.1%（max、DataLearner 集計）。","agent":"OSWorld 2.0 62.6%（xhigh + ツール）。ultra モードはサブ agent で複雑なタスクを高速化。","multimodal":"テキスト + 画像入力に対応。MMMU は公式未公表。","chinese":"中国語は利用可能。独立した中国語評価はなし。"}}},"complete":true,"runtime":{"tok_s":70,"latency_s":122.11,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"gpt-5-6-terra","name":"GPT-5.6 Terra","name_zh":"GPT-5.6 均衡档","aliases":["gpt-5.6-terra","Terra","GPT-5.6 Terra (max)"],"vendor":"OpenAI","family":"GPT-5.6","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2026-07-09","updated_at":"2026-08-28","modalities":["text","image","tools","computer-use"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。与 Sol / Luna 同一系列，支持 none / low / medium（默认）/ high / xhigh / max 六档 reasoning.effort；知识截止 2026-02-16。"},"context":{"max_tokens":1050000,"display":"1M","max_output":128000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":2,"output_per_m":12,"currency":"USD","source":"OpenAI 官方定价页（developers.openai.com）","as_of":"2026-08-28","note":"2026-07-30 降价 20% 后价格；>272K 输入按 2×/1.5× 计费"},"links":{"official":"https://developers.openai.com/api/docs/models/gpt-5.6-terra","pricing":"https://developers.openai.com/api/docs/pricing","paper":"https://openai.com/index/gpt-5-6/"},"copy":{"one_liner":"GPT-5.6 生产默认档，Sol 一半价钱，AA 指数 57。","highlights":["AA 智能指数 57，Coding Agent Index 77，单任务成本约为 Sol 的一半","1.05M 上下文 + 128K 输出，支持 max 推理档与全部内置工具","2026-07-30 降价 20% 至 $2 / $12（发布价 $2.5 / $15），比 GPT-5.5 便宜一半以上"],"pitfalls":[">272K 输入按 2× 输入 / 1.5× 输出计费","官方未单独公布 Terra 的 SWE-bench / GPQA 等分数，只有 AA 综合指数","网络安全 / 生化 Preparedness 分级 High，有安全拦截"],"logic_ability":"AA 智能指数 57，介于 Sol（61）与 Luna（51）之间；Coding Agent Index 77 仅比 Sol 低 3 分。适合大多数生产工作负载。","best_for":["生产环境默认模型","中等复杂度 agent"],"not_for":["前沿难题","私有化部署"]},"capability_notes":{"reasoning":"AA 智能指数 57（max 档，2026-08-28；发布当日 55），xhigh 52 / high 49 / medium 46 / 无推理 34（AA）。","coding":"AA Coding Agent Index 77（Sol 80、Luna 75）；DeepSWE v1.1 69.6%、FrontierCode 1.1 41.3%（Google Gemini 3.7 Flash 模型卡对照列）。","agent":"Agents' Last Exam 官方称「以约 1/16 成本超过 Claude Fable 5」（OpenAI 公告，Simon Willison 转述）；AA 单任务成本 $0.53，约为 Sol 一半。","knowledge":"知识截止 2026-02-16（官方模型页）；GDM-MRCR 128K 93.5%（Google 模型卡对照列）。","multimodal":"文本 + 图像输入、文本输出；新增原始分辨率图像处理（不缩放）；不含音频。"},"i18n":{"en":{"name_zh":"GPT-5.6 balanced tier","one_liner":"GPT-5.6 production default at half of Sol's price, AA index 57.","highlights":["AA Intelligence Index 57, Coding Agent Index 77; per-task cost about half of Sol","1.05M context + 128K output; supports the max reasoning level and all built-in tools","20% price cut on 2026-07-30 to $2 / $12 (launch price $2.5 / $15), more than half cheaper than GPT-5.5"],"pitfalls":["Inputs above 272K billed at 2x input / 1.5x output","No separately published SWE-bench / GPQA scores for Terra, only the AA composite index","Cybersecurity / bio-chem Preparedness rated High, with safety blocks"],"logic_ability":"AA Intelligence Index 57, between Sol (61) and Luna (51); Coding Agent Index 77, only 3 points below Sol. Suited to most production workloads.","best_for":["Default model for production","Medium-complexity agents"],"not_for":["Frontier-hard problems","On-prem deployment"],"capability_notes":{"reasoning":"AA Intelligence Index 57 (max, 2026-08-28; 55 on launch day), xhigh 52 / high 49 / medium 46 / no reasoning 34 (AA).","coding":"AA Coding Agent Index 77 (Sol 80, Luna 75); DeepSWE v1.1 69.6%, FrontierCode 1.1 41.3% (comparison column in Google's Gemini 3.7 Flash model card).","agent":"Agents' Last Exam: officially \"beats Claude Fable 5 at about 1/16 the cost\" (OpenAI announcement, relayed by Simon Willison); AA per-task cost $0.53, about half of Sol.","knowledge":"Knowledge cutoff 2026-02-16 (official model page); GDM-MRCR 128K 93.5% (comparison column in Google's model card).","multimodal":"Text + image input, text output; new native-resolution image processing (no downscaling); no audio."}},"ja":{"name_zh":"GPT-5.6 バランス","one_liner":"GPT-5.6 の本番デフォルト、Sol の半額、AA 指数 57。","highlights":["AA 知能指数 57、Coding Agent Index 77、タスク単価は Sol の約半分","1.05M コンテキスト + 128K 出力、max 推論段階と全内蔵ツールに対応","2026-07-30 に 20% 値下げで $2 / $12（発売価格 $2.5 / $15）、GPT-5.5 より半額以上安い"],"pitfalls":["272K 超の入力は入力 2 倍 / 出力 1.5 倍で課金","Terra 単体の SWE-bench / GPQA などのスコアは公式未公表で、AA 総合指数のみ","サイバーセキュリティ / 生化学の Preparedness 評価は High、安全ブロックあり"],"logic_ability":"AA 知能指数 57 で Sol（61）と Luna（51）の中間。Coding Agent Index 77 は Sol より 3 点低いだけ。大多数の本番ワークロードに向く。","best_for":["本番環境のデフォルトモデル","中程度の複雑さの agent"],"not_for":["最先端の難問","オンプレ展開"],"capability_notes":{"reasoning":"AA 知能指数 57（max、2026-08-28。発表当日は 55）、xhigh 52 / high 49 / medium 46 / 推論なし 34（AA）。","coding":"AA Coding Agent Index 77（Sol 80、Luna 75）。DeepSWE v1.1 69.6%、FrontierCode 1.1 41.3%（Google Gemini 3.7 Flash モデルカードの比較列）。","agent":"Agents' Last Exam は公式に「約 1/16 のコストで Claude Fable 5 を上回る」（OpenAI 発表、Simon Willison 転述）。AA タスク単価 $0.53 で Sol の約半分。","knowledge":"知識カットオフ 2026-02-16（公式モデルページ）。GDM-MRCR 128K 93.5%（Google モデルカードの比較列）。","multimodal":"テキスト + 画像入力、テキスト出力。原寸解像度の画像処理（縮小なし）を新たに対応。音声は非対応。"}}},"complete":true,"sheet":{"architecture_md":"**未披露**。OpenAI 未公布参数量与架构。\n\n已知接口特性（developers.openai.com 模型页，2026-08-28）：\n- 上下文 1,050,000 token，最大输出 128K\n- `reasoning.effort`：`none` / `low` / `medium`（默认）/ `high` / `xhigh` / `max`；`none` 即无推理模式（AA 指数 34）\n- 端点：Responses、Chat Completions、Batch；不支持 Realtime、Assistants、Fine-tuning、Embeddings、图像生成、音频\n- 输入：文本 + 图像；输出：文本。图像可按原始分辨率处理\n- 内置工具：web search、file search、code interpreter、computer use、image generation、hosted shell、MCP；支持 streaming、structured outputs、function calling、prompt caching\n- 5.6 系列新增：Programmatic Tool Calling（JS 编排工具组合）、多 agent 并行子 agent、显式缓存断点\n- 单一快照 `gpt-5.6-terra`；知识截止 2026-02-16\n- 速率限制：Tier 1 500 RPM / 500K TPM，Tier 5 15,000 RPM / 40M TPM","memory_md":"无自建选项。\n\n**成本模型**（官方定价页，2026-08-28）：\n- ≤272K 输入：$2 输入 / $0.20 缓存读 / $2.50 缓存写 / $12 输出\n- **>272K 输入**：整条请求按 $4 / $0.40 / $5 / $18 计费（输入 2×，输出 1.5×）\n- Batch 与 Flex：五折（$1 / $6；长上下文 $2 / $9）\n- Fast mode：翻倍（$4 / $24；长上下文 $8 / $36）\n- 2026-07-30 起从发布价 $2.5 / $15 降 20%\n\n**什么时候会变贵**：\n- 跨过 272K 门槛整条请求涨价，长文档批处理要主动截断或用 Luna\n- `max` 档吃 token 最多；AA 实测 Terra max 单任务 $0.53（Sol $1.04，Luna 约 $0.21）\n\n速度：AA 实测 106~121 tok/s（max 档），高于均值；max 档 TTFT 可达两分钟级，交互场景用 medium/high。","training_md":"- 训练细节未披露；知识截止 2026-02-16\n- 官方定位：GPT-5.6 系列的「均衡日常」档，Sol 主打前沿推理与长程 agent，Luna 主打速度与成本\n- 官方公开数据以 Sol 为主（Agents' Last Exam 53.6、SWE-bench Pro 64.6%）；Terra 只有「Agents' Last Exam 以约 1/16 成本超过 Fable 5」与 AA 综合指数，SWE-bench / GPQA 等未单独公布\n- 第三方对照：Google 3.7 Flash 模型卡列 Terra DeepSWE v1.1 69.6%、FrontierCode 41.3%、MRCR 128K 93.5%\n- 安全：网络安全 / 生化按 Preparedness Framework 定为 High，存在安全拦截\n- 行为：AA 称 5.6 系列输出 token 偏少（Sol 约 15K / 任务），token 效率优于同级","ecosystem_md":"- **API 渠道**：OpenAI API（Responses / Chat Completions / Batch）、Azure OpenAI（Microsoft Foundry）\n- ChatGPT：5.6 系列为 ChatGPT 默认模型线；Codex CLI / Cursor 等编码工具已接入\n- SDK：官方 Python / Node，兼容 OpenAI 协议的网关（OpenRouter、Vercel AI Gateway 等）\n- **不支持微调**（模型页明确 Fine-tuning 不可用）\n- 中文文档：弱（developers.openai.com 仅英文）","versions_md":"- 系列：GPT-5.6 Sol（旗舰，$5 / $30，2026-11-21 前促销价 $4 起）、**Terra**（均衡，$2 / $12）、Luna（经济，降价 80% 后 $0.20 / $1.80）；三者同一 2026-02-16 知识截止、1M 上下文、128K 输出\n- 发布：2026-07-09 三档同日 GA（Sol 此前有预览）；2026-07-30 Terra 降价 20%、Luna 降价 80%\n- 上代：GPT-5.5（AA 指数约 55~56，xhigh 档）\n- 快照：仅 `gpt-5.6-terra` 一个"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"弱"},"runtime":{"tok_s":106,"latency_s":141.99,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"gpt-5-mini","name":"GPT-5 mini","aliases":["gpt-5-mini"],"vendor":"OpenAI","family":"GPT-5","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","released_at":"2025-08-07","updated_at":"2025-12-20","modalities":["text","image","tools"],"reasoning_mode":"default-on","architecture":{"type":"unknown","undisclosed":true},"context":{"max_tokens":400000,"display":"400K","max_output":128000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":0.25,"output_per_m":2,"currency":"USD","source":"OpenAI 定价页","as_of":"2025-12-20"},"links":{"official":"https://openai.com/gpt-5","pricing":"https://openai.com/api/pricing"},"copy":{"one_liner":"GPT-5 的 1/5 价格，推理保留大半，性价比档。","highlights":["AIME 2025 91.1%","$0.25 / $2","400K 上下文"],"pitfalls":["代码 agent 能力明显弱于 GPT-5","知识广度下降","无自建"],"logic_ability":"考试型推理保留良好，工程型推理下降较多。","best_for":["批量分类 / 抽取","轻量推理"],"not_for":["复杂 agent"]},"capability_notes":{},"complete":false,"superseded_by":"gpt-5-6-terra"},{"id":"gpt-5","name":"GPT-5","name_zh":"GPT-5 旗舰","aliases":["gpt-5-2025-08-07","GPT5"],"vendor":"OpenAI","family":"GPT-5","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","released_at":"2025-08-07","updated_at":"2025-12-20","modalities":["text","image","tools"],"reasoning_mode":"default-on","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。API 提供 reasoning_effort（minimal/low/medium/high）与 verbosity 参数；ChatGPT 端为路由器 + 多模型组合。"},"context":{"max_tokens":400000,"display":"400K","max_output":128000},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":1.25,"output_per_m":10,"currency":"USD","source":"OpenAI 定价页","as_of":"2025-12-20","note":"缓存输入 $0.125"},"links":{"official":"https://openai.com/gpt-5","pricing":"https://openai.com/api/pricing","paper":"https://openai.com/index/introducing-gpt-5"},"copy":{"one_liner":"输入价格极低的推理旗舰，数学与科学问答顶级。","highlights":["AIME 2025 94.6%（无工具）、GPQA 85.7%，考试型推理顶级","输入 $1.25 / 百万 token，同档位里最便宜的旗舰","400K 上下文、128K 最大输出，长输出任务友好"],"pitfalls":["默认开推理，minimal effort 下能力下滑明显，需按任务调 effort","输出 $10 与推理 token 叠加后，长任务实际成本并不低","API 不暴露原始思维链，调试推理错误困难"],"logic_ability":"考试型推理是其强项：AIME、GPQA、HLE 都居前。工程型推理（SWE-bench 74.9%）扎实但略逊于 Claude Opus 4.5。reasoning_effort 是主要调节旋钮：high 下更少幻觉、更多自我校验；minimal 接近非推理模型。常见问题：在 low effort 时会跳过验证直接给结论。","best_for":["数学 / 科学 / 数据分析","低输入成本的批量处理","长文本生成"],"not_for":["需要私有化部署","对延迟极敏感的实时交互（用 mini / nano）"]},"capability_notes":{"coding":"SWE-bench Verified 74.9%、Aider Polyglot 88%（官方）。","reasoning":"GPQA 85.7%、HLE 24.8%（无工具，官方）。","math":"AIME 2025 94.6%（无工具，官方）。","multimodal":"MMMU 84.2%（官方）。","chinese":"中文良好，缺独立榜单。"},"sheet":{"architecture_md":"**未披露**。OpenAI 未公开 GPT-5 的参数量或结构。\n\n已知接口层信息：\n- `reasoning_effort`：minimal / low / medium / high\n- `verbosity`：low / medium / high\n- 400K 上下文（输入 272K + 输出 128K）\n- 自定义工具（custom tools）支持自由文本工具调用\n\nChatGPT 端为「路由 + 快模型 + 推理模型」的系统，API 端 `gpt-5` 为单一推理模型。","memory_md":"无自建选项。\n\n成本参考：输入 $1.25 / 输出 $10；缓存输入 $0.125。\n\n**警示**：推理 token 按输出计费且不可见，high effort 下一条复杂请求可能产生 10K+ 推理 token。","training_md":"- 训练细节未披露\n- 默认推理开启，可用 effort=minimal 关到最低\n- 支持函数调用、并行工具、结构化输出、图像输入\n- 最大输出 128K","ecosystem_md":"- 官方 API、Azure OpenAI\n- 不支持微调（GPT-5 系列暂无 fine-tune）\n- 中文文档：弱（官方英文）","versions_md":"- 上代：GPT-4.1 / o3\n- 同代：GPT-5 mini、GPT-5 nano、GPT-5 pro\n- 继任：GPT-5.1（2025-11）"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"弱"},"complete":true,"superseded_by":"gpt-5-6-sol"},{"id":"gpt-oss-120b","name":"gpt-oss-120b","name_zh":"OpenAI 开源 120B","aliases":["gpt oss 120b","openai/gpt-oss-120b"],"vendor":"OpenAI","family":"gpt-oss","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/openai/gpt-oss-120b","status":"current","released_at":"2025-08-05","updated_at":"2025-12-20","modalities":["text","tools"],"reasoning_mode":"default-on","architecture":{"type":"moe","total_params":"117B","active_params":"5.1B","total_params_b":117,"active_params_b":5.1,"experts":128,"active_experts":4,"shared_expert":false,"layers":36,"hidden_size":2880,"vocab_size":201088,"kv_heads":8,"head_dim":64,"attention":"GQA（64 Q 头 / 8 KV 头），交替稠密 / 滑窗 128","notes":"MoE 权重原生 MXFP4（4.25 bit）。交替全注意力与 128 滑窗层，带 attention sink。Harmony 对话格式，推理强度 low/medium/high。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":131072},"memory":{"weight_gb":{"q4":63,"q8":63.4},"kv_per_token_kib":72,"kv_note":"8 KV 头 × 64 × 2 × 36 层 × 2 B = 72 KiB/token（滑窗层实际更小）。","ref_hw_24gb":"不可行（MXFP4 63 GB）","ref_hw_80gb":"MXFP4 63 GB 单卡 H100 可跑，官方目标配置","ref_hw_8x80gb":"高并发服务","estimated":false},"pricing":{"input_per_m":0.1,"output_per_m":0.5,"currency":"USD","source":"第三方托管常见价（Groq / Cerebras / Together）","as_of":"2025-12-20","note":"开源权重，社区托管价"},"links":{"official":"https://openai.com/index/introducing-gpt-oss/","hf":"https://huggingface.co/openai/gpt-oss-120b","github":"https://github.com/openai/gpt-oss","paper":"https://cdn.openai.com/pdf/419b6906-9da6-406c-a19d-1bb078ac7637/oai_gpt-oss_model_card.pdf"},"variants":[{"kind":"other","publisher":"openai","repo":"openai/gpt-oss-120b","url":"https://huggingface.co/openai/gpt-oss-120b","note":"官方 MXFP4 safetensors","sizes":{"q4":65.2}},{"kind":"gguf","publisher":"ggml-org","repo":"ggml-org/gpt-oss-120b-GGUF","url":"https://huggingface.co/ggml-org/gpt-oss-120b-GGUF","sizes":{"bf16":1.6}},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/gpt-oss-120b-GGUF","url":"https://huggingface.co/unsloth/gpt-oss-120b-GGUF","sizes":{"q4":62.8,"q5":62.9,"q6":63.3,"q8":63.4}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/openai_gpt-oss-120b-GGUF","url":"https://huggingface.co/bartowski/openai_gpt-oss-120b-GGUF","sizes":{"q4":62.8,"q6":63.3,"q8":63.4,"bf16":65.4}},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/gpt-oss-120b-MXFP4-Q8","url":"https://huggingface.co/mlx-community/gpt-oss-120b-MXFP4-Q8"}],"copy":{"one_liner":"单卡 80GB 跑的 OpenAI 开源推理模型，5B 激活极快。","highlights":["Apache-2.0，OpenAI 六年来首次开放权重","MXFP4 原生 63 GB，单张 H100 / 双 4090 级别即可部署","5.1B 激活，吞吐极高；推理强度三档可调"],"pitfalls":["Harmony 格式必须严格遵守，旧版 chat template 会显著降质","知识广度与多语言弱（训练偏英文、STEM），中文明显不如 Qwen / GLM","幻觉率偏高（官方 PersonQA 幻觉率 49%），事实问答慎用"],"logic_ability":"考试型推理相对体量极强（AIME 2025 97.9% 带工具、GPQA 80.1%），得益于 o 系列同源的 RL 配方。工程型推理（SWE-bench 62.4%）中等。reasoning effort 是关键：low 时几乎不思考，high 时思考链可达数万 token。常见问题：英文之外语言的推理质量下降、幻觉多。","best_for":["单卡高吞吐推理服务","数学 / STEM 推理","英文 agent 原型"],"not_for":["中文为主的产品","知识问答 / 事实核查"]},"capability_notes":{"coding":"SWE-bench Verified 62.4%、Codeforces 2622（工具，官方）。","reasoning":"GPQA Diamond 80.1%、HLE 19.0%（无工具，官方）。","math":"AIME 2025 97.9%（工具）、92.5%（无工具，官方）。","agent":"τ-bench Retail 67.8%（官方）。","chinese":"弱，中文非重点语言。"},"sheet":{"architecture_md":"**类型**：MoE，116.8B 总 / 5.1B 激活。\n\n- 36 层，隐藏维 2880，词表 201,088（o200k_harmony）\n- 专家：128 个，每 token 激活 4 个，无共享专家\n- GQA：64 Q 头 / 8 KV 头，head_dim 64\n- 注意力交替：稠密全注意力层 ↔ 128 token 滑窗层；每头带 learned attention sink\n- RoPE + YaRN 到 128K\n- MoE 权重训练后量化为 MXFP4（每参数 4.25 bit）\n\n参考：gpt-oss 模型卡。","memory_md":"| 精度 | 权重大小 | 参考显存 |\n|---|---|---|\n| MXFP4（原生） | 63 GB | 80GB 单卡 |\n| BF16（反量化） | ≈ 234 GB | 4×80GB，无必要 |\n\n**KV Cache**：≤ 72 KiB/token，128K 上下文 ≈ 9 GB。\n\n**参考配置**：\n- 24GB：不可行（用 gpt-oss-20b）\n- 80GB H100：MXFP4 + 128K 上下文，官方目标\n- 2×RTX 4090 48GB：社区可行，需 offload","training_md":"- 预训练规模未详细披露（以英文 STEM / 代码为主）\n- 后训练：与 o 系列相同的 SFT + 高算力 RL\n- 推理默认开启，system prompt 里 `Reasoning: low|medium|high`\n- 工具：内置浏览、Python、自定义函数（Harmony 格式）\n- 最大输出 131K","ecosystem_md":"- HF：openai/gpt-oss-120b\n- 引擎：vLLM、SGLang、llama.cpp、Ollama、LM Studio、TensorRT-LLM；Groq / Cerebras 提供超高速托管\n- 微调：OpenAI 官方给出 LoRA 示例；Unsloth 支持\n- 中文文档：无（官方英文）","versions_md":"- 小版：gpt-oss-20b（21B/3.6B，16GB 可跑）\n- 变体：gpt-oss-safeguard（安全分类）"},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Ollama","TensorRT-LLM"],"finetune":"LoRA（官方示例）","zh_docs":"无"},"i18n":{"en":{"name_zh":"OpenAI open-weight 120B","one_liner":"OpenAI open reasoning model on one 80GB GPU; 5B active, very fast.","highlights":["Apache-2.0, OpenAI's first open weights in six years","MXFP4 native 63 GB; deploys on a single H100 / dual 4090 class","5.1B active, very high throughput; three adjustable reasoning levels"],"pitfalls":["The Harmony format must be followed strictly; an old chat template degrades quality noticeably","Weak knowledge breadth and multilinguality (training skews English, STEM); Chinese clearly below Qwen / GLM","High hallucination rate (official PersonQA hallucination rate 49%); use with care for factual QA"],"logic_ability":"Exam-style reasoning is very strong for its size (AIME 2025 97.9% with tools, GPQA 80.1%), thanks to an RL recipe shared with the o-series. Engineering reasoning (SWE-bench 62.4%) is middling. Reasoning effort is key: at low it barely thinks, at high the chain can reach tens of thousands of tokens. Common issues: reasoning quality drops outside English; frequent hallucinations.","best_for":["Single-GPU high-throughput inference services","Math / STEM reasoning","English agent prototypes"],"not_for":["Chinese-first products","Knowledge QA / fact checking"],"capability_notes":{"coding":"SWE-bench Verified 62.4%, Codeforces 2622 (tools, official).","reasoning":"GPQA Diamond 80.1%, HLE 19.0% (no tools, official).","math":"AIME 2025 97.9% (tools), 92.5% (no tools, official).","agent":"τ-bench Retail 67.8% (official).","chinese":"Weak; Chinese is not a focus language."}},"ja":{"name_zh":"OpenAI オープンウェイト 120B","one_liner":"80GB 単一 GPU で動く OpenAI オープン推論モデル、5B 活性。","highlights":["Apache-2.0、OpenAI が 6 年ぶりに公開したオープンウェイト","MXFP4 ネイティブ 63 GB、H100 1 枚 / 4090 2 枚クラスで展開可能","5.1B アクティブでスループットが極めて高い。推論強度は 3 段階で調整可能"],"pitfalls":["Harmony フォーマットを厳密に守る必要があり、旧版の chat template では品質が大きく低下","知識の広さと多言語性が弱い（学習は英語・STEM 寄り）、中国語は Qwen / GLM に明らかに劣る","ハルシネーション率が高め（公式 PersonQA ハルシネーション率 49%）、事実 QA は慎重に"],"logic_ability":"試験型推論はサイズに対して極めて強い（AIME 2025 ツールあり 97.9%、GPQA 80.1%）。o シリーズと同系統の RL レシピの恩恵。エンジニアリング型推論（SWE-bench 62.4%）は中程度。reasoning effort が鍵：low ではほぼ思考せず、high では思考チェーンが数万トークンに達する。よくある問題：英語以外の言語で推論品質が低下、ハルシネーションが多い。","best_for":["単一 GPU の高スループット推論サービス","数学 / STEM 推論","英語 agent のプロトタイプ"],"not_for":["中国語中心の製品","知識 QA / ファクトチェック"],"capability_notes":{"coding":"SWE-bench Verified 62.4%、Codeforces 2622（ツールあり、公式）。","reasoning":"GPQA Diamond 80.1%、HLE 19.0%（ツールなし、公式）。","math":"AIME 2025 97.9%（ツールあり）、92.5%（ツールなし、公式）。","agent":"τ-bench Retail 67.8%（公式）。","chinese":"弱い。中国語は重点言語ではない。"}}},"complete":true,"runtime":{"tok_s":152,"latency_s":0.87,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"id":"gpt-oss-20b","name":"gpt-oss-20b","name_zh":"OpenAI 开源 20B","aliases":["gpt oss 20b"],"vendor":"OpenAI","family":"gpt-oss","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/openai/gpt-oss-20b","status":"current","released_at":"2025-08-05","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"default-on","architecture":{"type":"moe","total_params":"20.9B","active_params":"3.6B","total_params_b":20.9,"active_params_b":3.6,"experts":32,"active_experts":4,"layers":24,"hidden_size":2880,"kv_heads":8,"head_dim":64,"attention":"GQA（64 Q / 8 KV，head_dim 64），滑窗 128 与全注意力逐层交替","undisclosed":false,"vocab_size":201088,"kv_layers":12,"notes":"24 层 = 12 × (滑窗 128 + 全注意力)。32 专家选 4，专家中间维 2880，带 attention sink。YaRN factor 32（原始 4K → 128K），theta 150000。MoE 权重原生 MXFP4，注意力 / 嵌入为 BF16。表内 kv_layers 指 12 层全注意力层。"},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":13.8,"q8":12.1,"q4":11.6},"kv_per_token_kib":24,"kv_note":"只有 12 层全注意力层随长度增长：8 KV 头 × 64 × 2（K/V）× 2 B = 2 KiB/层 → 24 KiB/token（BF16）。12 层滑窗层窗口 128，KV 固定可忽略。128K 上下文 KV ≈ 3 GB。表内「bf16」为 HF 官方 safetensors 13.8 GB（MoE 为 MXFP4、其余 BF16），完全反量化到 BF16 约 42 GB。","ref_hw_24gb":"官方 MXFP4 13.8 GB（GGUF Q4_K_M 11.6 GB），16GB 卡即可；24GB 卡 128K 上下文单并发","ref_hw_80gb":"过剩；可高并发","ref_hw_8x80gb":"过剩","estimated":false},"pricing":{"input_per_m":0.05,"output_per_m":0.2,"currency":"USD","source":"第三方托管常见价（OpenRouter）","as_of":"2025-12-20"},"links":{"official":"https://openai.com/index/introducing-gpt-oss/","hf":"https://huggingface.co/openai/gpt-oss-20b","github":"https://github.com/openai/gpt-oss"},"variants":[{"kind":"other","publisher":"openai","repo":"openai/gpt-oss-20b","url":"https://huggingface.co/openai/gpt-oss-20b","note":"官方 MXFP4 safetensors（MoE 4-bit、其余 BF16）","sizes":{"bf16":13.8}},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/gpt-oss-20b-GGUF","url":"https://huggingface.co/unsloth/gpt-oss-20b-GGUF","sizes":{"q4":11.6,"q5":11.7,"q6":12,"q8":12.1}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/openai_gpt-oss-20b-GGUF","url":"https://huggingface.co/bartowski/openai_gpt-oss-20b-GGUF","sizes":{"q4":11.7,"q5":11.7,"q6":12,"q8":12.1,"bf16":13.8}},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/gpt-oss-20b-MXFP4-Q8","url":"https://huggingface.co/mlx-community/gpt-oss-20b-MXFP4-Q8"},{"kind":"bnb","publisher":"unsloth","repo":"unsloth/gpt-oss-20b-unsloth-bnb-4bit","url":"https://huggingface.co/unsloth/gpt-oss-20b-unsloth-bnb-4bit"}],"copy":{"one_liner":"16GB 显存跑的推理模型，数学能力惊人。","highlights":["官方 MXFP4 13.8 GB，16GB 卡 / 笔记本可跑，仅 12 层全注意力 KV 24 KiB/token","AIME 2025 91.7（无工具）/ 98.7（工具）、GPQA 71.5，考试型推理远超体量","Apache-2.0，reasoning 三档可调，SWE-bench Verified 60.7"],"pitfalls":["知识与中文弱（MMLU 85.3、MMMLU 75.7），HLE 无工具仅 10.9","幻觉率高，tau-bench Airline 38.0、Aider Polyglot 34.2，agent 稳定性一般","Harmony 格式要求严格，chat template 不对会明显掉分"],"logic_ability":"考试型推理相对体量极强：AIME 2025 91.7（无工具）/ 98.7（工具）、AIME 2024 92.1、GPQA Diamond 71.5、Codeforces 2230（模型卡，high 档）。HLE 10.9 / 17.3（工具）显示前沿知识浅。工程型推理中等：SWE-bench Verified 60.7、tau-bench Retail 54.8 / Airline 38.0。reasoning 低档时推理链很短，分数明显下滑。","best_for":["本地推理助手","边缘 STEM 推理"],"not_for":["中文产品","知识问答"]},"capability_notes":{"coding":"SWE-bench Verified 60.7、Codeforces Elo 2230、Aider Polyglot 34.2（OpenAI 模型卡 arXiv 2508.10925，high）。","reasoning":"GPQA Diamond 71.5（无工具）/ 74.2（工具）、HLE 10.9 / 17.3（工具）（OpenAI 模型卡）。","math":"AIME 2025 91.7（无工具）/ 98.7（工具）、AIME 2024 92.1 / 96.0（OpenAI 模型卡）。","agent":"tau-bench Retail 54.8、Airline 38.0（OpenAI 模型卡）。","knowledge":"MMLU 85.3、MMMLU 75.7、HealthBench 42.5（OpenAI 模型卡）。","chinese":"中文弱，MMMLU 75.7，不建议中文产品主力。"},"ecosystem":{"engines":["llama.cpp","Ollama","vLLM","SGLang","MLX","Transformers"],"finetune":"官方指南 + Unsloth / TRL；LoRA 24GB 可行","zh_docs":"无"},"i18n":{"en":{"name_zh":"OpenAI open-weight 20B","one_liner":"A reasoning model that runs in 16GB VRAM with surprising math ability.","highlights":["Official MXFP4 13.8 GB, runs on a 16GB GPU / laptop; only 12 full-attention layers, KV 24 KiB/token","AIME 2025 91.7 (no tools) / 98.7 (tools), GPQA 71.5; exam-style reasoning far beyond its size","Apache-2.0, three adjustable reasoning levels, SWE-bench Verified 60.7"],"pitfalls":["Weak knowledge and Chinese (MMLU 85.3, MMMLU 75.7); HLE only 10.9 without tools","High hallucination rate; tau-bench Airline 38.0, Aider Polyglot 34.2, mediocre agent stability","Strict Harmony format requirements; a wrong chat template loses noticeable points"],"logic_ability":"Exam-style reasoning is very strong for its size: AIME 2025 91.7 (no tools) / 98.7 (tools), AIME 2024 92.1, GPQA Diamond 71.5, Codeforces 2230 (model card, high). HLE 10.9 / 17.3 (tools) shows shallow frontier knowledge. Engineering reasoning is middling: SWE-bench Verified 60.7, tau-bench Retail 54.8 / Airline 38.0. At low reasoning the chain is very short and scores drop markedly.","best_for":["Local reasoning assistants","Edge STEM reasoning"],"not_for":["Chinese products","Knowledge QA"],"capability_notes":{"coding":"SWE-bench Verified 60.7, Codeforces Elo 2230, Aider Polyglot 34.2 (OpenAI model card arXiv 2508.10925, high).","reasoning":"GPQA Diamond 71.5 (no tools) / 74.2 (tools), HLE 10.9 / 17.3 (tools) (OpenAI model card).","math":"AIME 2025 91.7 (no tools) / 98.7 (tools), AIME 2024 92.1 / 96.0 (OpenAI model card).","agent":"tau-bench Retail 54.8, Airline 38.0 (OpenAI model card).","knowledge":"MMLU 85.3, MMMLU 75.7, HealthBench 42.5 (OpenAI model card).","chinese":"Weak Chinese, MMMLU 75.7; not recommended as the primary model for Chinese products."}},"ja":{"name_zh":"OpenAI オープンウェイト 20B","one_liner":"16GB VRAM で動く推論モデル、数学能力が驚異的。","highlights":["公式 MXFP4 13.8 GB、16GB GPU / ノート PC で動作。全アテンション層は 12 層のみで KV 24 KiB/token","AIME 2025 91.7（ツールなし）/ 98.7（ツールあり）、GPQA 71.5、試験型推論はサイズをはるかに超える","Apache-2.0、reasoning は 3 段階で調整可能、SWE-bench Verified 60.7"],"pitfalls":["知識と中国語が弱い（MMLU 85.3、MMMLU 75.7）、HLE ツールなしはわずか 10.9","ハルシネーション率が高く、tau-bench Airline 38.0、Aider Polyglot 34.2 と agent の安定性は平凡","Harmony フォーマットの要件が厳格で、chat template が違うと明らかにスコアが落ちる"],"logic_ability":"試験型推論はサイズに対して極めて強い：AIME 2025 91.7（ツールなし）/ 98.7（ツールあり）、AIME 2024 92.1、GPQA Diamond 71.5、Codeforces 2230（モデルカード、high）。HLE 10.9 / 17.3（ツールあり）は最先端知識の浅さを示す。エンジニアリング型推論は中程度：SWE-bench Verified 60.7、tau-bench Retail 54.8 / Airline 38.0。reasoning が低段階だと推論チェーンが非常に短く、スコアが明らかに下がる。","best_for":["ローカル推論アシスタント","エッジでの STEM 推論"],"not_for":["中国語製品","知識 QA"],"capability_notes":{"coding":"SWE-bench Verified 60.7、Codeforces Elo 2230、Aider Polyglot 34.2（OpenAI モデルカード arXiv 2508.10925、high）。","reasoning":"GPQA Diamond 71.5（ツールなし）/ 74.2（ツールあり）、HLE 10.9 / 17.3（ツールあり）（OpenAI モデルカード）。","math":"AIME 2025 91.7（ツールなし）/ 98.7（ツールあり）、AIME 2024 92.1 / 96.0（OpenAI モデルカード）。","agent":"tau-bench Retail 54.8、Airline 38.0（OpenAI モデルカード）。","knowledge":"MMLU 85.3、MMMLU 75.7、HealthBench 42.5（OpenAI モデルカード）。","chinese":"中国語は弱く MMMLU 75.7。中国語製品の主力には非推奨。"}}},"complete":true,"runtime":{"tok_s":119,"latency_s":1.12,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"sheet":{"architecture_md":"**类型**：MoE，20.9B 总参 / 3.6B 激活，纯文本。\n\n| 项目 | 数值 |\n|---|---|\n| 层数 | 24 = 12 × (滑窗 128 + 全注意力) |\n| 隐藏维 | 2,880 |\n| 专家 | 32，每 token 选 4，专家中间维 2,880 |\n| 注意力 | 64 Q / 8 KV，head_dim 64，attention sink，bias |\n| RoPE | YaRN factor 32，原始 4,096，theta 150,000 |\n| 词表 | 201,088（o200k_harmony） |\n| 上下文 | 131,072 |\n| 量化 | MoE 权重 MXFP4，其余 BF16 |\n\n`config.json` 中 `model_type: gpt_oss`，`quant_method: mxfp4`。\n\n参考：openai/gpt-oss-20b config.json 与模型卡（arXiv 2508.10925）。","memory_md":"| 精度 | 权重大小 | 参考显存 | 说明 |\n|---|---|---|---|\n| MXFP4（官方） | 13.8 GB | 16GB | HF safetensors 3 分片 13.76 GB；unsloth GGUF F16 同为 13.8 GB |\n| Q8 | 12.1 GB | 16GB | unsloth GGUF Q8_0 |\n| Q4 | 11.6 GB | 16GB | unsloth GGUF Q4_K_M |\n| 全 BF16 反量化 | ≈42 GB | 80GB | 按 20.9B × 2 B 估算，无官方文件 |\n\n**KV Cache**：12 层全注意力层 × 2 KiB = 24 KiB/token；滑窗层固定可忽略。128K 上下文 KV ≈ 3 GB。\n\n**参考配置**：\n- 16GB 卡 / Mac：MXFP4 全量可跑\n- RTX 4090 24GB：128K 上下文单并发\n- 80GB：高并发\n\n警示：MXFP4 只有 Hopper / Blackwell 原生支持，其他 GPU 会在加载时反量化到 BF16（约 42 GB）或需用 GGUF。","training_md":"- 预训练 token 数：**未披露**（模型卡仅写「数万亿 token」，STEM / 代码为主，H100 训练）\n- Harmony 响应格式训练，需严格使用官方 chat template\n- reasoning effort 三档：low / medium / high，可在系统提示中设定\n- 支持函数调用、浏览、Python 工具、结构化输出\n- 训练即 MXFP4 感知，MoE 权重原生 4-bit","ecosystem_md":"- HF：openai/gpt-oss-20b（13.8 GB）；GGUF：unsloth/gpt-oss-20b-GGUF、ggml-org\n- 引擎：Transformers、vLLM、SGLang、llama.cpp、Ollama、LM Studio、MLX；官方另给 PyTorch / Triton / Metal 参考实现\n- 微调：官方 fine-tuning 指南，Unsloth / TRL 支持；LoRA 24GB 可行\n- 中文文档：无","versions_md":"- 同代：gpt-oss-120b（117B-A5.1B，同架构 36 层）\n- 后续：gpt-oss-safeguard（安全分类衍生）\n- 发布于 2025-08-05，为 OpenAI 自 GPT-2 后首个开放权重模型"}},{"id":"granite-4-2-30b","name":"Granite 4.2 30B","name_zh":"IBM Granite 4.2 30B","aliases":["granite-4.2-30b","ibm-granite/granite-4.2-30b"],"vendor":"IBM","family":"Granite 4","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/ibm-granite/granite-4.2-30b","status":"current","released_at":"2026-08-25","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"optional","architecture":{"type":"dense","total_params":"30B","total_params_b":30,"layers":64,"hidden_size":4096,"vocab_size":100352,"kv_heads":8,"head_dim":128,"attention":"GQA（32 Q / 8 KV）","notes":"稠密 Transformer，FFN 32768，RoPE theta 1e7；原生 128K，可扩 512K。基于 Granite-4.1-30B-Base 后训练，<think> 三档思考模式。 上下文：128K（可扩 512K）。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":60,"q8":32,"q4":17.4,"fp8":30.1},"kv_per_token_kib":256,"estimated":true,"ref_hw_24gb":"Q4 约 17 GB 可跑，KV 256 KiB/token 上下文有限","ref_hw_80gb":"BF16 60 GB 单卡 + 数万 token 上下文"},"links":{"official":"https://huggingface.co/ibm-granite/granite-4.2-30b","hf":"https://huggingface.co/ibm-granite/granite-4.2-30b"},"variants":[{"kind":"gguf","publisher":"ibm-granite","repo":"ibm-granite/granite-4.2-30b-GGUF","url":"https://huggingface.co/ibm-granite/granite-4.2-30b-GGUF","note":"官方 GGUF","sizes":{"q4":17.7,"q5":20.8,"q6":24,"q8":31.1,"bf16":58.6}},{"kind":"fp8","publisher":"ibm-granite","repo":"ibm-granite/granite-4.2-30b-fp8","url":"https://huggingface.co/ibm-granite/granite-4.2-30b-fp8","note":"官方 FP8","sizes":{"fp8":30.1}},{"kind":"nvfp4","publisher":"ibm-granite","repo":"ibm-granite/granite-4.2-30b-nvfp4","url":"https://huggingface.co/ibm-granite/granite-4.2-30b-nvfp4","note":"官方 NVFP4"},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/granite-4.2-30b-GGUF","url":"https://huggingface.co/bartowski/granite-4.2-30b-GGUF","sizes":{"q4":18,"q5":21,"q6":24.5,"q8":31.1,"bf16":58.6}},{"kind":"mlx","publisher":"lmstudio-community","repo":"lmstudio-community/granite-4.2-30b-MLX-4bit","url":"https://huggingface.co/lmstudio-community/granite-4.2-30b-MLX-4bit"}],"copy":{"one_liner":"IBM 30B 稠密推理模型，Apache-2.0，24GB 可跑。","highlights":["AIME 2025 89.2、LiveCodeBench v6 75.8、MMLU-Pro 77.6（官方）","Apache-2.0，企业友好，12 种语言含中文","Q4 约 17 GB，24GB 单卡本地"],"pitfalls":["刚发布 3 天，AA 智能指数 24 低于 Muse Glimmer（35）","SWE-bench Pro 33.3，编程 agent 弱","无视觉（Granite Vision 另出）"],"logic_ability":"考试型推理数学强（AIME 89.2）、GPQA 66.4 中游；工程型推理一般（SWE-bench Pro 33.3）。","best_for":["企业内网合规部署","24GB 本地推理 / 工具调用"],"not_for":["编程 agent 主力","多模态"]},"capability_notes":{"coding":"LiveCodeBench v6 75.77、SWE-bench Pro 33.29（官方）。","reasoning":"GPQA 66.41、MMLU-Pro 77.60（官方）。","math":"AIME 2025 89.17（官方）。","agent":"BFCL v4 61.39（官方）。","chinese":"官方 12 语言含中文。"},"ecosystem":{"engines":["vLLM","SGLang","Transformers","llama.cpp","Ollama"],"zh_docs":"无"},"i18n":{"en":{"name_zh":"IBM Granite 4.2 30B","one_liner":"IBM 30B dense reasoning model, Apache-2.0, runs on 24GB.","highlights":["AIME 2025 89.2, LiveCodeBench v6 75.8, MMLU-Pro 77.6 (official)","Apache-2.0, enterprise-friendly, 12 languages including Chinese","About 17 GB at Q4, local on a single 24GB GPU"],"pitfalls":["Released only 3 days ago; AA Intelligence Index 24, below Muse Glimmer (35)","SWE-bench Pro 33.3, weak as a coding agent","No vision (Granite Vision is a separate release)"],"logic_ability":"Exam-style reasoning strong in math (AIME 89.2), GPQA 66.4 mid-pack; engineering reasoning mediocre (SWE-bench Pro 33.3).","best_for":["Compliant deployment on enterprise intranets","24GB local inference / tool calling"],"not_for":["Primary coding agent","Multimodal"],"capability_notes":{"coding":"LiveCodeBench v6 75.77, SWE-bench Pro 33.29 (official).","reasoning":"GPQA 66.41, MMLU-Pro 77.60 (official).","math":"AIME 2025 89.17 (official).","agent":"BFCL v4 61.39 (official).","chinese":"Officially 12 languages including Chinese."}},"ja":{"name_zh":"IBM Granite 4.2 30B","one_liner":"IBM 30B Dense 推論モデル、Apache-2.0、24GB で動作。","highlights":["AIME 2025 89.2、LiveCodeBench v6 75.8、MMLU-Pro 77.6（公式）","Apache-2.0、企業向け、中国語を含む 12 言語","Q4 で約 17 GB、24GB 単一 GPU でローカル動作"],"pitfalls":["公開から 3 日で AA 知能指数 24、Muse Glimmer（35）を下回る","SWE-bench Pro 33.3、コーディング agent としては弱い","視覚なし（Granite Vision は別リリース）"],"logic_ability":"試験型推論は数学が強く（AIME 89.2）、GPQA 66.4 は中位。エンジニアリング型推論は平凡（SWE-bench Pro 33.3）。","best_for":["企業イントラネットでのコンプライアンス準拠展開","24GB ローカル推論 / ツール呼び出し"],"not_for":["コーディング agent の主力","マルチモーダル"],"capability_notes":{"coding":"LiveCodeBench v6 75.77、SWE-bench Pro 33.29（公式）。","reasoning":"GPQA 66.41、MMLU-Pro 77.60（公式）。","math":"AIME 2025 89.17（公式）。","agent":"BFCL v4 61.39（公式）。","chinese":"公式に中国語を含む 12 言語対応。"}}},"complete":false,"runtime":{"tok_s":75,"latency_s":0.92,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"grok-2","name":"Grok-2","aliases":["grok-2-1212","grok-2-beta","grok-2-mini","sus-column-r"],"vendor":"xAI","family":"Grok 2","superseded_by":"grok-3","released_at":"2024-08-13","modalities":["text","image","tools"],"reasoning_mode":"none","context":{"max_tokens":131072,"display":"128K","max_output":8192},"pricing":{"input_per_m":2,"output_per_m":10,"currency":"USD","source":"xAI API 定价（2024-10 开放）","as_of":"2024-10-21"},"links":{"official":"https://x.ai/news/grok-2","pricing":"https://x.ai/api"},"copy":{"one_liner":"xAI 首个追平 GPT-4 级的模型，X 平台内置。","highlights":["GPQA 56.0%、MMLU 87.5%，追平 GPT-4 Turbo","LMArena 上以 sus-column-r 匿名测试排名靠前","集成 Flux 图像生成，限制少"],"pitfalls":["API 2024-10 才开放","长上下文与代码能力一般","权重 2025-08 以 Grok 2.5 名义开源，但许可证受限"],"logic_ability":"与 GPT-4 Turbo / Claude 3.5 Sonnet 同档的推理，MATH 76.1%，无推理链。","best_for":["历史对照"],"not_for":["新项目"]}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。Colossus 20 万卡集群训练；Think 模式为推理版本，API 2025-04-09 开放。"},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"grok-3","name":"Grok 3","aliases":["grok-3-beta","grok-3-mini","grok-3-latest"],"vendor":"xAI","family":"Grok 3","superseded_by":"grok-4","released_at":"2025-02-17","modalities":["text","image","tools"],"reasoning_mode":"optional","context":{"max_tokens":131072,"display":"128K（宣称 1M）","max_output":16384},"pricing":{"input_per_m":3,"output_per_m":15,"currency":"USD","source":"xAI API 定价","as_of":"2025-04-09","note":"Grok 3 mini $0.30 / $0.50"},"links":{"official":"https://x.ai/news/grok-3","pricing":"https://x.ai/api"},"copy":{"one_liner":"首个 LMArena 破 1400 的模型，Think 版 AIME 93%。","highlights":["LMArena Elo 1402（\"chocolate\"），首个破 1400","Think 模式 AIME 2025 93.3%（cons@64）、GPQA 84.6%","DeepSearch 联网研究功能"],"pitfalls":["官方图表 cons@64 与他家 pass@1 对比引发争议","API 上下文 131K，非宣称的 1M","API 比产品晚近两个月"],"logic_ability":"非推理版 GPQA 75.4%、LCB 57.0%；Think 版考试型推理进入第一梯队，工程型无官方数据。","best_for":["历史对照"],"not_for":["新项目"]}},{"id":"grok-4-1-fast","name":"Grok 4.1 Fast","aliases":["grok-4-1-fast-reasoning"],"vendor":"xAI","family":"Grok 4","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","released_at":"2025-11-19","updated_at":"2025-12-20","modalities":["text","image","tools"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true},"context":{"max_tokens":2000000,"display":"2M"},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":0.2,"output_per_m":0.5,"currency":"USD","source":"xAI 定价页","as_of":"2025-12-20"},"links":{"official":"https://x.ai/news/grok-4-1-fast","pricing":"https://docs.x.ai/docs/models"},"copy":{"one_liner":"2M 上下文、$0.2 输入的工具调用特化档。","highlights":["2M 上下文","$0.20 / $0.50","τ²-bench 工具调用官方领先"],"pitfalls":["独立复测数据少","推理深度不如 Grok 4","无自建"],"logic_ability":"工具调用型 agent 强；深度推理中等。","best_for":["低价长上下文 agent"],"not_for":["顶级推理"]},"capability_notes":{},"complete":false,"superseded_by":"grok-4-3"},{"id":"grok-4-3","name":"Grok 4.3","name_zh":"Grok 4.3","aliases":["grok-4.3","Grok 4.3 High"],"vendor":"xAI","vendor_zh":"xAI","family":"Grok 4","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2026-05-01","updated_at":"2026-08-28","modalities":["text","image","video","tools"],"reasoning_mode":"optional","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。reasoning 可设 none / low / medium / high。发布日期以 2026-05-01 媒体报道为准，xAI 未给出精确日期。"},"context":{"max_tokens":1000000,"display":"1M"},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":1.25,"output_per_m":2.5,"currency":"USD","source":"xAI 模型定价页","as_of":"2026-08-28","note":"≥200K 上下文：$2.50 / $5.00；缓存输入 $0.20"},"runtime":{"tok_s":126,"latency_s":15.33,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"links":{"official":"https://x.ai/news/grok-amazon-bedrock","pricing":"https://docs.x.ai/docs/models"},"copy":{"one_liner":"xAI 低价档：1M 上下文、$1.25 / $2.50、可关推理。","highlights":["1M 上下文，输出价 $2.50，是 Grok 4.5 的 1/2.4","reasoning none 档可当非推理模型用，AA 实测 126 tok/s","支持视频输入与文档（PDF / 表格 / 幻灯片）生成"],"pitfalls":["AA 智能指数 38，远低于 Grok 4.5 / 4.6，复杂任务别指望","无独立官方公告页，基准数据零散","Arena 1442，与 Gemma 4 31B 开源模型相当"],"logic_ability":"中档；xAI 强调低幻觉率与性价比，推理深度可配。","best_for":["高并发低成本调用","长文档 / 大代码库检索","简单 agent 工具调用"],"not_for":["前沿推理与数学","私有化部署"]},"capability_notes":{},"complete":false,"i18n":{"en":{"name_zh":"Grok 4.3","one_liner":"xAI's budget tier: 1M context, $1.25 / $2.50, reasoning can be turned off.","highlights":["1M context, $2.50 output price, 1/2.4 of Grok 4.5","reasoning none tier works as a non-reasoning model; AA measured 126 tok/s","Supports video input and document (PDF / spreadsheet / slides) generation"],"pitfalls":["AA Intelligence Index 38, far below Grok 4.5 / 4.6; do not expect it on complex tasks","No standalone official announcement page; benchmark data is scattered","Arena 1442, on par with the open-source Gemma 4 31B"],"logic_ability":"Mid-tier; xAI emphasizes low hallucination rate and cost-effectiveness, with configurable reasoning depth.","best_for":["High-concurrency, low-cost calls","Long-document / large-codebase retrieval","Simple agent tool calling"],"not_for":["Frontier reasoning and math","Self-hosted deployment"],"capability_notes":{}},"ja":{"name_zh":"Grok 4.3","one_liner":"xAI の低価格帯。1M コンテキスト、$1.25 / $2.50、推論オフ可。","highlights":["1M コンテキスト、出力価格 $2.50 で Grok 4.5 の 1/2.4","reasoning none 設定で非推論モデルとして利用可能、AA 実測 126 tok/s","動画入力とドキュメント（PDF / 表 / スライド）生成に対応"],"pitfalls":["AA 知能指数 38 で Grok 4.5 / 4.6 を大きく下回り、複雑なタスクには期待できない","独立した公式発表ページがなく、ベンチマークデータが散在","Arena 1442、オープンソースの Gemma 4 31B と同等"],"logic_ability":"中堅クラス。xAI は低ハルシネーション率とコスパを強調、推論の深さは設定可能。","best_for":["高並列・低コストの呼び出し","長文ドキュメント / 大規模コードベース検索","シンプルなエージェントのツール呼び出し"],"not_for":["最先端の推論と数学","プライベート環境へのデプロイ"],"capability_notes":{}}}},{"id":"grok-4-5","name":"Grok 4.5","name_zh":"Grok 4.5","aliases":["grok-4.5","Grok 4.5 High"],"vendor":"xAI","vendor_zh":"xAI","family":"Grok 4","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2026-07-16","updated_at":"2026-08-28","modalities":["text","image","tools"],"reasoning_mode":"default-on","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。reasoning_effort 可选 low / medium / high（默认 high）。"},"context":{"max_tokens":500000,"display":"500K"},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":2,"output_per_m":6,"currency":"USD","source":"xAI 模型定价页","as_of":"2026-08-28","note":"≥200K 上下文：$4 / $12；缓存输入 $0.30"},"runtime":{"tok_s":52,"latency_s":14.21,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"links":{"official":"https://x.ai/news/grok-4-5","pricing":"https://docs.x.ai/docs/models"},"copy":{"one_liner":"xAI 面向代码与 agent 的旗舰，输出价仅 $6，Arena 1470。","highlights":["Terminal-Bench 2.1 83.3%、SWE-Bench Pro 64.7%，代码 agent 进入第一梯队","$2 / $6 定价，输出价是 Gemini 3.1 Pro 的一半","Arena 文本榜 1470，高于后续的 Grok 4.6（1461）"],"pitfalls":["发布不到一个月就被 Grok 4.6 取代为「推荐」模型，xAI 迭代节奏极快","官方公告只给代码 / agent 类基准，GPQA / HLE / AIME 等通用推理数据缺失","AA 实测 51.6 tok/s、TTFT 约 14 s，速度偏慢；仅文本 + 图像，无音视频"],"logic_ability":"工程型推理是主打：DeepSWE 1.0 62.0%、SWE Marathon 29.0%（公告称第一）、Terminal-Bench 2.1 83.3%、SWE-Bench Pro 64.7%。考试型推理官方未公布，AA 智能指数 56（与 Gemini 3.7 Flash 持平，低于 Grok 4.6 的 61）。失败模式：长程自主软件工程任务仍不及 GPT-5.5 / Opus 级；token 效率好但延迟高。","best_for":["代码 agent / IDE 集成","知识工作类长任务","预算敏感的旗舰级调用"],"not_for":["低延迟交互","音视频输入","私有化部署"]},"capability_notes":{"coding":"Terminal-Bench 2.1 83.3%、SWE-Bench Pro 64.7%、DeepSWE 1.0 62.0%（xAI 公告）。","reasoning":"官方未公布 GPQA / HLE；AA 智能指数 56。","agent":"SWE Marathon 29.0%（xAI 公告）。","multimodal":"文本 + 图像输入。"},"sheet":{"architecture_md":"**类型**：未披露。\n\n已知：\n- 500K 上下文，文本 + 图像输入，文本输出\n- `reasoning_effort`：low / medium / high（默认 high）；Grok 4.6 再加 xhigh\n- 面向代码、agent 与知识工作训练\n- 工具调用、结构化输出\n\n参考：x.ai/news/grok-4-5、docs.x.ai 模型页。","memory_md":"无自建选项。\n\n成本参考（2026-08-28 xAI 定价页）：\n\n| 档位 | 输入 $/M | 输出 $/M | 缓存输入 |\n|---|---|---|---|\n| <200K | 2.00 | 6.00 | 0.30 |\n| ≥200K | 4.00 | 12.00 | 0.60 |\n\n**提示**：AA 估算单次智能指数任务成本约 $0.31，token 效率优于同档闭源；但 TTFT 长，按并发设计要留超时。","training_md":"- 训练细节未披露\n- 默认开启推理，effort 三档可调\n- 官方基准（公告）：DeepSWE 1.0 62.0%、DeepSWE 1.1 53%、SWE Marathon 29.0%、Terminal-Bench 2.1 83.3%、SWE-Bench Pro 64.7%\n- 第三方：AA 智能指数 56，Arena 文本 1470","ecosystem_md":"- xAI API（OpenAI 兼容）、Grok Build、Cursor、OpenRouter、Amazon Bedrock\n- Grok App（iOS / Android / Web / X）内置\n- 不支持微调\n- 中文文档：无","versions_md":"- 上代：Grok 4.3（2026-04 末，$1.25 / $2.50，1M 上下文，现为低价档）、Grok 4.20（2026-03）\n- 后续：Grok 4.6（2026-08-12，同价，AA 指数 61，xAI 现推荐）\n- 更早：Grok 4（2025-07）、Grok 4.1 Fast（2025-11），均已被取代"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"无"},"complete":true,"i18n":{"en":{"name_zh":"Grok 4.5","one_liner":"xAI's coding & agent flagship, only $6 output price, Arena 1470.","highlights":["Terminal-Bench 2.1 83.3%, SWE-Bench Pro 64.7%; coding agents enter the top tier","$2 / $6 pricing; output price is half of Gemini 3.1 Pro","Arena text leaderboard 1470, higher than the later Grok 4.6 (1461)"],"pitfalls":["Replaced by Grok 4.6 as the \"recommended\" model less than a month after release; xAI iterates extremely fast","Official announcement only gives coding / agent benchmarks; general reasoning data such as GPQA / HLE / AIME is missing","AA measured 51.6 tok/s and TTFT about 14 s, on the slow side; text + image only, no audio/video"],"logic_ability":"Engineering-style reasoning is the focus: DeepSWE 1.0 62.0%, SWE Marathon 29.0% (claimed #1 in the announcement), Terminal-Bench 2.1 83.3%, SWE-Bench Pro 64.7%. Exam-style reasoning not officially published; AA Intelligence Index 56 (tied with Gemini 3.7 Flash, below Grok 4.6 at 61). Failure modes: long-horizon autonomous software engineering still falls short of GPT-5.5 / Opus level; good token efficiency but high latency.","best_for":["Coding agents / IDE integration","Long knowledge-work tasks","Budget-sensitive flagship-level calls"],"not_for":["Low-latency interaction","Audio/video input","Self-hosted deployment"],"capability_notes":{"coding":"Terminal-Bench 2.1 83.3%, SWE-Bench Pro 64.7%, DeepSWE 1.0 62.0% (xAI announcement).","reasoning":"GPQA / HLE not officially published; AA Intelligence Index 56.","agent":"SWE Marathon 29.0% (xAI announcement).","multimodal":"Text + image input."}},"ja":{"name_zh":"Grok 4.5","one_liner":"xAI のコーディング・エージェント旗艦。出力 $6、Arena 1470。","highlights":["Terminal-Bench 2.1 83.3%、SWE-Bench Pro 64.7%、コーディングエージェントで第一集団入り","$2 / $6 の価格設定、出力価格は Gemini 3.1 Pro の半分","Arena テキスト榜 1470、後継の Grok 4.6（1461）より高い"],"pitfalls":["公開から 1 か月足らずで Grok 4.6 に「推奨」モデルの座を譲り、xAI の更新ペースが極めて速い","公式発表はコーディング / エージェント系ベンチマークのみで、GPQA / HLE / AIME などの汎用推論データが欠けている","AA 実測 51.6 tok/s、TTFT 約 14 秒とやや遅い。テキスト + 画像のみで音声・動画は非対応"],"logic_ability":"エンジニアリング型推論が主力：DeepSWE 1.0 62.0%、SWE Marathon 29.0%（発表では 1 位と主張）、Terminal-Bench 2.1 83.3%、SWE-Bench Pro 64.7%。試験型推論は公式未公開、AA 知能指数 56（Gemini 3.7 Flash と同等、Grok 4.6 の 61 より低い）。弱点：長期の自律的ソフトウェアエンジニアリングタスクでは GPT-5.5 / Opus 級に及ばない。トークン効率は良いがレイテンシが高い。","best_for":["コーディングエージェント / IDE 統合","ナレッジワーク系の長時間タスク","予算重視のフラッグシップ級呼び出し"],"not_for":["低レイテンシの対話","音声・動画入力","プライベート環境へのデプロイ"],"capability_notes":{"coding":"Terminal-Bench 2.1 83.3%、SWE-Bench Pro 64.7%、DeepSWE 1.0 62.0%（xAI 発表）。","reasoning":"GPQA / HLE は公式未公開。AA 知能指数 56。","agent":"SWE Marathon 29.0%（xAI 発表）。","multimodal":"テキスト + 画像入力。"}}}},{"id":"grok-4-6","name":"Grok 4.6","name_zh":"Grok 4.6 加训版","aliases":["grok-4.6","Grok 4.6 High","grok-4.6-fast"],"vendor":"xAI","vendor_zh":"xAI","family":"Grok 4","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2026-08-12","updated_at":"2026-08-28","modalities":["text","image","tools"],"reasoning_mode":"default-on","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。xAI 称在 Grok 4.5 基座上追加训练与 agent 环境 RL，非更大模型。reasoning_effort：low / medium / high（默认）/ xhigh，推理不可关闭。知识截止 2026-02-01。"},"context":{"max_tokens":500000,"display":"500K"},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":2,"output_per_m":6,"currency":"USD","source":"xAI 模型定价页","as_of":"2026-08-28","note":"≥200K 上下文：$4 / $12；缓存输入 $0.50；Fast 变体价格翻倍"},"runtime":{"tok_s":58,"latency_s":42.98,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"links":{"official":"https://x.ai/news/grok-4-6","pricing":"https://docs.x.ai/docs/models"},"copy":{"one_liner":"Grok 4.5 的加训版，长程 agent 与知识工作更强，同价。","highlights":["AA 智能指数 61，GDPval-AA v2 1753，知识工作类基准领先","新增 xhigh 推理档，DeepSWE v1.1 65.9%、APEX-Agents 57.5%","与 4.5 同价 $2 / $6，xAI 当前推荐模型"],"pitfalls":["Terminal-Bench v3.0 仅 26%，自主软件工程类基准仍落后对手","Arena 文本 1461，反而低于 Grok 4.5 的 1470","无音视频输入；官方仍未公布 GPQA / HLE 等通用推理基准"],"logic_ability":"在 4.5 基础上强化长程 agent；知识工作类领先，纯软件工程类不占优。","best_for":["多步研究 / 分析 agent","可视化与交互式应用生成","长上下文代码库工作"],"not_for":["低延迟交互","私有化部署"]},"capability_notes":{"coding":"DeepSWE v1.1 65.9%、CursorBench v3.2 69.9%、FrontierCode v1.1 61.3%（xAI 公告）。","agent":"APEX-Agents 57.5%、Terminal-Bench v3.0 26%（xAI 公告）；Terminal-Bench v2.1 88.4%、τ³-Banking 50.7%、AA-Briefcase 1577（AA）。","reasoning":"AA 智能指数 61（high 档，2026-08-28，与 GPT-5.6 Sol 持平、次于 Opus 5 的 63）；较 Grok 4.5 +5 分。","knowledge":"GDPval-AA v2 1753（AA，仅次于 Opus 5）；知识截止 2026-02-01（docs.x.ai）。","multimodal":"文本 + 图像输入、文本输出；无音视频。"},"complete":true,"sheet":{"architecture_md":"**未披露**。xAI 未公布参数量与架构，仅说明 4.6 是 Grok 4.5 基座的追加训练版（改进优化器与训练配方），不是更大的模型。\n\n已知接口特性（docs.x.ai，2026-08-28）：\n- 上下文 500K；官方文档写「无文本输出上限」（未给出 max_output 数值）\n- `reasoning_effort`：`low` / `medium` / `high`（默认）/ `xhigh`；`xhigh` 为 4.6 起新增，**推理不可关闭**\n- 输入：文本 + 图像；输出：文本\n- 能力：function calling、structured outputs、prompt caching、context compaction；服务端工具 web search、X search、code execution\n- **不支持 Batch API**；速率限制 150 RPS / 50M TPM；区域 us-east-1、us-west-2\n- Fast 变体：同模型更快推理，价格翻倍\n- 知识截止 2026-02-01（模型页；开发者指南写 2026-01）","memory_md":"无自建选项。\n\n**成本模型**（docs.x.ai 模型页，2026-08-28）：\n- <200K 上下文：$2 输入 / $0.50 缓存读 / $6 输出\n- **≥200K 上下文：整条请求按 $4 / $1 / $12 计费**（翻倍）\n- Fast 变体价格翻倍；无 Batch 折扣\n- 与 Grok 4.5 同价；缓存读由 $0.30 涨到 $0.50\n\n**什么时候会变贵**：\n- 跨 200K 门槛整条翻倍，长上下文代码库工作要注意压缩（可用 context compaction）\n- xhigh 档 token 消耗最大\n- 优势在输出价：$6 仅为 Opus 5 的 1/4、GPT-5.6 Sol 的 1/5；AA 实测单任务 $0.84~0.94，比 Opus 5 低 60% 以上；agent 轮次约 53（Opus 5 约 103）、输入 token 约 0.5B（Opus 5 约 2.0B）\n\n速度：AA 实测 57.8 tok/s（低于均值），TTFT 约 43 s，不适合低延迟交互；需要速度选 Fast 变体。","training_md":"- 参数与数据规模未披露\n- 官方公开训练方法：在 Grok 4.5 上做延长补充训练——用 Grok 4.5 重新生成 SFT 轨迹（推理、agent harness、STEM / 软件工程 / 知识工作），配合高质量工程数据集、改进优化器与训练配方\n- Agentic RL：知识工作、通用编码，以及内核优化、Web 开发、CAD 等领域专用环境\n- 官方称擅长：长程 agent 任务、多步项目、可视化 / 交互式应用生成、自测与迭代修正\n- 官方数据：AA 指数 61、GDPval-AA v2 1753、CursorBench 69.9%、DeepSWE 65.9%、FrontierCode 61.3%、APEX-Agents 57.5%、Terminal-Bench v3.0 26%\n- 短板：Terminal-Bench v3.0 与自主软件工程类落后；Arena 文本 1461 低于 4.5 的 1470；未公布 GPQA / HLE","ecosystem_md":"- **API 渠道**：xAI API（`grok-4.6`，OpenAI / Anthropic 兼容协议）、Grok Build（x.ai/build 编码 agent）、Cursor\n- 网关：OpenRouter、Vercel AI Gateway、Cloudflare AI Gateway；官方文档未提及 Azure / Bedrock / Vertex\n- 发布首周 API 与 Grok Build 双倍额度\n- SDK：xai-sdk（Python）+ OpenAI 兼容 SDK\n- **不支持微调**\n- 中文文档：无（docs.x.ai 仅英文）","versions_md":"- 上代：Grok 4.5（同价 $2 / $6，AA 指数 56，Arena 1470）；更早 Grok 4.3、Grok 4.1 Fast、Grok 4\n- 本代：`grok-4.6`（2026-08-12）与 Fast 变体（价格翻倍）\n- 定位：xAI 当前推荐模型，「最智能且最快」的代码 / 对话模型\n- 迭代节奏：4.3 → 4.5 → 4.6 间 AA 指数 +23 分"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"无"},"i18n":{"en":{"name_zh":"Grok 4.6 (further-trained)","one_liner":"Further-trained Grok 4.5: stronger long-horizon agents and knowledge work.","highlights":["AA Intelligence Index 61, GDPval-AA v2 1753; leads on knowledge-work benchmarks","New xhigh reasoning tier; DeepSWE v1.1 65.9%, APEX-Agents 57.5%","Same $2 / $6 price as 4.5; xAI's current recommended model"],"pitfalls":["Terminal-Bench v3.0 only 26%; still behind rivals on autonomous software-engineering benchmarks","Arena text 1461, actually below Grok 4.5's 1470","No audio/video input; general reasoning benchmarks such as GPQA / HLE still not officially published"],"logic_ability":"Strengthens long-horizon agents on top of 4.5; leads on knowledge work, no edge on pure software engineering.","best_for":["Multi-step research / analysis agents","Visualization and interactive app generation","Long-context codebase work"],"not_for":["Low-latency interaction","Self-hosted deployment"],"capability_notes":{"coding":"DeepSWE v1.1 65.9%, CursorBench v3.2 69.9%, FrontierCode v1.1 61.3% (xAI announcement).","agent":"APEX-Agents 57.5%, Terminal-Bench v3.0 26% (xAI announcement); Terminal-Bench v2.1 88.4%, τ³-Banking 50.7%, AA-Briefcase 1577 (AA).","reasoning":"AA Intelligence Index 61 (high tier, 2026-08-28, tied with GPT-5.6 Sol, behind Opus 5 at 63); +5 over Grok 4.5.","knowledge":"GDPval-AA v2 1753 (AA, second only to Opus 5); knowledge cutoff 2026-02-01 (docs.x.ai).","multimodal":"Text + image input, text output; no audio/video."}},"ja":{"name_zh":"Grok 4.6 追加学習版","one_liner":"Grok 4.5 の追加学習版。長期エージェントとナレッジワークが強化、同価格。","highlights":["AA 知能指数 61、GDPval-AA v2 1753、ナレッジワーク系ベンチマークで首位","xhigh 推論レベルを新設、DeepSWE v1.1 65.9%、APEX-Agents 57.5%","4.5 と同じ $2 / $6、xAI の現行推奨モデル"],"pitfalls":["Terminal-Bench v3.0 はわずか 26%、自律的ソフトウェアエンジニアリング系ベンチマークでは競合に後れ","Arena テキスト 1461 で、Grok 4.5 の 1470 を下回る","音声・動画入力なし。GPQA / HLE などの汎用推論ベンチマークは依然公式未公開"],"logic_ability":"4.5 をベースに長期エージェントを強化。ナレッジワーク系では首位だが、純粋なソフトウェアエンジニアリングでは優位性なし。","best_for":["多段階のリサーチ / 分析エージェント","可視化・インタラクティブアプリ生成","長コンテキストのコードベース作業"],"not_for":["低レイテンシの対話","プライベート環境へのデプロイ"],"capability_notes":{"coding":"DeepSWE v1.1 65.9%、CursorBench v3.2 69.9%、FrontierCode v1.1 61.3%（xAI 発表）。","agent":"APEX-Agents 57.5%、Terminal-Bench v3.0 26%（xAI 発表）。Terminal-Bench v2.1 88.4%、τ³-Banking 50.7%、AA-Briefcase 1577（AA）。","reasoning":"AA 知能指数 61（high 設定、2026-08-28、GPT-5.6 Sol と同等、Opus 5 の 63 に次ぐ）。Grok 4.5 比 +5 点。","knowledge":"GDPval-AA v2 1753（AA、Opus 5 に次ぐ 2 位）。知識カットオフ 2026-02-01（docs.x.ai）。","multimodal":"テキスト + 画像入力、テキスト出力。音声・動画なし。"}}}},{"id":"grok-4","name":"Grok 4","aliases":["grok-4-0709"],"vendor":"xAI","family":"Grok 4","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","released_at":"2025-07-09","updated_at":"2025-12-20","modalities":["text","image","tools"],"reasoning_mode":"default-on","architecture":{"type":"unknown","undisclosed":true},"context":{"max_tokens":256000,"display":"256K"},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":3,"output_per_m":15,"currency":"USD","source":"xAI 定价页","as_of":"2025-12-20"},"links":{"official":"https://x.ai/news/grok-4","pricing":"https://docs.x.ai/docs/models"},"copy":{"one_liner":"考试型推理顶级的闭源旗舰，价格与 Claude Sonnet 同档但更慢。","highlights":["HLE 25.4%（无工具）、GPQA 87.5%","256K 上下文","Grok 4 Heavy 多 agent 变体"],"pitfalls":["延迟高、吞吐低","代码 agent 生态弱","无自建"],"logic_ability":"考试型推理顶级；工程型推理中上。","best_for":["科学推理","深度研究"],"not_for":["低延迟产品"]},"capability_notes":{},"complete":false,"superseded_by":"grok-4-5"},{"aliases":["Hunyuan-A13B-Instruct"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"optional","capability_notes":{"coding":"LiveCodeBench 63.9%（官方）。","reasoning":"GPQA-Diamond 71.2%（官方）。","math":"AIME 2025 76.8%（官方）。","agent":"BFCL v3 78.3、τ-Bench 54.7（官方）。","chinese":"中文强。"},"complete":false,"id":"hunyuan-a13b","name":"Hunyuan-A13B","name_zh":"混元 A13B · 80B-A13B","vendor":"Tencent","vendor_zh":"腾讯","family":"Hunyuan","license":"Tencent Hunyuan Community License","license_commercial":"restricted","weights_url":"https://huggingface.co/tencent/Hunyuan-A13B-Instruct","superseded_by":"hunyuan-hy3","released_at":"2025-06-27","architecture":{"type":"moe","total_params":"80B","active_params":"13B","total_params_b":80,"active_params_b":13,"experts":64,"active_experts":8,"shared_expert":true,"layers":32,"hidden_size":4096,"vocab_size":128167,"kv_heads":8,"head_dim":128,"attention":"GQA（32 Q 头 / 8 KV 头）","notes":"64 专家 top-8 + 1 共享；快慢思考双模式；256K 上下文；20T token。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K"},"memory":{"weight_gb":{"bf16":160,"q8":84.8,"q4":48.8,"fp8":80},"estimated":false,"kv_per_token_kib":128,"kv_note":"128 KiB/token。","ref_hw_24gb":"不可行（Q4 48.8 GB）","ref_hw_80gb":"Q4 / FP8 单卡 80GB 可跑","ref_hw_8x80gb":"BF16 高并发"},"links":{"official":"https://hunyuan.tencent.com/","github":"https://github.com/Tencent-Hunyuan/Hunyuan-A13B","paper":"https://arxiv.org/abs/2508.08088","hf":"https://huggingface.co/tencent/Hunyuan-A13B-Instruct"},"copy":{"one_liner":"13B 激活的 MoE，单卡 80GB 跑 256K 上下文与双模思考。","highlights":["快 / 慢思考切换（/no_think），256K 上下文","AIME 2025 76.8%、GPQA-Diamond 71.2%（官方）","单张 80GB 卡可部署，腾讯云原生"],"pitfalls":["许可禁止欧盟 / 英国 / 韩国使用，月活 1 亿限制","社区生态与量化明显少于 Qwen","BF16 160 GB，24GB 卡无解"],"logic_ability":"慢思考模式：AIME 2025 76.8%、GPQA-Diamond 71.2%、LiveCodeBench 63.9%（官方）。agent 榜 BFCL v3 78.3 突出。","best_for":["单卡 80GB 中文 agent","腾讯云生态"],"not_for":["欧盟合规场景","消费级显卡"]},"ecosystem":{"engines":["vLLM","SGLang","TensorRT-LLM","llama.cpp"],"finetune":"LLaMA-Factory 支持","zh_docs":"有"}},{"id":"hunyuan-hy3","name":"Hunyuan Hy3","name_zh":"腾讯混元 Hy3","aliases":["Hy3","tencent/Hy3","混元 Hy3"],"vendor":"Tencent","vendor_zh":"腾讯","family":"Hunyuan","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/tencent/Hy3","status":"current","released_at":"2026-07-06","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"295B","active_params":"21B","total_params_b":295,"active_params_b":21,"experts":192,"active_experts":8,"shared_expert":true,"layers":80,"hidden_size":4096,"vocab_size":120832,"kv_heads":8,"head_dim":128,"attention":"GQA（64 Q / 8 KV）","notes":"另带 3.8B 多 token 预测层（MTP）。快慢思考混合。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K"},"memory":{"weight_gb":{"bf16":598,"fp8":299,"q4":173,"q8":317.7},"kv_per_token_kib":320,"estimated":true,"ref_hw_80gb":"不可行","ref_hw_8x80gb":"FP8 8 卡（官方推荐 H20-3e 等大显存卡）"},"pricing":{"input_per_m":0.14,"output_per_m":0.58,"currency":"USD","source":"第三方托管常见价（OpenRouter）","as_of":"2026-08-28","note":"腾讯云官方价 ¥1 / ¥4 每百万 token"},"links":{"official":"https://www.tencent.com/en-us/articles/2202386.html","hf":"https://huggingface.co/tencent/Hy3"},"variants":[{"kind":"fp8","publisher":"tencent","repo":"tencent/Hy3-FP8","url":"https://huggingface.co/tencent/Hy3-FP8","note":"官方 FP8","sizes":{"fp8":299.9}},{"kind":"nvfp4","publisher":"RedHatAI","repo":"RedHatAI/Hy3-NVFP4-FP8","url":"https://huggingface.co/RedHatAI/Hy3-NVFP4-FP8"},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/Hy3-GGUF","url":"https://huggingface.co/bartowski/Hy3-GGUF","sizes":{"q4":182.2,"q5":212.8,"q6":257.2,"q8":317.7}},{"kind":"awq","publisher":"cyankiwi","repo":"cyankiwi/Hy3-AWQ-INT4","url":"https://huggingface.co/cyankiwi/Hy3-AWQ-INT4"}],"copy":{"one_liner":"腾讯 295B MoE 开源，SWE-bench 78%、GPQA 90。","highlights":["SWE-bench Verified 78%、GPQA Diamond 90.4%、Terminal-Bench 2.1 71.7%（官方）","Apache-2.0，21B 激活，256K 上下文","国产模型，中文与微信 / 腾讯生态原生"],"pitfalls":["295B 需 8 卡，官方还建议 H20-3e 级大显存","Arena 1456 排 59 名，对话体验不及榜单数字","文本 only，无视觉"],"logic_ability":"考试型推理顶级（GPQA 90.4），工程型推理同样一流（SWE-bench 78%、TB 71.7%），AA 智能指数 42 居开源前列。","best_for":["国产合规编程 / 办公 agent","腾讯云部署"],"not_for":["单机部署","多模态"]},"capability_notes":{"coding":"SWE-bench Verified 78%、Terminal-Bench 2.1 71.7%（官方）。","reasoning":"GPQA Diamond 90.4%（官方）。","chinese":"强，国产模型主场。"},"ecosystem":{"engines":["vLLM","SGLang","TensorRT-LLM"],"zh_docs":"有"},"complete":false,"runtime":{"tok_s":67,"latency_s":2.81,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"Tencent Hunyuan Hy3","one_liner":"Tencent's open-source 295B MoE: SWE-bench 78%, GPQA 90.","highlights":["SWE-bench Verified 78%, GPQA Diamond 90.4%, Terminal-Bench 2.1 71.7% (official)","Apache-2.0, 21B active, 256K context","Chinese domestic model, native to Chinese and the WeChat / Tencent ecosystem"],"pitfalls":["295B needs 8 GPUs; official guidance also recommends H20-3e-class large memory","Arena 1456, ranked 59th; chat experience falls short of the leaderboard numbers","Text only, no vision"],"logic_ability":"Top-tier exam-style reasoning (GPQA 90.4) and equally first-rate engineering reasoning (SWE-bench 78%, TB 71.7%); AA Intelligence Index 42, among the open-source leaders.","best_for":["China-compliant coding / office agents","Tencent Cloud deployment"],"not_for":["Single-machine deployment","Multimodal"],"capability_notes":{"coding":"SWE-bench Verified 78%, Terminal-Bench 2.1 71.7% (official).","reasoning":"GPQA Diamond 90.4% (official).","chinese":"Strong; home turf for a Chinese domestic model."}},"ja":{"name_zh":"Tencent 混元 Hy3","one_liner":"Tencent 295B MoE。SWE-bench 78%、GPQA 90。","highlights":["SWE-bench Verified 78%、GPQA Diamond 90.4%、Terminal-Bench 2.1 71.7%（公式）","Apache-2.0、21B アクティブ、256K コンテキスト","中国産モデルで、中国語と WeChat / Tencent エコシステムにネイティブ対応"],"pitfalls":["295B は 8 GPU が必要で、公式は H20-3e 級の大容量メモリを推奨","Arena 1456 で 59 位、対話体験はリーダーボードの数値に届かない","テキストのみ、ビジョン非対応"],"logic_ability":"試験型推論は最上位（GPQA 90.4）、エンジニアリング型推論も一流（SWE-bench 78%、TB 71.7%）。AA 知能指数 42 でオープンソース上位。","best_for":["中国のコンプライアンスに沿ったコーディング / オフィスエージェント","Tencent Cloud へのデプロイ"],"not_for":["単一マシンでのデプロイ","マルチモーダル"],"capability_notes":{"coding":"SWE-bench Verified 78%、Terminal-Bench 2.1 71.7%（公式）。","reasoning":"GPQA Diamond 90.4%（公式）。","chinese":"強い。中国産モデルのホームグラウンド。"}}}},{"aliases":["internlm2_5-20b-chat"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"math":"MATH 64.7%（官方）。","knowledge":"MMLU 73.5%（官方）。","agent":"工具调用与 Lagent 框架集成。","chinese":"中文强。"},"complete":false,"id":"internlm2-5-20b","name":"InternLM2.5-20B","name_zh":"书生·浦语 2.5 · 20B","vendor":"Shanghai AI Lab","vendor_zh":"上海人工智能实验室","family":"InternLM2.5","license":"Apache-2.0（权重需申请商用授权，免费）","license_commercial":"restricted","weights_url":"https://huggingface.co/internlm/internlm2_5-20b-chat","released_at":"2024-08-05","architecture":{"type":"dense","total_params":"19.9B","total_params_b":19.9,"layers":48,"hidden_size":6144,"vocab_size":92544,"kv_heads":8,"head_dim":128,"attention":"GQA（48 Q 头 / 8 KV 头）","notes":"1M 上下文能力（需 LMDeploy）；工具调用与 MindSearch 集成。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K（1M 需 LMDeploy）"},"memory":{"weight_gb":{"bf16":39.8,"q8":21.1,"q4":12},"estimated":false,"kv_per_token_kib":192,"kv_note":"192 KiB/token。","ref_hw_24gb":"Q4 12 GB 可跑 32K；BF16 39.8 GB 超 24GB","ref_hw_80gb":"BF16 单卡长上下文","ref_hw_8x80gb":"过剩"},"links":{"official":"https://internlm.intern-ai.org.cn/","github":"https://github.com/InternLM/InternLM","hf":"https://huggingface.co/internlm/internlm2_5-20b-chat"},"copy":{"one_liner":"上海 AI Lab 的 20B，主打百万上下文与工具调用。","highlights":["1M 上下文（LMDeploy），大海捞针接近满分","MATH 64.7%（官方）当年 20B 领先","LMDeploy / XTuner / OpenCompass 全套自研工具链"],"pitfalls":["商用需邮件申请授权（免费）","国际生态与社区量化较少","InternLM3 仅出 8B，20B 无后续"],"logic_ability":"20B 中上：MATH 64.7%、MMLU 73.5%（官方）。","best_for":["中文长文本","国产工具链部署"],"not_for":["需 Apache 无条件商用"]},"ecosystem":{"engines":["LMDeploy","vLLM","llama.cpp","Ollama"],"finetune":"XTuner / LLaMA-Factory","zh_docs":"有"}},{"id":"kimi-k2-0905","name":"Kimi K2 Instruct (0905)","name_zh":"月之暗面 K2（0905）","aliases":["kimi-k2-0905","Kimi K2"],"vendor":"Moonshot AI","vendor_zh":"月之暗面","family":"Kimi K2","license":"Modified MIT","license_commercial":"restricted","openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/moonshotai/Kimi-K2-Instruct-0905","status":"superseded","released_at":"2025-09-05","updated_at":"2025-12-20","modalities":["text","tools"],"reasoning_mode":"none","architecture":{"type":"moe","total_params":"1T","active_params":"32B","total_params_b":1000,"active_params_b":32,"experts":384,"active_experts":8,"shared_expert":true,"layers":61,"hidden_size":7168,"kv_heads":1,"head_dim":288,"attention":"MLA","undisclosed":false},"context":{"max_tokens":262144,"display":"256K"},"memory":{"weight_gb":{"bf16":2000,"fp8":1000,"q4":580},"kv_per_token_kib":70,"estimated":true,"ref_hw_8x80gb":"FP8 需 16×H100 或 8×H200"},"pricing":{"input_per_m":0.6,"output_per_m":2.5,"currency":"USD","source":"Moonshot 开放平台","as_of":"2025-12-20"},"links":{"official":"https://moonshotai.github.io/Kimi-K2/","hf":"https://huggingface.co/moonshotai/Kimi-K2-Instruct-0905","github":"https://github.com/MoonshotAI/Kimi-K2","paper":"https://arxiv.org/abs/2507.20534"},"copy":{"one_liner":"非思考的 1T agent 模型，编程 agent 与前端生成强。","highlights":["SWE-bench Verified 69.2%（官方）","256K 上下文","非思考：响应快"],"pitfalls":["1T 体量","Modified MIT 标注条款","深度推理用 Thinking 版"],"logic_ability":"工程型推理强，无思考链，考试型推理中等。","best_for":["编程 agent","前端生成"],"not_for":["任何单机"]},"capability_notes":{},"ecosystem":{"engines":["vLLM","SGLang","KTransformers"],"zh_docs":"有"},"complete":false,"superseded_by":"kimi-k2-6"},{"id":"kimi-k2-6","name":"Kimi K2.6","name_zh":"月之暗面 Kimi K2.6","aliases":["kimi-k2.6","K2.6","Kimi-K2.6"],"vendor":"Moonshot AI","vendor_zh":"月之暗面","family":"Kimi K2","license":"Modified MIT","license_commercial":"restricted","openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/moonshotai/Kimi-K2.6","status":"current","released_at":"2026-04-20","updated_at":"2026-08-28","modalities":["text","image","video","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"1T","active_params":"32B","total_params_b":1000,"active_params_b":32,"experts":384,"active_experts":8,"shared_expert":true,"layers":61,"hidden_size":7168,"vocab_size":160000,"kv_heads":1,"head_dim":288,"attention":"MLA（64 头）","notes":"与 Kimi K2 / K2.5 同构（DeepSeek-V3 式 MLA + 384 专家），加 MoonViT 视觉编码器（400M），新增视频输入。原生 INT4 QAT 发布。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K"},"memory":{"weight_gb":{"q4":595,"fp8":1030.9,"bf16":2053.2},"kv_per_token_kib":70,"kv_note":"MLA，与 K2 同：≈ 70 KiB/token。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行","ref_hw_8x80gb":"INT4 595 GB，8×H100 勉强，推荐 8×H200","estimated":false},"pricing":{"input_per_m":0.95,"output_per_m":4,"currency":"USD","source":"Kimi 开放平台","as_of":"2026-08-28","note":"缓存命中输入 $0.16"},"links":{"official":"https://www.kimi.com/","hf":"https://huggingface.co/moonshotai/Kimi-K2.6","github":"https://github.com/MoonshotAI/Kimi-K2.5","paper":"https://arxiv.org/abs/2602.02276","pricing":"https://platform.kimi.ai/docs/pricing/chat-k26"},"variants":[{"kind":"other","publisher":"moonshotai","repo":"moonshotai/Kimi-K2.6","url":"https://huggingface.co/moonshotai/Kimi-K2.6","note":"官方原生 INT4（compressed-tensors）","sizes":{"q4":595.2}},{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/Kimi-K2.6-NVFP4","url":"https://huggingface.co/nvidia/Kimi-K2.6-NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/Kimi-K2.6-GGUF","url":"https://huggingface.co/unsloth/Kimi-K2.6-GGUF","sizes":{"bf16":2053.2}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/moonshotai_Kimi-K2.6-GGUF","url":"https://huggingface.co/bartowski/moonshotai_Kimi-K2.6-GGUF"},{"kind":"fp8","publisher":"RedHatAI","repo":"RedHatAI/Kimi-K2.6-FP8-BLOCK","url":"https://huggingface.co/RedHatAI/Kimi-K2.6-FP8-BLOCK","sizes":{"fp8":1030.9}}],"copy":{"one_liner":"1T/32B 多模态 agent 模型，K3 之下的性价比选择。","highlights":["SWE-Bench Verified 80.2、SWE-Bench Pro 58.6、LiveCodeBench v6 89.6（官方）","Agent Swarm 可并行 300 子 agent / 4000 步；支持图像与视频输入","思考可关（Instant 模式），API 价仅 K3 的 1/3"],"pitfalls":["Modified MIT：月活 > 1 亿或月收入 > 2000 万美元须标注「Kimi K2.6」","INT4 595 GB 仍需 8×H100 以上","已被 K3 取代旗舰位置，官方主推 K3 / K2.7 Code"],"logic_ability":"GPQA 90.5、HLE-Full 34.7 / 54.0（工具）、AIME 2026 96.4、BrowseComp 83.2。LMArena 1461。","best_for":["预算受限的多模态 agent","K2 系列用户平滑升级"],"not_for":["单机部署","需要 1M 上下文（用 K3）"]},"capability_notes":{"coding":"SWE-Bench Verified 80.2、SWE-Bench Pro 58.6、Terminal-Bench 2.0 66.7、LiveCodeBench v6 89.6（官方）。","reasoning":"GPQA Diamond 90.5、HLE-Full 34.7 / 54.0（工具）（官方）。","math":"AIME 2026 96.4、HMMT Feb 2026 92.7（官方）。","agent":"BrowseComp 83.2、OSWorld-Verified 73.1、Toolathlon 50.0（官方）。","multimodal":"MMMU-Pro 79.4、MathVision 87.4（官方）。"},"ecosystem":{"engines":["vLLM","SGLang","KTransformers"],"finetune":"基本不可行","zh_docs":"有"},"complete":false,"runtime":{"tok_s":39,"latency_s":2.75,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"Moonshot Kimi K2.6","one_liner":"1T/32B multimodal agent model; the value pick below K3.","highlights":["SWE-Bench Verified 80.2, SWE-Bench Pro 58.6, LiveCodeBench v6 89.6 (official)","Agent Swarm can run 300 sub-agents / 4000 steps in parallel; supports image and video input","Thinking can be turned off (Instant mode); API price only 1/3 of K3"],"pitfalls":["Modified MIT: over 100M MAU or $20M monthly revenue requires displaying \"Kimi K2.6\"","INT4 595 GB still needs 8×H100 or more","Flagship spot taken by K3; officially K3 / K2.7 Code are the recommended models"],"logic_ability":"GPQA 90.5, HLE-Full 34.7 / 54.0 (with tools), AIME 2026 96.4, BrowseComp 83.2. LMArena 1461.","best_for":["Budget-constrained multimodal agents","Smooth upgrade for K2-series users"],"not_for":["Single-machine deployment","Needing 1M context (use K3)"],"capability_notes":{"coding":"SWE-Bench Verified 80.2, SWE-Bench Pro 58.6, Terminal-Bench 2.0 66.7, LiveCodeBench v6 89.6 (official).","reasoning":"GPQA Diamond 90.5, HLE-Full 34.7 / 54.0 (with tools) (official).","math":"AIME 2026 96.4, HMMT Feb 2026 92.7 (official).","agent":"BrowseComp 83.2, OSWorld-Verified 73.1, Toolathlon 50.0 (official).","multimodal":"MMMU-Pro 79.4, MathVision 87.4 (official)."}},"ja":{"name_zh":"Moonshot Kimi K2.6","one_liner":"1T/32B マルチモーダルエージェント。K3 の下のコスパ選択肢。","highlights":["SWE-Bench Verified 80.2、SWE-Bench Pro 58.6、LiveCodeBench v6 89.6（公式）","Agent Swarm で 300 サブエージェント / 4000 ステップを並列実行可能。画像・動画入力対応","思考オフ可（Instant モード）、API 価格は K3 の 1/3"],"pitfalls":["Modified MIT：MAU 1 億超または月収 2000 万ドル超の場合「Kimi K2.6」の表示が必要","INT4 595 GB でも 8×H100 以上が必要","フラッグシップの座は K3 に移り、公式の推奨は K3 / K2.7 Code"],"logic_ability":"GPQA 90.5、HLE-Full 34.7 / 54.0（ツールあり）、AIME 2026 96.4、BrowseComp 83.2。LMArena 1461。","best_for":["予算が限られたマルチモーダルエージェント","K2 シリーズユーザーのスムーズなアップグレード"],"not_for":["単一マシンでのデプロイ","1M コンテキストが必要な場合（K3 を使う）"],"capability_notes":{"coding":"SWE-Bench Verified 80.2、SWE-Bench Pro 58.6、Terminal-Bench 2.0 66.7、LiveCodeBench v6 89.6（公式）。","reasoning":"GPQA Diamond 90.5、HLE-Full 34.7 / 54.0（ツールあり）（公式）。","math":"AIME 2026 96.4、HMMT Feb 2026 92.7（公式）。","agent":"BrowseComp 83.2、OSWorld-Verified 73.1、Toolathlon 50.0（公式）。","multimodal":"MMMU-Pro 79.4、MathVision 87.4（公式）。"}}}},{"id":"kimi-k2-7-code","name":"Kimi K2.7 Code","name_zh":"月之暗面 Kimi K2.7 Code","aliases":["kimi-k2.7-code","K2.7 Code","Kimi-K2.7-Code"],"vendor":"Moonshot AI","vendor_zh":"月之暗面","family":"Kimi K2","license":"Modified MIT","license_commercial":"restricted","openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/moonshotai/Kimi-K2.7-Code","status":"current","released_at":"2026-06-11","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"default-on","architecture":{"type":"moe","total_params":"1T","active_params":"32B","total_params_b":1000,"active_params_b":32,"experts":384,"active_experts":8,"shared_expert":true,"layers":61,"hidden_size":7168,"vocab_size":160000,"kv_heads":1,"head_dim":288,"attention":"MLA（64 头）","notes":"基于 K2.6 的编程特化后训练版，结构完全相同；强制思考 + preserve_thinking，思考 token 比 K2.6 少约 30%。原生 INT4。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K"},"memory":{"weight_gb":{"q4":595},"kv_per_token_kib":70,"kv_note":"MLA，≈ 70 KiB/token。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行","ref_hw_8x80gb":"INT4 595 GB，推荐 8×H200","estimated":false},"pricing":{"input_per_m":0.95,"output_per_m":4,"currency":"USD","source":"Kimi 开放平台","as_of":"2026-08-28","note":"缓存命中 $0.19；highspeed 版 $1.9 / $8（约 180 tok/s）"},"links":{"official":"https://www.kimi.com/code","hf":"https://huggingface.co/moonshotai/Kimi-K2.7-Code","pricing":"https://platform.kimi.ai/docs/pricing/chat-k27-code"},"variants":[{"kind":"other","publisher":"moonshotai","repo":"moonshotai/Kimi-K2.7-Code","url":"https://huggingface.co/moonshotai/Kimi-K2.7-Code","note":"官方原生 INT4（compressed-tensors）","sizes":{"q4":595.2}},{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/Kimi-K2.7-Code-NVFP4","url":"https://huggingface.co/nvidia/Kimi-K2.7-Code-NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/Kimi-K2.7-Code-GGUF","url":"https://huggingface.co/unsloth/Kimi-K2.7-Code-GGUF"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/Kimi-K2.7-Code-4bit","url":"https://huggingface.co/mlx-community/Kimi-K2.7-Code-4bit"}],"copy":{"one_liner":"K2.6 编程特化版，长程编程任务成功率更高。","highlights":["Kimi Code Bench v2 62.0（K2.6 50.9）、MCP Mark Verified 81.1","思考 token 比 K2.6 少约 30%，highspeed 版 180–260 tok/s","AA 指数 43；与 K2.6 同权重规模，INT4 595 GB"],"pitfalls":["纯文本，不带视觉；思考与 preserve_thinking 强制开启不可关","Modified MIT 品牌标注条款同 K2.6","官方仅披露自家 Kimi Code Bench 等少量基准"],"logic_ability":"工程型推理专用：ProgramBench 53.6、MLS Bench Lite 35.1，官方对比仍低于 GPT-5.5 / Opus 4.8。","best_for":["Kimi Code CLI / Claude Code 后端","长程多文件编程"],"not_for":["通用对话与多模态","单机部署"]},"capability_notes":{"coding":"Kimi Code Bench v2 62.0、ProgramBench 53.6（官方）。","agent":"MCP Atlas 76.0、MCP Mark Verified 81.1（官方）。"},"ecosystem":{"engines":["vLLM","SGLang","KTransformers"],"finetune":"基本不可行","zh_docs":"有"},"complete":false,"runtime":{"tok_s":45,"latency_s":2.82,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"Moonshot Kimi K2.7 Code","one_liner":"Coding-specialized K2.6 with higher success on long-horizon coding tasks.","highlights":["Kimi Code Bench v2 62.0 (K2.6: 50.9), MCP Mark Verified 81.1","About 30% fewer thinking tokens than K2.6; highspeed version 180–260 tok/s","AA index 43; same weight size as K2.6, INT4 595 GB"],"pitfalls":["Text only, no vision; thinking and preserve_thinking are forced on and cannot be disabled","Modified MIT branding clause same as K2.6","Officially discloses only a few benchmarks such as its own Kimi Code Bench"],"logic_ability":"Dedicated to engineering reasoning: ProgramBench 53.6, MLS Bench Lite 35.1; official comparisons still below GPT-5.5 / Opus 4.8.","best_for":["Backend for Kimi Code CLI / Claude Code","Long-horizon multi-file coding"],"not_for":["General chat and multimodal","Single-machine deployment"],"capability_notes":{"coding":"Kimi Code Bench v2 62.0, ProgramBench 53.6 (official).","agent":"MCP Atlas 76.0, MCP Mark Verified 81.1 (official)."}},"ja":{"name_zh":"Moonshot Kimi K2.7 Code","one_liner":"K2.6 のコーディング特化版。長期コーディングタスクの成功率が向上。","highlights":["Kimi Code Bench v2 62.0（K2.6 は 50.9）、MCP Mark Verified 81.1","思考トークンは K2.6 より約 30% 少なく、highspeed 版は 180–260 tok/s","AA 指数 43。K2.6 と同じ重みサイズで INT4 595 GB"],"pitfalls":["テキストのみでビジョン非対応。思考と preserve_thinking は強制オンで無効化不可","Modified MIT のブランド表示条項は K2.6 と同じ","公式が開示するのは自社の Kimi Code Bench など少数のベンチマークのみ"],"logic_ability":"エンジニアリング型推論専用：ProgramBench 53.6、MLS Bench Lite 35.1。公式比較でも GPT-5.5 / Opus 4.8 には及ばない。","best_for":["Kimi Code CLI / Claude Code のバックエンド","長期のマルチファイルコーディング"],"not_for":["汎用対話とマルチモーダル","単一マシンでのデプロイ"],"capability_notes":{"coding":"Kimi Code Bench v2 62.0、ProgramBench 53.6（公式）。","agent":"MCP Atlas 76.0、MCP Mark Verified 81.1（公式）。"}}}},{"aliases":["Kimi-K2-Instruct","Kimi K2 (July 2025)"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"coding":"SWE-bench Verified 65.8%、LiveCodeBench v6 53.7%（官方）。","reasoning":"GPQA Diamond 75.1%、HLE 4.7%（官方）。","math":"AIME 2025 49.5%（官方）。","agent":"Tau2 retail 70.6 / airline 56.5 / telecom 65.8，Terminal-bench 30.0（官方）。","chinese":"中文顶级。"},"complete":false,"id":"kimi-k2-instruct","name":"Kimi K2 Instruct","name_zh":"Kimi K2（0711 原版）","vendor":"Moonshot AI","vendor_zh":"月之暗面","family":"Kimi K2","license":"Modified MIT License","license_commercial":"restricted","weights_url":"https://huggingface.co/moonshotai/Kimi-K2-Instruct","superseded_by":"kimi-k2-0905","released_at":"2025-07-11","architecture":{"type":"moe","total_params":"1T","active_params":"32B","total_params_b":1026,"active_params_b":32,"experts":384,"active_experts":8,"shared_expert":true,"layers":61,"hidden_size":7168,"vocab_size":163840,"kv_heads":64,"head_dim":192,"attention":"MLA（kv_lora_rank 512 + rope 64，64 头）","notes":"DeepSeek-V3 式结构放大：384 专家 top-8 + 1 共享；MuonClip 优化器；128K；非推理。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":2052,"q8":1087.6,"q4":620.8,"fp8":1026},"estimated":false,"kv_per_token_kib":68.6,"kv_note":"MLA ≈ 68.6 KiB/token。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行","ref_hw_8x80gb":"官方 FP8 ≈ 1 TB 需 16×80GB 或 8×H200；Q4 621 GB 8×80GB 勉强"},"pricing":{"input_per_m":0.6,"output_per_m":2.5,"currency":"USD","source":"Moonshot 官方 API（kimi-k2）","as_of":"2026-08-28"},"links":{"official":"https://moonshotai.github.io/Kimi-K2/","github":"https://github.com/MoonshotAI/Kimi-K2","paper":"https://arxiv.org/abs/2507.20534","hf":"https://huggingface.co/moonshotai/Kimi-K2-Instruct"},"copy":{"one_liner":"首个万亿参数开放权重，非推理 agent 模型的高峰。","highlights":["1T 总参 / 32B 激活，SWE-bench Verified 65.8%（官方）","非思考模式即达强 agent 能力，延迟低","MuonClip 优化器，15.5T token 训练零 loss spike"],"pitfalls":["Modified MIT：月活 1 亿或月收 2000 万美元以上需标注 Kimi K2","1T 权重本地几乎不可行","0905 版把上下文扩到 256K 并提升前端代码，建议直接用新版"],"logic_ability":"非推理模型中顶级：GPQA 75.1%、AIME 2025 49.5%、LiveCodeBench v6 53.7%、SWE-bench Verified 65.8%（官方）。工具调用与多步 agent 稳定，但没有思考链，深度数学不及推理模型。","best_for":["agent / 工具调用后端","代码 agent"],"not_for":["本地部署","竞赛数学"]},"ecosystem":{"engines":["vLLM","SGLang","KTransformers","TensorRT-LLM"],"finetune":"极困难","zh_docs":"有"}},{"id":"kimi-k2-thinking","name":"Kimi K2 Thinking","name_zh":"月之暗面 K2 推理版","aliases":["kimi-k2-thinking","K2 Thinking"],"vendor":"Moonshot AI","vendor_zh":"月之暗面","family":"Kimi K2","license":"Modified MIT","license_commercial":"restricted","openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/moonshotai/Kimi-K2-Thinking","status":"superseded","released_at":"2025-11-06","updated_at":"2025-12-20","modalities":["text","tools"],"reasoning_mode":"default-on","architecture":{"type":"moe","total_params":"1T","active_params":"32B","total_params_b":1000,"active_params_b":32,"experts":384,"active_experts":8,"shared_expert":true,"layers":61,"hidden_size":7168,"vocab_size":160000,"kv_heads":1,"head_dim":288,"attention":"MLA（64 头）","notes":"DeepSeek-V3 同构放大：384 专家 / 8 激活 + 1 共享，MLA。原生 INT4 QAT 发布（权重即 4 bit）。MuonClip 优化器。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K","max_output":65536},"memory":{"weight_gb":{"q4":594},"kv_per_token_kib":70,"kv_note":"MLA，与 DeepSeek-V3 同：≈ 70 KiB/token。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行","ref_hw_8x80gb":"INT4 权重 594 GB，8×H100 80GB 勉强（KV 余量小），推荐 8×H200","estimated":false},"pricing":{"input_per_m":0.6,"output_per_m":2.5,"currency":"USD","source":"Moonshot 开放平台","as_of":"2025-12-20","note":"缓存命中 $0.15"},"links":{"official":"https://moonshotai.github.io/Kimi-K2/thinking.html","hf":"https://huggingface.co/moonshotai/Kimi-K2-Thinking","github":"https://github.com/MoonshotAI/Kimi-K2","paper":"https://arxiv.org/abs/2507.20534","pricing":"https://platform.moonshot.ai/docs/pricing"},"copy":{"one_liner":"1T 参数、原生 INT4 的开源 agent 推理旗舰，HLE 开源第一。","highlights":["HLE（带工具）44.9%，发布时超过所有闭源模型","官方称可连续 200–300 次工具调用不漂移，agent 长任务稳定","原生 INT4 QAT：权重 594 GB 即为「全精度」，推理速度翻倍"],"pitfalls":["Modified MIT：月活 > 1 亿或月收入 > 2000 万美元的产品须在界面标注「Kimi K2」","1T 体量，8×H100 部署 KV 余量很小，实际要 H200","思考模式 token 消耗大，简单任务性价比差（用 K2-0905 非思考版）"],"logic_ability":"工程型与工具增强推理是核心：BrowseComp 60.2%、SWE-bench Verified 71.3%。考试型推理带工具时极强（AIME 99.1% with Python），无工具时略低于 GPT-5。思考链交错工具调用（interleaved thinking）是其特色。常见问题：无工具纯文本对话时偶尔过度冗长；数学题喜欢调 Python 而非心算。","best_for":["深度搜索 / 研究型 agent","长链工具调用","自建高端推理服务"],"not_for":["任何单机部署","简单对话（成本高）"]},"capability_notes":{"coding":"SWE-bench Verified 71.3%、LiveCodeBench v6 83.1%（官方）。","reasoning":"HLE 44.9%（工具）、GPQA 84.5%（官方）。","math":"AIME 2025 99.1%（Python）、94.5%（无工具，官方）。","agent":"BrowseComp 60.2%、τ²-bench 74.3%（官方）。","chinese":"中文顶级，BrowseComp-ZH 62.3%。"},"sheet":{"architecture_md":"**类型**：MoE，1.04T 总 / 32B 激活。\n\n- 61 层（含 1 层 Dense），隐藏维 7168，词表 160K\n- 专家：384 路由 + 1 共享，每 token 激活 8 个\n- 注意力：MLA，64 头，KV 潜向量 512 + RoPE 64\n- 与 DeepSeek-V3 同构但专家数 ×1.5、注意力头数减半（为降低长上下文推理开销）\n- 上下文 256K\n- 训练用 MuonClip 优化器解决 attention logit 爆炸\n- **原生 INT4 QAT**：后训练阶段做量化感知训练，MoE 部分权重以 INT4 发布\n\n参考：Kimi K2 技术报告（arXiv 2507.20534）。","memory_md":"| 精度 | 权重大小 | 参考显存 |\n|---|---|---|\n| INT4（原生） | 594 GB | 8×H200 |\n\n无 BF16 发布版本（INT4 即官方精度）。\n\n**KV Cache**：≈ 70 KiB/token（MLA）。\n\n**参考配置**：\n- 24GB / 80GB：不可行\n- 8×H100 80GB：合计 640 GB，权重后仅剩 ~46 GB 给 KV，并发受限\n- 8×H200 141GB：舒适","training_md":"- 预训练 15.5T token（K2 base）\n- 后训练：大规模 agentic 数据合成 + 联合 RL，再加 QAT\n- 思考默认开启，不可关（非思考用 Kimi-K2-Instruct-0905）\n- 工具调用：原生，支持交错思考\n- 最大输出 64K","ecosystem_md":"- HF：moonshotai/Kimi-K2-Thinking\n- 引擎：vLLM、SGLang、KTransformers（官方部署指南）\n- 微调：体量过大，社区几乎无\n- 中文文档：有","versions_md":"- 同系列：Kimi-K2-Instruct（2025-07）、Kimi-K2-Instruct-0905（256K）\n- 视觉：Kimi-VL 系列另计\n- 上代：Kimi K1.5（闭源）"},"ecosystem":{"engines":["vLLM","SGLang","KTransformers"],"finetune":"基本不可行","zh_docs":"有"},"complete":true,"superseded_by":"kimi-k3"},{"id":"kimi-k3","name":"Kimi K3","name_zh":"月之暗面 Kimi K3","aliases":["kimi-k3","K3","Kimi-K3"],"vendor":"Moonshot AI","vendor_zh":"月之暗面","family":"Kimi K3","license":"Kimi K3 License（Modified MIT）","license_commercial":"restricted","openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/moonshotai/Kimi-K3","status":"current","released_at":"2026-07-16","updated_at":"2026-08-28","modalities":["text","image","video","tools"],"reasoning_mode":"default-on","architecture":{"type":"hybrid","total_params":"2.8T","active_params":"104B","total_params_b":2800,"active_params_b":104,"experts":896,"active_experts":16,"shared_expert":true,"layers":93,"kv_layers":24,"hidden_size":7168,"vocab_size":163840,"kv_heads":1,"head_dim":288,"attention":"KDA 线性注意力 69 层 + Gated MLA 24 层（96 头）","notes":"首个开源 3T 级模型。93 层（1 层 Dense）中 69 层为 Kimi Delta Attention（线性，固定状态），24 层为带输出门的 MLA（kv_lora 512 + RoPE 64，按「1 KV 头 × 288」等价表示）。Attention Residuals（AttnRes）、SiTU-GLU 激活、Latent MoE（潜维 3584）；2 个共享专家。原生 MXFP4 权重 / MXFP8 激活 QAT。视觉编码器 MoonViT-V2（401M）。","undisclosed":false},"context":{"max_tokens":1048576,"display":"1M"},"memory":{"weight_gb":{"q4":1561,"fp8":2819.9},"kv_per_token_kib":27,"kv_note":"仅 24 层 MLA 存 per-token KV：576 × 2 B × 24 ≈ 27 KiB/token；69 层 KDA 为固定状态（每层 96 头 × 128 × 128 × 2 B ≈ 3 MiB，每请求约 207 MiB 常量）。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行","ref_hw_8x80gb":"MXFP4 权重 1,561 GB 已超单节点 8×H200（1,128 GB），需 16×H200 / 16×H100 双节点或 8×B300 级","estimated":false},"pricing":{"input_per_m":3,"output_per_m":15,"currency":"USD","source":"Kimi 开放平台","as_of":"2026-08-28","note":"缓存命中输入 $0.30"},"links":{"official":"https://www.kimi.com/","hf":"https://huggingface.co/moonshotai/Kimi-K3","github":"https://github.com/MoonshotAI/Kimi-K3","pricing":"https://platform.kimi.ai/docs/pricing/chat-k3"},"variants":[{"kind":"other","publisher":"moonshotai","repo":"moonshotai/Kimi-K3","url":"https://huggingface.co/moonshotai/Kimi-K3","note":"官方原生 MXFP4 权重","sizes":{"q4":1560.9}},{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/Kimi-K3-NVFP4","url":"https://huggingface.co/nvidia/Kimi-K3-NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/Kimi-K3-GGUF","url":"https://huggingface.co/unsloth/Kimi-K3-GGUF"},{"kind":"fp8","publisher":"RedHatAI","repo":"RedHatAI/Kimi-K3-FP8-BLOCK","url":"https://huggingface.co/RedHatAI/Kimi-K3-FP8-BLOCK","sizes":{"fp8":2819.9}}],"copy":{"one_liner":"2.8T 参数的开源多模态 agent 旗舰，LMArena 开源第一。","highlights":["2.8T/104B 激活，KDA 线性 + Gated MLA 混合，1M 上下文；LMArena 文本榜 1489（第 10，开源最高）、AA 指数 60","Terminal-Bench 2.1 88.3、BrowseComp 91.2、GPQA 93.5，官方对标 Claude Fable 5 / GPT-5.6 Sol 仅小幅落后","原生 MXFP4 QAT 发布，1.56 TB 即「全精度」；文本 + 图像 + 视频原生多模态"],"pitfalls":["Kimi K3 License：月活 > 1 亿或月收入 > 2000 万美元须在界面标注「Kimi K3」；MaaS 年收入 > 2000 万美元须另签协议","1.56 TB 权重超出单节点 8×H200，自建需双节点，几乎只有云厂商能跑","API 价 $3/$15 是 K2.6 的 3–4 倍；思考不可关，max 档 token 消耗大；多轮必须原样回传 reasoning_content"],"logic_ability":"工程型与 agent 推理是核心：Terminal-Bench 2.1 88.3、DeepSWE 67.5、FrontierSWE 81.2、SWE-Marathon 42.0，BrowseComp 91.2 为公开最高。考试型推理顶级（GPQA 93.5、HLE-Full 43.5 / 带工具 56.0），但 CritPt 23.4 明显低于 GPT-5.6 Sol。思考默认开启，reasoning_effort 可调 low/high/max。常见问题：low 档 AA 指数掉到 48，差距大；无工具纯文本对话偏冗长。","best_for":["深度搜索 / 研究型 agent","长程编程 agent（Kimi Code CLI）","视觉 + 长上下文的文档 / 视频理解","云厂商级自建"],"not_for":["任何单节点部署","简单对话（成本高，用 K2.6）","超大 MAU 产品且不愿标注品牌"]},"capability_notes":{"coding":"Terminal-Bench 2.1 88.3、DeepSWE 67.5、ProgramBench 77.8、SciCode 58.7、FrontierSWE 81.2（官方）。","reasoning":"GPQA Diamond 93.5、HLE-Full 43.5 / 56.0（工具）、CritPt 23.4（官方）。","math":"官方未披露 AIME；MathVision 94.3 / 97.8（Python）。","agent":"BrowseComp 91.2、MCP-Atlas 84.2、Toolathlon-Verified 76.5、OSWorld-Verified 84.8、τ³-Banking 33.4（官方）。","multimodal":"MMMU-Pro 81.6 / 83.4（工具）、Video-MME 90.0、OmniDocBench 91.1（官方）。","chinese":"中文顶级。"},"sheet":{"architecture_md":"**类型**：混合线性注意力 MoE，2.8T 总 / 104B 激活。\n\n- 93 层（含 1 层 Dense）：69 层 **Kimi Delta Attention（KDA）** + 24 层 **Gated MLA**，每 3 层 KDA 接 1 层 MLA\n- 隐藏维 7168，96 注意力头，词表 163,840\n- 专家：896 路由 + 2 共享，每 token 激活 16 个；Latent MoE（潜维 3584），专家中间维 3072\n- MLA：kv_lora 512，qk_nope 128 + RoPE 64，v_head 128，带输出门（mla_use_output_gate）\n- KDA：head_dim 128，短卷积核 4，full-rank gate\n- **Attention Residuals（AttnRes）**：每 12 层一个注意力残差块\n- 激活：SiTU-GLU\n- 上下文 1,048,576\n- 视觉：MoonViT-V2（401M），patch 14，2×2 merge\n- **原生 MXFP4**：从 SFT 阶段起做 QAT，权重 MXFP4 / 激活 MXFP8\n\n参考：Kimi K3 技术报告（GitHub）。","memory_md":"| 精度 | 权重大小 | 参考显存 |\n|---|---|---|\n| MXFP4（原生） | 1,561 GB（官方文件） | 16×H200 |\n\n无 BF16 发布版本（MXFP4 即官方精度）。\n\n**KV Cache**：≈ 27 KiB/token（仅 24 层 MLA）+ 每请求约 207 MiB KDA 固定状态。1M 上下文单请求 KV ≈ 27 GB，比 GLM-5.2 小 3 倍。\n\n**参考配置**：\n- 24GB / 80GB / 单节点 8×H100：不可行\n- 8×H200 141GB（1,128 GB）：权重放不下\n- 16×H200 双节点：可服务；官方部署指南基于 vLLM / SGLang / TokenSpeed","training_md":"- 预训练 token 数官方未披露\n- 后训练：agentic RL；从 SFT 起 MXFP4 量化感知训练\n- 思考默认开启，不可关；`reasoning_effort` low / high / max（默认 max）\n- 多轮与工具调用要求原样回传含 `reasoning_content` 的完整 assistant 消息（preserved thinking）\n- 评测统一 temperature 1.0、effort max\n- API 最大输出官方未披露（OpenRouter 标注 1M）","ecosystem_md":"- HF：moonshotai/Kimi-K3\n- 引擎：vLLM、SGLang、TokenSpeed（官方 recipes）\n- Agent 框架：Kimi Code CLI（官方推荐）\n- 微调：体量过大，社区几乎无\n- 中文文档：有","versions_md":"- 发布：2026-07-16 API；2026-07-26 开源权重\n- 上代：Kimi K2.5（2026-01，多模态）、K2.6（2026-04，1T/32B）、K2.7 Code（2026-06，编程特化）\n- K2.5 于 2026-08-31 下线"},"ecosystem":{"engines":["vLLM","SGLang","TokenSpeed"],"finetune":"基本不可行","zh_docs":"有"},"complete":true,"runtime":{"tok_s":36,"latency_s":8.6,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"Moonshot Kimi K3","one_liner":"2.8T-parameter open multimodal agent flagship; #1 open-source on LMArena.","highlights":["2.8T / 104B active, KDA linear + Gated MLA hybrid, 1M context; LMArena text 1489 (#10, highest open-source), AA index 60","Terminal-Bench 2.1 88.3, BrowseComp 91.2, GPQA 93.5; officially benchmarked against Claude Fable 5 / GPT-5.6 Sol with only a small gap","Released natively as MXFP4 QAT, 1.56 TB counts as \"full precision\"; native text + image + video multimodal"],"pitfalls":["Kimi K3 License: over 100M MAU or $20M monthly revenue requires displaying \"Kimi K3\" in the UI; MaaS with annual revenue over $20M needs a separate agreement","1.56 TB weights exceed a single 8×H200 node; self-hosting needs two nodes, so practically only cloud providers can run it","API price $3/$15 is 3–4× K2.6; thinking cannot be disabled and the max tier burns many tokens; multi-turn must return reasoning_content verbatim"],"logic_ability":"Engineering and agent reasoning are the core: Terminal-Bench 2.1 88.3, DeepSWE 67.5, FrontierSWE 81.2, SWE-Marathon 42.0; BrowseComp 91.2 is the highest published. Exam-style reasoning is top-tier (GPQA 93.5, HLE-Full 43.5 / 56.0 with tools), but CritPt 23.4 is clearly below GPT-5.6 Sol. Thinking is on by default; reasoning_effort adjustable low/high/max. Common issues: at the low tier the AA index drops to 48, a big gap; tool-free plain-text chat tends to be verbose.","best_for":["Deep search / research agents","Long-horizon coding agents (Kimi Code CLI)","Vision + long-context document / video understanding","Cloud-provider-scale self-hosting"],"not_for":["Any single-node deployment","Simple chat (costly; use K2.6)","Very high-MAU products unwilling to display the brand"],"capability_notes":{"coding":"Terminal-Bench 2.1 88.3, DeepSWE 67.5, ProgramBench 77.8, SciCode 58.7, FrontierSWE 81.2 (official).","reasoning":"GPQA Diamond 93.5, HLE-Full 43.5 / 56.0 (with tools), CritPt 23.4 (official).","math":"AIME not officially disclosed; MathVision 94.3 / 97.8 (Python).","agent":"BrowseComp 91.2, MCP-Atlas 84.2, Toolathlon-Verified 76.5, OSWorld-Verified 84.8, τ³-Banking 33.4 (official).","multimodal":"MMMU-Pro 81.6 / 83.4 (with tools), Video-MME 90.0, OmniDocBench 91.1 (official).","chinese":"Top-tier Chinese."}},"ja":{"name_zh":"Moonshot Kimi K3","one_liner":"2.8T のオープン多モーダルエージェント旗艦。LMArena OSS 1 位。","highlights":["2.8T / 104B アクティブ、KDA 線形 + Gated MLA ハイブリッド、1M コンテキスト。LMArena テキスト榜 1489（10 位、オープンソース最高）、AA 指数 60","Terminal-Bench 2.1 88.3、BrowseComp 91.2、GPQA 93.5。公式は Claude Fable 5 / GPT-5.6 Sol と比較し、わずかな差にとどまる","ネイティブ MXFP4 QAT で公開、1.56 TB が「フル精度」。テキスト + 画像 + 動画のネイティブマルチモーダル"],"pitfalls":["Kimi K3 License：MAU 1 億超または月収 2000 万ドル超は UI に「Kimi K3」を表示する必要あり。MaaS で年収 2000 万ドル超は別途契約が必要","1.56 TB の重みは単一 8×H200 ノードを超え、自前ホスティングには 2 ノード必要。実質クラウド事業者しか動かせない","API 価格 $3/$15 は K2.6 の 3–4 倍。思考は無効化不可で max 設定はトークン消費大。マルチターンでは reasoning_content をそのまま返す必要あり"],"logic_ability":"エンジニアリング型とエージェント推論が中核：Terminal-Bench 2.1 88.3、DeepSWE 67.5、FrontierSWE 81.2、SWE-Marathon 42.0、BrowseComp 91.2 は公開値で最高。試験型推論も最上位（GPQA 93.5、HLE-Full 43.5 / ツールあり 56.0）だが、CritPt 23.4 は GPT-5.6 Sol を明確に下回る。思考はデフォルトでオン、reasoning_effort は low/high/max で調整可。よくある問題：low 設定では AA 指数が 48 まで落ち差が大きい。ツールなしの純テキスト対話はやや冗長。","best_for":["ディープサーチ / リサーチ型エージェント","長期コーディングエージェント（Kimi Code CLI）","ビジョン + 長コンテキストのドキュメント / 動画理解","クラウド事業者規模の自前ホスティング"],"not_for":["あらゆる単一ノードデプロイ","シンプルな対話（コスト高、K2.6 を使う）","ブランド表示を望まない超大規模 MAU 製品"],"capability_notes":{"coding":"Terminal-Bench 2.1 88.3、DeepSWE 67.5、ProgramBench 77.8、SciCode 58.7、FrontierSWE 81.2（公式）。","reasoning":"GPQA Diamond 93.5、HLE-Full 43.5 / 56.0（ツールあり）、CritPt 23.4（公式）。","math":"AIME は公式未開示。MathVision 94.3 / 97.8（Python）。","agent":"BrowseComp 91.2、MCP-Atlas 84.2、Toolathlon-Verified 76.5、OSWorld-Verified 84.8、τ³-Banking 33.4（公式）。","multimodal":"MMMU-Pro 81.6 / 83.4（ツールあり）、Video-MME 90.0、OmniDocBench 91.1（公式）。","chinese":"中国語は最上位。"}}}},{"aliases":["Llama-2-70b-chat-hf","Llama 2 70B Chat"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"knowledge":"MMLU 68.9%（官方论文，基座）。","chinese":"中文弱，词表几乎无中文 token。"},"complete":false,"id":"llama-2-70b","name":"Llama 2 70B","name_zh":"Llama 2 · 70B","vendor":"Meta","family":"Llama 2","license":"Llama 2 Community License","license_commercial":"restricted","weights_url":"https://huggingface.co/meta-llama/Llama-2-70b-chat-hf","superseded_by":"llama-3-1-405b","released_at":"2023-07-18","architecture":{"type":"dense","total_params":"70B","total_params_b":69,"layers":80,"hidden_size":8192,"vocab_size":32000,"kv_heads":8,"head_dim":128,"attention":"GQA（64 Q 头 / 8 KV 头）","notes":"2023 年开源社区的基石模型；词表仅 32K，中文效率低；上下文 4K。","undisclosed":false},"context":{"max_tokens":4096,"display":"4K"},"memory":{"weight_gb":{"bf16":138,"q8":73.1,"q4":41.4},"estimated":false,"kv_per_token_kib":320,"kv_note":"8×128×2×80×2 B = 320 KiB/token；4K 上下文约 1.3 GB。","ref_hw_24gb":"Q4_K_M 41.4 GB 装不下，需两张 24GB 卡或 Q2 级量化","ref_hw_80gb":"Q4 单卡 80GB 可跑；BF16 需 2×80GB","ref_hw_8x80gb":"BF16 高并发服务"},"links":{"official":"https://ai.meta.com/llama/","paper":"https://arxiv.org/abs/2307.09288","hf":"https://huggingface.co/meta-llama/Llama-2-70b-chat-hf"},"copy":{"one_liner":"2023 年开源浪潮的起点，首个可商用的 70B 级开放权重。","highlights":["首个允许商用的主流 70B 开放权重，催生海量微调衍生","GQA + RLHF 对话版，当年开源对话质量标杆","llama.cpp / vLLM 等生态几乎围绕它建立"],"pitfalls":["上下文仅 4K，词表 32K，中文分词效率极低","能力已被后续每一代 8B 级模型全面超越","许可禁止月活 7 亿以上企业使用，且不许用其输出训练他模"],"logic_ability":"无推理链训练，逻辑与数学在今天看属入门水平（当年 GSM8K 56.8%）。多步推理容易失败，主要价值为历史坐标与微调底座。","best_for":["历史对照 / 学术复现","旧有微调流水线"],"not_for":["任何新项目","中文任务"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama"],"finetune":"LoRA 生态最成熟的老模型","zh_docs":"弱"}},{"aliases":["Meta-Llama-3.1-405B-Instruct","Llama 3.1 405B Instruct"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"coding":"HumanEval 89.0%（官方）；无 SWE-bench 官方数。","reasoning":"GPQA 50.7%（0-shot，官方模型卡）。","math":"MATH 73.8%（官方）；AIME 未报告。","knowledge":"MMLU 87.3%（官方）。","agent":"BFCL 88.5%（官方），原生 tool-calling 模板。","chinese":"官方 8 语言不含中文，中文可用但不强。"},"complete":true,"id":"llama-3-1-405b","name":"Llama 3.1 405B","name_zh":"Llama 3.1 · 405B","vendor":"Meta","family":"Llama 3.1","license":"Llama 3.1 Community License","license_commercial":"restricted","weights_url":"https://huggingface.co/meta-llama/Llama-3.1-405B-Instruct","superseded_by":"llama-4-maverick","released_at":"2024-07-23","architecture":{"type":"dense","total_params":"405.9B","total_params_b":405.9,"layers":126,"hidden_size":16384,"vocab_size":128256,"kv_heads":8,"head_dim":128,"attention":"GQA（128 Q 头 / 8 KV 头）","notes":"史上最大开放 Dense 模型；15.6T token；上下文 128K；官方另发 FP8 版。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":811.8,"fp8":405.9,"q8":430,"q4":243},"kv_per_token_kib":504,"kv_note":"8×128×2×126×2 B = 504 KiB/token；128K 上下文约 63 GB。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（单卡）","ref_hw_8x80gb":"FP8 官方版 8×H100 可服务（405 GB 权重 + KV）；BF16 需 16×80GB","estimated":true},"pricing":{"input_per_m":0.8,"output_per_m":0.8,"currency":"USD","source":"第三方托管常见价（DeepInfra）","as_of":"2026-08-28","note":"开源权重，社区托管价，供参考"},"links":{"official":"https://ai.meta.com/blog/meta-llama-3-1/","paper":"https://arxiv.org/abs/2407.21783","github":"https://github.com/meta-llama/llama-models","hf":"https://huggingface.co/meta-llama/Llama-3.1-405B-Instruct"},"copy":{"one_liner":"首个正面对标 GPT-4o 的开放权重旗舰，Dense 405B。","highlights":["2024 年 7 月开源首次在多项基准追平 GPT-4o / Claude 3.5 Sonnet","许可明确允许用其输出做蒸馏，成为大量小模型的教师","128K 上下文 + 原生工具调用格式，8 种语言官方支持"],"pitfalls":["Dense 405B 推理成本极高，FP8 也要 8×H100，性价比已被 MoE 碾压","无推理模式，数学 / 代码在 2025 年后明显落后","月活 7 亿以上企业需单独授权"],"logic_ability":"非推理模型的天花板级别：GPQA 50.7%、MATH 73.8%（官方）。逻辑链条稳定、指令遵循强，但遇到需要长思考的题目会直接给出错误答案而非停下来推演。今天 30B 级推理模型在数学 / 代码上已能超过它。","best_for":["蒸馏教师 / 合成数据","需要许可清晰的大模型私有化"],"not_for":["成本敏感推理","数学 / 竞赛代码"]},"sheet":{"architecture_md":"**类型**：Dense Transformer，405.9B 参数。\n\n- 126 层，隐藏维 16,384，FFN 53,248，词表 128,256\n- GQA：128 Query 头 / 8 KV 头，head_dim 128\n- RoPE θ=500,000，llama3 型 RoPE scaling（8K 原生 → 128K）\n- 预训练 15.6T token，训练 FLOPs 3.8×10²⁵，16K H100\n\n参考：《The Llama 3 Herd of Models》。","memory_md":"| 精度 | 权重大小 | 参考硬件 | 说明 |\n|---|---|---|---|\n| BF16 | 812 GB | 16×80GB | 官方 safetensors |\n| FP8 | 406 GB | 8×H100 | 官方 FP8 版（动态量化） |\n| Q4 | ≈ 243 GB（估） | 4×80GB 起 | GGUF 需拆分，社区 Q4_K_M 约 243 GB |\n\n**KV Cache**：504 KiB/token（BF16）。\n\n**参考配置**：\n- 8×H100 80GB：FP8 权重 406 GB，剩余 ~230 GB 给 KV，可服务 128K 少量并发\n- 单卡任何消费级：不可行","training_md":"- 预训练 15.6T token，多阶段退火 + 长上下文扩展至 128K\n- 后训练：SFT → 拒绝采样 → DPO，6 轮迭代\n- 原生支持 tool-calling（brave_search / wolfram_alpha / code_interpreter 内置模板）\n- 官方 8 语言：英、德、法、意、葡、印地、西、泰","ecosystem_md":"- HF：meta-llama/Llama-3.1-405B-Instruct、-FP8\n- 引擎：vLLM（FP8 推荐）、SGLang、TensorRT-LLM、llama.cpp（多机）\n- 微调：全参需 64+ 卡；LoRA / QLoRA 可在 8×80GB\n- 中文文档：弱","versions_md":"- 同系列：Llama 3.1 70B、8B（同日发布）\n- 后继：Llama 3.3 70B（2024-12，接近 405B 水平）→ Llama 4 Maverick（2025-04，MoE）\n- 衍生：Hermes 3 405B、Nemotron 系列"},"ecosystem":{"engines":["vLLM","SGLang","TensorRT-LLM","llama.cpp"],"finetune":"LoRA 可行，全参需集群","zh_docs":"弱"}},{"aliases":["Llama-3.1-8B-Instruct","Meta-Llama-3.1-8B-Instruct"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"coding":"HumanEval 72.6%（官方）。","reasoning":"GPQA 30.4%（官方模型卡）。","math":"MATH 51.9%（官方）。","knowledge":"MMLU 69.4%（官方）。","chinese":"弱。"},"complete":false,"id":"llama-3-1-8b","name":"Llama 3.1 8B","name_zh":"Llama 3.1 · 8B","vendor":"Meta","family":"Llama 3.1","license":"Llama 3.1 Community License","license_commercial":"restricted","weights_url":"https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct","superseded_by":"qwen3-8-27b","released_at":"2024-07-23","architecture":{"type":"dense","total_params":"8.03B","total_params_b":8.03,"layers":32,"hidden_size":4096,"vocab_size":128256,"kv_heads":8,"head_dim":128,"attention":"GQA（32 Q 头 / 8 KV 头）","notes":"128K 上下文（llama3 RoPE scaling，原生 8K）。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":16.1,"q8":8.5,"q4":4.9},"estimated":false,"kv_per_token_kib":128,"kv_note":"128 KiB/token；32K 上下文约 4 GB。","ref_hw_24gb":"BF16 16 GB 单卡可跑；Q4 4.9 GB 留大量 KV，可上 128K","ref_hw_80gb":"BF16 高并发","ref_hw_8x80gb":"过剩"},"pricing":{"input_per_m":0.03,"output_per_m":0.05,"currency":"USD","source":"第三方托管常见价（DeepInfra）","as_of":"2026-08-28","note":"开源权重，社区托管价"},"links":{"official":"https://ai.meta.com/blog/meta-llama-3-1/","paper":"https://arxiv.org/abs/2407.21783","hf":"https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct"},"copy":{"one_liner":"2024 年最流行的 8B 底座，微调与部署生态第一。","highlights":["消费级显卡 BF16 直接跑，Q4 不到 5 GB","128K 上下文 + 原生工具调用模板","衍生微调数量在 HF 上长期第一"],"pitfalls":["中文能力弱于同期 Qwen2 / GLM-4 9B","无推理模式，数学 GPQA 30.4% 属入门","许可限制（7 亿月活）与品牌标注要求"],"logic_ability":"入门级：GPQA 30.4%、MATH 51.9%（官方）。简单逻辑与格式化输出可靠，多步推理与复杂代码不稳。","best_for":["微调底座 / 教学","边缘设备英文助手"],"not_for":["中文","复杂推理"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama","MLX","TensorRT-LLM"],"finetune":"LoRA / 全参极友好","zh_docs":"弱"}},{"aliases":["Llama-3.1-Nemotron-70B-Instruct-HF"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"instruction":"Arena Hard 85.0、MT-Bench 8.98（官方）。","chinese":"弱。"},"complete":false,"id":"llama-3-1-nemotron-70b","name":"Llama-3.1-Nemotron-70B","name_zh":"Llama-3.1-Nemotron · 70B","vendor":"NVIDIA","family":"Nemotron","license":"Llama 3.1 Community License","license_commercial":"restricted","weights_url":"https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF","superseded_by":"nemotron-3-super-120b-a12b","released_at":"2024-10-15","architecture":{"type":"dense","total_params":"70.6B","total_params_b":70.6,"layers":80,"hidden_size":8192,"vocab_size":128256,"kv_heads":8,"head_dim":128,"attention":"GQA（64 Q 头 / 8 KV 头）","notes":"Llama 3.1 70B 用 HelpSteer2 + REINFORCE RLHF 对齐；128K。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":141.2,"q8":74.8,"q4":42.5},"estimated":false,"kv_per_token_kib":320,"kv_note":"320 KiB/token。","ref_hw_24gb":"Q4 需 2×24GB","ref_hw_80gb":"Q4 单卡；BF16 2×80GB","ref_hw_8x80gb":"高并发"},"links":{"official":"https://build.nvidia.com/nvidia/llama-3_1-nemotron-70b-instruct","paper":"https://arxiv.org/abs/2410.01257","hf":"https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF"},"copy":{"one_liner":"RLHF 强化版 Llama 3.1 70B，对话榜一度超 GPT-4o。","highlights":["Arena Hard 85.0、AlpacaEval 2 LC 57.6（官方），当时开源第一","Arena Elo 1267（2024-10-24，官方引用）","展示了奖励模型 + REINFORCE 的高效对齐"],"pitfalls":["回答冗长，「strawberry 有几个 r」式炫技","基础能力未变，数学 / 代码等于 Llama 3.1 70B","Llama 许可限制"],"logic_ability":"对话偏好榜强，硬推理与 Llama 3.1 70B 同级（无额外训练）。","best_for":["英文对话助手","对齐研究"],"not_for":["需简洁输出的场景","中文"]},"ecosystem":{"engines":["vLLM","llama.cpp","TensorRT-LLM"],"finetune":"LoRA 成熟","zh_docs":"弱"}},{"aliases":["Llama-3.3-70B-Instruct"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"coding":"HumanEval 88.4%（官方）。","reasoning":"GPQA Diamond 50.5%（CoT，官方模型卡）。","math":"MATH 77.0%（官方）。","knowledge":"MMLU 86.0%（官方）。","agent":"BFCL v2 77.3%（官方）。"},"complete":false,"id":"llama-3-3-70b","name":"Llama 3.3 70B","name_zh":"Llama 3.3 · 70B","vendor":"Meta","family":"Llama 3.3","license":"Llama 3.3 Community License","license_commercial":"restricted","weights_url":"https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct","superseded_by":"llama-4-maverick","released_at":"2024-12-06","architecture":{"type":"dense","total_params":"70.6B","total_params_b":70.6,"layers":80,"hidden_size":8192,"vocab_size":128256,"kv_heads":8,"head_dim":128,"attention":"GQA（64 Q 头 / 8 KV 头）","notes":"结构同 3.1 70B，仅后训练升级（在线 RL），接近 405B 表现。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":141.2,"q8":74.8,"q4":42.5},"estimated":false,"kv_per_token_kib":320,"kv_note":"320 KiB/token；32K 上下文约 10 GB。","ref_hw_24gb":"Q4_K_M 42.5 GB 需 2×24GB","ref_hw_80gb":"Q4 单卡 80GB 可跑 32K；BF16 2×80GB","ref_hw_8x80gb":"BF16 高并发服务"},"pricing":{"input_per_m":0.23,"output_per_m":0.4,"currency":"USD","source":"第三方托管常见价（DeepInfra）","as_of":"2026-08-28","note":"开源权重，社区托管价"},"links":{"official":"https://ai.meta.com/blog/future-of-ai-built-with-llama/","hf":"https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct"},"copy":{"one_liner":"用 70B 达到 405B 水平，2024 年末最划算的开源通用模型。","highlights":["官方称多数基准追平 3.1 405B，成本仅 1/6","128K 上下文，原生工具调用，8 语言","Q4 单张 80GB 卡即可服务"],"pitfalls":["无推理模式，2025 年后数学 / 代码榜被 DeepSeek / Qwen 拉开","中文不是官方语言，质量一般","Llama 许可对大企业与再分发有限制"],"logic_ability":"非推理模型中上：GPQA Diamond 50.5%、MATH 77.0%（官方）。指令遵循与对话稳定，但复杂数学与 agent 任务不及 2025 年推理模型。","best_for":["英文通用助手私有化","替代 405B 的低成本方案"],"not_for":["竞赛数学 / 深度代码","中文优先场景"]},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Ollama","TensorRT-LLM"],"finetune":"LoRA 成熟","zh_docs":"弱"}},{"aliases":["Meta-Llama-3-70B-Instruct"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"reasoning":"GPQA 0-shot 39.5%（Llama 3.1 模型卡对照列）。","knowledge":"MMLU 82.0%（官方）。"},"complete":false,"id":"llama-3-70b","name":"Llama 3 70B","name_zh":"Llama 3 · 70B","vendor":"Meta","family":"Llama 3","license":"Meta Llama 3 Community License","license_commercial":"restricted","weights_url":"https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct","superseded_by":"llama-3-3-70b","released_at":"2024-04-18","architecture":{"type":"dense","total_params":"70.6B","total_params_b":70.6,"layers":80,"hidden_size":8192,"vocab_size":128256,"kv_heads":8,"head_dim":128,"attention":"GQA（64 Q 头 / 8 KV 头）","notes":"词表扩至 128K（tiktoken）；15T token 预训练；上下文 8K。","undisclosed":false},"context":{"max_tokens":8192,"display":"8K"},"memory":{"weight_gb":{"bf16":141.2,"q8":74.8,"q4":42.5},"estimated":false,"kv_per_token_kib":320,"kv_note":"320 KiB/token；8K 上下文约 2.6 GB。","ref_hw_24gb":"Q4_K_M 42.5 GB 需 2×24GB 或 Q2/Q3","ref_hw_80gb":"Q4 单卡 80GB；BF16 2×80GB","ref_hw_8x80gb":"BF16 高并发"},"links":{"official":"https://ai.meta.com/blog/meta-llama-3/","github":"https://github.com/meta-llama/llama3","hf":"https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct"},"copy":{"one_liner":"2024 年上半年开源第一梯队，首次逼近 GPT-4 早期版本。","highlights":["15T token 预训练，同尺寸下当时英文能力最强","128K 词表大幅提升多语言与代码分词效率","发布即登 Arena 开源榜首，微调衍生极多"],"pitfalls":["上下文只有 8K，长文档需社区外推版","非工具调用原生格式，agent 用法要靠提示词","已被 3.1 / 3.3 70B 完全取代（同许可，更强）"],"logic_ability":"无 CoT 专训，但指令遵循与常识推理在当年开源中领先；数学 GSM8K 93%（官方），复杂多步推理明显不及 2025 年推理模型。","best_for":["历史对照","已有 Llama 3 微调链路的维护"],"not_for":["长文本","新项目（直接用 3.3）"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama","TensorRT-LLM"],"finetune":"LoRA / 全参成熟","zh_docs":"弱"}},{"id":"llama-4-maverick","name":"Llama 4 Maverick","aliases":["Llama-4-Maverick-17B-128E-Instruct"],"vendor":"Meta","family":"Llama 4","license":"Llama 4 Community License","license_commercial":"restricted","openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct","status":"superseded","released_at":"2025-04-05","updated_at":"2025-12-20","modalities":["text","image","tools"],"reasoning_mode":"none","architecture":{"type":"moe","total_params":"400B","active_params":"17B","total_params_b":400,"active_params_b":17,"experts":128,"active_experts":1,"shared_expert":true,"layers":48,"hidden_size":5120,"kv_heads":8,"head_dim":128,"attention":"GQA + iRoPE（交替无位置编码层）","undisclosed":false},"context":{"max_tokens":1048576,"display":"1M"},"memory":{"weight_gb":{"bf16":800,"fp8":400,"q4":245},"estimated":true,"ref_hw_8x80gb":"FP8 8×H100（官方目标）"},"pricing":{"input_per_m":0.15,"output_per_m":0.6,"currency":"USD","source":"第三方托管常见价","as_of":"2025-12-20"},"links":{"official":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/","hf":"https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct"},"copy":{"one_liner":"Meta 的原生多模态 MoE，1M 上下文，许可有 MAU 限制。","highlights":["原生图文","1M 上下文（宣称）","17B 激活，吞吐高"],"pitfalls":["月活 > 7 亿需单独授权，欧盟限制","推理 / 代码弱于同期中国开源模型","长上下文实测质量争议"],"logic_ability":"非推理模型，考试型与工程型均为中游。","best_for":["多模态英文对话"],"not_for":["推理 / 代码主力"]},"capability_notes":{},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp"],"zh_docs":"无"},"complete":false,"superseded_by":"muse-spark"},{"id":"llama-4-scout","name":"Llama 4 Scout","aliases":["Llama-4-Scout-17B-16E-Instruct"],"vendor":"Meta","family":"Llama 4","license":"Llama 4 Community License","license_commercial":"restricted","openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct","status":"superseded","released_at":"2025-04-05","updated_at":"2025-12-20","modalities":["text","image","tools"],"reasoning_mode":"none","architecture":{"type":"moe","total_params":"109B","active_params":"17B","total_params_b":109,"active_params_b":17,"experts":16,"active_experts":1,"shared_expert":true,"layers":48,"hidden_size":5120,"kv_heads":8,"head_dim":128,"attention":"GQA + iRoPE","undisclosed":false},"context":{"max_tokens":10000000,"display":"10M"},"memory":{"weight_gb":{"bf16":218,"fp8":109,"q4":65},"estimated":true,"ref_hw_80gb":"官方 int4 单 H100 可跑"},"pricing":{"input_per_m":0.08,"output_per_m":0.3,"currency":"USD","source":"第三方托管常见价","as_of":"2025-12-20"},"links":{"official":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/","hf":"https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct"},"copy":{"one_liner":"单卡 H100 可跑的多模态 MoE，宣称 10M 上下文。","highlights":["int4 单 H100","10M 上下文（宣称）","原生图文"],"pitfalls":["10M 上下文实际可用性远低于宣称","推理弱","许可限制"],"logic_ability":"非推理模型，能力中游。","best_for":["多模态 + 长文原型"],"not_for":["推理 / 代码"]},"capability_notes":{},"ecosystem":{"engines":["vLLM","llama.cpp"],"zh_docs":"无"},"complete":false,"superseded_by":"muse-glimmer-30b"},{"aliases":["MiMo-7B-RL"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"default-on","capability_notes":{"coding":"LiveCodeBench v5 57.8%（官方）。","reasoning":"GPQA Diamond 54.4%（官方）。","math":"AIME 2025 55.4%（官方）。"},"complete":false,"id":"mimo-7b","name":"MiMo-7B","name_zh":"小米 MiMo · 7B","vendor":"Xiaomi","vendor_zh":"小米","family":"MiMo","license":"MIT","license_commercial":true,"weights_url":"https://huggingface.co/XiaomiMiMo/MiMo-7B-RL","superseded_by":"mimo-v2-5","released_at":"2025-04-30","architecture":{"type":"dense","total_params":"7B","total_params_b":7.3,"layers":36,"hidden_size":4096,"vocab_size":151680,"kv_heads":8,"head_dim":128,"attention":"GQA（32 Q 头 / 8 KV 头）","notes":"从零预训练 25T token + MTP；推理 RL；上下文 32K。","undisclosed":false},"context":{"max_tokens":32768,"display":"32K"},"memory":{"weight_gb":{"bf16":14.6,"q8":7.7,"q4":4.2},"estimated":true,"kv_per_token_kib":144,"kv_note":"144 KiB/token。","ref_hw_24gb":"BF16 14.6 GB 单卡可跑","ref_hw_80gb":"高并发","ref_hw_8x80gb":"过剩"},"links":{"official":"https://github.com/XiaomiMiMo/MiMo","github":"https://github.com/XiaomiMiMo/MiMo","paper":"https://arxiv.org/abs/2505.07608","hf":"https://huggingface.co/XiaomiMiMo/MiMo-7B-RL"},"copy":{"one_liner":"小米首个开源推理模型，7B 数学超 o1-mini。","highlights":["AIME 2025 55.4%、LiveCodeBench v5 57.8%（官方）超 o1-mini","MIT 许可，从零预训练而非蒸馏","MTP 层可加速推理"],"pitfalls":["通用对话与知识弱（推理专精）","上下文 32K","思考链长，7B 优势被抵消"],"logic_ability":"7B 中推理最强档：AIME 2025 55.4%、GPQA 54.4%、LCB v5 57.8%（官方 RL 版）。","best_for":["边缘数学 / 代码推理","推理 RL 研究"],"not_for":["通用助手","知识问答"]},"ecosystem":{"engines":["vLLM（官方 fork）","SGLang","llama.cpp"],"finetune":"LoRA 可行","zh_docs":"有"}},{"id":"mimo-v2-5","name":"MiMo-V2.5","name_zh":"小米 MiMo-V2.5","aliases":["XiaomiMiMo/MiMo-V2.5","mimo-v2.5","MiMo V2.5"],"vendor":"Xiaomi","vendor_zh":"小米","family":"MiMo","license":"MIT","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/XiaomiMiMo/MiMo-V2.5","status":"current","released_at":"2026-04-22","updated_at":"2026-08-28","modalities":["text","image","video","audio","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"310B","active_params":"15B","total_params_b":310,"active_params_b":15,"experts":256,"active_experts":8,"shared_expert":false,"layers":48,"hidden_size":4096,"vocab_size":152576,"kv_heads":4,"head_dim":192,"attention":"混合：39 层 128 滑窗 + 9 层全注意力（5:1），64 Q 头，QK 192 / V 128","notes":"1 稠密 + 47 MoE 层；729M 视觉 + 261M 音频编码器，全模态输入。约 48T token FP8 预训练。","undisclosed":false},"context":{"max_tokens":1048576,"display":"1M"},"memory":{"weight_gb":{"bf16":620,"fp8":310,"q4":180,"q8":329.3},"estimated":true,"ref_hw_80gb":"不可行","ref_hw_8x80gb":"FP8 多卡"},"pricing":{"input_per_m":0.119,"output_per_m":0.238,"currency":"USD","source":"第三方托管常见价（OpenRouter）","as_of":"2026-08-28"},"links":{"official":"https://huggingface.co/XiaomiMiMo/MiMo-V2.5","hf":"https://huggingface.co/XiaomiMiMo/MiMo-V2.5"},"variants":[{"kind":"fp8","publisher":"XiaomiMiMo","repo":"XiaomiMiMo/MiMo-V2.5","url":"https://huggingface.co/XiaomiMiMo/MiMo-V2.5","note":"官方发布版（FP8）","sizes":{"fp8":315.7}},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/MiMo-V2.5-GGUF","url":"https://huggingface.co/unsloth/MiMo-V2.5-GGUF","sizes":{"q8":329.3,"bf16":619.6}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/MiMo-V2.5-GGUF","url":"https://huggingface.co/bartowski/MiMo-V2.5-GGUF","sizes":{"q4":188.8,"q5":221.4,"q6":267.6,"q8":329.3}}],"copy":{"one_liner":"小米 310B 全模态 MoE，15B 激活，MIT 许可，1M 上下文。","highlights":["MIT 许可，文本 / 图 / 视频 / 音频全模态输入","Terminal-Bench 2.0 65.8、SWE-bench Pro 56.1（官方）","1M 上下文，滑窗 5:1 混合注意力省 KV；AA 智能指数 38"],"pitfalls":["310B 总参，多卡部署","官方未给 GPQA / AIME / SWE-bench Verified 数字","Arena 文本榜未见 V2.5（Pro 版 1468）"],"logic_ability":"工程型推理强（TB 2.0 65.8），考试型推理官方数字缺；更强的 V2.5-Pro（1T-A42B）AA 43。","best_for":["全模态 agent","超长上下文低成本 API"],"not_for":["单机部署","需要权威推理评测背书的场景"]},"capability_notes":{"coding":"Terminal-Bench 2.0 65.8、SWE-bench Pro 56.1（官方）。","multimodal":"图片 / 视频 / 音频理解。","chinese":"强，国产模型。"},"ecosystem":{"engines":["vLLM","SGLang"],"zh_docs":"有"},"complete":false,"runtime":{"tok_s":68,"latency_s":2.82,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"Xiaomi MiMo-V2.5","one_liner":"Xiaomi's 310B omni-modal MoE, 15B active, MIT license, 1M context.","highlights":["MIT license; omni-modal input across text / image / video / audio","Terminal-Bench 2.0 65.8, SWE-bench Pro 56.1 (official)","1M context; 5:1 sliding-window hybrid attention saves KV; AA Intelligence Index 38"],"pitfalls":["310B total parameters; multi-GPU deployment","No official GPQA / AIME / SWE-bench Verified numbers","V2.5 not seen on the Arena text leaderboard (Pro version 1468)"],"logic_ability":"Strong engineering reasoning (TB 2.0 65.8); official exam-style reasoning numbers missing; the stronger V2.5-Pro (1T-A42B) scores AA 43.","best_for":["Omni-modal agents","Low-cost ultra-long-context API"],"not_for":["Single-machine deployment","Scenarios needing authoritative reasoning benchmark backing"],"capability_notes":{"coding":"Terminal-Bench 2.0 65.8, SWE-bench Pro 56.1 (official).","multimodal":"Image / video / audio understanding.","chinese":"Strong; Chinese domestic model."}},"ja":{"name_zh":"Xiaomi MiMo-V2.5","one_liner":"Xiaomi の 310B 全モーダル MoE。15B 活性、MIT、1M。","highlights":["MIT ライセンス、テキスト / 画像 / 動画 / 音声の全モーダル入力","Terminal-Bench 2.0 65.8、SWE-bench Pro 56.1（公式）","1M コンテキスト、5:1 スライディングウィンドウ混合アテンションで KV を節約。AA 知能指数 38"],"pitfalls":["総パラメータ 310B、マルチ GPU デプロイが必要","公式の GPQA / AIME / SWE-bench Verified の数値なし","Arena テキスト榜に V2.5 は見当たらない（Pro 版は 1468）"],"logic_ability":"エンジニアリング型推論は強い（TB 2.0 65.8）、試験型推論は公式数値が欠落。上位の V2.5-Pro（1T-A42B）は AA 43。","best_for":["全モーダルエージェント","超長コンテキストの低コスト API"],"not_for":["単一マシンでのデプロイ","権威ある推論評価の裏付けが必要な場面"],"capability_notes":{"coding":"Terminal-Bench 2.0 65.8、SWE-bench Pro 56.1（公式）。","multimodal":"画像 / 動画 / 音声理解。","chinese":"強い。中国産モデル。"}}}},{"aliases":["MiniMax-M1-80k","MiniMax-M1-40k"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"default-on","capability_notes":{"coding":"LiveCodeBench 65.0%、SWE-bench Verified 56.0%（官方）。","reasoning":"GPQA Diamond 70.0%、HLE 8.4%（官方）。","math":"AIME 2025 76.9%（官方）。","agent":"TAU-bench airline 62.0 / retail 63.5（官方）。","chinese":"中文强。"},"complete":false,"id":"minimax-m1","name":"MiniMax-M1","name_zh":"MiniMax-M1 · 456B-A46B","vendor":"MiniMax","vendor_zh":"MiniMax","family":"MiniMax-M1","license":"Apache-2.0","license_commercial":true,"weights_url":"https://huggingface.co/MiniMaxAI/MiniMax-M1-80k","superseded_by":"minimax-m2","released_at":"2025-06-16","architecture":{"type":"hybrid","total_params":"456B","active_params":"45.9B","total_params_b":456,"active_params_b":45.9,"experts":32,"active_experts":2,"shared_expert":false,"layers":80,"hidden_size":6144,"vocab_size":200064,"kv_heads":8,"head_dim":128,"attention":"混合：7 层 Lightning Attention（线性）+ 1 层 softmax GQA，循环","notes":"基于 MiniMax-Text-01；1M 上下文；CISPO 强化学习；80K 思考预算版。","undisclosed":false},"context":{"max_tokens":1000000,"display":"1M","max_output":80000},"memory":{"weight_gb":{"bf16":912,"q8":483.4,"q4":264.5,"fp8":456},"estimated":true,"kv_per_token_kib":40,"kv_note":"仅 1/8 的层是 softmax 注意力（10 层），KV ≈ 40 KiB/token，1M 上下文约 40 GB；线性层状态固定。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行","ref_hw_8x80gb":"BF16 912 GB 需 16×80GB；官方推荐 8×H800/H20 跑 FP8 / int8"},"pricing":{"input_per_m":0.4,"output_per_m":2.2,"currency":"USD","source":"MiniMax 官方 API（M1，≤200K 输入档）","as_of":"2026-08-28"},"links":{"official":"https://www.minimax.io/news/minimaxm1","github":"https://github.com/MiniMax-AI/MiniMax-M1","paper":"https://arxiv.org/abs/2506.13585","hf":"https://huggingface.co/MiniMaxAI/MiniMax-M1-80k"},"copy":{"one_liner":"首个开源混合线性注意力推理模型，1M 上下文。","highlights":["Lightning Attention 让 100K 生成的 FLOPs 仅为 R1 的 25%","1M 上下文 + 80K 思考输出，Apache-2.0","RL 成本仅 53.5 万美元（官方）"],"pitfalls":["456B 需多机部署，社区量化少","考试型推理弱于 R1-0528 / Qwen3-235B","被 M2（2025-10，更小更强）取代"],"logic_ability":"AIME 2025 76.9%、GPQA 70.0%、LiveCodeBench 65.0%、SWE-bench Verified 56.0%（官方，80k 版）。长上下文与 agent 是强项。","best_for":["超长上下文推理","长程 agent 任务"],"not_for":["本地部署","极致数学"]},"ecosystem":{"engines":["vLLM","SGLang","Transformers"],"finetune":"极困难","zh_docs":"有"}},{"id":"minimax-m2-7","name":"MiniMax-M2.7","name_zh":"MiniMax M2.7","aliases":["MiniMax M2.7","minimax-m2.7"],"vendor":"MiniMax","family":"MiniMax-M","license":"MiniMax Non-Commercial License（MIT-style）","license_commercial":false,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/MiniMaxAI/MiniMax-M2.7","status":"current","released_at":"2026-03-18","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"default-on","architecture":{"type":"moe","total_params":"229B","active_params":"10B","total_params_b":229,"active_params_b":10,"experts":256,"active_experts":8,"shared_expert":false,"layers":62,"hidden_size":3072,"vocab_size":200064,"kv_heads":8,"head_dim":128,"attention":"GQA（48 Q 头 / 8 KV 头）","notes":"与 M2 同构（62 层 GQA + 256 专家），3 个 MTP 模块，QK-norm。HF 发布版为 FP8（block 128×128）。预训练 29.2T token（NVIDIA 技术博客）。","undisclosed":false},"context":{"max_tokens":204800,"display":"200K"},"memory":{"weight_gb":{"fp8":230,"q4":133,"q8":243.1,"bf16":457.5},"kv_per_token_kib":248,"kv_note":"8 KV 头 × 128 × 2 × 62 层 × 2 B = 248 KiB/token。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（Q4 约 133 GB）","ref_hw_8x80gb":"FP8 230 GB 官方发布版，4×H100 起","estimated":true},"pricing":{"input_per_m":0.3,"output_per_m":1.2,"currency":"USD","source":"MiniMax 开放平台","as_of":"2026-08-28","note":"缓存读 $0.06 / 写 $0.375"},"links":{"official":"https://www.minimax.io/news/minimax-m27-en","hf":"https://huggingface.co/MiniMaxAI/MiniMax-M2.7","github":"https://github.com/MiniMax-AI/MiniMax-M2.7","pricing":"https://platform.minimax.io/docs/guides/pricing-paygo"},"variants":[{"kind":"fp8","publisher":"MiniMaxAI","repo":"MiniMaxAI/MiniMax-M2.7","url":"https://huggingface.co/MiniMaxAI/MiniMax-M2.7","note":"官方发布版（FP8）","sizes":{"fp8":230.1}},{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/MiniMax-M2.7-NVFP4","url":"https://huggingface.co/nvidia/MiniMax-M2.7-NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/MiniMax-M2.7-GGUF","url":"https://huggingface.co/unsloth/MiniMax-M2.7-GGUF","sizes":{"q8":243.1,"bf16":457.5}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/MiniMaxAI_MiniMax-M2.7-GGUF","url":"https://huggingface.co/bartowski/MiniMaxAI_MiniMax-M2.7-GGUF","sizes":{"q4":138.8,"q5":162.7,"q6":197.1,"q8":243.1}},{"kind":"awq","publisher":"cyankiwi","repo":"cyankiwi/MiniMax-M2.7-AWQ-4bit","url":"https://huggingface.co/cyankiwi/MiniMax-M2.7-AWQ-4bit"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/MiniMax-M2.7-4bit","url":"https://huggingface.co/mlx-community/MiniMax-M2.7-4bit"}],"copy":{"one_liner":"10B 激活的 agent 编程模型，但权重仅限非商用。","highlights":["SWE-Pro 56.22、Terminal Bench 2 57.0、GDPval-AA Elo 1495（发布时开源最高）","10B 激活低延迟，FP8 230 GB 4×H100 可跑","原生 Agent Teams 多 agent 协作"],"pitfalls":["许可为非商用：商用须 MiniMax 书面授权并标注「Built with MiniMax M2.7」，与 M2 的 MIT 不同","已被 M3 取代旗舰；多模态需用 M3","通用推理基准（GPQA / AIME）官方未披露"],"logic_ability":"工程型推理为主：SWE-Pro 56.22、Multi-SWE 52.7、NL2Repo 39.8、MLE Bench Lite 66.6% 奖牌率。","best_for":["非商用 / 研究自建的编程 agent","API 低成本 agent 服务"],"not_for":["商用自建（许可限制）","多模态"]},"capability_notes":{"coding":"SWE-Pro 56.22、Terminal Bench 2 57.0、SWE Multilingual 76.5、Multi-SWE 52.7（官方）。","agent":"Toolathon 46.3、GDPval-AA 1495、MM Claw 62.7（官方）。"},"ecosystem":{"engines":["SGLang","vLLM","Transformers"],"finetune":"社区少","zh_docs":"有"},"complete":false,"runtime":{"tok_s":61,"latency_s":1.56,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"MiniMax M2.7","one_liner":"10B-active agentic coding model, but weights are non-commercial only.","highlights":["SWE-Pro 56.22, Terminal Bench 2 57.0, GDPval-AA Elo 1495 (highest open-source at release)","10B active for low latency; FP8 230 GB runs on 4×H100","Native Agent Teams multi-agent collaboration"],"pitfalls":["Non-commercial license: commercial use requires written authorization from MiniMax and a \"Built with MiniMax M2.7\" notice, unlike M2's MIT","Flagship spot taken by M3; use M3 for multimodal","General reasoning benchmarks (GPQA / AIME) not officially disclosed"],"logic_ability":"Mainly engineering reasoning: SWE-Pro 56.22, Multi-SWE 52.7, NL2Repo 39.8, MLE Bench Lite 66.6% medal rate.","best_for":["Non-commercial / research self-hosted coding agents","Low-cost agent services via API"],"not_for":["Commercial self-hosting (license restriction)","Multimodal"],"capability_notes":{"coding":"SWE-Pro 56.22, Terminal Bench 2 57.0, SWE Multilingual 76.5, Multi-SWE 52.7 (official).","agent":"Toolathon 46.3, GDPval-AA 1495, MM Claw 62.7 (official)."}},"ja":{"name_zh":"MiniMax M2.7","one_liner":"10B アクティブのエージェント型コーディングモデル。重みは非商用限定。","highlights":["SWE-Pro 56.22、Terminal Bench 2 57.0、GDPval-AA Elo 1495（公開時オープンソース最高）","10B アクティブで低レイテンシ、FP8 230 GB で 4×H100 で動作","ネイティブの Agent Teams マルチエージェント協調"],"pitfalls":["ライセンスは非商用：商用には MiniMax の書面許諾と「Built with MiniMax M2.7」表示が必要で、M2 の MIT とは異なる","フラッグシップの座は M3 に移行。マルチモーダルは M3 を使う","汎用推論ベンチマーク（GPQA / AIME）は公式未開示"],"logic_ability":"エンジニアリング型推論が主：SWE-Pro 56.22、Multi-SWE 52.7、NL2Repo 39.8、MLE Bench Lite メダル率 66.6%。","best_for":["非商用 / 研究目的の自前コーディングエージェント","API 経由の低コストエージェントサービス"],"not_for":["商用の自前ホスティング（ライセンス制限）","マルチモーダル"],"capability_notes":{"coding":"SWE-Pro 56.22、Terminal Bench 2 57.0、SWE Multilingual 76.5、Multi-SWE 52.7（公式）。","agent":"Toolathon 46.3、GDPval-AA 1495、MM Claw 62.7（公式）。"}}}},{"id":"minimax-m2","name":"MiniMax-M2","name_zh":"MiniMax M2","aliases":["MiniMax M2","minimax-m2"],"vendor":"MiniMax","family":"MiniMax-M","license":"MIT","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/MiniMaxAI/MiniMax-M2","status":"superseded","released_at":"2025-10-27","updated_at":"2025-12-20","modalities":["text","tools"],"reasoning_mode":"default-on","architecture":{"type":"moe","total_params":"230B","active_params":"10B","total_params_b":230,"active_params_b":10,"experts":256,"active_experts":8,"shared_expert":false,"layers":62,"hidden_size":3072,"vocab_size":200064,"kv_heads":8,"head_dim":128,"attention":"GQA（48 Q 头 / 8 KV 头）","notes":"放弃 M1 的 Lightning Attention，回到全注意力 GQA + MoE。思考内容用 <think> 标签交错输出，官方要求保留历史思考。","undisclosed":false},"context":{"max_tokens":204800,"display":"200K","max_output":131072},"memory":{"weight_gb":{"bf16":460,"fp8":230,"q4":135},"kv_per_token_kib":248,"kv_note":"8 KV 头 × 128 × 2 × 62 层 × 2 B = 248 KiB/token。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（Q4 135 GB）","ref_hw_8x80gb":"BF16 460 GB 可部署，官方推荐 4×H100 FP8 起","estimated":true},"pricing":{"input_per_m":0.3,"output_per_m":1.2,"currency":"USD","source":"MiniMax 开放平台","as_of":"2025-12-20"},"links":{"official":"https://www.minimax.io/news/minimax-m2","hf":"https://huggingface.co/MiniMaxAI/MiniMax-M2","github":"https://github.com/MiniMax-AI/MiniMax-M2","pricing":"https://platform.minimax.io/docs/guides/pricing"},"copy":{"one_liner":"10B 激活的 agent 编程模型，AA 指数开源第一（发布时）。","highlights":["10B 激活带来的低延迟，Artificial Analysis 指数 61 为发布时开源最高","SWE-bench Verified 69.4%、Terminal-Bench 46.3%，编程 agent 表现接近闭源中档","MIT 许可，API 价 $0.3 / $1.2 约为 Claude Sonnet 的 8%"],"pitfalls":["<think> 交错输出必须完整回传历史，否则多轮 agent 质量明显下降","230B 总参数，Q4 也要 135 GB，自建门槛不低","知识与写作类任务弱于同期 GLM-4.6 / DeepSeek"],"logic_ability":"工程型推理为主：训练重点是「多文件编辑 → 运行 → 修复」的 agent 循环，BrowseComp 44%。考试型推理（AIME 78%）在开源里中游。思考不可关闭。常见问题：过度依赖工具反馈，无工具纯推理时较弱。","best_for":["编程 agent（Claude Code / Cursor 后端）","低延迟工具调用服务","自建 4×H100 级别推理"],"not_for":["单卡部署","写作 / 知识类产品"]},"capability_notes":{"coding":"SWE-bench Verified 69.4%、Multi-SWE 36.2%（官方）。","reasoning":"GPQA Diamond 78%、HLE 31.8%（工具，官方）。","math":"AIME 2025 78%（官方）。","agent":"Terminal-Bench 46.3%、τ²-bench 77.2%、BrowseComp 44.0%（官方）。","chinese":"中文良好。"},"sheet":{"architecture_md":"**类型**：MoE，230B 总 / 10B 激活。\n\n- 62 层，隐藏维 3072，词表 200,064\n- 专家：256 个，每 token 激活 8 个\n- GQA：48 Q 头 / 8 KV 头，head_dim 128\n- 全注意力（放弃 M1 的 Lightning Attention 混合）\n- 上下文 204,800\n\n参考：MiniMax-M2 GitHub / HF 模型卡。","memory_md":"| 精度 | 权重大小 | 参考显存 |\n|---|---|---|\n| BF16 | ≈ 460 GB | 8×80GB |\n| FP8 | ≈ 230 GB | 4×80GB |\n| Q4 | ≈ 135 GB（估） | 2×80GB |\n\n**KV Cache**：≈ 248 KiB/token，128K 上下文 ≈ 31 GB。\n\n**参考配置**：\n- 24GB / 80GB：不可行\n- 4×H100：FP8，官方推荐起步\n- 8×H100：BF16 舒适","training_md":"- 训练细节披露有限\n- 思考默认开启（<think>），交错输出，官方要求历史保留\n- 工具调用：原生，兼容 Anthropic / OpenAI 接口格式\n- 最大输出 131K","ecosystem_md":"- HF：MiniMaxAI/MiniMax-M2\n- 引擎：vLLM、SGLang（官方指南）、Transformers\n- 微调：社区少\n- 中文文档：有","versions_md":"- 上代：MiniMax-M1（Lightning Attention 混合，456B/46B）\n- 后续：M2.1（2025-12）"},"ecosystem":{"engines":["vLLM","SGLang"],"finetune":"社区少","zh_docs":"有"},"complete":true,"superseded_by":"minimax-m3"},{"id":"minimax-m3","name":"MiniMax-M3","name_zh":"MiniMax M3","aliases":["MiniMax M3","minimax-m3"],"vendor":"MiniMax","family":"MiniMax-M","license":"MiniMax Community License","license_commercial":"restricted","openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/MiniMaxAI/MiniMax-M3","status":"current","released_at":"2026-06-01","updated_at":"2026-08-28","modalities":["text","image","video","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"428B","active_params":"23B","total_params_b":428,"active_params_b":23,"experts":128,"active_experts":4,"shared_expert":true,"layers":60,"hidden_size":6144,"vocab_size":200064,"kv_heads":4,"head_dim":128,"attention":"MSA 块稀疏注意力（GQA 64 Q 头 / 4 KV 头）","notes":"MiniMax Sparse Attention（MSA）：在 GQA 之上用 4 头索引分支选 top-16 个 128-token 块做块稀疏注意力；前 3 层全注意力 + Dense MLP。专家 128 路由 + 1 共享，每 token 4 个，swigluoai 激活，QK-norm，半维 RoPE。7 个 MTP 模块。原生多模态（CLIP 式 ViT，32 层，2016px），从预训练第一步混合模态。","undisclosed":false},"context":{"max_tokens":1048576,"display":"1M"},"memory":{"weight_gb":{"bf16":854,"fp8":428,"q4":248,"q8":453.6},"kv_per_token_kib":120,"kv_note":"4 KV 头 × 128 × 2 × 60 层 × 2 B = 120 KiB/token；MSA 只减少注意力计算，KV 仍全量存储。1M 上下文单请求 ≈ 117 GB。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行","ref_hw_8x80gb":"BF16 854 GB 需 8×H200；FP8 428 GB 可 8×H100（KV 余量约 200 GB）","estimated":true},"pricing":{"input_per_m":0.3,"output_per_m":1.2,"currency":"USD","source":"MiniMax 开放平台","as_of":"2026-08-28","note":"标价 $0.6/$2.4「永久 5 折」后；缓存读 $0.06；输入 > 512K 计费 ×2"},"links":{"official":"https://www.minimax.io/blog/minimax-m3","hf":"https://huggingface.co/MiniMaxAI/MiniMax-M3","github":"https://github.com/MiniMax-AI/MiniMax-M3","paper":"https://arxiv.org/abs/2606.13392","pricing":"https://platform.minimax.io/docs/guides/pricing-paygo"},"variants":[{"kind":"fp8","publisher":"MiniMaxAI","repo":"MiniMaxAI/MiniMax-M3-MXFP8","url":"https://huggingface.co/MiniMaxAI/MiniMax-M3-MXFP8","note":"官方 MXFP8","sizes":{"fp8":443.7}},{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/MiniMax-M3-NVFP4","url":"https://huggingface.co/nvidia/MiniMax-M3-NVFP4"},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/MiniMax-M3-GGUF","url":"https://huggingface.co/bartowski/MiniMax-M3-GGUF","sizes":{"q4":261.3,"q5":305.3,"q6":369.4,"q8":453.6}},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/MiniMax-M3-GGUF","url":"https://huggingface.co/unsloth/MiniMax-M3-GGUF","sizes":{"q8":452.7,"bf16":852}},{"kind":"awq","publisher":"cyankiwi","repo":"cyankiwi/MiniMax-M3-AWQ-INT4","url":"https://huggingface.co/cyankiwi/MiniMax-M3-AWQ-INT4"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/MiniMax-M3-4bit","url":"https://huggingface.co/mlx-community/MiniMax-M3-4bit"}],"copy":{"one_liner":"428B/23B 稀疏注意力多模态开源模型，1M 上下文低价。","highlights":["MSA 稀疏注意力：1M 上下文下 prefill 9× / decode 15× 于 M2，每 token 计算降至 1/20","原生多模态（文本 + 图像 + 视频），SWE-Bench Pro 59.0 与 GPT-5.5 持平，API 价 $0.3/$1.2","思考可选：enabled / adaptive / disabled 三档，低延迟场景可关"],"pitfalls":["MiniMax Community License：商用须标注「Built with MiniMax M3」，年收入 > 2000 万美元须书面授权，非 MIT","428B 总参数，BF16 854 GB，比 M2 大近一倍；KV 120 KiB/token 不因 MSA 减少","纯文本编程弱于同期 GLM-5.2 / Kimi K3；LMArena 1442、AA 指数 45，在旗舰中偏低"],"logic_ability":"工程型推理为主：SWE-Bench Pro 59.0、Terminal-Bench 2.1 66.0、MCP Atlas 74.2，BrowseComp 83.5。考试型基准官方未披露（Z.ai 对比表中列 GPQA 93、HLE 37，属第三方引用）。思考默认可切 adaptive。常见问题：长程自主任务（PostTrainBench 37.1）落后 Opus 4.8；多模态强于纯文本推理。","best_for":["图像 / 视频 + 长上下文的 agent","低成本 1M 上下文服务","自建 8×H100 级多模态推理"],"not_for":["单卡部署","对纯文本编程要求最高（用 GLM-5.2 / K3）","不愿标注品牌的商用产品"]},"capability_notes":{"coding":"SWE-Bench Pro 59.0、Terminal-Bench 2.1 66.0、SWE-fficiency 34.8、KernelBench Hard 28.8（官方）。","reasoning":"官方未披露 GPQA / HLE；Z.ai GLM-5.2 对比表引用 GPQA 93、HLE 37。","agent":"BrowseComp 83.5、MCP Atlas 74.2、PostTrainBench 37.1（官方）。","multimodal":"原生图像 + 视频；官方博客未给出 MMMU 数值。","chinese":"中文良好。"},"sheet":{"architecture_md":"**类型**：MoE + 块稀疏注意力，428B 总 / 23B 激活，原生多模态。\n\n- 60 层（前 3 层全注意力 + Dense MLP 12288），隐藏维 6144，词表 200,064\n- 专家：128 路由 + 1 共享，每 token 激活 4 个，专家中间维 3072，sigmoid 路由 + 路由偏置\n- **MSA（MiniMax Sparse Attention）**：GQA 64 Q 头 / 4 KV 头，head_dim 128；4 头索引分支（维 128）对 128-token 块打分选 top-16 块；从第 4 层起启用\n- 半维 RoPE（rotary_dim 64），QK-norm（per-head），Gemma 式 norm，swigluoai 激活\n- 7 个 MTP 模块用于投机解码\n- 上下文 1,048,576，RoPE theta 5e6\n- 视觉：32 层 ViT（隐藏 1280，patch 14，最大 2016px，3D RoPE），投影到 6144\n\n参考：MSA 论文（arXiv 2606.13392）、HF 模型卡。","memory_md":"| 精度 | 权重大小 | 参考显存 |\n|---|---|---|\n| BF16 | 854 GB（官方文件） | 8×H200 |\n| FP8 | ≈ 428 GB（估） | 8×H100 |\n| Q4 | ≈ 248 GB（估） | 4×80GB |\n\n**KV Cache**：≈ 120 KiB/token。128K 上下文 ≈ 15 GB，1M ≈ 117 GB。MSA 降低的是注意力计算，不是 KV 存储。\n\n**参考配置**：\n- 24GB / 80GB：不可行\n- 8×H100：FP8 权重 + 有限并发\n- 8×H200：BF16 舒适\n- ROCm ATOM 提供 MXFP4 / MXFP8 路径","training_md":"- 从预训练第一步混合文本 / 图像 / 视频（mixed-modality training from step zero）\n- 预训练 token 数与 RL 细节官方未披露\n- 思考：`thinking` 参数 enabled / adaptive / disabled\n- 推荐 temperature 1.0、top_p 0.95\n- 最大输出官方未披露","ecosystem_md":"- HF：MiniMaxAI/MiniMax-M3；nvidia/MiniMax-M3-NVFP4\n- 引擎：SGLang、vLLM、Transformers、KTransformers、Unsloth、ROCm ATOM（官方均有指南）\n- 微调：Unsloth 指南；社区少\n- 中文文档：有","versions_md":"- 上代：MiniMax-M2（2025-10，230B/10B）、M2.1、M2.5、M2.7（2026-03，229B/10B，非商用许可）\n- 同期：MiniMax-H3（HF 已上架）\n- M2.5 及更早已在开放平台标记 Legacy"},"ecosystem":{"engines":["SGLang","vLLM","KTransformers","Transformers"],"finetune":"Unsloth 指南","zh_docs":"有"},"complete":true,"runtime":{"tok_s":118,"latency_s":1.13,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"MiniMax M3","one_liner":"428B/23B sparse-attention multimodal open model; cheap 1M context.","highlights":["MSA sparse attention: at 1M context, prefill 9× / decode 15× vs M2, per-token compute down to 1/20","Native multimodal (text + image + video); SWE-Bench Pro 59.0 on par with GPT-5.5; API price $0.3/$1.2","Optional thinking: enabled / adaptive / disabled; can be turned off for low-latency scenarios"],"pitfalls":["MiniMax Community License: commercial use requires a \"Built with MiniMax M3\" notice; annual revenue over $20M needs written authorization; not MIT","428B total parameters, BF16 854 GB, nearly double M2; KV 120 KiB/token is not reduced by MSA","Weaker at pure-text coding than contemporaries GLM-5.2 / Kimi K3; LMArena 1442 and AA index 45 are low among flagships"],"logic_ability":"Mainly engineering reasoning: SWE-Bench Pro 59.0, Terminal-Bench 2.1 66.0, MCP Atlas 74.2, BrowseComp 83.5. Exam-style benchmarks not officially disclosed (Z.ai's comparison table lists GPQA 93, HLE 37, a third-party citation). Thinking can default to adaptive. Common issues: long-horizon autonomous tasks (PostTrainBench 37.1) lag Opus 4.8; multimodal is stronger than pure-text reasoning.","best_for":["Image / video + long-context agents","Low-cost 1M-context services","Self-hosted 8×H100-class multimodal inference"],"not_for":["Single-GPU deployment","Highest demands on pure-text coding (use GLM-5.2 / K3)","Commercial products unwilling to display the brand"],"capability_notes":{"coding":"SWE-Bench Pro 59.0, Terminal-Bench 2.1 66.0, SWE-fficiency 34.8, KernelBench Hard 28.8 (official).","reasoning":"GPQA / HLE not officially disclosed; Z.ai GLM-5.2 comparison table cites GPQA 93, HLE 37.","agent":"BrowseComp 83.5, MCP Atlas 74.2, PostTrainBench 37.1 (official).","multimodal":"Native image + video; official blog gives no MMMU number.","chinese":"Good Chinese."}},"ja":{"name_zh":"MiniMax M3","one_liner":"428B/23B スパース注意の多モーダル OSS。1M を低価格で。","highlights":["MSA スパースアテンション：1M コンテキストで prefill 9 倍 / decode 15 倍（M2 比）、トークンあたり計算量は 1/20 に","ネイティブマルチモーダル（テキスト + 画像 + 動画）、SWE-Bench Pro 59.0 で GPT-5.5 と同等、API 価格 $0.3/$1.2","思考は enabled / adaptive / disabled の 3 段階で選択可、低レイテンシ用途ではオフにできる"],"pitfalls":["MiniMax Community License：商用は「Built with MiniMax M3」の表示が必要、年収 2000 万ドル超は書面許諾が必要。MIT ではない","総パラメータ 428B、BF16 854 GB で M2 の約 2 倍。KV 120 KiB/token は MSA でも減らない","純テキストのコーディングは同時期の GLM-5.2 / Kimi K3 より弱い。LMArena 1442、AA 指数 45 でフラッグシップの中では低め"],"logic_ability":"エンジニアリング型推論が主：SWE-Bench Pro 59.0、Terminal-Bench 2.1 66.0、MCP Atlas 74.2、BrowseComp 83.5。試験型ベンチマークは公式未開示（Z.ai の比較表に GPQA 93、HLE 37 があるが第三者引用）。思考はデフォルトを adaptive に切替可。よくある問題：長期の自律タスク（PostTrainBench 37.1）は Opus 4.8 に後れ、純テキスト推論よりマルチモーダルが強い。","best_for":["画像 / 動画 + 長コンテキストのエージェント","低コストの 1M コンテキストサービス","8×H100 級の自前マルチモーダル推論"],"not_for":["単一 GPU でのデプロイ","純テキストコーディングに最高水準を求める場合（GLM-5.2 / K3 を使う）","ブランド表示を望まない商用製品"],"capability_notes":{"coding":"SWE-Bench Pro 59.0、Terminal-Bench 2.1 66.0、SWE-fficiency 34.8、KernelBench Hard 28.8（公式）。","reasoning":"GPQA / HLE は公式未開示。Z.ai の GLM-5.2 比較表は GPQA 93、HLE 37 を引用。","agent":"BrowseComp 83.5、MCP Atlas 74.2、PostTrainBench 37.1（公式）。","multimodal":"ネイティブ画像 + 動画。公式ブログに MMMU の数値なし。","chinese":"中国語は良好。"}}}},{"aliases":["Mistral-7B-Instruct-v0.3","Mistral-7B-v0.1"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"knowledge":"MMLU 60.1%（v0.1 官方论文）。","chinese":"弱。"},"complete":false,"id":"mistral-7b","name":"Mistral 7B","name_zh":"Mistral · 7B","vendor":"Mistral AI","family":"Mistral","license":"Apache-2.0","license_commercial":true,"weights_url":"https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.3","superseded_by":"mistral-small-3-2","released_at":"2023-09-27","architecture":{"type":"dense","total_params":"7.25B","total_params_b":7.25,"layers":32,"hidden_size":4096,"vocab_size":32768,"kv_heads":8,"head_dim":128,"attention":"GQA（32 Q 头 / 8 KV 头）；v0.1 用 4K 滑窗","notes":"v0.1（2023-09）4K 滑窗 8K 上下文；v0.3（2024-05）词表 32,768 并支持函数调用，上下文 32K。","undisclosed":false},"context":{"max_tokens":32768,"display":"32K"},"memory":{"weight_gb":{"bf16":14.5,"q8":7.7,"q4":4.4},"estimated":false,"kv_per_token_kib":128,"kv_note":"128 KiB/token。","ref_hw_24gb":"BF16 14.5 GB 单卡随便跑","ref_hw_80gb":"高并发","ref_hw_8x80gb":"过剩"},"links":{"official":"https://mistral.ai/news/announcing-mistral-7b","paper":"https://arxiv.org/abs/2310.06825","github":"https://github.com/mistralai/mistral-inference","hf":"https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.3"},"copy":{"one_liner":"7B 击败 Llama 2 13B 的 2023 年经典，Apache-2.0。","highlights":["Apache-2.0，无任何商用限制","GQA + 滑窗注意力，推理快、显存省","衍生了 Zephyr、OpenHermes 等大量微调"],"pitfalls":["能力属 2023 年水平，早已被 Qwen / Gemma 小模型超过","中文弱，词表 32K","v0.1 与 v0.3 上下文与模板不同，易混淆"],"logic_ability":"入门级：无 CoT 训练，MMLU 60% 档。适合简单指令与分类，复杂推理不可靠。","best_for":["教学 / 微调实验","低资源英文任务"],"not_for":["新项目","中文"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama","MLX"],"finetune":"LoRA 极成熟","zh_docs":"弱"}},{"license":"Mistral Research License","license_commercial":"restricted","openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"dense","total_params":"123B","total_params_b":123,"undisclosed":false,"notes":"123B 稠密模型，权重以 MRL 许可公开（研究 / 非商用），商用需另行授权。"},"memory":{"weight_gb":{"bf16":246,"fp8":123,"q4":70},"kv_note":"128K 上下文 KV 占用大，单卡 80GB 需 q4 且限制上下文","ref_hw_80gb":"q4 可跑","ref_hw_8x80gb":"bf16 全上下文","estimated":true},"capability_notes":{},"complete":false,"id":"mistral-large-2","name":"Mistral Large 2","aliases":["mistral-large-2407","mistral-large-2411","Mistral-Large-Instruct-2407"],"vendor":"Mistral AI","family":"Mistral Large","superseded_by":"mistral-large-3","released_at":"2024-07-24","weights_url":"https://huggingface.co/mistralai/Mistral-Large-Instruct-2407","modalities":["text","tools"],"reasoning_mode":"none","context":{"max_tokens":131072,"display":"128K","max_output":8192},"pricing":{"input_per_m":2,"output_per_m":6,"currency":"USD","source":"Mistral 定价页（2024-09 降价后）","as_of":"2024-09-17","note":"发布价 $3 / $9"},"links":{"official":"https://mistral.ai/news/mistral-large-2407","hf":"https://huggingface.co/mistralai/Mistral-Large-Instruct-2407","pricing":"https://mistral.ai/pricing"},"copy":{"one_liner":"123B 稠密开放权重旗舰，代码与多语言追平 GPT-4o。","highlights":["MMLU 84.0%、HumanEval 92%，与 GPT-4o / Llama 3.1 405B 同档","123B 稠密可单节点部署，权重公开","80+ 编程语言、强多语言支持"],"pitfalls":["MRL 许可：商用需付费授权","无视觉（2411 版仍纯文本）","无推理模式，GPQA 未公布"],"logic_ability":"非推理模型中的中上水平，MATH 71.5%；函数调用与多轮遵循较好。","best_for":["历史对照","研究用途的本地部署"],"not_for":["未授权的商用","新项目"]}},{"id":"mistral-large-3","name":"Mistral Large 3","name_zh":"Mistral Large 3","aliases":["mistral-large-2512","Mistral-Large-3-675B-Instruct-2512"],"vendor":"Mistral AI","family":"Mistral Large","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512","status":"current","released_at":"2025-12-02","updated_at":"2026-08-28","modalities":["text","image","tools"],"reasoning_mode":"none","architecture":{"type":"moe","total_params":"675B","active_params":"41B","total_params_b":675,"active_params_b":41,"experts":128,"active_experts":4,"shared_expert":true,"layers":36,"hidden_size":4096,"vocab_size":131072,"kv_heads":32,"head_dim":128,"attention":"MLA（kv_lora_rank 256，32 头）","notes":"细粒度 MoE：128 路由专家 + 1 共享，每 token 4 专家。语言模型 673B/39B 激活 + 2.5B 视觉编码器。YaRN 扩到 1M 位置，官方宣称 256K。官方仅放出 FP8 / NVFP4 指令版与 BF16 基座。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K"},"memory":{"weight_gb":{"bf16":1350,"fp8":675,"q4":392,"q8":715.7},"estimated":true,"ref_hw_80gb":"不可行","ref_hw_8x80gb":"NVFP4 单节点 8×H100 / A100（官方）；FP8 需 8×H200 或 B200"},"pricing":{"input_per_m":0.5,"output_per_m":1.5,"currency":"USD","source":"Mistral 定价页","as_of":"2026-08-28"},"links":{"official":"https://mistral.ai/news/mistral-3/","hf":"https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512","pricing":"https://docs.mistral.ai/inference/pricing"},"variants":[{"kind":"fp8","publisher":"mistralai","repo":"mistralai/Mistral-Large-3-675B-Instruct-2512","url":"https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512","note":"官方发布版（FP8）"},{"kind":"nvfp4","publisher":"mistralai","repo":"mistralai/Mistral-Large-3-675B-Instruct-2512-NVFP4","url":"https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512-NVFP4","note":"官方 NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/Mistral-Large-3-675B-Instruct-2512-GGUF","url":"https://huggingface.co/unsloth/Mistral-Large-3-675B-Instruct-2512-GGUF","sizes":{"q4":407,"q5":478,"q6":553.3,"q8":715.7,"bf16":1347}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/mistralai_Mistral-Large-3-675B-Instruct-2512-GGUF","url":"https://huggingface.co/bartowski/mistralai_Mistral-Large-3-675B-Instruct-2512-GGUF","sizes":{"q4":411.4,"q5":480.8,"q6":554.9,"q8":715.7,"bf16":1347}}],"copy":{"one_liner":"Mistral 675B MoE 开源多模态，Apache-2.0，非推理。","highlights":["Apache-2.0，欧洲最大开源 MoE，3000 张 H200 从零训练","41B 激活，多模态 + 40 余种语言","API 便宜：$0.5 / $1.5"],"pitfalls":["非推理模型，AA 智能指数仅 16，已被 Medium 3.5 全面超越","675B 需整节点 8 卡，个人无法部署","发布 8 个月，Mistral 已转向 Medium 3.5 / Small 4"],"logic_ability":"非推理模型，考试型与工程型均为中游；LMArena 开源非推理类第 2（发布时官方说法）。","best_for":["需要 Apache-2.0 的大模型微调基座","多语言非推理对话"],"not_for":["推理 / 编程主力","单机部署"]},"capability_notes":{"chinese":"官方支持中文，独立数据少。"},"ecosystem":{"engines":["vLLM","SGLang","TensorRT-LLM"],"zh_docs":"无"},"complete":false,"runtime":{"tok_s":30,"latency_s":1.88,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"Mistral Large 3","one_liner":"Mistral's 675B MoE open multimodal model, Apache-2.0, non-reasoning.","highlights":["Apache-2.0, Europe's largest open-source MoE, trained from scratch on 3000 H200s","41B active, multimodal + 40-plus languages","Cheap API: $0.5 / $1.5"],"pitfalls":["Non-reasoning model, AA Intelligence Index only 16; fully surpassed by Medium 3.5","675B needs a full 8-GPU node; not deployable by individuals","8 months after release, Mistral has moved on to Medium 3.5 / Small 4"],"logic_ability":"Non-reasoning model; mid-pack on both exam-style and engineering reasoning; #2 among open-source non-reasoning models on LMArena (official claim at release).","best_for":["Large fine-tuning base requiring Apache-2.0","Multilingual non-reasoning chat"],"not_for":["Primary reasoning / coding model","Single-machine deployment"],"capability_notes":{"chinese":"Chinese officially supported; little independent data."}},"ja":{"name_zh":"Mistral Large 3","one_liner":"Mistral の 675B MoE 多モーダル。Apache-2.0、非推論。","highlights":["Apache-2.0、欧州最大のオープンソース MoE、3000 枚の H200 でゼロから学習","41B アクティブ、マルチモーダル + 40 超の言語","API が安い：$0.5 / $1.5"],"pitfalls":["非推論モデルで AA 知能指数はわずか 16、Medium 3.5 に全面的に抜かれた","675B はフル 8 GPU ノードが必要で、個人ではデプロイ不可","公開から 8 か月、Mistral はすでに Medium 3.5 / Small 4 に軸足を移している"],"logic_ability":"非推論モデル。試験型・エンジニアリング型とも中位。LMArena のオープンソース非推論部門 2 位（公開時の公式主張）。","best_for":["Apache-2.0 が必要な大型ファインチューニング基盤","多言語の非推論対話"],"not_for":["推論 / コーディングの主力","単一マシンでのデプロイ"],"capability_notes":{"chinese":"公式に中国語対応、独立データは少ない。"}}}},{"id":"mistral-medium-3-1","name":"Mistral Medium 3.1","aliases":["mistral-medium-2508"],"vendor":"Mistral AI","family":"Mistral Medium","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","released_at":"2025-08-12","updated_at":"2025-12-20","modalities":["text","image","tools"],"reasoning_mode":"none","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露；可私有化授权部署（4 卡起），但权重不公开。"},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":0.4,"output_per_m":2,"currency":"USD","source":"Mistral 定价页","as_of":"2025-12-20"},"links":{"official":"https://mistral.ai/news/mistral-medium-3","pricing":"https://mistral.ai/pricing"},"copy":{"one_liner":"欧洲厂商的中档闭源，可签约私有化部署。","highlights":["$0.4 / $2","企业私有化选项（合同）","欧盟数据合规"],"pitfalls":["分数落后同价位美中模型","非推理模型","权重不公开"],"logic_ability":"非推理模型，中游。","best_for":["欧盟合规场景"],"not_for":["榜单分数优先"]},"capability_notes":{},"complete":false,"superseded_by":"mistral-medium-3-5"},{"id":"mistral-medium-3-5","name":"Mistral Medium 3.5","name_zh":"Mistral Medium 3.5","aliases":["mistral-medium-3-5","Mistral-Medium-3.5-128B","mistral-medium-2604"],"vendor":"Mistral AI","family":"Mistral Medium","license":"Modified MIT License","license_commercial":"restricted","openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/mistralai/Mistral-Medium-3.5-128B","status":"current","released_at":"2026-04-29","updated_at":"2026-08-28","modalities":["text","image","tools"],"reasoning_mode":"optional","architecture":{"type":"dense","total_params":"128B","total_params_b":128,"layers":88,"hidden_size":12288,"vocab_size":131072,"kv_heads":8,"head_dim":128,"attention":"GQA（96 Q 头 / 8 KV 头），无滑窗","notes":"稠密 Transformer，YaRN RoPE（theta 1e6，factor 64）扩到 256K。自研视觉编码器（48 层，1664 维，1540px，可变分辨率 / 长宽比）。官方配套 EAGLE 投机解码器。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K"},"memory":{"weight_gb":{"bf16":256,"fp8":128,"q4":74,"q8":132.9},"kv_per_token_kib":352,"kv_note":"8 KV 头 × 128 × 2 × 88 层 × 2 B = 352 KiB/token，128K 上下文约 44 GB，长上下文显存占用不小。","ref_hw_24gb":"不可行","ref_hw_80gb":"NVFP4 / Q4 约 74 GB 勉强单卡，上下文很短；官方说法「最少 4 张 GPU」","ref_hw_8x80gb":"BF16 TP8 官方示例配置，长上下文高并发","estimated":true},"pricing":{"input_per_m":1.5,"output_per_m":7.5,"currency":"USD","source":"Mistral 定价页","as_of":"2026-08-28","note":"缓存输入九折"},"links":{"official":"https://mistral.ai/news/vibe-remote-agents-mistral-medium-3-5/","hf":"https://huggingface.co/mistralai/Mistral-Medium-3.5-128B","pricing":"https://docs.mistral.ai/inference/pricing"},"variants":[{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/Mistral-Medium-3.5-128B-NVFP4","url":"https://huggingface.co/nvidia/Mistral-Medium-3.5-128B-NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/Mistral-Medium-3.5-128B-GGUF","url":"https://huggingface.co/unsloth/Mistral-Medium-3.5-128B-GGUF","sizes":{"q4":74.9,"q5":88.3,"q6":102.6,"q8":132.9,"bf16":250.1}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/mistralai_Mistral-Medium-3.5-128B-GGUF","url":"https://huggingface.co/bartowski/mistralai_Mistral-Medium-3.5-128B-GGUF","sizes":{"q4":78.4,"q5":91.1,"q6":107.8,"q8":132.9}},{"kind":"awq","publisher":"cyankiwi","repo":"cyankiwi/Mistral-Medium-3.5-128B-AWQ-INT4","url":"https://huggingface.co/cyankiwi/Mistral-Medium-3.5-128B-AWQ-INT4"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/Mistral-Medium-3.5-128B-4bit","url":"https://huggingface.co/mlx-community/Mistral-Medium-3.5-128B-4bit"},{"kind":"other","publisher":"turboderp","repo":"turboderp/Mistral-Medium-3.5-128B-exl3","url":"https://huggingface.co/turboderp/Mistral-Medium-3.5-128B-exl3","note":"EXL3（ExLlamaV3）"}],"copy":{"one_liner":"Mistral 128B 稠密开源旗舰，SWE-bench 77.6%。","highlights":["SWE-bench Verified 77.6%、τ³-Telecom 91.4%（官方），Mistral 系最强工程模型","指令 / 推理 / 编程 / 视觉合一，reasoning effort 按请求可调","开放权重，256K 上下文，配 EAGLE 投机解码"],"pitfalls":["Modified MIT：公司上月全球合并营收超 2000 万美元即不得使用，需另签商业许可","128B 稠密，单卡 80GB 只能勉强 Q4 跑短上下文，实际要 4 卡起","API 输出 $7.5/M 偏贵；中文能力独立数据少"],"logic_ability":"工程型推理强（SWE-bench 77.6%、AA 智能指数 30 居开源中上），考试型推理官方未给 GPQA / AIME 数字，独立复测有限。reasoning=none 时退化为普通指令模型。","best_for":["自托管编程 agent / Vibe 类工具","多工具长任务 agent","欧洲合规场景"],"not_for":["大型企业（营收超阈值）自部署","24GB 单卡本地"]},"capability_notes":{"coding":"SWE-bench Verified 77.6%（官方）。","agent":"τ³-Telecom 91.4%（官方）。","multimodal":"自研视觉编码器，支持可变分辨率图片。","chinese":"官方多语言支持含中文，独立数据少。"},"sheet":{"architecture_md":"**类型**：稠密 Transformer，128B 参数。\n\n- 88 层，隐藏维 12288，FFN 28672，词表 131,072\n- GQA：96 Q 头 / 8 KV 头，head_dim 128，无滑窗\n- RoPE theta 1e6 + YaRN（factor 64）扩到 256K\n- 视觉编码器：从零训练的 Pixtral 系 48 层、1664 维、1540px、patch 14，2× 空间合并，支持可变图片尺寸与长宽比\n- 官方提供 EAGLE 投机解码器（Mistral-Medium-3.5-128B-EAGLE）\n\n参考：HF 模型卡 config.json。","memory_md":"| 精度 | 权重大小 | 参考显存 |\n|---|---|---|\n| BF16 | ≈ 256 GB | 4×80GB 起（官方「最少 4 张 GPU」） |\n| FP8 | ≈ 128 GB | 2×80GB |\n| NVFP4 / Q4（估算） | ≈ 74 GB | 单卡 80GB 极勉强，仅短上下文 |\n\n**KV Cache**：8 KV 头 × 128 × 2 × 88 层 × 2 B = 352 KiB/token；32K 上下文 ≈ 11 GB，128K ≈ 44 GB。\n\n**参考配置**：\n- 24GB：不可行\n- 80GB：Q4 权重 74 GB + 几 GB KV，仅原型验证\n- 4×80GB：FP8 + 长上下文，实际推荐起点\n- 8×80GB：官方 vLLM 示例 `--tensor-parallel-size 8`","training_md":"- 预训练规模未披露\n- 首个把指令 / 推理 / 编程合并进单一权重的 Mistral 旗舰\n- 推理：`reasoning_effort` 按请求 none / high；high 推荐 temperature 0.7、top_p 0.95\n- 原生函数调用、JSON 输出；官方主打多工具长时程 agent（τ³-Telecom 91.4%）\n- 官方公布：SWE-bench Verified 77.6%\n- 知识截止未披露","ecosystem_md":"- HF：mistralai/Mistral-Medium-3.5-128B（另有 EAGLE 版；NVIDIA 官方 NVFP4 版）\n- 引擎：vLLM（需 mistral_common ≥ 1.11.1、transformers ≥ 5.4）、SGLang、Transformers、llama.cpp、Ollama\n- 社区量化：Unsloth GGUF、AWQ-INT4\n- API：Mistral 平台 $1.5 / $7.5，缓存输入九折；Mistral Vibe 远程 agent 默认模型\n- 中文文档：无（官方英文）\n- 注意：发布初期修过配置，旧版 GGUF 长上下文质量下降","versions_md":"- 2604：Mistral Medium 3.5（本条）\n- 前代：Mistral Medium 3.1（2508，API-only）→ 被本条取代\n- 同代：Mistral Large 3（675B MoE，Apache-2.0）、Mistral Small 4（119B MoE）"},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Ollama","Transformers"],"finetune":"全量 / LoRA（社区）","zh_docs":"无"},"complete":true,"runtime":{"tok_s":147,"latency_s":2.25,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"Mistral Medium 3.5","one_liner":"Mistral's 128B dense open flagship, SWE-bench 77.6%.","highlights":["SWE-bench Verified 77.6%, τ³-Telecom 91.4% (official); Mistral's strongest engineering model","Instruct / reasoning / coding / vision in one; reasoning effort adjustable per request","Open weights, 256K context, with EAGLE speculative decoding"],"pitfalls":["Modified MIT: companies whose global consolidated revenue exceeded $20M last month may not use it without a separate commercial license","128B dense; a single 80GB GPU barely runs Q4 with short context; realistically 4 GPUs minimum","API output $7.5/M is pricey; little independent data on Chinese ability"],"logic_ability":"Strong engineering reasoning (SWE-bench 77.6%, AA Intelligence Index 30, upper-mid among open-source); exam-style reasoning has no official GPQA / AIME numbers and limited independent re-testing. With reasoning=none it degrades to a plain instruct model.","best_for":["Self-hosted coding agents / vibe-coding tools","Multi-tool long-task agents","European compliance scenarios"],"not_for":["Self-deployment by large enterprises (revenue over threshold)","Local 24GB single GPU"],"capability_notes":{"coding":"SWE-bench Verified 77.6% (official).","agent":"τ³-Telecom 91.4% (official).","multimodal":"In-house vision encoder supporting variable-resolution images.","chinese":"Official multilingual support includes Chinese; little independent data."}},"ja":{"name_zh":"Mistral Medium 3.5","one_liner":"Mistral の 128B 稠密オープン旗艦。SWE-bench 77.6%。","highlights":["SWE-bench Verified 77.6%、τ³-Telecom 91.4%（公式）、Mistral 系で最強のエンジニアリングモデル","指示 / 推論 / コーディング / ビジョンを統合、reasoning effort はリクエストごとに調整可","オープンウェイト、256K コンテキスト、EAGLE 投機的デコーディング対応"],"pitfalls":["Modified MIT：前月の全世界連結売上が 2000 万ドル超の企業は利用不可で、別途商用ライセンスが必要","128B 稠密、80GB 単一 GPU では Q4 で短コンテキストがやっと。実際は 4 GPU 以上","API 出力 $7.5/M はやや高い。中国語能力の独立データが少ない"],"logic_ability":"エンジニアリング型推論は強い（SWE-bench 77.6%、AA 知能指数 30 でオープンソース中上位）。試験型推論は公式の GPQA / AIME 数値がなく、独立再検証も限定的。reasoning=none では通常の指示モデルに退化。","best_for":["自前ホスティングのコーディングエージェント / Vibe 系ツール","マルチツールの長時間タスクエージェント","欧州のコンプライアンス用途"],"not_for":["大企業（売上が閾値超）の自前デプロイ","24GB 単一 GPU のローカル環境"],"capability_notes":{"coding":"SWE-bench Verified 77.6%（公式）。","agent":"τ³-Telecom 91.4%（公式）。","multimodal":"自社製ビジョンエンコーダ、可変解像度画像に対応。","chinese":"公式の多言語対応に中国語を含む。独立データは少ない。"}}}},{"aliases":["Mistral-Small-3.1-24B-Instruct-2503"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","image","tools"],"reasoning_mode":"none","capability_notes":{"coding":"HumanEval 88.41%（官方）。","reasoning":"GPQA Diamond 45.96%（5-shot CoT，官方）。","math":"MATH 69.3%（官方）。","multimodal":"MMMU 64.0%（官方）。","chinese":"一般。"},"complete":false,"id":"mistral-small-3-1","name":"Mistral Small 3.1","name_zh":"Mistral Small 3.1 · 24B","vendor":"Mistral AI","family":"Mistral Small","license":"Apache-2.0","license_commercial":true,"weights_url":"https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503","superseded_by":"mistral-small-3-2","released_at":"2025-03-17","architecture":{"type":"dense","total_params":"24B","total_params_b":24,"layers":40,"hidden_size":5120,"vocab_size":131072,"kv_heads":8,"head_dim":128,"attention":"GQA（32 Q 头 / 8 KV 头）","notes":"在 Small 3 基础上加入视觉编码器（Pixtral 式）与 128K 上下文。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":48,"q8":25.4,"q4":14.3},"estimated":false,"kv_per_token_kib":160,"kv_note":"160 KiB/token；32K 约 5 GB。","ref_hw_24gb":"Q4_K_M ≈14 GB，可跑 32K+；BF16 48 GB 需 2 卡","ref_hw_80gb":"BF16 单卡 + 大 KV","ref_hw_8x80gb":"过剩"},"pricing":{"input_per_m":0.1,"output_per_m":0.3,"currency":"USD","source":"Mistral 官方 API（mistral-small-latest）","as_of":"2026-08-28","note":"官方托管价"},"links":{"official":"https://mistral.ai/news/mistral-small-3-1","hf":"https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503"},"copy":{"one_liner":"24B 多模态 Apache-2.0，单卡跑图文 128K 上下文。","highlights":["Apache-2.0，图文多模态 + 128K，24GB 卡可跑 Q4","官方称超越 Gemma 3 / GPT-4o mini","函数调用与 JSON 输出稳定"],"pitfalls":["无推理模式，GPQA Diamond 46%","中文一般，欧洲语言更强","3.2 版（2025-06）修复了重复输出等问题，建议直接用 3.2"],"logic_ability":"非推理 24B 档中游：GPQA Diamond 45.96%、MATH 69.3%（官方）。日常逻辑够用，复杂推理让位给 Magistral / Qwen3。","best_for":["单卡多模态助手","低延迟函数调用"],"not_for":["深度推理","中文优先"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama"],"finetune":"LoRA 友好","zh_docs":"弱"}},{"id":"mistral-small-3-2","name":"Mistral Small 3.2","aliases":["mistral-small-2506"],"vendor":"Mistral AI","family":"Mistral Small","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/mistralai/Mistral-Small-3.2-24B-Instruct-2506","status":"superseded","released_at":"2025-06-20","updated_at":"2025-12-20","modalities":["text","image","tools"],"reasoning_mode":"none","architecture":{"type":"dense","total_params":"24B","total_params_b":24,"layers":40,"hidden_size":5120,"kv_heads":8,"head_dim":128,"attention":"GQA（32 Q / 8 KV）","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":48,"q8":25.5,"q4":14.5},"kv_per_token_kib":160,"estimated":false,"ref_hw_24gb":"Q4 14.5 GB 舒适"},"pricing":{"input_per_m":0.1,"output_per_m":0.3,"currency":"USD","source":"Mistral 定价页","as_of":"2025-12-20"},"links":{"official":"https://mistral.ai/news/mistral-small-3-1","hf":"https://huggingface.co/mistralai/Mistral-Small-3.2-24B-Instruct-2506"},"copy":{"one_liner":"Apache-2.0 的 24B 多模态通用模型，欧洲语言强。","highlights":["Apache-2.0，24GB 单卡","视觉输入","指令遵循较 3.1 明显改善"],"pitfalls":["非推理模型","中文弱于 Qwen","代码一般"],"logic_ability":"非推理模型，中游。","best_for":["单卡多语言助手"],"not_for":["推理密集"]},"capability_notes":{},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama"],"zh_docs":"无"},"complete":false,"superseded_by":"mistral-small-4"},{"id":"mistral-small-4","name":"Mistral Small 4","name_zh":"Mistral Small 4","aliases":["mistral-small-2603","Mistral-Small-4-119B-2603"],"vendor":"Mistral AI","family":"Mistral Small","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/mistralai/Mistral-Small-4-119B-2603","status":"current","released_at":"2026-03-16","updated_at":"2026-08-28","modalities":["text","image","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"119B","active_params":"6.5B","total_params_b":119,"active_params_b":6.5,"experts":128,"active_experts":4,"shared_expert":true,"layers":36,"hidden_size":4096,"vocab_size":131072,"kv_heads":32,"head_dim":128,"attention":"MLA（kv_lora_rank 256 + RoPE 64，q_lora_rank 1024，32 头，head_dim 128）","notes":"与 Large 3 同构的细粒度 MoE：128 路由 + 1 共享专家，每 token 选 4，专家中间维 2048，全部 36 层均为 MoE（first_k_dense_replace 0）。MLA：qk_nope 64 + qk_rope 64、v_head 128，YaRN factor 128（原始 8K → config 1M，官方上下文 256K）。Pixtral 视觉编码器 24 层、hidden 1024、patch 14、image 1540。合并 Magistral 推理、Pixtral 视觉、Devstral 编程于一体。官方权重为 FP8（静态激活量化）。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K"},"memory":{"weight_gb":{"bf16":238,"fp8":121,"q8":126,"q4":73.8},"kv_per_token_kib":22.5,"kv_note":"MLA 压缩缓存：(256 潜向量 + 64 RoPE) × 2 B = 640 B/层 × 36 层 ≈ 22.5 KiB/token。若引擎不做 MLA 吸收而缓存完整 K/V（32 头 × 128 × 2 × 2 B），则为 576 KiB/token。256K 上下文 MLA 缓存 ≈ 5.6 GB。","ref_hw_24gb":"不可行","ref_hw_80gb":"UD-Q4_K_M 73.8 GB 单卡勉强可载，仅短上下文；官方最低 4×H100 / 2×H200 / 1×B200","ref_hw_8x80gb":"官方 FP8 121 GB，4×H100 推荐配置，8×H100 高并发","estimated":false},"pricing":{"input_per_m":0.15,"output_per_m":0.6,"currency":"USD","source":"Mistral 定价页","as_of":"2026-08-28"},"links":{"official":"https://mistral.ai/news/mistral-small-4/","hf":"https://huggingface.co/mistralai/Mistral-Small-4-119B-2603","pricing":"https://docs.mistral.ai/inference/pricing"},"variants":[{"kind":"fp8","publisher":"mistralai","repo":"mistralai/Mistral-Small-4-119B-2603","url":"https://huggingface.co/mistralai/Mistral-Small-4-119B-2603","note":"官方发布版（FP8）","sizes":{"fp8":120.9}},{"kind":"nvfp4","publisher":"mistralai","repo":"mistralai/Mistral-Small-4-119B-2603-NVFP4","url":"https://huggingface.co/mistralai/Mistral-Small-4-119B-2603-NVFP4","note":"官方 NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/Mistral-Small-4-119B-2603-GGUF","url":"https://huggingface.co/unsloth/Mistral-Small-4-119B-2603-GGUF","sizes":{"q8":126.5,"bf16":238}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/mistralai_Mistral-Small-4-119B-2603-GGUF","url":"https://huggingface.co/bartowski/mistralai_Mistral-Small-4-119B-2603-GGUF","sizes":{"q4":72.6,"q5":85,"q6":102.8,"q8":126.5}},{"kind":"awq","publisher":"cyankiwi","repo":"cyankiwi/Mistral-Small-4-119B-2603-AWQ-4bit","url":"https://huggingface.co/cyankiwi/Mistral-Small-4-119B-2603-AWQ-4bit"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/Mistral-Small-4-119B-2603-4bit","url":"https://huggingface.co/mlx-community/Mistral-Small-4-119B-2603-4bit"}],"copy":{"one_liner":"119B MoE 仅 6.5B 激活，推理 / 视觉 / 编程三合一。","highlights":["Apache-2.0，指令 + 推理 + 多模态 + 编程 agent 单一权重，reasoning_effort 可切","6.5B 激活，官方称比 Small 3 吞吐 3 倍、端到端延迟降 40%","GPQA Diamond 71.2、AIME 2025 84、LiveCodeBench 64（high），输出比 gpt-oss-120b 短 20%"],"pitfalls":["名为 Small 实为 119B，FP8 121 GB，24GB 单卡跑不了，官方最低 4×H100","推理档输出很长（AIME 均 27.9K 字符），关思考后 GPQA 掉到 59.1、AIME 36","Agent 类官方未给 SWE-bench / Terminal-Bench / tau2 数字，IFBench 仅 48"],"logic_ability":"考试型推理中游：reasoning_effort=high 时 GPQA Diamond 71.2、AIME 2025 84（另一张图 83.8）、MMLU Pro 78，与 Magistral Medium 1.2 持平；instruct 模式下 GPQA 59.1、AIME 36、LiveCodeBench 32，差距很大。工程型推理：LiveCodeBench 64（high），继承 Devstral 但官方未给 SWE-bench 数字。长上下文 AA LCR 72。指令遵循偏弱（IFBench 48）。","best_for":["低成本 API 多模态 agent","80GB 单卡本地推理"],"not_for":["24GB 消费卡","高难度数学 / 推理"]},"capability_notes":{"coding":"LiveCodeBench 64（high）/ 32（instruct）（官方模型卡图表）；SWE-bench 未披露。","reasoning":"GPQA Diamond 71.2（high）/ 59.1（instruct）、MMLU Pro 78（官方模型卡图表）；HLE 未披露。","math":"AIME 2025 84（high，图 1）/ 83.8（图 2）/ 36（instruct）（官方模型卡图表）。","agent":"tau2 / Terminal-Bench 官方未披露；Collie 62.9。","instruction":"AllenAI IFBench 48（high）/ 35.7（instruct）、Arena Hard 58.3（官方模型卡图表）。","multimodal":"MMMU-Pro 60（high）/ 46.3（instruct）（官方模型卡图表）。","chinese":"官方支持中文，独立数据少。"},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Transformers"],"finetune":"mistral-finetune / Unsloth；全参需 ≥ 4×80GB","zh_docs":"无"},"complete":true,"runtime":{"tok_s":164,"latency_s":0.89,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"vendor_zh":"Mistral AI","sheet":{"architecture_md":"**类型**：细粒度 MoE + MLA，119B 总参 / 6.5B 激活，图像输入。\n\n| 项目 | 数值 |\n|---|---|\n| 层数 | 36（全部 MoE） |\n| 隐藏维 | 4,096 |\n| 专家 | 128 路由 + 1 共享，每 token 选 4，专家中间维 2,048 |\n| 注意力 | MLA：32 头，q_lora_rank 1024，kv_lora_rank 256，qk_nope 64 + qk_rope 64，v_head 128 |\n| RoPE | YaRN factor 128，原始 8,192，theta 1e4 |\n| 词表 | 131,072（Tekken） |\n| 上下文 | 官方 256K（config max_position 1,048,576） |\n| 视觉编码器 | Pixtral 24 层，hidden 1024，patch 14，image 1540，spatial merge 2 |\n\n`config.json` 中 `model_type: mistral3` / `text_config.model_type: mistral4`，`quant_method: fp8`。\n\n参考：mistralai/Mistral-Small-4-119B-2603 config.json。","memory_md":"| 精度 | 权重大小 | 参考显存 | 说明 |\n|---|---|---|---|\n| BF16 | 238 GB | 4×80GB | unsloth GGUF BF16（官方未发布 BF16） |\n| FP8 | 121 GB | 2×80GB | 官方 safetensors，3 分片 120.9 GB（另有 7 分片 consolidated 格式 141 GB） |\n| Q8 | 126 GB | 2×80GB | unsloth GGUF Q8_0 |\n| Q4 | 73.8 GB | 80GB 单卡 | unsloth UD-Q4_K_M；UD-Q4_K_S 69.5 GB |\n\n**KV Cache**：MLA 压缩缓存 ≈ 22.5 KiB/token（(256 + 64) × 2 B × 36 层）；非吸收实现 576 KiB/token。256K 上下文 MLA 缓存 ≈ 5.6 GB。\n\n**参考配置**：\n- 24GB：不可行\n- 80GB 单卡：Q4 勉强可载，短上下文\n- 4×H100 / 2×H200 / 1×B200：官方最低；4×H200 / 2×B200：官方推荐\n\n警示：high 推理档输出很长，预留输出 KV；vLLM 需 tensor-parallel ≥ 2。","training_md":"- 预训练 token 数：**未披露**\n- 单一权重同时覆盖指令、推理（Magistral）、视觉（Pixtral）、编程 agent（Devstral）\n- reasoning_effort：high（深度推理，推荐 temperature 0.7）/ none（快速响应）\n- 官方发布即 FP8（静态激活量化），无 BF16 官方权重\n- 官方评测：AA LCR 0.72 且平均输出仅 1.6K 字符","ecosystem_md":"- HF：mistralai/Mistral-Small-4-119B-2603（FP8 121 GB）；GGUF：unsloth/Mistral-Small-4-119B-2603-GGUF\n- 引擎：vLLM（官方推荐）、SGLang、llama.cpp、LM Studio、Transformers（main 分支）\n- 微调：mistral-finetune / Unsloth；MoE 全参需 ≥ 4×80GB\n- 中文文档：无","versions_md":"- 同代：Mistral Large 3（同构更大 MoE）、Magistral Medium 1.2 / Small 1.2（纯推理）、Devstral 2\n- 上代：Mistral Small 3.2（24B Dense）、Mistral Medium 3.1\n- 本代把 Small 从 24B Dense 改为 119B-A6.5B MoE，并合并推理 / 视觉 / 编程分支"},"i18n":{"en":{"name_zh":"Mistral Small 4","one_liner":"119B MoE with only 6.5B active; reasoning / vision / coding in one.","highlights":["Apache-2.0; instruct + reasoning + multimodal + coding agent in a single set of weights, reasoning_effort switchable","6.5B active; officially 3× the throughput of Small 3 with 40% lower end-to-end latency","GPQA Diamond 71.2, AIME 2025 84, LiveCodeBench 64 (high); outputs 20% shorter than gpt-oss-120b"],"pitfalls":["Called Small but actually 119B, FP8 121 GB; will not run on a 24GB single GPU, official minimum 4×H100","Reasoning-tier outputs are very long (AIME avg 27.9K chars); with thinking off GPQA drops to 59.1, AIME to 36","No official SWE-bench / Terminal-Bench / tau2 numbers for agent tasks; IFBench only 48"],"logic_ability":"Mid-pack exam-style reasoning: at reasoning_effort=high, GPQA Diamond 71.2, AIME 2025 84 (83.8 in another chart), MMLU Pro 78, on par with Magistral Medium 1.2; in instruct mode GPQA 59.1, AIME 36, LiveCodeBench 32, a large gap. Engineering reasoning: LiveCodeBench 64 (high); inherits Devstral but no official SWE-bench number. Long context AA LCR 72. Instruction following is weak (IFBench 48).","best_for":["Low-cost API multimodal agents","Local inference on an 80GB single GPU"],"not_for":["24GB consumer GPUs","Hard math / reasoning"],"capability_notes":{"coding":"LiveCodeBench 64 (high) / 32 (instruct) (official model card charts); SWE-bench not disclosed.","reasoning":"GPQA Diamond 71.2 (high) / 59.1 (instruct), MMLU Pro 78 (official model card charts); HLE not disclosed.","math":"AIME 2025 84 (high, chart 1) / 83.8 (chart 2) / 36 (instruct) (official model card charts).","agent":"tau2 / Terminal-Bench not officially disclosed; Collie 62.9.","instruction":"AllenAI IFBench 48 (high) / 35.7 (instruct), Arena Hard 58.3 (official model card charts).","multimodal":"MMMU-Pro 60 (high) / 46.3 (instruct) (official model card charts).","chinese":"Chinese officially supported; little independent data."}},"ja":{"name_zh":"Mistral Small 4","one_liner":"119B MoE、6.5B 活性。推論 / 視覚 / コーディングを統合。","highlights":["Apache-2.0、指示 + 推論 + マルチモーダル + コーディングエージェントを単一の重みで、reasoning_effort 切替可","6.5B アクティブ、公式は Small 3 比でスループット 3 倍、エンドツーエンド遅延 40% 減と主張","GPQA Diamond 71.2、AIME 2025 84、LiveCodeBench 64（high）、出力は gpt-oss-120b より 20% 短い"],"pitfalls":["Small の名だが実際は 119B、FP8 121 GB。24GB 単一 GPU では動かず、公式最低構成は 4×H100","推論設定の出力は非常に長い（AIME 平均 27.9K 文字）。思考オフでは GPQA 59.1、AIME 36 に低下","エージェント系は SWE-bench / Terminal-Bench / tau2 の公式数値なし、IFBench はわずか 48"],"logic_ability":"試験型推論は中位：reasoning_effort=high で GPQA Diamond 71.2、AIME 2025 84（別の図では 83.8）、MMLU Pro 78 で Magistral Medium 1.2 と同等。instruct モードでは GPQA 59.1、AIME 36、LiveCodeBench 32 と差が大きい。エンジニアリング型推論：LiveCodeBench 64（high）、Devstral を継承するが公式 SWE-bench 数値なし。長コンテキスト AA LCR 72。指示追従はやや弱い（IFBench 48）。","best_for":["低コスト API のマルチモーダルエージェント","80GB 単一 GPU でのローカル推論"],"not_for":["24GB コンシューマー GPU","高難度の数学 / 推論"],"capability_notes":{"coding":"LiveCodeBench 64（high）/ 32（instruct）（公式モデルカード図表）。SWE-bench 未開示。","reasoning":"GPQA Diamond 71.2（high）/ 59.1（instruct）、MMLU Pro 78（公式モデルカード図表）。HLE 未開示。","math":"AIME 2025 84（high、図 1）/ 83.8（図 2）/ 36（instruct）（公式モデルカード図表）。","agent":"tau2 / Terminal-Bench は公式未開示。Collie 62.9。","instruction":"AllenAI IFBench 48（high）/ 35.7（instruct）、Arena Hard 58.3（公式モデルカード図表）。","multimodal":"MMMU-Pro 60（high）/ 46.3（instruct）（公式モデルカード図表）。","chinese":"公式に中国語対応、独立データは少ない。"}}}},{"aliases":["Mixtral-8x22B-Instruct-v0.1"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"coding":"HumanEval 45.1%（官方基座）。","math":"GSM8K 78.6%（官方）。","knowledge":"MMLU 77.8%（官方）。","chinese":"弱。"},"complete":false,"id":"mixtral-8x22b","name":"Mixtral 8x22B","name_zh":"Mixtral · 8x22B","vendor":"Mistral AI","family":"Mixtral","license":"Apache-2.0","license_commercial":true,"weights_url":"https://huggingface.co/mistralai/Mixtral-8x22B-Instruct-v0.1","superseded_by":"mistral-large-3","released_at":"2024-04-17","architecture":{"type":"moe","total_params":"141B","active_params":"39B","total_params_b":141,"active_params_b":39,"experts":8,"active_experts":2,"shared_expert":false,"layers":56,"hidden_size":6144,"vocab_size":32768,"kv_heads":8,"head_dim":128,"attention":"GQA（48 Q 头 / 8 KV 头）","notes":"64K 上下文；原生函数调用；数学与多语言强化。","undisclosed":false},"context":{"max_tokens":65536,"display":"64K"},"memory":{"weight_gb":{"bf16":282,"q8":149.5,"q4":85.6},"estimated":false,"kv_per_token_kib":224,"kv_note":"224 KiB/token；64K 约 14 GB。","ref_hw_24gb":"不可行","ref_hw_80gb":"Q4_K_M 85.6 GB 略超单卡，需 2×80GB；BF16 4×80GB","ref_hw_8x80gb":"BF16 高并发"},"links":{"official":"https://mistral.ai/news/mixtral-8x22b","github":"https://github.com/mistralai/mistral-inference","hf":"https://huggingface.co/mistralai/Mixtral-8x22B-Instruct-v0.1"},"copy":{"one_liner":"Mistral 最后一款开源大 MoE，Apache-2.0 的 141B。","highlights":["Apache-2.0，2024 年初最大的无限制商用模型之一","64K 上下文 + 原生函数调用","39B 激活，吞吐远高于同期 70B Dense"],"pitfalls":["141B 总参显存压力大，至少 2×80GB","中文弱，词表 32K","被随后的 Llama 3 70B / Qwen2 72B 迅速超过"],"logic_ability":"非推理模型中游偏上：MMLU 77.8%、GSM8K 78.6%（官方）。多语言与函数调用是亮点，复杂推理一般。","best_for":["需 Apache-2.0 的大模型私有化","欧洲多语言"],"not_for":["中文","单卡部署"]},"ecosystem":{"engines":["vLLM","llama.cpp","TGI"],"finetune":"LoRA 可行，成本高","zh_docs":"弱"}},{"aliases":["Mixtral-8x7B-Instruct-v0.1"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"coding":"HumanEval 40.2%（官方基座）。","math":"GSM8K 74.4%（官方基座，8-shot）。","knowledge":"MMLU 70.6%（官方）。","chinese":"弱。"},"complete":true,"id":"mixtral-8x7b","name":"Mixtral 8x7B","name_zh":"Mixtral · 8x7B","vendor":"Mistral AI","family":"Mixtral","license":"Apache-2.0","license_commercial":true,"weights_url":"https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1","superseded_by":"mistral-small-3-2","released_at":"2023-12-11","architecture":{"type":"moe","total_params":"46.7B","active_params":"12.9B","total_params_b":46.7,"active_params_b":12.9,"experts":8,"active_experts":2,"shared_expert":false,"layers":32,"hidden_size":4096,"vocab_size":32000,"kv_heads":8,"head_dim":128,"attention":"GQA（32 Q 头 / 8 KV 头）","notes":"第一款广泛流行的开源 MoE：8 专家 top-2，FFN 14,336；上下文 32K。","undisclosed":false},"context":{"max_tokens":32768,"display":"32K"},"memory":{"weight_gb":{"bf16":93.4,"q8":49.5,"q4":26.4},"estimated":false,"kv_per_token_kib":128,"kv_note":"128 KiB/token（与 7B 相同，MoE 不增加 KV）；32K 约 4 GB。","ref_hw_24gb":"Q4_K_M 26.4 GB 略超 24GB，需 Q3 或 CPU 卸载部分专家","ref_hw_80gb":"BF16 93 GB 需 2×80GB；Q8 单卡 80GB","ref_hw_8x80gb":"高并发"},"pricing":{"input_per_m":0.24,"output_per_m":0.24,"currency":"USD","source":"第三方托管常见价（DeepInfra）","as_of":"2026-08-28","note":"开源权重，社区托管价"},"links":{"official":"https://mistral.ai/news/mixtral-of-experts","paper":"https://arxiv.org/abs/2401.04088","github":"https://github.com/mistralai/mistral-inference","hf":"https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1"},"copy":{"one_liner":"把 MoE 带进开源主流的里程碑，12.9B 激活打平 Llama 2 70B。","highlights":["首个大规模流行的开源稀疏 MoE，推理速度接近 13B Dense","Apache-2.0 无限制商用，当年 Arena 开源第一","32K 上下文，法德西意多语言"],"pitfalls":["总参 46.7B，显存按 47B 算，24GB 卡放不下 Q4","中文弱、无工具调用原生格式","2025 年 24B Dense（Mistral Small 3）已全面超过它"],"logic_ability":"2023 年末开源顶级：MMLU 70.6%、GSM8K 74.4%（官方基座数据），推理与 Llama 2 70B 相当。无 CoT 专训，复杂题目会错但格式稳定。","best_for":["MoE 研究 / 教学对照","已有部署的维护"],"not_for":["中文","代码 / 数学新项目"]},"sheet":{"architecture_md":"**类型**：稀疏 MoE Transformer，46.7B 总参 / 12.9B 激活。\n\n- 32 层，隐藏维 4096，词表 32,000\n- 每层 8 个专家 FFN（14,336），router top-2；无共享专家\n- GQA：32 Query 头 / 8 KV 头，head_dim 128；RoPE θ=1e6\n- 上下文 32,768（无滑窗）\n\n参考：《Mixtral of Experts》。","memory_md":"| 精度 | 权重大小 | 参考硬件 | 说明 |\n|---|---|---|---|\n| BF16 | 93.4 GB | 2×80GB | 官方 safetensors |\n| Q8 | ≈ 49.6 GB | 80GB 单卡 | GGUF Q8_0 |\n| Q4 | 26.4 GB | 32GB / 2×24GB | GGUF Q4_K_M（TheBloke） |\n\n**KV Cache**：128 KiB/token。\n\n**参考配置**：\n- RTX 4090 24GB：Q3_K_M 约 20 GB 或 Q4 + 专家 CPU 卸载\n- A100/H100 80GB：Q8 或 FP8，32K 上下文\n- 2×80GB：BF16","training_md":"- 预训练数据与 token 数未公开；多语言（英法德西意）\n- Instruct 版：SFT + DPO\n- 无原生 function-calling 模板（v0.1）\n- 官方推荐 temperature 0.7 以下","ecosystem_md":"- HF：mistralai/Mixtral-8x7B-Instruct-v0.1、-v0.1（基座）\n- 引擎：vLLM（MoE kernel 首个支持对象）、llama.cpp、Ollama、TGI\n- 微调：LoRA 可行，专家层微调需要 MoE 感知框架\n- 中文文档：弱","versions_md":"- 后继：Mixtral 8x22B（2024-04）→ Mistral Small 3 系列 Dense（2025）\n- 同期：Mistral 7B v0.2\n- Mistral 官方后来放弃开源 MoE 路线，转向 Dense 小模型与闭源 Large"},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama","TGI"],"finetune":"LoRA 可行","zh_docs":"弱"}},{"id":"muse-glimmer-30b","name":"Muse Glimmer 30B","name_zh":"Meta Muse Glimmer 30B","aliases":["Muse-Glimmer-30B","meta/muse-glimmer-30b","Muse Glimmer"],"vendor":"Meta","family":"Muse","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/meta-models/Muse-Glimmer-30B","status":"current","released_at":"2026-08-10","updated_at":"2026-08-28","modalities":["text","image","tools"],"reasoning_mode":"optional","architecture":{"type":"dense","total_params":"29.6B","total_params_b":29.6,"layers":52,"hidden_size":6656,"vocab_size":202048,"kv_heads":2,"head_dim":128,"attention":"GQA（32 Q 头 / 2 KV 头，16:1）","notes":"稠密因果 Transformer + 约 1.8B Perception Encoder 视觉编码器。词表 200K BPE + 2048 特殊 token。官方配 DFlash 投机解码 drafter。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":55,"q4":18,"q8":29.6,"fp8":34.4},"kv_per_token_kib":52,"kv_note":"2 KV 头 × 128 × 2 × 52 层 × 2 B = 52 KiB/token，128K 上下文仅约 6.5 GB，极省显存。","ref_hw_24gb":"官方 4-bit K-Quant 约 17–20 GB，24GB 单卡 / 32GB 官方目标配置","ref_hw_80gb":"BF16 55 GB 单卡全精度 + 长上下文","ref_hw_8x80gb":"无必要","estimated":false},"pricing":{"input_per_m":0.3,"output_per_m":1.1,"currency":"USD","source":"第三方托管常见价（OpenRouter）","as_of":"2026-08-28","note":"开源权重，社区托管价（Together / Fireworks / DeepInfra 等）"},"links":{"official":"https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model","hf":"https://huggingface.co/meta-models/Muse-Glimmer-30B"},"variants":[{"kind":"gguf","publisher":"meta-models","repo":"meta-models/Muse-Glimmer-30B-GGUF","url":"https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF","note":"官方 GGUF","sizes":{"q4":18.4}},{"kind":"fp8","publisher":"RedHatAI","repo":"RedHatAI/Muse-Glimmer-30B-FP8-block","url":"https://huggingface.co/RedHatAI/Muse-Glimmer-30B-FP8-block","sizes":{"fp8":34.4}},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/Muse-Glimmer-30B-GGUF","url":"https://huggingface.co/unsloth/Muse-Glimmer-30B-GGUF","sizes":{"q8":29.6,"bf16":55.7}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/Muse-Glimmer-30B-GGUF","url":"https://huggingface.co/bartowski/Muse-Glimmer-30B-GGUF","sizes":{"q4":17.3,"q5":20.1,"q6":23.4,"q8":32.3,"bf16":55.7}},{"kind":"awq","publisher":"cyankiwi","repo":"cyankiwi/Muse-Glimmer-30B-AWQ-INT4","url":"https://huggingface.co/cyankiwi/Muse-Glimmer-30B-AWQ-INT4"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/Muse-Glimmer-30B-4bit","url":"https://huggingface.co/mlx-community/Muse-Glimmer-30B-4bit"}],"copy":{"one_liner":"Meta 回归开源：30B 稠密 agent 模型，24GB 单卡跑。","highlights":["Apache-2.0，Meta 首个 Muse 系开源权重，从 Muse Spark 蒸馏","SWE-bench Verified 76.0%、GPQA Diamond 83.5、AIME 2026 94.7（官方）","4-bit 不到 20 GB，2 KV 头 KV 极省，配 DFlash 投机解码 RTX 5090 提速 3.1×"],"pitfalls":["刚发布不到一月，独立复测与长期稳定性待观察","上下文 128K，低于同级 Qwen / Gemma 的 256K+","中文能力官方称 100+ 语言但无中文专项数据"],"logic_ability":"考试型推理在 30B 级别顶尖（GPQA 83.5、AIME 2026 94.7、HLE 文本 22.0），工程型推理同样突出（SWE-bench Verified 76.0%、Terminal-Bench 2.1 51.7%），AA 智能指数 35 超过 Nemotron 3 Super。reasoning effort 可调，关闭思考时能力回落。","best_for":["24GB 单卡本地编程 agent","本地多模态助手 / 工具调用","MCP / 长任务 agent"],"not_for":["超长上下文（>128K）","中文为主产品的首选"]},"capability_notes":{"coding":"SWE-bench Verified 76.0%、SWE-bench Pro 51.2%、Terminal-Bench 2.1 51.7%（官方）。","reasoning":"GPQA Diamond 83.5、HLE 文本 22.0（官方）。","math":"AIME 2026 94.7（官方；AIME 2025 未披露）。","agent":"MCP-Atlas 75.5、OSWorld-Verified 65.9（官方）。","multimodal":"MMMU-Pro 74（官方）。","chinese":"官方支持 100+ 语言，中文未单独评测。"},"sheet":{"architecture_md":"**类型**：稠密 Transformer，约 29.6B 参数（含约 1.8B 视觉编码器）。\n\n- 语言模型 52 层，隐藏维 6656，词表 202,048（200K BPE + 2,048 特殊 token）\n- GQA：32 Q 头 / 2 KV 头，head_dim 128 —— 16:1 极端比例，KV Cache 非常小\n- 视觉：Meta Perception Encoder，图文输入、文本输出\n- 上下文 131,072+\n- 配套 DFlash 投机解码 drafter：RTX 5090 3.1×、M5 Max 1.8×、M4 Max 1.5×\n\n参考：HF 模型卡。","memory_md":"| 精度 | 权重大小 | 参考显存 |\n|---|---|---|\n| BF16 | 55 GB | 80GB 单卡 |\n| 4-bit K-Quant（官方两档） | ≈ 17–20 GB | 24GB / 32GB 消费卡 |\n\n**KV Cache**：2 KV 头 × 128 × 2 × 52 层 × 2 B = 52 KiB/token；128K 上下文 ≈ 6.5 GB。\n\n**参考配置**：\n- 24GB（4090 / 5090）：官方目标，4-bit + 数万 token 上下文\n- 32GB（5090）：4-bit 高档量化 + 满上下文\n- 80GB：BF16 全精度 + 128K\n- Mac M4 / M5 Max：官方 MLX 支持\n\n**实测**：RTX 5090 基线 74.9 tok/s，DFlash 投机 233.4 tok/s（官方）。","training_md":"- 预训练：从 Muse Spark logit 蒸馏，数据构成与 Spark 相近\n- 中期训练：扩上下文 + agent 数据 + 推理轨迹\n- 后训练：SFT + on-policy 蒸馏 + RL，覆盖通用 / 推理 / 编程 / agent\n- 推理强度可调；函数调用严格遵循 schema；强调错误诊断与重试\n- 官方对标 Gemma 4 31B、Qwen3.6 27B\n- 知识截止未披露","ecosystem_md":"- HF：meta-models/Muse-Glimmer-30B（BF16 + 两档 4-bit）；Meta AI Developer Center\n- 引擎：llama.cpp、MLX、ExecuTorch、Ollama、LM Studio、vLLM、SGLang（发布日原生支持）\n- 微调：Unsloth 支持\n- 托管：Together、Fireworks、OpenRouter、DeepInfra\n- 硬件伙伴：NVIDIA、AMD、Intel、Arm、Dell\n- 中文文档：无（官方英文）","versions_md":"- 2026-08-10：Muse Glimmer 30B（本条，唯一尺寸）\n- 闭源上游：Muse Spark（2026-04）→ 1.1（07-09）→ 1.2（08-05），API-only\n- 前代开源：Llama 4 Scout / Maverick（2025-04，Llama 社区许可）→ 被本条在实用层面取代"},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Ollama","MLX","LM Studio"],"finetune":"Unsloth / LoRA","zh_docs":"无"},"complete":true,"runtime":{"tok_s":102,"latency_s":0.84,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"Meta Muse Glimmer 30B","one_liner":"Meta returns to open source: 30B dense agent model that runs on a 24GB GPU.","highlights":["Apache-2.0; Meta's first open-weight Muse-series model, distilled from Muse Spark","SWE-bench Verified 76.0%, GPQA Diamond 83.5, AIME 2026 94.7 (official)","Under 20 GB at 4-bit; 2 KV heads make KV extremely lean; with DFlash speculative decoding, 3.1× faster on RTX 5090"],"pitfalls":["Released less than a month ago; independent re-testing and long-term stability still to be observed","128K context, below the 256K+ of same-tier Qwen / Gemma","Officially claims 100+ languages but no Chinese-specific data"],"logic_ability":"Top exam-style reasoning at the 30B level (GPQA 83.5, AIME 2026 94.7, HLE text 22.0), with equally standout engineering reasoning (SWE-bench Verified 76.0%, Terminal-Bench 2.1 51.7%); AA Intelligence Index 35 beats Nemotron 3 Super. Reasoning effort adjustable; capability drops when thinking is off.","best_for":["Local coding agents on a 24GB single GPU","Local multimodal assistants / tool calling","MCP / long-task agents"],"not_for":["Ultra-long context (>128K)","First choice for Chinese-first products"],"capability_notes":{"coding":"SWE-bench Verified 76.0%, SWE-bench Pro 51.2%, Terminal-Bench 2.1 51.7% (official).","reasoning":"GPQA Diamond 83.5, HLE text 22.0 (official).","math":"AIME 2026 94.7 (official; AIME 2025 not disclosed).","agent":"MCP-Atlas 75.5, OSWorld-Verified 65.9 (official).","multimodal":"MMMU-Pro 74 (official).","chinese":"Officially supports 100+ languages; Chinese not separately evaluated."}},"ja":{"name_zh":"Meta Muse Glimmer 30B","one_liner":"Meta が OSS 回帰。24GB GPU で動く 30B 稠密エージェント。","highlights":["Apache-2.0、Meta 初の Muse 系オープンウェイト、Muse Spark から蒸留","SWE-bench Verified 76.0%、GPQA Diamond 83.5、AIME 2026 94.7（公式）","4-bit で 20 GB 未満、2 KV ヘッドで KV が極めて省メモリ。DFlash 投機的デコーディングで RTX 5090 は 3.1 倍高速化"],"pitfalls":["公開から 1 か月未満で、独立再検証と長期的な安定性は要観察","コンテキスト 128K、同クラスの Qwen / Gemma の 256K+ を下回る","公式は 100+ 言語対応と謳うが中国語の個別データなし"],"logic_ability":"試験型推論は 30B クラス最上位（GPQA 83.5、AIME 2026 94.7、HLE テキスト 22.0）、エンジニアリング型推論も突出（SWE-bench Verified 76.0%、Terminal-Bench 2.1 51.7%）。AA 知能指数 35 で Nemotron 3 Super を上回る。reasoning effort 調整可、思考オフでは能力が低下。","best_for":["24GB 単一 GPU でのローカルコーディングエージェント","ローカルのマルチモーダルアシスタント / ツール呼び出し","MCP / 長時間タスクのエージェント"],"not_for":["超長コンテキスト（>128K）","中国語中心製品の第一選択"],"capability_notes":{"coding":"SWE-bench Verified 76.0%、SWE-bench Pro 51.2%、Terminal-Bench 2.1 51.7%（公式）。","reasoning":"GPQA Diamond 83.5、HLE テキスト 22.0（公式）。","math":"AIME 2026 94.7（公式。AIME 2025 は未開示）。","agent":"MCP-Atlas 75.5、OSWorld-Verified 65.9（公式）。","multimodal":"MMMU-Pro 74（公式）。","chinese":"公式に 100+ 言語対応、中国語は個別に評価されていない。"}}}},{"id":"muse-spark","name":"Muse Spark 1.2","name_zh":"Meta Muse Spark","aliases":["muse-spark-1.2","muse-spark-1.1","Muse Spark"],"vendor":"Meta","family":"Muse","license":"专有（Meta Model API 服务条款）","license_commercial":true,"openness":"api-only","weights_available":false,"status":"current","released_at":"2026-08-05","updated_at":"2026-08-28","modalities":["text","image","tools","computer-use"],"reasoning_mode":"default-on","architecture":{"type":"unknown","notes":"未披露。原生多模态推理模型，Contemplating 模式可编排多个并行 agent。","undisclosed":true},"context":{"max_tokens":1048576,"display":"1M"},"memory":{"weight_gb":{},"estimated":true},"pricing":{"input_per_m":1.25,"output_per_m":4.25,"currency":"USD","source":"Meta Model API 定价页","as_of":"2026-08-28","note":"Standard 档；缓存输入 $0.15；Contributor 档（允许 Meta 用数据训练）$0.10 / $0.20"},"links":{"official":"https://developer.meta.com/ai/models/muse-spark/","pricing":"https://developer.meta.com/ai/models/muse-spark/"},"copy":{"one_liner":"Meta 超级智能实验室闭源旗舰，1M 上下文，Arena 前十。","highlights":["LMArena 文本 1498（1.2 xHigh，第 4 名）","HLE 58%（Contemplating 模式，官方）","1M 上下文，编程 / 计算机使用优化，$1.25 / $4.25"],"pitfalls":["Meta 首个闭源旗舰，无权重","Contributor 低价档会拿数据训练，慎用","API 2026-07 才开放，生态与 SDK 尚新"],"logic_ability":"考试型推理顶级（HLE 58% 带多 agent 编排）；工程型推理官方图表未给数字。","best_for":["编程 agent / Muse Code","超长上下文研究任务"],"not_for":["自托管","数据敏感场景用 Contributor 档"]},"capability_notes":{"reasoning":"HLE 58%、FrontierScience Research 38%（Contemplating，官方）。"},"ecosystem":{"zh_docs":"无"},"complete":false,"i18n":{"en":{"name_zh":"Meta Muse Spark","one_liner":"Meta Superintelligence Labs' closed flagship, 1M context, Arena top 10.","highlights":["LMArena text 1498 (1.2 xHigh, #4)","HLE 58% (Contemplating mode, official)","1M context, optimized for coding / computer use, $1.25 / $4.25"],"pitfalls":["Meta's first closed flagship; no weights","The cheaper Contributor tier uses your data for training; use with caution","API only opened in 2026-07; ecosystem and SDK still new"],"logic_ability":"Top-tier exam-style reasoning (HLE 58% with multi-agent orchestration); official charts give no engineering-reasoning numbers.","best_for":["Coding agents / Muse Code","Ultra-long-context research tasks"],"not_for":["Self-hosting","Data-sensitive scenarios on the Contributor tier"],"capability_notes":{"reasoning":"HLE 58%, FrontierScience Research 38% (Contemplating, official)."}},"ja":{"name_zh":"Meta Muse Spark","one_liner":"Meta 超知能ラボのクローズド旗艦。1M、Arena トップ 10。","highlights":["LMArena テキスト 1498（1.2 xHigh、4 位）","HLE 58%（Contemplating モード、公式）","1M コンテキスト、コーディング / コンピュータ操作に最適化、$1.25 / $4.25"],"pitfalls":["Meta 初のクローズド旗艦で重みなし","低価格の Contributor 帯はデータが学習に使われるため注意","API 公開は 2026-07 で、エコシステムと SDK はまだ新しい"],"logic_ability":"試験型推論は最上位（マルチエージェント編成込みで HLE 58%）。エンジニアリング型推論は公式図表に数値なし。","best_for":["コーディングエージェント / Muse Code","超長コンテキストのリサーチタスク"],"not_for":["自前ホスティング","データ機密性の高い用途での Contributor 帯利用"],"capability_notes":{"reasoning":"HLE 58%、FrontierScience Research 38%（Contemplating、公式）。"}}}},{"id":"nemotron-3-nano-30b-a3b","name":"Nemotron 3 Nano 30B-A3B","aliases":["NVIDIA-Nemotron-3-Nano-30B-A3B","Nemotron Nano 3"],"vendor":"NVIDIA","family":"Nemotron 3","license":"NVIDIA Open Model License","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16","status":"current","released_at":"2025-12-15","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"optional","architecture":{"type":"hybrid","total_params":"31.6B","active_params":"3.2B","total_params_b":31.6,"active_params_b":3.2,"attention":"Mamba-2 + 稀疏注意力 混合 MoE","notes":"混合 Mamba-Transformer MoE；训练数据、配方开放。KV 未建模（Mamba 状态为主）。","undisclosed":false},"context":{"max_tokens":1048576,"display":"1M"},"memory":{"weight_gb":{"bf16":63,"q4":19,"q8":33.6,"fp8":32.7},"estimated":true,"ref_hw_24gb":"Q4 约 19 GB（估）","ref_hw_80gb":"BF16 舒适"},"links":{"official":"https://developer.nvidia.com/nemotron","hf":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16"},"variants":[{"kind":"fp8","publisher":"nvidia","repo":"nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8","url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8","note":"官方 FP8","sizes":{"fp8":32.7}},{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4","url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4","note":"官方 NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/Nemotron-3-Nano-30B-A3B-GGUF","url":"https://huggingface.co/unsloth/Nemotron-3-Nano-30B-A3B-GGUF","sizes":{"q4":24.6,"q5":26.1,"q6":33.5,"q8":33.6,"bf16":63.2}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/nvidia_Nemotron-3-Nano-30B-A3B-GGUF","url":"https://huggingface.co/bartowski/nvidia_Nemotron-3-Nano-30B-A3B-GGUF","sizes":{"q4":24.7,"q5":26.2,"q6":33.4,"q8":33.6,"bf16":63.2}},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit","url":"https://huggingface.co/mlx-community/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit"}],"copy":{"one_liner":"NVIDIA 开放数据与配方的混合 Mamba MoE，1M 上下文。","highlights":["训练数据 / 配方开放","Mamba 混合，长上下文吞吐高","3.2B 激活"],"pitfalls":["新架构，llama.cpp 支持滞后","新模型，独立评测少","NVIDIA 许可非 Apache（但允许商用）"],"logic_ability":"考试型推理相对体量良好，思考可开关；工程型待验证。","best_for":["长上下文低延迟推理","架构研究"],"not_for":["成熟生态依赖"]},"capability_notes":{"coding":"LiveCodeBench v6 68.3、SWE-bench（OpenHands）38.8、Terminal-Bench hard 8.5（官方模型卡）。","reasoning":"GPQA 73.0（无工具）/ 75.0（工具）、HLE 10.6 / 15.5（工具）、MMLU-Pro 78.3（官方模型卡）。","math":"AIME 2025 89.1（无工具）/ 99.2（工具）、MiniF2F pass@1 50.0（官方模型卡）。","agent":"tau2-bench 均值 49.0、BFCL v4 53.8（官方模型卡）。","instruction":"IFBench 71.5、Arena-Hard-V2 均值 67.7（官方模型卡）。","chinese":"官方语言列表为英德西法意日，不含中文；MMLU-ProX 59.5。"},"ecosystem":{"engines":["vLLM","TensorRT-LLM"],"zh_docs":"无"},"complete":false,"runtime":{"tok_s":227,"latency_s":1.19,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"name_zh":"NVIDIA Nemotron 3 Nano · 30B A3B","pricing":{"input_per_m":0.05,"output_per_m":0.2,"currency":"USD","source":"第三方托管常见价（OpenRouter）","as_of":"2026-08-28","note":"DeepInfra / Crusoe / Nebius / Novita 托管；缓存读取 $0.025"},"i18n":{"en":{"name_zh":"NVIDIA Nemotron 3 Nano · 30B A3B","one_liner":"NVIDIA's hybrid Mamba MoE with open data and recipe, 1M context.","highlights":["Training data / recipe open","Mamba hybrid; high long-context throughput","3.2B active"],"pitfalls":["New architecture; llama.cpp support lags","New model; few independent evaluations","NVIDIA license is not Apache (but allows commercial use)"],"logic_ability":"Good exam-style reasoning for its size, thinking switchable; engineering reasoning yet to be verified.","best_for":["Long-context low-latency inference","Architecture research"],"not_for":["Reliance on a mature ecosystem"],"capability_notes":{"coding":"LiveCodeBench v6 68.3, SWE-bench (OpenHands) 38.8, Terminal-Bench hard 8.5 (official model card).","reasoning":"GPQA 73.0 (no tools) / 75.0 (tools), HLE 10.6 / 15.5 (tools), MMLU-Pro 78.3 (official model card).","math":"AIME 2025 89.1 (no tools) / 99.2 (tools), MiniF2F pass@1 50.0 (official model card).","agent":"tau2-bench average 49.0, BFCL v4 53.8 (official model card).","instruction":"IFBench 71.5, Arena-Hard-V2 average 67.7 (official model card).","chinese":"Official language list is English, German, Spanish, French, Italian, Japanese; no Chinese. MMLU-ProX 59.5."}},"ja":{"name_zh":"NVIDIA Nemotron 3 Nano · 30B A3B","one_liner":"NVIDIA のデータ・レシピ公開 Mamba 混合 MoE。1M。","highlights":["学習データ / レシピを公開","Mamba ハイブリッドで長コンテキストのスループットが高い","3.2B アクティブ"],"pitfalls":["新アーキテクチャで llama.cpp の対応が遅れている","新モデルで独立評価が少ない","NVIDIA ライセンスは Apache ではない（商用は可）"],"logic_ability":"サイズの割に試験型推論は良好、思考はオン/オフ可。エンジニアリング型は要検証。","best_for":["長コンテキストの低レイテンシ推論","アーキテクチャ研究"],"not_for":["成熟したエコシステムへの依存"],"capability_notes":{"coding":"LiveCodeBench v6 68.3、SWE-bench（OpenHands）38.8、Terminal-Bench hard 8.5（公式モデルカード）。","reasoning":"GPQA 73.0（ツールなし）/ 75.0（ツールあり）、HLE 10.6 / 15.5（ツールあり）、MMLU-Pro 78.3（公式モデルカード）。","math":"AIME 2025 89.1（ツールなし）/ 99.2（ツールあり）、MiniF2F pass@1 50.0（公式モデルカード）。","agent":"tau2-bench 平均 49.0、BFCL v4 53.8（公式モデルカード）。","instruction":"IFBench 71.5、Arena-Hard-V2 平均 67.7（公式モデルカード）。","chinese":"公式の対応言語は英・独・西・仏・伊・日で、中国語を含まない。MMLU-ProX 59.5。"}}}},{"id":"nemotron-3-super-120b-a12b","name":"NVIDIA Nemotron 3 Super 120B-A12B","name_zh":"NVIDIA Nemotron 3 Super · 120B A12B","aliases":["Nemotron-3-Super-120B-A12B","nvidia/nemotron-3-super-120b-a12b","Nemotron 3 Super"],"vendor":"NVIDIA","family":"Nemotron 3","license":"NVIDIA Nemotron Open Model License","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16","status":"current","released_at":"2026-03-11","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"optional","architecture":{"type":"hybrid","total_params":"120B","active_params":"12B","total_params_b":120,"active_params_b":12,"experts":512,"active_experts":22,"shared_expert":true,"layers":88,"hidden_size":4096,"vocab_size":131072,"kv_heads":2,"head_dim":128,"attention":"混合：8 层 GQA 注意力（32 Q / 2 KV，head_dim 128）+ 40 层 Mamba-2 + 40 层 LatentMoE","notes":"config hybrid_override_pattern 共 88 层：40 M（Mamba-2，128 头 × 64，state 128）/ 40 E（LatentMoE：512 路由 + 1 共享专家选 22，latent 1024，专家中间维 2688）/ 8 *（注意力），无 RoPE 缩放（theta 1e4，注意力层主要靠 Mamba 提供位置信息）。1 层 MTP（*E）。训练用 NVFP4。config max_position 262,144，官方标称 1M（RULER@1M 91.75）。","undisclosed":false,"kv_layers":8},"context":{"max_tokens":1048576,"display":"1M","max_output":32768},"memory":{"weight_gb":{"bf16":247,"fp8":128,"q8":128,"q4":82.5},"kv_per_token_kib":8,"kv_note":"仅 8 层注意力有 KV：2 KV 头 × 128 × 2（K/V）× 2 B = 1 KiB/层 → 8 KiB/token（BF16）。40 层 Mamba-2 为固定 SSM 状态（128 头 × 64 × 128，FP32，约 4 MiB/层），每序列约 160 MB 与长度无关。1M 上下文 KV ≈ 8 GB。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（UD-Q4_K_M 82.5 GB 已超 80GB；官方 NVFP4 版可跑单 B200 / DGX Spark）","ref_hw_8x80gb":"官方最低 8×H100 BF16；FP8 128 GB 用 2×H100 即可载","estimated":false},"pricing":{"input_per_m":0.085,"output_per_m":0.4,"currency":"USD","source":"第三方托管常见价（OpenRouter）","as_of":"2026-08-28"},"links":{"official":"https://nvidianews.nvidia.com/news/nvidia-debuts-nemotron-3-family-of-open-models","hf":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16"},"variants":[{"kind":"fp8","publisher":"nvidia","repo":"nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8","url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8","note":"官方 FP8","sizes":{"fp8":128.4}},{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4","url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4","note":"官方 NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF","url":"https://huggingface.co/unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF","sizes":{"q8":128.5,"bf16":241.5}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/nvidia_Nemotron-3-Super-120B-A12B-GGUF","url":"https://huggingface.co/bartowski/nvidia_Nemotron-3-Super-120B-A12B-GGUF","sizes":{"q4":87,"q5":96.6,"q6":113.6,"q8":128.5}},{"kind":"awq","publisher":"cyankiwi","repo":"cyankiwi/NVIDIA-Nemotron-3-Super-120B-A12B-AWQ-4bit","url":"https://huggingface.co/cyankiwi/NVIDIA-Nemotron-3-Super-120B-A12B-AWQ-4bit"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/NVIDIA-Nemotron-3-Super-120B-A12B-4bit","url":"https://huggingface.co/mlx-community/NVIDIA-Nemotron-3-Super-120B-A12B-4bit"}],"copy":{"one_liner":"NVIDIA 120B 混合 Mamba-MoE，12B 激活，1M 上下文。","highlights":["AIME 2025 90.2、LiveCodeBench 81.2、GPQA 79.2、RULER@1M 91.75（官方）","仅 8 层注意力 + 2 KV 头，KV 8 KiB/token，1M 上下文 KV 仅约 8 GB","Nemotron 开放许可可商用，25T token 预训练数据 / 配方公开度高"],"pitfalls":["SWE-bench Verified 60.5、Terminal-Bench Core 2.0 31.0、tau2 均值 61.2，工程与 agent 中游","AA 智能指数 26；HLE 无工具 18.3、Arena-Hard-V2 73.9 低于 gpt-oss-120b","Mamba-MoE 混合架构需新版 vLLM（≥ 0.18.1）/ TRT-LLM，llama.cpp 支持滞后；Q4 也超 80GB"],"logic_ability":"考试型推理强：AIME 2025 90.21（无工具）、HMMT Feb25 93.67 / 94.73（工具）、GPQA 79.23 / 82.70（工具）、MMLU-Pro 83.73。HLE 18.26 / 22.82（工具）中等。工程型推理中游：LiveCodeBench 81.19 高，但 SWE-bench Verified 60.47（OpenHands）、Terminal-Bench Core 2.0 31.00 / hard 25.78、tau2 均值 61.15。长上下文是强项（RULER 256K 96.3、1M 91.75、AA-LCR 58.31）。thinking 可开关。","best_for":["超长上下文 RAG / 文档 agent","高吞吐推理服务"],"not_for":["24GB 消费卡","高难度编程 agent"]},"capability_notes":{"coding":"LiveCodeBench 81.19、SWE-bench Verified 60.47（OpenHands）/ 59.20（OpenCode）、Terminal-Bench Core 2.0 31.00、SciCode 42.05（官方模型卡）。","reasoning":"GPQA 79.23（无工具）/ 82.70（工具）、HLE 18.26 / 22.82（工具）、MMLU-Pro 83.73（官方模型卡）。","math":"AIME 2025 90.21（无工具）、HMMT Feb25 93.67 / 94.73（工具）（官方模型卡）。","agent":"tau2-bench 均值 61.15（Airline 56.25 / Retail 62.83 / Telecom 64.36）、BrowseComp（搜索）31.28、BIRD 41.80（官方模型卡）。","instruction":"IFBench 72.56、Arena-Hard-V2 73.88、Multi-Challenge 55.23（官方模型卡）。","chinese":"官方支持中文（7 种语言之一），MMLU-ProX 均值 79.36、WMT24++ 86.67（官方模型卡）。"},"ecosystem":{"engines":["vLLM","SGLang","TensorRT-LLM","Transformers","llama.cpp"],"finetune":"NeMo / Megatron 官方脚本；LoRA ≥ 2×80GB","zh_docs":"无"},"complete":true,"runtime":{"tok_s":142,"latency_s":1.79,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"vendor_zh":"英伟达","sheet":{"architecture_md":"**类型**：Mamba-2 / 注意力 / LatentMoE 混合，120B 总参 / 12B 激活，纯文本。\n\n| 项目 | 数值 |\n|---|---|\n| 层数 | 88 = 40 Mamba-2 + 40 LatentMoE + 8 注意力（pattern MEMEMEM*E…） |\n| 隐藏维 | 4,096 |\n| Mamba-2 | 128 头 × head_dim 64，state 128，conv kernel 4，expand 2 |\n| 专家 | 512 路由 + 1 共享（中间维 5,376），每 token 选 22，专家中间维 2,688，latent 1,024 |\n| 注意力 | 32 Q / 2 KV，head_dim 128，theta 1e4 |\n| 词表 | 131,072 |\n| 上下文 | 官方 1M（config max_position 262,144） |\n| MTP | 1 层（*E） |\n\n`config.json` 中 `model_type: nemotron_h`，`architectures: NemotronHForCausalLM`。\n\n参考：nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 config.json 与模型卡。","memory_md":"| 精度 | 权重大小 | 参考显存 | 说明 |\n|---|---|---|---|\n| BF16 | 247 GB | 8×H100 | 官方 safetensors，50 分片（官方最低配置） |\n| FP8 | 128 GB | 2×H100 | 官方 NVIDIA-Nemotron-3-Super-120B-A12B-FP8，26 分片 |\n| Q8 | 128 GB | 2×H100 | unsloth GGUF Q8_0 |\n| Q4 | 82.5 GB | 2×80GB | unsloth GGUF Q4_K_M（超过单张 80GB） |\n| NVFP4 | 官方另发 | 1×B200 / DGX Spark | 官方 NVFP4 版本 |\n\n**KV Cache**：仅 8 层注意力 × 1 KiB = 8 KiB/token；40 层 Mamba-2 状态固定约 160 MB/序列。1M 上下文 KV ≈ 8 GB。\n\n**参考配置**：\n- 24GB / 80GB 单卡：不可行\n- 2×H100：FP8 或 Q8 可载，长上下文余量大\n- 8×H100：官方 BF16 推荐\n\n警示：vLLM 需 ≥ 0.18.1，Transformers ≥ 5.3；Mamba 状态用 FP32 缓存。","training_md":"- 预训练：25T+ token（模型卡），预训练截止 2025-06，后训练截止 2026-02\n- 训练精度 NVFP4；带 MTP 头（另发 MTPv2 版本）\n- 数据与配方：Nemotron 系列公开预训练 / 后训练数据集与训练脚本（NeMo）\n- 思考模式可开关（reasoning on/off）\n- 支持 7 种语言：英、法、德、意、日、西、中","ecosystem_md":"- HF：nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16（247 GB）、-FP8（128 GB）、NVFP4；GGUF：unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF\n- 引擎：vLLM（≥ 0.18.1）、SGLang、TensorRT-LLM、Transformers（≥ 5.3）；llama.cpp 经 GGUF\n- 微调：NeMo / Megatron 官方脚本；社区 LoRA 需 ≥ 2×80GB\n- 中文文档：无","versions_md":"- 同代：Nemotron 3 Nano 30B-A3B（52 层同构小版）、Nemotron 3 Ultra（更大）\n- 后续：Nemotron 3.5 Lightning（已上线 OpenRouter）\n- 上代：Nemotron 2 Nano 9B / 12B（Mamba 混合 Dense）"},"i18n":{"en":{"name_zh":"NVIDIA Nemotron 3 Super · 120B A12B","one_liner":"NVIDIA's 120B hybrid Mamba-MoE, 12B active, 1M context.","highlights":["AIME 2025 90.2, LiveCodeBench 81.2, GPQA 79.2, RULER@1M 91.75 (official)","Only 8 attention layers + 2 KV heads; KV 8 KiB/token, about 8 GB of KV at 1M context","Nemotron open license allows commercial use; 25T-token pretraining data / recipe highly open"],"pitfalls":["SWE-bench Verified 60.5, Terminal-Bench Core 2.0 31.0, tau2 average 61.2; mid-pack on engineering and agents","AA Intelligence Index 26; HLE no-tools 18.3 and Arena-Hard-V2 73.9 below gpt-oss-120b","Hybrid Mamba-MoE needs recent vLLM (≥ 0.18.1) / TRT-LLM; llama.cpp support lags; even Q4 exceeds 80GB"],"logic_ability":"Strong exam-style reasoning: AIME 2025 90.21 (no tools), HMMT Feb25 93.67 / 94.73 (tools), GPQA 79.23 / 82.70 (tools), MMLU-Pro 83.73. HLE 18.26 / 22.82 (tools) is moderate. Engineering reasoning mid-pack: LiveCodeBench 81.19 is high, but SWE-bench Verified 60.47 (OpenHands), Terminal-Bench Core 2.0 31.00 / hard 25.78, tau2 average 61.15. Long context is a strength (RULER 256K 96.3, 1M 91.75, AA-LCR 58.31). Thinking switchable.","best_for":["Ultra-long-context RAG / document agents","High-throughput inference services"],"not_for":["24GB consumer GPUs","Hard coding agents"],"capability_notes":{"coding":"LiveCodeBench 81.19, SWE-bench Verified 60.47 (OpenHands) / 59.20 (OpenCode), Terminal-Bench Core 2.0 31.00, SciCode 42.05 (official model card).","reasoning":"GPQA 79.23 (no tools) / 82.70 (tools), HLE 18.26 / 22.82 (tools), MMLU-Pro 83.73 (official model card).","math":"AIME 2025 90.21 (no tools), HMMT Feb25 93.67 / 94.73 (tools) (official model card).","agent":"tau2-bench average 61.15 (Airline 56.25 / Retail 62.83 / Telecom 64.36), BrowseComp (search) 31.28, BIRD 41.80 (official model card).","instruction":"IFBench 72.56, Arena-Hard-V2 73.88, Multi-Challenge 55.23 (official model card).","chinese":"Chinese officially supported (one of 7 languages); MMLU-ProX average 79.36, WMT24++ 86.67 (official model card)."}},"ja":{"name_zh":"NVIDIA Nemotron 3 Super · 120B A12B","one_liner":"NVIDIA の 120B Mamba-MoE 混合。12B 活性、1M。","highlights":["AIME 2025 90.2、LiveCodeBench 81.2、GPQA 79.2、RULER@1M 91.75（公式）","アテンション層はわずか 8 層 + 2 KV ヘッド、KV 8 KiB/token で 1M コンテキストの KV は約 8 GB","Nemotron オープンライセンスで商用可、25T トークンの事前学習データ / レシピの公開度が高い"],"pitfalls":["SWE-bench Verified 60.5、Terminal-Bench Core 2.0 31.0、tau2 平均 61.2 でエンジニアリングとエージェントは中位","AA 知能指数 26。HLE ツールなし 18.3、Arena-Hard-V2 73.9 は gpt-oss-120b を下回る","Mamba-MoE ハイブリッドは新版 vLLM（≥ 0.18.1）/ TRT-LLM が必要で llama.cpp 対応は遅れ。Q4 でも 80GB 超"],"logic_ability":"試験型推論は強い：AIME 2025 90.21（ツールなし）、HMMT Feb25 93.67 / 94.73（ツールあり）、GPQA 79.23 / 82.70（ツールあり）、MMLU-Pro 83.73。HLE 18.26 / 22.82（ツールあり）は中程度。エンジニアリング型推論は中位：LiveCodeBench 81.19 は高いが、SWE-bench Verified 60.47（OpenHands）、Terminal-Bench Core 2.0 31.00 / hard 25.78、tau2 平均 61.15。長コンテキストが強み（RULER 256K 96.3、1M 91.75、AA-LCR 58.31）。thinking はオン/オフ可。","best_for":["超長コンテキストの RAG / ドキュメントエージェント","高スループットの推論サービス"],"not_for":["24GB コンシューマー GPU","高難度のコーディングエージェント"],"capability_notes":{"coding":"LiveCodeBench 81.19、SWE-bench Verified 60.47（OpenHands）/ 59.20（OpenCode）、Terminal-Bench Core 2.0 31.00、SciCode 42.05（公式モデルカード）。","reasoning":"GPQA 79.23（ツールなし）/ 82.70（ツールあり）、HLE 18.26 / 22.82（ツールあり）、MMLU-Pro 83.73（公式モデルカード）。","math":"AIME 2025 90.21（ツールなし）、HMMT Feb25 93.67 / 94.73（ツールあり）（公式モデルカード）。","agent":"tau2-bench 平均 61.15（Airline 56.25 / Retail 62.83 / Telecom 64.36）、BrowseComp（検索）31.28、BIRD 41.80（公式モデルカード）。","instruction":"IFBench 72.56、Arena-Hard-V2 73.88、Multi-Challenge 55.23（公式モデルカード）。","chinese":"公式に中国語対応（7 言語の 1 つ）、MMLU-ProX 平均 79.36、WMT24++ 86.67（公式モデルカード）。"}}}},{"id":"nemotron-3-ultra-550b-a55b","name":"NVIDIA Nemotron 3 Ultra 550B-A55B","name_zh":"NVIDIA Nemotron 3 Ultra","aliases":["Nemotron-3-Ultra-550B-A55B","nvidia/nemotron-3-ultra-550b-a55b","Nemotron 3 Ultra"],"vendor":"NVIDIA","family":"Nemotron 3","license":"OpenMDW License Agreement v1.1","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16","status":"current","released_at":"2026-06-04","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"optional","architecture":{"type":"hybrid","total_params":"550B","active_params":"55B","total_params_b":550,"active_params_b":55,"experts":512,"active_experts":22,"shared_expert":true,"layers":128,"hidden_size":8192,"vocab_size":131072,"kv_heads":2,"head_dim":128,"attention":"混合：10 层注意力（64 Q / 2 KV）+ 64 层 Mamba-2 + 54 层 LatentMoE","notes":"Mamba-2 / 注意力 / LatentMoE 混合 + 多 token 预测。另有 GenRM 奖励模型版。","undisclosed":false},"context":{"max_tokens":1048576,"display":"1M"},"memory":{"weight_gb":{"bf16":1100,"fp8":550,"q4":319,"q8":584.3},"estimated":true,"ref_hw_80gb":"不可行","ref_hw_8x80gb":"官方最低 8×B200 / 8×H200 / 16×H100（约 1.5 TB HBM）"},"pricing":{"input_per_m":0.5,"output_per_m":2.2,"currency":"USD","source":"第三方托管常见价（OpenRouter）","as_of":"2026-08-28"},"links":{"official":"https://nvidianews.nvidia.com/news/nvidia-debuts-nemotron-3-family-of-open-models","hf":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16"},"variants":[{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4","url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4","note":"官方 NVFP4"},{"kind":"fp8","publisher":"RedHatAI","repo":"RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-block","url":"https://huggingface.co/RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-FP8-block","sizes":{"fp8":563.3}},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/NVIDIA-Nemotron-3-Ultra-550B-A55B-GGUF","url":"https://huggingface.co/unsloth/NVIDIA-Nemotron-3-Ultra-550B-A55B-GGUF","sizes":{"q8":584.3,"bf16":1099}},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/Nemotron-3-Ultra-550B-A55B-4bit","url":"https://huggingface.co/mlx-community/Nemotron-3-Ultra-550B-A55B-4bit"}],"copy":{"one_liner":"NVIDIA 550B 开源旗舰，GPQA 87、LCB 89，1M 上下文。","highlights":["GPQA 87.0、LiveCodeBench v6 89.0、SWE-bench Verified 70.7%（官方）","OpenMDW 许可，权重 + 数据 + 配方开放","1M 上下文，混合 Mamba 架构长序列成本低"],"pitfalls":["需 8×B200 或 16×H100 级别，个人 / 小团队不可部署","AA 智能指数 38，低于 Hy3（42）、MiMo-V2.5-Pro（43）","llama.cpp / Ollama 支持依赖社区 GGUF"],"logic_ability":"考试型推理强（GPQA 87.0、MMLU-Pro 86.8），工程型推理不错（SWE-bench 70.7、Terminal-Bench 2.1 56.4）。","best_for":["企业级长时程 agent","超长上下文代码库分析"],"not_for":["单机部署","成本敏感场景"]},"capability_notes":{"coding":"LiveCodeBench v6 89.0、SWE-bench Verified 70.7、Terminal-Bench 2.1 56.4（官方）。","reasoning":"GPQA 87.0、MMLU-Pro 86.8（官方）。","agent":"TauBench 平均 70.9（官方）。","chinese":"官方支持中文（10 种语言之一）。"},"ecosystem":{"engines":["vLLM","SGLang","TensorRT-LLM"],"zh_docs":"无"},"complete":false,"runtime":{"tok_s":150,"latency_s":2.09,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"NVIDIA Nemotron 3 Ultra","one_liner":"NVIDIA's 550B open flagship: GPQA 87, LCB 89, 1M context.","highlights":["GPQA 87.0, LiveCodeBench v6 89.0, SWE-bench Verified 70.7% (official)","OpenMDW license; weights + data + recipe open","1M context; hybrid Mamba architecture keeps long-sequence cost low"],"pitfalls":["Needs 8×B200 or 16×H100 class; not deployable by individuals / small teams","AA Intelligence Index 38, below Hy3 (42) and MiMo-V2.5-Pro (43)","llama.cpp / Ollama support depends on community GGUFs"],"logic_ability":"Strong exam-style reasoning (GPQA 87.0, MMLU-Pro 86.8), decent engineering reasoning (SWE-bench 70.7, Terminal-Bench 2.1 56.4).","best_for":["Enterprise long-horizon agents","Ultra-long-context codebase analysis"],"not_for":["Single-machine deployment","Cost-sensitive scenarios"],"capability_notes":{"coding":"LiveCodeBench v6 89.0, SWE-bench Verified 70.7, Terminal-Bench 2.1 56.4 (official).","reasoning":"GPQA 87.0, MMLU-Pro 86.8 (official).","agent":"TauBench average 70.9 (official).","chinese":"Chinese officially supported (one of 10 languages)."}},"ja":{"name_zh":"NVIDIA Nemotron 3 Ultra","one_liner":"NVIDIA の 550B オープン旗艦。GPQA 87、LCB 89、1M。","highlights":["GPQA 87.0、LiveCodeBench v6 89.0、SWE-bench Verified 70.7%（公式）","OpenMDW ライセンス、重み + データ + レシピを公開","1M コンテキスト、ハイブリッド Mamba アーキテクチャで長系列コストが低い"],"pitfalls":["8×B200 または 16×H100 級が必要で、個人 / 小規模チームではデプロイ不可","AA 知能指数 38 で Hy3（42）、MiMo-V2.5-Pro（43）を下回る","llama.cpp / Ollama 対応はコミュニティ GGUF 頼み"],"logic_ability":"試験型推論は強い（GPQA 87.0、MMLU-Pro 86.8）、エンジニアリング型推論も良好（SWE-bench 70.7、Terminal-Bench 2.1 56.4）。","best_for":["エンタープライズの長期エージェント","超長コンテキストのコードベース分析"],"not_for":["単一マシンでのデプロイ","コスト重視の用途"],"capability_notes":{"coding":"LiveCodeBench v6 89.0、SWE-bench Verified 70.7、Terminal-Bench 2.1 56.4（公式）。","reasoning":"GPQA 87.0、MMLU-Pro 86.8（公式）。","agent":"TauBench 平均 70.9（公式）。","chinese":"公式に中国語対応（10 言語の 1 つ）。"}}}},{"aliases":["Nemotron-4-340B-Instruct"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"math":"GSM8K 92.3%（官方）。","knowledge":"MMLU 78.7%（官方）。","chinese":"弱。"},"complete":false,"id":"nemotron-4-340b","name":"Nemotron-4 340B","name_zh":"Nemotron-4 · 340B","vendor":"NVIDIA","family":"Nemotron-4","license":"NVIDIA Open Model License","license_commercial":true,"weights_url":"https://huggingface.co/nvidia/Nemotron-4-340B-Instruct","superseded_by":"nemotron-3-ultra-550b-a55b","released_at":"2024-06-14","architecture":{"type":"dense","total_params":"340B","total_params_b":340,"layers":96,"hidden_size":18432,"vocab_size":256000,"kv_heads":8,"head_dim":192,"attention":"GQA（96 Q 头 / 8 KV 头）","notes":"9T token；NeMo 格式发布（非 HF 原生）；4K 上下文；同发 Reward 模型。","undisclosed":false},"context":{"max_tokens":4096,"display":"4K"},"memory":{"weight_gb":{"bf16":680,"q8":360.4,"q4":197.2},"estimated":true,"kv_per_token_kib":576,"kv_note":"8×192×2×96×2 B = 576 KiB/token。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行","ref_hw_8x80gb":"BF16 680 GB 需 16×80GB；官方推荐 8×H200 或 FP8 8×H100"},"links":{"official":"https://blogs.nvidia.com/blog/nemotron-4-synthetic-data-generation-llm-training/","paper":"https://arxiv.org/abs/2406.11704","hf":"https://huggingface.co/nvidia/Nemotron-4-340B-Instruct"},"copy":{"one_liner":"NVIDIA 为合成数据而生的 340B Dense，许可允许用输出训练。","highlights":["Open Model License 明确允许用其输出训练商业模型","配套 Reward 模型，RewardBench 当年第一","98% 对齐数据为合成，展示合成数据流水线"],"pitfalls":["上下文仅 4K","NeMo 格式，HF 生态支持弱","340B Dense 部署成本极高，实用价值被 Llama 3.1 405B 覆盖"],"logic_ability":"非推理模型一线：MMLU 78.7%、GSM8K 92.3%（官方 Instruct）。","best_for":["合成数据生成","奖励模型研究"],"not_for":["长文本","日常部署"]},"ecosystem":{"engines":["NeMo","TensorRT-LLM"],"finetune":"NeMo 框架","zh_docs":"无"}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。强化学习训练的推理模型，API 提供 reasoning_effort 参数。"},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"o1","name":"OpenAI o1","aliases":["o1-2024-12-17","o1-preview","o1-preview-2024-09-12"],"vendor":"OpenAI","family":"o 系列","superseded_by":"o3","released_at":"2024-12-05","modalities":["text","image","tools"],"reasoning_mode":"default-on","context":{"max_tokens":200000,"display":"200K","max_output":100000},"pricing":{"input_per_m":15,"output_per_m":60,"currency":"USD","source":"OpenAI 定价页","as_of":"2024-12-17","note":"缓存输入 $7.5；o1-preview 同价"},"links":{"official":"https://openai.com/o1/","paper":"https://openai.com/index/openai-o1-system-card/","pricing":"https://openai.com/api/pricing"},"copy":{"one_liner":"首个正式版推理模型，AIME 2024 83.3%。","highlights":["GPQA 78.0%、AIME 2024 83.3%，开创推理模型范式","200K 上下文、100K 最大输出","正式版比 o1-preview 快、错误率降低 34%（官方）"],"pitfalls":["$15 / $60，推理 token 不可见但计费","不支持流式 / system 提示的早期限制","SWE-bench Verified 48.9%，工程能力不及后来 o3"],"logic_ability":"考试型推理跃迁：数学、物理、化学博士级问题上首次超过人类专家均值。工程型推理仍一般。","best_for":["历史对照"],"not_for":["新项目"]}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"o3-mini","name":"OpenAI o3-mini","aliases":["o3-mini-2025-01-31","o3-mini-high"],"vendor":"OpenAI","family":"o 系列","superseded_by":"o4-mini","released_at":"2025-01-31","modalities":["text","tools"],"reasoning_mode":"default-on","context":{"max_tokens":200000,"display":"200K","max_output":100000},"pricing":{"input_per_m":1.1,"output_per_m":4.4,"currency":"USD","source":"OpenAI 定价页","as_of":"2025-01-31","note":"缓存输入 $0.55"},"links":{"official":"https://openai.com/index/openai-o3-mini/","pricing":"https://openai.com/api/pricing"},"copy":{"one_liner":"廉价版推理模型，high 档 AIME 87.3%。","highlights":["$1.1 / $4.4，推理模型价格降至 o1 的 1/14","AIME 2024（high）87.3%、GPQA 79.7%","支持函数调用、结构化输出、三档 reasoning_effort"],"pitfalls":["不支持图像输入","SWE-bench Verified 49.3%，代理编程一般","发布 2.5 个月即被 o4-mini 替代"],"logic_ability":"考试型推理接近 o1，工程型推理中等；medium 档是速度与准确率的折中。","best_for":["历史对照"],"not_for":["新项目"]}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true,"notes":"未披露。强化学习规模化训练的推理模型，可在推理链中调用工具（搜索、Python、图像操作）。"},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{"coding":"SWE-bench Verified 69.1%、Aider Polyglot 81.3%、Codeforces 2706（官方）。","reasoning":"GPQA 83.3%、HLE 20.32%（无工具）/ 24.90%（Python + 浏览）（官方）。","math":"AIME 2024 91.6%、AIME 2025 88.9%（无工具，官方）。","agent":"Tau-bench airline 52.0%、retail 70.4%（官方）。","multimodal":"MMMU 82.9%、MathVista 86.8%（官方）。","chinese":"中文良好，缺独立榜单。"},"complete":true,"id":"o3","name":"OpenAI o3","aliases":["o3-2025-04-16","o3-pro"],"vendor":"OpenAI","family":"o 系列","superseded_by":"gpt-5","released_at":"2025-04-16","modalities":["text","image","tools"],"reasoning_mode":"default-on","context":{"max_tokens":200000,"display":"200K","max_output":100000},"pricing":{"input_per_m":2,"output_per_m":8,"currency":"USD","source":"OpenAI 定价页（2025-06-10 降价）","as_of":"2025-06-10","note":"发布价 $10 / $40，2025-06-10 降 80%；缓存输入 $0.5"},"links":{"official":"https://openai.com/index/introducing-o3-and-o4-mini/","paper":"https://openai.com/index/o3-o4-mini-system-card/","pricing":"https://openai.com/api/pricing"},"copy":{"one_liner":"首个能在思考中用工具的旗舰推理模型。","highlights":["AIME 2025 88.9%、GPQA 83.3%、HLE 20.3%（无工具）","推理链内调用搜索 / Python / 图像缩放，\"用图思考\"","SWE-bench Verified 69.1%，Codeforces 2706"],"pitfalls":["系统卡显示幻觉率高于 o1（PersonQA 33%）","推理 token 不可见，长任务成本难预估","2025-08 被 GPT-5 替代"],"logic_ability":"考试型与工程型推理在 2025 年上半年均为第一梯队。工具增强推理是关键：HLE 带 Python + 浏览可到 24.9%。问题：过度自信，会编造工具执行过程。","best_for":["数学 / 科学研究","多步工具调用的分析任务","历史对照"],"not_for":["低延迟场景","新项目"]},"sheet":{"architecture_md":"**未披露**。OpenAI 未公开 o3 的参数量或结构。\n\n已知接口层信息：\n- `reasoning_effort`：low / medium / high\n- 200K 上下文，100K 最大输出\n- 推理链中可自主调用工具：网页搜索、Python、文件分析、图像处理\n- 支持视觉输入，可在推理中裁剪 / 旋转图像\n- o3-pro（2025-06）为更多算力的版本","memory_md":"无自建选项。\n\n成本参考（2025-06-10 后）：输入 $2 / 输出 $8；缓存输入 $0.5。发布时为 $10 / $40。\n\n**警示**：推理 token 计入输出计费，high effort 下单次请求可达数万 token。","training_md":"- 训练细节未披露，官方称 RL 算力较 o1 提升约 10 倍\n- 默认推理开启，无法关闭\n- 支持函数调用、结构化输出、图像输入\n- 系统卡指出幻觉率高于 o1","ecosystem_md":"- 官方 API、Azure OpenAI\n- 不支持微调\n- 中文文档：弱（官方英文）","versions_md":"- 上代：o1（2024-12）\n- 同代：o4-mini、o3-pro（2025-06）\n- 2025-06-10 降价 80%\n- 继任：GPT-5（2025-08）"},"ecosystem":{"engines":[],"finetune":"不支持","zh_docs":"弱"}},{"license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","updated_at":"2026-08-28","architecture":{"type":"unknown","undisclosed":true},"memory":{"weight_gb":{},"estimated":false},"capability_notes":{},"complete":false,"id":"o4-mini","name":"OpenAI o4-mini","aliases":["o4-mini-2025-04-16","o4-mini-high"],"vendor":"OpenAI","family":"o 系列","superseded_by":"gpt-5-mini","released_at":"2025-04-16","modalities":["text","image","tools"],"reasoning_mode":"default-on","context":{"max_tokens":200000,"display":"200K","max_output":100000},"pricing":{"input_per_m":1.1,"output_per_m":4.4,"currency":"USD","source":"OpenAI 定价页","as_of":"2025-04-16","note":"缓存输入 $0.275"},"links":{"official":"https://openai.com/index/introducing-o3-and-o4-mini/","pricing":"https://openai.com/api/pricing"},"copy":{"one_liner":"小推理模型，AIME 2025 92.7% 超过 o3。","highlights":["AIME 2025 92.7%（无工具），数学超越 o3","$1.1 / $4.4，与 o3-mini 同价但带视觉与工具","SWE-bench Verified 68.1%，接近 o3"],"pitfalls":["GPQA 81.4%、HLE 14.3%，知识面弱于 o3","幻觉率偏高（PersonQA 48%）","4 个月后被 GPT-5 mini 替代"],"logic_ability":"考试型推理极强（数学、竞赛编程），知识型推理和长任务稳定性弱于 o3。","best_for":["历史对照"],"not_for":["新项目"]}},{"aliases":["Phi-3-medium-128k-instruct","Phi-3-medium-4k-instruct"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"math":"GSM8K 91.0%（官方）。","knowledge":"MMLU 78.0%（官方）。","chinese":"弱。"},"complete":false,"id":"phi-3-medium","name":"Phi-3 Medium","name_zh":"Phi-3 Medium · 14B","vendor":"Microsoft","family":"Phi-3","license":"MIT","license_commercial":true,"weights_url":"https://huggingface.co/microsoft/Phi-3-medium-128k-instruct","superseded_by":"phi-4","released_at":"2024-05-21","architecture":{"type":"dense","total_params":"14B","total_params_b":14,"layers":40,"hidden_size":5120,"vocab_size":32064,"kv_heads":10,"head_dim":128,"attention":"GQA（40 Q 头 / 10 KV 头）","notes":"4.8T「教科书级」合成 + 过滤数据；4K 与 128K（LongRoPE）两版。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":28,"q8":14.8,"q4":8.6},"estimated":false,"kv_per_token_kib":200,"kv_note":"200 KiB/token。","ref_hw_24gb":"BF16 28 GB 超 24GB；Q8 15 GB 可跑","ref_hw_80gb":"BF16 单卡","ref_hw_8x80gb":"过剩"},"links":{"official":"https://azure.microsoft.com/en-us/blog/new-models-added-to-the-phi-3-family-available-now/","paper":"https://arxiv.org/abs/2404.14219","hf":"https://huggingface.co/microsoft/Phi-3-medium-128k-instruct"},"copy":{"one_liner":"合成数据路线的 14B，MMLU 78% 逼近 70B 级。","highlights":["MIT 许可","MMLU 78.0% 当年 14B 最高（官方）","128K 上下文（LongRoPE）"],"pitfalls":["知识面窄、多语言与中文弱","对话自然度与指令遵循被批「刷榜」","被 Phi-4 取代"],"logic_ability":"考试型强：MMLU 78.0%、GSM8K 91.0%（官方）；开放式对话与常识弱。","best_for":["英文考试 / 教育类任务"],"not_for":["中文","开放对话"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama","ONNX Runtime"],"finetune":"LoRA 友好","zh_docs":"弱"}},{"aliases":["phi-4 14B"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"coding":"HumanEval 82.6%（官方）。","reasoning":"GPQA 56.1%（官方）。","math":"MATH 80.4%（官方）。","knowledge":"MMLU 84.8%（官方）。","chinese":"一般。"},"complete":false,"id":"phi-4","name":"Phi-4","name_zh":"Phi-4 · 14B","vendor":"Microsoft","family":"Phi-4","license":"MIT","license_commercial":true,"weights_url":"https://huggingface.co/microsoft/phi-4","released_at":"2024-12-12","architecture":{"type":"dense","total_params":"14.7B","total_params_b":14.7,"layers":40,"hidden_size":5120,"vocab_size":100352,"kv_heads":10,"head_dim":128,"attention":"GQA（40 Q 头 / 10 KV 头）","notes":"9.8T token，合成数据为主；tiktoken 词表；上下文 16K。","undisclosed":false},"context":{"max_tokens":16384,"display":"16K"},"memory":{"weight_gb":{"bf16":29.4,"q8":15.6,"q4":9.1},"estimated":false,"kv_per_token_kib":200,"kv_note":"200 KiB/token。","ref_hw_24gb":"Q8 15.6 GB / Q4 9.1 GB 单卡轻松；BF16 29.4 GB 超 24GB","ref_hw_80gb":"BF16 单卡","ref_hw_8x80gb":"过剩"},"links":{"official":"https://techcommunity.microsoft.com/blog/aiplatformblog/introducing-phi-4-microsoft%E2%80%99s-newest-small-language-model-specializing-in-comple/4357090","paper":"https://arxiv.org/abs/2412.08905","hf":"https://huggingface.co/microsoft/phi-4"},"copy":{"one_liner":"14B 拿到 GPQA 56%，合成数据路线的代表作。","highlights":["GPQA 56.1%、MATH 80.4%（官方），超越 GPT-4o-mini","MIT 许可，24GB 卡 Q8 可跑","后续衍生 Phi-4-reasoning / mini / multimodal"],"pitfalls":["上下文仅 16K","事实知识与多语言弱，中文一般","指令遵循 IFEval 63% 偏弱（官方）"],"logic_ability":"考试型推理在 14B 中极强：GPQA 56.1%、MATH 80.4%、HumanEval 82.6%（官方）。但没有思考链，复杂多步与开放式任务不稳。","best_for":["STEM 问答 / 教育","本地轻量推理"],"not_for":["长文本","中文对话"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama","ONNX Runtime"],"finetune":"LoRA 友好","zh_docs":"弱"}},{"aliases":["Qwen1.5-72B-Chat"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"knowledge":"MMLU 77.5%（官方）。","math":"GSM8K 79.5%（官方）。","chinese":"C-Eval 84.1%（官方），当年中文开源最强。"},"complete":false,"id":"qwen1-5-72b","name":"Qwen1.5-72B","name_zh":"通义千问 1.5 · 72B","vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen1.5","license":"Tongyi Qianwen License","license_commercial":"restricted","weights_url":"https://huggingface.co/Qwen/Qwen1.5-72B-Chat","superseded_by":"qwen2-5-72b","released_at":"2024-02-05","architecture":{"type":"dense","total_params":"72.3B","total_params_b":72.3,"layers":80,"hidden_size":8192,"vocab_size":152064,"kv_heads":64,"head_dim":128,"attention":"MHA（64 头，无 GQA）","notes":"仍用 MHA，KV 缓存极大；上下文 32K；首个并入 HF transformers 的 Qwen 版本。","undisclosed":false},"context":{"max_tokens":32768,"display":"32K"},"memory":{"weight_gb":{"bf16":144.6,"q8":76.6,"q4":44.2},"estimated":false,"kv_per_token_kib":2560,"kv_note":"MHA：64×128×2×80×2 B = 2.5 MiB/token，32K 上下文要 80 GB KV！","ref_hw_24gb":"不可行","ref_hw_80gb":"Q4 权重可放，但 KV 巨大只能短上下文","ref_hw_8x80gb":"BF16 + 长上下文"},"links":{"official":"https://qwenlm.github.io/blog/qwen1.5/","github":"https://github.com/QwenLM/Qwen1.5","hf":"https://huggingface.co/Qwen/Qwen1.5-72B-Chat"},"copy":{"one_liner":"Qwen 系列进入 HF 主线的第一代 72B，中文开源当年最强。","highlights":["2024 年初中文开源综合第一，Arena 开源前列","全尺寸（0.5B–72B）统一模板，生态起步","32K 上下文，词表 152K 中文效率高"],"pitfalls":["MHA 无 GQA，KV 每 token 2.5 MiB，长文本显存爆炸","自有许可，月活 1 亿以上需授权","Qwen2 起改用 GQA 且更强，无理由再用 1.5"],"logic_ability":"2024 年初开源一线：MMLU 77.5%、GSM8K 79.5%（官方基座）。无 CoT 训练，复杂推理弱。","best_for":["历史对照"],"not_for":["长上下文","新项目"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama"],"finetune":"LoRA 成熟","zh_docs":"有"}},{"aliases":["Qwen2.5-72B-Instruct"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"coding":"LiveCodeBench 55.5%、HumanEval 86.6%（官方）。","reasoning":"GPQA 49.0%（官方）。","math":"MATH 83.1%（官方）。","knowledge":"MMLU-Pro 71.1%（官方）。","agent":"工具调用 Hermes 格式，Qwen-Agent 支持。","chinese":"中文顶级，长期作为中文开源基准。"},"complete":true,"id":"qwen2-5-72b","name":"Qwen2.5-72B","name_zh":"通义千问 2.5 · 72B","vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen2.5","license":"Qwen License","license_commercial":"restricted","weights_url":"https://huggingface.co/Qwen/Qwen2.5-72B-Instruct","superseded_by":"qwen3-235b-a22b-thinking-2507","released_at":"2024-09-19","architecture":{"type":"dense","total_params":"72.7B","total_params_b":72.7,"layers":80,"hidden_size":8192,"vocab_size":152064,"kv_heads":8,"head_dim":128,"attention":"GQA（64 Q 头 / 8 KV 头）","notes":"18T token 预训练；128K 上下文（YaRN），最大输出 8K；结构与 Qwen2-72B 相同。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":8192},"memory":{"weight_gb":{"bf16":145.4,"q8":77.1,"q4":47.4},"estimated":false,"kv_per_token_kib":320,"kv_note":"8×128×2×80×2 B = 320 KiB/token；32K 上下文约 10 GB。","ref_hw_24gb":"Q4_K_M 47.4 GB 需 2×24GB；单卡只能 IQ2 级","ref_hw_80gb":"Q4 / FP8 单卡 80GB 可跑 32K；BF16 2×80GB","ref_hw_8x80gb":"BF16 高并发服务"},"pricing":{"input_per_m":0.13,"output_per_m":0.4,"currency":"USD","source":"第三方托管常见价（DeepInfra）","as_of":"2026-08-28","note":"开源权重，社区托管价"},"links":{"official":"https://qwenlm.github.io/blog/qwen2.5/","github":"https://github.com/QwenLM/Qwen2.5","paper":"https://arxiv.org/abs/2412.15115","hf":"https://huggingface.co/Qwen/Qwen2.5-72B-Instruct"},"copy":{"one_liner":"2024 下半年开源 Dense 王者，无数微调与蒸馏的底座。","highlights":["发布时综合超越 Llama 3.1 405B（官方多项基准）","18T token 预训练，数学 / 代码专项数据注入","DeepSeek-R1 蒸馏、大量学术工作的默认底座"],"pitfalls":["72B 用 Qwen License（1 亿月活限制），32B 及以下才是 Apache","无思考模式，2025 年推理榜落后于 Qwen3 / R1","最大输出仅 8K，长生成需改配置"],"logic_ability":"非推理模型顶级：MATH 83.1%、GPQA 49.0%、LiveCodeBench 55.5%、MMLU-Pro 71.1%（官方）。逻辑清晰、指令遵循极稳，多轮与结构化输出可靠；但没有内在的长思考能力，竞赛题会自信地给错答案。","best_for":["中英双语通用私有化","微调 / 蒸馏底座","结构化输出与工具调用"],"not_for":["竞赛数学 / 深度代码 agent","单张 24GB 卡"]},"sheet":{"architecture_md":"**类型**：Dense Transformer，72.7B 参数（70.0B 非 embedding）。\n\n- 80 层，隐藏维 8192，FFN 29,568，词表 152,064\n- GQA：64 Query 头 / 8 KV 头，head_dim 128；QKV bias\n- RoPE θ=1e6；原生 32K，YaRN factor 4 → 131,072\n- 与 Qwen2-72B 结构相同，差异全在数据与训练\n\n参考：Qwen2.5 技术报告（arXiv 2412.15115）。","memory_md":"| 精度 | 权重大小 | 参考硬件 | 说明 |\n|---|---|---|---|\n| BF16 | 145.4 GB | 2×80GB | 官方 safetensors |\n| Q8 | ≈ 77 GB | 80GB 单卡 | GGUF Q8_0 |\n| Q4 | 47.4 GB | 2×24GB / 80GB | GGUF Q4_K_M（官方 + bartowski） |\n| AWQ/GPTQ int4 | ≈ 41 GB | 48GB | 官方量化版 |\n\n**KV Cache**：320 KiB/token（BF16）。\n\n**参考配置**：\n- 2×RTX 4090：Q4_K_M 拆分，上下文 8–16K\n- A100/H100 80GB：官方 AWQ / FP8，32K 上下文\n- 2×80GB：BF16","training_md":"- 预训练 18T token（Qwen2 为 7T），数学与代码合成数据大量加入\n- 后训练：百万级 SFT + DPO + GRPO 两阶段 RL\n- 工具调用：Hermes 风格 JSON；支持 system prompt 灵活角色\n- 长上下文：YaRN 需在 config 中手动启用（rope_scaling）","ecosystem_md":"- HF：Qwen/Qwen2.5-72B-Instruct、-AWQ、-GPTQ-Int4/Int8、-GGUF\n- 引擎：vLLM、SGLang、llama.cpp、Ollama、TensorRT-LLM、MLX\n- 微调：LLaMA-Factory、Unsloth、ms-swift、Axolotl 全支持\n- 中文文档：有（官方文档 qwen.readthedocs.io 中英双语）","versions_md":"- 同系列：Qwen2.5-32B（Apache-2.0）、Qwen2.5-Coder-32B、Qwen2.5-Math\n- 后继：QwQ-32B（推理）→ Qwen3-235B-A22B（2025-04）→ Qwen3 2507 系列\n- 上代：Qwen2-72B（2024-06）"},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Ollama","TensorRT-LLM","MLX"],"finetune":"LoRA / 全参均友好","zh_docs":"有"}},{"aliases":["Qwen2.5-Coder-32B-Instruct"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"coding":"HumanEval 92.7%、LiveCodeBench 31.4%、Aider 73.7%（官方）。","chinese":"中文注释与需求理解好。"},"complete":false,"id":"qwen2-5-coder-32b","name":"Qwen2.5-Coder-32B","name_zh":"通义千问 2.5 Coder · 32B","vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen2.5-Coder","license":"Apache-2.0","license_commercial":true,"weights_url":"https://huggingface.co/Qwen/Qwen2.5-Coder-32B-Instruct","superseded_by":"qwen3-coder-480b-a35b","released_at":"2024-11-12","architecture":{"type":"dense","total_params":"32.5B","total_params_b":32.5,"layers":64,"hidden_size":5120,"vocab_size":152064,"kv_heads":8,"head_dim":128,"attention":"GQA（40 Q 头 / 8 KV 头）","notes":"5.5T 代码为主数据续训；支持 FIM 补全；128K（YaRN）。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":65,"q8":34.5,"q4":19.9},"estimated":false,"kv_per_token_kib":256,"kv_note":"256 KiB/token。","ref_hw_24gb":"Q4_K_M 19.9 GB，RTX 4090 可跑 8–16K","ref_hw_80gb":"BF16 65 GB 单卡","ref_hw_8x80gb":"过剩"},"links":{"official":"https://qwenlm.github.io/blog/qwen2.5-coder-family/","github":"https://github.com/QwenLM/Qwen2.5-Coder","paper":"https://arxiv.org/abs/2409.12186","hf":"https://huggingface.co/Qwen/Qwen2.5-Coder-32B-Instruct"},"copy":{"one_liner":"首个代码能力对标 GPT-4o 的开源模型，单卡可跑。","highlights":["官方称 HumanEval / EvalPlus / Aider 等对标 GPT-4o","Apache-2.0，24GB 卡 Q4 可跑，本地 Copilot 首选","支持 FIM 补全与仓库级上下文"],"pitfalls":["非 agent 训练，SWE-bench 类多步任务弱","无推理模式，算法题不如 QwQ / Qwen3","2025 年被 Qwen3-Coder 与 Devstral 取代"],"logic_ability":"代码补全 / 生成极强（HumanEval 92.7%、Aider 73.7%，官方），通用推理与数学是 32B Dense 中上水平。","best_for":["本地代码补全 / IDE 插件","代码微调底座"],"not_for":["复杂 agent 工作流","通用对话"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama","MLX"],"finetune":"LoRA 友好","zh_docs":"有"}},{"aliases":["Qwen2-72B-Instruct"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","capability_notes":{"coding":"HumanEval 86.0%（官方 Instruct）。","math":"MATH 59.7%（官方 Instruct）。","knowledge":"MMLU 82.3%（官方基座）。","chinese":"C-Eval 91.0%（官方）。"},"complete":false,"id":"qwen2-72b","name":"Qwen2-72B","name_zh":"通义千问 2 · 72B","vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen2","license":"Tongyi Qianwen License","license_commercial":"restricted","weights_url":"https://huggingface.co/Qwen/Qwen2-72B-Instruct","superseded_by":"qwen2-5-72b","released_at":"2024-06-06","architecture":{"type":"dense","total_params":"72.7B","total_params_b":72.7,"layers":80,"hidden_size":8192,"vocab_size":152064,"kv_heads":8,"head_dim":128,"attention":"GQA（64 Q 头 / 8 KV 头）","notes":"Qwen 首次全系 GQA；上下文 128K（YaRN）；27 语言。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K"},"memory":{"weight_gb":{"bf16":145.4,"q8":77.1,"q4":47.4},"estimated":false,"kv_per_token_kib":320,"kv_note":"320 KiB/token。","ref_hw_24gb":"Q4_K_M 47.4 GB 需 2×24GB","ref_hw_80gb":"Q4 单卡 80GB；BF16 2×80GB","ref_hw_8x80gb":"BF16 高并发"},"links":{"official":"https://qwenlm.github.io/blog/qwen2/","github":"https://github.com/QwenLM/Qwen2","paper":"https://arxiv.org/abs/2407.10671","hf":"https://huggingface.co/Qwen/Qwen2-72B-Instruct"},"copy":{"one_liner":"2024 年中期开源榜首之一，Qwen 全面转 GQA + 128K。","highlights":["发布时 Open LLM Leaderboard 第一，超过 Llama 3 70B","128K 上下文，Needle 测试全绿","72B 仍保留自有许可，其余尺寸 Apache-2.0"],"pitfalls":["72B 许可非 Apache（1 亿月活限制）","无推理模式，代码与数学被 2.5 大幅刷新","官方 3 个月后即被 Qwen2.5 取代"],"logic_ability":"非推理模型一线：MMLU 82.3%（基座）、MATH 59.7%、HumanEval 86.0%（Instruct，官方）。","best_for":["历史对照 / 已有部署"],"not_for":["新项目（直接用 2.5 / 3）"]},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Ollama"],"finetune":"LoRA 成熟","zh_docs":"有"}},{"id":"qwen3-235b-a22b-thinking-2507","name":"Qwen3-235B-A22B-Thinking-2507","name_zh":"通义千问 3 · 235B 推理版","aliases":["Qwen3 235B Thinking","qwen3-235b-a22b-thinking"],"vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen3","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/Qwen/Qwen3-235B-A22B-Thinking-2507","status":"superseded","released_at":"2025-07-25","updated_at":"2025-12-20","modalities":["text","tools"],"reasoning_mode":"default-on","architecture":{"type":"moe","total_params":"235B","active_params":"22B","total_params_b":235,"active_params_b":22,"experts":128,"active_experts":8,"shared_expert":false,"layers":94,"hidden_size":4096,"vocab_size":151936,"kv_heads":4,"head_dim":128,"attention":"GQA（64 Q 头 / 4 KV 头）","notes":"无共享专家；QK-Norm；原生 262K 上下文，YaRN 可扩至 1M（需引擎支持）。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K","max_output":81920},"memory":{"weight_gb":{"bf16":470,"fp8":235,"q8":250,"q4":142},"kv_per_token_kib":188,"kv_note":"GQA 4 KV 头 × 128 × 2 × 94 层 × 2 B ≈ 188 KiB/token。128K 上下文单请求 ≈ 23 GB。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（Q4 142 GB）；2×80GB Q4 可跑但上下文受限","ref_hw_8x80gb":"BF16 / FP8 舒适，官方 FP8 版本推荐","estimated":true},"pricing":{"input_per_m":0.7,"output_per_m":8.4,"currency":"USD","source":"阿里云百炼（国际站）","as_of":"2025-12-20","note":"开源权重，此为官方托管价；第三方托管更便宜"},"links":{"official":"https://qwen.ai/","hf":"https://huggingface.co/Qwen/Qwen3-235B-A22B-Thinking-2507","github":"https://github.com/QwenLM/Qwen3","paper":"https://arxiv.org/abs/2505.09388"},"copy":{"one_liner":"Apache-2.0 的 MoE 推理旗舰，中文与数学突出。","highlights":["Apache-2.0，权重可下、商用无附加条件","AIME 2025 92.3%，开源模型里少数能跟上闭源旗舰的数学能力","22B 激活参数，吞吐远优于同能力的 Dense 模型"],"pitfalls":["Thinking 版默认输出大量思考 token，短任务成本高（用 Instruct-2507 版）","235B 权重，单卡 80GB 装不下 Q4，实际门槛 2×80GB 起","无共享专家 + 128 专家，对推理引擎的专家并行要求高"],"logic_ability":"考试型推理非常强（AIME、HMMT、GPQA），且喜欢写很长的思考链——对简单问题也会「想很久」，这是它的主要坑。工程型推理（SWE-bench）弱于同期闭源模型，也弱于专门的 Qwen3-Coder。中文逻辑题与阅读理解在开源里顶级。","best_for":["数学 / 逻辑 / 科学推理","中文复杂任务","多卡自建推理服务"],"not_for":["单卡本地部署","对延迟敏感的对话产品（思考太长）"]},"capability_notes":{"coding":"LiveCodeBench v6 74.1%（官方）；工程任务用 Qwen3-Coder。","reasoning":"GPQA Diamond 81.1%、HLE 18.2%（官方）。","math":"AIME 2025 92.3%、HMMT 83.9%（官方）。","agent":"BFCL-v3 71.9%、TAU2 ~ 70（官方）。","chinese":"C-Eval / 中文 SimpleQA 开源顶级。"},"sheet":{"architecture_md":"**类型**：MoE，235B 总 / 22B 激活。\n\n- 94 层，隐藏维 4096，词表 151,936\n- 专家：128 个，每 token 激活 8 个，**无共享专家**\n- 注意力：GQA，64 Query 头 / 4 KV 头，head_dim 128，QK-Norm\n- 上下文：原生 262,144；YaRN factor 4 可到 1M（vLLM/SGLang 需显式配置）\n- 与 Instruct-2507 共享结构，仅后训练不同\n\n参考：Qwen3 技术报告（arXiv 2505.09388）。","memory_md":"| 精度 | 权重大小 | 参考显存 | 说明 |\n|---|---|---|---|\n| BF16 | ≈ 470 GB | 8×80GB | 官方 safetensors |\n| FP8 | ≈ 235 GB | 4×80GB | 官方 FP8 版本 |\n| Q8 | ≈ 250 GB（估） | 4×80GB | GGUF Q8_0 |\n| Q4 | ≈ 142 GB（估） | 2×80GB | GGUF Q4_K_M（unsloth） |\n\n**KV Cache**：≈ 188 KiB/token（BF16）。128K 上下文 ≈ 23 GB；Thinking 模式输出常达 30K+，请留足。\n\n**参考配置**：\n- 24GB：不可行\n- 单卡 80GB：不可行\n- 8×H100：BF16 舒适，FP8 可服务较高并发","training_md":"- 预训练 36T token，119 语言\n- 三阶段预训练（通用 → 推理密集 → 长上下文）\n- 后训练：长 CoT 冷启动 → 推理 RL → 通用 RL；2507 版为增强版后训练\n- **默认思考且不可关闭**（关闭思考请用 Instruct-2507）\n- 工具调用：Hermes 风格 / OpenAI 兼容，支持 Qwen-Agent\n- 最大输出 81,920","ecosystem_md":"- HF：Qwen/Qwen3-235B-A22B-Thinking-2507（+ FP8 版）\n- 推理引擎：vLLM、SGLang、TensorRT-LLM、llama.cpp（GGUF）、Ollama\n- 微调：LLaMA-Factory / ms-swift 支持 LoRA，全参需多节点\n- 中文文档：有（官方中文文档完整）","versions_md":"- 同系列：Qwen3-235B-A22B-Instruct-2507（非思考）、Qwen3-32B、Qwen3-30B-A3B-2507\n- 后续：Qwen3-Next-80B-A3B（新架构）、Qwen3-Max（仅 API）\n- 被替代：Qwen3-235B-A22B（2025-04 混合思考版）"},"ecosystem":{"engines":["vLLM","SGLang","TensorRT-LLM","llama.cpp","Ollama"],"finetune":"LoRA 友好，全参需多节点","zh_docs":"有"},"complete":true,"superseded_by":"qwen3-8-2-4t-a95b"},{"aliases":["Qwen3-235B-A22B (April 2025)","Qwen3 235B"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"optional","capability_notes":{"coding":"LiveCodeBench v5 70.7%（思考，官方）。","reasoning":"GPQA Diamond 71.1%（思考，官方）。","math":"AIME 2025 81.5%（思考，官方）。","agent":"BFCL v3 70.8%（官方）。","chinese":"中文顶级。"},"complete":false,"id":"qwen3-235b-a22b","name":"Qwen3-235B-A22B","name_zh":"通义千问 3 · 235B-A22B（原版）","vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen3","license":"Apache-2.0","license_commercial":true,"weights_url":"https://huggingface.co/Qwen/Qwen3-235B-A22B","superseded_by":"qwen3-235b-a22b-thinking-2507","released_at":"2025-04-29","architecture":{"type":"moe","total_params":"235B","active_params":"22B","total_params_b":235,"active_params_b":22,"experts":128,"active_experts":8,"shared_expert":false,"layers":94,"hidden_size":4096,"vocab_size":151936,"kv_heads":4,"head_dim":128,"attention":"GQA（64 Q 头 / 4 KV 头）+ QK-Norm","notes":"原版混合思考（/think 与 /no_think 同一权重）；原生 32K，YaRN 至 128K；128 专家 top-8，无共享专家。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":32768},"memory":{"weight_gb":{"bf16":470,"q8":249.1,"q4":142.2,"fp8":235},"estimated":false,"kv_per_token_kib":188,"kv_note":"4×128×2×94×2 B = 188 KiB/token，MoE 大模型中最省。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（单卡）；Q4 142 GB 需 2×80GB","ref_hw_8x80gb":"BF16 470 GB 或 FP8 235 GB 高并发"},"pricing":{"input_per_m":0.13,"output_per_m":0.6,"currency":"USD","source":"第三方托管常见价（DeepInfra）","as_of":"2026-08-28","note":"开源权重，社区托管价"},"links":{"official":"https://qwenlm.github.io/blog/qwen3/","github":"https://github.com/QwenLM/Qwen3","paper":"https://arxiv.org/abs/2505.09388","hf":"https://huggingface.co/Qwen/Qwen3-235B-A22B"},"copy":{"one_liner":"Qwen3 首发旗舰，一套权重同时支持思考与非思考。","highlights":["混合思考首创：/think 与 /no_think 软开关","22B 激活对标 DeepSeek-R1 / o1，Apache-2.0","119 种语言，KV 极省（4 KV 头）"],"pitfalls":["混合模式导致两种模式都略逊于专训版，2507 拆成 Instruct / Thinking 两套后各自更强","原生 32K，128K 需 YaRN","235B 总参需多卡"],"logic_ability":"思考模式：AIME 2025 81.5%、GPQA 71.1%、LiveCodeBench 70.7%（官方）。非思考模式退化为强指令模型。混合训练使它在 agent 长任务上稳定性略逊于 2507 Thinking。","best_for":["单模型两用（对话 + 推理）","多语言"],"not_for":["单卡部署","极致推理（用 2507 Thinking）"]},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Ollama","MLX"],"finetune":"LoRA 可行（MoE 感知框架）","zh_docs":"有"}},{"id":"qwen3-30b-a3b-thinking-2507","name":"Qwen3-30B-A3B-Thinking-2507","name_zh":"通义千问 3 · 30B 小 MoE 推理版","aliases":["Qwen3 30B A3B","qwen3-30b-a3b-thinking"],"vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen3","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/Qwen/Qwen3-30B-A3B-Thinking-2507","status":"superseded","released_at":"2025-07-30","updated_at":"2025-12-20","modalities":["text","tools"],"reasoning_mode":"default-on","architecture":{"type":"moe","total_params":"30.5B","active_params":"3.3B","total_params_b":30.5,"active_params_b":3.3,"experts":128,"active_experts":8,"shared_expert":false,"layers":48,"hidden_size":2048,"vocab_size":151936,"kv_heads":4,"head_dim":128,"attention":"GQA（32 Q 头 / 4 KV 头）","notes":"每 token 只算 3.3B，CPU / Apple Silicon 也能有可用速度。原生 262K 上下文。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K","max_output":81920},"memory":{"weight_gb":{"bf16":61,"fp8":31,"q8":32.5,"q4":18.6},"kv_per_token_kib":96,"kv_note":"4 KV 头 × 128 × 2 × 48 层 × 2 B = 96 KiB/token。","ref_hw_24gb":"Q4_K_M 18.6 GB，剩 ~5 GB 给 KV ≈ 40K 上下文；生成速度 100+ tok/s","ref_hw_80gb":"BF16 61 GB 可跑，KV 余量 ~15 GB","ref_hw_8x80gb":"过剩","estimated":false},"pricing":{"input_per_m":0.08,"output_per_m":0.3,"currency":"USD","source":"第三方托管常见价","as_of":"2025-12-20","note":"开源权重，社区托管价"},"links":{"official":"https://qwen.ai/","hf":"https://huggingface.co/Qwen/Qwen3-30B-A3B-Thinking-2507","github":"https://github.com/QwenLM/Qwen3","paper":"https://arxiv.org/abs/2505.09388"},"copy":{"one_liner":"3B 激活的小 MoE，单卡跑出 32B 级推理，速度快 3 倍。","highlights":["激活仅 3.3B，4090 上 100+ tok/s，Mac M 系列也可用","AIME 2025 85%，推理分超过 Dense 的 Qwen3-32B","原生 256K 上下文，KV 仅 96 KiB/token"],"pitfalls":["30B 总参数仍要 18.6 GB（Q4）显存，不是「3B 模型」","小激活量导致知识广度与指令细节遵循弱于 Dense 32B","Thinking 版不可关闭思考，短问答成本高，用 Instruct-2507"],"logic_ability":"考试型推理相对其体量非常强，思考链长度控制好于 235B 版。工程型推理（多文件代码修改）明显弱，容易丢失约束。事实知识覆盖不足，幻觉高于 Dense 32B。适合「快、便宜、要有推理能力」的场景，不适合做知识库问答主力。","best_for":["本地推理助手","高吞吐批量推理","边缘 / 笔记本部署"],"not_for":["知识密集型问答","复杂 agent 编排"]},"capability_notes":{"coding":"LiveCodeBench v6 66.0%（官方）。","reasoning":"GPQA Diamond 73.4%（官方）。","math":"AIME 2025 85.0%（官方）。","agent":"BFCL-v3 72.4%（官方）。","chinese":"中文良好。"},"sheet":{"architecture_md":"**类型**：MoE，30.5B 总 / 3.3B 激活。\n\n- 48 层，隐藏维 2048，词表 151,936\n- 专家：128 个，每 token 激活 8 个，无共享专家\n- GQA：32 Q 头 / 4 KV 头，head_dim 128\n- 上下文：原生 262,144\n\n参考：Qwen3 技术报告。","memory_md":"| 精度 | 权重大小 | 参考显存 | 说明 |\n|---|---|---|---|\n| BF16 | 61 GB | 80GB | 官方 |\n| FP8 | ≈ 31 GB | 48GB | 官方 FP8 版 |\n| Q8 | 32.5 GB | 48GB | GGUF |\n| Q4 | 18.6 GB | 24GB | GGUF Q4_K_M |\n\n**KV Cache**：96 KiB/token，128K 上下文 ≈ 12 GB。\n\n**参考配置**：\n- 24GB：Q4，~40K 上下文\n- 80GB：BF16，128K 上下文可行\n- Mac 64GB 统一内存：Q8 舒适","training_md":"- 与 Qwen3 系列共享预训练（36T token）\n- 2507 版后训练增强推理深度\n- 思考默认且不可关\n- 工具调用支持\n- 最大输出 81,920","ecosystem_md":"- HF：Qwen/Qwen3-30B-A3B-Thinking-2507（+ FP8、GGUF）\n- 引擎：vLLM、SGLang、llama.cpp、Ollama、MLX\n- 微调：LoRA 友好（MoE 全参微调需注意路由稳定）\n- 中文文档：有","versions_md":"- 姊妹：Qwen3-30B-A3B-Instruct-2507（非思考）、Qwen3-Coder-30B-A3B\n- 更大：Qwen3-235B-A22B-2507\n- 新架构：Qwen3-Next-80B-A3B"},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Ollama","MLX"],"finetune":"LoRA 友好","zh_docs":"有"},"complete":true,"superseded_by":"qwen3-8-27b"},{"id":"qwen3-32b","name":"Qwen3-32B","name_zh":"通义千问 3 · 32B","aliases":["Qwen3 32B"],"vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen3","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/Qwen/Qwen3-32B","status":"superseded","released_at":"2025-04-29","updated_at":"2025-12-20","modalities":["text","tools"],"reasoning_mode":"optional","architecture":{"type":"dense","total_params":"32.8B","total_params_b":32.8,"layers":64,"hidden_size":5120,"vocab_size":151936,"kv_heads":8,"head_dim":128,"attention":"GQA（64 Q 头 / 8 KV 头）","notes":"Dense 模型；原生 32K，YaRN 至 128K。混合思考：/think 与 /no_think 软开关。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":32768},"memory":{"weight_gb":{"bf16":65.6,"fp8":33,"q8":34.8,"q4":19.8},"kv_per_token_kib":256,"kv_note":"8 KV 头 × 128 × 2 × 64 层 × 2 B = 256 KiB/token。32K 上下文 ≈ 8 GB。","ref_hw_24gb":"Q4_K_M 19.8 GB 可加载，留 ~4 GB 给 KV，约 8–16K 上下文","ref_hw_80gb":"BF16 65.6 GB 单卡可跑，KV 余量约 14 GB","ref_hw_8x80gb":"过剩；用于高并发服务","estimated":false},"pricing":{"input_per_m":0.1,"output_per_m":0.3,"currency":"USD","source":"第三方托管常见价（DeepInfra 等）","as_of":"2025-12-20","note":"开源权重，此为社区托管价，标注供参考"},"links":{"official":"https://qwen.ai/","hf":"https://huggingface.co/Qwen/Qwen3-32B","github":"https://github.com/QwenLM/Qwen3","paper":"https://arxiv.org/abs/2505.09388"},"copy":{"one_liner":"单卡 24GB 能跑的 Dense 全能型，混合思考可开关。","highlights":["Q4 量化 19.8 GB，RTX 4090 / 3090 单卡可跑","同一权重 /think 与 /no_think 切换，无需两套模型","Apache-2.0，微调生态最完整的 30B 档之一"],"pitfalls":["原生上下文 32K，128K 需 YaRN 且会略降短文本质量","24GB 卡上 Q4 只剩 ~4 GB 给 KV，长对话会 OOM","已被 2507 系列 MoE 小模型（30B-A3B）在多数榜上超过，但 Dense 仍更稳"],"logic_ability":"开启思考后考试型推理接近 2025 年初的 o1 水平（AIME 2025 72.9%），关闭思考则退化为常规 32B 指令模型。Dense 结构让它在多轮对话与指令遵循上比 30B-A3B 更稳，但吞吐慢 3–4 倍。常见问题：思考模式对简单问题也会展开长链。","best_for":["单卡本地部署","私有化微调底座","中英双语通用助手"],"not_for":["高吞吐批量（用 30B-A3B）","超长文档"]},"capability_notes":{"coding":"LiveCodeBench v5 ~65%（思考，官方）；日常代码可用。","reasoning":"GPQA Diamond 68.4%（思考，官方）。","math":"AIME 2025 72.9%（思考，官方）。","agent":"BFCL-v3 70.3%（官方）。","chinese":"中文顶级的 32B 档。"},"sheet":{"architecture_md":"**类型**：Dense Transformer，32.8B 参数（含 embedding），31.2B 非 embedding。\n\n- 64 层，隐藏维 5120，词表 151,936\n- GQA：64 Query 头 / 8 KV 头，head_dim 128；QK-Norm\n- 上下文：原生 32,768；YaRN factor 4 到 131,072\n- Tied embeddings：否\n\n参考：Qwen3 技术报告。","memory_md":"| 精度 | 权重大小 | 参考显存 | 说明 |\n|---|---|---|---|\n| BF16 | 65.6 GB | 80GB 单卡 | 官方 safetensors |\n| FP8 | ≈ 33 GB | 48GB | 官方 FP8 版 |\n| Q8 | 34.8 GB | 48GB | GGUF Q8_0 |\n| Q4 | 19.8 GB | 24GB | GGUF Q4_K_M |\n\n**KV Cache**：256 KiB/token（BF16）。\n\n**参考配置**：\n- RTX 4090 24GB：Q4_K_M，上下文 8–16K，单并发\n- A100/H100 80GB：BF16，32K 上下文，中等并发\n- 8×80GB：高并发服务","training_md":"- 预训练 36T token\n- 后训练：长 CoT 冷启动 → 推理 RL → 思考模式融合 → 通用 RL\n- 思考默认开启，`enable_thinking=False` 或 `/no_think` 关闭\n- 工具调用：Hermes 格式，Qwen-Agent\n- 最大输出 32K（思考建议 38K 总长）","ecosystem_md":"- HF：Qwen/Qwen3-32B、Qwen3-32B-FP8、GGUF（Qwen 官方 + unsloth）\n- 引擎：vLLM、SGLang、llama.cpp、Ollama、LM Studio、MLX\n- 微调：LLaMA-Factory、Unsloth、ms-swift 均一等支持\n- 中文文档：有","versions_md":"- 同系列：Qwen3-30B-A3B（MoE，更快）、Qwen3-14B、Qwen3-8B\n- 2507 更新未覆盖 32B（仅 235B 与 30B-A3B）\n- 上代：Qwen2.5-32B"},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Ollama","MLX"],"finetune":"LoRA / 全参均友好","zh_docs":"有"},"complete":true,"superseded_by":"qwen3-8-27b"},{"id":"qwen3-8-2-4t-a95b","name":"Qwen3.8-2.4T-A95B","name_zh":"通义千问 3.8 · 2.4T-A95B（开源 Max 级）","aliases":["Qwen3.8 2.4T A95B","Qwen3.8-Max（API 版）","qwen3.8-max"],"vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen3.8","license":"Qwen3.8-Max License","license_commercial":"restricted","openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B","status":"current","released_at":"2026-08-13","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"default-on","architecture":{"type":"moe","total_params":"2.4T","active_params":"95B","total_params_b":2400,"active_params_b":95,"experts":512,"active_experts":10,"shared_expert":true,"layers":92,"kv_layers":23,"hidden_size":8192,"vocab_size":248320,"kv_heads":4,"head_dim":256,"attention":"Gated DeltaNet × 3 + Gated Attention（GQA，64 Q / 4 KV）× 1 交错","notes":"92 层 = 23 × (3 × Gated DeltaNet → MoE + 1 × Gated Attention → MoE)。512 路由专家 + 1 共享专家，每 token 激活 10 路由 + 1 共享；专家中间维 2,048。DeltaNet 128 V 头 / 16 QK 头、head_dim 128；全注意力 64 Q / 4 KV、head_dim 256，RoPE 占 1/4。自带 1 层 MTP。表内 kv_heads / head_dim 指 23 层全注意力层。开源权重为纯文本；API 版 Qwen3.8-Max 在同一底座上加了视觉输入、非思考模式与默认 1M 上下文。 上下文：256K（可扩至 1M；API 版默认 1M）。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K","max_output":131072},"memory":{"weight_gb":{"bf16":4890,"fp8":2500,"q8":2600,"q4":1310},"kv_per_token_kib":92,"kv_note":"仅 23 层全注意力层有 KV：4 KV 头 × 256 × 2 × 2 B = 4 KiB/层 → 92 KiB/token（BF16）。69 层 DeltaNet 固定递归状态约 0.6 GB / 序列。256K 上下文 KV ≈ 23 GB。q4 一栏为 unsloth UD-IQ4_XS（1.31 TB），无 Q4_K_M。","ref_hw_24gb":"不可行","ref_hw_80gb":"不可行（单卡）","ref_hw_8x80gb":"不可行（单机 8×80GB = 640 GB，远小于 FP8 2.5 TB）；需多节点 H200 / B200 集群，或 IQ4_XS 1.31 TB 于 ≥ 2 台 8×H200","estimated":false},"pricing":{"input_per_m":2,"output_per_m":6,"currency":"USD","source":"阿里云 Model Studio 国际站 qwen3.8-max 定价","as_of":"2026-08-28","note":"此为 API 产品 Qwen3.8-Max 的价格（同一底座 + 视觉 / 1M 上下文 / 内置工具）；缓存命中输入 $0.25；中国站更便宜。开源权重本身无官方托管价"},"links":{"official":"https://qwen.ai/","hf":"https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B","github":"https://github.com/QwenLM/Qwen3.8","pricing":"https://www.alibabacloud.com/help/en/model-studio/models"},"variants":[{"kind":"fp8","publisher":"Qwen","repo":"Qwen/Qwen3.8-2.4T-A95B-FP8","url":"https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B-FP8","note":"官方 FP8","sizes":{"fp8":2496.1}},{"kind":"nvfp4","publisher":"RedHatAI","repo":"RedHatAI/Qwen3.8-2.4T-A95B-NVFP4","url":"https://huggingface.co/RedHatAI/Qwen3.8-2.4T-A95B-NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/Qwen3.8-2.4T-A95B-GGUF","url":"https://huggingface.co/unsloth/Qwen3.8-2.4T-A95B-GGUF","note":"2.4T 级仅 unsloth 提供 GGUF（UD-IQ4_XS 等），无 Q4_K_M","sizes":{"q8":2600.2,"bf16":4893.2}}],"copy":{"one_liner":"首个可下载的 Max 级 Qwen，2.4T MoE，自定义许可。","highlights":["Qwen 首次开放 Max 级权重：2.4T / 95B 激活，官方榜 Terminal Bench 2.1 86.6、GPQA Diamond 92.6","Artificial Analysis 智能指数 58，Agentic 指数一度全球第一；Text Arena 1479 分（qwen3.8-max）","3:1 DeltaNet 混合注意力，KV 仅 92 KiB/token，长上下文显存远小于同规模 GQA"],"pitfalls":["非 Apache-2.0：Qwen3.8-Max License 对 >1 亿 MAU / 月收入 >$2000 万要求展示模型名，MaaS 年收入 >$5000 万需另签许可","开源权重为纯文本且思考不可关闭；视觉、非思考、1M 默认上下文与内置工具仅 API 版 Qwen3.8-Max 有","BF16 4.89 TB / FP8 2.5 TB，自建门槛为多节点集群；官方榜单数字来自 API 版 Qwen3.8-Max，开源权重复现可能有差"],"logic_ability":"闭源旗舰级：GPQA Diamond 92.6、HLE 43.6（无工具）/ 56.2（有工具），Artificial Analysis 智能指数 58 仅次于 Anthropic 与 OpenAI 旗舰。强项在长程 Agent：Terminal Bench 2.1 86.6、PaperBench 93.0、DeepSWE 1.1 56.6，均远超上代 Qwen3.7-Max。思考强制开启，reasoning_effort xhigh 默认；AA 实测约 21 tok/s、输出极其冗长，不适合低延迟场景。常见问题：Agents' Last Exam Pass@1 仅 27.0，超长多步任务仍会失误。","best_for":["超大规模自建旗舰推理（有集群）","长程 Coding / Office Agent","通过 Qwen3.8-Max API 低价替代闭源旗舰"],"not_for":["单机部署","需要视觉输入的开源自建（用 27B）","MaaS 大厂（许可限制）"]},"capability_notes":{"coding":"Terminal Bench 2.1 86.6、SWE-bench Pro 67.7、DeepSWE 1.1 56.6、NL2Repo 55.9（官方模型卡，Qwen3.8-Max）。","reasoning":"GPQA Diamond 92.6、HLE 43.6 / 56.2（有工具）（官方模型卡）。","math":"AIME 官方未披露。","agent":"Toolathlon Verified 72.5、WideSearch 81.9、Agents' Last Exam 27.0 Pass@1、CoWorkBench 74.8（官方模型卡）。","instruction":"IFBench 82.8（官方模型卡）。","chinese":"中文旗舰级；PLawBench 73.2。"},"sheet":{"architecture_md":"**类型**：MoE + 混合注意力，2.4T 总参数 / 95B 激活。\n\n| 项目 | 数值 |\n|---|---|\n| 层数 | 92 = 23 × (3 × DeltaNet → MoE + 1 × Gated Attention → MoE) |\n| 隐藏维 | 8,192 |\n| 专家 | 512 路由 + 1 共享，每 token 激活 10 路由 + 1 共享 |\n| 专家中间维 | 2,048（共享专家同） |\n| 全注意力层 | 23 层，64 Q 头 / 4 KV 头，head_dim 256 |\n| 线性注意力层 | 69 层，128 V 头 / 16 QK 头，head_dim 128 |\n| RoPE | partial_rotary_factor 0.25（64 维），theta 1e7 |\n| 词表 | 248,320（padded） |\n| 原生上下文 | 262,144，可扩至 1,010,000 |\n| MTP | 1 层 |\n\n`config.json`：`model_type: qwen3_5_moe_text`，`Qwen3_5MoeForCausalLM`，纯文本。\n\n**与 Qwen3.8-Max 的关系**：模型卡称 Qwen3.8-Max 是「基于 Qwen3.8-2.4T-A95B 的官方版本」，额外具备视觉输入、非思考模式、默认 1M 上下文与内置工具。API 名 `qwen3.8-max`。","memory_md":"| 精度 | 权重大小 | 说明 |\n|---|---|---|\n| BF16 | 4.89 TB | 官方 safetensors，213 分片 |\n| FP8 | ≈ 2.5 TB | 官方 Qwen3.8-2.4T-A95B-FP8 仓库 |\n| Q8 | 2.6 TB | unsloth GGUF Q8_0 |\n| IQ4_XS | 1.31 TB | unsloth UD-IQ4_XS |\n| IQ2_XXS | 657 GB | unsloth UD-IQ2_XXS |\n| IQ1_S | 508 GB | unsloth UD-IQ1_S（质量损失大） |\n\n**KV Cache**：92 KiB/token（BF16，仅 23 层全注意力层）；DeltaNet 状态约 0.6 GB / 序列。256K ≈ 23 GB。\n\n**参考配置**：\n- 24GB / 80GB 单卡：不可行\n- 单机 8×80GB：不可行（640 GB < 最小可用量化）\n- 多节点：FP8 需 ≥ 3–4 台 8×H200（141GB）级别；IQ4_XS 可在 2 台 8×H200 装下但精度有损\n\n警示：模型卡未给官方硬件建议；上述节点数为按文件大小换算，未含 KV 与并发余量。","training_md":"- 预训练 token 数：**未披露**\n- 后训练：模型卡仅写 Pre-training & Post-training，带 MTP 训练\n- 思考模式：**强制开启，不可关闭**（非思考仅 API 版 Qwen3.8-Max 提供）\n- `reasoning_effort`：xhigh（默认）/ medium / low；preserve-thinking 默认开\n- 最大输出：推理内容 262,144，最终回复建议 131,072\n- 推荐采样：temperature 1.0、top_p 0.95、top_k 20、min_p 0\n- 官方榜单说明：Terminal Bench 2.1 用 Claude Code 评测（avg@10），对比模型取已公开最佳分","ecosystem_md":"- HF：Qwen/Qwen3.8-2.4T-A95B、Qwen3.8-2.4T-A95B-FP8；GGUF：unsloth/Qwen3.8-2.4T-A95B-GGUF\n- 引擎：SGLang、vLLM、TokenSpeed（官方推荐生产用）、Transformers ≥ 4.57.3\n- 微调：全参不现实；LoRA 需多节点\n- API：阿里云 Model Studio `qwen3.8-max`（$2 / $6，缓存 $0.25），OpenAI / Anthropic / DashScope 三种兼容接口\n- 中文文档：有","versions_md":"- API 版：Qwen3.8-Max（2026-08-03 GA，7 月 19 日 WAIC 预览）\n- 同代开源：Qwen3.8-27B（Apache-2.0，多模态）\n- 上代：Qwen3.7-Max（API-only）、Qwen3-Max（API-only）\n- 更早开源旗舰：Qwen3-235B-A22B-2507 系列"},"ecosystem":{"engines":["SGLang","vLLM","TokenSpeed","Transformers"],"finetune":"多节点 LoRA","zh_docs":"有"},"complete":true,"runtime":{"tok_s":25,"latency_s":2.53,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"Qwen 3.8 · 2.4T-A95B (open Max-class)","one_liner":"First downloadable Max-class Qwen, 2.4T MoE, custom license.","highlights":["Qwen's first open Max-class weights: 2.4T / 95B active; official leaderboard Terminal Bench 2.1 86.6, GPQA Diamond 92.6","Artificial Analysis Intelligence Index 58; Agentic Index was briefly #1 worldwide; Text Arena 1479 (qwen3.8-max)","3:1 DeltaNet hybrid attention, KV only 92 KiB/token; long-context memory far below same-size GQA"],"pitfalls":["Not Apache-2.0: the Qwen3.8-Max License requires displaying the model name above 100M MAU / $20M monthly revenue; MaaS with annual revenue over $50M needs a separate license","Open weights are text-only and thinking cannot be disabled; vision, non-thinking, 1M default context and built-in tools are only in the API version Qwen3.8-Max","BF16 4.89 TB / FP8 2.5 TB; self-hosting requires a multi-node cluster; official leaderboard numbers come from the API Qwen3.8-Max, and open-weight reproduction may differ"],"logic_ability":"Closed-flagship class: GPQA Diamond 92.6, HLE 43.6 (no tools) / 56.2 (with tools); Artificial Analysis Intelligence Index 58, behind only the Anthropic and OpenAI flagships. Strength is long-horizon agents: Terminal Bench 2.1 86.6, PaperBench 93.0, DeepSWE 1.1 56.6, all far above the previous Qwen3.7-Max. Thinking is forced on, reasoning_effort defaults to xhigh; AA measured about 21 tok/s with extremely verbose output, unsuitable for low-latency scenarios. Common issue: Agents' Last Exam Pass@1 only 27.0; very long multi-step tasks still fail.","best_for":["Very-large-scale self-hosted flagship inference (with a cluster)","Long-horizon coding / office agents","Cheap replacement for closed flagships via the Qwen3.8-Max API"],"not_for":["Single-machine deployment","Open self-hosting that needs vision input (use 27B)","Large MaaS providers (license restriction)"],"capability_notes":{"coding":"Terminal Bench 2.1 86.6, SWE-bench Pro 67.7, DeepSWE 1.1 56.6, NL2Repo 55.9 (official model card, Qwen3.8-Max).","reasoning":"GPQA Diamond 92.6, HLE 43.6 / 56.2 (with tools) (official model card).","math":"AIME not officially disclosed.","agent":"Toolathlon Verified 72.5, WideSearch 81.9, Agents' Last Exam 27.0 Pass@1, CoWorkBench 74.8 (official model card).","instruction":"IFBench 82.8 (official model card).","chinese":"Flagship-class Chinese; PLawBench 73.2."}},"ja":{"name_zh":"Qwen 3.8 · 2.4T-A95B（オープン Max 級）","one_liner":"初のダウンロード可能な Max 級 Qwen。2.4T MoE、独自ライセンス。","highlights":["Qwen 初の Max 級オープンウェイト：2.4T / 95B アクティブ、公式榜で Terminal Bench 2.1 86.6、GPQA Diamond 92.6","Artificial Analysis 知能指数 58、Agentic 指数は一時世界 1 位。Text Arena 1479（qwen3.8-max）","3:1 DeltaNet ハイブリッドアテンション、KV わずか 92 KiB/token、長コンテキストの VRAM は同規模 GQA よりはるかに少ない"],"pitfalls":["Apache-2.0 ではない：Qwen3.8-Max License は MAU 1 億超 / 月収 2000 万ドル超でモデル名表示を要求、MaaS で年収 5000 万ドル超は別途ライセンスが必要","オープンウェイトはテキストのみで思考は無効化不可。ビジョン、非思考、1M デフォルトコンテキスト、内蔵ツールは API 版 Qwen3.8-Max 限定","BF16 4.89 TB / FP8 2.5 TB で自前ホスティングはマルチノードクラスタが前提。公式榜の数値は API 版 Qwen3.8-Max 由来で、オープンウェイトでの再現は差が出る可能性あり"],"logic_ability":"クローズド旗艦級：GPQA Diamond 92.6、HLE 43.6（ツールなし）/ 56.2（ツールあり）、Artificial Analysis 知能指数 58 で Anthropic と OpenAI の旗艦に次ぐ。強みは長期エージェント：Terminal Bench 2.1 86.6、PaperBench 93.0、DeepSWE 1.1 56.6 で、いずれも前世代 Qwen3.7-Max を大きく上回る。思考は強制オン、reasoning_effort はデフォルト xhigh。AA 実測約 21 tok/s で出力が極めて冗長、低レイテンシ用途には不向き。よくある問題：Agents' Last Exam Pass@1 はわずか 27.0、超長い多段階タスクでは依然ミスが出る。","best_for":["超大規模な自前旗艦推論（クラスタあり）","長期のコーディング / オフィスエージェント","Qwen3.8-Max API 経由でクローズド旗艦を低価格に代替"],"not_for":["単一マシンでのデプロイ","ビジョン入力が必要なオープン自前運用（27B を使う）","大手 MaaS 事業者（ライセンス制限）"],"capability_notes":{"coding":"Terminal Bench 2.1 86.6、SWE-bench Pro 67.7、DeepSWE 1.1 56.6、NL2Repo 55.9（公式モデルカード、Qwen3.8-Max）。","reasoning":"GPQA Diamond 92.6、HLE 43.6 / 56.2（ツールあり）（公式モデルカード）。","math":"AIME は公式未開示。","agent":"Toolathlon Verified 72.5、WideSearch 81.9、Agents' Last Exam 27.0 Pass@1、CoWorkBench 74.8（公式モデルカード）。","instruction":"IFBench 82.8（公式モデルカード）。","chinese":"中国語は旗艦級。PLawBench 73.2。"}}}},{"id":"qwen3-8-27b","name":"Qwen3.8-27B","name_zh":"通义千问 3.8 · 27B","aliases":["Qwen3.8 27B","qwen3.8-27b"],"vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen3.8","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/Qwen/Qwen3.8-27B","status":"current","released_at":"2026-08-14","updated_at":"2026-08-28","modalities":["text","image","video","tools","computer-use"],"reasoning_mode":"default-on","architecture":{"type":"hybrid","total_params":"27B","total_params_b":27,"layers":64,"kv_layers":16,"hidden_size":5120,"vocab_size":248320,"kv_heads":4,"head_dim":256,"attention":"Gated DeltaNet（线性注意力）× 3 + Gated Attention（GQA，24 Q / 4 KV）× 1 交错","notes":"Dense 主干但注意力为混合结构：64 层 = 16 × (3 × Gated DeltaNet → FFN + 1 × Gated Attention → FFN)。DeltaNet 48 V 头 / 16 QK 头、head_dim 128；Gated Attention 24 Q 头 / 4 KV 头、head_dim 256，RoPE 仅占 1/4（64 维）。FFN 中间维 17,408。自带 1 层 MTP 头。视觉编码器 27 层、hidden 1152、patch 16。表内 kv_heads / head_dim 指 16 层全注意力层。 上下文：256K（YaRN 可扩至 1M）。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K","max_output":131072},"memory":{"weight_gb":{"bf16":55.6,"fp8":30.9,"q8":29,"q4":16.5},"kv_per_token_kib":64,"kv_note":"仅 16 层全注意力层有 KV：4 KV 头 × 256 × 2（K/V）× 2 B = 4 KiB/层 → 64 KiB/token（BF16）。另有 48 层 DeltaNet 的固定递归状态（48 头 × 128 × 128 × FP32），每序列约 150 MB，与长度无关。256K 上下文 KV ≈ 16 GB。","ref_hw_24gb":"UD-Q4_K_M 16.5 GB 可加载，剩约 6 GB 给 KV，约 64K 上下文单并发；视觉编码器另占少量显存","ref_hw_80gb":"BF16 55.6 GB 单卡可跑，KV 余量约 20 GB ≈ 256K 上下文；或 FP8 30.9 GB 留更多并发","ref_hw_8x80gb":"过剩；用于高并发 / 长上下文服务","estimated":false},"pricing":{"input_per_m":0.289,"output_per_m":2.4,"currency":"USD","source":"第三方托管参考价（OpenRouter）","as_of":"2026-08-28","note":"开源权重，官方 Qwen Cloud 托管端点尚未上线定价；此为社区托管价仅供参考"},"links":{"official":"https://qwen.ai/","hf":"https://huggingface.co/Qwen/Qwen3.8-27B","github":"https://github.com/QwenLM/Qwen3.8"},"variants":[{"kind":"fp8","publisher":"Qwen","repo":"Qwen/Qwen3.8-27B-FP8","url":"https://huggingface.co/Qwen/Qwen3.8-27B-FP8","note":"官方 FP8","sizes":{"fp8":30.9}},{"kind":"nvfp4","publisher":"unsloth","repo":"unsloth/Qwen3.8-27B-NVFP4","url":"https://huggingface.co/unsloth/Qwen3.8-27B-NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/Qwen3.8-27B-GGUF","url":"https://huggingface.co/unsloth/Qwen3.8-27B-GGUF","sizes":{"q8":29,"bf16":54.7}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/Qwen3.8-27B-GGUF","url":"https://huggingface.co/bartowski/Qwen3.8-27B-GGUF","sizes":{"q4":17.8,"q5":20.8,"q6":23.5,"q8":29.1,"bf16":54.7}},{"kind":"awq","publisher":"cyankiwi","repo":"cyankiwi/Qwen3.8-27B-AWQ-INT4","url":"https://huggingface.co/cyankiwi/Qwen3.8-27B-AWQ-INT4"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/Qwen3.8-27B-4bit","url":"https://huggingface.co/mlx-community/Qwen3.8-27B-4bit"}],"copy":{"one_liner":"24GB 单卡能跑的原生多模态 27B，Apache-2.0。","highlights":["Apache-2.0，Q4 仅 16.5 GB，RTX 4090 / 5090 单卡可跑，原生图像与视频输入","3:1 DeltaNet 混合注意力，KV 仅 64 KiB/token，256K 原生上下文可在 80GB 单卡跑满","官方榜上整体超过闭源 Qwen3.7-Plus：SWE-bench Pro 61.7、OSWorld-Verified 84.3、LiveCodeBench v6 90.3"],"pitfalls":["思考默认开启且非常「话痨」（AA 评为 very verbose），推理 token 成本与延迟高","混合 DeltaNet 结构较新，llama.cpp / MLX 需最新版本，旧版引擎与部分微调框架不支持","官方未披露预训练 token 数与后训练细节；SWE-bench Verified、AIME 等常用榜未给官方数"],"logic_ability":"27B 档目前最强的开源逻辑推理之一：GPQA Diamond 89.2、HLE 30.8（无工具）、LiveCodeBench v6 90.3，均已接近或超过上一代 Qwen3.7-Plus 闭源 API。Artificial Analysis 智能指数 52，在同体量开源模型中排第一。reasoning_effort 提供 xhigh / medium / low 三档，xhigh 下推理链很长；medium 档可明显缩短但考试型分数会下降。Agent 场景（Terminal Bench 2.1 73.0、DeepSWE 1.1 42.2）在 30B 档大幅领先。常见问题：思考输出冗长、简单问题也会展开长链。","best_for":["单卡本地多模态助手 / 桌面 Agent","私有化 Coding Agent 底座","视觉文档、图表理解（OmniDocBench 91.1）"],"not_for":["对首 token 延迟敏感的实时对话","无法升级推理引擎的老旧部署环境"]},"capability_notes":{"coding":"SWE-bench Pro 61.7、LiveCodeBench v6 90.3、Terminal Bench 2.1 73.0（官方模型卡）。","reasoning":"GPQA Diamond 89.2、HLE 30.8（官方模型卡）。","math":"MathVision 90.0（无代码解释器）/ 94.6（有）；AIME 官方未披露。","agent":"OSWorld-Verified 84.3、AndroidWorld 81.9、Agents' Last Exam 20.4 Pass@1（官方模型卡）。","multimodal":"原生图像 + 视频；OmniDocBench 1.5 91.1、RealWorldQA 85.9、CharXiv RQ 83.7（官方模型卡）。","instruction":"IFBench 79.5（官方模型卡）。","chinese":"中文顶级的 27B 档，Qwen 系列一贯优势。"},"sheet":{"architecture_md":"**类型**：Dense 主干 + 混合注意力（Gated DeltaNet / Gated Attention），27B 参数，原生多模态。\n\n| 项目 | 数值 |\n|---|---|\n| 层数 | 64 = 16 × (3 × DeltaNet → FFN + 1 × Gated Attention → FFN) |\n| 隐藏维 | 5,120 |\n| FFN 中间维 | 17,408 |\n| 全注意力层 | 16 层，24 Q 头 / 4 KV 头，head_dim 256 |\n| 线性注意力层 | 48 层，48 V 头 / 16 QK 头，head_dim 128，conv kernel 4 |\n| RoPE | partial_rotary_factor 0.25（64 维），theta 1e7，M-RoPE 交错 |\n| 词表 | 248,320（padded） |\n| 原生上下文 | 262,144，YaRN 可扩至 1,000,000 |\n| MTP | 1 层多 token 预测头 |\n| 视觉编码器 | 27 层，hidden 1152，patch 16，spatial merge 2，temporal patch 2 |\n\n`config.json` 中 `model_type: qwen3_5`，`architectures: Qwen3_5ForConditionalGeneration`。\n\n参考：Qwen/Qwen3.8-27B 模型卡与 config.json。","memory_md":"| 精度 | 权重大小 | 参考显存 | 说明 |\n|---|---|---|---|\n| BF16 | 55.6 GB | 80GB 单卡 | 官方 safetensors，18 分片 |\n| FP8 | 30.9 GB | 48GB | 官方 Qwen3.8-27B-FP8 仓库总大小 |\n| Q8 | 29 GB | 48GB | unsloth GGUF Q8_0 |\n| Q4 | 16.5 GB | 24GB | unsloth GGUF UD-Q4_K_M |\n\n**KV Cache**：只有 16 层全注意力层产生 KV，64 KiB/token（BF16）；DeltaNet 层为固定大小递归状态（约 150 MB / 序列）。256K 上下文 KV ≈ 16 GB。\n\n**参考配置**：\n- RTX 4090 / 5090 24GB：UD-Q4_K_M，约 64K 上下文单并发\n- A100 / H100 80GB：BF16 可跑满 256K；FP8 可做中等并发\n- 8×80GB：高并发服务\n\n警示：视觉编码器与视频帧 token 会额外吃显存；思考模式输出很长，预留输出 KV。","training_md":"- 预训练 token 数：**未披露**（模型卡仅写 Pre-training & Post-training）\n- 思考模式默认开启，可按请求关闭；`reasoning_effort` 支持 xhigh（默认）/ medium / low\n- Preserve-thinking：多轮对话默认保留上一轮推理上下文\n- 最大输出：推理内容最长 262,144，最终回复建议 131,072\n- 推荐采样：temperature 1.0、top_p 0.95、top_k 20、min_p 0\n- 训练时带 MTP，可用于投机解码","ecosystem_md":"- HF：Qwen/Qwen3.8-27B、Qwen3.8-27B-FP8；GGUF：unsloth/Qwen3.8-27B-GGUF（发布 24 小时内 HF 热度第 3）\n- 引擎：SGLang、vLLM、TokenSpeed、Transformers（≥ 5.8 dev）；llama.cpp / Ollama / LM Studio 需支持 qwen3_5 混合结构的新版本\n- 微调：Unsloth 已适配；LLaMA-Factory / ms-swift 跟进中\n- 中文文档：有","versions_md":"- 同代：Qwen3.8-2.4T-A95B（开源 Max 级 MoE）、Qwen3.8-Max（API）\n- 上代：Qwen3.6-27B（官方对比列）、Qwen3.7-Plus（API）\n- 更早：Qwen3-32B、Qwen3-30B-A3B-2507、Qwen3-Next-80B-A3B\n- 本代无 Qwen3.8 小体量 MoE 开源版（截至 2026-08-28 HF 集合仅 27B 与 2.4T-A95B 两档）"},"ecosystem":{"engines":["SGLang","vLLM","TokenSpeed","llama.cpp","Transformers"],"finetune":"Unsloth 已适配；全参需 ≥ 2×80GB","zh_docs":"有"},"complete":true,"runtime":{"tok_s":47,"latency_s":3.77,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"Qwen 3.8 · 27B","one_liner":"Natively multimodal 27B that runs on a 24GB single GPU, Apache-2.0.","highlights":["Apache-2.0; Q4 only 16.5 GB, runs on a single RTX 4090 / 5090; native image and video input","3:1 DeltaNet hybrid attention, KV only 64 KiB/token; full 256K native context fits on a single 80GB GPU","Official leaderboard beats the closed Qwen3.7-Plus overall: SWE-bench Pro 61.7, OSWorld-Verified 84.3, LiveCodeBench v6 90.3"],"pitfalls":["Thinking on by default and very \"chatty\" (rated very verbose by AA); high reasoning-token cost and latency","Hybrid DeltaNet architecture is new; llama.cpp / MLX need the latest versions; older engines and some fine-tuning frameworks unsupported","No official pretraining token count or post-training details; no official numbers on common leaderboards such as SWE-bench Verified and AIME"],"logic_ability":"One of the strongest open-source logical reasoners at the 27B tier: GPQA Diamond 89.2, HLE 30.8 (no tools), LiveCodeBench v6 90.3, all near or above the previous-generation closed Qwen3.7-Plus API. Artificial Analysis Intelligence Index 52, #1 among open models of the same size. reasoning_effort offers xhigh / medium / low; reasoning chains are very long at xhigh; medium shortens them noticeably but exam-style scores drop. Agent scenarios (Terminal Bench 2.1 73.0, DeepSWE 1.1 42.2) lead the 30B tier by a wide margin. Common issue: verbose thinking output; even simple questions trigger long chains.","best_for":["Single-GPU local multimodal assistants / desktop agents","Base for self-hosted coding agents","Visual document and chart understanding (OmniDocBench 91.1)"],"not_for":["Real-time chat sensitive to first-token latency","Legacy deployments that cannot upgrade the inference engine"],"capability_notes":{"coding":"SWE-bench Pro 61.7, LiveCodeBench v6 90.3, Terminal Bench 2.1 73.0 (official model card).","reasoning":"GPQA Diamond 89.2, HLE 30.8 (official model card).","math":"MathVision 90.0 (no code interpreter) / 94.6 (with); AIME not officially disclosed.","agent":"OSWorld-Verified 84.3, AndroidWorld 81.9, Agents' Last Exam 20.4 Pass@1 (official model card).","multimodal":"Native image + video; OmniDocBench 1.5 91.1, RealWorldQA 85.9, CharXiv RQ 83.7 (official model card).","instruction":"IFBench 79.5 (official model card).","chinese":"Top-tier Chinese at the 27B level, a consistent Qwen strength."}},"ja":{"name_zh":"Qwen 3.8 · 27B","one_liner":"24GB GPU で動くネイティブ多モーダル 27B。Apache-2.0。","highlights":["Apache-2.0、Q4 でわずか 16.5 GB、RTX 4090 / 5090 単一 GPU で動作。ネイティブ画像・動画入力","3:1 DeltaNet ハイブリッドアテンション、KV わずか 64 KiB/token、256K ネイティブコンテキストを 80GB 単一 GPU でフルに使える","公式榜で全体的にクローズドの Qwen3.7-Plus を上回る：SWE-bench Pro 61.7、OSWorld-Verified 84.3、LiveCodeBench v6 90.3"],"pitfalls":["思考はデフォルトオンで非常に「おしゃべり」（AA 評価は very verbose）、推論トークンのコストとレイテンシが高い","ハイブリッド DeltaNet 構造は新しく、llama.cpp / MLX は最新版が必要。旧版エンジンや一部ファインチューニングフレームワークは未対応","事前学習トークン数や事後学習の詳細は公式未開示。SWE-bench Verified、AIME などの定番榜も公式数値なし"],"logic_ability":"27B クラスで現在最強級のオープンソース論理推論：GPQA Diamond 89.2、HLE 30.8（ツールなし）、LiveCodeBench v6 90.3 で、いずれも前世代クローズドの Qwen3.7-Plus API に迫るか上回る。Artificial Analysis 知能指数 52 で同規模オープンモデル中 1 位。reasoning_effort は xhigh / medium / low の 3 段階、xhigh では推論チェーンが非常に長く、medium は大幅に短縮できるが試験型スコアは低下。エージェント用途（Terminal Bench 2.1 73.0、DeepSWE 1.1 42.2）は 30B クラスで大きくリード。よくある問題：思考出力が冗長で、簡単な質問でも長いチェーンを展開する。","best_for":["単一 GPU のローカルマルチモーダルアシスタント / デスクトップエージェント","プライベート環境のコーディングエージェント基盤","ビジュアルドキュメント・図表理解（OmniDocBench 91.1）"],"not_for":["初回トークン遅延に敏感なリアルタイム対話","推論エンジンを更新できない古いデプロイ環境"],"capability_notes":{"coding":"SWE-bench Pro 61.7、LiveCodeBench v6 90.3、Terminal Bench 2.1 73.0（公式モデルカード）。","reasoning":"GPQA Diamond 89.2、HLE 30.8（公式モデルカード）。","math":"MathVision 90.0（コードインタプリタなし）/ 94.6（あり）。AIME は公式未開示。","agent":"OSWorld-Verified 84.3、AndroidWorld 81.9、Agents' Last Exam 20.4 Pass@1（公式モデルカード）。","multimodal":"ネイティブ画像 + 動画。OmniDocBench 1.5 91.1、RealWorldQA 85.9、CharXiv RQ 83.7（公式モデルカード）。","instruction":"IFBench 79.5（公式モデルカード）。","chinese":"27B クラスで中国語は最上位、Qwen シリーズ一貫の強み。"}}}},{"aliases":["Qwen3 8B"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"optional","capability_notes":{"coding":"LiveCodeBench v5 57.5%（思考，技术报告）。","reasoning":"GPQA Diamond 62.0%（思考，技术报告）。","math":"AIME 2025 67.3%（思考，技术报告）。","chinese":"中文好。"},"complete":false,"id":"qwen3-8b","name":"Qwen3-8B","name_zh":"通义千问 3 · 8B","vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen3","license":"Apache-2.0","license_commercial":true,"weights_url":"https://huggingface.co/Qwen/Qwen3-8B","superseded_by":"qwen3-8-27b","released_at":"2025-04-29","architecture":{"type":"dense","total_params":"8.2B","total_params_b":8.2,"layers":36,"hidden_size":4096,"vocab_size":151936,"kv_heads":8,"head_dim":128,"attention":"GQA（32 Q 头 / 8 KV 头）+ QK-Norm","notes":"混合思考；原生 32K，YaRN 至 128K。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":32768},"memory":{"weight_gb":{"bf16":16.4,"q8":8.7,"q4":5,"fp8":8.2},"estimated":false,"kv_per_token_kib":144,"kv_note":"144 KiB/token。","ref_hw_24gb":"BF16 16.4 GB 单卡可跑 32K；Q4 5 GB 留大量 KV","ref_hw_80gb":"高并发","ref_hw_8x80gb":"过剩"},"pricing":{"input_per_m":0.035,"output_per_m":0.138,"currency":"USD","source":"第三方托管常见价（DeepInfra）","as_of":"2026-08-28","note":"开源权重，社区托管价"},"links":{"official":"https://qwenlm.github.io/blog/qwen3/","github":"https://github.com/QwenLM/Qwen3","paper":"https://arxiv.org/abs/2505.09388","hf":"https://huggingface.co/Qwen/Qwen3-8B"},"copy":{"one_liner":"8B 里带思考模式的全能选手，本地部署首选之一。","highlights":["思考模式下数学 / 代码超过 Qwen2.5-14B 乃至部分 32B","Apache-2.0，消费卡 BF16 直接跑","119 语言，工具调用与 MCP 生态完善"],"pitfalls":["思考链长，8B 吞吐优势被抵消","原生 32K，128K 需 YaRN 且质量下降","知识量受尺寸限制，幻觉多"],"logic_ability":"思考模式：AIME 2025 67.3%、GPQA 62.0%、LiveCodeBench v5 57.5%（Qwen3 技术报告）。8B 中推理最强档，但知识型任务仍受限。","best_for":["本地 / 边缘推理助手","微调底座"],"not_for":["知识密集问答","高并发低延迟"]},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Ollama","MLX"],"finetune":"LoRA / 全参极友好","zh_docs":"有"}},{"id":"qwen3-coder-480b-a35b","name":"Qwen3-Coder-480B-A35B-Instruct","name_zh":"通义千问 3 Coder · 480B","aliases":["Qwen3 Coder","qwen3-coder-plus"],"vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen3-Coder","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct","status":"current","released_at":"2025-07-22","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"none","architecture":{"type":"moe","total_params":"480B","active_params":"35B","total_params_b":480,"active_params_b":35,"experts":160,"active_experts":8,"layers":62,"hidden_size":6144,"kv_heads":8,"head_dim":128,"attention":"GQA（96 Q / 8 KV）","undisclosed":false},"context":{"max_tokens":262144,"display":"256K"},"memory":{"weight_gb":{"bf16":960,"fp8":480,"q4":280,"q8":510.4},"estimated":true,"ref_hw_8x80gb":"FP8 8×H100 可服务"},"pricing":{"input_per_m":1,"output_per_m":5,"currency":"USD","source":"阿里云百炼国际站（≤32K 档）","as_of":"2025-12-20","note":"阶梯计价"},"links":{"official":"https://qwenlm.github.io/blog/qwen3-coder/","hf":"https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct","github":"https://github.com/QwenLM/Qwen3-Coder"},"variants":[{"kind":"fp8","publisher":"Qwen","repo":"Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8","url":"https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8","note":"官方 FP8","sizes":{"fp8":482.1}},{"kind":"nvfp4","publisher":"nvidia","repo":"nvidia/Qwen3-Coder-480B-A35B-Instruct-NVFP4","url":"https://huggingface.co/nvidia/Qwen3-Coder-480B-A35B-Instruct-NVFP4"},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/Qwen3-Coder-480B-A35B-Instruct-GGUF","url":"https://huggingface.co/unsloth/Qwen3-Coder-480B-A35B-Instruct-GGUF","sizes":{"q4":290.1,"q5":340.5,"q6":394.1,"q8":510.4,"bf16":960.4}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/Qwen_Qwen3-Coder-480B-A35B-Instruct-GGUF","url":"https://huggingface.co/bartowski/Qwen_Qwen3-Coder-480B-A35B-Instruct-GGUF","sizes":{"q4":290.8,"q5":340.6,"q6":394.2,"q8":510.4}},{"kind":"awq","publisher":"QuantTrio","repo":"QuantTrio/Qwen3-Coder-480B-A35B-Instruct-AWQ","url":"https://huggingface.co/QuantTrio/Qwen3-Coder-480B-A35B-Instruct-AWQ"},{"kind":"mlx","publisher":"mlx-community","repo":"mlx-community/Qwen3-Coder-480B-A35B-Instruct-4bit","url":"https://huggingface.co/mlx-community/Qwen3-Coder-480B-A35B-Instruct-4bit"}],"copy":{"one_liner":"Apache-2.0 的编程 agent 专用 MoE，非思考模式。","highlights":["SWE-bench Verified 69.6%（官方）","Apache-2.0，Qwen Code CLI 配套","256K 上下文，YaRN 到 1M"],"pitfalls":["480B 体量，自建门槛高","非思考，复杂推理弱","已有 30B-A3B 小版可单卡"],"logic_ability":"工程型推理专精，长时程 agent RL 训练；考试型推理非重点。","best_for":["编程 agent"],"not_for":["通用推理"]},"capability_notes":{"coding":"SWE-bench Verified 69.6（官方博客）；SWE-bench Pro 38.7、Terminal-Bench 2.0 23.9（HF 模型卡）。","agent":"长时程 agent RL 训练，Qwen Code / CLINE 工具调用；Evasion Bench 78.16（HF 模型卡）。","reasoning":"非思考模式，官方未给 GPQA / HLE。","math":"官方未披露。","chinese":"中文顶级（Qwen 系列）。"},"ecosystem":{"engines":["vLLM","SGLang"],"zh_docs":"有"},"complete":false,"runtime":{"tok_s":66,"latency_s":2.97,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"Qwen3 Coder · 480B","one_liner":"Apache-2.0 MoE dedicated to coding agents, non-thinking mode.","highlights":["SWE-bench Verified 69.6% (official)","Apache-2.0, paired with Qwen Code CLI","256K context, up to 1M with YaRN"],"pitfalls":["480B size; high self-hosting barrier","Non-thinking; weak on complex reasoning","A smaller 30B-A3B version already runs on a single GPU"],"logic_ability":"Specialized in engineering reasoning with long-horizon agent RL training; exam-style reasoning is not the focus.","best_for":["Coding agents"],"not_for":["General reasoning"],"capability_notes":{"coding":"SWE-bench Verified 69.6 (official blog); SWE-bench Pro 38.7, Terminal-Bench 2.0 23.9 (HF model card).","agent":"Long-horizon agent RL training, Qwen Code / CLINE tool calling; Evasion Bench 78.16 (HF model card).","reasoning":"Non-thinking mode; no official GPQA / HLE.","math":"Not officially disclosed.","chinese":"Top-tier Chinese (Qwen series)."}},"ja":{"name_zh":"Qwen3 Coder · 480B","one_liner":"Apache-2.0 のコーディングエージェント専用 MoE。非思考モード。","highlights":["SWE-bench Verified 69.6%（公式）","Apache-2.0、Qwen Code CLI と組み合わせ","256K コンテキスト、YaRN で 1M まで"],"pitfalls":["480B 規模で自前ホスティングの敷居が高い","非思考で複雑な推論は弱い","単一 GPU で動く小型の 30B-A3B 版がすでにある"],"logic_ability":"エンジニアリング型推論に特化、長期エージェント RL 学習。試験型推論は重点外。","best_for":["コーディングエージェント"],"not_for":["汎用推論"],"capability_notes":{"coding":"SWE-bench Verified 69.6（公式ブログ）。SWE-bench Pro 38.7、Terminal-Bench 2.0 23.9（HF モデルカード）。","agent":"長期エージェント RL 学習、Qwen Code / CLINE のツール呼び出し。Evasion Bench 78.16（HF モデルカード）。","reasoning":"非思考モード、公式の GPQA / HLE なし。","math":"公式未開示。","chinese":"中国語は最上位（Qwen シリーズ）。"}}}},{"id":"qwen3-max","name":"Qwen3-Max","name_zh":"通义千问 3 Max","aliases":["qwen3-max","qwen3-max-2025-09-23"],"vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen3","license":"Proprietary","license_commercial":true,"openness":"api-only","weights_available":false,"status":"superseded","released_at":"2025-09-23","updated_at":"2025-12-20","modalities":["text","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"1T+","undisclosed":true,"notes":"官方仅称「超过 1T 参数 MoE」，细节未披露；仅 API。"},"context":{"max_tokens":262144,"display":"256K"},"memory":{"weight_gb":{},"estimated":false},"pricing":{"input_per_m":1.2,"output_per_m":6,"currency":"USD","source":"阿里云百炼国际站（≤32K 档）","as_of":"2025-12-20","note":"阶梯计价"},"links":{"official":"https://qwen.ai/","pricing":"https://www.alibabacloud.com/help/en/model-studio/models"},"copy":{"one_liner":"Qwen 系列仅 API 的 1T+ 旗舰，中文与 agent 强。","highlights":["1T+ MoE","SWE-bench Verified 69.6%（官方）","中文 / 阿里云生态"],"pitfalls":["仅 API，不开源（与 Qwen3 开源系列区分）","国际站可用性 / 合规按地区","参数细节未披露"],"logic_ability":"工程型与中文推理强；考试型推理 Thinking 版另计。","best_for":["中文企业应用","阿里云生态"],"not_for":["需要权重"]},"capability_notes":{},"complete":false,"superseded_by":"qwen3-8-2-4t-a95b"},{"id":"qwen3-next-80b-a3b-thinking","name":"Qwen3-Next-80B-A3B-Thinking","name_zh":"通义千问 3 Next · 80B 混合注意力","aliases":["Qwen3 Next","qwen3-next-80b"],"vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"Qwen3-Next","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking","status":"superseded","released_at":"2025-09-11","updated_at":"2025-12-20","modalities":["text","tools"],"reasoning_mode":"default-on","architecture":{"type":"hybrid","total_params":"80B","active_params":"3B","total_params_b":80,"active_params_b":3,"experts":512,"active_experts":10,"shared_expert":true,"layers":48,"kv_layers":12,"hidden_size":2048,"vocab_size":151936,"kv_heads":2,"head_dim":256,"attention":"Gated DeltaNet ×3 + Gated Attention ×1（混合，3:1）","notes":"48 层中 36 层为 Gated DeltaNet（线性注意力），12 层为门控全注意力（16 Q 头 / 2 KV 头，head_dim 256）。512 专家（10 路由 + 1 共享）。多 token 预测。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K","max_output":32768},"memory":{"weight_gb":{"bf16":160,"fp8":80,"q4":47},"kv_per_token_kib":24,"kv_note":"仅 12 层全注意力计 KV：2 × 256 × 2 × 12 × 2 B = 24 KiB/token；DeltaNet 层状态与序列长度无关。","ref_hw_24gb":"不可行（Q4 47 GB）","ref_hw_80gb":"FP8 80 GB 边缘；Q4 47 GB 舒适，256K 上下文仅需 6 GB KV","ref_hw_8x80gb":"BF16 高并发","estimated":true},"pricing":{"input_per_m":0.15,"output_per_m":1.5,"currency":"USD","source":"第三方托管常见价","as_of":"2025-12-20","note":"开源权重，社区托管价"},"links":{"official":"https://qwen.ai/blog?id=4074cca80393150c248e508aa62983f9cb7d27cd","hf":"https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking","github":"https://github.com/QwenLM/Qwen3-Next","paper":"https://qwen.ai/blog?id=4074cca80393150c248e508aa62983f9cb7d27cd"},"copy":{"one_liner":"线性注意力混合架构，长上下文吞吐比 235B 高 10 倍。","highlights":["Gated DeltaNet 3:1 混合，32K+ 上下文下解码吞吐是 Qwen3-32B 的 10 倍以上","训练成本仅 235B 的 1/10，推理分数接近 235B-Thinking","KV Cache 24 KiB/token，256K 上下文只要 6 GB"],"pitfalls":["新算子：需要 vLLM / SGLang 最新版 + 定制 kernel（flash-linear-attention、causal-conv1d），llama.cpp 支持滞后","3B 激活 + 线性注意力，精确回忆长文中的细节（needle 类任务）弱于全注意力模型","512 专家极端稀疏，量化后质量波动大于常规 MoE"],"logic_ability":"考试型推理接近 Qwen3-235B-Thinking（AIME 87.8%、GPQA 77.2%），且长上下文推理成本低得多。工程型推理与长文精确检索是弱项（线性注意力天然的记忆压缩）。思考不可关闭。适合「长输入、短输出、要有推理」的任务。","best_for":["超长上下文高吞吐推理","长文档批处理","架构研究（DeltaNet 落地样本）"],"not_for":["需要精确长文回忆","llama.cpp / 本地量化生态"]},"capability_notes":{"coding":"LiveCodeBench v6 68.7%（官方）。","reasoning":"GPQA Diamond 77.2%、HLE 13.2%（官方）。","math":"AIME 2025 87.8%（官方）。","agent":"BFCL-v3 72.0%（官方）。","chinese":"中文良好。"},"sheet":{"architecture_md":"**类型**：Hybrid（线性 + 全注意力）MoE，80B 总 / 3B 激活。\n\n- 48 层 = 12 × [3 × (Gated DeltaNet → MoE) + 1 × (Gated Attention → MoE)]\n- Gated DeltaNet：线性注意力，32 V 头 / 16 QK 头，head_dim 128，状态固定大小\n- Gated Attention：16 Q 头 / 2 KV 头，head_dim 256，输出门控，仅前 25% 维度用 RoPE\n- MoE：512 专家，激活 10 路由 + 1 共享，专家中间维 512\n- Zero-centered RMSNorm，MTP 多 token 预测\n- 上下文 262K，YaRN 可到 1M\n\n参考：Qwen3-Next 官方博客 / Gated DeltaNet 论文（arXiv 2412.06464）。","memory_md":"| 精度 | 权重大小 | 参考显存 |\n|---|---|---|\n| BF16 | ≈ 160 GB | 2×80GB |\n| FP8 | ≈ 80 GB | 80GB（边缘）|\n| Q4 | ≈ 47 GB（估） | 48GB / 80GB |\n\n**KV Cache**：仅 12 层全注意力 ≈ 24 KiB/token。256K 上下文 ≈ 6 GB。DeltaNet 层状态固定（每层 ~ 数 MB）。\n\n**参考配置**：\n- 24GB：不可行\n- 80GB：Q4/FP8 + 256K 上下文可行\n- 8×80GB：BF16 高并发","training_md":"- 预训练 15T token（Qwen3 语料子集）\n- 后训练：Thinking 版由 Qwen3-235B-Thinking 蒸馏 + RL\n- 思考默认且不可关（Instruct 版无思考）\n- 工具调用支持\n- 最大输出 32K","ecosystem_md":"- HF：Qwen/Qwen3-Next-80B-A3B-Thinking（+ FP8）\n- 引擎：vLLM ≥ 0.10.2、SGLang ≥ 0.5.2（需 flash-linear-attention）；llama.cpp 支持后续跟进\n- 微调：Transformers 主线支持，LoRA 可行\n- 中文文档：有","versions_md":"- 姊妹：Qwen3-Next-80B-A3B-Instruct\n- 官方称 Qwen3.5 将沿用该架构"},"ecosystem":{"engines":["vLLM","SGLang"],"finetune":"LoRA 可行（需新版 Transformers）","zh_docs":"有"},"complete":true,"superseded_by":"qwen3-8-27b"},{"aliases":["QwQ 32B","Qwen QwQ"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text","tools"],"reasoning_mode":"default-on","capability_notes":{"coding":"LiveCodeBench 63.4%（官方）。","math":"AIME 2024 79.5%（官方）。","agent":"BFCL 66.4%（官方）。","chinese":"中文推理好。"},"complete":false,"id":"qwq-32b","name":"QwQ-32B","name_zh":"通义千问 QwQ · 32B","vendor":"Alibaba","vendor_zh":"阿里巴巴","family":"QwQ","license":"Apache-2.0","license_commercial":true,"weights_url":"https://huggingface.co/Qwen/QwQ-32B","superseded_by":"qwen3-32b","released_at":"2025-03-05","architecture":{"type":"dense","total_params":"32.5B","total_params_b":32.5,"layers":64,"hidden_size":5120,"vocab_size":152064,"kv_heads":8,"head_dim":128,"attention":"GQA（40 Q 头 / 8 KV 头）","notes":"基于 Qwen2.5-32B 用 RL 训练的推理模型；131K 上下文（YaRN）；始终输出 <think>。","undisclosed":false},"context":{"max_tokens":131072,"display":"128K","max_output":32768},"memory":{"weight_gb":{"bf16":65,"q8":34.5,"q4":19.9},"estimated":false,"kv_per_token_kib":256,"kv_note":"256 KiB/token；思考链常达 10K+，KV 预留要大。","ref_hw_24gb":"Q4_K_M 19.9 GB 可加载，但长思考易 OOM，建议 IQ4_XS + 16K","ref_hw_80gb":"BF16 单卡，32K 上下文","ref_hw_8x80gb":"高并发"},"links":{"official":"https://qwenlm.github.io/blog/qwq-32b/","github":"https://github.com/QwenLM/Qwen2.5","hf":"https://huggingface.co/Qwen/QwQ-32B"},"copy":{"one_liner":"32B 追平 DeepSeek-R1 的推理模型，单卡可跑。","highlights":["官方基准与 671B 的 R1 相当（AIME24 79.5%、LCB 63.4%）","Apache-2.0，24GB 卡 Q4 可跑","RL 阶段加入工具调用，agent 能力比蒸馏版好"],"pitfalls":["无法关闭思考，简单问题也输出长链","思考经常上万 token，吞吐与显存压力大","采样参数敏感（需 temp 0.6、top-p 0.95），否则死循环"],"logic_ability":"考试型推理极强：AIME 2024 79.5%、LiveCodeBench 63.4%（官方）。数学与算法题可靠，但通用知识与写作弱于同尺寸 Qwen2.5，且常出现过度思考。","best_for":["本地数学 / 算法求解","推理蒸馏数据生成"],"not_for":["低延迟对话","知识问答 / 写作"]},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Ollama","MLX"],"finetune":"LoRA 可行，建议保留思考格式","zh_docs":"有"}},{"id":"step-3-7-flash","name":"Step 3.7 Flash","name_zh":"阶跃星辰 Step 3.7 Flash","aliases":["Step-3.7-Flash","stepfun-ai/Step-3.7-Flash","step-3.7-flash"],"vendor":"StepFun","vendor_zh":"阶跃星辰","family":"Step","license":"Apache-2.0","license_commercial":true,"openness":"open-weights","weights_available":true,"weights_url":"https://huggingface.co/stepfun-ai/Step-3.7-Flash","status":"current","released_at":"2026-05-29","updated_at":"2026-08-28","modalities":["text","image","video","tools"],"reasoning_mode":"optional","architecture":{"type":"moe","total_params":"198B","active_params":"11B","total_params_b":198,"active_params_b":11,"experts":288,"active_experts":8,"shared_expert":true,"layers":45,"hidden_size":4096,"vocab_size":128896,"kv_heads":8,"head_dim":128,"attention":"混合：全注意力 / 512 滑窗交替，GQA（64 Q / 8 KV）","notes":"196B 语言骨干 + 1.8B 视觉编码器；42 层 MoE，共享专家维 1280。推理强度 low / medium / high。","undisclosed":false},"context":{"max_tokens":262144,"display":"256K"},"memory":{"weight_gb":{"bf16":396,"fp8":198,"q4":115,"q8":213.1},"estimated":true,"ref_hw_80gb":"不可行（官方最低 120GB 统一内存，如 Mac Studio / DGX Station）","ref_hw_8x80gb":"FP8 多卡"},"pricing":{"input_per_m":0.2,"output_per_m":1.15,"currency":"USD","source":"StepFun 开放平台定价","as_of":"2026-08-28","note":"缓存命中输入 $0.04"},"links":{"official":"https://huggingface.co/stepfun-ai/Step-3.7-Flash","hf":"https://huggingface.co/stepfun-ai/Step-3.7-Flash"},"variants":[{"kind":"fp8","publisher":"stepfun-ai","repo":"stepfun-ai/Step-3.7-Flash-FP8","url":"https://huggingface.co/stepfun-ai/Step-3.7-Flash-FP8","note":"官方 FP8","sizes":{"fp8":212.5}},{"kind":"nvfp4","publisher":"stepfun-ai","repo":"stepfun-ai/Step-3.7-Flash-NVFP4","url":"https://huggingface.co/stepfun-ai/Step-3.7-Flash-NVFP4","note":"官方 NVFP4"},{"kind":"gguf","publisher":"stepfun-ai","repo":"stepfun-ai/Step-3.7-Flash-GGUF","url":"https://huggingface.co/stepfun-ai/Step-3.7-Flash-GGUF","note":"官方 GGUF","sizes":{"q8":213.1,"bf16":401}},{"kind":"gguf","publisher":"unsloth","repo":"unsloth/Step-3.7-Flash-GGUF","url":"https://huggingface.co/unsloth/Step-3.7-Flash-GGUF","sizes":{"q8":209.4,"bf16":394}},{"kind":"gguf","publisher":"bartowski","repo":"bartowski/Step-3.7-Flash-GGUF","url":"https://huggingface.co/bartowski/Step-3.7-Flash-GGUF","sizes":{"q4":121.6,"q5":142,"q6":171.8,"q8":212}},{"kind":"awq","publisher":"cyankiwi","repo":"cyankiwi/Step-3.7-Flash-AWQ-INT4","url":"https://huggingface.co/cyankiwi/Step-3.7-Flash-AWQ-INT4"}],"copy":{"one_liner":"阶跃 198B 视觉 MoE，11B 激活，图 / 视频 / 编程 agent。","highlights":["Terminal-Bench 2.1 59.5、SWE-bench Pro 56.3、HLE 带工具 48.1（官方）","Apache-2.0，图片 + 视频理解 + 编程 agent 一体","官方放 BF16 / FP8 / NVFP4 / GGUF，Mac 128GB 可跑"],"pitfalls":["198B 总参，80GB 单卡跑不了","官方未给 SWE-bench Verified / GPQA 等常见指标","StepFun 海外生态与工具链较弱"],"logic_ability":"工程型推理不错（TB 2.1 59.5），考试型推理官方主要给 HLE 带工具 48.1，纯推理指标缺。","best_for":["视觉 + 搜索 agent","Mac / 大内存工作站本地"],"not_for":["24GB / 80GB 单卡","纯数学竞赛"]},"capability_notes":{"coding":"Terminal-Bench 2.1 59.5、SWE-bench Pro 56.3（官方）。","reasoning":"HLE 带工具 48.1（官方）。","multimodal":"图片 + 视频理解，V* 95.3（官方）。","chinese":"强，国产模型。"},"ecosystem":{"engines":["vLLM","SGLang","llama.cpp","Transformers"],"zh_docs":"有"},"complete":false,"runtime":{"tok_s":86,"latency_s":2.78,"source":"Artificial Analysis 官方 API 中位数（2026-08-28）"},"i18n":{"en":{"name_zh":"StepFun Step 3.7 Flash","one_liner":"StepFun's 198B vision MoE, 11B active; image / video / coding agents.","highlights":["Terminal-Bench 2.1 59.5, SWE-bench Pro 56.3, HLE with tools 48.1 (official)","Apache-2.0; image + video understanding + coding agent in one","Official BF16 / FP8 / NVFP4 / GGUF releases; runs on a 128GB Mac"],"pitfalls":["198B total parameters; will not run on an 80GB single GPU","No official common metrics such as SWE-bench Verified / GPQA","StepFun's overseas ecosystem and tooling are weak"],"logic_ability":"Decent engineering reasoning (TB 2.1 59.5); for exam-style reasoning the official figure is mainly HLE with tools 48.1, pure reasoning metrics missing.","best_for":["Vision + search agents","Local use on Macs / large-memory workstations"],"not_for":["24GB / 80GB single GPUs","Pure math competitions"],"capability_notes":{"coding":"Terminal-Bench 2.1 59.5, SWE-bench Pro 56.3 (official).","reasoning":"HLE with tools 48.1 (official).","multimodal":"Image + video understanding, V* 95.3 (official).","chinese":"Strong; Chinese domestic model."}},"ja":{"name_zh":"StepFun Step 3.7 Flash","one_liner":"StepFun 198B 視覚 MoE、11B 活性。画像/動画/コード。","highlights":["Terminal-Bench 2.1 59.5、SWE-bench Pro 56.3、HLE ツールあり 48.1（公式）","Apache-2.0、画像 + 動画理解 + コーディングエージェントを一体化","公式が BF16 / FP8 / NVFP4 / GGUF を公開、128GB の Mac で動作"],"pitfalls":["総パラメータ 198B、80GB 単一 GPU では動かない","SWE-bench Verified / GPQA などの定番指標は公式未開示","StepFun の海外エコシステムとツールチェーンが弱い"],"logic_ability":"エンジニアリング型推論は良好（TB 2.1 59.5）。試験型推論は公式が主に HLE ツールあり 48.1 のみで、純粋な推論指標が欠落。","best_for":["ビジョン + 検索エージェント","Mac / 大容量メモリのワークステーションでのローカル利用"],"not_for":["24GB / 80GB 単一 GPU","純粋な数学コンテスト"],"capability_notes":{"coding":"Terminal-Bench 2.1 59.5、SWE-bench Pro 56.3（公式）。","reasoning":"HLE ツールあり 48.1（公式）。","multimodal":"画像 + 動画理解、V* 95.3（公式）。","chinese":"強い。中国産モデル。"}}}},{"aliases":["Yi-1.5-34B-Chat"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"knowledge":"MMLU 76.8%（官方基座）。","math":"GSM8K 78.5%（官方基座）。","chinese":"中文好。"},"complete":false,"id":"yi-1-5-34b","name":"Yi-1.5-34B","name_zh":"零一万物 Yi-1.5 · 34B","vendor":"01.AI","vendor_zh":"零一万物","family":"Yi-1.5","license":"Apache-2.0","license_commercial":true,"weights_url":"https://huggingface.co/01-ai/Yi-1.5-34B-Chat","released_at":"2024-05-13","architecture":{"type":"dense","total_params":"34.4B","total_params_b":34.4,"layers":60,"hidden_size":7168,"vocab_size":64000,"kv_heads":8,"head_dim":128,"attention":"GQA（56 Q 头 / 8 KV 头）","notes":"Yi 续训 500B token；4K（另有 16K / 32K 版）。","undisclosed":false},"context":{"max_tokens":4096,"display":"4K（16K/32K 版另发）"},"memory":{"weight_gb":{"bf16":68.8,"q8":36.5,"q4":20.7},"estimated":false,"kv_per_token_kib":240,"kv_note":"240 KiB/token。","ref_hw_24gb":"Q4_K_M 20.7 GB 可跑","ref_hw_80gb":"BF16 单卡","ref_hw_8x80gb":"过剩"},"links":{"official":"https://01.ai/","github":"https://github.com/01-ai/Yi-1.5","hf":"https://huggingface.co/01-ai/Yi-1.5-34B-Chat"},"copy":{"one_liner":"Yi 的续训升级版，代码与数学补课。","highlights":["Apache-2.0","官方称接近 Llama 3 70B 的多数基准","24GB 卡 Q4 可跑"],"pitfalls":["默认 4K 上下文","零一万物 2024 年底转向闭源与 API，系列停更","被 Qwen2.5-32B 全面超过"],"logic_ability":"34B 中游：MMLU 76.8%、GSM8K 78.5%（官方基座）。","best_for":["中英双语微调"],"not_for":["长文本","新项目"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama"],"finetune":"LoRA 成熟","zh_docs":"有"}},{"aliases":["Yi-34B-Chat","Yi-34B-200K"],"openness":"open-weights","weights_available":true,"status":"superseded","updated_at":"2026-08-28","modalities":["text"],"reasoning_mode":"none","capability_notes":{"knowledge":"MMLU 76.3%（官方基座）。","chinese":"C-Eval 81.4%（官方）。"},"complete":false,"id":"yi-34b","name":"Yi-34B","name_zh":"零一万物 Yi · 34B","vendor":"01.AI","vendor_zh":"零一万物","family":"Yi","license":"Apache-2.0（2024-03 起，原为 Yi License）","license_commercial":true,"weights_url":"https://huggingface.co/01-ai/Yi-34B-Chat","superseded_by":"yi-1-5-34b","released_at":"2023-11-02","architecture":{"type":"dense","total_params":"34.4B","total_params_b":34.4,"layers":60,"hidden_size":7168,"vocab_size":64000,"kv_heads":8,"head_dim":128,"attention":"GQA（56 Q 头 / 8 KV 头）","notes":"Llama 结构，3T 中英 token；Chat 版 4K，另有 200K 长文本版。","undisclosed":false},"context":{"max_tokens":4096,"display":"4K（200K 版另发）"},"memory":{"weight_gb":{"bf16":68.8,"q8":36.5,"q4":20.7},"estimated":false,"kv_per_token_kib":240,"kv_note":"240 KiB/token。","ref_hw_24gb":"Q4_K_M 20.7 GB，24GB 卡可跑短上下文","ref_hw_80gb":"BF16 68.8 GB 单卡","ref_hw_8x80gb":"过剩"},"links":{"official":"https://01.ai/","github":"https://github.com/01-ai/Yi","paper":"https://arxiv.org/abs/2403.04652","hf":"https://huggingface.co/01-ai/Yi-34B-Chat"},"copy":{"one_liner":"2023 年末登顶 HF 开源榜的中英双语 34B。","highlights":["发布时 Open LLM Leaderboard 第一，超 Llama 2 70B","200K 上下文版当年罕见","2024-03 改为 Apache-2.0"],"pitfalls":["Chat 版仅 4K 上下文","早期因张量命名照搬 Llama 引发争议","被 Yi-1.5 与 Qwen 系列超过"],"logic_ability":"2023 年末开源一线：MMLU 76.3%、C-Eval 81.4%（官方基座）。无 CoT 训练。","best_for":["历史对照","中英双语微调实验"],"not_for":["新项目"]},"ecosystem":{"engines":["vLLM","llama.cpp","Ollama"],"finetune":"LoRA 成熟（Llama 结构）","zh_docs":"有"}}],"scores":[{"model_id":"gemini-3-pro","key":"arena_text","value":1501,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"gemini-3-pro","key":"swe_verified","value":76.2,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-3-pro","key":"gpqa_diamond","value":91.9,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-3-pro","key":"hle","value":37.5,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-3-pro","key":"aime_2025","value":95,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-3-pro","key":"tau2_bench","value":85.4,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-3-pro","key":"terminal_bench","value":54.2,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-3-pro","key":"mmmu","value":81,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-opus-4-5","key":"arena_text","value":1470,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"claude-opus-4-5","key":"swe_verified","value":80.9,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-opus-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-opus-4-5","key":"gpqa_diamond","value":87,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-opus-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-opus-4-5","key":"tau2_bench","value":88.9,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-opus-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-opus-4-5","key":"terminal_bench","value":59.3,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-opus-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-opus-4-5","key":"mmmu","value":80.7,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-opus-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-5-1","key":"arena_text","value":1461,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"gpt-5-1","key":"swe_verified","value":76.3,"unit":"percent","source":"OpenAI 发布公告","source_url":"https://openai.com/index/gpt-5-1/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-5-1","key":"gpqa_diamond","value":88.1,"unit":"percent","source":"OpenAI 发布公告","source_url":"https://openai.com/index/gpt-5-1/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-5-1","key":"hle","value":26.5,"unit":"percent","source":"OpenAI 发布公告","source_url":"https://openai.com/index/gpt-5-1/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-5","key":"arena_text","value":1440,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"gpt-5","key":"swe_verified","value":74.9,"unit":"percent","source":"OpenAI GPT-5 发布公告","source_url":"https://openai.com/index/introducing-gpt-5/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-5","key":"gpqa_diamond","value":85.7,"unit":"percent","source":"OpenAI GPT-5 发布公告","source_url":"https://openai.com/index/introducing-gpt-5/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-5","key":"hle","value":24.8,"unit":"percent","source":"OpenAI GPT-5 发布公告","source_url":"https://openai.com/index/introducing-gpt-5/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-5","key":"aime_2025","value":94.6,"unit":"percent","source":"OpenAI GPT-5 发布公告","source_url":"https://openai.com/index/introducing-gpt-5/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-5","key":"tau2_bench","value":81.1,"unit":"percent","source":"OpenAI GPT-5 发布公告","source_url":"https://openai.com/index/introducing-gpt-5/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-5","key":"mmmu","value":84.2,"unit":"percent","source":"OpenAI GPT-5 发布公告","source_url":"https://openai.com/index/introducing-gpt-5/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-3-flash","key":"arena_text","value":1470,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"gemini-3-flash","key":"swe_verified","value":78,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/flash/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-3-flash","key":"gpqa_diamond","value":90.4,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/flash/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-3-flash","key":"hle","value":33.7,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/flash/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-3-flash","key":"terminal_bench","value":51.5,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/flash/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-3-flash","key":"mmmu","value":81.2,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/flash/","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-sonnet-4-5","key":"arena_text","value":1450,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"claude-sonnet-4-5","key":"swe_verified","value":77.2,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-sonnet-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-sonnet-4-5","key":"gpqa_diamond","value":83.4,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-sonnet-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-sonnet-4-5","key":"aime_2025","value":87,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-sonnet-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-sonnet-4-5","key":"tau2_bench","value":84.7,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-sonnet-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-sonnet-4-5","key":"terminal_bench","value":50,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-sonnet-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-sonnet-4-5","key":"mmmu","value":77.8,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-sonnet-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"grok-4","key":"arena_text","value":1430,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"grok-4","key":"swe_verified","value":75,"unit":"percent","source":"xAI 发布公告","source_url":"https://x.ai/news/grok-4","as_of":"2025-12-20","evidence":"official"},{"model_id":"grok-4","key":"gpqa_diamond","value":87.5,"unit":"percent","source":"xAI 发布公告","source_url":"https://x.ai/news/grok-4","as_of":"2025-12-20","evidence":"official"},{"model_id":"grok-4","key":"hle","value":25.4,"unit":"percent","source":"xAI 发布公告","source_url":"https://x.ai/news/grok-4","as_of":"2025-12-20","evidence":"official"},{"model_id":"grok-4","key":"aime_2025","value":91.7,"unit":"percent","source":"xAI 发布公告","source_url":"https://x.ai/news/grok-4","as_of":"2025-12-20","evidence":"official"},{"model_id":"grok-4-1-fast","key":"arena_text","value":1440,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"grok-4-1-fast","key":"hle","value":17.6,"unit":"percent","source":"xAI 发布公告","source_url":"https://x.ai/news/grok-4-1-fast","as_of":"2025-12-20","evidence":"official"},{"model_id":"grok-4-1-fast","key":"tau2_bench","value":93,"unit":"percent","source":"xAI 发布公告","source_url":"https://x.ai/news/grok-4-1-fast","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-opus-4-1","key":"arena_text","value":1447,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"claude-opus-4-1","key":"swe_verified","value":74.5,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-opus-4-1","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-opus-4-1","key":"gpqa_diamond","value":80.9,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-opus-4-1","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-opus-4-1","key":"aime_2025","value":78,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-opus-4-1","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-opus-4-1","key":"terminal_bench","value":43.3,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-opus-4-1","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-opus-4-1","key":"mmmu","value":77.1,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-opus-4-1","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-haiku-4-5","key":"arena_text","value":1400,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"swe_verified","value":73.3,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-haiku-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-haiku-4-5","key":"gpqa_diamond","value":73,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-haiku-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-haiku-4-5","key":"aime_2025","value":80.6,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-haiku-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-haiku-4-5","key":"tau2_bench","value":79.6,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-haiku-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-haiku-4-5","key":"terminal_bench","value":41,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-haiku-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-haiku-4-5","key":"mmmu","value":70,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-haiku-4-5","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-2-5-pro","key":"arena_text","value":1451,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"gemini-2-5-pro","key":"swe_verified","value":63.8,"unit":"percent","source":"Google 模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-2-5-pro","key":"livecodebench","value":69,"unit":"percent","source":"Google 模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-2-5-pro","key":"gpqa_diamond","value":86.4,"unit":"percent","source":"Google 模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-2-5-pro","key":"hle","value":21.6,"unit":"percent","source":"Google 模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-2-5-pro","key":"aime_2025","value":88,"unit":"percent","source":"Google 模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-2-5-pro","key":"mmmu","value":82,"unit":"percent","source":"Google 模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-2-5-flash","key":"arena_text","value":1400,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"gemini-2-5-flash","key":"swe_verified","value":60.4,"unit":"percent","source":"Google 模型页","source_url":"https://deepmind.google/models/gemini/flash/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-2-5-flash","key":"gpqa_diamond","value":82.8,"unit":"percent","source":"Google 模型页","source_url":"https://deepmind.google/models/gemini/flash/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-2-5-flash","key":"hle","value":11,"unit":"percent","source":"Google 模型页","source_url":"https://deepmind.google/models/gemini/flash/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-2-5-flash","key":"aime_2025","value":72,"unit":"percent","source":"Google 模型页","source_url":"https://deepmind.google/models/gemini/flash/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemini-2-5-flash","key":"mmmu","value":79.7,"unit":"percent","source":"Google 模型页","source_url":"https://deepmind.google/models/gemini/flash/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-5-mini","key":"arena_text","value":1400,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"gpt-5-mini","key":"gpqa_diamond","value":82.3,"unit":"percent","source":"OpenAI 模型页","source_url":"https://platform.openai.com/docs/models/gpt-5-mini","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-5-mini","key":"hle","value":16.7,"unit":"percent","source":"OpenAI 模型页","source_url":"https://platform.openai.com/docs/models/gpt-5-mini","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-5-mini","key":"aime_2025","value":91.1,"unit":"percent","source":"OpenAI 模型页","source_url":"https://platform.openai.com/docs/models/gpt-5-mini","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-5-mini","key":"mmmu","value":81.6,"unit":"percent","source":"OpenAI 模型页","source_url":"https://platform.openai.com/docs/models/gpt-5-mini","as_of":"2025-12-20","evidence":"official"},{"model_id":"mistral-medium-3-1","key":"arena_text","value":1370,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"qwen3-max","key":"arena_text","value":1428,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"qwen3-max","key":"swe_verified","value":69.6,"unit":"percent","source":"Qwen 官方博客","source_url":"https://qwen.ai/blog?id=qwen3-max","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-max","key":"tau2_bench","value":74.8,"unit":"percent","source":"Qwen 官方博客","source_url":"https://qwen.ai/blog?id=qwen3-max","as_of":"2025-12-20","evidence":"official"},{"model_id":"deepseek-v3-2","key":"arena_text","value":1424,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"deepseek-v3-2","key":"swe_verified","value":73.1,"unit":"percent","source":"DeepSeek-V3.2 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.2","as_of":"2025-12-20","evidence":"official"},{"model_id":"deepseek-v3-2","key":"gpqa_diamond","value":82.4,"unit":"percent","source":"DeepSeek-V3.2 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.2","as_of":"2025-12-20","evidence":"official"},{"model_id":"deepseek-v3-2","key":"hle","value":25.1,"unit":"percent","source":"DeepSeek-V3.2 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.2","as_of":"2025-12-20","evidence":"official"},{"model_id":"deepseek-v3-2","key":"aime_2025","value":93.1,"unit":"percent","source":"DeepSeek-V3.2 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.2","as_of":"2025-12-20","evidence":"official"},{"model_id":"deepseek-v3-2","key":"tau2_bench","value":80.3,"unit":"percent","source":"DeepSeek-V3.2 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.2","as_of":"2025-12-20","evidence":"official"},{"model_id":"deepseek-v3-2","key":"terminal_bench","value":37.7,"unit":"percent","source":"DeepSeek-V3.2 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.2","as_of":"2025-12-20","evidence":"official"},{"model_id":"kimi-k2-thinking","key":"arena_text","value":1436,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"kimi-k2-thinking","key":"swe_verified","value":71.3,"unit":"percent","source":"Moonshot 发布页","source_url":"https://moonshotai.github.io/Kimi-K2/thinking.html","as_of":"2025-12-20","evidence":"official"},{"model_id":"kimi-k2-thinking","key":"livecodebench","value":83.1,"unit":"percent","source":"Moonshot 发布页","source_url":"https://moonshotai.github.io/Kimi-K2/thinking.html","as_of":"2025-12-20","evidence":"official"},{"model_id":"kimi-k2-thinking","key":"gpqa_diamond","value":84.5,"unit":"percent","source":"Moonshot 发布页","source_url":"https://moonshotai.github.io/Kimi-K2/thinking.html","as_of":"2025-12-20","evidence":"official"},{"model_id":"kimi-k2-thinking","key":"hle","value":23.9,"unit":"percent","source":"Moonshot 发布页","source_url":"https://moonshotai.github.io/Kimi-K2/thinking.html","as_of":"2025-12-20","evidence":"official"},{"model_id":"kimi-k2-thinking","key":"aime_2025","value":94.5,"unit":"percent","source":"Moonshot 发布页","source_url":"https://moonshotai.github.io/Kimi-K2/thinking.html","as_of":"2025-12-20","evidence":"official"},{"model_id":"kimi-k2-thinking","key":"tau2_bench","value":74.3,"unit":"percent","source":"Moonshot 发布页","source_url":"https://moonshotai.github.io/Kimi-K2/thinking.html","as_of":"2025-12-20","evidence":"official"},{"model_id":"kimi-k2-thinking","key":"terminal_bench","value":47.1,"unit":"percent","source":"Moonshot 发布页","source_url":"https://moonshotai.github.io/Kimi-K2/thinking.html","as_of":"2025-12-20","evidence":"official"},{"model_id":"kimi-k2-0905","key":"arena_text","value":1420,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"kimi-k2-0905","key":"swe_verified","value":69.2,"unit":"percent","source":"Kimi K2 模型卡","source_url":"https://huggingface.co/moonshotai/Kimi-K2-Instruct-0905","as_of":"2025-12-20","evidence":"official"},{"model_id":"kimi-k2-0905","key":"gpqa_diamond","value":75.1,"unit":"percent","source":"Kimi K2 模型卡","source_url":"https://huggingface.co/moonshotai/Kimi-K2-Instruct-0905","as_of":"2025-12-20","evidence":"official"},{"model_id":"kimi-k2-0905","key":"tau2_bench","value":70,"unit":"percent","source":"Kimi K2 模型卡","source_url":"https://huggingface.co/moonshotai/Kimi-K2-Instruct-0905","as_of":"2025-12-20","evidence":"official"},{"model_id":"kimi-k2-0905","key":"terminal_bench","value":44.5,"unit":"percent","source":"Kimi K2 模型卡","source_url":"https://huggingface.co/moonshotai/Kimi-K2-Instruct-0905","as_of":"2025-12-20","evidence":"official"},{"model_id":"glm-4-6","key":"arena_text","value":1420,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"glm-4-6","key":"swe_verified","value":68,"unit":"percent","source":"GLM-4.6 模型卡","source_url":"https://huggingface.co/zai-org/GLM-4.6","as_of":"2025-12-20","evidence":"official"},{"model_id":"glm-4-6","key":"livecodebench","value":82.8,"unit":"percent","source":"GLM-4.6 模型卡","source_url":"https://huggingface.co/zai-org/GLM-4.6","as_of":"2025-12-20","evidence":"official"},{"model_id":"glm-4-6","key":"gpqa_diamond","value":81,"unit":"percent","source":"GLM-4.6 模型卡","source_url":"https://huggingface.co/zai-org/GLM-4.6","as_of":"2025-12-20","evidence":"official"},{"model_id":"glm-4-6","key":"aime_2025","value":93.9,"unit":"percent","source":"GLM-4.6 模型卡","source_url":"https://huggingface.co/zai-org/GLM-4.6","as_of":"2025-12-20","evidence":"official"},{"model_id":"glm-4-6","key":"tau2_bench","value":75.9,"unit":"percent","source":"GLM-4.6 模型卡","source_url":"https://huggingface.co/zai-org/GLM-4.6","as_of":"2025-12-20","evidence":"official"},{"model_id":"glm-4-6","key":"terminal_bench","value":40.5,"unit":"percent","source":"GLM-4.6 模型卡","source_url":"https://huggingface.co/zai-org/GLM-4.6","as_of":"2025-12-20","evidence":"official"},{"model_id":"glm-4-5-air","key":"arena_text","value":1380,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"glm-4-5-air","key":"swe_verified","value":57.6,"unit":"percent","source":"GLM-4.5 模型卡","source_url":"https://huggingface.co/zai-org/GLM-4.5-Air","as_of":"2025-12-20","evidence":"official"},{"model_id":"glm-4-5-air","key":"gpqa_diamond","value":75,"unit":"percent","source":"GLM-4.5 模型卡","source_url":"https://huggingface.co/zai-org/GLM-4.5-Air","as_of":"2025-12-20","evidence":"official"},{"model_id":"glm-4-5-air","key":"aime_2025","value":89.4,"unit":"percent","source":"GLM-4.5 模型卡","source_url":"https://huggingface.co/zai-org/GLM-4.5-Air","as_of":"2025-12-20","evidence":"official"},{"model_id":"glm-4-5-air","key":"tau2_bench","value":68,"unit":"percent","source":"GLM-4.5 模型卡","source_url":"https://huggingface.co/zai-org/GLM-4.5-Air","as_of":"2025-12-20","evidence":"official"},{"model_id":"minimax-m2","key":"arena_text","value":1395,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"minimax-m2","key":"swe_verified","value":69.4,"unit":"percent","source":"MiniMax-M2 模型卡","source_url":"https://huggingface.co/MiniMaxAI/MiniMax-M2","as_of":"2025-12-20","evidence":"official"},{"model_id":"minimax-m2","key":"gpqa_diamond","value":78,"unit":"percent","source":"MiniMax-M2 模型卡","source_url":"https://huggingface.co/MiniMaxAI/MiniMax-M2","as_of":"2025-12-20","evidence":"official"},{"model_id":"minimax-m2","key":"aime_2025","value":78,"unit":"percent","source":"MiniMax-M2 模型卡","source_url":"https://huggingface.co/MiniMaxAI/MiniMax-M2","as_of":"2025-12-20","evidence":"official"},{"model_id":"minimax-m2","key":"tau2_bench","value":77.2,"unit":"percent","source":"MiniMax-M2 模型卡","source_url":"https://huggingface.co/MiniMaxAI/MiniMax-M2","as_of":"2025-12-20","evidence":"official"},{"model_id":"minimax-m2","key":"terminal_bench","value":46.3,"unit":"percent","source":"MiniMax-M2 模型卡","source_url":"https://huggingface.co/MiniMaxAI/MiniMax-M2","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-235b-a22b-thinking-2507","key":"arena_text","value":1420,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"qwen3-235b-a22b-thinking-2507","key":"livecodebench","value":74.1,"unit":"percent","source":"Qwen3-235B-A22B-Thinking-2507 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3-235B-A22B-Thinking-2507","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-235b-a22b-thinking-2507","key":"gpqa_diamond","value":81.1,"unit":"percent","source":"Qwen3-235B-A22B-Thinking-2507 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3-235B-A22B-Thinking-2507","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-235b-a22b-thinking-2507","key":"hle","value":18.2,"unit":"percent","source":"Qwen3-235B-A22B-Thinking-2507 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3-235B-A22B-Thinking-2507","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-235b-a22b-thinking-2507","key":"aime_2025","value":92.3,"unit":"percent","source":"Qwen3-235B-A22B-Thinking-2507 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3-235B-A22B-Thinking-2507","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-235b-a22b-thinking-2507","key":"tau2_bench","value":70,"unit":"percent","source":"Qwen3-235B-A22B-Thinking-2507 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3-235B-A22B-Thinking-2507","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-next-80b-a3b-thinking","key":"arena_text","value":1390,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"qwen3-next-80b-a3b-thinking","key":"livecodebench","value":68.7,"unit":"percent","source":"Qwen3-Next 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-next-80b-a3b-thinking","key":"gpqa_diamond","value":77.2,"unit":"percent","source":"Qwen3-Next 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-next-80b-a3b-thinking","key":"hle","value":13.2,"unit":"percent","source":"Qwen3-Next 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-next-80b-a3b-thinking","key":"aime_2025","value":87.8,"unit":"percent","source":"Qwen3-Next 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-next-80b-a3b-thinking","key":"tau2_bench","value":60,"unit":"percent","source":"Qwen3-Next 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-coder-480b-a35b","key":"arena_text","value":1400,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"qwen3-coder-480b-a35b","key":"swe_verified","value":69.6,"unit":"percent","source":"Qwen3-Coder 博客","source_url":"https://qwenlm.github.io/blog/qwen3-coder/","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-coder-480b-a35b","key":"terminal_bench","value":37.5,"unit":"percent","source":"Qwen3-Coder 博客","source_url":"https://qwenlm.github.io/blog/qwen3-coder/","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-32b","key":"arena_text","value":1360,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"qwen3-32b","key":"livecodebench","value":65.7,"unit":"percent","source":"Qwen3 技术报告","source_url":"https://arxiv.org/abs/2505.09388","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-32b","key":"gpqa_diamond","value":68.4,"unit":"percent","source":"Qwen3 技术报告","source_url":"https://arxiv.org/abs/2505.09388","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-32b","key":"aime_2025","value":72.9,"unit":"percent","source":"Qwen3 技术报告","source_url":"https://arxiv.org/abs/2505.09388","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-30b-a3b-thinking-2507","key":"arena_text","value":1370,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"qwen3-30b-a3b-thinking-2507","key":"livecodebench","value":66,"unit":"percent","source":"Qwen3-30B-A3B-Thinking-2507 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3-30B-A3B-Thinking-2507","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-30b-a3b-thinking-2507","key":"gpqa_diamond","value":73.4,"unit":"percent","source":"Qwen3-30B-A3B-Thinking-2507 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3-30B-A3B-Thinking-2507","as_of":"2025-12-20","evidence":"official"},{"model_id":"qwen3-30b-a3b-thinking-2507","key":"aime_2025","value":85,"unit":"percent","source":"Qwen3-30B-A3B-Thinking-2507 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3-30B-A3B-Thinking-2507","as_of":"2025-12-20","evidence":"official"},{"model_id":"deepseek-r1-0528","key":"arena_text","value":1415,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"deepseek-r1-0528","key":"swe_verified","value":57.6,"unit":"percent","source":"DeepSeek-R1-0528 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-R1-0528","as_of":"2025-12-20","evidence":"official"},{"model_id":"deepseek-r1-0528","key":"livecodebench","value":73.3,"unit":"percent","source":"DeepSeek-R1-0528 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-R1-0528","as_of":"2025-12-20","evidence":"official"},{"model_id":"deepseek-r1-0528","key":"gpqa_diamond","value":81,"unit":"percent","source":"DeepSeek-R1-0528 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-R1-0528","as_of":"2025-12-20","evidence":"official"},{"model_id":"deepseek-r1-0528","key":"hle","value":17.7,"unit":"percent","source":"DeepSeek-R1-0528 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-R1-0528","as_of":"2025-12-20","evidence":"official"},{"model_id":"deepseek-r1-0528","key":"aime_2025","value":87.5,"unit":"percent","source":"DeepSeek-R1-0528 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-R1-0528","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-oss-120b","key":"arena_text","value":1340,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"gpt-oss-120b","key":"swe_verified","value":62.4,"unit":"percent","source":"gpt-oss 模型卡","source_url":"https://openai.com/index/introducing-gpt-oss/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-oss-120b","key":"gpqa_diamond","value":80.1,"unit":"percent","source":"gpt-oss 模型卡","source_url":"https://openai.com/index/introducing-gpt-oss/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-oss-120b","key":"hle","value":19,"unit":"percent","source":"gpt-oss 模型卡","source_url":"https://openai.com/index/introducing-gpt-oss/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-oss-120b","key":"aime_2025","value":92.5,"unit":"percent","source":"gpt-oss 模型卡","source_url":"https://openai.com/index/introducing-gpt-oss/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-oss-120b","key":"tau2_bench","value":67.8,"unit":"percent","source":"gpt-oss 模型卡","source_url":"https://openai.com/index/introducing-gpt-oss/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-oss-20b","key":"arena_text","value":1300,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"gpt-oss-20b","key":"swe_verified","value":60.7,"unit":"percent","source":"gpt-oss 模型卡","source_url":"https://openai.com/index/introducing-gpt-oss/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-oss-20b","key":"gpqa_diamond","value":71.5,"unit":"percent","source":"gpt-oss 模型卡","source_url":"https://openai.com/index/introducing-gpt-oss/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-oss-20b","key":"hle","value":10.9,"unit":"percent","source":"gpt-oss 模型卡","source_url":"https://openai.com/index/introducing-gpt-oss/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gpt-oss-20b","key":"aime_2025","value":91.7,"unit":"percent","source":"gpt-oss 模型卡","source_url":"https://openai.com/index/introducing-gpt-oss/","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemma-3-27b","key":"arena_text","value":1338,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"gemma-3-27b","key":"livecodebench","value":29.7,"unit":"percent","source":"Gemma 3 技术报告","source_url":"https://arxiv.org/abs/2503.19786","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemma-3-27b","key":"gpqa_diamond","value":42.4,"unit":"percent","source":"Gemma 3 技术报告","source_url":"https://arxiv.org/abs/2503.19786","as_of":"2025-12-20","evidence":"official"},{"model_id":"gemma-3-27b","key":"mmmu","value":64.9,"unit":"percent","source":"Gemma 3 技术报告","source_url":"https://arxiv.org/abs/2503.19786","as_of":"2025-12-20","evidence":"official"},{"model_id":"llama-4-maverick","key":"arena_text","value":1330,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"llama-4-maverick","key":"livecodebench","value":43.4,"unit":"percent","source":"Llama 4 模型卡","source_url":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/","as_of":"2025-12-20","evidence":"official"},{"model_id":"llama-4-maverick","key":"gpqa_diamond","value":69.8,"unit":"percent","source":"Llama 4 模型卡","source_url":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/","as_of":"2025-12-20","evidence":"official"},{"model_id":"llama-4-maverick","key":"mmmu","value":73.4,"unit":"percent","source":"Llama 4 模型卡","source_url":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/","as_of":"2025-12-20","evidence":"official"},{"model_id":"llama-4-scout","key":"arena_text","value":1300,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"llama-4-scout","key":"livecodebench","value":32.8,"unit":"percent","source":"Llama 4 模型卡","source_url":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/","as_of":"2025-12-20","evidence":"official"},{"model_id":"llama-4-scout","key":"gpqa_diamond","value":57.2,"unit":"percent","source":"Llama 4 模型卡","source_url":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/","as_of":"2025-12-20","evidence":"official"},{"model_id":"llama-4-scout","key":"mmmu","value":69.4,"unit":"percent","source":"Llama 4 模型卡","source_url":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/","as_of":"2025-12-20","evidence":"official"},{"model_id":"devstral-small-2","key":"swe_verified","value":68,"unit":"percent","source":"Mistral 发布公告","source_url":"https://mistral.ai/news/devstral-2-vibe-cli","as_of":"2025-12-20","evidence":"official"},{"model_id":"mistral-small-3-2","key":"arena_text","value":1320,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"nemotron-3-nano-30b-a3b","key":"livecodebench","value":68.3,"unit":"percent","source":"NVIDIA 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16","as_of":"2025-12-20","evidence":"official"},{"model_id":"nemotron-3-nano-30b-a3b","key":"gpqa_diamond","value":73.7,"unit":"percent","source":"NVIDIA 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16","as_of":"2025-12-20","evidence":"official"},{"model_id":"nemotron-3-nano-30b-a3b","key":"hle","value":10.6,"unit":"percent","source":"NVIDIA 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16","as_of":"2025-12-20","evidence":"official"},{"model_id":"nemotron-3-nano-30b-a3b","key":"aime_2025","value":89.1,"unit":"percent","source":"NVIDIA 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16","as_of":"2025-12-20","evidence":"official"},{"model_id":"nemotron-3-nano-30b-a3b","key":"tau2_bench","value":49,"unit":"percent","source":"NVIDIA 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16","as_of":"2025-12-20","evidence":"official"},{"model_id":"claude-fable-5","key":"arena_text","value":1507,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-fable-5","key":"aa_index","value":62,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-fable-5","key":"terminal_bench","value":88,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-fable-5-mythos-5","as_of":"2026-08-28","evidence":"official"},{"model_id":"claude-opus-5","key":"arena_text","value":1492,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-opus-5","key":"aa_index","value":63,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-sonnet-5","key":"aa_index","value":55,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-sonnet-5","key":"terminal_bench","value":80.4,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-sonnet-5","as_of":"2026-08-28","evidence":"official"},{"model_id":"claude-opus-4-8","key":"arena_text","value":1481,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-opus-4-8","key":"swe_verified","value":88.6,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-opus-4-8","as_of":"2026-08-28","evidence":"official"},{"model_id":"claude-opus-4-8","key":"gpqa_diamond","value":93.6,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-opus-4-8","as_of":"2026-08-28","evidence":"official"},{"model_id":"claude-opus-4-8","key":"terminal_bench","value":74.6,"unit":"percent","source":"Anthropic 发布公告","source_url":"https://www.anthropic.com/news/claude-opus-4-8","as_of":"2026-08-28","evidence":"official"},{"model_id":"gpt-5-6-sol","key":"arena_text","value":1482,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-sol","key":"aa_index","value":61,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-sol","key":"terminal_bench","value":88.8,"unit":"percent","source":"OpenAI 发布公告","source_url":"https://openai.com/index/gpt-5-6/","as_of":"2026-08-28","evidence":"official"},{"model_id":"gpt-5-6-terra","key":"aa_index","value":57,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-luna","key":"aa_index","value":51,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-5","key":"arena_text","value":1482,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-5","key":"terminal_bench","value":88,"unit":"percent","source":"OpenAI 发布公告","source_url":"https://openai.com/index/introducing-gpt-5-5/","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemini-3-1-pro","key":"arena_text","value":1487,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-1-pro","key":"aa_index","value":48,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-1-pro","key":"swe_verified","value":80.6,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemini-3-1-pro","key":"gpqa_diamond","value":94.3,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemini-3-1-pro","key":"hle","value":44.4,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemini-3-1-pro","key":"terminal_bench","value":68.5,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemini-3-1-pro","key":"mmmu","value":80.5,"unit":"percent","source":"Google 官方模型页","source_url":"https://deepmind.google/models/gemini/pro/","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemini-3-7-flash","key":"arena_text","value":1490,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-7-flash","key":"aa_index","value":56,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-5-flash-lite","key":"arena_text","value":1458,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-5-flash-lite","key":"aa_index","value":37,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"arena_text","value":1451,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"aa_index","value":30,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"livecodebench","value":80,"unit":"percent","source":"Gemma 4 官方模型卡","source_url":"https://ai.google.dev/gemma/docs/core/model_card_4","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemma-4-31b","key":"gpqa_diamond","value":84.3,"unit":"percent","source":"Gemma 4 官方模型卡","source_url":"https://ai.google.dev/gemma/docs/core/model_card_4","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemma-4-31b","key":"hle","value":19.5,"unit":"percent","source":"Gemma 4 官方模型卡","source_url":"https://ai.google.dev/gemma/docs/core/model_card_4","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemma-4-31b","key":"tau2_bench","value":76.9,"unit":"percent","source":"Gemma 4 官方模型卡","source_url":"https://ai.google.dev/gemma/docs/core/model_card_4","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemma-4-31b","key":"mmmu","value":76.9,"unit":"percent","source":"Gemma 4 官方模型卡","source_url":"https://ai.google.dev/gemma/docs/core/model_card_4","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemma-4-26b-a4b","key":"arena_text","value":1438,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-26b-a4b","key":"aa_index","value":26,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-26b-a4b","key":"livecodebench","value":77.1,"unit":"percent","source":"Gemma 4 官方模型卡","source_url":"https://ai.google.dev/gemma/docs/core/model_card_4","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemma-4-26b-a4b","key":"gpqa_diamond","value":82.3,"unit":"percent","source":"Gemma 4 官方模型卡","source_url":"https://ai.google.dev/gemma/docs/core/model_card_4","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemma-4-26b-a4b","key":"hle","value":8.7,"unit":"percent","source":"Gemma 4 官方模型卡","source_url":"https://ai.google.dev/gemma/docs/core/model_card_4","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemma-4-26b-a4b","key":"tau2_bench","value":68.2,"unit":"percent","source":"Gemma 4 官方模型卡","source_url":"https://ai.google.dev/gemma/docs/core/model_card_4","as_of":"2026-08-28","evidence":"official"},{"model_id":"gemma-4-26b-a4b","key":"mmmu","value":73.8,"unit":"percent","source":"Gemma 4 官方模型卡","source_url":"https://ai.google.dev/gemma/docs/core/model_card_4","as_of":"2026-08-28","evidence":"official"},{"model_id":"grok-4-5","key":"arena_text","value":1470,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-5","key":"aa_index","value":56,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-6","key":"arena_text","value":1461,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-6","key":"aa_index","value":61,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-3","key":"arena_text","value":1442,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-3","key":"aa_index","value":38,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-27b","key":"arena_text","value":1436,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-27b","key":"aa_index","value":52,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-27b","key":"livecodebench","value":90.3,"unit":"percent","source":"Qwen3.8-27B 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3.8-27B","as_of":"2026-08-28","evidence":"official"},{"model_id":"qwen3-8-27b","key":"gpqa_diamond","value":89.2,"unit":"percent","source":"Qwen3.8-27B 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3.8-27B","as_of":"2026-08-28","evidence":"official"},{"model_id":"qwen3-8-27b","key":"hle","value":30.8,"unit":"percent","source":"Qwen3.8-27B 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3.8-27B","as_of":"2026-08-28","evidence":"official"},{"model_id":"qwen3-8-27b","key":"terminal_bench","value":73,"unit":"percent","source":"Qwen3.8-27B 模型卡","source_url":"https://huggingface.co/Qwen/Qwen3.8-27B","as_of":"2026-08-28","evidence":"official"},{"model_id":"qwen3-8-2-4t-a95b","key":"arena_text","value":1479,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-2-4t-a95b","key":"aa_index","value":58,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-2-4t-a95b","key":"gpqa_diamond","value":92.6,"unit":"percent","source":"Qwen3.8-2.4T-A95B 模型卡（Qwen3.8-Max 列）","source_url":"https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B","as_of":"2026-08-28","evidence":"official"},{"model_id":"qwen3-8-2-4t-a95b","key":"hle","value":43.6,"unit":"percent","source":"Qwen3.8-2.4T-A95B 模型卡（Qwen3.8-Max 列）","source_url":"https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B","as_of":"2026-08-28","evidence":"official"},{"model_id":"qwen3-8-2-4t-a95b","key":"terminal_bench","value":86.6,"unit":"percent","source":"Qwen3.8-2.4T-A95B 模型卡（Qwen3.8-Max 列）","source_url":"https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B","as_of":"2026-08-28","evidence":"official"},{"model_id":"deepseek-v4-pro","key":"arena_text","value":1462,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-pro","key":"aa_index","value":53,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-pro","key":"swe_verified","value":80.6,"unit":"percent","source":"DeepSeek-V4-Pro 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813","as_of":"2026-08-28","evidence":"official"},{"model_id":"deepseek-v4-pro","key":"livecodebench","value":93.5,"unit":"percent","source":"DeepSeek-V4-Pro 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813","as_of":"2026-08-28","evidence":"official"},{"model_id":"deepseek-v4-pro","key":"gpqa_diamond","value":90.1,"unit":"percent","source":"DeepSeek-V4-Pro 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813","as_of":"2026-08-28","evidence":"official"},{"model_id":"deepseek-v4-pro","key":"hle","value":42.7,"unit":"percent","source":"DeepSeek-V4-Pro 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813","as_of":"2026-08-28","evidence":"official"},{"model_id":"deepseek-v4-pro","key":"terminal_bench","value":87.9,"unit":"percent","source":"DeepSeek-V4-Pro 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813","as_of":"2026-08-28","evidence":"official"},{"model_id":"deepseek-v4-flash","key":"arena_text","value":1436,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-flash","key":"aa_index","value":52,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-flash","key":"swe_verified","value":79,"unit":"percent","source":"DeepSeek-V4-Flash 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731","as_of":"2026-08-28","evidence":"official"},{"model_id":"deepseek-v4-flash","key":"livecodebench","value":91.6,"unit":"percent","source":"DeepSeek-V4-Flash 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731","as_of":"2026-08-28","evidence":"official"},{"model_id":"deepseek-v4-flash","key":"gpqa_diamond","value":88.1,"unit":"percent","source":"DeepSeek-V4-Flash 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731","as_of":"2026-08-28","evidence":"official"},{"model_id":"deepseek-v4-flash","key":"hle","value":37.8,"unit":"percent","source":"DeepSeek-V4-Flash 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731","as_of":"2026-08-28","evidence":"official"},{"model_id":"deepseek-v4-flash","key":"terminal_bench","value":82.7,"unit":"percent","source":"DeepSeek-V4-Flash 模型卡","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731","as_of":"2026-08-28","evidence":"official"},{"model_id":"glm-5-2","key":"arena_text","value":1472,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-2","key":"aa_index","value":53,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-2","key":"gpqa_diamond","value":91.2,"unit":"percent","source":"GLM-5.2 模型卡","source_url":"https://huggingface.co/zai-org/GLM-5.2","as_of":"2026-08-28","evidence":"official"},{"model_id":"glm-5-2","key":"hle","value":40.5,"unit":"percent","source":"GLM-5.2 模型卡","source_url":"https://huggingface.co/zai-org/GLM-5.2","as_of":"2026-08-28","evidence":"official"},{"model_id":"glm-5-2","key":"terminal_bench","value":81,"unit":"percent","source":"GLM-5.2 模型卡","source_url":"https://huggingface.co/zai-org/GLM-5.2","as_of":"2026-08-28","evidence":"official"},{"model_id":"glm-5-3","key":"arena_text","value":1484,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3","key":"aa_index","value":60,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3-flash","key":"arena_text","value":1469,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3-flash","key":"aa_index","value":57,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3-flash","key":"terminal_bench","value":84.3,"unit":"percent","source":"GLM-5.3-Flash 模型卡","source_url":"https://huggingface.co/zai-org/GLM-5.3-Flash","as_of":"2026-08-28","evidence":"official"},{"model_id":"kimi-k3","key":"arena_text","value":1489,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k3","key":"aa_index","value":60,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k3","key":"gpqa_diamond","value":93.5,"unit":"percent","source":"Kimi K3 模型卡","source_url":"https://huggingface.co/moonshotai/Kimi-K3","as_of":"2026-08-28","evidence":"official"},{"model_id":"kimi-k3","key":"hle","value":43.5,"unit":"percent","source":"Kimi K3 模型卡","source_url":"https://huggingface.co/moonshotai/Kimi-K3","as_of":"2026-08-28","evidence":"official"},{"model_id":"kimi-k3","key":"terminal_bench","value":88.3,"unit":"percent","source":"Kimi K3 模型卡","source_url":"https://huggingface.co/moonshotai/Kimi-K3","as_of":"2026-08-28","evidence":"official"},{"model_id":"kimi-k2-6","key":"arena_text","value":1461,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-6","key":"swe_verified","value":80.2,"unit":"percent","source":"Kimi K2.6 模型卡","source_url":"https://huggingface.co/moonshotai/Kimi-K2.6","as_of":"2026-08-28","evidence":"official"},{"model_id":"kimi-k2-6","key":"livecodebench","value":89.6,"unit":"percent","source":"Kimi K2.6 模型卡","source_url":"https://huggingface.co/moonshotai/Kimi-K2.6","as_of":"2026-08-28","evidence":"official"},{"model_id":"kimi-k2-6","key":"gpqa_diamond","value":90.5,"unit":"percent","source":"Kimi K2.6 模型卡","source_url":"https://huggingface.co/moonshotai/Kimi-K2.6","as_of":"2026-08-28","evidence":"official"},{"model_id":"kimi-k2-6","key":"hle","value":34.7,"unit":"percent","source":"Kimi K2.6 模型卡","source_url":"https://huggingface.co/moonshotai/Kimi-K2.6","as_of":"2026-08-28","evidence":"official"},{"model_id":"kimi-k2-6","key":"terminal_bench","value":66.7,"unit":"percent","source":"Kimi K2.6 模型卡","source_url":"https://huggingface.co/moonshotai/Kimi-K2.6","as_of":"2026-08-28","evidence":"official"},{"model_id":"kimi-k2-7-code","key":"aa_index","value":43,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"arena_text","value":1442,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"aa_index","value":45,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"terminal_bench","value":66,"unit":"percent","source":"MiniMax M3 官方博客","source_url":"https://www.minimax.io/blog/minimax-m3","as_of":"2026-08-28","evidence":"official"},{"model_id":"minimax-m2-7","key":"terminal_bench","value":57,"unit":"percent","source":"MiniMax M2.7 模型卡","source_url":"https://huggingface.co/MiniMaxAI/MiniMax-M2.7","as_of":"2026-08-28","evidence":"official"},{"model_id":"mistral-medium-3-5","key":"aa_index","value":30,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-medium-3-5","key":"swe_verified","value":77.6,"unit":"percent","source":"Mistral Medium 3.5 模型卡","source_url":"https://huggingface.co/mistralai/Mistral-Medium-3.5-128B","as_of":"2026-08-28","evidence":"official"},{"model_id":"mistral-large-3","key":"aa_index","value":16,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-small-4","key":"aa_index","value":20,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-small-4","key":"livecodebench","value":64,"unit":"percent","source":"Mistral Small 4 模型卡","source_url":"https://huggingface.co/mistralai/Mistral-Small-4-119B-2603","as_of":"2026-08-28","evidence":"official"},{"model_id":"mistral-small-4","key":"gpqa_diamond","value":71.2,"unit":"percent","source":"Mistral Small 4 模型卡","source_url":"https://huggingface.co/mistralai/Mistral-Small-4-119B-2603","as_of":"2026-08-28","evidence":"official"},{"model_id":"mistral-small-4","key":"aime_2025","value":84,"unit":"percent","source":"Mistral Small 4 模型卡","source_url":"https://huggingface.co/mistralai/Mistral-Small-4-119B-2603","as_of":"2026-08-28","evidence":"official"},{"model_id":"muse-glimmer-30b","key":"aa_index","value":35,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-glimmer-30b","key":"swe_verified","value":76,"unit":"percent","source":"Muse Glimmer 模型卡","source_url":"https://huggingface.co/meta-models/Muse-Glimmer-30B","as_of":"2026-08-28","evidence":"official"},{"model_id":"muse-glimmer-30b","key":"gpqa_diamond","value":83.5,"unit":"percent","source":"Muse Glimmer 模型卡","source_url":"https://huggingface.co/meta-models/Muse-Glimmer-30B","as_of":"2026-08-28","evidence":"official"},{"model_id":"muse-glimmer-30b","key":"hle","value":22,"unit":"percent","source":"Muse Glimmer 模型卡","source_url":"https://huggingface.co/meta-models/Muse-Glimmer-30B","as_of":"2026-08-28","evidence":"official"},{"model_id":"muse-glimmer-30b","key":"terminal_bench","value":51.7,"unit":"percent","source":"Muse Glimmer 模型卡","source_url":"https://huggingface.co/meta-models/Muse-Glimmer-30B","as_of":"2026-08-28","evidence":"official"},{"model_id":"muse-spark","key":"arena_text","value":1498,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-spark","key":"aa_index","value":57,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-spark","key":"hle","value":58,"unit":"percent","source":"Meta Muse Spark 发布公告","source_url":"https://ai.meta.com/blog/introducing-muse-spark-msl/","as_of":"2026-08-28","evidence":"official"},{"model_id":"nemotron-3-super-120b-a12b","key":"aa_index","value":26,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-super-120b-a12b","key":"swe_verified","value":60.5,"unit":"percent","source":"NVIDIA Nemotron 3 Super 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16","as_of":"2026-08-28","evidence":"official"},{"model_id":"nemotron-3-super-120b-a12b","key":"livecodebench","value":81.2,"unit":"percent","source":"NVIDIA Nemotron 3 Super 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16","as_of":"2026-08-28","evidence":"official"},{"model_id":"nemotron-3-super-120b-a12b","key":"gpqa_diamond","value":79.2,"unit":"percent","source":"NVIDIA Nemotron 3 Super 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16","as_of":"2026-08-28","evidence":"official"},{"model_id":"nemotron-3-super-120b-a12b","key":"hle","value":18.3,"unit":"percent","source":"NVIDIA Nemotron 3 Super 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16","as_of":"2026-08-28","evidence":"official"},{"model_id":"nemotron-3-super-120b-a12b","key":"aime_2025","value":90.2,"unit":"percent","source":"NVIDIA Nemotron 3 Super 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16","as_of":"2026-08-28","evidence":"official"},{"model_id":"nemotron-3-super-120b-a12b","key":"tau2_bench","value":61.2,"unit":"percent","source":"NVIDIA Nemotron 3 Super 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16","as_of":"2026-08-28","evidence":"official"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"aa_index","value":38,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"swe_verified","value":70.7,"unit":"percent","source":"NVIDIA Nemotron 3 Ultra 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16","as_of":"2026-08-28","evidence":"official"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"livecodebench","value":89,"unit":"percent","source":"NVIDIA Nemotron 3 Ultra 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16","as_of":"2026-08-28","evidence":"official"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"gpqa_diamond","value":87,"unit":"percent","source":"NVIDIA Nemotron 3 Ultra 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16","as_of":"2026-08-28","evidence":"official"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"terminal_bench","value":56.4,"unit":"percent","source":"NVIDIA Nemotron 3 Ultra 模型卡","source_url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16","as_of":"2026-08-28","evidence":"official"},{"model_id":"hunyuan-hy3","key":"arena_text","value":1456,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2026-08-28","evidence":"independent"},{"model_id":"hunyuan-hy3","key":"aa_index","value":42,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"hunyuan-hy3","key":"swe_verified","value":78,"unit":"percent","source":"Hy3 模型卡","source_url":"https://huggingface.co/tencent/Hy3","as_of":"2026-08-28","evidence":"official"},{"model_id":"hunyuan-hy3","key":"gpqa_diamond","value":90.4,"unit":"percent","source":"Hy3 模型卡","source_url":"https://huggingface.co/tencent/Hy3","as_of":"2026-08-28","evidence":"official"},{"model_id":"hunyuan-hy3","key":"terminal_bench","value":71.7,"unit":"percent","source":"Hy3 模型卡","source_url":"https://huggingface.co/tencent/Hy3","as_of":"2026-08-28","evidence":"official"},{"model_id":"step-3-7-flash","key":"aa_index","value":31,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"step-3-7-flash","key":"terminal_bench","value":59.5,"unit":"percent","source":"Step-3.7-Flash 模型卡","source_url":"https://huggingface.co/stepfun-ai/Step-3.7-Flash","as_of":"2026-08-28","evidence":"official"},{"model_id":"mimo-v2-5","key":"aa_index","value":38,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"command-a-plus","key":"aa_index","value":23,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"granite-4-2-30b","key":"aa_index","value":24,"unit":"index","source":"Artificial Analysis Intelligence Index v3","source_url":"https://artificialanalysis.ai/leaderboards/models","as_of":"2026-08-28","evidence":"independent"},{"model_id":"granite-4-2-30b","key":"livecodebench","value":75.8,"unit":"percent","source":"Granite 4.2 30B 模型卡","source_url":"https://huggingface.co/ibm-granite/granite-4.2-30b","as_of":"2026-08-28","evidence":"official"},{"model_id":"granite-4-2-30b","key":"gpqa_diamond","value":66.4,"unit":"percent","source":"Granite 4.2 30B 模型卡","source_url":"https://huggingface.co/ibm-granite/granite-4.2-30b","as_of":"2026-08-28","evidence":"official"},{"model_id":"granite-4-2-30b","key":"aime_2025","value":89.2,"unit":"percent","source":"Granite 4.2 30B 模型卡","source_url":"https://huggingface.co/ibm-granite/granite-4.2-30b","as_of":"2026-08-28","evidence":"official"},{"model_id":"gpt-4-turbo","key":"gpqa_diamond","value":48,"unit":"percent","source":"OpenAI GPT-4o 发布公告（GPT-4T 列）","source_url":"https://openai.com/index/hello-gpt-4o/","as_of":"2024-05-13","evidence":"official"},{"model_id":"gpt-4-turbo","key":"mmmu","value":63.1,"unit":"percent","source":"OpenAI GPT-4o 发布公告（GPT-4T 列）","source_url":"https://openai.com/index/hello-gpt-4o/","as_of":"2024-05-13","evidence":"official"},{"model_id":"gpt-4o","key":"swe_verified","value":33.2,"unit":"percent","source":"OpenAI GPT-4o 发布公告 / SWE-bench Verified 博客（未核实项已移除）","source_url":"https://openai.com/index/hello-gpt-4o/","as_of":"2024-08-13","evidence":"official"},{"model_id":"gpt-4o","key":"gpqa_diamond","value":53.6,"unit":"percent","source":"OpenAI GPT-4o 发布公告 / SWE-bench Verified 博客（未核实项已移除）","source_url":"https://openai.com/index/hello-gpt-4o/","as_of":"2024-08-13","evidence":"official"},{"model_id":"gpt-4o","key":"mmmu","value":69.1,"unit":"percent","source":"OpenAI GPT-4o 发布公告 / SWE-bench Verified 博客（未核实项已移除）","source_url":"https://openai.com/index/hello-gpt-4o/","as_of":"2024-08-13","evidence":"official"},{"model_id":"gpt-4o-mini","key":"gpqa_diamond","value":40.2,"unit":"percent","source":"OpenAI GPT-4o mini 发布公告 / simple-evals（GPQA）","source_url":"https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/","as_of":"2024-07-18","evidence":"official"},{"model_id":"gpt-4o-mini","key":"mmmu","value":59.4,"unit":"percent","source":"OpenAI GPT-4o mini 发布公告 / simple-evals（GPQA）","source_url":"https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/","as_of":"2024-07-18","evidence":"official"},{"model_id":"o1","key":"swe_verified","value":48.9,"unit":"percent","source":"OpenAI o1-2024-12-17 发布表（o1 and new tools for developers）","source_url":"https://openai.com/index/o1-and-new-tools-for-developers/","as_of":"2024-12-17","evidence":"official"},{"model_id":"o1","key":"gpqa_diamond","value":75.7,"unit":"percent","source":"OpenAI o1-2024-12-17 发布表（o1 and new tools for developers）","source_url":"https://openai.com/index/o1-and-new-tools-for-developers/","as_of":"2024-12-17","evidence":"official"},{"model_id":"o1","key":"mmmu","value":77.3,"unit":"percent","source":"OpenAI o1-2024-12-17 发布表（o1 and new tools for developers）","source_url":"https://openai.com/index/o1-and-new-tools-for-developers/","as_of":"2024-12-17","evidence":"official"},{"model_id":"o3-mini","key":"swe_verified","value":49.3,"unit":"percent","source":"OpenAI o3-mini 发布公告（high）","source_url":"https://openai.com/index/openai-o3-mini/","as_of":"2025-01-31","evidence":"official"},{"model_id":"o3-mini","key":"gpqa_diamond","value":79.7,"unit":"percent","source":"OpenAI o3-mini 发布公告（high）","source_url":"https://openai.com/index/openai-o3-mini/","as_of":"2025-01-31","evidence":"official"},{"model_id":"gpt-4-5","key":"swe_verified","value":38,"unit":"percent","source":"OpenAI GPT-4.5 发布公告","source_url":"https://openai.com/index/introducing-gpt-4-5/","as_of":"2025-02-27","evidence":"official"},{"model_id":"gpt-4-5","key":"gpqa_diamond","value":71.4,"unit":"percent","source":"OpenAI GPT-4.5 发布公告","source_url":"https://openai.com/index/introducing-gpt-4-5/","as_of":"2025-02-27","evidence":"official"},{"model_id":"gpt-4-5","key":"mmmu","value":74.4,"unit":"percent","source":"OpenAI GPT-4.5 发布公告","source_url":"https://openai.com/index/introducing-gpt-4-5/","as_of":"2025-02-27","evidence":"official"},{"model_id":"gpt-4-1","key":"swe_verified","value":54.6,"unit":"percent","source":"OpenAI GPT-4.1 发布公告","source_url":"https://openai.com/index/gpt-4-1/","as_of":"2025-04-14","evidence":"official"},{"model_id":"gpt-4-1","key":"gpqa_diamond","value":66.3,"unit":"percent","source":"OpenAI GPT-4.1 发布公告","source_url":"https://openai.com/index/gpt-4-1/","as_of":"2025-04-14","evidence":"official"},{"model_id":"gpt-4-1","key":"mmmu","value":74.8,"unit":"percent","source":"OpenAI GPT-4.1 发布公告","source_url":"https://openai.com/index/gpt-4-1/","as_of":"2025-04-14","evidence":"official"},{"model_id":"o3","key":"swe_verified","value":69.1,"unit":"percent","source":"OpenAI o3 / o4-mini 发布公告（无工具）","source_url":"https://openai.com/index/introducing-o3-and-o4-mini/","as_of":"2025-04-16","evidence":"official"},{"model_id":"o3","key":"gpqa_diamond","value":83.3,"unit":"percent","source":"OpenAI o3 / o4-mini 发布公告（无工具）","source_url":"https://openai.com/index/introducing-o3-and-o4-mini/","as_of":"2025-04-16","evidence":"official"},{"model_id":"o3","key":"hle","value":20.3,"unit":"percent","source":"OpenAI o3 / o4-mini 发布公告（无工具）","source_url":"https://openai.com/index/introducing-o3-and-o4-mini/","as_of":"2025-04-16","evidence":"official"},{"model_id":"o3","key":"aime_2025","value":88.9,"unit":"percent","source":"OpenAI o3 / o4-mini 发布公告（无工具）","source_url":"https://openai.com/index/introducing-o3-and-o4-mini/","as_of":"2025-04-16","evidence":"official"},{"model_id":"o3","key":"mmmu","value":82.9,"unit":"percent","source":"OpenAI o3 / o4-mini 发布公告（无工具）","source_url":"https://openai.com/index/introducing-o3-and-o4-mini/","as_of":"2025-04-16","evidence":"official"},{"model_id":"o4-mini","key":"swe_verified","value":68.1,"unit":"percent","source":"OpenAI o3 / o4-mini 发布公告（无工具）","source_url":"https://openai.com/index/introducing-o3-and-o4-mini/","as_of":"2025-04-16","evidence":"official"},{"model_id":"o4-mini","key":"gpqa_diamond","value":81.4,"unit":"percent","source":"OpenAI o3 / o4-mini 发布公告（无工具）","source_url":"https://openai.com/index/introducing-o3-and-o4-mini/","as_of":"2025-04-16","evidence":"official"},{"model_id":"o4-mini","key":"hle","value":14.3,"unit":"percent","source":"OpenAI o3 / o4-mini 发布公告（无工具）","source_url":"https://openai.com/index/introducing-o3-and-o4-mini/","as_of":"2025-04-16","evidence":"official"},{"model_id":"o4-mini","key":"aime_2025","value":92.7,"unit":"percent","source":"OpenAI o3 / o4-mini 发布公告（无工具）","source_url":"https://openai.com/index/introducing-o3-and-o4-mini/","as_of":"2025-04-16","evidence":"official"},{"model_id":"o4-mini","key":"mmmu","value":81.6,"unit":"percent","source":"OpenAI o3 / o4-mini 发布公告（无工具）","source_url":"https://openai.com/index/introducing-o3-and-o4-mini/","as_of":"2025-04-16","evidence":"official"},{"model_id":"claude-3-opus","key":"gpqa_diamond","value":50.4,"unit":"percent","source":"Anthropic Claude 3 发布公告","source_url":"https://www.anthropic.com/news/claude-3-family","as_of":"2024-03-04","evidence":"official"},{"model_id":"claude-3-opus","key":"mmmu","value":59.4,"unit":"percent","source":"Anthropic Claude 3 发布公告","source_url":"https://www.anthropic.com/news/claude-3-family","as_of":"2024-03-04","evidence":"official"},{"model_id":"claude-3-5-sonnet","key":"swe_verified","value":49,"unit":"percent","source":"Anthropic 3.5 Sonnet（20241022）公告","source_url":"https://www.anthropic.com/news/3-5-models-and-computer-use","as_of":"2024-10-22","evidence":"official"},{"model_id":"claude-3-5-sonnet","key":"gpqa_diamond","value":65,"unit":"percent","source":"Anthropic 3.5 Sonnet（20241022）公告","source_url":"https://www.anthropic.com/news/3-5-models-and-computer-use","as_of":"2024-10-22","evidence":"official"},{"model_id":"claude-3-5-sonnet","key":"mmmu","value":70.4,"unit":"percent","source":"Anthropic 3.5 Sonnet（20241022）公告","source_url":"https://www.anthropic.com/news/3-5-models-and-computer-use","as_of":"2024-10-22","evidence":"official"},{"model_id":"claude-3-7-sonnet","key":"swe_verified","value":62.3,"unit":"percent","source":"Anthropic Claude 3.7 Sonnet 公告（GPQA/MMMU 为 64K 扩展思考）","source_url":"https://www.anthropic.com/news/claude-3-7-sonnet","as_of":"2025-02-24","evidence":"official"},{"model_id":"claude-3-7-sonnet","key":"gpqa_diamond","value":78.2,"unit":"percent","source":"Anthropic Claude 3.7 Sonnet 公告（GPQA/MMMU 为 64K 扩展思考）","source_url":"https://www.anthropic.com/news/claude-3-7-sonnet","as_of":"2025-02-24","evidence":"official"},{"model_id":"claude-3-7-sonnet","key":"mmmu","value":75,"unit":"percent","source":"Anthropic Claude 3.7 Sonnet 公告（GPQA/MMMU 为 64K 扩展思考）","source_url":"https://www.anthropic.com/news/claude-3-7-sonnet","as_of":"2025-02-24","evidence":"official"},{"model_id":"claude-opus-4","key":"swe_verified","value":72.5,"unit":"percent","source":"Anthropic Claude 4 公告（GPQA/AIME/MMMU 为扩展思考）","source_url":"https://www.anthropic.com/news/claude-4","as_of":"2025-05-22","evidence":"official"},{"model_id":"claude-opus-4","key":"gpqa_diamond","value":79.6,"unit":"percent","source":"Anthropic Claude 4 公告（GPQA/AIME/MMMU 为扩展思考）","source_url":"https://www.anthropic.com/news/claude-4","as_of":"2025-05-22","evidence":"official"},{"model_id":"claude-opus-4","key":"aime_2025","value":75.5,"unit":"percent","source":"Anthropic Claude 4 公告（GPQA/AIME/MMMU 为扩展思考）","source_url":"https://www.anthropic.com/news/claude-4","as_of":"2025-05-22","evidence":"official"},{"model_id":"claude-opus-4","key":"terminal_bench","value":43.2,"unit":"percent","source":"Anthropic Claude 4 公告（GPQA/AIME/MMMU 为扩展思考）","source_url":"https://www.anthropic.com/news/claude-4","as_of":"2025-05-22","evidence":"official"},{"model_id":"claude-opus-4","key":"mmmu","value":76.5,"unit":"percent","source":"Anthropic Claude 4 公告（GPQA/AIME/MMMU 为扩展思考）","source_url":"https://www.anthropic.com/news/claude-4","as_of":"2025-05-22","evidence":"official"},{"model_id":"claude-sonnet-4","key":"swe_verified","value":72.7,"unit":"percent","source":"Anthropic Claude 4 公告（GPQA/AIME/MMMU 为扩展思考）","source_url":"https://www.anthropic.com/news/claude-4","as_of":"2025-05-22","evidence":"official"},{"model_id":"claude-sonnet-4","key":"gpqa_diamond","value":75.4,"unit":"percent","source":"Anthropic Claude 4 公告（GPQA/AIME/MMMU 为扩展思考）","source_url":"https://www.anthropic.com/news/claude-4","as_of":"2025-05-22","evidence":"official"},{"model_id":"claude-sonnet-4","key":"aime_2025","value":70.5,"unit":"percent","source":"Anthropic Claude 4 公告（GPQA/AIME/MMMU 为扩展思考）","source_url":"https://www.anthropic.com/news/claude-4","as_of":"2025-05-22","evidence":"official"},{"model_id":"claude-sonnet-4","key":"terminal_bench","value":35.5,"unit":"percent","source":"Anthropic Claude 4 公告（GPQA/AIME/MMMU 为扩展思考）","source_url":"https://www.anthropic.com/news/claude-4","as_of":"2025-05-22","evidence":"official"},{"model_id":"claude-sonnet-4","key":"mmmu","value":74.4,"unit":"percent","source":"Anthropic Claude 4 公告（GPQA/AIME/MMMU 为扩展思考）","source_url":"https://www.anthropic.com/news/claude-4","as_of":"2025-05-22","evidence":"official"},{"model_id":"gemini-1-0-pro","key":"mmmu","value":47.9,"unit":"percent","source":"Gemini 1.0 技术报告（MMMU val, pass@1）","source_url":"https://arxiv.org/abs/2312.11805","as_of":"2023-12-06","evidence":"official"},{"model_id":"gemini-1-5-pro","key":"gpqa_diamond","value":59.1,"unit":"percent","source":"Google Gemini 1.5 Pro-002 发布","source_url":"https://developers.googleblog.com/en/updated-production-ready-gemini-models-reduced-15-pro-pricing-increased-rate-limits-and-more/","as_of":"2024-09-24","evidence":"official"},{"model_id":"gemini-1-5-pro","key":"mmmu","value":65.9,"unit":"percent","source":"Google Gemini 1.5 Pro-002 发布","source_url":"https://developers.googleblog.com/en/updated-production-ready-gemini-models-reduced-15-pro-pricing-increased-rate-limits-and-more/","as_of":"2024-09-24","evidence":"official"},{"model_id":"gemini-1-5-flash","key":"gpqa_diamond","value":51,"unit":"percent","source":"Google Gemini 1.5 Flash-002 发布","source_url":"https://developers.googleblog.com/en/updated-production-ready-gemini-models-reduced-15-pro-pricing-increased-rate-limits-and-more/","as_of":"2024-09-24","evidence":"official"},{"model_id":"gemini-1-5-flash","key":"mmmu","value":62.3,"unit":"percent","source":"Google Gemini 1.5 Flash-002 发布","source_url":"https://developers.googleblog.com/en/updated-production-ready-gemini-models-reduced-15-pro-pricing-increased-rate-limits-and-more/","as_of":"2024-09-24","evidence":"official"},{"model_id":"gemini-2-0-flash","key":"swe_verified","value":51.8,"unit":"percent","source":"Google Gemini 2.0 Flash GA 表（LiveCodeBench v5；SWE-bench 51.8 来自 2024-12 开发者博客，含代码执行工具）","source_url":"https://developers.googleblog.com/en/gemini-2-family-expands/","as_of":"2025-02-05","evidence":"official"},{"model_id":"gemini-2-0-flash","key":"livecodebench","value":34.5,"unit":"percent","source":"Google Gemini 2.0 Flash GA 表（LiveCodeBench v5；SWE-bench 51.8 来自 2024-12 开发者博客，含代码执行工具）","source_url":"https://developers.googleblog.com/en/gemini-2-family-expands/","as_of":"2025-02-05","evidence":"official"},{"model_id":"gemini-2-0-flash","key":"gpqa_diamond","value":60.1,"unit":"percent","source":"Google Gemini 2.0 Flash GA 表（LiveCodeBench v5；SWE-bench 51.8 来自 2024-12 开发者博客，含代码执行工具）","source_url":"https://developers.googleblog.com/en/gemini-2-family-expands/","as_of":"2025-02-05","evidence":"official"},{"model_id":"gemini-2-0-flash","key":"mmmu","value":71.7,"unit":"percent","source":"Google Gemini 2.0 Flash GA 表（LiveCodeBench v5；SWE-bench 51.8 来自 2024-12 开发者博客，含代码执行工具）","source_url":"https://developers.googleblog.com/en/gemini-2-family-expands/","as_of":"2025-02-05","evidence":"official"},{"model_id":"grok-2","key":"gpqa_diamond","value":56,"unit":"percent","source":"xAI Grok-2 公告（GPQA 未标注 Diamond）","source_url":"https://x.ai/news/grok-2","as_of":"2024-08-13","evidence":"official"},{"model_id":"grok-2","key":"mmmu","value":66.1,"unit":"percent","source":"xAI Grok-2 公告（GPQA 未标注 Diamond）","source_url":"https://x.ai/news/grok-2","as_of":"2024-08-13","evidence":"official"},{"model_id":"grok-3","key":"arena_text","value":1402,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-02-19","evidence":"independent"},{"model_id":"grok-3","key":"livecodebench","value":57,"unit":"percent","source":"xAI Grok 3 公告（非推理 pass@1；LCB 2024-10-01~2025-02-01）","source_url":"https://x.ai/news/grok-3","as_of":"2025-02-19","evidence":"official"},{"model_id":"grok-3","key":"gpqa_diamond","value":75.4,"unit":"percent","source":"xAI Grok 3 公告（非推理 pass@1；LCB 2024-10-01~2025-02-01）","source_url":"https://x.ai/news/grok-3","as_of":"2025-02-19","evidence":"official"},{"model_id":"grok-3","key":"mmmu","value":73.2,"unit":"percent","source":"xAI Grok 3 公告（非推理 pass@1；LCB 2024-10-01~2025-02-01）","source_url":"https://x.ai/news/grok-3","as_of":"2025-02-19","evidence":"official"},{"model_id":"llama-3-1-405b","key":"gpqa_diamond","value":49,"unit":"percent","source":"Llama 3.3 模型卡（GPQA Diamond CoT 对照列）","source_url":"https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct","as_of":"2024-07-23","evidence":"official"},{"model_id":"llama-3-1-8b","key":"gpqa_diamond","value":31.8,"unit":"percent","source":"Llama 3.3 模型卡（GPQA Diamond CoT 对照列）","source_url":"https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct","as_of":"2024-07-23","evidence":"official"},{"model_id":"llama-3-3-70b","key":"gpqa_diamond","value":50.5,"unit":"percent","source":"Llama 3.3 模型卡","source_url":"https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct","as_of":"2024-12-06","evidence":"official"},{"model_id":"mistral-small-3-1","key":"gpqa_diamond","value":46,"unit":"percent","source":"Mistral Small 3.1 模型卡","source_url":"https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503","as_of":"2025-03-17","evidence":"official"},{"model_id":"mistral-small-3-1","key":"mmmu","value":64,"unit":"percent","source":"Mistral Small 3.1 模型卡","source_url":"https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503","as_of":"2025-03-17","evidence":"official"},{"model_id":"devstral-small","key":"swe_verified","value":46.8,"unit":"percent","source":"Devstral 模型卡（OpenHands）","source_url":"https://huggingface.co/mistralai/Devstral-Small-2505","as_of":"2025-05-21","evidence":"official"},{"model_id":"qwen2-5-72b","key":"livecodebench","value":55.5,"unit":"percent","source":"Qwen2.5 技术报告表6（LiveCodeBench 2305-2409；GPQA 未标注 Diamond）","source_url":"https://arxiv.org/abs/2412.15115","as_of":"2024-09-19","evidence":"official"},{"model_id":"qwen2-5-72b","key":"gpqa_diamond","value":49,"unit":"percent","source":"Qwen2.5 技术报告表6（LiveCodeBench 2305-2409；GPQA 未标注 Diamond）","source_url":"https://arxiv.org/abs/2412.15115","as_of":"2024-09-19","evidence":"official"},{"model_id":"qwen2-5-coder-32b","key":"livecodebench","value":31.4,"unit":"percent","source":"Qwen2.5-Coder 博客（LiveCodeBench 2024.07-2024.11）","source_url":"https://qwenlm.github.io/blog/qwen2.5-coder-family/","as_of":"2024-11-12","evidence":"official"},{"model_id":"qwq-32b","key":"livecodebench","value":63.4,"unit":"percent","source":"QwQ-32B 博客（LiveCodeBench 24.08-25.02）","source_url":"https://qwenlm.github.io/blog/qwq-32b/","as_of":"2025-03-05","evidence":"official"},{"model_id":"qwen3-235b-a22b","key":"livecodebench","value":70.7,"unit":"percent","source":"Qwen3 技术报告表11（思考模式，LiveCodeBench v5）","source_url":"https://qwenlm.github.io/blog/qwen3/","as_of":"2025-04-29","evidence":"official"},{"model_id":"qwen3-235b-a22b","key":"gpqa_diamond","value":71.1,"unit":"percent","source":"Qwen3 技术报告表11（思考模式，LiveCodeBench v5）","source_url":"https://qwenlm.github.io/blog/qwen3/","as_of":"2025-04-29","evidence":"official"},{"model_id":"qwen3-235b-a22b","key":"aime_2025","value":81.5,"unit":"percent","source":"Qwen3 技术报告表11（思考模式，LiveCodeBench v5）","source_url":"https://qwenlm.github.io/blog/qwen3/","as_of":"2025-04-29","evidence":"official"},{"model_id":"qwen3-8b","key":"livecodebench","value":57.5,"unit":"percent","source":"Qwen3 技术报告表17（思考模式，LiveCodeBench v5）","source_url":"https://arxiv.org/abs/2505.09388","as_of":"2025-05-14","evidence":"official"},{"model_id":"qwen3-8b","key":"gpqa_diamond","value":62,"unit":"percent","source":"Qwen3 技术报告表17（思考模式，LiveCodeBench v5）","source_url":"https://arxiv.org/abs/2505.09388","as_of":"2025-05-14","evidence":"official"},{"model_id":"qwen3-8b","key":"aime_2025","value":67.3,"unit":"percent","source":"Qwen3 技术报告表17（思考模式，LiveCodeBench v5）","source_url":"https://arxiv.org/abs/2505.09388","as_of":"2025-05-14","evidence":"official"},{"model_id":"deepseek-v2","key":"livecodebench","value":32.5,"unit":"percent","source":"DeepSeek-V2 模型卡（LiveCodeBench 0901-0401）","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V2-Chat","as_of":"2024-05-06","evidence":"official"},{"model_id":"deepseek-coder-v2","key":"livecodebench","value":43.4,"unit":"percent","source":"DeepSeek-Coder-V2 论文表4（LiveCodeBench 1201-0601）","source_url":"https://arxiv.org/abs/2406.11931","as_of":"2024-06-17","evidence":"official"},{"model_id":"deepseek-v3","key":"livecodebench","value":40.5,"unit":"percent","source":"DeepSeek-V3 模型卡（LiveCodeBench Pass@1-COT）","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3","as_of":"2024-12-26","evidence":"official"},{"model_id":"deepseek-v3","key":"gpqa_diamond","value":59.1,"unit":"percent","source":"DeepSeek-V3 模型卡（LiveCodeBench Pass@1-COT）","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3","as_of":"2024-12-26","evidence":"official"},{"model_id":"deepseek-r1","key":"swe_verified","value":49.2,"unit":"percent","source":"DeepSeek-R1 模型卡（LiveCodeBench Pass@1-COT）","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-R1","as_of":"2025-01-20","evidence":"official"},{"model_id":"deepseek-r1","key":"livecodebench","value":65.9,"unit":"percent","source":"DeepSeek-R1 模型卡（LiveCodeBench Pass@1-COT）","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-R1","as_of":"2025-01-20","evidence":"official"},{"model_id":"deepseek-r1","key":"gpqa_diamond","value":71.5,"unit":"percent","source":"DeepSeek-R1 模型卡（LiveCodeBench Pass@1-COT）","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-R1","as_of":"2025-01-20","evidence":"official"},{"model_id":"deepseek-v3-0324","key":"swe_verified","value":45.4,"unit":"percent","source":"DeepSeek-V3-0324 / V3.1 模型卡对照列","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3-0324","as_of":"2025-03-24","evidence":"official"},{"model_id":"deepseek-v3-0324","key":"livecodebench","value":49.2,"unit":"percent","source":"DeepSeek-V3-0324 / V3.1 模型卡对照列","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3-0324","as_of":"2025-03-24","evidence":"official"},{"model_id":"deepseek-v3-0324","key":"gpqa_diamond","value":68.4,"unit":"percent","source":"DeepSeek-V3-0324 / V3.1 模型卡对照列","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3-0324","as_of":"2025-03-24","evidence":"official"},{"model_id":"deepseek-v3-0324","key":"terminal_bench","value":13.3,"unit":"percent","source":"DeepSeek-V3-0324 / V3.1 模型卡对照列","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3-0324","as_of":"2025-03-24","evidence":"official"},{"model_id":"deepseek-v3-1","key":"swe_verified","value":66,"unit":"percent","source":"DeepSeek-V3.1 模型卡（思考模式；LiveCodeBench 2408-2505）","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.1","as_of":"2025-08-21","evidence":"official"},{"model_id":"deepseek-v3-1","key":"livecodebench","value":74.8,"unit":"percent","source":"DeepSeek-V3.1 模型卡（思考模式；LiveCodeBench 2408-2505）","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.1","as_of":"2025-08-21","evidence":"official"},{"model_id":"deepseek-v3-1","key":"gpqa_diamond","value":80.1,"unit":"percent","source":"DeepSeek-V3.1 模型卡（思考模式；LiveCodeBench 2408-2505）","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.1","as_of":"2025-08-21","evidence":"official"},{"model_id":"deepseek-v3-1","key":"hle","value":15.9,"unit":"percent","source":"DeepSeek-V3.1 模型卡（思考模式；LiveCodeBench 2408-2505）","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.1","as_of":"2025-08-21","evidence":"official"},{"model_id":"deepseek-v3-1","key":"aime_2025","value":88.4,"unit":"percent","source":"DeepSeek-V3.1 模型卡（思考模式；LiveCodeBench 2408-2505）","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.1","as_of":"2025-08-21","evidence":"official"},{"model_id":"deepseek-v3-1","key":"terminal_bench","value":31.3,"unit":"percent","source":"DeepSeek-V3.1 模型卡（思考模式；LiveCodeBench 2408-2505）","source_url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.1","as_of":"2025-08-21","evidence":"official"},{"model_id":"glm-4-5","key":"swe_verified","value":64.2,"unit":"percent","source":"GLM-4.5 技术报告表4/5（LiveCodeBench 2407-2501；仅报告 AIME 24）（未核实项已移除）","source_url":"https://arxiv.org/abs/2508.06471","as_of":"2025-07-28","evidence":"official"},{"model_id":"glm-4-5","key":"livecodebench","value":72.9,"unit":"percent","source":"GLM-4.5 技术报告表4/5（LiveCodeBench 2407-2501；仅报告 AIME 24）（未核实项已移除）","source_url":"https://arxiv.org/abs/2508.06471","as_of":"2025-07-28","evidence":"official"},{"model_id":"glm-4-5","key":"gpqa_diamond","value":79.1,"unit":"percent","source":"GLM-4.5 技术报告表4/5（LiveCodeBench 2407-2501；仅报告 AIME 24）（未核实项已移除）","source_url":"https://arxiv.org/abs/2508.06471","as_of":"2025-07-28","evidence":"official"},{"model_id":"glm-4-5","key":"hle","value":14.4,"unit":"percent","source":"GLM-4.5 技术报告表4/5（LiveCodeBench 2407-2501；仅报告 AIME 24）（未核实项已移除）","source_url":"https://arxiv.org/abs/2508.06471","as_of":"2025-07-28","evidence":"official"},{"model_id":"glm-4-5","key":"terminal_bench","value":37.5,"unit":"percent","source":"GLM-4.5 技术报告表4/5（LiveCodeBench 2407-2501；仅报告 AIME 24）（未核实项已移除）","source_url":"https://arxiv.org/abs/2508.06471","as_of":"2025-07-28","evidence":"official"},{"model_id":"kimi-k2-instruct","key":"swe_verified","value":65.8,"unit":"percent","source":"Kimi K2 模型卡（LiveCodeBench v6）","source_url":"https://huggingface.co/moonshotai/Kimi-K2-Instruct","as_of":"2025-07-11","evidence":"official"},{"model_id":"kimi-k2-instruct","key":"livecodebench","value":53.7,"unit":"percent","source":"Kimi K2 模型卡（LiveCodeBench v6）","source_url":"https://huggingface.co/moonshotai/Kimi-K2-Instruct","as_of":"2025-07-11","evidence":"official"},{"model_id":"kimi-k2-instruct","key":"gpqa_diamond","value":75.1,"unit":"percent","source":"Kimi K2 模型卡（LiveCodeBench v6）","source_url":"https://huggingface.co/moonshotai/Kimi-K2-Instruct","as_of":"2025-07-11","evidence":"official"},{"model_id":"kimi-k2-instruct","key":"hle","value":4.7,"unit":"percent","source":"Kimi K2 模型卡（LiveCodeBench v6）","source_url":"https://huggingface.co/moonshotai/Kimi-K2-Instruct","as_of":"2025-07-11","evidence":"official"},{"model_id":"kimi-k2-instruct","key":"aime_2025","value":49.5,"unit":"percent","source":"Kimi K2 模型卡（LiveCodeBench v6）","source_url":"https://huggingface.co/moonshotai/Kimi-K2-Instruct","as_of":"2025-07-11","evidence":"official"},{"model_id":"kimi-k2-instruct","key":"terminal_bench","value":30,"unit":"percent","source":"Kimi K2 模型卡（LiveCodeBench v6）","source_url":"https://huggingface.co/moonshotai/Kimi-K2-Instruct","as_of":"2025-07-11","evidence":"official"},{"model_id":"phi-4","key":"gpqa_diamond","value":56.1,"unit":"percent","source":"Phi-4 模型卡","source_url":"https://huggingface.co/microsoft/phi-4","as_of":"2024-12-12","evidence":"official"},{"model_id":"llama-3-1-nemotron-70b","key":"arena_text","value":1267,"unit":"elo","source":"LMArena Text (style control)","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2024-10-24","evidence":"independent"},{"model_id":"minimax-m1","key":"swe_verified","value":56,"unit":"percent","source":"MiniMax-M1 模型卡（80k；LiveCodeBench 24/8-25/5）","source_url":"https://huggingface.co/MiniMaxAI/MiniMax-M1-80k","as_of":"2025-06-16","evidence":"official"},{"model_id":"minimax-m1","key":"livecodebench","value":65,"unit":"percent","source":"MiniMax-M1 模型卡（80k；LiveCodeBench 24/8-25/5）","source_url":"https://huggingface.co/MiniMaxAI/MiniMax-M1-80k","as_of":"2025-06-16","evidence":"official"},{"model_id":"minimax-m1","key":"gpqa_diamond","value":70,"unit":"percent","source":"MiniMax-M1 模型卡（80k；LiveCodeBench 24/8-25/5）","source_url":"https://huggingface.co/MiniMaxAI/MiniMax-M1-80k","as_of":"2025-06-16","evidence":"official"},{"model_id":"minimax-m1","key":"hle","value":8.4,"unit":"percent","source":"MiniMax-M1 模型卡（80k；LiveCodeBench 24/8-25/5）","source_url":"https://huggingface.co/MiniMaxAI/MiniMax-M1-80k","as_of":"2025-06-16","evidence":"official"},{"model_id":"minimax-m1","key":"aime_2025","value":76.9,"unit":"percent","source":"MiniMax-M1 模型卡（80k；LiveCodeBench 24/8-25/5）","source_url":"https://huggingface.co/MiniMaxAI/MiniMax-M1-80k","as_of":"2025-06-16","evidence":"official"},{"model_id":"hunyuan-a13b","key":"livecodebench","value":63.9,"unit":"percent","source":"Hunyuan-A13B 模型卡","source_url":"https://huggingface.co/tencent/Hunyuan-A13B-Instruct","as_of":"2025-06-27","evidence":"official"},{"model_id":"hunyuan-a13b","key":"gpqa_diamond","value":71.2,"unit":"percent","source":"Hunyuan-A13B 模型卡","source_url":"https://huggingface.co/tencent/Hunyuan-A13B-Instruct","as_of":"2025-06-27","evidence":"official"},{"model_id":"hunyuan-a13b","key":"aime_2025","value":76.8,"unit":"percent","source":"Hunyuan-A13B 模型卡","source_url":"https://huggingface.co/tencent/Hunyuan-A13B-Instruct","as_of":"2025-06-27","evidence":"official"},{"model_id":"mimo-7b","key":"livecodebench","value":57.8,"unit":"percent","source":"MiMo-7B-RL 模型卡（LiveCodeBench v5）","source_url":"https://huggingface.co/XiaomiMiMo/MiMo-7B-RL","as_of":"2025-04-30","evidence":"official"},{"model_id":"mimo-7b","key":"gpqa_diamond","value":54.4,"unit":"percent","source":"MiMo-7B-RL 模型卡（LiveCodeBench v5）","source_url":"https://huggingface.co/XiaomiMiMo/MiMo-7B-RL","as_of":"2025-04-30","evidence":"official"},{"model_id":"mimo-7b","key":"aime_2025","value":55.4,"unit":"percent","source":"MiMo-7B-RL 模型卡（LiveCodeBench v5）","source_url":"https://huggingface.co/XiaomiMiMo/MiMo-7B-RL","as_of":"2025-04-30","evidence":"official"},{"model_id":"claude-fable-5","key":"aa_index","value":62,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-fable-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-fable-5","key":"gpqa_diamond","value":92.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-fable-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-fable-5","key":"hle","value":55.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-fable-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-fable-5","key":"scicode","value":60.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-fable-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-fable-5","key":"tb_hard","value":62.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-fable-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-fable-5","key":"tau2_telecom","value":98.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-fable-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-fable-5","key":"ifbench","value":63.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-fable-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"aa_index","value":30,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"gpqa_diamond","value":67.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"hle","value":10.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"aime_2025","value":83.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"livecodebench","value":61.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"scicode","value":43.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"tb_hard","value":27.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"tau2_telecom","value":54.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"mmmu_pro","value":58.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"ifbench","value":54.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-opus-5","key":"aa_index","value":63,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-opus-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-opus-5","key":"gpqa_diamond","value":93.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-opus-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-opus-5","key":"hle","value":54.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-opus-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-opus-5","key":"scicode","value":55.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-opus-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-opus-5","key":"mmmu_pro","value":84.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-opus-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-sonnet-5","key":"aa_index","value":55,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-sonnet-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-sonnet-5","key":"gpqa_diamond","value":91.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-sonnet-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-sonnet-5","key":"hle","value":41.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-sonnet-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-sonnet-5","key":"scicode","value":53.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-sonnet-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-sonnet-5","key":"mmmu_pro","value":77.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/claude-sonnet-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"command-a-plus","key":"aa_index","value":23,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/command-a-plus","as_of":"2026-08-28","evidence":"independent"},{"model_id":"command-a-plus","key":"gpqa_diamond","value":76.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/command-a-plus","as_of":"2026-08-28","evidence":"independent"},{"model_id":"command-a-plus","key":"hle","value":12,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/command-a-plus","as_of":"2026-08-28","evidence":"independent"},{"model_id":"command-a-plus","key":"scicode","value":37.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/command-a-plus","as_of":"2026-08-28","evidence":"independent"},{"model_id":"command-a-plus","key":"tb_hard","value":25,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/command-a-plus","as_of":"2026-08-28","evidence":"independent"},{"model_id":"command-a-plus","key":"tau2_telecom","value":80.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/command-a-plus","as_of":"2026-08-28","evidence":"independent"},{"model_id":"command-a-plus","key":"mmmu_pro","value":63.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/command-a-plus","as_of":"2026-08-28","evidence":"independent"},{"model_id":"command-a-plus","key":"ifbench","value":73.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/command-a-plus","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-flash","key":"aa_index","value":52,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/deepseek-v4-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-flash","key":"gpqa_diamond","value":90.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/deepseek-v4-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-flash","key":"hle","value":38.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/deepseek-v4-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-flash","key":"scicode","value":49.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/deepseek-v4-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-pro","key":"aa_index","value":53,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/deepseek-v4-pro","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-pro","key":"gpqa_diamond","value":92.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/deepseek-v4-pro","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-pro","key":"hle","value":41,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/deepseek-v4-pro","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-pro","key":"scicode","value":49.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/deepseek-v4-pro","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-1-pro","key":"aa_index","value":48,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-1-pro-preview","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-1-pro","key":"gpqa_diamond","value":94.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-1-pro-preview","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-1-pro","key":"hle","value":47,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-1-pro-preview","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-1-pro","key":"scicode","value":58.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-1-pro-preview","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-1-pro","key":"tb_hard","value":53.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-1-pro-preview","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-1-pro","key":"tau2_telecom","value":95.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-1-pro-preview","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-1-pro","key":"mmmu_pro","value":82.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-1-pro-preview","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-1-pro","key":"ifbench","value":77.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-1-pro-preview","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-5-flash-lite","key":"aa_index","value":37,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-5-flash-lite","key":"gpqa_diamond","value":83.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-5-flash-lite","key":"hle","value":18.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-5-flash-lite","key":"scicode","value":40.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-5-flash-lite","key":"mmmu_pro","value":79,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-7-flash","key":"aa_index","value":56,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-7-flash","key":"gpqa_diamond","value":94.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-7-flash","key":"hle","value":47.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-7-flash","key":"scicode","value":56.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-7-flash","key":"mmmu_pro","value":85.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemini-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-26b-a4b","key":"aa_index","value":26,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-26b-a4b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-26b-a4b","key":"gpqa_diamond","value":79.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-26b-a4b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-26b-a4b","key":"hle","value":19.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-26b-a4b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-26b-a4b","key":"scicode","value":40,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-26b-a4b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-26b-a4b","key":"tb_hard","value":13.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-26b-a4b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-26b-a4b","key":"tau2_telecom","value":43.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-26b-a4b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-26b-a4b","key":"mmmu_pro","value":69.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-26b-a4b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-26b-a4b","key":"ifbench","value":72.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-26b-a4b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"aa_index","value":30,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-31b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"gpqa_diamond","value":85.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-31b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"hle","value":23.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-31b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"scicode","value":43.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-31b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"tb_hard","value":36.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-31b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"tau2_telecom","value":59.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-31b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"mmmu_pro","value":73.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-31b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"ifbench","value":75.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gemma-4-31b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-2","key":"aa_index","value":53,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-2","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-2","key":"gpqa_diamond","value":89.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-2","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-2","key":"hle","value":41.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-2","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-2","key":"scicode","value":50.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-2","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-2","key":"tb_hard","value":50.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-2","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-2","key":"tau2_telecom","value":99.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-2","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-2","key":"ifbench","value":73.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-2","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3-flash","key":"aa_index","value":57,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-3-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3-flash","key":"gpqa_diamond","value":91.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-3-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3-flash","key":"hle","value":39.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-3-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3-flash","key":"scicode","value":46.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-3-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3","key":"aa_index","value":60,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3","key":"gpqa_diamond","value":91.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3","key":"hle","value":42.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3","key":"scicode","value":56.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/glm-5-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-5","key":"aa_index","value":56,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-5","key":"gpqa_diamond","value":93.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-5","key":"hle","value":45.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-5","key":"scicode","value":56.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-5","key":"tb_hard","value":60.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-5","key":"tau2_telecom","value":93.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-5","key":"mmmu_pro","value":79.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-5","key":"ifbench","value":75.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-luna","key":"aa_index","value":52,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-luna","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-luna","key":"gpqa_diamond","value":91.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-luna","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-luna","key":"hle","value":39.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-luna","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-luna","key":"scicode","value":52.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-luna","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-luna","key":"mmmu_pro","value":78.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-luna","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-sol","key":"aa_index","value":61,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-sol","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-sol","key":"gpqa_diamond","value":94.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-sol","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-sol","key":"hle","value":49.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-sol","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-sol","key":"scicode","value":56.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-sol","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-sol","key":"tb_hard","value":65.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-sol","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-sol","key":"tau2_telecom","value":85.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-sol","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-sol","key":"mmmu_pro","value":83.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-sol","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-sol","key":"ifbench","value":72.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-sol","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-terra","key":"aa_index","value":57,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-terra","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-terra","key":"gpqa_diamond","value":92.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-terra","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-terra","key":"hle","value":42.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-terra","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-terra","key":"scicode","value":53.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-terra","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-terra","key":"tb_hard","value":57.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-terra","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-terra","key":"tau2_telecom","value":86.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-terra","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-terra","key":"mmmu_pro","value":80.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-terra","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-terra","key":"ifbench","value":71.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-5-6-terra","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-120b","key":"aa_index","value":24,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-120b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-120b","key":"gpqa_diamond","value":78.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-120b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-120b","key":"hle","value":19.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-120b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-120b","key":"aime_2025","value":93.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-120b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-120b","key":"livecodebench","value":87.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-120b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-120b","key":"scicode","value":38.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-120b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-120b","key":"tb_hard","value":23.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-120b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-120b","key":"tau2_telecom","value":65.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-120b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-120b","key":"ifbench","value":69,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-120b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-20b","key":"aa_index","value":15,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-20b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-20b","key":"gpqa_diamond","value":68.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-20b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-20b","key":"hle","value":11,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-20b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-20b","key":"aime_2025","value":89.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-20b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-20b","key":"livecodebench","value":77.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-20b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-20b","key":"scicode","value":34.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-20b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-20b","key":"tb_hard","value":10.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-20b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-20b","key":"tau2_telecom","value":60.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-20b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-20b","key":"ifbench","value":65.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/gpt-oss-20b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"granite-4-2-30b","key":"aa_index","value":24,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/granite-4-2-30b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"granite-4-2-30b","key":"gpqa_diamond","value":64.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/granite-4-2-30b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"granite-4-2-30b","key":"hle","value":11.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/granite-4-2-30b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"granite-4-2-30b","key":"scicode","value":36.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/granite-4-2-30b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-3","key":"aa_index","value":38,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-3","key":"gpqa_diamond","value":90.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-3","key":"hle","value":37.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-3","key":"scicode","value":47.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-3","key":"tb_hard","value":37.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-3","key":"tau2_telecom","value":97.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-3","key":"mmmu_pro","value":78.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-3","key":"ifbench","value":81.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-5","key":"aa_index","value":56,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-5","key":"gpqa_diamond","value":93.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-5","key":"hle","value":42.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-5","key":"scicode","value":54.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-5","key":"mmmu_pro","value":80.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-6","key":"aa_index","value":61,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-6","key":"gpqa_diamond","value":94.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-6","key":"hle","value":42.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-6","key":"scicode","value":53.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/grok-4-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"hunyuan-hy3","key":"aa_index","value":42,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/hy3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"hunyuan-hy3","key":"gpqa_diamond","value":89.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/hy3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"hunyuan-hy3","key":"hle","value":33.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/hy3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"hunyuan-hy3","key":"scicode","value":47.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/hy3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-6","key":"aa_index","value":45,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-6","key":"gpqa_diamond","value":91.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-6","key":"hle","value":37.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-6","key":"scicode","value":53.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-6","key":"tb_hard","value":43.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-6","key":"tau2_telecom","value":95.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-6","key":"mmmu_pro","value":79.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-6","key":"ifbench","value":76,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-7-code","key":"aa_index","value":43,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-7-code","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-7-code","key":"gpqa_diamond","value":89.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-7-code","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-7-code","key":"hle","value":35,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-7-code","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-7-code","key":"scicode","value":47.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-7-code","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-7-code","key":"tb_hard","value":44.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-7-code","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-7-code","key":"tau2_telecom","value":90.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-7-code","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-7-code","key":"ifbench","value":63.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k2-7-code","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k3","key":"aa_index","value":60,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k3","key":"gpqa_diamond","value":93.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k3","key":"hle","value":46.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k3","key":"scicode","value":58.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k3","key":"mmmu_pro","value":80.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/kimi-k3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mimo-v2-5","key":"aa_index","value":38,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mimo-v2-5-0424","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mimo-v2-5","key":"gpqa_diamond","value":84.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mimo-v2-5-0424","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mimo-v2-5","key":"hle","value":27.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mimo-v2-5-0424","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mimo-v2-5","key":"scicode","value":43.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mimo-v2-5-0424","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mimo-v2-5","key":"tb_hard","value":41.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mimo-v2-5-0424","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mimo-v2-5","key":"tau2_telecom","value":90.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mimo-v2-5-0424","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mimo-v2-5","key":"mmmu_pro","value":75.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mimo-v2-5-0424","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mimo-v2-5","key":"ifbench","value":67.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mimo-v2-5-0424","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m2-7","key":"aa_index","value":39,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m2-7","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m2-7","key":"gpqa_diamond","value":87.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m2-7","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m2-7","key":"hle","value":29.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m2-7","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m2-7","key":"scicode","value":47,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m2-7","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m2-7","key":"tb_hard","value":39.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m2-7","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m2-7","key":"tau2_telecom","value":84.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m2-7","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m2-7","key":"ifbench","value":75.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m2-7","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"aa_index","value":45,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"gpqa_diamond","value":92.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"hle","value":39,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"scicode","value":45.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"tb_hard","value":42.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"tau2_telecom","value":88.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"mmmu_pro","value":78.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"ifbench","value":82.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/minimax-m3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-large-3","key":"aa_index","value":16,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-large-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-large-3","key":"gpqa_diamond","value":68,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-large-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-large-3","key":"hle","value":4.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-large-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-large-3","key":"aime_2025","value":38,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-large-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-large-3","key":"livecodebench","value":46.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-large-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-large-3","key":"scicode","value":36.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-large-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-large-3","key":"tb_hard","value":15.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-large-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-large-3","key":"tau2_telecom","value":24.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-large-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-large-3","key":"mmmu_pro","value":55.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-large-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-large-3","key":"ifbench","value":36.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-large-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-medium-3-5","key":"aa_index","value":30,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-medium-3-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-medium-3-5","key":"gpqa_diamond","value":74.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-medium-3-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-medium-3-5","key":"hle","value":13.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-medium-3-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-medium-3-5","key":"scicode","value":39.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-medium-3-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-medium-3-5","key":"tb_hard","value":33.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-medium-3-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-medium-3-5","key":"tau2_telecom","value":94.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-medium-3-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-medium-3-5","key":"mmmu_pro","value":64.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-medium-3-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-medium-3-5","key":"ifbench","value":68.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-medium-3-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-small-4","key":"aa_index","value":20,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-small-4","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-small-4","key":"gpqa_diamond","value":76.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-small-4","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-small-4","key":"hle","value":9.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-small-4","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-small-4","key":"scicode","value":38,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-small-4","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-small-4","key":"tb_hard","value":17.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-small-4","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-small-4","key":"tau2_telecom","value":41.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-small-4","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-small-4","key":"mmmu_pro","value":56.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-small-4","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-small-4","key":"ifbench","value":48.2,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/mistral-small-4","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-glimmer-30b","key":"aa_index","value":35,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/muse-glimmer","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-glimmer-30b","key":"gpqa_diamond","value":83.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/muse-glimmer","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-glimmer-30b","key":"hle","value":22,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/muse-glimmer","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-glimmer-30b","key":"scicode","value":43.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/muse-glimmer","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-glimmer-30b","key":"mmmu_pro","value":74.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/muse-glimmer","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-spark","key":"aa_index","value":44,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/muse-spark","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-spark","key":"gpqa_diamond","value":88.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/muse-spark","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-spark","key":"hle","value":40.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/muse-spark","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-spark","key":"scicode","value":51.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/muse-spark","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-spark","key":"tb_hard","value":45.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/muse-spark","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-spark","key":"tau2_telecom","value":91.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/muse-spark","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-spark","key":"mmmu_pro","value":80.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/muse-spark","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-spark","key":"ifbench","value":75.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/muse-spark","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-nano-30b-a3b","key":"aa_index","value":15,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-nano-30b-a3b","key":"gpqa_diamond","value":75.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-nano-30b-a3b","key":"hle","value":11.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-nano-30b-a3b","key":"aime_2025","value":91,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-nano-30b-a3b","key":"livecodebench","value":74.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-nano-30b-a3b","key":"scicode","value":29.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-nano-30b-a3b","key":"tb_hard","value":13.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-nano-30b-a3b","key":"tau2_telecom","value":40.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-nano-30b-a3b","key":"ifbench","value":71.1,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-super-120b-a12b","key":"aa_index","value":26,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-super-120b-a12b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-super-120b-a12b","key":"gpqa_diamond","value":80,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-super-120b-a12b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-super-120b-a12b","key":"hle","value":20.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-super-120b-a12b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-super-120b-a12b","key":"scicode","value":36,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-super-120b-a12b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-super-120b-a12b","key":"tb_hard","value":28.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-super-120b-a12b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-super-120b-a12b","key":"tau2_telecom","value":67.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-super-120b-a12b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-super-120b-a12b","key":"ifbench","value":71.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-super-120b-a12b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"aa_index","value":38,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-ultra-550b-a55b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"gpqa_diamond","value":86.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-ultra-550b-a55b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"hle","value":28.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-ultra-550b-a55b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"scicode","value":39.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-ultra-550b-a55b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"tb_hard","value":36.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-ultra-550b-a55b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"tau2_telecom","value":83.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-ultra-550b-a55b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"ifbench","value":81.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-ultra-550b-a55b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-2-4t-a95b","key":"aa_index","value":58,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-8-2-4t-a95b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-2-4t-a95b","key":"gpqa_diamond","value":93.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-8-2-4t-a95b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-2-4t-a95b","key":"hle","value":42.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-8-2-4t-a95b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-2-4t-a95b","key":"scicode","value":51.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-8-2-4t-a95b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-27b","key":"aa_index","value":52,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-8-27b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-27b","key":"gpqa_diamond","value":90.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-8-27b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-27b","key":"hle","value":33.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-8-27b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-27b","key":"scicode","value":44.7,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-8-27b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-27b","key":"mmmu_pro","value":76.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-8-27b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-coder-480b-a35b","key":"aa_index","value":18,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-coder-480b-a35b-instruct","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-coder-480b-a35b","key":"gpqa_diamond","value":61.8,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-coder-480b-a35b-instruct","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-coder-480b-a35b","key":"hle","value":4.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-coder-480b-a35b-instruct","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-coder-480b-a35b","key":"aime_2025","value":39.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-coder-480b-a35b-instruct","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-coder-480b-a35b","key":"livecodebench","value":58.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-coder-480b-a35b-instruct","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-coder-480b-a35b","key":"scicode","value":35.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-coder-480b-a35b-instruct","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-coder-480b-a35b","key":"tb_hard","value":18.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-coder-480b-a35b-instruct","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-coder-480b-a35b","key":"tau2_telecom","value":43.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-coder-480b-a35b-instruct","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-coder-480b-a35b","key":"ifbench","value":40.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/qwen3-coder-480b-a35b-instruct","as_of":"2026-08-28","evidence":"independent"},{"model_id":"step-3-7-flash","key":"aa_index","value":31,"unit":"index","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/step-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"step-3-7-flash","key":"gpqa_diamond","value":80.9,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/step-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"step-3-7-flash","key":"hle","value":21.4,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/step-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"step-3-7-flash","key":"scicode","value":40,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/step-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"step-3-7-flash","key":"tb_hard","value":35.6,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/step-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"step-3-7-flash","key":"tau2_telecom","value":98.5,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/step-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"step-3-7-flash","key":"mmmu_pro","value":75.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/step-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"step-3-7-flash","key":"ifbench","value":67.3,"unit":"percent","source":"Artificial Analysis (v4.1.1)","source_url":"https://artificialanalysis.ai/models/step-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-fable-5","key":"terminal_bench","value":84.6,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/claude-fable-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-fable-5","key":"tau3_banking","value":38.1,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/claude-fable-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"terminal_bench","value":44.2,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"tau3_banking","value":9.3,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/claude-4-5-haiku-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-opus-5","key":"terminal_bench","value":89.1,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/claude-opus-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-opus-5","key":"tau3_banking","value":42.1,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/claude-opus-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-sonnet-5","key":"terminal_bench","value":80.5,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/claude-sonnet-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-sonnet-5","key":"tau3_banking","value":37.3,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/claude-sonnet-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"command-a-plus","key":"terminal_bench","value":22.8,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/command-a-plus","as_of":"2026-08-28","evidence":"independent"},{"model_id":"command-a-plus","key":"tau3_banking","value":6,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/command-a-plus","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-flash","key":"terminal_bench","value":78.7,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/deepseek-v4-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-flash","key":"tau3_banking","value":39.4,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/deepseek-v4-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-pro","key":"terminal_bench","value":78.7,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/deepseek-v4-pro","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-pro","key":"tau3_banking","value":39.6,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/deepseek-v4-pro","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-1-pro","key":"terminal_bench","value":73.8,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/gemini-3-1-pro-preview","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-1-pro","key":"tau3_banking","value":21.4,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/gemini-3-1-pro-preview","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-5-flash-lite","key":"terminal_bench","value":53.6,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-5-flash-lite","key":"tau3_banking","value":17.5,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-7-flash","key":"terminal_bench","value":85.8,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/gemini-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-7-flash","key":"tau3_banking","value":32.8,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/gemini-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-26b-a4b","key":"terminal_bench","value":39,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/gemma-4-26b-a4b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-26b-a4b","key":"tau3_banking","value":12,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/gemma-4-26b-a4b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"terminal_bench","value":43.4,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/gemma-4-31b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"tau3_banking","value":14.8,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/gemma-4-31b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-2","key":"terminal_bench","value":77.9,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/glm-5-2","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-2","key":"tau3_banking","value":34.6,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/glm-5-2","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3-flash","key":"terminal_bench","value":84.3,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/glm-5-3-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3-flash","key":"tau3_banking","value":47.2,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/glm-5-3-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3","key":"terminal_bench","value":83.9,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/glm-5-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3","key":"tau3_banking","value":50.3,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/glm-5-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-5","key":"terminal_bench","value":84.3,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/gpt-5-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-5","key":"tau3_banking","value":39,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/gpt-5-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-luna","key":"terminal_bench","value":80.9,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/gpt-5-6-luna","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-luna","key":"tau3_banking","value":31.1,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/gpt-5-6-luna","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-sol","key":"terminal_bench","value":88,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/gpt-5-6-sol","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-sol","key":"tau3_banking","value":44.3,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/gpt-5-6-sol","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-terra","key":"terminal_bench","value":88,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/gpt-5-6-terra","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-terra","key":"tau3_banking","value":40.2,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/gpt-5-6-terra","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-120b","key":"terminal_bench","value":26.2,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/gpt-oss-120b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-120b","key":"tau3_banking","value":12.8,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/gpt-oss-120b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-20b","key":"terminal_bench","value":13.9,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/gpt-oss-20b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-20b","key":"tau3_banking","value":7,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/gpt-oss-20b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"granite-4-2-30b","key":"terminal_bench","value":26.6,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/granite-4-2-30b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"granite-4-2-30b","key":"tau3_banking","value":14.4,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/granite-4-2-30b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-3","key":"terminal_bench","value":39.7,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/grok-4-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-3","key":"tau3_banking","value":12.4,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/grok-4-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-5","key":"terminal_bench","value":81.6,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/grok-4-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-5","key":"tau3_banking","value":42.1,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/grok-4-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-6","key":"terminal_bench","value":88.4,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/grok-4-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-6","key":"tau3_banking","value":50.7,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/grok-4-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"hunyuan-hy3","key":"terminal_bench","value":64.4,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/hy3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"hunyuan-hy3","key":"tau3_banking","value":22.9,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/hy3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-6","key":"terminal_bench","value":65.9,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/kimi-k2-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-6","key":"tau3_banking","value":23.3,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/kimi-k2-6","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-7-code","key":"terminal_bench","value":67.4,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/kimi-k2-7-code","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-7-code","key":"tau3_banking","value":20.2,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/kimi-k2-7-code","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k3","key":"terminal_bench","value":85,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/kimi-k3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k3","key":"tau3_banking","value":46,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/kimi-k3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mimo-v2-5","key":"terminal_bench","value":63.7,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/mimo-v2-5-0424","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mimo-v2-5","key":"tau3_banking","value":8.7,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/mimo-v2-5-0424","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m2-7","key":"terminal_bench","value":55.4,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/minimax-m2-7","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m2-7","key":"tau3_banking","value":9.9,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/minimax-m2-7","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"terminal_bench","value":65.2,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/minimax-m3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"tau3_banking","value":15.3,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/minimax-m3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-large-3","key":"terminal_bench","value":12,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/mistral-large-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-large-3","key":"tau3_banking","value":5.8,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/mistral-large-3","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-medium-3-5","key":"terminal_bench","value":50.6,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/mistral-medium-3-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-medium-3-5","key":"tau3_banking","value":15.1,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/mistral-medium-3-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-small-4","key":"terminal_bench","value":21,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/mistral-small-4","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-small-4","key":"tau3_banking","value":4.9,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/mistral-small-4","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-glimmer-30b","key":"terminal_bench","value":51.7,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/muse-glimmer","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-glimmer-30b","key":"tau3_banking","value":23.5,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/muse-glimmer","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-spark","key":"terminal_bench","value":62.2,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/muse-spark","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-nano-30b-a3b","key":"terminal_bench","value":6.7,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-nano-30b-a3b","key":"tau3_banking","value":6,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-nano-30b-a3b-reasoning","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-super-120b-a12b","key":"terminal_bench","value":38.6,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-super-120b-a12b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-super-120b-a12b","key":"tau3_banking","value":10.3,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-super-120b-a12b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"terminal_bench","value":53.9,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-ultra-550b-a55b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"tau3_banking","value":14.2,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/nvidia-nemotron-3-ultra-550b-a55b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-2-4t-a95b","key":"terminal_bench","value":82,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/qwen3-8-2-4t-a95b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-2-4t-a95b","key":"tau3_banking","value":49.1,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/qwen3-8-2-4t-a95b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-27b","key":"terminal_bench","value":79.8,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/qwen3-8-27b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-27b","key":"tau3_banking","value":48,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/qwen3-8-27b","as_of":"2026-08-28","evidence":"independent"},{"model_id":"step-3-7-flash","key":"terminal_bench","value":39.3,"unit":"percent","source":"Artificial Analysis · Terminal-Bench 2.1","source_url":"https://artificialanalysis.ai/models/step-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"step-3-7-flash","key":"tau3_banking","value":12,"unit":"percent","source":"Artificial Analysis · τ³-Bench Banking","source_url":"https://artificialanalysis.ai/models/step-3-7-flash","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-fable-5","key":"arena_zh","value":1555,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-haiku-4-5","key":"arena_zh","value":1434,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-opus-5","key":"arena_zh","value":1564,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-sonnet-5","key":"arena_zh","value":1503,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-flash","key":"arena_zh","value":1480,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v4-pro","key":"arena_zh","value":1500,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-1-pro","key":"arena_zh","value":1531,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-5-flash-lite","key":"arena_zh","value":1485,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemini-3-7-flash","key":"arena_zh","value":1554,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-26b-a4b","key":"arena_zh","value":1481,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gemma-4-31b","key":"arena_zh","value":1476,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-2","key":"arena_zh","value":1515,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"glm-5-3","key":"arena_zh","value":1530,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-5","key":"arena_zh","value":1529,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-luna","key":"arena_zh","value":1478,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-sol","key":"arena_zh","value":1538,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-5-6-terra","key":"arena_zh","value":1517,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-120b","key":"arena_zh","value":1377,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"gpt-oss-20b","key":"arena_zh","value":1350,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-3","key":"arena_zh","value":1479,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-5","key":"arena_zh","value":1513,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"grok-4-6","key":"arena_zh","value":1539,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"hunyuan-hy3","key":"arena_zh","value":1463,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k2-6","key":"arena_zh","value":1526,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"kimi-k3","key":"arena_zh","value":1531,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mimo-v2-5","key":"arena_zh","value":1483,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m2-7","key":"arena_zh","value":1443,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"minimax-m3","key":"arena_zh","value":1467,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-large-3","key":"arena_zh","value":1430,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"mistral-medium-3-5","key":"arena_zh","value":1448,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-glimmer-30b","key":"arena_zh","value":1477,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"muse-spark","key":"arena_zh","value":1536,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-nano-30b-a3b","key":"arena_zh","value":1347,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-super-120b-a12b","key":"arena_zh","value":1399,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"nemotron-3-ultra-550b-a55b","key":"arena_zh","value":1458,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-27b","key":"arena_zh","value":1508,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-8-2-4t-a95b","key":"arena_zh","value":1538,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"qwen3-coder-480b-a35b","key":"arena_zh","value":1404,"unit":"elo","source":"LMArena Text · Chinese (style control)","source_url":"https://arena.ai/leaderboard/text/chinese","as_of":"2026-08-28","evidence":"independent"},{"model_id":"deepseek-v3-2","key":"arena_zh","value":1440,"unit":"elo","source":"LMArena Text · Chinese","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"qwen3-max","key":"arena_zh","value":1445,"unit":"elo","source":"LMArena Text · Chinese","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"kimi-k2-thinking","key":"arena_zh","value":1442,"unit":"elo","source":"LMArena Text · Chinese","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"glm-4-6","key":"arena_zh","value":1430,"unit":"elo","source":"LMArena Text · Chinese","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"gemini-3-pro","key":"arena_zh","value":1490,"unit":"elo","source":"LMArena Text · Chinese","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"claude-opus-4-5","key":"arena_zh","value":1462,"unit":"elo","source":"LMArena Text · Chinese","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"gpt-5","key":"arena_zh","value":1445,"unit":"elo","source":"LMArena Text · Chinese","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"qwen3-235b-a22b-thinking-2507","key":"arena_zh","value":1425,"unit":"elo","source":"LMArena Text · Chinese","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"claude-sonnet-4-5","key":"arena_zh","value":1440,"unit":"elo","source":"LMArena Text · Chinese","source_url":"https://lmarena.ai/leaderboard/text","as_of":"2025-12-20","evidence":"independent"},{"model_id":"claude-opus-5","key":"hle","value":53,"unit":"percent","source":"Artificial Analysis · Opus 5 评测文章","source_url":"https://artificialanalysis.ai/articles/opus-5","as_of":"2026-08-28","evidence":"independent"},{"model_id":"claude-opus-5","key":"terminal_bench","value":89,"unit":"percent","source":"Artificial Analysis · Opus 5 评测文章","source_url":"https://artificialanalysis.ai/articles/opus-5","as_of":"2026-08-28","evidence":"independent"}]}