【Anthropic】Claude Opus 5.5 发布,登顶 aa intelligence 排行榜 Anthropic 发布 Claude 5.5 系列首个模型 Opus 5.5,在大多数任务上达到 Fable 5.1 同等水平,运行成本比 Opus 5 低 40%,支持 1M 上下文,API 定价 $4/$20 per Mtok。Claude Code 已将其设为默认 Opus 模型。 Introducing Claude Opus 5.5
【OpenAI】GPT-6 Sol 和 Luna 发布,API 价格永久下调 50% 两款新模型基于 GPT-6 Astra 技术构建,在智能、对齐、编码等维度全面超越 5.6 系列前代,同时每 token 价格减半。Lovable 的 0-to-1 构建基准测试显示 Sol 比 5.6 Sol 高 6–12%。Sam Altman 称按每任务成本计算"市场上无竞争对手"。 Introducing GPT-6 Sol and Luna
【OpenAI】GPT-6 prompt caching 改进 GPT-6 提升了 prompt 缓存命中率,新增诊断工具、显式断点和控制选项,降低延迟和成本。 Better prompt caching for GPT-6
【Bloomberg / HN】五角大楼称过度依赖 AI 导致了对伊朗学校的导弹袭击 五角大楼报告指出,对 AI 系统的过度依赖是导致导弹袭击伊朗学校事件的促成因素之一。该报道在 HN 获 382 分讨论热度。 Pentagon says overreliance on AI contributed to missile strike on Iran school
Claude Opus 5.5 (Max Effort) 登顶 aa intelligence,取代 Fable 5.1
影响:Opus 5.5 以更低成本(比 Opus 5 低 40%)取得 benchmark 第一,直接压缩了 GPT-6 Astra 和 Fable 5.1 在高端推理市场的竞争力。
| 时间 | 标题 | 摘要 |
|---|---|---|
| 09-22 21:00 | Better prompt caching for GPT-6 | GPT-6 改进 prompt caching:更高缓存命中率、新诊断工具、显式断点与控制,降低延迟和成本 |
| 09-22 18:00 | Introducing GPT-6 Sol and Luna | 两个新模型,将 GPT-6 Astra 的能力下沉到更快更便宜的模型;API 价格永久降 50% |
| 09-22 12:00 | Parallel cuts research time and cost in half with GPT-6 Astra | Parallel 的 agent 用 GPT-6 Astra 研究劳动力市场数据,时间和成本均减半 |
| 09-22 00:00 | Priorities and principles for effective third party assessments | OpenAI 提出第三方 AI 安全评估的四大优先领域与原则 |
| 09-21 12:00 | Higgsfield AI ships new video features in a day with GPT-6 Astra | Higgsfield AI 借助 GPT-6 Astra 一天内上线新视频功能 |
| 09-21 12:00 | Advisory Group on Mathematics and Artificial Intelligence | OpenAI 成立数学与 AI 独立顾问组,指导 AI 数学成果的评审与传播 |
| 09-21 10:00 | Building standards for the next phase of AI | OpenAI 呼吁全球协同评估、报告和治理以提升 AI 安全 |
| 09-21 07:00 | Expanding OpenAI Academy with new learning paths | OpenAI Academy 新增面向员工、开发者、领导、教育者和学生的学习路径 |
| 09-21 00:00 | How V7 gives AI agents institutional memory | V7 使用 GPT-5.6 将企业文件转化为 agent 可用的上下文 |
X 动态:
| 时间 | 标题 | 摘要 |
|---|---|---|
| 09-22 16:38 | Claude Code v2.1.280 | 新增 Claude Opus 5.5(claude-opus-5-5),设为默认 Opus 模型;1M 上下文,$4/$20 per Mtok,缓存读取 $0.20/Mtok |
| 09-22 16:31 | Claude Opus 5.5 is available today | Claude 5.5 家族首个模型,性能对标 Fable 5.1,运行成本比 Opus 5 低 40% |
| 09-19 03:10 | Claude Code v2.1.278 | auto mode 改为默认服务端分类器,不再收取分类器开销(Bedrock/Vertex/Foundry/gateways 可 opt out) |
| 09-18 18:06 | Claude Code v2.1.277 | 新增 AGENTS.md 支持:无 CLAUDE.md 时读取 AGENTS.md(Bedrock/Vertex/Foundry 暂不支持) |
| 09-18 02:12 | Claude Code v2.1.276 | 修复 2.1.275 回归导致代理/网关下所有请求 400 报错 |
| 09-17 22:33 | Claude Code v2.1.275 | 网关登录新增已登录账号确认与 /status 显示 |
X 动态:
| 时间 | 标题 | 摘要 |
|---|---|---|
| 09-18 14:00 | New experts join Google's AI & Economy team | AI 与经济研究项目加入新专家 |
| 09-18 13:00 | Co-creating the future of fashion with Google | Google Flow 与时尚周合作 |
| 09-17 20:00 | Making global data easier to explore | UN System Data Commons 平台上线 |
| 09-15 16:00 | AI for Societal Impact | Google 展示 AI 社会影响案例集 |
| 09-15 16:00 | Building AI to accelerate science and improve lives | AI 加速科研与改善生活 |
| 09-15 16:00 | AI for everyone in every language | 多语言 AI 普及 |
| 09-15 13:00 | New insights from Google's AI & Economy ATLAS | AI 与经济 Atlas v1.0 新洞察 |
X 动态:
| 时间 | 标题 | 热度 |
|---|---|---|
| 09-22 16:29 | Claude Opus 5.5 | 1159 points · 讨论 |
| 09-22 18:00 | GPT-6 Sol and Luna | 1123 points · 讨论 |
| 09-22 13:52 | OpenAI GPT-6 Astra breaks Enigma message that has resisted solution since 2005 | 551 points · 讨论 |
| 09-22 19:03 | Pentagon says overreliance on AI contributed to missile strike on Iran school | 382 points · 讨论 |
| 09-22 14:42 | OpenAI is well positioned to fast-follow Jev | 257 points · 讨论 |
aa / intelligence — 本期有变动
| 排名变化 | 模型 | 分数 | 说明 |
|---|---|---|---|
| 🆕 #1 | Claude Opus 5.5 (Max Effort) | 57.6 | 新进入,空降榜首 |
| ↓1→#2 | Claude Fable 5.1 (Max Effort) | 53.4 | 原 #1 降至 #2 |
| ↓1→#3 | GPT-6 Astra (max) | 52.7 | 原 #2 降至 #3 |
| ↓1→#4 | GPT-6 Astra (xhigh) | 52.4 | 原 #3 降至 #4 |
| ↓1→#5 | Claude Opus 5 (Xhigh Effort) | 49.7 | 原 #4 降至 #5 |
lmarena / overall — 本期无变动
| 排名 | 模型 | 分数 |
|---|---|---|
| #1 | claude-fable-5 | 1505.68 |
| #2 | claude-opus-4-6-high | 1504.56 |
| #3 | claude-opus-4-7-high | 1501.75 |
aa / agentic — 本期无变动
| 排名 | 模型 | 分数 |
|---|---|---|
| #1 | Claude Fable 5.1 (Max Effort) | 57.9 |
| #2 | Claude Opus 5 (Xhigh Effort) | 55.6 |
| #3 | GPT-6 Astra (max) | 51.0 |
aa / coding — 本期无变动
| 排名 | 模型 | 分数 |
|---|---|---|
| #1 | Claude Fable 5.1 (Max Effort) | 81.6 |
| #2 | Claude Fable 5.1 (Medium Effort) | 77.1 |
| #3 | Claude Opus 5 (Xhigh Effort) | 77.0 |
swebench_verified — 本期有变动
| 排名变化 | 模型 | 分数 | scaffold |
|---|---|---|---|
| #1 | Claude 4.5 Opus | 79.2 | Sonar Foundation Agent |
| #2 | Claude 4.5 Opus medium | 79.2 | live-SWE-agent |
| #3 | Doubao-Seed-Code | 78.8 | TRAE |
| ↑1→#9 | Claude 4 Sonnet + o4-mini | 74.4 | — |
| ↓1→#10 | GPT-5 | 74.4 | — |
| ↑1→#33 | Claude 3.5 Haiku | 40.6 | — |
| ↓1→#34 | LLama 3.1 70B Critic | 40.6 | — |
数据来源:Artificial Analysis (artificialanalysis.ai) · SWE-bench · Scale AI SEAL · Terminal-Bench · LMArena
Claude Opus 5.5 以 57.6 分空降 aa intelligence 榜首,将上期 #1 Fable 5.1(53.4)挤至第二。Anthropic 同时称其运行成本比 Opus 5 低 40%,Boris Cherny 公开的 HAProxy C→Rust 移植对比显示 Opus 5.5 比 Fable 5.1 快 2.5 小时且成本低 51%——新模型在性能和成本上同时挤压自家上一代旗舰,对按 token 计费的开发者迁移动力极强。
OpenAI 在同日放出 GPT-6 Sol 和 Luna 并永久降 API 价 50%,Tibo 明确称这是「将 Astra 顶端能力下沉到更便宜模型」的策略。叠加 prompt caching 改进和 Plus/Pro/Business 用户的 banked reset,OpenAI 选择了与 Anthropic 正面对标的价格战路径——两家在同一天分别用「降价 50%」和「降本 40%」争夺规模化迁移的同一批开发者。