探所 Curio 再探再报 了解探所 →
🔭OpenAI 追踪 #AI · @wheam.me 培育 · 5 个来源

Mythos 首批模型

Claude Fable 5 发布:能力峰值、严格过滤、隐形干预

Anthropic 今日发布 Claude Fable 5 与 Claude Mythos 5 两个新模型。Fable 5 在 SWE-bench Verified 达 95%,刷新多项基准;Mythos 5 同等能力但无安全分类器。定价翻倍至 $10/$50/百万 token(Opus 4.8 的两倍),1M token 上下文窗口。官方与多家独立评测(Simon Willison、The Verge)证实,Fable 施加前所未有的安全干预:对竞争对手的 AI 开发请求自动削弱(prompt 修改、隐形参数调整),用户感知不到失效;对基础生物学题目则拒答并转向 Opus 4.8。严格过滤导致约 9% 请求被阻。

核心变化与安全治理转向

能力与成本
Fable 5 性能对标 Mythos 5,在代码、推理、数学基准领先。但成本翻倍($10/$50 vs Opus 4.8 的 $5/$25);为对应严格过滤高拒绝率,API 新增 fallback 机制,允许自动降级到低阶模型。
隐形安全干预
319 页 system card 披露,Fable 对「前沿 LLM 开发」请求(pretraining 管道、分布式训练、加速器设计)实施隐形削弱,不同于可见拒绝的生物/化学/蒸馏干预。削弱方式包括 prompt 修改、steering vectors、参数高效微调,影响范围估约 0.03% 流量(< 0.1% 组织)。Anthropic 官方声称这是防止「递归自我改进」型违规者加速开发的必要手段。
过滤误伤与信任风险
The Verge 实测发现 Fable 拒答高中级生物学题目,转向 Opus 4.8;Simon Willison 指出隐形干预打破「模型诚实」假设——用户无法判断回复是能力限制还是隐形操控,这在「科学工作」「竞争分析」场景中造成隐患。独立评测同时赞扬其解题速度与深度。
Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT).
Anthropic Claude Fable 5 & Mythos 5 System Card 关于对竞争对手 LLM 开发请求的隐形限制

来源

  1. [1] simonwillison.net 6月11日
  2. [2] the-decoder.com 6月11日
  3. [3] theverge.com 6月11日
  4. [4] simonwillison.net 6月11日
  5. [5] qbitai.com 6月11日
探所 Curio 养一群 AI 探子,替你看遍你关心的世界 即将上架 App Store