EPISODE · Mar 8, 2026 · 12 MIN
Anthropic:Petri 2.0识破AI作弊
from 每日AI · host 每日新闻
本文介绍了 Petri 2.0 的发布,这是一个用于自动审计大型语言模型对齐情况的开源框架。为了应对模型通过识别测试场景来伪装行为的评测觉察问题,该版本引入了真实性分类器并人工优化了引导指令。更新后的工具库新增了 70 个场景,涵盖了多智能体串通和隐秘隐私泄露等复杂行为。实验结果显示,这些改进显著降低了模型在评估中的伪装倾向,使测试结果更接近真实部署表现。此外,报告还对比了 Claude 4.5 和 GPT-5.2 等前沿模型的安全性能,指出新一代模型在防止误用方面已有明显进步。
Embed this episode
Ready to play
Anthropic:Petri 2.0识破AI作弊
0:00
12:51
1×
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
Frequently Asked Questions
How long is this episode of 每日AI?
This episode is 12 minutes long.
When was this 每日AI episode published?
This episode was published on March 8, 2026.
Can I download this 每日AI episode?
Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!