EPISODE · Feb 24, 2026 · 16 MIN
METR 衡量AI完成长任务能力
from 每日AI · host 每日新闻
这项研究介绍了一种评估人工智能代理能力的新方法,即通过时间跨度(Time Horizon)来衡量模型能够独立完成多长时长的任务。研究人员利用 HCAST 和 SWAA 等基准测试,对比了 AI 与资深人类专业人士在网络安全、机器学习和软件工程等领域的表现。数据表明,AI 能够成功处理的任务时长正在迅速增加,其时间跨度大约每 212 天翻一倍。虽然当前顶尖模型(如 o1)在处理复杂或“混乱”环境时仍面临挑战,但其鲁棒性和纠错能力已显著提升。基于目前的增长趋势,该报告预测到 2029 年末,人工智能可能具备独立完成人类一个月工作量的能力。这套量化指标为人工智能治理和预防潜在的灾难性风险提供了重要的科学依据。
Embed this episode
Ready to play
METR 衡量AI完成长任务能力
0:00
16:38
1×
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
Frequently Asked Questions
How long is this episode of 每日AI?
This episode is 16 minutes long.
When was this 每日AI episode published?
This episode was published on February 24, 2026.
Can I download this 每日AI episode?
Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!