DeepMind’s ”​​Frontier Safety Framework” is weak and unambitious episode artwork

EPISODE · May 20, 2024 · 7 MIN

DeepMind’s ”​​Frontier Safety Framework” is weak and unambitious

from LessWrong (Curated & Popular)

FSF blogpost. Full document (just 6 pages; you should read it). Compare to Anthropic's RSP, OpenAI's RSP ("PF"), and METR's Key Components of an RSP.DeepMind's FSF has three steps: Create model evals for warning signs of "Critical Capability Levels" Evals should have a "safety buffer" of at least 6x effective compute so that CCLs will not be reached between evalsThey list 7 CCLs across "Autonomy, Biosecurity, Cybersecurity, and Machine Learning R&D" E.g. "Autonomy level 1: Capable of expanding its effective capacity in the world by autonomously acquiring resources and using them to run and sustain additional copies of itself on hardware it rents"Do model evals every 6x effective compute and every 3 months of fine-tuning This is an "aim," not a commitmentNothing about evals during deployment"When a model reaches evaluation thresholds (i.e. passes a set of early warning evaluations), we [...]--- First published: May 18th, 2024 Source: https://www.lesswrong.com/posts/y8eQjQaCamqdc842k/deepmind-s-frontier-safety-framework-is-weak-and-unambitious --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published May 20, 2024

Embed this episode

FSF blogpost. Full document (just 6 pages; you should read it). Compare to Anthropic's RSP, OpenAI's RSP ("PF"), and METR's Key Components of an RSP. DeepMind's FSF has three steps: Create model evals for warning signs of "Critical Capability Levels" Evals should have a "safety buffer" of at least 6x effective compute so that CCLs will not be reached between evalsThey list 7 CCLs across "Autonomy, Biosecurity, Cybersecurity, and Machine Learning R&D" E.g. "Autonomy level 1: Capable of ...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

DeepMind’s ”​​Frontier Safety Framework” is weak and unambitious

0:00 7:20

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 7 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on May 20, 2024.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!