“AGI Safety and Alignment at Google DeepMind:A Summary of Recent Work ” by Rohin Shah, Seb Farquhar, Anca Dragan episode artwork

EPISODE · Aug 21, 2024 · 18 MIN

“AGI Safety and Alignment at Google DeepMind:A Summary of Recent Work ” by Rohin Shah, Seb Farquhar, Anca Dragan

from LessWrong (Curated & Popular)

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.We wanted to share a recap of our recent outputs with the AF community. Below, we fill in some details about what we have been working on, what motivated us to do it, and how we thought about its importance. We hope that this will help people build off things we have done and see how their work fits with ours. Who are we?We’re the main team at Google DeepMind working on technical approaches to existential risk from AI systems. Since our last post, we’ve evolved into the AGI Safety & Alignment team, which we think of as AGI Alignment (with subteams like mechanistic interpretability, scalable oversight, etc.), and Frontier Safety (working on the Frontier Safety Framework, including developing and running dangerous capability evaluations). We’ve also been growing since our last post: by 39% last year [...] ---Outline:(00:32) Who are we?(01:32) What have we been up to?(02:16) Frontier Safety(02:38) FSF(04:05) Dangerous Capability Evaluations(05:12) Mechanistic Interpretability(08:54) Amplified Oversight(09:23) Theoretical Work on Debate(10:32) Empirical Work on Debate(11:37) Causal Alignment(12:47) Emerging Topics(14:57) Highlights from Our Collaborations(17:07) What are we planning next?--- First published: August 20th, 2024 Source: https://www.lesswrong.com/posts/79BPxvSsjzBkiSyTq/agi-safety-and-alignment-at-google-deepmind-a-summary-of --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Aug 21, 2024

Embed this episode

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.We wanted to share a recap of our recent outputs with the AF community. Below, we fill in some details about what we have been working on, what motivated us to do it, and how we thought about its importance. We hope that this will help people build off things we have done and see how their work fits with ours. Who are we? We’re the main team at Google DeepMind working on technical approaches to existentia...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

“AGI Safety and Alignment at Google DeepMind:A Summary of Recent Work ” by Rohin Shah, Seb Farquhar, Anca Dragan

0:00 18:39

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 18 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on August 21, 2024.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!