LLMs for Alignment Research: a safety priority? episode artwork

EPISODE · Apr 6, 2024 · 20 MIN

LLMs for Alignment Research: a safety priority?

from LessWrong (Curated & Popular)

A recent short story by Gabriel Mukobi illustrates a near-term scenario where things go bad because new developments in LLMs allow LLMs to accelerate capabilities research without a correspondingly large acceleration in safety research.This scenario is disturbingly close to the situation we already find ourselves in. Asking the best LLMs for help with programming vs technical alignment research feels very different (at least to me). LLMs might generate junk code, but you can keep pointing out the problems with the code, and the code will eventually work. This can be faster than doing it myself, in cases where I don't know a language or library well; the LLMs are moderately familiar with everything.When I try to talk to LLMs about technical AI safety work, however, I just get garbage.I think a useful safety precaution for frontier AI models would be to make them more useful for [...]The original text contained 8 footnotes which were omitted from this narration. --- First published: April 4th, 2024 Source: https://www.lesswrong.com/posts/nQwbDPgYvAbqAmAud/llms-for-alignment-research-a-safety-priority --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Apr 6, 2024

Embed this episode

A recent short story by Gabriel Mukobi illustrates a near-term scenario where things go bad because new developments in LLMs allow LLMs to accelerate capabilities research without a correspondingly large acceleration in safety research. This scenario is disturbingly close to the situation we already find ourselves in. Asking the best LLMs for help with programming vs technical alignment research feels very different (at least to me). LLMs might generate junk code, but you can keep pointing o...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

LLMs for Alignment Research: a safety priority?

0:00 20:46

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 20 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on April 6, 2024.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!