“AI companies aren’t really using external evaluators” by Zach Stein-Perlman episode artwork

EPISODE · May 24, 2024 · 7 MIN

“AI companies aren’t really using external evaluators” by Zach Stein-Perlman

from LessWrong (Curated & Popular)

New blog: AI Lab Watch. Subscribe on Substack.Many AI safety folks think that METR is close to the labs, with ongoing relationships that grant it access to models before they are deployed. This is incorrect. METR (then called ARC Evals) did pre-deployment evaluation for GPT-4 and Claude 2 in the first half of 2023, but it seems to have had no special access since then.[1] Other model evaluators also seem to have little access before deployment.Frontier AI labs' pre-deployment risk assessment should involve external model evals for dangerous capabilities.[2] External evals can improve a lab's risk assessment and—if the evaluator can publish its results—provide public accountability.The evaluator should get deeper access than users will get. To evaluate threats from a particular deployment protocol, the evaluator should get somewhat deeper access than users will — then the evaluator's failure to elicit dangerous capabilities is stronger evidence [...]The original text contained 5 footnotes which were omitted from this narration.--- First published: May 24th, 2024 Source: https://www.lesswrong.com/posts/WjtnvndbsHxCnFNyc/ai-companies-aren-t-really-using-external-evaluators --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published May 24, 2024

Embed this episode

New blog: AI Lab Watch. Subscribe on Substack. Many AI safety folks think that METR is close to the labs, with ongoing relationships that grant it access to models before they are deployed. This is incorrect. METR (then called ARC Evals) did pre-deployment evaluation for GPT-4 and Claude 2 in the first half of 2023, but it seems to have had no special access since then.[1] Other model evaluators also seem to have little access before deployment. Frontier AI labs' pre-deployment risk assessm...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

“AI companies aren’t really using external evaluators” by Zach Stein-Perlman

0:00 7:42

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 7 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on May 24, 2024.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!