"More information about the dangerous capability evaluations we did with GPT-4 and Claude." by Beth Barnes episode artwork

EPISODE · Mar 21, 2023 · 14 MIN

"More information about the dangerous capability evaluations we did with GPT-4 and Claude." by Beth Barnes

from LessWrong (Curated & Popular)

https://www.lesswrong.com/posts/4Gt42jX7RiaNaxCwP/more-information-about-the-dangerous-capability-evaluationsCrossposted from the AI Alignment Forum. May contain more technical jargon than usual.This is a linkpost for https://evals.alignment.org/blog/2023-03-18-update-on-recent-evals/[Written for more of a general-public audience than alignment-forum audience. We're working on a more thorough technical report.]We believe that capable enough AI systems could pose very large risks to the world. We don’t think today’s systems are capable enough to pose these sorts of risks, but we think that this situation could change quickly and it’s important to be monitoring the risks consistently. Because of this, ARC is partnering with leading AI labs such as Anthropic and OpenAI as a third-party evaluator to assess potentially dangerous capabilities of today’s state-of-the-art ML models. The dangerous capability we are focusing on is the ability to autonomously gain resources and evade human oversight.We attempt to elicit models’ capabilities in a controlled environment, with researchers in-the-loop for anything that could be dangerous, to understand what might go wrong before models are deployed. We think that future highly capable models should involve similar “red team” evaluations for dangerous capabilities before the models are deployed or scaled up, and we hope more teams building cutting-edge ML systems will adopt this approach. The testing we’ve done so far is insufficient for many reasons, but we hope that the rigor of evaluations will scale up as AI systems become more capable.As we expected going in, today’s models (while impressive) weren’t capable of autonomously making and carrying out the dangerous activities we tried to assess. But models are able to succeed at several of the necessary components. Given only the ability to write and run code, models have some success at simple tasks involving browsing the internet, getting humans to do things for them, and making long-term plans – even if they cannot yet execute on this reliably.

Episode metadata supplied by the publisher feed · Published Mar 21, 2023

Embed this episode

https://www.lesswrong.com/posts/4Gt42jX7RiaNaxCwP/more-information-about-the-dangerous-capability-evaluations Crossposted from the AI Alignment Forum. May contain more technical jargon than usual. This is a linkpost for https://evals.alignment.org/blog/2023-03-18-update-on-recent-evals/ [Written for more of a general-public audience than alignment-forum audience. We're working on a more thorough technical report.] We believe that capable enough AI systems could pose very large risks to the ...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

"More information about the dangerous capability evaluations we did with GPT-4 and Claude." by Beth Barnes

0:00 14:16

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 14 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on March 21, 2023.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!