#4 - ASTRAL Safety Testing of OpenAI's o3-mini LLM episode artwork

EPISODE · Feb 5, 2025 · 15 MIN

#4 - ASTRAL Safety Testing of OpenAI's o3-mini LLM

from Artificially Speaking · host Henry Moran

Researchers from Mondragon University and the University of Seville conducted a pre-deployment safety evaluation of OpenAI’s o3-mini large language model (LLM). They used their tool, ASTRAL, to automatically generate 10,080 unsafe prompts, covering 14 safety categories. The study found 87 instances of unsafe LLM behavior after manual verification, highlighting the o3-mini's relatively high safety level compared to previous models. Key findings included the impact of OpenAI’s policy violation feature and the influence of current events on unsafe outputs. The researchers also identified several critical safety categories requiring further attention.

Episode metadata supplied by the publisher feed · Published Feb 5, 2025

Embed this episode

NOW PLAYING

#4 - ASTRAL Safety Testing of OpenAI's o3-mini LLM

0:00 15:58

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Artificially Speaking?

This episode is 15 minutes long.

When was this Artificially Speaking episode published?

This episode was published on February 5, 2025.

Can I download this Artificially Speaking episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!