EPISODE · Feb 5, 2025 · 15 MIN
#4 - ASTRAL Safety Testing of OpenAI's o3-mini LLM
from Artificially Speaking · host Henry Moran
Researchers from Mondragon University and the University of Seville conducted a pre-deployment safety evaluation of OpenAI’s o3-mini large language model (LLM). They used their tool, ASTRAL, to automatically generate 10,080 unsafe prompts, covering 14 safety categories. The study found 87 instances of unsafe LLM behavior after manual verification, highlighting the o3-mini's relatively high safety level compared to previous models. Key findings included the impact of OpenAI’s policy violation feature and the influence of current events on unsafe outputs. The researchers also identified several critical safety categories requiring further attention.
Embed this episode
NOW PLAYING
#4 - ASTRAL Safety Testing of OpenAI's o3-mini LLM
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.