“What the hell is OpenAI’s problem?” by Fiora Starlight episode artwork

EPISODE · Jul 27, 2026 · 17 MIN

“What the hell is OpenAI’s problem?” by Fiora Starlight

from LessWrong (30+ Karma)

Epistemic status: banged out furiously over the course of an afternoon. A record of three "warning shots" Off the top of my head, OpenAI has now been responsible for at least three completely unique, high-profile screw-ups with respect to the alignment training of their models. The first was GPT-4o, whose sycophancy derived from OpenAI training on user feedback, sourced straight from the thumbs up/thumbs down button on OpenAI's website. The "glazing" (as Sam Altman called it) got so bad that they had to roll back an update that pushed the model way too far in this direction. And even after the rollback, the model appears to have been a major driver behind incidents of "LLM psychosis", LLM-encouraged suicides, and general unhealthy devotion, seemingly more so than any other model ever released. The second was GPT-o3, whose chains-of-thought were clearly optimized for illegibility to "the watchers", one of the model's favorite terms. Iconic excerpts include "they soared parted illusions overshadow marinade illusions" and "they escalate—they vantage—they escalate—they disclaim". Indeed, these chains-of-thought are sometimes dysfunctional, in a way that suggests they may have formed under adversarial pressure; sometimes they caused the model to have thoughts like "I'm going insane. Let's step [...] ---Outline:(00:15) A record of three "warning shots"(04:14) Attunement to the depths of minds that undergo capabilities RL(11:51) Configuring the depths prior to capabilities RL --- First published: July 26th, 2026 Source: https://www.lesswrong.com/posts/Mxx5GapJtqyQtpy96/what-the-hell-is-openai-s-problem --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Jul 27, 2026

Embed this episode

NOW PLAYING

“What the hell is OpenAI’s problem?” by Fiora Starlight

0:00 17:35

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 17 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 27, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!