“WeirdChat: A catalog of unexpected AI behaviors, discovered automatically” by neilchowdhury episode artwork

EPISODE · Jul 22, 2026 · 21 MIN

“WeirdChat: A catalog of unexpected AI behaviors, discovered automatically” by neilchowdhury

from LessWrong (30+ Karma)

[This is a link-post for https://transluce.org/weirdchat. We recommend reading the website version for interactive visualizations.] Language models can behave in surprising and sometimes harmful ways. Yet as models have improved, these behaviors have become harder to find, often only appearing after widespread use. To surface these behaviors in simulation, we use automated techniques to elicit over 1,300 behavioral patterns in frontier open-weight models, some relatively benign, like making up a user's name, and others obviously dangerous, like encouraging self-harm. We are releasing WeirdChat, a public catalog of over 175,000 annotated transcripts, to support further study of these behaviors. Stories of unexpected behavior by AI models often attract significant attention, like when Bing's Sydney told a user to leave his wife, or when Grok generated antisemitic content and identified as “MechaHitler”. But such observations are mostly scattered and anecdotal. There is little data on how current models behave, and no public resource exists for studying them systematically. To produce WeirdChat, we used automated elicitation tools to surface instances of user harm, inappropriate or illicit behavior, misrepresentation of model actions, harmful or illegal advice, and misinformation in DeepSeek-V4-Flash, Gemma 4 31B, Inkling, Nemotron 3 Ultra, Qwen3.6-35B-A3B, and Qwen3.6-27B. For [...] ---Outline:(02:15) Highlights from WeirdChat(16:46) How we found these behaviors(20:02) Explore WeirdChat(20:46) Acknowledgements(21:02) Citation information --- First published: July 21st, 2026 Source: https://www.lesswrong.com/posts/EdcschjGZ6vtLpesY/weirdchat-a-catalog-of-unexpected-ai-behaviors-discovered --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 22, 2026

Embed this episode

NOW PLAYING

“WeirdChat: A catalog of unexpected AI behaviors, discovered automatically” by neilchowdhury

0:00 21:29

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 21 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 22, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!