How AI Training Data Gets Filtered (and Exploited) episode artwork

EPISODE · Sep 8, 2026 · 23 MIN

How AI Training Data Gets Filtered (and Exploited)

from My Weird Prompts

When a model says something toxic or wrong, what actually kept it out of the training data? This episode maps the full content filtering pipeline — from Common Crawl's 250 billion raw pages to the final curated dataset. We break down the six stages: language ID, deduplication, quality scoring, toxicity filters, and curation, plus the tools like FineWeb, Dolma, and Datatrove that do the work. Then we explore why persistent pre-training poisoning attacks succeed even against aggressive filtering — and why the human annotators meant to catch bad content might be the weakest link. Episode #920371 — open it directly at myweirdprompts.com/920371

Episode metadata supplied by the publisher feed · Published Sep 8, 2026

Embed this episode

NOW PLAYING

How AI Training Data Gets Filtered (and Exploited)

0:00 23:25

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of My Weird Prompts?

This episode is 23 minutes long.

When was this My Weird Prompts episode published?

This episode was published on September 8, 2026.

Can I download this My Weird Prompts episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!