EPISODE · Jul 15, 2026 · 21 MIN
Why Speech-to-Text Still Fails at Its Own Name
from My Weird Prompts
When OpenAI's Whisper transcribed its own name as "Wispr," it exposed the fundamental flaw in speech-to-text: models that hear perfectly but understand nothing. This episode unpacks why homophone errors, dropped negations, and hallucinated punctuation survive even low word-error rates — and explores two competing architectural solutions. We compare the two-pass pipeline (Whisper + LLM cleanup) against unified multimodal models like GPT-4o that process audio and reasoning in a single pass. Which approach actually eliminates the need for human review? And what are the hidden failure modes of each? If you dictate more than a few hundred words a day, this episode will change how you think about voice input. Episode #494803 — open it directly at myweirdprompts.com/494803
Embed this episode
NOW PLAYING
Why Speech-to-Text Still Fails at Its Own Name
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.