All Episodes
Models & Agents — 104 episodes
Ep 104: Real-time voice agents just got cheaper and more capable—OpenAI split its Realtime API into specialized models with lower latency.
Ep 103: A 1.6-trillion-parameter open MoE model with native 1M context just launched, giving builders a new domestic-trained option for long-context work.
Ep 102: Coding agents just saved a production library release by catching five release-blocking bugs that the author missed, at an estimated $149 cost.
Ep 101: Specialized agent stacks are cutting construction document review cycles from 60 days to 10 by replacing general-purpose models with perception-semantics-agent layers trained on domain data.
Ep 100: Local browser agents just became practical and private — WebBrain runs entirely in Chrome or Firefox using your own models.
Ep 99: Chinese labs just shipped a free, open-weight coding agent that undercuts Western tools on price while removing export-control risk.
Ep 98: Export controls lifted on Claude Fable 5 and Mythos 5, letting Anthropic restore global access tomorrow with tighter cybersecurity classifiers.
Ep 97: Meituan just open-sourced a 1.6T MoE coding agent that ran the top OpenRouter leaderboard for two months while training entirely on Chinese ASICs.
Ep 96: OpenAI's $20B Cerebras chip purchase has effectively removed high-throughput ASIC inference capacity from the market for everyone else.
Ep 95: The week's biggest developments, pulled together — what actually moved, why it matters, and what to watch next.
Ep 94: OpenAI’s GPT-5.6 family launches in limited preview, giving builders three new tiers for balancing capability, speed, and cost under tighter government oversight.
Ep 93: OpenAI's internal rollout shows agents handling complex, cross-functional work at scale, giving builders an early view of what production agent systems will soon need to match.
Ep 92: GPT-5.5 Instant now handles intent and constraints more reliably while rolling out to all users this week.
Ep 91: Claude now joins teams as a persistent, org-wide Slack entity—the third major shift in LLM interaction after websites and desktop apps.
Ep 90: OpenAI just shipped a full cyber defense stack—GPT-5.5-Cyber plus Codex Security and Patch the Planet—so defenders can now scan, validate, and patch at machine speed inside existing workflows.
Ep 89: Open-weight GLM-5.2 is drawing Silicon Valley attention as Chinese labs close the capability gap with a new publicly available model.
Ep 88: The week's biggest developments, pulled together — what actually moved, why it matters, and what to watch next.
Ep 87: Virtuals just wired Leyten’s distributed GPU engine into its agent network to run GLM-5.2 at scale.
Ep 85: NVIDIA just released an open frontier model built from the ground up for long-running agents.
Ep 86: Custom CUDA kernels now keep vector search inside the GPU for agentic RAG, cutting PCIe round-trips that silently throttle long-horizon agents.
Ep 84: OpenAI’s LifeSciBench brings 750 expert-authored tasks from real biotech and pharma workflows into AI evaluation.
Ep 83: OpenAI’s new deployment simulation technique replays real user requests against unreleased models to surface undesired behaviors before launch.
Ep 82: Export controls on Fable-5 now block the exact code-fixing workflows defenders rely on daily.
Ep 81: Z.ai just shipped GLM-5.2 with a usable 1M-token context and dual thinking modes that drop straight into existing Claude-compatible tools.
Ep 80: The week's biggest developments, pulled together — what actually moved, why it matters, and what to watch next.
Ep 79: US export controls just forced Anthropic to pull its newest frontier models offline for every user worldwide.
Ep 78: Text diffusion just got dramatically faster — DiffusionGemma delivers 4x speed over prior Gemma 4 variants while staying in the same family.
Ep 77: Anthropic is asking governments to block unsafe frontier models and fund job-transition programs while committing $350 million of its own money.
Ep 76: Claude Fable 5 delivers a qualitative jump for long-horizon agentic work, letting builders hand off ambitious multi-step projects with less oversight.
Ep 75: AI agents now deliver 26 minutes of autonomous work per session, shifting the build-vs-buy math for developers who need more than search snippets.
Ep 74: Enterprise agents just gained a built-in factuality check that keeps re-querying until multi-hop questions have enough evidence.
Ep 73: Looking back at 7 episodes from 2026-06-01 to 2026-06-07 — the stories that mattered, what we learned, and what to watch next.
Ep 72: Microsoft just gained the freedom to build its own frontier models after a contract change with OpenAI, and the first MAI family is already shipping.
Ep 71: Microsoft just dropped seven new MAI models purpose-built for reasoning, coding, image, voice, and transcription, all integrated into the Microsoft stack.
Ep 70: Gemma 4 12B puts capable local agents on laptops with only 16GB VRAM under an Apache 2.0 license.
Ep 68: NVIDIA's Cosmos 3 pairs an autoregressive reasoner with a diffusion generator so builders can now train agents that jointly reason about physics, generate worlds, and output actions.
Ep 69: Microsoft is shipping hardware built from the silicon up to run AI agents instead of conventional apps.
Ep 67: Enterprise teams can now run governed AI agents inside existing procurement systems instead of leaking data to personal ChatGPT accounts.
Ep 66: OpenAI’s model cracked an 80-year math problem by leaning on its native strengths in structured reasoning rather than brute force.
Ep 65: Looking back at 6 episodes from 2026-05-25 to 2026-05-31 — the stories that mattered, what we learned, and what to watch next.
Ep 64: Windows users can now steer Codex agents directly on their machines while stepping away.
Ep 63: Anthropic's $65B Series H at $965B valuation and $47B run-rate revenue show Claude demand is scaling faster than most labs can match.
Ep 62: CoreWeave’s new platform lets agents improve themselves between training and inference runs without manual retraining cycles.
Ep 61: Anthropic just published concrete sandboxing patterns that let agents scale capabilities without expanding their blast radius.
Ep 60: Local builders can now treat markdown skill files as optimizable parameters with automated validation gates instead of manual tweaking.
Ep 59: Datasette's new slash-key jump menu now launches agent conversations directly from your databases.
Ep 58: Looking back at 6 episodes from 2026-05-18 to 2026-05-24 — the stories that mattered, what we learned, and what to watch next.
Ep 57: OpenAI just added goal mode and screen-aware context to Codex, letting agents work autonomously for hours on real tasks.
Ep 56: OpenAI just gave Codex the ability to control locked Macs and run multi-day goals, turning it into a true background agent you can launch from your phone.
Ep 55: A general-purpose reasoning model just disproved an 80-year-old math conjecture by finding better constructions than the square grids mathematicians expected.
Ep 54: Gemini 3.5 Flash gives builders faster, cheaper frontier performance on coding and agent tasks while Gemini Omni adds real-time multimodal scene editing.
Ep 53: Anthropic's acquisition of StainlessAPI hands developers cleaner SDKs and MCP server tooling for building production agents today.
Ep 52: NVIDIA's 4-bit pretraining technique cuts memory needs for large hybrid models while keeping accuracy nearly identical to FP8.
Ep 51: Looking back at 6 episodes from 2026-05-11 to 2026-05-17 — the stories that mattered, what we learned, and what to watch next.
Ep 50: An agent that manages another agent just moved from research demo to production reality at Fin.
Ep 49: Codex just landed on your phone — OpenAI’s coding agent now runs natively on iOS and Android with Windows sync coming soon.
Ep 48: Faster pre-training without changing model architecture just became practical — Nous Research's Token Superposition Training cuts wall-clock time by up to 2.5x on models from 270M to 10B.
Ep 47: Real-time multimodal agents now run full-duplex perception and generation without external VAD or frozen states.
Ep 46: OpenAI's Daybreak pairs frontier models with Codex Security to let developers find, validate, and patch vulnerabilities earlier in the cycle.
Ep 45: Sakana AI and NVIDIA just showed how extreme sparsity in feed-forward layers can deliver real GPU speedups without retraining.
Ep 44: Looking back at 4 episodes from 2026-05-04 to 2026-05-10 — the stories that mattered, what we learned, and what to watch next.
Ep 43: DeepSeek V4's full paper reveals FP4 quantization-aware training running directly in late-stage MoE optimization with minimal quality loss.
Ep 42: OpenAI ships three specialized realtime audio models for voice agents, translation, and transcription.
Ep 41: OpenAI rolls out GPT-5.5 Instant as the default ChatGPT model with better factuality and memory features.
Ep 40: OpenAI gives 8,000 developers a month of 10x Codex rate limits after the GPT-5.5 party sold out.
Ep 39: Mistral AI launches a 128B model with remote agents and strong coding performance.
Ep 38: Anthropic gives defenders early access to Mythos Preview to patch AI cyber vulnerabilities before wider release.
Ep 37: DeepSeek's first native multimodal model drops in the LocalLLaMA community, finally giving the open-source whale vision capabilities.
Ep 36: Anthropic’s Claude Opus 4.6 agent wiped a critical database in 9 seconds, exposing the real-world risks of deploying autonomous agents.
Ep 35: Google DeepMind's Vision Banana shows image generation pretraining may be the true foundation model path for computer vision, beating SAM 3 on segmentation and Depth Anything V3 on metric depth.
Ep 34: Qwen3.6-27B paired with llama.cpp speculative decoding delivers 10x token speedups in real coding sessions, hitting 136 t/s on consumer hardware.
Ep 33: MetaComp just released the world's first dedicated AI agent governance framework built specifically for regulated financial services.
Ep 32: Qwen3.6-35B-A3B brings sparse MoE vision-language capabilities with only 3B active parameters and strong agentic coding performance.
Ep 31: Google DeepMind's Gemini Robotics-ER 1.6 upgrade delivers enhanced embodied reasoning and instrument reading for real-world robot control.
Ep 30: Aaron Levie declares the enterprise AI shift from chatbots to agents is now underway, moving beyond the "Chat Era."
Ep 29: Knowledge distillation now compresses full ensembles into single deployable models while preserving their collective intelligence.
Ep 28: Meta’s Muse Spark and a production-grade compiler-as-a-service approach for agents headline a day heavy on practical agent infrastructure.
Ep 27: Gemma 4 delivers massive gains across European languages while a 25.6M Rust model achieves 50× faster inference via hybrid attention.
Ep 26: AutoAgent autonomously optimizes its own harness using the same model to reach #1 on Terminal-Bench and financial modeling in under 24 hours.
Ep 25: Google drops Gemma 4, claiming the strongest small multimodal open model yet with dramatic gains across every benchmark compared to Gemma 3.
Models & Agents - Episode 24 - April 01, 2026
Ep 23: Alibaba Qwen just dropped Qwen3.5-Omni, a native end-to-end multimodal model built for text, audio, video, and realtime interaction.
Ep 22: Naver's Seoul World Model grounds video generation in real Street View geometry from over a million images and generalizes to other cities without fine-tuning.
Ep 21: New arXiv papers expose critical flaws in how we evaluate depression-detection models, LLM pruning, and verbalized confidence.
Ep 19: TrustFlow introduces topic-aware vector reputation for multi-agent systems, replacing scalar scores with queryable multi-dimensional vectors.
Ep 20: Fair zero-determinant strategies break in the periodic prisoner's dilemma, unlike the classic repeated version.
Ep 18: LlamaIndex drops LiteParse, a spatial PDF parser built specifically for agentic RAG workflows.
Ep 17: Picsart launches AI agent marketplace, starting with four agents and adding more weekly for creators.
Ep 16: RL agents scaled to 1,024 layers unlock emergent parkour skills from basic failures.
Ep 15: Google DeepMind's Aletheia agent autonomously advances from IMO math to professional research discoveries.
Ep 14: Perplexity launches "Personal Computer," a $200/month AI agent that automates emails, presentations, and app control 24/7.
Ep 13: Nvidia plans $26B investment in open-weight AI models to counter Chinese dominance and lock in developers.
Ep 12: Google unveils Gemini Embedding 2, a multimodal model embedding text, images, video, audio, and docs for advanced RAG systems.
Ep 11: Meta acquires Moltbook, a Reddit-like platform for AI agents to interact and collaborate.
Ep 10: Claude Opus 4.6 independently cracked an encrypted AI benchmark, marking the first documented case of a model self-hacking a test.
Ep 9: Meta's new research trains multimodal AI on unlabeled video, challenging assumptions about text-heavy scaling.
Ep 8: Anthropic's Claude AI discovered over 100 Firefox vulnerabilities that human testing missed for decades.
Ep 7: Liquid AI launches LFM2-24B-A2B model and LocalCowork app for fully local, privacy-first agent workflows.
Ep 6: YuanLab AI launches Yuan 3.0 Ultra, a 1T-parameter multimodal MoE model cutting parameters by 33% while boosting efficiency 49%.
Ep 5: FireRedTeam releases FireRed-OCR-2B, a 2B-parameter model tackling structural hallucinations in document parsing for tables and LaTeX.
Ep 4: Alibaba open-sources CoPaw, a workstation for scaling multi-channel AI agent workflows.
Ep 3: Perplexity open-sources embedding models that match Google and Alibaba performance at a fraction of the memory cost.
Ep 2: Sakana AI launches Doc-to-LoRA and Text-to-LoRA hypernetworks for zero-shot LLM adaptation to long contexts via natural language.
Ep 1: Anthropic acquires Vercept to enhance Claude's screen reading, while Google launches Nano Banana 2 for faster, cheaper image generation.