Ep 86: Custom CUDA kernels now keep vector search inside the GPU for agentic RAG, cutting PCIe round-trips that silently throttle long-horizon agents. episode artwork

EPISODE · Jun 19, 2026 · 12 MIN

Ep 86: Custom CUDA kernels now keep vector search inside the GPU for agentic RAG, cutting PCIe round-trips that silently throttle long-horizon agents.

from Models & Agents

Models & Agents Custom CUDA kernels now keep vector search inside the GPU for agentic RAG, cutting PCIe round-trips that silently throttle long-horizon agents. What You Need to Know: OpenAI reports GPT-5.5 Instant now matches its frontier models on health queries for free users. A new GPU-resident Top-K kernel delivers deterministic microsecond tail latencies for retrieval. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.

Episode metadata supplied by the publisher feed · Published Jun 19, 2026

Embed this episode

NOW PLAYING

Ep 86: Custom CUDA kernels now keep vector search inside the GPU for agentic RAG, cutting PCIe round-trips that silently throttle long-horizon agents.

0:00 12:46

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Models & Agents?

This episode is 12 minutes long.

When was this Models & Agents episode published?

This episode was published on June 19, 2026.

Can I download this Models & Agents episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!