EPISODE · Jun 19, 2026 · 12 MIN
Ep 86: Custom CUDA kernels now keep vector search inside the GPU for agentic RAG, cutting PCIe round-trips that silently throttle long-horizon agents.
from Models & Agents
Models & Agents Custom CUDA kernels now keep vector search inside the GPU for agentic RAG, cutting PCIe round-trips that silently throttle long-horizon agents. What You Need to Know: OpenAI reports GPT-5.5 Instant now matches its frontier models on health queries for free users. A new GPU-resident Top-K kernel delivers deterministic microsecond tail latencies for retrieval. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production. 📝 Full show notes, transcript & sources: read the episode page 🌐 Part of the Nerra Network — explore every show at nerranetwork.com.
Embed this episode
NOW PLAYING
Ep 86: Custom CUDA kernels now keep vector search inside the GPU for agentic RAG, cutting PCIe round-trips that silently throttle long-horizon agents.
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.