EPISODE · May 22, 2026 · 38 MIN
Self-Hosting AI: Scaling Is the Real Problem
from Domesticating AI · host SoyPete Tech
AI is easy to use — but hard to scale.In this episode of Domesticating AI, we’re joined by Daniel Dowler (Red Hat) to break down what actually happens when you move from calling APIs to running AI systems yourself.Recorded on April 21stMost developers interact with AI through APIs — fast, simple, and pay-per-token. But behind the scenes, those systems rely on GPU scheduling, batching, and infrastructure that doesn’t behave like traditional software.We cover:Why GPU scaling is fundamentally different from CPU scalingWhy tools like vLLM are becoming the default for high-performance inferenceHow Ray and Kubernetes fit into real-world AI systemsWhat parallelism (tensor, data, expert) actually means in practiceWhen self-hosting AI makes senseWhen APIs are still the better choiceClaude Opus 4.7https://www.anthropic.com/news/claude-opus-4-7Qwen 3.6 (Alibaba)https://qwen.ai/researchKimi K2.6 (community discussion)https://www.reddit.com/r/LocalLLaMA/s/kvRWb7uJgMvLLM → https://github.com/vllm-project/vllmRay → https://github.com/ray-project/rayKubernetes → https://kubernetes.ioKueue → https://kueue.sigs.k8s.ioLiteLLM → https://github.com/BerriAI/litellmKServe → https://kserve.github.ioDaniel DowlerPlatform engineer at Red Hat focused on Kubernetes and AI infrastructure. Daniel works on how modern systems support real workloads, including GPU scheduling, distributed inference, and scaling AI in production environments. He recently spoke at Machine Learning Utah on AI infrastructure and clustering.You don’t scale AI with replicas.You scale it by managing scarce compute.Subscribe on Spotify or Apple, and follow us on YouTube.👉 Keep your AI on a leash.🧠 News🔗 Tools & Tech Mentioned👤 Guest🎯 Key Takeaway
Embed this episode
Ready to play
Self-Hosting AI: Scaling Is the Real Problem
No transcript for this episode yet
Similar Episodes
No similar episodes found.