Junchen Jiang, Tensormesh CEO on Faster, Cheaper LLM Inference episode artwork

EPISODE · Jul 20, 2026 · 57 MIN

Junchen Jiang, Tensormesh CEO on Faster, Cheaper LLM Inference

from Venture with Grace · host Grace Gong

Junchen Jiang is the Co-Founder and CEO of Tensormesh, an AI infrastructure company building the first commercial platform for KV cache-accelerated LLM inference. ~~~~~~~~~~~~~This episode is brought to you by Nebius — the ultimate cloud for AI innovators.Nebius provides AI infrastructure you can count on, combining reliability and speed with flexibility and engineering support unmatched by hyperscalers.AI leaders like Meta, Shopify, and Higgsfield already partner with Nebius to run their AI workloads. Plus, venture-backed startups can save up to $150,000 on compute costs when they apply for access. Visit nebius.com or nebius.com/startups to learn more~~~~~~~~~~~~~Built on years of research at the University of Chicago, UC Berkeley, and Carnegie Mellon, Tensormesh helps organizations reduce AI inference latency and GPU costs by up to 10x while keeping models and data on their own infrastructure. The company emerged from stealth with $4.5 million in seed funding led by Laude Ventures. Junchen is also an Associate Professor of Computer Science at the University of Chicago, where he leads research on large-scale AI systems and co-created LMCache and CacheBlend—breakthrough technologies that have become foundational components of modern LLM infrastructure. His work earned the ACM EuroSys 2025 Best Paper Award and has been adopted by organizations including Bloomberg, Red Hat, Redis, WEKA, and Tencent. He is also a recipient of the Google Faculty Award and the Carnegie Mellon Best Computer Science Ph.D. Dissertation Award.TopicsWhy LLM Inference, Not Training, Is Becoming the Biggest Bottleneck in Enterprise AIHow KV Cache Optimization Can Reduce AI Latency and GPU Costs by Up to 10xBuilding the Next Generation of Open-Source AI Infrastructure: From LMCache Research to Tensormesh#ArtificialIntelligence #LLM #AIInfrastructure #MachineLearning #Startups

Episode metadata supplied by the publisher feed · Published Jul 20, 2026

Embed this episode

NOW PLAYING

Junchen Jiang, Tensormesh CEO on Faster, Cheaper LLM Inference

0:00 57:55

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Venture with Grace?

This episode is 57 minutes long.

When was this Venture with Grace episode published?

This episode was published on July 20, 2026.

Can I download this Venture with Grace episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!