Vectors, AI, and the Infrastructure Fury Ahead episode artwork

EPISODE · May 23, 2025 · 25 MIN

Vectors, AI, and the Infrastructure Fury Ahead

from Shared Everything · host Jeff Denworth, Nicole Hemsoth Prickett

00:00 – IntroductionNicole introduces Jeff Denworth, reminiscing about the Big Data era (~2010–2014).01:15 – Big Data to Big MetadataJeff reflects on the Big Data era (Hadoop, analytics, NoSQL).Today's valuations (Snowflake, Databricks) suggest Big Data's continued relevance.02:11 – The Rise of Big MetadataJeff describes the shift from Big Data to Big Metadata.AI creates new categories and applications, rapidly driving data infrastructure demands.Example: Nvidia’s rapid growth due to deep learning-driven workloads.05:01 – Synthetic Data and Metadata ExplosionJeff notes social networks using synthetic data to circumvent privacy regulations.Metadata types include large-scale data catalogs to manage exabytes of data (e.g., OpenAI).08:14 – Dynamic Data CatalogsVAST Database as an example of transactional and analytical infrastructure.Benefits of SQL queries replacing traditional file operations for faster data handling.09:50 – Metadata Evolves with VectorsExplanation of embeddings, vector databases, and similarity search.AI-driven understanding of unstructured data via vectors.11:56 – Massive Scale of Vector DatabasesRough estimate: ~40 trillion vectors per 100 petabytes of data.Challenges with conventional vector databases at massive scale (cost, memory, speed).13:22 – Future Scale Problems and AI-driven Data EngineeringRetrieval-Augmented Generation (RAG) increases vector database scale needs.Nvidia's data flywheel concept accelerates embedding and data engineering automation.15:48 – Predicting Infrastructure Needs (Two-Year Outlook)Jeff predicts AI models will significantly improve data engineering within two years.Enterprises need vector databases capable of transactional, real-time performance.18:10 – Future-Proofing Infrastructure (Five-Year Outlook)Jeff expects AI-driven automation to impact all business processes (factories, back-office).Businesses must be prepared for rapid scaling and foundational AI-driven changes.21:14 – Industries Leading the AI Infrastructure RaceAI adoption speed varies by industry—highest "fury" is in software development.Banks and trading firms leverage AI differently: profit efficiency vs. alpha-seeking.23:55 – Cloud vs. On-Premises Infrastructure ChoicesJeff sees hybrid approaches prevailing; decision-making depends on enterprise-specific needs.Introduces idea of "agentic workforce" prompted by Jensen Huang's statement (100M AI agents).24:31 – Agent Ownership and Future ConsequencesRaises profound questions about ownership and management of AI agents in business.Jeff notes limited current customer recognition of these deeper implications.25:56 – Closing RemarksNicole and Jeff conclude by noting broad societal implications of AI-driven changes.Emphasis on importance of continued discussions around big metadata.

Episode metadata supplied by the publisher feed · Published May 23, 2025

Embed this episode

In this episode of Shared Everything, Nicole Hemsoth Prickett chats with Jeff Denworth, co-founder of VAST Data, about why the Big Data hype of Hadoop is now ancient history—and how the real story is quietly exploding around metadata. They explore how the seemingly humble metadata concept has transformed into something vast and complicated, driven by the rise of AI, vectors, and deep-learning embeddings. Jeff reveals why traditional infrastructure struggles at exabyte scales, how vector databases underpin retrieval-augmented generative AI (RAG), and what companies can do right now to future-proof their systems for an increasingly agent-driven world. It’s a sharp, fast-moving tour through an AI landscape on the verge of metadata-induced chaos, raising deep questions about who (or what) will own the enterprise workforce of tomorrow.

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

Vectors, AI, and the Infrastructure Fury Ahead

0:00 25:57

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Shared Everything?

This episode is 25 minutes long.

When was this Shared Everything episode published?

This episode was published on May 23, 2025.

Can I download this Shared Everything episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!