166: Data Processing Fundamentals and Building a Unified Execution Engine Featuring Pedro Pedreira of Meta episode artwork

EPISODE · Nov 29, 2023 · 1H 12M

166: Data Processing Fundamentals and Building a Unified Execution Engine Featuring Pedro Pedreira of Meta

from The Data Stack Show · host Rudderstack

Highlights from this week’s conversation include:The concept of composable at a lower level of data infrastructure (1:28)New architectures and components that allow developers to build databases (3:44)Pedro's background and experience in data infrastructure (6:18)The Spectrum of Latency and Analytics (12:59)Different Query Engines for Different Use Cases (16:32)Vectorized vs Code Gen Data Processing (19:33)Vectorization and Code Generation (21:21)Examples of Vectorized Engines (24:33)Rewriting Execution Engine in C++ (27:22)Different Organization of Presto and Spark (33:17)Arrow and its Extensions (37:15)The similarities between analytics and ML (44:33)Offline feature engineering and data preprocessing for training (48:00)Dialect and semantic differences in using Velox for different engines (50:01)The convergence of dialects (52:23)Challenges of substrate and semantics (53:18)Future plans for Velox (58:09)The discussion on evolving Parquet (1:03:38)The integration of the relational model and the tensor model (1:07:29)The Data Stack Show is a weekly podcast powered by RudderStack, the CDP for developers. Each week we’ll talk to data engineers, analysts, and data scientists about their experience around building and maintaining data infrastructure, delivering data and data products, and driving better outcomes across their businesses with data.RudderStack helps businesses make the most out of their customer data while ensuring data privacy and security. To learn more about RudderStack visit rudderstack.com. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Episode metadata supplied by the publisher feed · Published Nov 29, 2023

Embed this episode

This week on The Data Stack Show, Eric and Kostas chat with Pedro Pedreira, a Software Engineer at Meta (formerly Facebook). During the conversation, Pedro discusses the intricacies of data infrastructure and the concept of composability. He delves into the specifics of data latency, explaining the spectrum of requirements from low-latency dashboards to high-latency batch queries. Pedro also discusses the concept of vectorization in database operations and the benefits it brings. The conversation shifts to the development and usage of certain libraries and the goal of converging dialects in data computation. Don’t miss this great episode!

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

166: Data Processing Fundamentals and Building a Unified Execution Engine Featuring Pedro Pedreira of Meta

0:00 1:12:16

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of The Data Stack Show?

This episode is 1 hour and 12 minutes long.

When was this The Data Stack Show episode published?

This episode was published on November 29, 2023.

Can I download this The Data Stack Show episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!