Real-Time Data Pipelines Made Easy with Structured Streaming in Apache Spark | Databricks episode artwork

EPISODE · Feb 18, 2024 · 4 MIN

Real-Time Data Pipelines Made Easy with Structured Streaming in Apache Spark | Databricks

from Higher Signal: Get Smarter. Faster. · host Higher Signal

Get the slides: https://www.datacouncil.ai/talks/building-real-time-data-pipelines-made-easy-with-structured-streaming-in-apache-spark?utm_source=youtube&utm_medium=social&utm_campaign=%20-%20DEC-SF-18%20Slides%20DownloadABOUT THE TALK:Structured Streaming is the next generation of distributed, streaming processing in Apache Spark. Developers can write a query written in their language of choice (Scala/Java/Python/R) using powerful high-level APIs (DataFrames / Datasets / SQL) and apply that same query to both static datasets and streaming data. In case of streaming, Spark will automatically create an incremental execution plan that automatically handles late, out-of-order data and ensures end-to-end exactly-once fault-tolerance guarantees. In this practical session, I will walk through a concrete streaming ETL example where – in less than 10 lines – you can read raw, unstructured data from Kafka data, transform it and write it out as a structured table ready for batch and ad-hoc queries on up-to-the-last-minute data. I will give a quick glimpse of advanced features like event-time based aggregations, stream-stream joins and arbitrary stateful operations.ABOUT THE SPEAKER:Tathagata is a committer and PMC to the Apache Spark project and a Software Engineer at Databricks. He is the lead developer of Spark Streaming, and now focuses primarily on Structured Streaming. Previously, he was a member of the AMPLab, UC Berkeley as a graduate student researcher where he conducted research on data-center frameworks and networks with Scott Shenker and Ion Stoica.ABOUT DATA COUNCIL: Data Council (https://www.datacouncil.ai/) is a community and conference series that provides data professionals with the learning and networking opportunities they need to grow their careers. Make sure to subscribe to our channel for more videos, including DC_THURS, our series of live online interviews with leading data professionals from top open source projects and startups. FOLLOW DATA COUNCIL:Twitter: https://twitter.com/DataCouncilAI LinkedIn: https://www.linkedin.com/company/datacouncil-ai

Episode metadata supplied by the publisher feed · Published Feb 18, 2024

Embed this episode

NOW PLAYING

Real-Time Data Pipelines Made Easy with Structured Streaming in Apache Spark | Databricks

0:00 4:29

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Higher Signal: Get Smarter. Faster.?

This episode is 4 minutes long.

When was this Higher Signal: Get Smarter. Faster. episode published?

This episode was published on February 18, 2024.

Can I download this Higher Signal: Get Smarter. Faster. episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!