Summit 2014: Taboola's experience with Apache Spark / Tal Sliwowicz episode artwork

EPISODE · Apr 21, 2014

Summit 2014: Taboola's experience with Apache Spark / Tal Sliwowicz

from רברס עם פלטפורמה

At taboola we are getting a constant feed of data (many billions of user events a day) and are using Apache Spark together with Cassandra for both real time data stream processing as well as offline data processing. We'd like to share our experience with these cutting edge technologies. Apache Spark is an open source project - Hadoop-compatible computing engine that makes big data analysis drastically faster, through in-memory computing, and simpler to write, through easy APIs in Java, Scala and Python. This project was born as part of a PHD work in UC Berkley's AMPLab (part of the BDAS - pronounced "Bad Ass") and turned into an incubating Apache project with more active contributors than Hadoop. Surprisingly, Yahoo! are one of the biggest contributors to the project and already have large production clusters of Spark on YARN. Spark can run either standalone cluster, or using either Apache mesos and ZooKeeper or YARN and can run side by side with Hadoop/Hive on the same data. One of the biggest benefits of Spark is that the API is very simple and the same analytics code can be used for both streaming data and offline data processing. MP3

Episode metadata supplied by the publisher feed · Published Apr 21, 2014

Embed this episode

NOW PLAYING

Summit 2014: Taboola's experience with Apache Spark / Tal Sliwowicz

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

When was this רברס עם פלטפורמה episode published?

This episode was published on April 21, 2014.

Can I download this רברס עם פלטפורמה episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!