Data Provenance and Reproducibility with Pachyderm episode artwork

EPISODE · Feb 3, 2017 · 40 MIN

Data Provenance and Reproducibility with Pachyderm

from Data Skeptic

Versioning isn't just for source code. Being able to track changes to data is critical for answering questions about data provenance, quality, and reproducibility. Daniel Whitenack joins me this week to talk about these concepts and share his work on Pachyderm. Pachyderm is an open source containerized data lake. During the show, Daniel mentioned the Gopher Data Science github repo as a great resource for any data scientists interested in the Go language. Although we didn't mention it, Daniel also did an interesting analysis on the 2016 world chess championship that complements our recent episode on chess well. You can find that post here Supplemental music is Lee Rosevere's Let's Start at the Beginning.   Thanks to Periscope Data for sponsoring this episode. More about them at periscopedata.com/skeptics      

Episode metadata supplied by the publisher feed · Published Feb 3, 2017

Embed this episode

Ready to play

Data Provenance and Reproducibility with Pachyderm

0:00 40:11

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Data Skeptic?

This episode is 40 minutes long.

When was this Data Skeptic episode published?

This episode was published on February 3, 2017.

Can I download this Data Skeptic episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!