More powerful deep learning with transformers (Ep. 84) (Rebroadcast) episode artwork

EPISODE · Nov 27, 2019 · 37 MIN

More powerful deep learning with transformers (Ep. 84) (Rebroadcast)

from Data Science at Home · host Francesco Gadaleta

Some of the most powerful NLP models like BERT and GPT-2 have one thing in common: they all use the transformer architecture. Such architecture is built on top of another important concept already known to the community: self-attention. In this episode I explain what these mechanisms are, how they work and why they are so powerful. Don't forget to subscribe to our Newsletter or join the discussion on our Discord server   References Attention is all you need  https://arxiv.org/abs/1706.03762 The illustrated transformer  https://jalammar.github.io/illustrated-transformer Self-attention for generative models  http://web.stanford.edu/class/cs224n/slides/cs224n-2019-lecture14-transformers.pdf

Episode metadata supplied by the publisher feed · Published Nov 27, 2019

Embed this episode

NOW PLAYING

More powerful deep learning with transformers (Ep. 84) (Rebroadcast)

0:00 37:44

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Data Science at Home?

This episode is 37 minutes long.

When was this Data Science at Home episode published?

This episode was published on November 27, 2019.

Can I download this Data Science at Home episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!