#58 Bagaço: A pretraining dataset for European Portuguese episode artwork

EPISODE · Feb 22, 2026

#58 Bagaço: A pretraining dataset for European Portuguese

from Duarte O.Carmo's articles

Let's say your goal is to train a Large Language Model only on European Portuguese. Where do you start? What datasets are out there? What websites are being scraped for the large black box? Bagaço - named after the popular Portuguese moonshine - is a small step in that direction. In June …

Episode metadata supplied by the publisher feed · Published Feb 22, 2026

Embed this episode

NOW PLAYING

#58 Bagaço: A pretraining dataset for European Portuguese

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

When was this Duarte O.Carmo's articles episode published?

This episode was published on February 22, 2026.

Can I download this Duarte O.Carmo's articles episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!