Ep 48: Faster pre-training without changing model architecture just became practical — Nous Research's Token Superposition Training cuts wall-clock time by up to 2.5x on models from 270M to 10B. episode artwork

EPISODE · May 14, 2026 · 8 MIN

Ep 48: Faster pre-training without changing model architecture just became practical — Nous Research's Token Superposition Training cuts wall-clock time by up to 2.5x on models from 270M to 10B.

from Models & Agents

Models & Agents Faster pre-training without changing model architecture just became practical — Nous Research's Token Superposition Training cuts wall-clock time by up to 2.5x on models from 270M to 10B. What You Need to Know: Nous Research introduced Token Superposition Training, a two-phase method that averages token embeddings early then switches back to standard prediction. Open-source builders shipped a full cinematic video pipeline that runs end-to-end on a single AMD MI300X. ... AI Disclosure: This podcast is curated by Patrick but uses AI-generated voice synthesis for audio production.

Episode metadata supplied by the publisher feed · Published May 14, 2026

Embed this episode

NOW PLAYING

Ep 48: Faster pre-training without changing model architecture just became practical — Nous Research's Token Superposition Training cuts wall-clock time by up to 2.5x on models from 270M to 10B.

0:00 8:31

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Models & Agents?

This episode is 8 minutes long.

When was this Models & Agents episode published?

This episode was published on May 14, 2026.

Can I download this Models & Agents episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!