INTERVIEW: Polysemanticity w/ Dr. Darryl Wright episode artwork

EPISODE · Jan 22, 2024 · 45 MIN

INTERVIEW: Polysemanticity w/ Dr. Darryl Wright

from Into AI Safety · host Jacob Haimes

Darryl and I discuss his background, how he became interested in machine learning, and a project we are currently working on investigating the penalization of polysemanticity during the training of neural networks.Check out a diagram of the decoder task used for our research!01:46 - Interview begins02:14 - Supernovae classification08:58 - Penalizing polysemanticity20:58 - Our "toy model"30:06 - Task description32:47 - Addressing hurdles39:20 - Lessons learnedLinks to all articles/papers which are mentioned throughout the episode can be found below, in order of their appearance.ZooniverseBlueDot ImpactAI Safety SupportZoom In: An Introduction to CircuitsMNIST dataset on PapersWithCodeClusterability in Neural NetworksCIFAR-10 datasetEffective Altruism GlobalCLIP (blog post)Long Term Future FundEngineering Monosemanticity in Toy Models

Episode metadata supplied by the publisher feed · Published Jan 22, 2024

Embed this episode

Ready to play

INTERVIEW: Polysemanticity w/ Dr. Darryl Wright

0:00 45:09

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Into AI Safety?

This episode is 45 minutes long.

When was this Into AI Safety episode published?

This episode was published on January 22, 2024.

Can I download this Into AI Safety episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!