Optimize custom machine learning operations with Metal tensors episode artwork

EPISODE · Jun 12, 2026 · 7 MIN

Optimize custom machine learning operations with Metal tensors

from Podkey WWDC 2026

A Podkey summary of Optimize custom machine learning operations with Metal tensors, from WWDC 2026.A lot of this week’s Apple graphics and AI story really comes down to one idea: getting more model work closer to the hardware, with less fuss for developers. The big pieces are the M5’s new neural accelerator inside each shader core, broader quantized tensor support in TensorOps, and a pretty practical path for building things like FlashAttention and dropping them into real models. It’s fairly technical stuff, but the throughline is simple enough: faster inference, less wasted memory movement, and fewer awkward hand-built code paths.The M5 adds AI help right inside the GPU coresQuantization gets broader and a lot more usableA scale plane keeps quantized tensors togetherCooperative tensors cut down on memory trafficHow FlashAttention fits into all thisCustom Metal kernels can plug into real model workflowsThe dequantization trade-off is speed versus extra movementThis podcast was created with Podkey. Make your own at https://podkey.fm

Episode metadata supplied by the publisher feed · Published Jun 12, 2026

Embed this episode

NOW PLAYING

Optimize custom machine learning operations with Metal tensors

0:00 7:12

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Podkey WWDC 2026?

This episode is 7 minutes long.

When was this Podkey WWDC 2026 episode published?

This episode was published on June 12, 2026.

Can I download this Podkey WWDC 2026 episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!