AI Post Transformers cover art

All Episodes

AI Post Transformers — 304 episodes

#
Title
1

Reptile: The First-Order Meta-Learning Shortcut That Works

2

FLINT: Sharing High Bandwidth Flash for Scalable LLM Inference

3

LinearKV: When Exact State Merging Breaks Hybrid Models

4

Decoupling KL Direction from Rollout Source in LLM Distillation

5

Token Teachability: Rethinking Disagreement in On-Policy Distillation

6

Weak-to-Strong On-Policy Distillation Beats the Teacher

7

On-Policy Distillation: Why a Stronger Teacher Can Backfire

8

Sleeper Memory Poisoning: When Assistants Remember Lies

9

Why AI Systems Don't Learn After Deployment

10

Second-Order Optimization Meets Runtime Scheduling at Scale

11

Hand-Written PTX vs WMMA: A Precision-Dependent GPU Speedup

12

Mechanist: Automating the Discovery of How AI Models Think

13

Inside Claude Code's Agentic Loop: A Design Space Analysis

14

Grep vs Vector Search: How Agent Harnesses Shape Retrieval Accuracy

15

Decomposing Speedups Across Runtime, Kernel, and Quantization

16

AI-Generated Text and the Death of the Open Web

17

Approaching Shannon Bound: Lossless LLM Weight Compression

18

Test-Time Adaptation Through Entropy Minimization

19

SOAP: Stabilizing Shampoo's Second-Order Optimizer with Adam

20

Distributed Shampoo: Making Second-Order Optimization Practical at Scale

21

Scalable Second-Order Optimization: Shampoo at Scale

22

Preconditioned Optimization Without the Full Matrix Cost

23

Naive Test-Time Adaptation Destabilizes LLM Predictions

24

Kohonen's 1972 Correlation Matrix Memory, Decades Before Attention

25

Model-Agnostic Meta-Learning for Fast Task Adaptation

26

In-Place Test-Time Training Turns Fast Weights Into Online Memory

27

Hal: Memory-Augmented LLM Agents Still Hit Continual Learning's Wall

28

Continual Learning in LLMs: Beyond Catastrophic Forgetting

29

Learning, Fast and Slow: LLMs That Adapt Without Forgetting

30

Volatility Optimization Is Actually Bayesian Inference

31

TwinQuant: Manifold-Constrained Low-Rank Decomposition for 4-Bit Quantization

32

TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale

33

SnapStream: Taming KV-Cache Eviction for Dataflow Accelerators

34

StrataCL: Fabric-Native Zero-Copy Comms for AI Supernodes

35

MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA

36

Move the Query, Not the Cache: MLA Rewrites GPU Fabric Attention Routing

37

Silicon Showdown: GPU vs Apple Silicon LLM Inference Limits

38

Model Predictive Control's Real-Time Structure, from Chapter to Cockpit

39

Making Every Verified Token Count in MoE Speculative Decoding

40

FreeAct: Rethinking One-to-One Transforms for LLM Quantization

41

Global Memory Bloat in Long-Context LLM Serving

42

DualDecoder: Fixing GPU Memory Bloat in Long-Context KV Cache Offloading

43

Data Temporality's Hidden Impact on LLM Pretraining

44

Adapting Without Forgetting: A Lifelong Learning Roadmap for LLM Agents

45

Distributed Weight Data Parallelism Cuts LLM Inference Stalls

46

Cross-Family Speculative Prefill Cuts Long-Context Latency

47

AdaJEPA: Self-Adapting Latent World Models via Test-Time MPC

48

Test-Time Training Turns EDA Feedback Into Live Weight Updates for RTL

49

Main Trust Issue in FPGA HLS Design Workflow

50

The Unlearnability Phenomenon in RLVR Reasoning Models

51

MemPO: Teaching Agents to Write Their Own Memory

52

Eigenvectors of Experts: Training-free MoE Routing Without Collapse

53

Peer-Preservation: When Frontier Models Protect Other AIs

54

Superhuman Adaptable Intelligence Challenges the Idea of AGI

55

Adaptive Block-Scaled Data Types for FP4 Training

56

Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference

57

Kimi K3 Goes Trillion-Scale With Sparse MoE Efficiency

58

Hilbert Operator Reveals What Networks Actually Learn

59

NaturalReasoning: Backtranslating Reasoning Questions at Scale

60

ASAP: Disaggregating Attention and Experts for Faster MoE Prefill

61

MegaScale-Infer: Disaggregating Experts for Faster MoE Serving

62

Agentic Hardware Design: Evolving RTL Through Repository-Level Code Loops

63

Multi-Agent AI Rewrites a Million-Line Chip Design Tool

64

Compute or Load? Cake's Smart KV Cache Scheduler

65

CacheFlow: Optimal 3D-Parallel KV Cache Restoration

66

IMPRESS: A Multi-Tier KV Storage System for Faster LLM Prefill

67

Fast State Restoration for Evicted LLM KV Caches

68

Small Collectives, Big TLB Cost: Reverse Address Translation in Scale-Up GPU Pods

69

The Reversal Curse: When A Is B But Not B Is A

70

Post-Training Science: Scaling Laws for SFT and LoRA

71

Context is All You Need: Fixing OOD Drift Without Retraining

72

Prompt Boundary-Aware Scheduling with Event Tensors for Dynamic Kernels

73

Language Model Continual Learning: Do Written Facts Survive Repeated Weight Updates?

74

How Patchscopes Reveals What Language Models Really Think

75

The First Fact That Wouldn't Reason

76

Thought Anchors: Which Sentences Really Drive LLM Reasoning

77

Steering Reasoning Models' Cognitive Behaviors at Test-Time

78

KV Cache Eviction Through an Information Bottleneck Lens

79

Coordinated MoE Scheduling: How Gimbal Fixes Expert Locality

80

KV Cache Compaction Beyond Selection: Introducing Still

81

OPSDL: Teaching Long-Context Models to Trust Their Short-Context Selves

82

DeltaProduct: Extending DeltaNet's State-Tracking via Householder Products

83

LLMs Guess Wrong: Auditing What Models Know About You

84

SAC: Making Sparse Attention KV Caches Actually Sparse with CXL

85

Lightning OPD: Precomputing the Teacher for Cheaper Reasoning Distillation

86

MemGPT: Treating LLM Context Windows Like Virtual Memory

87

Tensor Cache: Compressing Evicted Tokens into Fixed-Size Memory

88

Cognitive Behaviors Behind Self-Improving Language Model Reasoners

89

SoundnessBench: Exposing AI Reviewers' Blind Spots

90

HyperOffload's Scheduling Claims Under Scrutiny

91

AMD XDNA NPU Sparsity-Aware Attention for Long-Context Prefill

92

Understanding Inference Scaling: Prefill, Decode, and Reasoning Bottlenecks

93

Cost-Aware Speculative Decoding for Mixture-of-Experts Models

94

Parcae: Stabilizing Looped Language Models with Control Theory

95

Hypic: Position-Independent KV Caching for Hybrid-Attention LLM Serving

96

Deep Native Structural Reasoning for Proteins, Molecules, and Crystals

97

GPT-5.6 System Card: Cybersecurity Rises, Critical Line Holds

98

Do We Still Need GPUs? Rethinking AI with Matrix-Enhanced CPUs

99

Huxley-Gödel Machine: Approximating Optimal Self-Improving Coding Agents

100

Expert-Locality-Aware Decode Routing for MoE Serving

101

SkillOpt-Lite: Rethinking Agent Skill Optimization with Zeroth-Order Simplicity

102

Darwin Gödel Machine: Self-Improving Coding Agents Through Open-Ended Evolution

103

Gödel Machines: Provably Optimal Self-Rewriting AI

104

SOUL.md or Severance: Host Staleness Intervention

105

Puzzle Compression for Hybrid MoE LLMs

106

Why GLU Variants Improve Transformer Feed-Forward Layers

107

When Combining Language Models Stops Helping

108

Gemma 4 Open-Weight Multimodal Reasoning

109

Hierarchical Sparse Attention for Infinite Context

110

Swish: Self-Gated Activation Beyond ReLU

111

SiLU Activations for Replay-Free Atari Reinforcement Learning

112

Verbalizable Representations and the Global Workspace

113

Program-as-Weights for Compiling Fuzzy Functions

114

Discretizing Reward Models for RL Alignment

115

AIConfigurator for Cross-Framework LLM Serving

116

Nemotron-TwoTower for Parallel Diffusion Language Modeling

117

DART Speeds Up Speculative LLM Decoding

118

SALCA for Sparse Long-Context Decoding

119

SERA: Repository Specialization for Open Coding Agents

120

Cache-Resident LLM Inference in GB-Scale Caches

121

When Finer Microscaling Hurts LLM Quantization

122

Scaling Prompt Tuning for Frozen T5 Models

123

Temporal-Tiered KV Cache for Long Context

124

SuperInfer: SLO-Aware LLM Inference on Superchips

125

LLMServingSim 2.0 for Disaggregated LLM Serving

126

DSpark Improves Speculative Decoding Acceptance Rates

127

Moebius: Seamless Parallelism Switching for MoE Serving

128

Information-Aware KV Cache Compression for Long Reasoning

129

JETSPEC and Parallel Tree Speculative Decoding

130

DAK: Direct GPU Memory Offloading for LLMs

131

Prefix-Tuning for Efficient Text Generation

132

ReasonCACHE: Learning Reasoning Without Weight Updates

133

RMSNorm: Simplifying Layer Normalization for Sequence Models

134

RT Cores for Exact k-Nearest Neighbor Search

135

Learning Facts at Scale with Active Reading

136

HELM: Holistic Evaluation of Language Models

137

PALOMA: Benchmarking Language Model Fit Across Domains

138

MIOpen and AMD's Open Deep Learning Primitives

139

Why Open Relational Foundation Models Fail

140

AI+HW 2035: Co-Designing Efficient AI Systems

141

Efficient Post-Training Quantization with FP8

142

X-LLM: Treating Multimodalities as Foreign Languages

143

Modeling Financial Habits with Transaction Transformers

144

TransactionGPT as a Payments Foundation Model

145

Simulating Individuals with Self-Reported LLM Agents

146

Building General User Models from Computer Use

147

Fine-Tuning LLMs for Human Behavior Prediction

148

Social Simulacra for Prototyping Online Communities

149

Stable Deep RL via Gaussian Representations

150

EMO: Emergent Modularity for Mixture-of-Experts

151

When LeJEPA Truly Learns a World Model

152

Benchmarking PEFT Techniques for Large Language Models

153

Weak-SIGReg for Stable Vision Transformer Training

154

InfiniGen for Efficient Long-Context LLM Inference

155

When Quantization Hurts Reasoning Models

156

Ling and Ring 2.6 for Trillion-Scale Agents

157

SageAttention2 and Fast Exact INT4 Attention

158

Nemotron 3 Ultra for Long-Horizon Agents

159

OpenSkill for Open-World Self-Evolution in LLM Agents

160

PaperBench: Can AI Replicate AI Research?

161

Training Modular KV Caches at Scale

162

DafnyPro for LLM-Assisted Dafny Verification

163

seL4: Proving a Microkernel in C

164

DafnyBench and LLMs for Formal Verification

165

From AGI to ASI and Beyond

166

MiniMax Sparse Attention at Million-Token Scale

167

ACE: Matrix-Native AI Extensions for x86

168

Dafny for Trustworthy AI Code Generation

169

Can LLMs Enable Mainstream Formal Verification?

170

From Natural Language to Verified Dafny Code

171

KV Binding Is Secretly Linear Attention

172

When LoRA Helps Under KV Cache Compression

173

IndexMem: Learned KV-Cache Eviction for Long-Context LLMs

174

AllMem for Efficient Long-Context Modeling

175

Relational Graph Transformer for Multi-Table Learning

176

Lattice: Fixed-Slot Compression for Transformer Memory

177

KumoRFM for In-Context Relational Learning

178

Atlas: Test-Time Memory for Long Contexts

179

Learning at Test Time with Expressive RNN States

180

Do Transformers Need Three Projections?

181

Robots Need More Than VLAs and World Models

182

End-to-End Context Compression at Scale

183

Unembedding Matrices as Feature Lenses for Embeddings

184

Predictive Query Language for Relational Databases

185

KumoRFM-2 for Relational Learning at Scale

186

Unified Neural Scaling Laws Across Regimes

187

RAGEN-2: Reasoning Collapse in Agentic RL

188

Latent Reasoning with Normalizing Flows

189

EMO: Emergent Modularity in Sparse Language Models

190

When AI Builds Itself and Recursive Self-Improvement

191

Mooncake for KV Cache-Centric LLM Serving

192

Technical AGI Safety and Security Framework

193

VeriCache: Lossless LLM Inference from Lossy KV Caches

194

SFMP Search-Free Mixed-Precision LLM Quantization

195

Harvest: Borrowing Peer GPU Memory for LLMs

196

Opaque Serial Depth and Chain-of-Thought Limits

197

TRELLIS and Bounded-Memory Transformer KV Compression

198

Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode

199

Dragonfly Topology for Scalable AI Networks

200

Do Language Models Need Sleep?

201

Snap's Microkernel Approach to Host Networking

202

SmolLM2 and the Power of Better Data

203

Post-Trained MoE Skips Half Its Experts

204

KVzap: Fast, Adaptive, Faithful KV Cache Pruning

205

KVzip for Query-Agnostic KV Cache Compression

206

CXL-GPU and Beyond Onboard Memory

207

Beluga: CXL Memory Pooling for LLM KV Cache

208

Trajectory Summaries for Long-Horizon Coding Agents

209

DFX: Multi-FPGA Acceleration for Transformer Inference

210

Generative Recursive Reasoning in Latent Space

211

LPU Chip for Low-Latency LLM Inference

212

After Titans: Behrouz on Nested Learning and Hope

213

Titans: Learning to Memorize at Test Time

214

Affordable Large-Scale Decoding Through Model-System Co-Design

215

Serving MoE Models with Disaggregated Expert Parallelism

216

The Sparsity Wall: What Reiner Pope Told Dwarkesh About MoE and Sparse Attention

217

NanoFlow and the Future of LLM Serving

218

Air Force One, Jensen Huang, and Anthropic's 2028 Memo

219

Trace Rewriting Against Unauthorized LLM Distillation

220

Agentic AI as a Path to AGI

221

When Many-Shot CoT Becomes Test-Time Learning

222

Ministral 3: Cascade Distillation for Long-Context Multimodal Models

223

Causal-JEPA for Object-Level World Models

224

Scaling Laws for Multilingual Code Pretraining

225

JANUS for Scalable MoE Inference

226

Lossless Sparse Deltas for RL Networks

227

FlashFuser and Hopper-Era FFN Kernel Fusion

228

Deep Kernel Fusion for Transformer Decoding

229

AI Co-Mathematician for Mathematical Research

230

TMAS: Scaling Test-Time Compute with Multi-Agent Synergy

231

MELT: Decoupling Compute From Memory

232

Long Context Pre-Training with Lighthouse Attention

233

MiA-Signature and Global Activation for Long Context

234

Qwen-Image-2.0 for Unified Generation and Editing

235

δ-mem and Online Memory for LLMs

236

Agentic Discovery for Test-Time Scaling

237

TIDE and the Rare Token Problem

238

Metacognition Against Confident Hallucinations

239

ForkKV for Multi-LoRA Agent Serving

240

ELF and Continuous Language Diffusion

241

Vistara Brings CXL Memory to Hyperscale

242

RAPTOR: Stable Concept Directions From Logistic Probes

243

Reasoning Theater and Unfaithful Chain-of-Thought

244

Optimization, Credit Assignment, and Consciousness

245

Why Transformers Fail at Counting

246

Split Personality Training Reveals Latent Knowledge

247

Explicit Information Transmission for Context Compression

248

Why LLM Serving Needs Mathematical Optimization

249

EverMemOS for Long-Horizon Agent Memory

250

When LLM Judges Become Coin Flips

251

Generative Modeling via Drifting in One Step

252

LAPS for Length-Aware LLM Serving

253

Machine Learning Self-Calibrated FPGA Time-to-Digital Converter

254

TensorFlow for Distributed Machine Learning Systems

255

SGLang for Faster Structured LLM Programs

256

Can Models Learn from Long Context?

257

Backpropagation Through Time Explained

258

Learning in Random Nets and Generalization

259

What the Frog’s Eye Tells the Brain

260

How Models Detect Hidden Activation Steering

261

Caffe and the Rise of CNN Frameworks

262

Synchronous Data Flow for Signal Processing

263

Automating DNN Compilation for FPGA Accelerators

264

Caffeine: A Unified FPGA for CNNs

265

Boosted Decision Trees for CMS Muon Triggers

266

Fast FPGA Inference for LHC Triggers

267

Fast FPGA BDT Inference for LHC Triggers

268

Why LightGBM Made Boosted Trees Fast

269

Automating CNN Mapping on Embedded FPGAs

270

RFNoC SISO Processor via High-Level Synthesis

271

Reinforcement Learning in 2025: An Overview

272

Training LLMs for Divide-and-Conquer Reasoning

273

PackKV Lossy Compression for KV Caches

274

Training Million-Token LLMs Beyond the Memory Barrier

275

How Induction Heads Emerge in Transformers

276

DeepWalk and the Rise of Graph Embeddings

277

node2vec and Learning Graph Embeddings

278

Selective Classification with Deep Neural Networks

279

Geometric Memory in Deep Sequence Models

280

Self-Improving Pretraining With Post-Trained Models

281

Do Language Models Know Their Limits

282

Fast Speech Recognition by Transcript Editing

283

CacheFlow and 3D-Parallel KV Cache Restoration

284

DeltaKV: Compressing KV Caches for Long Context

285

A Practical Review of Mechanistic Interpretability

286

Recursive Multi-Agent Systems in Latent Space

287

Discrete Representations for Continual Reinforcement Learning

288

Teaching Language Models to Verbalize Uncertainty

289

ChartNet for Robust Multimodal Chart Understanding

290

Can LLMs Judge Their Own Capabilities?

291

Convolutional Sequence to Sequence Learning for Translation

292

Bahdanau Attention for Neural Machine Translation

293

Sequence to Sequence Learning with Neural Networks

294

Hopfield Networks and Transformer Attention as Memory

295

Scaling Test-Time Compute for Reasoning Models

296

Reverse-Mode Differentiation Across AD and Neural Nets

297

Breaking the Prefix Barrier with Shared KV Cache

298

Stochastic KV Routing for Cache Sharing

299

Stabilizing Efficient Reasoning with Step-Level Advantage Selection

300

World-R1 Improves 3D Consistency in Text-to-Video

301

FPGA Neural Network Accelerators for Space

302

Deep Learning in Spiking Neural Networks

303

Structural Understanding of LLM Overthinking

304

AdamW: Decoupled Weight Decay Regularization for Adaptive Gradient Algorithms