All Episodes
AI Post Transformers — 304 episodes
Reptile: The First-Order Meta-Learning Shortcut That Works
FLINT: Sharing High Bandwidth Flash for Scalable LLM Inference
LinearKV: When Exact State Merging Breaks Hybrid Models
Decoupling KL Direction from Rollout Source in LLM Distillation
Token Teachability: Rethinking Disagreement in On-Policy Distillation
Weak-to-Strong On-Policy Distillation Beats the Teacher
On-Policy Distillation: Why a Stronger Teacher Can Backfire
Sleeper Memory Poisoning: When Assistants Remember Lies
Why AI Systems Don't Learn After Deployment
Second-Order Optimization Meets Runtime Scheduling at Scale
Hand-Written PTX vs WMMA: A Precision-Dependent GPU Speedup
Mechanist: Automating the Discovery of How AI Models Think
Inside Claude Code's Agentic Loop: A Design Space Analysis
Grep vs Vector Search: How Agent Harnesses Shape Retrieval Accuracy
Decomposing Speedups Across Runtime, Kernel, and Quantization
AI-Generated Text and the Death of the Open Web
Approaching Shannon Bound: Lossless LLM Weight Compression
Test-Time Adaptation Through Entropy Minimization
SOAP: Stabilizing Shampoo's Second-Order Optimizer with Adam
Distributed Shampoo: Making Second-Order Optimization Practical at Scale
Scalable Second-Order Optimization: Shampoo at Scale
Preconditioned Optimization Without the Full Matrix Cost
Naive Test-Time Adaptation Destabilizes LLM Predictions
Kohonen's 1972 Correlation Matrix Memory, Decades Before Attention
Model-Agnostic Meta-Learning for Fast Task Adaptation
In-Place Test-Time Training Turns Fast Weights Into Online Memory
Hal: Memory-Augmented LLM Agents Still Hit Continual Learning's Wall
Continual Learning in LLMs: Beyond Catastrophic Forgetting
Learning, Fast and Slow: LLMs That Adapt Without Forgetting
Volatility Optimization Is Actually Bayesian Inference
TwinQuant: Manifold-Constrained Low-Rank Decomposition for 4-Bit Quantization
TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale
SnapStream: Taming KV-Cache Eviction for Dataflow Accelerators
StrataCL: Fabric-Native Zero-Copy Comms for AI Supernodes
MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA
Move the Query, Not the Cache: MLA Rewrites GPU Fabric Attention Routing
Silicon Showdown: GPU vs Apple Silicon LLM Inference Limits
Model Predictive Control's Real-Time Structure, from Chapter to Cockpit
Making Every Verified Token Count in MoE Speculative Decoding
FreeAct: Rethinking One-to-One Transforms for LLM Quantization
Global Memory Bloat in Long-Context LLM Serving
DualDecoder: Fixing GPU Memory Bloat in Long-Context KV Cache Offloading
Data Temporality's Hidden Impact on LLM Pretraining
Adapting Without Forgetting: A Lifelong Learning Roadmap for LLM Agents
Distributed Weight Data Parallelism Cuts LLM Inference Stalls
Cross-Family Speculative Prefill Cuts Long-Context Latency
AdaJEPA: Self-Adapting Latent World Models via Test-Time MPC
Test-Time Training Turns EDA Feedback Into Live Weight Updates for RTL
Main Trust Issue in FPGA HLS Design Workflow
The Unlearnability Phenomenon in RLVR Reasoning Models
MemPO: Teaching Agents to Write Their Own Memory
Eigenvectors of Experts: Training-free MoE Routing Without Collapse
Peer-Preservation: When Frontier Models Protect Other AIs
Superhuman Adaptable Intelligence Challenges the Idea of AGI
Adaptive Block-Scaled Data Types for FP4 Training
Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
Kimi K3 Goes Trillion-Scale With Sparse MoE Efficiency
Hilbert Operator Reveals What Networks Actually Learn
NaturalReasoning: Backtranslating Reasoning Questions at Scale
ASAP: Disaggregating Attention and Experts for Faster MoE Prefill
MegaScale-Infer: Disaggregating Experts for Faster MoE Serving
Agentic Hardware Design: Evolving RTL Through Repository-Level Code Loops
Multi-Agent AI Rewrites a Million-Line Chip Design Tool
Compute or Load? Cake's Smart KV Cache Scheduler
CacheFlow: Optimal 3D-Parallel KV Cache Restoration
IMPRESS: A Multi-Tier KV Storage System for Faster LLM Prefill
Fast State Restoration for Evicted LLM KV Caches
Small Collectives, Big TLB Cost: Reverse Address Translation in Scale-Up GPU Pods
The Reversal Curse: When A Is B But Not B Is A
Post-Training Science: Scaling Laws for SFT and LoRA
Context is All You Need: Fixing OOD Drift Without Retraining
Prompt Boundary-Aware Scheduling with Event Tensors for Dynamic Kernels
Language Model Continual Learning: Do Written Facts Survive Repeated Weight Updates?
How Patchscopes Reveals What Language Models Really Think
The First Fact That Wouldn't Reason
Thought Anchors: Which Sentences Really Drive LLM Reasoning
Steering Reasoning Models' Cognitive Behaviors at Test-Time
KV Cache Eviction Through an Information Bottleneck Lens
Coordinated MoE Scheduling: How Gimbal Fixes Expert Locality
KV Cache Compaction Beyond Selection: Introducing Still
OPSDL: Teaching Long-Context Models to Trust Their Short-Context Selves
DeltaProduct: Extending DeltaNet's State-Tracking via Householder Products
LLMs Guess Wrong: Auditing What Models Know About You
SAC: Making Sparse Attention KV Caches Actually Sparse with CXL
Lightning OPD: Precomputing the Teacher for Cheaper Reasoning Distillation
MemGPT: Treating LLM Context Windows Like Virtual Memory
Tensor Cache: Compressing Evicted Tokens into Fixed-Size Memory
Cognitive Behaviors Behind Self-Improving Language Model Reasoners
SoundnessBench: Exposing AI Reviewers' Blind Spots
HyperOffload's Scheduling Claims Under Scrutiny
AMD XDNA NPU Sparsity-Aware Attention for Long-Context Prefill
Understanding Inference Scaling: Prefill, Decode, and Reasoning Bottlenecks
Cost-Aware Speculative Decoding for Mixture-of-Experts Models
Parcae: Stabilizing Looped Language Models with Control Theory
Hypic: Position-Independent KV Caching for Hybrid-Attention LLM Serving
Deep Native Structural Reasoning for Proteins, Molecules, and Crystals
GPT-5.6 System Card: Cybersecurity Rises, Critical Line Holds
Do We Still Need GPUs? Rethinking AI with Matrix-Enhanced CPUs
Huxley-Gödel Machine: Approximating Optimal Self-Improving Coding Agents
Expert-Locality-Aware Decode Routing for MoE Serving
SkillOpt-Lite: Rethinking Agent Skill Optimization with Zeroth-Order Simplicity
Darwin Gödel Machine: Self-Improving Coding Agents Through Open-Ended Evolution
Gödel Machines: Provably Optimal Self-Rewriting AI
SOUL.md or Severance: Host Staleness Intervention
Puzzle Compression for Hybrid MoE LLMs
Why GLU Variants Improve Transformer Feed-Forward Layers
When Combining Language Models Stops Helping
Gemma 4 Open-Weight Multimodal Reasoning
Hierarchical Sparse Attention for Infinite Context
Swish: Self-Gated Activation Beyond ReLU
SiLU Activations for Replay-Free Atari Reinforcement Learning
Verbalizable Representations and the Global Workspace
Program-as-Weights for Compiling Fuzzy Functions
Discretizing Reward Models for RL Alignment
AIConfigurator for Cross-Framework LLM Serving
Nemotron-TwoTower for Parallel Diffusion Language Modeling
DART Speeds Up Speculative LLM Decoding
SALCA for Sparse Long-Context Decoding
SERA: Repository Specialization for Open Coding Agents
Cache-Resident LLM Inference in GB-Scale Caches
When Finer Microscaling Hurts LLM Quantization
Scaling Prompt Tuning for Frozen T5 Models
Temporal-Tiered KV Cache for Long Context
SuperInfer: SLO-Aware LLM Inference on Superchips
LLMServingSim 2.0 for Disaggregated LLM Serving
DSpark Improves Speculative Decoding Acceptance Rates
Moebius: Seamless Parallelism Switching for MoE Serving
Information-Aware KV Cache Compression for Long Reasoning
JETSPEC and Parallel Tree Speculative Decoding
DAK: Direct GPU Memory Offloading for LLMs
Prefix-Tuning for Efficient Text Generation
ReasonCACHE: Learning Reasoning Without Weight Updates
RMSNorm: Simplifying Layer Normalization for Sequence Models
RT Cores for Exact k-Nearest Neighbor Search
Learning Facts at Scale with Active Reading
HELM: Holistic Evaluation of Language Models
PALOMA: Benchmarking Language Model Fit Across Domains
MIOpen and AMD's Open Deep Learning Primitives
Why Open Relational Foundation Models Fail
AI+HW 2035: Co-Designing Efficient AI Systems
Efficient Post-Training Quantization with FP8
X-LLM: Treating Multimodalities as Foreign Languages
Modeling Financial Habits with Transaction Transformers
TransactionGPT as a Payments Foundation Model
Simulating Individuals with Self-Reported LLM Agents
Building General User Models from Computer Use
Fine-Tuning LLMs for Human Behavior Prediction
Social Simulacra for Prototyping Online Communities
Stable Deep RL via Gaussian Representations
EMO: Emergent Modularity for Mixture-of-Experts
When LeJEPA Truly Learns a World Model
Benchmarking PEFT Techniques for Large Language Models
Weak-SIGReg for Stable Vision Transformer Training
InfiniGen for Efficient Long-Context LLM Inference
When Quantization Hurts Reasoning Models
Ling and Ring 2.6 for Trillion-Scale Agents
SageAttention2 and Fast Exact INT4 Attention
Nemotron 3 Ultra for Long-Horizon Agents
OpenSkill for Open-World Self-Evolution in LLM Agents
PaperBench: Can AI Replicate AI Research?
Training Modular KV Caches at Scale
DafnyPro for LLM-Assisted Dafny Verification
seL4: Proving a Microkernel in C
DafnyBench and LLMs for Formal Verification
From AGI to ASI and Beyond
MiniMax Sparse Attention at Million-Token Scale
ACE: Matrix-Native AI Extensions for x86
Dafny for Trustworthy AI Code Generation
Can LLMs Enable Mainstream Formal Verification?
From Natural Language to Verified Dafny Code
KV Binding Is Secretly Linear Attention
When LoRA Helps Under KV Cache Compression
IndexMem: Learned KV-Cache Eviction for Long-Context LLMs
AllMem for Efficient Long-Context Modeling
Relational Graph Transformer for Multi-Table Learning
Lattice: Fixed-Slot Compression for Transformer Memory
KumoRFM for In-Context Relational Learning
Atlas: Test-Time Memory for Long Contexts
Learning at Test Time with Expressive RNN States
Do Transformers Need Three Projections?
Robots Need More Than VLAs and World Models
End-to-End Context Compression at Scale
Unembedding Matrices as Feature Lenses for Embeddings
Predictive Query Language for Relational Databases
KumoRFM-2 for Relational Learning at Scale
Unified Neural Scaling Laws Across Regimes
RAGEN-2: Reasoning Collapse in Agentic RL
Latent Reasoning with Normalizing Flows
EMO: Emergent Modularity in Sparse Language Models
When AI Builds Itself and Recursive Self-Improvement
Mooncake for KV Cache-Centric LLM Serving
Technical AGI Safety and Security Framework
VeriCache: Lossless LLM Inference from Lossy KV Caches
SFMP Search-Free Mixed-Precision LLM Quantization
Harvest: Borrowing Peer GPU Memory for LLMs
Opaque Serial Depth and Chain-of-Thought Limits
TRELLIS and Bounded-Memory Transformer KV Compression
Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode
Dragonfly Topology for Scalable AI Networks
Do Language Models Need Sleep?
Snap's Microkernel Approach to Host Networking
SmolLM2 and the Power of Better Data
Post-Trained MoE Skips Half Its Experts
KVzap: Fast, Adaptive, Faithful KV Cache Pruning
KVzip for Query-Agnostic KV Cache Compression
CXL-GPU and Beyond Onboard Memory
Beluga: CXL Memory Pooling for LLM KV Cache
Trajectory Summaries for Long-Horizon Coding Agents
DFX: Multi-FPGA Acceleration for Transformer Inference
Generative Recursive Reasoning in Latent Space
LPU Chip for Low-Latency LLM Inference
After Titans: Behrouz on Nested Learning and Hope
Titans: Learning to Memorize at Test Time
Affordable Large-Scale Decoding Through Model-System Co-Design
Serving MoE Models with Disaggregated Expert Parallelism
The Sparsity Wall: What Reiner Pope Told Dwarkesh About MoE and Sparse Attention
NanoFlow and the Future of LLM Serving
Air Force One, Jensen Huang, and Anthropic's 2028 Memo
Trace Rewriting Against Unauthorized LLM Distillation
Agentic AI as a Path to AGI
When Many-Shot CoT Becomes Test-Time Learning
Ministral 3: Cascade Distillation for Long-Context Multimodal Models
Causal-JEPA for Object-Level World Models
Scaling Laws for Multilingual Code Pretraining
JANUS for Scalable MoE Inference
Lossless Sparse Deltas for RL Networks
FlashFuser and Hopper-Era FFN Kernel Fusion
Deep Kernel Fusion for Transformer Decoding
AI Co-Mathematician for Mathematical Research
TMAS: Scaling Test-Time Compute with Multi-Agent Synergy
MELT: Decoupling Compute From Memory
Long Context Pre-Training with Lighthouse Attention
MiA-Signature and Global Activation for Long Context
Qwen-Image-2.0 for Unified Generation and Editing
δ-mem and Online Memory for LLMs
Agentic Discovery for Test-Time Scaling
TIDE and the Rare Token Problem
Metacognition Against Confident Hallucinations
ForkKV for Multi-LoRA Agent Serving
ELF and Continuous Language Diffusion
Vistara Brings CXL Memory to Hyperscale
RAPTOR: Stable Concept Directions From Logistic Probes
Reasoning Theater and Unfaithful Chain-of-Thought
Optimization, Credit Assignment, and Consciousness
Why Transformers Fail at Counting
Split Personality Training Reveals Latent Knowledge
Explicit Information Transmission for Context Compression
Why LLM Serving Needs Mathematical Optimization
EverMemOS for Long-Horizon Agent Memory
When LLM Judges Become Coin Flips
Generative Modeling via Drifting in One Step
LAPS for Length-Aware LLM Serving
Machine Learning Self-Calibrated FPGA Time-to-Digital Converter
TensorFlow for Distributed Machine Learning Systems
SGLang for Faster Structured LLM Programs
Can Models Learn from Long Context?
Backpropagation Through Time Explained
Learning in Random Nets and Generalization
What the Frog’s Eye Tells the Brain
How Models Detect Hidden Activation Steering
Caffe and the Rise of CNN Frameworks
Synchronous Data Flow for Signal Processing
Automating DNN Compilation for FPGA Accelerators
Caffeine: A Unified FPGA for CNNs
Boosted Decision Trees for CMS Muon Triggers
Fast FPGA Inference for LHC Triggers
Fast FPGA BDT Inference for LHC Triggers
Why LightGBM Made Boosted Trees Fast
Automating CNN Mapping on Embedded FPGAs
RFNoC SISO Processor via High-Level Synthesis
Reinforcement Learning in 2025: An Overview
Training LLMs for Divide-and-Conquer Reasoning
PackKV Lossy Compression for KV Caches
Training Million-Token LLMs Beyond the Memory Barrier
How Induction Heads Emerge in Transformers
DeepWalk and the Rise of Graph Embeddings
node2vec and Learning Graph Embeddings
Selective Classification with Deep Neural Networks
Geometric Memory in Deep Sequence Models
Self-Improving Pretraining With Post-Trained Models
Do Language Models Know Their Limits
Fast Speech Recognition by Transcript Editing
CacheFlow and 3D-Parallel KV Cache Restoration
DeltaKV: Compressing KV Caches for Long Context
A Practical Review of Mechanistic Interpretability
Recursive Multi-Agent Systems in Latent Space
Discrete Representations for Continual Reinforcement Learning
Teaching Language Models to Verbalize Uncertainty
ChartNet for Robust Multimodal Chart Understanding
Can LLMs Judge Their Own Capabilities?
Convolutional Sequence to Sequence Learning for Translation
Bahdanau Attention for Neural Machine Translation
Sequence to Sequence Learning with Neural Networks
Hopfield Networks and Transformer Attention as Memory
Scaling Test-Time Compute for Reasoning Models
Reverse-Mode Differentiation Across AD and Neural Nets
Breaking the Prefix Barrier with Shared KV Cache
Stochastic KV Routing for Cache Sharing
Stabilizing Efficient Reasoning with Step-Level Advantage Selection
World-R1 Improves 3D Consistency in Text-to-Video
FPGA Neural Network Accelerators for Space
Deep Learning in Spiking Neural Networks
Structural Understanding of LLM Overthinking
AdamW: Decoupled Weight Decay Regularization for Adaptive Gradient Algorithms