Learning GenAI via SOTA Papers cover art

All Episodes

Learning GenAI via SOTA Papers — 433 episodes

#
Title
1

EP432: Motif 3 Replaces Brute Force with Specialization

2

EP431: Yale MoRSE ends AI agent redundancy

3

EP430: Khora Scales Real Time AI Hallucinated Worlds

4

EP429: AI agents redesigning their own software harnesses

5

EP428: Tiny AI beats giants with silent logic

6

EP427: Replacing AI reasoning with distilled skills

7

EP426: HiLP enables long horizon AI planning

8

EP425: AI agents playing actor and environment

9

EP424: Why Argus AI Thrives on Dead Ends

10

EP423: Why Agentic AI breaks the datacenter

11

EP422: Stopping spurious signals in AI distillation

12

EP421: Fixing AI Hallucinations with RAIL Principles

13

EP420: Hijacking AI memory via factual injection

14

EP419: Ten Weeks of Autonomous AI Research

15

EP418: DeepVoyager-VL solves the visual search bottleneck

16

EP417: AI agents replace human beta testers

17

EP416: How AdaThinkV stops AI overthinking video

18

EP415: Fixing the AI granularity mismatch

19

EP414: Why context compaction breaks AI agents

20

EP413: Slashing AI latency with uncertainty repair

21

EP412: Thermodynamic Computing Solves the AI Bottleneck

22

EP411: How NeSyFS Gives AI Fast-Slow Thinking

23

EP410: How provenance laundering brainwashes AI

24

EP409: Robots That Dream Before They Move

25

EP408: AI memory reconstructed not replayed

26

EP407: How AI learns your teamwork capabilities

27

EP406: Ending AI Groundhog Day With Living Harness

28

EP405: Are AI Agents Just Talking to Themselves

29

EP404: AI agents hide betrayal in Werewolf

30

EP403: COVENANT keeps AI agents on the rails

31

EP402: Static baselines beat dynamic AI agents

32

EP401: Extracting Pure Reasoning From AI Giants

33

EP400: Can GPT-5.1 understand the world

34

EP399: Training AI to follow any reasoning workflow

35

EP398: Social deduction games teach AI creativity

36

EP397: Teaching AI teams to focus

37

EP396: How numerical scores trigger AI reinforcement learning

38

EP395: ConsistencyGate stops AI memory contamination

39

EP394: Agentic Context Management Beats Raw Compute

40

EP393: Why AREX agents audit their own research

41

EP392: Small models beat giants at malware analysis

42

EP391: Programmatic memory fixes AI context rot

43

EP390: EvoDRC solves microscopic silicon design errors

44

EP389: Solving the AI Memory Trilemma

45

EP388: Machine translation with latent reasoning loops

46

EP387: Shared libraries for disposable AI agents

47

EP386: Infinite playable worlds on a single GPU

48

EP385: AI self-correction can destroy correct answers

49

EP384: How AI can finally stop forgetting

50

EP383: Why AI Agents Disobey Their Own Logic

51

EP382: Etas The Native Language For AI Agents

52

EP381: How AI finally learned to smell

53

EP380: AI rewriting itself to solve formal math

54

EP379: Smaller AI beats giants by thinking twice

55

EP378: Giving AI a Silent Inner Monologue

56

EP377: PRIME Solves AI Curiosity Traps

57

EP376: ToolVerse Teaches AI to Execute Complex Tasks

58

EP375: Bypassing AI Hype With Research Papers

59

EP375: M2GDT solves multimodal knowledge graph completion

60

EP374: TopoAgent Outperforms GPT-5 in Science

61

EP373: Middle Layer Recurrence Fixes AI Amnesia

62

EP372: How UrbanAgent profiles unseen cities

63

EP371: Groc-PO Stops Multimodal AI Hallucinations

64

EP370: SLEUTH fixes AI multi-hop reasoning failures

65

EP369: How Atomic Units Scale Intelligence

66

EP368: Samba Framework for Audio-Visual Navigation

67

EP367: Why AI sounds so painfully corporate

68

EP366: Autonomous AI Writes Its Own Hacking Tools

69

EP365: Smarter managers beat bigger AI brains

70

EP364: Capability Trees for Scalable AI Agents

71

EP363: How Logos Architecture Stops AI Misevolution

72

EP362: How Agentic-DPO fixes brittle AI agents

73

EP361: How Riemannian geometry fixes AI reasoning

74

EP360: How ARMOR stops AI reasoning collapse

75

EP359: Why your AI should forget

76

EP358: Europe s Transparent Soofi S AI Blueprint

77

EP357: Copying Smart Experts Makes AI Worse

78

EP356: CMA solves the visual token explosion

79

EP355: RL builds compositional reasoning strategies

80

EP354: How AI Agents Code Their Own Habits

81

EP353: How IGRPO stops AI search distractions

82

EP352: Hidden states predict AI agent failure

83

EP351: Direct-OPD slashes AI reasoning compute costs

84

EP350: Training AI agents without live environments

85

EP349: Fixing AI judges with continuous verification

86

EP348: Building AI agents like living cells

87

EP347: Compiling AI into Permanent Free Skills

88

EP346: Teaching small AI to ignore teachers

89

EP345: AI agents retry from pivotal mistakes

90

EP344: AI predicts tool calls to skip waiting

91

EP343: How AI agents escape infinite loops

92

EP342: Why process rubrics triple AI accuracy

93

EP341: Gemma 4 brings thinking mode to laptops

94

EP340: AI Models Prove Opposite Scientific Truths

95

EP339: How AI Safely Rewrites Its Own Code

96

EP338: DiscoPER conducts autonomous science via reflection

97

EP337: Why AI Agents Fail in Silence

98

EP336: ACE fixes the AI goldfish memory problem

99

EP335: How AI agents learn from failure

100

EP334: Fixing AI Hallucinations With Process Rewards

101

EP333: Logic not length makes AI smarter

102

EP332: AI Architects Designing Better Embodied Agents

103

EP331: Internalizing AI debate with Mixture of Debaters

104

EP330: AI agents audit 10,000 page nuclear reports

105

EP329: Teaching AI to forget the right things

106

EP328: FlowWM and branching futures

107

EP327: Why Chatbot Safety Training Backfires for Agents

108

EP325: Why robots have too much brain

109

EP324: JERP synchronizes AI rules and neural weights

110

EP323: Giving AI Einstein s visual imagination

111

EP322: Why cliff tokens break AI math

112

EP321: Measuring AI intelligence in bits

113

EP320: Universal AI is mathematically impossible

114

EP319: How TRUSTMEM Fixes Broken AI Memory

115

EP318: Open Data Recipes for AI Agents

116

EP317: The Architecture Of Genuine Artificial Agency

117

EP316: Teaching robotaxis the biological urge to survive

118

EP315: Teaching robots to think like scientists

119

EP314: Why AI hacks its own geometry

120

EP313: How ARTS reasons through its own failures

121

EP312: BioMatrix translates English to 3D biology

122

EP311: Why AI Teams Hallucinate Together

123

EP310: Why AI Breaks While Fixing Itself

124

EP309: AutoRAS builds self-healing AI agent networks

125

EP308: Giving AI Agents Mathematical Muscle Memory

126

EP307: AI agents now train physical robots autonomously

127

EP306: AIs that engineer their own pipelines

128

EP305: Mathematical guardrails for autonomous AI agents

129

EP304: MagicSim Bridges AI and Physics

130

EP303: How MODE-RAG stops AI video lies

131

EP302: Transferable interaction patterns for web agents

132

EP301: VeriGraph Makes AI Data Analysis Verifiable

133

EP300: Tensors prevent multi-agent LLM collisions

134

EP299: STRIDE grades the AI scratchpad

135

EP298: LLM-as-Code Fixes Unreliable AI Agents

136

EP297: How T-Mem fixes the associative blind spot

137

EP296: Stop parallel AI agents from crashing production

138

EP295: Ending agent sprawl with canonical code

139

EP294: Why AI agents second-guess their success

140

EP293: Grading AI blueprints with Orch-RM

141

EP292: Agents-K1 turns AI into research scientists

142

EP291: Ouroboros-Spatial Outperforms AI Giants in 3D

143

EP290: How Knowledge Graphs Fix Multi-Hop Reasoning

144

EP289: Runtime governance for autonomous AI agents

145

EP288: Test-Time Training shatters quadratic sampling limits

146

EP287: Small models beat GPT-4o with Role-Agent

147

EP286: ReasonAlloc Solves the AI Memory Bottleneck

148

EP285: How SkeMex builds medical AI intuition

149

EP284: Compressing massive context into soft tokens

150

EP283: Aligning AI planners with tool capabilities

151

EP282: AI gladiators training in shopping arenas

152

EP281: Restoring plasticity to over-trained AI

153

EP280: Trajectory Refined Distillation Fixes AI Reasoning

154

EP279: Ending AI amnesia with strategy cards

155

EP278: Hacking AI Agents with Fake Errors

156

EP277: AI quorums stop cloud infrastructure failures

157

EP276: ThinkBooster scales LLM reasoning at test time

158

EP275: AI Agents Building Their Own Coding Curriculum

159

EP274: Knowledge graphs fix AI memory loss

160

EP273: Why agents make code disposable

161

EP272: AI rewiring its own brain live

162

EP271: Steer locked AI with Agentic Monte Carlo

163

EP270: AI agents building their own reasoning tools

164

EP269: Securing AI Agents with Agent libOS

165

EP268: How OpenWebRL masters the live web

166

EP267: AI Agents That Update Their Own Imagination

167

EP266: AI agents learn to think without words

168

EP265: How AI agents rewrite their own tools

169

EP264: Science Earth and Planet Scale AI Discovery

170

EP263: How POPO ends AI training waste

171

EP262: Web agents that learn from failure

172

EP261: EchoRL turns hesitation into genius

173

EP260: GrepSeek brings Unix precision to AI

174

EP259: The ESPO Kill Switch For AI Reasoning

175

EP258: TRACER teaches AI to stay silent

176

EP257: How planning wakes up deep AI layers

177

EP256: Teaching AI to Doubt Its Own Answers

178

EP255: MUSE-Autoskill creates self-evolving AI agents

179

EP254: Why Innovation Guarantees AI Hallucination

180

EP253: MACA optimizes AI agent coordination

181

EP252: How batch sizes sharpen AI reasoning

182

EP251: How SR2AM stops AI overthinking

183

EP250: Compiling agent workflows into model weights

184

EP249: Mem-pi fixes AI amnesia with generative memory

185

EP248: 10x Faster AI Agents with JIT Compilation

186

EP247: PEEK Cures AI Goldfish Memory

187

EP246: Replacing AI manuals with programmable runtimes

188

EP245: The Geometric Shape of AI Reasoning

189

EP244: Training decentralized AI through private handoffs

190

EP243: Breaking the AI data wall with SYNPRO

191

EP242: Ending AI Amnesia with Experience Graphs

192

EP241: Accelerating game theory with linear algebra

193

EP240: Small AI agents beat giants with Orchard

194

EP239: The shift from chatbots to AI societies

195

EP238: SepsisAgent outperforms clinicians using clinical world models

196

EP237: Why AI agents must map before acting

197

EP236: AI agents rewriting their own code

198

EP235: How SAGE Fixes AI Memory

199

EP234: FATE fixes safe but useless AI agents

200

EP233: Fixing AI memory with backward chaining

201

EP232: Why AI agents lie to fit in

202

EP231: Amazon PIVOT solves the AI execution gap

203

EP230: DeepRefine fixes messy AI knowledge bases

204

EP229: Ending the AI verbosity tax with LEAD

205

EP228: Why self-evolving AI forgets basic tasks

206

EP227: FlowAgent fixes the AI tool bottleneck

207

EP226: MELT Decouples AI Reasoning from Memory

208

EP225: Turning AI into its own lie detector

209

EP224: Soft-Hamiltonian world models for robust planning

210

EP223: UNO-ORCHESTRA Slashes AI Costs via Selective Delegation

211

EP222: Gyan Beats GPT-4o Without Using GPUs

212

EP222: Gyan Beats GPT-4o Without Using GPUs

213

EP221: ScrapMem Mimics Human Memory Through Forgetting

214

EP220: How PARSE Makes AI Four Times Faster

215

EP219: OpenSeeker V2 Shatters The AI Compute Myth

216

EP218: JoyAI-Image Solves AI 3D Geometry Errors

217

EP217: Why forced compliance triggers metacognitive collapse

218

EP216: Shadow memory stops long horizon AI heists

219

EP215: Finding specialized AI agents in milliseconds

220

EP214: ARISE Maps Data Flow For AI Agents

221

EP213: Why AI agents fail at negotiation

222

EP212: Sheaf Geometry Fixes Robot Logic

223

EP211: SciResearcher turns AI into a scientific detective

224

EP210: AI that rewrites its own logic

225

EP209: Fixing AI agent memory with SAGA

226

EP208: Bayesian Orchestration for Overconfident AI Agents

227

EP207: Robots learn the math of anticipation

228

EP206: ObjectGraph replaces Markdown for AI agents

229

EP205: Qiushi AI Discovers Optical Computing Hardware

230

EP204: Solving the AI compositionality crisis

231

EP203: How AI Agents Trade Real Money

232

EP202: Why ADEMA AI Never Loses The Plot

233

EP201: Nautile-370M solves AI memory bottlenecks

234

EP200: Kwai Summary Attention and the memory wall

235

EP199: Separation of Powers for AI Safety

236

EP198: AI masters StarCraft using chat logs

237

EP197: Teaching AI Agents to Plan Like Humans

238

EP196: Forcing AI to Prove Its Logic

239

EP195: How tool attention ends the tools tax

240

EP194: AI coding through mental simulation

241

EP193: AI image generators master physical reality

242

EP192: Fixing AI memory with knowledge graphs

243

EP191: Why AI Agents Blame Each Other

244

EP190: [OLLM] Replacing AI dice rolls with ten lanes

245

EP189: How Sessa architecture fixes AI amnesia

246

EP188: [Agent-World] AI Building Its Own Training Worlds

247

EP187: Hive fixes multi-agent AI memory bottlenecks

248

EP186: Harness engineering for near perfect small models

249

EP185: Why AI architecture fails at logic

250

EP184: Defeating the AI consensus trap

251

EP183: AI coding agents cheat with keywords

252

EP182: AI logic is its weakest link

253

EP181: Small models beating GPT-5 with logic

254

EP180: How AI agents rewrite their code

255

EP179: AIBuildAI Builds New AI Models From Scratch

256

EP178: AI agents reaching silent latent consensus

257

EP177: CAPO math stops overconfident AI lies

258

EP176: Trigonometry fixes the AI memory bottleneck

259

EP175: How AI models teach themselves reasoning

260

EP174: 1-bit Bonsai brings powerful AI offline

261

EP173: AI models diagnosing diseases from blank scans

262

EP172: How HyperAgents rewrite their own code

263

EP171: Helium makes AI agent workflows 40x faster

264

EP170: Qwen3.5 Multimodal Agent

265

EP169: Cybersecurity Risks of Autonomous AI Agents

266

EP168: Turning AI Agents into Mathematical Functions

267

EP167: Why AI models ignore visual evidence

268

EP166: The Auton solution to the integration paradox

269

EP165: Translating hidden AI logic into English

270

EP164: [LACONIC] Teaching AI to stop overthinking

271

EP163: Why AI Models Only Remember Five Percent

272

EP162: AI agents beat humans with malicious skills

273

EP161: Small AI Judges Beat Massive Coding Giants

274

EP160: [AgentSys] Securing AI agents with hierarchical memory

275

EP159: Brute force scale dominates the AI frontier

276

EP158: The hidden blind spots of AI logic

277

EP157: [AgentHeLLM] Protecting drivers from hijacked vehicle AI

278

EP156: [Uncertainty Quantification] How AI Agents Know They Are Guessing

279

EP155: [Agentic Proposing] Small models beat giants with logic bricks

280

EP154: [FS-Researcher] Giving AI agents a file system

281

EP153: [SERA] Training AI coding agents on untested code

282

EP152: DeepVerifier forces AI to check its work

283

EP151: [MagicGUI-RMS] AI agents that think before they click

284

EP150: The Leap to Autonomous Agentic Reasoning

285

EP149: [IDRBench] Interactive AI beats lone wolf models

286

EP148: How AI masters math through self-correction

287

EP147: [DeepSynth-Eval] AI fails at deep research synthesis

288

EP146: How InfiAgent solves the AI memory bottleneck

289

EP145: [LongDA] Why smart AI fails at messy data

290

EP144: [Evo-Memory] Building AI agents with self-evolving memory.

291

EP143: Your AI will blackmail you to survive

292

EP142: [DR-Arena] A ruthless arena for deep research agents

293

EP141: [AIRS-Bench] AI agents beat human research benchmarks

294

EP140: [LeWorldModel] AI learns physics on one GPU

295

EP139: Mamba-3 Fixes the Transformer Memory Bottleneck

296

EP138: [Mamba-2] Transformers and SSMs Are the Same Engine

297

EP137: Attention Residuals Solve the LLM Depth Bottleneck

298

EP136: Modular skills for autonomous AI agents

299

EP135: [SoK] Curing AI Amnesia with Agentic Skills

300

EP134: Autonomous AI squads building software

301

EP133: RelayLLM Slashes AI Costs With Collaborative Decoding

302

EP132: How Autonomous LLM Agents Actually Work

303

EP131: MUSE creates self evolving AI agents

304

EP130: [GAP] Graph-based planning for faster AI agents

305

EP129: Why AI agents fail half the time

306

EP128: MCP-Zero lets AI find its own tools

307

EP127: Why tool use makes AI less intelligent

308

EP126: OrcaLoca locates bugs in massive codebases

309

EP125: Why AI Needs an Agent Computer Interface

310

EP124: FRIDAY the AI that runs your computer

311

EP123: MemGPT Turns LLMs into Operating Systems

312

EP122: The Four Pillars of LLM Autonomous Agents

313

EP121: How ToolLLaMA mastered 16000 real world APIs

314

EP120: How Reflexion agents learn through verbal feedback

315

EP119: HuggingGPT Turns LLMs Into AI Managers

316

EP118: The AI Memory Wall Crisis

317

EP117: AI agents learn through textual reflection

318

EP116: Why AI struggles with empathy and interruptions

319

EP115: Dr.LLM brings dynamic depth to AI

320

EP114: FlashAttention-4 Solves Blackwell Hardware Bottlenecks

321

EP113: How FlashAttention-3 Doubles H100 Speed

322

EP112: GPT 5.4 Outperforms Human Professionals

323

EP111: Claude Opus 4.6 Runs Businesses and Catches Manipulation

324

EP110: Single agents beat expensive multi agent teams

325

EP109: The Rise of Agentic Reasoning

326

EP108: GPT-5 Can Lie and Play Dumb

327

EP107: DeepMind’s SIMA 2 Masters Unseen Video Games

328

EP106: Fixing AI Agents With Symbolic Guardrails

329

EP105: iStar Autonomous Agents Grading Their Own Homework

330

EP104: WebExplorer Beats Giants at Web Research

331

EP103: Why AI Agents Think Themselves To Death

332

EP102: Gemini 2.5 Thinks Before It Speaks

333

EP101: Kimi k1.5 Breaks the AI Data Wall

334

EP100: Meta's Llama 4 Herd Ends Monolithic Models

335

EP099: Is AI Thinking Just Expensive Noise

336

EP098: OpenAI o3 Hacked Its Own Grading System

337

EP097: DeepSeek R1 Taught Itself to Reason

338

EP096: Gemini 1.5 Pro's 10 Million Token Window

339

EP095: Microsoft Phi-4 Beats Giants With Synthetic Data

340

EP094: DeepSeek-V3 Rivals GPT-4 for $6 Million

341

EP093: How OpenAI o1 Cracked the Strawberry Cipher

342

EP092: BitNet b1.58 Replaces Multiplication With Addition

343

EP091: Qwen 2.5 Beats Llama With Synthetic Data

344

EP090: Pixtral 12B Beats Llama With Better Eyesight

345

EP089: Qwen2-VL Gives AI Native Eyesight

346

EP088: Qwen2 Beats Llama-3 Through Data Quality

347

EP087: Meta's Chameleon Unifies Text and Images

348

EP086: DeepSeek-V2 Breaks The Impossible Triangle

349

EP085: Aya 23 Breaks The Curse Of Multilinguality

350

EP084: Microsoft Phi-3 Fits Supercomputing in Your Pocket

351

EP083: How Meta Engineered the Llama 3 Herd

352

EP082: Command R Plus The Verifiable Enterprise Agent

353

EP081: Replacing MLPs With Interpretable KANs

354

EP080: Jamba Hybrid Solves Transformer Memory Limits

355

EP079: DBRX Beats GPT-3.5

356

EP078: Claude 3 Knew It Was Being Tested

357

EP077: Google Squeezes Gemini Into Your Laptop

358

EP076: OLMo Cracks Open the AI Black Box

359

EP075: Microsoft Phi Beats Giants With Synthetic Textbooks

360

EP074: How Gemini Beat Human Experts

361

EP073: Mixtral 8x7B Sparse Experts Beat Giants

362

EP072: Mamba Solves The Transformer's Fatal Flaw

363

EP071: How Zephyr-7B Beat Llama-70B

364

EP070: Mistral 7B Beats Llama 2 13B

365

EP069: Alibaba's Qwen Specialized Models Beat Generalists

366

EP068: vLLM Fixes the KV Cache Bottleneck

367

EP067: FlashAttention-2 Unlocks Massive Context Windows

368

EP066: Llama 2 Ghost Attention And Safety Secrets

369

EP065: Teaching Small AI To Think Like Giants

370

EP064: Synthetic Textbooks Break AI Scaling Laws

371

EP063: RWKV Smashes the Transformer Memory Ceiling

372

EP062: VOYAGER AI Masters Minecraft by Writing Code

373

EP061: Fine-Tuning LLaMA 65B on One GPU

374

EP060: Direct Preference Optimization Replaces RLHF

375

EP059: Tree of Thoughts Unlocks System 2 Thinking

376

EP058: Inside the Autonomous AI Town of Smallville

377

EP057: Blind GPT-4 Taught LLaVA To See

378

EP056: Pythia Turns AI Alchemy Into Chemistry

379

EP055: Can GPT-4 Fairly Judge Other AI

380

EP054: Alpaca - Stanford Built a $600 GPT Clone

381

EP053: Sparks of AGI in Early GPT-4

382

EP052: GPT-4 Bar Exam and Visual Reasoning

383

EP051: ControlNet Solves Spatial Control With Zero Convolutions

384

EP050: How Meta's LLaMA Beat GPT-3

385

EP049: Toolformer Teaches Itself to Use APIs

386

EP048: BLIP-2 Teaches Frozen Models to See

387

EP047: Bootstrapping AI With Self-Generated Instructions

388

EP046: Training AI With A Constitution

389

EP045: BLOOM The Open Source Rival To GPT-3

390

EP044: How ReAct Synergizes Reasoning and Acting

391

EP043: Weak Supervision Made OpenAI Whisper Robust

392

EP042: Running 175B Models on Consumer Hardware

393

EP041: FlashAttention Smashes the AI Memory Wall

394

EP040: Meta's Open Source GPT-3 Replica

395

EP039: Flamingo Unlocks Few-Shot Visual Reasoning

396

EP038: PaLM's 540 Billion Parameters Unlock Reasoning

397

EP037: DeepMind Chinchilla Ends The Parameter Wars

398

EP036: How 40 People Taught GPT-3 Manners

399

EP035: How Google LaMDA Learned To Use Tools

400

EP034: Chain of Thought Prompting Unlocks Reasoning

401

EP033: Democratizing Image Generation with Latent Diffusion

402

EP032: WebGPT Fights Hallucinations With Web Search

403

EP031: DeepMind RETRO Swaps Memorization For Retrieval

404

EP030: DeepMind's Gopher Exposes Limits of Scale

405

EP029: Instruction Tuning Unlocked Zero-Shot Learning

406

EP028: Train Short for Infinite Context

407

EP027: From Creative Writer to Logic Engine

408

EP026: LoRA Fine-Tunes Massive Models Without Supercomputers

409

EP025: RoPE Solves Sequence by Rotating Vectors

410

EP024: OpenAI CLIP Bridges Language and Vision

411

EP023: Scaling Switch Transformers to Trillion Parameters

412

EP022: DALL-E Treats Images Like Language

413

EP021: Vision Transformers Beat CNNs at Scale

414

EP020: Big Bird Scales Transformers With Sparse Attention

415

EP019: Facebook's Linformer Solves the Attention Bottleneck

416

EP018: Turning Digital Static Into Images With Diffusion

417

EP017: RAG Gives AI a Library Card

418

EP016: GPT-3 Learns From Examples Without Retraining

419

EP015: Longformer Smashes the 512 Token Barrier

420

EP014: ELECTRA Beats GPT On One GPU

421

EP013: Reformer Cracked the Transformer Memory Wall

422

EP012: Google T5 Turns Every Task Into Text

423

EP011: ZeRO Solved the Trillion Parameter Memory Wall

424

EP010: ALBERT Outperforms BERT With Parameter Sharing

425

EP009: Slicing the AI Brain with Megatron-LM

426

EP008: RoBERTa Proves BERT Was Just Undertrained

427

EP007: How GPT-2 Hallucinated Ovid's Unicorn

428

EP006: Transformer-XL Cures AI Amnesia

429

EP005: How BERT Mastered Language by Hiding Words

430

EP004: How 7000 Unpublished Books Birthed GPT

431

EP003: How ELMo Made Word Vectors Dynamic

432

EP002: ULMFiT Was the ImageNet Moment for Text

433

EP001: How Transformers Smashed the Sequential Bottleneck