Best AI papers explained cover art

All Episodes

Best AI papers explained — 815 episodes

#
Title
1

Recursive Experiential–Working Memory Evolution for Long-Horizon Agent Harnesses

2

TailSFT: Filtered Fine-Tuning Improves Post-Training Performance

3

SPADE: Self-Play in Adaptive Synthetic Executable Environments

4

Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

5

Impression Share Prediction: An Offline Evaluation Task for Ranking Systems

6

Q-Learning with World Models

7

Conformal Language Modeling via Posterior Sampling

8

BoNVoyage: Learning Better Rewards without Ranking

9

Demystifying Agent Skills: Why They Work—Until They Don’t

10

Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence

11

Predicting Neural Scaling Laws without Training: A Data Manifold Oracle

12

Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing

13

Overcoming the Incentive Collapse Paradox

14

Position: Modular Memory is the Key to Continual Learning Agents

15

Harness RL is Meta-Learning: Training to Self-Improve at Test Time

16

Escaping the Nash Trap: Structural Estimation and Alignment of Strategic Reasoning in Large Language Models

17

When Does LeJEPA Learn a World Model?

18

Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems

19

Do you really need to pretrain Q-functions for online RL fine-tuning?

20

The Evolution of Digital Search: From Blue Links to Delegated Decision-Making

21

Ask, Don’t Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

22

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning

23

Understanding Reasoning from Pretraining to Post-Training

24

A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior

25

Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference

26

Rethinking the Evaluation of Harness Evolution for Agents

27

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning

28

Position: Interpretability can be actionable

29

High-accuracy sampling for diffusion models and log-concave distributions

30

Causal Inference with Video Features as Treatments

31

What Does Thompson Sampling Optimize?

32

Globally Convergent Offline Reinforcement Learning with Smoothed Bellman Residual Minimization

33

LLM-as-a-Verifier: A General-Purpose Verification Framework

34

How Much Do Language Models Memorize?

35

Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering

36

Position: Agents Should Invoke External Tools ONLY When Epistemically Necessary

37

From conversations to mechanisms: aligning advertiser Incentives in ai-powered product recommendations

38

Is one layer enough? Training a single transformer layer can match full-parameter RL training

39

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

40

Language Generation with Feedback: Queries and Mistakes

41

Quantifying Theoretical AI Alignment Guarantees: Receiver-Utility Bounds in Bayesian Persuasion

42

SPIRAL: Learning to search and aggregate

43

Qwen-AgentWorld: Language World Models for General Agents

44

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?

45

SuperThoughts: Reasoning Tokens in Superposition

46

First-Explore PPO : Learning Meta-Exploration with Proximal Policy Optimization

47

Self-Distillation for Data-Scarce Language Model Pretraining

48

Meta-Harness for Agent-State Construction

49

ExpRL: Using Reference Solutions as Rewards for LLM Mid-Training

50

Valid Inference with Synthetic Data via Task Exchangeability

51

GRPO is Secretly a Process Reward Model

52

Agentic Interactions

53

A Unifying View of Attention Sinks: Two Algorithms, Two Solutions

54

From AGI to ASI

55

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings

56

Critical Batch Size for LLM Policy Optimization

57

Self-supervised User Profile Generation for Personalization

58

From Augmentation to Reconstruction: Guiding the AI Disruption to the Good Place

59

Self-Distilled Agentic Reinforcement Learning

60

Subliminal Learning Is Steering Vector Distillation

61

Subsidizing Sequential Search

62

Meta-Harness: End-to-End Optimization of Model Harnesses

63

Self-Improving Language Models with Bidirectional Evolutionary Search

64

Generative Modeling via Drifting

65

Instance-Optimal Estimation with Multiple LLM Judges on a Budget

66

Robust AI Personalization Will Require a Human Context Protocol

67

Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning

68

Position: The Pre/Post-Training Boundary Should Govern IP in Industry–Academia ML Collaborations

69

MEMO: Memory as a Model

70

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces

71

General Preference Reinforcement Learning

72

Explaining and Preventing Alignment Collapse in Iterative RLHF

73

Curriculum Learning-Guided Progressive Distillation in Large Language Models

74

Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents

75

How Much Should a Conversational Recommender System Converse?

76

FUSE: Ensembling Verifiers with Zero Labeled Data

77

EVOLM: Self-Evolving Language Models through Co-Evolved Discriminative Rubrics

78

Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity

79

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies

80

Adaptive Querying with AI Persona Priors

81

Rethinking the Role of LLMs in Time Series Forecasting

82

Robust Representation Learning through Explicit Environment Modeling

83

Magentic Marketplace: An Open-Source Environment for studying Agentic Markets

84

Hyperloop Transformers

85

Scaling Self-Play with Self-Guidance

86

RL Token: Bootstrapping Online RL with Vision-Language-Action Models

87

Agentic Data Environments

88

AI organizations are more effective but less aligned than individual agents

89

Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context

90

Distortion of AI alignment revisited: RLHF is a decent utilitarian aligner

91

Llms get lost in multi-turn conversation

92

Transformers are inherently succint

93

The Coasean Singularity? Demand, Supply, and Market Design with AI Agents

94

Demystifying the unreasonable effectiveness of online alignment methods

95

Specialization after generalization: towards understanding test-time training in foundation models

96

Exploration and Exploitation Errors Are Measurable for Language Model Agents

97

A Mechanistic Analysis of Looped Reasoning Language Models

98

Sample Complexity of Autoregressive Reasoning: Chain-of-Thought vs. End-to-End

99

Why AI systems don’t learn and what to do about it

100

The Illusion of Learning from Observational Data: An Empirical Bayes Perspective

101

Ads in AI chatbots? An analysis of how large language models navigate conflicts of interest

102

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models

103

LLM Evaluation as Tensor Completion: Low-Rank Efficiency and Uncertainty Quantification

104

Neural Computers

105

How AI Aggregation Affects Knowledge

106

World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry

107

In-Place Test-Time Training

108

Test-Time Scaling Makes Overtraining Compute-Optimal

109

AI Agent Prevalence and Data Quality Across Multiple Online Sample Providers

110

POLCA: Stochastic Generative Optimization with LLM

111

Agentic Markets: Equilibrium Effects of Improving Consumer Search

112

One Model, Two Markets: Bid-Aware Generative Recommendation

113

How Well Do LLMs Predict Human Behavior? A Measure of their Pretrained Knowledge

114

Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum

115

Agentic AI and the next intelligence explosion

116

Understanding Behavior Cloning with Action Quantization

117

HyperAgents: : Open-Ended Metacognitive Self-Improvement for Any Computable Task

118

Harness design for long-running application development \ Anthropic

119

Reasonably reasoning AI agents can avoid game-theoretic failures in zero-shot, provably

120

How Log-Barrier Helps Exploration in Policy Optimization

121

The Finetuner’s Fallacy: When to Pretrain with Your Finetuning Data

122

TURNWISE: The Gap between Single- and Multi-turn Language Model Capabilities

123

Temporal Straightening for Latent Planning

124

Fine-Tuning Strategies for Preserving In-Context Learning in Linear Attention

125

LLMs Can Learn to Reason Via Off-Policy RL

126

Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning

127

Provable and practical in-context policy optimization for self-improvement

128

Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models

129

Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights

130

AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization

131

∇−reasoner: LLM reasoning via test-time gradient descent in latent space

132

Inference for Regression with Variables Generated by AI or Machine Learning

133

Fast KV Compaction via Attention Matching

134

Position: stop anthropomorphizing intermediate tokens as reasoning/thinking traces!

135

Code World Models for General Game Playing

136

Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought

137

Task Descriptors Help Transformers Learn Linear Models In-Context

138

Equivalence of Context and Parameter Updates in Modern Transformer Blocks

139

Learning without training: The implicit dynamics of in-context learning

140

Causal Identification from Counterfactual Data: Completeness and Bounding Results

141

Is Cosine-Similarity of Embeddings Really About Similarity?

142

Diffusion LLMs are Natural Adversaries for any LLM

143

Are you going to finish that? A Practical Study of the Partial Token Problem

144

Language Models Struggle to Use Representations Learned In-Context

145

LLMs are Bayesian, In Expectation, Not in Realization

146

Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs

147

LLMs Can Learn to Reason Via Off-Policy RL

148

Test-Time Training with KV Binding Is Secretly Linear Attention

149

Unified Latents (UL): How to train your latents

150

Spectral Bellman Method: Unifying RL Representation and Exploration

151

Prescriptive Scaling Reveals the Evolution of Language Model Capabilities

152

Experiential Reinforcement Learning

153

Learning Personalized Agents from Human Feedback

154

Learning to summarize user information for personalized RLHF

155

Intrinsic Credit Assignment for Long Horizon Interaction

156

Learning to Continually Learn via Meta-learning Agentic Memory Designs

157

Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models

158

PAD: Personalized Alignment of LLMs at Decoding-Time

159

The Reward Model Selection Crisis in Personalized Alignment

160

Causal-JEPA: Learning World Models through Object-Level Latent Interventions

161

How Sampling Shapes LLM Alignment: From One-Shot Optima to Iterative Dynamics

162

Deriving neural scaling laws from the statistics of natural language

163

Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL

164

Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL

165

Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning

166

Owning the AI Pareto Frontier — Jeff Dean

167

Learning to Reason in 13 Parameters

168

Nearly Optimal Active Preference Learning and Its Application to LLM Alignment

169

Language Model Circuits Are Sparse in the Neuron Basis

170

Rethinking the Trust Region in LLM Reinforcement Learning

171

Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward

172

Self-distillation enables continual learning

173

Maximum Likelihood Reinforcement Learning

174

In-Context Algorithm Emulation in Fixed-Weight Transformers

175

PPI-SVRG: Unifying Prediction-Powered Inference and Variance Reduction for Semi-Supervised Optimization

176

When Models Don’t Collapse: On the Consistency of Iterative MLE

177

An orthogonal learner for individualized outcomes In markov decision processes

178

Shaping capabilities with token-level data filtering

179

Self-Improving Pretraining: using post-trained models to pretrain better models

180

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success

181

Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning

182

GameTalk: Training LLMs for Strategic Multi-Turn Conversation

183

Reinforcement Learning via Self-Distillation

184

Self-Supervised Contrastive Learning is Approximately Supervised Contrastive Learning

185

On the alignment between supervised and self-supervised contrastive learning

186

Rethinking the value of multi-agent work-flow: a strong single agent baseline

187

Greedy Sampling Is Provably Efficient for RLHF

188

A Generalization Theory for Zero-Shot Prediction

189

Learning to Discover at Test Time

190

How Does the Pretraining Distribution Shape In-Context Learning? Task Selection, Generalization, and Robustness

191

Highlighting What Matters: Promptable Embeddings for Attribute-Focused Retrieval

192

Activation Reward Models for Few-Shot Model Alignment

193

Reward is enough: LLMs are in-context reinforcement learners

194

Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO

195

The End of Reward Engineering: How LLMs Are Redefining Multi-Agent Coordination

196

PRL: Process Reward Learning Improves LLMs’ Reasoning Ability and Broadens the Reasoning Boundary

197

Coverage Improvement and Fast Convergence of On-policy Preference Learning

198

Stagewise Reinforcement Learning and the Geometry of the Regret Landscape

199

Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

200

Learning Latent Action World Models In The Wild

201

From Unstructured Data to Demand Counterfactuals: Theory and Practice

202

In-context reinforcement learning through bayesian fusion of context and value prior

203

Digital RedQueen: Adversarial Program Evolution in Core War with LLMs

204

Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings

205

Representation-Based Exploration for Language Models: from test-time to post-training

206

NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation

207

RelayLLM: Efficient Reasoning via Collaborative Decoding

208

A Unified Definition of Hallucination, Or: It’s the World Model, Stupid

209

Deep sequence models tend to memorize geometrically; it is unclear why.

210

From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence

211

Diffusion Language Models are Provably Optimal Parallel Samplers

212

Universal Reasoning Model

213

Recursive language models

214

Adapting fast and slow: transportable circuits for few shot learning

215

Position: Probabilistic Modelling is Sufficient for Causal Inference

216

End-to-End Test-Time Training for Long Context

217

Parallel Token Generation for Language Models

218

Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning

219

Activation oracles: training and evaluating llms as general-purpose activation explainers

220

Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning

221

Joint-Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction

222

Monitoring Monitorability/ OpenAI

223

Detailed Balance in Large Language Model-Driven Agents

224

Learning to reason in LLMs by expectation maximization

225

Exploratory Causal Inference in SAEnce

226

Detailed balance in large language model-driven agents

227

The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding

228

Adaptation of Agentic AI

229

Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning

230

Let’s (not) just put things in Context: Test-Time Training for Long-Context LLMs

231

TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models

232

What’s In My Human Feedback? Learning Interpretable Descriptions of Preference Data

233

Bolmo: Byteifying the Next Generation of Language Models

234

What happened with sparse autoencoders?

235

What Matters Right Now in Mechanistic Interpretability

236

CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning

237

Self-Improving AI and Human Co-Improvement for Safer Co-Superintelligence

238

Towards a Science of Scaling Agent Systems / Google Deepmind

239

Emergent hierarchical reasoning in LLMs through reinforcement learning

240

AI revolution finally comes to Relational foundational models for structured data

241

REFRAG: Rethinking RAG based Decoding

242

Provable Long-Range Benefits of Next-Token Prediction

243

Jeff Dean on TPUs, AI Research, and Funding

244

Latent Debate: surrogate framework for Interpreting LLM Thinking

245

Distribution-calibrated inference time compute for thinking llm-as-a-judge

246

Principled RL for diffusion LLMs emerges from sequence level perspective

247

Algorithmic Thinking Theory

248

On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models

249

Natural language actor-critic: Scalable off-policy learning in language space

250

Beyond the Transformer: Titans, MIRAS, and the Future of Infinite Context

251

On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference

252

The Universal Weight Subspace Hypothesis

253

Stabilizing Reinforcement Learning with LLMs: Formulation and Practices

254

Benchmarking In-context Experiential Learning Through Repeated Product Recommendations

255

Training LLMs for Honesty via Confessions

256

STOIC REASONER: Dual-Mode Transformers that Compress to Think and Decompress to Speak

257

E-GEO: A Testbed for Generative Engine Optimization in E-Commerce

258

1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities

259

Treatment Effect Estimation for Optimal Decision-Making

260

Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems

261

Debugging misaligned completions with sparse-autoencoder latent attribution

262

Building Effective AI Agents \ Anthropic

263

How to Correctly Report LLM-as-a-Judge Evaluations

264

In-Context Learning with Hypothesis-Class Guidance

265

Selecting Belief-State Approximations in Simulators with Latent States

266

Latent Collaboration in Multi-Agent Systems

267

CausalPFN: Amortized Causal Effect Estimation via In-Context Learning

268

DELTA: How Does RL Unlock and Transfer New Algorithms in LLMs?

269

Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing

270

Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs

271

Ilya Sutskever – We're moving from the age of scaling to the age of research

272

Cognitive Foundations for Reasoning and Their Manifestation in LLMs

273

Natural emergent misalignment from reward hacking in production RL

274

Evolution Strategies at the Hyperscale

275

The Path Not Taken: RLVR Provably Learns Off the Principals

276

Back to Basics: Let Denoising Generative Models Denoise

277

LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization

278

Black-Box On-Policy Distillation of Large Language Models

279

Solving a million step LLM task with zero errors

280

Not All Thoughts Matter: Selective Attention for Efficient Reasoning

281

Sample-Efficient Parametric Learning from Natural Language

282

Bayesian Optimization in Language space: An Eval-Efficient AI Self-Improvement Framework

283

Context Engineering: Sessions, Memory

284

The Era of Agentic Organization: Learning to Organize with Language Models

285

Understanding neural networks through sparse circuits

286

Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning

287

Multi-Agent Evolve: LLM Self-Improvement Through Co-Evolution

288

LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics

289

PREFDISCO: Evaluating Proactive Personalization through Interactive Preference Discovery

290

Reusing pre-training data at test time is a compute multiplier

291

Scaling Agent Learning via Experience Synthesis

292

Continuous Autoregressive Language Models

293

Toward a Theory of Agents as Tool-Use Decision-Makers

294

Nested Learning: The Illusion of Deep Learning Architectures

295

GST-UNet: A Neural Framework for Spatiotemporal Causal Inference with Time-Varying Confounding

296

Beyond a million tokens: benchmarking and enhancing long-term memory in llms

297

Agentic Economic Modeling

298

Emergent Introspective Awareness in Large Language Models

299

Can Large reasoning models self-train?

300

ALITA-G: Self-Evolving Generative Agent for Agent Generation

301

Self-improving LLM agents at test-time

302

Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization

303

Language models are injective and hence invertible

304

ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory

305

RLAD: Training LLMs to Discover Abstractions

306

How to Train Your Advisor: Steering Black-Box LLMs with ADVISOR MODELS

307

Self-improving LLM agents at Test-Time

308

KL-Regularized Reinforcement Learning is designed to Mode Collapse

309

How do LLMs use their depth?

310

Thought Communication in Multiagent Collaboration

311

Reasoning with Sampling: Base Models Outperform RL

312

Continual Learning via Sparse Memory Finetuning

313

Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences

314

The Coverage Principle: How Pre-Training Enables Post-Training

315

The Era of Real-World Human Interaction: RL from User Conversations

316

Agent Learning via Early Experience

317

Demystifying the Mechanisms Behind Emergent Exploration in Goal-conditioned RL

318

Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior

319

A Definition of AGI

320

Provably Learning from Language Feedback

321

In-Context Learning for Pure Exploration

322

On the Role of Preference Variance in Preference Optimization

323

Training LLM Agents to Empower Humans

324

Richard Sutton Declares LLMs a Dead End

325

Demystifying Reinforcement Learning in Agentic Reasoning

326

Emergent coordination in multi-agent language models

327

Learning-to-measure: in-context active feature acquisition

328

Andrej Karpathy's insights: AGI, Intelligence, and Evolution

329

Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data

330

Representation-Based Exploration for Language Models: From Test-Time to Post-Training

331

The attacker moves second: stronger adaptive attacks bypass defenses against LLM jail- Breaks and prompt injections

332

When can in-context learning generalize out of task distribution?

333

The Art of Scaling Reinforcement Learning Compute for LLMs

334

A small number of samples can poison LLMs of any size

335

Dual Goal Representations

336

Welcome to the Era of Experience

337

Value Flows: Flow-Based Distributional Reinforcement Learning

338

Self-Adapting Language Models

339

The Markovian Thinker

340

Moloch’s Bargain: emergent misalignment when LLMs compete for audiences

341

Transformer Predictor Dynamics and Task Diversity

342

Base models know how to reason, thinking models learn when

343

Spectrum tuning: Post-training for distributional coverage and in-context steerability

344

Understanding Prompt Tuning and In-Context Learning via Meta-Learning

345

MLPs Learn In-Context on Regression and Classification tasks

346

Is Pre-Training Truly Better than Meta-Learning?

347

Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

348

Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs

349

Learning dynamics of LLM finetuning

350

Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF

351

OpenAI Agent Builder and n8n: Orchestrating Reasoning Versus Automating Process

352

Training Agents Inside of Scalable World Models

353

Small Language Models are the Future of Agentic AI

354

Activation Steering in Generative Settings via Contrastive Causal Mediation Analysis

355

Eliciting Secret Knowledge from Language Models

356

Temporal difference flow

357

Personalized reasoning: just-in-time personalization and why LLMs fail at it

358

Prompt Curriculum Learning for Efficient LLM Post-Training

359

Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning

360

Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward

361

Learning to summarize user information for personalized reinforcement learning from human feedback

362

Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

363

LIMI: Less is More for Agency

364

LoRA Without Regret

365

Actor-Critic without Actor: Critic-Guided Denoising for RL

366

DELTA-Code: How Does RL Unlock and Transfer New Programming Algorithms in LLMs?

367

Linear Transformers Implicitly Discover Unified Numerical Algorithms

368

Regularizing Extrapolation in Causal Inference

369

DoubleGen - Debiased Generative Modeling of Counterfactuals

370

What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT

371

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision

372

Learning without training: The implicit dynamics of in-context learning

373

Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model

374

Open Problems in Mechanistic Interpretability

375

Maestro: Joint Graph & Config Optimization for Reliable AI Agents

376

Thought Anchors: Which LLM Reasoning Steps Matter?

377

RL's Razor: Why Online RL Forgets Less

378

Why Language Models Hallucinate

379

ALFA: Aligning LLMs to Ask Good Questions A Case Study in Clinical Reasoning

380

Sample Efficient Preference Alignment in LLMs via Active Exploration

381

Adventures in Demand Analysis Using AI

382

Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

383

On the Theoretical Limitations of Embedding-Based Retrieval

384

Performance Prediction for Large Systems via Text-to-Text Regression

385

Demystifying the Visual Quality Paradox in Multimodal Large Language Models

386

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

387

Compute-Optimal Scaling for Value-Based Deep RL

388

LLM-based Conversational Recommendation Agents with Collaborative Verbalized Experience

389

Signal and Noise: Evaluating Language Model Benchmarks

390

Breaking Feedback Loops in Recommender Systems with Causal Inference

391

RAG is Dead, Context Engineering is King: Building Reliable AI Systems

392

A Survey of Personalization: From RAG to Agent

393

Facilitating the Adoption of Causal Infer-ence Methods Through LLM-Empowered Co-Pilot

394

Performance Prediction for Large Systems via Text-to-Text Regression

395

Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

396

DINOv3: Vision Models for Self-Supervised Learning

397

Agent Lightning: Training Any AI Agents with Reinforcement Learning

398

Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier

399

From Model Weights to Agent Workflows: Charting the New Frontier of Optimization in Large Language Models

400

Is Chain-of-Thought Reasoning a Mirage?

401

Agentic Web: Weaving the Next Web with AI Agents

402

The Assimilation-Accommodation Gap in LLM Intelligence

403

The Minimalist AI Kernel: A New Frontier in Reasoning

404

Statistical Rigor for Interpretable AI

405

Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value

406

A foundation model to predict and capture human cognition

407

Generative Recommendation with Semantic IDs: A Practitioner’s Handbook

408

Hierarchical Reasoning Model

409

Test-time Offline Reinforcement Learning on Goal-related Experience

410

Interpreting Chain of Thought: A Walkthrough and Discussion

411

The wall confronting large language models

412

COLLABLLM: LLMs From Passive to Collaborative

413

A decade's battle on dataset bias: are we there yet?

414

GEPA: Generative Feedback for AI System Optimization

415

From AI-Curious to AI-First: Engineering Production AI Systems

416

Context Engineering: Beyond Simple Prompting to LLM Architecture

417

Agentic Misalignment: LLMs as Insider Threats

418

Small Language Models: Future of Agentic AI

419

Learning without training: The implicit dynamics of in-context learning

420

Inverse Scaling in Test-Time Compute

421

LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra

422

Microsoft's Blueprint: AI, Quantum, and the Agentic Future

423

Zuckerberg's AI Vision Analyzed

424

Inside Claude: Scaling, Agency, and Interpretability

425

Personalized language modeling from personalized human feedback

426

Position: Empowering Time Series Reasoning with Multimodal LLMs

427

An empirical risk minimization approach for offline inverse RL and Dynamic Discrete Choice models

428

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities

429

The Invisible Leash: Why RLVR May Not Escape Its Origin

430

Language Model Personalization via Reward Factorization

431

Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions

432

Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective

433

Soft Best-of-n Sampling for Model Alignment

434

On Temporal Credit Assignment and Data-Efficient Reinforcement Learning

435

Bradley–Terry and Multi-Objective Reward Modeling Are Complementary

436

Probing Foundation Models for World Models

437

GenAI-Powered Statistical Inference (with Unstructured Data)

438

Interpretable Reward Modeling with Active Concept Bottlenecks

439

PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications

440

A Collectivist, Economic Perspective on AI

441

Textual Bayes: Quantifying Uncertainty in LLM-Based Systems

442

The Winner's Curse in Data-Driven Decisions

443

SPIRAL: Self-Play for Reasoning Through Zero-Sum Games

444

Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence

445

Aligning Learning and Endogenous Decision-Making

446

Reliable Statistical Inference with Synthetic Data from Large Language Models

447

Multi-Turn Reinforcement Learning from Human Preference Feedback

448

Provably Learning from Language Feedback

449

Markets with Heterogeneous Agents: Dynamics and Survival of Bayesian vs. No-Regret Learners

450

Why Neural Network Can Discover Symbolic Structures with Gradient-based Training: An Algebraic and Geometric Foundation

451

Causal Abstraction with Lossy Representations

452

The Winner's Curse in Data-Driven Decisions

453

Embodied AI Agents: Modeling the World

454

Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence

455

What Has a Foundation Model Found? Inductive Bias Reveals World Models

456

Language Bottleneck Models: A Framework for Interpretable Knowledge Tracing and Beyond

457

Learning to Explore: An In-Context Learning Approach for Pure Exploration

458

Human-AI Matching: The Limits of Algorithmic Search

459

Uncertainty Quantification Needs Reassessment for Large-language Model Agents

460

Bayesian Meta-Reasoning for Robust LLM Generalization

461

General Intelligence Requires Reward-based Pretraining

462

Deep Learning is Not So Mysterious or Different

463

AI Agents Need Authenticated Delegation

464

Probabilistic Modelling is Sufficient for Causal Inference

465

Not All Explanations for Deep Learning Phenomena Are Equally Valuable

466

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

467

Extrapolation by Association: Length Generalization Transfer in Transformers

468

Uncovering Causal Hierarchies in Language Model Capabilities

469

Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers

470

Improving Treatment Effect Estimation with LLM-Based Data Augmentation

471

LLM Numerical Prediction Without Auto-Regression

472

Why in-context learning models are good few-shot learners?

473

Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina∗

474

The Logic of Machines: The AI Reasoning Debate

475

Layer by Layer: Uncovering Hidden Representations in Language Models

476

Causal Attribution Analysis for Continuous Outcomes

477

Training a Generally Curious Agent

478

Estimation of Treatment Effects Under Nonstationarity via Truncated Difference-in-Q’s

479

Strategy Coopetition Explains the Emergence and Transience of In-Context Learning

480

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

481

Agentic Supernet for Multi-agent Architecture Search

482

Sample Complexity and Representation Ability of Test-time Scaling Paradigms

483

Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

484

LLMs Get Lost In Multi-Turn Conversation

485

PromptPex: Automatic Test Generation for Prompts

486

General Agents Need World Models

487

The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models

488

Decisions With Algorithms

489

Adapting, fast and slow: Causal Approach to Few-Shot Sequence Learning

490

Conformal Arbitrage for LLM Objective Balancing

491

Simulation-Based Inference for Adaptive Experiments

492

Agents as Tool-Use Decision-Makers

493

Quantitative Judges for Large Language Models

494

Self-Challenging Language Model Agents

495

Learning to Explore: An In-Context Learning Approach for Pure Exploration

496

How Bidirectionality Helps Language Models Learn Better via Dynamic Bottleneck Estimation

497

A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models

498

Simplifying Bayesian Optimization Via In-Context Direct Optimum Sampling

499

Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models

500

IPO: Interpretable Prompt Optimization for Vision-Language Models

501

Evolutionary Prompt Optimization discovers emergent multimodal reasoning strategies

502

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?

503

Diffusion Guidance Is a Controllable Policy Improvement Operator

504

Alita: Generalist Agent With Self-Evolution

505

A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement Learning

506

Learning Compositional Functions with Transformers from Easy-to-Hard Data

507

Preference Learning with Response Time

508

Accelerating RL for LLM Reasoning with Optimal Advantage Regression

509

Algorithms for reliable decision-making need causal reasoning

510

Belief Attribution as Mental Explanation: The Role of Accuracy, Informativity, and Causality

511

Distances for Markov chains from sample streams

512

When and Why LLMs Fail to Reason Globally

513

IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis

514

No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered Inference

515

Accelerating RL for LLM Reasoning with Optimal Advantage Regression

516

Statistical Inference for Online Algorithms

517

Prismatic Synthesis for Diverse LLM Reasoning Data

518

Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents

519

The Agentic Economy

520

Statistics for Large Language Models

521

Efficient Bayes-Adaptive Reinforcement Learning using Sample-Based Search

522

Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning

523

Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL

524

Value-Guided Search for Efficient Chain-of-Thought Reasoning

525

Shallow Preference Signals: Large Language model aligns even better without truncated data?

526

Gaming Tool Preferences in Agentic LLMs

527

Partner Modelling Emerges in Recurrent Agents (But Only When It Matters)

528

LLM Populations Form Social Conventions and Collective Bias

529

LLM Generated Persona is a Promise with a Catch

530

Large Language Models for Digital Twin Simulation

531

From RL Distillation to Autonomous LLM Agents

532

Prompting, Auto-Prompting, and Human-AI Communication

533

Textual Gradients for LLM Optimization

534

Large Language Models as Markov Chains

535

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation

536

Selective induction heads: how transformers select causal structures in context

537

The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains

538

How Transformers Learn Causal Structure with Gradient Descent

539

Planning anything with rigor: general-purpose zero-shot planning with llm-based formalized programming

540

Automated Design of Agentic Systems

541

What’s the Magic Word? A Control Theory of LLM Prompting

542

BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling

543

RL with KL penalties is better viewed as Bayesian inference

544

Asymptotics of Language Model Alignment

545

Qwen 2.5, RL, and Random Rewards

546

Theoretical guarantees on the best-of-n alignment policy

547

Score Matching Enables Causal Discovery of Nonlinear Additive Noise Models

548

Improved Techniques for Training Score-Based Generative Models

549

Your Pre-trained LLM is Secretly an Unsupervised Confidence Calibrator

550

AlphaEvolve: A coding agent for scientific and algorithmic discovery

551

Harnessing the Universal Geometry of Embeddings

552

Goal Inference using Reward-Producing Programs in a Novel Physics Environment

553

Trial-Error-Explain In-Context Learning for Personalized Text Generation

554

Reinforcement Learning for Reasoning in Large Language Models with One Training Example

555

Test-Time Reinforcement Learning (TTRL)

556

Interpreting Emergent Planning in Model-Free Reinforcement Learning

557

Agentic Reward Modeling_Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems

558

Beyond Reward Hacking: Causal Rewards for Large LanguageModel Alignment

559

Learning How Hard to Think: Input-Adaptive Allocation of LM Computation

560

Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval

561

UFT: Unifying Supervised and Reinforcement Fine-Tuning

562

Understanding High-Dimensional Bayesian Optimization

563

Inference time alignment in continuous space

564

Efficient Test-Time Scaling via Self-Calibration

565

Conformal Prediction via Bayesian Quadrature

566

Predicting from Strings: Language Model Embeddings for Bayesian Optimization

567

Self-Evolving Curriculum for LLM Reasoning

568

Online Decision-Focused Learning in Dynamic Environments

569

FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain

570

Reward Shaping from Confounded Offline Data

571

Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning

572

Understanding Best-of-N Language Model Alignment

573

Maximizing Acquisition Functions for Bayesian Optimization - and its relation to Gradient Descent

574

Bayesian Prompt Ensembles: Model Uncertainty Estimation for Black-Box Large Language Models

575

Prompting Strategies for Enabling Large Language Models to Infer Causation from Correlation

576

The Parallel Knowledge Gradient Method for Batch Bayesian Optimization

577

FunBO: Discovering Acquisition Functions for Bayesian Optimization with FunSearch

578

Automated Social Science: A Structural Causal Model-Based Approach

579

Causal Interpretation of Transformer Self-Attention

580

A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment

581

Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs

582

Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation

583

Prompts from Reinforcement Learning (PRL)

584

Logits are All We Need to Adapt Closed Models

585

Large Language Models Are (Bayesian) Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning

586

Inference-Time Intervention: Eliciting Truthful Answers from a Language Model

587

From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models

588

LLM In-Context Learning as Kernel Regression

589

Personalizing LLMs via Decode-Time Human Preference Optimization

590

Almost Surely Safe LLM Inference-Time Alignment

591

Survey of In-Context Learning Interpretation and Analysis

592

From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models

593

LLM In-Context Learning as Kernel Regression

594

Where does In-context Learning Happen in Large Language Models?

595

Auto-Differentiating Any LLM Workflow: A Farewell to Manual Prompting

596

metaTextGrad: Learning to learn with language models as optimizers

597

Semantic Operators: A Declarative Model for Rich, AI-based Data Processing

598

Isolated Causal Effects of Language

599

Sleep-time Compute: Beyond Inference Scaling at Test-time

600

J1: Incentivizing Thinking in LLM-as-a-Judge

601

ShiQ: Bringing back Bellman to LLMs

602

Policy Learning with a Natural Language Action Space: A Causal Approach

603

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models

604

End-to-End Learning for Stochastic Optimization: A Bayesian Perspective

605

TEXTGRAD: Automatic Differentiation via Text

606

Steering off Course: Reliability Challenges in Steering Language Models

607

Past-Token Prediction for Long-Context Robot Policies

608

Recovering Coherent Event Probabilities from LLM Embeddings

609

Systematic Meta-Abilities Alignment in Large Reasoning Models

610

Predictability Shapes Adaptation: An Evolutionary Perspective on Modes of Learning in Transformers

611

Efficient Exploration for LLMs

612

Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation

613

Bayesian Concept Bottlenecks with LLM Priors

614

Transformers for In-Context Reinforcement Learning

615

Evaluating Large Language Models Across the Lifecycle

616

Active Ranking from Human Feedback with DopeWolfe

617

Optimal Designs for Preference Elicitation

618

Dual Active Learning for Reinforcement Learning from Human Feedback

619

Active Learning for Direct Preference Optimization

620

Active Preference Optimization for RLHF

621

Test-Time Alignment of Diffusion Models without reward over-optimization

622

Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback

623

GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

624

Advantage-Weighted Regression: Simple and Scalable Off-Policy RL

625

Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective

626

Transformers can be used for in-context linear regression in the presence of endogeneity

627

Bayesian Concept Bottlenecks with LLM Priors

628

In-Context Parametric Inference: Point or Distribution Estimators?

629

Enough Coin Flips Can Make LLMs Act Bayesian

630

Bayesian Scaling Laws for In-Context Learning

631

Posterior Mean Matching Generative Modeling

632

Can Generative AI Solve Your In-Context Learning Problem? A Martingale Perspective

633

Dynamic Search for Inference-Time Alignment in Diffusion Models

634

Is In-Context Learning in Large Language Models Bayesian? A Martingale Perspective

635

Leaked Claude Sonnet 3.7 System Instruction tuning

636

Converging Predictions with Shared Information

637

Test-Time Alignment Via Hypothesis Reweighting

638

Rethinking Diverse Human Preference Learning through Principal Component Analysis

639

Active Statistical Inference

640

Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework

641

AI-Powered Bayesian Inference

642

Can Unconfident LLM Annotations Be Used for Confident Conclusions?

643

Predictions as Surrogates: Revisiting Surrogate Outcomes in the Age of AI

644

Learn then Test: Calibrating Predictive Algorithms to Achieve Risk Control

645

How to Evaluate Reward Models for RLHF

646

LLMs as Judges: Survey of Evaluation Methods

647

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

648

Limits to scalable evaluation at the frontier: LLM as Judge won’t beat twice the data

649

Stratified Prediction-Powered Inference for Hybrid Language Model Evaluation

650

Accelerating Unbiased LLM Evaluation via Synthetic Feedback

651

Prediction-Powered Statistical Inference Framework

652

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

653

RM-R1: Reward Modeling as Reasoning

654

Reexamining the Aleatoric and Epistemic Uncertainty Dichotomy

655

Decoding Claude Code: Terminal Agent for Developers

656

Emergent Strategic AI Equilibrium from Pre-trained Reasoning

657

Benefiting from Proprietary Data with Siloed Training

658

Advantage Alignment Algorithms

659

Asymptotic Safety Guarantees Based On Scalable Oversight

660

What Makes a Reward Model a Good Teacher? An Optimization Perspective

661

Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems

662

Identifiable Steering via Sparse Autoencoding of Multi-Concept Shifts

663

You Are What You Eat - AI Alignment Requires Understanding How Data Shapes Structure and Generalisation

664

Interplay of LLMs in Information Retrieval Evaluation

665

Trade-Offs Between Tasks Induced by Capacity Constraints Bound the Scope of Intelligence

666

Toward Efficient Exploration by Large Language Model Agents

667

Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT

668

Self-Consuming Generative Models with Curated Data

669

Bootstrapping Language Models with DPO Implicit Rewards

670

DeepSeek-Prover-V2: Advancing Formal Reasoning

671

THINKPRM: Data-Efficient Process Reward Models

672

Societal Frameworks and LLM Alignment

673

Risks from Multi-Agent Advanced AI

674

Causality-Aware Alignment for Large Language Model Debiasing

675

Reward Models Evaluate Consistency, Not Causality

676

Causal Rewards for Large Language Model Alignment

677

Sycophancy to subterfuge: Investigating reward-tampering in large language models

678

Bidirectional AI Alignment

679

Why Do Multi-Agent LLM Systems Fail?

680

LLMs as Greedy Agents: RL Fine-tuning for Decision-Making

681

LLM Feedback Loops and the Lock-in Hypothesis

682

Representational Alignment Drives Effective Teaching and Learning

683

Adaptive Parallel Reasoning with Language Models

684

AI: Rewiring the Flow of Ideas and Human Knowledge

685

Learning and Equilibrium with Ranking Feedback

686

Designing Human-AI Collaboration: A Sufficient-Statistic Approach

687

GOAT: Generative Adversarial Training for Human-AI Coordination

688

π0.5: Generalization in Robotic Manipulation via Diverse Data

689

NoWag: Unified Compression for Large Language Models

690

Optimal Tool Calls in Language Model Reasoning

691

Data Selection for Empirical Risk Minimization

692

LoRe: Low-Rank Reward Modeling for Personalized LLMs

693

ParaPO: Reducing Language Model Verbatim Reproduction

694

Test-Time RL: Self-Evolving LLMs via Majority Voting Rewards

695

Tina: Tiny LoRA Reasoning Models

696

Evaluating large language models in theory of mind tasks

697

QUEST: Quality Sampling for Machine Translation

698

Offline Preference Learning via Simulated Trajectory Feedback

699

Reasoning Elicitation in Language Models via Counterfactual Feedback

700

Eliciting Human Preferences with Language Models

701

Sub-Optimal Data for Human-in-the-Loop Reinforcement Learning

702

γ-Bench: Evaluating LLMs in Multi-Agent Games

703

DRAFT: Self-Driven LLM Tool Mastery via Documentation Refinement

704

Optimal Prediction Sets for Enhanced Human-AI Accuracy

705

Self-Correction via Reinforcement Learning for Language Models

706

Tractable Multi-Agent Reinforcement Learning through Behavioral Economics

707

Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement

708

Iterative Nash Policy Optimization for Language Model Alignment

709

SycEval: Benchmarking LLM Sycophancy in Mathematics and Medicine

710

Stack AI: Democratizing Enterprise AI Development

711

Evaluating Modern Recommender Systems: Challenges and Future Directions

712

AI in the Enterprise: Seven Lessons from Frontier Companies by OpenAI

713

Discussion: Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

714

AI Agent Protocols and Human Preference

715

Cross-Environment Cooperation for Zero-Shot Multi-Agent Coordination

716

Sutton and Silver: The Era of Experience: Learning Beyond Human Data

717

Sample, Don't Search: Rethinking Test-Time Alignment for Language Models

718

AI Agents: Echoes of Past Technology Pivots?

719

Minimalist LLM Reasoning: Rejection Sampling to Reinforcement

720

Securing the Model Context Protocol in Enterprise Environments

721

Improving Multi-Turn Tool Use with Reinforcement Learning

722

Cultural Knowledge Conservation and Control in Large Language Models

723

Data Quality, Repetition, and Scaling of Language Models

724

Compute-Optimal Scaling Laws for Language Models Revisited

725

Concise Reasoning via Reinforcement Learning

726

Throughput Limits for LLM Inference and AI Agent Scheduling

727

RL Post-training Amplifies Pretraining Behaviors in Language Models

728

Fast Adaptation of Behavioral Foundation Models

729

Proprietary Reward Models: Sustaining Advantage in Agentic AI

730

Why Multi-Agent LLM Systems Fail: A Comprehensive Study

731

Play2Prompt: Zero-Shot Tool Instruction Optimization via Tool Play

732

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems

733

API and GUI Agents: Divergence, Convergence, and Hybrid Approaches

734

AI, Chess, and Competitive Advantage: Substitution and Complementation

735

Knowledge of the Firm and Replication of Technology

736

Firm Resources and Sustained Competitive Advantage

737

Evaluating Pharmaceutical Marketing to Physicians with Panel Data

738

Theory of the firm in the era of Agents

739

Large Language Models: An Applied Econometric Framework

740

Evaluating the World Model Implicit in a Generative Model

741

Machine Learning for Hypothesis Generation in Social Science

742

Active Learning for Moral Preference Elicitation: Challenges and Nuances

743

Gradient-Based Surveys for Nonparametric Discrete Choice Experiments

744

Explainable Data-driven Share-of-choice Product Line Design Optimization

745

The More You Ask, the Less You Get: When Additional Questions Hurt External Validity

746

Conjoint topics from Handbook of Marketing Analytics: Methods and Applications

747

Choice-Based Conjoint Analysis: Methods and Applications

748

Beyond Conjoint Analysis: The Future of Preference Measurement

749

An Optimization Framework for Adaptive Questionnaire Design

750

Adaptive Self-Explication of Multiattribute Preferences

751

Conjoint Analysis: Methods, Applications, and Recent Developments

752

Current Issues and a “Wish List” for Conjoint Analysis

753

Ellipsoidal Methods for Adaptive Choice-Based Conjoint Analysis

754

Adaptive Polyhedral Methods for Conjoint Analysis

755

MSL: Enhancing LLM Recommenders via Masked Softmax Loss

756

Self-Supervised Deep Reinforcement Learning for Optimal Question Ranking

757

Adaptive Language Elicitation for Latent Information Discovery

758

LLM Persona Bias: Promise and Peril in Simulation

759

AutoTools: Automating Tool Use for Large Language Models

760

Tool Learning with Large Language Models: A Comprehensive Survey

761

All Roads Lead to Likelihood: RL for Fine-Tuning Value

762

ATLAS: Tuning Agents via Critical Step Learning

763

Thinking Faster by Writing Less: Chain of Draft Reasoning

764

Meta Plan Optimization for Boosting LLM Agents

765

L1: Length Controlled Reasoning with Reinforcement Learning

766

WikiBigEdit: Benchmarking Lifelong Knowledge Editing in LLMs

767

PLAN-AND-ACT: LLM Agent Planning with Synthetic Data

768

SEARCH-R1: LLMs Learn to Reason and Search via Reinforcement Learning

769

The Theory of the Firm: Information, Incentives, and Organization

770

Four Formalizable Theories of the Firm

771

Efficient Tool Use with Chain-of-Abstraction Reasoning

772

CodeTool: Process Supervision for Enhanced LLM Tool Invocation

773

Evaluating LLM Agents in Multi-Turn Conversations: A Survey

774

Epistemic Alignment in User-LLM Knowledge Delivery

775

MCP is (not) all you need

776

AI, Human Skills, and Competitive Advantage in Chess

777

Inference-Time Scaling for Generalist Reward Modeling

778

Optimal Pure Exploration in Linear Bandits via Sampling

779

Presidential Address: The Economist as Designer in the Innovation Process for Socially Impactful Digital Products

780

Emergent Symbolic Mechanisms for Reasoning in Large Language Models

781

Inference-Time Alignment: Coverage, Scaling, and Optimality

782

Sharpe Ratio-Guided Active Learning for Preference Optimization

783

Active Learning for Adaptive In-Context Prompt Design

784

Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

785

On the Biology of a Large Language Model

786

Async-TB: Asynchronous Trajectory Balance for Scalable LLM RL

787

Instacart's Economics Team: A Hybrid Role in Tech

788

Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework

789

Why MCP won

790

SWEET-RL: Training LLM Agents for Collaborative Reasoning

791

TheoryCoder: Bilevel Planning with Synthesized World Models

792

Driving Forces in AI: Scaling to 2025 and Beyond (Jason Wei, OpenAI)

793

Expert Demonstrations for Sequential Decision Making under Heterogeneity

794

TextGrad: Backpropagating Language Model Feedback for Generative AI Optimization

795

MemReasoner: Generalizing Language Models on Reasoning-in-a-Haystack Tasks

796

RAFT: In-Domain Retrieval-Augmented Fine-Tuning for Language Models

797

Inductive Biases for Exchangeable Sequence Modeling

798

InverseRLignment: LLM Alignment via Inverse Reinforcement Learning

799

Prompt-OIRL: Offline Inverse RL for Query-Dependent Prompting

800

Alignment from Demonstrations for Large Language Models

801

Q♯: Distributional RL for Optimal LLM Post-Training

802

Scaling Test-Time Compute Without Verification or RL is Suboptimal

803

Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

804

Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

805

Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

806

Revisiting Superficial Alignment Hypothesis

807

Diagnostic uncertainty: teaching language Models to describe open-ended uncertainty

808

Language Model Personalization via Reward Factorization

809

How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach

810

Can Large Language Models Extract Customer Needs as well as Professional Analysts?

811

Spurlens: finding spurious correlations in Multimodal llms

812

Improving test-time search with backtrack- Ing Improving test-time search with backtrack- Ing against in-context value verifiersagainst in-context value verifiers

813

Adaptive elicitation of latent information Using natural language

814

Document Valuation in LLM Summaries: A Cluster Shapley Approach

815

s1: simple test time scaling