EPISODE · Jun 22, 2026
Fine-Tuning LLMs for Human Behavior Prediction
from AI Post Transformers
This episode explores a 2025 study on fine-tuning large language models to predict how people respond in social science experiments, asking whether trained models can simulate new studies more reliably than prompting alone. It explains how the researchers built SOCSCI210, a dataset of 2.9 million responses from more than 400,000 participants across 210 TESS experiments, and why standardizing those studies into respondent-condition-question-answer records is central to the method. The discussion breaks down the paper’s evaluation criteria, including out-of-distribution generalization, distribution matching via Wasserstein distance, normalized individual accuracy, and treatment-effect recovery, to show the difference between sounding plausible and preserving real experimental patterns. Listeners would find it interesting because it treats LLMs not as chatbots but as possible “wind tunnels” for testing study designs in advance, while also confronting the risk that a convincing simulator could still get causal effects wrong. Sources: 1. Finetuning LLMs for Human Behavior Prediction in Social Science Experiments — Akaash Kolluri, Shengguang Wu, Joon Sung Park, Michael S. Bernstein, 2025 http://arxiv.org/abs/2509.05830 2. Out of One, Many: Using Language Models to Simulate Human Samples — Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, David Wingate, 2022 https://scholar.google.com/scholar?q=Out+of+One%2C+Many%3A+Using+Language+Models+to+Simulate+Human+Samples 3. LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals — Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Michael S. Bernstein, et al., 2024 https://scholar.google.com/scholar?q=LLM+Agents+Grounded+in+Self-Reports+Enable+General-Purpose+Simulation+of+Individuals 4. Large Language Models Show Human-like Social Desirability Biases in Survey Responses — Aadesh Salecha, Molly E. Ireland, Shashanka Subrahmanya, Joao Sedoc, Lyle H. Ungar, Johannes C. Eichstaedt, 2024 https://scholar.google.com/scholar?q=Large+Language+Models+Show+Human-like+Social+Desirability+Biases+in+Survey+Responses 5. Finetuning LLMs for Human Behavior Prediction in Social Science Experiments — Akaash Kolluri, Shengguang Wu, Joon Sung Park, Michael S. Bernstein, 2025 https://scholar.google.com/scholar?q=Finetuning+LLMs+for+Human+Behavior+Prediction+in+Social+Science+Experiments 6. Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies — Gati Aher, Rosa I. Arriaga, Adam Tauman Kalai, 2022 https://scholar.google.com/scholar?q=Using+Large+Language+Models+to+Simulate+Multiple+Humans+and+Replicate+Human+Subject+Studies 7. Can Large Language Models Replace Human Subjects? A Large-Scale Replication of Scenario-Based Experiments in Psychology and Management — Ziyan Cui, Ning Li, Huaikang Zhou, 2024 https://scholar.google.com/scholar?q=Can+Large+Language+Models+Replace+Human+Subjects%3F+A+Large-Scale+Replication+of+Scenario-Based+Experiments+in+Psychology+and+Management 8. Using Large Language Models to Create AI Personas for Replication, Generalization and Prediction of Media Effects: An Empirical Test of 133 Published Experimental Research Findings — Leo Yeykelis, Kaavya Pichai, James J. Cummings, Byron Reeves, 2024 https://scholar.google.com/scholar?q=Using+Large+Language+Models+to+Create+AI+Personas+for+Replication%2C+Generalization+and+Prediction+of+Media+Effects%3A+An+Empirical+Test+of+133+Published+Experimental+Research+Findings 9. This human study did not involve human subjects: Validating LLM simulations as behavioral evidence — Jessica Hullman, David Broska, Huaman Sun, Aaron Shaw, 2026 https://scholar.google.com/scholar?q=This+human+study+did+not+involve+human+subjects%3A+Validating+LLM+simulations+as+behavioral+evidence 10. Centaur: a Foundation Model of Human Cognition — Marcel Binz et al., 2024 https://scholar.google.com/scholar?q=Centaur%3A+a+Foundation+Model+of+Human+Cognition 11. Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions — Joseph Suh, Erfan Jahanparast, Suhong Moon, Minwoo Kang, Serina Chang, 2025 https://scholar.google.com/scholar?q=Language+Model+Fine-Tuning+on+Scaled+Survey+Data+for+Predicting+Distributions+of+Public+Opinions 12. Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions — Matthias Orlikowski, Jiaxin Pei, Paul Rottger, Philipp Cimiano, David Jurgens, Dirk Hovy, 2025 https://scholar.google.com/scholar?q=Beyond+Demographics%3A+Fine-tuning+Large+Language+Models+to+Predict+Individuals%27+Subjective+Text+Perceptions 13. Large Language Models that Replace Human Participants Can Harmfully Misportray and Flatten Identity Groups — Angelina Wang, Jamie Morgenstern, John P. Dickerson, 2025 https://scholar.google.com/scholar?q=Large+Language+Models+that+Replace+Human+Participants+Can+Harmfully+Misportray+and+Flatten+Identity+Groups 14. Beyond Believability: Accurate Human Behavior Simulation with Fine-Tuned LLMs — Yuxuan Lu et al., 2025 https://scholar.google.com/scholar?q=Beyond+Believability%3A+Accurate+Human+Behavior+Simulation+with+Fine-Tuned+LLMs 15. The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models — Marlene Lutz et al., 2025 https://arxiv.org/abs/2507.16076 16. Prompt Fairness: Sub-group Disparities in LLMs — Meiyu Zhong, Noel Teku, Ravi Tandon, 2025 https://arxiv.org/abs/2511.19956 17. Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment — Bryan Chen Zhengyu Tan et al., 2026 https://arxiv.org/abs/2604.12851 18. Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM — Zizhao Hu, Mohammad Rostami, Jesse Thomason, 2026 https://arxiv.org/abs/2603.18507 19. Causality for Large Language Models — Anpeng Wu et al., 2024 https://arxiv.org/abs/2410.15319 20. AI Post Transformers: PaperBench: Can AI Replicate AI Research? — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-17-paperbench-can-ai-replicate-ai-research-862944.mp3 21. AI Post Transformers: When Many-Shot CoT Becomes Test-Time Learning — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-15-when-many-shot-cot-becomes-test-time-lea-c25bfe.mp3 Interactive Visualization: Fine-Tuning LLMs for Human Behavior Prediction
Embed this episode
NOW PLAYING
Fine-Tuning LLMs for Human Behavior Prediction
No transcript for this episode yet
Similar Episodes
No similar episodes found.