EPISODE · Jan 8, 2026 · 12 MIN
A Unified Definition of Hallucination, Or: It’s the World Model, Stupid
from Best AI papers explained · host Enoch H. Kang
Researchers from institutions like Carnegie Mellon and Stanford propose a unified definition of hallucination in large language models by framing it as a failure of internal world modeling. Traditionally, the term referred to scattered issues like translation errors or unverified summaries, but this framework suggests that all hallucinations are simply mismatches between model outputs and a reference truth. By defining a Reference World Model, researchers can specify exactly what counts as "true" across diverse domains like chess, web navigation, or medical QA. This structured approach helps distinguish between genuine world-modeling errors and other failures like poor planning or instruction following. Ultimately, this perspective enables the creation of larger-scale, synthetic benchmarks that programmatically test a model's ability to maintain factual consistency within complex, evolving environments.
Embed this episode
NOW PLAYING
A Unified Definition of Hallucination, Or: It’s the World Model, Stupid
No transcript for this episode yet
Similar Episodes
No similar episodes found.