Model is the Product | Common Corpus, Mid-Training, Open Science | Pierre-Carl Langlais, Pleias episode artwork

EPISODE · May 16, 2026 · 2H 4M

Model is the Product | Common Corpus, Mid-Training, Open Science | Pierre-Carl Langlais, Pleias

from GroundZero AI Talks · host Himanshu Dubey

Pierre-Carl Langlais (aka Alexandar Doria) is Co-founder of Pleias. We'd discussed about pre-training recipes, common corpus, mid-training, agentic systems, good post-training and everything AI.TIMESTAMPS00:00:00 - TEASER00:01:12 - INTRO00:02:03 - Who is Alexander Doria [Pierre-Carl Langlais]? 00:04:10 - Early career: From humanities to AI research00:07:50 - Meeting influential people in computational humanities00:10:00 - How the idea of Pleias came about00:13:30 - Building Pleias: Infrastructure and compute challenges in Europe00:17:06 - Team structure and work culture at Pleias00:19:06 - What is "open science" and why it matters00:21:53 - Big announcement: OpenSynthetic initiative00:25:25 - Synthetic data experiments and surprising results00:28:11 - "The Model is the Product" - explained00:31:56 - Implications for companies building on top of models00:35:25 - Differentiation in a world of shared base models00:38:40 - Common Corpus: Origins and development00:44:12 - The lack of open, legally clear datasets00:47:03 - Anthropic's use of Common Corpus for mechanistic interpretability00:50:20 - What makes good post-training?00:54:00 - Reasoning under 400M parameters in SLMs00:56:35 - Generalist scaling is stalling - where are the diminishing returns?00:59:40 - Will specialization always win over scale?01:02:00 - Opinionated and task-specialized models01:06:29 - How inference cost drops change monetization models01:09:12 - New value layers beyond token marketplaces01:11:38 - Major technical obstacles to embedding workflows in models01:13:40 - How smaller labs can compete on training infrastructure01:15:36 - Should startups raise capital for AI training?01:17:16 - What new capabilities do models need for orchestration?01:19:50 - Designing verifier functions for agentic models01:22:17 - RL in domains with weak or delayed rewards01:24:50 - Multi-step training loops: Draft, verify, refine, backtrack01:26:38 - The scarcity of agentic data and bootstrapping solutions01:29:32 - Making agent training tractable at scale01:31:44 - What is mid-training and why it matters01:34:55 - Deployment, use cases, and hybrid model architectures01:37:37 - Human-in-the-loop for regulated domains01:39:48 - Advice for startups positioning in this transition01:41:58 - Europe's structural challenges in AI01:45:52 - Tokenizers: The overlooked competitive frontier01:49:59 - Training LLMs on personal data and dead languages01:52:12 - World models and JEPA architectures01:53:50 - Building agentic systems: Stack and RL environments01:55:34 - The art of training good RL models01:58:49 - Trivia: Underrated habits and mindsets in research02:00:09 - AI Twitter community and its impact02:01:40 - Advice for folks starting in AI research02:03:27 - Final thoughts and wrap-up

Episode metadata supplied by the publisher feed · Published May 16, 2026

Embed this episode

NOW PLAYING

Model is the Product | Common Corpus, Mid-Training, Open Science | Pierre-Carl Langlais, Pleias

0:00 2:04:40

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of GroundZero AI Talks?

This episode is 2 hours and 4 minutes long.

When was this GroundZero AI Talks episode published?

This episode was published on May 16, 2026.

Can I download this GroundZero AI Talks episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!