Jev: ChatGPT Co-Creator’s Answer to RLHF
4 months before TypeSafe AI launched Jev, founder Diogo Almeida stood up at AI Council and laid out the problem Jev was built to solve. He never names the product in this talk. He says only that TypeSafe is going all in on automation and getting ready to ship in a couple of months, then spends 36 minutes on why he thinks the rest of the field is pointed the wrong way. TypeSafe came out of stealth on September 15, 2026 with $40M led by DCVC and launched Jev, a model that outputs typed decisions with probabilities for software to act on. Almeida spent four and a half years at OpenAI, where he co-authored the InstructGPT paper and the GPT-4 Technical Report and worked on the RLHF research behind ChatGPT. His argument here starts from that work. His read is that roughly all production LLMs are trained with RLHF or some variation of it, and RLHF optimizes for human preference. Human preference is a different target from doing the task correctly. That gap is why AI looks superhuman on assistance work and falls apart on automation work that looks easier. He walks through mode collapse, Yann LeCun's doomed-LLM argument, why a 1.3B-parameter InstructGPT model was preferred to 175B GPT-3 at following instructions, and where he thinks Sutton's bitter lesson is incomplete. He closes on the question that started TypeSafe, which is what it would take to treat language models as a primitive that software calls, the way it calls a database or an API. Recorded live at AI Council 2026 in San Francisco, May 12-14. SPEAKER Diogo Almeida, Co-founder and CEO, TypeSafe AI, InstructGPT and GPT-4 co-author X: https://x.com/CompleteSkeptic LinkedIn: https://www.linkedin.com/in/diogomda/ TypeSafe AI: https://typesafe.ai/ Jev: https://typesafe.ai/blog/introducing-system-one-models-and-jev CHAPTERS 0:00 AI is too good to be true and too bad to be useful 0:51 Four and a half years at OpenAI 2:11 Where is the economic revolution? 5:23 Assistance work versus automation work 6:09 But aren't coding agents automation? 9:43 When is a result too good to be true? 11:18 Sutton's bitter lesson, and the bitterest lesson 13:49 What RLHF is and what it costs you 14:44 Yann LeCun's doomed-LLM argument 15:49 Mode collapse, RLHF's deal with the devil 18:18 Is software engineering easier than drive-thrus? 19:49 Why a model can't be trusted with stakes 20:42 Is ChatGPT cooked? 25:33 Type-safe language models, going beyond strings 27:13 The FLAN lesson, when the whole field is wrong 29:40 Can one model do assistance and automation? 31:08 What TypeSafe is building 33:42 Summary 35:00 Was RLHF AGI, and what was missing [2026 - DAY 1 - INFERENCE SYSTEMS] Sign up for our "No BS" Newsletter to get the latest technical data & AI content: https://aicouncil.com/newsletter ABOUT AI COUNCIL: AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools. FIND US: Website: https://aicouncil.com/ LinkedIn: https://www.linkedin.com/company/aicouncilconf/ X: https://x.com/aicouncilconf
Read Video · Transcripción e Insights
Este episodio tiene una transcripción completa + análisis IA
Cuenta gratuita · sin tarjeta · 150 créditos al registrarte, suficiente para este episodio
- 📄 Transcripción completa con marcas de tiempo
- ✨ Resumen IA, palabras clave y mapa mental
- 💡 Puntos clave y citas destacadas
- Interlocutor 1
Línea de tiempo del episodio
Speaker introduces himself as a former OpenAI researcher and frames the talk's central paradox of AI being simultaneously overhyped and underdelivering.
- The speaker spent four and a half years at OpenAI and co-authored GPT-4, InstructGPT, ChatGPT, and RLHF, giving him insider credibility to critique the field.
- He argues there is a massive divide between LLM overpromise and LLM underdeliver, and that benchmarks like GPQA saturating at human level did not translate into real-world danger or utility.
- He notes that even true believers have shifted goalposts from automating economically valuable work to merely generating $100 billion in revenue or profit.
The core thesis: today's LLMs are optimized for human-in-the-loop assistance tasks, not machine-in-the-loop automation.
- Tasks where AI looks superhuman (chat, coding assistance) are all human-in-the-loop assistance, while tasks where it fails (drive-throughs, customer service actions) are automation tasks.
- Customer service chatbots are deliberately not allowed to take actions because models are not reliable enough, throwing users into a labyrinth of documentation instead.
- Coding agents are also assistance rather than automation because code is arguably a language for humans, not machines, and the workflow still involves humans in the loop.
The speaker argues that as precision requirements rise, LLM utility falls, which is the opposite of how traditional machine learning behaves.
- Traditional ML is highly structured and integrates with automation, so its utility skyrockets as precision needs increase, unlike LLMs which degrade.
- The field exhibits a Stockholm syndrome where people accept that LLMs are unpredictable like humans rather than reliable like software.
- Almost all of today's LLMs are optimized for the perception of utility rather than true utility, which explains the overpromise/underdeliver divide.
Conceptos clave
- RLHF— Reinforcement Learning from Human Feedback, the algorithm behind ChatGPT that the speaker argues is the root cause of AI's overpromise/underdeliver divide.
- assistance vs automation— The core distinction the speaker uses to explain why AI excels at some tasks and fails at others.
- mode collapse— RLHF's tendency to drop minority modes, producing plausibly correct but unreliable outputs.
Citas destacadas
I believe that the answer to this is simply that all of the stuff in the left were human in the loop assistance tasks and the things on the right were like machine in the loop automation tasks.
💡— This is the central thesis that explains the entire overpromise/underdeliver divide in AI, reframing the problem as one of task optimization rather than model capability.
I don't believe drive-throughs are automated yet. Someone correct me if I'm wrong, but like there's lots of economic incentive to do so.
🤯— A surprising fact that a simple, economically incentivized automation task remains unsolved, highlighting AI's limitations in real-world automation.
Acciones a tomar
🧠AI Strategy
Most AI products today are optimized for assistance, not automation, which limits their economic impact.
Audit your AI use cases this week and label each as assistance or automation; prioritize automation tasks that have clear, measurable outcomes.
RLHF models are optimized for human preference, making them unreliable for high-stakes decisions.
Identify one decision in your workflow that currently involves AI and add a human verification step before any irreversible action.
🛠️Product Development
LLMs are resistant to layering; building on top of them without controlling the optimization is risky.
If you're building an AI product, map out where the optimization happens and ensure you have control or visibility into it.
Type safety for language models could enable reliable automation.
Explore integrating type-safe interfaces or structured outputs into your AI prototypes to reduce errors.
La transcripción y los insights son generados por IA y pueden contener errores. La precisión depende de la calidad del audio y la claridad de los hablantes — si algo no parece correcto, el audio original es siempre la fuente de verdad.
Episodios y videos listos para leer
Episodios de podcast

50. Self-Checkout
The Economics of Everyday Things
22 jun 202618:48EN
The gift and power of emotional courage | Susan David (re-release)
TED Talks Daily
29 jul 202619:12EN
Architecting AI for the Power Grid with Hitachi Vantara's Chief Strategy Officer Simon Ninan
AI with Arun Show
10 jul 202619:14EN
Conflicto de roles: familia, trabajo y ocio | Experiencias compartidas | 419
Entiende Tu Mente
18 feb 202623:35ES
商业小样17 | 哪些政府持有比特币?
商业就是这样
12 ene 202510:59ZH
Paula Wehmeyer - VP Marketing und Talent bei General Intuition & Ex General Catalyst Partner (#40)
Mannheim Investment Club (MIC) Podcast
18 sept 202655:19DE
Videos

Why Graph Engineering will 10x your Claude/Codex
Greg Isenberg
3 ago 202626:29EN
Uber Sells the Dream, Waymo Logs the Autonomous Miles
The Road to Autonomy
28 feb 202653:54EN
Sam Altman on Building OpenAI & Betting on the Impossible
David Senra
23 ago 20261:18:17EN
#40 Karl IV. und sein goldenes Prag
99 mal Geschichte
8 ene 20261:03:51DE
Firma bez šéfov: Funguje to? - Money Talk 113 s Ferom Baníkom
Milan Dubec
4 ago 202654:10SK