Jev: ChatGPT Co-Creator’s Answer to RLHF
4 months before TypeSafe AI launched Jev, founder Diogo Almeida stood up at AI Council and laid out the problem Jev was built to solve. He never names the product in this talk. He says only that TypeSafe is going all in on automation and getting ready to ship in a couple of months, then spends 36 minutes on why he thinks the rest of the field is pointed the wrong way. TypeSafe came out of stealth on September 15, 2026 with $40M led by DCVC and launched Jev, a model that outputs typed decisions with probabilities for software to act on. Almeida spent four and a half years at OpenAI, where he co-authored the InstructGPT paper and the GPT-4 Technical Report and worked on the RLHF research behind ChatGPT. His argument here starts from that work. His read is that roughly all production LLMs are trained with RLHF or some variation of it, and RLHF optimizes for human preference. Human preference is a different target from doing the task correctly. That gap is why AI looks superhuman on assistance work and falls apart on automation work that looks easier. He walks through mode collapse, Yann LeCun's doomed-LLM argument, why a 1.3B-parameter InstructGPT model was preferred to 175B GPT-3 at following instructions, and where he thinks Sutton's bitter lesson is incomplete. He closes on the question that started TypeSafe, which is what it would take to treat language models as a primitive that software calls, the way it calls a database or an API. Recorded live at AI Council 2026 in San Francisco, May 12-14. SPEAKER Diogo Almeida, Co-founder and CEO, TypeSafe AI, InstructGPT and GPT-4 co-author X: https://x.com/CompleteSkeptic LinkedIn: https://www.linkedin.com/in/diogomda/ TypeSafe AI: https://typesafe.ai/ Jev: https://typesafe.ai/blog/introducing-system-one-models-and-jev CHAPTERS 0:00 AI is too good to be true and too bad to be useful 0:51 Four and a half years at OpenAI 2:11 Where is the economic revolution? 5:23 Assistance work versus automation work 6:09 But aren't coding agents automation? 9:43 When is a result too good to be true? 11:18 Sutton's bitter lesson, and the bitterest lesson 13:49 What RLHF is and what it costs you 14:44 Yann LeCun's doomed-LLM argument 15:49 Mode collapse, RLHF's deal with the devil 18:18 Is software engineering easier than drive-thrus? 19:49 Why a model can't be trusted with stakes 20:42 Is ChatGPT cooked? 25:33 Type-safe language models, going beyond strings 27:13 The FLAN lesson, when the whole field is wrong 29:40 Can one model do assistance and automation? 31:08 What TypeSafe is building 33:42 Summary 35:00 Was RLHF AGI, and what was missing [2026 - DAY 1 - INFERENCE SYSTEMS] Sign up for our "No BS" Newsletter to get the latest technical data & AI content: https://aicouncil.com/newsletter ABOUT AI COUNCIL: AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools. FIND US: Website: https://aicouncil.com/ LinkedIn: https://www.linkedin.com/company/aicouncilconf/ X: https://x.com/aicouncilconf
Leggi video · Trascrizione e analisi
Questo episodio ha una trascrizione completa e analisi AI
Account gratuito · nessuna carta necessaria · 150 crediti all'iscrizione, sufficienti per sbloccare questo episodio
- 📄 Trascrizione completa con timestamp
- ✨ Riepilogo AI, parole chiave e mappa mentale
- 💡 Conclusioni e citazioni chiave
- Parlante 1
Cronologia dell'episodio
Speaker introduces himself as a former OpenAI researcher and frames the talk's central paradox of AI being simultaneously overhyped and underdelivering.
- The speaker spent four and a half years at OpenAI and co-authored GPT-4, InstructGPT, ChatGPT, and RLHF, giving him insider credibility to critique the field.
- He argues there is a massive divide between LLM overpromise and LLM underdeliver, and that benchmarks like GPQA saturating at human level did not translate into real-world danger or utility.
- He notes that even true believers have shifted goalposts from automating economically valuable work to merely generating $100 billion in revenue or profit.
The core thesis: today's LLMs are optimized for human-in-the-loop assistance tasks, not machine-in-the-loop automation.
- Tasks where AI looks superhuman (chat, coding assistance) are all human-in-the-loop assistance, while tasks where it fails (drive-throughs, customer service actions) are automation tasks.
- Customer service chatbots are deliberately not allowed to take actions because models are not reliable enough, throwing users into a labyrinth of documentation instead.
- Coding agents are also assistance rather than automation because code is arguably a language for humans, not machines, and the workflow still involves humans in the loop.
The speaker argues that as precision requirements rise, LLM utility falls, which is the opposite of how traditional machine learning behaves.
- Traditional ML is highly structured and integrates with automation, so its utility skyrockets as precision needs increase, unlike LLMs which degrade.
- The field exhibits a Stockholm syndrome where people accept that LLMs are unpredictable like humans rather than reliable like software.
- Almost all of today's LLMs are optimized for the perception of utility rather than true utility, which explains the overpromise/underdeliver divide.
Concetti chiave
- RLHF— Reinforcement Learning from Human Feedback, the algorithm behind ChatGPT that the speaker argues is the root cause of AI's overpromise/underdeliver divide.
- assistance vs automation— The core distinction the speaker uses to explain why AI excels at some tasks and fails at others.
- mode collapse— RLHF's tendency to drop minority modes, producing plausibly correct but unreliable outputs.
Citazioni rilevanti
I believe that the answer to this is simply that all of the stuff in the left were human in the loop assistance tasks and the things on the right were like machine in the loop automation tasks.
💡— This is the central thesis that explains the entire overpromise/underdeliver divide in AI, reframing the problem as one of task optimization rather than model capability.
I don't believe drive-throughs are automated yet. Someone correct me if I'm wrong, but like there's lots of economic incentive to do so.
🤯— A surprising fact that a simple, economically incentivized automation task remains unsolved, highlighting AI's limitations in real-world automation.
Conclusioni applicabili
🧠AI Strategy
Most AI products today are optimized for assistance, not automation, which limits their economic impact.
Audit your AI use cases this week and label each as assistance or automation; prioritize automation tasks that have clear, measurable outcomes.
RLHF models are optimized for human preference, making them unreliable for high-stakes decisions.
Identify one decision in your workflow that currently involves AI and add a human verification step before any irreversible action.
🛠️Product Development
LLMs are resistant to layering; building on top of them without controlling the optimization is risky.
If you're building an AI product, map out where the optimization happens and ensure you have control or visibility into it.
Type safety for language models could enable reliable automation.
Explore integrating type-safe interfaces or structured outputs into your AI prototypes to reduce errors.
Trascrizione e analisi sono generate dall'AI e possono contenere errori. L'accuratezza dipende dalla qualità dell'audio e dalla chiarezza dei parlanti: in caso di dubbi, l'audio originale resta la fonte attendibile.
Episodi e video pronti da leggere
Episodi podcast

He Couldn't Walk Away || How Jetha Devapura Built Sri Lanka's Biggest Crisis Line
The Giving Habit
15 lug 202655:53EN
The shape-shifting sounds of the accordion | Maria Telesheva
TED Talks Daily
27 ago 202613:22EN
Bad Maps and Good Intentions; Sophie Radice on the trials and tribulations of life beyond the comfort zone S5 E11
How to have Extraordinary Relationships
26 mag 202657:25EN
深度加分題:不是所有思考都有幫助,小心反芻思考的內耗迴圈
心情加分站
24 giu 202628:58ZH
#74 Deutschland: Die Idee einer Nation
Wer wir sind und warum das nicht klappte ...
8 set 20261:06:37DE
Café com Deus Pai | 21 de maio
Café Com Deus Pai | Podcast oficial
21 mag 20264:53PT
Video

Loop Engineering from First Principles — Kyle Mistele, HumanLayer
AI Engineer
25 lug 202617:40EN
Why the AI’s honeymoon is ending (and tech workers are feeling it) | Noam Segal
Lenny's Podcast
12 lug 20261:36:29EN
Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history
Dwarkesh Patel
4 giu 20244:32:07EN
Firma bez šéfov: Funguje to? - Money Talk 113 s Ferom Baníkom
Milan Dubec
4 ago 202654:10SK
2026/08/24(一) 輝達伺服器傳漲價15%:AI成本暴增,成本誰吸收?
財女珍妮
24 ago 202630:06ZH-Hant