Jev: ChatGPT Co-Creator’s Answer to RLHF
4 months before TypeSafe AI launched Jev, founder Diogo Almeida stood up at AI Council and laid out the problem Jev was built to solve. He never names the product in this talk. He says only that TypeSafe is going all in on automation and getting ready to ship in a couple of months, then spends 36 minutes on why he thinks the rest of the field is pointed the wrong way. TypeSafe came out of stealth on September 15, 2026 with $40M led by DCVC and launched Jev, a model that outputs typed decisions with probabilities for software to act on. Almeida spent four and a half years at OpenAI, where he co-authored the InstructGPT paper and the GPT-4 Technical Report and worked on the RLHF research behind ChatGPT. His argument here starts from that work. His read is that roughly all production LLMs are trained with RLHF or some variation of it, and RLHF optimizes for human preference. Human preference is a different target from doing the task correctly. That gap is why AI looks superhuman on assistance work and falls apart on automation work that looks easier. He walks through mode collapse, Yann LeCun's doomed-LLM argument, why a 1.3B-parameter InstructGPT model was preferred to 175B GPT-3 at following instructions, and where he thinks Sutton's bitter lesson is incomplete. He closes on the question that started TypeSafe, which is what it would take to treat language models as a primitive that software calls, the way it calls a database or an API. Recorded live at AI Council 2026 in San Francisco, May 12-14. SPEAKER Diogo Almeida, Co-founder and CEO, TypeSafe AI, InstructGPT and GPT-4 co-author X: https://x.com/CompleteSkeptic LinkedIn: https://www.linkedin.com/in/diogomda/ TypeSafe AI: https://typesafe.ai/ Jev: https://typesafe.ai/blog/introducing-system-one-models-and-jev CHAPTERS 0:00 AI is too good to be true and too bad to be useful 0:51 Four and a half years at OpenAI 2:11 Where is the economic revolution? 5:23 Assistance work versus automation work 6:09 But aren't coding agents automation? 9:43 When is a result too good to be true? 11:18 Sutton's bitter lesson, and the bitterest lesson 13:49 What RLHF is and what it costs you 14:44 Yann LeCun's doomed-LLM argument 15:49 Mode collapse, RLHF's deal with the devil 18:18 Is software engineering easier than drive-thrus? 19:49 Why a model can't be trusted with stakes 20:42 Is ChatGPT cooked? 25:33 Type-safe language models, going beyond strings 27:13 The FLAN lesson, when the whole field is wrong 29:40 Can one model do assistance and automation? 31:08 What TypeSafe is building 33:42 Summary 35:00 Was RLHF AGI, and what was missing [2026 - DAY 1 - INFERENCE SYSTEMS] Sign up for our "No BS" Newsletter to get the latest technical data & AI content: https://aicouncil.com/newsletter ABOUT AI COUNCIL: AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools. FIND US: Website: https://aicouncil.com/ LinkedIn: https://www.linkedin.com/company/aicouncilconf/ X: https://x.com/aicouncilconf
Read Video · Транскрипция и инсайты
У этого выпуска есть полная расшифровка + AI-анализ
Бесплатный аккаунт · без карты · 150 кредитов при регистрации, достаточно для этого эпизода
- 📄 Полная транскрипция с временными метками
- ✨ AI-резюме, ключевые слова и ментальная карта
- 💡 Ключевые тезисы и цитаты
- Говорящий 1
Таймлайн эпизода
Speaker introduces himself as a former OpenAI researcher and frames the talk's central paradox of AI being simultaneously overhyped and underdelivering.
- The speaker spent four and a half years at OpenAI and co-authored GPT-4, InstructGPT, ChatGPT, and RLHF, giving him insider credibility to critique the field.
- He argues there is a massive divide between LLM overpromise and LLM underdeliver, and that benchmarks like GPQA saturating at human level did not translate into real-world danger or utility.
- He notes that even true believers have shifted goalposts from automating economically valuable work to merely generating $100 billion in revenue or profit.
The core thesis: today's LLMs are optimized for human-in-the-loop assistance tasks, not machine-in-the-loop automation.
- Tasks where AI looks superhuman (chat, coding assistance) are all human-in-the-loop assistance, while tasks where it fails (drive-throughs, customer service actions) are automation tasks.
- Customer service chatbots are deliberately not allowed to take actions because models are not reliable enough, throwing users into a labyrinth of documentation instead.
- Coding agents are also assistance rather than automation because code is arguably a language for humans, not machines, and the workflow still involves humans in the loop.
The speaker argues that as precision requirements rise, LLM utility falls, which is the opposite of how traditional machine learning behaves.
- Traditional ML is highly structured and integrates with automation, so its utility skyrockets as precision needs increase, unlike LLMs which degrade.
- The field exhibits a Stockholm syndrome where people accept that LLMs are unpredictable like humans rather than reliable like software.
- Almost all of today's LLMs are optimized for the perception of utility rather than true utility, which explains the overpromise/underdeliver divide.
Ключевые понятия
- RLHF— Reinforcement Learning from Human Feedback, the algorithm behind ChatGPT that the speaker argues is the root cause of AI's overpromise/underdeliver divide.
- assistance vs automation— The core distinction the speaker uses to explain why AI excels at some tasks and fails at others.
- mode collapse— RLHF's tendency to drop minority modes, producing plausibly correct but unreliable outputs.
Знаковые цитаты
I believe that the answer to this is simply that all of the stuff in the left were human in the loop assistance tasks and the things on the right were like machine in the loop automation tasks.
💡— This is the central thesis that explains the entire overpromise/underdeliver divide in AI, reframing the problem as one of task optimization rather than model capability.
I don't believe drive-throughs are automated yet. Someone correct me if I'm wrong, but like there's lots of economic incentive to do so.
🤯— A surprising fact that a simple, economically incentivized automation task remains unsolved, highlighting AI's limitations in real-world automation.
Действенные выводы
🧠AI Strategy
Most AI products today are optimized for assistance, not automation, which limits their economic impact.
Audit your AI use cases this week and label each as assistance or automation; prioritize automation tasks that have clear, measurable outcomes.
RLHF models are optimized for human preference, making them unreliable for high-stakes decisions.
Identify one decision in your workflow that currently involves AI and add a human verification step before any irreversible action.
🛠️Product Development
LLMs are resistant to layering; building on top of them without controlling the optimization is risky.
If you're building an AI product, map out where the optimization happens and ensure you have control or visibility into it.
Type safety for language models could enable reliable automation.
Explore integrating type-safe interfaces or structured outputs into your AI prototypes to reduce errors.
Транскрипция и инсайты создаются автоматически и могут содержать ошибки. Точность зависит от качества звука и чёткости речи дикторов — если что-то выглядит неверно, исходная запись всегда остаётся главным источником.
Эпизоды и видео готовы к чтению
Эпизоды подкастов

Ep #82 Polar Explorers
Case by Case
9 мая 2024 г.38:29EN
(Preview) Doom Debates Go Mainstream, AI Religion and the Economic Future, Several Vectors of the China Question
Sharp Tech with Ben Thompson
18 сент. 2026 г.33:19EN
How video games can level up the way you learn | Kris Alexander
TED Talks Daily
7 сент. 2026 г.14:25EN
Радио-Т 1026
Радио-Т
15 авг. 2026 г.RU
绿点小样 | 3M公司:能赚钱的可持续项目才可持续
商业就是这样
25 мая 2025 г.11:20ZH
#23 「人類好きなの?」
朝井リョウ・加藤千恵 信頼できない語り手
10 июл. 2026 г.56:14JA
Видео

Steve Jobs' 2005 Stanford Commencement Address (with intro by President John Hennessy)
Stanford
14 мая 2008 г.22:10EN
Tesla, Figure and the Fight Over Humanoid Robot Hands | Scott Walter, RoboStrategy
Core Matter
18 сент. 2026 г.1:23:21EN
Finally, some truth about loop engineering
NeetCode
28 июл. 2026 г.17:18EN
«Современный урок по ФГОС: требования, этапы, цифровые решения»
ЯКласс
11 апр. 2023 г.1:39:24RU
#32 Die Schlacht von Worringen - Der Freiheitskampf der Kölner
99 mal Geschichte
12 нояб. 2025 г.53:29DE