Home / Founders / Directory / Diogo Almeida
Diogo Almeida
Co-founder and CEO · TypeSafe AI
Calibrated models OpenAI alum
Almeida is training AI models whose answers go to software, each with a confidence score that tells the program when to act on its own and when to hand the case to a person.
Diogo Almeida is co-founder and chief executive of TypeSafe AI, a lab in San Francisco, California, whose first model, Jev, returns answers that other software can act on, each with a measure of confidence. He spent four years at OpenAI improving ChatGPT’s responses and is a co-author of the InstructGPT paper, which showed how reinforcement learning from human feedback (RLHF) trains a language model to follow instructions. Earlier he worked at Google Brain. He left OpenAI in 2024 and started TypeSafe with Erik Gafni and Sasha Sheng.
A critique of RLHF from the inside
Almeida argues that the training method he helped build made models too sure of themselves. “We’ve been optimizing for humans and we’re super human at pleasing humans,” he told Forbes. In a CEO Insider interview he said RLHF training makes a model drop the minority cases to concentrate on what gets rewarded. “The model becomes hyperconfident,” he said, which is a problem for any program that has to decide when to trust an answer.
He assumed another lab must already be working on the idea and found that none was. “I thought the whole project would take a week. I was unbelievably wrong,” he said. TypeSafe spent two years in stealth. He has said that if an AI winter came and he had not done everything he could to prevent it, he would have considered himself personally responsible.
Models whose customer is code
TypeSafe calls Jev a system one model. Its name refers to the Jevons paradox, and it is tuned for intelligence per dollar. Almeida says an API should never refuse a request the way a chatbot does. “In an API, a refusal is a type error,” he said, because a dependency that refuses at random breaks the software built on top of it.
He gives insurance underwriting as an example. Asked whether a property has a history of fires, Jev reads the evidence and returns a probability that a program can set a threshold on. “An LLM doesn’t have the nuance to communicate that,” he told Forbes.
