Models

TypeSafe AI's Jev Model Offers a Different Path to Machine Intelligence

Diogo Almeida, an OpenAI researcher who helped develop RLHF, has launched TypeSafe AI with a transformer model that abandons language generation entirely in favor of producing calibrated probabilities—and developers are responding with enthusiasm.

·5 min read
A new kind of AI model from a ChatGPT inventor is thrilling developers
A new kind of AI model from a ChatGPT inventor is thrilling developers

Diogo Almeida spent years at OpenAI building ChatGPT and pioneering reinforcement learning from human feedback (RLHF), the training method widely credited with enabling today's AI era. Yet he found himself frustrated by what the technology could actually accomplish in practice.

We have lightning in a bottle, and yet it is not useful. I've been battling that problem since then. It took me a while to come to the conclusion: The problem is we are optimizing for human language … We have been super good at human language for four years, but it's not useful for automation because computers speak a different language.

Diogo Almeida

That dissatisfaction led Almeida to leave OpenAI two years ago and establish TypeSafe AI. This week, the company unveiled Jev, a transformer-based model that represents a fundamental departure from the large language model paradigm. Rather than generating text, Jev outputs probabilities—what TypeSafe calls "calibrated decisions."

The shift away from language-based outputs creates several practical advantages. The model operates at dramatically lower cost and speed compared to traditional LLMs. Because outputs are predefined by users, the system cannot produce hallucinations. Output tokens carry no cost, while input tokens are priced per billion rather than per million.

Demand for the model has been so strong that TypeSafe temporarily lost the ability to serve API requests. Developers are finding Jev particularly valuable for software automation tasks, where it functions as a cheaper and more dependable way to embed intelligence into applications.

Real-world performance gains

Pranit Sharma, a software engineer at Vercel, tested Jev against OpenAI's ChatGPT Luna 5.6 for a safety classifier that reviews commands. When Vercel switched from Luna to Jev, the system delivered results five to 18 times faster while improving accuracy.

Nikhil Mudholkar, CTO at Bryo AI, compared Jev to Google's Gemini for classifying business emails. While Gemini proved slightly more accurate, it cost 10 to 20 times more. What impressed Mudholkar most was Jev's confidence scoring. He noted that "it is the only one that hands back a real probability which makes it ideal for automating workflows!!" according to his testing.

Beyond replacement: augmentation and monitoring

Beyond simply replacing LLMs in specific applications, Jev can serve as a safeguard against model misbehavior. Using one AI system to monitor another quickly becomes expensive, but Almeida argues that deploying Jev for this purpose makes economic sense. He envisions users leveraging Jev to examine LLM agent traces and block jailbreak attempts.

Armin Ronacher, CTO of Earendil (which maintains the open source model harness Pi), explained the practical implications: "At the end of the day, it delegates the hallucination problem a little bit to the user. The user has to say, okay, if this only comes back with 50% probability, maybe this is a coin toss, and I disregard it. But if it's 95%, sure, then I can do something with it."

Ronacher also identified model routing as another promising application. Determining whether a particular workload needs a specific model would normally require an expensive LLM call. Jev's combination of low cost and speed makes real-time routing feasible.

The philosophy behind the name

Almeida named the model after William Stanley Jevons, a 19th-century economist known for the Jevons paradox—the observation that falling commodity costs often lead to increased consumption. The parallel is intentional: as the cost of intelligence drops, deployment should become ubiquitous.

We think that there's just going to be smart software all over the place in a way that's emergent and distributed … much more like the early internet than you know like the mega apps that people are trying to build right now.

Diogo Almeida

Training approach and architecture

Almeida has remained guarded about Jev's internal architecture, though outside observers suspect it builds on an open-weight LLM foundation. TypeSafe describes Jev as a "System One model," emphasizing intuition over reasoning and tailored to specific tasks. The training process relies exclusively on synthetic data using what Almeida calls "reinforcement learning from calibrated decisions."

We made an early bet that we will be making all of our data, and that has been one of the best bets I've ever made in my life — better than our launch, in my opinion, better than RLHF.

Diogo Almeida

Almeida elaborated on this commitment: "Half of [our company] is a lab that basically owns this entire subfield of statistically well-understood synthetic data, and that is now my life joy."

Competition on the horizon

Currently, Jev operates without direct competitors in its category. However, Ronacher expects that once the model's utility becomes clear, other companies will follow. He suggested that competitors have been slow to emerge partly because LLMs remain inexpensive and heavily subsidized. "We should have seen this earlier in many ways, but presumably because the LLMs are so cheap and subsidized, you often don't have to be creative yet," he said.

TypeSafe plans to develop additional versions of Jev across new modalities. When asked whether TypeSafe functions as a frontier lab, Almeida pushed back against that framing. "The main product of frontier labs is fear or hype. I would like our main product to be intelligence…[but we are] not a lab in the sense of, you know, like bet on infinite wealth, or a religion, or building God in a data center, or whatever is the thing of today."