AI GUIDEPartner content

Jev: the model that solves tasks in 70 ms and doesn’t make up answers

TypeSafe has released Jev: the model doesn’t generate text but produces typed decisions with probabilities in 70-500 ms; input costs $42 per billion tokens, while output is free. What it can do and where it still doesn’t work.

Affiliate link: your price stays the same and the project earns a commission.

On 15 September 2026, TypeSafe AI released Jev, the first public model in the System One class. This is not a chatbot: the model does not generate text; it produces typed decisions with probabilities, doing so in 70-500 milliseconds. The company’s founder, Diogo Almeida, previously worked at OpenAI and contributed to the creation of methods that formed the basis of ChatGPT, while over the past two years the team developed a new architecture in stealth mode.

There are plenty of questionable figures in the claims, so let’s examine them in order.

How System One differs from a language model

A conventional LLM builds an answer sequentially, one token at a time: this is flexible, but expensive and slow when the model has to make hundreds of small decisions in succession within a program. Jev works differently. A single request produces all decisions in parallel, in the form of predefined structures with calibrated probabilities and confidence scores. The training method is called RLCD (Reinforcement Learning for Calibrated Decisions): not human preferences, but honest probabilities for tasks whose results can be checked programmatically.

Giving up free-form strings has two consequences. First, an error in the response type is impossible in principle, hence the developers’ claim that the model does not hallucinate in the usual sense. Second, every response comes with a probability. If the model solves a task correctly in 95% of cases, it honestly reports when it falls into the remaining 5. Without this, automating a process is impossible: a system that does not know when it makes mistakes is useful only under human supervision.

The model’s name is not accidental either. Jev is named after economist William Stanley Jevons, who proposed the idea that increased resource efficiency leads to increased consumption of that resource. The System One class is named after Kahneman’s fast, intuitive thinking. The name is a promise: every drop in the cost of intelligence opens up an order of magnitude more application scenarios.

Speed and money

Frontier models respond in 3-329 seconds. Jev fits within 70-500 milliseconds, making it 40-200 times faster at a comparable level of intelligence on structured tasks. The model charges $0,042 per million input tokens, or $42 per billion, and charges nothing for output at all: the developers call it too cheap to count.

ParameterFrontier LLMJev
Response time3-329 sec70-500 ms
Input cost$0,20-10 per million tokens$0,042 per million tokens
Output costapproximately 5 times more expensive than inputfree
Confidenceavailable when explicitly requested, but unstablecalibrated for every output
Formattext that needs to be parsedready-made typed values

On TypeSafe’s homepage, the figures are more aggressive: 193,6 times faster and 444,6 times cheaper. This is the result of the company’s own workflow tests, and the company itself notes that such gains are considered an upper bound for real-world results; the benchmark is the average of two strong models. For the intelligence comparison, they used GPT-5.6 Terra, acknowledging it as the most comparable to Jev.

What was built during the first day

The first projects appeared within 24-48 hours of the release, and they all share the same design: a game or simulator sends a simplified state of the world, the model chooses one of the permitted actions, the state is updated, and the cycle repeats.

  • In Minecraft, an agent named Linknberi controls a character: the model receives the time of day, health, nearby enemies, and a list of available actions. Two minutes of operation, using approximately 150 000 tokens, cost one cent. At night, when zombies appear, the agent leaves on its own; it does not need a separate “run away” command.
  • The car simulator was not built as an autonomous-driving system. There is a car, a road, other road users, and four actions: accelerate, brake, maintain speed, and turn. The cycle is so fast that the car drives in real time, and the entire prototype was built without collecting driving datasets.
  • Subway Surfers tests reaction speed: the character must jump, duck, or turn as soon as an obstacle appears. A slow response is useless here, making such games an honest test for the System 1 approach.
  • The developer assembled a drone obstacle course in about 15 minutes and estimated the entire experiment at 10 cents. The drone maintains its position, speed, and distances from obstacles, choosing from five actions: forward, turn, up, down, and hover.
  • TypeSafe itself has a Doom demo on its blog, where ten requests per second cost approximately $7 per hour, as well as wiki racing: moving between Wikipedia articles via links, with a choice of hundreds of options at every step.

The list is less impressive than the general principle: the same model, not fine-tuned for a specific scenario, tracks a changing state and reacts to it. Specialized game bots existed long before Jev, but they were trained separately for each game. Here, there is no need to change the rules, and that is the whole point of the new category.

Where it still doesn’t work

All the demonstrations are built on simulations. A real car would require cameras, sensors, control software, and several layers of safety; a robot or drone would require stabilization systems and emergency limits. Until TypeSafe’s claims are confirmed across a broad range of tasks, it is too early to draw conclusions about operation in real cars and actual drones.

There is also a limitation inherent to the model itself: Jev cannot generate text, so it cannot serve as a conversational partner or writing assistant. Its domain is classification, routing, evaluation, field extraction, branching within code, scoring large volumes of data, and checking other people’s answers. When a decision needs to be made hundreds or thousands of times in succession, request cost and speed become part of the product architecture, and here they matter more than increasing the model’s size.

What to do in practice

  • Request early access from TypeSafe and run your own workflows: the service is available in early access, and the queue is being cleared gradually.
  • Open evals.typesafe.ai: it contains the workflow tests themselves, with example requests, explanations of discrepancies, and the full text of the tasks.
  • Calculate the economics before integration. If the model answers thousands of structured questions per day, the difference between expensive input and free output already offsets the connection cost in the first week.
  • Test the promise of no hallucinations on your own data: the types are mathematically guaranteed, but the calibration of probabilities must be confirmed on your task distribution.
  • Monitor broad validation. The first day showed that speed is in demand, but the real test is weeks of operation on other people’s systems, not one day of enthusiastic experimentation.

Compare models before you start

The service sets its plans, limits and model catalog. If they differ from this article, contact us so we can update it and record a new review date.

Browse models

Affiliate link: your price stays the same and the project earns a commission.

Jev model System One TypeSafe AI fast AI models structured AI decisions RLCD AI pricing decision automation

SEO Mind42 editorial team

We explore SEO and neural networks in practice: test services on our own projects, verify prices and limits against primary sources, and share things you can put to use the same day.

📚 Reference guide to SEO and AI 🔄 Materials are updated 🕐 Updated: 4 October 2026

Related reading

All in this section →