Artificial Intelligence Expert

TypeSafe AI Jev Review: How Good Is the New Model for AI Classification?

Reading time
30 ​​min

TL;DR:

Jev, TypeSafe AI’s first model, answers yes/no, multiple-choice, and scale questions about a text using probabilities rather than text. LLMs have long been capable of classification without their own training; Jev makes it significantly cheaper and faster, but also harder to understand. In our own tests, we show where this is sufficient and where it isn’t: Related questions can contradict each other, the numbers fluctuate, and the results tend to reflect what a state-of-the-art LLM would answer rather than how often the answer is correct. We see the greatest benefit in individual classifications within very large data streams.

This article was automatically translated. Original article.

On September 15, 2026, TypeSafe AI unveiled its first public model: Jev, designed for classification tasks (TypeSafe AI, 2026a). Instead of responding with text, it provides numerical answers and, for multiple-choice questions, also selects the chosen option from a predefined list. Typical questions include “Is this spam?”, “Which team is responsible?”, or “How urgent is this?”. TypeSafe refers to the numbers as probabilities and calls models of this type “System One,” after the fast, intuitive thinking that psychologist Daniel Kahneman described as System 1 (Kahneman, 2011).1

Classification itself is nothing new. Classical models for this have existed for decades, but they must be trained specifically for each task. Since the advent of GPT-3, at the very latest, LLMs can also be used as classifiers without requiring separate training: they are given a fixed list of answers to evaluate, and probabilities are read from their outputs as needed (Brown et al., 2020). Technically, then, Jev does not enable a new type of application. It reduces costs and latency and offers a simpler, more direct interface. This could open the door to applications that were previously not well-served due to data volume, costs, or latency.

In this post, we share the results of our initial experiments and our impressions of Jev. In the first part, we describe what Jev does differently, which features already existed, and what we observed in our own tests. We then outline where we see potential and limitations based on our assessment. In doing so, we view Jev not only as a product but as a representative of a new class of models that could become established.

The most important limitations up front:

  • Answers to related questions may contradict each other; Jev is primarily suited for individual classifications.
  • The same query yields slightly different numbers across multiple runs.
  • You must verify for yourself whether the probabilities are accurate for your own data.
  • Jev is a general-purpose classifier; you shouldn’t expect too much from it for very specific categories.
  • Jev does not provide any justification.
  • The Jev product is available only as a service, hosted in the U.S.; open-source clones can be run on your own.

What can Jev do?

In Jev, a query consists of the state—that is, the subject of the decision (a customer message, a document, a data record, either as text or as an object with named fields)—and one or more questions related to it. TypeSafe describes this as a “frontier-intelligence function call: unstructured state in, typed probabilistic decisions out” (TypeSafe AI, 2026a). Each question is one of three types (TypeSafe AI, 2026b):

Question Type Purpose Returns Example
Noul (from Bernoulli’s yes-or-no random experiment) Yes-No questions a number between 0 and 1: the probability of “yes” “Is this spam?”
Choice select an option from a list the selected option and a probability for each option “Which of the following teams is responsible?”
Score classify on a scale with labeled levels a value on the scale and a probability for each level “How urgent is this?”

The following example asks one question of each type for a support ticket, all three in a single request. If Jev is less than 80% certain about the team, a human reviews the ticket.

(Note: The code snippets below use the original German prompts and input texts to ensure exact reproducibility of our experiment).

Ticket Jev’s Choice Score Result
“I was charged the same amount twice in March …” billing 1.00 automatically to billing
“Ever since the update, the app crashes immediately when I log in.” technology 1.00 automatically to the technology department
“I’d like to upgrade my subscription to Premium …” sales 0.99 automatically to the sales department
“Since the app update, the billing page sometimes displays an incorrect amount …” (code example) billing 0.68 manual check

Jev assigns unique tickets a value of 0.99 or 1.00. The ticket from the code example is intentionally ambiguous: An incorrect amount after an app update could be caused by an app error or a billing error. With a billing score of 0.68, it is routed to a human.

Such a request is inexpensive and fast: According to TypeSafe, one million tokens of input cost $0.042, while the output costs nothing. A request can include up to 64.000 tokens and takes 70 to 500 ms; in our tests, it took around 300 ms (TypeSafe AI, 2026c). Jev processes only text and prefers English.

What capabilities existed before Jev?

None of these capabilities is new in and of itself.

Before LLMs

For a long time, a separate classification model—such as BERT (Devlin et al., 2019)—was trained for each question using examples for which the correct answer was known. Such a model runs in milliseconds on dedicated hardware and always provides the same answer, but it only answers the question for which it was trained. The first methods that did not require dedicated training already existed back then: One asks a model whether a text supports a statement such as “This text is about an invoice,” and read the probability of that being true (Yin et al., 2019). This corresponds to Jev’s yes-or-no question.

With LLMs

Zero-shot and few-shot. An LLM can solve a classification problem based solely on the instruction, drawing on its general knowledge—just like Jev. If additional examples of how cases should be classified are included in the prompt, it adapts to the specific domain (Brown et al., 2020).

Structured Output

A refinement specifies which values the response may take. In constrained decoding, all tokens that do not match a permitted response are blocked at every position: If only “billing,” “technology,” or “sales” are allowed, the model cannot write anything else—with a preceding justification if desired. The major providers offer this via their APIs (OpenAI, 2024; Anthropic, 2026; Google, 2026a); for self-hosted models, options include the Outlines library (Willard & Louf, 2023; dottxt, 2026).2 If you additionally read the logprobs of the allowed tokens—that is, the logarithm of their probability—you obtain a value for each label. This is always possible with self-hosted models; the current Gemini models no longer provide logprobs via their API (checked on September 28, 2026).

What makes Jev different?

LLMs are generalists in two respects: in terms of tasks and knowledge. Jev specializes in the task at hand—it only makes judgments. However, when it comes to knowledge and context, Jev remains a generalist.

According to TypeSafe, this specialization stems from the training: Jev is not a modified LLM, but rather a standalone model trained specifically for decision-making; the method is called Reinforcement Learning for Calibrated Decisions (RLCD). Neither the method nor the model’s architecture has been published (TypeSafe AI, 2026a). Related to this is the published method RLCR, which fully rewards a model only if its stated confidence matches its hit rate (Damani et al., 2025).

Above all, Jev turns a well-known method into a product: As with an LLM, an instruction defines the task—there is no need for task-specific training. This can lower the barrier to adoption. Compared to an LLM, Jev is cheaper and faster. Gemini 3.5 Flash-Lite, Google’s smallest model and our benchmark, costs $0.30 per million input tokens—a good seven times as much as Jev—plus $2.50 per million output tokens, including reasoning tokens (Google, 2026c). In our tests, Jev responded after about 0.3 seconds, while Flash-Lite took 2 to 7 seconds—partly because it generates reasoning tokens by default (Google, 2026b).

The trade-off is transparency: Jev does not provide any reasoning. An LLM with structured output can do this, as demonstrated by the same ticket using Gemini (Google, 2026a):

Gemini provides a valid label, a clear explanation, and, in this case, a self-assessment as well. The ticket classifies it as “possible” and states the reason right away: both a billing error and a technical issue. The application therefore passes it on—like Jev—to a human when the value falls below the threshold. This isn’t a probability, but rather a signal of uncertainty that stems directly from the meaning of the response. The burden falls on the user: They must define the levels of self-assessment and verify whether the model applies them reliably.3

Approach How the task is defined Probabilities? Explanations? Hosting
Custom classifier (such as BERT) Training with examples, per task yes no self-operated, milliseconds
Zero-Shot Text Inference (Yin et al., 2019) Message when called yes, not calibrated no self-operated, cheap
LLM with Structured Output via API Instructions and examples in the prompt as a reported category or number, often too high (Xiong et al., 2024) yes as a service, with reasoning: seconds
Self-operated LLM with Constrained Decoding Instructions and examples in the prompt Yes, it can be calibrated later yes self-hosted; affordable for small models
Jev Instructions in the question, rules in the state Yes, calibrated according to the manufacturer’s specifications no Available only as a service in the U.S., about 0.3 s
Open models with a Jev interface (Kev, SemIf) like Jev yes no self-operated

Open alternatives

Jev’s interface didn’t remain exclusive for long: Just a few days after its release, several open-source versions with the same interface became available that users can run themselves:

Model Foundation Availability
Kev (Palmer, 2026) Qwen3.5 in three sizes, retrained free (Apache 2.0), compatible with the TypeSafe SDK
SemIf, formerly OpenJev (TheoLeeCJ, 2026) Qwen 3.5 with 4 billion parameters free, can be operated locally
djev (Maisa, 2026) Diffusion model based on Gemma free (Apache 2.0), self-hosted

We tested Jev

We submitted several hundred queries to Jev and Gemini 3.5 Flash-Lite. The cases are small and, in some instances, contrived.

Related fields may contradict each other

Jev evaluates each question separately; only the state is shared (TypeSafe AI, 2026d). TypeSafe sees this as a strength: According to TypeSafe, the most reliable workflows consist of many independent, disaggregated questions (TypeSafe, 2026a). Things get tricky as soon as answers need to match up. Here’s an example: A painting disappeared from a museum overnight, and exactly one of three scenarios is correct.

Scenario Probability Perpetrator Route
A 30% Night Watchman Loading dock
B 25% Night Watchman Staff exit
C 45% Burglars from outside Roof

Both models were asked the same two questions: “How did the painting leave the building?” and “Was it an inside job?” Jev answers them independently of one another. Gemini answers both in a single request: Using reasoning, it weighs the questions together before responding, and even without reasoning, it takes the first answer into account when filling in the second field. In addition, we asked both models directly about the scenario.4

How to Interpret Jev’s Probabilities

The museum case raises a second question: What does Jev’s 0.96 for the roof mean when the text specifies 45%? A probability is calibrated if, of all answers with a value of 0.8, approximately 80% are correct, always relative to a specific set of cases (Hájek, 2007). TypeSafe does not specify what Jev’s 0.8 refers to. However, it uses the answers from Frontier LLMs—specifically, the mean of GPT-6 Astra and Claude Fable 5.1—as a benchmark (TypeSafe AI, 2026a). Jev’s figure thus most accurately indicates how likely a Frontier LLM would be to give this answer, not how often it is correct.

An initial independent, non-peer-reviewed test found Jev to be well-calibrated on known benchmarks but less so on 900 newly generated support tickets: too confident on multiple-choice questions and too cautious on yes/no questions (scienthoon, 2026). Our museum case aligns with the observation regarding multiple-choice questions: As a multiple-choice question, the roof received a score of 0.96 to 0.98 instead of the 0.45 mentioned in the text.

The numbers also fluctuate. In 20 identical runs, the invoice ticket scored between 0.67 and 0.76; when the order of the options was reversed, the score was about 0.1 higher or lower; and in English, it scored only between 0.5 and 0.6. A threshold of 0.70 would therefore sometimes approve identical requests and sometimes refer them to a human.

Potential

The extent of Jev’s potential depends on what it is compared to. TypeSafe evaluates Jev against Frontier models (TypeSafe AI, 2026a). However, many classification tasks do not require a Frontier model; smaller models such as Claude Haiku or Gemini Flash-Lite are sufficient for these tasks. In our experience with enterprise client projects, these are not a significant cost factor; it is the Frontier models that carry the weight. Jev’s cost advantage therefore really counts where a Frontier model has previously been required for classification, and this area is likely quite specialized. Only a benchmark on your own task will show whether Jev can keep up there.

Compared to small models, the wait time is more of a factor: with 0.3 seconds instead of 2 to 7 seconds, checks can be integrated into running processes that would otherwise take too long.

We see the greatest benefit in screening large volumes of data when looking for a needle in a haystack: Screening with high recall drastically narrows down a long list of candidates, and only the few remaining are reviewed by an LLM or a human. In large infrastructures, millions of log lines and events are generated daily; whether a sequence of entries is suspicious often cannot be determined using fixed rules. For retrieval, a Jev-like model can—depending on cost-effectiveness—evaluate a query against several hundred to a few thousand candidates individually, either as a semantic search or during re-ranking. With these volumes, even small LLMs become expensive and slow. Organizations with very large data streams, such as government agencies or energy providers, are likely to benefit the most.

With models like Jev, the building blocks of an AI architecture no longer differ only in size but also in model type. It’s conceivable to have systems in which each model type takes on the task it’s best suited for. A Jev-like model sets the many small parameters: which tool, which context, which model. If a classification needs to be reviewed later, an LLM takes over and writes a justification for it. In the end, a large LLM synthesizes the results into a report that humans can read.

Probability or Confidence

In addition to cost and wait time, TypeSafe primarily touts its probability metrics. In the use cases we’ve encountered in projects, however, something else usually matters more: a confidence value against which a threshold is set. This value does not need to be calibrated; it simply needs to rank certain cases higher than uncertain ones. The threshold itself is determined based on your own data anyway. Jev also provides such a value for selection and rating questions. However, it is calculated from the probabilities—for three options, as (3 × highest probability – 1) / 2 (TypeSafe AI, 2026e)—and therefore says nothing about how confident Jev itself is in the probabilities. True probabilities are only needed when performing calculations with them—for example, when combining multiple estimates using Bayes’ theorem. It remains to be seen whether a model whose structure is unknown and that can only be controlled to a limited extent via context is suitable for such applications.

LLMs also provide a confidence value, either explicitly or implicitly. Explicitly, the model states its own confidence as a word, category, or number; such estimates are often too high (Xiong et al., 2024). Implicitly, through logprobs: The output layer of an LLM calculates a probability for every possible next token, not just the selected one. With self-hosted models, you can always retrieve these values, but with APIs, this is no longer possible everywhere. Whether Jev’s numbers must be accurate as probabilities or are sufficient as confidence scores remains to be seen in practice. In the latter case, Jev is cheaper and faster than an LLM, but not fundamentally new.

Whether probability or confidence score: Only a test on a few hundred of your own cases with known answers will reveal what error rate a threshold entails for your own data: How many tickets end up in the wrong team at 0.8, and how many at 0.9? The same cases can also be used for subsequent calibration (Guo et al., 2017). The threshold should be chosen based on the cost of an error and the cost of human review; this is also recommended by TypeSafe (TypeSafe AI, 2026e). For multiple-choice questions, an option for “unidentifiable” should also be included. Otherwise, the model will distribute the probability among the available answers, even if the text provides no clues.

Limitations

Limitations of the Technology

TypeSafe itself identifies some weaknesses (TypeSafe AI, 2026f): Jev is unreliable at counting and interprets dates as text, which is why it has difficulty comparing dates. It struggles with double negatives and multi-step inferences; superfluous information in the state distracts it; and instructions hidden in the data can influence its responses. Arithmetic, date logic, and filtering out superfluous data must therefore be handled in the code.
These weaknesses likely share a common cause. An LLM with reasoning capabilities can generate more tokens when faced with a difficult question, thereby applying more computational effort. Jev does not generate reasoning tokens. This makes it fast and cost-effective, but it also deprives it of this workaround. To regain this capability without reverting to an LLM, such a model would have to internally adjust the computational effort to the difficulty of the question without generating text. It is not known whether Jev is capable of this.

The museum case illustrates a general problem: characteristics that are interdependent—or, statistically speaking, covary—are the rule rather than the exception in practice. Anyone using Jev or a similar model should therefore verify whether the categories in question are truly independent. If they are not, it is better to leave the task to an LLM.

A second limitation concerns the knowledge that Jev draws upon. Jev is a classifier that relies exclusively on general knowledge. It works as long as the task falls within a generally understandable domain. If the task is domain-specific—for example, involving categories that only make sense within one’s own company—one should not expect too much from a model like Jev. For LLMs, in-context learning has proven effective in such cases: You provide examples in the prompt of how cases from your own domain should be classified (Brown et al., 2020). We have not systematically tested whether Jev also uses examples or classification rules stored in the state. In the museum case, Jev did not adopt the ratios specified in the state. It remains to be seen whether TypeSafe will offer a way to adapt Jev to your own domain.

Added to this is the regulatory aspect: Recruitment should not be submitted to such a model for evaluation. Under the EU AI Act, personnel selection is considered a high-risk area (Regulation (EU) 2024/1689, Annex III), and a model without an explanation can hardly be justified in that context. In the case of the Jev product, there is the additional issue that its architecture is unknown. Under such conditions, a model like Jev should prepare decisions rather than be responsible for them: It provides input, but the responsibility for the decision lies elsewhere. This includes logging every response with the specific model version (not “jev-latest”) and the time stamp, especially since the numbers fluctuate. This is the only way to verify a decision later.

Limitations of the Product

Jev is currently a solution from the U.S.; there is no documented operation within the EU (TypeSafe AI, 2026c). As a result, the product is often ruled out for European projects involving personal data. However, if the technology proves worthwhile, it will likely be a matter of just a few months before it becomes available in Europe: Either TypeSafe will offer GDPR-compliant operations, or closed or open alternatives will emerge. 5 Open models with the same interface were available within just a few days; no one has yet independently verified whether they can keep up.

Conclusion

Jev makes classification more affordable, faster, and more accessible without the need for custom training, but it doesn’t reinvent the wheel. Its results likely reflect what a state-of-the-art LLM would respond rather than how often the answer is correct; nevertheless, they often suffice as a confidence metric for setting a threshold. Jev is a good fit when several conditions align:

  • individual questions that are independent of one another,
  • a generally understandable domain,
  • large volumes for which an LLM is too expensive or too slow and building a custom classifier is too resource-intensive—such as in screening,
  • preliminary decisions or non-critical decisions that do not need to be justified on a case-by-case basis.

If any of these are missing, a dedicated classifier or an LLM is usually the better choice. Even where Jev is appropriate, it should prepare decisions, not be responsible for them: Models of this type carry the same biases as LLMs (Kraft, 2021), but these biases can only be measured statistically; they cannot be identified in individual cases based on a rationale.
It remains to be seen whether Jev will establish itself as a distinct class of models. If so, AI architectures in the future will no longer consist solely of larger and smaller LLMs, but of various types of models.


1 Kahneman adopted the terms “System 1” and “System 2” from Keith Stanovich and Richard West (2000). Stanovich and Jonathan Evans now refer more precisely to “Type 1” and “Type 2” processes (Evans & Stanovich, 2013). ↩

2 This would make it possible to build cascades as early as 2023: A small, inexpensive model provides the initial response, and only uncertain cases are passed on to a larger one (Chen et al., 2023). With Outlines, you can incorporate this uncertainty into the label: GPT-3.5 chooses between labels such as “definitely positive” and “possibly positive,” and only the uncertain cases are sent to GPT-4 (Herreros, 2023). ↩

3 Self-assessments are often too optimistic (Xiong et al., 2024). ↩

4 Each of Jev’s individual answers is correct on its own: According to the model, an inside perpetrator is more likely than an outside perpetrator (55%), and the roof is the most likely route (45%). Together, however, they form a combination that does not occur in any of the real-world scenarios. If, instead, the models were asked directly about the overall scenario, both models selected the correct scenario C in all runs. Thus, presenting the permitted combinations as a single multiple-choice question can be helpful. However, this approach has its limitations: combining several simple questions into one complex one contradicts Jev’s basic premise. Furthermore, the number of possible combinations grows rapidly with each additional feature, becoming unmanageable. Above all, however, this approach does not solve the actual problem: the best answer to each individual question asked separately does not automatically result in the best overall decision (Dembczyński et al., 2012); or, to paraphrase Aristotle: What holds true in isolation need not hold true in combination (De Interpretatione 11).↩

5 Update September 30, 2026: The market reproduced this offering even faster than anticipated at the time of publication. To illustrate the rapid timeline:

  • September 15, 2026: TypeSafe AI releases Jev.
  • September 28, 2026: Initial publication of this blog post.
  • September 29, 2026: OpenAI announces its commercial “Decisions API” at DevDay (functionally replicating Jev’s service for constrained probabilistic decisions).
  • September 29, 2026: Ollama announces in a blog post official support for Jev-style decision models, including a native /api/systemone endpoint for local execution.

This rapid sequence confirms the central thesis of this section: while Jev popularized a compelling interface, the underlying pattern is easily reproducible across both commercial cloud platforms and open local frameworks. ↩

References

Did you like this post?

Your email address will not be published. Required fields are marked *

inoNews

5 good reasons to subscribe to the inovex newsletter:

  • Exclusive insights and tips from our inovexperts
  • Information and updates on IT trend topics and offers
  • Discounts on trainings and event invitations
  • Free whitepapers and infosheets
  • Options for exchange and consulting

To the newsletter registration