{"id":70278,"date":"2026-09-29T10:49:27","date_gmt":"2026-09-29T08:49:27","guid":{"rendered":"https:\/\/www.inovex.de\/?p=70278"},"modified":"2026-10-02T13:05:28","modified_gmt":"2026-10-02T11:05:28","slug":"typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification","status":"publish","type":"post","link":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/","title":{"rendered":"TypeSafe AI Jev Review: How Good Is the New Model for AI Classification?"},"content":{"rendered":"<p><em>This article was automatically translated. <a href=\"https:\/\/www.inovex.de\/de\/blog\/typesafe-ai-jev-im-test-wie-gut-ist-das-neue-modell-fuer-ki-klassifikation\/\">Original article.<\/a><\/em><\/p>\n<p>On September 15, 2026, TypeSafe AI unveiled its first public model: Jev, designed for classification tasks (TypeSafe AI, 2026a). Instead of responding with text, it provides numerical answers and, for multiple-choice questions, also selects the chosen option from a predefined list. Typical questions include \u201cIs this spam?\u201d, \u201cWhich team is responsible?\u201d, or \u201cHow urgent is this?\u201d. TypeSafe refers to the numbers as probabilities and calls models of this type \u201cSystem One,\u201d after the fast, intuitive thinking that psychologist Daniel Kahneman described as System 1 (Kahneman, 2011).<sup id=\"fnref-1\"><a href=\"#fn-1\">1<\/a><\/sup><!--more--><\/p>\n<p>Classification itself is nothing new. Classical models for this have existed for decades, but they must be trained specifically for each task. Since the advent of GPT-3, at the very latest, LLMs can also be used as classifiers without requiring separate training: they are given a fixed list of answers to evaluate, and probabilities are read from their outputs as needed (Brown et al., 2020). <strong>Technically, then, Jev does not enable a new type of application. It reduces costs and latency and offers a simpler, more direct interface. This could open the door to applications that were previously not well-served due to data volume, costs, or latency.<\/strong><\/p>\n<p>In this post, we share the results of our initial experiments and our impressions of Jev. In the first part, we describe what Jev does differently, which features already existed, and what we observed in our own tests. We then outline where we see potential and limitations based on our assessment. In doing so, we view Jev not only as a product but as a representative of a new class of models that could become established.<\/p>\n<p><strong>The most important limitations up front:<\/strong><\/p>\n<ul>\n<li>Answers to related questions may contradict each other; Jev is primarily suited for individual classifications.<\/li>\n<li>The same query yields slightly different numbers across multiple runs.<\/li>\n<li>You must verify for yourself whether the probabilities are accurate for your own data.<\/li>\n<li>Jev is a general-purpose classifier; you shouldn\u2019t expect too much from it for very specific categories.<\/li>\n<li>Jev does not provide any justification.<\/li>\n<li>The Jev product is available only as a service, hosted in the U.S.; open-source clones can be run on your own.<\/li>\n<\/ul>\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_88 counter-hierarchy ez-toc-counter ez-toc-custom ez-toc-container-direction\">\n<div class=\"ez-toc-title-container\"><p class=\"ez-toc-title\" style=\"cursor:inherit\"><\/p>\n<\/div><nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#What-can-Jev-do\" >What can Jev do?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#What-capabilities-existed-before-Jev\" >What capabilities existed before Jev?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-3\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#What-makes-Jev-different\" >What makes Jev different?<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-4\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#Open-alternatives\" >Open alternatives<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-5\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#We-tested-Jev\" >We tested Jev<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-6\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#Related-fields-may-contradict-each-other\" >Related fields may contradict each other<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-7\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#How-to-Interpret-Jevs-Probabilities\" >How to Interpret Jev\u2019s Probabilities<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-8\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#Potential\" >Potential<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-9\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#Probability-or-Confidence\" >Probability or Confidence<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-10\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#Limitations\" >Limitations<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-11\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#Limitations-of-the-Technology\" >Limitations of the Technology<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-12\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#Limitations-of-the-Product\" >Limitations of the Product<\/a><\/li><\/ul><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-13\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#Conclusion\" >Conclusion<\/a><\/li><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-14\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#References\" >References<\/a><\/li><\/ul><\/nav><\/div>\n<h2><span class=\"ez-toc-section\" id=\"What-can-Jev-do\"><\/span>What can Jev do?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>In Jev, a query consists of the state\u2014that is, the subject of the decision (a customer message, a document, a data record, either as text or as an object with named fields)\u2014and one or more questions related to it. TypeSafe describes this as a \u201cfrontier-intelligence function call: unstructured state in, typed probabilistic decisions out\u201d (TypeSafe AI, 2026a). Each question is one of three types (TypeSafe AI, 2026b):<\/p>\n<table>\n<thead>\n<tr>\n<th>Question Type<\/th>\n<th>Purpose<\/th>\n<th>Returns<\/th>\n<th>Example<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Noul (from Bernoulli&#8217;s yes-or-no random experiment)<\/td>\n<td>Yes-No questions<\/td>\n<td>a number between 0 and 1: the probability of \u201cyes\u201d<\/td>\n<td>\u201cIs this spam?\u201d<\/td>\n<\/tr>\n<tr>\n<td>Choice<\/td>\n<td>select an option from a list<\/td>\n<td>the selected option and a probability for each option<\/td>\n<td>\u201cWhich of the following teams is responsible?\u201d<\/td>\n<\/tr>\n<tr>\n<td>Score<\/td>\n<td>classify on a scale with labeled levels<\/td>\n<td>a value on the scale and a probability for each level<\/td>\n<td>\u201cHow urgent is this?\u201d<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>The following example asks one question of each type for a support ticket, all three in a single request. If Jev is less than 80% certain about the team, a human reviews the ticket.<\/p>\n<p>(Note: The code snippets below use the original German prompts and input texts to ensure exact reproducibility of our experiment).<\/p>\n<pre style=\"background-color: #f6f8fa; padding: 12px; font-size: 13px; line-height: 1.35;\"><code>from typesafe_sdk import Choice, Noul, Score, TypeSafeClient\r\n\r\nclient = TypeSafeClient(model=\"jev-1.13.0\")\r\n\r\nticket = {\r\n    \"betreff\": \"Rechnung falsch?\", # \"subject\": \"Incorrect invoice?\"\r\n    \"text\": (\"Seit dem App-Update zeigt die Rechnungsseite manchmal \"\r\n             \"einen falschen Betrag an. Wurde mir zu viel berechnet?\"),\r\n<span class=\"hljs-comment\">           # \"Since the app update, the billing page sometimes shows an incorrect amount. Was I overcharged?\"<\/span> \r\n}\r\n\r\nantworten = client.system_one(\r\n    state={\"ticket\": ticket},\r\n    questions={\r\n        \"team\": Choice(\r\n            instructions=\"Welches Team soll dieses Ticket bearbeiten?\", <span class=\"hljs-comment\"># \"Which team should handle this ticket?\" # \"Which team should handle this ticket?\"<\/span>\r\n            criteria={\"abrechnung\": \"Zahlungen, Rechnungen\", \"technik\": \"Fehler, App\", \"vertrieb\": \"Neuk\u00e4ufe\"}, # billing: payments, invoices; tech: bugs, app; sales: new purchases\r\n        ),\r\n        \"dringlichkeit\": Score(\r\n            instructions=\"Wie dringend ist dieses Ticket?\", <span class=\"hljs-comment\"># \"How urgent is this ticket?\"<\/span>\r\n            criteria=[\"Nicht dringend\", \"Etwas dringend\", \"Sehr dringend\"], # [\"Not urgent\", \"Somewhat urgent\", \"Very urgent\"]\r\n        ),\r\n        \"spam\": Noul(instructions=\"Diese Nachricht ist Spam\"), # \"This message is spam\"\r\n    },\r\n).answers\r\n\r\nprint(antworten[\"team\"].probabilities)   # {'abrechnung': 0.68, 'technik': 0.32, 'vertrieb': 0.0}\r\nprint(antworten[\"dringlichkeit\"].score)  # 1.03\r\nprint(antworten[\"spam\"].noul)            # 0.04\r\n\r\nteam = antworten[\"team\"].choice\r\n<span class=\"hljs-comment\"># If confidence is below 80%, route to \"manual_review\"<\/span> \r\nziel = team if antworten[\"team\"].probabilities[team] &gt;= 0.8 else \"manuelle_pruefung\"<\/code><\/pre>\n<table>\n<thead>\n<tr>\n<th>Ticket<\/th>\n<th>Jev\u2019s Choice<\/th>\n<th>Score<\/th>\n<th>Result<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>\u201cI was charged the same amount twice in March \u2026\u201d<\/td>\n<td>billing<\/td>\n<td>1.00<\/td>\n<td>automatically to billing<\/td>\n<\/tr>\n<tr>\n<td>\u201cEver since the update, the app crashes immediately when I log in.\u201d<\/td>\n<td>technology<\/td>\n<td>1.00<\/td>\n<td>automatically to the technology department<\/td>\n<\/tr>\n<tr>\n<td>\u201cI&#8217;d like to upgrade my subscription to Premium \u2026\u201d<\/td>\n<td>sales<\/td>\n<td>0.99<\/td>\n<td>automatically to the sales department<\/td>\n<\/tr>\n<tr>\n<td>\u201cSince the app update, the billing page sometimes displays an incorrect amount \u2026\u201d (code example)<\/td>\n<td>billing<\/td>\n<td>0.68<\/td>\n<td>manual check<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Jev assigns unique tickets a value of 0.99 or 1.00. The ticket from the code example is intentionally ambiguous: An incorrect amount after an app update could be caused by an app error or a billing error. With a billing score of 0.68, it is routed to a human.<\/p>\n<p>Such a request is inexpensive and fast: According to TypeSafe, one million tokens of input cost $0.042, while the output costs nothing. A request can include up to 64.000 tokens and takes 70 to 500 ms; in our tests, it took around 300 ms (TypeSafe AI, 2026c). Jev processes only text and prefers English.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"What-capabilities-existed-before-Jev\"><\/span>What capabilities existed before Jev?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>None of these capabilities is new in and of itself.<\/p>\n<p><strong>Before LLMs<\/strong><\/p>\n<p>For a long time, a separate classification model\u2014such as BERT (Devlin et al., 2019)\u2014was trained for each question using examples for which the correct answer was known. Such a model runs in milliseconds on dedicated hardware and always provides the same answer, but it only answers the question for which it was trained. The first methods that did not require dedicated training already existed back then: One asks a model whether a text supports a statement such as \u201cThis text is about an invoice,\u201d and read the probability of that being true (Yin et al., 2019). This corresponds to Jev\u2019s yes-or-no question.<\/p>\n<p><strong>With LLMs<\/strong><\/p>\n<p>Zero-shot and few-shot. An LLM can solve a classification problem based solely on the instruction, drawing on its general knowledge\u2014just like Jev. If additional examples of how cases should be classified are included in the prompt, it adapts to the specific domain (Brown et al., 2020).<\/p>\n<p><strong>Structured Output<\/strong><\/p>\n<p>A refinement specifies which values the response may take. In constrained decoding, all tokens that do not match a permitted response are blocked at every position: If only \u201cbilling,\u201d \u201ctechnology,\u201d or \u201csales\u201d are allowed, the model cannot write anything else\u2014with a preceding justification if desired. The major providers offer this via their APIs (OpenAI, 2024; Anthropic, 2026; Google, 2026a); for self-hosted models, options include the Outlines library (Willard &amp; Louf, 2023; dottxt, 2026).<sup id=\"fnref-2\"><a href=\"#fn-2\">2<\/a><\/sup> If you additionally read the logprobs of the allowed tokens\u2014that is, the logarithm of their probability\u2014you obtain a value for each label. This is always possible with self-hosted models; the current Gemini models no longer provide logprobs via their API (checked on September 28, 2026).<\/p>\n<h2><span class=\"ez-toc-section\" id=\"What-makes-Jev-different\"><\/span>What makes Jev different?<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>LLMs are generalists in two respects: in terms of tasks and knowledge. Jev specializes in the task at hand\u2014it only makes judgments. However, when it comes to knowledge and context, Jev remains a generalist.<\/p>\n<p>According to TypeSafe, this specialization stems from the training: Jev is not a modified LLM, but rather a standalone model trained specifically for decision-making; the method is called Reinforcement Learning for Calibrated Decisions (RLCD). Neither the method nor the model\u2019s architecture has been published (TypeSafe AI, 2026a). Related to this is the published method RLCR, which fully rewards a model only if its stated confidence matches its hit rate (Damani et al., 2025).<\/p>\n<p>Above all, Jev turns a well-known method into a product: As with an LLM, an instruction defines the task\u2014there is no need for task-specific training. This can lower the barrier to adoption. Compared to an LLM, Jev is cheaper and faster. Gemini 3.5 Flash-Lite, Google\u2019s smallest model and our benchmark, costs $0.30 per million input tokens\u2014a good seven times as much as Jev\u2014plus $2.50 per million output tokens, including reasoning tokens (Google, 2026c). In our tests, Jev responded after about 0.3 seconds, while Flash-Lite took 2 to 7 seconds\u2014partly because it generates reasoning tokens by default (Google, 2026b).<\/p>\n<p>The trade-off is transparency: Jev does not provide any reasoning. An LLM with structured output can do this, as demonstrated by the same ticket using Gemini (Google, 2026a):<\/p>\n<pre style=\"background-color: #f6f8fa; border: 1px solid #e1e4e8; border-radius: 6px; padding: 16px; font-family: SFMono-Regular, Consolas, 'Liberation Mono', Menlo, monospace; font-size: 13px; line-height: 1.4;\"><code>from typing import Literal\r\nfrom google import genai\r\nfrom pydantic import BaseModel, Field\r\n\r\nclass Einordnung(BaseModel):\r\n    <span class=\"hljs-comment\"># Order: first reasoning, then label, then self-assessment<\/span> \r\n    begruendung: str = Field(description=\"Kurze Begr\u00fcndung, vor dem Label geschrieben\")\r\n<span class=\"hljs-comment\">    # \"Short reasoning, written before the label\"<\/span> \r\n    team: Literal[\"abrechnung\", \"technik\", \"vertrieb\"] # [\"billing\", \"tech\", \"sales\"]\r\n    sicherheit: Literal[\"eindeutig\", \"moeglich\"] = Field( # [\"unambiguous\", \"possible\"]\r\n        description=(\r\n            \"'eindeutig' nur, wenn das Ticket zweifelsfrei zu genau einem Team geh\u00f6rt; \"\r\n            \"'moeglich', sobald es Aspekte aus mehr als einem Team-Bereich enth\u00e4lt\"\r\n        )\r\n<span class=\"hljs-comment\">        # \"'unambiguous' only if the ticket undoubtedly belongs to exactly one team; \"<\/span> \r\n<span class=\"hljs-comment\">        # \"'possible' as soon as it contains aspects from more than one team area\"<\/span>\r\n    )\r\n\r\nclient = genai.Client()  # requires a GEMINI_API_KEY\r\nantwort = client.models.generate_content(\r\n    model=\"gemini-3.5-flash-lite\",\r\n    contents=f\"Welches Team soll dieses Ticket bearbeiten?\\n{ticket}\", <span class=\"hljs-comment\"># \"Which team should handle this ticket?\"<\/span>\r\n    config={\r\n        \"response_mime_type\": \"application\/json\",\r\n        \"response_schema\": Einordnung,\r\n        \"temperature\": 0.0,  # Default would be 1.0\r\n        \"seed\": 42,          # for more reproducible runs\r\n    },\r\n)\r\n\r\neinordnung = Einordnung.model_validate_json(antwort.text)\r\nprint(einordnung.team)        # abrechnung (billing)\r\nprint(einordnung.sicherheit)  # moeglich (possible)\r\nprint(einordnung.begruendung) # Das Ticket betrifft sowohl einen Abrechnungsfehler als auch ein technisches Problem nach einem App-Update.\r\n                                                  # (The ticket concerns both a billing error and a technical issue after an app update.)\r\n\r\n<span class=\"hljs-comment\"># The self-assessment determines whether a human reviews the ticket<\/span> \r\nziel = einordnung.team if einordnung.sicherheit == \"eindeutig\" else \"manuelle_pruefung\"<\/code><\/pre>\n<p>Gemini provides a valid label, a clear explanation, and, in this case, a self-assessment as well. The ticket classifies it as \u201cpossible\u201d and states the reason right away: both a billing error and a technical issue. The application therefore passes it on\u2014like Jev\u2014to a human when the value falls below the threshold. This isn\u2019t a probability, but rather a signal of uncertainty that stems directly from the meaning of the response. The burden falls on the user: They must define the levels of self-assessment and verify whether the model applies them reliably.<sup id=\"fnref-3\"><a href=\"#fn-3\">3<\/a><\/sup><\/p>\n<table>\n<thead>\n<tr>\n<th>Approach<\/th>\n<th>How the task is defined<\/th>\n<th>Probabilities?<\/th>\n<th>Explanations?<\/th>\n<th>Hosting<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Custom classifier (such as BERT)<\/td>\n<td>Training with examples, per task<\/td>\n<td>yes<\/td>\n<td>no<\/td>\n<td>self-operated, milliseconds<\/td>\n<\/tr>\n<tr>\n<td>Zero-Shot Text Inference (Yin et al., 2019)<\/td>\n<td>Message when called<\/td>\n<td>yes, not calibrated<\/td>\n<td>no<\/td>\n<td>self-operated, cheap<\/td>\n<\/tr>\n<tr>\n<td>LLM with Structured Output via API<\/td>\n<td>Instructions and examples in the prompt<\/td>\n<td>as a reported category or number, often too high (Xiong et al., 2024)<\/td>\n<td>yes<\/td>\n<td>as a service, with reasoning: seconds<\/td>\n<\/tr>\n<tr>\n<td>Self-operated LLM with Constrained Decoding<\/td>\n<td>Instructions and examples in the prompt<\/td>\n<td>Yes, it can be calibrated later<\/td>\n<td>yes<\/td>\n<td>self-hosted; affordable for small models<\/td>\n<\/tr>\n<tr>\n<td>Jev<\/td>\n<td>Instructions in the question, rules in the state<\/td>\n<td>Yes, calibrated according to the manufacturer&#8217;s specifications<\/td>\n<td>no<\/td>\n<td>Available only as a service in the U.S., about 0.3 s<\/td>\n<\/tr>\n<tr>\n<td>Open models with a Jev interface (Kev, SemIf)<\/td>\n<td>like Jev<\/td>\n<td>yes<\/td>\n<td>no<\/td>\n<td>self-operated<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2><span class=\"ez-toc-section\" id=\"Open-alternatives\"><\/span><b>Open alternatives<\/b><span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Jev&#8217;s interface didn&#8217;t remain exclusive for long: Just a few days after its release, several open-source versions with the same interface became available that users can run themselves:<\/p>\n<table>\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Foundation<\/th>\n<th>Availability<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Kev (Palmer, 2026)<\/td>\n<td>Qwen3.5 in three sizes, retrained<\/td>\n<td>free (Apache 2.0), compatible with the TypeSafe SDK<\/td>\n<\/tr>\n<tr>\n<td>SemIf, formerly OpenJev (TheoLeeCJ, 2026)<\/td>\n<td>Qwen 3.5 with 4 billion parameters<\/td>\n<td>free, can be operated locally<\/td>\n<\/tr>\n<tr>\n<td>djev (Maisa, 2026)<\/td>\n<td>Diffusion model based on Gemma<\/td>\n<td>free (Apache 2.0), self-hosted<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2><span class=\"ez-toc-section\" id=\"We-tested-Jev\"><\/span>We tested Jev<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>We submitted several hundred queries to Jev and Gemini 3.5 Flash-Lite. The cases are small and, in some instances, contrived.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Related-fields-may-contradict-each-other\"><\/span>Related fields may contradict each other<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Jev evaluates each question separately; only the state is shared (TypeSafe AI, 2026d). TypeSafe sees this as a strength: According to TypeSafe, the most reliable workflows consist of many independent, disaggregated questions (TypeSafe, 2026a). Things get tricky as soon as answers need to match up. Here\u2019s an example: A painting disappeared from a museum overnight, and exactly one of three scenarios is correct.<\/p>\n<table>\n<thead>\n<tr>\n<th>Scenario<\/th>\n<th>Probability<\/th>\n<th>Perpetrator<\/th>\n<th>Route<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>A<\/td>\n<td>30%<\/td>\n<td>Night Watchman<\/td>\n<td>Loading dock<\/td>\n<\/tr>\n<tr>\n<td>B<\/td>\n<td>25%<\/td>\n<td>Night Watchman<\/td>\n<td>Staff exit<\/td>\n<\/tr>\n<tr>\n<td>C<\/td>\n<td>45%<\/td>\n<td>Burglars from outside<\/td>\n<td>Roof<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Both models were asked the same two questions: \u201cHow did the painting leave the building?\u201d and \u201cWas it an inside job?\u201d Jev answers them independently of one another. Gemini answers both in a single request: Using reasoning, it weighs the questions together before responding, and even without reasoning, it takes the first answer into account when filling in the second field. In addition, we asked both models directly about the scenario.4<\/p>\n<h3><span class=\"ez-toc-section\" id=\"How-to-Interpret-Jevs-Probabilities\"><\/span>How to Interpret Jev\u2019s Probabilities<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>The museum case raises a second question: What does Jev\u2019s 0.96 for the roof mean when the text specifies 45%? A probability is calibrated if, of all answers with a value of 0.8, approximately 80% are correct, always relative to a specific set of cases (H\u00e1jek, 2007). TypeSafe does not specify what Jev\u2019s 0.8 refers to. However, it uses the answers from Frontier LLMs\u2014specifically, the mean of GPT-6 Astra and Claude Fable 5.1\u2014as a benchmark (TypeSafe AI, 2026a). Jev\u2019s figure thus most accurately indicates how likely a Frontier LLM would be to give this answer, not how often it is correct.<\/p>\n<p>An initial independent, non-peer-reviewed test found Jev to be well-calibrated on known benchmarks but less so on 900 newly generated support tickets: too confident on multiple-choice questions and too cautious on yes\/no questions (scienthoon, 2026). Our museum case aligns with the observation regarding multiple-choice questions: As a multiple-choice question, the roof received a score of 0.96 to 0.98 instead of the 0.45 mentioned in the text.<\/p>\n<p>The numbers also fluctuate. In 20 identical runs, the invoice ticket scored between 0.67 and 0.76; when the order of the options was reversed, the score was about 0.1 higher or lower; and in English, it scored only between 0.5 and 0.6. A threshold of 0.70 would therefore sometimes approve identical requests and sometimes refer them to a human.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Potential\"><\/span>Potential<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>The extent of Jev\u2019s potential depends on what it is compared to. TypeSafe evaluates Jev against Frontier models (TypeSafe AI, 2026a). However, many classification tasks do not require a Frontier model; smaller models such as Claude Haiku or Gemini Flash-Lite are sufficient for these tasks. In our experience with enterprise client projects, these are not a significant cost factor; it is the Frontier models that carry the weight. Jev\u2019s cost advantage therefore really counts where a Frontier model has previously been required for classification, and this area is likely quite specialized. Only a benchmark on your own task will show whether Jev can keep up there.<\/p>\n<p>Compared to small models, the wait time is more of a factor: with 0.3 seconds instead of 2 to 7 seconds, checks can be integrated into running processes that would otherwise take too long.<\/p>\n<p>We see the greatest benefit in screening large volumes of data when looking for a needle in a haystack: Screening with high recall drastically narrows down a long list of candidates, and only the few remaining are reviewed by an LLM or a human. In large infrastructures, millions of log lines and events are generated daily; whether a sequence of entries is suspicious often cannot be determined using fixed rules. For retrieval, a Jev-like model can\u2014depending on cost-effectiveness\u2014evaluate a query against several hundred to a few thousand candidates individually, either as a semantic search or during re-ranking. With these volumes, even small LLMs become expensive and slow. Organizations with very large data streams, such as government agencies or energy providers, are likely to benefit the most.<\/p>\n<p>With models like Jev, the building blocks of an AI architecture no longer differ only in size but also in model type. It\u2019s conceivable to have systems in which each model type takes on the task it\u2019s best suited for. A Jev-like model sets the many small parameters: which tool, which context, which model. If a classification needs to be reviewed later, an LLM takes over and writes a justification for it. In the end, a large LLM synthesizes the results into a report that humans can read.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Probability-or-Confidence\"><\/span>Probability or Confidence<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>In addition to cost and wait time, TypeSafe primarily touts its probability metrics. In the use cases we\u2019ve encountered in projects, however, something else usually matters more: a confidence value against which a threshold is set. This value does not need to be calibrated; it simply needs to rank certain cases higher than uncertain ones. The threshold itself is determined based on your own data anyway. Jev also provides such a value for selection and rating questions. However, it is calculated from the probabilities\u2014for three options, as (3 \u00d7 highest probability &#8211; 1) \/ 2 (TypeSafe AI, 2026e)\u2014and therefore says nothing about how confident Jev itself is in the probabilities. True probabilities are only needed when performing calculations with them\u2014for example, when combining multiple estimates using Bayes\u2019 theorem. It remains to be seen whether a model whose structure is unknown and that can only be controlled to a limited extent via context is suitable for such applications.<\/p>\n<p>LLMs also provide a confidence value, either explicitly or implicitly. Explicitly, the model states its own confidence as a word, category, or number; such estimates are often too high (Xiong et al., 2024). Implicitly, through logprobs: The output layer of an LLM calculates a probability for every possible next token, not just the selected one. With self-hosted models, you can always retrieve these values, but with APIs, this is no longer possible everywhere. Whether Jev\u2019s numbers must be accurate as probabilities or are sufficient as confidence scores remains to be seen in practice. In the latter case, Jev is cheaper and faster than an LLM, but not fundamentally new.<\/p>\n<p>Whether probability or confidence score: Only a test on a few hundred of your own cases with known answers will reveal what error rate a threshold entails for your own data: How many tickets end up in the wrong team at 0.8, and how many at 0.9? The same cases can also be used for subsequent calibration (Guo et al., 2017). The threshold should be chosen based on the cost of an error and the cost of human review; this is also recommended by TypeSafe (TypeSafe AI, 2026e). For multiple-choice questions, an option for \u201cunidentifiable\u201d should also be included. Otherwise, the model will distribute the probability among the available answers, even if the text provides no clues.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Limitations\"><\/span>Limitations<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<h3><span class=\"ez-toc-section\" id=\"Limitations-of-the-Technology\"><\/span>Limitations of the Technology<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>TypeSafe itself identifies some weaknesses (TypeSafe AI, 2026f): Jev is unreliable at counting and interprets dates as text, which is why it has difficulty comparing dates. It struggles with double negatives and multi-step inferences; superfluous information in the state distracts it; and instructions hidden in the data can influence its responses. Arithmetic, date logic, and filtering out superfluous data must therefore be handled in the code.<br \/>\nThese weaknesses likely share a common cause. An LLM with reasoning capabilities can generate more tokens when faced with a difficult question, thereby applying more computational effort. Jev does not generate reasoning tokens. This makes it fast and cost-effective, but it also deprives it of this workaround. To regain this capability without reverting to an LLM, such a model would have to internally adjust the computational effort to the difficulty of the question without generating text. It is not known whether Jev is capable of this.<\/p>\n<p>The museum case illustrates a general problem: characteristics that are interdependent\u2014or, statistically speaking, covary\u2014are the rule rather than the exception in practice. Anyone using Jev or a similar model should therefore verify whether the categories in question are truly independent. If they are not, it is better to leave the task to an LLM.<\/p>\n<p>A second limitation concerns the knowledge that Jev draws upon. Jev is a classifier that relies exclusively on general knowledge. It works as long as the task falls within a generally understandable domain. If the task is domain-specific\u2014for example, involving categories that only make sense within one\u2019s own company\u2014one should not expect too much from a model like Jev. For LLMs, in-context learning has proven effective in such cases: You provide examples in the prompt of how cases from your own domain should be classified (Brown et al., 2020). We have not systematically tested whether Jev also uses examples or classification rules stored in the state. In the museum case, Jev did not adopt the ratios specified in the state. It remains to be seen whether TypeSafe will offer a way to adapt Jev to your own domain.<\/p>\n<p>Added to this is the regulatory aspect: Recruitment should not be submitted to such a model for evaluation. Under the EU AI Act, personnel selection is considered a high-risk area (<a href=\"https:\/\/eur-lex.europa.eu\/eli\/reg\/2024\/1689\/oj\">Regulation (EU) 2024\/1689, Annex III<\/a>), and a model without an explanation can hardly be justified in that context. In the case of the Jev product, there is the additional issue that its architecture is unknown. Under such conditions, a model like Jev should prepare decisions rather than be responsible for them: It provides input, but the responsibility for the decision lies elsewhere. This includes logging every response with the specific model version (not \u201cjev-latest\u201d) and the time stamp, especially since the numbers fluctuate. This is the only way to verify a decision later.<\/p>\n<h3><span class=\"ez-toc-section\" id=\"Limitations-of-the-Product\"><\/span>Limitations of the Product<span class=\"ez-toc-section-end\"><\/span><\/h3>\n<p>Jev is currently a solution from the U.S.; there is no documented operation within the EU (TypeSafe AI, 2026c). As a result, the product is often ruled out for European projects involving personal data. However, if the technology proves worthwhile, it will likely be a matter of just a few months before it becomes available in Europe: Either TypeSafe will offer GDPR-compliant operations, or closed or open alternatives will emerge. <a id=\"fnref-5\" href=\"#fn-5\">5<\/a> Open models with the same interface were available within just a few days; no one has yet independently verified whether they can keep up.<\/p>\n<h2><span class=\"ez-toc-section\" id=\"Conclusion\"><\/span>Conclusion<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<p>Jev makes classification more affordable, faster, and more accessible without the need for custom training, but it doesn\u2019t reinvent the wheel. Its results likely reflect what a state-of-the-art LLM would respond rather than how often the answer is correct; nevertheless, they often suffice as a confidence metric for setting a threshold. Jev is a good fit when several conditions align:<\/p>\n<ul>\n<li>individual questions that are independent of one another,<\/li>\n<li>a generally understandable domain,<\/li>\n<li>large volumes for which an LLM is too expensive or too slow and building a custom classifier is too resource-intensive\u2014such as in screening,<\/li>\n<li>preliminary decisions or non-critical decisions that do not need to be justified on a case-by-case basis.<\/li>\n<\/ul>\n<p>If any of these are missing, a dedicated classifier or an LLM is usually the better choice. Even where Jev is appropriate, it should prepare decisions, not be responsible for them: Models of this type carry the same biases as LLMs (Kraft, 2021), but these biases can only be measured statistically; they cannot be identified in individual cases based on a rationale.<br \/>\nIt remains to be seen whether Jev will establish itself as a distinct class of models. If so, AI architectures in the future will no longer consist solely of larger and smaller LLMs, but of various types of models.<\/p>\n<hr \/>\n<p id=\"fn-1\"><small><sup>1<\/sup> Kahneman adopted the terms \u201cSystem 1\u201d and \u201cSystem 2\u201d from Keith Stanovich and Richard West (2000). Stanovich and Jonathan Evans now refer more precisely to \u201cType 1\u201d and \u201cType 2\u201d processes (Evans &amp; Stanovich, 2013). <a href=\"#fnref-1\">\u21a9<\/a><\/small><\/p>\n<p id=\"fn-2\"><small><sup>2<\/sup> This would make it possible to build cascades as early as 2023: A small, inexpensive model provides the initial response, and only uncertain cases are passed on to a larger one (Chen et al., 2023). With Outlines, you can incorporate this uncertainty into the label: GPT-3.5 chooses between labels such as \u201cdefinitely positive\u201d and \u201cpossibly positive,\u201d and only the uncertain cases are sent to GPT-4 (<a href=\"https:\/\/www.linkedin.com\/posts\/dr-ivan-herreros-b64a204_cascaded-llms-for-sentiment-labelling-a-activity-7140970944268873730-J5L2?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAADJIWygBa4JaZ_avecM70QCAW3JPhnQVam0\" target=\"_blank\" rel=\"noopener\">Herreros, 2023<\/a>). <a href=\"#fnref-2\">\u21a9<\/a><\/small><\/p>\n<p id=\"fn-3\"><small><sup>3<\/sup> Self-assessments are often too optimistic (Xiong et al., 2024). <a href=\"#fnref-3\">\u21a9<\/a><\/small><\/p>\n<p id=\"fn-4\"><small><sup>4 <\/sup>Each of Jev\u2019s individual answers is correct on its own: According to the model, an inside perpetrator is more likely than an outside perpetrator (55%), and the roof is the most likely route (45%). Together, however, they form a combination that does not occur in any of the real-world scenarios. If, instead, the models were asked directly about the overall scenario, both models selected the correct scenario C in all runs. Thus, presenting the permitted combinations as a single multiple-choice question can be helpful. However, this approach has its limitations: combining several simple questions into one complex one contradicts Jev\u2019s basic premise. Furthermore, the number of possible combinations grows rapidly with each additional feature, becoming unmanageable. Above all, however, this approach does not solve the actual problem: the best answer to each individual question asked separately does not automatically result in the best overall decision (Dembczy\u0144ski et al., 2012); or, to paraphrase Aristotle: What holds true in isolation need not hold true in combination (De Interpretatione 11).<a href=\"#fnref-4\">\u21a9<\/a><\/small><\/p>\n<p id=\"fn-5\"><small><sup>5<\/sup> <strong>Update September 30, 2026:<\/strong> The market reproduced this offering even faster than anticipated at the time of publication. To illustrate the rapid timeline:<\/small><\/p>\n<ul>\n<li><small><strong>September 15, 2026:<\/strong> TypeSafe AI releases Jev.<\/small><\/li>\n<li><small><strong>September 28, 2026:<\/strong> Initial publication of this blog post.<\/small><\/li>\n<li><small><strong>September 29, 2026:<\/strong> OpenAI announces its commercial <a href=\"https:\/\/openai.com\/index\/devday-2026-recap\/\" target=\"_blank\" rel=\"noopener\">&#8220;Decisions API&#8221;<\/a> at DevDay (functionally replicating Jev&#8217;s service for constrained probabilistic decisions).<\/small><\/li>\n<li><small><strong>September 29, 2026:<\/strong> Ollama announces <a href=\"https:\/\/ollama.com\/blog\/ollama-now-supports-jev-style-decision-models\" target=\"_blank\" rel=\"noopener\">in a blog post<\/a> official support for Jev-style decision models, including a native <code>\/api\/systemone<\/code> endpoint for local execution.<\/small><\/li>\n<\/ul>\n<p><small>This rapid sequence confirms the central thesis of this section: while Jev popularized a compelling interface, the underlying pattern is easily reproducible across both commercial cloud platforms and open local frameworks. <a href=\"#fnref-5\">\u21a9<\/a><\/small><\/p>\n<h2><span class=\"ez-toc-section\" id=\"References\"><\/span>References<span class=\"ez-toc-section-end\"><\/span><\/h2>\n<ul>\n<li>\n<p data-path-to-node=\"6,0,0\"><b data-path-to-node=\"6,0,0\" data-index-in-node=\"0\">Anthropic. (2026).<\/b> <i data-path-to-node=\"6,0,0\" data-index-in-node=\"19\">Structured outputs<\/i> (Dokumentation). <a class=\"ng-star-inserted\" href=\"https:\/\/docs.anthropic.com\/en\/docs\/build-with-claude\/structured-outputs?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/docs.anthropic.com\/en\/docs\/build-with-claude\/structured-outputs<\/a><\/p>\n<\/li>\n<li><b data-path-to-node=\"8,0,0\" data-index-in-node=\"0\">Aristoteles.<\/b> <i data-path-to-node=\"8,0,0\" data-index-in-node=\"13\">De Interpretatione (Lehre vom Satz)<\/i> (E. Rolf, Trans., 1995). Meiner Verlag. (Original written approx. 350 v. Chr.)<\/li>\n<li>\n<p data-path-to-node=\"6,1,0\"><b data-path-to-node=\"6,1,0\" data-index-in-node=\"0\">Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., \u2026 Amodei, D. (2020).<\/b> Language models are few-shot learners. <i data-path-to-node=\"6,1,0\" data-index-in-node=\"300\">Advances in Neural Information Processing Systems (NeurIPS 2020)<\/i>, 33, 1877\u20131901. <a class=\"ng-star-inserted\" href=\"https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,2,0\"><b data-path-to-node=\"6,2,0\" data-index-in-node=\"0\">Chen, L., Zaharia, M., &amp; Zou, J. (2023).<\/b> FrugalGPT: How to use large language models while reducing cost and improving performance. <i data-path-to-node=\"6,2,0\" data-index-in-node=\"132\">arXiv preprint arXiv:2305.05176<\/i>. <a class=\"ng-star-inserted\" href=\"https:\/\/arxiv.org\/abs\/2305.05176?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/arxiv.org\/abs\/2305.05176<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,3,0\"><b data-path-to-node=\"6,3,0\" data-index-in-node=\"0\">Damani, M., Puri, I., Slocum, S., Shenfeld, I., Choshen, L., Kim, Y., &amp; Andreas, J. (2025).<\/b> Beyond binary rewards: Training LMs to reason about their uncertainty. <i data-path-to-node=\"6,3,0\" data-index-in-node=\"163\">arXiv preprint arXiv:2507.16806<\/i>. <a class=\"ng-star-inserted\" href=\"https:\/\/arxiv.org\/abs\/2507.16806?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/arxiv.org\/abs\/2507.16806<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,4,0\"><b data-path-to-node=\"6,4,0\" data-index-in-node=\"0\">Dembczy\u0144ski, K., Waegeman, W., Cheng, W., &amp; H\u00fcllermeier, E. (2012).<\/b> On label dependence and loss minimization in multi-label classification. <i data-path-to-node=\"6,4,0\" data-index-in-node=\"141\">Machine Learning<\/i>, 88(1-2), 5\u201345. <a class=\"ng-star-inserted\" href=\"https:\/\/doi.org\/10.1007\/s10994-012-5292-0?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/doi.org\/10.1007\/s10994-012-5292-0<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,5,0\"><b data-path-to-node=\"6,5,0\" data-index-in-node=\"0\">Devlin, J., Chang, M.-W., Lee, K., &amp; Toutanova, K. (2019).<\/b> BERT: Pre-training of deep bidirectional transformers for language understanding. <i data-path-to-node=\"6,5,0\" data-index-in-node=\"141\">NAACL 2019<\/i>, 4171\u20134186. <a class=\"ng-star-inserted\" href=\"https:\/\/aclanthology.org\/N19-1423\/?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/aclanthology.org\/N19-1423\/<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,6,0\"><b data-path-to-node=\"6,6,0\" data-index-in-node=\"0\">dottxt. (2026).<\/b> <i data-path-to-node=\"6,6,0\" data-index-in-node=\"16\">Outlines<\/i> (Software, GitHub Repository). <a class=\"ng-star-inserted\" href=\"https:\/\/github.com\/dottxt-ai\/outlines?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/dottxt-ai\/outlines<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,7,0\"><b data-path-to-node=\"6,7,0\" data-index-in-node=\"0\">Evans, J. S. B. &amp; Stanovich, K. E. (2013).<\/b> Dual-process theories of higher cognition: Advancing the debate. <i data-path-to-node=\"6,7,0\" data-index-in-node=\"108\">Perspectives on Psychological Science<\/i>, 8(3), 223\u2013241. <a class=\"ng-star-inserted\" href=\"https:\/\/doi.org\/10.1177\/1745691612460685?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/doi.org\/10.1177\/1745691612460685<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,8,0\"><b data-path-to-node=\"6,8,0\" data-index-in-node=\"0\">Google. (2026a).<\/b> <i data-path-to-node=\"6,8,0\" data-index-in-node=\"17\">Structured outputs<\/i> (Gemini API Dokumentation). <a class=\"ng-star-inserted\" href=\"https:\/\/ai.google.dev\/gemini-api\/docs\/structured-output?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/ai.google.dev\/gemini-api\/docs\/structured-output<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,9,0\"><b data-path-to-node=\"6,9,0\" data-index-in-node=\"0\">Google. (2026b).<\/b> <i data-path-to-node=\"6,9,0\" data-index-in-node=\"17\">Gemini thinking \/ Interactions API<\/i> (Gemini API Dokumentation). <a class=\"ng-star-inserted\" href=\"https:\/\/ai.google.dev\/gemini-api\/docs\/thinking?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/ai.google.dev\/gemini-api\/docs\/thinking<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,10,0\"><b data-path-to-node=\"6,10,0\" data-index-in-node=\"0\">Google. (2026c).<\/b> <i data-path-to-node=\"6,10,0\" data-index-in-node=\"17\">Gemini Developer API pricing<\/i> (Pricing list, accessed September 24, 2026). <a class=\"ng-star-inserted\" href=\"https:\/\/ai.google.dev\/pricing?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/ai.google.dev\/pricing<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,11,0\"><b data-path-to-node=\"6,11,0\" data-index-in-node=\"0\">Guo, C., Pleiss, G., Sun, Y., &amp; Weinberger, K. Q. (2017).<\/b> On calibration of modern neural networks. <i data-path-to-node=\"6,11,0\" data-index-in-node=\"100\">ICML 2017<\/i>, 1321\u20131330. <a class=\"ng-star-inserted\" href=\"https:\/\/proceedings.mlr.press\/v70\/guo17a.html?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/proceedings.mlr.press\/v70\/guo17a.html<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,12,0\"><b data-path-to-node=\"6,12,0\" data-index-in-node=\"0\">H\u00e1jek, A. (2007).<\/b> The reference class problem is your problem too. <i data-path-to-node=\"6,12,0\" data-index-in-node=\"67\">Synthese<\/i>, 156(3), 563\u2013585. <a class=\"ng-star-inserted\" href=\"https:\/\/doi.org\/10.1007\/s11229-006-9090-y?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/doi.org\/10.1007\/s11229-006-9090-y<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,13,0\"><b data-path-to-node=\"6,13,0\" data-index-in-node=\"0\">Herreros, I. (2023).<\/b> <i data-path-to-node=\"6,13,0\" data-index-in-node=\"21\">Cascaded LLMs for sentiment labelling with Outlines<\/i> [LinkedIn-Post]. LinkedIn. <a class=\"ng-star-inserted\" href=\"https:\/\/www.linkedin.com\/posts\/dr-ivan-herreros-b64a204_cascaded-llms-for-sentiment-labelling-a-activity-7140970944268873730-J5L2?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/www.linkedin.com\/posts\/dr-ivan-herreros-b64a204_cascaded-llms-for-sentiment-labelling-a-activity-7140970944268873730-J5L2<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,14,0\"><b data-path-to-node=\"6,14,0\" data-index-in-node=\"0\">Kahneman, D. (2011).<\/b> <i data-path-to-node=\"6,14,0\" data-index-in-node=\"21\">Thinking, fast and slow<\/i>. Farrar, Straus and Giroux.<\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,15,0\"><b data-path-to-node=\"6,15,0\" data-index-in-node=\"0\">Kahneman, D. &amp; Tversky, A. (1973).<\/b> On the psychology of prediction. <i data-path-to-node=\"6,15,0\" data-index-in-node=\"68\">Psychological Review<\/i>, 80(4), 237\u2013251. <a class=\"ng-star-inserted\" href=\"https:\/\/doi.org\/10.1037\/h0034747?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/doi.org\/10.1037\/h0034747<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,16,0\"><b data-path-to-node=\"6,16,0\" data-index-in-node=\"0\">Kraft, A. (2021, 9. Dezember).<\/b> <i data-path-to-node=\"6,16,0\" data-index-in-node=\"31\">Social Bias in gro\u00dfen KI-Modellen und wie man damit umgeht<\/i>. inovex Blog. <a class=\"ng-star-inserted\" href=\"https:\/\/www.inovex.de\/de\/blog\/social-bias-in-grossen-ki-modellen\/?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/www.inovex.de\/de\/blog\/social-bias-in-grossen-ki-modellen\/<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,17,0\"><b data-path-to-node=\"6,17,0\" data-index-in-node=\"0\">Maisa. (2026).<\/b> <i data-path-to-node=\"6,17,0\" data-index-in-node=\"15\">djev<\/i> (Software, GitHub Repository).<\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,18,0\"><b data-path-to-node=\"6,18,0\" data-index-in-node=\"0\">OpenAI. (2024, 6. August).<\/b> <i data-path-to-node=\"6,18,0\" data-index-in-node=\"27\">Introducing Structured Outputs in the API<\/i>. OpenAI Blog. <a class=\"ng-star-inserted\" href=\"https:\/\/openai.com\/index\/introducing-structured-outputs-in-the-api\/?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/openai.com\/index\/introducing-structured-outputs-in-the-api\/<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,19,0\"><b data-path-to-node=\"6,19,0\" data-index-in-node=\"0\">Palmer, J. (2026).<\/b> <i data-path-to-node=\"6,19,0\" data-index-in-node=\"19\">Kev: Open decision models<\/i> (Software, README). GitHub. <a class=\"ng-star-inserted\" href=\"https:\/\/github.com\/?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/github.com\/<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,20,0\"><b data-path-to-node=\"6,20,0\" data-index-in-node=\"0\">scienthoon. (2026).<\/b> <i data-path-to-node=\"6,20,0\" data-index-in-node=\"20\">jev-ood-calibration<\/i> (Software, README). GitHub. Not peer-reviewed.<\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,21,0\"><b data-path-to-node=\"6,21,0\" data-index-in-node=\"0\">Stanovich, K. E. &amp; West, R. F. (2000).<\/b> Individual differences in reasoning: Implications for the rationality debate? <i data-path-to-node=\"6,21,0\" data-index-in-node=\"117\">Behavioral and Brain Sciences<\/i>, 23(5), 645\u2013665. <a class=\"ng-star-inserted\" href=\"https:\/\/doi.org\/10.1017\/S0140525X00003435?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/doi.org\/10.1017\/S0140525X00003435<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,22,0\"><b data-path-to-node=\"6,22,0\" data-index-in-node=\"0\">TheoLeeCJ. (2026).<\/b> <i data-path-to-node=\"6,22,0\" data-index-in-node=\"19\">SemIf (fr\u00fcher OpenJev)<\/i> (Software). GitHub.<\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,23,0\"><b data-path-to-node=\"6,23,0\" data-index-in-node=\"0\">TypeSafe AI (2026a).<\/b> <i data-path-to-node=\"6,23,0\" data-index-in-node=\"21\">Introducing System One Models &amp; Jev<\/i>. <a class=\"ng-star-inserted\" href=\"https:\/\/typesafe.ai\/blog\/introducing-system-one-models-and-jev?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/typesafe.ai\/blog\/introducing-system-one-models-and-jev<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,24,0\"><b data-path-to-node=\"6,24,0\" data-index-in-node=\"0\">TypeSafe AI (2026b).<\/b> <i data-path-to-node=\"6,24,0\" data-index-in-node=\"21\">The Three Primitives<\/i>. <a class=\"ng-star-inserted\" href=\"https:\/\/docs.typesafe.ai\/primitives?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/docs.typesafe.ai\/primitives<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,25,0\"><b data-path-to-node=\"6,25,0\" data-index-in-node=\"0\">TypeSafe AI (2026c).<\/b> <i data-path-to-node=\"6,25,0\" data-index-in-node=\"21\">Models<\/i>. <a class=\"ng-star-inserted\" href=\"https:\/\/docs.typesafe.ai\/models?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/docs.typesafe.ai\/models<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,26,0\"><b data-path-to-node=\"6,26,0\" data-index-in-node=\"0\">TypeSafe AI (2026d).<\/b> <i data-path-to-node=\"6,26,0\" data-index-in-node=\"21\">Asking Parallel Questions<\/i> (Cookbook). <a class=\"ng-star-inserted\" href=\"https:\/\/docs.typesafe.ai\/cookbooks\/parallel_questions?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/docs.typesafe.ai\/cookbooks\/parallel_questions<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,27,0\"><b data-path-to-node=\"6,27,0\" data-index-in-node=\"0\">TypeSafe AI (2026e).<\/b> <i data-path-to-node=\"6,27,0\" data-index-in-node=\"21\">Confidence<\/i> (Dokumentation). <a class=\"ng-star-inserted\" href=\"https:\/\/docs.typesafe.ai\/confidence?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/docs.typesafe.ai\/confidence<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,28,0\"><b data-path-to-node=\"6,28,0\" data-index-in-node=\"0\">TypeSafe AI (2026f).<\/b> <i data-path-to-node=\"6,28,0\" data-index-in-node=\"21\">Known Limitations<\/i> (Dokumentation). <a class=\"ng-star-inserted\" href=\"https:\/\/docs.typesafe.ai\/limitations?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/docs.typesafe.ai\/limitations<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,29,0\"><b data-path-to-node=\"6,29,0\" data-index-in-node=\"0\">Verordnung (EU) 2024\/1689 (KI-Verordnung). (2024).<\/b> <i data-path-to-node=\"6,29,0\" data-index-in-node=\"51\">Amtsblatt der Europ\u00e4ischen Union<\/i>, L 2024\/1689. <a class=\"ng-star-inserted\" href=\"http:\/\/data.europa.eu\/eli\/reg\/2024\/1689\/oj?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">http:\/\/data.europa.eu\/eli\/reg\/2024\/1689\/oj<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,30,0\"><b data-path-to-node=\"6,30,0\" data-index-in-node=\"0\">Willard, B. T. &amp; Louf, R. (2023).<\/b> Efficient guided generation for large language models. <i data-path-to-node=\"6,30,0\" data-index-in-node=\"89\">arXiv preprint arXiv:2307.09702<\/i>. <a class=\"ng-star-inserted\" href=\"https:\/\/arxiv.org\/abs\/2307.09702?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/arxiv.org\/abs\/2307.09702<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,31,0\"><b data-path-to-node=\"6,31,0\" data-index-in-node=\"0\">Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., Liu, Y., &amp; Liu, Q. (2024).<\/b> Can LLMs express their uncertainty? An empirical evaluation of confidence elicitation in LLMs. <i data-path-to-node=\"6,31,0\" data-index-in-node=\"165\">ICLR 2024<\/i>. <a class=\"ng-star-inserted\" href=\"https:\/\/openreview.net\/forum?id=1TqD4fF5yJ&amp;utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/openreview.net\/forum?id=1TqD4fF5yJ<\/a><\/p>\n<\/li>\n<li>\n<p data-path-to-node=\"6,32,0\"><b data-path-to-node=\"6,32,0\" data-index-in-node=\"0\">Yin, W., Hay, J., &amp; Roth, D. (2019).<\/b> Benchmarking zero-shot text classification: Datasets, evaluation and entailment approach. <i data-path-to-node=\"6,32,0\" data-index-in-node=\"127\">EMNLP-IJCNLP 2019<\/i>, 3914\u20133923. <a class=\"ng-star-inserted\" href=\"https:\/\/aclanthology.org\/D19-1395\/?utm_source=gemini\" target=\"_blank\" rel=\"noopener\">https:\/\/aclanthology.org\/D19-1395\/<\/a><\/p>\n<\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>This article was automatically translated. Original article. On September 15, 2026, TypeSafe AI unveiled its first public model: Jev, designed for classification tasks (TypeSafe AI, 2026a). Instead of responding with text, it provides numerical answers and, for multiple-choice questions, also selects the chosen option from a predefined list. Typical questions include \u201cIs this spam?\u201d, \u201cWhich [&hellip;]<\/p>\n","protected":false},"author":464,"featured_media":70277,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"inline_featured_image":false,"ep_exclude_from_search":false,"footnotes":""},"tags":[],"service":[473],"level":[1280],"coauthors":[{"id":464,"display_name":"Nicolas Werner","user_nicename":"nwerner"},{"id":465,"display_name":"Ivan Herreros Alonso","user_nicename":"iherreros-alonso"}],"class_list":["post-70278","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","service-artificial-intelligence-en"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.5 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>TypeSafe AI Jev Review: How Good Is the New Model for AI Classification? - inovex GmbH<\/title>\n<meta name=\"description\" content=\"How does the System-One model Jev perform in classification tasks? A real-world test of confidence, cost, and speed.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"TypeSafe AI Jev Review: How Good Is the New Model for AI Classification? - inovex GmbH\" \/>\n<meta property=\"og:description\" content=\"How does the System-One model Jev perform in classification tasks? A real-world test of confidence, cost, and speed.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/\" \/>\n<meta property=\"og:site_name\" content=\"inovex GmbH\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/inovexde\" \/>\n<meta property=\"article:published_time\" content=\"2026-09-29T08:49:27+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-10-02T11:05:28+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.inovex.de\/wp-content\/uploads\/Jev_TypeSafe_AI.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1500\" \/>\n\t<meta property=\"og:image:height\" content=\"880\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Nicolas Werner, Ivan Herreros Alonso\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/www.inovex.de\/wp-content\/uploads\/Jev_TypeSafe_AI-1024x601.png\" \/>\n<meta name=\"twitter:creator\" content=\"@inovexgmbh\" \/>\n<meta name=\"twitter:site\" content=\"@inovexgmbh\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Nicolas Werner, Ivan Herreros Alonso\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"24 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\\\/\"},\"author\":{\"name\":\"Nicolas Werner\",\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/#\\\/schema\\\/person\\\/dbf2a7628b5fc2149d389512f71d91af\"},\"headline\":\"TypeSafe AI Jev Review: How Good Is the New Model for AI Classification?\",\"datePublished\":\"2026-09-29T08:49:27+00:00\",\"dateModified\":\"2026-10-02T11:05:28+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\\\/\"},\"wordCount\":4494,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.inovex.de\\\/wp-content\\\/uploads\\\/Jev_TypeSafe_AI.png\",\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\\\/\",\"url\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\\\/\",\"name\":\"TypeSafe AI Jev Review: How Good Is the New Model for AI Classification? - inovex GmbH\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.inovex.de\\\/wp-content\\\/uploads\\\/Jev_TypeSafe_AI.png\",\"datePublished\":\"2026-09-29T08:49:27+00:00\",\"dateModified\":\"2026-10-02T11:05:28+00:00\",\"description\":\"How does the System-One model Jev perform in classification tasks? A real-world test of confidence, cost, and speed.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\\\/#primaryimage\",\"url\":\"https:\\\/\\\/www.inovex.de\\\/wp-content\\\/uploads\\\/Jev_TypeSafe_AI.png\",\"contentUrl\":\"https:\\\/\\\/www.inovex.de\\\/wp-content\\\/uploads\\\/Jev_TypeSafe_AI.png\",\"width\":1500,\"height\":880},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"TypeSafe AI Jev Review: How Good Is the New Model for AI Classification?\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/#website\",\"url\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/\",\"name\":\"inovex GmbH\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/#organization\",\"name\":\"inovex GmbH\",\"url\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/www.inovex.de\\\/wp-content\\\/uploads\\\/2021\\\/03\\\/inovex-logo-16-9-1.png\",\"contentUrl\":\"https:\\\/\\\/www.inovex.de\\\/wp-content\\\/uploads\\\/2021\\\/03\\\/inovex-logo-16-9-1.png\",\"width\":1921,\"height\":1081,\"caption\":\"inovex GmbH\"},\"image\":{\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/#\\\/schema\\\/logo\\\/image\\\/\"},\"sameAs\":[\"https:\\\/\\\/www.facebook.com\\\/inovexde\",\"https:\\\/\\\/x.com\\\/inovexgmbh\",\"https:\\\/\\\/www.instagram.com\\\/inovexlife\\\/\",\"https:\\\/\\\/www.linkedin.com\\\/company\\\/inovex\",\"https:\\\/\\\/www.youtube.com\\\/channel\\\/UC7r66GT14hROB_RQsQBAQUQ\"]},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/#\\\/schema\\\/person\\\/dbf2a7628b5fc2149d389512f71d91af\",\"name\":\"Nicolas Werner\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/b480bc2f7cc503268b8d04a93557f65cc5cd09840ce62f5f745102f73a3f1b29?s=96&d=retro&r=g8439c5f586fe0e81da8a2c2813b77254\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/b480bc2f7cc503268b8d04a93557f65cc5cd09840ce62f5f745102f73a3f1b29?s=96&d=retro&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/b480bc2f7cc503268b8d04a93557f65cc5cd09840ce62f5f745102f73a3f1b29?s=96&d=retro&r=g\",\"caption\":\"Nicolas Werner\"},\"url\":\"https:\\\/\\\/www.inovex.de\\\/en\\\/blog\\\/author\\\/nwerner\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"TypeSafe AI Jev Review: How Good Is the New Model for AI Classification? - inovex GmbH","description":"How does the System-One model Jev perform in classification tasks? A real-world test of confidence, cost, and speed.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/","og_locale":"en_US","og_type":"article","og_title":"TypeSafe AI Jev Review: How Good Is the New Model for AI Classification? - inovex GmbH","og_description":"How does the System-One model Jev perform in classification tasks? A real-world test of confidence, cost, and speed.","og_url":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/","og_site_name":"inovex GmbH","article_publisher":"https:\/\/www.facebook.com\/inovexde","article_published_time":"2026-09-29T08:49:27+00:00","article_modified_time":"2026-10-02T11:05:28+00:00","og_image":[{"width":1500,"height":880,"url":"https:\/\/www.inovex.de\/wp-content\/uploads\/Jev_TypeSafe_AI.png","type":"image\/png"}],"author":"Nicolas Werner, Ivan Herreros Alonso","twitter_card":"summary_large_image","twitter_image":"https:\/\/www.inovex.de\/wp-content\/uploads\/Jev_TypeSafe_AI-1024x601.png","twitter_creator":"@inovexgmbh","twitter_site":"@inovexgmbh","twitter_misc":{"Written by":"Nicolas Werner, Ivan Herreros Alonso","Est. reading time":"24 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#article","isPartOf":{"@id":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/"},"author":{"name":"Nicolas Werner","@id":"https:\/\/www.inovex.de\/en\/#\/schema\/person\/dbf2a7628b5fc2149d389512f71d91af"},"headline":"TypeSafe AI Jev Review: How Good Is the New Model for AI Classification?","datePublished":"2026-09-29T08:49:27+00:00","dateModified":"2026-10-02T11:05:28+00:00","mainEntityOfPage":{"@id":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/"},"wordCount":4494,"commentCount":0,"publisher":{"@id":"https:\/\/www.inovex.de\/en\/#organization"},"image":{"@id":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#primaryimage"},"thumbnailUrl":"https:\/\/www.inovex.de\/wp-content\/uploads\/Jev_TypeSafe_AI.png","inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/","url":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/","name":"TypeSafe AI Jev Review: How Good Is the New Model for AI Classification? - inovex GmbH","isPartOf":{"@id":"https:\/\/www.inovex.de\/en\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#primaryimage"},"image":{"@id":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#primaryimage"},"thumbnailUrl":"https:\/\/www.inovex.de\/wp-content\/uploads\/Jev_TypeSafe_AI.png","datePublished":"2026-09-29T08:49:27+00:00","dateModified":"2026-10-02T11:05:28+00:00","description":"How does the System-One model Jev perform in classification tasks? A real-world test of confidence, cost, and speed.","breadcrumb":{"@id":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#primaryimage","url":"https:\/\/www.inovex.de\/wp-content\/uploads\/Jev_TypeSafe_AI.png","contentUrl":"https:\/\/www.inovex.de\/wp-content\/uploads\/Jev_TypeSafe_AI.png","width":1500,"height":880},{"@type":"BreadcrumbList","@id":"https:\/\/www.inovex.de\/en\/blog\/typesafe-ai-jev-review-how-good-is-the-new-model-for-ai-classification\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.inovex.de\/en\/"},{"@type":"ListItem","position":2,"name":"TypeSafe AI Jev Review: How Good Is the New Model for AI Classification?"}]},{"@type":"WebSite","@id":"https:\/\/www.inovex.de\/en\/#website","url":"https:\/\/www.inovex.de\/en\/","name":"inovex GmbH","description":"","publisher":{"@id":"https:\/\/www.inovex.de\/en\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.inovex.de\/en\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/www.inovex.de\/en\/#organization","name":"inovex GmbH","url":"https:\/\/www.inovex.de\/en\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.inovex.de\/en\/#\/schema\/logo\/image\/","url":"https:\/\/www.inovex.de\/wp-content\/uploads\/2021\/03\/inovex-logo-16-9-1.png","contentUrl":"https:\/\/www.inovex.de\/wp-content\/uploads\/2021\/03\/inovex-logo-16-9-1.png","width":1921,"height":1081,"caption":"inovex GmbH"},"image":{"@id":"https:\/\/www.inovex.de\/en\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/inovexde","https:\/\/x.com\/inovexgmbh","https:\/\/www.instagram.com\/inovexlife\/","https:\/\/www.linkedin.com\/company\/inovex","https:\/\/www.youtube.com\/channel\/UC7r66GT14hROB_RQsQBAQUQ"]},{"@type":"Person","@id":"https:\/\/www.inovex.de\/en\/#\/schema\/person\/dbf2a7628b5fc2149d389512f71d91af","name":"Nicolas Werner","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/b480bc2f7cc503268b8d04a93557f65cc5cd09840ce62f5f745102f73a3f1b29?s=96&d=retro&r=g8439c5f586fe0e81da8a2c2813b77254","url":"https:\/\/secure.gravatar.com\/avatar\/b480bc2f7cc503268b8d04a93557f65cc5cd09840ce62f5f745102f73a3f1b29?s=96&d=retro&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/b480bc2f7cc503268b8d04a93557f65cc5cd09840ce62f5f745102f73a3f1b29?s=96&d=retro&r=g","caption":"Nicolas Werner"},"url":"https:\/\/www.inovex.de\/en\/blog\/author\/nwerner\/"}]}},"_links":{"self":[{"href":"https:\/\/www.inovex.de\/en\/wp-json\/wp\/v2\/posts\/70278","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.inovex.de\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.inovex.de\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.inovex.de\/en\/wp-json\/wp\/v2\/users\/464"}],"replies":[{"embeddable":true,"href":"https:\/\/www.inovex.de\/en\/wp-json\/wp\/v2\/comments?post=70278"}],"version-history":[{"count":6,"href":"https:\/\/www.inovex.de\/en\/wp-json\/wp\/v2\/posts\/70278\/revisions"}],"predecessor-version":[{"id":70459,"href":"https:\/\/www.inovex.de\/en\/wp-json\/wp\/v2\/posts\/70278\/revisions\/70459"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.inovex.de\/en\/wp-json\/wp\/v2\/media\/70277"}],"wp:attachment":[{"href":"https:\/\/www.inovex.de\/en\/wp-json\/wp\/v2\/media?parent=70278"}],"wp:term":[{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.inovex.de\/en\/wp-json\/wp\/v2\/tags?post=70278"},{"taxonomy":"service","embeddable":true,"href":"https:\/\/www.inovex.de\/en\/wp-json\/wp\/v2\/service?post=70278"},{"taxonomy":"level","embeddable":true,"href":"https:\/\/www.inovex.de\/en\/wp-json\/wp\/v2\/level?post=70278"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/www.inovex.de\/en\/wp-json\/wp\/v2\/coauthors?post=70278"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}