Home → Jev, TypeSafe's "System One" model, tested in Node.js: 400 support tickets, accuracy, latency, cost and calibration

Jev, TypeSafe's "System One" model, tested in Node.js: 400 support tickets, accuracy, latency, cost and calibration

By · Node.js & JavaScript developer
Published October 2, 2026

Jev is the first model from TypeSafe AI, released in early access on September 15, 2026. It isn't a chat model. It never produces text, so there's nothing to parse and nothing to hallucinate outside the answers you define. You give it a state (text or JSON) and named, typed questions, and it returns an answer for each one with probabilities:

  • choice: pick one of your labels, with a probability per label.
  • noul: a yes/no question, answered with the probability of yes.
  • score: rate on an ordered rubric, with a probability per level.

TypeSafe calls this a "System One" model, built for fast decisions inside software, such as routing, triage and guards, rather than for conversation. Its marketing claims it is 193 times faster and 444 times cheaper than LLMs on such tasks. Instead of repeating that, this post measures it on one realistic job: routing customer support messages.

Setup: @typesafe-ai/sdk 0.6.0 (the official, MIT-licensed SDK) on node:24-slim (Node 24.21), models jev-latest and jev-preview, tested on October 2, 2026 with an early-access account. The data is the Banking77 test set from PolyAI (CC BY 4.0): real customer messages labelled with one of 77 intents. The test uses 10 intents that a support team would route differently and that are easy to confuse, with all 40 test messages for each, 400 in total.

Calling Jev from Node.js

npm install @typesafe-ai/sdk
import { TypeSafeClient, choice } from '@typesafe-ai/sdk';

const client = new TypeSafeClient(); // reads TYPESAFE_API_KEY

const route = choice('Which support team should handle this banking customer message?', {
  card_arrival: 'Customer is waiting for a new card to arrive',
  lost_or_stolen_card: 'Card was lost or stolen',
  compromised_card: 'Card details may have been used by someone else (fraud)',
  card_not_working: 'Card does not work at all',
  declined_card_payment: 'A card payment was declined',
  pending_card_payment: 'A card payment is still pending',
  request_refund: 'Customer wants a refund for a purchase',
  transfer_not_received_by_recipient: 'A bank transfer was sent but the recipient did not get it',
  pending_transfer: 'A bank transfer is still pending',
  failed_transfer: 'A bank transfer failed',
});

const { answers, usage } = await client.systemOne({
  state: 'I made a transfer yesterday but my friend says the money is not there.',
  questions: { intent: route },
});

console.log(answers.intent.choice, answers.intent.confidence, usage.input_tokens);
transfer_not_received_by_recipient 0.99 512

The descriptions are optional (null is allowed), but they're where you explain your labels, the way you would in an LLM prompt. The client defaults to jev-latest, a 10-second timeout per attempt and two retries on 408, 429 and 5xx responses.

The SDK's types follow your questions. A choice answer's .choice is typed as the union of your label keys, so a typo or an impossible label is a compile error. With tsc --strict:

import { TypeSafeClient, choice, noul } from '@typesafe-ai/sdk';
const client = new TypeSafeClient();
const { answers } = await client.systemOne({
  state: 'My card was declined',
  questions: { team: choice('Which team?', { cards: null, transfers: null }), urgent: noul('Urgent?') },
});
const team: 'cards' | 'transfers' = answers.team.choice;   // inferred from the criteria keys
const p: number = answers.urgent.noul;
const wrong: 'refunds' = answers.team.choice;              // should be a type error
console.log(team, p, wrong, answers.team.typo);            // .typo should be a type error too
types-check.ts(9,7): error TS2322: Type '"cards" | "transfers"' is not assignable to type '"refunds"'.
types-check.ts(10,42): error TS2339: Property 'typo' does not exist on type 'ChoiceResponse<{ readonly cards: null; readonly transfers: null; }>'.

Both deliberate mistakes were caught, and the valid lines compiled.

The benchmark

The script sends each message as one request, 8 at a time, with retries switched off so failures would be counted rather than hidden. It records the answer, the confidence, the latency and the token usage:

const client = new TypeSafeClient({ retry: { maxRetries: 0 } });

async function classify(row) {
  const t = performance.now();
  const { answers, usage } = await client.systemOne({ state: row.text, questions: { intent: route } });
  return { ...row, predicted: answers.intent.choice, confidence: answers.intent.confidence,
    ms: performance.now() - t, tokens: usage.input_tokens };
}

For comparison, the same 400 messages, the same intent descriptions and the same order went to Llama 3.2 3B, running locally in Ollama 0.35.0 at temperature 0. Its output was restricted with a JSON schema to one of the 10 labels plus a self-reported confidence. That's the "just use a small LLM with structured output" approach many teams would try first. It isn't a speed comparison with a hosted model: it ran on the CPU of a 4-core Docker VM, so look at its accuracy and confidence rather than its latency.

400 messages, 10 intentsjev-latestjev-previewLlama 3.2 3B (local, CPU)
Accuracy92.8% (371)92.5% (370)85.3% (341)
Errors / failed requests000
Latency p50 / p90 / p99267 / 301 / 360 ms275 / 329 / 426 ms763 / 942 / 1,521 ms
400 requests, wall time13.4 s (concurrency 8)14.2 s (concurrency 8)317 s (one at a time)
Input tokens per request511511204 (different tokenizer)
Cost of the whole run$0.0086$0.0086your own hardware

The cost works out to about $0.021 per 1,000 decisions at the published $0.042 per million input tokens, since output is free. Most of the 511 tokens are the ten label descriptions, which are sent with every message, so short labels are cheaper. At 32 requests in parallel, all 400 finished in 3.9 seconds, about 100 per second, with no errors or rate limiting. The median stayed at 270 ms, while the 99th percentile rose to 673 ms.

Repeating 50 of the messages gave the same answer all 50 times, but confidence values differed by up to 0.08 between runs. Over the full 400, a second run scored 369 instead of 371: two borderline answers flipped.

Confidence you can actually use

Accuracy alone undersells the interesting part. Each answer includes a confidence value, and the probabilities over all labels always added up to exactly 1. If those values are honest, you can let the model handle confident cases and send the rest to a person. Grouping jev-latest's answers by confidence:

confidence   answers   avg confidence   actually correct
0.4–0.5          7            0.467              0.429
0.5–0.6          9            0.541              0.222
0.6–0.7          5            0.634              0.600
0.7–0.8         10            0.748              0.600
0.8–0.9         15            0.861              0.800
0.9–1.0        354            0.994              0.975

Overall, the confidence was close to the real hit rate (an expected calibration error of 0.032). The useful property: low confidence really does mean "probably wrong", so the low-confidence answers are the ones worth a human look. Choosing a threshold:

Automate when confidence ≥Jev: automatedJev: correctLlama 3.2 3B: automatedLlama 3.2 3B: correct
0 (everything)100%92.8%100%85.3%
0.8092.3%96.7%99.0%85.9%
0.9088.5%97.5%74.8%89.6%
0.9586.0%98.0%1.0%100%
0.9979.0%98.7%1.0%100%

With Jev, a 0.9 threshold handled 88.5% of tickets automatically at 97.5% accuracy and sent the uncertain 11.5% to a person. The LLM's self-reported confidence was nearly useless for this: almost every answer said 0.8 or 0.9, so a threshold either changed nothing or excluded almost everything. Asking an LLM "how sure are you?" doesn't produce a probability.

const { intent } = (await client.systemOne({ state: message, questions: { intent: route } })).answers;

if (intent.confidence >= 0.9) {
  await assignTicket(intent.choice);
} else {
  // Show the two most likely teams to whoever triages it
  const top = Object.entries(intent.probabilities).sort((a, b) => b[1] - a[1]).slice(0, 2);
  await sendToTriage(message, top);
}

One detail: confidence isn't just the highest probability. It differed from it by up to 0.07, so use confidence for thresholds and probabilities for ranking alternatives.

Where it gets it wrong

Almost half of jev-latest's 29 mistakes (13) were about transfers: 10 "transfer not received by recipient" messages were routed to "pending transfer", and 3 "failed transfer" messages were too. The mistakes made with the highest confidence:

1.000 [card_not_working → declined_card_payment] My card was declined today when eating and I need to know what's wrong.
1.000 [transfer_not_received_by_recipient → pending_transfer] When will my funds transfer?
0.990 [pending_transfer → pending_card_payment] I would like to know why my payment is still pending, can you help?
0.990 [pending_card_payment → declined_card_payment] Why has my payment not gone through?
0.970 [card_arrival → lost_or_stolen_card] How do I locate my card?
0.960 [failed_transfer → pending_transfer] Why hasn't my transfer been made?

Read them as a support agent would, and most are arguable. "My card was declined" sounds like a declined payment, "How do I locate my card?" could easily be a lost card, and "When will my funds transfer?" doesn't say whether anything was received. Part of what remains is ambiguity in the dataset's own labels, and no threshold catches those: a few wrong answers come back with confidence 1.0. If a decision is expensive to get wrong, keep a way for people to correct it.

Several questions in one call

All questions in a request are answered in one pass, so extra questions are nearly free. One message, first with one question, then with four, 20 runs each:

import { TypeSafeClient, choice, noul, score } from '@typesafe-ai/sdk';

const client = new TypeSafeClient();
const message = "I sent £400 to my landlord on Monday and he still hasn't got it. This is the second time, I'm really fed up.";

const team = choice('Which team should handle this?', {
  cards: 'Card problems: lost, stolen, not working, declined',
  transfers: 'Bank transfers: pending, failed, not received',
  refunds: 'Refunds and disputes',
});

const one = { team };
const four = {
  team,
  urgent: noul('Does the customer need a reply today?'),
  frustration: score('How frustrated is the customer?', ['calm', 'slightly annoyed', 'frustrated', 'angry']),
  mentionsAmount: noul('Does the message mention a specific amount of money?'),
};
1 question:  median 265 ms, 372 input tokens
4 questions: median 261 ms, 454 input tokens
team: transfers (1.00) { transfers: 1, cards: 0, refunds: 0 }
urgent: 0.72 | mentionsAmount: 0.99
frustration: 2.09 → "frustrated" (confidence 0.91) { '0': 0, '1': 0, '2': 0.91, '3': 0.09 }

Four questions took the same time as one and added 82 input tokens. A score answer is an expected value (2.09 here, between "frustrated" and "angry") together with the probability of each level, and legend maps the levels back to your rubric.

Is it worth using?

For decisions with a fixed set of answers, such as routing, tagging, triage, "does this need a human?" and moderation-style yes/no checks, Jev did what it promises in this test. It was more accurate than a small LLM with structured output, fast and flat in latency, very cheap, and its confidence values can actually be used to choose what to automate. The typed SDK removes the usual JSON-parsing and "the model invented a new label" code.

Things to keep in mind:

  • It's early access, from a company that launched two weeks ago. There are two model versions (jev-latest and jev-preview) that behaved almost identically here. Expect changes, and pin the model name if you need repeatable results.
  • It only answers questions you define. It won't write a reply, summarize or extract free-form fields. Use it next to an LLM, not instead of one.
  • Your data goes to an external API, like any hosted model. Check the terms before sending customer messages. Cloudflare also offers Jev on Workers AI as typesafe/jev, at the same price, with zero data retention according to its model page.
  • Measure on your own data. 92.8% on these ten Banking77 intents says nothing definite about your labels. A benchmark like the one above takes a few dozen lines and runs in seconds.
  • The vendor's "193× faster, 444× cheaper" compares against large LLMs. Against a small local model, the honest summary from this test is: more accurate, and with confidence values that mean something. Speed depends too much on hardware to compare with a model running on your own CPU.

Sources & further reading

About Code with Node.js

This is a personal blog and reference point of a Node.js developer.

I write and explain how different Node and JavaScript aspects work, as well as research popular and cool packages, and of course fail time to time.