Back to Blog
Vitalii•

Jev, and a New Type of AI Processing with Reasoning Classifiers

Most of what we do with LLMs in production is not chat. It is a decision: is this ticket urgent? Does this answer follow from the source? Should this request go to the expensive model? We ask a model that generates paragraphs to produce one word, then parse that word back out of a string.

A new class of models is questioning that pattern. Three releases in the last few weeks point the same way: Jev from TypeSafe AI, the open Kev-4B adapter, and Span-01 from Respan. Each one is a reasoning classifier, meaning a model that reads a state and returns calibrated probabilities instead of text.

The idea: "System One" models

TypeSafe describes Jev as its first System One Model, a model built for fast, structured decisions inside software rather than for conversation. Their pitch, in the founder's words: "Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."

What makes this different from a normal LLM call:

  • Typed outputs. The model returns values from a schema (enums, booleans, numbers), not free text. There is nothing to parse and no malformed JSON to retry.
  • Parallel decoding. Answers are produced in one forward pass instead of token by token, so latency stops growing with the length of the answer.
  • Calibrated confidence. Every answer comes with a probability, so you can auto-approve at 0.98 and send 0.6 to a human.
  • Different training target. TypeSafe trains with what they call RLCD (Reinforcement Learning for Calibrated Decisions) rather than optimizing for human-preferred prose.

What TypeSafe reports for Jev

These are vendor-reported numbers, so treat them as a claim to verify on your own data:

Reported
End-to-end latency 70–500 ms (vs. 3–329 s for LLMs on the same tasks)
Speedup 40×–200× on System One tasks
Input price $0.042 per 1M tokens, output free
Type errors Zero, by construction

Intended uses are routing, scoring, classification, extraction, guardrails and conditional logic inside workflows.

Kev-4B: the same shape, open and small

The most useful part of the story for engineers is that you can look inside one of these. jaredpalmer/kev-4b on Hugging Face is an Apache-2.0 adapter that takes a document plus a set of typed questions and returns a probability distribution over the options for each question, in a single forward pass with no text generated.

Under the hood:

  • Base model: Qwen3.5-4B-Base, a hybrid of 24 Gated DeltaNet (linear attention) layers and 8 full-attention layers.
  • Adapter: LoRA rank 16 (about 33.8M trainable parameters) plus a pointer head that does the classification.
  • Trained on public decision records, policy minimal-pairs and programmatically generated rule structures, then fine-tuned on harder policy and developer-tooling data.

Reported results include 0.803 accuracy on its hard-v1 test, 0.756 on devtools-v1, and 0.838 accuracy with a Brier score of 0.224 on a locked out-of-domain set. The number we find most interesting is calibration: only 0.9% of predictions made at ≥0.9 confidence were wrong on that out-of-domain data. For a classifier, knowing when it doesn't know is the feature.

Span-01: reasoning classifiers for observability

Respan's Span-01 applies the approach to evaluating AI systems. LLM-as-judge is accurate but slow and costly, so most teams sample a small fraction of production traces. Span-01 returns present, absent or not_observable probabilities for many behaviors at once ("hyper-parallel definition branches") in one forward pass, and it takes behavior definitions it hasn't seen before.

Respan reports an overall behavior-detection F1 of 84.3, against 81.5 for OpenAI's GPT-6 Luna and 71.5 for Jev 1.13.0. Per-domain F1 ranges from 0.779 (jailbreak and prompt injection) to 1.000 (privacy and secrets). Pricing is $0.02 per 1M input tokens with free output, plus a free Lite tier. Note that this is Respan's own benchmark of a model built for exactly that task, and it shows a generalist System One model losing to a specialist, which is what you would expect.

Why this matters

Our view, beyond what the three announcements say:

  1. Separate "deciding" from "writing." Generation is the expensive, slow, non-deterministic part. If your step only needs a label, paying for generation is waste.
  2. Probabilities are a control surface. Thresholds, human-review queues and A/B rollouts all become simple comparisons instead of prompt tweaks.
  3. Cost changes what you can do. At cents per million tokens and sub-second latency, you can classify every request, every trace, every row, rather than a sample.
  4. They are narrow by design. These models don't replace your LLM. They sit around it as routers, verifiers and guardrails, and the LLM still writes the answer.

Caveats

  • Nearly all numbers above come from the vendors' own benchmarks. Run your own labeled set before adopting one.
  • Calibration can degrade out of domain, so keep monitoring it.
  • Typed outputs guarantee the format, not that the decision is right. Zero type errors is not zero wrong answers.
  • Both Jev and Span-01 are hosted, closed models. Kev-4B is the option if you need to self-host.

Where this is going

For comparing AI services and models, which is what we do here, this suggests a new axis. Alongside "which model is smartest" and "which is cheapest per token," we should be asking what it costs and how fast it is to make a decision with each. We'll be tracking this category as it develops.

Sources: TypeSafe AI, "Introducing System One Models and Jev" · jaredpalmer/kev-4b on Hugging Face · Respan, "Introducing Span-01"