Jev, and a New Type of AI Processing with Reasoning Classifiers
Most of what we do with LLMs in production is not chat. It is a decision: is this ticket urgent? Does this answer follow from the source? Should this request go to the expensive model? We ask a model that generates paragraphs to produce one word, then parse that word back out of a string.
A new class of models is questioning that pattern. Three releases in the last few weeks point the same way: Jev from TypeSafe AI, the open Kev-4B adapter, and Span-01 from Respan. Each one is a reasoning classifier, meaning a model that reads a state and returns calibrated probabilities instead of text.
The idea: "System One" models
TypeSafe describes Jev as its first System One Model, a model built for fast, structured decisions inside software rather than for conversation. Their pitch, in the founder's words: "Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."
What makes this different from a normal LLM call:
- Typed outputs. The model returns values from a schema (enums, booleans, numbers), not free text. There is nothing to parse and no malformed JSON to retry.
- Parallel decoding. Answers are produced in one forward pass instead of token by token, so latency stops growing with the length of the answer.
- Calibrated confidence. Every answer comes with a probability, so you can auto-approve at 0.98 and send 0.6 to a human.
- Different training target. TypeSafe trains with what they call RLCD (Reinforcement Learning for Calibrated Decisions) rather than optimizing for human-preferred prose.
What TypeSafe reports for Jev
These are vendor-reported numbers, so treat them as a claim to verify on your own data:
| Reported | |
|---|---|
| End-to-end latency | 70–500 ms (vs. 3–329 s for LLMs on the same tasks) |
| Speedup | 40×–200× on System One tasks |
| Input price | $0.042 per 1M tokens, output free |
| Type errors | Zero, by construction |
Intended uses are routing, scoring, classification, extraction, guardrails and conditional logic inside workflows.
Kev-4B: the same shape, open and small
The most useful part of the story for engineers is that you can look inside one of these. jaredpalmer/kev-4b on Hugging Face is an Apache-2.0 adapter that takes a document plus a set of typed questions and returns a probability distribution over the options for each question, in a single forward pass with no text generated.
Under the hood:
- Base model: Qwen3.5-4B-Base, a hybrid of 24 Gated DeltaNet (linear attention) layers and 8 full-attention layers.
- Adapter: LoRA rank 16 (about 33.8M trainable parameters) plus a pointer head that does the classification.
- Trained on public decision records, policy minimal-pairs and programmatically generated rule structures, then fine-tuned on harder policy and developer-tooling data.
Reported results include 0.803 accuracy on its hard-v1 test, 0.756 on devtools-v1, and 0.838 accuracy with a Brier score of 0.224 on a locked out-of-domain set. The number we find most interesting is calibration: only 0.9% of predictions made at ≥0.9 confidence were wrong on that out-of-domain data. For a classifier, knowing when it doesn't know is the feature.
Span-01: reasoning classifiers for observability
Respan's Span-01 applies the approach to evaluating AI systems. LLM-as-judge is accurate but slow and costly, so most teams sample a small fraction of production traces. Span-01 returns present, absent or not_observable probabilities for many behaviors at once ("hyper-parallel definition branches") in one forward pass, and it takes behavior definitions it hasn't seen before.
Respan reports an overall behavior-detection F1 of 84.3, against 81.5 for OpenAI's GPT-6 Luna and 71.5 for Jev 1.13.0. Per-domain F1 ranges from 0.779 (jailbreak and prompt injection) to 1.000 (privacy and secrets). Pricing is $0.02 per 1M input tokens with free output, plus a free Lite tier. Note that this is Respan's own benchmark of a model built for exactly that task, and it shows a generalist System One model losing to a specialist, which is what you would expect.
Why this matters
Our view, beyond what the three announcements say:
- Separate "deciding" from "writing." Generation is the expensive, slow, non-deterministic part. If your step only needs a label, paying for generation is waste.
- Probabilities are a control surface. Thresholds, human-review queues and A/B rollouts all become simple comparisons instead of prompt tweaks.
- Cost changes what you can do. At cents per million tokens and sub-second latency, you can classify every request, every trace, every row, rather than a sample.
- They are narrow by design. These models don't replace your LLM. They sit around it as routers, verifiers and guardrails, and the LLM still writes the answer.
Caveats
- Nearly all numbers above come from the vendors' own benchmarks. Run your own labeled set before adopting one.
- Calibration can degrade out of domain, so keep monitoring it.
- Typed outputs guarantee the format, not that the decision is right. Zero type errors is not zero wrong answers.
- Both Jev and Span-01 are hosted, closed models. Kev-4B is the option if you need to self-host.
Where this is going
For comparing AI services and models, which is what we do here, this suggests a new axis. Alongside "which model is smartest" and "which is cheapest per token," we should be asking what it costs and how fast it is to make a decision with each. We'll be tracking this category as it develops.
Sources: TypeSafe AI, "Introducing System One Models and Jev" · jaredpalmer/kev-4b on Hugging Face · Respan, "Introducing Span-01"