Blog Posts

First open-source model GPT-OSS by OpenAI: AI pricing and API availability

GPT-OSS is the first open-weight AI model from OpenAI, with high performance capabilities and reasoning. The models aim to provide enterprise-grade AI functionality that can run efficiently on consumer hardware or single GPUs, making top AI industry expertise from the OpenAI Team accessible to developers, researchers, and organizations of all sizes. ### **Open Source Benefits** - **License**: Apache 2.0 - fully commercial-friendly - **Cost**: $0 for model weights and local inference - **Customization**: Full model modification and fine-tuning rights - **Distribution**: Can be redistributed and integrated into commercial products ## **Pricing & Availability** ### Web Interface OpenAI provides a dedicated web interface where you can test GPT-OSS models at no cost. The platform offers immediate access to both 20B and 120B models without requiring API keys or technical setup. **Features:** - Free access to all GPT-OSS model variants - No registration or API keys required - Interactive chat interface with adjustable parameters - Real-time model comparison capabilities **Try it now:** [gpt-oss.com](https://gpt-oss.com/) ### **Cloud Inference Pricing (API)** #### **GPT-OSS 20B Pricing** | Provider | Inference | Context | Input Price | Output Price | Throughput | |----------|-------|---------|-------------|--------------|------------| | OpenRouter | Fireworks | 131K | $0.05 | $0.20 | 456.0tps | | OpenRouter | NovitaAI | 131K | $0.05 | $0.20 | 268.8tps | | OpenRouter | Groq | 131K | $0.10 | $0.50 | 10,850tps | | Cloudflare AI | @cf/openai/gpt-oss-20b | - | $0.20 | $0.30 | - | #### **GPT-OSS 120B Pricing** | Provider | Inference | Context | Input Price | Output Price | Throughput | |----------|-------|---------|-------------|--------------|------------| | OpenRouter | Baseten | 131K | $0.10 | $0.50 | 997.4tps | | OpenRouter | NovitaAI | 131K | $0.10 | $0.50 | 95.02tps | | OpenRouter | Fireworks | 131K | $0.15 | $0.60 | 233.3tps | | OpenRouter | Together | 131K | $0.15 | $0.60 | 177.7tps | | OpenRouter | Parasail | 131K | $0.15 | $0.60 | 102.9tps | | OpenRouter | Groq | 131K | $0.15 | $0.75 | 954.8tps | | OpenRouter | Cerebras | 131K | $0.25 | $0.69 | 3,512tps | | Cloudflare AI | @cf/openai/gpt-oss-120b | - | $0.35 | $0.75 | - | ## **Use Cases** ### **1. Enterprise AI on Private Infrastructure** - Complete data privacy, no API costs, full control - Deploy 120B model on single H100 GPU - Enterprise-grade reasoning with zero data leaving premises ### **2. Local Development and Prototyping** - API rate limits, offline development, cost predictability - 20B model running on 16GB consumer hardware - Rapid prototyping with production-ready AI capabilities --- **References:** 1. [Introducing GPT-OSS](https://openai.com/index/introducing-gpt-oss/) - OpenAI Official Announcement 2. [Welcome OpenAI GPT-OSS](https://huggingface.co/blog/welcome-openai-gpt-oss) - Hugging Face Technical Deep Dive 3. [GPT-OSS 20B on OpenRouter](https://openrouter.ai/openai/gpt-oss-20b) - Pricing and Specifications 4. [GPT-OSS Web Interface](https://gpt-oss.com/) - Free Web Interface 5. [GPT-OSS on Ollama](https://ollama.com/library/gpt-oss) - Local Installation **Tags**: #OpenAI #GPT-OSS #OpenSource #AI #MoE #Reasoning #LocalAI #Enterprise

Vitalii•

Jev, and a New Type of AI Processing with Reasoning Classifiers

Most of what we do with LLMs in production is not chat. It is a decision: *is this ticket urgent? Does this answer follow from the source? Should this request go to the expensive model?* We ask a model that generates paragraphs to produce one word, then parse that word back out of a string. A new class of models is questioning that pattern. Three releases in the last few weeks point the same way: **Jev** from TypeSafe AI, the open **Kev-4B** adapter, and **Span-01** from Respan. Each one is a *reasoning classifier*, meaning a model that reads a state and returns calibrated probabilities instead of text. ## The idea: "System One" models TypeSafe describes Jev as its first *System One Model*, a model built for fast, structured decisions inside software rather than for conversation. Their pitch, in the founder's words: "Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out." What makes this different from a normal LLM call: - **Typed outputs.** The model returns values from a schema (enums, booleans, numbers), not free text. There is nothing to parse and no malformed JSON to retry. - **Parallel decoding.** Answers are produced in one forward pass instead of token by token, so latency stops growing with the length of the answer. - **Calibrated confidence.** Every answer comes with a probability, so you can auto-approve at 0.98 and send 0.6 to a human. - **Different training target.** TypeSafe trains with what they call RLCD (Reinforcement Learning for Calibrated Decisions) rather than optimizing for human-preferred prose. ## What TypeSafe reports for Jev These are vendor-reported numbers, so treat them as a claim to verify on your own data: | | Reported | |---|---| | End-to-end latency | 70–500 ms (vs. 3–329 s for LLMs on the same tasks) | | Speedup | 40×–200× on System One tasks | | Input price | $0.042 per 1M tokens, output free | | Type errors | Zero, by construction | Intended uses are routing, scoring, classification, extraction, guardrails and conditional logic inside workflows. ## Kev-4B: the same shape, open and small The most useful part of the story for engineers is that you can look inside one of these. `jaredpalmer/kev-4b` on Hugging Face is an Apache-2.0 adapter that takes a document plus a set of typed questions and returns a probability distribution over the options for each question, in a single forward pass with no text generated. Under the hood: - Base model: Qwen3.5-4B-Base, a hybrid of 24 Gated DeltaNet (linear attention) layers and 8 full-attention layers. - Adapter: LoRA rank 16 (about 33.8M trainable parameters) plus a pointer head that does the classification. - Trained on public decision records, policy minimal-pairs and programmatically generated rule structures, then fine-tuned on harder policy and developer-tooling data. Reported results include 0.803 accuracy on its `hard-v1` test, 0.756 on `devtools-v1`, and 0.838 accuracy with a Brier score of 0.224 on a locked out-of-domain set. The number we find most interesting is calibration: only 0.9% of predictions made at ≥0.9 confidence were wrong on that out-of-domain data. For a classifier, knowing when it doesn't know is the feature. ## Span-01: reasoning classifiers for observability Respan's Span-01 applies the approach to evaluating AI systems. LLM-as-judge is accurate but slow and costly, so most teams sample a small fraction of production traces. Span-01 returns `present`, `absent` or `not_observable` probabilities for many behaviors at once ("hyper-parallel definition branches") in one forward pass, and it takes behavior definitions it hasn't seen before. Respan reports an overall behavior-detection F1 of 84.3, against 81.5 for OpenAI's GPT-6 Luna and 71.5 for Jev 1.13.0. Per-domain F1 ranges from 0.779 (jailbreak and prompt injection) to 1.000 (privacy and secrets). Pricing is $0.02 per 1M input tokens with free output, plus a free Lite tier. Note that this is Respan's own benchmark of a model built for exactly that task, and it shows a generalist System One model losing to a specialist, which is what you would expect. ## Why this matters Our view, beyond what the three announcements say: 1. **Separate "deciding" from "writing."** Generation is the expensive, slow, non-deterministic part. If your step only needs a label, paying for generation is waste. 2. **Probabilities are a control surface.** Thresholds, human-review queues and A/B rollouts all become simple comparisons instead of prompt tweaks. 3. **Cost changes what you can do.** At cents per million tokens and sub-second latency, you can classify every request, every trace, every row, rather than a sample. 4. **They are narrow by design.** These models don't replace your LLM. They sit around it as routers, verifiers and guardrails, and the LLM still writes the answer. ## Caveats - Nearly all numbers above come from the vendors' own benchmarks. Run your own labeled set before adopting one. - Calibration can degrade out of domain, so keep monitoring it. - Typed outputs guarantee the *format*, not that the *decision* is right. Zero type errors is not zero wrong answers. - Both Jev and Span-01 are hosted, closed models. Kev-4B is the option if you need to self-host. ## Where this is going For comparing AI services and models, which is what we do here, this suggests a new axis. Alongside "which model is smartest" and "which is cheapest per token," we should be asking what it costs and how fast it is to make a *decision* with each. We'll be tracking this category as it develops. **Sources:** [TypeSafe AI, "Introducing System One Models and Jev"](https://typesafe.ai/blog/introducing-system-one-models-and-jev) · [jaredpalmer/kev-4b on Hugging Face](https://huggingface.co/jaredpalmer/kev-4b) · [Respan, "Introducing Span-01"](https://www.respan.ai/blog/introducing-span-1)

Vitalii•