Jev¶
Summary
Jev is the first "System One" model from TypeSafe AI, a startup co-founded by CEO Diogo Almeida (ex-OpenAI) that emerged from stealth on 2026-09-15 with a $40M seed. Jev cannot chat, write, or explain, and that is the product. You POST unstructured state plus typed questions, and it returns typed decisions with calibrated probabilities in one parallel round trip. TypeSafe quotes 70-500 ms, at $0.042 per million input tokens with output free.
Three primitives cover everything. choice picks from up to 255 options and returns the full probability distribution. score places the input on an ordered rubric of up to 10 levels. noul is a calibrated yes/no from 0 to 1.
The bet: most LLM calls in software are small judgments (route this, score that, gate this). Those should be bounded decisions with explicit uncertainty, not parsed prose.
Fast-moving (checked 2026-09-25)
Jev launched ten days ago. Direct signups opened to everyone on 2026-09-20 and were paused on 2026-09-22 under demand. Existing accounts work, and Vercel AI Gateway, OpenRouter and Cloudflare Workers AI offer Jev without a TypeSafe account. Re-verify versions, limits and access before relying on this page.
Key Facts¶
| Fact | Value |
|---|---|
| Vendor | TypeSafe AI (founders Diogo Almeida, Erik Gafni, Sasha Sheng); $40M seed led by DCVC, reported ~$200M valuation |
| Latest Version | jev-1.13.0 (2026-09-15); jev-latest and jev-preview aliases both resolve to it |
| Availability | Hosted API only, closed weights. Direct signups paused since 2026-09-22; gateways open. |
| API | POST https://api.typesafe.ai/v1/systemone, GET /v1/models |
| Pricing | $0.042/MTok input, output free; $5 signup credit (while signups were open) |
| Latency | 70-500 ms end to end (vendor); ~150-500 ms observed by integrators |
| Limits | ~32K tokens for state + longest question, ~64K total; 255 choice options; 10 score levels; text/JSON input only |
| Official SDKs | typesafe-sdk 0.7.1 (Python, 2026-09-21), @typesafe-ai/sdk 0.6.0 (TS, 2026-09-15), both MIT |
| Gateways | Vercel typesafe-ai/jev, OpenRouter typesafe/jev-1.13 (Decisions API, alpha), Cloudflare typesafe/jev |
| Framework support | Pydantic AI TypeSafeModel, LangChain langchain-typesafe, AI SDK 7 experimental_evaluate, MCP (@jkudish/jev-mcp) |
| Self-hosted analogue | AnyJev (Nokia Applied Research, Apache-2.0) over open LLMs |
How It Fits¶
This diagram shows where Jev sits in a program: the LLM or your code produces state, Jev returns typed answers, and your code owns the branch.
flowchart LR
S["state<br/>ticket, tool call, page, trace"] --> J["Jev via api.typesafe.ai<br/>or a gateway"]
Q["typed questions<br/>noul / choice / score"] --> J
J -->|"probabilities + confidence"| C{"your code<br/>thresholds"}
C -->|"confident, low stakes"| ACT["act"]
C -->|"unsure"| LLM["escalate to an LLM"]
C -->|"low confidence or high stakes"| HUM["human or hard enforcement"]
For the full request sequence, the component map, and the abstention flowchart, see Explanation.
Evaluation¶
- Why it's better: Jev replaces a generative LLM call plus a parse-and-validate loop for closed-outcome judgments.
- It is type-safe by construction: output can't break the declared schema. That is a structural guarantee, not an empirical one.
- Its answers carry calibrated confidence for abstention policies.
- Its latency and cost fall in the range where high-volume and real-time uses become practical.
- An independent SREGym-Lite evaluation showed +8 points of agent reliability (40% -> 48% pass rate) using Jev as a plan ranker and submission gate.
- When it fits: closed outcome sets a human could label quickly. Examples: classification and routing, scoring and ranking, guardrails and gates, retrieval relevance, corpus-scale labeling, and real-time action selection over text-represented state. The sweet spot is high frequency, low consequence per call, and fuzzy semantics that a hand-written
ifcan't express. - When it does not fit:
- Anything generative: no prose, code, or summaries.
- Multi-step hidden reasoning: split it into atomic questions or use a reasoning model.
- Arithmetic, counting, date comparison, and exact string operations: do them in code.
- Extracting unknown values: Jev only chooses among candidates you supply.
- Irrelevant context degrades accuracy, so send minimal state.
| Pros | Cons |
|---|---|
| Type errors impossible by construction; answers carry calibrated probabilities and confidence | Closed weights, hosted only; direct signups paused since 2026-09-22 |
| 70-500 ms end to end; $0.042/MTok input, output free | Cannot generate; every use must decompose into choice/score/noul questions |
| All questions in one round trip, evaluated in parallel and in isolation (no context rot) | "Cannot hallucinate" ≠ always right: it can pick a wrong valid option with high confidence |
| Code owns the branching: thresholds, escalations, and abstention live in your program | Prompt injection can shift its probabilities (TypeSafe's own limitations page; arXiv 2609.28613) |
| First-class support in Pydantic AI, LangChain, AI SDK, OpenRouter, Cloudflare within ten days | Calibration is self-reported and must be re-verified on your traffic; no model changelog |
| Four access routes, three of them without a TypeSafe account | Naming traps: jevtypesafeai.com (reseller), typesafe-ai (PyPI shim), jev-mcp (different npm project) |
- Common use cases:
- Support triage and ticket routing, content moderation gates.
- LLM guardrails: prompt-injection screening, and tool-call risk classification before execution.
- Model routing for agents (fast vs frontier), retrieval relevance filtering ahead of generation.
- Resume and lead scoring, corpus-scale semantic enrichment.
- SRE agent decision support (SREGym's jev_plan/jev_submit pattern).
- Licensing & commercial use: closed weights, hosted API. Pay per token: $0.042/MTok input, output free, confirmed by TypeSafe's blog, the OpenRouter listing, Vercel AI Gateway and independent write-ups. TypeSafe itself cautions it "can't prove it isn't subsidized". The unofficial site jevtypesafeai.com resells keys at $0.25-$0.42/M; that is not official pricing.
- Ecosystem & data connections: Jev's integrations fall into five groups. Versions and dates are in Reference: SDKs and Integrations.
- Official SDKs:
typesafe-sdkand@typesafe-ai/sdk, plus TypeSafe's LLM-baseline adaptersystem-one-adapter. - Frameworks: Pydantic AI
TypeSafeModel(typesafe:jev-latest); LangChainTypeSafeClassifierwith experimentalModelRouterMiddleware/AutoModeMiddleware; AI SDK 7experimental_evaluate. - Gateways: Vercel AI Gateway, the OpenRouter Decisions API, and Cloudflare Workers AI.
- Community:
@jkudish/jev-mcp(eleven MCP tools),awesome-jev(curated list; it warns that a listing is not an endorsement),pi-jevmodel routers, and SREGym's optional Jev decision support. - Self-hosted: AnyJev (Nokia Applied Research, Apache-2.0) turns open LLMs into Jev-style deciders. Simpler clones (
von,mini-jev,verdict,go-system-one) mimic the interface without the calibration work.
- Official SDKs:
- Compatibility & requirements: an HTTPS API. State is a string, JSON object, or JSON array (text only). The Python SDK needs Python 3.10+ and the TS SDK Node 20+.
- Latest versions: model
jev-1.13.0(2026-09-15) is the only published version; third-party tools reference an earlierjev-1.12. TypeSafe publishes per-version "jaggedness" pages but no model changelog. Responses report the versioned ID, so log it and pin it. - Alternatives:
- LLM structured-output / JSON-schema modes: flexible and generative, but no calibration guarantee.
- Classic ML classifiers: cheap, but they need labeled training data per task.
- Embedding similarity: finds related text but doesn't decide relevance.
- AnyJev: self-hosted and Apache-2.0. With 100-300 labels per question, its L2 heads on Qwen3-4B to 32B score 0.771-0.799 on LocalLLaMA/typed-decisions, against Jev's published 0.727. The label-free L0 and L1 levels trail it (Qwen3-32B L1: 0.699).
- Simpler raw-logit clones.
- Migration & lock-in risks: questions and criteria are program logic, so version them like code for replay. Lock-in is real but softer than it looks. The interface (state + typed questions + distributions) is reproduced by AnyJev over open LLMs and adapted by the AI SDK's provider-neutral
experimental_evaluate, and three gateways serve the same model. Pin model versions; thresholds tuned on one version may not transfer. Cloudflare serves only a moving alias. - Community health: fast-growing and contested.
- TypeSafe ships unusually honest nuance: subsidy caveats, eval-bias disclosures, and per-version jaggedness pages.
- Coverage spans the Latent Space founder interview, independent deep dives (flaviocopes, Daily Dose of DS), major-framework integrations within a week, and security research on prompt injection within two weeks.
- The signup pause signals demand outrunning capacity.
- Watch: independent calibration studies are still thin.
FAQ¶
- Is Jev just a small LLM? TypeSafe says no. It describes a new architecture with a parallel sampler that emits all answers in one hardware-aware query, trained with RLCD rather than RLHF/RLVR. The FAQ on the announcement post rebuts the "smaller LLM" framing directly. The architecture is not published in detail, so this remains a vendor claim.
- Can it replace my LLM calls? No. It replaces the judgment-shaped subset: closed-outcome questions that need a calibrated answer. Anything generative, arithmetic, or multi-step stays with your LLM and your code. The architectural trade (DDDS): the LLM plans and writes; Jev routes, gates, checks, and escalates.
- Why "Jev"? After William Stanley Jevons and the Jevons paradox: cheaper intelligence, like cheaper steam-engine coal, increases total demand. The model-class name comes from Kahneman's System 1 / System 2 distinction.
- How do I get access while signups are paused? Use Vercel AI Gateway (
typesafe-ai/jev), OpenRouter's Decisions API (typesafe/jev-1.13), or Cloudflare Workers AI (typesafe/jev). See How-to Guides: Getting Access. - Can I run it locally or offline? Not Jev itself: its weights are closed and it is hosted only. For a self-hosted equivalent, AnyJev reproduces the interface over open LLMs with debiasing and calibration. Its labelled L2 heads beat Jev's published accuracy on one benchmark, at the cost of labels per question. The simpler clones mimic the API shape without calibration.
Topic Map¶
- How-to Guides: access routes, curl/SDK/gateway/framework recipes, patterns (routing, guardrails, SRE gates), rollout procedure, AnyJev self-hosting, troubleshooting
- Reference: endpoints, request/response schema, limits, model IDs, access routes, pricing, SDK matrix, timeline, source discrepancies
- Explanation: System One thesis, RLCD, primitives, architecture diagrams, evidence, AnyJev, threat model, agent-loop placement
Related Topics¶
- LLM Fundamentals: structured outputs and inference context
- Zero Data Retention: the retention terms to demand before sending regulated state
- OpenClaw and Hermes Agent: agent harnesses where Jev-style tool gating and routing apply
- LLM Wiki: the agent-maintained knowledge pattern (same domain)
- DFlash2: non-autoregressive speedups on the generative side of inference (different domain)
- Jev vs AnyJev vs Structured Outputs: decision guide between hosted Jev, the open AnyJev heads, and LLM structured outputs
- Unified Tools Catalogue: registered under Inference, Serving & Fine-tuning
Sources¶
URLs below appeared in search results or were fetched on 2026-09-25. Many hosts are unreachable from the research environment, so some were verified by search only.
- Introducing System One Models & Jev (TypeSafe AI, official): RLCD, pricing, evidence, nuance disclosures
- TypeSafe docs: primitives, confidence, patterns, models, API reference
- typesafe-sdk-python (PyPI) and @typesafe-ai/sdk: official SDKs; wire schema generated from the OpenAPI spec
- system-one-adapter-python (TypeSafe): LLM-backed drop-in for workflow-eval baselines
- TypeSafe on X: available to everyone and signups paused
- AIwire: TypeSafe AI emerges from stealth with $40M and Forbes: $200M valuation
- Jev (AI model), Wikipedia
- Pydantic AI: TypeSafe (Jev): limits, aliases, jaggedness, integrator observations
- OpenRouter: Jev 1.13 and OpenRouter Jev guide
- Cloudflare Workers AI: Jev
- Vercel: What is Jev: AI Gateway route
typesafe-ai/jev - langchain-typesafe (PyPI): LangChain integration
- Jev, clearly explained (Daily Dose of DS): independent technical walkthrough, abstention policies, failure modes, rollout procedure
- Jev: System One Models for Prod, Not God (Latent Space): founder interview on RLCD vs RLHF/RLVR, Jevons naming, calibration vision
- Jev deep dive (flaviocopes) and getting an API key
- Jev + SREGym-Lite (sregym.com) and SREGym on GitHub: independent agent-reliability experiment (MIT)
- Decision Hijacking (arXiv 2609.28613), Check Point analysis, VentureBeat: prompt-injection research
- AnyJev (Nokia Applied Research): open-source Jev-style decisions over open LLMs (Apache-2.0); levels, results_exit; PyPI
- jev-mcp (jkudish): Jev as MCP tools, multi-provider
- awesome-jev (yibie): curated ecosystem list
- agent-beacon (Asymptote Labs): cross-harness agent memory layer (MIT); earlier reported Jev use
- evals.typesafe.ai: workflow evals site
- von (wfzyx) and mini-jev (r-ms): local, non-Jev reimplementations of the interface
- jevtypesafeai.com: UNOFFICIAL community/demo and reseller site (naming trap; $0.25-$0.42/M is reseller pricing)
Questions¶
- Does RLCD-trained calibration hold on third-party traffic at scale? So far there are only self-reported numbers, one SREGym experiment and one published accuracy figure (0.727).
- When will direct signups reopen, and will TypeSafe publish an SLA? The 2026-09-22 post gives no date. (Rate limits are now published: 250,000 tokens/s and 1,200 requests/min, adjusted dynamically, per the models page.)
- What are each gateway's (Vercel, OpenRouter, Cloudflare) data-retention terms for Jev
statepayloads? TypeSafe's own policy publishes no fixed retention window and offers ZDR to enterprise customers only (see Explanation). - Does OpenRouter proxy to TypeSafe or host the model itself? Will Cloudflare offer pinned versions?
- Can TypeSafe's promised robustness improvements close the prompt-injection gap shown by arXiv 2609.28613?
- How does Jev behave as a prospective safety reviewer for destructive agent actions? (SREGym lists it as future work it has not run)
- Will TypeSafe publish a model changelog or deprecation policy for pinned versions like
jev-1.13.0? None was found as of 2026-09-27.