Skip to content

Jev

Summary

Jev is the first "System One" model from TypeSafe AI, a startup co-founded by CEO Diogo Almeida (ex-OpenAI) that emerged from stealth on 2026-09-15 with a $40M seed. Jev cannot chat, write, or explain, and that is the product. You POST unstructured state plus typed questions, and it returns typed decisions with calibrated probabilities in one parallel round trip. TypeSafe quotes 70-500 ms, at $0.042 per million input tokens with output free.

Three primitives cover everything. choice picks from up to 255 options and returns the full probability distribution. score places the input on an ordered rubric of up to 10 levels. noul is a calibrated yes/no from 0 to 1.

The bet: most LLM calls in software are small judgments (route this, score that, gate this). Those should be bounded decisions with explicit uncertainty, not parsed prose.

Fast-moving (checked 2026-09-25)

Jev launched ten days ago. Direct signups opened to everyone on 2026-09-20 and were paused on 2026-09-22 under demand. Existing accounts work, and Vercel AI Gateway, OpenRouter and Cloudflare Workers AI offer Jev without a TypeSafe account. Re-verify versions, limits and access before relying on this page.

Key Facts

Fact Value
Vendor TypeSafe AI (founders Diogo Almeida, Erik Gafni, Sasha Sheng); $40M seed led by DCVC, reported ~$200M valuation
Latest Version jev-1.13.0 (2026-09-15); jev-latest and jev-preview aliases both resolve to it
Availability Hosted API only, closed weights. Direct signups paused since 2026-09-22; gateways open.
API POST https://api.typesafe.ai/v1/systemone, GET /v1/models
Pricing $0.042/MTok input, output free; $5 signup credit (while signups were open)
Latency 70-500 ms end to end (vendor); ~150-500 ms observed by integrators
Limits ~32K tokens for state + longest question, ~64K total; 255 choice options; 10 score levels; text/JSON input only
Official SDKs typesafe-sdk 0.7.1 (Python, 2026-09-21), @typesafe-ai/sdk 0.6.0 (TS, 2026-09-15), both MIT
Gateways Vercel typesafe-ai/jev, OpenRouter typesafe/jev-1.13 (Decisions API, alpha), Cloudflare typesafe/jev
Framework support Pydantic AI TypeSafeModel, LangChain langchain-typesafe, AI SDK 7 experimental_evaluate, MCP (@jkudish/jev-mcp)
Self-hosted analogue AnyJev (Nokia Applied Research, Apache-2.0) over open LLMs

How It Fits

This diagram shows where Jev sits in a program: the LLM or your code produces state, Jev returns typed answers, and your code owns the branch.

flowchart LR
    S["state<br/>ticket, tool call, page, trace"] --> J["Jev via api.typesafe.ai<br/>or a gateway"]
    Q["typed questions<br/>noul / choice / score"] --> J
    J -->|"probabilities + confidence"| C{"your code<br/>thresholds"}
    C -->|"confident, low stakes"| ACT["act"]
    C -->|"unsure"| LLM["escalate to an LLM"]
    C -->|"low confidence or high stakes"| HUM["human or hard enforcement"]

For the full request sequence, the component map, and the abstention flowchart, see Explanation.

Evaluation

  • Why it's better: Jev replaces a generative LLM call plus a parse-and-validate loop for closed-outcome judgments.
    • It is type-safe by construction: output can't break the declared schema. That is a structural guarantee, not an empirical one.
    • Its answers carry calibrated confidence for abstention policies.
    • Its latency and cost fall in the range where high-volume and real-time uses become practical.
    • An independent SREGym-Lite evaluation showed +8 points of agent reliability (40% -> 48% pass rate) using Jev as a plan ranker and submission gate.
  • When it fits: closed outcome sets a human could label quickly. Examples: classification and routing, scoring and ranking, guardrails and gates, retrieval relevance, corpus-scale labeling, and real-time action selection over text-represented state. The sweet spot is high frequency, low consequence per call, and fuzzy semantics that a hand-written if can't express.
  • When it does not fit:
    • Anything generative: no prose, code, or summaries.
    • Multi-step hidden reasoning: split it into atomic questions or use a reasoning model.
    • Arithmetic, counting, date comparison, and exact string operations: do them in code.
    • Extracting unknown values: Jev only chooses among candidates you supply.
    • Irrelevant context degrades accuracy, so send minimal state.
Pros Cons
Type errors impossible by construction; answers carry calibrated probabilities and confidence Closed weights, hosted only; direct signups paused since 2026-09-22
70-500 ms end to end; $0.042/MTok input, output free Cannot generate; every use must decompose into choice/score/noul questions
All questions in one round trip, evaluated in parallel and in isolation (no context rot) "Cannot hallucinate" ≠ always right: it can pick a wrong valid option with high confidence
Code owns the branching: thresholds, escalations, and abstention live in your program Prompt injection can shift its probabilities (TypeSafe's own limitations page; arXiv 2609.28613)
First-class support in Pydantic AI, LangChain, AI SDK, OpenRouter, Cloudflare within ten days Calibration is self-reported and must be re-verified on your traffic; no model changelog
Four access routes, three of them without a TypeSafe account Naming traps: jevtypesafeai.com (reseller), typesafe-ai (PyPI shim), jev-mcp (different npm project)
  • Common use cases:
    • Support triage and ticket routing, content moderation gates.
    • LLM guardrails: prompt-injection screening, and tool-call risk classification before execution.
    • Model routing for agents (fast vs frontier), retrieval relevance filtering ahead of generation.
    • Resume and lead scoring, corpus-scale semantic enrichment.
    • SRE agent decision support (SREGym's jev_plan/jev_submit pattern).
  • Licensing & commercial use: closed weights, hosted API. Pay per token: $0.042/MTok input, output free, confirmed by TypeSafe's blog, the OpenRouter listing, Vercel AI Gateway and independent write-ups. TypeSafe itself cautions it "can't prove it isn't subsidized". The unofficial site jevtypesafeai.com resells keys at $0.25-$0.42/M; that is not official pricing.
  • Ecosystem & data connections: Jev's integrations fall into five groups. Versions and dates are in Reference: SDKs and Integrations.
    • Official SDKs: typesafe-sdk and @typesafe-ai/sdk, plus TypeSafe's LLM-baseline adapter system-one-adapter.
    • Frameworks: Pydantic AI TypeSafeModel (typesafe:jev-latest); LangChain TypeSafeClassifier with experimental ModelRouterMiddleware / AutoModeMiddleware; AI SDK 7 experimental_evaluate.
    • Gateways: Vercel AI Gateway, the OpenRouter Decisions API, and Cloudflare Workers AI.
    • Community: @jkudish/jev-mcp (eleven MCP tools), awesome-jev (curated list; it warns that a listing is not an endorsement), pi-jev model routers, and SREGym's optional Jev decision support.
    • Self-hosted: AnyJev (Nokia Applied Research, Apache-2.0) turns open LLMs into Jev-style deciders. Simpler clones (von, mini-jev, verdict, go-system-one) mimic the interface without the calibration work.
  • Compatibility & requirements: an HTTPS API. State is a string, JSON object, or JSON array (text only). The Python SDK needs Python 3.10+ and the TS SDK Node 20+.
  • Latest versions: model jev-1.13.0 (2026-09-15) is the only published version; third-party tools reference an earlier jev-1.12. TypeSafe publishes per-version "jaggedness" pages but no model changelog. Responses report the versioned ID, so log it and pin it.
  • Alternatives:
    • LLM structured-output / JSON-schema modes: flexible and generative, but no calibration guarantee.
    • Classic ML classifiers: cheap, but they need labeled training data per task.
    • Embedding similarity: finds related text but doesn't decide relevance.
    • AnyJev: self-hosted and Apache-2.0. With 100-300 labels per question, its L2 heads on Qwen3-4B to 32B score 0.771-0.799 on LocalLLaMA/typed-decisions, against Jev's published 0.727. The label-free L0 and L1 levels trail it (Qwen3-32B L1: 0.699).
    • Simpler raw-logit clones.
  • Migration & lock-in risks: questions and criteria are program logic, so version them like code for replay. Lock-in is real but softer than it looks. The interface (state + typed questions + distributions) is reproduced by AnyJev over open LLMs and adapted by the AI SDK's provider-neutral experimental_evaluate, and three gateways serve the same model. Pin model versions; thresholds tuned on one version may not transfer. Cloudflare serves only a moving alias.
  • Community health: fast-growing and contested.
    • TypeSafe ships unusually honest nuance: subsidy caveats, eval-bias disclosures, and per-version jaggedness pages.
    • Coverage spans the Latent Space founder interview, independent deep dives (flaviocopes, Daily Dose of DS), major-framework integrations within a week, and security research on prompt injection within two weeks.
    • The signup pause signals demand outrunning capacity.
    • Watch: independent calibration studies are still thin.

FAQ

  • Is Jev just a small LLM? TypeSafe says no. It describes a new architecture with a parallel sampler that emits all answers in one hardware-aware query, trained with RLCD rather than RLHF/RLVR. The FAQ on the announcement post rebuts the "smaller LLM" framing directly. The architecture is not published in detail, so this remains a vendor claim.
  • Can it replace my LLM calls? No. It replaces the judgment-shaped subset: closed-outcome questions that need a calibrated answer. Anything generative, arithmetic, or multi-step stays with your LLM and your code. The architectural trade (DDDS): the LLM plans and writes; Jev routes, gates, checks, and escalates.
  • Why "Jev"? After William Stanley Jevons and the Jevons paradox: cheaper intelligence, like cheaper steam-engine coal, increases total demand. The model-class name comes from Kahneman's System 1 / System 2 distinction.
  • How do I get access while signups are paused? Use Vercel AI Gateway (typesafe-ai/jev), OpenRouter's Decisions API (typesafe/jev-1.13), or Cloudflare Workers AI (typesafe/jev). See How-to Guides: Getting Access.
  • Can I run it locally or offline? Not Jev itself: its weights are closed and it is hosted only. For a self-hosted equivalent, AnyJev reproduces the interface over open LLMs with debiasing and calibration. Its labelled L2 heads beat Jev's published accuracy on one benchmark, at the cost of labels per question. The simpler clones mimic the API shape without calibration.

Topic Map

  • How-to Guides: access routes, curl/SDK/gateway/framework recipes, patterns (routing, guardrails, SRE gates), rollout procedure, AnyJev self-hosting, troubleshooting
  • Reference: endpoints, request/response schema, limits, model IDs, access routes, pricing, SDK matrix, timeline, source discrepancies
  • Explanation: System One thesis, RLCD, primitives, architecture diagrams, evidence, AnyJev, threat model, agent-loop placement

Sources

URLs below appeared in search results or were fetched on 2026-09-25. Many hosts are unreachable from the research environment, so some were verified by search only.

Questions

  • Does RLCD-trained calibration hold on third-party traffic at scale? So far there are only self-reported numbers, one SREGym experiment and one published accuracy figure (0.727).
  • When will direct signups reopen, and will TypeSafe publish an SLA? The 2026-09-22 post gives no date. (Rate limits are now published: 250,000 tokens/s and 1,200 requests/min, adjusted dynamically, per the models page.)
  • What are each gateway's (Vercel, OpenRouter, Cloudflare) data-retention terms for Jev state payloads? TypeSafe's own policy publishes no fixed retention window and offers ZDR to enterprise customers only (see Explanation).
  • Does OpenRouter proxy to TypeSafe or host the model itself? Will Cloudflare offer pinned versions?
  • Can TypeSafe's promised robustness improvements close the prompt-injection gap shown by arXiv 2609.28613?
  • How does Jev behave as a prospective safety reviewer for destructive agent actions? (SREGym lists it as future work it has not run)
  • Will TypeSafe publish a model changelog or deprecation policy for pinned versions like jev-1.13.0? None was found as of 2026-09-27.