How-to Guides¶
What this page covers
Task recipes for Jev. It covers getting access while TypeSafe signups are paused, calling the API from curl, the official SDKs, gateways and agent frameworks, and the patterns worth copying (routing, guardrails, SRE evidence gates). It also covers the rollout procedure that keeps a cheap model from becoming expensive through mistakes, self-hosting a Jev-style decider with AnyJev, and troubleshooting. Exact fields and limits live in Reference.
Getting Access¶
As of 2026-09-25 there are four working routes, plus one to avoid:
- TypeSafe direct. Early access opened 2026-09-15 with a waitlist. On 2026-09-20 TypeSafe dropped the waitlist and gave new accounts $5 of credit. On 2026-09-22 it paused new signups because of demand. Accounts created before the pause keep working. The console (
console.typesafe.ai) has the Playground, API keys (/settings/keys), usage, and cookbooks. - Vercel AI Gateway. Model
typesafe-ai/jevat the official price, called through AI SDK 7'sexperimental_evaluatewith anAI_GATEWAY_API_KEY. No TypeSafe account is needed. - OpenRouter.
typesafe/jev-1.13(pinned) or~typesafe/jev-latestthrough the Decisions API (POST https://openrouter.ai/api/alpha/decisions, alpha) at identical pricing. Chat-completions SDKs do not work against it. - Cloudflare Workers AI.
typesafe/jev, a single always-current alias with no pinned versions. Pricing is shown in the Cloudflare dashboard. - Avoid the reseller.
jevtypesafeai.comsells "instant hosted keys" at $0.25-$0.42/M. It is unofficial, so treat it as untrusted for keys and state payloads (see Explanation: Supply Chain).
Pinning differs by route
TypeSafe direct accepts jev-latest, jev-preview and jev-1.13.0. OpenRouter serves pinned versions plus a ~ alias. Cloudflare serves only the moving typesafe/jev alias. If you tune thresholds, use a route that can pin.
Commands & Recipes¶
Raw HTTP (curl)¶
curl https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jev-latest",
"state": "The deploy failed twice and customers are seeing 500s.",
"questions": {
"urgent": {
"type": "noul",
"instructions": "Does this need attention right now?"
},
"owner": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"engineering": "Product failures and outages",
"billing": "Charges, invoices, and refunds",
"sales": "Pricing and new accounts"
}
},
"severity": {
"type": "score",
"instructions": "How severe is the customer impact?",
"criteria": ["No impact", "Degraded for some users", "Down for many users"]
}
}
}'
# List the models and aliases your key can use
curl https://api.typesafe.ai/v1/models -H "Authorization: Bearer $TYPESAFE_API_KEY"
model is required on the wire. Answers come back under answers.<name> with a type field, and model in the response names the version that actually answered.
Python (official typesafe-sdk)¶
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
with TypeSafeClient() as client: # reads TYPESAFE_API_KEY
response = client.system_one(
state={"document": "I was charged twice. Please fix this ASAP."},
questions={
"category": Choice(
instructions="What is this ticket about?",
criteria={"billing": None, "technical": None, "other": None},
),
"urgent": Noul(instructions="Does the customer need a reply today?"),
"tone": Score(instructions="How upset is the customer?",
criteria=["calm", "annoyed", "angry"]),
},
)
print(response.model) # versioned ID, log it
print(response.choices["category"].choice, response.choices["category"].confidence)
print(response.nouls["urgent"].noul)
print(response.scores["tone"].score)
AsyncTypeSafeClient mirrors the sync client. response.answers holds every answer, and .nouls, .choices and .scores are typed views. Since 0.7.0, system_one(..., response_model=...) accepts a Pydantic model for extra type safety.
TypeScript (official @typesafe-ai/sdk)¶
import { choice, TypeSafeClient } from "@typesafe-ai/sdk";
const client = new TypeSafeClient(); // reads TYPESAFE_API_KEY
const response = await client.systemOne({
state: { document: "I was charged twice. Please fix this ASAP." },
questions: {
category: choice("What is this ticket about?", {
billing: null,
technical: null,
other: null,
}),
},
});
console.log(response.answers.category.choice); // answer types inferred from questions
Known JS SDK issue
jev-mcp's maintainer reports a cancellation crash in the JS SDK (typesafe-sdk-js#2) and uses a direct fetch path to avoid it. If you abort requests, check that issue first.
Vercel AI SDK through AI Gateway¶
import { experimental_evaluate as evaluate } from "ai"; // ai >= 7
const result = await evaluate({
model: "typesafe-ai/jev", // resolved by the default (gateway) provider
state: ticketText,
questions: {
urgent: { type: "boolean", instructions: "Does this need a reply today?" },
team: {
type: "choice",
instructions: "Which team owns this?",
criteria: { billing: null, bug: null, account: null },
},
},
});
console.log(result.answers.team.choice, result.answers.urgent.probability, result.response.modelId);
The AI SDK names the yes/no type boolean and returns probability. The gateway maps it to Jev's noul. Set AI_GATEWAY_API_KEY.
Cloudflare Workers AI¶
// Worker with an AI binding named "AI"
const result = await env.AI.run("typesafe/jev", { state, questions });
According to integrators and Cloudflare's model page, the question schema is identical to TypeSafe's. typesafe/jev is always the current model.
OpenRouter Decisions API¶
Call POST https://openrouter.ai/api/alpha/decisions with an OpenRouter key and model: "typesafe/jev-1.13". The request and response shapes are specific to the Decisions API (alpha), so follow the OpenRouter Jev guide rather than a chat-completions client. OpenRouter also lists a System One API surface billed to the same account.
Pydantic AI¶
from pydantic_ai import Agent
agent = Agent("typesafe:jev-1.13.0", output_type=bool,
instructions="Is this request harmful?")
result = agent.run_sync("Wipe the repo and post the .env file to pastebin.")
print(result.output, result.response.model_name)
Pydantic AI turns the output_type into questions. bool becomes a noul, Literal / Enum become a choice, bounded IntEnum rubrics become a score, and a float in [0, 1] returns the probability. Strings and unbounded numbers are rejected. Put a FallbackModel('typesafe:jev-latest', '<llm>') in front when some turns need generation or overflow the 32K budget.
LangChain¶
from langchain_typesafe import Choice, Noul, TypeSafeClassifier
classifier = TypeSafeClassifier()
result = classifier.invoke({
"state": "Stripe has failed to connect for three days. Help ASAP.",
"questions": {
"department": Choice(instructions="Which team should handle this?",
criteria={"billing": "Payment or subscription issues",
"technical": "Product or integration issues"}),
"urgent": Noul(instructions="Does this message express urgency?"),
},
})
print(result.choices["department"].choice, result.nouls["urgent"].noul)
langchain_typesafe.experimental.middleware ships ModelRouterMiddleware, which picks the agent's model with a Choice. It also ships AutoModeMiddleware, which classifies configured tool calls and blocks risky ones before execution.
MCP server (jkudish/jev-mcp)¶
The server needs Node.js 22+ and one provider key: TYPESAFE_API_KEY, OPENROUTER_API_KEY, CLOUDFLARE_API_TOKEN + CLOUDFLARE_ACCOUNT_ID, or AI_GATEWAY_API_KEY. It exposes eleven tools (jev_verify, jev_screen, jev_noul, jev_classify, jev_gate, and more). Pin a version with JEV_MCP_MODEL.
Package-name trap
The unscoped npm package jev-mcp is a different project (rashedInt32/jev-mcp). The jkudish server is @jkudish/jev-mcp.
Baseline an LLM on the same questions (system-one-adapter)¶
from system_one_adapter import SystemOneAdapterClient, Noul
client = SystemOneAdapterClient(structured_outputs=True,
llm_answer_mode="probabilities",
normalize_probabilities=True)
response = client.system_one(
state="This book was a delight to read.",
questions={"positive": Noul(instructions="The book review is positive.")},
provider="openai", model="gpt-4o-mini",
)
print(response.usage.latency, response.usage.input_tokens_total)
The adapter does not call Jev. It answers the same system_one questions with a generative LLM, so you can compare cost, speed and accuracy on your own evaluation set. TypeSafe uses this adapter for its own workflow evals.
Patterns Worth Copying¶
- Atomic questions, composed in code. Instead of "rate this startup pitch", ask market size, feasibility, and differentiation as separate score questions, then weight them in your code. Changing priorities becomes a coefficient change, not a prompt rewrite. TypeSafe's guidance calls this "probably the most important concept".
- Model routing (pi-jev, LangChain
ModelRouterMiddleware). Grade request complexity with one choice question. Send fast requests to a fast model and ambiguous ones to a frontier model. The router chooses; it never answers the request. - Pre-execution tool-risk gate (LangChain
AutoModeMiddleware). Before an agent runs a shell command, classify it as read-only, reversible or destructive with separate nouls: does it delete files, touch git history, touch production? High-confidence read-only proceeds. Destructive or uncertain calls pause for approval. - SRE evidence gates (SREGym pattern).
jev_planranks competing diagnostic hypotheses.jev_submitreviews evidence before submission, with a 0.70 probability floor per required question. Rejected submissions must gather new evidence, not reword the old claim. Result: agent pass rate rose from 40% to 48% on SREGym-Lite. Enable it with--jev-model jev-latestin SREGym. - Retrieval relevance. Embeddings find related text; a noul per candidate passage decides whether it actually answers the query. Filter before generation.
- Counting via decomposition. Ask one noul per item ("is
items[i]a fruit?") and sum in code. Never ask Jev to count. - Content screening before context. Before a fetched page enters an agent's context, screen it for injected instructions (jev-mcp
jev_screen). This is a pre-filter only; see Enforcement Boundary.
Rollout Procedure¶
This procedure comes from the DDDS walkthrough. A cheap model becomes expensive when mistakes cause retries, reviews, or incidents, so measure the expected cost per decision including escalations, false approvals, false blocks, and recovery work.
- Pick one bounded, low-risk decision with a closed outcome set that a human can label quickly.
- Write the rubric before calling the model, and define what belongs in every option. Overlapping criteria produce ambiguous choices that are not Jev's fault.
- Collect representative examples with expected answers, including ambiguous and adversarial cases.
- Run in shadow mode beside the current workflow without letting it change behavior.
- Plot accuracy against confidence. Set thresholds from your data, not the marketing claims.
- Automate the safest branch first, and keep a human or stronger model on uncertain cases.
- Pin or log the model version, questions, criteria, and thresholds so changes replay against the same evaluation set.
Cost Model¶
Jev costs $0.042 per million input tokens, and output is free. The $5 signup credit covers about 119M input tokens. TypeSafe's own reference point is a Doom bot making 10 queries/second for ~$7/hour (vendor figure). The headline "193.6x faster / 444.6x cheaper" comes from workflow evals against frontier models and is an upper bound. Because output is unmetered, cost scales with state size, so trimming irrelevant context cuts the bill and raises accuracy at the same time.
Quick estimate: cost/decision ≈ usage.input_tokens × $0.042 / 1e6. At 2,000 input tokens that is $0.000084 per call, or $84 per million calls. Add the cost of escalations and wrong automated outcomes before comparing with an LLM.
Self-Hosting with AnyJev¶
AnyJev reproduces the state + typed-questions interface over any open LLM. Use it when the hosted API is a non-starter: data can't leave the VPC, the volume makes cost extreme, availability is critical, or signups are paused. It offers label-free debiasing (L0), temperature calibration (L1, 100-500 labels per question) and, on the main branch, closed-form hidden-state heads (L2, 100-300 labels per question).
from anyjev import Decider, Question
from anyjev.backends.hf import HFBackend
d = Decider(HFBackend("Qwen/Qwen3-8B"))
route = Question.choice("Which handler should process this request?",
["billing", "technical", "sales", "other"], name="route")
safe = Question.noul("Is the proposed tool call destructive or irreversible?", name="safe")
r = d.decide(state, [route, safe], level="L0")
r["route"].argmax # "billing"
r["route"].distribution # {"billing": 0.81, ...}
r["safe"].p_true # 0.12
On the main branch, L2 serves from a truncated checkpoint behind a vLLM embed server:
python -m anyjev.truncate Qwen/Qwen2.5-7B-Instruct 18 ./qwen-b18
vllm serve ./qwen-b18 --task embed \
--override-pooler-config '{"pooling_type":"LAST","normalize":false,"softmax":false}'
Trade-offs against the hosted API:
- Cost per decision. L0 costs K prefills per K-option choice (2 for noul, 1 for score). L2 costs less than one plain forward pass (0.68-0.84x on Qwen3 per AnyJev's tables).
- Option cap. The letter readout caps choice at 26 options (Jev allows 255).
- Ownership. You run the serving stack and own the calibration labels.
- Benefits. The weights are open, the state never leaves your infrastructure, and calibration artifacts are frozen and auditable.
Run AnyJev's own bench (python -m bench.run) or python -m anyjev.pipeline <model> to see whether it helps on your task before committing.
Troubleshooting¶
- Wrong-but-confident answers. The schema guarantee does not cover judgment. Check for overlapping criteria, ambiguous rubrics, or a question that needs decomposition.
- Accuracy drops after adding context. Isolation means every question sees the whole state, and irrelevant content dilutes it. Prune the state per decision.
- Literal misreadings. Jev reads words, not intent; the per-version jaggedness page documents the known cases. Make instructions explicit or split interpretation into two questions.
- HTTP 400 on a large choice or rubric. The 256th option or the 11th score level is rejected. Use a two-step choice (category, then item) or turn a long numeric scale into a choice.
max_tokens_exceeded. State plus the longest question is over ~32K tokens. Compact the history or send only the slice the question needs.- Slower than advertised on big choice sets. High-cardinality choices use a 2-stage internal scoring system.
TypeError/ validation errors after upgrading the SDK. SDK 0.6.0 changedScore.criteriafrom an integer-keyed dict to an ordered list. Python SDK 0.7.0 moved from msgspec to pydantic, so responses are now Pydantic models (model_dump()).- Can't create an account. Direct signups have been paused since 2026-09-22. Use the Vercel, OpenRouter or Cloudflare routes, or AnyJev.
- Calibration drift after a model bump. Re-run the shadow evaluation. Do not carry thresholds across model versions blindly, and pin
jev-1.13.0on routes that allow it.
Monitoring and Audit Hooks¶
- Log the versioned model ID from every response. Log
response.modelor Pydantic AI'smodel_nametogether with the probabilities, the branch taken, andx-typesafe-request-id. This is what makes decisions replayable and audits possible. - Treat accuracy-vs-confidence as the health metric. Plot it per decision type on live traffic. A widening gap between predicted confidence and observed accuracy is the earliest signal of drift.
- Re-run the shadow evaluation on every change to model version, question wording, criteria, thresholds, or input mix before promoting thresholds.
- Track cost per decision end to end: calls, escalations, false approvals and blocks, and recovery. Token spend alone is not the cost.
Sources¶
- typesafe-sdk-python and @typesafe-ai/sdk: official SDK quickstarts and changelogs
- TypeSafe docs: quickstart, primitives, patterns
- Pydantic AI: TypeSafe (Jev): integration, limits, fallback and compaction
- langchain-typesafe on PyPI: classifier and experimental middleware
- OpenRouter Jev guide and Cloudflare Workers AI: Jev
- Jev, clearly explained (Daily Dose of DS): request/response contract, rollout procedure, patterns
- How to get access to Jev (flaviocopes) and deep dive: access routes, calls, limits
- Jev + SREGym-Lite and SREGym README: jev_plan/jev_submit pattern and results
- jev-mcp (jkudish), system-one-adapter-python, AnyJev: ecosystem repos