Explanation¶
What this page explains
How Jev works and why it is built that way. The page covers the System One thesis (judgment is not generation), RLCD training behind calibrated confidence, the three question primitives, and the parallel-in-isolation evaluation design. It also covers what independent evaluations and security research do and don't confirm, and how to place a decision model safely in an agent loop. Field-level facts are in Reference.
The System One Thesis¶
TypeSafe's co-founder and CEO, Diogo Almeida, previously worked at OpenAI on the instruction-following research behind ChatGPT. He frames generation and judgment as different workloads. Each LLM training paradigm optimizes toward its own North Star, and each North Star bends the model differently:
| Training paradigm | North Star | Pathology |
|---|---|---|
| RLHF | Human preference on generated text | Mode collapse and mode drop; miscalibration "poisons" the probability space because confident style wins rewards |
| RLVR | Programmatically verifiable outputs | Superhuman on benchmarks, "fractal" jaggedness elsewhere |
| RLCD (Reinforcement Learning for Calibrated Decisions) | Calibrated decisions for programmatic use: answers whose confidence tracks observed accuracy | New; the claim is epistemically honest probabilities on System One tasks |
Jev gives up string generation entirely. The output space is fixed at request time, which makes the type-safety guarantee structural rather than empirical. A response is a distribution over options you declared, so a malformed value is not a representable outcome. TypeSafe calls a type error "mathematically impossible". The sharp edge of that claim is that the model can still choose a wrong valid option with high confidence. Schema safety ≠ judgment quality.
The names come from two sources. "System One" is Kahneman's System 1 from Thinking, Fast and Slow: fast intuition, the layer under System 2 reasoning models. "Jev" is William Stanley Jevons, whose paradox says cheaper intelligence unlocks proportionally more demand. TypeSafe's stated mission is "machine-native AI": intelligence consumed by software rather than by people.
Request/Response Model¶
This sequence shows one system_one call end to end, from the SDK through alias resolution to the typed answers your code branches on.
sequenceDiagram
autonumber
participant App as Your code
participant SDK as typesafe-sdk / AI SDK / gateway
participant API as api.typesafe.ai /v1/systemone
participant M as Jev (jev-1.13.0)
App->>SDK: system_one(state, questions)
SDK->>API: POST {model: jev-latest, state, questions}
Note over API: resolve alias to a versioned model<br/>check budgets: 32K state + longest question, 64K total<br/>255 options max, 10 score levels max
API->>M: one pass over the shared state
Note over M: every question answered in parallel<br/>and in isolation, no text decoded
M-->>API: distribution per question
API-->>SDK: {model: jev-1.13.0, answers, usage}
SDK-->>App: typed views (nouls, choices, scores)
Note over App: code branches on probability and confidence<br/>act, escalate, or abstain
Key semantics, from the docs and independent verification:
- One round trip, N questions. All questions ride in a single call. TypeSafe reports that response time barely moves as questions are added.
- Parallel and in isolation. Every question is evaluated independently against the same state, so one answer never becomes context for another. There is no context rot and no cross-question contamination. Independent tests found no batching effect beyond sampling noise.
- Question names never reach the model. Your code branches on the keys in
questions, but the model sees onlyinstructionsandcriteria. Put the full question ininstructions, or in labels and criteria descriptions. - The response names the model.
modelin the response is the version that answered (e.g.jev-1.13.0), even when you sentjev-latest. - Code owns the branch. Jev returns estimates. Your program compares them against thresholds and decides to act, escalate, or abstain.
The Three Primitives¶
| Primitive | Question shape | Returns | Constraints |
|---|---|---|---|
choice |
Which of these N options? | winning label, per-option probabilities, confidence | Up to 255 options. High-cardinality choices use a 2-stage internal system (score options independently, then choose), which is occasionally slower. |
score |
Where on this ordered rubric? | fractional expected score, level legend, full probability distribution, confidence | Up to 10 levels, indexed from 0 in list order |
noul |
Is this true? | probability 0.0-1.0 | TypeSafe's Boolean decision type. criteria.true / criteria.false are optional. |
A single call can mix all three freely. The DDDS walkthrough's choice answer shows the contract:
{
"type": "choice",
"choice": "billing",
"probabilities": { "billing": 0.52, "technical": 0.46, "sales": 0.02 },
"confidence": 0.18
}
The winner was 6 points ahead. The distribution, not the label, is the automation signal.
Type Safety and Calibration¶
Jev offers two orthogonal guarantees that are often conflated:
- Schema safety (absolute). The output cannot violate the declared types. This follows from the constrained output space, not from intelligence.
- Calibration (empirical, per deployment). RLCD training aims to make predicted confidence track observed accuracy. That relationship must be measured on your traffic, per decision type, and re-checked after any change to model version, questions, or input distribution.
Because only the first is guaranteed, every automated branch needs an abstention policy scaled to consequence. This flowchart shows the canonical policy (DDDS, SREGym, and the Pydantic AI fallback all follow this shape):
flowchart TD
Q["Jev answer<br/>noul p, or choice/score + confidence"] --> R{"Consequence of<br/>a wrong action?"}
R -->|low| T1{"confidence >= your<br/>high threshold?"}
R -->|high or irreversible| H["Hard enforcement decides<br/>permissions, allowlist, tests"]
T1 -->|yes| A["Act automatically<br/>log model ID + probabilities"]
T1 -->|no| T2{"confidence >= your<br/>low threshold?"}
T2 -->|yes| E["Escalate to an LLM<br/>FallbackModel or stronger model"]
T2 -->|no| HU["Route to a human"]
H --> HU
The thresholds come from your own accuracy-vs-confidence plots, live in code, and are versioned with the model ID they were tuned on.
Internal Architecture¶
Jev's internals are not published in detail. TypeSafe discloses three things:
- A new model architecture. Jev is not a downsized LLM; the FAQ explicitly rebuts "is Jev just a smaller LLM".
- A parallel sampler. It emits all outputs in one hardware-aware query instead of generating tokens sequentially.
- RLCD training.
The efficiency story is structural. With no autoregressive decoding, output tokens are too cheap to be worth metering, and latency does not scale with the number of questions. Third-party descriptions call it "non-autoregressive". Treat any deeper detail as speculation until TypeSafe publishes it.
This diagram maps the parts you can observe: your client, the access routes, the hosted API, and the model behind it.
flowchart LR
subgraph Client["Your application"]
Q["questions map<br/>noul / choice / score"]
S["state<br/>string / JSON"]
B["branch logic<br/>thresholds + abstention"]
end
subgraph Routes["Access routes"]
SDK["typesafe-sdk<br/>@typesafe-ai/sdk"]
FW["Pydantic AI, LangChain,<br/>jev-mcp"]
GW["Vercel AI Gateway<br/>typesafe-ai/jev"]
OR["OpenRouter Decisions API<br/>typesafe/jev-1.13"]
CF["Cloudflare Workers AI<br/>typesafe/jev"]
end
subgraph TypeSafe["TypeSafe hosted service"]
API["api.typesafe.ai<br/>/v1/systemone, /v1/models"]
AL["alias resolution<br/>jev-latest, jev-preview"]
PS["parallel sampler<br/>one pass, all questions"]
W["Jev weights, RLCD-trained<br/>closed"]
end
Q --> SDK
S --> SDK
Q --> FW
FW --> SDK
SDK --> API
GW --> API
OR -.->|"provider route"| API
CF -.->|"hosted copy"| W
API --> AL --> PS --> W
PS -->|"answers + usage"| B
The dotted edges mark routes whose hosting arrangement is not public. Cloudflare hosts "the same model" per its model page. Whether OpenRouter proxies to TypeSafe or hosts the model itself is not documented in OpenRouter's listing or Jev guide (as of the 2026-09-25 check).
Component Breakdown¶
| Component | Role | Notes |
|---|---|---|
state input |
The context under judgment | String, JSON object, or JSON array, shared by all questions; no multimodal input |
questions map |
Typed decisions to evaluate | { type, instructions, criteria }. Atomic questions work best: decompose multi-factor judgments and combine the answers in code. |
| Parallel sampler | One-pass evaluation of all questions | Replaces autoregressive decoding; 70-500 ms end to end |
| RLCD training stack | Calibrated confidence | The differentiator against RLHF/RLVR models |
| Versioned model IDs | jev-1.13.0, jev-latest, jev-preview |
Responses report the answering version. Log it, and pin thresholds to it. |
| Token budgets | ~32K for state + longest question (~150K English characters); ~64K for state + all questions | Irrelevant context measurably hurts accuracy, so send minimal state |
| Official SDKs | typesafe-sdk (Python), @typesafe-ai/sdk (TS) |
Typed questions and answers; 10 s default timeout; 2 retries by default |
Workflow evals (evals.typesafe.ai) |
TypeSafe's benchmark: fixed compute graphs scored against an average of large frontier models | Source of the 193.6x faster / 444.6x cheaper headline claims |
Benchmarks and Evidence¶
TypeSafe's workflow evals (self-reported)¶
TypeSafe introduced a new evaluation type. A correct compute graph is assumed (the "workflow" lives in code), and models are scored against the average predictions of the largest external frontier models running the same workflow. The reference is the average of GPT-6 Astra and Fable 5.1, per TypeSafe's launch post (confirmed via search listing and TrueFoundry's summary, 2026-09-27). Jev claims the cost/accuracy Pareto frontier across almost two orders of magnitude.
TypeSafe discloses these caveats itself:
- Its own model-capabilities team built the workflows. Bias is possible, though the workflows are not in the training distribution.
- The reference answers bias toward OpenAI/Anthropic models, which undercounts DeepSeek and Jev alike.
- The headline multipliers are "on the higher end of real world gains".
- Speed was measured from West Coast laptops.
- Pricing sustainability is unproven: TypeSafe "can't prove it isn't subsidized".
LLM baselines ran through TypeSafe's own structured-decision wrapper (system-one-adapter). TypeSafe argues this is the most accurate way to get decisions from LLMs, though it is slower and costlier.
Independent evaluation: SREGym-Lite (September 2026)¶
SREGym, an MIT-licensed SRE incident benchmark, integrated Jev as decision support for a Codex-harness agent on gpt-5.6-luna, using two gates:
jev_planranks 3-5 competing diagnostic hypotheses via choice and score questions.jev_submitreviews the evidence before a diagnosis or mitigation is submitted. Every required question must reach probability >= 0.70, and a rejected submission forces new evidence gathering.
The integration ships in the SREGym repo behind --jev-model.
- Result: 24/50 vs 20/50 attempts (40% -> 48%) across 10 problems × 5 attempts. internal-traffic-policy went from 0/5 to 3/5; two problems regressed.
- Where it helped: distinguishing a causal mechanism from believable noise (OpenTelemetry errors, unrelated workloads). Jev added "useful friction before premature diagnosis".
- Where it failed: it accepted evidence of current functionality without testing the invariant that makes a repair durable, for example fixing a pod but leaving
maxUnavailable: 100%. It also could not rescue a missing hypothesis: if the right test never enters the candidate set, ranking cannot recover it. - Caveats: a small sample from a single source, self-published by the benchmark's authors; pass rate only (time-to-diagnosis was untested).
Integrator observations¶
Pydantic AI's docs report several behaviors seen while running Jev behind agents. Jev re-picks a tool whose result is already in the history. It will propose tool calls whose required arguments the text does not state. Asked whether it can answer, rather than what the text calls for, it hands off nearly everything. These are integrator reports, not TypeSafe claims, but they are the best early field data on agent use.
Ecosystem signal¶
Earlier write-ups described Beacon (Asymptote Labs, MIT) as using Jev to score agent traces across 20+ coding-agent harnesses. The current Beacon README (checked 2026-09-27) does not mention Jev, so this note no longer counts Beacon as an adopter. A clearer adoption signal is that within ten days of launch Pydantic AI, LangChain, the Vercel AI SDK, OpenRouter and Cloudflare all added first-class support.
What Jev Cannot Do (per TypeSafe's own jaggedness disclosures)¶
TypeSafe publishes these per model version (for example model-jaggedness/jev-1.13) and revises them as models change:
- No generation of any kind: no reasoning chains, no explanations.
- No arithmetic, counting, or date comparison. Dates are read as text, not as ordered quantities. No precise string manipulation: it cannot compare
#FF4B0Ato another hex color. - Literal reading. It "answers the question you wrote, not the one you meant". Scoping words, negations and implied conditions are read at face value.
- Indirection hurts. Questions about a property of a property, or ones that need several hops, lose accuracy.
- Adversarial text can move answers. Jev treats the state as data, not as hostile. Injected instructions, misleading framing, or text that argues for its own classification can shift the result. TypeSafe says it expects to improve this.
- Structural invariants don't hold. A question and its negation need not sum to one. The same question asked as a noul and as a choice gives numbers that do not compare.
- No unknown-value extraction: it chooses only among the candidates you supply.
- Text-only state; multimodal is future work.
- Calibration numbers are self-reported, and independent replication is thin as of 2026-09-25.
Latency and Cost Anatomy¶
A fixed output space is fast for a structural reason. Autoregressive generation pays for each output token in sequence. Jev's parallel sampler emits all answers in one hardware-aware pass, so latency is dominated by one forward pass over the state, not by question count or answer length. The consequences:
- The 70-500 ms latency range is roughly flat in the number of questions (docs: "adding questions barely changes the response time"). Integrators observe ~150-500 ms, with ~180 ms typical.
- Output is unmetered ("currently free of charge"), so the bill scales with state tokens only.
- Cost per decision is dominated by engineering: rubric design, shadow evaluation, and threshold maintenance. TypeSafe's own guidance is to measure end-to-end cost per decision, including escalations and false outcomes.
- The exception to flat latency is very high-cardinality choice (up to 255 options). It uses a 2-stage internal process (independent option scoring, then an explicit choice), which is visibly slower.
The Open Counterfactual: AnyJev (Nokia Applied Research, September 2026)¶
AnyJev (Apache-2.0, by researchers at Nokia and Tencent Hunyuan; not affiliated with TypeSafe) reproduces the Jev interface over open LLMs. It reads decisions straight off next-token log-probabilities, with no generation and no fine-tuning, and fixes the two biases that make raw logit readouts unusable:
- Cyclic-shift marginalization. A K-option list is shown in K rotations so every option sits at every position once, and the results are combined in log space. This removes position bias.
- Prior correction. The model's label prior is estimated without labels and divided out. The default is batch calibration at strength 0.75; the opt-in content-free probe uses
N/Ainputs. In the README's example, a spam noul reads P(Yes) = 0.62 on the content but 0.70 onN/A, so the model leans Yes regardless. Dividing out the prior flips the judgment to 0.41.
Every decision carries its level, and require="L1" refuses weaker ones:
| Level | Needs | Does |
|---|---|---|
raw |
nothing | Restricted softmax over label tokens. This is what the simple clones (von, mini-jev) do. |
L0 |
nothing | Removes position bias and divides out the prior. Not calibrated. |
L1 |
100-500 labels per question | Temperature scaling on top of L0; calibrated within its distribution |
L2 (main branch, 0.1.0 unreleased) |
100-300 labels per question | A closed-form head on a hidden state partway down the network, one prompt per state, cheaper than a full forward pass |
Measured on Qwen3-8B with BANKING77 20-way (300 items):
- The order-flip rate falls from 0.230 to 0.073.
- Accuracy rises from 0.747 to 0.803.
- ECE falls from 0.240 to 0.095 (L1).
- The operationally decisive number: the share of traffic auto-decidable at <=5% error rises from 7.7% (raw) to 52.0% (L1).
Against Jev's published number. On LocalLLaMA/typed-decisions (20 questions, 2,000 held-out decisions), AnyJev's current tables (regenerated 2026-09-22) quote Jev at 0.727 accuracy, a figure published by its authors and not rerun. Qwen3-32B + L1 reaches 0.699, 2.8 points behind. The new L2 heads pass it: Qwen3-4B scores 0.786, Qwen3-8B 0.771, Qwen3-30B-A3B 0.799 and Qwen3-32B 0.798, with pooled ECE 0.03-0.05. A Qwen3-1.7B truncated to 64% depth reaches 0.730. A 421M model fine-tuned on the task (Laya) scores 0.768.
Treat this as first-party: the benchmark is AnyJev's own. L2 needs labels per question and per model. Jev needs none.
AnyJev's README states its own limits:
- L0 is not a free win everywhere. It lowers coverage at 5% risk on one prompt-injection split, and the batch prior hurts when the true label marginal is skewed.
- Calibration makes uncertainty legible, not smaller. On Minesweeper and maze games no readout beats random.
- The letter readout caps choice at 26 options.
- Only Qwen rows are measured so far.
- L1 and L2 do not survive distribution shift without refitting. L2 heads adapt to rewordings from ~30 unlabelled requests.
Source Discrepancies¶
These are now tracked in Reference: Source Discrepancies: pricing versus the reseller, headline multipliers, latency and context figures, the "cannot hallucinate" wording, and independent accuracy numbers.
Output Integrity: Schema Safety ≠ Judgment Safety¶
- What is guaranteed. The response cannot violate the declared output schema. A malformed value is not a representable outcome; TypeSafe calls a 0% type-error rate structural ("mathematically impossible"), not an empirical measurement. Downstream code will never see an undeclared option, a string where a number was declared, or a hallucinated key.
- What is not guaranteed. The winning option may not be correct. Jev can select a wrong valid option with high confidence. In the DDDS walkthrough's billing/technical example, billing won by 6 points with 0.18 confidence, which is exactly the case the confidence field exists to catch.
- The engineering consequence is the abstention policy in Type Safety and Calibration. Small consequence with high confidence: act. Middling confidence: confirm, or escalate to a stronger model. Low confidence: route to a human.
- Calibration drift is the silent failure mode. The confidence-accuracy relationship can break when the model version changes, when questions or criteria change, or when input traffic drifts. Re-measure after any of those. SREGym's failure case, a repair accepted as complete without testing the durability invariant, is the canonical example of confident-but-wrong at the boundary.
Supply Chain¶
Jev has closed weights and is hosted only. You are trusting TypeSafe's service, plus whatever route you take to reach it.
| Artifact | Risk posture | Control |
|---|---|---|
api.typesafe.ai (official) |
Primary trust anchor. Pricing may be subsidized (TypeSafe's own disclosure), and signups were paused on 2026-09-22 under load. | Pin model versions (jev-1.13.0, not jev-latest) for replayable behavior, and log the versioned ID from every response |
Gateway routes (Vercel typesafe-ai/jev, OpenRouter typesafe/jev-1.13, Cloudflare typesafe/jev) |
Each gateway is a second trust anchor with its own logging and retention posture. Cloudflare cannot pin versions, and OpenRouter's Decisions API is alpha. | Check each gateway's data terms independently. Prefer routes that pin when thresholds matter. |
jevtypesafeai.com (UNOFFICIAL) |
Community demo and reseller selling "instant hosted keys" at 6-10x official pricing; the name mimics the vendor | Do not send sensitive state or keys through it; it is not TypeSafe. The official domain is typesafe.ai. |
Package-name lookalikes (typesafe-ai on PyPI, unscoped jev-mcp on npm) |
Third-party shims or unrelated projects under plausible names | Install typesafe-sdk, @typesafe-ai/sdk, and @jkudish/jev-mcp explicitly; verify the repository link on the registry page |
| Ecosystem packages (awesome-jev entries, pi-jev routers, clones) | Third-party code at varying maturity. awesome-jev itself warns that a listing is not an endorsement. | Review before wiring into agent loops; do not assume a clone inherits Jev's calibration |
| AnyJev (Nokia, Apache-2.0, self-hosted) | You own the stack: model weights, calibration labels, serving infrastructure. Calibration artifacts are frozen and auditable. | Standard model-supply-chain hygiene on the underlying LLM; re-run calibration checks on every model or distribution change |
- No official offline story. TypeSafe offers no self-hosted or weights-export path, and availability, rate limits, and deprecation policy are the vendor's to change. Keep the interface shim thin (one client module). The credible escape hatch is AnyJev: the same interface over your own open LLM, with state never leaving your VPC. The costs are K prefills per choice at L0 (or labels per question for L2), a 26-option cap, and owning the calibration data yourself.
- No known CVEs as of 2026-09-27 (web search found none); the product launched on 2026-09-15. The first public adversarial research appeared in late September (see Enforcement Boundary). The absence of CVEs reflects the product's age, not audit depth. Re-check before relying on it in production.
Privacy and Data Flow¶
- Every call sends your data out, by design. The full state goes to TypeSafe's hosted service, or to the gateway you chose. For triage and moderation use cases that means user content, tickets, resumes, or logs. Classify what flows through: the ~32K/64K-token budget is generous enough to over-share by accident. Send the minimum state per decision, which also improves accuracy.
- No fixed retention window is published. TypeSafe's privacy policy lets it retain personal data while reasonably needed for service or business purposes, with no fixed period for standard accounts, and says input is not used to train or fine-tune models; zero data retention is offered to enterprise customers on request (checked 2026-09-28 via search listing; the same reading appears in the KorWF-Pi verification PR, 2026-09-21). The official adapter sets
store=Falsewhen calling LLM providers, but that says nothing about TypeSafe's own retention. Before wiring in regulated data (HR screening, support tickets with PII), get written retention and processing terms from TypeSafe and from any gateway in the path. The ZDR configurations in Zero Data Retention are the pattern to demand. - Questions and criteria are sensitive configuration. Your choice criteria encode business taxonomy and policy, and they transit the same channel. Treat rubric leakage as a minor business-logic exposure.
- Downstream flow. Jev outputs feed code branches, so an automated rejection or escalation is a decision about a person (resume scoring, moderation). Keep humans on the low-confidence path, and log the probability with the decision for auditability.
Enforcement Boundary: Scores Inform, Never Enforce¶
Jev is well suited to screening for prompt injection, policy violations, and risky tool calls. It is unsuitable as the enforcement mechanism: its score should gate nothing by itself. Permissions, sandboxes, allowlists, and tests must enforce the exact rules.
The correct composition is Jev as a fast semantic pre-filter in front of hard enforcement. A high-confidence read-only call proceeds to the permission check; anything destructive or uncertain pauses. That way a miscalibrated or adversarially steered score can only change latency and friction, never authority.
Late-September research confirms the concern:
- TypeSafe's own limitations page says adversarial text "can move the answer".
- Pydantic AI's docs state that "a guard built on Jev belongs alongside deterministic checks, not instead of them".
- An arXiv paper, Decision Hijacking: Prompt Injection Attacks on Jev's Typed Probabilistic Decisions (2609.28613), found that malicious content shifts action probabilities but rarely makes Jev select the attacker's target. Adaptive attacks that use score feedback roughly double the attacker-target probability, so do not expose raw probabilities to untrusted callers.
- Check Point's analysis argues a decision model "breaks like any other language model" under injection.
- VentureBeat flags a shared-failure risk: when Jev both evaluates an agent and makes decisions during its execution, one piece of injected text can sway both checkpoints.
Threat Model Summary¶
| Threat | Vector | Impact | Likelihood | Control |
|---|---|---|---|---|
| Confident wrong decision | Ambiguous rubric or distribution drift | Bad automated outcomes (misrouting, wrong approval) | Medium | Abstention thresholds from your own calibration plots; shadow mode first |
| Calibration drift | Model version bump, question edits, traffic shift | Thresholds silently stop meaning what they meant | Medium | Pin version IDs; re-run the shadow eval on every change |
| Reseller or key theft | jevtypesafeai.com and similar lookalike sites or packages |
Key or state leakage to an unknown third party | Low-Medium | Official domain (typesafe.ai) and official packages only; treat resellers as untrusted |
| State over-collection | A generous budget invites dumping full context | PII or regulatory exposure to a hosted API | Medium | Minimal state per decision; written retention terms for regulated data |
| Score steering (prompt injection) | Attacker-controlled text in state; adaptive attacks that read the returned probabilities |
The semantic filter says "safe" for malicious input | Medium (documented) | Jev pre-filters, hard enforcement decides; do not return raw probabilities to untrusted parties |
| Shared failure across checkpoints | The same model judges and gates the same agent | One injection defeats both layers | Medium | Diversify checkpoint mechanisms (deterministic checks, different models) |
| Vendor lock-in or availability | Closed weights, hosted only, signups paused under load | A service or pricing change breaks automated branches | Medium (early) | Thin interface shim; versioned questions and criteria for replay; LLM or AnyJev fallback path |
| Overlapping criteria | Bad rubric design | Ambiguous categories misrouted, which looks like model error | High (design-time) | Review rubrics before calling the model; criteria are program logic, so test them |
Agent-Loop Placement Guidance¶
Where Jev sits in an agent loop changes its risk profile. There are three placements, in order of increasing caution:
- Advisory (lowest risk). Jev ranks, sorts, or labels, and a human or the LLM consumes the output as context. SREGym's
jev_plantest ranking and trace scoring are examples. A wrong answer wastes attention, not authority. - Gated decision (medium). Jev's probability is one input to a branch that also checks hard conditions: a confidence floor, evidence requirements, confirmation prompts. SREGym's
jev_submitis the reference design: every required question must clear 0.70, and a rejection forces new evidence rather than rewording. - Autonomous enforcement (highest risk). Jev's output alone triggers consequential actions. Avoid this placement. Combine Jev with hard enforcement (permissions, allowlists, sandboxes) so a miscalibrated or steered score can add friction but never authority.
Rules of thumb for agent builders:
- Give Jev the smallest state slice per decision; accuracy and privacy improve together.
- Keep separate questions per risk dimension instead of one composite judgment: deletion, git history, and production reachability each get their own question.
- Log every gate outcome with its probability, so incidents can be replayed against the versioned model that made them.
- In agent frameworks, cap loops (
UsageLimits(request_limit=...)in Pydantic AI). Jev tends to re-pick a tool whose result is already in the history.
Sources¶
- Introducing System One Models & Jev (TypeSafe AI): RLCD, workflow evals, evidence and nuance; the type-safety claim and its disclosed limits
- TypeSafe docs: primitives, confidence, quickstart, patterns, models
- typesafe-sdk-python: wire schema generated from the OpenAPI spec
- Pydantic AI: TypeSafe (Jev): jaggedness summary, integrator observations, limits
- Jev, clearly explained (Daily Dose of DS): response contract, abstention policies, schema-vs-judgment precision
- Jev + SREGym-Lite: independent experiment and confident-but-wrong failure cases
- Latent Space: Jev (Diogo Almeida): RLCD rationale, the mode-collapse argument, calibration framing
- Decision Hijacking (arXiv 2609.28613), Check Point: A decision model breaks like any other language model, VentureBeat: prompt injection can influence the verdict
- AnyJev (Nokia Applied Research): levels, results_exit, results_bench
- evals.typesafe.ai: workflow eval methodology
- jevtypesafeai.com: the unofficial reseller site, documented here as a supply-chain trap