JEV: WHAT IT

IS AND WHAT

YOU CAN BUILD

2 SRC

KE

2 sources Updated September 20, 2026

Jev: what it is and what you can build

Jev is TypeSafe AI's model for fast, bounded decisions. Give it relevant context and questions with defined answers; it returns choices, rubric scores and probabilities. Its practical uses include document classification, support routing, agent evaluation, retrieval, adaptive interfaces and attention management. Application code still owns policy, arithmetic and actions, while generative models supply prose and deeper reasoning.

Read the complete deep dive: What is Jev, and what can you do with it? The guide explains the three primitives, examines real builds and GitHub examples, and shows how to design a useful first experiment. Evidence is current to September 20, 2026, five days after the public launch; demonstrations and narrow benchmarks do not establish broad production reliability.

The evals discussion reinforces atomic criteria and shared-state batching; confidence thresholds still require task-specific validation.

Guides

Insights

  • Marc Klingen highlights batching atomic eval questions against shared state. Use separate failure-mode criteria; calibrate action and review thresholds on representative examples rather than treating a probability as proof of correctness. (from typesafe jev classification model for evals)

Model and workflow

  • Choice selects among supplied options, Score measures against ordered criteria, and Noul returns a yes/no proposition's probability. They are not interchangeable confidence scales. (from jev research)
  • Independent judgments can share one request. Keep aggregation, calculations, permission checks and actions in ordinary code. (from jev research)
  • A constrained output cannot invent an undeclared category, but can still choose the wrong category. (from jev research)

Applications

  • DocJev classifies and splits document packets, with published inputs and benchmark limitations; json-render composes interfaces from prepared component candidates. (from jev research)
  • Newsjack screens news for brand relevance; Jevmail categorizes inbox messages; Browser Use combines Jev action selection with generative text when needed. (from jev research)
  • Jev can evaluate narrow properties of agent traces and flag potentially stale memories, provided independent labels and review paths establish the quality boundary. (from jev research)

Evidence and limitations

  • A multilingual retrieval evaluation supports combining embedding and Jev rankings, while standalone Jev reranking does not reliably beat embeddings. (from jev research)
  • Skillful's failed holdout recall gate shows that a selector cannot choose the right skill when candidate retrieval omits it. (from jev research)
  • Test unknowns, contradictory evidence and missing domain rules; reported model confidence is not a correctness guarantee. (from jev research)