知識がなくても始められる、AIと共にある豊かな毎日。
AI News & Trends

What Is Jev? A Decision Model That Returns Answers and Probabilities, Placed in a 3D Printing Workflow

swiftwand

Ask an AI assistant whether a file is safe to print and you usually get a paragraph back. It starts with “in general this should be fine,” lists a few conditions, and ends by suggesting you double-check. That is a kind answer for a person to read, but an awkward one for a program. What the program needs is a yes or a no, and how sure the model is. Jev, released by TypeSafe as an early-access model on September 15, 2026, is built so that this mismatch never happens in the first place.

This article contains affiliate links. As an Amazon Associate, we earn from qualifying purchases.

Jev does not write replies, code, or explanations of its reasoning. You define a typed question, and it returns the answer it picked plus probabilities. In this article we call a model that skips text generation and returns typed answers with probabilities a decision model.

A note on our position. What follows is drawn mainly from TypeSafe’s official docs and blog, the official announcements of the vendors that followed, and a paper on arXiv. Press coverage and our own measurements are labeled where they appear. This article does not contain results from calling Jev ourselves: request and response shapes are taken from the official examples, and speed and cost multipliers are reported as the vendor’s own claims. In the second half we split a 3D printing workflow into files and settings before printing, printer state during a print, camera images during a print, and the decision to stop, and map which of those Jev can read.

忍者AdMax

Where Text Answers Create Extra Work

When a chat-style AI makes a decision for you, the extra work appears after the answer, not before it.

  • Parsing. The answer is prose, so you have to write code that extracts the part that means “yes.” Even when you ask for JSON, a preamble sentence or an unrequested field can sneak in, and you must decide what to do when extraction fails.
  • Variation. Ask twice and the wording changes. Whether “this should be fine” is more confident than “probably OK” cannot be read from the text. Any rule you write to map phrases to numbers becomes a new source of error.
  • Confidence. Automation eventually needs thresholds such as “stop above this value” or “send to a person in this range,” and thresholds need numbers. You can ask a chat model for a number, but you still have to check how well that number matches reality.

In 3D printing these costs are concrete: reading a model license to decide whether prints can be sold, checking slicer settings before starting a print, or reading printer state mid-print to decide whether to stop. In every case the end product is a branch in code, not a paragraph. Decision models are designed to return that branch value directly.

What Jev Returns

A “System One model”

The Jev decision model is what TypeSafe calls a System One model, after the fast, intuitive “System 1” thinking that Daniel Kahneman popularized in Thinking, Fast and Slow. The official blog says the name Jev comes from the economist William Stanley Jevons.

The docs are explicit: a System One model does not write replies, produce code, or explain its reasoning. So you cannot swap Jev in as the model behind a coding agent such as Claude Code. Jev is not a replacement for the language model that drives an agent; it is a decision component that an agent or an ordinary program calls.

Inputs and outputs

You send Jev the content to judge (called state in the API) and a set of named questions. The state can be a string, a JSON object, or an array of text values. Images, audio, and video are not accepted; the docs ask you to pre-process non-text inputs into text or structured fields first.

Each question comes back with an answer and probabilities. The model ID is jev-1.13.0, with the aliases jev-latest and jev-preview. You can use it from a Playground, an HTTP API (POST /v1/systemone), a Python SDK (typesafe-sdk), a JavaScript/TypeScript SDK, and an agent skill that helps coding agents write Jev code; the skill installs as a plugin in Claude Code and also works with Codex and others, and its repository is MIT licensed. Since September 16, 2026, Jev can also be called through Vercel’s AI Gateway (via the AI SDK 7 evaluate API), and on September 21 the gateway added support for TypeSafe clients and an HTTP API.

Training and what “calibrated” means

TypeSafe calls its training method RLCD (Reinforcement Learning for Calibrated Decisions): the model is trained to return decisions and calibrated probabilities instead of text. TypeSafe introduces its founder Diogo Almeida as a co-inventor of RLHF. Note that TypeSafe is not the only one using the RLCD name; Cloudflare, covered below, says it used a method of the same name for its own models.

Calibration is a property of many predictions taken together. If you collect all the questions where the model said 0.8, about 80 percent of them should turn out to be “yes.” TypeSafe’s docs say this plainly: calibration is measured over a population of predictions and does not guarantee that any single answer is right. A 0.95 answer can still be wrong.

Three Question Types: noul, choice, score

Questions in Jev are typed. There are three types, which TypeSafe’s API calls noul, choice, and score.

TypeWhat it asksWhat comes back
noulA yes/no question such as “may the prints be sold?”A single number from 0 to 1: the probability that the answer is yes
choicePick one option from a list you define (up to 255 options)The chosen option, the probability of every option, and a confidence value
scoreRate on a scale you define (2 to 10 levels)A probability-weighted value that can land between levels, plus confidence

Vendors name the yes/no type differently: Vercel’s AI Gateway API calls it boolean, and OpenAI calls it predicate. We use TypeSafe’s noul here.

Confidence is not the probability of being right

choice and score answers carry a confidence value between 0 and 1; noul answers do not. Confidence is derived from the probability distribution: it is 1 when all probability sits on one answer and 0 when it is spread evenly. It tells you how concentrated the distribution is, not how likely the answer is to be correct. Ollama’s docs, for the local version of this API, state that confidence is “not calibrated correctness,” and the llama.cpp docs say the probabilities are not guaranteed to be calibrated on your data. Mixing up probability and confidence leads to badly designed thresholds.

The official request and response example

Here is the choice example from TypeSafe’s API reference, unchanged. It routes a support message to billing, technical, or sales.

{
  "state": "Help! My payouts have been failing for 3 days.",
  "model": "jev-latest",
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this?",
      "criteria": {
        "billing": "Payments, invoicing, refunds",
        "technical": "Bugs, outages, integrations",
        "sales": "Pricing, upgrades, new accounts"
      }
    }
  }
}

The documented response for this request looks like this:

{
  "model": "jev-1.13.0",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "billing",
      "probabilities": { "billing": 0.88, "technical": 0.12, "sales": 0.0 },
      "confidence": 0.81
    }
  },
  "usage": { "input_tokens": 318, "output_tokens": 34 }
}

All numbers are from the official example, not from a call we made. Three things stand out. The answer comes back under the key you named, so there is nothing to parse. The probabilities of the options that lost are returned too, so you can see how close the runner-up was. And while the request asked for the alias jev-latest, the response names the exact version, jev-1.13.0, which you can log for later audits.

TypeSafe’s docs compute choice confidence as (p_max − 1/n) / (1 − 1/n), that is, how far the top probability is from an even split. Plugging in 0.88 with three options gives 0.82, slightly different from the 0.81 shown in the example. The docs do not explain the gap, so we left the example as published.

You can put several questions in one request. According to the docs, Jev reads the state once and evaluates every question against it in parallel. No cap on the number of questions is documented; the token budget is the limit.

Pricing and Limits: You Pay for Input, Output Is Free

Values listed in the official docs as of October 7, 2026:

ItemValue
Model IDjev-1.13.0 (aliases jev-latest, jev-preview)
Input price$0.042 per 1M tokens
Output priceFree
Per-request limit64k tokens (state plus the longest question: 32k tokens)
Rate limits100,000 tokens/s and 80 requests/s
choice optionsUp to 255 per question
score levels2 to 10
Input typesText only (string, JSON object, array of text values)

At 1,000 input tokens per call, 1,000 calls add up to 1M tokens, or about $0.04 in input charges. Real token counts depend on the length of your state and questions, so treat this as a way to estimate rather than a quote. The response in the official example also reports output tokens, but only input is billed.

The docs say rate limits are adjusted dynamically and may change without notice; higher limits are handled through individual contracts or enterprise plans. Billing is prepaid credits, and the customer agreement says purchased credits expire within 12 months of purchase unless the order says otherwise. The docs do not state an amount for any sign-up credit, and the agreement leaves promotional credits to TypeSafe’s discretion, so we do not quote a figure.

One caveat on limits: Cloudflare’s comparison lists Jev’s context as 32k, while TypeSafe’s docs show two numbers, 64k per request and 32k for the state plus the longest question. Quote only one and you change what the limit refers to. On data, the docs and privacy policy say Jev does not train on customer requests or responses, and a zero-data-retention (ZDR) option is available for enterprise customers. A retention period for standard plans is not stated.

What Is New, and What Is Not

Several features that look unique to Jev are not. Separating them makes the real novelty clearer.

Typed output alone is not new

OpenAI’s Structured Outputs already generates responses that follow a JSON schema you supply. The guarantee has conditions: only a subset of JSON Schema is supported, and a refusal or hitting the output token limit can produce a response that does not match the schema, which the caller must handle. OpenAI itself draws the line this way: use Structured Outputs to generate an object that follows your own schema, and the Decisions API when you need typed answers with probabilities. The difference is the probability. A decision model returns a probability for each option and is trained so those probabilities are calibrated; Structured Outputs enforces shape, not confidence.

Some decision models are close cousins of log probabilities

Returning probabilities is not a brand-new technique either. Some open decision models turn the logits of candidate answers into probabilities with a softmax, and Cloudflare says its prototype used the log probabilities produced by a large language model. The difference lies in whether the model is trained to calibrate those probabilities. For Jev, TypeSafe describes a new architecture and a parallel sampler that outputs all probabilities at once rather than generating tokens one by one, and the FAQ on its launch post says Jev is not an LLM. The architecture is not published, so we do not guess at it from later open models.

For local users there is one more catch: as of October 7, 2026, Ollama’s OpenAI-compatible API does not support log probabilities. To get decision-model probabilities from Ollama you need its dedicated decision endpoint, described below.

Versus classifiers trained per task

Classifiers trained on labeled data for one task, such as image-based defect detectors, have been around for years. The paper analyzing Jev also notes that supervised classifiers need task-specific labeled data. The advantage of a decision model is that writing a question turns it into a classifier, which matters most for small decisions where you cannot collect training data or where the question changes often. For a high-volume task with plenty of labels, a purpose-trained classifier is still a strong option. Decision models cover the long tail of decisions that never justified training.

What “does not hallucinate” means

TypeSafe’s blog argues that hallucination and type safety are intrinsically related. The claim is that output never leaves the predefined types and options, so schema matching is guaranteed, and the blog notes that this figure is derived from the guarantee rather than measured. Reading it as “Jev is never wrong” is a mistake: Jev can still pick the wrong option among valid ones. Taiwan’s iThome made the same point: a correct type does not mean a correct judgment.

Speed and cost multipliers are self-reported

According to the blog, Jev responds in 70 to 500 ms and is 40 to 200 times faster than LLMs of comparable intelligence on System One-shaped questions. Most published evaluations were run from laptops on the US West Coast, where the service is also hosted, so calls from elsewhere add round-trip time. The “193.6× faster, 444.6× cheaper” figure on the homepage comes from an internal workflow evaluation that used the average answer of GPT-6 Astra and Claude Fable 5.1 as the reference, and the blog says it is likely on the high end. For contrast, Cloudflare’s own measurement puts Jev’s median latency at 524.1 ms. Who measures, and from where, changes the number.

So what is actually new?

The novelty is in the combination: a model that writes no text at all, trained to calibrate its probabilities, offered as a typed question API that bills only for input. Each piece has precedents, but shipping them together as one product prompted several vendors to release models and endpoints of the same shape within weeks.

If you want to see how language models turn text into tokens and token scores into probabilities, a hands-on book on LLM internals makes the difference between generation and decision models easier to follow.

USD 37.68 on Amazon.com (as of 2026/09/22)

Same-Type Models from Several Vendors in Three Weeks

Within three weeks of Jev’s launch, other vendors shipped decision models of the same shape, or endpoints to call them.

DateProviderWhat was released
2026-09-15TypeSafeJev (early access)
2026-09-16VercelJev callable from AI Gateway (TypeSafe clients and HTTP API added on 09-21)
2026-09-18Bespoke LabsNimble (9B) decision model weights on Hugging Face
2026-09-28OllamaDecision endpoint /v1/systemone in 0.35
2026-10-01strands-labsStrands Decider (weights registered on Hugging Face on 09-30)
2026-10-01CloudflareClef (27B) and Clef-flash (9B)
2026-10-01OllamaClef and Clef-flash support merged (shipped in 0.35.1)
2026-10-02llama.cpp/v1/systemone added to the server
2026-10-06OpenAIDecisions API (public beta)

Dates for Bespoke Labs, Ollama, strands-labs, and llama.cpp are UTC dates from GitHub and Hugging Face records. The followers also describe themselves relative to Jev: Ollama’s blog post is titled around “Jev-style decision models,” and Cloudflare reports its models against a “Jev Decision Index.”

Ollama: decision models on your own PC

Ollama 0.35 (released September 28, 2026, UTC) added /v1/systemone, an endpoint based on the Jev API, so decision models can run locally. One request can ask 1 to 64 questions. The listed models are Bespoke Labs’ 9B Nimble and Together AI’s experimental tev1 in 4B and 0.8B sizes. From 0.35.1, Ollama also supports Cloudflare’s Clef and Clef-flash, which can read images. The software and weights are free; the GPU and electricity are yours.

Strands Decider: a small model with the text head removed

strands-labs, the project space around AWS’s open-source agent SDK Strands Agents, published Strands Decider on October 1, 2026 (UTC); the weights were registered on Hugging Face on September 30. It takes Qwen3.5-2B-Base, discards the language-modeling head, and replaces it with a small pointer head that points at options, for 1.9B parameters under Apache 2.0. It accepts images without image-specific training. The announcement appeared on the official Strands Agents blog; we found no announcement on AWS’s own site.

Cloudflare Clef: decision models that read images

On October 1, 2026, Cloudflare released Clef (27B) and Clef-flash (9B) on Workers AI, with weights on Hugging Face under Apache 2.0. Clef is built on a frozen Qwen3.8-27B and Clef-flash on a frozen Qwen3.5-9B; each runs a single prefill pass and then scores the valid options in parallel. Cloudflare says it used Brier loss to refine calibration.

The biggest difference from Jev is image input, and Cloudflare itself notes that Jev only does text classification today. Clef-flash on Workers AI costs $0.09 per 1M input tokens. Cloudflare says Clef is 2.5 times and Clef-flash 13 times faster than Jev at median latency, again by its own measurement. It also says Clef follows the same API shape as Jev, so an existing Jev integration can move by changing the endpoint and model name. The confidence formula, however, may differ: TypeSafe defines it per question type, while Ollama computes it from the entropy of the distribution. Re-measure any confidence thresholds when you switch.

OpenAI Decisions API: public beta, gpt-6-luna only

OpenAI released the Decisions API as a public beta on October 6, 2026. The only model so far is gpt-6-luna, and the types are predicate, choice, and score. Images are accepted only as inline base64 data URLs. Pricing is $0.10 per 1M input tokens with no charge for output or cache, though regional processing premiums and long-context multipliers apply. OpenAI says it is 10 times faster than the Responses API, again a vendor claim, and plans general availability in the coming weeks, so terms may change. Its request shape differs from Jev’s (field names, an array of questions, type names), so Jev code cannot be pointed at OpenAI unchanged.

What they share

Input-only pricing and typed answers with probabilities are common to Jev, OpenAI, and Clef-flash on Workers AI. The main differences are whether a model can read images and whether it can run locally. Jev and OpenAI are API-only; Clef-flash, Nimble, and Strands Decider also run locally. Clef models and OpenAI read images (Strands Decider accepts them through its base model); Jev and Nimble handle text only.

How 2,170 Public Projects Use It

A paper has already looked at early use. “Jev in the Wild” (arXiv 2609.30216, submitted September 24, 2026) by researchers at Sun Yat-sen University and CUHK analyzed 2,170 public GitHub projects using Jev as of September 22. By type, 81.0 percent used choice, 72.2 percent noul, and 45.4 percent score; projects can use several types, so the total exceeds 100 percent. Of the 2,170, 1,865 were created in the week after launch and 305 were existing repositories that integrated Jev.

Read these numbers with care. The paper is not peer reviewed, and each project’s purpose and types were labeled by a GPT-6 Luna Max agent that read the repositories, so the percentages are themselves LLM classifications. Even so, choice being most common suggests many people are wiring in branches with more than two outcomes, which is relevant to 3D printing too.

Where It Fits in a 3D Printing Workflow

A print involves several “may I print?” and “should I stop?” decisions. Sorting each by whether the AI receives text or an image shows where Jev fits.

StepWhat the AI receivesExample questionTypeCan Jev read it?
Before printing (file)License text and creator notes on the download pageMay the prints be sold?noulYes (text)
Before printing (settings)Filament, nozzle and bed temperature, build plateIs this combination OK to start, and what is wrong?noul and choiceYes (text); compare numbers with rules
During printing (state)JSON with temperatures, targets, progress, errorsContinue, notify, or stop?choiceYes (JSON)
During printing (image)Camera photosWhich defect is visible?choiceNo; use Clef-flash or similar
Deciding to stopProbabilities from the steps aboveDid the probability cross the threshold?(handled in code)A design question, not a decision

Before printing: license text

Model sites attach a license to each design: Creative Commons variants, GPL, licenses the site provides, and sometimes extra conditions in the description. Whether prints can be sold or remixes shared is a judgment on text, exactly the kind of input Jev reads. An AI judgment is not legal advice, though. Its job is to narrow down which files a person should read in full; anything that lands in the middle goes to a human, and to the creator if needed. TypeSafe also says English is the primary training language and CJK text is handled but not as well, so test with your own content. We tested this step in the license article in this series.

Before printing: slicer settings

Bambu Studio publishes official filament profiles on GitHub, including recommended nozzle temperature ranges and bed temperatures per plate. That gives you text data to judge “may I start with this filament, plate, and temperature?” TypeSafe’s list of Jev 1.13 weaknesses matters here: numerical precision, arithmetic, date ordering, closeness of RGB or hex colors, a bias toward the first option, accuracy loss from irrelevant content in the state, and steering text in the state. Its advice is to keep arithmetic in code. So checking whether a temperature is in range belongs in a rule, while reading inconsistent filament names or plate names fits a choice question. See the slicer settings article for the test.

During printing: printer state

Klipper printers expose temperatures, targets, progress, and errors through the Moonraker API. Put that JSON in the state, ask “stop, notify, or continue?” as a choice, and you get a probability for each option plus confidence. A decision model does not replace the printer’s own safety features such as thermal runaway protection, and adding one does not make unattended printing safe. We walk through the design in the printer state article.

During printing: camera images are out of Jev’s reach

Spaghetti, layer shifts, and parts coming off the bed are visible in a photo, but Jev does not accept images. Cloudflare’s Clef-flash fills that gap: the weights are Apache 2.0, and Ollama 0.35.1 or later can run it locally with images. On our RTX 5070 Ti (16,303 MiB), the Q8_0 build used 13,112 MiB of VRAM under the Linux build of Ollama 0.35.1 in Docker (one measurement). The Windows build of Ollama 0.35.1 failed on every Clef-flash decision we tried; a fix was merged on October 7, 2026, but no release with the fix was out that day. Results are in the defect classification article.

Deciding to stop: cut probabilities with a threshold

Because answers come back as probabilities, the code can decide: stop above 0.9, notify between 0.3 and 0.9. Where to cut depends on the cost of each mistake, which the threshold article works out from material cost. Calibration holds across many predictions, not for each one, so check how well the probabilities match your own data before you rely on a threshold.

Our slicer test used the official Bambu Lab P1S profiles. If you print on the same machine, the answer table from that test doubles as a check for your own settings.

USD 499.99 on Amazon.com (as of 2026/09/21)

Limits to Know Before You Use It

  • Early access. The official site still labels Jev early access as of October 7, 2026. The waitlist was reported to have been removed in late September, but we found no official announcement.
  • Text only. No images, audio, or video. Cover image decisions with a model that reads images.
  • “No hallucination” means staying within types and options. It can still choose the wrong option.
  • Speed and cost multipliers are self-reported. Check who measured 40–200×, 193.6×, and 444.6×, from where, and against what.
  • English first. The vendor says Japanese and other CJK text is handled but not as well. Test with your own content.
  • Published weaknesses. Numerical precision, arithmetic, date comparison, color proximity, option-order bias, irrelevant content, and steering text. Keep calculations in code.
  • Calibration is a population property, and confidence is not correctness. Validate thresholds on your data.
  • Rate limits can change without notice. If you poll printer state frequently, build in retries.
  • Not a substitute for printer safety features. Treat stop decisions as an aid.

Most of these come straight from the vendor’s own docs. Reading the relevant pages before you integrate avoids most of the mistakes.

Summary: What to Do Before Putting Jev in Your Workflow

Jev is a decision model that writes no text and returns only typed answers with probabilities. Typed output, input-only pricing, and probabilities each exist elsewhere, but after Jev bundled them into one API, Bespoke Labs, Ollama, strands-labs, Cloudflare, and OpenAI shipped models or endpoints of the same shape within three weeks. For a 3D printing workflow, work through these steps:

  1. Split decisions into text and images. License text, settings, and state JSON suit Jev. Camera images need an image-capable decision model such as Clef-flash, or a trained detector.
  2. Write numeric comparisons and calculations as rules. Do not ask the model whether a temperature is in range; give it only the judgments that are hard to write as rules.
  3. Pick the question type. noul for yes/no, choice for three or more branches, score for degree. Use choice or score if you need confidence.
  4. Set thresholds from the cost of mistakes and check them on your data. A probability is not a per-answer guarantee.
  5. Keep safety features in charge. Stop decisions are an aid; keep a path that notifies a person.

With a decision model in the loop, what you expect from AI shifts from “explain this to me” to “give me the value for this branch.” A good first step is to list the decisions in your own workflow that need a branch value, not an explanation.

Sources

ブラウザだけでできる本格的なAI画像生成【ConoHa AI Canvas】
ABOUT ME
swiftwand
swiftwand
AIを使って、毎日の生活をもっと快適にするアイデアや将来像を発信しています。 初心者にもわかりやすく、すぐに取り入れられる実践的な情報をお届けします。 Sharing ideas and visions for a better daily life with AI. Practical tips that anyone can start using right away.
記事URLをコピーしました