ai-ml

AI Decision Models: Jev, Laya, Clef and More

October 7, 2026

AI Decision Models: Jev, Laya, Clef and More

AI decision models answer typed questions, such as "which team?", "how urgent?" or "yes or no?", with a probability for every answer instead of writing text.

TypeSafe AI named the category "System One models" when it launched Jev on September 15, 2026.1 Within three weeks, Cloudflare, Liquid AI, OpenAI and many open-source projects followed.

The Agentic Edge: AI agents and automation, explained for the people who decide, fund and ship.

TL;DR

Decision models are built for the small, repeated judgments inside a business process: routing a ticket, scoring a lead, approving an invoice. They are designed to be fast and cheap, and they return a confidence number your systems can act on.

You can try several for free today, including open models on Ollama. A recent independent benchmark is encouraging but mixed: a small classifier trained on your own labels can still match or beat them.2

What happened

September 15, 2026. TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, released Jev "in early access." Its launch post says Jev "gives up string generation" and returns "typed probabilistic decisions."1

September 29. OpenAI announced a Decisions API at DevDay. The OpenAI Developer Community's DevDay roundup says it is in "limited preview" and "uses Luna to classify inputs, route requests, or choose an action from predefined answers."3

The same week. Liquid AI released d1, a hosted decision model that works with TypeSafe's software kits and has a free text-only tier, d1:free.4

October 1. Cloudflare released Clef and Clef-Flash, open-source decision models under the Apache 2.0 license, hosted on its Workers AI platform.5

Meanwhile, open source exploded. A community list, awesome-system-one, tracked 27 decision models as of October 2, 2026. It warns that their "benchmarks are self-reported and not directly comparable."6

What is a decision model, in plain English?

A large language model (LLM), like ChatGPT, writes text for people to read. A decision model never writes text. You give it an input, such as a support ticket, and a set of questions with fixed answer options.7

It returns the chosen answer plus a probability. TypeSafe's three question types are Choice (pick an option), Score (rate on a scale) and Noul (is this statement true?).7

Why the confidence number matters: your software can act on its own when the model is confident and send the case to a person when it is not. TypeSafe describes this as setting "the thresholds for when it acts autonomously and when it asks for review."8

How a decision model fits a workflow: input, typed questions, probabilities, then your threshold decides to act or escalate.

Figure 1: how a decision model fits into a workflow. Example values are illustrative. Built from TypeSafe's docs and Cloudflare's Clef post.

The decision-model landscape

Landscape of AI decision models in October 2026: hosted (Jev, Liquid AI d1, OpenAI Decisions API, Cloudflare Clef on Workers AI) and open (Clef, Clef-Flash, Laya, Tev1, Nimble, Kev, GLiNER2.5-Decide).

Figure 2: the main models by access type. Not a complete list.

Hosted models (pay per use)

  • TypeSafe Jev (docs, OpenRouter listing). The model that started the category. Price: $0.042 per million input tokens, with output "FREE."1 TypeSafe has not published Jev's backbone or detailed architecture.2
  • Liquid AI d1 (OpenRouter listing). Uses the same question format as Jev, with a free text-only version and a paid version that accepts images.4 OpenRouter lists it at $0.04 per million input tokens.9
  • OpenAI Decisions API. Limited preview, built on GPT-6 Luna. OpenAI had not published pricing as of October 7, 2026.36
  • Cloudflare Clef on Workers AI. Hosted versions of Cloudflare's open models. Cloudflare says Clef reads images, unlike Jev, and has a 64k context window against Jev's 32k.5

Open models you can download

  • Cloudflare Clef (27 billion parameters) and Clef-Flash (9 billion). Apache 2.0. On Ollama as clef and clef-flash.510
  • Laya by Convai Innovations (GitHub). Very small: 421 million parameters for English and 322 million for 100+ languages. On Ollama as laya.11
  • Tev1 by Together AI. An experimental 4B model; its weights license is still "being finalized." On Ollama as tev1.1012
  • Bespoke Nimble by Bespoke Labs (9B). On Ollama as nimble.10
  • Kev by Jared Palmer. Sizes from 0.8B to 27B, Apache 2.0, built to work with TypeSafe's software kit unchanged.13
  • GLiNER2.5-Decide by Fastino. A 340M-parameter classifier for intent routing, sentiment and triage, Apache 2.0.14

A parameter, in one line: a rough measure of model size. Smaller models are cheaper and faster to run, and Kev's smallest 0.8B version "runs on a laptop."13

Where to run them locally

Ollama, an app for running AI models on your own computer, now has a "Decision" model category. On October 7, 2026, that category listed five models: clef-flash, clef, laya, tev1 and nimble. Community uploads with similar names also appear in its search.10

Screenshot of Ollama's model search for "decision", showing clef-flash, clef, laya and tev1 with the Decision tag.

Screenshot: ollama.com model search, captured October 7, 2026.

Each one is a one-line download in Ollama. Check the model page for the minimum version: the Laya page says it "requires Ollama 0.40.0 or later."11 There is also Ollaya, a separate open-source runner for decision models that says it is "not affiliated with or endorsed by Ollama or TypeSafe."15

What the benchmarks actually show

There are two kinds of evidence: vendor benchmarks and independent research. This post uses one recent independent study. The two do not always agree.

Vendor claims

TypeSafe's homepage claims Jev is "193.6x Faster, 444.6x Cheaper" than LLMs on its own workflow tests.8 Its launch post adds the caveat that "we expect that these are on the higher end of real world gains," and that the tests were built by its own team, "so some bias could exist."1

Screenshot of TypeSafe's "Workflow Intelligence vs. Cost" chart, plotting Jev against frontier LLMs on accuracy and average cost per workflow.

Screenshot: typesafe.ai homepage, captured October 7, 2026. Vendor-reported results.

Cloudflare published its own comparison. In its table, Clef beats Jev on most of the 10 benchmarks it lists, and Clef-Flash's median latency is 38.8 milliseconds against 524.1 for Jev.5

Screenshot of Cloudflare's benchmark table comparing Clef, Clef-Flash, Jev, DiffusionGemma Jev, Kev 9B and Laya on 10 decision benchmarks.

Screenshot: Cloudflare blog, captured October 7, 2026. Vendor-reported results, measured by Cloudflare.

An independent benchmark

On September 29, 2026, researchers Amir Rafe and Subasish Das posted a study on arXiv comparing eight decision-model checkpoints, including Jev, with trained classifiers and general LLMs. It is a preprint, and its arXiv page lists no peer-reviewed publication yet.2

Screenshot of Table 5 from arXiv paper 2610.00346, showing accuracy for Jev, Laya, Kev, decider-2b, this-that-model, Nimble-9B and Qwen comparators.

Screenshot: arXiv 2610.00346, Table 5 (accuracy), captured October 7, 2026.

Four findings matter for business readers:2

  1. Jev led the decision models on accuracy. It scored 0.732 on the typed-decisions workflow set and 0.913 on the CLINC-150 intent set.
  2. A small trained classifier is still hard to beat. A BGE-small classifier trained on each task's labels scored 0.702 and 0.948. The authors conclude that "small trained classifiers are the most accurate on intents."
  3. Labels matter. "Swapping yes and no flips 50.5 answers per hundred for Jev." How the answer options are named can change the answer.
  4. A two-stage setup can cut the bill by more than half. An intent-trained first stage that escalates hard cases to Jev "matches its accuracy at 0.43 of its cost" under the study's assumption of a fully used GPU.

The same study measured speed and cost. Jev answered in a median of 138 milliseconds, at $0.0079 per thousand decisions. The open Laya multilingual model took 32 milliseconds, at $0.0009 per thousand, on a rented cloud GPU.2

Screenshot of Table 8 from arXiv paper 2610.00346, showing cost per thousand decisions and latency for each model.

Screenshot: arXiv 2610.00346, Table 8 (cost and latency), captured October 7, 2026.

Note the gap between sources: Laya scores very low in Cloudflare's table and 0.366 on typed decisions in the independent study.52 Its own model card reports different tasks.11 That is the clearest reason to test on your own data.

Use cases by role

These are our suggestions, built on the examples vendors give. Cloudflare shows a support ticket being checked for urgency, routed to a team and scored for severity, and an invoice being classified as paid, overdue or draft.516 TypeSafe lists routing, scoring, "map-reducing over big data," and checking LLM outputs, including detecting jailbreaks.1

Decision-model use cases by role: support, sales, finance, security, ops, marketing, HR and AI agent owners.

Figure 3: use cases by role. Our suggestions, informed by vendor examples.

  • Customer support. Ask three questions of every ticket: is it urgent, which team owns it, and how severe is it? Auto-route the confident ones.
  • Sales and RevOps. Score inbound leads and tag each email's intent, such as demo request, pricing question or churn risk.
  • Finance and accounts payable. Classify invoice status and decide "approve" or "send for review." Keep a person on anything above a value limit.
  • Security and IT. Cloudflare says it has been testing Clef with its Threat Intelligence team to classify website domains. Clef took 2.2 seconds against 4.7 seconds for gpt-oss-120b in the same workflow.5
  • Operations and delivery. Route internal requests, flag work at risk of missing its deadline, and score backlog items for priority.
  • Marketing and communications. Tag sentiment and topics across reviews or social posts, and screen drafts against your content policy.
  • HR and people teams. Route employee questions to the right policy owner. Do not use these models for hiring or performance decisions about individuals.
  • AI agent owners. Before an AI agent takes an action, ask a decision model whether the action fits your policy, and escalate when it is unsure.

Why it matters

For executives

Our read: teams often use a large, expensive model for small yes-or-no judgments. Decision models aim to make those judgments a cheap line item. The bigger win may be control, because every decision comes with a confidence score you can audit.

For product managers

Look for what TypeSafe calls "smart if-statements" in your product: places where code makes a judgment that is too fuzzy for fixed rules.1 TypeSafe's docs recommend splitting a big judgment into small questions and combining the answers in code.7

For project and delivery leads

These are easy to pilot. Pick one high-volume decision, run a model next to your current process for two weeks, and compare. Set a confidence threshold and send everything below it to a person.

What to do this week

  1. List your repeated decisions. Find three judgments your team makes hundreds of times a week, such as ticket routing, lead scoring or invoice checks.
  2. Gather 200 past examples with the correct answer for one of them. This becomes your private test.
  3. Try a free option. Ask a technical colleague to test an open model from Ollama or the free Liquid d1 tier against your examples.
  4. Compare against a simple classifier. The independent study shows a small trained model can match these on narrow tasks.2
  5. Decide your threshold. Agree what confidence level lets the system act alone, and who reviews the rest.

Questions to ask your team or vendor

  1. What accuracy does the model reach on our examples, not public benchmarks?
  2. How often does it change its answer if we reword the question?
  3. What happens to inputs that fit none of the options?
  4. Where is the data processed, and is it used for training?
  5. What does it cost per thousand decisions at our volume?
  6. Can we run it ourselves, and what is the license?
  7. Who reviews low-confidence cases, and how quickly?

Risks and what we don't know yet

Benchmarks are mostly self-reported. The community tracker says vendor benchmarks are "not directly comparable."6 Independent research is still early: the study cited here is a preprint, and by its own review most earlier studies covered fewer models or test conditions.2

"Can't hallucinate" is not "can't be wrong." TypeSafe says Jev "can't hallucinate" because it always answers in the allowed format.1 It can still pick the wrong option, and the independent study found that, at a threshold set for 5% risk, Jev still accepted 0.310 of out-of-scope requests.2

Early products. Jev launched in early access and OpenAI's API is in limited preview.13 Prices may change: TypeSafe says "We can't prove it isn't subsidized."1

Decisions about people. Automated decisions about employees, customers' credit or similar areas can be regulated. Keep a person responsible. Not legal advice.

Our view: this is one of the most practical AI trends of the year for business teams. Start with one narrow, high-volume decision and measure it.

Bottom line

Decision models turn a vague "use AI" goal into something measurable: one decision, one confidence score, one threshold. Several are free to try today.

Test them on your own examples before trusting any benchmark, and keep a person on the cases the model is unsure about.

Not financial or legal advice.

References

Footnotes

  1. Introducing System One Models & Jev — Diogo Almeida, TypeSafe AI, September 15, 2026. Source for the launch date, early access, pricing, "gives up string generation", "can't hallucinate", the workflow-eval caveats, use cases and the subsidy statement. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11

  2. Benchmarking System One decision models against trained classifiers and language models for automated decision gates — Amir Rafe and Subasish Das, arXiv preprint 2610.00346, submitted September 29, 2026. Source for Tables 5 and 8, the accuracy, cost and latency figures, the yes/no swap result, the out-of-scope result and the cascade result. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10

  3. DevDay 2026 announcements and developer resources — OpenAI Developer Community, September 29, 2026. Source for the Decisions API's limited preview and description. ↩ ↩2 ↩3

  4. Decision models — Liquid AI documentation, accessed October 7, 2026. Source for d1, d1:free and TypeSafe SDK compatibility. ↩ ↩2

  5. Introducing Clef: our open-source decision models, and new RL fine-tuning platform — Michelle Chen, Alex Reneau and Kevin Flansburg, Cloudflare Blog, October 1, 2026. Source for Clef and Clef-Flash, Apache 2.0, Workers AI, the vision and context comparison, the benchmark and latency table, the support-ticket example and the threat-intelligence example. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8

  6. awesome-system-one — community-maintained list, last checked October 2, 2026. Source for the count of 27 models, the self-reported benchmark warning, and the note that OpenAI's price was not yet published. ↩ ↩2 ↩3 ↩4

  7. Introduction — TypeSafe AI documentation, accessed October 7, 2026. Source for the Choice, Score and Noul question types and the advice to decompose questions. ↩ ↩2 ↩3

  8. TypeSafe AI homepage — accessed October 7, 2026. Source for "193.6x Faster, 444.6x Cheaper", the threshold quote and the Pareto chart screenshot. ↩ ↩2

  9. Liquid AI: d1 — OpenRouter, accessed October 7, 2026. Source for the d1 price listing. ↩

  10. Ollama Decision models, Ollama search: decision and the clef, clef-flash, tev1 and nimble pages — Ollama, accessed October 7, 2026. Source for the Decision category, the five models and their sizes. ↩ ↩2 ↩3 ↩4 ↩5

  11. convaiinnovations/laya — Hugging Face model card, and laya on Ollama, accessed October 7, 2026. Source for the sizes, languages and the Ollama version requirement. ↩ ↩2 ↩3 ↩4

  12. Tev1-4B-experimental — Together AI on Hugging Face, accessed October 7, 2026. Source for the base model and license status. ↩

  13. Kev — Jared Palmer on GitHub, accessed October 7, 2026. Source for sizes, the "runs on a laptop" quote, license and TypeSafe SDK compatibility. ↩ ↩2

  14. GLiNER2.5-Decide — Fastino on Hugging Face, accessed October 7, 2026. Source for the size, license and use cases. ↩

  15. Ollaya — GitHub, accessed October 7, 2026. Source for the independence statement. ↩

  16. Cloudflare/clef — Hugging Face model card, accessed October 7, 2026. Source for the invoice-status example. ↩

Frequently Asked Questions

A model that answers typed questions about an input, such as choosing an option, scoring on a scale or answering yes or no, and returns a probability instead of generated text. TypeSafe calls them "System One models."1