Cloudflare Clef: Fast Decision Models for AI Agents
Cloudflare has shipped two models that make choices instead of writing prose. Clef and its smaller sibling Clef-flash are "decision models" for AI agents. Cloudflare's main argument for them is speed, measured against the current reference point in this small category, TypeSafe AI's Jev.
The company also makes a bolder claim. In its words, "a human does not necessarily need to be in the loop for agentic decisions anymore."
What a decision model does
A decision model does not answer with a paragraph. It takes an input, picks from a fixed set of options and attaches a probability to each one. Cloudflare's example is a customer support message. Clef rates how urgent the message is and which team should handle it. Code further down the pipeline then acts on that output. It can route the ticket, escalate it, or pass it to a person.
According to Cloudflare, agents can "programmatically gather context, make decisions, and take actions on tasks, or defer to a human when needed." The name comes from music notation, where a clef sets the pitch of each line on a staff. The model plays a similar role by setting the frame for whatever action follows. The name also sounds a bit like "Jev," which is probably intentional.
The category sits between two familiar tools. Large language models can reason and call tools, but their outputs vary and they can be slow. Classic classifiers are fast, but they need retraining for every new category. Cloudflare positions Clef as a direct rival to Jev and keeps its API fully compatible, so customers can switch with little effort.
The speed numbers
Cloudflare says Clef and Clef-flash are faster than all relevant competing decision models across 43 benchmarks. Its reported median latencies:
- Clef-flash: about 39 milliseconds
- Clef: about 209 milliseconds
- Jev: just over 524 milliseconds
Both models run on Cloudflare's own infrastructure. The company says this gives them an extra advantage because they sit close to its edge data centers.
Cloudflare's threat intelligence team is already testing Clef to classify websites. In one example, Clef gave a domain a 95 percent probability of being a fashion site and 85 percent of being an online shop. The probability that it was a phishing site was under one percent. Fetching, rendering and classifying the site took 2.2 seconds. Cloudflare's fastest general-purpose language model needed 4.7 seconds for the same job and returned only two categories.
Clef also accepts images, while Jev handles only text so far. Its 64,000-token context window is twice the size of Jev's. Clef also scores ahead on the Jev Decision Index, although those results come from Cloudflare's own benchmarks.
Built on Qwen
Clef is based on Qwen3.8-27B, and Clef-flash on the smaller Qwen3.5-9B. Cloudflare leaves the base models unchanged during training. It uses its own synthetic data to train additional components, which derive answer options and probabilities from the models' internal computations.
Cloudflare also uses its own version of Reinforcement Learning for Calibrated Decisions (RLCD), the method TypeSafe used to train Jev. RLCD trains a model to answer several questions about one input in a single call. The goal is that the assigned probabilities match how often the answers are actually correct. Before this, Cloudflare had experimented with DiffusionGemma to extract fixed decision values from a language model's internals.
Fine-tuning for customers
Cloudflare is also launching a reinforcement learning service so customers can adapt Clef to their own tasks. At first, forward deployed engineers will run the fine-tuning together with customers. A self-service platform is planned for later.
The process has three steps. Customers log requests through AI Gateway to build a dataset. They then evaluate those requests in containers that serve as an RL sandbox. Finally, a new trainer component deploys the tuned model on Workers AI. Custom models run on technology from Replicate, which Cloudflare acquired in late 2025.
Internally, Cloudflare plans to use Clef to review abuse reports, sort support requests and separate useful bots from harmful ones. Both models are available on Workers AI and on Hugging Face under the Apache-2.0 license.
A niche that filled up fast
TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, popularized the approach when it introduced Jev in mid-September. TypeSafe markets Jev as a model "without hallucinations." That only guarantees that Jev stays inside the predefined answer options. It can still pick the wrong one. In late September, OpenAI followed with a Decisions API built on GPT-6 Luna, which also accepts text or images as context.
Cloudflare mainly runs a global network for content delivery, DNS and security services, and its own AI models have drawn little attention so far. Its recent AI news has focused on giving website owners control over access. In July, it let site operators block or allow AI bots based on their purpose.
Our Take
Several launches in a few weeks, plus AWS's open decision model for agents, suggest decision models are becoming a standard building block. For teams that run agents, a bounded set of outputs with probabilities is much easier to wire into code, test and audit than free-form text.
The "no human in the loop" claim needs some caution. A model that cannot invent answers can still choose the wrong one, as TypeSafe's own caveat shows. Calibration is what matters here: a 95 percent score only helps if it is right about 95 percent of the time. As our coverage of AI bookkeeping that still needs oversight showed, better accuracy rarely removes the need to review high-stakes actions.
Two things are worth watching. One is whether independent tests check calibration and not only latency. The other is whether the open Apache-2.0 weights draw developers away from Jev.
