Releases

Cloudflare Releases Clef, Its First Open-Source Decision Model

Cloudflare logo above the word Clef on a dark orange and blue background.
Cloudflare Clef, an open-source decision model designed for fast and structured AI decisions.

Cloudflare has released Clef and Clef-flash, its first open-source decision models.

Both models are available on Cloudflare Workers AI, while their weights are also available publicly under the Apache 2.0 license, meaning developers can download them and experiment with them locally.

You might be wondering what a decision model actually is, since it works very differently from a normal LLM. The short answer: it makes decisions.

Decision Models

Most AI models we talk about are generative models such as ChatGPT and Claude. You give them an input, and the model starts to generate text token by token.

A decision model, on the other hand, works differently. Instead of asking the model to write an answer, you give it some information and a set of possible decisions, and that little guy chooses one. It’s very similar to a student solving a multiple-choice question.

Let’s say your website’s 404 traffic suddenly spikes, but you’re not sure if you really need bot protection or not. So you ask Clef a simple question:

Should we enable bot protection? Yes or no?

Instead of generating a whole paragraph explaining its reasoning, Clef returns probabilities for the allowed answers. Your application can then use Clef’s decision to take the next step, such as enabling bot protection if that’s how you programmed it.

Cloudflare describes this as one of the main advantages of decision models: there is no free-form answer that your application has to parse before doing something with it.

Clef and Clef-flash

There are basically two versions of Clef, and one of them is Clef-flash, which is faster and has a smaller parameter count.

Clef is the larger 27B model, designed for higher-precision decisions. Clef-flash is a smaller 9B model, designed for situations where latency matters more.

Both models have a context window of around 64K tokens and can work with more than just text. They can accept text and images as input thanks to a built-in vision encoder.

So the main differences between the two models are their parameter count, speed, and the trade-off between decision accuracy and latency.

Why Not Just Use an LLM?

You can technically ask a normal AI chatbot to make the same decision. If you tell it to answer only Yes or No, most modern LLMs will probably do exactly that.

The difference is mostly about how the model handles the task. A normal LLM is designed to generate text, while Clef is specifically designed to choose between a set of possible answers.

Instead of simply generating Yes, Clef can score the available options and give you something like:

Yes: 92%
No: 8%

That becomes more useful when your application needs to make lots of decisions, compare confidence between different options, or make those decisions with lower latency.

And if you’re wondering, yes, a normal LLM can also give you percentages if you ask for them. For example, it could say Yes: 85% and No: 15%. The difference is that those numbers can just be part of the generated response and don’t necessarily represent the model’s actual decision score. A decision model like Clef is specifically built to score the available choices directly, so those probabilities are part of how it makes the decision.

Clef Shines in Speed

Cloudflare reports a median latency of around 209 ms for Clef and 38.8 ms for Clef-flash. Typesafe’s Jev came in at around 524 ms on the same setup, which makes Clef roughly 2.5× faster and Clef-flash roughly 13× faster.

Jev isn’t the only rival, though. In the same table, DiffusionGemma Jev (84 ms) and Kev-9B (51 ms) are faster than the full Clef, and Laya (5.8 ms) is faster than both Clef models, although its accuracy is far lower in Cloudflare’s tests.

These are Cloudflare’s own benchmark runs, so treat them as a starting point and test on your own workload.

Clef Can Run Fully Locally

Cloudflare didn’t keep Clef locked inside Workers AI. The weights for both Clef and Clef-flash are publicly available on Hugging Face under the Apache 2.0 license, so developers can download the models and experiment with them locally.

Clef uses Qwen3.8-27B as its backbone, while Clef-flash uses the smaller Qwen3.5-9B backbone. Cloudflare then post-trained both models specifically for decision-making workloads.

Fine-Tuning Service

Cloudflare is also offering a reinforcement-learning fine-tuning service for Clef. For now it’s a hands-on service with Cloudflare’s own engineers, with a self-serve platform planned for later. The idea is that companies can adapt the model to their own decision-making workloads instead of relying only on the general model.

For example, a company could fine-tune Clef for its own support categories, security policies, internal workflows, or any other task where it needs the model to make specific decisions.

Conclusion

Clef might be a great tool for a larger AI ecosystem and AI agents. You might use it for specialized operations. Your LLM might generate information, Clef can decide something, and the agent takes an action, and so on.

Sources

Cloudflare: Introducing Clef

Cloudflare Workers AI: Clef

Clef on Hugging Face

Updated Oct 4, 2026