# Cloudflare Releases Clef, Its First In-House Decision Models, on Workers AI

Cloudflare's Workers AI team published two open-weight decision models that return probabilities for typed questions instead of generated text, alongside an RL fine-tuning service it is recruiting design partners for.

Canonical URL: https://freelancenews.online/news/cloudflare-releases-clef-its-first-in-house-decision-models-on-ba3c7cb2
Published: 2026-10-02T18:18:50.491Z
Updated: 2026-10-02T18:18:50.491Z
Source published: 2026-10-01T00:00:00.000Z
Event date: Not established
Review status: source-reviewed
Review method: Automated comparison against retrieved source text; not independent fact-checking.

## Report

Cloudflare has released Clef and Clef-flash, described in its changelog as the first models trained by the Cloudflare Workers AI team and the company's first open-source decision models. Both are available on Workers AI under the identifiers @cf/cloudflare/clef and @cf/cloudflare/clef-flash, and the weights are published on Hugging Face under the Apache 2.0 license.

Text generation is not what these models do. Clef, per the changelog, takes an input state plus a set of typed questions and, for each permitted answer, produces a probability, meaning an agent gets a structured decision usable directly. Cloudflare's stated position is that nothing free-form must be parsed and no reasoning tokens must be awaited, which is central to how it contrasts with generative models employed for classification or routing.

Cloudflare says the System One interface is what the API follows, so switching an existing Jev integration to Clef requires only a change of endpoint and model. Up to 64 questions, drawn from three types, fit in one request: noul, a yes/no question that yields the probability of yes; choice, selecting one option from a caller-defined set and returning the chosen option along with a per-option probability and a confidence value; and score, rating against an ordered rubric and returning a probability-weighted score plus a per-level probability.

The changelog's example shows the shape of a call: a state string describing a failing checkout, a noul question about whether the request is urgent, and a choice question assigning the ticket to billing, technical or sales. The response exposes the urgency probability and the highest-probability team. Cloudflare lists support triage, threat intelligence, trust and safety scoring, agent guardrails and visual classification as intended uses.

Turning to performance, Cloudflare reports a median speed advantage of 2.5x for Clef over Jev across 43 benchmark runs, with Clef-flash reaching 13x. A Clef model, it further says, ranked highest on 7 of 10 decision benchmarks, beating Jev and other open decision models, and on Typesafe's own workflow evals Clef outperformed Jev in three of four areas: invoice processing, customer service and security incidents. Vendor-reported figures are what these are; for full results the changelog directs readers to the Hugging Face model card, and the supplied text cites no independent evaluation.

For a compound workflow, one concrete comparison appears: with Browser Run, a domain was fetched, rendered and classified by Clef in 2.2 seconds, versus 4.7 seconds for gpt-oss-120b in the same workflow. A single reported scenario, not a general speed claim, is what that represents, and the hardware, network conditions or sample of domains behind it are not described in the changelog.

Clef also accepts up to four images alongside the state through a vision encoder, which Cloudflare contrasts with text-only decision models. Hosting on Workers AI means requests run on GPUs across Cloudflare's network, which the company says keeps the network round trip short and lets Clef sit directly in an agent's request path before handing off to an LLM on the same platform for action.

Access is through the Workers AI binding via env.AI.run() or the REST API at /ai/run, and Cloudflare says AI Gateway can be used with these endpoints. The changelog directs readers to the Clef and Clef-flash model pages and to pricing, but the supplied text does not state what that pricing is, nor does it give a release date beyond the changelog entry itself.

Cloudflare is also launching a reinforcement learning fine-tuning service to tune Clef for specific workloads and is recruiting design partners. The changelog does not describe the service's mechanics, cost, availability or timeline, so it should be treated as an announced offering rather than a generally available product.

For freelancers and small studios building agent features, the practical appeal is the shape of the output rather than raw model quality. A probability per allowed answer is directly usable in branching logic, so a triage or moderation step can be wired into an existing pipeline without writing a parser for model prose or budgeting latency for reasoning tokens. The System One compatibility claim matters for anyone who already built against Jev, because it suggests a migration path that does not require rewriting the surrounding integration.

The tradeoff is scope. A decision model only answers the questions it is given, so the quality of a routing or scoring system shifts onto how well the caller defines the question set, the option criteria and the rubric. The changelog's own example embeds the team definitions in the request, which means that logic lives in application code and must be maintained there. Generative models remain the better fit when the task requires explanation, drafting or open-ended output, which is why Cloudflare frames Clef as a step before an LLM rather than a replacement.

The open-weight release under Apache 2.0 is the part with the longest shelf life for this audience. It means the models can be self-hosted or adapted without depending on Workers AI availability, subject to the usual caveats about running inference yourself. Cloudflare's own hosting pitch, however, is explicitly about latency and proximity to users, so the two paths trade operational control against the speed advantage the company is claiming.

What remains unknown is substantial. The supplied text gives no independent benchmark, no pricing, no rate limits, no context or state size limits beyond the 64-question cap and four-image cap, and no detail on the fine-tuning service. The speed and accuracy comparisons come from Cloudflare and, in the Typesafe case, from evals Typesafe itself publishes, which the changelog cites as the basis for a favorable comparison rather than as neutral ground.

The honest reading is that Clef is a credible addition to the small category of structured decision models, released with unusually permissive licensing and a documented API surface, but its headline numbers are vendor-reported and its fine-tuning service is not yet described in enough detail to plan around. Developers evaluating it should test it against their own question sets and rubrics rather than the published benchmarks, since the model's usefulness depends heavily on how those are written.

## Key points

- Clef and Clef-flash are Cloudflare's first in-house Workers AI models and its first open-source decision models, with weights on Hugging Face under Apache 2.0.
- The models return probabilities for typed questions (noul, choice, score) rather than generated text, with up to 64 questions per request and support for up to four images.
- Cloudflare reports Clef at 2.5x and Clef-flash at 13x faster than Jev at the median across 43 benchmark runs, and a Clef model scoring highest on 7 of 10 decision benchmarks.
- The API follows the System One interface, so Cloudflare says existing Jev integrations can switch by changing the endpoint and model.
- A reinforcement learning fine-tuning service is announced with design partner recruitment, but no mechanics, pricing or availability are described.

## Practical implications — editorial interpretation

For developers building agent routing, triage or moderation steps, Clef's probability-per-answer output can be wired into branching logic without parsing model prose or waiting on reasoning tokens, and the System One compatibility claim offers a low-friction path for anyone already on Jev. The Apache 2.0 weights also allow self-hosting if Workers AI availability or pricing becomes a constraint.

## Limitations and unknowns

All speed and accuracy figures come from Cloudflare's changelog, with the Typesafe comparison drawn from Typesafe's own evals; no independent evaluation is cited. The supplied text gives no pricing, rate limits, state size limits or release date, and the RL fine-tuning service is described only as an announcement with design partner signup. The 2.2-second versus 4.7-second Browser Run comparison is a single reported workflow with no stated conditions.

## Sources

- [1] developers.cloudflare.com: Changelog
  https://developers.cloudflare.com/changelog/post/2026-10-01-clef-workers-ai/
  Retrieved: 2026-10-02T18:18:21.249Z

## Claim references

- Clef and Clef-flash are the first models trained by the Cloudflare Workers AI team and are available on Workers AI, with weights open-sourced under Apache 2.0 on Hugging Face. [source 1]
- Clef reads an input state and typed questions and returns a probability for every allowed answer, with no free-form output to parse and no reasoning tokens. [source 1]
- Cloudflare reports Clef at 2.5x faster than Jev at the median and Clef-flash 13x faster across 43 benchmark runs. [source 1]
- Across 10 decision benchmarks a Clef model scores highest on 7, and on Typesafe's workflow evals Clef beats Jev in three of four areas. [source 1]
- Clef follows the System One API, allowing an existing Jev integration to switch by changing the endpoint and model, with up to 64 questions per request in three types. [source 1]
- Paired with Browser Run, Clef fetched, rendered and classified a domain in 2.2 seconds versus 4.7 seconds for gpt-oss-120b in the same workflow. [source 1]
- Clef accepts up to four images alongside the state via a vision encoder, and is accessible through the Workers AI binding or the REST API at /ai/run. [source 1]
