Skip to content
OTFotf
All posts

Can you swap a Jev decision call for Clef on Workers AI?

D
DaveAuthor
7 min read
Can you swap a Jev decision call for Clef on Workers AI?

Cloudflare says Clef and Clef-flash are fully Jev-API compatible, so you can make the swap extremely easily. That one claim is what makes Cloudflare’s 1 Oct 2026 blog post worth a careful read for anyone already calling a Jev-style decision model. Whether it holds for your application is something only a trial on your own data can answer; we would test this way before trusting it in a workflow.

The post introduces two Cloudflare-trained decision models hosted on Workers AI, weights open-sourced under Apache 2.0 on Hugging Face, and a fine-tuning offer. In short, it is a second vendor’s model offering Jev’s API, running on a platform many builders already use. For background on Jev itself, see our System One and Jev overview; this post focuses on the additions Cloudflare lists. Builders on Cloudflare runtimes may also know our post on Worker Previews isolating Durable Objects, which covers a separate CI concern.

What a decision call does

Cloudflare describes a decision model as one that classifies inputs and gives back answers with a type and a probability attached. Its example is a support message, asked whether it is urgent and which team should own it. Your code can then route the ticket, trigger an escalation, or defer to a human. The last option matters: a result is one input to a workflow, and a human can take over when the odds or context call for it.

Cloudflare also describes its own use: its Threat Intelligence team has been testing Clef, paired with Browser Run, to classify website domains. For one domain it cites likelihoods such as 95% fashion website, 85% ecommerce, and under 1% phishing. It says Clef needed 2.2s to fetch, render, and classify, while gpt-oss-120b, its fastest general LLM, needed 4.7s in the same workflow and produced only two classifications. That is Cloudflare’s internal anecdote, not a controlled test.

The page shows a request with a state and several questions, using question types written noul (spelled that way on Cloudflare’s page), choice, and score. I’m not reproducing that request here because the captured page’s formatting does not preserve a dependable verbatim command block. The boundary that matters here is the one Cloudflare states: Clef is fully API-compatible with Jev and its outputs are strictly typed.

What Cloudflare says differs from Jev

Cloudflare names two model differences. Clef has a vision encoder for classifying image content, while the page says Jev does text classification today. Clef also has a 64k context window, against Jev’s 32k. Both are Cloudflare’s statements, not measurements made here.

Sharing an API can make an experiment simple at the call boundary, but it does not show that your decisions will come out the same. Your schemas, thresholds, inputs, and handling of uncertain results still shape what the application does. The ease-of-swap claim is Cloudflare’s; checking your own call sites and outcomes is our advice, not something the page offers.

Two people slide a module into a device while spare modules wait in a tray.

Same component. Web and mobile. One codebase.

The free, open-source SDK gives you components that work the same on web and mobile — one codebase. github.com/otf-kit/sdk

Get the free SDK

Reading Cloudflare’s benchmark tables

Cloudflare publishes results for Clef and Clef-flash next to Jev and other models. These are vendor-reported figures from Cloudflare’s own tables, run by Cloudflare and not independently verified here. They give comparison points, not a forecast for your workload.

In Cloudflare’s own table, Clef reports 98.47 on BFCL case exact and Jev reports 95.75. The same table goes the other way on When2Call accuracy: Jev reports 80.97 and Clef 72.37. For BRIGHT nDCG@10, Cloudflare reports Jev at 47.52 and Clef at 45.91. Rows like these show why a headline about leading an index says little about which model suits one particular decision.

Cloudflare’s table for Typesafe’s evaluation suite also has Jev ahead on Agent trace observability: Cloudflare reports Jev at 71.6 and Clef-flash at 69.8. In that suite Cloudflare reports Clef-flash at 77 for customer service and Jev at 76.0. Cloudflare says its models beat Jev in three of four areas there; the rows themselves are what to read.

Cloudflare also says the Clef models were faster across the 43 benchmarks it ran, with one exception, a model Cloudflare calls very fast but weaker on quality in its earlier tables. In its latency table the median is 209.3 ms for Clef, 38.8 ms for Clef-flash, and 524.1 ms for Jev. It credits Workers AI’s edge GPUs and describes low network latency. Those are Cloudflare’s findings and Cloudflare’s explanation. We would test this way: send identical real inputs to each model and compare the outcomes your application depends on, rather than relying on a vendor table.

Data handling, weights, and fine-tuning

Cloudflare says it does not read, store, or train on your requests or responses, unless you choose to use its fine-tuning product. It also says the weights are open-sourced on Hugging Face under Apache 2.0, so developers can run them locally. Whether that fits your own data requirements is yours to review; the review is our advice, not a separate Cloudflare feature.

On fine-tuning, Cloudflare says it is starting with a hands-on service run with its forward-deployed engineer team. A self-serve platform for capturing data, fine-tuning, and redeploying a model on Cloudflare comes later. The post lays out that order; it does not describe a self-serve workflow you can use today.

Cloudflare also suggests pairing Clef with one of its LLMs on Workers AI: Clef makes a structured decision in an agent’s hot path and the LLM takes the action. That is a layout the page proposes, not a requirement to use both models.

What Cloudflare says is under the hood

Cloudflare says Clef uses Qwen as its base model, running a prefill-only pass and then scores every valid schema option in parallel. Because that decision step is non-autoregressive, it says, there is no intermediate text generated token by token. It names Qwen3.8-27B as the frozen backbone for Clef and Qwen3.5-9B for Clef-flash, tuned with rank-256 low-rank adapters. Training combined label-smoothed cross-entropy with a Brier loss, on Cloudflare’s own synthetic datasets. These are descriptions of the vendor’s method; we have not examined the models.

Two people feed the same blank cards into two matching compact devices side by side.

A swap test plan (our advice, not Cloudflare’s)

  1. Pin the exact model name your application calls, so the target is explicit.
  2. Put the decision request behind a thin adapter in your repository, so two interchangeable implementations can be compared side by side.
  3. Worth checking, in our view: run Jev and Clef on the same labelled samples from your own workload, and record the outcomes that matter to your product.
  4. Treat Cloudflare’s tables as a starting point only, and read its data-handling claim against your own requirements.
  5. Before switching, decide what a low-probability result does. Cloudflare’s own example includes deferring to a human; make that path deliberate.

Gaps in Cloudflare’s announcement

The post does not state a Workers AI price, rate limits, model-size or latency guarantees beyond its tables, a GA or beta status, or a migration guide. It covers API compatibility and reports benchmark results. It gives no migration procedure and makes no promise that every application will behave identically after a swap.

Where a decision call would live

A decision call has to sit inside some app. The Arcade Kit page lists a Live Games Showcase Kit for a one-time $99, built on React Native and Expo, with a five-tab games store, hero carousels, and a player profile. The page says the source is on GitHub and the demo stays free, so you can click through the live demo before buying. Arcade contains no Clef or any other decision model; you would add that call yourself. The pricing page counts it among three Live full-stack kits, the other two being Fitness & Wellness and the SaaS Dashboard kit.

FAQ

Is Clef API-compatible with Jev?

Cloudflare says Clef and Clef-flash are fully Jev-API compatible and calls the swap easy. Our advice is to run your own samples before you switch.

Can Clef classify images?

Cloudflare says Clef has a vision encoder for image classification. Its page describes Jev as text-only today.

Can I fine-tune Clef myself today?

Cloudflare outlines a hands-on fine-tuning service first and a self-serve platform later. It does not say that the self-serve platform is available now.

Does Cloudflare give a Workers AI price or rate limit?

No. The page does not state a Workers AI price or rate limits.

Sources

Primary source, read on 3 Oct 2026: https://blog.cloudflare.com/clef-decision-models/

OTF pages read on 3 Oct 2026: https://otf-kit.dev/pricing ; https://otf-kit.dev/templates/arcade-kit ; https://otf-kit.dev/templates/saas-dashboard ; https://otf-kit.dev/templates/fitness-kit ; https://arcade-preview.otf-kit.dev/

Internal posts linked, read on 3 Oct 2026: https://otf-kit.dev/blog/typesafe-system-one-jev-builders ; https://otf-kit.dev/blog/cloudflare-workers-previews-adlc

cloudflareworkers-aidecision-models
OTF SDK + Kits

Buy once, own the code. Ship with the agent you already use.

  • Free, open-source SDK — same component, web and mobile
  • Paid kits include AI configs + 40+ tested prompts — your agent reads the whole project
  • $99/kit or $149 for everything. No subscription, no sandbox limit.
Need more than components?

Full-stack kits.
Pay once, own the code.

Auth, database, and payments already connected — so you ship product, not setup. Or take every kit in the Bundle.

Everything Bundle — $149See full pricing

Get the free AI configs pack

Pre-tuned AI configs for Cursor, Claude, and Lovable — drop them in and your AI tool instantly understands your project.

No spam. Unsubscribe any time.

Prefer the free SDK? Star it on GitHub →