Complete AI Training

AI news ·

Cloudflare releases clef, a 27B open-source multimodal model for turning states and schemas into decisions

Cloudflare released Clef, a 27-billion-parameter model that scores business decisions in a single pass under 210 ms median latency. It returns typed answers and probability distributions for routing, condition checks, and document reading without generating tokens.

Share

Cloudflare has released Clef, a 27-billion-parameter multimodal model that turns a block of text or images and a set of typed questions into structured decisions. The model is built specifically for business workflows - routing requests, checking conditions, applying policies, and reading documents - and it does this in a single forward pass rather than generating tokens one by one.

Clef ships under an Apache 2.0 license and tops Cloudflare's Decision Index leaderboard. It handles a 64K context window, double the 32K that the company's earlier Jev model supports, and accepts text, JSON, or images as input. You define the questions and the possible answers; the model scores every option for every question simultaneously and returns probabilities.

How the decision API works

Decision models use an endpoint that takes a state (the text or structured data to judge), optional base64-encoded images, and a list of up to 64 questions. Each question specifies a type - choice for picking from options, noul for yes/no, or score for rating on a scale - along with instructions and criteria. The model returns the chosen answer, a probability distribution, and a confidence score between 0 and 1 that reflects how concentrated the probabilities are.

The API supports PNG, JPEG, and WebP images. URLs and data URLs are not accepted. For choice questions, you can define 2 to 26 options. If two options tie, the model follows the order you defined them in, so placing a preferred option first acts as a tiebreaker.

What teams can build with it

The model is designed for four core task patterns. For routing, you supply destinations and rules; Clef returns the chosen destination and each option's probability. For condition checks, it answers yes or no with a probability. Policy application works by providing rules and allowed outcomes, and the model returns a typed decision. Document reading accepts a photo of a page plus fields to extract, returning field-by-field decisions.

Cloudflare also highlights tool call moderation - checking an agent's planned action before it executes - and image-based decisions like processing screenshots, receipts, and forms alongside text.

Performance across benchmarks

Clef delivers a median latency of 209.3 ms and a p95 of 238.6 ms. A smaller variant, Clef-flash, cuts median latency to 38.8 ms with a p95 of 122.4 ms. Both are faster than Jev, which runs at 524.1 ms median and 536.0 ms p95.

On classification benchmarks, Clef scores 98.5% exact accuracy on BFCL, 94.2 macro-F1 on BANKING77, and 97.4 macro-F1 on CLINC150+OOS. Clef-flash leads on API-Bank at 93.1% and on a home appliance simulator at 97.7%. On four end-to-end business workflows - invoice processing, customer service, security incidents, and agent trace observability - Clef and Clef-flash perform close to or above Jev, with Clef reaching 76.3% on customer service and 68.5% on agent trace observability.

Why this matters for customer support, healthcare, insurance, and product teams

For teams that run high-volume decision pipelines - triaging support tickets, checking claim eligibility, flagging security incidents - Clef's single-pass architecture means decisions arrive in roughly 200 milliseconds rather than the half-second or more that autoregressive models need. The structured output format eliminates parsing headaches: you get a typed answer and a probability distribution, not free text that requires regex or secondary classification. Insurance and healthcare teams processing forms and receipts can send images directly to the model without a separate OCR step. Product development teams building agentic systems can use the tool call moderation pattern to add a safety check before an agent acts, with probabilities that make gating logic straightforward.

Share