Cloudflare has introduced Clef and Clef-flash, two AI models built for rapid structured decisions rather than open-ended text generation.
The models answer bounded questions such as yes/no calls, multiple-choice selection or prioritising a list of options, the kind of fast routing decisions agentic systems and customer-support workflows rely on.
Taking on Jev
Cloudflare positions the pair as rivals to Jev, the decision model from TypeSafe AI that became the fastest adopted model in Vercel AI Gateway's history.
Unlike Jev, Clef and Clef-flash can also process images and video, so a decision can be triggered by a camera frame rather than just text.
Full Clef runs on a 27-billion-parameter Qwen3.8 base while Clef-flash uses a 9-billion-parameter Qwen3.5 base, both with a 64,000-token context window.
Cloudflare reported a median response time of about 209 milliseconds for Clef, versus 524 milliseconds for Jev, with Clef-flash at roughly 39 milliseconds.
Workers AI
Both models are hosted on Workers AI, Cloudflare's serverless platform for running AI inference on its edge network without dedicated GPUs, priced at $0.24 per million input tokens for Clef and $0.09 for Clef-flash, above Jev's $0.042 rate.
Cloudflare published the weights on Hugging Face under an Apache 2.0 licence for local deployment, though The Register noted Clef-flash needs a GPU with at least 41GB of video memory and full Clef needs roughly 85GB for a single concurrent request at full context length.
The models are also compatible with Jev's API. This lets developers swap endpoints without reworking existing integrations, a direct bid for the workflow-automation niche.
This is where smaller, cheaper decision models are emerging as an alternative to general-purpose large language models for tasks requiring typed, structured outputs rather than free text.