Cloudflare Releases Clef Decision Models and RL Fine-Tuning Platform
Cloudflare has officially released Clef and Clef-flash, open-source decision models designed to accelerate deterministic routing and classification within agentic AI workflows. Hosted on Workers AI and distributed under the Apache 2.0 license on Hugging Face, the models represent a strategic move toward efficient, structured decision-making that complements traditional large language models. The accompanying launch introduces a new reinforcement learning fine-tuning platform, enabling enterprises to adapt Clef to highly specific operational requirements. Decision models differ fundamentally from generative LLMs by producing bounded, structured outputs with calculated probabilities rather than open-ended text. This architecture allows autonomous agents to programmatically evaluate inputs, trigger escalations, or route tasks without human intervention. Cloudflare developed Clef to optimize workflows such as threat intelligence classification, customer support triage, and bot detection. By eliminating token-by-token generation, the models deliver near-instant classifications at a fraction of the compute cost associated with conventional generative models. Technically, Clef leverages a frozen Qwen base model enhanced with a specialized two-stage attention routing process. Instead of autoregressive text generation, it derives schema-bound choices directly from internal representations, enabling parallel scoring of valid classifications. The architecture supports a 64,000-token context window and includes a vision encoder for multimodal input, distinguishing it from text-only competitors like Typesafe AI’s Jev. Internal benchmarks against the Jev Decision Index demonstrate competitive accuracy across classification, tool retrieval, and safety evaluation metrics. Clef-flash optimizes this foundation for latency-sensitive applications, achieving median inference times under 40 milliseconds while maintaining robust precision. Both models are fully API-compatible with Jev, allowing seamless integration into existing infrastructure. To address enterprise-specific classification needs, Cloudflare is launching a reinforcement learning service that allows customers to fine-tune Clef using proprietary datasets. The platform, initially supported by forward-deployed engineers before transitioning to a self-serve interface, leverages Cloudflare AI Gateway for data capture, isolated container sandboxes for training environments, and Bring Your Own Model deployment capabilities. This fine-tuning approach prioritizes domain accuracy over general-purpose performance, enabling organizations to train classifiers on historical network data, trust-and-safety submissions, or specialized support tickets. The release positions Cloudflare to expand its AI infrastructure offerings, emphasizing edge-hosted inference and deterministic agent orchestration. By integrating Clef with Workers AI, developers can deploy decision models directly within application hot paths, significantly reducing network latency and operational overhead. The models operate under strict data privacy guarantees, with Cloudflare confirming that client requests and responses are neither stored nor used for training unless explicitly opted into the fine-tuning service. Organizations seeking to integrate structured decision logic into autonomous systems can now access Clef and Clef-flash through the Workers AI platform or self-host the open-source weights. The accompanying reinforcement learning tooling establishes a foundation for customized model optimization, reinforcing Cloudflare’s focus on scalable, edge-native AI infrastructure.
