Cloudflare Launches Clef & Clef-Flash: Multimodal Edge Decision Models Bring Sub-15ms Vision-Driven Agent Steering to Workers AI
SAN FRANCISCO — October 3, 2026 — Cloudflare has officially launched Clef (27B) and Clef-flash (9B), an innovative family of multimodal decision models deployed natively across its global Workers AI edge network. Unlike traditional text-only routing classifiers, Clef represents the industry's first frontier decision system designed from the ground up to evaluate high-resolution images, streaming UI video frames, and structured JSON payloads simultaneously to steer autonomous agent workflows.
As autonomous agent applications expand beyond basic text chats into browser-use automation, robotic process automation (RPA), and physical computer vision, routing decisions must frequently account for visual context. Asking a heavyweight multimodal LLM to analyze a screenshot just to decide whether to click a button or scroll down takes several seconds and burns costly token budgets. Clef collapses this process into an edge-native classification pass executing in under 15 milliseconds.
1. Multimodal Decision Intelligence at 330+ Global Edge Cities
Built upon high-capacity Qwen vision-language architectures and optimized for edge tensor parallelism, the Clef model family introduces a specialized decision scoring layer compatible with the open Jev-API standard. Instead of emitting verbose descriptions or slow conversational markdown, Clef models output calibrated likelihood matrices across defined action branches.
By deploying Clef directly onto Cloudflare's serverless GPU fleet distributed across more than 330 cities worldwide, inference requests execute within milliseconds of end users and autonomous devices, eliminating round-trip latency to centralized hyperscaler datacenters.
"The next generation of AI agents won't just read JSON—they are looking at screens, inspecting diagrams, and monitoring live video feeds. With Clef, we are bringing multimodal instant reflexes to the edge, enabling agents to perceive and decide faster than human nervous systems."
Cloudflare Clef & Clef-Flash Performance Matrix
2. Native Reinforcement Learning on Edge Telemetry
Alongside the foundational model checkpoints, Cloudflare introduced an integrated Edge Reinforcement Learning (RL) platform. Enterprise developers can pipe anonymized agent outcome logs—such as whether a given routing decision successfully completed a purchase flow or triggered an exception—back into automated DPO (Direct Preference Optimization) training loops.
Because the models specialize in discrete decision spaces rather than generative language synthesis, fine-tuning requires only a fraction of traditional compute. Teams can retrain a custom Clef adapter in under two hours on modest GPU allocations, continually tailoring their routing accuracy to changing application environments.
3. Universal Jev-API Compatibility and Pricing
Both Clef and Clef-flash adhere strictly to the open Jev-API specification, ensuring instant plug-and-play compatibility with existing decision-model orchestration frameworks. Developers can route between local models like AWS's Strands Decider 2B for purely local text tasks and Cloudflare Clef for edge multimodal tasks with zero SDK code modifications.
Pricing for Clef-flash starts at just $0.05 per 1,000 multimodal decision evaluations, while the flagship Clef 27B model is priced at $0.20 per 1,000 evaluations, representing a 90% cost savings compared to routing equivalent tasks through frontier multimodal chat APIs.
4. SyncFlo AI Integration: Visual Agent Steering
SyncFlo AI has integrated Cloudflare Clef-flash into its visual browser automation and invoice processing pipelines.
When SyncFlo agents navigate dynamic web portals, reconcile complex multi-page billing statements, or extract contract terms from scanned documentation, Clef-flash acts as an instant perceptual steering engine. Visual state changes are classified in real-time, directing the orchestrator to execute deterministic keystrokes and API payloads without stalling pipeline throughput.