Cloudflare released Clef-omni on October 9 as a hosted Workers AI decision model that can examine text alongside an image, audio clip or video. For an application developer, the useful test is narrow: give the model the ticket state and a fixed set of questions, inspect its answers, and leave the routing decision with a person. This guide follows Cloudflare's current model documentation and launch example. BIG CHANGE has not run the API or measured its accuracy on support tickets.
The model ID is @cf/cloudflare/clef-omni on Workers AI. It is a decision model: the caller defines typed questions and allowed choices, then receives scores for those choices. It does not write a support reply. Cloudflare has also released open weights, but the steps below use its hosted endpoint. Cloudflare's speed and benchmark figures are its own claims; they do not establish how well this workflow will perform on your tickets.
Prepare access and a bounded question set
You need a Cloudflare account with Workers AI access, its account ID, a Workers AI API token and either Python with requests installed or a tool that can make the same HTTP request. Cloudflare's REST setup guide shows where to copy the account ID and create the token in the dashboard. A manually created token needs Workers AI - Read and Workers AI - Edit permissions. Keep the token out of the ticket and out of source control.
Sending a real ticket or attachment sends customer content to Cloudflare for processing. Cloudflare's Workers AI data-usage page says it does not use that content to train models offered on Workers AI or improve services without explicit consent; it also says content may be stored when Workers AI is used with a storage service such as R2 or KV. Before sending customer tickets, check your applicable Cloudflare agreement and your organization's customer-data rules, and minimize or redact personal information where practical. Cloudflare's no-training statement alone does not establish permission to send a particular customer's data.
Start with text so you can inspect the response before adding media. Cloudflare's own support example uses the state Checkout has been failing for every customer for the last hour and asks three questions: whether it is urgent (noul, its yes/no type), which team should handle it (choice), and how severe the impact is (score). This is a documented example, not an observed incident. The API accepts a string or structured object/array as state. The questions map must contain 1 to 64 question IDs; each has a type and instructions, while choice and score use criteria to define the allowed responses.
Send the text request
Set CLOUDFLARE_AUTH_TOKEN in your environment and replace the account ID below. This Python request adapts the exact endpoint and fields in Cloudflare's model example:
import os
import requests
account_id = "your-account-id"
token = os.environ["CLOUDFLARE_AUTH_TOKEN"]
payload = {
"model": "clef-omni",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {
"urgent": {
"type": "noul",
"instructions": "Is this support request urgent?",
},
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Payments, invoices, and refunds",
"technical": "Outages, errors, and configuration",
"sales": "Plans and upgrades",
},
},
"severity": {
"type": "score",
"instructions": "How severe is the customer impact?",
"criteria": ["No impact", "Minor", "Major", "Critical"],
},
},
}
response = requests.post(
f"https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef-omni",
headers={"Authorization": f"Bearer {token}"},
json=payload,
timeout=60,
)
response.raise_for_status()
print(response.json())For a successful check, look for the response's answers entries under the same urgent, team and severity IDs. Cloudflare describes the first as a probability of urgency, the second as a selected team with probabilities for its options, and the third as a probability-weighted score whose lowest level is zero. The API also documents model and usage response fields. Inspect the full JSON returned to your account before mapping it into an application. A probability is the model's output for the supplied question and state, not a guarantee that a real ticket belongs in that queue. Have a person compare the ticket and any attachment with the proposed route before taking action.
Add a photo, recording or video
Cloudflare's launch request puts optional media in separate images, audio and videos arrays, as embedded base64 data: URLs. For example, its sample uses data:image/png;base64,..., data:audio/mpeg;base64,... and data:video/mp4;base64,.... To add one local file to the Python payload before requests.post, encode it and append the data URL to the matching array:
import base64
from pathlib import Path
def add_media(payload, field, file_path, mime_type):
raw = Path(file_path).read_bytes()
encoded = base64.b64encode(raw).decode("ascii")
payload.setdefault(field, []).append(f"data:{mime_type};base64,{encoded}")
# Examples: use only the files present in the ticket, with their real MIME types.
# add_media(payload, "images", "photo.png", "image/png")
# add_media(payload, "audio", "recording.mp3", "audio/mpeg")
# add_media(payload, "videos", "clip.mp4", "video/mp4")Run add_media before the POST. Keep state factual and describe what each attachment is; ask only questions the available ticket and media can answer. Cloudflare's launch example uses a unit photo, sound recording and fan video to ask three yes/no questions about visible and audible details. A support application can keep its own ticket fields in state and its own allowed queues in criteria, but should check those definitions against its real workflow. The service does not accept remote media URLs.
Check limits before encoding. The model page allows up to four PNG, JPEG or WebP images, each at most 4 MiB and 16 megapixels, with 8 MiB total decoded images. It allows up to four audio clips, each at most 8 MiB and 300 seconds, and two videos, each at most 16 MiB and 60 seconds. Audio and video together have a 16 MiB decoded limit. Video is sampled at two frames per second; its soundtrack is used with the frames when every video in the request has one. If an attachment is too large or too long, have a person select a relevant smaller excerpt or use a text-only check and review the original media separately. Do not treat a cropped or shortened file as complete evidence.
Check cost, context and failure paths
Cloudflare currently lists Clef-omni at $0.15 per million input tokens and says it does not charge output tokens. Media is converted to input tokens at that rate. Its model page describes roughly 780 tokens per audio minute and up to roughly 15,400 tokens per video minute at maximum resolution, plus audio tokens when the video has sound. Image cost depends on the resized image and is capped at 1,024 tokens per image. These are billing rules, not a fixed price per attachment.
The hosted model has a 64,000-token context. Media and questions count toward it. If they exceed the window, Cloudflare says the request fails; otherwise long text in state may be truncated to fit. A returned answer therefore does not prove that the model considered every line of a long ticket. When a request fails, inspect the HTTP error and reduce or remove attachments within the documented limits before retrying. When it succeeds, check that the expected question IDs are present, retain the original ticket for review and compare the proposed route with the evidence. The model's probabilities may help a reviewer focus attention, but a service guarantee or safe automatic threshold would require evaluation on your own labeled tickets.
The big change
Cloudflare's October 9 release adds audio and video to the hosted Clef decision model path already used for typed questions over text and images. A developer can submit one ticket state with supported media and receive answers keyed to a fixed schema through a single Workers AI request. What remains for the application team is deciding which questions belong in that schema, testing performance on its own tickets and requiring human approval before a response or routing action.
Sources & further reading
- Cloudflare Clef-omni model documentation, checked October 10, 2026. The primary reference for the hosted model ID, API fields and example, media and context limits, response shape and listed token price. Documentation describes service behavior; BIG CHANGE did not execute the request.
- Cloudflare's Clef-omni launch post, October 9, 2026. Establishes the launch date and provides the embedded-media request example. Performance and benchmark figures are Cloudflare's own claims.
- Cloudflare Workers AI REST setup, last updated September 15, 2026 and checked October 10. Shows how to obtain the account ID and API token and identifies the permissions required for a manually created token.
- Cloudflare Workers AI data usage, last updated April 21, 2026 and checked October 10. Explains Cloudflare's processing of customer content, its stated training and service-improvement restriction absent explicit consent, possible storage when a separate storage service is used, and the applicable agreement boundary. It does not determine whether a developer may submit any particular ticket or recording.



