Discuss your project
AI Assessment

Clef: test the routing of requests before automating your support

Billing, a technical problem or a sales request: directing a message is a decision to test. With Clef, prepare the categories, count errors according to their business consequences, and plan for human intervention.

Demandes de support orientées vers trois catégories avec une voie de reprise humaine

Your support inbox receives three messages: 'I can't find my invoice,' 'The download fails,' and 'Can you present your offer to me?' Before writing a reply, it is necessary to decide to which team to forward them. This first step is a classification: choosing a category from a defined list. Clef, announced by Cloudflare on October 1, encourages treating this choice as a task in its own right. Here is how to prepare its evaluation for a fictional support team, with visible errors and human intervention.

Start with the expected decision

In our demonstration example, the categories are billing, technical, and commercial. A fourth category, out of scope, accommodates requests that this organization does not know how to handle. Alongside these categories, the application has a processing status: accepted or to be reviewed. The category describes the request; the status describes what you authorize to be done with it.

This separation helps to understand an ambiguous request: "My subscription is paid, but I can no longer access my files." The model can suggest the technical category, while your policy requires a verification because the message also concerns a payment. You do not need to force a single label to carry the entire complexity of the intervention.

Define what remains outside the first pilot: refund, account closure, rights modification, or response sent to the client. The classifier proposed here suggests a direction. It does not trigger any of these actions. For a first trial, people continue to work normally while you compare the suggestions to the decisions made.

What Cloudflare makes available

In the announcement of October 1, 2026, Cloudflare presents Clef and Clef-flash as decision models with structured outputs, available in Workers AI. Adaptation to client tasks starts with human support; a self-service fine-tuning platform is announced for later. Fine-tuning involves adapting a model based on examples of a particular task.

The official Clef model card on Hugging Face, consulted on October 6, displays an Apache-2.0 license and describes the model's artifacts. This does not demonstrate its operation on your hardware. The sheet also contains an unusual type in an example: we do not take this example as a validated integration contract.

This information opens up two avenues for testing to qualify: a hosted service or running the available weights. The choice depends on your data constraints, your hardware, and the operating time. Before promising a reproducible local test, it would be necessary to set the revision of the artifacts, the dependencies, and the hardware configuration, then run a test. This article presents neither an executed installation nor a billed call.

The comparisons published in the announcement are those of the supplier. Our trial decision is not based on their ranking: the interesting result would be better routing of your requests, with an acceptable error cost. Several pages from the same publisher do not constitute an independent confirmation of quality for French-language support.

Write a simple contract for the application

The contract below is an internal Partitech proposal, independent of the official Clef schema. It describes what your application must receive and verify. No field name is presented as a parameter of the provider's API.

Conceptual element Role in the pilot Check to be scheduled
Demo ID Link input and decision Presence, uniqueness in the batch
Authorized text Content to classify Non-empty, bounded size, absence of secrets
Version of the categories Define the possible choices Version recognized by the application
Proposed category Suggested orientation Value present in the allowed list
Possible score Help assess acceptance Expected format and documented interpretation
Processing status Accept or request a review Business rule distinct from the model

You can check the contract with a few synthetic messages. For "The demonstration invoice cannot be found," the reference label chosen for the exercise would be billing. For "The download of my demonstration document is interrupted," it would be technical. An advertising message without a request would receive out of scope. These examples are used to check the wiring and categories; they do not measure quality on real customers.

Plan for rejection cases: missing text, unknown category, unreadable score, incomplete response, and exceeded deadline. The expected behavior is a human takeover or an explicit failure, never a silent routing to a default category. A well-formed response can still contain a wrong decision; format and quality controls remain separate.

Authorized entry, proposed category, and acceptance business rule with human review

Create a test set that reveals errors

The following protocol is proposed by Partitech and has not been executed. Start with an annotation guideline: what each class means, how to handle messages with multiple topics, and when to consider something out of scope? Have ambiguous cases reviewed by a second person. If annotators do not agree, the model cannot resolve the imprecision of your organization on its own.

Then separate three uses of the data. One batch is potentially used to adapt the model. Another is used to choose the settings and thresholds. The final test set remains frozen until the comparison. If you adjust an instruction after seeing its errors on this last set, it becomes a development set; then prepare a new independent evaluation.

Also avoid copies between batches: two messages from the same thread, a reproduced email template, or nearly identical variants can give an impression of generalization. For our fictitious support, separating by conversation and reviewing duplicates would be relevant checks. Keep the authorized messages, minimize identifying information, and do not send any customer data without an already established framework.

Include the real difficulties: very short sentences, approximate French, requests mixing two subjects and new vocabulary. Keep their frequency and nature. You will thus be able to explain if a method mainly succeeds with obvious requests and defers all others to a person.

Count the errors according to their consequences

A confusion matrix is a table that crosses the expected category with the proposed one. It allows seeing where the system makes mistakes. In our example, classifying a business request as technical wastes time; classifying an access problem as business can further delay its resolution. The same number of errors therefore does not produce the same business cost.

Define this cost with the people who handle the requests. For the pilot, you could count the number of reassignments, the review time, and the misdirected urgent cases. If you use a numerical scale, indicate that it is a choice of your team and keep its reasons. A weighting decided after seeing favorable results would make the comparison difficult to defend.

A high score does not automatically mean a 'reliable decision.' Calibration describes the agreement between a reported confidence level and the frequency of correct decisions in comparable cases. To check it, group the decisions by score intervals, count the cases, and compare confidence and observed correctness. An interval with very few examples provides little information: display its count rather than giving it a definitive verdict.

Abstention also deserves a measurement. If you require a review for all difficult messages, the accepted routing decisions may seem excellent while the human team retains most of the work. Count the share of accepted requests, the share to be reviewed, and the rework time. Examine the results by category and by type of difficulty.

Compare before adapting

Your starting point, called the baseline, can be a simple rule: a few billing terms, a list of technical errors, and a fallback in case of conflict. Add a generalist model with bounded output if useful. Compare these solutions to Clef on the same inputs, with the same categories and the same validation rules. A specialized adaptation is only justified after this comparison.

Proposed measure Simple rules General-purpose model Clef Clef-flash
Errors by category To measure To measure To measure To measure
Business cost of errors To measure To measure To measure To measure
Part to review To measure To measure To measure To measure
Response time To measure To measure To measure To measure
Cost and recovery time To measure To measure To measure To measure

Measure the time for the complete process, including validation and any human intervention. Note the configuration, the model version, and the date. Do not replace an empty box with a result from the announcement: this table should describe your trial. A faster-to-predict solution may still be less useful if it requires more corrections.

Allow a pilot that can be stopped

Start by observing, without automatically modifying the support queues. Compare suggestion and human guidance, then examine the disagreements. They may reveal a model error, a wrong reference, or a category that needs clarification. Address these causes before running a test again with an explicit protocol.

Decide in advance the acceptance criteria: tolerated critical errors, maximum review workload, and covered categories. Plan for immediate return to human processing if the outputs become invalid, if critical errors increase, or if the vocabulary changes significantly. For this example, the next useful step is therefore to stabilize the categories and the test set. The choice of model comes after this business agreement.

Sources and verification date

Primary pages reopened and read on October 6, 2026: Cloudflare, launch of Clef, October 1 ; Cloudflare, official Clef model card on Hugging Face. The synthetic corpus, the contracts, and the tables are Partitech proposals. No benchmark of our own, training, or hosted call was conducted for this article.

Share this article