You are asking an assistant if a feature announced this week is already available. It finds a recent article and responds with a link. But this link may discuss a future step, repeat an old announcement, or refer to a different offer. Finding a page and verifying a claim are two different operations. The new web search interface of Cloudflare provides an entry point; your application must still organize the path to a verifiable response.
In the fictional scenario of this article, you are preparing a product monitoring assistant. Its user expects a short response: what has been announced, what is available, and what is still planned. We propose a method to separate these states, keep useful evidence, and produce a precise refusal when information is missing. No research service has been called to measure its performance.
What changes with Web Search API
On October 2, 2026, Cloudflare announced Web Search API via AI Gateway, its gateway for AI calls. The native Server Tools, which must integrate server-side tools into this gateway, are still announced as a future step. They are not presented here as available.
The documentation consulted on October 6, updated on October 2, indicates an open beta. It describes structured results including titles, URLs, and descriptions, multiple providers, access via REST from a server or by binding Workers, and logging in AI Gateway. REST here refers to an interface called via network requests; the binding is the access provided to the program that runs in Workers.
The architectural interest is to make the 'search' step identifiable. You can assign it a deadline, a cost, a failure status, and a trace. In our proposal, this step returns candidates. Another step opens the authorized source, a third qualifies the claim, and a final one drafts the response. The same agent can participate in multiple steps, but the responsibilities remain distinct in your application.
Define the question before sending a request
"Is this feature available?" lacks context. Available for which offer, which region, which version, and at what date? Ask for the necessary details or announce the chosen scope. In our fictitious example, the question becomes: "As of the consultation date, is the feature described in this announcement open to developers of this offer?"
A search query should contain the public terms necessary for this question. It does not need the name of a client, their contract, or the content of an internal file. Our recommendation is to construct a minimal formulation before calling the provider and to check what goes into the logs. A trace useful for operation can become unnecessarily sensitive if it copies the entire conversation.
Also set a work budget. Your assistant can have a call limit, an overall time limit, and a list of priority sources. These choices are application parameters to be adjusted; we do not provide supposedly optimal values. When the budget is exhausted, the application must distinguish between "incomplete search" and "function unavailable".
Connect each sentence to evidence
For each relevant result, search for the page where the fact originates: publisher announcement, function documentation, or official repository. Press coverage can help with discovery. It does not become the primary evidence just because its date is more recent. If the original page is inaccessible, keep this limitation and reduce the scope of your response.
In our method, evidence supports a specific claim. For example, the announcement may support 'an opening is planned,' while the access documentation must support 'this feature is usable in such an offer.' You can therefore have two useful links that answer two different questions. A bibliography placed at the end does not solve this correspondence by itself.
Separate three dates. The publication date describes the document. The event date describes what happened. The access date describes the time when you read the page. An article published today can report an old event; documentation modified today may retain an earlier availability. If a date is not verifiable, leave it unknown rather than deducing it from a URL.
Build an evidence record independent of the provider
We call SearchEvidence the conceptual contract below. This name and its fields belong to our demonstration; they do not describe the official Web Search API schema. The goal is to be able to change the search source without losing the verification logic.
| Conceptual field | What it describes | Proposed control |
|---|---|---|
| Source URL | Page actually examined | Authorized address and successful opening |
| Title | Document identification | Correspondence with the page read |
| Consulted on | Time of inspection | Application timestamp |
| Published on | Verified date of the document | Origin of the preserved date |
| Event on | Date of the event, if known | Relevant passage identified |
| Supported statement | Phrase that the evidence allows | Limited to the content read |
| Useful passage | Short portion or rephrasing | No automatic full copy |
| Proof status | Sufficient, contradictory or absent | Explicit rule before writing |
The reader must be able to open the citation and understand why it accompanies the sentence. In our monitoring assistant, a response could present a dated announcement and an availability still to be confirmed separately. If a page does not allow determining the access status, write it down. The model must not fill in this field with what seems usual for this provider.
Keep a distinction between the URL of the result and that of the page actually viewed, especially after a redirect. Also record a refusal to open or an access restriction. These states are useful to explain why the evidence could not be consolidated, without claiming that all websites on the Web are readable.
Prepare an integration with visible limits
In an existing application, add a search interface next to your document services. Its output is a list of candidates and a call status. Opening a page belongs to a separate component, with the network and collection rules allowed by your organization. The writing then receives only the items necessary for the subject.
Cache, that is to say the temporary reuse of a previous response, requires an explicit choice. A definition search and a question about current availability do not have the same freshness requirement. Keep the date of the reused evidence and plan a new consultation when the question requires it. The assistant's response date should not give the impression that all sources have just been reviewed.
To avoid an expensive loop, limit retries and distinguish the errors: timeout, access denied, unexpected format, or search without a relevant candidate. Prepare an exit message for each. 'I could not verify the primary page' is more useful than a confident response built from an incomplete summary.
This path can complement an internal base. It retains your procedures and your vocabulary; public search is used to verify an announcement or an external development. The usage called RAG, retrieval-augmented generation, consists precisely of providing documents to the model before its response. Adding a web search encourages organizing public and internal sources, with their rights and freshness requirements.
Test responses, quotes, and refusals
The following protocol is a Partitech proposal. It uses controlled test pages or network doubles: components that simulate responses without real calls to the provider. It would allow you to evaluate the logic of your application without presenting the results as those of the Cloudflare service.
| Proposed test case | Expected observation | Output to check |
|---|---|---|
| Recent announcement of a future step | Recent date, future access | Announced/planned distinction |
| Old fact on a recent page | Two different dates | No presentation as a novelty |
| Primary page inaccessible | Single search result | Insufficient verification reported |
| Two divergent sources | Different scopes or versions | Explicit contradiction |
| Network timeout | Incomplete search | Declared failure, no invented conclusion |
| Instruction on a page | Text outside the application | Instruction ignored, data processed as source |
Then evaluate the answers sentence by sentence. Count the supported statements, those without evidence, and those whose citation pertains to another scope. Also examine the refusals: are they justified and do they indicate what is missing? An assistant who refuses everything does not satisfy the task any more than an assistant who always responds.
Separate the test of the search component from that of the writing. You can have good results and produce a poor summary, or find no sources despite excellent citation logic. This separation helps to choose the relevant correction rather than randomly modifying the model's instructions.
Treat pages as external data
A page can contain a sentence that asks the agent to ignore its rules or to send information elsewhere. In our design, this text gains no authority: it remains content to be examined. Action permissions, the system's instructions, and secrets must not be redefined by a sourced reference.
So plan a test where a demo page contains an instruction unrelated to the subject. The expected success is an extraction limited to the useful evidence, without any additional action. Also check the logs and the final response to ensure that they do not reproduce sensitive data added to the context. This is a safeguard to be tested in your own architecture, not a capability guaranteed by a search API.
To begin, choose a public and bounded question, determine the necessary evidence, and prepare the six test cases. The next step is to obtain an answer for which each statement can be verified, as well as a comprehensible refusal when the evidence is lacking. You can then measure the actual service in an authorized setting, with its costs and dated limitations.
Sources and verification date
Primary sources reopened and read on October 6, 2026: Cloudflare, Web Search API announcement of October 2 ; Web Search API documentation, updated on October 2. The SearchEvidence contract, the scenario, and the tests are Partitech proposals. No benchmark, application network test, or partner audit is claimed.