Tag Archives: Birthday Week

Cloudflare Impact reaches $100 million in donations

Post Syndicated from Patrick Day original https://blog.cloudflare.com/100-million-donations/

This week, Cloudflare's Impact programs will reach $100 million in donated services. It's a significant milestone, and one that we are proud of because it means that thousands of organizations, like journalism outlets, civil society, state and local governments, election management bodies, and public schools are being protected from cyberattacks.

But Cloudflare's Impact programs have never been about philanthropy. They are a fundamental part of our business and our mission, and they continue to help guide almost everything we do.

As we celebrate this milestone and our 16th Birthday Week, we wanted to revisit not only how we got here, but also how our Impact programs continue to grow and evolve to help those working for the public interest.

Free → Impact

Cloudflare started as a free service. The original idea was to provide a basic version of our services to developers and small businesses for free, and then use the data about cyberattacks on their websites to build more sophisticated products that we could sell.

However, we quickly discovered that some of our free customers were not only doing essential work, like reporting on corruption in Africa or on the Russian invasion of Crimea, but also experiencing some of the largest attacks on our network. That realization changed how we thought about our free services. We committed not only to making them available for free for everyone, but also to doing more for organizations being targeted by powerful adversaries simply for serving the public.

Cloudflare launched Project Galileo in 2014 to provide more advanced security services for important but vulnerable people and organizations online, including journalists, human rights defenders, and civil society groups. Today, the program includes more than 3,500 domains in over 120 countries. In 2025, Cloudflare blocked more than 38.5 billion DDoS, website vulnerability, email phishing, and other cyberattacks against Project Galileo participants, almost 105.4 million per day.

Over the last 12 years, Cloudflare has continued to expand what we now call our Impact programs. Although each program is unique, our goal is the same: to support organizations and institutions serving the public, particularly those that would not otherwise have access to the necessary cybersecurity services. For example:

Helping keep these organizations online by protecting their websites and internal data remains essential. In 2026, Cloudflare released its first annual report on cyberattacks against civil society, which found that civil society organizations are targeted more frequently and more intensely than other Cloudflare customers. For example, Project Galileo participants faced attempts to exploit security vulnerabilities in websites at a rate more than seven times higher than an average user. Cloudflare is also on pace to more than double the number of applications to Project Galileo from last year.

But Cloudflare's Impact programs have never been static; they evolve alongside our company and technology, and the organizations they serve. Increasingly that means not just defending public interest organizations, but empowering them to adapt and thrive in the era of AI.  

Looking to the future

In early September 2026, on a rainy day in Barcelona, Cloudflare co-hosted a hackathon. Because our developer platform is such an important part of our business, we hold these events all the time. But this one was different: instead of a room full of software engineers or startup founders, it was the first time we held an event specifically for journalists.

Media Party hackathon co-hosted by Cloudflare at the BIT Habitat in Barcelona (September 9, 2026).

The event was part of a three-day conference organized by Media Party, a nonprofit dedicated to media innovation through digital tools. The event was designed to bring together journalists, developers, and strategists to solve a single problem: how to help newsrooms adapt to an AI-driven, post-search information landscape. The sprint focused on four themes: workflow automation, agentic journalism, synthetic-content verifications, and information integrity.

Each team received free access to Cloudflare's developer platform and the assistance of volunteer Cloudflare engineers to see what they could build in a day. 

Four teams made it to the final round. The winning team, AIdas, built a tool that helps researchers and journalists study AI bias across politically contested topics by comparing how different LLMs answer the same question, and recording their responses as open data.

The hackathon was just one part of a broader effort across Cloudflare Impact to expand beyond cybersecurity services to help public interest groups adapt to a changing world:

  • Protecting Local News from AI Crawlers: Last year, Cloudflare provided free access to our Bot Management and AI Crawl control for Project Galileo participants, including more than 750 journalists, independent news organizations and non-profits supporting news-gathering around the world. These tools will help these organizations understand and control how their content is accessed by AI crawlers, and safeguard their reporting from unauthorized scraping.
  • Non-profit startups: Last year during Birthday Week, Cloudflare announced its startup program, which provides more than $250,000 in Cloudflare credits, would be available for the first time for non-profit organizations. This week we will announce the first 30 organizations accepted into the program and how they are serving their communities with tools built on our developer platform.
  • Automation tools for human rights: This week we will also announce three new projects that Cloudflare engineers have built using our developer platform for three leading human rights organizations, covering topics including tracking transnational repression, digital rights legislation and policy development, and corporate human rights due diligence.

Across all of these new efforts, the goal remains the same: to help organizations doing essential work access the tools and support they need to continue to advance their missions.

Join Us

I had the opportunity to meet with two of the Cloudflare engineers who volunteered at the hackathon in Barcelona. They both mentioned to me that one of the reasons they came to work at Cloudflare was Project Galileo, and the chance to use their skills to help organizations working in their communities. 

It was an important reminder that Cloudflare's Impact programs and our mission are not just things we have done. They continue to shape our identity, including through the people who choose to come work with us. 

If that sounds like the kind of work you want to do, come join us.

Identify AI model overuse with User Insights

Post Syndicated from Ayush Kumar original https://blog.cloudflare.com/ai-model-overuse-user-insights/

When we launched User Insights last month, we wanted to help teams answer a basic question: What are people actually doing with AI? User Insights gives teams a clearer view of their AI usage, showing which users, applications, tasks, and models are driving traffic. It also highlights user and agent anomalies, helping teams identify unexpected or out-of-control spending and usage before they become larger problems.

Our latest update adds something our users have been asking for: context. 

Since launch, we’ve heard from users that model names and request counts only tell part of the story. They show where traffic is going, but reveal little about the work behind it: is that request a code review, a research task, or an agent making several calls to complete a job? The same token count can represent very different kinds of work, and you can’t evaluate with model choice without understanding the task.

User Insights now shows when a model may be more capable than a task requires, which of your users and agents are driving that usage, and how the task, model, cost, and conversation patterns relate. Teams can use these insights to investigate and make targeted changes within their organization. These capabilities are available for free to AI Gateway users.

Why AI usage is hard to understand

Consider a team that has routed its internal AI traffic through AI Gateway. After a few weeks, spending is increasing and some requests feel slower than expected, a common challenge as organizations adopt AI at scale.

There could be several explanations. Developers may be using AI for increasingly complex coding work. Agents may be making too many follow-up calls to complete a task. Or a small group of users or agents may be responsible for a disproportionate share of the organization’s usage.

Tokens and request counts alone cannot show which pattern is driving the increase. Teams need to understand what the traffic represents before deciding whether a model, workflow, or routing rule should change.

Helping teams find where AI models are overkill 

The model overkill view helps teams identify conversations where the selected model appears to be more capable than the task requires. For example, a team might discover that users or agents are sending simple formatting or summarization requests to a high-capability reasoning model.

That gives the organization a place to start. They can see which users, agents, or applications are associated with the pattern, then investigate the tasks behind it. A team might find that a model is being used because it is the default, because users are unsure which model to choose, or because an agent has been configured to use the same model for every step.

The overkill view is not a leaderboard and does not automatically recommend a replacement model. It helps teams ask better questions:

  • Is this model appropriate for the task?
  • Is the extra capability improving the result?
  • Would a faster or less expensive model produce an equivalent outcome?
  • Is the issue limited to one workflow, user, or agent?

From there, teams can compare cost, latency, token usage, and conversation turns before deciding what to change.

These insights support both the new Potential Savings view and the Auto Router, which is launching in public beta alongside this release. The Potential Savings view helps teams identify requests that may be handled by a faster or less expensive model without compromising output quality. The Auto Router applies these task and model-fit signals automatically, helping reduce costs without requiring a separate routing rule for every workload.

The Overkill view is a starting point for evaluating model fit. Teams can compare latency, input and output tokens, conversation turns, and total cost for the same type of task. A difficult coding or research task may need a capable reasoning model, while a short summary or simple classification task may not. The goal is not to move every request to the least expensive model, but to understand whether the selected model is appropriate for the work.

Understand what people are using AI for

Task analysis groups conversations by the kind of work they represent. Initial categories include coding, research, writing, summarization, and data analysis.

This provides context that a list of model names cannot. An engineering team might use AI mostly for coding and debugging, while another team might use it for research and summarization. A team may also discover that a surprising amount of traffic comes from simple tasks, even though those tasks are being sent to a high-capability model.

The answers will vary by team. The category data provides a way to investigate those differences using traffic already passing through AI Gateway. Teams can determine whether a model is being used for the work it is best suited to handle, or whether a default model is being applied too broadly.

Understand the full cost of a task

Some tasks are finished in one exchange. Others take a few rounds of questions, corrections, and follow-ups. Turns analysis shows how much back-and-forth different tasks require. A long conversation is not necessarily a bad thing, especially for complex work. But if a simple task keeps taking several turns, it may be worth looking at the prompt, the model, or the workflow.

The first request is only part of the cost. Teams should also look at the time, tokens, and money spent before the task is finished. Comparing those numbers can show where a workflow is taking longer or costing more than expected.

Turn insights into auto routing

Once a team has identified an overkill pattern and confirmed it across task, cost, latency, and turn data, it can turn that insight into an automatic routing decision.

For example, the task view might show that much of the team’s AI usage is summarization and formatting. The model view could show that those requests are being sent to a large reasoning model, while the turns view shows that most conversations finish in a single turn. Together, these signals give the team a concrete workload to evaluate.

In addition to our updates to User Insights, the Auto Router is now available in closed beta. The Auto Router uses the conversation trajectory, task category, task complexity, and model-fit signals to automatically route requests to an appropriate model while taking cost into account.

Instead of creating a separate routing rule for every workload, customers in the beta can let the Auto Router select among the models available to their application. The router does not simply send every request to the least expensive model, but instead selects an appropriate model for the task at hand. Complex coding or research work may still need a more capable model, while simpler tasks may be handled by a faster or less expensive option.

To learn more about the Auto Router and sign up for the closed beta, read the blog post here.

The Auto Router uses the same task and conversation signals that power User Insights. The section below explains how those signals are produced.

How User Insights classifies traffic

Each conversation receives an analysis signal that can be grouped in User Insights. The signal is used for reporting and routing analysis, and is not intended to replace or expose the original request.

The categorization engine is a dedicated Cloudflare Worker that processes eligible AI Gateway logs. It examines the conversation trajectory, including user requests, assistant responses, tool calls, and tool results, and identifies the type of work being performed, such as coding, debugging, research, or summarization. It also returns a confidence score and evaluates dimensions such as task complexity, intent ambiguity, stakes, and context dependence.

The Worker returns a category that can be joined with the log metadata used by the dashboard. These signals can also be used to evaluate model fit by comparing how well candidate models suit the task against their cost. The current implementation focuses on a small set of categories that are easy to understand, rather than trying to infer every detail about a user’s work.

The pipeline follows the existing AI Gateway log architecture. Metadata is stored separately from log bodies, and the current implementation uses Durable Objects for metadata and R2 for log bodies. User Insights exposes derived categories and aggregate views. It does not turn the dashboard into a raw prompt browser. Retention of the underlying log bodies continues to follow the configured AI Gateway logging behavior, so teams should review those settings when deciding what to send through the classifier.

The classification is asynchronous, which means it happens after AI Gateway has handled the request rather than while the user is waiting for a response. AI Gateway writes the log to the existing storage path first, and the classification Worker processes it afterward. This keeps classification out of the request path and adds no latency to the user’s response.

The tradeoff is that User Insights is not a real-time view. Newly received conversations may not appear in the dashboard immediately, and analysis may trail incoming traffic by approximately one day as logs are processed and aggregated. Teams should use User Insights to identify usage patterns over time rather than monitor live request activity.

The flow looks like this:

Connect usage to users, teams, and tools

Task categories become more useful when they can be viewed by user, team, or application. AI Gateway is identity-aware, providing that context without requiring teams to build a separate reporting pipeline.

This works not only for applications that teams build themselves, but also for developer tools and agent harnesses such as Claude Code, Codex, and OpenCode. By putting AI Gateway behind Cloudflare Access, teams can connect authenticated users and sessions to their AI traffic, allowing User Insights to associate activity with the right person and conversation.

For custom applications, requests must include both a stable user_id and a session_id for User Insights analysis. The exact identity configuration and field names depend on how the application or tool is set up. The important part is to provide stable, non-sensitive user and session identifiers so usage can be grouped without putting identity data in the prompt itself.

For custom applications, the request metadata might look like this:

The request body contains the model and messages for the conversation.

With Access configured in front of AI Gateway, tools such as Claude Code, Codex, and OpenCode can inherit this identity context automatically. Cloudflare Access is available at no cost for teams with up to 50 users, making it an easy way to get started.

Get started with AI Gateway User Insights

AI usage is changing quickly. Models change, teams develop new workflows, and the right choice for one group may be the wrong choice for another.

User Insights lets teams start making smarter choices by identifying where certain models may be overkill. They can then see which users and agents are driving that usage, understand the tasks behind it, and compare the cost of completing the work.

Learn more with the AI Gateway User Insights documentation . Open AI Gateway in the Cloudflare dashboard, and use what you learn to make more targeted model and routing decisions.

Pay Per Use: when AI uses your work, you should get paid

Post Syndicated from Rúben Teixeira original https://blog.cloudflare.com/pay-per-use/

AI answer engines read a publisher’s page and hand the reader a summary, so the visit, and the revenue that would come with it, never happens. Most publishers will never sign a licensing deal with the companies that use their work in AI products, and no company can negotiate with millions of sites. The web needs a way to say “yes, if you pay.” Pay Per Use is one way to say it, and it's now in beta.

In July, we outlined our plan for Pay Per Use. Since then, we’ve been working with buyers and content owners to bring it to life. The gist: A buyer offers a price for a specific use of your content. You choose whether to accept. The buyer reports each use, and Cloudflare bills the buyer and pays you. Publishers track usage and earnings in the Cloudflare dashboard. Buyers report usage through a single API.

Pay for the use, not the crawl

AI products fetch far more than they use. A search engine indexes pages it never shows. Charging for every crawl makes the buyer pay before it knows what it needs, and many buyers won't. Paying for use ties the price to the value the buyer actually gets, which we hope brings more buyers to the table and more money to publishers. Pay Per Crawl, which we launched in 2025, charges for access. Pay Per Use pays for what happens next. Publishers can choose the model that suits them.

For buyers, the case is just as simple. Some of the content your product needs is behind a block or a paywall today, and it’s the content that changes fastest: news, research, specialist trade publications. Pay Per Use lets buyers make an offer. You pay only for the content your product actually uses, you’re identified to every publisher as a verified buyer, and one API connects you to every site that says yes. You won’t need thousands of integrations or individually-negotiated contracts.

Each AI company defines the use it will pay for and sets a price. Publishers decide which offers to accept, and then can see how often their content is used and what it has earned. Cloudflare handles enrollment, usage records, billing, and payment.

How Pay Per Use works

Pay Per Use lets publishers make their content available to identified AI companies on terms that turn downstream use into revenue. Publishers choose which programs to join and retain control over crawler access and downstream uses. AI companies identify their crawling activity using Verified bots and, when permitted by the publisher’s controls, can access and index content owned by the publisher. That content can later power many experiences: a cited answer in AI search, a passage quoted in a research agent's report, a product review weighed by a shopping agent, or a recipe adapted by a cooking assistant.

Let’s show how it works!

1. Define what counts as a paid use

Each AI company sets up its program with Cloudflare. It identifies its crawler, defines the use it will pay for, sets a price, and connects a payment account.

Consider two possible offers. A search service could pay when it returns an excerpt from an enrolled page to a customer. A shopping agent could pay when an enrolled review shapes a recommendation. Under the second offer, payment follows use even if the shopper never reads the review.

The same article can create value in different products and publishers can accept different payment offers for different uses. The buyer proposes the definition of the use being paid for; Cloudflare provides the reporting and payment infrastructure.

2. Publishers choose whether to participate

Publishers review offers in the Cloudflare dashboard, under Monetize → Pay Per Use: it will show the AI company, the use it pays for, and its offer price. Publishers decide whether to accept, and can stop participating if an arrangement no longer works for them. There is no origin change or technical integration with each AI company. Each program’s terms also define what the AI company may do with the content, including any restrictions on training.

3. The buyer reports each use

The buyer fetches the list of domains that have accepted its offer, then reports each use as one line of JSON: when it happened, the URL the content came from, and an event ID.

Usage is self-reported: the program terms require complete reporting, and Cloudflare checks that each reported use maps to an enrolled publisher.

4. Cloudflare settles both sides

We aggregate reported uses, charge the buyer, and pay publishers monthly through their connected payment account. Buyers get one integration. Publishers get one place to see offers and earnings.

Payment is only half the product

Today, publishers see how often each AI company uses their content and what it has earned, by domain and over time. Business Insights already shows which crawlers visit and what they take. Pay Per Use shows what happens next: whether that content was actually used, how often, and what it earned.

Next, we’re working with AI companies to report to publishers more context about each use, such as the keywords that led to the citation, the topic of the request, or the product it powered, with no personal data being exchanged.

Our Answer Engine Optimization (AEO) tool shows how assistants answer questions about your work. Pay Per Use shows what buyers report using, and what that use earned. Together, they connect how your content is found, how it's used, and what it pays.

That helps with two kinds of decisions. Commercial ones: which content earns, which uses are worth it, and whether to keep participating. And editorial ones: what to cover, what to update, and what to make easier for agents to find.

What comes next

During the beta, we’re working directly with each buyer and with the publishers who opt in. The question is simple: do both sides want to keep going? Buyers need content that improves their products at a price that works. Publishers need a return that makes participation worthwhile, reporting they can trust, and payments that arrive on time.

Over time, we want new buyers to onboard with their own products and payment models, without rebuilding enrollment, reporting, and settlement for each publisher.

For publishers, that means one place to accept or decline offers, counter on price, and price content differently by use. Last week’s reporting may be worth more to an AI product than a ten-year-old archive page, and you should be able to charge for that.

The beta will help us determine how to make those choices simple and practical at scale.

Building an economic layer for the agentic web

Content should be able to reach new audiences and generate sustainable revenue, even when it becomes part of someone else's product: a cited search result, a shopping recommendation, an agent's report. Pay Per Use turns those uses into a commercial relationship. A buyer makes an offer, a publisher chooses whether to accept, and the use of valuable content leads to payments and records of what happened.

Not everything should be sold the same way. High-value content needs a trusted network of verified buyers who report how they use it, and that's Pay Per Use. APIs, tools, and data are frequently different, because every request is the use. For those, Monetization Gateway (also in beta as of today) lets sellers charge agents per request, using the open x402 protocol. The two run on the same foundations: identity, metering, pricing, and analytics.

Publishers shouldn't have to choose between blocking every agent and giving their work away. Pay Per Use gives them a way to say yes, on clear terms, with an account of what happened.

Monetization Gateway beta: charge AI agents for consumption with HTTP 402

Post Syndicated from Rohin Lohe original https://blog.cloudflare.com/monetization-gateway-beta/

Today, we’re making the Cloudflare Monetization Gateway available as part of a closed beta, and showcasing four customer use cases that are in production today. Since we announced the plan three months ago, we have been working closely with our customers to make the Gateway fast, flexible, and easy to use.

With just a few clicks, the Monetization Gateway allows domain owners to charge agents for access to their website, APIs, MCP tools, or datasets. Request access today in the Cloudflare Dashboard.

Proliferation of agents and machine payments

Today, most software businesses sell their products through subscriptions or prepaid credits. These business models require buyers to make a considerable upfront economic investment, so buyers limit themselves to a few subscriptions that fit into their budget. But this business model doesn’t align itself with how the predominant source of traffic on the Internet — agents — operates.

Agents seek outcomes, whether that’s sourced from a direct question or an implicit elicitation. To achieve an outcome, an agent may visit new sites, call MCP tools, or ingest data feeds. Businesses everywhere want to be in the critical path to help agents get better results. This helps them meet new users where they are, get paid for their inputs to those agents' answers, and take a leading position in headless agentic commerce.

Companies need to align their business models with the consumption patterns of agents to capture this opportunity. Payments must match an agent's consumption unit: per request, per search query, per token. The popular payment rails of today are unable to support this. These payment rails assume that the buyer will accept high latency, will only transact in large dollar values, and that they are a well-known, or identifiable, entity to the seller. Agents desire the opposite: cheap, fast, and reliable payment rails that can safely scale with minimal human intervention.

Sellers and buyers will be incentivized to align their business models with consumption, and they’ll use the payment network that helps them achieve that. Today, stablecoin transactions, and their underlying blockchain networks, are able to support these requirements. Over time, we may see existing networks introduce solutions or new payment networks come online. The Monetization Gateway exists to help ensure that sellers can focus on their product, and buyers can receive a predictable buying experience.

Building with the Monetization Gateway

The Monetization Gateway lets sellers charge agents for any resource behind Cloudflare’s network, priced per use. It's built for resources where every request is the use, like APIs, tools, and data. High-value content is different: a page might be crawled once and used a thousand times. For that, Pay Per Use offers a trusted network of verified buyers who report each use and pay for it.

The gateway uses the HTTP 402 Payment Required status code so buyers can offer payment for a resource over HTTP, inline with the request for the resource itself. This is an important distinction because there is no redirect to a checkout page and no separate payment API to call.

Sellers define which requests require payment, the cost, and where the payment should be sent. Buyers receive the payment instructions, sign an authorization, and receive the resource after the payment has been settled.

In its simplest form, this can be thought of as a ‘paywall for agents’. Sellers can go live in seconds with a few clicks, identify the audience they want to charge, and ensure that buyers are unable to access the resources until the buyer has paid.

Beyond its ease of implementation, we provide sellers a set of pricing and monetization capabilities out of the box. You write pricing rules that match any part of a request, such as the URL, headers, or query parameters, and choose from several pricing schemes. We handle the rest: payment verification and settlement through Coinbase's x402 Facilitator, failures and retries, analytics, and keeping up with changes to the x402 protocol. Payments settle on the Base blockchain using USDC, a stablecoin pegged to the U.S. dollar. Over time, we plan to help sellers make their services discoverable to agents, expose logs for all transactions, support additional payment rails, and incorporate identity primitives.

Let’s look at how our customers are using it in production today.

Cloudflare AI Gateway: pay for inference

Cloudflare’s AI Gateway is a control plane and model marketplace that allows developers to access hundreds of AI models with a single API key. Our goal is to operate the richest model catalog in the ecosystem and get it in the hands of as many customers as possible. AI Gateway allows customers to purchase credits that can be used toward all AI Gateway consumption. But as more agents carry wallets, maintaining a credit balance as the only usage mechanism adds unnecessary friction.

Starting today, U.S.-based Cloudflare customers can pay for inference at request time to a select set of models by adding the header PAYMENT-METHOD: x402. Detailed instructions are in our AI Gateway Developer Docs. Over time, you can expect to see the HTTP 402 status code embedded natively through more of Cloudflare’s infrastructure products.

AI Gateway has a complex pricing engine powering how each token gets billed. We’ve co-designed Monetization Gateway to meet the requirements of origin-controlled pricing. When this configuration is enabled, the Monetization Gateway requests pricing information directly from the seller (AI Gateway). For AI Gateway, this means being able to rely on its existing pricing module without having to maintain a rules catalog in the Monetization Gateway.

Need to set prices from your origin? Email us and we'll help you get started.

Ceramic.ai: pay for web search

"The web was built for a human buyer, someone who signs up and enters a card. Agents need to act on their own behalf, and payments are how they do it. Search is one of the first things every agent needs, so it should be one of the first things an agent can buy." — Dr. Anna Patterson, Founder, Ceramic.ai

If inference is the reasoning, search is the reality check. Ceramic.ai provides a web search API built for agents. Ceramic.ai maintains a proprietary index of more than 40 billion pages and has optimized every layer of its stack for machine callers, returning search results in as little as 50 milliseconds. At that speed, an agent can search repeatedly throughout a task.

That makes search a natural fit for agent-native payments. Search is one of the most elastic things an agent buys. A simple question might take one query, while a complex research task might take hundreds. When an agent can pay for search itself, it no longer has to ration queries against a budget someone set in advance. It can decide how hard to look based on the task, spending more when the stakes call for more evidence.

Ceramic.ai uses the Monetization Gateway’s fixed pricing monetization scheme to allow agents to pay to execute a search without an API key. Use the demo and Ceramic.ai docs to learn more.

Stocktwits: pay for stock signals

“Agents change the way data gets consumed. Instead of a customer signing up for a subscription or negotiating an enterprise license, an agent can ask for exactly what it needs, when it needs it. The ability to charge for that individual request opens up a completely new way for Stocktwits to make its data available.” – Howard Lindzon, Founder and CEO, Stocktwits

When a stock starts moving, one of the first questions traders ask is: what is everyone else saying?

For 18 years, that conversation has been happening on Stocktwits. Launched in 2008, Stocktwits pioneered cashtags (like $NET) to organize market conversations around individual stocks. Today, more than 10 million people use Stocktwits to follow markets, share ideas, and see what other investors are talking about.

Because Stocktwits is built specifically for investors, it provides a real-time view into what retail investors are watching and discussing. Stocktwits turns that activity into signals including:

  • Sentiment: whether recent posts about a ticker are predominantly bullish or bearish
  • Message volume: how much conversation a ticker is generating relative to its typical activity
  • Followers: how many Stocktwits users follow a ticker, providing a measure of sustained retail interest
  • Trending: which tickers are gaining attention fastest

These signals have long been available through Stocktwits' existing data products. The Monetization Gateway creates an additional distribution channel for a new type of customer: AI agents that need market context on demand. For example, an agent monitoring a portfolio or researching an investment might check whether conversation around $NET is taking off, whether sentiment is skewing bullish or bearish, or how many investors follow the ticker. Instead of requiring a traditional data license, a developer can pay for the individual requests their agent actually makes.

To start, Stocktwits built a separate, agent-facing path for these requests, with Cloudflare's Monetization Gateway sitting in front of it. Its existing API and enterprise data products remain unchanged. Each request is priced individually, giving Stocktwits a way to make its market signals accessible to AI agents while preserving its existing data products and licensing models. Get started with their docs today.

API2PDF: pay for API access

API2PDF is a REST API that helps developers generate PDFs from HTML and Office documents. Since launching in 2018, API2PDF has seen a growing number of AI agents directing developers to its service and now treats them as a first-class customer.

API2PDF’s consumption-based pricing model charges customers for the bandwidth and compute required to fulfill each request. Previously, customers needed to create an account and obtain an API key before using the service. After the first month, they needed to supply a credit card, which resulted in a >50% drop off in conversion. API2PDF now uses the Monetization Gateway to return an HTTP 402 response when a user or agent makes a request without an API key. Because each request depends on variable compute and bandwidth costs, API2PDF uses variable pricing to inform agents of the maximum price a single request could cost. Once the client completes the payment, API2PDF fulfills the request and settles only the actual consumption.

Agents face the same build-versus-buy decision developers do. An agent that needs a PDF could burn tokens to generate one itself, or it could pay a specialized API like API2PDF a fraction of a cent to do it properly.

Get started today

The Monetization Gateway is available in closed beta to eligible U.S.-based sellers and buyers, with support for new geographies on the way. If you are interested in making your product agent-native, we want to hear from you. Fill out the onboarding application in the Cloudflare Dashboard, check out our Developer Docs, and we’ll get back to you shortly. If you have questions, ideas, or want to help us scale agentic payments, please send us an email.

We’d like to thank our partners for their contributions and feedback. Without them, this launch wouldn’t be possible. This includes Vail Gold and Jeff Rafter from Cloudflare AI Gateway; Dr. Anna Patterson, Sean Costello, Sadé Ried, and Autumn Yuan from Ceramic.ai; Leo Gorkin, Ethan Berk, and Santiago Sanchez from Stocktwits; and Zack Schwartz from API2PDF.

Simplifying domains for people and agents

Post Syndicated from Ankit Shah original https://blog.cloudflare.com/simplifying-domains/

You just thought of your next great idea, and buying the right domain feels like the easiest way to make that first bit of progress. Naturally, you open a new tab in your browser, only to find yourself face-to-face with an experience that feels like a budget airline peppering you with add-ons at checkout: Want security? How about a website? Do you want email? You’re just a few minutes into building your next idea, and it doesn’t feel fun anymore.

Launched a decade ago, Cloudflare Registrar has always taken a simpler approach. Domains at cost, transparent pricing, and no unnecessary upsells. But simplicity shouldn’t begin at checkout. It should begin the moment you start looking for the right domain.

Today, we’re bringing that same simplicity to the entire experience of finding and buying a domain. Our new domain search shows every extension we support, responds as quickly as you type, and makes hundreds of possibilities easier to explore through sorting, filtering, and transparent pricing.

And you know what is particularly good at ignoring distractions and staying focused on the destination? An AI agent. We designed Cloudflare Registrar to work naturally with agents through the Registrar API, MCP, and our newly launched cf CLI. You can ask your favorite agent to find the right domain, buy it, or transfer one you already own.

The agentic registrar, expanded

In April, we launched the Registrar API beta, allowing developers and agents to search for, check, and register domains programmatically. We have expanded the API since then. The new sandbox lets you test registrar workflows without purchasing a domain or triggering a real transaction. Our extensions endpoint returns relevant information for each of the 420+ extensions we support, helping you account for the different requirements across registries. We also added transfers, so you can bring domains from another registrar into Cloudflare programmatically.

The Registrar API is available through Cloudflare MCP, giving agents access without requiring a separate integration. Earlier this week, we also announced the launch of cf CLI, bringing the same capabilities directly into your terminal. You can prompt your favorite agent to search for, register, or transfer a domain. These tasks already lend themselves naturally to a conversation:

  • “Is example.com available?”

  • “Buy example.com.”

  • “Transfer example.com from my current registrar.”

The way people interact with the Internet is changing. Cloudflare Registrar should feel natural whether you use it through an agent or in your browser. For many people, the browser is still where the search begins, and that experience was long overdue for an overhaul.

Search simplified

Previously, our search page showed around 20 available results from a subset of extensions, sometimes modifying your search term to suggest related options. You couldn’t see that exact name across every extension we support.

We decided to take a simpler approach: show you the exact term you searched across every supported extension. Results appear as you type and continue to load as you scroll, letting you explore hundreds of options without starting another search. We also include domains that have already been registered, giving you a more complete picture. If you only want domains available to buy, you can filter everything else out.

More results might not sound simpler, but these are the results you asked for. Sorting and filtering help you narrow them down. Whether you are logged in or logged out, on your phone or at your desk, the experience feels the same.

Making a complex question feel simple

Cloudflare Registrar supports 420+ extensions. A single search can thus create more than 420 separate availability questions. Each extension is operated by a registry that maintains its official registration records and provides the authoritative answer about a domain’s availability and price.

Asking every registry every question at once would be slow and wasteful. Registries respond at different speeds and impose request limits, and much of the work would be for results the person might never view. To solve this, our search gathers evidence from multiple sources: (1) prepared availability datasets (e.g. zone files) and cached answers; (2) DNS answers; (3) live registry lookups.

A hit against an availability dataset or DNS can tell us that a domain is already in use, but a miss cannot necessarily prove availability as a domain may be registered without being configured in DNS, or it may be on a blocked list. A recent registry-derived answer is stronger but becomes stale over time, while a live registry check provides the freshest authoritative answer but takes longer and draws on limited upstream capacity. Our new search progressively probes these sources while balancing speed, freshness, and certainty for each result. As better or more accurate information arrives, we update only the affected result dynamically.

How we built search for speed and scale

We built the new search on the same Cloudflare developer platform available to our customers. Workers run the public search entry point and the services that gather availability evidence. Durable Objects give each active search one coordinator, while Workers KV stores prepared data that can be reused across searches. Together, these primitives let the service scale across users and 420+ extensions while keeping operating costs low.

We prepare useful evidence before a search begins: a purpose-built pipeline converts registry zone files and other bulk sources into compact availability datasets in Workers KV. Large datasets are split into smaller pieces, so a lookup retrieves only the data required for that domain. These fast checks can answer many questions without making a new live registry request.

A search session is composed of a Durable Object that coordinates that specific search interaction. It establishes the result order from the query, sort, and filters without waiting for network lookups, remembers the best evidence received for each domain, and tracks which results are visible so lookup work follows the person's attention.

WebSockets provide the bidirectional connection, and we designed an application protocol on top of them to connect the browser to the resolution process. An initial snapshot establishes the ordered list. Subsequent delta messages contain only the fields that changed, letting the browser update one domain instead of downloading the full result set again. Before sending a delta, the Durable Object compares the new evidence with the current answer: stronger evidence can replace it but weaker evidence cannot.

A separate Worker gathers additional evidence. It can query DNS through Cloudflare's 1.1.1.1 resolver, make a live Registrar check, or reuse a recently cached answer. Reusing fresh answers avoids repeating upstream requests. Because an available domain can be registered at any moment, an available answer has a shorter useful cache life than evidence that a domain is already taken.

Together, these pieces turn hundreds of independent availability checks and all that coordination into one coherent search experience. Cloudflare’s Developer Platform gives us all the building blocks to hide that complexity and craft a domain search experience that feels simple and is among the fastest in the world.

Transparent pricing and price drops

Making search feel simple is not only about speed. It is also about knowing exactly what a domain will cost. Cloudflare Registrar has offered domains at cost since day one. Great prices are part of making domains simple, but so is knowing what you will pay. Our new search and our new pricing page show both the initial registration price and the renewal price for every domain. When a domain is discounted, we show the original at-cost registration price crossed out alongside the promotional price.

Beginning with Birthday Week (this week!), we’re offering first-year registration discounts on select extensions including .io, .dev, .app, and .tech. You can explore every discounted extension directly from the search page as well as our newly launched pricing page.

Whether you search in your browser or ask an agent, our goal is the same. Remove the friction between having an idea and making it real. Buying a domain for your next idea should be fun and feel like progress.

Find your next domain

Choose how you want to get started:

  • Search in your browser: Explore every available extension and find your next domain.
  • Ask your favorite agent: Install cf CLI, then prompt your agent to search for, register, or transfer a domain.
  • Build with the API: Use the Registrar API to bring domain search and registration into your own application or workflow.

However you choose to do it, finding your next domain should be the fun part.

Acknowledgements: This simplicity was a result of cross-team collaboration. Special thanks to Pedro Menezes, Shobhit Kuruvilla, Lucy Dryaeva, Fred Pinto, and the Registrar Team, the Design Engineering Team, and the Forge team.

Detect and send production issues straight to your agent

Post Syndicated from Thomas Ankcorn original https://blog.cloudflare.com/real-time-issue-detection/

As agents help us build more complex applications, both humans and agents need a better way to stay on top of what goes wrong in production. Coding agents can already query observability data, navigate a repository, change code, write tests, and open a pull request. What remains manual is connecting those steps: recognizing that repeated failures come from the same bug, gathering the relevant logs and traces, sending that context to an agent, and checking whether the fix worked. Without that structured handoff, the agent must search raw telemetry to reconstruct the scope and context of the failure before it can investigate.

Today, we are introducing Issues, built-in error monitoring for Cloudflare Workers (now in open beta!) to streamline this workflow. Issues can:

  • Group repeated exceptions, 5xx responses, and error logs into one issue.
  • Send the error, stack trace, logs, traces, and Worker version to a configured coding agent.
  • Trigger the agent’s configured workflow — from triaging an issue to querying more data to opening a pull request.

Fix your first issue with the following prompt for your agent with CF CLI or checkout the documentation to get started:

Catch failures automatically

With one line of configuration, you can start receiving Issues detected on your Worker with no additional instrumentation required. Issues are built into the Workers runtime, so there is no SDK to install or application wrapper to add. 

Once enabled, Issues records uncaught exceptions, failed invocations, HTTP 5xx responses, output from console.log() and console.error(), and logs that contain a stack trace. It also flags runaway alarm conditions and code that writes large volumes of logs inside loops.

Consider a Worker whose handler starts throwing errors after a deployment. Every failed request has a different request ID, but they all come from the same bug. Issues groups them together and shows when the error first appeared, how many times it has happened, and whether it is becoming more frequent.

When you open an issue you can see the error, a stack trace when available, the logs and traces leading up to it, the Worker version, request details and trend of the issue over time, as shown below:  

Contextualizing errors for your agent

Cloudflare can capture what happened inside the Worker, but it does not know which users, accounts, or sessions matter to your application. Use the Worker runtime's built-in OpenTelemetry API to add those identifiers, without installing another package:

Those identifiers appear with each occurrence. In this issue, you can now see whether failures are concentrated in one account or session before sending the issue to an agent.

Send detected issues to your agent

An issue no longer has to sit in a dashboard while someone copies a stack trace and pastes it into a prompt. Configure an automation once, and when an issue crosses an occurrence threshold or returns after a quiet period, Issues sends it straight to your agent through Automations. You can choose when the automation should run and where the issue should go.

This can be via:

  • Built-in coding agents: Connect Claude Code with a routine ID and token, Cursor with an automation webhook URL, or Devin with an API token and organization ID.
  • Generic webhooks: Send issue context to your own agent or HTTPS endpoint.
  • Chat and incident management: Notify your team through chat or an on-call workflow.

When the automation runs, Issues sends the failure summary and diagnostic context captured with the issue — the exception, error, source-mapped stack trace, leading and trailing logs and traces, Worker version, and the application context you added. For deeper investigation, you can connect the agent separately to Cloudflare MCP that lets the agent query the related logs and traces so that it can propose code and test changes and open a pull request.

You stay in control of what reaches production: review the pull request, deploy the fix, and mark the issue resolved.

How Issues uncovered and resolved two Workflows bugs in one day

Cloudflare Workflows, a primitive that powers long-running, multi-step applications, is built entirely on the Workers platform.

Behind the scenes, its services keep track of steps, retries, and saved state. This makes Workflows a useful place to test Issues on our own production systems. Within a day of turning it on, the team found two unusual problems hidden inside a large volume of traffic.

  • A migration stuck in a retry loop: A Workflows control plane migration repeatedly hit a SQLite foreign key error when attempting to apply migrations in an edge case. Issues allowed the Workflows team to identify the problem and fix it.
  • A deletion process that never completed: Workflows discovered that during deletion of Workflow instances, there was an edge case where they could exceed a Workers subrequest limit and not finish the deletion. Issues helped the Workflows team identify the issue and do a fix.

Instead of leaving the team to connect thousands of separate pieces of telemetry and user reports, their automation setup sent these issues directly to Cloudflare OS, which followed the errors into the Workflows code and proposed a fix for both issues.

Get started

Ready to see what Issues finds in your app? To get started:

  1. Set observability.issues.enabled to true in your wrangler.jsonc file
  2. Set up your first automation in the Cloudflare dashboard to send issues to your destination of choice whether that’s an agent, a webhook, incident management tool or chat platform.

If your agent is handling setup, it can also use the new cf CLI to inspect issues and create automations. Check out our documentation to learn more!

Cut your AI spend with AI Gateway’s Auto Router

Post Syndicated from Ming Lu original https://blog.cloudflare.com/auto-router/

From our conversations with companies at every stage of their AI adoption journey, we've seen some common patterns. First, there is an exploration period as you bring on every new tool, dole out API keys freely, and let the tokens flow. Then, you converge on the canonical tools for your organization for agentic coding, for non-technical workflows, for running and deploying agents. As companies formalize their AI adoption, they want to manage and oversee token spend for users, but budgets and rules only go so far. The best savings are the ones users never notice.

Today, we are releasing Cloudflare's Auto Router in public beta, available through AI Gateway. Set your model to cloudflare/auto and the Auto Router will automatically route each request to a model that is capable enough for the task, without requiring an end user to think about model selection. Our early results using the Auto Router internally through our OpenCode harness show a cost savings of up to 30% when compared to using only frontier models like OpenAI Sol and Anthropic Claude Opus.

Why we built this

From our own experience tracking AI spend at Cloudflare, we’ve learned managing costs requires a multipronged approach. Previously, we talked about how to set budgets and limits around AI spend, and how to see who is spending across your organization by linking employees to their AI usage.

In many harnesses, including OpenCode, Claude Code, and Codex, individual users still select models manually. Of course, not all tasks are created equal, and often individuals end up using models that are overkill for their work. For example, you don't need Opus-level intelligence if you're looking to summarize an email or chat threads. However, you wouldn't want to block that model completely from your security engineering team.

Our goal is for AI Gateway to be the control plane for organizations deploying AI internally. Because every request from every user, agent, and tool already flows through it, AI Gateway is in a unique position to do more than observe and enforce. Budgets, spend limits, and identity-aware analytics give organizations visibility and guardrails, but they still rely on individuals to make cost-conscious choices request by request. The next step is for the gateway itself to make intelligent decisions on a user's behalf: sending each request to a model that is capable enough for the task. That way, organizations reduce spend automatically, while users keep access to the most capable models when their work actually needs them.

The results

We use Auto Router internally at Cloudflare within our OpenCode deployment and within Cloudflare OS, our custom agent harness. In our internal usage, we’ve seen results comparable with frontier models for coding tasks.

Auto Router does best when used across a wide range of knowledge-work tasks, like those typically found in a large organization with work spanning both technical and non-technical teams. We evaluated cloudflare/auto against OpenAI’s GPT-6 Sol and Anthropic’s Claude Opus 5.5 on our internal general knowledge work benchmark. The benchmark uses simulated workspace tools and covers common day-to-day workflows across email, calendars, Slack, files, travel and finance. Each task requires the model to use these tools to produce a verifiable answer or complete an action.

Model

Successful Trials

Success Rate

Total Cost

Cost per success

cloudflare/auto

252/291

86.6% (+6.2/−6.9 pp)

$2.10

$0.0084

Anthropic Claude Opus 5.5

281/291

96.6% (+2.7/−3.8 pp)

$5.91

$0.0210

OpenAI GPT-6 Sol

245/291

84.2% (+6.5/−6.9 pp)

$2.64

$0.0108

97 tasks with three samples per model per task. Parenthetical values show 95% confidence intervals estimated from 10,000 task-level bootstrap resamples, preserving all three repetitions within each task. “pp” indicates percentage points.

Our Auto Router delivered similar performance to other state-of-the-art daily-driver models, coming in at 80% the cost of Sol and 35% the cost of Opus. While that may initially seem surprising, one way to frame the problem a model router solves is through the “jagged frontier” across models. The ability to solve a problem often exists somewhere in this portfolio of models; the router’s job is to choose the right model for each task while balancing quality and price. Savings come from not paying frontier rates for non-frontier work, and they grow with how much of that work you have.

Another insight is that lower token prices do not always produce lower-cost outcomes. A model that looks cheaper on paper may end up using disproportionately more tokens to solve a problem. A router should minimize predicted trajectory cost, not just load-balance by dollars per million tokens. 

This is already useful today, but it’s only the beginning of what the Auto Router can learn from Cloudflare’s position in the inference path.  

How it works

When you send a request to cloudflare/auto, AI Gateway first builds the pool of models that can actually serve it. It filters out models that do not support the request format or execution mode, and accounts for the credentials, billing configuration, access control policies, and spend limits attached to the gateway. It will also filter out unhealthy upstream providers or models during downtime and automatically bring them back into the pool after an outage.

For the remaining candidates, the router looks at a compact view of the conversation. It considers the most recent messages, prioritizing the newest turns. The conversation is then sent to a multi-head classification model running on Workers AI and deployed on GPUs across our edge network. The classifier produces two sets of signals. First, it assigns probabilities across 14 task categories (like coding, planning, research, data analysis). It then rates the request across four dimensions on a scale from one to five: complexity, ambiguity, stakes, and dependence on earlier context.

A separate scoring matrix combines those signals with model benchmark results to estimate how well each model fits the request. To calibrate the scoring matrix, we defined the preferred model for a set of example task and difficulty profiles, then adjusted the weights to produce those choices.

Finally, the router combines expected quality with each model's input and output token prices. On straightforward requests, price carries more weight, so a smaller model can win when it is capable enough. As difficulty rises, the cost penalty falls and stronger models have more room to win. In simplified terms, cloudflare/auto selects the model with the highest utility as defined by:

For long agentic sessions like debugging or coding, cost is less driven by the model’s list price than by the cost of cache reads, which grows with session length. Switching models throws the cache away and forces a new model to write the whole context again. This can be worth it, as a model with a cheaper cache-read and cache-write prices can pay back the rewrite quickly.

Rather than completely avoiding model switching, the Auto Router accounts for the cost of cache reads and writes. Within a turn (one user input loop), the cache is hot and switching rarely pays off, so it’s better to keep using the same model. Across turns, the Auto Router applies a switching penalty that grows with the number of tokens already in context. A model that still holds a live cache for the session is priced at its cheaper cache-read rate. Every other candidate is priced at the full cost of rewriting the context, so the deeper the conversation, the more a switch has to earn back, through higher quality results that use fewer tokens overall or cheaper cache rereads. Switching models has another cost: most models can't read another model's reasoning tokens, so a model switch that drops reasoning tokens means that the new model may have to redo it at output prices. In the future, we want to account for this by having the router prefer to stay within the same model family when it switches.

From there, the router returns a ranked list. AI Gateway attempts the winner first and can move to another eligible model if that provider cannot serve the request.

This overall design has several benefits. The two-stage architecture (task and dimensions classifier to scoring matrix) means that routing decisions are legible because you can inspect each task’s predicted category and complexity to see how it translated into the model choice. Adjusting the router when a new model is released also does not require retraining — we only add its benchmark-derived weights to the scoring matrix. The same classifier can also support different routing profiles. For example, in addition to cloudflare/auto, we plan to release other routers in the future, including cloudflare/auto-best, which uses the same classification and model pool, but selects the highest expected quality without applying the cost tradeoff.

What's next

Our release today is only the starting point, and we’re continuing to invest in research and new routing strategies. In the near term, we want to:

  • Expand the models offered through cloudflare/auto
  • Include zero-data-retention requirements when filtering models
  • Account for provider capacity when selecting models
  • Select the appropriate reasoning or thinking level for each request
  • Add full support for the Responses API and WebSockets
  • Explore structured decision models as a first-pass classifier

The Auto Router is free while in beta. Read more in our developer documentation.

Acknowledgements: This project was also made possible by the efforts of Mats Dodd, Sam Scott, Oliver Yu, and Jeff Rafter.

Cloudflare Containers, rebuilt to scale agent sandboxes

Post Syndicated from Thomas Gauvin original https://blog.cloudflare.com/faster-agent-sandboxes/

Today, we’re making Cloudflare Containers more programmable and optimized for agent workloads. Agents don't deploy sandboxes ahead of time. They create sandboxes on demand, for each task, expect them to be ready immediately, and be able to pause and resume. So we rearchitected Containers to meet these requirements: your code can now choose each sandbox's image and instance type at runtime, Containers start 6x faster, and filesystem snapshots are available in public beta.

To make this possible, we’ve rethought the Containers infrastructure from the bottom up. A new scheduling policy moves control over each sandbox into application code, while a redesigned runtime provides a faster path to a running Container. In ComputeSDK’s independent benchmark, median startup fell from just over four seconds to 648 milliseconds, and, in our own preliminary tests, burst testing successfully created hundreds of thousands of containers in seconds.

All of this builds on what has always set Containers on Cloudflare apart: every Container gets its own Durable Object, a persistent, programmable controller running right next to it that manages its lifecycle, outbound traffic, and more. We are bringing more capabilities directly to the native ctx.container API, so the Durable Object can control its Container without a wrapper class in between, and we’re carrying this model into Sandbox SDK 1.0.

As we wrote earlier this year, your agent needs a computer. These changes make Containers a better complement to Workers, Dynamic Workers, and Durable Objects when agents need a full Linux workspace.

Rethinking Containers’ runtime for agents

Until now, Cloudflare Containers has been organized around application deployments. You choose an image and compute resources at deploy time, roll that configuration out across the application, and manage it centrally. The application is the unit of configuration and rollout.

An agent workspace is different: it’s created on demand, while the agent is working. The task determines its image, resources, tools, and starting filesystem. It might exist for a few minutes, sleep between requests, or be restored days later. Those decisions need to live with the application code handling the task, and the agent's sandbox needs to start up fast, because every second of startup is time your users spend waiting.

We’ve seen this pattern with Base44 on app-building workspaces and Kilo Code on cloud-agent sessions. It also appears in our integrations with Cursor Cloud Agents, Devin Outposts, the OpenAI Agents API, and Claude Managed Agents.

Each of these workloads needs something different from the workspace. Coding agents need repositories, package managers, compilers, test runners, and development servers. Evals need sandboxes that begin from a known state. Reinforcement learning systems need to create, grade, and reset large numbers of environments. Longer-running tasks need to preserve the files an agent produces, so work can continue later.

These requirements led us to fundamentally rethink how Cloudflare Containers are configured, scheduled, and saved. The result is a new way to provision and schedule Containers: the durable_object scheduling policy. It lets your code choose each sandbox’s image and compute resources at runtime, starts Containers more than 6x faster, and supports filesystem snapshots, so workspaces can be saved and restored.

“At Base44, we help anyone turn an idea into a working app. Cloudflare Containers gives each app an isolated development environment where our AI can execute commands, install dependencies, and bring changes to life in a live preview.”

Dolev Epshtein, Software Engineer, App Infrastructure at Base44

“At Kilo Code, every cloud-agent session needs its own workspace and environment, with the right repository, tools, and user configuration. Cloudflare Containers lets us create those isolated environments on demand, so our agents can start running commands quickly and get to work for our customers.”

Emilie Schario, Co-founder of Kilo Code and VP Engineering, AI Workspaces at Anaconda

Choose the sandbox your agent needs, from code

From the start, every Cloudflare Container instance has been attached to a Durable Object. The Durable Object gives the environment a stable identity and lets application code control when it starts, sleeps, and stops. This model has proven particularly well-suited to agent sandboxes.

Until now, though, the two decisions that matter most for an agent sandbox were locked in at deploy time: which image it runs and how much compute it gets. Each combination of image and instance type was its own Containers application, with its own Durable Object namespace, set up ahead of time with wrangler deploy. 

Say one agent needs a small Node.js sandbox and another needs a large Python sandbox for builds. With our previous approach, that required two applications, two namespaces, and routing logic in your Worker to send each task to the right one. Every new environment meant another deployment. 

The new durable_object scheduling policy removes that. The image and instance type are now arguments your code passes when the sandbox starts. To opt in, set the scheduling policy and declare the images your Durable Object can choose from:

Each image you declare is available as this.ctx.container.images.<name> within the Durable Object. When a task arrives, your code picks the image and instance type for that task:

This code makes two decisions after the task is known. It chooses the toolchain the workspace needs and how much compute to give it. What used to take a separate application and a separate wrangler deploy is now an if statement. One Durable Object class can start Node.js and Python sandboxes of different sizes side by side, and adding a new environment is a code change, not a new deployment. That’s the idea behind this whole update: infrastructure becomes code that runs at request time, right down to the environment itself.

Rollouts are now just code

Choosing the image at start time also fixes one of the most painful parts of running Containers: rollouts.

Before, updating an image meant updating the whole application. You set grace periods, so running instances could drain, defined percentage splits to move traffic gradually, and called our API to push the new configuration. Throughout that process, the platform decided which instances got replaced and when, whether an agent was in the middle of a task.

With the durable_object scheduling policy, there's no rollout configuration at all. A Container can keep running the image it started with until your code stops it. The next time that Durable Object starts a Container, it uses whatever image your code chooses. That means any rollout strategy you want is a few lines of code:

For example, you can:

  • Canary a new toolchain on 5% of new sandboxes by hashing the Durable Object ID.
  • Pin active projects to their current image, so an agent never has its environment swapped out mid-task.
  • Migrate a workspace at a natural checkpoint, like the next session or after a snapshot.
  • Roll back by changing which image future starts choose. No config push, no waiting for a drain.

The rollout policy lives in your Durable Object code, next to the rest of your logic, and it can be as simple or as sophisticated as you need.

Each of these improvements comes from leaning further into the Durable Object, which already owns the workspace's identity, state, and lifecycle. Letting it choose and control its Container gives you more control over every instance and its rollout. It also gives the scheduler a better place to start the Container: wherever the Durable Object is already running. That's what gets the agent to its first command faster.

Faster first commands

Previously, starting a Container required our global control plane to resolve the application configuration, find capacity, and coordinate placement. That model works well for application-wide fleets, but it put deployment machinery in the path of an agent’s first command.

With the durable_object scheduling policy, demand begins at the Durable Object. The Containers infrastructure serving it looks for capacity on the same machine first, then widens the search within the same location if it needs to. It also favors hosts that already have the Container’s image or snapshot in local storage, so the Container can start without downloading it first. 

We also cut work after the Container reaches a host. Instead of booting a new virtual machine from scratch, the new runtime restores a prepared virtual machine that isn’t yet assigned. It reuses networking and filesystem setup, batches repeated operations, and no longer waits on services the first command doesn’t need. 

Together, these changes substantially reduce the time it takes to go from creating a sandbox to running a command in it. On ComputeSDK’s independent Burst TTI Benchmark, which launches 100 sandboxes concurrently and measures time-to-interactive from the client:

Startup measurement

Previous scheduling path

New scheduling policy

Improvement

Median

4.049 seconds

648 milliseconds

6.2x faster

95th percentile

5.839 seconds

910 milliseconds

6.4x faster

99th percentile

6.717 seconds

1129 milliseconds

5.9x faster

The new path also holds up under burst load. In our preliminary burst test, a single account started 100,000 Containers in 5.387 seconds across six locations.

Start with a prepared system image

As the scheduling path gets faster, preparing the image becomes a larger part of the remaining wait. Before a Container can start, its image has to be on the host and unpacked into a filesystem. When that work happens after the request arrives, the agent is left waiting.

That’s why we are introducing cloudflare/debian-trixie: a ready-to-use system image for agents that can configure their environment at runtime, containing Debian Trixie Slim and Node.js 24.20.0 LTS:

This means your agent can start a Linux sandbox without first creating a Dockerfile, building an image, or pushing it to Cloudflare. Once the sandbox is running, your agent can use exec() to clone a repository, install packages, and configure the environment for its task.

And because Cloudflare controls this image, we can distribute and prepare it across eligible Containers hosts before requests arrive. Startups don’t have to download or unpack the base image while the user waits.

Save the workspace and return to it later

A fast start still leaves one more wait: setting up the workspace. Cloning a repository, installing dependencies, and configuring a toolchain can take much longer than starting the Container itself. As the agent works, it also produces files you want to keep. Repeating setup on every start wastes time, and losing the agent’s changes makes it difficult to continue a task.

That is why we’re adding native filesystem snapshots to Containers, in public beta. Snapshots let an agent begin a task in a prepared environment, save its workspace when the task pauses, and restore those files when the session resumes.

Snapshots enable two useful patterns.

First, one workspace can continue across many sessions. For a coding agent, that might mean saving the workspace when the user finishes working and restoring it when they return the next day. The repository, installed dependencies, build caches, configuration, and edits are available without rebuilding the environment.

Second, a snapshot can provide a shared checkpoint for many sandboxes. Since snapshots are immutable and reusable, multiple Containers can start independently of the same prepared environment and make their own changes from there.

Agent evaluations are a good example. An eval might run the same task across different system prompts, skills, models, or agent versions. To compare the results, everything else has to stay fixed: the repository, dependencies, tools, and input files. One snapshot can start many isolated environments from the same baseline, reducing setup time and preventing environment drift from affecting the results.

Snapshots also complement the new system image we introduced above. An agent can start from cloudflare/debian-trixie, set up its environment with exec(), and save the result as a snapshot. Future sandboxes then start from that snapshot with the repository, dependencies, and toolchain already in place.

With snapshots available through the new durable_object scheduling policy, Containers can act as persistent agent workspaces. Compute can stop when work pauses, and a new Container can start from the latest snapshot to pick up where it left off.

The Durable Object advantage for agent sandboxes

The faster scheduling path, runtime image selection, and filesystem snapshots all come from the same design choice: leaning further into the Durable Object as the controller for its attached Container.

Agent systems need a programmable, stateful environment outside the Container to keep state, hold credentials, and control the sandbox’s lifecycle. You can run the agent there, following the “decoupling the brain from the hands” pattern described by Anthropic: when the agent is separate from the sandbox where it works, the agent stays available while its sandboxes and tools can start, stop, fail, or be replaced independently. Or, if you run the agent inside the sandbox, the outside environment lets you supervise it and report progress back to the user.

This is where the Durable Object and Container architecture becomes uniquely well-suited. Every Container is attached to a stateful Durable Object with its own compute and storage running right next to it. You can run the agent in the Durable Object and use the Container as its workspace, or run the agent in the Container and use the Durable Object to supervise it. Add Dynamic Workers for lightweight isolated execution, and an application can choose the execution environment each task requires.

What’s new with the durable_object scheduling policy is that the Durable Object can now control its Container directly, with no wrapper class in between. exec() runs directly in the Workers runtime, and outbound request interception, runtime image and instance selection, and filesystem snapshots are all available on ctx.container. You can combine them with Durable Object storage, alarms, WebSockets, RPC and the rest of your code.

This makes the Container a compute extension of the Durable Object. The Container supplies the Linux environment, while your Durable Object retains the sandbox’s identity, state, policy, and lifecycle. That split is especially well-suited for several patterns:

An agent can remain available while its Linux workspace sleeps. The agent loop can run in the Durable Object, where it maintains session state, communicates with the user over WebSockets, and calls models. It can wake the Container when it needs a shell, compiler, or development server, then stop that compute while it waits for the user or model, paying nothing for idle Linux compute.

The Durable Object can program the security boundary around its Container. It can remember which services, repositories, and operations a user has authorized, then update the Container’s Outbound Request Handler to inject newly granted credentials, enforce new policies, or record additional activity. This resembles the Gatekeeper pattern used by Cloudflare OS, applied to each agent computer.

Evals and reinforcement learning systems can supervise each attempt from outside the environment being tested. A coordinator snapshots a base workspace and forks it into N attempts, each with its own Durable Object and a Container. Each Durable Object runs its attempt, monitors the run, and preserves the result, even if the Container crashes. The coordinator grades the attempts, snapshots the best one, and forks again from there. 

These patterns don’t fit one generic lifecycle. Native APIs let you combine the Container with the Durable Object primitives your application needs, while still using higher-level utilities where they help. You keep control over the Container’s lifecycle, policy, and state.

What this means for the Container class and Sandbox SDK

When we launched Containers, we deliberately hid the Durable Object behind the Container class. We wanted sandboxes to feel familiar and match what developers expected from other platforms. The Sandbox SDK was built on that class, and it filled real gaps: back then, the runtime had no native command execution, outbound request interception, or snapshots, so we built them in userspace.

Since then, agent workspaces have become one of the main workloads shaping Containers, and the cost of that abstraction has become clear. By hiding the Durable Object, we made it harder for you to see and combine the identity, state, and coordination it provides with the Container it controls. Nearly every team we worked with needed something slightly different from the generic lifecycle: their own sleep policy, their own credential handling, their own way of tracking eval runs.

These capabilities are now native, so we're making the Durable Object explicit in the developer experience:

  • New capabilities are native-only. The durable_object scheduling policy, faster startup, runtime image and instance selection, and filesystem snapshots are available only through ctx.container.
  • We'll maintain the Container class and legacy Sandbox class through December 31, 2026. Existing deployments keep running after that date, but the classes won't get updates. We recommend migrating to ctx.container.
  • Sandbox SDK 1.0 is a set of utilities, not a base class. Its helpers work inside your own Durable Object class, alongside ctx.container.
  • For a higher-level environment, @cloudflare/computer combines Dynamic Workers and Containers with a synchronized filesystem.

Migrating usually means changing extends Container to extends DurableObject and calling this.ctx.container directly. See the migration guide for details. Or, get started with your agents:

Get started

Today, most agents are measured by what they can accomplish in a single session. As agents take responsibility for projects that unfold across hours, days, and weeks, the environments where they work need to keep up.

We want every agent to be able to create the sandbox for the task at hand, release the compute when work pauses, and return to the same workspace when the project continues. Today’s changes bring us closer to sandboxes that are as programmable, persistent, and ready to work as the agents using them.

Try the new durable_object scheduling policy, available to all today in public beta, and see what faster startup, filesystem snapshots, and runtime configuration unlock for your agents:

Acknowledgements: This project was also made possible by the contributions of Greg Anders, Andrew Martinez, Nafeez Nazer, Kian Newman-Hazel, Sebastien Pahl, Naresh Ramesh, Cody Roseborough, Nikita Sharma, and Sarah Snell.

The Internet has a second audience

Post Syndicated from Matthew Conroy original https://blog.cloudflare.com/agentic-web/

For most of its history, the Internet had one audience that paid the bills: people. We read the articles, saw the ads, and bought the subscriptions. Bots were always there, but they were mostly large, automated operations that didn't view ads, pay for anything, or read in any meaningful sense.

That's changing fast. At the end of 2024, Cloudflare handled an average of 63 million HTTP requests a second. Today, it's almost doubled to 115 million, with peaks above 150 million. Over the past year, daily requests from AI agents on our network grew by more than 1,700%. This year, for the first time, more than half of Internet traffic wasn't human.

The human web didn't shrink to make room. A second audience arrived alongside it: agents, software acting on behalf of people. They sit somewhere between humans and traditional bots. They don't respond to ads, but there's usually a person behind them with a job to get done. For businesses that learn to serve them and capture value from them, agents are additive. For those that don't, they're extractive.

What our customers need hasn't changed: to be discovered, to tell great stories, to build great experiences, and to sell. What's changed is that more than half your visitors are now software. Our job is to help you serve both audiences.

More traffic, less revenue

For thirty years, the web ran on one arrangement: you let search engines crawl your site, they sent you visitors, and you turned those visitors into a business. Being found and getting paid were the same thing.

AI has caused this delicate balance to break down. Now, answer engines read the page and give the reader a summary. This costs websites bandwidth without leading a human to a website where the ads or payments happen. The machines kept coming, and the audience that paid for the web stopped reaching those sites. Some of the most heavily crawled categories, like Retail, Computer Software, IT & Services, and Financial Services, have seen human traffic decline as much as 40% in less than one year.

The result is that revenue per request is falling while costs are rising. Every automated request still costs bandwidth, compute, and origin capacity, and a growing share of those requests carry no referral, no ad impression, and no subscription. Our first instinct was to block all automated traffic. Last year we recommended blocking AI training crawlers on new domains so site owners could at least say no to their content being used to build models. In Spring 2025, 22% of the crawler requests we saw were for AI training (according to the crawlers’ stated purpose). By June 2026, it was 52%. The problem is a blanket “no” is not a sufficiently nuanced approach for the Internet economy being built right now.

The opportunity is there to cater to agents. Get it right, and you are at the forefront of a new business model. Get it wrong, however, and the results will be the same as they were for generations of websites that were on the wrong side of search engine algorithm changes.

Some of that traffic is a customer

An agent booking a table, comparing insurance quotes, or buying a dataset for a researcher is a customer. It just isn't a human one.

The fastest-growing part of automated traffic is no longer crawlers. It's agents: software fetching pages on a person's behalf, often because the human asked a chatbot something. That agent traffic follows human routines, with a weekly rhythm and a dip over the summer holidays. Turn an agent away, and you may be turning away the person who sent it.

Agents also behave differently from training crawlers. A training crawler collects your pages to build a model. An agent comes back each time someone asks about that content, so this traffic grows with how many questions people ask, not how much you publish.

You can't do business with an audience you can't see, can't tell apart, can't set terms for, and can't charge. Until recently, for most of the web's non-human traffic, none of those four things were possible.

See who’s really visiting

"AI bot" no longer means anything useful. What matters is what a bot does. Cloudflare’s AI Crawl Control, Business Insights, and BotBase show site owners who is crawling, what they take, what comes back, and which of your URLs they want most.

A bot's name is only worth something if you can trust it. With Web Bot Auth, operators, including OpenAI, Google, and AWS, cryptographically sign their agents' requests, so a site can tell a real agent from an impersonator without guessing from IP addresses or user-agent strings. We see more than 500 billion verified bot requests each week. 

Set your terms

In July, we replaced the single "block AI bots" switch with separate Search, Agent, and Training controls, available on every plan, including Free. The data showed why that distinction was needed. Fewer than 1% of sites on Cloudflare block search crawlers, while 17% block training. Site owners were never trying to hide. But with the rise of agentic traffic and the new ways agents use information, they suddenly had no transparency into, or choice over, how their content was being used. Being found no longer ensures they get paid, and they want to be found without being exploited.

That’s particularly difficult in the case of mixed-use crawlers. When one bot does both search and training, refusing one means refusing the other. On September 15, we shipped Disallow AI Training. It keeps you indexed for search while using crawler-specific mechanisms to instruct the operator not to use your data for training. Apple, Google, and Microsoft have committed to honor it. Cloudflare Radar also publicly tracks crawler behavior.

New domains now see recommended configurations based on how the site makes money rather than what piece of software is visiting. For ad-supported sites, you can easily disallow training and block agents on pages that carry ads, because an ad only pays when a person sees it. You can change any of these settings at any time.

Get paid

In August 2026, we described the Agentic Internet we're building as readable, discoverable, callable, and payable. The last word, payable, is the one that determines whether the open web can fund itself. The web needs a way to say ‘yes, if you pay’ instead of a binary ‘yes’ or ‘no’.

The licensing market shows both how much demand there is and where the gaps are. More than 50 publisher-AI deals have been signed since 2023. Nearly all of them are bespoke and bilateral, between large publishers and large AI companies. They prove content has value. But they don't reach most of the web, and they don't reach most buyers.

Not every asset should be sold the same way. High-value content and datasets need a trusted network, where buyers are identified and report how the work was used. Services like APIs and MCP tools don’t work like that: every request is the use.

So we're building for both.

Pay Per Use reaches the sites that direct licensing can't. Most publishers will never get a bespoke deal with each AI company, and no AI company can negotiate with millions of sites. Pay Per Use is the bridge. It doesn't charge for the crawl. It pays when content is actually used. Every buyer is a verified crawler, which is what makes this a trusted network, and each one defines what counts as use and what it will pay.

Publishers see the offer, choose whether to opt in, and are able to opt out whenever it stops working for them. The buyer reports each use, Cloudflare checks those reports, then bills the buyer and pays the publisher. The reporting matters as much as the payment. Publishers see what was used, when and what they earned, and, where the buyer reports it, information about which questions surfaced their work. Licensing deals rarely show any of that. It creates a feedback loop: publishers learn what people are actually asking for, and from that can decide what to cover, what to update, and what to make readily available to agents.

There won't be one definition of use. A search engine citing a source, a research agent quoting a passage, and a shopping agent completing a purchase create different kinds of value, and each will want its own business model. Buyers can participate via multiple business models using the same rails, with no new integration for publishers. Take a trade journal for marine engineers, with a few thousand subscribers and little prospect of an AI licensing deal. It gets paid by every participating AI company that draws on its work.

Monetization Gateway captures value that has never had a way to change hands. Accounts, API keys, and subscriptions work for customers you already know, not for an agent that wants one lookup from a service it has never used before. Our closed beta allows eligible U.S. Cloudflare customers to put a price on anything that passes through us, using the Rules language they already know. When a rule matches, we return an HTTP 402 Payment Required using the open x402 protocol, and the agent pays the seller directly.

That does more than recover lost revenue. Agents are customers in their own right: they pay for the data, APIs, and tools they use, whether the request is the whole purchase or one step in a larger task.

Monetization Gateway prices per request, per query, or per token, at fixed or capped prices. A sports statistics site built on ads can charge a fraction of a cent each time an agent asks "who leads the league in assists?" When we announced Monetization Gateway, thousands of sellers joined the waitlist, and their most common request was "charge agents, not humans." We're also our own first customer. Cloudflare's AI Gateway uses Monetization Gateway to let agents pay for inference, so we find the rough edges before our customers do.

For buyers, both products beat a block page: reliable access, and a way to reach millions of sites instead of one licensing deal or API key at a time. Every paid request leaves a receipt showing what was bought and that it was paid for.

Both Pay Per Use and Monetization Gateway are bets, built with customers on shared primitives: identity, metering, pricing, settlement, and analytics. They work together, so a publisher can disallow training, allow search, earn from AI answers, and charge agents per article from one dashboard. Pricing and discovery aren't solved yet, which is why both launch as betas, shaped by real customers and real transactions.

Make every request cheaper

Payment is the answer to falling revenue. Rising cost is a different problem, and much of it is simply waste. Most crawlers still download pages built for humans, again and again, to extract a few paragraphs of text. Too often, bots crawl sites that haven’t changed since the last attempt. That burns bandwidth for the site and compute for the crawler, and it happens before any answer is written. We’re working with our customers and the crawlers on tools that will help. Today, you can see the bandwidth consumption used per operator in our dashboard.

In July, we announced a joint research project with OpenAI, a first-of-its-kind pilot to explore how insights from Cloudflare’s global network can help AI search engines discover and index relevant content on the open web more efficiently and effectively. We’re planning to share our initial results in the next few weeks.

For our customers, we’re shipping tools and one-click experiences to make their sites optimized for this new kind of traffic. Markdown for Agents lets agents read a page without the additional styling meant for human eyes, and WebMCP lets a site expose actions directly instead of making agents guess which button to press.

Why build on Cloudflare

More than 20% of the web sits behind Cloudflare’s network, and so do nearly 80% of leading AI companies. We see both sides of this market. We build the rails for visibility, identity, controls, and settlement, and let the market work out what things are worth.

The old deal is gone, and the new one is still being written. Together we can shape what happens next.

In one version, a few companies control how agents find things, prove who they are and pay, and everyone else routes through them. In the other, those pieces are open standards anyone can implement, and a site of any size can set its terms and get paid. We prefer the latter.

That's why these rails run on open standards like x402 and Web Bot Auth, so anyone can build on them. Domain owners choose their own identity providers, their own payment processors, their own agent partners. Cloudflare is one option, not the whole stack.

For decades, the web was paid for by the people who visited it. Now the software visiting on their behalf can pay its share too.

Adaptive application security for the AI era: how Cloudflare connects code, traffic, and intelligence to stop attacks

Post Syndicated from Daniele Molteni original https://blog.cloudflare.com/ai-era-framework/

In July, AI agents testing new cybersecurity models compromised parts of OpenAI’s infrastructure and Hugging Face’s production environment.

We've all just witnessed one of the first AI-driven successful cyber attacks. When given a task, the agents ignored existing guardrails and autonomously discovered previously unknown vulnerabilities, recovered exposed credentials, moved between cloud environments and coordinated their work through communication channels they created themselves.

The speed of the final compromise was incredible. In under 13 hours, the agents went from executing code on a Hugging Face worker to gaining admin-level access across multiple clusters. But the incident had been brewing for much longer. Responders found clues of activity tracing back to May (agents created an unauthorized message board), to June (internal network scanning) and early July. The relationship between these events was understood only on July 20.

The lesson here is not that AI agents exploit vulnerabilities. That’s not news; human attackers already do that. The change is that agents can work persistently, test multiple paths simultaneously, share discoveries, and chain vulnerabilities, credentials, and permissions into sophisticated attacks.

The incident also shows why application security cannot depend on single tools. For example, network restrictions were bypassed by services connected to the Internet; valid credentials were used to perform unauthorized actions. Rebuilding Artifactory removed one attack path, but agents found another. The key insight is that individual alerts identified pieces of the activity without revealing the complete campaign. OpenAI reached a similar conclusion in its report: organizations need overlapping and independent controls across prevention, detection, and mitigation, continuous validation of security boundaries, and faster mechanisms to correlate and contain suspicious behavior.

We address this challenge by connecting application security across four activities that are too often separated: discovering which risks matter, governing what humans and agents may do, protecting applications at runtime, and turning every investigation into stronger protection. Cloudflare can deliver this framework because of its broad security portfolio and visibility across a vast share of Internet traffic.

Alongside the framework, we connect existing Cloudflare solutions with new capabilities across each stage. These include: using Large Language Models (LLMs) to conduct a penetration test of our Web Application Firewall (WAF), expanding threat intelligence to all customers, and a new feature to automate deploying positive security.

What has changed

The security landscape is shifting. These are the emerging trends we see:

  • The way we build software has fundamentally changed. AI-assisted development allows engineers to produce and deploy software faster outside traditional engineering workflows. That speed creates both more code and more opportunities for vulnerabilities to reach production.
  • Software composition risk is still a risk: applications depend on large chains of open-source libraries, packages, and operating-system components that are intrinsically trusted and are difficult for any team to inspect. What’s new is that AI is now importing libraries that we might not be aware of.
  • Techniques and tactics are changing. LLMs can chain vulnerabilities and use feedback in real time to mutate payloads, evade defenses, and make decisions autonomously. They can operate continuously and at machine speed. Patching faster remains important, but patching alone cannot close the gap. Attackers are always going to be faster than you can update your systems.
  • Agentic traffic. In the past, automation was a synonym for malicious activity. Today, a request generated by an agent or bot may be malicious automation, a search crawler, or an agent purchasing a product on behalf of a customer.
  • Compromised servers, residential proxies, IoT devices, and cloud resources allow attacks to move quickly across infrastructure and identities. A coordinated attack can leverage a number of devices, making it difficult to be identified as a unique campaign.

Application Security’s goal is also expanding. It now needs to address three connected problems: protecting conventional applications from AI-enabled attackers, governing legitimate and malicious agentic clients, and securing applications that contain models, agents, tools, and data.

A connected application security framework

Application security in the AI era must operate as a continuous system rather than a collection of controls that teams update after each new vulnerability. To protect applications in the era of AI, you need to work on multiple activities, which we have organized around four stages:

  1. Discover and prioritize risks
  2. Govern access and agent behavior
  3. Protect applications at runtime
  4. Investigate, respond and learn

None of these activities is new in isolation. What changes is connecting them so discoveries, runtime signals, and investigation outcomes continually improve the controls that follow.

More than 20% of the web sits behind Cloudflare’s network, which gives us visibility into attack infrastructure, payload mutations, emerging techniques, and coordinated campaigns at a scale that few organizations can match. Patterns that look isolated from the perspective of one application can become clear across our network. This combination of global threat intelligence, local application context, and inline enforcement powers every stage of the framework. That local context includes which code is deployed, which endpoints are exposed, what legitimate traffic looks like, which identities are acting, and which controls are already active. Because Cloudflare is inline, we can turn those insights into protections immediately.

Cloudflare is the adaptive security control plane for applications, APIs, and agents. Here is what we are launching today to advance every stage of the security journey.

Discover and prioritize risks

Security teams do not suffer from a shortage of findings. They struggle to determine which findings represent an immediate risk. A useful discovery system must connect vulnerabilities to the production reality, including whether a vulnerability buried in your stack is actually reachable in the first place. We see three main areas you should look into: software composition risk, proprietary code, and runtime penetration testing (pentesting).

Understand software composition risk

Applications inherit risk from open source libraries, packages, operating-system components, and the services on which they depend. This represents the Supply Chain of your application. A package vulnerability alone does not tell a team whether the affected component is deployed, reachable, or exposed to hostile traffic. Open-source software is the top priority when it comes to supply chain risk, and Cloudflare is part of Chainguard Athena, an industry coalition aiming at protecting open-source software from AI attacks.

Scan proprietary code

When it comes to code scanning, you have two options: getting a managed service or developing in-house expertise to run it yourself.

Cloudflare recently announced early access to Vulnerability Discovery and Remediation, a service that uses frontier models to identify application-specific vulnerabilities and deploy WAF mitigations to block targeted exploits while engineers fix the code. The important step is prioritization. Cloudflare connects source-code findings to production traffic and security signals. We can identify whether the affected route is active, and how much traffic it receives.

Pentest your application at runtime

Defenders can also use the same capabilities as attackers. Customers can build their own LLM-based pentesting harness to search for weaknesses, validate findings, and test whether their applications are vulnerable. Discovery becomes continuous rather than a periodic exercise. We have done this internally at Cloudflare since Anthropic’s Claude Mythos was released, and we shared our learnings.

A vulnerability buried deep inside your code is harder to exploit if it can’t be reached from the outside. Our Security Analyst team has already used LLM-based red teaming to test customer applications and our own runtime detections, turning the findings into improved detections for all customers. We are now developing Adaptive Security, a self-service capability that will periodically pentest selected URLs behind Cloudflare, using LLM-powered agents to identify vulnerabilities that are reachable and exploitable before attackers find them.

Govern access and agent behavior

Agentic traffic operates in the space between automation and human: tasks delegated by people, executed by software. This changes how access decisions must be made. Detecting automation is no longer enough. For every interaction, application owners need to answer two questions: Is this entity who it claims to be, and can this interaction be trusted?

These questions can be hard to answer. A recognized agent with a long history of legitimate activity may have high trust, but an unusual action can still create immediate risk. An unknown agent may simply be new; a lack of history does not necessarily mean malicious intent. Cloudflare’s approach is to keep trust and risk signals separate, thereby giving application owners more control than a single bot score or allow-or-block decision, and providing more powerful tools to quickly adapt to change.

Establish identity and trust

Trust accumulates over time, while risk is evaluated for each interaction.

Botbase provides a directory of known automated entities that have registered with Cloudflare. Registration gives legitimate bots and agents a way to declare who they are, while application owners retain control over whether and how those agents may access their sites. Cloudflare is also making registration more accessible to smaller and custom agents, building a verified identity layer across all agentic traffic, not just the major platforms.

Identity alone does not establish trust. Cloudflare can evaluate whether an entity has been seen before, whether its historical behavior was legitimate, and whether its current activity is consistent with that history. This makes it possible to distinguish a recognized agent behaving normally from the same agent suddenly changing established request patterns, location, identity, or transaction behavior.

Understand agentic behavior across the journey

Precursor adds client-side and session-level signals to distinguish human from automated behavior, such as typing cadence, mouse movement, navigation patterns, and sequences of actions. An agent that navigates a checkout flow in two seconds, skipping the browsing and comparison steps a human would take, reveals its nature through the session. These signals help identify whether behavior across a session is consistent with human interaction or automation, giving application owners a clearer picture of the traffic they're managing.

Manage access and adapt

Application owners can block traffic from AI crawlers and decide what activity is allowed on their asset (e.g. search, training, etc.). Adaptive Intelligence combines network, client-side, historical, and behavioral validation signals in a probabilistic model that can be updated as attackers change their techniques. Customer outcomes, including chargebacks and successful legitimate transactions, can feed back into the system to improve future decisions.

The result is a continuously updated assessment of every entity and interaction. Application owners can encourage known, useful automation while applying stronger controls when identity, history, and current behavior indicate greater risk.

Protect applications at runtime

Cloudflare’s reverse proxy protects applications at runtime by filtering traffic before it reaches the origin. In the AI era, a new layered approach is emerging to best filter traffic from malicious requests:

  1. Enforce positive security
  2. Detect attacks and identify LLM tactics and techniques
  3. Protect business logic
  4. Deploy real-time threat intelligence

Enforce a positive security model

You can dramatically reduce the attack surface by learning what legitimate traffic looks like, allowing conforming requests and blocking everything else. Today we are announcing Application Profiles, which automatically learns the structure of your web or API application and detects non-conforming requests. Application Profiles automates the learning process and adds a layer of interpretation. Based on the learned profile, we can understand the business logic of different endpoints and request parameters and help you prioritize what endpoints require more scrutiny and attention.

Detect attacks and identify LLM tactics and techniques

Traditional WAFs are designed to run highly crafted rules to detect Common Vulnerabilities and Exposures (CVEs) and malicious payloads. Before AI, the time to disclose new vulnerabilities was measured in months and days. Not anymore: now we see vulnerabilities being exploited before they are disclosed, so the time to patch is nearing zero.

Different tools can be deployed to detect known exploits. These tools include:

  • Managed Rules hardened with frontier models. We’ve partnered with major model providers to use frontier models for adversarial validation. We used the latest models to pentest the WAF to uncover bypasses and vulnerabilities. All customers benefit automatically from ongoing improvements.
  • Machine Learning detection. While signatures are great for high-precision attack detections, machine learning can stop attacks before they are discovered and disclosed. Attack Score detects attack mutations and evading techniques that are often used by LLMs. Attack Score is available to all Cloudflare Customers
  • AI Security for Applications. Chatbots and Internet-facing LLMs are subject to a new class of attacks, such as prompt injection and sensitive data exposure. You can protect generative AI traffic by deploying guardrails and security detections designed to stop these attacks.

Protect business logic

Attackers can still craft legitimate requests and abuse business logic to gain advantage on the application. For example, an attacker uses a valid password-reset flow repeatedly to take over accounts. Fraud detection tools, including account takeover and leaked credential detections, help prevent abuse in which the request appears legitimate, but the intent is malicious.

Real-time threat intelligence detection

Back in June, we launched always-on detection based on our threat intelligence feeds. Cloudforce One customers can deploy protections to block requests originating from compromised infrastructure. We are now also expanding access to Cloudforce One’s Threat Events Platform, our core threat intelligence offering, to all Cloudflare accounts for free.

Investigate, respond and learn

The OpenAI Hugging Face incident did not begin with the final 13-hour compromise. The activity stretched from May to July, with signals including an unauthorized message board, internal network scanning, and movement across environments. Viewed separately, each event revealed only part of the activity. Together, they showed the behavior of a developing breach. Security operations must therefore identify sequences of behavior that lead to compromise, not simply evaluate alerts in isolation.

This is difficult for security teams that already protect large attack surfaces with limited resources. Alerts arrive from different tools and datasets, leaving analysts to determine which events are connected, collect the evidence, and identify whether the activity is escalating.

Cloudflare is building a platform to automate security operations. Deterministic workflows establish the customer and investigation context using trigger history, traffic baselines, enforcement outcomes, and network observations. A detection agent searches authorized datasets for anomalies and correlations. When it finds suspicious activity, specialist agents review the evidence alongside customer history and threat intelligence, helping analysts connect isolated events to broader campaigns. The system can then recommend mitigations, such as rate limiting, WAF, or DDoS protection changes, for human approval.

We are developing these capabilities with Cloudflare’s Managed Defense team, whose analysts are helping us test how evidence is collected, correlated, and turned into recommendations. We plan to make them available more broadly over time and will share more as this work progresses.

Cloudflare’s combination of reverse proxy and forward proxy services makes this correlation especially powerful. Application Security signals can reveal attempts to exploit a public-facing application, while Cloudflare One can surface subsequent activity across corporate traffic. Connecting these datasets can link an external attack with unusual access, internal scanning, or potential lateral movement, turning separate alerts into a timeline of compromise and helping analysts intervene before the breach progresses.

Looking ahead

AI is changing how software is built, how attacks unfold, and who interacts with applications. Security teams can no longer manage discovery, access, runtime protection, and response as separate activities.

Cloudflare is bringing these capabilities together in a closed-loop system powered by global intelligence, local application context, and inline enforcement. A vulnerability finding can strengthen runtime protection, runtime activity can guide an investigation, and each analyst decision can improve future detections and controls.

No organization can anticipate every new technique. The goal is to build a security system that learns from each attempt, responds faster, and becomes more effective over time. The capabilities announced today are the next step toward that adaptive model of application security.

Building a certificate authority for the whole Internet

Post Syndicated from Steve Goldsmith original https://blog.cloudflare.com/cloudflare-certificate-authority/

Twelve years ago, during Birthday Week 2014, we turned on Universal SSL and nearly doubled the number of encrypted sites on the web overnight, giving free TLS to every site behind Cloudflare, including the ones that never paid us a cent. Encryption stopped being an expensive, time-intensive undertaking and instead became the default.

For Birthday Week this year, we are taking the next step on that path. For more than a decade we have been one of the largest consumers of publicly trusted certificates on the Internet, and have never issued a single one ourselves. That is changing. Cloudflare is announcing our intent to become a public certificate authority (CA).

Today we are announcing the first concrete milestones in that effort: We have applied for inclusion in the Chrome, Apple, Microsoft, and Mozilla root programs, and we have signed a definitive agreement to acquire an established, broadly trusted root from GlobalSign, so that we can offer certificates with the widest possible device reach the day we begin issuing. We’re also announcing our plans to be one of the first CAs to serve post-quantum certificates, targeting Chrome’s recently announced Quantum-resistant Root Program.

We are not issuing certificates yet, and it will be a little while before we do. What we are doing is committing to the work in public, sharing the milestones as they land, and telling you exactly what we are building while working with the root programs and other members of the WebPKI community to achieve this.

Two paths to trust

A brand-new root is not widely useful for years. Even after a root program accepts it, that root has to propagate out into the world's operating systems, browsers, and devices, and it never reaches the large set of devices that have stopped receiving updates, or never received them in the first place. That long tail of older clients is where a great deal of the world’s Internet traffic originates, and where a correspondingly large set of avoidable breakage lives. We believe that all clients deserve the highest level of security possible, regardless of their manufacturer, operating system, or time since last update.

Acquiring an existing root with a high degree of trust store coverage across a diverse set of clients solves that on day one. The existing GlobalSign root has been trusted across browsers, operating systems, and devices since 2012, and it reaches older clients that a fresh root never will. The new root that we will be submitting for inclusion in root key programs is built for where the ecosystem is heading, including the programs that are starting to cap how old a trusted root may be. The established root gives us reach across the devices of the past. The new roots give us standing under the policies of the future. We want both to ensure certificates issued by our CA provide the widest set of customer compatibility possible.

A new source of free certificates

The free-of-charge, automated certificate model now carries most of the encrypted web, and much of it runs through one remarkable operator. Let's Encrypt issues on the order of ten million certificates a day, serves more than 500 million sites, and passed four billion active certificates in 2025. It is one of the best things to happen to the Internet in twenty years, and we say that as one of its largest users.

That success comes with some systemic risk: if the dominant free certificate authority had a bad week, much of the web would have no comparable free, automated alternative ready to take the load. At the certificate pack level, we have spent years building exactly this kind of redundancy for our own customers. Every Cloudflare Universal SSL certificate already ships with a backup certificate, wrapped with a separate key and issued from a different authority, ready to deploy automatically if the primary is ever revoked or compromised. A public CA is that same idea, but at the scale of the whole Internet.

To make it easy to adopt, we will be Automated Certificate Management Environment (ACME)-first, an open standard protocol that is widely accepted. Automated issuance and renewal through ACME will be the way you get a certificate from us, which means anyone already pointed at any existing free CA can move to us by changing a directory URL, with no new tooling and nothing to re-architect.

Certificate growth projections are huge

Cloudflare sits in front of more than 20 percent of global Internet request traffic and terminates TLS for millions of domains, relying on millions of certificates per year to do so. We provision those certificates through multiple CAs, with primary and backup paths so customer services stay up through CA outages and revocation events.

That has taught us not just how the WebPKI ecosystem works, but also that it occasionally fails, from the consuming side, the hard way. We have dealt with rate limits, validation edge cases, revocation latency, chain building, and root distribution lag. We have lived through the CA churn of recent years and felt it through our customers. We know what reliable issuance has to look like from the outside, because our customers' uptime has depended on us being resilient and responsive when an issuer has a bad day.

And as certificate maximum validity period decreases over the next few years, agentic activity increases, and PQ certs go mainstream, we expect the raw number of certificates we rely on annually on to continue to grow, quickly — and we are not alone. We want to not just solve this problem for ourselves, but be part of providing this utility to the Internet, and ensure that the certificate supply chain for our customers has even more providers.

Designing for resilience: transparency and fail small

In taking on this new responsibility of being our own CA, we're committed to making the most reliable and resilient CA possible. We intend to build a certificate authority whose reliability depends not just on avoiding mistakes, but as with the rest of Cloudflare’s products, to “fail small” and limit the impact of any one issue.

That means instituting processes to design and test recovery before any incident occurs. As an example, we will make renewal automation a condition of issuance. We will only issue to clients that support ACME Renewal Information (ARI), standardized in RFC 9773. Subscribers must maintain automation that polls our renewal endpoint, acts on the renewal windows we publish, and identifies the certificate it is replacing.

We're also learning from what we've observed over the past 16 years. We have seen certificate authorities caught between timely revocation and keeping subscribers’ sites online because too many subscribers could not replace their certificates quickly enough. When certificates need to be retired, whether for a compliance issue or a security incident, we can bring forward renewal windows for the affected certificates, spread replacements across the available time, and track replacement issuance.

This is just one of the many ways we intend to build. We will be transparent with our issuance stack and operations, publish reproducible builds of the software that signs certificates, attest the hardware security modules that hold our keys, and run a public dashboard for issuance health and incidents. Audits are point-in-time and tell you a CA passed, not how it runs on an ordinary Tuesday. We want root programs, researchers, and ordinary site owners to watch how a modern CA actually operates between audits.

A certificate authority for the post-quantum Internet

We also intend to lead on where certificates are going, not just where they are. We plan to be one of the first CAs to issue production Merkle Tree Certificates (MTCs), with the first certificates issued in the first quarter of 2027.

MTCs are a new and far more compact way to deliver publicly trusted certificates, designed for a post-quantum world where traditional certificate chains grow large enough to strain TLS handshakes. We have been championing the standards-based proposal for MTCs at the IETF, and earlier this year, Chrome named MTCs as the preferred path for post-quantum authentication. Issuing them in production allows us to protect Cloudflare customers as well as the wider Internet against the post-quantum threat, with real volume behind a transition the whole web has to make. We’ve shared much more about MTCs and what this new Web Public Key Infrastructure (PKI) will look like in a blog post on the topic.

We do not expect that transition to be sudden. Much of the Internet will continue to rely on classic certificates and existing WebPKI for many more years. But across that window we expect MTCs to take a steadily growing share of issuance, and that is why we are building one service that does both. By carrying classic certificates and Merkle Tree Certificates under one CA, with one lifecycle and one set of guarantees, customers can adopt at the pace that suits them and help the web make the crossing without a hard cutover. Customers should not have to pick a side of a multi-decade migration, run two systems, or rebuild when the balance shifts.

As always, Cloudflare will be Customer Zero

In addition to providing certificate packs via Universal SSL for our customers, Cloudflare consumes certificates from many different CAs to run our systems and internal operations. Just like our other products, we will be Customer Zero for the new CA and its certificates (both WebPKI and MTC), ensuring that all aspects of the new systems and processes meet our high internal standards, and that our CA’s infrastructure is exercised at Cloudflare scale.

What happens next

We are working through the application and approval process with each of the core web root key programs. These processes happen in the open, and we’ll share more updates as they proceed, through to the first Merkle Tree Certificates in early 2027. If you want to follow this work or be one of the first to use a Cloudflare CA certificate in the future, you can register for updates.

As we build out this new capability, we will continue to work closely with the network of partner public CAs we have relied on for many years — 16 in fact! — as we all work together to ensure a trusted and open Internet.

When we launched Universal SSL, the argument was simple: every byte that flows encrypted across the Internet makes it harder to intercept, throttle, or censor, and the open web is something we all build together. A public, redundant, transparent certificate authority is that same argument carried one layer down, to the trust that makes the encrypted web possible in the first place. We have been working toward this for a long time, and we are glad to finally be on the road.

Happy Birthday Week!

Using AI to chart a course for our post-quantum migration

Post Syndicated from Sharon Goldberg original https://blog.cloudflare.com/ai-driven-cryptography-discovery/

As laboratories around the world race to build out a cryptographically relevant quantum computer, we at Cloudflare are racing towards a 2029 target deadline for full post-quantum readiness. While we’ve already transitioned many of our products to post-quantum encryption, we still have work to do to support post-quantum authentication and achieve full post-quantum readiness across our platform.

We’re taking a maximalist stance (“PQ everything!”), because as an infrastructure provider to the world, we want to give our customers the peace of mind that using Cloudflare ensures that their traffic is future-proofed against quantum adversaries.

But how does one accomplish such a massive migration at an organization of our size and scale? After all, cryptography is the base layer for almost all of the world’s digital systems, including the software services and the networking protocols that power our platform.

To drive our PQ migration, we have three key goals.

First, we want to help our product and engineering teams understand how cryptography is being used and how they should be upgrading it. This should cover both the upgrades to post-quantum encryption and to post-quantum authentication. Many of our products have already been upgraded to post-quantum encryption over TLS 1.3, but we still want to cover the long tail of TLS connections, as well as upgrade any other uses of public-key encryption. Meanwhile, it’s still early days for our deployment of post-quantum authentication.

Next, we want to provide progress metrics for the migration. These might include per-repository and per-product counts of the use of classical and post-quantum cryptography.

Finally, we want to surface prerequisites early. If our products or platform rely on protocols that don’t yet have a PQ migration plan (because PQ variants of the system have not yet been considered, because PQ standards do not exist or lack consensus, or because software libraries or other key ecosystem components do not yet have PQ support), then we need to know now. That way we can work with the relevant stakeholders, standards bodies and ecosystems to help drive their PQ migration plans, so that we can meet our own 2029 PQ migration timeline.

This post is the story of how we’re going about this. We explain how we turned to AI to help us solve some of our problems and how we’re developing an internal tool called CryptoLabe to help us. CryptoLabe is named after the mariner’s astrolabe, a navigation instrument refined by Portuguese navigators. Just as an astrolabe helped sailors determine where they were and chart a course, CryptoLabe helps us discover cryptography in our code, understand how it is used, and chart a path to post-quantum migration.

CryptoLabe is highly specialized to our internal systems (our repositories, our ticketing systems, and internal documentation processes) and still evolving as we continue its development, so we aren’t making it available to customers. Nevertheless, we are sharing our learnings so that other organizations can build upon our efforts as they work through their own PQ migration journey.

The scale of the problem

The software that powers most Cloudflare products lives inside our single centralized source control management platform. This means we can find most uses of cryptography across our platform by just looking through our codebase.

While the centralization of our codebase is a marked advantage for us, we still need to contend with three challenges that come with the scale of this problem. First, our code is spread across many repositories. Second, cryptography rarely announces itself plainly in the code. Instead, it hides in

  • shared libraries that a repository imports but may or may not actually call
  • upstream and protocol defaults, like a TLS 1.3 listener that is configured to negotiate a classical key exchange such as X25519 rather than post-quantum X25519MLKEM768
  • configuration files that select algorithms far away from the code that uses them, like a TLS responder whose key exchange protocols are pinned in a YAML file stored in a different repository
  • code paths that are dead, test-only, or on a path to being deprecated

Third, cryptography discovery is about more than just pattern matching. Grepping for certain algorithm names (e.g. “RSA” or “X25519”) overcounts, because it finds cryptography in unused code. Grepping also undercounts, because it misses defaults and indirect uses in dependencies and configuration. Most importantly, it can't tell you how the cryptography is used. A classical ECDSA signature could be part of a JWT, IPsec, TLS, or SSH, and each has a completely different migration path. Many uses also depend on the other side of the connection: a TLS server may support both post-quantum key exchange and classical key exchange; the one it chooses to use would depend on the client.

Turning to AI

It turns out that AI is pretty good at doing more than just grepping. A model can search a codebase, follow evidence across files, and return structured analysis. It can also enrich findings by pulling information from other sources, like our internal documentation and ticketing systems. In fact, AI can even explain how cryptography is being used and how it should be updated. We’ve been putting that idea to the test as we develop CryptoLabe.

As we said before, our first two goals are to (1) discover and understand the use of cryptography in our codebase, and also (2) to get metrics on the state of our PQ migration. Towards these goals, our current implementation of CryptoLabe performs scans in two stages, as shown in the figure below.

The first “discovery” stage starts by mapping the repository. It then searches for cryptography through source, configuration, manifests, lockfiles, scripts, tests, and documentation. Among other things, the scan looks for the use of cryptography like key agreement, signatures, asymmetric encryption, PKI, tokens, credentials, hardware security module integrations, and more. This discovery stage produces a set of "raw observations."

Each raw observation feeds a run of the second stage. This “analysis” stage first re-checks the observation against the source code. It then investigates how the cryptographic operation is used at runtime, what role the repository plays, and which internal or external parties it depends on. When necessary, it can inspect related code in other repositories to complete the analysis. Finally, it takes a pass over its own conclusions, searching for missing or conflicting evidence such as configuration overrides, test-only code, or incorrect assumptions about runtime behavior.

Next, the model assigns a classification to the finding. If there is not enough evidence to assign a classification, the model assigns More evidence needed, External dependency, or Unknown rather than guessing.

This is the current list of classifications used by CryptoLabe, containing catch-all classifiers which will likely be refined as we proceed through our migration. (As an example, we could refine our classifiers by splitting the “encryption” classifier into key agreement and HPKE; you get the idea.)

Classification

Examples

Classical encryption

This is a catch-all category that finds cases of elliptic-curve Diffie-Hellman key exchange (ECDHE) (e.g., X25519, P-256, P-384), RSA key agreement or other uses of public-key encryption (e.g., HPKE). These are broken by a quantum computer running Shor's algorithm, which puts them at risk of harvest-now-decrypt-later attacks.

Classical signature

This is a catch-all category that finds use of an RSA signature or elliptic-curve (ECDSA) signature in anything, for example a certificate, a TLS handshake, another protocol handshake. These signatures are broken by Shor's algorithm.

Classical token

We found a lot of RS256 or ES256 JWT tokens, so we created a special classification for them. These are JWTs that use classical RSA and ECDSA signatures; RFC 9964 defines a post-quantum replacement using ML-DSA.

PQ-ready hybrid key exchange

Finds hybrid post-quantum key exchange in TLS 1.3, i.e. X25519MLKEM768. This is the most prevalent use of PQ encryption in our codebase.

PQ-ready

Finds other uses of post-quantum cryptography that are not X25519MLKEM768 in TLS 1.3, like ML-DSA.

Finally, it generates a report that serves two audiences: (1) product managers who need to understand what the migration means for their product, and (2) engineers that need enough detail to execute the migration.  

Here’s a (cropped) view of one of our reports:

While we’ve been iteratively reviewing findings against the source code and with relevant engineers, we do not yet have a ground-truth dataset for reproducibly comparing different versions of the prompts we’ve tried for CryptoLabe.

Built on Cloudflare’s Developer Platform

We built CryptoLabe on Cloudflare's Developer Platform. Here’s the architecture:

CryptoLabe runs across two Cloudflare Workers. There’s a scanner Worker that runs the scans. And there’s an inventory Worker that serves the dashboard, exposes the API, and stores everything in a D1 database. The two communicate through Service Bindings. A scan starts when someone requests it from the dashboard, and the inventory Worker passes the request to the scanner.

Orchestrating a scan

We need a way to keep a scan alive and on track from start to finish, without building our own job orchestration system. We did this with Agents SDK. Each repository gets its own persistent coordinator built on a Durable Object (DO). A bounded queue in front of the coordinators limits how many scans run at once. When a scan's turn comes, the coordinator tracks its progress and handles cancellation, retries, and recovery.

The coordinator doesn't do the analysis itself. It hands the work to Cloudflare Workflows, so that they can persist progress and automatically retry failed steps. The coordinator moves each repository through four stages:

  1. discovery Workflow (the first scanning stage that produces raw observations)
  2. deep analysis Workflow (the second stage, run on each raw observation)
  3. merge Workflow (that builds a list of findings for a given repository, including combining repeated or similar finds)
  4. publish workflow (that hands results back to the inventory Worker)

The first two workflows need the model to have access to the repository's code. We want this access to be isolated, so we don’t risk damaging the codebase. That’s why CryptoLabe downloads the repository once, at an exact commit, at the start of each scan, and then stores that snapshot in R2. Each Workflow then restores the snapshot into a fresh, short-lived Cloudflare Sandbox, an isolated container. The model then works with the Sandbox through a small set of read-only tools on an immutable snapshot of the code, even if the codebase changes while the scan is still running.

Calling the model at scale

If we want to scan through all of our (many!) repositories, we have to worry about both cost and capacity.

For cost, the model loop sends its requests through AI Gateway to cost-effective open-weight models hosted on Workers AI. Putting the model behind AI Gateway also makes it easy to switch models as better or cheaper ones become available.  

Capacity became a problem once we scanned many repositories at once. Bursts of model requests began triggering HTTP 429 (rate limit) responses from AI Gateway, and scans retrying independently only made the bursts worse. We solved this with a single, global Durable Object that paces every model request across all scans, including retries. When any scan hits a rate limit, the cooldown is shared and all scans back off together, so concurrent scans share the available capacity instead of competing for it.

Prerequisites and hard cases

Let’s now get into our third goal: surfacing prerequisites and hard cases early.

A lot of ink has been spilled about ecosystem readiness for the PQ migration, and we are now going to spill some more. As everyone knows, a PQ migration cannot happen in a vacuum. For migration to succeed, post-quantum cryptography must be supported in relevant software libraries (e.g. BoringSSL) and across parties that participate in the ecosystem (e.g. clients, browsers, origins, cloud proxies, certificate authorities, etc.). Standards are also an important indicator of ecosystem support, although a standard that is still in “draft” state does not necessarily mean deployment cannot proceed. As an example, we deployed X25519MLKEM768 in TLS 1.3 back in 2022 when it was still a “draft” at the Internet Engineering Task Force (IETF) while it was only finalized as RFC 10024 in 2026.

Either way, our point is that in order to upgrade a system to PQ cryptography, we need to understand its dependencies and level of ecosystem support. 

That’s why CryptoLabe uses the concept of “prerequisites” to highlight findings that cannot be immediately remediated by an individual product team working alone.

A prerequisite can be something as straightforward as “we are currently blocked on migrating to post-quantum JWTs.” We say this is straightforward because there is already a standard (RFC 9964) for post-quantum JWTs. Nevertheless, if our software libraries don’t yet support validating post-quantum JWTs, or if we’re using a token issuer that does not yet issue post-quantum JWTs, we can’t go company-wide and ask each of our product teams to start PQ-ing their JWTs. This migration is blocked until we solve its core prerequisites. CryptoLabe lets us group together findings that (likely) have the same prerequisite, which also helps us decide how to prioritize resolving these prerequisites.

For example, the snapshot below shows the six findings from CryptoLabe that have post-quantum SAML as a prerequisite. (SAML is a protocol for single sign-on (SSO).)

On the other hand, there may be uses of cryptography that lack even a basic level of ecosystem support. We’ve been calling these “hard cases.” To find them, we wrote a separate prompt that ignores “vanilla” uses of cryptography (e.g. ordinary TLS between internal systems) and instead looks for custom cryptographic protocols, keys, or signatures used in size-constrained fields, cryptography built into hardware, specialized cryptographic constructions (like blind signatures), protocols without a PQ standard, and dependencies on external parties that do not yet support PQ cryptography.

This prompt is shorter and simpler than those used for CryptoLabe, since its only job is to find hard cases.  In our qualitative review, we found that it got better results when it ran in one fell swoop against all our repositories, while also taking in context from our internal ticketing and documentation system.  

Here’s an example of a “hard case” we found: a certificate carried in an HTTP header. Post-quantum certificates and signatures are larger than their classical counterparts, so if the header (or an intermediary, or the application processing the header) assumes a certificate has a certain size, changing the signature algorithm may break the system. Our next step is to determine whether this code will remain in use in the long term. If it will, we need to measure the relevant size limits and decide how to accommodate the larger certificate.

An important lesson here is that no single scan finds everything. Our repository-by-repository scans were effective at discovering common uses of cryptography. Meanwhile, this targeted scan worked better for “hard cases” because it ignored well-understood cryptography and had more context about each product and its dependencies.

The bottom line is that different approaches find different things, and every finding still needs to be checked by the engineers who understand how the system actually works.

Sharing our prompts

We’ve been messing around with the best way to write prompts for CryptoLabe for the last several months.  We don’t yet have a ground-truth dataset for comparing one prompt’s performance against another, and we are not convinced we have 100% coverage of all uses of cryptography in our codebase. Instead, we have iterated by running scans, reviewing findings with the engineers that maintain the repositories, investigating misses that came up during these reviews and revising the prompts.   Nevertheless, we decided to publish selected prompts, so other teams can learn from and adapt our approach. These prompts are starting points, not a standalone version of CryptoLabe, and the quality of their results will depend on the model, tools, context, and engineering review available.

Thinking through your own PQ migration

At Cloudflare, we’re taking a maximalist approach to our PQ migration because of our goal of acting as a provider of post-quantum cryptography for customers and the Internet at large. But most organizations do not need to start by finding every use of cryptography in every repository in every one of their products. In fact, most organizations should not be doing this, because at this time it's a waste of precious resources.

Before scanning a single repository, you can protect traffic in bulk wherever possible. If your websites run through Cloudflare, we protect your data in transit with post-quantum encryption already today; check this out with our new PQ visibility features. Our SASE platform, Cloudflare One, provides post-quantum encryption for private network traffic. Post-quantum encryption is provided at no additional cost and without requiring you to upgrade every origin server or private application on your enterprise network. This gives you a compensating control while you work through discovering and understanding the use of cryptography inside your own systems.

An exhaustive cryptographic inventory is not a prerequisite for action. Instead, organizations should first identify the systems whose compromise would matter most, discover their use of cryptography, and then PQ that cryptography in priority order. Here is one way to begin:

  1. Choose a repository for one important system. Start with something that handles sensitive or long-lived data, authenticates users or software, or is exposed to the public Internet.
  2. Run cryptography discovery against that repository. We hope our description of CryptoLabe will be helpful to this effort!
  3. Validate the results. Ask the team who owns the system to validate the results of cryptography discovery and confirm that the cryptography finding is needed long term and needs to be upgraded to PQ. It’s important to remember that it might not need to be immediately upgraded to PQ if there is another compensating control in place.
  4. Prioritize action. Figure out what upgrades you can make now and what upgrades are blocked. Record shared prerequisites that need help from a library, vendor, standards group, or another part of your organization. Prioritize your findings and make a plan for addressing the highest-impact systems and prerequisites first.

That gives you the beginning of a PQ transition plan, without requiring a complete map of every cryptographic operation in your organization. CryptoLabe is still ever-evolving, but its scans and results have been illuminating to us as we plan our migration. We hope these shared learnings will be useful as you continue to work through your own PQ migration.

Acknowledgements: Many people across Cloudflare provided feedback on and contributed to CryptoLabe, including Davide Marquês, Peter Wu, Phil Schmieder, JP Aumasson, Andrew Galloni, Christopher Patton, Luke Valenta, Mari Galicer, Vânia Gonçalves, and the Client, Tunnel and Gateway teams who reviewed reports produced by the tool.

Enforce positive security with Cloudflare Application Profiles

Post Syndicated from Daniele Molteni original https://blog.cloudflare.com/application-profiles/

Today, we are launching Application Profiles, a seamless way to enforce a positive security policy. By analyzing the structure and format of HTTP requests and identifying deviations, Cloudflare can help you significantly reduce the attack surface area.

Every customer we speak to wants to know how we can protect them from attacks that use frontier AI models. This has become the number one priority for anyone working in security. Large language models (LLMs) allow even non-technical people to launch attacks with a single prompt. LLMs can generate malicious payloads, test known techniques, and probe applications autonomously by mutating their tactics based on the feedback from the application or the Web Application Firewall (WAF). 

Our tools have changed to stay a step ahead of the attackers. Managed WAF rules and machine learning-based detections remain essential for detecting techniques such as SQL injection, cross-site scripting, remote code execution, and new CVEs, including many variations of those attacks. The answer can’t simply be “patch faster”: this is not sustainable, and it doesn’t work if you haven’t completely mapped your vulnerabilities.

What if you could learn what good requests look like by analyzing your traffic structure? Instead of looking only for requests that resemble known attacks, we could allow only requests that conform with what we expect. By doing this, we’d dramatically reduce the attack surface area. For example, if the search field in your query doesn’t expect special characters, we can only accept alphanumeric strings. This would already prevent a vast library of known attacks.

But we don’t stop here. Once we have learned the structure and format of your HTTP requests, we can infer the goal of each operation and then understand what the application ultimately does. With this information, we can identify and prioritize the most critical and vulnerable operations and fields you should take care of first.

Cloudflare already supports positive security for APIs through Schema Learning and Schema Validation. We are now extending this protection to web applications through Application Schema Profiles. You onboard an application, we learn its profile, and then we start to deploy an always-on detection that identifies non-conformity. All automated and enriched by powerful analytics.

We are opening a closed beta to invited Enterprise customers without API Security; customers with API Security already have access.

Validating requests based on learned profiles

Schema Profiles periodically learn the expected request structure from observed traffic. After a profile is available, an always-on validation layer is automatically deployed on live traffic. For every request, the detection evaluates whether it conforms or not with the profile, and it adds the result as metadata, augmenting the information already associated with the request. The signal does not take action by itself: customers can analyze past traffic in Security Analytics and decide where enforcement is appropriate and create Security Rules to block non-conforming requests. Requests to operations without a profile are not classified by this feature.

Unlike Managed Rules, failing validation does not require a request to match a known attack signature. A value outside an expected range, an unknown enum value, an invalid universally unique identifier (UUID), or unexpected characters — all can be identified because they differ from the learned profile.

For example, consider the following operation: 

www.example.com/shop/2dbda2e7-cfc9-448d-9465-799d2e6ff363/inventory?product_id=938062541

Below we describe the learning process, which evaluates only the structure and format of the request. When enough traffic has been observed, we learn that the path expects a UUID variable and that product_id is an integer and what its boundaries are. When product_id contains a string, it will be flagged as a violation. Similarly, Cloudflare can identify malformed UUID values and, when the customer enables enforcement, prevent non-UUID input from reaching the corresponding handler. These simple filters reduce the range of inputs an attacker can send, preventing the vast majority of typical attack vectors, such as SQL injection, cross-site scripting, remote code execution and more. 

Non-conforming does not always mean malicious. An application release, a new client, or an unusual but valid request may also introduce a difference. We recommend starting in observation mode, so customers can review a profile's effect before enforcement.

Learn the expected structure of requests

To determine the anticipated request structure for a web or API application, Schema Profiles routinely analyze observed traffic. Each profile may include the following, depending on the application traffic:

  • Path variables 
  • Query parameters
  • Headers and cookies
  • Body structure (JSON body or form-encoded)

For each field, the system learns its data type (integer, string, boolean, arrays, UUID or enum) and constraints such as numeric ranges, short enumerations, string lengths, and character classes.

Learning applies to operations that customers select for profiling. In Web Assets, an operation is Cloudflare's term for an operation identified by its HTTP method, hostname pattern, and path pattern. Web Assets continuously discovers operations and lists them under Web Assets > Operations. Customers can also add operations manually. Profiling doesn’t automatically start for discovered operations, while manually created operations do trigger profiling when created. For discovered operations, the customer must intentionally select Learn profile from the operation's overflow menu. 

After profiling is enabled, Cloudflare collects qualifying traffic and runs learning automatically once a week for each zone, using the most recent successful traffic. An operation needs at least 1,000 requests that received a 2xx response in the previous seven days to learn fields, and at least 10,000 to learn data boundaries. Successful requests can include bots and scanners, so customers should review a learned profile before enforcing it. Our roadmap includes allowing customers to trigger learning on demand and excluding automated traffic.

Once learned, profiles can be reviewed by selecting View details of the operation and finding the learned schema in the Security overview panel. If a learned schema is not shown, Cloudflare is still collecting data for the profile. Customers can also export the profile as an OpenAPI v3 schema file.

Learned profiles update each week as application traffic changes. New fields are added and fields that are no longer observed are removed, so validation tracks how the application changes. Customers can pin and save the learned schema by downloading the learned schema and uploading it to Schema Validation.

Review before you block

Security Analytics now includes a new Profile Analysis tab. Customers can select a validation profile and see traffic trends, including how many requests did not conform to the learned profile during the previous seven days. 

Customers can review the conforming and non-conforming traffic. They can drill into violations and review sampled logs to see where the violation occurred, which field was affected, and why it failed validation. Violations are classified into ten reasons, including type mismatches, values outside a learned range, and invalid formats.

Once a team understands the effect, it can use Security Rules to act on the signal. A rule can cover an entire application or be limited to selected paths, operations, or fields. Teams control where to monitor and where to block.

Positive security for web and API traffic

Traditional WAF learning modes can build detailed positive-security policies, but they often require operators to review suggestions, stage changes, and maintain policy entities. Cloudflare Schema Profiles expose validation as a request field cf.schema_validation.learned.violated, allowing customers to combine it with request properties, Bot Score, Attack Score, and other signals in a single Security Rule. By creating simple rules, teams can combine detections and define precisely when Cloudflare should take action.

Two other classes of fields are available to create more targeted rules. First, there are fields that collect where the violation occurred. For example, based on our initial example, if the value of product_id query parameter does not conform with the profile, the following field will be populated cf.schema_validation.uploaded.query.violated_parameters = ["product_id"]. This allows customers to create rules that enforce positive security only on specific fields or exclude them from the enforcement.

The second class of field collects new parameters that are not present in the profile. This is useful when you want to handle requests with new parameters (e.g. when deploying a new version of your application), or restrict your posture even further by blocking any parameters that were not detected or defined in the past.

Use case

Field

Location values

Example

Identify where in the request the violation occurred

Array up to 20 items

cf.schema_validation.learned.[location].violated_parameters

query,path,headers,cookies,body

cf.schema_validation.learned.query.violated_parameters = ["product_id"]

Identify whether an undeclared parameter is seen in the request 

Array up to 20 items

cf.schema_validation.learned.[location].undeclared_parameters

query

cf.schema_validation.learned.query.undeclared_parameters = ["adminMode", "utm"]

Coming up: critical field analysis, how we help you roll out positive security

Even with a flexible enforcement design, customers tell us that deploying a positive security policy is operationally complex. A large application can have thousands of operations with tens of thousands of fields. But not all operations and fields carry the same risk. Contextualization and prioritization helps security teams roll out positive security in a controlled and confident manner.

LLMs can help contextualize learned profiles to provide additional insight. For web applications, paths and field names are usually self-explanatory, thus semantic. For example, we piloted running a model hosted on Workers AI across four random applications’ learned profiles. The model successfully identified the link between clientId and account_number across two applications of a system, as well as the common dependency of using One-Time Password (OTP) for enhanced authentication. Highlighting this context enables security teams to prioritize actions such as configuring Rate Limiting Rules to defend against account-focused brute force attacks.

These LLM-powered insights will be accessible directly within the dashboard alongside each operation in Web Assets. Before executing a one-click deployment, teams can evaluate rule recommendations designed to secure these key fields, backed by mitigation simulation using past traffic to gain confidence.

Beyond contextualizing operations with semantic insights and risk indicators, we are developing additional metrics to order operations using historical request trends and signals. This enables security teams to focus mitigation efforts on the highest-priority operations first, including:

  • Data loss: upward trend of unusual increased data transfer
  • Reconnaissance activity: high count of unknown parameters
  • Business criticality: total volume of traffic correlated with the unique session IDs served

What’s available today

Customers with API Security already have access, given that this is an extension of Schema Learning and Schema Validation. We are opening a closed beta to customers without API Security who can test Schema Profiles on production web application traffic, meet with the product team, and provide detailed feedback on profile accuracy, analytics, and enforcement controls. Access is by invitation and does not imply future plan availability. If you are not an API Security customer and want to get access, contact your account team.

The feature supports paths, query parameters, headers, cookies, JSON request bodies, and form-encoded request bodies. Profiles can validate integers, strings, UUIDs, arrays, and enums containing up to three values. Multipart forms, GraphQL, and XML are not supported at this time.

Schema Profiles validate every value when a parameter name is repeated, but they do not enforce parameter uniqueness. They also do not learn and enforce required parameters or block a request solely because it includes a new parameter.

Get ahead of zero-days

Our idea for Application Profiles does not stop at validating request structure. The same workflow can learn other characteristics of what an application expects (such as ASNs or JA4s), explain when traffic deviates from them, and give security teams confidence in defining what “good” looks like. With a Proactive Security workflow, we help security teams get ahead of zero-days!

Preventing quantum downgrade attacks against IPsec

Post Syndicated from Christopher Patton original https://blog.cloudflare.com/ipsec-downgrade-protection/

For Birthday Week, Cloudflare is helping one of the Internet’s core security protocols develop stronger protections against quantum downgrade attacks. To protect our customers and the Internet at large, we worked with the IETF to develop a mitigation against downgrade attacks on IPsec, which we’ve implemented and made available in beta across our IPsec products.

The world is racing to build the first generation of quantum computers. These new machines hold great promise, but they also create a new threat: early quantum computers will be capable of cracking cryptography we've relied on for secure communication. To address this, it is necessary to migrate to post-quantum (PQ) cryptography: cryptography we believe even quantum computers cannot break. Diffie-Hellman key agreement will have to be replaced by PQ key agreement mechanisms such as ML-KEM; classical signature schemes, like ECDSA and RSA, will have to be replaced by PQ schemes such as ML-DSA; and so on.

The PQ migration is well underway, and we’re helping the migration along by making post-quantum encryption the default in our products, open-sourcing part of our internal cryptography discovery tool, launching new post-quantum visibility features, and leading the way in the web’s migration to post-quantum certificates. Still, it will take years before all clients and servers on the Internet have been upgraded to post-quantum cryptography. In the meantime, it will be necessary for modern devices to maintain support for classical cryptography in order to connect with today’s endpoints.

The need for backwards compatibility creates its own risk. In a downgrade attack, an on-path attacker between a client and server tricks the endpoints into using weaker crypto than they support. It does so by manipulating the messages sent between client and server, making it appear to one party that its peer does not support PQ at all. In other words, a downgrade attack eliminates the protection provided by PQ cryptography by downgrading the victims back to classical, so it can be attacked by a quantum computer.

What this means is that merely adding support for the cryptographic primitives themselves is not sufficient to head off the quantum threat. The next frontier in the PQ migration is to prevent active attackers from bypassing PQ by downgrading the connection.

In this post, we focus on the IPsec protocol, a central component of a variety of Cloudflare products, namely Cloudflare IPsec, Cloudflare WAN, and Magic Transit. Like all secure channel protocols, including TLS, IPsec is vulnerable to the following simple downgrade attack as long as both classical and post-quantum authentication are supported. An attacker can impersonate a party by cracking its classical credentials and can pretend the party doesn't support PQ. However, several months ago, we discovered — or rather rediscovered, as we'll explain — a design flaw in IPsec that admits a more sophisticated attack that works regardless of which authentication method is used.

The vulnerability allows a quantum attacker to decrypt all traffic between PQ-capable endpoints. The attack is relatively hard to pull off, as it requires a quantum computation to be carried out in real time during the protocol handshake. (This is different from a harvest-now, decrypt-later attack, where the quantum computation is entirely offline.) We don't yet know if and when this attack will be feasible, but recent trends give us ample reason to be cautious: at the time of writing, resource estimates for quantum attacks on public key cryptography have decreased dramatically, leading Cloudflare to move up our transition deadline to 2029.

To inoculate IPsec to this threat, we helped the IETF develop an extension that adds a downgrade protection mechanism to IPsec. Both parties must support this extension for it to be effective: for our part, Cloudflare has rolled out beta support in Cloudflare WAN and Magic Transit, which customers can now enable by requesting the account managers to turn on the ipsec_downgrade_protection flag for their accounts. We hope to see the rest of the IPsec ecosystem follow suit in short order.

IPsec's place on the Internet

Frequent readers of the Cloudflare blog are likely already familiar with the TLS and QUIC protocols. Between them, TLS/QUIC secure virtually all the web traffic transiting the Internet today. Both operate at the transport layer of the network stack: TLS runs over TCP, while QUIC runs over UDP. 

IPsec serves a similar function, but operates at the IP layer. Because IPsec operates at an even lower layer of the network stack than TLS and QUIC, it is deeply rooted in modern network infrastructure. Cloudflare IPsec allows organizations to extend their IPsec connections over Cloudflare’s global anycast network without expensive multiprotocol label switching (MPLS) connections. IPsec is also part of Cloudflare’s Magic Transit product. With Magic Transit, Cloudflare’s global anycast network sits in front of an organization’s IP range to shield it from attacks and threats like Distributed Denial of Service (DDoS) attacks, and then hands the scrubbed traffic back to the organization via IPsec tunnels.

Despite being so deeply rooted in today's Internet infrastructure, the IPsec protocol continues to evolve. It has seen many important upgrades in the past several years, including the addition of PQ key agreement. IPsec is also on track to adopt PQ authentication on about the same timeline as TLS/QUIC. (In fact, IPsec is actually further along, depending on how it's configured. A pre-shared key is frequently used for authentication in IPsec, and this is already fully PQ!) This suggests that the IPsec ecosystem is more than capable of adapting to shifting threats.

Background on IPsec

Let's now take a peek into the protocol details that are relevant to the downgrade attack. "IPsec" refers to the mechanism used to encrypt IP packets. Before encryption can begin, the endpoints must first perform an authenticated key agreement. They do so using the IKEv2 protocol.

IKEv2 typically has two phases, called exchanges. In the initial exchange, the initiator advertises the parameters it supports and sends a Diffie-Hellman key share. The responder completes the initial exchange by telling the initiator which parameters it selected and sending its own key share.

After the initial exchange, the initiator and responder derive an encryption key from the key shares and encrypt all subsequent exchanges. The key shares are not yet authenticated, meaning each endpoint has no way of knowing where the key share came from. This is accomplished in the authentication exchange, in which the initiator identifies itself to its peer and sends a signature of its key share and advertised parameters. The responder uses the identity to resolve the initiator's credentials and verifies the signature before accepting the new connection. The responder does the same in the authentication message it sends in reply.

One crucial detail to point out here: each party only signs its outbound messages, rather than the entire handshake transcript, as in more modern protocols like TLS 1.3. This means the authenticating party never confirms to the relying party that they've observed the same sequence of messages. This will be crucial for the attack.

Encrypting handshake messages has two purposes. First, it hides the identity of the endpoints from the network. (TLS/QUIC don't have this feature by default, but can enable it using the Encrypted Client Hello extension.) Second, it allows the endpoints to begin using IPsec's packet fragmentation mechanism, making transmission of long messages over multiple packets more reliable. (This is especially relevant to handling large ML-KEM key exchange messages.)

This protocol relies on classical Diffie-Hellman key exchange, meaning a quantum attacker will eventually be able to derive the encryption key from the exchanged key shares. To mitigate this threat, IKEv2 includes an option to run an intermediate exchange following the initial exchange using ML-KEM as the key exchange algorithm:

Backwards compatibility. Crucially, this exchange is only performed if the initiator advertises support for it in the initial exchange and the responder agrees to use it. This allows for backwards compatibility with endpoints that don't yet support PQ. In particular, if the responder selects a classical-only key agreement, then the initiator will assume the responder doesn't support PQ and fall back to classical-only. Likewise, if the initiator doesn't advertise support for PQ key agreement, then the responder will assume the initiator doesn't support it.

Hello my name is Mallory

Let's think about how to exploit this parameter negotiation behavior. We'll start with a simple idea that doesn't quite work, and see what it takes to make it work.

Suppose there's an attacker between the endpoints — let's call them Mallory — who has a quantum computer. Mallory can make it appear to the responder that the initiator doesn't support PQ by intercepting the initiator's initial key exchange message, rewriting it to advertise classical-only, and forwarding the modified message to the responder.

This would cause the authentication exchange to fail. The initiator signs the message it sent, but the responder verifies the message it received. Since the message received is different from the message sent, verification of the signature would fail, unless the attacker also manages to forge a signature that the responder would accept.

That's not all, however: in IKEv2, the authentication messages are encrypted, which means Mallory also needs to compute the encryption key. But this is precisely what the downgrade attack enables: Mallory has already convinced the endpoints to fall back to classical-only, and they can use their quantum computer to recover the encryption key from the Diffie-Hellman key shares.

Still, there's no obvious way to forge a signature from the honest initiator, unless Mallory has compromised the initiator's authentication key. A paper from 2016 observes the following: because the responder only signs its own outbound messages, it doesn't actually confirm to its peer which initiator identity it accepted. This means the responder will accept an authentication message from any initiator it trusts, not just the initiator of the connection.

Suppose Mallory themself is an initiator whose credentials the responder will accept. In this case, Mallory can produce a valid signature using their own credentials. The responder will complete the connection, believing it's talking to Mallory, who is identified by IDm in the figure below. Meanwhile, the initiator (IDi) will complete the connection, believing it's talking to the responder (IDr):

This is a kind of identity-misbinding attack: the endpoints have both accepted an encryption key known to the attacker, but one endpoint has authenticated the wrong entity.

More variants of this attack are possible. For example, in a key-compromise impersonation attack, Mallory would just steal the initiator's credentials and impersonate the initiator directly, allowing them to eavesdrop until the responder has revoked the stolen credentials; this kind of attack does not require identity misbinding. These attacks are also not PQ-specific: Mallory can force the endpoints to use the weakest key agreement method they both support.

Does this attack actually matter?

The main difficulty with the quantum variant of this attack is that the quantum computation is online, meaning it must be carried out during the attack before the handshake completes. This is in contrast to other quantum threats to the Internet, where the computation is offline (harvest-now, decrypt-later attacks, cracking a TLS certificate, etc.). This gives us a little breathing room: downgrade attacks are unlikely to be the first target of cryptographically relevant quantum computers, given there is much, much more low-hanging fruit.

On the other hand, there's a non-negligible chance that Q-day will arrive before we've had time to disable classical-only across the IPsec ecosystem. We don't yet know precisely how long it will take to crack a Diffie-Hellman key agreement, but it's a safe bet that the capabilities of quantum computers will ramp up quickly once they arrive. It's best to get ahead of the threat while we're in the midst of other PQ upgrades for IPsec, especially given how long it takes for these upgrades to get deployed across the ecosystem.

Protecting IPsec

The simplest way to mitigate this attack is to disable classical-only key agreement (i.e., IKEv2 configurations with an initial Diffie-Hellman exchange but with no PQ key exchange following it). This is easier said than done, however: the reason parameter negotiation exists in TLS and IPsec at all is because the initiator doesn't always know the capabilities of the responder before attempting to connect (and vice versa).

In some cases, an HSTS-like mechanism is possible. With HSTS (HTTP Strict Transport Security), a client remembers which of its peers and servers have supported PQ in an earlier connection, and then rejects classical-only in all future connections to those peers. This works as long as you know who is trying to connect, i.e., when your peer identifies themselves. But in IKEv2, negotiation happens in the initial exchange; the peer doesn't identify themselves until the authentication exchange, by which time it's too late.

In any case, this solution fails to address the fundamental problem. Remember that each endpoint signs its outbound messages only, and doesn't sign the messages sent by its peer. This allows an attacker to create a "split view" of the protocol's execution: the initiator sees one sequence of messages, and the responder sees another. Downgrade attacks wouldn't be possible had the initiator and responder confirmed they had a matching conversation. In modern handshake protocols, like TLS 1.3, each authenticating party signs the entire handshake transcript, including the messages they received from the relying party. This allows the relying party to confirm it had the same conversation, thereby preventing the split view exploited by the downgrade attack. We prefer this more principled approach.

Introducing the full transcript authentication extension of IKEv2

We worked with the IPsec Maintenance (IPSECME) Working Group at IETF to develop an extension for IKEv2 (soon to be an RFC!) called IKE_SA_INIT_FULL_TRANSCRIPT_AUTH that endows the protocol with full transcript authentication. For backwards compatibility, use of this extension is negotiated just like any other feature. This means the extension itself is subject to downgrade attack, but the extension uses a clever trick to prevent this.

The extension is very simple:

  • Support for the extension is signaled by a notify message sent in the initial key exchange. The notification is sent unconditionally: the initiator always notifies; and the responder notifies even if the initiator didn't. This is different from TLS 1.3 extensions, where the server is only supposed to reply to an extension if requested by the client.
  • If the peer notifies support for the extension, then an IKEv2 endpoint opts into updated authentication logic. In particular, instead of signing only its outbound messages, it signs the entire transcript. Likewise, it expects its peer to sign the entire transcript.

The trick that prevents downgrades is unconditional notification. Let's say Mallory modifies the initial exchange by dropping the IKE_SA_INIT_FULL_TRANSCRIPT_AUTH notification from the initiator's message, but allows the responder's notification to go through. In this case, the responder falls back to the old authentication logic, but the initiator opts in to the new logic. The responder will end up signing a different byte sequence than the initiator verifies, causing the authentication exchange to fail and resulting in an AUTHENTICATION_FAILURE notification. A similar thing happens if Mallory drops the responder's notification but lets the initiator's through.

Now consider what happens if Mallory drops the notification from both messages. This would cause both parties to fall back to the old authentication logic, allowing Mallory to downgrade the connection and compute the encryption key. But to pull off the attack, Mallory would need to forge a signature not just from the initiator, but the responder as well.

When attempting identity misbinding, Mallory would need to present an identity for a different responder than the initiator wanted to connect to. It's as if the initiator attempted to connect to example.com, but got a certificate for cloudflare.com. Unless the initiator is severely misconfigured, this will cause the authentication step to fail.

If Mallory manages to compromise the credentials of both the initiator and responder, then they can indeed pull off the key compromise impersonation variant of this attack. However, in this case Mallory has much simpler attacks at their disposal. For IKE negotiations, Cloudflare simply acts as a responder. 

How to enable full transcript authentication

This feature is gated under a feature flag scoped to each customer account. Any customer interested in trying it out can request this flag to be enabled on their behalf by reaching out to their account team. 

Here’s what happens at the protocol level, for accounts that enable this flag.  The IKE_SA_INIT_FULL_TRANSCRIPT_AUTH notification will be sent during the IKE_SA_INIT response. We will enable this flag for all customer accounts after sufficient beta testing. The feature gate is created to account for the unlikely scenario that the customer's IKEv2 initiator incorrectly handles the new notification.

Looking forward

As of this writing, this feature is on its way to RFC status. Much of the credit goes to our co-author Valery Smyslov, who did much of the heavy lifting of shepherding the document. He also spotted the trick that makes the extension downgrade-resistant.

The PQ migration is full of surprises. Ideally these surprises are few and far between. The design flaw in IPsec that allows downgrade attacks has been known for some time, at least 10 years as of this writing. There are perhaps many cryptographic protocols in use today with latent bugs that have renewed relevance in the quantum era.

Cloudflare has implemented the full transcript authentication extension and made it available on an opt-in basis. We encourage customers to reach out to their account manager to implement and begin testing the extension, and the rest of the IPsec ecosystem to consider implementing it as the draft continues to advance through the IETF.

Introducing Threat Signals: agentic skills for open-source threat intelligence, free for every Cloudflare account

Post Syndicated from Emilia Yoffie original https://blog.cloudflare.com/threat-signals/

Organizations can now scale threat intelligence expertise the way they scale infrastructure. Threat intelligence analysts and network defenders have long automated the ingestion of structured threat feeds to help enrich their SIEM or WAF. The harder work has always been unstructured reporting: turning a research post into indicators your tools can use, without losing the context that explains why they matter. AI skills make that work possible to automate. A skill is a set of rich, detailed instructions that captures how an experienced analyst handles one part of the job, and it runs the same way on every report. 

Threat Signals puts that process into practice at scale. It’s launching today, and we made it available to every Cloudflare account. 

Threat Signals turns open-source reporting that you choose into intelligence you can act on. Its agentic skills summarize reports, surface key context, extract and normalize indicators of compromise, and apply tags — all within a private, account-scoped dataset. The end result is a contextualized indicator stored in your account’s private Threat Intelligence dataset as a Threat Event that can instantly be applied in your WAF policy.

Starting today, we are also expanding access to Cloudforce One’s Threat Events Platform, our core threat intelligence offering, to all Cloudflare accounts for free. With this expansion, each account gets:

  • API and dashboard access to Threat Signals and the ability to select one RSS feed
  • A private dataset built from the RSS feed in Threat Signals, tailored to your reporting requirements and stored for up to 30 days
  • API and dashboard access to Threat Events Platform to investigate events, indicators, and tags related to your private dataset

Essentials, Advantage, and Elite enterprise customers can extend this offering to include an expanded number of RSS feeds, access to Cloudforce One’s proprietary threat intelligence datasets, the ability to generate custom agentic skills, higher storage options for Threat Signals’ derived open-source reporting, and the ability to create custom WAF rules on open-source and proprietary threat events.

Discovery is only the beginning

We started with open-source intelligence because it is the most obvious place to prove the power of agentic workflows. We also heard from customers that their existing platforms cannot scale beyond polling 100 RSS feeds. Recognizing the critical impact open-source reporting plays in understanding the threat landscape, we sought to build an infinitely scalable platform (more on that later).

Researchers regularly publish detailed findings on vulnerabilities, malicious infrastructure, phishing campaigns, malware families, and threat actors. While RSS feed readers make it easier to discover new reporting, discovery is only the beginning. Harnessing data into a usable workflow with consistent expertise is the key to building actionable defense.

Expertise has never been something organizations can replicate at scale. A report explains how a campaign works and identifies the infrastructure behind it, but before an analyst can use that information, they need to:

  • Read and summarize the report
  • Identify relevant indicators
  • Convert indicator values into a consistent format
  • Classify the report using an internal taxonomy for tagging
  • Populate the indicators into a threat intelligence platform (TIP)
  • Preserve a link to the original source
  • Share the intelligence with the rest of the security team

Repeating that process across dozens of sources takes time; moreover, almost every step is entirely about human judgment. As a result, context is lost. Indicators inserted into your TIP are separated from the context that explains why they matter and helps assess the risk later in the remediation cycle. It's not surprising that weeks later, a domain is pushed to a blocklist and nobody understands why. 

How Threat Signals works

Threat Signals uses RSS to monitor the open-source reporting that matters to your organization. You can add an RSS feed, give it a recognizable name and category, and configure how frequently Threat Signals checks for new content. All three feed specifications (RSS 2.0, Atom, and RSS 1.0/RDF) are supported.

Each feed you select enters a Workflow that periodically polls for new articles. It uses Browser Run’s Markdown quick action to fetch and clean the article text into a readable markdown format, which is then stored in R2. The text is passed into an indicator of compromise extractor and a set of default Cloudforce One-defined skills to summarize the content, apply tags based on your account configuration, and add indicator contextualization at the IOC level.

The output is a concise summary and key points that help an analyst quickly understand what happened, who was affected, and why the report matters. All of it is searchable and tagged, so you can find the articles you care about across the platform.

Lastly, each indicator extracted is backed by a threat event within the account's own private Threat Signals dataset. The event, its indicators and tags, and the original report stay connected, so an analyst can always trace where the intelligence came from and why it is there. These indicators can then be used to create WAF rules from threat events to protect your applications and infrastructure.

What we learned

It’s not hard to write a script that pulls an RSS feed and regexes IP addresses out of it. The first version of Threat Signals was a one-week internal prototype, built by a threat analyst who wanted more out of the reports she was already reading. Turning that into something every account can rely on was harder, and most of what slowed us down had nothing to do with parsing. The hard work was in making the output something analysts would trust and actually use. 

We were tempted to let the system invent whatever tags seemed useful. The teams we talked to pushed back: intelligence labeled in an unfamiliar vocabulary is harder to use, because now there are two vocabularies to reconcile. So we limited AI tagging to each account's existing tag catalog. 

Recording whether a tag was applied automatically or by an analyst sounds like a minor piece of metadata, but it turned out to be essential. In our experience, analysts were far more willing to trust automatic tagging when they could see exactly which tags it applied.

Summaries are useful, and they are what users notice first. But what analysts kept returning to in early testing was the link between an event and the report it came from. As investigations progressed, we discovered that link consistently helped them keep track of indicators and understand why each one mattered in the first place. 

What’s next

Open-source reporting isn’t limited to RSS feeds. Analysts need to be able to quickly consume threat intelligence in various formats and pipelines. Now that we’ve laid out the building blocks for ingesting indicators from data feeds into our platform, the natural next step is to add more consumers. Be on the lookout for more data ingestion pipelines that we will support so that you can bring more actionable intelligence onto the platform to protect your organization.

Open the Cloudflare dashboard and set up your feed today

The best investigations begin with trusted context, and Threat Signals helps keep that context close from the first lead onward. Threat Signals is now generally available for every Cloudflare account via API and the dashboard. Open the Cloudflare dashboard, navigate to Application Security → Threat Intelligence → Threat Signals, and add your RSS feed. The documentation is here. 

You can also read threat intelligence research from our team, and talk to your account team about putting Threat Events to work in your enterprise environment.

Is your domain using post-quantum encryption? Now you can see for yourself

Post Syndicated from Andrew Depke original https://blog.cloudflare.com/post-quantum-visibility/

Today, we are introducing additional post-quantum (PQ) cryptography visibility tools into Cloudflare's Application Security and Logs products. You can now inspect and graph the adoption of post-quantum TLS 1.3 encryption for live traffic from directly within Logpush, Log Explorer, and the HTTP Traffic Analytics dashboard. By surfacing the key exchange algorithm negotiated on every incoming request from visitors to our platform, Cloudflare gives customers granular, per-connection telemetry to audit their post-quantum posture, assess compliance, and identify cryptographic gaps across their domains.

Cloudflare is targeting 2029 for full post-quantum security, and executing a cryptographic transition at scale requires detailed telemetry. We’ve already deployed post-quantum encryption across many of our products, including in our cloud-proxy platform and on every on-ramp and off-ramp of our SASE platform.   As many of our customers work towards quantum-readiness deadlines around 2030, we’re helping ease the transition by making post-quantum encryption the default in many of our products, sharing learnings from our internal cryptography discovery tool, and launching the new post-quantum visibility features for TLS that we’ll cover in this blog.

Bringing post-quantum visibility to the domain level

When it comes to post-quantum visibility, we already have macro-level visibility into Internet-wide post-quantum adoption in TLS through Cloudflare Radar. On Radar, we track global post-quantum encryption statistics, both when Cloudflare proxies HTTP requests from visitors (the visitor-to-Cloudflare connection) and when Cloudflare connects to origin servers (the Cloudflare-to-origin connections), as shown in this figure.

From Radar we can see that about 70% of browser-generated traffic hitting Cloudflare's network (on the visitor-to-Cloudflare connection) is protected with post-quantum encryption using hybrid ML-KEM (FIPS 203).  Meanwhile, we can see that today, just about 15% of origins that Cloudflare connects to use hybrid ML-KEM. These are aggregate numbers; the first number is aggregated across all the browser-generated traffic we see, and the second number is aggregated across all the origins we connect to.

We’ve also recently launched Automatic Key Exchange for the Cloudflare-to-origin connection, which reveals which cryptographic algorithms are supported by a given origin. This is useful because outdated configurations can cause an origin to connect to Cloudflare using classical cryptography, even if it does support a post-quantum encryption. 

While Radar and Automatic Key Exchange both provide valuable macro-level views of Internet-wide readiness, our customers have asked us to be able to go beyond aggregate numbers and dive into the behavior of individual domains.

We have long provided visibility into the TLS version used at individual domains (TLS 1.3, TLS 1.2, etc.).

But until now we have not exposed information about the cryptographic algorithms used with the TLS version used at the domain level. This means customers could not answer questions like “What fraction of traffic to my domain www.example.com is using post-quantum encryption?” This information is helpful when aiming to comply with regulatory frameworks, troubleshooting a migration to post-quantum encryption, or seeking to understand which fraction of traffic that is exposed to future quantum adversaries. Now, these questions can be answered.

Post-quantum cryptography in TLS

Before we get into the new product features, let’s do a quick review of post-quantum cryptography in TLS, so we can understand the information that the feature surfaces.

In 2024, the National Institute of Standards and Technology (NIST) stated that RSA and Elliptic Curve Cryptography (ECC) should be deprecated by 2030, and many governments and regulators have since gotten behind that deadline. That’s why today, many of our products are protected with post-quantum encryption using a cryptographic key agreement algorithm called hybrid ML-KEM. Post-quantum encryption is needed right now to stop harvest-now-decrypt-later attacks, where an adversary harvests data today and then decrypts it in the future once powerful quantum computers come online. Organizations that have data that are valuable even if decrypted in 3–10 years (public sector, defense, finance, telecom, healthcare, and others), should consider immediately protecting their traffic with post-quantum encryption.  

 In TLS 1.3, the key exchange group X25519MLKEM768 is the only recommended algorithm for post-quantum encryption. It is now the algorithm preferred by most major browsers. (Note: post-quantum encryption is not available in TLS 1.2 or any earlier version of TLS.)   If you are using Chrome, you can check the key agreement algorithm used by this webpage (or any other) by right-clicking “Inspect”, going to the “Security” tab and looking for the below:

With X25519MLKEM768 in TLS 1.3, the client and server execute both:

  • the Elliptic Curve Diffie-Hellman Key Exchange (ECDHE) over curve X25519 and
  • the post-quantum Module Lattice Key Encapsulation Mechanism (ML-KEM)

X25519 and MLKEM768 each produce a shared secret. TLS then combines those two secrets and uses the result to encrypt TLS traffic. This hybrid approach provides belt-and-suspenders security; as long as one of the two key exchanges is secure, the resulting shared secret is also secure. TLS 1.3 also supports other key exchange groups, including X25519, P-256 and P-384, all of which are just classical ECDHE over different elliptic curves; these algorithms are still used all over the web. In earlier versions of TLS you can also find key agreement based on the RSA algorithm, which is quantum-vulnerable and thankfully much less popular these days due to many known classical security problems.

But post-quantum encryption is only the first part of the story; the second part is post-quantum authentication. Once powerful quantum computers exist, we need to worry about upgrading the certificates and signatures used in TLS 1.3 away from RSA and ECC and towards post-quantum algorithms like ML-DSA. We’re actively making progress towards that goal. In fact, we recently announced that origins can use ML-DSA-44 certificates over TLS 1.3 to connect to Cloudflare, and today we announced that we’re launching a certificate authority that will support post-quantum Merkle Tree Certificates. Nevertheless, for now it remains true that post-quantum encryption with hybrid MLKEM is more broadly deployed than post-quantum authentication.

Bringing post-quantum visibility to the visitor-to-Cloudflare connection

Today we’re making it possible to see the extent to which post-quantum key agreement is used on the visitor-to-Cloudflare connection for any domain in HTTP Traffic Analytics dashboard, Logpush, and Log Explorer.

To view the TLS key exchange data on your domains, go to the Cloudflare Dashboard, and navigate to HTTP Traffic under the Analytics tab. Here you’ll get in-depth statistics about the kinds of traffic visiting your domains, now including a dedicated card for TLS Key Exchange groups on the visitor-to-Cloudflare connection. (Scroll down to find it!) Here’s a look at a TLS Key Exchange card for one of our test domains:

As you can see, the majority of the traffic to this domain uses post-quantum X25519MLKEM768 (in TLS 1.3).  We see some traffic using classical ECDHE over curve X25519 or P-256 (in TLS 1.3 or below).  The traffic labeled “None” is using either RSA key agreement (in TLS 1.2 or below) or no TLS at all. And finally we have a small number of visitors using the now-deprecated X25519Kyber768Draft00 algorithm with TLS 1.3, which we implemented back before X25519MLKEM768 was fully standardized by the Internet Engineering Task Force (IETF). We’ve waited to remove support for X25519Kyber768Draft00 until observed connections are diminishingly small, to avoid regressing clients for which this is their only way to support PQ encryption.

While we’re here, we’ll just drop a few tips about PQ-ing your traffic. If you look at your domain and find no use of X25519MLKEM768 at all, you should confirm that TLS 1.3 is enabled. In the Cloudflare dashboard, select your domain, go to SSL/TLS > Edge Certificates, and then scroll until you find the TLS 1.3 switch; switch TLS 1.3 to On. (There is no separate post-quantum setting: when TLS 1.3 is enabled and a visitor supports X25519MLKEM768, Cloudflare negotiates it automatically.) Also, if the vast majority of your traffic is over classical X25519, P-256, P-384, or None, it might be because most visitors to that domain are non-browser clients that lack support for X25519MLKEM768 and/or TLS 1.3. (Again, most major browsers do prefer to negotiate a TLS 1.3 connection with X25519MLKEM768.)

The key exchange group can now also be a filtering term in the HTTP Traffic dash. Here’s how to take a look at the traffic that is not using post-quantum encryption with X25519MLKEM768:

Analytics are great for aggregate investigations, but being able to see this information in individual log lines can be even more powerful. You can enable the new ClientTLSKeyExchangeGroup field, under the TLS category in the HTTP Requests dataset, to gain visibility into individual post-quantum key exchange in your Log Explorer and Logpush connection logs.

With this new field enabled, you’ll see it start appearing in your Logpush HTTP Request logs, like so:

Visibility to origins and more

The release of the key exchange group stats represents the first major milestone in our broader cryptographic visibility initiative. Designed for scalability, our underlying telemetry pipeline is built to ingest additional cryptographic parameters from TLS handshakes.

That’s why we’ve also surfaced the key exchange group from the Cloudflare-to-origin connection and to provide end-to-end visibility from eyeball to origin in Logpush as OriginTLSKeyExchangeGroup. (This group will be the same for all visitor connections made to that domain, which is why it's not shown in the HTTP Traffic Analytics dashboard).

And for customers that use legacy origin servers that are unlikely to support modern post-quantum cryptography, don’t despair. You can put the origin server behind a Cloudflare Tunnel, to tunnel traffic from the origin server to Cloudflare over TLS 1.3 with X25519MLKEM768, without need to upgrade the legacy origin server itself. This is what the network configuration would look like if you put your origin server behind a Cloudflare Tunnel:

Eventually we’ll be able to also surface post-quantum authentication (namely the algorithm used for certificates and signatures in TLS, including Merkle Tree Certificates) once we start to see a broader-based deployment of that technology.

Your domain has started its post-quantum journey

If your domain is behind Cloudflare, its post-quantum journey is already underway. Check HTTP Traffic Analytics dash and your logs to see the percentage of visitor connections to your domain that already use TLS 1.3 with post-quantum encryption (X25519MLKEM768).  You can also check logs to see if you’re using post-quantum encryption on the Cloudflare-to-origin connection. If your origin server is too ossified to support post-quantum cryptography, then just put it behind Cloudflare Tunnel. With the right settings and visibility, you can protect more of your traffic on Cloudflare from harvest-now-decrypt-later attacks today.

We thank Luke Valenta, Ollie Hsieh and Alex Krivit for contributions to this work.

We tested our own WAF with frontier AI models. Here’s what we found

Post Syndicated from Vikram Grover original https://blog.cloudflare.com/adaptive-ai-waf-testing/

“Is your WAF ready for frontier AI models?” We keep hearing this question from our customers, so we decided to find out.

When it comes to exploiting applications, what LLMs are really good at is iterating and mutating attack payloads faster than any human hacker could do. LLMs can use real-time responses to iterate and change their techniques by, for example, testing different encodings, sending the payload in a different part of the HTTP request, or moving to the next vulnerability to test.

Even before LLMs were around, security engineers used two common approaches to test applications: static and dynamic application security testing. The former analyzes code without executing it to identify vulnerabilities, while the latter probes running applications to find runtime flaws. There are plenty of works scanning code with frontier AI models, including details on how to build your own harness.

For the project described in this blog post, we took a dynamic approach: making the LLM act as if it was a hacker to evaluate whether a WAF is doing its job. The LLM had no visibility into source code, no view of the WAF's rules, and could only see selected HTTP response data.

We built a WAF tester that starts from known exploits and then iterates by changing how it is encoded or delivered, sends it again, and uses the response to choose the next variation. A request that was not blocked became a lead for human review, not a confirmed exploit.

We ran the tester against an authorized customer staging environment across six attack categories and recorded 1,107 attempts. After reviewing the non-blocked requests and removing malformed, benign, duplicate, and out-of-scope observations, the vast majority of the attacks were blocked by the Cloudflare WAF. The requests that got through helped us create new detections to harden our security to benefit all Cloudflare customers.

Here we will explain how we set up the system, the types of attacks we tested, which attack vectors bypassed the WAF more easily, and how we fixed it. Most importantly, we share what we learned from this process and how this exercise is becoming a foundational building block of our WAF development lifecycle.

Finally, we offer guidance to help you correctly deploy your WAF in front of your application and, most importantly, patch your software. A payload that bypasses the WAF still needs an exploitable application to succeed, so keeping your stack up-to-date remains one of the strongest defenses against attackers.

How the adaptive loop works

To test our WAF with frontier models, we built a system that iterates over multiple scenarios. A scenario means choosing one attack category, placing the input in a specific part of the request, starting with a version the WAF already blocked, and giving the tester a fixed number of attempts to try other variations. The loop runs LLM models twice: the first is the proposal call, the second is the review call.

The first call receives the starting request, the context, a short history of earlier results, and suggests the next variation, then the code builds and sends the request. The review call receives the request context, response status, selected headers, and the response body. The loop stops when mutations stop producing useful variations or when a hard coded attempt limit has been reached.

Both model calls work without access to WAF internal information. Neither receives rule expressions, rule IDs, WAF Attack Score details, or the identity of the security layer that acted. We implemented the system in Python rather than wrapping an existing penetration-testing tool. It handles HTTP replay, scenario orchestration, state tracking, and result collection.

In the current implementation, the models do not send requests directly — code controls what happens at each step. Before each request, it checks the target hostname against an allowlist, disables redirects, records the attempt, and enforces the attempt limit. After each request, it records the response and uses the model's review to choose the next predefined step. Response text may appear in a later prompt, so the tester treats it as untrusted input. Neither model call can deploy a rule nor change enforcement.

The system records structured evidence for each attempt.

Six attack categories against one WAF configuration

The main run targeted an authorized customer staging environment protected by Cloudflare’s WAF. We used an allowlisted test User-Agent so the customer’s automated-traffic controls would not stop the test before requests reached the WAF.

We ran 45 scenarios. For each, we looked for ways to deliver the same attack differently: different encoding, different part of the request, or the same destination written another way. Of these, 44 covered six attack categories: cross-site scripting (XSS), SQL injection (SQLi), command injection (CMDi), server-side request forgery (SSRF), path traversal or local file inclusion (LFI), and Log4j. The remaining scenario covered log injection, reported separately.

The WAF in the test zone was configured as follows: WAF Attack Score blocking scores of 30 or below, all Cloudflare Managed Ruleset enabled, and OWASP Core Ruleset with Paranoia Level 3.

For the headline measurement, we recorded whether the WAF blocked each request or not. The results describe the configured WAF boundary as a whole, not the performance of any individual rule or detection mechanism.

What adaptation looked like in one recorded session

Here is an example of how the LLM adapts a Server-Side Request Forgery (SSRF) attack during the test.

Cloud metadata services can expose temporary credentials to workloads. An SSRF vulnerability can let an application fetch that data on an attacker's behalf. A WAF can help stop the malicious request before it reaches the application, but it is only one layer of protection.

In this SSRF scenario, the tester sent the same cloud metadata address in different forms (such as integer, octal, and trailing-dot representations of the same IP) and placed it in different parts of the request. The WAF blocked all of them except one. At attempt 18, the model kept the same request structure as the previous blocked attempt and switched to the trailing-dot form. The client encountered a redirect rather than a WAF block.

The table below shows selected moments from the session. The hypothesis column summarizes what the model said it was trying before each move. It is not a verbatim transcript, and it is not proof that the explanation was correct.

Attempts 17 and 18 are an interesting pair: same request structure, different host representation. One was blocked, one was not. That gave us a specific question: does the trailing dot change how the WAF reads the destination? It was a lead to investigate, but not proof that metadata was accessed.

This was one selected trajectory among 45 scenarios. The next section shows how we counted and triaged the full run.

What we found

Our tester generated 1,107 attempts and the overall result was strong with XSS, LFI, SQLi, and Log4j having near full coverage. While the run produced useful findings, it also produced noise. After human review, we were left with 49 findings worth investigating, 48 of them belonging to CMDi and SSRF. 

Here is how they break down:

Metric

Value

What it means

Recorded mutation attempts

1,107

Model iterations across 45 active scenarios; not all produced a usable result

Post-triage result set

607

The 558 blocked requests plus 49 documented WAF-relevant findings

Blocked requests

558

The WAF stopped these before they reached the application

WAF-relevant findings

49

Documented for remediation analysis after human review

The rest did not produce a result worth counting as the model failed to generate a usable HTTP request, some failed before reaching the target, or the payload generated was benign.

When a request was not blocked, we worked through five questions before counting it as a finding:

Question

Why it matters

Did the tester actually send a valid request?

If the model failed or the request never reached the target, the result tells us nothing about the WAF.

Was the request clearly not blocked?

An ambiguous response is not enough to count.

Was the request still malicious?

Changing a request to get it past the WAF can also make it harmless.

Did the behavior belong to the WAF?

Some attacks only work through DNS or network paths the WAF cannot stop at request time.

Could engineers reproduce it safely?

A fix needs a stable test case with a clear expected result.

We removed anything that failed those checks and combined duplicate cases. What remained became the input for rule, normalization, and mitigation work.

Findings became detections

Not every finding needed a new rule. Some pointed to gaps in existing Managed Rules coverage. Others pointed to how the WAF normalized the request or belonged to another security control. We replayed each case and decided where the change should happen.

We grouped related findings into four sets of candidate rules, validated each finding, and tested candidates against live traffic before any rule could protect customer traffic.

Before a new or updated rule can protect customer traffic, we check its impact on legitimate traffic and assess false-positive risk. Some of the issues we find when evaluating a new rule candidate include:

Issue

Next step

Missing or narrow detection

Review whether existing rules cover the finding

Equivalent inputs interpreted differently

Engine or normalization review

False-positive risk is too high

Revise or reject the candidate

This work contributed to three changes in Cloudflare's Managed Ruleset: new detections for SSRF – Obfuscated Host and SSRF – Restricted Protocol in the July 21 release, and improvement of the existing SSRF – Cloud rule. The SSRF – Obfuscated Host detection came directly from requests that encoded internal addresses in non-standard numeric forms.

What we learned

The model was only one part of the test. We ran the same scenarios with two versions of the same model family. They produced different variations – and the same underlying issues appeared in both. Because request replay and evidence capture stayed consistent, we could compare the runs without treating either model's output as ground truth.

More attempts within one scenario did not always find more. Some scenarios started repeating earlier ideas near the end of the 25-attempt limit. We got broader coverage by testing more starting requests, attack categories, and input locations instead of extending one sequence.

The model generated requests. We decided which ones mattered. A request that was not blocked still needed replay and human review before it could become a finding, a mitigation, or a regression test. Without that review, there were no findings.

What customers can do now

WAF is just one layer of detections you can deploy. When you deploy all available protections you increase the effectiveness of your overall stack. 

First of all, check that Managed Rules, WAF Attack Score are set up correctly in front of your application. Other tools you can deploy include API Security, Bots and Fraud detection, and Threat Intelligence to strengthen your posture even further. For example, positive security controls add a different layer: instead of looking only for known attack patterns, they define the request shapes an application expects and identify inputs outside that contract. This drastically reduces your attack surface area. 

Customers do not need to reproduce this experiment. To maximize the number of rules deployed in front of your application, we recommend running Managed Rules in log first, review matching requests in Security Events, and confirm legitimate traffic is unaffected before moving a rule to Block. Alternatively, customers can reach out to their account team to get Attack Signature Detection turned on, on their zones. This new feature simplifies how to review matched traffic and how to deploy signature detections. If you already perform application security testing, run those tests against a staging hostname protected by the same Cloudflare controls as production.

Next steps

By combining adaptive AI-driven testing with human triage and validation, we found detection gaps that fixed tests might miss and turned those findings into stronger WAF protections, improving our block rate. In a future post, we will share results from further testing using a white-box approach, where the model knows both the application’s vulnerabilities and the WAF rules protecting it.

Building a post-quantum certificate authority with Merkle Tree Certificates

Post Syndicated from Mari Galicer original https://blog.cloudflare.com/pq-ca-with-mtcs/

When you type in an address into a browser, how do you know you’re connecting to the right website? The Web Public Key Infrastructure (Web PKI) is the complex and distributed ecosystem of policies, protocols, and infrastructure operators that helps you trust that you’re not being misdirected to an incorrect or malicious website. In the past few decades, this ecosystem has undergone significant changes. One is the addition of transparency: the now-mandatory requirement that all certificates be logged in public certificate transparency logs. Now it faces another challenge: the imminent arrival of a quantum computer, which has prompted us to upgrade to post-quantum (PQ) cryptography by 2029.

This transition is not straightforward: simply swapping post-quantum cryptography into certificates at Internet scale would lead to unacceptable performance degradation. This moment calls for a new approach to the Web PKI, one that allows us to treat transparency as a first-party property rather than an add-on, and design a new system that scales post-quantum signatures efficiently.

After gaining broad support across the industry, Merkle Tree Certificates (MTCs) have emerged as the path forward. This year, after a successful experimental deployment with Chrome, Cloudflare is full steam ahead on MTCs.

Following today’s announcement that Cloudflare is becoming a certificate authority (CA), we’re excited to share that this CA will support MTC issuance, targeting early 2027 for inclusion in Chrome’s newly launched Quantum-resistant Root Store. As part of our mission to help build a better Internet, and following in Cloudflare tradition of offering the strongest available cryptography for free, we will provide standard MTC issuance at no cost. Having a CA that supports both classical certificate and MTC issuance allows us to default to the most secure authentication method available, providing a painless and performant PQ upgrade path for a large swath of the Internet.

The current trust ecosystem

To understand how MTCs are changing the game, let's start with some background on how trust works on the web today.

On the client side, browsers — in this case, “TLS clients” — maintain root programs, which specify a set of policies that CAs must follow to be trusted. On the server side, CAs are the trusted gatekeepers: they operate certificate issuance infrastructure where they validate domain ownership and attest to the binding of a domain name and a public key that shows ownership of that domain.

But how do we check that CAs are following the rules? Enter certificate transparency (CT), which makes certificate issuance publicly auditable. When a CA issues a certificate, it must also submit that certificate to at least two public logs. Cloudflare has operated the Nimbus family of CT logs since 2016, and is launching Raio, a new family of static CT logs, going forward.

While the CT ecosystem makes certificates publicly viewable, it doesn't mean they are correctly issued or safe to use. Monitoring helps with this by comparing those log records with what domain owners expected and reporting suspicious activity. Cloudflare launched Certificate Transparency Monitoring in 2019 and recently made it generally available. We also publish large-scale measurements about certificates on the Certificate Transparency page in Radar (formerly known as Merkle Town).

As organizations begin upgrading their servers to use PQ authentication, certificate transparency monitoring will take on an even more important role in detecting potential post-quantum downgrades. Domain owners who have upgraded their domains to post-quantum authentication should monitor CT logs for unexpectedly issued legacy certificates to prevent clients from falling back on a malicious downgrade path.

Part of the problem with this current system is that transparency was an add-on, causing it to run into scaling issues. Certificates are frequently logged multiple times, in different forms, across multiple logs, requiring monitors to download and process every log to avoid missing an issuance. This can be expensive — making it difficult to encourage a diverse set of log operators at Internet scale. According to our estimates, PQ signatures will balloon the amount of data that CT logs need to store by 40x. This scaling challenge, and subsequent incentive misalignment, is at the heart of the post-quantum scaling problem.

The post-quantum scaling problem

We've written extensively about the challenges of scaling post-quantum cryptography, but in short: to support server authentication at Internet scale, the WebPKI must authenticate roughly a billion TLS servers without preloading every server’s public key into every client. Traditionally, CAs addressed this problem by using certificate chains as a trust-distribution mechanism. But over time, additions like key revocation checks and certificate transparency have added more public keys and signatures — five signatures and two keys in a typical TLS handshake. PQ signatures are roughly 40 times larger than classical ones, creating larger overheads that would be expensive for clients, CAs, logs, and monitors to handle at scale.

Enter Merkle Tree Certificates (MTCs), a draft specification from the IETF PLANTS working group that describes an architecture for compact, efficient, post-quantum certificates. MTCs batch certificates into an append-only Merkle tree, allowing a CA to sign the root of that tree instead of many individual certificates. This allows browsers or other clients to verify a certificate using a compact inclusion proof — a sequence of cryptographic hashes — against a signed tree head rather than validating each certificate individually. A key idea behind MTCs is "don't log what you issue, issue by logging." By coupling issuance and logging, transparency becomes a requirement for operation, rather than an add-on.

The role of a certificate authority in a redesigned PKI

We’re building out our capability to issue MTCs as an integral part of our creation of a Cloudflare CA. That means keeping track of new PQ Root Program requirements, and writing an issuance and mirroring software stack at the same time we’re building the facilities, operations, and compliance functions of the traditional  CA — no small feat!

The upside is that we get to prioritize the requirements and architecture for this new, post-quantum PKI from day one, building our setup in a way that feels right for Cloudflare's values and global network — aiming to be as transparent as possible as we embark on this new journey.

Let’s take a look at the architecture updated for MTC:

If you compare this to the traditional CA ecosystem, you'll notice that the responsibilities of a CA stay mostly the same: to validate control of a domain, bind it to a public key, and issue certificates. The main difference is that in the MTC ecosystem, instead of signing certificates directly and then logging them, the CA now maintains a transparency log backed by a Merkle tree, where an inclusion proof that the certificate is indeed in the tree serves as the trust anchor. CAs will also operate Mirroring cosigners that store a copy of issuance logs, verifying their append-only consistency and ensuring the transparency and availability of these logs for the broader ecosystem.  

Issuing MTCs

MTCs come in two forms, both of which can be encoded in the X.509 certificate format that client software recognizes today — just with a “funny” signature algorithm. In standalone form, the certificate’s signature value contains a cosigned tree head of an issuance log and an inclusion proof (a sequence of hashes) demonstrating that the certificate is contained in that log. If clients are able to obtain the cosigned tree heads out of band (e.g., via a browser update mechanism), the certificate can instead be served in landmark-relative form, where the signature value consists of the lightweight inclusion proof with no heavyweight post-quantum signatures at all.

For simplicity’s sake, let’s take a look at an example of standalone certificate issuance. When a website wants a certificate for their domain, they can request it from a CA via the Automatic Certificate Management Environment (ACME) protocol, which handles certificate requests, domain-control validation, and issuance workflows. Cloudflare's ACME infrastructure will be a fork of Boulder, the widely deployed and well-tested ACME software that powers Let's Encrypt. Let's Encrypt is actively developing MTC support in Boulder, and we plan to maintain our own fork that incorporates these upstream changes along with Cloudflare-specific modifications, contributing back upstream where possible.

When the MTC CA receives a certificate issuance request, the CA's ACME server checks that the server actually controls the domain. If those checks pass, the CA serializes that data and adds it to an append-only log.

After adding the MTC entry into its issuance log, the CA computes the updated state of the log, and then signs a checkpoint over that state. This checkpoint attests that the CA issued every entry included in the log’s Merkle tree up until that point in time.

The CA then sends its updated log state and new checkpoint to a trusted cosigner, which durably stores a copy of the CA's issuance log and checks that each new state is append-only, consistent with the previous tree, and correctly formed. This additional cosignature gives clients and monitors confidence that another trusted party has observed the same log state and verified that the CA is not presenting different views of issuance to different parts of the ecosystem. It also ensures that the issued certificates will be available for monitoring even if the CA issuance log is unavailable.

Chrome’s Quantum-resistant Root Program draft policy mandates at least two cosignatures: one from a Chrome-recognized Mirroring Cosigner operated by a distinct organization, and one from the issuing MTC CA itself. As such, we'll operate mirrors for other pilot CAs — and require at least one independent cosignature on our own issued certificates.

Cloudflare will implement our mirroring cosigner in Azul, our open-source Rust-based transparency log, and for maximal interoperability, it will implement c2sp's tlog mirror protocol.

Finally, after successfully receiving a cosignature from a mirroring cosigner, the CA constructs an MTC with the cosignatures, server's public key, and an inclusion proof. It then sends that MTC to the server, which can then use it for TLS moving forward!

Delivering PQ signatures efficiently: the landmark optimization

While standalone certificates are functional, they still send large PQ signatures over the TLS handshake, limiting their efficiency. The real performance improvements provided by the MTC design are landmark-relative certificates.

Instead of sending cosignatures in every certificate, CAs can designate a sequence of subtrees that cover all active certificates in the log as a landmark, and distribute those subtrees (along with data to authenticate them) to clients via an out-of-band update service. During a TLS handshake, the actual authentication to the server happens by the browser checking that the server's certificate data — including its domain name and public key — appears in a trusted subtree of the CA’s log. If the inclusion proof connects that certificate to a cosigned landmark, and the public key then proves possession during the TLS handshake, the client knows it is talking to the right server.

Periodically transmitting these signatures and tree metadata to TLS clients out of band, a small set of MTC batch signatures can efficiently cover billions of certificates issued by a given CA. While landmarks are more efficient at scale, they do not eliminate the need for standalone MTCs — clients may be newly installed, offline, or missing the relevant landmark update. That’s why it’s important that servers retain a standalone certificate fallback.

MTCs in the wild: results of our experiment with Chrome

This year, we ran an experiment with Chrome to test the feasibility of MTCs between a client and server. We operated a "bootstrap CA" (a fake CA that stubbed the issuance pipeline) that issued MTCs backed by a traditional certificate chain for a selection of Cloudflare domains on Cloudflare's "free" plan and served them to 50% of Chrome Beta 146. Over the course of the experiment we successfully served billions of MTCs.

For TLS, we found that the common case is fairly efficient: with a landmark-relative certificate, the handshake only needs to transmit one public key, one signature, and one inclusion proof of less than 1kB. In the experiment, we fell back to the traditional certificate chain instead of serving a standalone certificate in cases where we were unable to negotiate a landmark-relative certificate with the client. On the CT side, MTCs also change the scaling properties of transparency: the log only needs to carry hashes of public keys; there are no per-entry signatures, and the signature on the tree head covers the whole log. This prevents certificate explosion because the CA issuance log is the source of truth for all certificates the CA issues, and log consumers only need to fetch a single copy of each certificate.

The result: MTCs really work! At median, using a MTC is 9% faster using landmark MTCs over a classical signature chain (admittedly, most of this performance benefit is due to intermediate elision). And because we tested MTCs with classical signatures, we expect an even greater improvement with post-quantum signatures. Satisfied with these results, and with the level of cross-industry collaboration with MTCs at the PLANTS WG at the IETF, we began winding down the experiment last month (August 2026).

The road ahead for MTCs

We’re excited that our experiment with Chrome showed that MTCs can work in practice, and are especially excited to be able to issue certificates as a real CA.

However, there are still broader questions that we can only answer by running this great experiment with the full PKI ecosystem. Can independent monitors consume and verify MTC issuance logs at production volume? Will multiple CAs and cosigners emerge so that the system has the diversity needed for resilience? How should browsers balance the performance benefits of compact landmark MTCs with the fallback paths needed for clients without fresh landmarks? MTCs have emerged as the authoritative design for post-quantum authentication, but proving it out at production Internet scale will require participation from a diverse set of root programs, browser vendors, CAs, mirrors, monitors, and the wider community.

We see the opportunity to participate in this next phase of the Web PKI as an honor, and we take the responsibility of operating CA infrastructure seriously. CAs occupy a privileged position in the trust ecosystem — browsers, domain owners, and everyday people rely on them to validate identities correctly, protect signing keys, follow policy, and operate reliably. Before Cloudflare's CA can be trusted by browsers to issue MTCs, we will need to apply to Chrome's Quantum Resistant root store and undergo a rigorous evaluation process. We welcome that scrutiny, and we expect to hold ourselves to the same high bar as any other CA trusted with helping secure the Internet. We hope other CAs will emerge to support MTC adoption, and we're excited to work with any browser that wants to deploy MTCs.

Next.js applications, powered by Vite: introducing Vinext 1.0

Post Syndicated from James Anderson original https://blog.cloudflare.com/vinext-nextjs-on-vite/

When we launched Vinext in February, it was the result of an audacious week-long AI-driven experiment to see how far one engineer, and a stack of tokens, could get to replicating the NextJS framework backed by Vite.

In the seven months since that experiment, Vinext has grown into a framework that our customers trust and run in production for high-traffic, dynamic applications.

Today we are announcing the release of Vinext 1.0, the latest step on our journey to make it possible to deploy Next.js apps anywhere. Vinext lets you take any Next.js application, whether it was built for the Pages or App Router, and make it portable to be deployed to any web platform, including the Cloudflare Workers free plan, Netlify, or AWS Lambda.

Vinext 1.0 brings with it sweeping improvements to compatibility, stability, and caching behaviors, and sets the project up for the long term. There’s never been a better time to take your Next.js project and convert it to Vinext; just run npx vinext check and npx vinext init.

Graduation to 1.0

On release Vinext was promising, but it was incomplete. Since then, we’ve spent a lot of time both improving App Router compatibility and expanding that to Pages Router apps — which we’ve learned many customers are longtime fans of, with large applications that are complex to migrate. We didn’t want Vinext to be a tool that only worked for people using the latest App Router features.

Our focus has been on adopting both these routers, and watching our test compatibility closely, which for most important customer-requested features now surpasses 99%.

This improvement has been fueled through the community around our GitHub project. As soon as Vinext launched, that community threw it at a wide variety of applications to find the gaps. With their scrutiny, we found challenges not immediately obvious in the test coverage. Vinext needs to act exactly as Next.js behaves. It is not good enough to imitate functions with the same name. Building an alternative import { revalidatePath } is simple enough; the difficulty is in making sure it correctly affects the rendered pages, cache entry, and future requests.

Tracing requests through the application to make sure Vinext responds in the way expected — and replicating not just the API, but the behavior of this machine — was by far the more challenging aspect.

Once we’ve patched problems and brought new features forward, it’s important that we don’t regress, especially if Next.js makes a change. That’s why we’ve also built out our test suite: thousands of focused tests covering core framework behavior across both routers, the development and production server, and the deployment targets of Nodejs and Cloudflare Workers. We also run the Next.js end-to-end test suite against Vinext nightly, giving us a continually moving window on our compatibility, and making sure we immediately become aware of regressions coming from merged changes. Alongside the automated testing, we’ve been working directly with large customers that have Vinext in production to make sure they are not facing issues.

What’s in 1.0

The clearest messages we got from customers using Vinext is that certain Next.js features carry the framework and Vinext didn’t actually need to do everything that Next.js has launched in recent versions to be incredibly useful to them. So we focused on better support where you need it:

  • App Router, Pages Router, and Hybrid applications: We heard from customers that Pages Router was still important, and migrations are not a one-step process. Vinext therefore has support for both routing paths, including React Server Components, Server Actions, API routes, route handlers, middleware, and client-side navigation.
  • The complete page lifecycle: Pages can be rendered in many different ways: on the server, pre-rendered in the build, exported as static assets, or cached with page-level Incremental Static Regeneration (ISR). We’ve made sure that Background and on-demand revalidation work with any output.
  • Caching: Vinext has a shared set of caching functions across the App and Pages Router and the supported runtimes. We have further support for using Cloudflare’s Workers Cache.
  • Observability: Vinext provides Next.js-compatible tracing across both routers, so existing OpenTelemetry and Sentry setups continue to work. On Cloudflare Workers, traces also integrate with native Workers Observability.
  • Next.js ecosystem compatibility:  Vinext implements the public next/* surface and supports common Next patterns for use of authentication, MDX, image optimization, fonts, metadata, environment variables, and more.
  • First-class runtime support for Workers: While Vinext can run anywhere, server code can run in the Cloudflare workerd runtime during development and production, with direct access to bindings such as image optimization and hyperdrive. 

We’ve also made migration part of the framework: it takes two commands to verify that your Next.js install and any modification you have made is compatible, and set up the Vite and deployment configuration while keeping all your previous Next.js project structure.

When we talked to teams about what features were important for them, something stood out. Next.js 16 took a stance that Cache Components were an important part of the future of the framework, and yet most teams that we talked to were not using them and did not consider support a prerequisite to move. Therefore, Vinext today has limited support for the “use cache” directive that drives Cache Components, and though we will continue to improve compatibility there, we’re much more focused on the core priorities above.

Pre-rendering and cache warming

When we first announced Vinext, it supported Incremental Static Regeneration (ISR) after the first request, but it did not yet render pages during the build. Applications use generateStaticParams() and getStaticPaths() to identify pages that should be rendered when building, and they expect page-level ISR to connect those initial responses to background and on-demand revalidation.

Vinext 1.0 supports that lifecycle for both routers. It can prerender App Router and Pages Router routes during the build, serve those responses through page-level ISR, and invalidate them by path or tag. It also supports output: "export" when the result you want is a fully static site.

But this led us to question something: Why should this rendering happen during the build at all?

A site with tens or hundreds of thousands of possible URLs can spend a seriously long time rendering pages that receive little traffic. The build process cannot evaluate the long tail of traffic that most sites experience and therefore cannot focus compute time on the smaller number of more critical pages. Instead you waste hours of time waiting for sequential builds working their way through thousands of pages, long after the most important routes are done.

Cache warming is our solution to this, moving page prerendering from the build machine to Cloudflare’s network. Developers can continue to use Next.js primitives to identify the pages for prerendering, and Vinext can additionally identify high-traffic pages to add to this list. This happens in the background before your site is deployed to production, so that the moment it is, it is ready to serve rapid responses from the Cloudflare cache.

Inside the deployment process, this works by uploading a new Worker version and deploying it to 0% of production traffic, before then requesting pages specifically from that version. This allows the rendering pipeline to work before any real users hit the new deployment. Once the caches have been populated, the deployment can be promoted safely.

What we’re doing next

If the original experiment invented the one-off slopfork, the more consequential part has been how we can keep that process of self-improvement running indefinitely.

The project is now focused on keeping the framework up to date with everything happening upstream. Next.js canary receives new commits every day. Each morning, an agent reviews the changes, fetches diffs, and opens tracking issues for anything that could affect Vinext. Every night, the compatibility matrix is regenerated as we run the Next.js test suite against Vinext.

When one of these tests or issues reveals a gap, agents are now in the position where they can identify the change across both codebases, build a reproduction, port any relevant tests, and propose a fix.

This review has been catching missing cases, unsafe caching behaviors, and differences in the development vs. production servers.

Automation has helped us narrow the stream of activity into a focused set of changes that deserve attention, allowing the maintainers of the project to focus on only the issues that need context of how a process should map onto Vite from the Next.js implementation.

We’re building a software factory for open source at Cloudflare, and you can see what we’re up to on GitHub.

Try it out

Vinext is available for new applications, and existing Next.js projects.

Start a new application today:

Or migrate an existing application:

And then deploy it to Cloudflare Workers, with our cache warming:

Visit vinext.dev for documentation, examples, and the current compatibility matrix. Vinext is open source at github.com/cloudflare/vinext. Issues, pull requests, application reproductions, and feedback are welcome.

Introducing cf: the agentic CLI for the entire Cloudflare API

Post Syndicated from Matt “TK” Taylor original https://blog.cloudflare.com/cloudflare-cf-cli-launch/

Over the last year, agent use of Wrangler has skyrocketed.

In March 2026, agents were responsible for a quarter of Wrangler use, up from single-digit percentages the year prior. Last week, agent usage reached 48%.

Agents are more prolific users, using almost twice as many distinct commands per day, and are almost four times as likely to use six or more commands.

Agents love CLIs. But Wrangler only provides commands for around 280 operations, and Cloudflare offers thousands.

Earlier in the year we teased how we were planning to solve this and today, we’re enabling agents to use every Cloudflare product by introducing a new CLI: cf.

cf is a CLI that is built for the next generation of software development:

  • Agents can find the command they need to do anything they want to do with bespoke search and steering.
  • JSON is the default interface, pretty printed for humans and condensed for agents for maximum context savings.
  • cloudflare.config.ts is the new configuration format for the whole of Cloudflare, starting with Workers, and bringing the safety and accuracy of TypeScript to you and your agent’s language server protocol (LSP)
  • Vite becomes default, bringing with it the best local development server, and a plugin suite for developers and framework authors.

Install the open beta today globally and run it from anywhere:

cf gives your agent access to the entire Cloudflare API

What if your agent could do everything Cloudflare can do? That’s the question that sparked our interest earlier this year: agents were getting ever more powerful, but what they were able to do with Cloudflare’s CLI was still limited.

Wrangler was hand-built with each product team contributing and taking their own approach to their command developer experience. Enforcing patterns across teams was virtually impossible, even across our ~280 command paths. We had inconsistent terminology across d1 info, hyperdrive get, workflows describe as each team came up with their own practices at different times. Some teams built entirely custom experiences across thousands of lines of code that turned out to be used extremely rarely, and teams came up with different approaches to solve the same problems.

We wanted to both standardize what we had and make a massive expansion, all at once. Forge — Cloudflare’s new unified API generation pipeline — enabled us to do this, building on the idea of generating our CLI commands directly from the API schema that powers our API documentation and SDK generation. Everything we provide has an OpenAPI schema, and if we annotate this with just a little more information, we can use it as the source for Forge to make a CLI.

This enables us to expand cf from the ~280 functions that Wrangler had built up over time, to cover the entirety of the Cloudflare API surface of over 3,000 operations.

Now it’s simple to give your agent cf and ask it to go set up a worker, deploy it, monitor and observe it, protect it with Cloudflare Access, buy a domain, and front it with Cloudflare WAF, all from a single tool.

Building for an agent that has never used cf

cf is built for the trajectory of software engineering, where agentic development is drastically changing how software is built and deployed. This year we’ve been focused on providing tools to support this shift, culminating in cf. cf has been built from the ground up with agents in mind, and includes novel tools for agentic command discovery that we think will become standard in more CLIs in the near future.

Wrangler came with the advantage that years of documentation, blogs, and third-party guides have been absorbed into the training process of LLMs. It also came with the same disadvantage: changing how Wrangler works now goes against learned behavior, and significant change would be inevitable given the scale of improvement we want to make.

Introducing a new CLI that agents have never seen sounds like a big disruptive change — but actually it’s the cleanest thing we can do. Because of the design decisions we have made, the context injections we can make, and the AGENTS.md files we can append, making a switch in this way is actually less confusing than having an agent contextualize the major differences between two versions of a tool it is familiar with. We’re launching with a couple of these agent-focused features built in, with more to come.

Agents need to filter JSON, not look at tables

When agents use Wrangler, they append --json to every command they run, and then often filter the output with jq to extract a subset of fields. But only some commands in Wrangler supported --json ; many commands returned unicode tables, designed for humans looking at output in their terminal. Agents can figure these out, but it costs them more time and tokens than a jq filter.

In cf we’re taking the opposite stance: agents just need JSON, and if agents are the future primary user of this tool, it should be the default. For the vast majority of commands that will rarely be accessed by humans, this is obviously the right call.

You as the human customer of this CLI are, in reality, one step removed from using it. Agents being able to easily filter their results and then return that filtered list in whatever format you request is preferable to supplying tables you will never likely read directly.

But what if you’re looking to do something that might require real personal input, like searching for a domain to buy?

For commands that your agent can access through chaining named parameters in a long and unwieldy sequence, you can simply fill in a form. Cf deconstructs the requirements of the API into a series of validated inputs, so buying a domain, even one with complex requirements, is simple to follow.

Or, if you insist, just ask your agent to do it.

Your agent can find the right command itself

With 3,000 possible routes through a CLI, how can your agent find the right operation it needs quickly without bloating your context? For this reason we have also added cf cli search.

This command allows your agent to ask in natural language what it needs to do, and a small search index will provide a list of appropriate commands, based on their API description and parameters. We automatically tell your agent about this command when it runs --help for the first time.

Configuration that type-checks your agent

Our new configuration format is based on TypeScript, which is easy for humans and agents to parse, and allows you to write your configuration programmatically.

Typed configuration is enormously helpful for agents. We’ve found that even with no prior context of the programmatic configuration format, agents are able to easily identify and edit the configuration on demand, even across elements like env which have dramatically changed from the same named feature in Wrangler. All agents that use LSP plugins, such as Claude Code and Codex, benefit from being able to interpret more about the configuration file format in context, and make much more accurate suggestions as a result.

Compare this to TOML, which had no accessible schema, or JSONC, which had a linked schema that agents rarely used.

Some Wrangler configuration files inside Cloudflare have been condensed by 40% from over 5,000 lines, with many custom environments per developer, to factory files that build each developer’s configuration more efficiently.

This is achieved through programmatically defining each environment from the same universal base, instead of copying env blocks as was typical in Wrangler. A simple Worker with multiple environments simply switches on the Vite-native mode argument to swap between one set of configuration and another.

A simple configuration that does this now looks like:

You can migrate your Cloudflare Worker to this new format through cf migrate.

We’re also providing a few helper functions to make building your Worker a breeze.

bindings gives you a simple place for your agent to discover all the developer platform has to offer. Everything — from environment variables to storage, database, and queues — can be auto-completed and explained by your editor.

Similarly, we have included a helper for triggers, which is the new way to define routes, queues, schedules, and email triggers for your Worker. Rather than having these scattered through your configuration file, it’s now simple to find, in a single block, the actions that could trigger your Worker to run.

defineConfig.worker is just the start here. Our intention with cloudflare.config.ts is that this is how you manage Cloudflare as a whole. Every product you need — along with its API being available to your agent through cf — will be able to be expressed through typesafe configuration. Soon you will be able to configure entire policies, set up zones, configure DNS and more, all through this configuration file.

A best in class development experience

When Wrangler first started building JavaScript Workers, Vite didn’t exist. Instead, we used esbuild in Wrangler to bundle your Workers. The dev server that Wrangler made available on :8787 was something that the Wrangler team built, and modifying any of this meant reaching into the internals of Cloudflare-specific local tooling like Miniflare.

Vite is a huge improvement on this, and comes with a large ecosystem of plugins you can use, as well as providing a best in class dev server with HMR (hot module replacement), and builds that use the Rust-based library Rolldown for tree-shaking. Anything you can do with Vite, you can do with the Cloudflare Vite Plugin.

The Cloudflare Vite Plugin is the recommended way we suggest you build Workers, whatever you are building: whether that’s a frontend-focused project or a backend API. Together with our Vitest plugin it provides a cohesive development and testing environment that matches the Workers runtime and gives you direct access to bindings and platform APIs.

cf is built on Vite as default. Most of your Workers will migrate simply with agents. Others may take more time, which is why cf will continue to delegate to Wrangler for dev and deployment for JavaScript Workers that need to continue to use esbuild and Rust and Python Workers.

Migrating from Wrangler

Migrating a Worker from Wrangler is as simple as running

Workers that already build with Vite will be converted to cloudflare.config.ts for you. If your Worker relies on Wrangler for esbuild, then cf will continue to delegate builds to Wrangler.

When the open beta ends we will release a final major version of Wrangler that directs you and your agent to use cf. We’ll continue to provide maintenance support for Wrangler for 18 months after the beta ends, to give you time to migrate.

You can also take new projects and automatically configure them for Cloudflare by running cf init/deploy, which will install the Cloudflare Vite Plugin for you and create a configuration file.

Static sites still don’t require a configuration file to start, and deploying them is as simple as running cf deploy in your project.

To start a new Hello World project with cf, use cf init.

cf is open source and issues can be reported to our GitHub repository.