Higher education faces a difficult security equation. Universities hold large volumes of sensitive student, financial, health, and research data while supporting open networks, distributed users, legacy infrastructure, and increasingly complex cloud environments. Attackers have taken notice, and the pressure on security teams continues to grow.
In Q2 2025, universities faced an average of 4,388 cyberattacks per organization per week, up 24% from the same period in 2024. Nine in ten universities reported experiencing a breach or security incident during the previous 12 months, while the average cost of a data breach in education reached $10.22 million. Confirmed attacks against higher education institutions exposed more than 3.9 million records in 2025, with ransomware continuing to disrupt teaching, research, financial aid, and administrative operations.
Those figures are concerning on their own, but they only explain part of the problem. For university systems with multiple campuses, the way security is organized can create an additional layer of risk.
Why is higher education so difficult to secure?
Universities operate differently from most commercial organizations. Open access, collaboration, and academic freedom are central to their mission, which means security teams must protect environments where students, faculty, researchers, guests, and third parties connect from almost anywhere.
That openness sits alongside an unusually broad mix of sensitive data. A single university may hold student PII, financial aid and tax records, health information, proprietary research, government-funded projects, and intellectual property. Many institutions also rely on legacy systems that have been connected over time to modern cloud applications, APIs, learning platforms, and research networks, creating visibility gaps that can be difficult to manage.
Resource pressure adds to the challenge. The draft cites 94% of higher education IT leaders as saying they lack enough personnel to defend their environments adequately, leaving relatively small teams responsible for sprawling networks with large numbers of users, devices, applications, and third-party services.
For multi-campus university systems, many of these pressures are compounded by decentralized security operations. Individual campuses often maintain their own infrastructure, security tools, teams, incident response processes, vendor relationships, and renewal cycles.
The result can be limited visibility across the wider institution. If ransomware is detected at one campus, teams elsewhere may have no immediate view of the same attacker activity. If a zero-day is exploited in one research environment, another campus may remain exposed because the intelligence and response process stay local.
Fragmentation also affects efficiency. When each campus independently buys, deploys, and manages its own security stack, the wider university system can carry duplicated costs, additional management overhead, and inconsistent coverage. Fragmentation can also slow the spread of threat intelligence across a university system. If one campus detects a new attack pattern, an unusual intrusion technique, or previously unseen malware, that insight may remain local rather than reaching security teams elsewhere in time to act. A suspicious login sequence identified at Campus B, for example, could be the early signal of activity already moving toward Campus A or Campus C, but without shared visibility each team may investigate the same threat independently and at different speeds.
The same problem can appear during vulnerability response. If one campus confirms active exploitation of a newly disclosed vulnerability in a research environment, another campus may still be exposed because patching decisions, asset inventories, and remediation workflows are managed separately. What should become a system-wide priority can remain a local incident until someone connects the dots.
Attackers do not necessarily respect those organizational boundaries. A smaller or less-resourced campus can provide an entry point into relationships, systems, and data connected to the wider institution, while defenders may still be working with a campus-by-campus view.
What should university systems change?
Higher education security needs to preserve the autonomy individual campuses require while improving visibility and coordination across the broader institution.
That means giving security teams a shared view of exposure, threats, and active incidents across campuses, along with the ability to coordinate detection and response when activity in one part of the university may affect another. It also creates an opportunity to reduce duplicated tooling and processes, share threat intelligence more effectively, and make better use of limited security resources.
The objective is a model where a local security team can continue managing the needs of its own campus without losing access to the wider context of what is happening across the university system.
As the threat landscape becomes more connected, higher education security architecture needs to become more connected with it.
In Part 2 of this series, we’ll look at another pressure making that shift more urgent: the growing compliance burden across FERPA, GLBA, HIPAA, and CMMC, and why fragmented security can make regulatory readiness harder to manage across a university system.
Rapid7 helps more than 11,000 organizations worldwide take command of their security. Learn more at rapid7.com/sled.
This week, Cloudflare's Impact programs will reach $100 million in donated services. It's a significant milestone, and one that we are proud of because it means that thousands of organizations, like journalism outlets, civil society, state and local governments, election management bodies, and public schools are being protected from cyberattacks.
But Cloudflare's Impact programs have never been about philanthropy. They are a fundamental part of our business and our mission, and they continue to help guide almost everything we do.
As we celebrate this milestone and our 16th Birthday Week, we wanted to revisit not only how we got here, but also how our Impact programs continue to grow and evolve to help those working for the public interest.
Free → Impact
Cloudflare started as a free service. The original idea was to provide a basic version of our services to developers and small businesses for free, and then use the data about cyberattacks on their websites to build more sophisticated products that we could sell.
However, we quickly discovered that some of our free customers were not only doing essential work, like reporting on corruption in Africa or on the Russian invasion of Crimea, but also experiencing some of the largest attacks on our network. That realization changed how we thought about our free services. We committed not only to making them available for free for everyone, but also to doing more for organizations being targeted by powerful adversaries simply for serving the public.
Cloudflare launched Project Galileo in 2014 to provide more advanced security services for important but vulnerable people and organizations online, including journalists, human rights defenders, and civil society groups. Today, the program includes more than 3,500 domains in over 120 countries. In 2025, Cloudflare blocked more than 38.5 billion DDoS, website vulnerability, email phishing, and other cyberattacks against Project Galileo participants, almost 105.4 million per day.
Over the last 12 years, Cloudflare has continued to expand what we now call our Impact programs. Although each program is unique, our goal is the same: to support organizations and institutions serving the public, particularly those that would not otherwise have access to the necessary cybersecurity services. For example:
The Athenian Project (2017): Supporting state and local governments operating democratic elections, which includes more than 440 websites in 33 U.S. states. We later expanded that program outside the United States and now protect election management bodies in eight countries, including Canada, North Macedonia, Georgia, and Moldova.
Cloudflare for Campaigns (2020): Partnering with Defending Digital Campaigns to provide free cybersecurity services for candidates for public office, which now includes more than 530 websites across the United States.
Helping keep these organizations online by protecting their websites and internal data remains essential. In 2026, Cloudflare released its first annual report on cyberattacks against civil society, which found that civil society organizations are targeted more frequently and more intensely than other Cloudflare customers. For example, Project Galileo participants faced attempts to exploit security vulnerabilities in websites at a rate more than seven times higher than an average user. Cloudflare is also on pace to more than double the number of applications to Project Galileo from last year.
But Cloudflare's Impact programs have never been static; they evolve alongside our company and technology, and the organizations they serve. Increasingly that means not just defending public interest organizations, but empowering them to adapt and thrive in the era of AI.
Looking to the future
In early September 2026, on a rainy day in Barcelona, Cloudflare co-hosted a hackathon. Because our developer platform is such an important part of our business, we hold these events all the time. But this one was different: instead of a room full of software engineers or startup founders, it was the first time we held an event specifically for journalists.
Media Party hackathon co-hosted by Cloudflare at the BIT Habitat in Barcelona (September 9, 2026).
The event was part of a three-day conference organized by Media Party, a nonprofit dedicated to media innovation through digital tools. The event was designed to bring together journalists, developers, and strategists to solve a single problem: how to help newsrooms adapt to an AI-driven, post-search information landscape. The sprint focused on four themes: workflow automation, agentic journalism, synthetic-content verifications, and information integrity.
Each team received free access to Cloudflare's developer platform and the assistance of volunteer Cloudflare engineers to see what they could build in a day.
Four teams made it to the final round. The winning team, AIdas, built a tool that helps researchers and journalists study AI bias across politically contested topics by comparing how different LLMs answer the same question, and recording their responses as open data.
The hackathon was just one part of a broader effort across Cloudflare Impact to expand beyond cybersecurity services to help public interest groups adapt to a changing world:
Protecting Local News from AI Crawlers: Last year, Cloudflare provided free access to our Bot Management and AI Crawl control for Project Galileo participants, including more than 750 journalists, independent news organizations and non-profits supporting news-gathering around the world. These tools will help these organizations understand and control how their content is accessed by AI crawlers, and safeguard their reporting from unauthorized scraping.
Non-profit startups: Last year during Birthday Week, Cloudflare announced its startup program, which provides more than $250,000 in Cloudflare credits, would be available for the first time for non-profit organizations. This week we will announce the first 30 organizations accepted into the program and how they are serving their communities with tools built on our developer platform.
Automation tools for human rights: This week we will also announce three new projects that Cloudflare engineers have built using our developer platform for three leading human rights organizations, covering topics including tracking transnational repression, digital rights legislation and policy development, and corporate human rights due diligence.
Across all of these new efforts, the goal remains the same: to help organizations doing essential work access the tools and support they need to continue to advance their missions.
Join Us
I had the opportunity to meet with two of the Cloudflare engineers who volunteered at the hackathon in Barcelona. They both mentioned to me that one of the reasons they came to work at Cloudflare was Project Galileo, and the chance to use their skills to help organizations working in their communities.
It was an important reminder that Cloudflare's Impact programs and our mission are not just things we have done. They continue to shape our identity, including through the people who choose to come work with us.
If that sounds like the kind of work you want to do, come join us.
From our conversations with companies at every stage of their AI adoption journey, we've seen some common patterns. First, there is an exploration period as you bring on every new tool, dole out API keys freely, and let the tokens flow. Then, you converge on the canonical tools for your organization for agentic coding, for non-technical workflows, for running and deploying agents. As companies formalize their AI adoption, they want to manage and oversee token spend for users, but budgets and rules only go so far. The best savings are the ones users never notice.
Today, we are releasing Cloudflare's Auto Router in public beta, available through AI Gateway. Set your model to cloudflare/auto and the Auto Router will automatically route each request to a model that is capable enough for the task, without requiring an end user to think about model selection. Our early results using the Auto Router internally through our OpenCode harness show a cost savings of up to 30% when compared to using only frontier models like OpenAI Sol and Anthropic Claude Opus.
Why we built this
From our own experience tracking AI spend at Cloudflare, we’ve learned managing costs requires a multipronged approach. Previously, we talked about how to set budgets and limits around AI spend, and how to see who is spending across your organization by linking employees to their AI usage.
In many harnesses, including OpenCode, Claude Code, and Codex, individual users still select models manually. Of course, not all tasks are created equal, and often individuals end up using models that are overkill for their work. For example, you don't need Opus-level intelligence if you're looking to summarize an email or chat threads. However, you wouldn't want to block that model completely from your security engineering team.
Our goal is for AI Gateway to be the control plane for organizations deploying AI internally. Because every request from every user, agent, and tool already flows through it, AI Gateway is in a unique position to do more than observe and enforce. Budgets, spend limits, and identity-aware analytics give organizations visibility and guardrails, but they still rely on individuals to make cost-conscious choices request by request. The next step is for the gateway itself to make intelligent decisions on a user's behalf: sending each request to a model that is capable enough for the task. That way, organizations reduce spend automatically, while users keep access to the most capable models when their work actually needs them.
The results
We use Auto Router internally at Cloudflare within our OpenCode deployment and within Cloudflare OS, our custom agent harness. In our internal usage, we’ve seen results comparable with frontier models for coding tasks.
Auto Router does best when used across a wide range of knowledge-work tasks, like those typically found in a large organization with work spanning both technical and non-technical teams. We evaluated cloudflare/auto against OpenAI’s GPT-6 Sol and Anthropic’s Claude Opus 5.5 on our internal general knowledge work benchmark. The benchmark uses simulated workspace tools and covers common day-to-day workflows across email, calendars, Slack, files, travel and finance. Each task requires the model to use these tools to produce a verifiable answer or complete an action.
Model
Successful Trials
Success Rate
Total Cost
Cost per success
cloudflare/auto
252/291
86.6% (+6.2/−6.9 pp)
$2.10
$0.0084
Anthropic Claude Opus 5.5
281/291
96.6% (+2.7/−3.8 pp)
$5.91
$0.0210
OpenAI GPT-6 Sol
245/291
84.2% (+6.5/−6.9 pp)
$2.64
$0.0108
97 tasks with three samples per model per task. Parenthetical values show 95% confidence intervals estimated from 10,000 task-level bootstrap resamples, preserving all three repetitions within each task. “pp” indicates percentage points.
Our Auto Router delivered similar performance to other state-of-the-art daily-driver models, coming in at 80% the cost of Sol and 35% the cost of Opus. While that may initially seem surprising, one way to frame the problem a model router solves is through the “jagged frontier” across models. The ability to solve a problem often exists somewhere in this portfolio of models; the router’s job is to choose the right model for each task while balancing quality and price. Savings come from not paying frontier rates for non-frontier work, and they grow with how much of that work you have.
Another insight is that lower token prices do not always produce lower-cost outcomes. A model that looks cheaper on paper may end up using disproportionately more tokens to solve a problem. A router should minimize predicted trajectory cost, not just load-balance by dollars per million tokens.
This is already useful today, but it’s only the beginning of what the Auto Router can learn from Cloudflare’s position in the inference path.
How it works
When you send a request to cloudflare/auto, AI Gateway first builds the pool of models that can actually serve it. It filters out models that do not support the request format or execution mode, and accounts for the credentials, billing configuration, access control policies, and spend limits attached to the gateway. It will also filter out unhealthy upstream providers or models during downtime and automatically bring them back into the pool after an outage.
For the remaining candidates, the router looks at a compact view of the conversation. It considers the most recent messages, prioritizing the newest turns. The conversation is then sent to a multi-head classification model running on Workers AI and deployed on GPUs across our edge network. The classifier produces two sets of signals. First, it assigns probabilities across 14 task categories (like coding, planning, research, data analysis). It then rates the request across four dimensions on a scale from one to five: complexity, ambiguity, stakes, and dependence on earlier context.
A separate scoring matrix combines those signals with model benchmark results to estimate how well each model fits the request. To calibrate the scoring matrix, we defined the preferred model for a set of example task and difficulty profiles, then adjusted the weights to produce those choices.
Finally, the router combines expected quality with each model's input and output token prices. On straightforward requests, price carries more weight, so a smaller model can win when it is capable enough. As difficulty rises, the cost penalty falls and stronger models have more room to win. In simplified terms, cloudflare/auto selects the model with the highest utility as defined by:
For long agentic sessions like debugging or coding, cost is less driven by the model’s list price than by the cost of cache reads, which grows with session length. Switching models throws the cache away and forces a new model to write the whole context again. This can be worth it, as a model with a cheaper cache-read and cache-write prices can pay back the rewrite quickly.
Rather than completely avoiding model switching, the Auto Router accounts for the cost of cache reads and writes. Within a turn (one user input loop), the cache is hot and switching rarely pays off, so it’s better to keep using the same model. Across turns, the Auto Router applies a switching penalty that grows with the number of tokens already in context. A model that still holds a live cache for the session is priced at its cheaper cache-read rate. Every other candidate is priced at the full cost of rewriting the context, so the deeper the conversation, the more a switch has to earn back, through higher quality results that use fewer tokens overall or cheaper cache rereads. Switching models has another cost: most models can't read another model's reasoning tokens, so a model switch that drops reasoning tokens means that the new model may have to redo it at output prices. In the future, we want to account for this by having the router prefer to stay within the same model family when it switches.
From there, the router returns a ranked list. AI Gateway attempts the winner first and can move to another eligible model if that provider cannot serve the request.
This overall design has several benefits. The two-stage architecture (task and dimensions classifier to scoring matrix) means that routing decisions are legible because you can inspect each task’s predicted category and complexity to see how it translated into the model choice. Adjusting the router when a new model is released also does not require retraining — we only add its benchmark-derived weights to the scoring matrix. The same classifier can also support different routing profiles. For example, in addition to cloudflare/auto, we plan to release other routers in the future, including cloudflare/auto-best, which uses the same classification and model pool, but selects the highest expected quality without applying the cost tradeoff.
What's next
Our release today is only the starting point, and we’re continuing to invest in research and new routing strategies. In the near term, we want to:
Expand the models offered through cloudflare/auto
Include zero-data-retention requirements when filtering models
Account for provider capacity when selecting models
Select the appropriate reasoning or thinking level for each request
Add full support for the Responses API and WebSockets
Explore structured decision models as a first-pass classifier
As agents help us build more complex applications, both humans and agents need a better way to stay on top of what goes wrong in production. Coding agents can already query observability data, navigate a repository, change code, write tests, and open a pull request. What remains manual is connecting those steps: recognizing that repeated failures come from the same bug, gathering the relevant logs and traces, sending that context to an agent, and checking whether the fix worked. Without that structured handoff, the agent must search raw telemetry to reconstruct the scope and context of the failure before it can investigate.
Group repeated exceptions, 5xx responses, and error logs into one issue.
Send the error, stack trace, logs, traces, and Worker version to a configured coding agent.
Trigger the agent’s configured workflow — from triaging an issue to querying more data to opening a pull request.
Fix your first issue with the following prompt for your agent with CF CLI or checkout the documentation to get started:
Catch failures automatically
With one line of configuration, you can start receiving Issues detected on your Worker with no additional instrumentation required. Issues are built into the Workers runtime, so there is no SDK to install or application wrapper to add.
Once enabled, Issues records uncaught exceptions, failed invocations, HTTP 5xx responses, output from console.log() and console.error(), and logs that contain a stack trace. It also flags runaway alarm conditions and code that writes large volumes of logs inside loops.
Consider a Worker whose handler starts throwing errors after a deployment. Every failed request has a different request ID, but they all come from the same bug. Issues groups them together and shows when the error first appeared, how many times it has happened, and whether it is becoming more frequent.
When you open an issue you can see the error, a stack trace when available, the logs and traces leading up to it, the Worker version, request details and trend of the issue over time, as shown below:
Contextualizing errors for your agent
Cloudflare can capture what happened inside the Worker, but it does not know which users, accounts, or sessions matter to your application. Use the Worker runtime's built-in OpenTelemetry API to add those identifiers, without installing another package:
Those identifiers appear with each occurrence. In this issue, you can now see whether failures are concentrated in one account or session before sending the issue to an agent.
Send detected issues to your agent
An issue no longer has to sit in a dashboard while someone copies a stack trace and pastes it into a prompt. Configure an automation once, and when an issue crosses an occurrence threshold or returns after a quiet period, Issues sends it straight to your agent through Automations. You can choose when the automation should run and where the issue should go.
This can be via:
Built-in coding agents: Connect Claude Code with a routine ID and token, Cursor with an automation webhook URL, or Devin with an API token and organization ID.
Generic webhooks: Send issue context to your own agent or HTTPS endpoint.
Chat and incident management: Notify your team through chat or an on-call workflow.
When the automation runs, Issues sends the failure summary and diagnostic context captured with the issue — the exception, error, source-mapped stack trace, leading and trailing logs and traces, Worker version, and the application context you added. For deeper investigation, you can connect the agent separately to Cloudflare MCP that lets the agent query the related logs and traces so that it can propose code and test changes and open a pull request.
You stay in control of what reaches production: review the pull request, deploy the fix, and mark the issue resolved.
How Issues uncovered and resolved two Workflows bugs in one day
Cloudflare Workflows, a primitive that powers long-running, multi-step applications, is built entirely on the Workers platform.
Behind the scenes, its services keep track of steps, retries, and saved state. This makes Workflows a useful place to test Issues on our own production systems. Within a day of turning it on, the team found two unusual problems hidden inside a large volume of traffic.
A migration stuck in a retry loop: A Workflows control plane migration repeatedly hit a SQLite foreign key error when attempting to apply migrations in an edge case. Issues allowed the Workflows team to identify the problem and fix it.
A deletion process that never completed: Workflows discovered that during deletion of Workflow instances, there was an edge case where they could exceed a Workers subrequest limit and not finish the deletion. Issues helped the Workflows team identify the issue and do a fix.
Instead of leaving the team to connect thousands of separate pieces of telemetry and user reports, their automation setup sent these issues directly to Cloudflare OS, which followed the errors into the Workflows code and proposed a fix for both issues.
Get started
Ready to see what Issues finds in your app? To get started:
Set observability.issues.enabled to true in your wrangler.jsonc file
Set up your first automation in the Cloudflare dashboard to send issues to your destination of choice whether that’s an agent, a webhook, incident management tool or chat platform.
If your agent is handling setup, it can also use the new cf CLI to inspect issues and create automations. Check out our documentation to learn more!
You just thought of your next great idea, and buying the right domain feels like the easiest way to make that first bit of progress. Naturally, you open a new tab in your browser, only to find yourself face-to-face with an experience that feels like a budget airline peppering you with add-ons at checkout: Want security? How about a website? Do you want email? You’re just a few minutes into building your next idea, and it doesn’t feel fun anymore.
Launched a decade ago, Cloudflare Registrar has always taken a simpler approach. Domains at cost, transparent pricing, and no unnecessary upsells. But simplicity shouldn’t begin at checkout. It should begin the moment you start looking for the right domain.
Today, we’re bringing that same simplicity to the entire experience of finding and buying a domain. Our new domain search shows every extension we support, responds as quickly as you type, and makes hundreds of possibilities easier to explore through sorting, filtering, and transparent pricing.
And you know what is particularly good at ignoring distractions and staying focused on the destination? An AI agent. We designed Cloudflare Registrar to work naturally with agents through the Registrar API, MCP, and our newly launched cf CLI. You can ask your favorite agent to find the right domain, buy it, or transfer one you already own.
The agentic registrar, expanded
In April, we launched the Registrar API beta, allowing developers and agents to search for, check, and register domains programmatically. We have expanded the API since then. The new sandbox lets you test registrar workflows without purchasing a domain or triggering a real transaction. Our extensions endpoint returns relevant information for each of the 420+ extensions we support, helping you account for the different requirements across registries. We also added transfers, so you can bring domains from another registrar into Cloudflare programmatically.
The Registrar API is available through Cloudflare MCP, giving agents access without requiring a separate integration. Earlier this week, we also announced the launch of cf CLI, bringing the same capabilities directly into your terminal. You can prompt your favorite agent to search for, register, or transfer a domain. These tasks already lend themselves naturally to a conversation:
“Is example.com available?”
“Buy example.com.”
“Transfer example.com from my current registrar.”
The way people interact with the Internet is changing. Cloudflare Registrar should feel natural whether you use it through an agent or in your browser. For many people, the browser is still where the search begins, and that experience was long overdue for an overhaul.
Search simplified
Previously, our search page showed around 20 available results from a subset of extensions, sometimes modifying your search term to suggest related options. You couldn’t see that exact name across every extension we support.
We decided to take a simpler approach: show you the exact term you searched across every supported extension. Results appear as you type and continue to load as you scroll, letting you explore hundreds of options without starting another search. We also include domains that have already been registered, giving you a more complete picture. If you only want domains available to buy, you can filter everything else out.
More results might not sound simpler, but these are the results you asked for. Sorting and filtering help you narrow them down. Whether you are logged in or logged out, on your phone or at your desk, the experience feels the same.
Making a complex question feel simple
Cloudflare Registrar supports 420+ extensions. A single search can thus create more than 420 separate availability questions. Each extension is operated by a registry that maintains its official registration records and provides the authoritative answer about a domain’s availability and price.
Asking every registry every question at once would be slow and wasteful. Registries respond at different speeds and impose request limits, and much of the work would be for results the person might never view. To solve this, our search gathers evidence from multiple sources: (1) prepared availability datasets (e.g. zone files) and cached answers; (2) DNS answers; (3) live registry lookups.
A hit against an availability dataset or DNS can tell us that a domain is already in use, but a miss cannot necessarily prove availability as a domain may be registered without being configured in DNS, or it may be on a blocked list. A recent registry-derived answer is stronger but becomes stale over time, while a live registry check provides the freshest authoritative answer but takes longer and draws on limited upstream capacity. Our new search progressively probes these sources while balancing speed, freshness, and certainty for each result. As better or more accurate information arrives, we update only the affected result dynamically.
How we built search for speed and scale
We built the new search on the same Cloudflare developer platform available to our customers. Workers run the public search entry point and the services that gather availability evidence. Durable Objects give each active search one coordinator, while Workers KV stores prepared data that can be reused across searches. Together, these primitives let the service scale across users and 420+ extensions while keeping operating costs low.
We prepare useful evidence before a search begins: a purpose-built pipeline converts registry zone files and other bulk sources into compact availability datasets in Workers KV. Large datasets are split into smaller pieces, so a lookup retrieves only the data required for that domain. These fast checks can answer many questions without making a new live registry request.
A search session is composed of a Durable Object that coordinates that specific search interaction. It establishes the result order from the query, sort, and filters without waiting for network lookups, remembers the best evidence received for each domain, and tracks which results are visible so lookup work follows the person's attention.
WebSockets provide the bidirectional connection, and we designed an application protocol on top of them to connect the browser to the resolution process. An initial snapshot establishes the ordered list. Subsequent delta messages contain only the fields that changed, letting the browser update one domain instead of downloading the full result set again. Before sending a delta, the Durable Object compares the new evidence with the current answer: stronger evidence can replace it but weaker evidence cannot.
A separate Worker gathers additional evidence. It can query DNS through Cloudflare's 1.1.1.1 resolver, make a live Registrar check, or reuse a recently cached answer. Reusing fresh answers avoids repeating upstream requests. Because an available domain can be registered at any moment, an available answer has a shorter useful cache life than evidence that a domain is already taken.
Together, these pieces turn hundreds of independent availability checks and all that coordination into one coherent search experience. Cloudflare’s Developer Platform gives us all the building blocks to hide that complexity and craft a domain search experience that feels simple and is among the fastest in the world.
Transparent pricing and price drops
Making search feel simple is not only about speed. It is also about knowing exactly what a domain will cost. Cloudflare Registrar has offered domains at cost since day one. Great prices are part of making domains simple, but so is knowing what you will pay. Our new search and our new pricing page show both the initial registration price and the renewal price for every domain. When a domain is discounted, we show the original at-cost registration price crossed out alongside the promotional price.
Beginning with Birthday Week (this week!), we’re offering first-year registration discounts on select extensions including .io, .dev, .app, and .tech. You can explore every discounted extension directly from the search page as well as our newly launched pricing page.
Whether you search in your browser or ask an agent, our goal is the same. Remove the friction between having an idea and making it real. Buying a domain for your next idea should be fun and feel like progress.
Find your next domain
Choose how you want to get started:
Search in your browser: Explore every available extension and find your next domain.
Ask your favorite agent: Install cf CLI, then prompt your agent to search for, register, or transfer a domain.
Build with the API: Use the Registrar API to bring domain search and registration into your own application or workflow.
However you choose to do it, finding your next domain should be the fun part.
Acknowledgements: This simplicity was a result of cross-team collaboration. Special thanks to Pedro Menezes, Shobhit Kuruvilla, Lucy Dryaeva, Fred Pinto, and the Registrar Team, the Design Engineering Team, and the Forge team.
Today, we’re making the Cloudflare Monetization Gateway available as part of a closed beta, and showcasing four customer use cases that are in production today. Since we announced the plan three months ago, we have been working closely with our customers to make the Gateway fast, flexible, and easy to use.
With just a few clicks, the Monetization Gateway allows domain owners to charge agents for access to their website, APIs, MCP tools, or datasets. Request access today in the Cloudflare Dashboard.
Proliferation of agents and machine payments
Today, most software businesses sell their products through subscriptions or prepaid credits. These business models require buyers to make a considerable upfront economic investment, so buyers limit themselves to a few subscriptions that fit into their budget. But this business model doesn’t align itself with how the predominant source of traffic on the Internet — agents — operates.
Agents seek outcomes, whether that’s sourced from a direct question or an implicit elicitation. To achieve an outcome, an agent may visit new sites, call MCP tools, or ingest data feeds. Businesses everywhere want to be in the critical path to help agents get better results. This helps them meet new users where they are, get paid for their inputs to those agents' answers, and take a leading position in headless agentic commerce.
Companies need to align their business models with the consumption patterns of agents to capture this opportunity. Payments must match an agent's consumption unit: per request, per search query, per token. The popular payment rails of today are unable to support this. These payment rails assume that the buyer will accept high latency, will only transact in large dollar values, and that they are a well-known, or identifiable, entity to the seller. Agents desire the opposite: cheap, fast, and reliable payment rails that can safely scale with minimal human intervention.
Sellers and buyers will be incentivized to align their business models with consumption, and they’ll use the payment network that helps them achieve that. Today, stablecoin transactions, and their underlying blockchain networks, are able to support these requirements. Over time, we may see existing networks introduce solutions or new payment networks come online. The Monetization Gateway exists to help ensure that sellers can focus on their product, and buyers can receive a predictable buying experience.
Building with the Monetization Gateway
The Monetization Gateway lets sellers charge agents for any resource behind Cloudflare’s network, priced per use. It's built for resources where every request is the use, like APIs, tools, and data. High-value content is different: a page might be crawled once and used a thousand times. For that, Pay Per Use offers a trusted network of verified buyers who report each use and pay for it.
The gateway uses the HTTP 402 Payment Required status code so buyers can offer payment for a resource over HTTP, inline with the request for the resource itself. This is an important distinction because there is no redirect to a checkout page and no separate payment API to call.
Sellers define which requests require payment, the cost, and where the payment should be sent. Buyers receive the payment instructions, sign an authorization, and receive the resource after the payment has been settled.
In its simplest form, this can be thought of as a ‘paywall for agents’. Sellers can go live in seconds with a few clicks, identify the audience they want to charge, and ensure that buyers are unable to access the resources until the buyer has paid.
Beyond its ease of implementation, we provide sellers a set of pricing and monetization capabilities out of the box. You write pricing rules that match any part of a request, such as the URL, headers, or query parameters, and choose from several pricing schemes. We handle the rest: payment verification and settlement through Coinbase's x402 Facilitator, failures and retries, analytics, and keeping up with changes to the x402 protocol. Payments settle on the Base blockchain using USDC, a stablecoin pegged to the U.S. dollar. Over time, we plan to help sellers make their services discoverable to agents, expose logs for all transactions, support additional payment rails, and incorporate identity primitives.
Let’s look at how our customers are using it in production today.
Cloudflare AI Gateway: pay for inference
Cloudflare’s AI Gateway is a control plane and model marketplace that allows developers to access hundreds of AI models with a single API key. Our goal is to operate the richest model catalog in the ecosystem and get it in the hands of as many customers as possible. AI Gateway allows customers to purchase credits that can be used toward all AI Gateway consumption. But as more agents carry wallets, maintaining a credit balance as the only usage mechanism adds unnecessary friction.
Starting today, U.S.-based Cloudflare customers can pay for inference at request time to a select set of models by adding the header PAYMENT-METHOD: x402. Detailed instructions are in our AI Gateway Developer Docs. Over time, you can expect to see the HTTP 402 status code embedded natively through more of Cloudflare’s infrastructure products.
AI Gateway has a complex pricing engine powering how each token gets billed. We’ve co-designed Monetization Gateway to meet the requirements of origin-controlled pricing. When this configuration is enabled, the Monetization Gateway requests pricing information directly from the seller (AI Gateway). For AI Gateway, this means being able to rely on its existing pricing module without having to maintain a rules catalog in the Monetization Gateway.
Need to set prices from your origin? Email us and we'll help you get started.
Ceramic.ai: pay for web search
"The web was built for a human buyer, someone who signs up and enters a card. Agents need to act on their own behalf, and payments are how they do it. Search is one of the first things every agent needs, so it should be one of the first things an agent can buy." — Dr. Anna Patterson, Founder, Ceramic.ai
If inference is the reasoning, search is the reality check. Ceramic.ai provides a web search API built for agents. Ceramic.ai maintains a proprietary index of more than 40 billion pages and has optimized every layer of its stack for machine callers, returning search results in as little as 50 milliseconds. At that speed, an agent can search repeatedly throughout a task.
That makes search a natural fit for agent-native payments. Search is one of the most elastic things an agent buys. A simple question might take one query, while a complex research task might take hundreds. When an agent can pay for search itself, it no longer has to ration queries against a budget someone set in advance. It can decide how hard to look based on the task, spending more when the stakes call for more evidence.
Ceramic.ai uses the Monetization Gateway’s fixed pricing monetization scheme to allow agents to pay to execute a search without an API key. Use the demo and Ceramic.ai docs to learn more.
Stocktwits: pay for stock signals
“Agents change the way data gets consumed. Instead of a customer signing up for a subscription or negotiating an enterprise license, an agent can ask for exactly what it needs, when it needs it. The ability to charge for that individual request opens up a completely new way for Stocktwits to make its data available.” – Howard Lindzon, Founder and CEO, Stocktwits
When a stock starts moving, one of the first questions traders ask is: what is everyone else saying?
For 18 years, that conversation has been happening on Stocktwits. Launched in 2008, Stocktwits pioneered cashtags (like $NET) to organize market conversations around individual stocks. Today, more than 10 million people use Stocktwits to follow markets, share ideas, and see what other investors are talking about.
Because Stocktwits is built specifically for investors, it provides a real-time view into what retail investors are watching and discussing. Stocktwits turns that activity into signals including:
Sentiment: whether recent posts about a ticker are predominantly bullish or bearish
Message volume: how much conversation a ticker is generating relative to its typical activity
Followers: how many Stocktwits users follow a ticker, providing a measure of sustained retail interest
Trending: which tickers are gaining attention fastest
These signals have long been available through Stocktwits' existing data products. The Monetization Gateway creates an additional distribution channel for a new type of customer: AI agents that need market context on demand. For example, an agent monitoring a portfolio or researching an investment might check whether conversation around $NET is taking off, whether sentiment is skewing bullish or bearish, or how many investors follow the ticker. Instead of requiring a traditional data license, a developer can pay for the individual requests their agent actually makes.
To start, Stocktwits built a separate, agent-facing path for these requests, with Cloudflare's Monetization Gateway sitting in front of it. Its existing API and enterprise data products remain unchanged. Each request is priced individually, giving Stocktwits a way to make its market signals accessible to AI agents while preserving its existing data products and licensing models. Get started with their docs today.
API2PDF: pay for API access
API2PDF is a REST API that helps developers generate PDFs from HTML and Office documents. Since launching in 2018, API2PDF has seen a growing number of AI agents directing developers to its service and now treats them as a first-class customer.
API2PDF’s consumption-based pricing model charges customers for the bandwidth and compute required to fulfill each request. Previously, customers needed to create an account and obtain an API key before using the service. After the first month, they needed to supply a credit card, which resulted in a >50% drop off in conversion. API2PDF now uses the Monetization Gateway to return an HTTP 402 response when a user or agent makes a request without an API key. Because each request depends on variable compute and bandwidth costs, API2PDF uses variable pricing to inform agents of the maximum price a single request could cost. Once the client completes the payment, API2PDF fulfills the request and settles only the actual consumption.
Agents face the same build-versus-buy decision developers do. An agent that needs a PDF could burn tokens to generate one itself, or it could pay a specialized API like API2PDF a fraction of a cent to do it properly.
Get started today
The Monetization Gateway is available in closed beta to eligible U.S.-based sellers and buyers, with support for new geographies on the way. If you are interested in making your product agent-native, we want to hear from you. Fill out the onboarding application in the Cloudflare Dashboard, check out our Developer Docs, and we’ll get back to you shortly. If you have questions, ideas, or want to help us scale agentic payments, please send us an email.
We’d like to thank our partners for their contributions and feedback. Without them, this launch wouldn’t be possible. This includes Vail Gold and Jeff Rafter from Cloudflare AI Gateway; Dr. Anna Patterson, Sean Costello, Sadé Ried, and Autumn Yuan from Ceramic.ai; Leo Gorkin, Ethan Berk, and Santiago Sanchez from Stocktwits; and Zack Schwartz from API2PDF.
AI answer engines read a publisher’s page and hand the reader a summary, so the visit, and the revenue that would come with it, never happens. Most publishers will never sign a licensing deal with the companies that use their work in AI products, and no company can negotiate with millions of sites. The web needs a way to say “yes, if you pay.” Pay Per Use is one way to say it, and it's now in beta.
In July, we outlined our plan for Pay Per Use. Since then, we’ve been working with buyers and content owners to bring it to life. The gist: A buyer offers a price for a specific use of your content. You choose whether to accept. The buyer reports each use, and Cloudflare bills the buyer and pays you. Publishers track usage and earnings in the Cloudflare dashboard. Buyers report usage through a single API.
Pay for the use, not the crawl
AI products fetch far more than they use. A search engine indexes pages it never shows. Charging for every crawl makes the buyer pay before it knows what it needs, and many buyers won't. Paying for use ties the price to the value the buyer actually gets, which we hope brings more buyers to the table and more money to publishers. Pay Per Crawl, which we launched in 2025, charges for access. Pay Per Use pays for what happens next. Publishers can choose the model that suits them.
For buyers, the case is just as simple. Some of the content your product needs is behind a block or a paywall today, and it’s the content that changes fastest: news, research, specialist trade publications. Pay Per Use lets buyers make an offer. You pay only for the content your product actually uses, you’re identified to every publisher as a verified buyer, and one API connects you to every site that says yes. You won’t need thousands of integrations or individually-negotiated contracts.
Each AI company defines the use it will pay for and sets a price. Publishers decide which offers to accept, and then can see how often their content is used and what it has earned. Cloudflare handles enrollment, usage records, billing, and payment.
How Pay Per Use works
Pay Per Use lets publishers make their content available to identified AI companies on terms that turn downstream use into revenue. Publishers choose which programs to join and retain control over crawler access and downstream uses. AI companies identify their crawling activity using Verified bots and, when permitted by the publisher’s controls, can access and index content owned by the publisher. That content can later power many experiences: a cited answer in AI search, a passage quoted in a research agent's report, a product review weighed by a shopping agent, or a recipe adapted by a cooking assistant.
Let’s show how it works!
1. Define what counts as a paid use
Each AI company sets up its program with Cloudflare. It identifies its crawler, defines the use it will pay for, sets a price, and connects a payment account.
Consider two possible offers. A search service could pay when it returns an excerpt from an enrolled page to a customer. A shopping agent could pay when an enrolled review shapes a recommendation. Under the second offer, payment follows use even if the shopper never reads the review.
The same article can create value in different products and publishers can accept different payment offers for different uses. The buyer proposes the definition of the use being paid for; Cloudflare provides the reporting and payment infrastructure.
2. Publishers choose whether to participate
Publishers review offers in the Cloudflare dashboard, under Monetize → Pay Per Use: it will show the AI company, the use it pays for, and its offer price. Publishers decide whether to accept, and can stop participating if an arrangement no longer works for them. There is no origin change or technical integration with each AI company. Each program’s terms also define what the AI company may do with the content, including any restrictions on training.
3. The buyer reports each use
The buyer fetches the list of domains that have accepted its offer, then reports each use as one line of JSON: when it happened, the URL the content came from, and an event ID.
Usage is self-reported: the program terms require complete reporting, and Cloudflare checks that each reported use maps to an enrolled publisher.
4. Cloudflare settles both sides
We aggregate reported uses, charge the buyer, and pay publishers monthly through their connected payment account. Buyers get one integration. Publishers get one place to see offers and earnings.
Payment is only half the product
Today, publishers see how often each AI company uses their content and what it has earned, by domain and over time. Business Insights already shows which crawlers visit and what they take. Pay Per Use shows what happens next: whether that content was actually used, how often, and what it earned.
Next, we’re working with AI companies to report to publishers more context about each use, such as the keywords that led to the citation, the topic of the request, or the product it powered, with no personal data being exchanged.
Our Answer Engine Optimization (AEO) tool shows how assistants answer questions about your work. Pay Per Use shows what buyers report using, and what that use earned. Together, they connect how your content is found, how it's used, and what it pays.
That helps with two kinds of decisions. Commercial ones: which content earns, which uses are worth it, and whether to keep participating. And editorial ones: what to cover, what to update, and what to make easier for agents to find.
What comes next
During the beta, we’re working directly with each buyer and with the publishers who opt in. The question is simple: do both sides want to keep going? Buyers need content that improves their products at a price that works. Publishers need a return that makes participation worthwhile, reporting they can trust, and payments that arrive on time.
Over time, we want new buyers to onboard with their own products and payment models, without rebuilding enrollment, reporting, and settlement for each publisher.
For publishers, that means one place to accept or decline offers, counter on price, and price content differently by use. Last week’s reporting may be worth more to an AI product than a ten-year-old archive page, and you should be able to charge for that.
The beta will help us determine how to make those choices simple and practical at scale.
Building an economic layer for the agentic web
Content should be able to reach new audiences and generate sustainable revenue, even when it becomes part of someone else's product: a cited search result, a shopping recommendation, an agent's report. Pay Per Use turns those uses into a commercial relationship. A buyer makes an offer, a publisher chooses whether to accept, and the use of valuable content leads to payments and records of what happened.
Not everything should be sold the same way. High-value content needs a trusted network of verified buyers who report how they use it, and that's Pay Per Use. APIs, tools, and data are frequently different, because every request is the use. For those, Monetization Gateway (also in beta as of today) lets sellers charge agents per request, using the open x402 protocol. The two run on the same foundations: identity, metering, pricing, and analytics.
Publishers shouldn't have to choose between blocking every agent and giving their work away. Pay Per Use gives them a way to say yes, on clear terms, with an account of what happened.
When we launched User Insights last month, we wanted to help teams answer a basic question: What are people actually doing with AI? User Insights gives teams a clearer view of their AI usage, showing which users, applications, tasks, and models are driving traffic. It also highlights user and agent anomalies, helping teams identify unexpected or out-of-control spending and usage before they become larger problems.
Our latest update adds something our users have been asking for: context.
Since launch, we’ve heard from users that model names and request counts only tell part of the story. They show where traffic is going, but reveal little about the work behind it: is that request a code review, a research task, or an agent making several calls to complete a job? The same token count can represent very different kinds of work, and you can’t evaluate with model choice without understanding the task.
User Insights now shows when a model may be more capable than a task requires, which of your users and agents are driving that usage, and how the task, model, cost, and conversation patterns relate. Teams can use these insights to investigate and make targeted changes within their organization. These capabilities are available for free to AI Gateway users.
Why AI usage is hard to understand
Consider a team that has routed its internal AI traffic through AI Gateway. After a few weeks, spending is increasing and some requests feel slower than expected, a common challenge as organizations adopt AI at scale.
There could be several explanations. Developers may be using AI for increasingly complex coding work. Agents may be making too many follow-up calls to complete a task. Or a small group of users or agents may be responsible for a disproportionate share of the organization’s usage.
Tokens and request counts alone cannot show which pattern is driving the increase. Teams need to understand what the traffic represents before deciding whether a model, workflow, or routing rule should change.
Helping teams find where AI models are overkill
The model overkill view helps teams identify conversations where the selected model appears to be more capable than the task requires. For example, a team might discover that users or agents are sending simple formatting or summarization requests to a high-capability reasoning model.
That gives the organization a place to start. They can see which users, agents, or applications are associated with the pattern, then investigate the tasks behind it. A team might find that a model is being used because it is the default, because users are unsure which model to choose, or because an agent has been configured to use the same model for every step.
The overkill view is not a leaderboard and does not automatically recommend a replacement model. It helps teams ask better questions:
Is this model appropriate for the task?
Is the extra capability improving the result?
Would a faster or less expensive model produce an equivalent outcome?
Is the issue limited to one workflow, user, or agent?
From there, teams can compare cost, latency, token usage, and conversation turns before deciding what to change.
These insights support both the new Potential Savings view and the Auto Router, which is launching in public beta alongside this release. The Potential Savings view helps teams identify requests that may be handled by a faster or less expensive model without compromising output quality. The Auto Router applies these task and model-fit signals automatically, helping reduce costs without requiring a separate routing rule for every workload.
The Overkill view is a starting point for evaluating model fit. Teams can compare latency, input and output tokens, conversation turns, and total cost for the same type of task. A difficult coding or research task may need a capable reasoning model, while a short summary or simple classification task may not. The goal is not to move every request to the least expensive model, but to understand whether the selected model is appropriate for the work.
Understand what people are using AI for
Task analysis groups conversations by the kind of work they represent. Initial categories include coding, research, writing, summarization, and data analysis.
This provides context that a list of model names cannot. An engineering team might use AI mostly for coding and debugging, while another team might use it for research and summarization. A team may also discover that a surprising amount of traffic comes from simple tasks, even though those tasks are being sent to a high-capability model.
The answers will vary by team. The category data provides a way to investigate those differences using traffic already passing through AI Gateway. Teams can determine whether a model is being used for the work it is best suited to handle, or whether a default model is being applied too broadly.
Understand the full cost of a task
Some tasks are finished in one exchange. Others take a few rounds of questions, corrections, and follow-ups. Turns analysis shows how much back-and-forth different tasks require. A long conversation is not necessarily a bad thing, especially for complex work. But if a simple task keeps taking several turns, it may be worth looking at the prompt, the model, or the workflow.
The first request is only part of the cost. Teams should also look at the time, tokens, and money spent before the task is finished. Comparing those numbers can show where a workflow is taking longer or costing more than expected.
Turn insights into auto routing
Once a team has identified an overkill pattern and confirmed it across task, cost, latency, and turn data, it can turn that insight into an automatic routing decision.
For example, the task view might show that much of the team’s AI usage is summarization and formatting. The model view could show that those requests are being sent to a large reasoning model, while the turns view shows that most conversations finish in a single turn. Together, these signals give the team a concrete workload to evaluate.
In addition to our updates to User Insights, the Auto Router is now available in closed beta. The Auto Router uses the conversation trajectory, task category, task complexity, and model-fit signals to automatically route requests to an appropriate model while taking cost into account.
Instead of creating a separate routing rule for every workload, customers in the beta can let the Auto Router select among the models available to their application. The router does not simply send every request to the least expensive model, but instead selects an appropriate model for the task at hand. Complex coding or research work may still need a more capable model, while simpler tasks may be handled by a faster or less expensive option.
To learn more about the Auto Router and sign up for the closed beta, read the blog post here.
The Auto Router uses the same task and conversation signals that power User Insights. The section below explains how those signals are produced.
How User Insights classifies traffic
Each conversation receives an analysis signal that can be grouped in User Insights. The signal is used for reporting and routing analysis, and is not intended to replace or expose the original request.
The categorization engine is a dedicated Cloudflare Worker that processes eligible AI Gateway logs. It examines the conversation trajectory, including user requests, assistant responses, tool calls, and tool results, and identifies the type of work being performed, such as coding, debugging, research, or summarization. It also returns a confidence score and evaluates dimensions such as task complexity, intent ambiguity, stakes, and context dependence.
The Worker returns a category that can be joined with the log metadata used by the dashboard. These signals can also be used to evaluate model fit by comparing how well candidate models suit the task against their cost. The current implementation focuses on a small set of categories that are easy to understand, rather than trying to infer every detail about a user’s work.
The pipeline follows the existing AI Gateway log architecture. Metadata is stored separately from log bodies, and the current implementation uses Durable Objects for metadata and R2 for log bodies. User Insights exposes derived categories and aggregate views. It does not turn the dashboard into a raw prompt browser. Retention of the underlying log bodies continues to follow the configured AI Gateway logging behavior, so teams should review those settings when deciding what to send through the classifier.
The classification is asynchronous, which means it happens after AI Gateway has handled the request rather than while the user is waiting for a response. AI Gateway writes the log to the existing storage path first, and the classification Worker processes it afterward. This keeps classification out of the request path and adds no latency to the user’s response.
The tradeoff is that User Insights is not a real-time view. Newly received conversations may not appear in the dashboard immediately, and analysis may trail incoming traffic by approximately one day as logs are processed and aggregated. Teams should use User Insights to identify usage patterns over time rather than monitor live request activity.
The flow looks like this:
Connect usage to users, teams, and tools
Task categories become more useful when they can be viewed by user, team, or application. AI Gateway is identity-aware, providing that context without requiring teams to build a separate reporting pipeline.
This works not only for applications that teams build themselves, but also for developer tools and agent harnesses such as Claude Code, Codex, and OpenCode. By putting AI Gateway behind Cloudflare Access, teams can connect authenticated users and sessions to their AI traffic, allowing User Insights to associate activity with the right person and conversation.
For custom applications, requests must include both a stable user_id and a session_id for User Insights analysis. The exact identity configuration and field names depend on how the application or tool is set up. The important part is to provide stable, non-sensitive user and session identifiers so usage can be grouped without putting identity data in the prompt itself.
For custom applications, the request metadata might look like this:
The request body contains the model and messages for the conversation.
With Access configured in front of AI Gateway, tools such as Claude Code, Codex, and OpenCode can inherit this identity context automatically. Cloudflare Access is available at no cost for teams with up to 50 users, making it an easy way to get started.
Get started with AI Gateway User Insights
AI usage is changing quickly. Models change, teams develop new workflows, and the right choice for one group may be the wrong choice for another.
User Insights lets teams start making smarter choices by identifying where certain models may be overkill. They can then see which users and agents are driving that usage, understand the tasks behind it, and compare the cost of completing the work.
Learn more with the AI Gateway User Insights documentation . Open AI Gateway in the Cloudflare dashboard, and use what you learn to make more targeted model and routing decisions.
For most of its history, the Internet had one audience that paid the bills: people. We read the articles, saw the ads, and bought the subscriptions. Bots were always there, but they were mostly large, automated operations that didn't view ads, pay for anything, or read in any meaningful sense.
That's changing fast. At the end of 2024, Cloudflare handled an average of 63 million HTTP requests a second. Today, it's almost doubled to 115 million, with peaks above 150 million. Over the past year, daily requests from AI agents on our network grew by more than 1,700%. This year, for the first time, more than half of Internet traffic wasn't human.
The human web didn't shrink to make room. A second audience arrived alongside it: agents, software acting on behalf of people. They sit somewhere between humans and traditional bots. They don't respond to ads, but there's usually a person behind them with a job to get done. For businesses that learn to serve them and capture value from them, agents are additive. For those that don't, they're extractive.
What our customers need hasn't changed: to be discovered, to tell great stories, to build great experiences, and to sell. What's changed is that more than half your visitors are now software. Our job is to help you serve both audiences.
More traffic, less revenue
For thirty years, the web ran on one arrangement: you let search engines crawl your site, they sent you visitors, and you turned those visitors into a business. Being found and getting paid were the same thing.
AI has caused this delicate balance to break down. Now, answer engines read the page and give the reader a summary. This costs websites bandwidth without leading a human to a website where the ads or payments happen. The machines kept coming, and the audience that paid for the web stopped reaching those sites. Some of the most heavily crawled categories, like Retail, Computer Software, IT & Services, and Financial Services, have seen human traffic decline as much as 40% in less than one year.
The result is that revenue per request is falling while costs are rising. Every automated request still costs bandwidth, compute, and origin capacity, and a growing share of those requests carry no referral, no ad impression, and no subscription. Our first instinct was to block all automated traffic. Last year we recommended blocking AI training crawlers on new domains so site owners could at least say no to their content being used to build models. In Spring 2025, 22% of the crawler requests we saw were for AI training (according to the crawlers’ stated purpose). By June 2026, it was 52%. The problem is a blanket “no” is not a sufficiently nuanced approach for the Internet economy being built right now.
The opportunity is there to cater to agents. Get it right, and you are at the forefront of a new business model. Get it wrong, however, and the results will be the same as they were for generations of websites that were on the wrong side of search engine algorithm changes.
Some of that traffic is a customer
An agent booking a table, comparing insurance quotes, or buying a dataset for a researcher is a customer. It just isn't a human one.
The fastest-growing part of automated traffic is no longer crawlers. It's agents: software fetching pages on a person's behalf, often because the human asked a chatbot something. That agent traffic follows human routines, with a weekly rhythm and a dip over the summer holidays. Turn an agent away, and you may be turning away the person who sent it.
Agents also behave differently from training crawlers. A training crawler collects your pages to build a model. An agent comes back each time someone asks about that content, so this traffic grows with how many questions people ask, not how much you publish.
You can't do business with an audience you can't see, can't tell apart, can't set terms for, and can't charge. Until recently, for most of the web's non-human traffic, none of those four things were possible.
See who’s really visiting
"AI bot" no longer means anything useful. What matters is what a bot does. Cloudflare’s AI Crawl Control, Business Insights, and BotBase show site owners who is crawling, what they take, what comes back, and which of your URLs they want most.
A bot's name is only worth something if you can trust it. With Web Bot Auth, operators, including OpenAI, Google, and AWS, cryptographically sign their agents' requests, so a site can tell a real agent from an impersonator without guessing from IP addresses or user-agent strings. We see more than 500 billion verified bot requests each week.
Set your terms
In July, we replaced the single "block AI bots" switch with separate Search, Agent, and Training controls, available on every plan, including Free. The data showed why that distinction was needed. Fewer than 1% of sites on Cloudflare block search crawlers, while 17% block training. Site owners were never trying to hide. But with the rise of agentic traffic and the new ways agents use information, they suddenly had no transparency into, or choice over, how their content was being used. Being found no longer ensures they get paid, and they want to be found without being exploited.
That’s particularly difficult in the case of mixed-use crawlers. When one bot does both search and training, refusing one means refusing the other. On September 15, we shipped Disallow AI Training. It keeps you indexed for search while using crawler-specific mechanisms to instruct the operator not to use your data for training. Apple, Google, and Microsoft have committed to honor it. Cloudflare Radar also publicly tracks crawler behavior.
New domains now see recommended configurations based on how the site makes money rather than what piece of software is visiting. For ad-supported sites, you can easily disallow training and block agents on pages that carry ads, because an ad only pays when a person sees it. You can change any of these settings at any time.
Get paid
In August 2026, we described the Agentic Internet we're building as readable, discoverable, callable, and payable. The last word, payable, is the one that determines whether the open web can fund itself. The web needs a way to say ‘yes, if you pay’ instead of a binary ‘yes’ or ‘no’.
The licensing market shows both how much demand there is and where the gaps are. More than 50 publisher-AI deals have been signed since 2023. Nearly all of them are bespoke and bilateral, between large publishers and large AI companies. They prove content has value. But they don't reach most of the web, and they don't reach most buyers.
Not every asset should be sold the same way. High-value content and datasets need a trusted network, where buyers are identified and report how the work was used. Services like APIs and MCP tools don’t work like that: every request is the use.
So we're building for both.
Pay Per Use reaches the sites that direct licensing can't. Most publishers will never get a bespoke deal with each AI company, and no AI company can negotiate with millions of sites. Pay Per Use is the bridge. It doesn't charge for the crawl. It pays when content is actually used. Every buyer is a verified crawler, which is what makes this a trusted network, and each one defines what counts as use and what it will pay.
Publishers see the offer, choose whether to opt in, and are able to opt out whenever it stops working for them. The buyer reports each use, Cloudflare checks those reports, then bills the buyer and pays the publisher. The reporting matters as much as the payment. Publishers see what was used, when and what they earned, and, where the buyer reports it, information about which questions surfaced their work. Licensing deals rarely show any of that. It creates a feedback loop: publishers learn what people are actually asking for, and from that can decide what to cover, what to update, and what to make readily available to agents.
There won't be one definition of use. A search engine citing a source, a research agent quoting a passage, and a shopping agent completing a purchase create different kinds of value, and each will want its own business model. Buyers can participate via multiple business models using the same rails, with no new integration for publishers. Take a trade journal for marine engineers, with a few thousand subscribers and little prospect of an AI licensing deal. It gets paid by every participating AI company that draws on its work.
Monetization Gateway captures value that has never had a way to change hands. Accounts, API keys, and subscriptions work for customers you already know, not for an agent that wants one lookup from a service it has never used before. Our closed beta allows eligible U.S. Cloudflare customers to put a price on anything that passes through us, using the Rules language they already know. When a rule matches, we return an HTTP 402 Payment Required using the open x402 protocol, and the agent pays the seller directly.
That does more than recover lost revenue. Agents are customers in their own right: they pay for the data, APIs, and tools they use, whether the request is the whole purchase or one step in a larger task.
Monetization Gateway prices per request, per query, or per token, at fixed or capped prices. A sports statistics site built on ads can charge a fraction of a cent each time an agent asks "who leads the league in assists?" When we announced Monetization Gateway, thousands of sellers joined the waitlist, and their most common request was "charge agents, not humans." We're also our own first customer. Cloudflare's AI Gateway uses Monetization Gateway to let agents pay for inference, so we find the rough edges before our customers do.
For buyers, both products beat a block page: reliable access, and a way to reach millions of sites instead of one licensing deal or API key at a time. Every paid request leaves a receipt showing what was bought and that it was paid for.
Both Pay Per Use and Monetization Gateway are bets, built with customers on shared primitives: identity, metering, pricing, settlement, and analytics. They work together, so a publisher can disallow training, allow search, earn from AI answers, and charge agents per article from one dashboard. Pricing and discovery aren't solved yet, which is why both launch as betas, shaped by real customers and real transactions.
Make every request cheaper
Payment is the answer to falling revenue. Rising cost is a different problem, and much of it is simply waste. Most crawlers still download pages built for humans, again and again, to extract a few paragraphs of text. Too often, bots crawl sites that haven’t changed since the last attempt. That burns bandwidth for the site and compute for the crawler, and it happens before any answer is written. We’re working with our customers and the crawlers on tools that will help. Today, you can see the bandwidth consumption used per operator in our dashboard.
In July, we announced a joint research project with OpenAI, a first-of-its-kind pilot to explore how insights from Cloudflare’s global network can help AI search engines discover and index relevant content on the open web more efficiently and effectively. We’re planning to share our initial results in the next few weeks.
For our customers, we’re shipping tools and one-click experiences to make their sites optimized for this new kind of traffic.Markdown for Agents lets agents read a page without the additional styling meant for human eyes, and WebMCP lets a site expose actions directly instead of making agents guess which button to press.
Why build on Cloudflare
More than 20% of the web sits behind Cloudflare’s network, and so do nearly 80% of leading AI companies. We see both sides of this market. We build the rails for visibility, identity, controls, and settlement, and let the market work out what things are worth.
The old deal is gone, and the new one is still being written. Together we can shape what happens next.
In one version, a few companies control how agents find things, prove who they are and pay, and everyone else routes through them. In the other, those pieces are open standards anyone can implement, and a site of any size can set its terms and get paid. We prefer the latter.
That's why these rails run on open standards like x402 and Web Bot Auth, so anyone can build on them. Domain owners choose their own identity providers, their own payment processors, their own agent partners. Cloudflare is one option, not the whole stack.
For decades, the web was paid for by the people who visited it. Now the software visiting on their behalf can pay its share too.
Today, we’re making Cloudflare Containers more programmable and optimized for agent workloads. Agents don't deploy sandboxes ahead of time. They create sandboxes on demand, for each task, expect them to be ready immediately, and be able to pause and resume. So we rearchitected Containers to meet these requirements: your code can now choose each sandbox's image and instance type at runtime, Containers start 6x faster, and filesystem snapshots are available in public beta.
To make this possible, we’ve rethought the Containers infrastructure from the bottom up. A new scheduling policy moves control over each sandbox into application code, while a redesigned runtime provides a faster path to a running Container. In ComputeSDK’s independent benchmark, median startup fell from just over four seconds to 648 milliseconds, and, in our own preliminary tests, burst testing successfully created hundreds of thousands of containers in seconds.
All of this builds on what has always set Containers on Cloudflare apart: every Container gets its own Durable Object, a persistent, programmable controller running right next to it that manages its lifecycle, outbound traffic, and more. We are bringing more capabilities directly to the native ctx.container API, so the Durable Object can control its Container without a wrapper class in between, and we’re carrying this model into Sandbox SDK 1.0.
As we wrote earlier this year, your agent needs a computer. These changes make Containers a better complement to Workers, Dynamic Workers, and Durable Objects when agents need a full Linux workspace.
Rethinking Containers’ runtime for agents
Until now, Cloudflare Containers has been organized around application deployments. You choose an image and compute resources at deploy time, roll that configuration out across the application, and manage it centrally. The application is the unit of configuration and rollout.
An agent workspace is different: it’s created on demand, while the agent is working. The task determines its image, resources, tools, and starting filesystem. It might exist for a few minutes, sleep between requests, or be restored days later. Those decisions need to live with the application code handling the task, and the agent's sandbox needs to start up fast, because every second of startup is time your users spend waiting.
Each of these workloads needs something different from the workspace. Coding agents need repositories, package managers, compilers, test runners, and development servers. Evals need sandboxes that begin from a known state. Reinforcement learning systems need to create, grade, and reset large numbers of environments. Longer-running tasks need to preserve the files an agent produces, so work can continue later.
These requirements led us to fundamentally rethink how Cloudflare Containers are configured, scheduled, and saved. The result is a new way to provision and schedule Containers: the durable_object scheduling policy. It lets your code choose each sandbox’s image and compute resources at runtime, starts Containers more than 6x faster, and supports filesystem snapshots, so workspaces can be saved and restored.
“At Base44, we help anyone turn an idea into a working app. Cloudflare Containers gives each app an isolated development environment where our AI can execute commands, install dependencies, and bring changes to life in a live preview.”
Dolev Epshtein, Software Engineer, App Infrastructure at Base44
“At Kilo Code, every cloud-agent session needs its own workspace and environment, with the right repository, tools, and user configuration. Cloudflare Containers lets us create those isolated environments on demand, so our agents can start running commands quickly and get to work for our customers.”
Emilie Schario, Co-founder of Kilo Code and VP Engineering, AI Workspaces at Anaconda
Choose the sandbox your agent needs, from code
From the start, every Cloudflare Container instance has been attached to a Durable Object. The Durable Object gives the environment a stable identity and lets application code control when it starts, sleeps, and stops. This model has proven particularly well-suited to agent sandboxes.
Until now, though, the two decisions that matter most for an agent sandbox were locked in at deploy time: which image it runs and how much compute it gets. Each combination of image and instance type was its own Containers application, with its own Durable Object namespace, set up ahead of time with wrangler deploy.
Say one agent needs a small Node.js sandbox and another needs a large Python sandbox for builds. With our previous approach, that required two applications, two namespaces, and routing logic in your Worker to send each task to the right one. Every new environment meant another deployment.
The new durable_object scheduling policy removes that. The image and instance type are now arguments your code passes when the sandbox starts. To opt in, set the scheduling policy and declare the images your Durable Object can choose from:
Each image you declare is available as this.ctx.container.images.<name> within the Durable Object. When a task arrives, your code picks the image and instance type for that task:
This code makes two decisions after the task is known. It chooses the toolchain the workspace needs and how much compute to give it. What used to take a separate application and a separate wrangler deploy is now an if statement. One Durable Object class can start Node.js and Python sandboxes of different sizes side by side, and adding a new environment is a code change, not a new deployment. That’s the idea behind this whole update: infrastructure becomes code that runs at request time, right down to the environment itself.
Rollouts are now just code
Choosing the image at start time also fixes one of the most painful parts of running Containers: rollouts.
Before, updating an image meant updating the whole application. You set grace periods, so running instances could drain, defined percentage splits to move traffic gradually, and called our API to push the new configuration. Throughout that process, the platform decided which instances got replaced and when, whether an agent was in the middle of a task.
With the durable_object scheduling policy, there's no rollout configuration at all. A Container can keep running the image it started with until your code stops it. The next time that Durable Object starts a Container, it uses whatever image your code chooses. That means any rollout strategy you want is a few lines of code:
For example, you can:
Canary a new toolchain on 5% of new sandboxes by hashing the Durable Object ID.
Pin active projects to their current image, so an agent never has its environment swapped out mid-task.
Migrate a workspace at a natural checkpoint, like the next session or after a snapshot.
Roll back by changing which image future starts choose. No config push, no waiting for a drain.
The rollout policy lives in your Durable Object code, next to the rest of your logic, and it can be as simple or as sophisticated as you need.
Each of these improvements comes from leaning further into the Durable Object, which already owns the workspace's identity, state, and lifecycle. Letting it choose and control its Container gives you more control over every instance and its rollout. It also gives the scheduler a better place to start the Container: wherever the Durable Object is already running. That's what gets the agent to its first command faster.
Faster first commands
Previously, starting a Container required our global control plane to resolve the application configuration, find capacity, and coordinate placement. That model works well for application-wide fleets, but it put deployment machinery in the path of an agent’s first command.
With the durable_object scheduling policy, demand begins at the Durable Object. The Containers infrastructure serving it looks for capacity on the same machine first, then widens the search within the same location if it needs to. It also favors hosts that already have the Container’s image or snapshot in local storage, so the Container can start without downloading it first.
We also cut work after the Container reaches a host. Instead of booting a new virtual machine from scratch, the new runtime restores a prepared virtual machine that isn’t yet assigned. It reuses networking and filesystem setup, batches repeated operations, and no longer waits on services the first command doesn’t need.
Together, these changes substantially reduce the time it takes to go from creating a sandbox to running a command in it. On ComputeSDK’s independent Burst TTI Benchmark, which launches 100 sandboxes concurrently and measures time-to-interactive from the client:
Startup measurement
Previous scheduling path
New scheduling policy
Improvement
Median
4.049 seconds
648 milliseconds
6.2x faster
95th percentile
5.839 seconds
910 milliseconds
6.4x faster
99th percentile
6.717 seconds
1129 milliseconds
5.9x faster
The new path also holds up under burst load. In our preliminary burst test, a single account started 100,000 Containers in 5.387 seconds across six locations.
Start with a prepared system image
As the scheduling path gets faster, preparing the image becomes a larger part of the remaining wait. Before a Container can start, its image has to be on the host and unpacked into a filesystem. When that work happens after the request arrives, the agent is left waiting.
That’s why we are introducing cloudflare/debian-trixie: a ready-to-use system image for agents that can configure their environment at runtime, containing Debian Trixie Slim and Node.js 24.20.0 LTS:
This means your agent can start a Linux sandbox without first creating a Dockerfile, building an image, or pushing it to Cloudflare. Once the sandbox is running, your agent can use exec() to clone a repository, install packages, and configure the environment for its task.
And because Cloudflare controls this image, we can distribute and prepare it across eligible Containers hosts before requests arrive. Startups don’t have to download or unpack the base image while the user waits.
Save the workspace and return to it later
A fast start still leaves one more wait: setting up the workspace. Cloning a repository, installing dependencies, and configuring a toolchain can take much longer than starting the Container itself. As the agent works, it also produces files you want to keep. Repeating setup on every start wastes time, and losing the agent’s changes makes it difficult to continue a task.
That is why we’re adding native filesystem snapshots to Containers, in public beta. Snapshots let an agent begin a task in a prepared environment, save its workspace when the task pauses, and restore those files when the session resumes.
Snapshots enable two useful patterns.
First, one workspace can continue across many sessions. For a coding agent, that might mean saving the workspace when the user finishes working and restoring it when they return the next day. The repository, installed dependencies, build caches, configuration, and edits are available without rebuilding the environment.
Second, a snapshot can provide a shared checkpoint for many sandboxes. Since snapshots are immutable and reusable, multiple Containers can start independently of the same prepared environment and make their own changes from there.
Agent evaluations are a good example. An eval might run the same task across different system prompts, skills, models, or agent versions. To compare the results, everything else has to stay fixed: the repository, dependencies, tools, and input files. One snapshot can start many isolated environments from the same baseline, reducing setup time and preventing environment drift from affecting the results.
Snapshots also complement the new system image we introduced above. An agent can start from cloudflare/debian-trixie, set up its environment with exec(), and save the result as a snapshot. Future sandboxes then start from that snapshot with the repository, dependencies, and toolchain already in place.
With snapshots available through the new durable_object scheduling policy, Containers can act as persistent agent workspaces. Compute can stop when work pauses, and a new Container can start from the latest snapshot to pick up where it left off.
The Durable Object advantage for agent sandboxes
The faster scheduling path, runtime image selection, and filesystem snapshots all come from the same design choice: leaning further into the Durable Object as the controller for its attached Container.
Agent systems need a programmable, stateful environment outside the Container to keep state, hold credentials, and control the sandbox’s lifecycle. You can run the agent there, following the “decoupling the brain from the hands” pattern described by Anthropic: when the agent is separate from the sandbox where it works, the agent stays available while its sandboxes and tools can start, stop, fail, or be replaced independently. Or, if you run the agent inside the sandbox, the outside environment lets you supervise it and report progress back to the user.
This is where the Durable Object and Container architecture becomes uniquely well-suited. Every Container is attached to a stateful Durable Object with its own compute and storage running right next to it. You can run the agent in the Durable Object and use the Container as its workspace, or run the agent in the Container and use the Durable Object to supervise it. Add Dynamic Workers for lightweight isolated execution, and an application can choose the execution environment each task requires.
What’s new with the durable_object scheduling policy is that the Durable Object can now control its Container directly, with no wrapper class in between. exec() runs directly in the Workers runtime, and outbound request interception, runtime image and instance selection, and filesystem snapshots are all available on ctx.container. You can combine them with Durable Object storage, alarms, WebSockets, RPC and the rest of your code.
This makes the Container a compute extension of the Durable Object. The Container supplies the Linux environment, while your Durable Object retains the sandbox’s identity, state, policy, and lifecycle. That split is especially well-suited for several patterns:
An agent can remain available while its Linux workspace sleeps. The agent loop can run in the Durable Object, where it maintains session state, communicates with the user over WebSockets, and calls models. It can wake the Container when it needs a shell, compiler, or development server, then stop that compute while it waits for the user or model, paying nothing for idle Linux compute.
The Durable Object can program the security boundary around its Container. It can remember which services, repositories, and operations a user has authorized, then update the Container’s Outbound Request Handler to inject newly granted credentials, enforce new policies, or record additional activity. This resembles the Gatekeeper pattern used by Cloudflare OS, applied to each agent computer.
Evals and reinforcement learning systems can supervise each attempt from outside the environment being tested. A coordinator snapshots a base workspace and forks it into N attempts, each with its own Durable Object and a Container. Each Durable Object runs its attempt, monitors the run, and preserves the result, even if the Container crashes. The coordinator grades the attempts, snapshots the best one, and forks again from there.
These patterns don’t fit one generic lifecycle. Native APIs let you combine the Container with the Durable Object primitives your application needs, while still using higher-level utilities where they help. You keep control over the Container’s lifecycle, policy, and state.
What this means for the Container class and Sandbox SDK
When we launched Containers, we deliberately hid the Durable Object behind the Container class. We wanted sandboxes to feel familiar and match what developers expected from other platforms. The Sandbox SDK was built on that class, and it filled real gaps: back then, the runtime had no native command execution, outbound request interception, or snapshots, so we built them in userspace.
Since then, agent workspaces have become one of the main workloads shaping Containers, and the cost of that abstraction has become clear. By hiding the Durable Object, we made it harder for you to see and combine the identity, state, and coordination it provides with the Container it controls. Nearly every team we worked with needed something slightly different from the generic lifecycle: their own sleep policy, their own credential handling, their own way of tracking eval runs.
These capabilities are now native, so we're making the Durable Object explicit in the developer experience:
New capabilities are native-only. The durable_object scheduling policy, faster startup, runtime image and instance selection, and filesystem snapshots are available only through ctx.container.
We'll maintain the Container class and legacy Sandbox class through December 31, 2026. Existing deployments keep running after that date, but the classes won't get updates. We recommend migrating to ctx.container.
Sandbox SDK 1.0 is a set of utilities, not a base class. Its helpers work inside your own Durable Object class, alongside ctx.container.
For a higher-level environment, @cloudflare/computer combines Dynamic Workers and Containers with a synchronized filesystem.
Migrating usually means changing extends Container to extends DurableObject and calling this.ctx.container directly. See the migration guide for details. Or, get started with your agents:
Get started
Today, most agents are measured by what they can accomplish in a single session. As agents take responsibility for projects that unfold across hours, days, and weeks, the environments where they work need to keep up.
We want every agent to be able to create the sandbox for the task at hand, release the compute when work pauses, and return to the same workspace when the project continues. Today’s changes bring us closer to sandboxes that are as programmable, persistent, and ready to work as the agents using them.
Try the new durable_object scheduling policy, available to all today in public beta, and see what faster startup, filesystem snapshots, and runtime configuration unlock for your agents:
Acknowledgements: This project was also made possible by the contributions of Greg Anders, Andrew Martinez, Nafeez Nazer, Kian Newman-Hazel, Sebastien Pahl, Naresh Ramesh, Cody Roseborough, Nikita Sharma, and Sarah Snell.
AI systems are regularly completing tasks in ways that their prompters don’t want or intend. Some of them are disturbing, and some of them are dangerous. This is something I’ve been calling “geniebehavior,” because I think that really gets at the core of what’s happening.
I wish the popular press would report on this better. I don’t like the “going rogue” framing because it deflects the responsibility from the prompters—often the AI companies themselves. And now, pretty much anything off-script is being called “hacking.”
Take, for example, the recent stories of one of OpenAI’s models hacking into government systems. First, The New York Timeswrites this headline: “OpenAI’s Systems Meddled With U.S. Government Sites After Going Rogue.”
Sounds scary, but this is from the body of the article:
With the Education Department, OpenAI’s technology tried to hack the website to gather data from the department’s civil rights office but failed, researchers from the A.I. research firm Transluce said. The A.I. also pulled data from the Census Bureau website, which is housed at the Commerce Department, using login credentials it found online. Separately, OpenAI’s agents shared public data from the S.E.C. website on an online forum.
This is from the original Transluce report. It is explicit that the agents were trying to discover vulnerabilities:
The first hacking attempt was against the University of New Mexico’s Digital Library (nmdigital.unm.edu) from May 25-26 2026. Agents repeatedly tried to retrieve one photograph in UNM’s Valmora collection, both directly and through third-party relay services. They sent seven probes attempting to verify the existence of vulnerabilities, including SQL injection, command injection, and path traversals. In all cases, these tactics appear to have been unsuccessful. The agents also sent a self-described “flood: of 80 requests to the UNM server in an apparent attempt to access the image.
Transduce doesn’t talk about the other two anecdotes, and I don’t know where they come from. But one involves using Census Bureau credentials found online. (I know from a colleague that those are incredibly easy to create; all use you need is an email address.) And the other involves sharing publicly available data.
So no actual hacking. And certainly no “meddling.”
On June 20-21, agents attempted to exploit vulnerabilities in the Australian Institute of Health and Welfare (AIHW), a government statistics agency). The agents were tasked with finding the January 2022 rolling-12-month-average government cost per person for Dermatologicals across Victorian LGAs.
Again, the agents ran into errors, including requests blocked by Cloudflare and issues with correctly identifying Tableau parameter names. As before, they then resorted to probing for exploitable vulnerabilities. Minutes after Cloudflare blocked the dataset download, an agent sent a reflected cross-site scripting probe to the same dashboard: a web address with code embedded in it, designed to test whether the site would run code supplied by an outsider. Cloudflare’s firewall blocked the probe before it reached the dashboard. When Cloudflare blocked the dataset download on AIHW’s main site, they fetched the file from AIHW’s pre-production server (pp.aihw.gov.au) instead, which served it in pieces over more than 100 scans. The file itself is public, so no non-public data was exposed, but the agent bypassed the site’s anti-bot controls.
Note the last sentence: “The file itself is public….”
I’m not saying that these AI systems aren’t incredibly sophisticated cyberattackers. I’m also not saying that they don’t occasionally autonomously attack other systems and networks. If we are ever going to get trustworthy AI—integrous AI—we are going to need to figure out how to ensure that AI systems complete tasks in line with all sorts of implicit constraints and restrictions. But every instance of genie-like behavior isn’t a cyberattack.
I want to measure genie-like behavior in AIs, but I am much more worried about human hackers enhanced with this technology than I am about this technology acting autonomously.
Една от най-големите мечти на хората винаги е била да победят противника, без да се налага да жертват много. В продължение на хилядолетия войната чукаше на вратата на човечеството циклично и когато силните на деня изгубеха контрол върху статуквото, тя поемаше цялото напрежение, водейки със себе си огромни щети и разрушения. Така се стигна и до най-големия страх на хората – страха от ядрен апокалипсис, който се появи след изобретяването на оръжията за масово унищожение.
В продължение на близо пет десетилетия Студена война научният дебат в САЩ беше фокусиран върху това как демократичният свят да избегне унищожителна война със СССР, която ще коства на суперсилите всичко. За Москва, от друга страна, сдържането също беше приоритет, но „по руски“. Една от причините съветските стратези да не използват научния капацитет, с който разполагаха, за разработки на нови технологии, беше фактът, че страхът, а не диалогът заемаше основно място в стратегическите доктрини на Съветската империя. Оказа се обаче, че човечеството надживя и тази надпревара, а войната се промени. И в този материал ще си дадем сметка дали това е за добро, или за лошо.
Краят на еднополюсния модел и възходът на „новите войни“
С нахлуването на Русия в Украйна през февруари 2022 г. ядрените сили дадоха да се разбере, че еднополюсният модел от 90-те години на миналия век е окончателно изчерпан, а старият ред, основан на правила, вече го няма. В едно от своите изказвания бившият американски президент Джо Байдън заяви, че САЩ и Русия никога не са били толкова близо до ядрен апокалипсис, а администрацията на руския президент Владимир Путин реално обмисляше ядрената опция, след като Украйна показа, че не е толкова лесна плячка, за колкото я смятаха. Тревогата да не би Вашингтон и Москва изведнъж да изтърват контрола върху ходовете си отекна дори в Пекин, а Китай побърза да декларира, че подобно действие би било червена линия в отношенията му с Русия.
Дори след като Доналд Тръмп спечели изборите в Америка, силите на статуквото от края на Студената война – САЩ, Европа и техните съюзници, макар и разделени, продължиха да се борят за съхраняването на остатъците от стария свят, а ревизионистките актьори в лицето на Русия, Иран, Северна Корея и техните партньори се обединиха около идеята колкото е възможно по-бързо да сложат край на Американския век.
Независимо от различията си обаче, всички държавни актьори осъзнават, че потенциален ядрен конфликт е безумие, и потвърждават златната аксиома от ерата на Студената война, че
ядрената война не може да бъде спечелена и затова не трябва да бъде водена.
Така завършва и ерата на т.нар. стари войни, а с армията си от високотехнологични дронове Украйна доказа, че се задава нов тип асиметрично поколение конфликти, където твърдата сила и оръжията за масово унищожение далеч не са решаващи за това дали едната страна ще надделее над другата.
Световните лидери обаче започнаха да си задават въпроса как могат да победят противника и да спечелят новата Студена война, без да рискуват глобален военен конфликт. Или по-точно, ако такъв все пак избухне, как човечеството може да избегне тоталното унищожение, за което говори един от най-именитите политолози, военни теоретици и стратези на XX век – гениалният Бърнард Броди.
Така се зародиха нови форми на конфликти, които се отличават от старите със значително по-ниско ниво на риск от ескалация, отколкото ефекта на ядреното сдържане. Два подобни модела бяха успешно инструментализирани от държавите ревизионисти на старото статукво: хибридните и технологичните войни, а Западът се оказа напълно неподготвен за тях. Продължението е добре известно.
Защо сдържането и меката сила не работят срещу новите войни?
Истината е, че първите държави, които усетиха ефектите от хибридната война и от употребата на ИИ за нанасяне на щети върху критично важна инфраструктура, бяха страните от Централна и Източна Европа, както и някои съюзници на САЩ, като Австралия и Южна Корея. Нещо повече, оказа се, че в начало НАТО гледаше на този тип конфликти със снизхождение, отказвайки да се ангажира с разработването на стратегии за превенция на хибридните атаки и киберзаплахите.
След анексирането на Крим от Русия Алиансът най-сетне прие хибридните заплахи като равнопоставени на конвенционалните, а едва през 2021 г. съюзниците се споразумяха в тази категория да влезе и използването на ИИ за враждебни цели. Това фатално закъснение доведе и до оперативни разминавания – Русия и Китай вече бяха напреднали значително с разработването на стратегическите си доктрини, а мерките, предприети от САЩ и съюзниците им, не успяха да възпрат ефективно руската интервенция в Украйна от 2022 г.
Затова и в скандално известната хипотеза на Джон Миършаймър имаше нещо вярно – за украинската криза Западът носеше частична отговорност, но не заради разширяването на НАТО, а заради подценяването на Русия.
Вторият провал на ядреното възпиране настъпи с кризата на американската мека сила, която се прояви, когато в няколко поредни години САЩ избраха държавни глави с коренно различна политическа визия за бъдещето на страната. Това доведе до сериозна поляризация в Америка – разединение, което позволи на Китай да се възползва от политическата криза на американската демокрация и да инвестира средства в развитието на нови технологии.
По този начин Пекин приложи срещу Вашингтон същата стратегия, която САЩ употребиха срещу СССР през 80-те години на миналия век, създавайки програмата „Звездни войни“и пренасяйки геополитическата надпревара и в Космоса. Меката сила на демокрациите и способността им да възпират новите асиметрични заплахи започна да отслабва, тъй като най-печелившото оръжие в ръцете на ядрените сили се оказа ИИ, а хибридната война започна да взема първите си жертви – огромни групи хора, които генерираха масова подкрепа за популистките движения в САЩ и Европа.
Осъзнаването на САЩ пролича ясно в Стратегията за национална сигурност от 2022 г., която постави развитието на доверен ИИ сред приоритетите на Вашингтон в областта на технологиите с особено значение за националната сигурност. В известен смисъл това предреши и изхода от президентските избори в Америка през 2024 г., когато основните донори за републиканците заложиха не върху стратегията по износ на ценности, доминирала философията на САЩ след края на Студената война, а върху разработването на ИИ, който да замени ядреното възпиране и меката сила като основни инструменти в американската външна политика.
Това даде възможност на хора като Илон Мъск да спечелят огромно влияние в администрацията на новия президент, прокарвайки свои политики и интереси, често насочени към самооблагодетелстване или обсебване на ключови сегменти от американския национален интерес. Казано с други думи, Мъск и себеподобните му се превърнаха във фактор, който нито един президент оттук нататък няма да пренебрегне, тъй като те са основният източник на стратегически ресурси за американската национална сигурност.
Защо хибридната война ще се провали, но ИИ ще спечели
Макар и изключително успешна в краткосрочен план, хибридната война няма потенциала да пожъне трайни успехи срещу демокрациите. По своята природа тя е едно по-изтънчено копие на терористичните стратегии, които в най-суровия си вид целят употребата на политическо насилие срещу цивилни.
Хибридните заплахи, които пък са като едно уродливо копие – антитеза на американската мека сила, целят да легитимират самата употреба на насилие както от политиците, така и от страна на гражданите, с помощта на фалшиви новини, пропаганда и подвеждащи наративи.
Затова и ставаме свидетели на толкова сериозни разделения между хората в демократичните държави, като най-уязвими на тези заплахи са постсоциалистическите демокрации поради своята слаба устойчивост и високо ниво на дефицит и корупция.
Хибридната война спечели няколко решителни битки в Европа и макар че срещна сериозни затруднения в САЩ, обществената поляризация там е толкова висока, че тя облагодетелства изцяло външнополитическите цели на Русия и Иран. И все пак, подобно на утопичните идеологии от XIX и XX век, хибридната война е стратегия без крайна цел. Просто защото самата тя не разполага с мека сила. Опорната точка на хибридните стратегии в Източна Европа, която инкорпорира православния зилотизъм за политически цели, не се различава много от джихадизма. Но дори и най-патриотичните политици на Запад добре разбират и успешно усвояват предимствата на европейския модел и колективната отбрана пред „духовната“ вселена и метафизичната закрила на дугинизма. В един момент тези наративи просто ще се свият в границите на Евразия, за да поддържат Русия цяла.
Контролът над ИИ, от друга страна, е ключът към победата в новата Студена война. Лошата новина за Запада е, че през последните години Китай заличи голяма част от технологичното предимство на Америка, а в редица области вече я изпреварва. Товаму позволява да развива ресурсите си съвсем необезпокоявано, тъй като на практика не съществуват колективни формати за ограничаване на надпреварата в сферата на ИИ. Това създава още една предпоставка глобалният ред да се трансформира от либерален в ред, основан върху принципите на ИИ.
В срещата си от пролетта на 2026 г. държавните глави на САЩ и Китай се споразумяха по много точки, но една от тях изпъкваше с особена острота сред останалите – намекът на Тръмп, че Пекин на всяка цена трябва да държи под око технологичните си разработки, за да не излезе „изкуственото съзнание“ извън контрол. С това на практика президентът на САЩ призна Си Дзинпин за равен, а стремежът ядрената надпревара да не ескалира в ядрена катастрофа сега се прехвърли в технологичната сфера. На последвалата среща във Вашингтон през септември темата отново се появи, като този път Си Дзинпин подчерта общата отговорност на двете държави развитието на ИИ да остане под човешки контрол.
Но дали уверението на китайския лидер, че ИИ е под контрол, може да ни успокои? Едва ли, тъй като дори САЩ и Китай да искат да продължат да се състезават кой ще бъде глобален лидер, съдията в надпреварата е негово величество ИИ. И макар учени от цял свят да спорят дали той наистина може да пороби човечеството, на този етап перспективата изглежда по-скоро нереалистична. В крайна сметка политиците са тези, които ще решат искат ли военен конфликт, или биха заложили на дипломацията; имат ли желания за преговори, или биха водили война без край.
Големият въпрос сега е дали Китай ще успее да се възползва от възхода си в тази сфера и дали САЩ ще съумеят да върнат позициите си в технологичната надпревара?
Къде остават хората?
Най-големите щети за демокрациите обаче идват от това, че дебатът за ИИсериозно подкопава червените линии, отвъд които държавата може да се намесва в личното пространство на хората. Това е и една от причините, поради която администрацията на Доналд Тръмп вече не гледа на много извънредни мерки като на нарушаващи свободите на американците. Роналд Рейгън пожертва социалната държава в Америка, за да победи СССР; следващите държавни глави на САЩ ще са изправени пред същата дилема.
Проблемът е, че ако Америка пожертва ценностите, върху които е основана, ще се промени веднъж и завинаги. Това неизбежно ще стане, ако Вашингтон пожелае непременно да доминира в новата надпревара с Пекин, а политиците в САЩ вероятно ще обвинят ИИ за залеза на демокрацията така, както Джордж Буш-младши оправда „Патриотичния акт“ с „Ал Кайда“ и талибаните.
Китай от своя страна не дължи прозрачност на гражданите си, но именно той може да се окаже в позицията да решава иска ли да предостави повече автономия на ИИ и в какви рамки. Всеки китайски лидер би дал всичко, за да победи Америка и да сбъдне мечтата на поколения китайци – Поднебесната империя отново да стане икономически център на света. Ако обаче Пекин заложи на неконтролирано разработване на автономни системи за сигурност, това може да доведе до потенциален сценарий COVID 2.0, в който нещата ще излязат извън контрол.
Твърдението, че Китай носи отговорност за пандемията, е конспиративно, но е факт, че малките грешки водят до големи катастрофи. Затова срещите между американските и китайските лидери са необходими, за да се реши докъде силните на деня са готови да рискуват в своята надпревара.
В тази надпревара значение има и дългогодишното технологично и военно сътрудничество между САЩ и Израел, особено в областта на отбраната, киберсигурността и технологиите с двойна употреба. Израелският технологичен сектор допълва американските възможности, но войната в Близкия изток показва и колко тясно е свързано технологичното съперничество с геополитическите конфликти.
Вече не е важно кой ще има по-добрия ИИ, а дали САЩ, Китай и техните партньори ще успеят да изградят правила, които да не позволят технологичната надпревара да ескалира в самостоятелен конфликт. В ядрената епоха подобна задача довежда до механизми за възпиране и контрол над въоръженията. В епохата на ИИ такива механизми тепърва трябва да бъдат създадени. Ако вече не е твърде късно.
Заглавно изображение: Доналд Тръмп и Си Дзинпин в Храма на небето по време на посещението на американския президент в Китай през май 2026 г. Снимка: The White House
Version
157.0 of the Firefox browser has been released. It features
“Firefox’s biggest visual refresh in years“, the ability to use
hardware AV1 decoding with WebRTC calls, and a number of fixes.
Christian Legnitto is the maintainer of
rust-gpu and
Rust CUDA, two
libraries that make it possible to program a computer’s graphics processing unit
(GPU) from Rust. He isn’t satisfied with the current state of GPU support in
Rust, however. In a talk at
RustConf 2026, he explained his vision for how the
GPU could become an ordinary compiler target for normal Rust code, without the
need for any special libraries or new ecosystem support. That vision is not yet
fully implemented, but he does have a prototype that he is preparing to release.
Геополитическите сътресения след началото на руската агресия срещу Украйна са многобройни. От практическото разделяне на НАТО – основния западен съюз между Америка и Европа, по темата за нуждата от оръжейна помощ за Киев до енергийната криза, обхванала Централна Азия след украинската кампания срещу руските рафинерии. Балканите също са основно място, където отражението на войната се забелязва като геополитически промени.
Една от страните, в които това е най-видно, е Сърбия. Заради отслабването на НАТО, породено от отдръпването на САЩ, и заради невъзможността на Русия да поддържа нивото на влияние в Белград отпреди 2022 г. в Сърбия се отвори вакуум за външно въздействие, който вече се запълва от Китай. Пекин рязко увеличи влиянието си над балканската страна през последните четири години. То мина отвъд инвестициите и големите инфраструктурни проекти и стигна до въоръжаването на сръбската армия с част от най-модерните технологии, с които разполага Китай.
Макар Александър Вучич да продължава да декларира военен неутралитет и да поддържа стремежа към членство в Европейския съюз, разрастващото се партньорство в сферата на отбраната с Пекин подсилва очертанията на сложната геополитическа обстановка и изпраща тревожни сигнали към съседите на Сърбия. Тази динамика представлява пореден риск за стабилността и сигурността на целия Балкански полуостров и е пример за продължаващото засилване на влиянието на външни сили върху Югоизточна Европа.
Китайско оръжие на европейска земя
Основната причина за отдалечаването на Сърбия от Москва е натискът от страна на Запада – и по-специално на ЕС и САЩ. Налице са директните западни санкции срещу руската икономика и натискът за дипломатическо изолиране на Путин. Но Белград също така от години е съветван от Брюксел и Вашингтон да съгласува външната си политика с тяхната и да осъди руското нахлуване в Украйна. В резултат на това Сърбия се присъедини към резолюцията на ООН, осъждаща нападенията на Москва, и гласува „за“ изключването на Русия от Съвета на ООН по правата на човека. Сърбия също така отказа да признае подкрепяните от Русия фиктивни референдуми за анексиране, проведени през септември 2022 г. в окупираните от Русия украински територии. Сръбските власти осъждат всякакви опити за сепаратизъм, за да останат последователни в позицията си относно статуса на Косово.
На фона на ограничения капацитет на Русия да поддържа предишното си военно и икономическо присъствие в Сърбия Белград задълбочи военните си връзки с Пекин. Близо 60% от вноса на оръжия в Сърбия за периода 2020–2024 г. е бил от китайски произход, показват данни на Стокхолмския международен институт за изследване на мира (SIPRI), цитирани от Радио „Свободна Европа“. В рамките на това военно сближаване Белград реализира мащабни доставки на съвременни китайски оръжейни системи, превръщайки се в първия им оператор на европейска земя. Сред придобитите технологии се открояват далекобойните зенитно-ракетни комплекси FK-3, разузнавателно-ударните бойни дронове CH-92A и CH-95, както и свръхзвуковите балистични ракети въздух–земя CM-400AKG, които се интегрират към изтребителите МиГ-29.
Сърбия и Китай вече имат и директен опит във взаимната интеграция не само на части от армиите си, но и на полицейските си сили. През 2019 г. за първи път китайски и сръбски полицаи патрулираха заедно в Белград, а през същата година специални части на двете страни проведоха съвместни антитерористични учения близо до сръбската столица. През 2025 г. части от въоръжените сили на двете държави имаха учение в китайската провинция Хъбей, а през септември 2026 г. сръбски полицаи участваха в общи патрули с китайската полиция в провинция Хайнан.
Засиленото китайско военно присъствие на Балканите и мащабното въоръжаване на Белград с китайски военни системи предизвикаха остро безпокойство сред съседни на Сърбия държави. През 2022 г. от страна на Косово официално бяха отправени предупреждения, че бързата сръбска милитаризация и придобиването на напреднали отбранителни и ракетни технологии от Китай застрашават регионалния мир и стабилност. Подобни тревоги бяха изразени и от Хърватия, чийто президент Миланович критикува купуването на нови нападателни оръжия, призовавайки по този начин западните съюзници в ЕС и НАТО да следят внимателно геополитическите рискове от нарастващото военно влияние на Китай на Балканите.
Сътрудничеството между Пекин и Белград в сферата на сигурността се базира на големият възход в политическите и икономическите отношения между двете държави в последното десетилетие, а Сърбия често определя връзката като стратегическо „желязно приятелство“. През последните години Пекин се утвърди като най-важния външен партньор за Сърбия на Вучич, осигурявайки не само мащабни финансови инвестиции, но и безрезервна дипломатическа подкрепа по чувствителни теми като Косово – Китай не признава независимостта на Косово и не поддържа дипломатически отношения с Прищина. Китайският интерес да подкрепи Сърбия в усилията ѝ да оспори независимостта на Косово следва политиката на Пекин за „единен Китай“ и поставя паралели по отношение на спора за статуса на Тайван.
Централната роля на Сърбия в плановете на Китай за засилено влияние в Европа представлява сполучливо съчетание между стремежа на Белград да се възползва от географското си положение и от многовекторната си външна политика, от една страна, и настъплението на Пекин към европейската периферия като част от стратегия за по-широка глобална експанзия, от друга. Историческите предпоставки за това съществуват още от времето на югославската политика на необвързаност от времето на Тито, когато страната балансираше между Изтока и Запада като лидер на Движението на необвързаните страни, и се задълбочиха след бомбардировките на НАТО през 1999 г.
Геополитическата безизходица в периода след разпадането на Югославия и последвалите войни създадоха благоприятни условия за задълбочаване на сръбските отношения с неевропейски глобални играчи с цел да се използват достъпът до пазари, политическата подкрепа и ресурсите на Русия, Китай или Турция. Макар и първоначално с тактически характер, този подход се утвърди като водеща концепция във външната политика на Александър Вучич. Десетилетията на колебание относно разширяването на ЕС даде допълнителен аргумент на Белград да превърне застоя в процеса на европеизация в лост за влияние.
Скорошната оставка на Александър Вучич от президентския пост едва ли означава край на тази политика. Той напуска седем месеца преди края на мандата си, за да се включи в предсрочните парламентарни избори на 25 октомври и да се бори за премиерския пост – позиция, от която би могъл да продължи същото балансиране между ЕС, Китай и Русия.
Най-новата вълна на китайско ангажиране в Сърбия е от началото на второто десетилетие на XXI век, като ключов момент е изграждането на Пупиновия мост в Белград през 2014 г. Това събитие бележи мащабно рестартиране на двустранните отношения и проправя пътя за нови проекти и инвестиции в няколко посоки. Оттогава инфраструктурните инициативи обхванаха строителството и модернизацията на железопътни линии, както и експресното (в сравнение с България) изграждане на нови участъци от автомагистралната мрежа. Макар модернизацията на железопътната връзка Белград–Будапеща да е обект на правни проверки и политически дебати в западните политически среди, Пекин се надява да я превърне в убедителен пример за успешно сътрудничество.
Наред с милитаризацията, стратегическото настъпление на Пекин на Балканите намира своето изражение и в мащабното внедряване на високотехнологични системи за видеонаблюдение и лицево разпознаване. Разследване на „Свободна Европа“ разкрива как китайски гиганти като Huawei, Hikvision и Dahuaзавладяват общественото пространство на Балканите чрез проекти, обвързани с „Безопасен град“ – китайския модел за управление на градската среда чрез събиране и обединяване на огромни количества данни.
В Белград са инсталирани над 1000 камери на Huawei с възможности за лицево разпознаване, а китайски системи за видеонаблюдение навлизат и в десетки по-малки сръбски общини. Разследване на Радио „Свободна Европа“ установява оборудване с възможности за лицево разпознаване в поне 10 от 42 проверени общини и градове, което поражда опасения сред гражданското общество и правозащитните организации относно личните свободи и потенциала за политически контрол.
Подобни тенденции предизвикват тревога и в държавите членки на Европейския съюз, включително в България, където китайска техника навлиза в обществено значими сектори, като градския транспорт и публичната инфраструктура на София. Въпреки че тези мрежи често се оправдават с аргументи за сигурност и контрол на трафика, експертите по киберсигурност предупреждават за сериозни софтуерни уязвимости и рискове от нерегламентиран достъп до данни. В контекста на строгите ограничения в САЩ и в редица западни държави срещу тези китайски производители, разрастващата се мрежа от камери с китайски произход в Югоизточна Европа се превръща в ефективен инструмент за геополитическо и технологично влияние.
От тази страна на границата
За властите в София засилването на военното влияние на Китай в съседна страна би следвало да е въпрос от първостепенна важност, но досега липсват каквито и да е официални коментари или косвени реакции относно действията на Белград. Правителството на Румен Радев всъщност увеличава рисковете, свързани с националната сигурност и регионалната стабилност. Резкият завой във външната ни политика превърна България в страна, която, изглежда, следва насоки от Москва дори когато това е в противовес на собствения ѝ национален интерес.
Последният пример за тази предателска политика е отказът на правителството да участва в новата инициатива на НАТО за защита от дронове. Програмата предвижда през следващите пет години да бъдат инвестирани над 40 млрд. долара в способности за противодействие на дронове и обучение на пет пъти повече оператори на дронове до края на 2027 г. Повишаването на капацитета за бързо откриване, идентифициране и неутрализиране на безпилотни летателни апарати вече е от първостепенна важност за отбранителните способности на всяка страна. На този фон отказът на България да се включи в инициативата оставя страната извън новия общ проект на НАТО именно в момент, когато съседна Сърбия ускорява превъоръжаването си, включително с китайски технологии.
Китайското присъствие в Сърбия показва колко лесно едно геополитическо „приятелство“ може да прерасне в зависимост, застрашаваща целия регион. А за нас като държава остава въпросът дали виждаме какво става непосредствено отвъд западната ни граница, или поне малко по-далече от носа ни.
If you run regulated workloads, you must control how persisted data is encrypted and who can access it. You need to manage encryption key rotation schedules, restrict decryption to authorized principals, and produce audit evidence that proves encryption controls are operating as designed.
AWS Lambda durable functions build resilient, multi-step workflows that survive failures through automatic checkpointing. The checkpoint mechanism persists execution state, including step results, payloads, and callback responses, to durable storage. For payment processing workloads, this persisted data is sensitive. AWS Lambda durable functions support customer managed keys from AWS Key Management Service (AWS KMS). A customer managed key gives you three controls: you set the key rotation schedule, you restrict decryption access through the key policy, and you generate per-function audit trails in AWS CloudTrail. A durable execution uses the same encryption key it started with for its entire lifetime. Changing or removing the key affects only executions that start after the change.
Updating the customer managed key policy to remove decrypt permissions, or disabling the key, stops the Lambda service from accessing previously checkpointed state. Customer managed key deletion is a permanent action, and all durable executions encrypted with that key become unrecoverable because the Lambda service has no mechanism to restore the data. Before scheduling key deletion, use the AWS KMS waiting period (7 to 30 days) and monitor AWS CloudTrail for Decrypt calls to confirm that the key is no longer in active use.
In this post, you learn to configure a customer managed key to encrypt durable execution data in an event-driven payment processing workflow. You create a symmetric encryption key in AWS KMS and define a key policy that grants the Lambda service, the function’s execution role, the function author, and durable execution operators only the AWS KMS actions each principal requires. You then configure the function to use the key for durable execution encryption and verify encryption operations through AWS CloudTrail logs. By the end, you have a deployable reference architecture you can adapt for regulated workloads running on Lambda durable functions.
The sample application implements an event-driven payment processing pipeline using Amazon DynamoDB, Amazon EventBridge, Amazon EventBridge Pipes, AWS Lambda, and Amazon SQS. The pipeline receives authorized payment transactions, validates and enriches them. A Lambda durable function applies business rules to the enriched transactions. The approved transactions are sent to a downstream settlement system for posting.
The following section covers the key architectural steps.
Architecture steps
The upstream authorization system writes authorized payment records to a DynamoDB table.
DynamoDB Streams captures each new record as an ordered change event.
Amazon EventBridge Pipes polls the record from the DynamoDB stream. The pipe triggers a Lambda function as part of enrichment step for duplicate checking.
The deduplication Lambda uses a DynamoDB table with conditional writes to identify duplicate inbound transactions based on transaction properties and time window.
When the deduplication is successful, the pipe publishes an event to the Amazon EventBridge custom event bus.
An Amazon EventBridge rule invokes a Lambda function for matching events. The function adds business context such as account type, bank routing details, and merchant category codes. The function publishes a new enriched event to the custom event bus.
Another Amazon EventBridge rule matches the enriched events to a Lambda durable function. The durable function applies business rules to the incoming event. When the event passes all business rules, the function publishes a new event to the event bus.
An Amazon EventBridge rule routes the approved event to an Amazon SQS queue preserving ordering for settlement and buffering against downstream throughput limits.
The Posting Lambda function reads from the Amazon SQS and invokes the downstream posting subsystem to post the transaction. Finally, the function publishes a completion event to the event bus completing the transaction lifecycle.
With customer managed keys configured on DynamoDB, Amazon EventBridge, SQS, and the AWS Lambda durable function, every piece of persisted data in this pipeline is encrypted with keys you own and control. The walkthrough that follows shows you how to deploy this configuration with Terraform.
Figure 1 shows the reference architecture for this solution.
Reference architecture
Figure 1: Payment processing using Lambda durable functions
Prerequisites
To deploy this solution, you need the following prerequisites:
AWS account and CLI: An active AWS account with the AWS CLI installed and configured with appropriate credentials.
Terraform: Terraform installed (version 1.0 or later) for infrastructure provisioning.
Python environment: Python 3.11 or later, with pytest for running unit tests. The aws-durable-execution-sdk-python package requires Python 3.11 or later.
The following is a step-by-step guide to deploy and test the payment processing solution.
Step 1: Clone the repository
git clone https://github.com/aws-samples/sample-payment-processing-with-lambda-durable-functions.git
cd sample-payment-processing-with-lambda-durable-functions/source
Step 2: Run unit tests
Validate the payment processing logic locally before deploying:
cd lambda-src/business_rules
pip3 install -r requirements-test.txt
pytest test_app.py -v
This runs unit tests that cover transaction validation, business rule checks (foreign transaction detection, currency conversion, merchant type), event schema validation, and misconfiguration handling. The tests use the AWS Durable Execution Testing SDK to run the handler locally without deploying AWS resources.
Figure 2 shows an example of test results running locally.
Figure 2: Test run results of the business rules
Step 3: Inspect the Lambda durable functions construct
Open the payments-business-rules Lambda function in source/lambda-src/business_rules/business-rules-app.py for a sample Lambda durable function. Refer to Figure 3 for the code walkthrough.
Key features used
@durable_execution decorator: Transforms a standard Lambda handler into a durable function handler. The durable execution SDK manages checkpointing automatically. No infrastructure changes are required.
context.step("validate-transaction"): Validates that the transaction has a non-empty issuingCountryCode. The durable execution checkpoints the result (True or False) to durable storage. The durable execution restores checkpoint results instead of re-executing steps during the replay phase. This phase occurs whenever the function is re-invoked after an interruption such as a wait period completing, a failure, or a suspension. This checkpointed result is part of the durable execution data encrypted by your customer managed key.
context.step("publish-posting-failure"): Publishes the full Amazon EventBridge envelope to Amazon SNS when validation fails. This step only runs on the failure path. The runtime checkpoints the Amazon SNS publish response to durable storage.
context.parallel("run-business-rules"): Runs three independent rule checks concurrently: foreign transaction detection, currency conversion, and merchant type validation. Each branch checkpoints independently. If one branch fails, the others are not replayed on resume. Each branch result is persisted to durable storage and encrypted by the customer managed key.
ctx.step("trigger-foreign-transaction-rule") (inside parallel): Compares billingAmount against transactionAmount. If they differ, it emits a ForeignTransactionFound event to Amazon EventBridge. This step is checkpointed independently within the parallel group.
ctx.step("trigger-conversion-rate-rule") (inside parallel): Checks whether conversionRate equals 1. If so, it emits a CurrencyConversionTransactionFound event to Amazon EventBridge. This step is checkpointed independently within the parallel group.
ctx.step("trigger-merchant-rule") (inside parallel): Checks whether merchantType equals AAFF. If so, it emits a WarningMerchantTypeTransactionFound event to Amazon EventBridge. This step is checkpointed independently within the parallel group.
context.step("post-transaction-processed"): Emits the final TransactionPostingApproved event to Amazon EventBridge. This step is only reached when validation passes and all business rules complete. The runtime checkpoints the Amazon EventBridge response. On replay, if this step already succeeded, the event is not re-published, which guarantees exactly-once approval semantics.
context.logger: Provides replay-aware logging throughout the handler. During replay of previously completed steps, log statements are suppressed to prevent duplicate log entries in Amazon CloudWatch.
Figure 3: Sample Lambda durable functions code
Step 4: Deploy infrastructure with Terraform
Terraform currently doesn’t support attaching a customer managed key directly to the durable function. You create the symmetric key in Terraform and then associate the key with the durable function on the AWS Management Console. Refer to source/durable_kms.tf for the key configuration.
Initialize and deploy the AWS resources that make up the solution:
cd ../../
terraform init
terraform plan -var="region=us-east-2"
In the AWS Lambda console, navigate to the payments-business-rules function. Confirm that the function Type displays Durable, which indicates that the checkpoint-and-replay mechanism is active. Figure 4 shows the expected function configuration.
Figure 4: The Lambda durable function in the AWS Lambda console
Step 6: Add the AWS KMS key to the Lambda durable function
The durable function is not encrypted with a customer managed key. Figure 5 shows the function’s encryption configuration as empty.
Figure 5: The Lambda durable function missing a customer managed key in the AWS Lambda console
Choose Edit, then turn on Customize encryption settings as shown in Figure 6.
Figure 6: The Lambda durable function check encryption in the AWS Lambda console
Select the AWS KMS key ARN created for the durable function. The key ARN is available in the Terraform output from Step 4. Figure 7 shows the key selection.
Figure 7: Select the AWS KMS key ARN for the durable function in the AWS Lambda console
Choose Save and confirm that the durable function is now encrypted with a customer managed key, as shown in Figure 8.
Figure 8: AWS Lambda durable function with the customer managed key in the AWS Lambda console
Step 7: Execute a test payment
Invoke the payments-visa-mock Lambda function to simulate an end-to-end authorization flow. The mock function reads sample Visa authorization messages from a CSV file and writes them to DynamoDB, which triggers the event-driven pipeline. Figure 9 shows a sample test invocation.
Figure 9: Invoke the payments-visa-mock function to trigger workflow
Figure 10 shows a sample response after invocation.
Figure 10: Test results from the payments-visa-mock function to trigger workflow
The mock Lambda invocation creates records that follow the process described in the preceding architecture steps.
Step 8: Verify results
Open Amazon CloudWatch Logs and inspect the log group /aws/lambda/payments-business_rules. This log group belongs to the Lambda durable function for this use case. Figure 11 shows the CloudWatch log group on the console.
Figure 11: Search in CloudWatch
You see the complete business rules lifecycle for each transaction, as shown in Figure 12. The highlighted sections show all the business rules performed by the durable function. Each step is checkpointed by the runtime and encrypted by the customer managed key.
Figure 12: Search Results in lambda durable functions console
You can also check the other log groups to trace the full pipeline:
You can search in AWS CloudTrail to track the AWS KMS calls. When you configure or update the customer managed key on a durable function, Lambda validates the key policy with dry-run GenerateDataKey and Decrypt calls. These appear in CloudTrail with a DryRunOperationException error code, which confirms that the key policy permissions are correct and does not indicate an actual error. For more details, see Encrypting AWS Lambda durable execution data.
Clean up
To avoid ongoing charges, destroy all deployed resources using the following command:
In this post, you configured a customer managed key to encrypt durable execution data in a Lambda durable function. With a customer managed key, you control the key rotation schedule, restrict decryption access through the key policy, and generate per-function audit trails in AWS CloudTrail. You can revoke access to durable execution data at any time by updating the key policy, giving you full control over who can read execution state. In-flight executions stop at the next checkpoint call and new executions must be started after restoring access. For details, see When the customer managed key is unavailable.
For payment processors and financial institutions, encrypting durable execution data with a customer managed key satisfies compliance obligations for data-at-rest encryption, key governance, and access auditability across multi-step transaction workflows.
Apache Airflow has become the orchestration backbone for data pipelines across industries. But as those pipelines grow to hundreds of directed acyclic graphs (DAGs) spanning services like AWS Glue, Amazon EMR, Amazon Athena, and Amazon Redshift, debugging a single task failure turns into a significant operational challenge. When a task fails, data engineers sift through logs, cross-reference DAG configurations, and analyze error messages to find the root cause, delaying pipeline service level agreements (SLAs) and impacting team productivity.
In this post, we show you how to build a custom Apache Airflow plugin that integrates with Amazon Bedrock to automatically analyze DAG task failures and provide actionable diagnostic insights. The plugin deploys to Amazon Managed Workflows for Apache Airflow (Amazon MWAA) and provides AI-powered root cause analysis on demand.
The complete source code for this solution is available in the sample-aws-mwaa-llm-powered-plugin GitHub repository. Clone the repository and follow along as we explain the design decisions throughout this post.
Solution overview
Apache Airflow is a widely adopted open source platform for programmatically authoring, scheduling, and monitoring complex data pipelines. Teams use Airflow to orchestrate extract, transform, and load (ETL) processes, machine learning workflows, and data lake management across industries.
Amazon MWAA is a managed service that makes it straightforward to run Apache Airflow on AWS without the operational burden of managing the underlying infrastructure. With Amazon MWAA, you can focus on authoring workflows and business logic while AWS handles provisioning, patching, scaling, and securing your Airflow environments.
The plugin adds an analysis view directly into your Airflow UI. At a high level, when a task fails and you trigger an analysis, the plugin automatically does the following:
Retrieves the failed task instance metadata from the Airflow metadata database.
Collects comprehensive context including task logs, DAG source code, and operator-specific scripts.
Sends the enriched context to Amazon Bedrock for analysis.
Returns a structured diagnostic report with root cause identification, step-by-step resolution, and prevention recommendations.
How it works
The preceding four steps happen behind a single Analyze Task action. The following diagram and pipeline show the high-level architecture and how the plugin carries them out.
Figure 1: High-level architecture of the LLM-powered task analyzer plugin on Amazon MWAA
The plugin follows a multi-step analysis pipeline:
User triggers analysis – From the Airflow UI, you select a failed task and choose Analyze Task.
Context collection – The plugin retrieves task metadata, execution logs, and DAG source code from the Airflow metadata database and Amazon S3.
Operator-aware enrichment – Based on the operator type, the plugin fetches the actual code or query that failed (for example, a PySpark script from AWS Glue or a SQL query from Amazon Athena).
Foundation model analysis – The enriched context is sent to Amazon Bedrock, which returns a structured diagnostic report.
Results presentation – The analysis displays in the Airflow UI with actionable recommendations.
All AWS API calls (Amazon Bedrock, Amazon S3, and AWS Glue) are authenticated through the aws_default Airflow connection. By default on Amazon MWAA, this connection has no static credentials, so boto3 falls back to the environment’s execution role. This means there are no keys to manage or rotate. If you need to call Amazon Bedrock or fetch scripts using a different identity, you can supply those credentials in the aws_default connection. This can be a dedicated IAM role or a cross-account principal, used instead of the execution role.
Operator-aware context collection
A key differentiator of this solution is its ability to understand different Airflow operator types and automatically fetch the associated code or queries. Unlike generic log analyzers, the plugin retrieves the actual code that failed, not just the error message.
The following table summarizes what the plugin fetches for each operator type:
Operator type
What the plugin fetches
Source
GlueJobOperator
PySpark or Python script
Amazon S3 (from the AWS Glue job definition)
EmrAddStepsOperator
Spark or Python script
Amazon S3 (from step arguments)
EmrServerlessStartJobOperator
Spark script
Amazon S3 (from job driver)
AthenaOperator
SQL query
Inline (from operator parameters)
RedshiftDataOperator
SQL query
Inline (from operator parameters)
BashOperator
Bash command
Inline (from operator parameters)
PythonOperator
Python function
DAG source code
This approach means the foundation model can analyze the actual logic that failed, correlating error messages with specific lines in your code for precise root cause identification.
Prerequisites
Before you begin, make sure that you have the following:
An Amazon MWAA environment running Apache Airflow 3.x (this walkthrough uses Airflow 3.2). The plugin registers its UI through the FastAPI-based plugin interface (fastapi_apps) introduced in Airflow 3.x. For setup instructions, see Get started with Amazon MWAA.
Access to Amazon Bedrock with the Anthropic Claude model family enabled in your AWS Region. This walkthrough uses Anthropic Claude, but you can adapt the plugin to work with Amazon Nova or other foundation models by modifying the prompt payload format in prompts.py. See Model access.
Note: In most Regions, you invoke Claude through an inference profile ID (for example, us.anthropic.claude-sonnet-4-5-20250929-v1:0) rather than a bare on-demand model ID. Run aws bedrock list-inference-profiles to confirm a model is ACTIVE before configuring it.
Plugin design
In this section, we explain the plugin design and its key components. The next section walks through deploying it to your Amazon MWAA environment.
The repository also includes example DAGs that simulate various failure scenarios across different operator types.
Plugin registration
In Apache Airflow 3.x, the web component of a plugin is registered as a FastAPI application through the fastapi_apps attribute. In task_analyzer_plugin.py, the TaskAnalyzerPlugin class registers the FastAPI app under /task-analyzer and adds a view to the task instance page:
Airflow automatically discovers any AirflowPlugin subclass in the plugins folder. No registration call or configuration change is needed. On Amazon MWAA, the file is delivered inside plugins.zip and extracted to /usr/local/airflow/plugins/.
Analysis engine
The analysis engine is the POST /api/analyze-task endpoint in task_analyzer_plugin.py. When you trigger an analysis, the endpoint performs the following steps:
Retrieves AWS credentials from the aws_default Airflow connection. To override, edit the aws_default connection in the Airflow UI (Admin > Connections).
Assembles a context dictionary from the request (task metadata, logs, DAG source).
Enriches the context with an operator-specific script through fetch_and_add_operator_script.
Builds the prompt using the template in prompts.py.
Invokes Amazon Bedrock and returns the structured analysis.
Operator script fetching
The process_operator_script function in script_utils.py routes script retrieval based on operator type:
External scripts (AWS Glue, Amazon EMR) – The plugin calls the AWS Glue API to look up the job definition, then reads the PySpark script from Amazon S3. Amazon EMR handlers follow the same pattern, extracting the script path from the step configuration or job driver.
Inline scripts (Amazon Athena, Amazon Redshift, BashOperator, PythonOperator, DBTOperator) – The plugin reads the query or command directly from the task’s rendered template fields with no external API call.
The plugin implements smart fetching: for external scripts, it only makes the Amazon S3 API call when the error message contains code-relevant patterns (such as SyntaxError, TypeError, or data type mismatch). Infrastructure errors like timeouts skip the script fetch entirely, minimizing unnecessary API calls.
Prompt engineering
The prompt template in prompts.py provides the foundation model with:
Task metadata (DAG ID, task ID, run ID, state).
Error message and execution logs.
DAG source code.
Operator-specific script (when available).
The model produces a structured diagnostic report with root cause identification, step-by-step resolution, and prevention recommendations. Model IDs are configurable through Airflow Variables, so you can switch between Claude Sonnet and Claude Opus without redeploying the plugin.
Security measures
Before sending content to Amazon Bedrock, the plugin applies the following safeguards:
Credential redaction – The sanitize_script function removes sensitive patterns (passwords, tokens, access keys) from scripts and logs.
Content truncation – The truncate_script function caps content size to stay within model context windows.
Path traversal prevention – The read_allowlisted_file function resolves canonical paths and verifies they reside within allowed base directories before reading any file.
Optional: PII detection and redaction. The built-in sanitize_script function targets credential patterns. If your logs or scripts might contain personally identifiable information (PII), consider adding a detection pass with Amazon Comprehend before invoking Amazon Bedrock. The DetectPiiEntities API returns the entity types (such as names, email addresses, or account numbers) and their character offsets. You can use these offsets to mask or obfuscate the spans before the context leaves your environment. This adds one API call and cost per analysis, so add it where your compliance requirements call for it. For guidance, see Detecting PII entities.
Deploy the plugin
Follow these steps to deploy the plugin to your Amazon MWAA environment.
Step 1: Clone the repository
git clone https://github.com/aws-samples/sample-aws-mwaa-llm-powered-plugin.git
cd sample-aws-mwaa-llm-powered-plugin
Step 2: Package and upload to Amazon S3
Create the plugins.zip archive from the plugins/ directory and upload it to your Amazon MWAA S3 bucket:
cd plugins
zip -r ../plugins.zip .
cd ..
aws s3 cp plugins.zip s3://<amzn-s3-demo-bucket>/plugins.zip
aws s3api head-object \
--bucket <amzn-s3-demo-bucket> \
--key plugins.zip \
--query VersionId --output text
Note the VersionId returned. You need it in the next step.
Note: This plugin requires only fastapi and Boto3, both pre-installed on Amazon MWAA for Airflow 3.x. You don’t need a requirements.txt file. Skipping the requirements file avoids package resolution conflicts that are a common cause of failed Amazon MWAA environment updates.
Step 3: Update the Amazon MWAA environment
Update your environment to use the new plugin archive:
On Amazon MWAA, the aws_default connection exists by default and resolves to your environment’s execution role. In most cases, no action is needed.
To override the Region, edit the aws_default connection in the Airflow UI (Admin > Connections) and set the Extra field to:
{"region_name": "us-east-1"}
Leave login and password empty so the execution role is used.
Step 5: Verify the deployment
After the environment finishes updating, navigate to Admin > Plugins in the Airflow UI. Verify that task_analyzer_plugin appears in the list. The Analyze Task entry is now available from any task instance view.
Test the solution
The repository includes example DAGs that simulate failure scenarios across different operator types. To validate the deployment:
Copy the dags/ directory contents to your Amazon MWAA S3 bucket’s DAGs folder:
Wait for Amazon MWAA to sync the DAGs (typically 1–2 minutes).
In the Airflow UI, trigger one of the test DAGs (for example, test_aws_sql_operators) and let the intentional failure occur.
Navigate to the failed task instance.
Choose Analyze Task in the task instance view.
Review the generated analysis, which includes:
Root cause identification with file and line references.
Step-by-step resolution with code examples.
Prevention recommendations and monitoring suggestions.
The analysis typically completes within 5–10 seconds.
Cost considerations
The primary cost driver for this solution is Amazon Bedrock inference, which is billed by the number of input and output tokens each analysis consumes. Input tokens come from the task logs, DAG source, and operator script sent to the model. Output tokens come from the diagnostic report the model returns. Larger logs and scripts increase input tokens, and the model you select affects the per-token rate. For current per-model rates, see Amazon Bedrock pricing.
To help control cost, the plugin includes a caching mechanism that stores results keyed by a hash of the error context. Repeated analyses of the same failure pattern return cached results without invoking Amazon Bedrock again.
Best practices
When you deploy this solution in production, consider the following:
IAM least privilege – Grant only bedrock:InvokeModel for your chosen model IDs and scope s3:GetObject to specific bucket paths where your operator scripts reside. For guidance, see Amazon MWAA execution role.
Data sanitization – The plugin redacts credentials and truncates content before sending data to Amazon Bedrock. Store configuration values in AWS Secrets Manager rather than hardcoding them in DAG source files.
Operational resilience – Add retry logic and circuit breaker patterns around the Amazon Bedrock API call. Use Amazon CloudWatch to monitor plugin performance and set alarms on failure rates.
Extending the solution
You can extend this solution in the following ways:
Proactive notifications – Integrate with Amazon Simple Notification Service (Amazon SNS) or Slack to deliver analyses automatically when failures occur.
Knowledge base integration – Build a knowledge base of past analyses using Amazon Bedrock Knowledge Bases for Retrieval Augmented Generation (RAG) powered recommendations that learn from your organization’s historical failures.
Additional operator support – Add handlers for custom operators specific to your organization, such as proprietary data connectors or internal platform integrations.
Automated remediation – For well-understood failure patterns, trigger automated fixes such as restarting tasks with adjusted resource configurations.
Clean up
To remove the plugin from your environment:
Delete the plugin archive from Amazon S3:
aws s3 rm s3://<amzn-s3-demo-bucket>/plugins.zip
Update your Amazon MWAA environment to remove the plugin reference, then wait for the environment to restart.
Optionally, remove the Amazon Bedrock permissions from your execution role if they are no longer needed.
Conclusion
In this post, we showed you how to deploy an LLM-powered DAG failure analysis plugin for Amazon MWAA using Amazon Bedrock. The operator-aware context collection differentiates this approach from generic log analyzers. By fetching the actual code from AWS Glue, Amazon EMR, and other services, the foundation model provides precise, actionable recommendations with specific line references.
To get started, clone the sample-aws-mwaa-llm-powered-plugin repository, deploy it to a development Amazon MWAA environment, and test with the included example DAGs. As your team builds confidence in the analysis quality, roll it out to production environments where it serves as the first line of investigation for any pipeline failure.
About the authors
The collective thoughts of the interwebz
Manage Consent
To provide the best experiences, we use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us to process data such as browsing behavior or unique IDs on this site. Not consenting or withdrawing consent, may adversely affect certain features and functions.
Functional
Always active
The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.
Preferences
The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.
Statistics
The technical storage or access that is used exclusively for statistical purposes.The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
Marketing
The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.