Tag Archives: Developers

Building an evidence-grounded agentic security operations harness on Cloudflare

Post Syndicated from Deanna Tran original https://blog.cloudflare.com/agentic-security-operations/

Security alerts rarely arrive one at a time. A single alert can cause a spike across the environment, requiring a human analyst to decide which alerts are related and what they mean. When multiple arrive at the same time, it can quickly overwhelm even a seasoned security analyst. Enter the alert paradox. Now, our built-in, multi-AI-agent security operations harness can handle more of this work at Cloudflare scale.

Our Cloudflare Managed Defense AI agent harness speeds up the process of gathering data, connecting and aggregating detections, and accounting for missing sources while new alerts continue to arrive. To further analyze context, we make use of the OpenAI Daybreak Defense Network and our partnership with Anthropic. Cloudflare uses approved OpenAI Daybreak and Anthropic models, including GPT-5.6 Cyber and Mythos, for deeper model-backed analysis. Initial analysis and scoring is done with Clef, Cloudflare’s open-source decision model.

Collecting evidence to understand what each alert means requires a lot of time. Even in a highly sophisticated Security Information and Event Management (SIEM), too much is still left for a human to review. Human analysts must gather data, connect and aggregate detections, and account for missing sources while new alerts continue to arrive.

Think about every time a human analyst reviews an alert: "Which should we silence? Which action should we take? Which alert should we ignore? Which should we resolve as false positives? Which should we resolve as true positives? Which should trigger our incident team?" We address this predicament with our AI agent strategy. Our approach reduces all of those questions and gives Managed Defense Analysts a quick and consolidated view, directly providing insight into related alerts, admitted evidence, visible gaps, and recommended next steps. The result: cutting back on the time needed to analyze, and creating a hyper focus on actually getting security alerts resolved and mitigations deployed.

Why a single agent fails

Our first prototype showed the limits of one general-purpose agent. We provided the AI agent the whole investigation. It produced useful analysis, but it also hallucinated claims the evidence did not support. Telemetry, detector descriptions, policies, and threat intelligence were flattened into one prompt, which caused their distinct roles to merge together.

We saw three recurring problems with our first single-shot AI agent harness:

  • Context became authority. A detection is a hypothesis, not proof that an exploit succeeded or an attack occurred. A broad AI agent can blur that distinction.
  • Scope drifted. An AI agent can query the wrong account, time range, or source. You can’t rely on a language model prompt to be a boundary.
  • Failure disappeared. If a lookup times out, the result may not distinguish "not checked" from "checked and not found."

To address these challenges we moved evidence collection and scope enforcement into application code, before model analysis begins.

Recon first, inference second

It's tempting to put an AI agent at every step. The front half of our harness has none.

Before we even call inference, deterministic code runs a fixed set of reconnaissance workflows with versioned API calls. It collects the customer's identity, detection history, traffic baseline, enforcement outcome, and network observations. Each piece of data is stored with its source, version, and timestamp.

Cloudflare sees both the request and the action applied to it. That lets the investigation connect the behavior that triggered an alert with both the control that fired and its outcome.

The fixed recon snapshot also makes evaluation reproducible. If AI agents fetch their own data, two runs may disagree because their inputs changed. Here, the same snapshot can be replayed, so differences between specialist AI agents’ findings come from interpretation rather than retrieval.

Filter noise early

Most alerts are not incidents. The same rule often fires repeatedly on a known traffic pattern, and paging Managed Defense Analysts every time makes it easier to miss a real security incident.

We needed a lightweight triage model to compare each alert with its reconnaissance data: Has this event been detected for the customer before? What did Managed Defense Analysts decide previously? Does the traffic look consistent with normal human behavior? Alerts scored with a high likelihood to be false positives skip analysis by the specialist AI agents.

Clef, running on Workers AI, was the perfect fit for this type of fast agentic reasoning. 

Known high-volume noise is deterministically classified as passive when it arrives. It remains available as context but does not enter the active queue.

Specialist AI agents handle the investigation

For alerts that need deeper review, a coordinator AI agent runs four specialist AI agents in parallel:

  • Traffic analysis reviews request behavior, historical changes, and enforcement.
  • Customer context reviews earlier alerts, dispositions, and Managed Defense Analysts’ decisions.
  • Global telemetry compares the activity with privacy-preserving Internet-wide signals.
  • Threat intelligence checks indicators already admitted to the alert or case.

A synthesis AI agent combines their typed findings into one advisory; it can’t fetch new evidence or choose a classification outside the approved vocabulary. Keeping each task narrow makes unsupported claims easier to catch, and recommendations easier to audit.

Global context without customer data

A security tool knows what happened inside the environment it is deployed in, but little about the world beyond it. Cloudflare compares an alert with patterns seen across its global network.

For example, an IP may be targeting one site, scanning thousands of sites, or appearing for the first time. Those patterns carry different weights. To preserve customer privacy, the global telemetry specialist works only with aggregates; it never receives another customer's individual records or identity.

This view combines features from Cloudflare's CDN, WAF, DDoS, Turnstile, Rate Limiting, and Cloudforce One threat intelligence. The synthesis AI agent weighs both global reputation and customer history. This ensures a globally common pattern can still be used on a per-customer basis, but does not automatically imply a widespread campaign for all customers.

History is evidence

Every evaluation is conscious of what came before it: the alert, the pattern, and which customer. The recon dossier records the alert's own track record including how many times this service alert has fired, how many of those were dispositioned as false positives, and what the Managed Defense Analyst concluded. An attack pattern that has been benign every time your Managed Defense Analysts have seen it is a very different object from the first sighting of something new, and the specialist AI agents are told which one they are looking at. Approved background context is retrieved from previous alerts and cases, so yesterday's conclusions are carried into today's decision, instead of being rebuilt from scratch.

Our system aggregates related alerts into a consolidated case. In each case, we store evidence, findings, and recommendations. Our system deterministically joins and correlates this data, but we leave it to a Managed Defense Analyst to confirm the actual scope. Over time, a case can connect network, application, and Zero Trust evidence together, while tracking the source of the evidence for additional context and reference purposes.

From evidence to decision

Before analysis, the system creates a versioned evidence package with the subject, scope, time anchor, admitted evidence, policy versions, sources, and coverage gaps. Specialists must cite items in that package. Application code checks that every citation exists, belongs to the investigation, and supports the attached claim. Invalid findings are corrected or recorded as limitations.

We lean on Clef a second time to score our evidence. Is the collected evidence enough for a decision? Does any of our collected evidence contradict? Based on this evidence, Clef picks from a deterministically reduced list of attack classifications and dispositions.

Cloudflare's developer platform runs this process. Application code on Workers admits evidence and validates results; Workflows coordinates each stage and saves completed work before the next begins, so a failed stage reuses evidence and findings that already passed validation instead of starting over. D1 keeps investigation and advisory state, R2 holds bounded context and evidence artifacts. Case-chat state persists in Durable Objects, uses Flue, and is enriched with AI Search.

Finally, an LLM-powered agent produces an advisory report using terms that Managed Defense Analysts already work with: affected surface, enforcement outcome, relevant controls, and next step. Managed Defense Analysts can inspect evidence, investigate further, revise the recommendation, or group alerts into a case. Application code fixes the customer scope before any model sees results, and gives each specialist only the evidence it needs. The model never receives authority to cross tenant boundaries or act for the Managed Defense Analysts.

Handling incomplete evidence

At network scale, a source will sometimes fail. A comparison may time out, metadata may be missing, or a threat intelligence lookup may return no match. The system keeps evidence already collected and records the gap.

The advisory distinguishes three states:

  • Not checked
  • Checked, with no matching result
  • Checked, with evidence supporting absence

If global telemetry is unavailable, the system can describe what is unusual for the customer but cannot say whether the pattern is widespread. When the evidence is insufficient, it makes no classification or disposition recommendation.

Remediation

A useful recommendation should lead to a solution, rather than a ticket. The advisory might suggest a rate limiting rule for an abusive path, a WAF custom rule for a signature, or a DDoS protection change. For fully managed customers, Managed Defense Analysts can apply the suggested rules; other customers will receive their recommendations in the dashboard and through their chosen alert path.

In the end, the Managed Defense Analyst remains responsible for the decision and any mitigation. Each alert and case includes the evidence behind the AI agent’s recommendation, which empowers the Managed Defense Analyst to reach a conclusion by accepting or updating the AI agent’s advice.

What comes next

Managed Defense Analysts remain responsible for judgment. The AI agent harness handles more of the repetitive work: assembling investigations, connecting related events, and showing the evidence behind each recommendation. Over the next few quarters, we plan to add a Custom Managed level with more flexibility for each organization.

We also plan to explore continuous AI agents that monitor Cloudflare traffic and surface patterns that fixed rules and thresholds may miss.

The early beta is available in Cloudflare Managed Defense for eligible application-security alerts and cases. If you already use Cloudflare WAF, DDoS protection, Magic Transit, or another supported product, talk to your enterprise account team about adding Managed Defense.

Protected Quick Tunnels: simple accountless authentication for your next dev project

Post Syndicated from Nikita Cano original https://blog.cloudflare.com/protected-quick-tunnels/

We launched Quick Tunnels in 2021 to give developers an easy way to share their latest service, application, or project running in their local development environment. A lot has changed since then, but the core use case remains the same.

Your coding agent has just finished the feature. The dev server is up on localhost:5173, and before you ask, the agent offers to let you try it on your phone. It runs one command and hands you a link:

That command starts a Quick Tunnel. cloudflared, Cloudflare's lightweight connector, publishes your local service at a random trycloudflare.com URL. No account, no domain, no cost. Agents now use Quick Tunnels for the same reason people do: they are the shortest path from a local port to a URL.

The catch has always been the same. Anyone with the link can open it.

Starting with cloudflared 2026.9.3, you can add --allowed-mail to the command, and your Quick Tunnel only lets in the email addresses and domains you choose. Visitors prove they own one of those addresses with a one-time PIN from Cloudflare Access. Nobody, on either side, needs a Cloudflare account.

Agents made Quick Tunnels more popular than ever

Agents that write code need somewhere to show you the result. Agents that live on a Mac mini at home need to be reachable from your phone. Model Context Protocol servers on a laptop need a public endpoint before a hosted assistant can call them. Each of these needs a URL, and a Quick Tunnel produces one from a single command an agent can run by itself. There is no signup form for it to get stuck on. Add --output json and every log line becomes a JSON object, so the agent can pick out the URL without scraping text.

Since agents took off, Cloudflare Tunnel and Quick Tunnels adoption has grown exponentially. On September 18, 2026, a link to the Quick Tunnels page climbed to the top of Hacker News and gathered more than 800 points and 300 comments. The thread reads like a catalog of agent workflows. One person's AI had found Quick Tunnels on its own to publish a site it had just built. Another called them "insanely helpful when doing agentic work on the go."

And one commenter asked this post answers: "how long until someone's agent sets up a tunnel for the world to see one's most sensitive, private and embarrassing information or insecure work-in-progress app?"

Control who can access your service

Pass an email address to --allowed-mail:

Alice opens the URL, enters her email address, types in the code sent to her inbox, and reaches your app. Anyone else is stopped before a single request reaches your machine. You still don't create a DNS record, write a configuration file, or open a dashboard.

To let in more people, repeat the flag or allow an entire domain:

If you leave out --allowed-mail, nothing changes. Public Quick Tunnels behave exactly as they always have.

To change who can get in, stop cloudflared and start a new tunnel. Access ends for everyone the moment the process exits.

For a stable hostname or richer rules, such as identity provider groups, use Cloudflare Tunnel with Cloudflare Access. To reach an agent at home from your own devices without any public URL and establish bidirectional connectivity, use Cloudflare Mesh.

Make it your agent's default

Because protection is a single flag, agents can use it as easily as people can. Add one line to the instructions file your coding agent reads, such as AGENTS.md:

From then on, the previews your agent shares should open only for you. Agents don't always follow instructions, so check what it ran: cloudflared prints whether a tunnel uses email authentication and how many rules it holds, without printing the addresses.

Start a protected tunnel from Wrangler

If you build on Workers, you can start the same kind of tunnel from the latest version of wrangler:

Wrangler supports repeated flags, comma-separated values, and wildcard domains, and it removes --allowed-mail values from its debug logs.

Cloudflare verifies the email. Your machine decides who gets in.

When someone opens a protected URL, they land on the Cloudflare Access sign-in page. They enter their email address, then the one-time PIN sent to that mailbox. Email sign-in is built for people using a browser.

That step answers one question only: does this person control this email address? It doesn't decide whether they're welcome. cloudflared makes that decision on your machine by comparing the verified address with the rules you typed.

Where does a policy live when there is no account?

Separating those two questions is the core of the design. Authentication proves who a visitor is. Authorization decides whether that visitor gets in. Every Cloudflare product that enforces access rules keeps the authorization half in the same place: your Cloudflare account. A Quick Tunnel doesn't have one. So the hard part was never sending someone a code. It was deciding where the guest list should live.

We started with four requirements. The design had to:

  • Keep Quick Tunnels accountless, because a signup step would defeat the point of a one-command tunnel.
  • Leave the request path for public Quick Tunnels untouched.
  • Avoid a central policy lookup on every request after a visitor signs in.
  • Protect the privacy of the email addresses developers type into their terminals.

Our first idea was to put a Cloudflare Access application in front of every Quick Tunnel hostname. Access already checks visitors before traffic reaches cloudflared, so reusing it looked like the shortest path. But hundreds of thousands of Quick Tunnels can be running at once, many for only a few minutes, and each would need its own application and policy. With no account to own them, we would have had to invent a new namespace and route applications dynamically, just to store a list that lives for an afternoon.

Our second idea was to build the whole flow. cloudflared would hold the rules, and a Tunnel service would send and check the codes. The authorization half of this idea was good: each connector checks its own list, which scales naturally and keeps the rules on the developer's machine. The authentication half was not. Sending a code is the easy part of email login. The hard parts are getting email delivered, stopping abuse, building secure challenges, managing sessions, and serving a sign-in page that is accessible and translated, then operating all of it safely for years. Cloudflare Access has already solved those problems.

So we kept the best half of each idea. Access verifies that the visitor controls the email address. A small authentication broker running on Cloudflare Workers turns that verified identity into a short-lived, signed handoff. The broker is stateless by design. It stores no tunnel policies, no visitor sessions, and no identity records, and it never sees a tunnel's guest list. cloudflared checks the handoff and makes the authorization decision itself, in memory, against the rules you typed.

The result is the property we cared about most: your guest list never leaves your machine. Cloudflare learns that a tunnel requires email authentication. It doesn't learn who you invited.

Following a request through a protected Quick Tunnel

A protected tunnel is created the same accountless way as a public one. The only extra thing cloudflared sends is the authentication mode, never your rules. If the service doesn't confirm that mode, cloudflared refuses to start rather than hand you a public URL by mistake.

The first time a visitor opens the URL:

  1. cloudflared sees a request with no session. It redirects the browser to login.trycloudflare.com with a random, single-use state tied to that browser and valid for 10 minutes.
  2. Cloudflare Access sends a one-time PIN to the visitor's email address and verifies it.
  3. The broker checks the Access identity and returns a short-lived, signed assertion bound to the tunnel hostname and to that state. The browser delivers it in a form POST, so it never lands in a URL, browser history, or logs.
  4. cloudflared verifies the assertion, uses up the state, and checks the email against your rules. On a match, it creates a local session and sends the visitor to the page they asked for. Otherwise, the visitor gets a generic response that reveals nothing about the list.
  5. Later requests use that session for up to four hours (less if the visitor's Access sign-in expires sooner), or until you stop cloudflared. There is no central lookup and no policy service.

The session cookie holds a random value and an expiry time, and nothing about who the visitor is. cloudflared strips authentication credentials before forwarding requests, so your app never sees them and never has to implement a login flow. If any check fails, the request never reaches your local service. A protected tunnel never falls back to public mode.

Built by interns

Protected Quick Tunnels were shipped by two interns: Hugo Vicente on product and Alessandro Frigerio on engineering. They took it from the product requirements to the authentication broker to the cloudflared release. That's how internships work at Cloudflare: interns own real problems and deliver solutions to production.

Try it on your next demo

Email protection for Quick Tunnels is free, like Quick Tunnels themselves. Install or update cloudflared, start your local server, and add the --allowed-mail flag:

Setup details, matching rules, and limits are in the Quick Tunnels documentation.

The next time you or your agent shares what you're building, the link will only open for the people you chose.

Building for good: How civil society organizations are automating on Cloudflare

Post Syndicated from Allie Funk original https://blog.cloudflare.com/civil-society-automation/

Tracking how governments target dissidents living in exile. Helping people in crisis find mental health support. Advocating for legislation that protects free expression online. These are a few examples of how some of the world's leading organizations are building the future of non-profit work with Cloudflare.

AI is changing how people do their work. The goal of Cloudflare Impact is to help ensure that non-profit organizations are among the first to benefit. Today, we’re sharing what dozens of civil society organizations have built using our developer services with more than $7.5 million of Cloudflare credits. These stories show what is possible when AI applications are accessible, secure, and affordable to build and run.

From "keep us secure" to "help us build"

We believe a better Internet is one that allows people to express themselves online and access a diverse range of viewpoints. A key part of Cloudflare's mission has been making security services available for everyone and helping ensure that individuals and organizations working for the public interest are not forced offline by those more powerful. Today, Project Galileo, which provides free cybersecurity services to civil society organizations, protects more than 3,400 domains across more than 120 countries.

Through these partnerships, organizations have shared with us how their needs have evolved from not only wanting to secure existing applications, but wanting to build new ones. AI has allowed non-technical teams to design, build, and scale tools tailored specifically for their workstreams and to advance their mission.

This opportunity is arriving at a challenging moment for the sector. Many organizations report operating under financial strain after major reductions in government funding contributed to layoffs and closing of programs. As groups rebuild their work in a new environment, some have reported using AI to help do more with less. But adoption remains ad hoc. In a CIVICUS survey, more than half of civil society respondents viewed privacy concerns as a barrier to use. These groups hold sensitive data, like the location of activists or the identities of anonymous sources, and are disproportionately at risk of cyberattacks, according to our own Cloudflare data. Additionally, the cost of building and running complex automation workflows can be prohibitive. Among those surveyed by CIVICUS, 48% cited financial constraints limiting their adoption.

Cloudflare helps address these concerns. Our developer services incorporate security and privacy protections from the outset. By running on our global network, applications are automatically protected from distributed denial-of-service attacks and attempts to gain unauthorized access to internal systems. Civil society groups can also build and run applications more cost effectively. Our lightweight and serverless architecture scales with demand without requiring organizations to provision or pay for idle infrastructure. Workers AI also removes the need to operate dedicated GPU infrastructure, while AI Gateway provides rate limits, observability into how teams are using AI, and intelligent model routing to reduce unnecessary model calls and keep costs under control.

Awarding $7.5 million for the next generation of civil society work

During Birthday Week last year, we expanded Cloudflare for Startups to include non-profit and public interest organizations, providing each of them with up to $250,000 in credits for our developer services. We are proud to announce that we selected 30 organizations to participate in our first non-profit, startup cohort, and we have been thrilled to watch their ideas and tools designed to tackle problems in climate science, humanitarian aid, mental health, civic engagement, and education come to life.

Here is what a few of these non-profits have built.

  • LebTown — local news, rebuilt on automation: LebTown is an independent non-profit newsroom covering Lebanon County, Pennsylvania. Until this year, the editorial team ran its entire operation across two Google Docs, a Google Calendar, Discord, and Gmail. LebTown is replacing that with a custom-built editorial management system that follows stories from pitch to publication, tracks audience reach, and lets reporters log community impact directly from their app or via Discord. The team has also built tools that convert county real-estate transfer records from PDF exports into structured data, and a real-time copyeditor agent that integrates with LebTown's WordPress site.

  • Kaya Guides — scaling mental health support: Kaya Guides built a WhatsApp-based mental health application that pairs people experiencing depression in India with trained lay counselors, making support accessible without the cost or waitlists of traditional therapy. As of September 2026, the application has supported 10,439 people, with 2,325 currently enrolled. To keep pace with demand, the team migrated its care management system to Cloudflare and now processes around 500,000 WhatsApp messages a month. They built a live AI system that listens to counseling calls and gives counselors real-time feedback to keep sessions on track. According to Kaya, these tools have helped develop a program that is more consistent and self-correcting as it scales, without significant additional cost or complexity.
  • The Snorkelling Society — mapping the world’s snorkeling sites. The Snorkelling Society, is a UK non-profit building a web and mobile application for the global snorkeling community called SnorkelMap. This tool will help people discover new places to snorkel, explore information about different locations, and allow users to contribute their experiences. Because the application is built and maintained by volunteers, the Snorkeling Society uses Cloudflare to secure storage to allow for community contributions and uploaded images while also safeguarding against cyberattacks.

Working together to build automation tools for human rights

We also heard from larger civil society organizations who wanted to build complex workflows and welcomed extra engineering support. These projects included several variations of the same problem: large amounts of manually-collected, dispersed information that needed to be integrated, reviewed, and structured in order for effective analysis to be conducted. The data involved was also sensitive, like the names of human rights abuse victims, and the risk of unauthorized access, data leakage, or hallucinations in an output could have serious consequences.

Cloudflare not only provided free access to its developer service to support the development of each of these tools, but also assembled a team of volunteer engineers — product experts, front-end designers, back-end builders, and security advisors — to design and build them. Questions we’re exploring include where automation is helpful, which models align best with each task, how to incorporate the necessary security controls, and where a person must remain involved.

  • Freedom House — tracking transnational repression: Freedom House was founded in 1941 to advance democracy and freedom globally. They created and maintain the world's most comprehensive database of transnational repression incidents, which is when governments reach across borders to silence dissent among diaspora and exile communities. Manually identifying incidents and coding patterns of transnational repression — which involves looking at thousands of potential cases — is labor intensive and time consuming. To help automate this process, we are creating an interactive dashboard that can ingest, structure, and flag potential cases of transnational repression from public reporting. The tool is fine-tuned on the organization’s methodology to improve accuracy, with Freedom House involved in every step of the process. Given the sensitivity of the topic, security and data minimization have been priorities in every decision. Automating this initial stage in the research process can help staff focus on deeper analysis, producing reports, and working with policymakers to address transnational repression.

Prototype of Freedom House’s tool tracking transnational repression

  • Global Network Initiative — protecting free expression and privacy online: The Global Network Initiative (GNI) is a membership organization of civil society groups, academics, investors, and tech companies (including Cloudflare) working to advance free expression and privacy online. To do this, GNI tracks and responds to a rapidly evolving landscape of regulatory proposals, policy developments, and legal trends across dozens of jurisdictions, from online transparency requirements in the European Union to data localization mandates in the Asia Pacific region. We are building a dashboard that identifies these proposals and recommends opportunities for advocacy based on GNI’s mission and previous work. It surfaces, describes, and categorizes relevant news, calls for comment, and legislative activity. The tool also helps manage each opportunity by tracking the internal review process and alerting staff of upcoming deadlines.

Prototype of GNI’s dashboard that tracks the status of engagement opportunities

  • Article One — assessing human rights risk: Article One is a specialized strategy and management consultancy that advises companies on understanding and mitigating the human rights impacts of their policies, products, and operations. This involves reviewing and summarizing large amounts of documentation, from factory audits to country-context reports. We are working with Article One to build a risk assessment tool that processes and structures this information, identifies salient human rights risks, and generates draft recommendations that staff can review and refine. AI helps structure information, but Article One makes all judgments.

What’s next? Apply to join our second cohort

It’s incredible to see how civil society organizations are evolving their work in an era of AI. Cloudflare is excited to play a small role in this process. Working directly alongside these organizations not only helps advance their missions, but also helps inform how we think about the security and privacy in high-risk environments and how we can scale similar programs moving forward.

Our first non-profit startup cohort shows how automation can help the next generation of community service organizations use AI and automation to serve the public, and how they can do this securely and affordably. 

We’re excited to announce that starting today we are officially opening our startup program to our second  cohort of non-profit organizations.

If your organization is interested, apply here and select the non-profit checkbox. We’d love to build with you!

Updates on our pledge to make Cloudflare features accessible to everyone

Post Syndicated from Justin Hutchings original https://blog.cloudflare.com/enterprise-for-all-update/

A year ago, Cloudflare CTO Dane Knecht announced our intention to make every Cloudflare feature available to everyone. Cloudflare launched an Enterprise tier years ago when larger customers came to us looking for procurement options beyond a credit card, like invoices, custom contracts, and dedicated support. Those offerings met a customer need but over time, a two-tier system developed where some of our most advanced and powerful features were only available to Enterprise customers. Our goal was to close that gap.

Today, teams of every size use Cloudflare, from Fortune 100 enterprises to small businesses, open-source projects, and individuals. Across the platform, we’re committed to ensuring that every user or team can make use of all of Cloudflare’s capabilities in a way that helps their organization thrive.

The underlying philosophy is that Cloudflare should offer products suitable for our most demanding customers — and make those capabilities available to everyone. Large or small, every customer would prefer not to have to call support. Building products that are easy to buy, configure, and consume means more of our products in use and a step closer to a better Internet for everybody.

Every generally available (GA) feature we launched this week that is available on an Enterprise plan is also available to Pay-as-you-go customers, and most are available on the free tier. Where our plans differ, it's in how much you can use, not what you can use. While we haven’t yet met our goal that every feature be available to everyone, in the year since Dane’s announcement, we’ve made great progress.

Here are a few products and features making the transition today from Enterprise to everyone.

Logpush and Logpush Transformers now available to all plans

Flexibility on pushing logs to third parties and how logs are formatted expanded this week from Enterprise-only to all customers.

Logpush delivers Cloudflare logs to storage, security, and analytics destinations, helping customers monitor traffic, investigate issues, and analyze their data using existing tools. Previously available only to Enterprise customers, Logpush is now available to Free, Pro, and Business customers through self-service, pay-as-you-go pricing. Datasets available to Logpush have been expanding as well. We’ve recently added account-scoped firewall events, WebSocket analytics and per-zone post-quantum visibility.

Transformers is also becoming generally available to all customers. With Transformers, customers can use SQL to filter unnecessary records, redact sensitive information, enrich events, and reformat logs before delivery without operating a separate extraction, transformation and loading (ETL) pipeline. Together, Logpush and Transformers give every customer greater control over how their Cloudflare data is prepared and delivered.

In addition, Custom Dashboards which let customers create personalized views highlighting the metrics most critical to them, is now available to all customers.

New tools for managing Cloudflare at scale

Expanding RBAC

Over the last year, we’ve dramatically expanded the availability of Role-Based Access Control (RBAC) across all Cloudflare products and for all customers. Today, nearly all products have RBAC roles available at the account and zone level. Recently, Workers joined R2 and Access in having RBAC roles available at the individual resource level as well, so Administrators can decide who on their team gets specific access to individual Workers.

Multiple Accounts

While fine-grained RBAC lets customers manage subsets of an account, this setup still relies on a small number of super administrators making choices about who gets access to what. Centralized authority works great when your problem space is small, but as the number of teams and projects being managed on Cloudflare grows, it can turn into an organizational bottleneck.

The single account model is excellent in its simplicity, but it can start to feel a little crowded for customers maintaining hundreds or thousands of zones, workers, and storage products. That’s why we’ve been expanding our capabilities around managing multiple accounts.

New Account button

Last month, we quietly launched the New Account button on the dashboard that, for the first time, lets users create additional accounts directly. The response has been overwhelmingly positive, and we’re seeing thousands of customers branching out into additional accounts every week. When you use this button, it creates a new, free, Cloudflare account that you can use to segment your open source projects, or segment the work of multiple teams in your organization. Each of these accounts is independently billed, so you can segment spending across multiple cost-centers directly. Safeguards are in place to prevent fraud and abuse.

New Accounts for Enterprises

While the New Account button is for everyone, for the time being, we recommend that Enterprise customers reach out to their account team to get new accounts provisioned instead. This lets you reuse your existing enterprise agreement and subscriptions across all of your accounts. There is no preset limit on how many accounts an enterprise can request. We will be adding additional features in the future that make this process self-serve for enterprises too.

Organizations

Once you’ve created multiple accounts, how do you organize and track them all? Organizations allow customers to group accounts together with a single analytics and shared configuration surface. It’s in beta for Enterprise customers now, will be GA in October, and will be rolling out to free accounts in early 2027. Adding your multiple accounts to a single organization makes managing them easier by providing a unified surface for visibility and management. Organizations provide shared administrators with unified analytics and audit logging as well as shared WAF, Gateway, and Access IdP configurations.

Enterprises are eligible for exactly one organization. We limit enterprises to a single organization, so there’s a single pane of glass that shows all the company’s assets in one place. This makes life easier, so you can invite the CISO, CTO, or other executive stakeholders and give them unified visibility. If you’re an Enterprise customer and haven’t tried organizations yet, you can set one up directly as long as you are a super administrator of at least one account and nobody else has already created the organization. If the organization has already been started, talk to the other Cloudflare administrators in your company to get your accounts added to it. This process ensures that there’s never an elevation of privilege as we layer on this new management plane.

Terraform and Tags

Once a customer has created multiple accounts, an organization to manage them, and set RBAC rules for the products and resources they contain, they need to be able to manage them in a way that’s auditable and repeatable. Terraform lets customers use Infrastructure as Code to manage everything using version-controlled code rather than clicking on the dashboard in a way that may not be repeatable. In the last year, Cloudflare has made dramatic progress creating a Terraform provider that is built programmatically, so it’s always up-to-date with the latest version of the Cloudflare API. Terraform, like the other features mentioned in this post, is available to all customers, Enterprise and not.

Additionally, Resource Tagging lets customers apply key value tags to a very broad set of resources within the Accounts and Organizations. Today tags can be produced interactively or via API and are useful for organizing resources in the dashboard. In the future we intend to make tags useful in billing and access control scenarios and to be manageable via Terraform.

How we use it all at Cloudflare

With the increasing menu of enterprise-ready options for everyone, one of the top questions we get is “What does Cloudflare do internally?” Within Cloudflare, we create accounts per team, or per service, depending on the nature of the team. We then use Terraform to manage account access and production configuration, giving teams a peer-reviewed, auditable path for changes. Because the scope of each account is narrow, we can grant broader permissions to the engineers responsible for that account while keeping the blast radius contained. This lets teams grow their accounts organically without bottlenecking on a small number of central administrators, and it makes operational work like on-call response faster and safer.

Every account at Cloudflare lives within Cloudflare’s organization, which provides our security team with administrative access to every account within the organization, as well as analytics, policy management, and shared configurations. This makes it easier to align every account in the organization to our security standards. Our teams have the right blend of autonomy and centralized control to go fast.

Enabling teams to quickly sort, organize, and filter their resources is critical in our production environments. While it’s still early, Resource Tagging is enabled internally and teams have begun to roll out tags to make finding the WAF rule, R2 bucket, etc. that they need to interact with easier.

More features for everyone

We launched support for the Authentik identity provider (IdP), SCIM Audit logging, and SCIM 2.0 Group Sync. MCP Server Portals moved into general availability. All these features were once in some way Enterprise-only. Even network management is going self-serve: the Network Overview page and Unified Routing both recently became available for all.

Starting with free

Solving big problems starts with first ensuring they aren’t getting any larger. This year, as part of Code Orange: Fail Small, we announced a commitment to rolling new code out by traffic cohort, starting with our free customers. As a result, today we are committed to introducing no new Enterprise-only features. Naturally there will be some carve-outs for things like Cloudflare for Government that are inherently Enterprise-oriented in nature.

Other progress for free and pay-as-you-go customers

Beyond making previously enterprise-only features available to everyone, we’ve also done a lot of work to make Cloudflare more powerful and accessible for everyone

Billable Usage Dashboard and API

In August, we introduced the billable usage dashboard and API which lets non-Enterprise customers see how much they’ve spent and download their consumption data to use offline directly or through third-party tools like Vantage. We also introduced budget alerts, which are on by default to prevent unpleasant billing surprises. We're prototyping hard spending caps now, with early availability in Q4 2026. Because Enterprise customers have dramatically more variation on contract terms and how they pay, this experience is not yet available to Enterprise customers, but we are hard at work and expect to have an announcement in 2027.

Higher limits available to all customers

Over the past year we’ve increased limits across Cloudflare products. We’re constantly working to increase these defaults, and keep our front door as open as possible to people building the next big thing.

Looking to the future

Between exposing formerly enterprise-only features to everyone and increasing the power of features that were already available to everyone, Cloudflare is committed to building the most powerful and accessible platform for customers large and small without the need for a contract. We still have much work to do on Dane’s pledge from a year ago, but we are committed to getting there and are delighted to be able to highlight our progress over the last year.

Take advantage of these new offerings

  • Create additional accounts to partition the concerns of your organization.
  • Use RBAC to define security policies at the zone and account level.
  • If you’re an Enterprise customer, create an Organization and onboard these accounts. For other customers, we’ll see you in early 2027.
  • Use Terraform to manage the state across your whole organization.  
  • Attend Cloudflare Connect next month to learn more about everything discussed here and meet the team that built it.

Introducing Clef: our open-source decision models, and new RL fine-tuning platform

Post Syndicated from Michelle Chen original https://blog.cloudflare.com/clef-decision-models/

Over the last few weeks, there has been lots of buzz around decision models such as Typesafe AI’s Jev System One model. While classifier models have been around for some time, Jev introduces a new decision model concept into the world of AI — a model that produces bounded structured outputs cheaply, quickly and consistently that can be added into a workflow when a decision is required. These models are capable enough to work over any set of inputs without constantly retraining the model to incorporate new classification categories. This contrasts with the world of Large Language Models (LLMs), which are largely non-deterministic, but are open-ended enough to reason and generate text and tool calls for agentic workloads. 

Today, we’re releasing two Cloudflare-trained decision models, Clef and Clef-flash, hosted on Workers AI. Clef is currently the leader when evaluated against the Jev Decision Index, you can view full results on the live benchmark demo site. These models are smarter, faster, and fully Jev-API compatible, so you can experiment with these hosted models easily. We’re fully open-sourcing these models on Hugging Face under an Apache 2.0 license for you to run locally and experiment with yourselves. 

Lastly, we’re excited to debut our new reinforcement learning (RL) product, which allows customers to fine-tune Clef to suit their use cases as well.

What is a decision model?

A decision model makes classifications to help agents decide how to act, based on certain probabilities. For example, you can pass in a customer support message (inputs) and ask if it is urgent and which team should handle it. A decision model will return typed answers with probabilities (outputs), which your code can use to route the ticket, trigger an escalation, or defer to a human. This means that a human does not necessarily need to be in the loop for agentic decisions anymore — agents can programmatically gather context, make decisions, and take actions on tasks, or defer to a human when needed.

Specifically at Cloudflare, we’ve been testing our new Clef model on our Threat Intelligence team to help us classify website domains. By giving a domain to Clef (with Browser Run) it can quickly identify categories that the domain falls under — for example, it might classify a domain with a 95% chance it is a fashion website, 85% ecommerce, <1% phishing, etc. This classification took our Clef model 2.2s to fetch, render, and classify the website. In contrast, our fastest general LLM gpt-oss-120b took 4.7s in the same workflow, and only returned two classifications. As a user, you can imagine how a 2x savings in latency and results can help us improve our threat intelligence workflows and be faster in identifying malicious or legitimate domains. Generalize this to any use case where you need to make quick programmatic decisions, and you unlock powerful agentic workflows that are able to autonomously decide, reason, and execute.

In music theory, a clef is a symbol placed at the beginning of a musical staff that assigns specific pitch names to the lines and spaces. A decision model is analogous to a music clef because it helps define the domain of the context and the subsequent notes (actions) that follow it. We chose Clef as the name of our family of decision models, as it serves similar purposes, and the CF hearkens to Cloudflare.

How is Clef different from other decision models?

Although the market is getting increasingly saturated with decision models, Clef has some unique properties that make us excited to release it to the public. First, it has a vision encoder so it’s able to take in images and classify visual content. This is different from Jev, which only does text classification today. Secondly, our model has a 64k context window (compared to Jev’s 32k), which allows users to squeeze more input state for the model to classify against.

Third, our model is accurate and powerful, scoring competitively against other decision models on the market across various quality benchmarks. We shortlisted some evaluations below that are important for decision-making as defined by the Jev Decision Index and scored some of the more popular models on the market for it. Check out the table below for benchmarks, or view the scores on our live decision index demo site:

Benchmark

Clef

Clef-flash

Jev

DiffusionGemma Jev

Kev 9B

Laya

BFCL · case exact

98.47

98.76

95.75

96.52

94.51

38.13

ToolRet · nDCG@10

69.19

66.43

65.28

61.21

64.26

12.69

API-Bank · accuracy

91.93

93.11

88.19

83.66

56.30

11.41

Home appliances · case exact

82.95

97.73

52.27

42.05

25.00

0.00

When2Call · accuracy

72.37

65.58

80.97

75.44

49.62

11.94

BANKING77 · macro-F1

94.20

90.93

79.74

74.28

84.83

14.29

CLINC150+OOS · macro-F1

97.43

66.77

89.27

83.49

79.03

3.19

BRIGHT · nDCG@10

45.91

39.26

47.52

42.94

38.53

19.90

Amazon ESCI · macro-F1

57.48

57.39

55.21

53.37

49.22

24.40

PhishNChips · accuracy

79.60

75.05

62.55

85.35

50.75

50.15

We also ran benchmarks across Typesafe’s own eval suite and our Clef models fared well, beating Jev in 3 out of 4 areas. Notably, our Clef-flash performs exceptionally well, given how much faster it is.

Workflow

Clef

Clef-flash

Jev

Invoice processing

64.7

57.1

61.8

Customer service

76.3

77

76.0

Security incidents

62.9

61.7

61.7

Agent trace observability

68.5

69.8

71.6

Across the 43 eval benchmarks that we ran, our Clef models beat the decision models on latency (except for Laya which is very fast but trades off quality in the benchmarks above):

Benchmark

Clef

Clef-flash

Jev

DiffusionGemma Jev

Kev-9B

Laya

Median latency  · ms

209.3

38.8

524.1

84.4

51.4

5.8

p95 latency · ms

238.6

122.4

536.0

211.2

187.9

222.5

On top of the latency benefits from the model itself, our Clef models are hosted on Workers AI. Because they are hosted on Cloudflare’s infrastructure, we’re able to take advantage of our GPUs at the edge, leading to low network latency and faster decisions. This means that you could put Clef into the hot path for agents to make decisions and combine that with one of our LLMs on Workers AI to take action. 

Clef also produces strictly typed outputs similar to Jev and is fully API-compatible, so you can make the swap extremely easily. The larger Clef model is your more powerful precision model, while the Clef-Flash model is great for latency-critical decisions. The models are enterprise-ready with our guarantee that we don’t read, store, or train on your requests or responses (unless you want to use our fine-tuning product, which we go into below). You can get started with the Clef models today, starting with our developer documentation or play around with the open-source model on the Hugging Face repo.

If you’d like help tuning Clef for a specific workload, we are also offering fine-tuning services — first as a hands-on partner with our forward-deployed engineer (FDE) team, and then later as a self-serve fine-tuning platform for customers to train and redeploy the model onto Cloudflare.

How we trained Clef

In the same week that Jev came out, we posted about some experiments we had with our own homegrown decision model. Our demo goes into how we adapted the DiffusionGemma model to output deterministic probabilities by exposing the logprobs that are generated by a large language model. Our initial approach built upon independent research by Matt Mastracci, who has been active in the machine learning (ML) community with sharing new ideas and pull requests to vLLM inference engine to make DiffusionGemma support stronger.

Clef builds upon this concept, but uses a different base model as the backbone. We currently use Qwen as the base model and post-trained it to suit decision model use cases. During inference, Clef uses Qwen for a prefill-only pass, then scores the valid schema choices in parallel. The decision step is non-autoregressive, so there’s no intermediate text to generate token by token, making Clef significantly faster than autoregressive LLMs. Rather than generating intermediate text to produce structured answers, Clef and Clef-flash derive schema choices directly from internal backbone representations. This approach relies on a specialized two-stage attention routing process: every valid choice extracts context relevant to the prompt, allowing individual field parameters to cross-attend with other fields and back to the original payload prior to scoring. By leveraging a lexical prior, the model preserves semantic intent across options. Ultimately, the architecture unites option-specific evidence routing, joint cross-field attention, and schema-bound scoring.

By freezing Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, we jointly optimized the routing head alongside rank-256 low-rank adapters. Our post-training utilizes label-smoothed cross-entropy for valid schema outputs paired with a Brier loss to refine probability calibration. This training leverages our own internal synthetic datasets permutating field orders, prompts, and schema structures. We also developed Reinforcement Learning for Calibrated Decisions (RLCD) to serve as a secondary optimization target, granting partial credit to adjacent ordinal choices, rewarding fully precise record outputs, and applying a reference penalty to prevent distribution shift, giving us better accuracy and generalization.

This means that we were able to achieve a few novel things with Clef: we improved accuracy of the model in classification, constrained it to output only probabilities instead of text generation, and made it faster than Jev and the base Qwen models.

How fine-tuning can extend the capabilities of Clef

We heard a lot of internal use cases that required fine-tuning our Clef model to be built into our agentic workflows at Cloudflare. For example, internal teams want a classifier model to be able to evaluate Trust & Safety submissions, help us triage Cloudflare Support requests, or even to be built-in to our Bot products to decide if a crawler is a good bot or bad bot.

These use cases are incredibly specific and we have had many years of labelled decisions that we could use to train a specific classifier. When you fine-tune a model, you may give up some general purpose performance in exchange for higher accuracy in a specific domain.. Because Cloudflare has more than 15 years of network data across different domains, we can fine-tune a model to fit these specific use cases which is more accurate and faster than our generic Clef model. We’re working with internal teams already to figure out how we can post-train Clef to create powerful ML models that boost our impact and improve workflows across Cloudflare. These internal teams and use cases are the next remit of our new FDE fine-tuning team and basis for our reinforcement learning (RL) product.

Our new RL service

We are offering a service to help customers fine-tune Clef to suit their workloads with our hands-on FDE team. From that, we’ll learn from our hands-on experiences to build a self-serve platform that customers can use to capture data, fine-tune, and redeploy the model, all on Cloudflare.

This has actually been a long time coming — we’ve been building our AI platform to have the right primitives where we could be building a custom RL product. The interest in Jev shows the need for a fast, small, specific, classifier model, and we chose this to be our niche to start experimenting with RL environments.

To do this, we leverage the primitives that we already have built on our Cloudflare platform:

  • Cloudflare AI Gateway – pass all your AI traffic through AI Gateway and automatically create a dataset of requests for your use case
  • Cloudflare Workers AI – generate rollouts against the base Clef model
  • Cloudflare Containers – RL sandbox for scoring and replaying agent actions
  • [NEW] Trainer – update weights of fine-tuned Clef model
  • Cloudflare Workers AI + BYO Model – redeploy the fine-tuned model on Workers AI

This combines a few work-in-progress pieces of the AI Platform that we’ve been working on, including AI Gateway that captures your AI traffic so you can leverage your own request/response data, Containers for RL Sandboxes, and Workers AI’s Bring Your Own Model (Cog) work that has been progressing since our acquisition of Replicate. 

Try it out today

We’re excited to launch our first Cloudflare-trained ML model from the Workers AI team today. We’re still early here and have a lot more improvements in store, but it is a wonderful first showcase of the hard work we’ve been doing on the AI Platform team. We believe that Clef has the ability to disrupt the way we use agents, which fits naturally into Cloudflare’s mission of being the agent cloud.

If you have specific use cases and are already customers of these products — we’d love to chat with you and be design partners as we experiment in this space.

Try out the Clef models hosted on Workers AI, download the weights on Hugging Face if you’d like to explore for yourself, and reach out if you have fine-tuning use cases you’d like us to help with.

Our ML team has been growing in impact, from model optimizations to model training research. If you’re interested in joining our mission, check out our open roles. 

We want you to build the next Git platform on Cloudflare

Post Syndicated from Dina Kozlov original https://blog.cloudflare.com/next-git-platform-on-cloudflare/

GitHub was built for a world where humans write code, organize it into repositories, and collaborate through branches, commits, issues, and pull requests.

But the next generation of software is going to be built differently because it is going to be built by a different kind of developer: agents.

Agents are already writing more code than ever before — they’re fixing bugs, building features, writing tests, reviewing changes, updating dependencies, and doing the routine maintenance required to keep an application running.

So in this new world where you have hundreds, or even thousands, of agents working on the same codebase at the same time, what does the foundation look like?

How do agents know what other agents are working on? What happens when they make conflicting changes? How do you review everything they produce? How do you keep track of not just what changed, but why a change was made?

And so the burning question is: What does the next GitHub look like?

We want you to help us answer it, by building it out.

Earlier this year, we launched Artifacts, a versioned filesystem that speaks Git and can scale to millions of repositories. From the start, we designed Artifacts as a set of programmable primitives that developers could use to build their own products, workflows, and abstractions.

Artifacts provides the foundation: repositories that can be created and forked programmatically, versioned storage for code and agent context, and the Git operations agents already know how to use.

With that foundation in place, you can focus on the layer above it: how agents coordinate their work, how changes are reviewed and merged, and what the developer experience should look like when hundreds or thousands of agents are working on the same codebase.

That is the layer we want you to build.

Now that Artifacts is in open beta, we’re holding a competition to see who can build the next Git platform on Cloudflare using Workers and Artifacts.

Artifacts is in open beta. Here’s why you should build on it

When we launched Artifacts, our goal was to make it possible to create a repository for every agent, session, task, or user — and to do that at the scale agents require.

Since then, we’ve seen developers use Artifacts in a range of ways: Vibe-coding platforms are using it to store the projects their users create. Developers are using it to persist the code and context from agent sessions. Others are creating isolated repositories, so multiple agents can safely work from the same starting point and compare or merge the results later.

Here are some new capabilities we’ve added since the initial launch.

Deploy Artifacts repos to Workers

You can now connect an Artifacts repository to a Worker through Workers Builds. When you or an agent pushes code to the Artifacts repository, Cloudflare will build the project and, for the production branch, deploy the updated Worker. Pushes to other branches automatically create or update Workers Previews, giving you an isolated, shareable version of your Worker where you can test changes before they go live.

You can connect an existing Worker to an Artifacts repository or start a new project and automatically store it in Artifacts.

Manage Artifacts directly from Workers

You can interact with Artifacts repositories directly from a Worker using an Artifacts binding to create or fork repos, inspect files and commits, and issue repo-scoped Git tokens. This makes your Git workflow programmable. When a new task arrives, a Worker can fork the project for an agent, read the files it needs for context, and give it a repository to work in. When the agent pushes a change, your automation can inspect the result and start a review. You define those steps in code to fit how your agents work.

For example, here’s how to fork a project for a new agent task and read its AGENTS.md for instructions:

React to every change with event subscriptions

Artifacts publishes events whenever a repository is created, imported, forked, deleted, pushed to, cloned, or fetched. You can subscribe to these events to decide what happens next: run CI, kick off a code review agent, or deploy a change.

For example, you can subscribe to Artifacts push events and have a Worker start a code review workflow for each push. The Worker passes the repository, branch, and new commit to the Workflow, giving a review agent the context it needs to inspect the change:

Data jurisdiction for Artifacts repos

You can now choose where Artifacts stores and processes your repository data. Set a U.S. or EU jurisdiction when you create a namespace, and every repository created in that namespace will automatically follow the same restriction.

View Artifacts metrics

You can now see metrics for your Artifacts repositories in the Cloudflare dashboard. For each repository, you can now see total operations, pulls, pushes, errors, and error rate, helping you understand how the repository is being used and spot failures. You can also query Artifacts metrics directly to build your own dashboards or monitoring.

Pricing

Artifacts pricing is based on repository operations and the amount of data stored. We will begin billing for Artifacts usage on October 15, 2026.

Competition: Build the next Git platform on Cloudflare

We want you to build your vision for the Git platform of the agentic era using Cloudflare Workers and Artifacts.

You could rethink repositories, branches, pull requests, worktrees, code review, and merge conflicts — or build new ways to preserve agent context, compare multiple changes at the same time, and decide which one should ship.

We aren’t looking for GitHub as it exists today with agents added on top. At a minimum, we want to see multiple agents working on changes concurrently. Beyond that, we want you to get creative — what you think comes next.

How to enter

Submit:

  • A 5-10 minute video demonstrating what you built, what it enables agents and developers to do, and how it works
  • A link to the source code, which must be provided under a permissive open source license (MIT, Apache, BSD)
  • Instructions for running or trying the project

Deadline

Submissions are open until October 14, 2026.

Why should you participate?

We’ll select the top three projects and fly up to two members from each team to San Francisco to attend Cloudflare Connect and show what they built.

The first-place team will also receive $25,000 in Cloudflare credits, along with invitations to the VIP speaker dinner on Monday night at Connect.

Get started

Artifacts is available in open beta to customers on the Workers Paid plan.

Get started with your coding agent: copy the prompt below to set up your first Artifacts repository and start pushing code to it.

You can view or create the Artifacts repositories in the dashboard or if you’re looking to learn more, check out the documentation.

Cloudflare OS: your company’s agent workspace, managed for you

Post Syndicated from Phillip Jones original https://blog.cloudflare.com/managed-cloudflare-os/

Cloudflare OS gives everyone in your organization an agent workspace that knows how your company works and connects to its data and systems. Today, we're opening the waitlist for fully managed Cloudflare OS deployments.

If I asked you to prepare for an important customer meeting later today, what would you do? You might learn how your company typically runs customer meetings, review the account in your CRM, check recent support tickets and product usage, then turn it into a short presentation to review with the group. Now imagine doing that another 100 times this month.

Every team has work like this. With Cloudflare OS, you can ask your agent to handle the work for you, build a tool for your team, or move between the two as the work evolves.

Last month, we announced Cloudflare OS and shared the open source repository. Since then, thousands of organizations have started using it to work with company data, produce docs and slides, build tools for their teams, and automate work with agents.

With a few clicks in the Cloudflare dashboard, you’ll be able to launch your organization’s own agent workspace. Just tell us what custom domain you want to use, what Cloudflare Access policies apply, and which AI Gateway to connect. We’ll handle the rest.

Cloudflare OS, managed for you

Every company has its own terminology, procedures, systems, and requirements. We made Cloudflare OS open source so you can customize it around how your company works.

You can already deploy Cloudflare OS into your own Cloudflare account from the open-source repository. That gives you full control, but it also means someone has to configure the deployment, operate it, and keep it up to date.

With the fully managed option, you decide who can access Cloudflare OS, which organizational skills and context are available, and which systems it can reach. You can leave the rest to us.

If you want Cloudflare OS fully managed for your organization, join the waitlist and we’ll reach out.

What’s new in Cloudflare OS

We’ve also spent the last month expanding what people and agents can do in Cloudflare OS. Here are a few highlights.

Mount Git repos and work with code

When we launched Cloudflare OS, we focused first on work outside software development: creating documents and slides, automating tasks, and building collaborative tools. Agents could write code for an app, but they could not work with code in an existing Git repository.

You can now connect an existing GitHub repository to Cloudflare OS. Ask your agent to explore the codebase, fix a bug, add a feature, or open a pull request. It can search and edit files, review its changes, create commits, and push them to GitHub.

Work across Google Workspace

For many organizations, work starts and ends in Google Workspace. Decisions live in email threads, context lives in Google Drive, analysis happens in Sheets, and teams coordinate through Calendar. Agents need to do work across those systems too.

We’ve made significant improvements to the Google Workspace Gatekeeper (a service-specific Worker that sits between Cloudflare OS and an external service). Cloudflare OS can now read and research Gmail threads, create drafts, and send emails. You can also connect your entire Google Drive, a specific folder, or an individual doc or sheet.

Export work in the formats your team uses

Work often needs to move into the formats your team already uses. Finance may need an Excel spreadsheet, and a report may need to become a PDF before sending to a customer.

The built-in document, presentation, and spreadsheet experiences can now export work to familiar formats. Depending on what you create, you can export to Microsoft Excel (.xlsx), CSV, PDF, Markdown, or HTML. Microsoft Word (.docx) and PowerPoint (.pptx) export is coming soon.

Tools you build can also define their own export formats. Tell the agent what you need, like “let me download this schedule as a calendar file (.ics)”, and it’ll add the option to the tool’s export menu.

Sign up for the waitlist

Cloudflare OS is open source and available today. You can check out the source code or deploy it into your own Cloudflare account.

If you want Cloudflare OS fully managed for your organization, join the waitlist and we’ll reach out with more information.

AI Search is now generally available

Post Syndicated from Gabriel Massadas original https://blog.cloudflare.com/ai-search-ga/

Cloudflare’s AI Search combines Workers AI, Vectorize, R2, and Browser Run into a fully managed index and retrieval pipeline. Since we launched AI Search over a year ago, we’ve seen developers use it to power a wide range of search use cases, from searching internal documentation to powering search for their websites. We use AI Search ourselves to power search on our own blog and developer docs.

Starting today, AI Search is generally available. And as part of it, we've expanded and improved our support for multimodal formats beyond text, adding native image embeddings, optical character recognition (OCR) for PDFs, and support for larger files.

As part of general availability, we’ll start billing for AI Search on November 1, 2026, and continue to offer a generous free tier on all Workers plans.

New: multimodal embedding and retrieval

An image is more than the sentence used to describe it. Product texture, screenshot state, chart relationships, document layout, and fine visual detail can all disappear when pixels are compressed into a caption.

AI Search now preserves both signals: it embeds image pixels directly for visual retrieval while retaining captions for textual understanding. To keep these richer representations efficient, AI Search leverages Matryoshka Representation Learning (MRL), allowing smaller embeddings to retain useful information while keeping storage manageable and search fast.

Although we originally supported retrieval over images, the implementation was naive: we would perform object detection, generate a caption, and then embed that text. This made images searchable, but only through the details captured in the caption. Now, we do both — caption-based understanding and native image retrieval.

Native multimodal retrieval is available today with the Qwen3-VL-Embedding model. At query time, AI Search checks whether your instance’s embedding model supports images. If it does, a query image is embedded directly by that model, landing in the same vector space as your indexed images and text.

If your embedding model is text-only, you can still query with an image. AI Search converts the query image to text with ToMarkdown and searches using the resulting caption. This gives every model basic multimodal support, while models with native image support get the full visual signal.

Caption-based understanding

A small bird with black-and-white markings perched among golden fruit and green leaves

Native Image Retrieval

Can match details omitted from the caption: the geometry of the bird’s white eyebrow stripe, its yellow-green plumage, the mixture of smooth and weathered fruit, the leaves’ deeply ribbed texture, and the image’s warm palette and shallow-focus composition.

The caption is a compressed interpretation of the image. Capturing every potentially useful detail requires long or specialized captions, written with the eventual search query in mind. Native image embeddings preserve visual characteristics without requiring the caption to anticipate which details matter.

This enables searches that are difficult to express precisely with words. You can describe an image you want to find, provide another image to locate visually similar results, or combine both, such as “a bird with similar markings” or “a bird perched on a leafy branch with plums.” It is useful for product discovery, screenshot matching, charts, diagrams, scanned documents, and other collections where color, texture, composition, or spatial relationships matter.

Here’s how a query moves through AI Search. First, the query is optionally rewritten, then embedded (image queries are embedded directly by multimodal models, or captioned first by text-only models). Vector and keyword search run in parallel, and results are fused and optionally reranked. The top chunks are returned, or passed to a generation model to write an answer.

Bigger files and OCR for scanned documents

AI Search now accepts your text files (Markdown, HTML, CSV, JSON and similar) and PDFs up to 10 MiB, up from 4 MiB. Many PDFs are really scanned images with no extractable text. For those, turn on OCR and AI Search reads the text from each page before chunking and embedding it. OCR is available to every account and is billed under the new AI Search pricing as image processing ingestion tokens.

Now in GA: billing and pricing for AI Search

During our August 2026 Agents Week, we announced preview pricing for AI Search. With the product going GA today, we’re announcing billing for AI Search that is going live on November 1, 2026. We’ll send a reminder email before billing is enabled.

AI Search pricing is designed, so you can estimate your bill before you index a single file. You pay for three things: the content you ingest, the data you store, and the queries you run. The work in between (parsing, chunking, embedding with Workers AI models, keyword indexing, and reranking) is included. There are no instance hours, capacity units, or monthly minimums to size up front, and small projects fit inside the free monthly allotment.

Ingestion pricing is based on one rate per token with whichever Workers AI embedding model you pick, and tokens are counted the same way for every model. Switching from a text-only embedding model to a multimodal one doesn't change what you pay to ingest, unless you are also processing images (add-on fee). Storage is priced on the size of data in your indices. Querying is priced based on the type of query (semantic vs. full-text) and how many queries you send.

Estimating your cost comes down to how much content you index, how much you store, and how many queries you expect. Here's the pricing we announced in preview, with one tweak that we’re making: on the free monthly allotment, you will receive 1,000 semantic queries and 1,000 full-text queries (instead of a shared pool of 2,000 queries).

Pricing

Free monthly allotment (all Workers plans)

Ingestion

Base Ingestion

$0.75 / 1M tokens

5M tokens †

Image processing (add-on)

+$0.50 / 1M tokens

5M tokens †

Storage

Stored data

$2.00 / GB-month

10 GB

Query

Semantic (hybrid and vector search)

$0.75 / 1k queries

1,000 queries

Full-text

$0.10 / 1k queries

1,000 queries

Embedding and Reranking

Ingestion and query

Free with select Workers AI models; third-party billed separately

N/A

† A single pool of 5M ingestion tokens per month, covering any file type currently supported (e.g., text, images).

What’s next

Multimodal embedding support is just the first step; we’re building an ingestion pipeline to support full video and audio processing to allow our customers to search their rich media assets.

We’re also refactoring the keyword search engine so that it scales better with your content requirements, particularly for when you have big data stores where the current implementation has limits.

Finally, we’re developing better and simpler ways to enable AI Search and create indexes for websites already running on Cloudflare. This ensures AI agents can discover, explore, and consume content more easily and efficiently.

Stay tuned for these follow-up announcements and more.

AI Search is now generally available to enable and use today. Get started with our new multimodal embeddings, bigger files, OCR features, and our new managed instance pricing. Check out the AI Search developer docs for more information.

Introducing Cloudflare Basin: an open, serverless data platform, now generally available

Post Syndicated from Marc Selwan original https://blog.cloudflare.com/cloudflare-basin/

During Birthday Week 2025, we announced the Cloudflare Data Platform, a suite of products that ingest, store, and query your analytical data. Today, we’re announcing that the platform is generally available, and we’re giving it a new name: Cloudflare Basin.

Basin is a serverless data analytics platform built on Apache Iceberg, the open standard for data lakes, and R2 Object Storage. The Basin family includes:

  • Basin Pipelines, formerly Cloudflare Pipelines, receives events from Workers, HTTP, or Cloudflare Logpush, transforms them with SQL, and writes them as Apache Iceberg tables or files in R2.
  • Basin Catalog, formerly R2 Data Catalog, manages Iceberg metadata and automatically maintains tables to keep them fast and cost-efficient.
  • Basin SQL, formerly R2 SQL, is our serverless, distributed SQL engine for querying Apache Iceberg tables directly on Cloudflare.

Basin brings an end-to-end analytics platform to the Developer Platform, enabling you to collect data from a variety of sources, such as apps, infrastructure, devices, and other Cloudflare services, then query it to answer analytical questions.

We set out to build a data platform last year when we saw two fundamental developments that changed how modern data applications were being built. First, Apache Iceberg emerged as the standard open table format, making data portable across nearly every major query engine. Second, we started seeing developers bring their analytics data to R2 where the lack of egress charges made it practical and cost-efficient to actually access their data from different tools, teams, regions, and cloud providers.

So, when we launched Basin in open beta, developers — including our billing and infrastructure teams at Cloudflare — immediately started adopting these services for a variety of use cases including using real-time data to optimize e-commerce sites, long-term storage and reporting of billing metrics, and ingesting and querying telemetry from Cloudflare’s infrastructure to measure and improve utilization and efficiency.

“We moved our entire company's data pipeline to Basin Pipelines, Catalog, and SQL, replacing a complex AWS S3 and Athena setup with a cleaner, serverless architecture that reliably handles all of our event data. -Dax Raad, Co-Founder, Anomaly

Our early adopters taught us a great deal about what it means to run an analytics platform on the edge. We spent the past year improving Basin around three specific areas that hone in on what makes analytics in Developer Platform unique: speed, openness, and cost efficiency.

Basin is built for speed — whether it’s about getting started or executing large queries. You can create a Basin Catalog, set up a Pipeline to ingest data, and query it with Basin SQL in seconds. This matters as we see more data applications being built from prompts to coding agents, which would otherwise have to wait and poll for resources or data. As datasets grow, Basin Catalog compacts metadata and data files to reduce I/O and generates statistics for query planning. Basin SQL uses those statistics to split queries into smaller tasks and distribute them across Workers, keeping queries fast and consistent as datasets scale.

A large driving force in the development and adoption of Basin is the continued growth we are seeing in the open Apache Iceberg ecosystem. Developers have been rallying around the radical idea that you should own and be in control of your own data — separating the storage layer from the compute layer and allowing you to use the right query engine for the job. With Basin, you can read and write your data using any Iceberg-compatible engine, including PyIceberg, DuckDB, Snowflake, and Apache Spark. That kind of data portability is only possible with free egress, which allows developers to access their data in Cloudflare from the wide variety of tools in the ecosystem, regardless of region or cloud.

“Bobsled is a data product platform that the world's most advanced data teams use to build and distribute AI-ready data to partners, vendors and customers," said Julien Grobbelaar, Head of Platform at Bobsled. "Basin allows us to build data products that can be made accessible in any region of every major data and AI platform, all at production-grade reliability and a fraction of the cost thanks to zero egress fees.” 

In addition to free egress, our serverless architecture allows us to offer further cost efficiencies to our customers with usage-based pricing. You are only billed when Basin ingests, processes, or queries your data. Developers can build out analytics for hobby projects at little to no cost, while our pricing scales economically for larger enterprise use cases. There are no hourly charges or separate infrastructure costs to worry about.

If you are ready to get started, refer to the Basin tutorial for a step-by-step guide on how to use Basin Pipelines to deliver events to an Apache Iceberg table managed by Basin Catalog, and query them with Basin SQL. Read on to learn more about Basin and where we are going next.

Why Basin?

We launched these products last year as the Cloudflare Data Platform, which has served us well for the first year of availability. For our GA launch, we decided we needed a new name that encompasses our ambitions for the platform, links together all the products, and is a bit punchier.

A basin is where rivers from many sources come together to a single point. We felt that Basin perfectly captures how the platform is used: Pipelines brings data into Basin Catalog while Basin SQL makes it instantly queryable. A fun fact: roughly 20% of Earth’s land drains into endorheic basins, much like over 20% of the web sits behind Cloudflare’s network. 

Today, Basin is made up of three products: Pipelines, Catalog, and SQL, covering ingestion, storage, and querying, and will expand over time with more products managing the rest of the analytical data lifecycle.

Basin Pipelines

Before you can query your data, your events need to be ingested, structured to a schema, and written to object storage. This is the role of Basin Pipelines. It accepts events through HTTP endpoints or Workers bindings, processes them according to a SQL query, and delivers them to Basin Catalog as Apache Iceberg tables or R2 as JSON or Parquet files.

Since our beta launch, users have created tens of thousands of Pipelines for a wide variety of use cases. For example, a common pattern we see is using Pipelines to transform Cloudflare HTTP logs before storing them:

Doing this work during ingestion can significantly reduce the storage footprint, reduce noise from dynamic data sources, and can help prevent sensitive or unnecessary values from being written.

Since the beta, we have greatly expanded the scalability of Pipelines: we now support ingesting up to 3GB/s per stream. We’ve also expanded the feature set and integration with other Cloudflare systems:

  • Cloudflare Logpush integration: you can transform Cloudflare logs with SQL and store them as compressed Parquet files or Iceberg tables, ready to query with Basin SQL or another engine.
  • Worker bindings are schema-aware. Running wrangler types generates TypeScript types from a stream's schema, catching missing fields and type mismatches before deployment.
  • Data quality errors are visible. The dashboard and GraphQL API surface dropped events and distinguish missing fields, type mismatches, parse failures, and null values.
  • The entire ingestion path can be infrastructure as code. Terraform resources cover the catalog, stream, sink, and the SQL that connects them.

Next we plan to expand Pipelines capabilities even further, including:

  • Custom partitioning when writing to Basin Catalog
  • Schema migrations, and updatable configuration and Pipelines SQL
  • Support for Iceberg V3, including the Variant type for efficient querying of semi-structured data
  • Stateful processing to support workloads such as streaming aggregations, joins, and incrementally updated materialized views

Basin Catalog

Basin Catalog was the first product we launched in the family last year. Since then, we’ve seen thousands of developers use Basin Catalog for simple use cases such as giving DuckDB a structured way to access analytics data in R2, all the way to developers building complete enterprise data sharing platforms, fully taking advantage of zero egress fees and easy-to-use APIs.

Basin Catalog is the easiest way to get started with Apache Iceberg. Just run:

You instantly get a fully managed Apache Iceberg REST catalog that automatically performs routine maintenance required to keep those tables performant and healthy.

When we announced the Data Platform, Basin Catalog had just added automatic compaction. Since then, it has evolved to maintain healthy tables as your data scales:

  • Per-table compaction policies let you choose target file sizes based on each table's access pattern.
  • Automatic snapshot expiration removes old Iceberg snapshots according to a retention policy, while preserving a minimum number of recent snapshots.
  • Unreferenced data-file cleanup reclaims storage when snapshots expire, without requiring a separate Spark maintenance job.
  • Manifest optimization consolidates and clusters fragmented manifests by partition before compaction, reducing metadata I/O during query planning.

We have some exciting features in the works for Basin Catalog including:

  • A new way for compaction to efficiently sort and cluster data for improved query performance
  • More granular auth controls for namespaces and tables
  • Jurisdiction support to adhere to data sovereignty and compliance requirements

Basin SQL

Basin SQL is our serverless, distributed query engine for Apache Iceberg tables stored in Basin Catalog. It’s designed for reading large datasets and automatically scales across Cloudflare's global network. There are no clusters or resources to provision, just a readily available API for you and your agents to immediately start querying your data.

At beta launch, Basin SQL was great at filtering and exploring large event and time-series tables. Over the last year, Basin SQL has evolved to support hundreds of functions including:

  • Standard and approximate aggregations, GROUP BY, HAVING, and schema-discovery commands
  • More than 190 scalar and aggregate functions across strings, timestamps, regular expressions, cryptography, statistics, arrays, maps, and structs
  • CASE expressions, common table expressions, casting, arithmetic, and EXPLAIN
  • Inner, outer, semi, and anti joins; subqueries; self-joins; and multi-table queries
  • DISTINCT, UNION, INTERSECT, and EXCEPT
  • Window functions, QUALIFY, grouping sets, rollups, and cubes
  • A suite of JSON functions

Suppose your Pipeline delivers application events into one table and account data into another. You can now join those tables, aggregate activity by customer, rank the results with a window function, and filter the ranking in one query:

You can run Basin SQL from Wrangler or the API, or open the built-in editor in the Cloudflare dashboard. The editor provides syntax highlighting and autocomplete, a browser for namespaces and tables, query statistics and plans, and exportable results. It makes the path from a new table to a useful answer a matter of seconds.

The team isn’t stopping here and is currently working on:

  • Advanced statistics and adaptive scheduling to improve performance and efficiency of queries
  • Full data definition language (DDL) support directly from Basin SQL
  • Iceberg V3 support including support for the VARIANT and geospatial types

What comes next

Our future vision is that data infrastructure is completely abstracted away. Storage formats, products, and resources are just implementation details — important ones that help enable the important outcomes — but tend to get in the way. We’re building towards a platform where developers start with questions rather than CREATE statements or CLI commands. We’ve laid the foundation for that vision, and now we’re building towards that vision including:

  • Support for the latest Apache Iceberg spec across the entire platform, unlocking more flexible ways to use your data
  • Push-button ingestion sources and destinations, with zero-configuration connections across Cloudflare's developer and observability products
  • Advanced adaptive table-maintenance strategies in Basin Catalog that automatically organize data around real query patterns
  • Continued expansion of SQL compatibility, performance, and observability for increasingly complex analytical workloads
  • More ways to continuously process data in real-time and trigger actions based on the signals within the data
  • Tools for adhering to data compliance and sovereignty rules across the platform

We will continue to build with open standards: using open formats and protocols, contributing improvements to the projects we depend on, and making sure your data remains available to the broader ecosystem.

Get started

Basin Pipelines, Basin Catalog, and Basin SQL are generally available today. You can use them together as an end-to-end platform or adopt the parts that fit your existing architecture.

Follow the getting started tutorial to ingest events, create an Apache Iceberg table in Basin Catalog, and query it with Basin SQL. Visit the Basin documentation for product guides, pricing, limits, and integrations.

Existing Cloudflare Pipelines, R2 Data Catalog, and R2 SQL configurations will continue to work.

We are excited to see what you build. Share your feedback with us in the Cloudflare Developer Discord.

Identify AI model overuse with User Insights

Post Syndicated from Ayush Kumar original https://blog.cloudflare.com/ai-model-overuse-user-insights/

When we launched User Insights last month, we wanted to help teams answer a basic question: What are people actually doing with AI? User Insights gives teams a clearer view of their AI usage, showing which users, applications, tasks, and models are driving traffic. It also highlights user and agent anomalies, helping teams identify unexpected or out-of-control spending and usage before they become larger problems.

Our latest update adds something our users have been asking for: context. 

Since launch, we’ve heard from users that model names and request counts only tell part of the story. They show where traffic is going, but reveal little about the work behind it: is that request a code review, a research task, or an agent making several calls to complete a job? The same token count can represent very different kinds of work, and you can’t evaluate with model choice without understanding the task.

User Insights now shows when a model may be more capable than a task requires, which of your users and agents are driving that usage, and how the task, model, cost, and conversation patterns relate. Teams can use these insights to investigate and make targeted changes within their organization. These capabilities are available for free to AI Gateway users.

Why AI usage is hard to understand

Consider a team that has routed its internal AI traffic through AI Gateway. After a few weeks, spending is increasing and some requests feel slower than expected, a common challenge as organizations adopt AI at scale.

There could be several explanations. Developers may be using AI for increasingly complex coding work. Agents may be making too many follow-up calls to complete a task. Or a small group of users or agents may be responsible for a disproportionate share of the organization’s usage.

Tokens and request counts alone cannot show which pattern is driving the increase. Teams need to understand what the traffic represents before deciding whether a model, workflow, or routing rule should change.

Helping teams find where AI models are overkill 

The model overkill view helps teams identify conversations where the selected model appears to be more capable than the task requires. For example, a team might discover that users or agents are sending simple formatting or summarization requests to a high-capability reasoning model.

That gives the organization a place to start. They can see which users, agents, or applications are associated with the pattern, then investigate the tasks behind it. A team might find that a model is being used because it is the default, because users are unsure which model to choose, or because an agent has been configured to use the same model for every step.

The overkill view is not a leaderboard and does not automatically recommend a replacement model. It helps teams ask better questions:

  • Is this model appropriate for the task?
  • Is the extra capability improving the result?
  • Would a faster or less expensive model produce an equivalent outcome?
  • Is the issue limited to one workflow, user, or agent?

From there, teams can compare cost, latency, token usage, and conversation turns before deciding what to change.

These insights support both the new Potential Savings view and the Auto Router, which is launching in public beta alongside this release. The Potential Savings view helps teams identify requests that may be handled by a faster or less expensive model without compromising output quality. The Auto Router applies these task and model-fit signals automatically, helping reduce costs without requiring a separate routing rule for every workload.

The Overkill view is a starting point for evaluating model fit. Teams can compare latency, input and output tokens, conversation turns, and total cost for the same type of task. A difficult coding or research task may need a capable reasoning model, while a short summary or simple classification task may not. The goal is not to move every request to the least expensive model, but to understand whether the selected model is appropriate for the work.

Understand what people are using AI for

Task analysis groups conversations by the kind of work they represent. Initial categories include coding, research, writing, summarization, and data analysis.

This provides context that a list of model names cannot. An engineering team might use AI mostly for coding and debugging, while another team might use it for research and summarization. A team may also discover that a surprising amount of traffic comes from simple tasks, even though those tasks are being sent to a high-capability model.

The answers will vary by team. The category data provides a way to investigate those differences using traffic already passing through AI Gateway. Teams can determine whether a model is being used for the work it is best suited to handle, or whether a default model is being applied too broadly.

Understand the full cost of a task

Some tasks are finished in one exchange. Others take a few rounds of questions, corrections, and follow-ups. Turns analysis shows how much back-and-forth different tasks require. A long conversation is not necessarily a bad thing, especially for complex work. But if a simple task keeps taking several turns, it may be worth looking at the prompt, the model, or the workflow.

The first request is only part of the cost. Teams should also look at the time, tokens, and money spent before the task is finished. Comparing those numbers can show where a workflow is taking longer or costing more than expected.

Turn insights into auto routing

Once a team has identified an overkill pattern and confirmed it across task, cost, latency, and turn data, it can turn that insight into an automatic routing decision.

For example, the task view might show that much of the team’s AI usage is summarization and formatting. The model view could show that those requests are being sent to a large reasoning model, while the turns view shows that most conversations finish in a single turn. Together, these signals give the team a concrete workload to evaluate.

In addition to our updates to User Insights, the Auto Router is now available in closed beta. The Auto Router uses the conversation trajectory, task category, task complexity, and model-fit signals to automatically route requests to an appropriate model while taking cost into account.

Instead of creating a separate routing rule for every workload, customers in the beta can let the Auto Router select among the models available to their application. The router does not simply send every request to the least expensive model, but instead selects an appropriate model for the task at hand. Complex coding or research work may still need a more capable model, while simpler tasks may be handled by a faster or less expensive option.

To learn more about the Auto Router and sign up for the closed beta, read the blog post here.

The Auto Router uses the same task and conversation signals that power User Insights. The section below explains how those signals are produced.

How User Insights classifies traffic

Each conversation receives an analysis signal that can be grouped in User Insights. The signal is used for reporting and routing analysis, and is not intended to replace or expose the original request.

The categorization engine is a dedicated Cloudflare Worker that processes eligible AI Gateway logs. It examines the conversation trajectory, including user requests, assistant responses, tool calls, and tool results, and identifies the type of work being performed, such as coding, debugging, research, or summarization. It also returns a confidence score and evaluates dimensions such as task complexity, intent ambiguity, stakes, and context dependence.

The Worker returns a category that can be joined with the log metadata used by the dashboard. These signals can also be used to evaluate model fit by comparing how well candidate models suit the task against their cost. The current implementation focuses on a small set of categories that are easy to understand, rather than trying to infer every detail about a user’s work.

The pipeline follows the existing AI Gateway log architecture. Metadata is stored separately from log bodies, and the current implementation uses Durable Objects for metadata and R2 for log bodies. User Insights exposes derived categories and aggregate views. It does not turn the dashboard into a raw prompt browser. Retention of the underlying log bodies continues to follow the configured AI Gateway logging behavior, so teams should review those settings when deciding what to send through the classifier.

The classification is asynchronous, which means it happens after AI Gateway has handled the request rather than while the user is waiting for a response. AI Gateway writes the log to the existing storage path first, and the classification Worker processes it afterward. This keeps classification out of the request path and adds no latency to the user’s response.

The tradeoff is that User Insights is not a real-time view. Newly received conversations may not appear in the dashboard immediately, and analysis may trail incoming traffic by approximately one day as logs are processed and aggregated. Teams should use User Insights to identify usage patterns over time rather than monitor live request activity.

The flow looks like this:

Connect usage to users, teams, and tools

Task categories become more useful when they can be viewed by user, team, or application. AI Gateway is identity-aware, providing that context without requiring teams to build a separate reporting pipeline.

This works not only for applications that teams build themselves, but also for developer tools and agent harnesses such as Claude Code, Codex, and OpenCode. By putting AI Gateway behind Cloudflare Access, teams can connect authenticated users and sessions to their AI traffic, allowing User Insights to associate activity with the right person and conversation.

For custom applications, requests must include both a stable user_id and a session_id for User Insights analysis. The exact identity configuration and field names depend on how the application or tool is set up. The important part is to provide stable, non-sensitive user and session identifiers so usage can be grouped without putting identity data in the prompt itself.

For custom applications, the request metadata might look like this:

The request body contains the model and messages for the conversation.

With Access configured in front of AI Gateway, tools such as Claude Code, Codex, and OpenCode can inherit this identity context automatically. Cloudflare Access is available at no cost for teams with up to 50 users, making it an easy way to get started.

Get started with AI Gateway User Insights

AI usage is changing quickly. Models change, teams develop new workflows, and the right choice for one group may be the wrong choice for another.

User Insights lets teams start making smarter choices by identifying where certain models may be overkill. They can then see which users and agents are driving that usage, understand the tasks behind it, and compare the cost of completing the work.

Learn more with the AI Gateway User Insights documentation . Open AI Gateway in the Cloudflare dashboard, and use what you learn to make more targeted model and routing decisions.

Simplifying domains for people and agents

Post Syndicated from Ankit Shah original https://blog.cloudflare.com/simplifying-domains/

You just thought of your next great idea, and buying the right domain feels like the easiest way to make that first bit of progress. Naturally, you open a new tab in your browser, only to find yourself face-to-face with an experience that feels like a budget airline peppering you with add-ons at checkout: Want security? How about a website? Do you want email? You’re just a few minutes into building your next idea, and it doesn’t feel fun anymore.

Launched a decade ago, Cloudflare Registrar has always taken a simpler approach. Domains at cost, transparent pricing, and no unnecessary upsells. But simplicity shouldn’t begin at checkout. It should begin the moment you start looking for the right domain.

Today, we’re bringing that same simplicity to the entire experience of finding and buying a domain. Our new domain search shows every extension we support, responds as quickly as you type, and makes hundreds of possibilities easier to explore through sorting, filtering, and transparent pricing.

And you know what is particularly good at ignoring distractions and staying focused on the destination? An AI agent. We designed Cloudflare Registrar to work naturally with agents through the Registrar API, MCP, and our newly launched cf CLI. You can ask your favorite agent to find the right domain, buy it, or transfer one you already own.

The agentic registrar, expanded

In April, we launched the Registrar API beta, allowing developers and agents to search for, check, and register domains programmatically. We have expanded the API since then. The new sandbox lets you test registrar workflows without purchasing a domain or triggering a real transaction. Our extensions endpoint returns relevant information for each of the 420+ extensions we support, helping you account for the different requirements across registries. We also added transfers, so you can bring domains from another registrar into Cloudflare programmatically.

The Registrar API is available through Cloudflare MCP, giving agents access without requiring a separate integration. Earlier this week, we also announced the launch of cf CLI, bringing the same capabilities directly into your terminal. You can prompt your favorite agent to search for, register, or transfer a domain. These tasks already lend themselves naturally to a conversation:

  • “Is example.com available?”

  • “Buy example.com.”

  • “Transfer example.com from my current registrar.”

The way people interact with the Internet is changing. Cloudflare Registrar should feel natural whether you use it through an agent or in your browser. For many people, the browser is still where the search begins, and that experience was long overdue for an overhaul.

Search simplified

Previously, our search page showed around 20 available results from a subset of extensions, sometimes modifying your search term to suggest related options. You couldn’t see that exact name across every extension we support.

We decided to take a simpler approach: show you the exact term you searched across every supported extension. Results appear as you type and continue to load as you scroll, letting you explore hundreds of options without starting another search. We also include domains that have already been registered, giving you a more complete picture. If you only want domains available to buy, you can filter everything else out.

More results might not sound simpler, but these are the results you asked for. Sorting and filtering help you narrow them down. Whether you are logged in or logged out, on your phone or at your desk, the experience feels the same.

Making a complex question feel simple

Cloudflare Registrar supports 420+ extensions. A single search can thus create more than 420 separate availability questions. Each extension is operated by a registry that maintains its official registration records and provides the authoritative answer about a domain’s availability and price.

Asking every registry every question at once would be slow and wasteful. Registries respond at different speeds and impose request limits, and much of the work would be for results the person might never view. To solve this, our search gathers evidence from multiple sources: (1) prepared availability datasets (e.g. zone files) and cached answers; (2) DNS answers; (3) live registry lookups.

A hit against an availability dataset or DNS can tell us that a domain is already in use, but a miss cannot necessarily prove availability as a domain may be registered without being configured in DNS, or it may be on a blocked list. A recent registry-derived answer is stronger but becomes stale over time, while a live registry check provides the freshest authoritative answer but takes longer and draws on limited upstream capacity. Our new search progressively probes these sources while balancing speed, freshness, and certainty for each result. As better or more accurate information arrives, we update only the affected result dynamically.

How we built search for speed and scale

We built the new search on the same Cloudflare developer platform available to our customers. Workers run the public search entry point and the services that gather availability evidence. Durable Objects give each active search one coordinator, while Workers KV stores prepared data that can be reused across searches. Together, these primitives let the service scale across users and 420+ extensions while keeping operating costs low.

We prepare useful evidence before a search begins: a purpose-built pipeline converts registry zone files and other bulk sources into compact availability datasets in Workers KV. Large datasets are split into smaller pieces, so a lookup retrieves only the data required for that domain. These fast checks can answer many questions without making a new live registry request.

A search session is composed of a Durable Object that coordinates that specific search interaction. It establishes the result order from the query, sort, and filters without waiting for network lookups, remembers the best evidence received for each domain, and tracks which results are visible so lookup work follows the person's attention.

WebSockets provide the bidirectional connection, and we designed an application protocol on top of them to connect the browser to the resolution process. An initial snapshot establishes the ordered list. Subsequent delta messages contain only the fields that changed, letting the browser update one domain instead of downloading the full result set again. Before sending a delta, the Durable Object compares the new evidence with the current answer: stronger evidence can replace it but weaker evidence cannot.

A separate Worker gathers additional evidence. It can query DNS through Cloudflare's 1.1.1.1 resolver, make a live Registrar check, or reuse a recently cached answer. Reusing fresh answers avoids repeating upstream requests. Because an available domain can be registered at any moment, an available answer has a shorter useful cache life than evidence that a domain is already taken.

Together, these pieces turn hundreds of independent availability checks and all that coordination into one coherent search experience. Cloudflare’s Developer Platform gives us all the building blocks to hide that complexity and craft a domain search experience that feels simple and is among the fastest in the world.

Transparent pricing and price drops

Making search feel simple is not only about speed. It is also about knowing exactly what a domain will cost. Cloudflare Registrar has offered domains at cost since day one. Great prices are part of making domains simple, but so is knowing what you will pay. Our new search and our new pricing page show both the initial registration price and the renewal price for every domain. When a domain is discounted, we show the original at-cost registration price crossed out alongside the promotional price.

Beginning with Birthday Week (this week!), we’re offering first-year registration discounts on select extensions including .io, .dev, .app, and .tech. You can explore every discounted extension directly from the search page as well as our newly launched pricing page.

Whether you search in your browser or ask an agent, our goal is the same. Remove the friction between having an idea and making it real. Buying a domain for your next idea should be fun and feel like progress.

Find your next domain

Choose how you want to get started:

  • Search in your browser: Explore every available extension and find your next domain.
  • Ask your favorite agent: Install cf CLI, then prompt your agent to search for, register, or transfer a domain.
  • Build with the API: Use the Registrar API to bring domain search and registration into your own application or workflow.

However you choose to do it, finding your next domain should be the fun part.

Acknowledgements: This simplicity was a result of cross-team collaboration. Special thanks to Pedro Menezes, Shobhit Kuruvilla, Lucy Dryaeva, Fred Pinto, and the Registrar Team, the Design Engineering Team, and the Forge team.

Cut your AI spend with AI Gateway’s Auto Router

Post Syndicated from Ming Lu original https://blog.cloudflare.com/auto-router/

From our conversations with companies at every stage of their AI adoption journey, we've seen some common patterns. First, there is an exploration period as you bring on every new tool, dole out API keys freely, and let the tokens flow. Then, you converge on the canonical tools for your organization for agentic coding, for non-technical workflows, for running and deploying agents. As companies formalize their AI adoption, they want to manage and oversee token spend for users, but budgets and rules only go so far. The best savings are the ones users never notice.

Today, we are releasing Cloudflare's Auto Router in public beta, available through AI Gateway. Set your model to cloudflare/auto and the Auto Router will automatically route each request to a model that is capable enough for the task, without requiring an end user to think about model selection. Our early results using the Auto Router internally through our OpenCode harness show a cost savings of up to 30% when compared to using only frontier models like OpenAI Sol and Anthropic Claude Opus.

Why we built this

From our own experience tracking AI spend at Cloudflare, we’ve learned managing costs requires a multipronged approach. Previously, we talked about how to set budgets and limits around AI spend, and how to see who is spending across your organization by linking employees to their AI usage.

In many harnesses, including OpenCode, Claude Code, and Codex, individual users still select models manually. Of course, not all tasks are created equal, and often individuals end up using models that are overkill for their work. For example, you don't need Opus-level intelligence if you're looking to summarize an email or chat threads. However, you wouldn't want to block that model completely from your security engineering team.

Our goal is for AI Gateway to be the control plane for organizations deploying AI internally. Because every request from every user, agent, and tool already flows through it, AI Gateway is in a unique position to do more than observe and enforce. Budgets, spend limits, and identity-aware analytics give organizations visibility and guardrails, but they still rely on individuals to make cost-conscious choices request by request. The next step is for the gateway itself to make intelligent decisions on a user's behalf: sending each request to a model that is capable enough for the task. That way, organizations reduce spend automatically, while users keep access to the most capable models when their work actually needs them.

The results

We use Auto Router internally at Cloudflare within our OpenCode deployment and within Cloudflare OS, our custom agent harness. In our internal usage, we’ve seen results comparable with frontier models for coding tasks.

Auto Router does best when used across a wide range of knowledge-work tasks, like those typically found in a large organization with work spanning both technical and non-technical teams. We evaluated cloudflare/auto against OpenAI’s GPT-6 Sol and Anthropic’s Claude Opus 5.5 on our internal general knowledge work benchmark. The benchmark uses simulated workspace tools and covers common day-to-day workflows across email, calendars, Slack, files, travel and finance. Each task requires the model to use these tools to produce a verifiable answer or complete an action.

Model

Successful Trials

Success Rate

Total Cost

Cost per success

cloudflare/auto

252/291

86.6% (+6.2/−6.9 pp)

$2.10

$0.0084

Anthropic Claude Opus 5.5

281/291

96.6% (+2.7/−3.8 pp)

$5.91

$0.0210

OpenAI GPT-6 Sol

245/291

84.2% (+6.5/−6.9 pp)

$2.64

$0.0108

97 tasks with three samples per model per task. Parenthetical values show 95% confidence intervals estimated from 10,000 task-level bootstrap resamples, preserving all three repetitions within each task. “pp” indicates percentage points.

Our Auto Router delivered similar performance to other state-of-the-art daily-driver models, coming in at 80% the cost of Sol and 35% the cost of Opus. While that may initially seem surprising, one way to frame the problem a model router solves is through the “jagged frontier” across models. The ability to solve a problem often exists somewhere in this portfolio of models; the router’s job is to choose the right model for each task while balancing quality and price. Savings come from not paying frontier rates for non-frontier work, and they grow with how much of that work you have.

Another insight is that lower token prices do not always produce lower-cost outcomes. A model that looks cheaper on paper may end up using disproportionately more tokens to solve a problem. A router should minimize predicted trajectory cost, not just load-balance by dollars per million tokens. 

This is already useful today, but it’s only the beginning of what the Auto Router can learn from Cloudflare’s position in the inference path.  

How it works

When you send a request to cloudflare/auto, AI Gateway first builds the pool of models that can actually serve it. It filters out models that do not support the request format or execution mode, and accounts for the credentials, billing configuration, access control policies, and spend limits attached to the gateway. It will also filter out unhealthy upstream providers or models during downtime and automatically bring them back into the pool after an outage.

For the remaining candidates, the router looks at a compact view of the conversation. It considers the most recent messages, prioritizing the newest turns. The conversation is then sent to a multi-head classification model running on Workers AI and deployed on GPUs across our edge network. The classifier produces two sets of signals. First, it assigns probabilities across 14 task categories (like coding, planning, research, data analysis). It then rates the request across four dimensions on a scale from one to five: complexity, ambiguity, stakes, and dependence on earlier context.

A separate scoring matrix combines those signals with model benchmark results to estimate how well each model fits the request. To calibrate the scoring matrix, we defined the preferred model for a set of example task and difficulty profiles, then adjusted the weights to produce those choices.

Finally, the router combines expected quality with each model's input and output token prices. On straightforward requests, price carries more weight, so a smaller model can win when it is capable enough. As difficulty rises, the cost penalty falls and stronger models have more room to win. In simplified terms, cloudflare/auto selects the model with the highest utility as defined by:

For long agentic sessions like debugging or coding, cost is less driven by the model’s list price than by the cost of cache reads, which grows with session length. Switching models throws the cache away and forces a new model to write the whole context again. This can be worth it, as a model with a cheaper cache-read and cache-write prices can pay back the rewrite quickly.

Rather than completely avoiding model switching, the Auto Router accounts for the cost of cache reads and writes. Within a turn (one user input loop), the cache is hot and switching rarely pays off, so it’s better to keep using the same model. Across turns, the Auto Router applies a switching penalty that grows with the number of tokens already in context. A model that still holds a live cache for the session is priced at its cheaper cache-read rate. Every other candidate is priced at the full cost of rewriting the context, so the deeper the conversation, the more a switch has to earn back, through higher quality results that use fewer tokens overall or cheaper cache rereads. Switching models has another cost: most models can't read another model's reasoning tokens, so a model switch that drops reasoning tokens means that the new model may have to redo it at output prices. In the future, we want to account for this by having the router prefer to stay within the same model family when it switches.

From there, the router returns a ranked list. AI Gateway attempts the winner first and can move to another eligible model if that provider cannot serve the request.

This overall design has several benefits. The two-stage architecture (task and dimensions classifier to scoring matrix) means that routing decisions are legible because you can inspect each task’s predicted category and complexity to see how it translated into the model choice. Adjusting the router when a new model is released also does not require retraining — we only add its benchmark-derived weights to the scoring matrix. The same classifier can also support different routing profiles. For example, in addition to cloudflare/auto, we plan to release other routers in the future, including cloudflare/auto-best, which uses the same classification and model pool, but selects the highest expected quality without applying the cost tradeoff.

What's next

Our release today is only the starting point, and we’re continuing to invest in research and new routing strategies. In the near term, we want to:

  • Expand the models offered through cloudflare/auto
  • Include zero-data-retention requirements when filtering models
  • Account for provider capacity when selecting models
  • Select the appropriate reasoning or thinking level for each request
  • Add full support for the Responses API and WebSockets
  • Explore structured decision models as a first-pass classifier

The Auto Router is free while in beta. Read more in our developer documentation.

Acknowledgements: This project was also made possible by the efforts of Mats Dodd, Sam Scott, Oliver Yu, and Jeff Rafter.

Next.js applications, powered by Vite: introducing Vinext 1.0

Post Syndicated from James Anderson original https://blog.cloudflare.com/vinext-nextjs-on-vite/

When we launched Vinext in February, it was the result of an audacious week-long AI-driven experiment to see how far one engineer, and a stack of tokens, could get to replicating the NextJS framework backed by Vite.

In the seven months since that experiment, Vinext has grown into a framework that our customers trust and run in production for high-traffic, dynamic applications.

Today we are announcing the release of Vinext 1.0, the latest step on our journey to make it possible to deploy Next.js apps anywhere. Vinext lets you take any Next.js application, whether it was built for the Pages or App Router, and make it portable to be deployed to any web platform, including the Cloudflare Workers free plan, Netlify, or AWS Lambda.

Vinext 1.0 brings with it sweeping improvements to compatibility, stability, and caching behaviors, and sets the project up for the long term. There’s never been a better time to take your Next.js project and convert it to Vinext; just run npx vinext check and npx vinext init.

Graduation to 1.0

On release Vinext was promising, but it was incomplete. Since then, we’ve spent a lot of time both improving App Router compatibility and expanding that to Pages Router apps — which we’ve learned many customers are longtime fans of, with large applications that are complex to migrate. We didn’t want Vinext to be a tool that only worked for people using the latest App Router features.

Our focus has been on adopting both these routers, and watching our test compatibility closely, which for most important customer-requested features now surpasses 99%.

This improvement has been fueled through the community around our GitHub project. As soon as Vinext launched, that community threw it at a wide variety of applications to find the gaps. With their scrutiny, we found challenges not immediately obvious in the test coverage. Vinext needs to act exactly as Next.js behaves. It is not good enough to imitate functions with the same name. Building an alternative import { revalidatePath } is simple enough; the difficulty is in making sure it correctly affects the rendered pages, cache entry, and future requests.

Tracing requests through the application to make sure Vinext responds in the way expected — and replicating not just the API, but the behavior of this machine — was by far the more challenging aspect.

Once we’ve patched problems and brought new features forward, it’s important that we don’t regress, especially if Next.js makes a change. That’s why we’ve also built out our test suite: thousands of focused tests covering core framework behavior across both routers, the development and production server, and the deployment targets of Nodejs and Cloudflare Workers. We also run the Next.js end-to-end test suite against Vinext nightly, giving us a continually moving window on our compatibility, and making sure we immediately become aware of regressions coming from merged changes. Alongside the automated testing, we’ve been working directly with large customers that have Vinext in production to make sure they are not facing issues.

What’s in 1.0

The clearest messages we got from customers using Vinext is that certain Next.js features carry the framework and Vinext didn’t actually need to do everything that Next.js has launched in recent versions to be incredibly useful to them. So we focused on better support where you need it:

  • App Router, Pages Router, and Hybrid applications: We heard from customers that Pages Router was still important, and migrations are not a one-step process. Vinext therefore has support for both routing paths, including React Server Components, Server Actions, API routes, route handlers, middleware, and client-side navigation.
  • The complete page lifecycle: Pages can be rendered in many different ways: on the server, pre-rendered in the build, exported as static assets, or cached with page-level Incremental Static Regeneration (ISR). We’ve made sure that Background and on-demand revalidation work with any output.
  • Caching: Vinext has a shared set of caching functions across the App and Pages Router and the supported runtimes. We have further support for using Cloudflare’s Workers Cache.
  • Observability: Vinext provides Next.js-compatible tracing across both routers, so existing OpenTelemetry and Sentry setups continue to work. On Cloudflare Workers, traces also integrate with native Workers Observability.
  • Next.js ecosystem compatibility:  Vinext implements the public next/* surface and supports common Next patterns for use of authentication, MDX, image optimization, fonts, metadata, environment variables, and more.
  • First-class runtime support for Workers: While Vinext can run anywhere, server code can run in the Cloudflare workerd runtime during development and production, with direct access to bindings such as image optimization and hyperdrive. 

We’ve also made migration part of the framework: it takes two commands to verify that your Next.js install and any modification you have made is compatible, and set up the Vite and deployment configuration while keeping all your previous Next.js project structure.

When we talked to teams about what features were important for them, something stood out. Next.js 16 took a stance that Cache Components were an important part of the future of the framework, and yet most teams that we talked to were not using them and did not consider support a prerequisite to move. Therefore, Vinext today has limited support for the “use cache” directive that drives Cache Components, and though we will continue to improve compatibility there, we’re much more focused on the core priorities above.

Pre-rendering and cache warming

When we first announced Vinext, it supported Incremental Static Regeneration (ISR) after the first request, but it did not yet render pages during the build. Applications use generateStaticParams() and getStaticPaths() to identify pages that should be rendered when building, and they expect page-level ISR to connect those initial responses to background and on-demand revalidation.

Vinext 1.0 supports that lifecycle for both routers. It can prerender App Router and Pages Router routes during the build, serve those responses through page-level ISR, and invalidate them by path or tag. It also supports output: "export" when the result you want is a fully static site.

But this led us to question something: Why should this rendering happen during the build at all?

A site with tens or hundreds of thousands of possible URLs can spend a seriously long time rendering pages that receive little traffic. The build process cannot evaluate the long tail of traffic that most sites experience and therefore cannot focus compute time on the smaller number of more critical pages. Instead you waste hours of time waiting for sequential builds working their way through thousands of pages, long after the most important routes are done.

Cache warming is our solution to this, moving page prerendering from the build machine to Cloudflare’s network. Developers can continue to use Next.js primitives to identify the pages for prerendering, and Vinext can additionally identify high-traffic pages to add to this list. This happens in the background before your site is deployed to production, so that the moment it is, it is ready to serve rapid responses from the Cloudflare cache.

Inside the deployment process, this works by uploading a new Worker version and deploying it to 0% of production traffic, before then requesting pages specifically from that version. This allows the rendering pipeline to work before any real users hit the new deployment. Once the caches have been populated, the deployment can be promoted safely.

What we’re doing next

If the original experiment invented the one-off slopfork, the more consequential part has been how we can keep that process of self-improvement running indefinitely.

The project is now focused on keeping the framework up to date with everything happening upstream. Next.js canary receives new commits every day. Each morning, an agent reviews the changes, fetches diffs, and opens tracking issues for anything that could affect Vinext. Every night, the compatibility matrix is regenerated as we run the Next.js test suite against Vinext.

When one of these tests or issues reveals a gap, agents are now in the position where they can identify the change across both codebases, build a reproduction, port any relevant tests, and propose a fix.

This review has been catching missing cases, unsafe caching behaviors, and differences in the development vs. production servers.

Automation has helped us narrow the stream of activity into a focused set of changes that deserve attention, allowing the maintainers of the project to focus on only the issues that need context of how a process should map onto Vite from the Next.js implementation.

We’re building a software factory for open source at Cloudflare, and you can see what we’re up to on GitHub.

Try it out

Vinext is available for new applications, and existing Next.js projects.

Start a new application today:

Or migrate an existing application:

And then deploy it to Cloudflare Workers, with our cache warming:

Visit vinext.dev for documentation, examples, and the current compatibility matrix. Vinext is open source at github.com/cloudflare/vinext. Issues, pull requests, application reproductions, and feedback are welcome.

Introducing cf: the agentic CLI for the entire Cloudflare API

Post Syndicated from Matt “TK” Taylor original https://blog.cloudflare.com/cloudflare-cf-cli-launch/

Over the last year, agent use of Wrangler has skyrocketed.

In March 2026, agents were responsible for a quarter of Wrangler use, up from single-digit percentages the year prior. Last week, agent usage reached 48%.

Agents are more prolific users, using almost twice as many distinct commands per day, and are almost four times as likely to use six or more commands.

Agents love CLIs. But Wrangler only provides commands for around 280 operations, and Cloudflare offers thousands.

Earlier in the year we teased how we were planning to solve this and today, we’re enabling agents to use every Cloudflare product by introducing a new CLI: cf.

cf is a CLI that is built for the next generation of software development:

  • Agents can find the command they need to do anything they want to do with bespoke search and steering.
  • JSON is the default interface, pretty printed for humans and condensed for agents for maximum context savings.
  • cloudflare.config.ts is the new configuration format for the whole of Cloudflare, starting with Workers, and bringing the safety and accuracy of TypeScript to you and your agent’s language server protocol (LSP)
  • Vite becomes default, bringing with it the best local development server, and a plugin suite for developers and framework authors.

Install the open beta today globally and run it from anywhere:

cf gives your agent access to the entire Cloudflare API

What if your agent could do everything Cloudflare can do? That’s the question that sparked our interest earlier this year: agents were getting ever more powerful, but what they were able to do with Cloudflare’s CLI was still limited.

Wrangler was hand-built with each product team contributing and taking their own approach to their command developer experience. Enforcing patterns across teams was virtually impossible, even across our ~280 command paths. We had inconsistent terminology across d1 info, hyperdrive get, workflows describe as each team came up with their own practices at different times. Some teams built entirely custom experiences across thousands of lines of code that turned out to be used extremely rarely, and teams came up with different approaches to solve the same problems.

We wanted to both standardize what we had and make a massive expansion, all at once. Forge — Cloudflare’s new unified API generation pipeline — enabled us to do this, building on the idea of generating our CLI commands directly from the API schema that powers our API documentation and SDK generation. Everything we provide has an OpenAPI schema, and if we annotate this with just a little more information, we can use it as the source for Forge to make a CLI.

This enables us to expand cf from the ~280 functions that Wrangler had built up over time, to cover the entirety of the Cloudflare API surface of over 3,000 operations.

Now it’s simple to give your agent cf and ask it to go set up a worker, deploy it, monitor and observe it, protect it with Cloudflare Access, buy a domain, and front it with Cloudflare WAF, all from a single tool.

Building for an agent that has never used cf

cf is built for the trajectory of software engineering, where agentic development is drastically changing how software is built and deployed. This year we’ve been focused on providing tools to support this shift, culminating in cf. cf has been built from the ground up with agents in mind, and includes novel tools for agentic command discovery that we think will become standard in more CLIs in the near future.

Wrangler came with the advantage that years of documentation, blogs, and third-party guides have been absorbed into the training process of LLMs. It also came with the same disadvantage: changing how Wrangler works now goes against learned behavior, and significant change would be inevitable given the scale of improvement we want to make.

Introducing a new CLI that agents have never seen sounds like a big disruptive change — but actually it’s the cleanest thing we can do. Because of the design decisions we have made, the context injections we can make, and the AGENTS.md files we can append, making a switch in this way is actually less confusing than having an agent contextualize the major differences between two versions of a tool it is familiar with. We’re launching with a couple of these agent-focused features built in, with more to come.

Agents need to filter JSON, not look at tables

When agents use Wrangler, they append --json to every command they run, and then often filter the output with jq to extract a subset of fields. But only some commands in Wrangler supported --json ; many commands returned unicode tables, designed for humans looking at output in their terminal. Agents can figure these out, but it costs them more time and tokens than a jq filter.

In cf we’re taking the opposite stance: agents just need JSON, and if agents are the future primary user of this tool, it should be the default. For the vast majority of commands that will rarely be accessed by humans, this is obviously the right call.

You as the human customer of this CLI are, in reality, one step removed from using it. Agents being able to easily filter their results and then return that filtered list in whatever format you request is preferable to supplying tables you will never likely read directly.

But what if you’re looking to do something that might require real personal input, like searching for a domain to buy?

For commands that your agent can access through chaining named parameters in a long and unwieldy sequence, you can simply fill in a form. Cf deconstructs the requirements of the API into a series of validated inputs, so buying a domain, even one with complex requirements, is simple to follow.

Or, if you insist, just ask your agent to do it.

Your agent can find the right command itself

With 3,000 possible routes through a CLI, how can your agent find the right operation it needs quickly without bloating your context? For this reason we have also added cf cli search.

This command allows your agent to ask in natural language what it needs to do, and a small search index will provide a list of appropriate commands, based on their API description and parameters. We automatically tell your agent about this command when it runs --help for the first time.

Configuration that type-checks your agent

Our new configuration format is based on TypeScript, which is easy for humans and agents to parse, and allows you to write your configuration programmatically.

Typed configuration is enormously helpful for agents. We’ve found that even with no prior context of the programmatic configuration format, agents are able to easily identify and edit the configuration on demand, even across elements like env which have dramatically changed from the same named feature in Wrangler. All agents that use LSP plugins, such as Claude Code and Codex, benefit from being able to interpret more about the configuration file format in context, and make much more accurate suggestions as a result.

Compare this to TOML, which had no accessible schema, or JSONC, which had a linked schema that agents rarely used.

Some Wrangler configuration files inside Cloudflare have been condensed by 40% from over 5,000 lines, with many custom environments per developer, to factory files that build each developer’s configuration more efficiently.

This is achieved through programmatically defining each environment from the same universal base, instead of copying env blocks as was typical in Wrangler. A simple Worker with multiple environments simply switches on the Vite-native mode argument to swap between one set of configuration and another.

A simple configuration that does this now looks like:

You can migrate your Cloudflare Worker to this new format through cf migrate.

We’re also providing a few helper functions to make building your Worker a breeze.

bindings gives you a simple place for your agent to discover all the developer platform has to offer. Everything — from environment variables to storage, database, and queues — can be auto-completed and explained by your editor.

Similarly, we have included a helper for triggers, which is the new way to define routes, queues, schedules, and email triggers for your Worker. Rather than having these scattered through your configuration file, it’s now simple to find, in a single block, the actions that could trigger your Worker to run.

defineConfig.worker is just the start here. Our intention with cloudflare.config.ts is that this is how you manage Cloudflare as a whole. Every product you need — along with its API being available to your agent through cf — will be able to be expressed through typesafe configuration. Soon you will be able to configure entire policies, set up zones, configure DNS and more, all through this configuration file.

A best in class development experience

When Wrangler first started building JavaScript Workers, Vite didn’t exist. Instead, we used esbuild in Wrangler to bundle your Workers. The dev server that Wrangler made available on :8787 was something that the Wrangler team built, and modifying any of this meant reaching into the internals of Cloudflare-specific local tooling like Miniflare.

Vite is a huge improvement on this, and comes with a large ecosystem of plugins you can use, as well as providing a best in class dev server with HMR (hot module replacement), and builds that use the Rust-based library Rolldown for tree-shaking. Anything you can do with Vite, you can do with the Cloudflare Vite Plugin.

The Cloudflare Vite Plugin is the recommended way we suggest you build Workers, whatever you are building: whether that’s a frontend-focused project or a backend API. Together with our Vitest plugin it provides a cohesive development and testing environment that matches the Workers runtime and gives you direct access to bindings and platform APIs.

cf is built on Vite as default. Most of your Workers will migrate simply with agents. Others may take more time, which is why cf will continue to delegate to Wrangler for dev and deployment for JavaScript Workers that need to continue to use esbuild and Rust and Python Workers.

Migrating from Wrangler

Migrating a Worker from Wrangler is as simple as running

Workers that already build with Vite will be converted to cloudflare.config.ts for you. If your Worker relies on Wrangler for esbuild, then cf will continue to delegate builds to Wrangler.

When the open beta ends we will release a final major version of Wrangler that directs you and your agent to use cf. We’ll continue to provide maintenance support for Wrangler for 18 months after the beta ends, to give you time to migrate.

You can also take new projects and automatically configure them for Cloudflare by running cf init/deploy, which will install the Cloudflare Vite Plugin for you and create a configuration file.

Static sites still don’t require a configuration file to start, and deploying them is as simple as running cf deploy in your project.

To start a new Hello World project with cf, use cf init.

cf is open source and issues can be reported to our GitHub repository.

Supporting native Rust in Workers with the new Emscripten target for wasm-bindgen

Post Syndicated from Guy Bedford original https://blog.cloudflare.com/rust-workers-emscripten-target/

Today we’re announcing the first public experimental preview of a feature to better support native Rust code and even Tokio-based applications just running natively on Workers: first-class support for the Emscripten wasm32-unknown-emscripten Rust compiler target on the wasm-bindgen open source toolchain and Cloudflare’s Rust Workers.

wasm-bindgen is the open source toolchain powering Rust-based WebAssembly applications on our V8-based Workers Runtime. Enabling the Emscripten target for wasm-bindgen has been a long-term effort, first initiated by Google over a year ago, and then further reviewed and supported by the Cloudflare engineers maintaining wasm-bindgen.

While still in pre-release, we’re excited to share the new workflow possibilities this work enables in running native wasm-bindgen Rust applications with Emscripten on the web, Node.js, and on Cloudflare’s global Workers platform.

In testing we’ve been able to see significantly improved library compatibility for Rust Workers. To illustrate the sort of capabilities supported, we were able to get a Rust-native Minecraft server (Pumpkin) running inside of a Durable Object with TCP ingress, using real TCP sockets via Tokio. See the end of this post for a full description of this port.

Emscripten is an open-source WebAssembly compiler toolchain initially created by Mozilla and currently maintained by Google engineers, which makes it possible to run and bridge native code with the web platform, including supporting and virtualizing platform features such as timers, file system operations, sockets, and other native functionality.

Since Cloudflare Workers supports Web Platform APIs and Node.js compatibility, we are able to support Emscripten on Workers using its Node.js compilation flags, fully virtualizing native platform features such as timers, file system operations, and sockets on top of our existing Node.js APIs. And with our work on Tokio support, Rust code building on top of Tokio’s async runtime ecosystem can now also be fully integrated into JavaScript-based host environments with this Emscripten target.

We’ve made these current experimental patchsets available with example applications to try out today, including:

Supporting the wasm32-unknown-emscripten target in wasm-bindgen

Mitch Foley on Google’s Portable Toolchains team first encountered the need for Rust WebAssembly toolchains to interoperate with C++ when another internal team at Google was interested in using wasm-bindgen to interface with their JavaScript. The goal was for wasm-bindgen to drive the build and produce the companion JavaScript, while Emscripten’s linker (wasm-ld) would be able to link in any C++ dependencies needed.

Internally, Google uses Emscripten in C++ codebases to generate JavaScript along with the rest of their applications, allowing the C++ and JS to interact with one another. While Emscripten was capable of linking in Rust code as a dependency, the thing it couldn’t do was provide a robust binding system between Rust and JavaScript like wasm-bindgen does.

The problem was both tools assume they are in charge of loading and interacting with JavaScript and generating the final JS and Wasm output. Because of this conflict, Google internal users would have to commit to using one or the other toolchain and never both, and adding a second set of tooling would have doubled the support surface for Google’s Portable Toolchains team.

Mitch and his colleague Yifan Yang crafted a plan to have them work cooperatively: Emscripten would continue to drive the build, load the Wasm module, and provide the companion JS, while wasm-bindgen would produce a smaller portable version of its JavaScript bindings in a format that could be directly included in Emscripten’s library system.

This wasn’t just a technical problem — both Emscripten and the wasm-bindgen maintainers had to support this plan and be willing to maintain integration tests that depended on the other. With feedback from Google’s Portable Toolchains team, Google’s Wasm Tools team, and Cloudflare engineers, this finally resulted in the required changes landing in both projects.

As a result of these efforts, the wasm-bindgen and Emscripten toolchains now have seamless interoperability under the new -sWASM_BINDGEN configuration:

  1. C++ Emscripten code driven by Emscripten’s compiler can be built against static wasm-bindgen Rust code, fully supporting the wasm-bindgen bindings layer alongside the Emscripten bindings layer.
  2. Rust applications using wasm-bindgen and driven by Rust’s compiler can now be built for the Emscripten target, fully supporting Emscripten’s bindings layer alongside wasm-bindgen’s bindings layer.

See the wasm-bindgen Emscripten documentation page for more information about using this target.

Rust library support for Emscripten

After landing the wasm32-unknown-emscripten target support in wasm-bindgen, early prototypes by Cloudflare engineers demonstrated clear success for supporting this target on Cloudflare Workers. Many libraries worked out of the box, even including low-level systems libraries, since Emscripten already supports the target_family = unix in Rust.

Some low-level systems libraries that were unaware of Emscripten required patching, for example libc, socket2, and Mio. Even for these low-level libraries, patches primarily involved adding the Emscripten target to the existing platform gates, for example to explicitly allow target_os = “emscripten”, in place of Wasm platform gates.

Overall we posted these target support patches fairly infrequently, and they were mostly trivial when needed, with the Rust library maintainers very receptive to reviewing the support.

That said, one of the critical foundational libraries that needed significant support was Tokio.

Tokio support

Cloudflare Workers are single-threaded and hosted within a JS event loop, while Tokio async is designed around blocking operations being supported through threaded parking semantics. The two models are clearly not compatible with each other — a blocking operation such as a pending socket read or epoll wait cannot block the shared JS event loop.

To get around this, we had two options: WebAssembly JavaScript Promise Integration or to modify Tokio to support event loop runtime integration.

To support Tokio building for Emscripten, we have contributed full Tokio support patchsets that are being reviewed upstream, with the first target support patch for wasm32-unknown-emscripten already landed upstream in Tokio. With these patchsets, we’ve been able to fully support both approaches in Cloudflare Workers. In our Rust Workers Tokio examples, these patches are required directly pending further integration upstream.

Supporting JSPI

WebAssembly JavaScript Promise Integration (JSPI) maps directly onto Tokio’s existing parking semantics. This is because JSPI allows a blocking Wasm call to suspend the WebAssembly stack on a synchronous operation and return control to the JS event loop, exactly as one would expect of a park.

But when a new Wasm call is made while an earlier one is suspended, JSPI doesn’t break: it instead supports having a new WebAssembly stack being entered while arbitrary existing stacks are suspended at the same time.

The problem here, though, is that Rust itself isn’t aware that its stack is being switched out from under it. Tokio’s runtime context is tracked via a thread local, but a JSPI stack switch is not a thread switch, so the suspended and new stack still share the same thread-local runtime context. As a result, the Tokio runtime context still thinks it is in the parked context when a new Wasm call enters, and then panics because the runtime is already entered.

Supporting fully reentrant JSPI therefore requires careful thread local handling to split context between JSPI context switches. To support this, the thread-local context itself must be swapped on each JSPI enter, exit, suspend, and resume. In effect this is cooperative time-multiplexed threading, with each suspended stack carrying its own runtime context.

With our pre-release Tokio patchset, we were able to implement and verify this model. Finalizing the design and upstreaming it is ongoing in collaboration with the Tokio and Emscripten teams.

Adding an event loop runtime to Tokio

The other runtime approach is full event loop integration. Event loops are of course a common paradigm in native UI applications, so the idea of supporting an event loop runtime for Tokio was certainly not something completely unfamiliar to the maintainers in discussions we had around the Emscripten target support.

The question was rather how to design an event loop for Tokio in such a way that it could work across native applications (including Windows and macOS), as well as for WebAssembly embeddings in JavaScript hosts.

If we could design a general LocalEventLoop runtime for Tokio, we could solve this problem more generally.

A Tokio runtime does two things in a loop: (1) it polls tasks until nothing is ready, and then (2) it waits. The wait is what makes it a runtime rather than a library: the thread parks inside the I/O driver until a socket becomes readable, a timer expires, or another thread wakes it.

But when the host already has its own event loop, Tokio's loop cannot run inside it without blocking the host's. The way around this is to split Tokio's loop in two with both parts able to integrate with the host: (1) becomes an explicit drive() operation that runs one batch of ready tasks and returns, and (2) is replaced by a wake, so that instead of parking, the runtime tells the host it has work and the host calls drive() when it is ready to.

Our proposed LocalEventLoop design for Tokio is a LocalRuntime whose wait has been replaced by a wake. It is built with a standard std::task::Waker that the host owns, which instead of being used to signal that a future should be polled soon, is used by the Tokio runtime to signal that the runtime itself should be driven soon.

Consider the example of a socket read under a regular Tokio runtime:

This would correspond to the following call diagram:

Instead, with LocalEventLoop we can write:

Here, spawn_local returns immediately, with the host event loop taking responsibility for driving the spawned task to completion via el.drive() calls from the host. When stream data is unavailable, the runtime simply returns, handing back control to the host. Once the socket becomes readable, Tokio uses the host_waker to signal that the runtime needs driving.

This LocalEventLoop flow corresponds to the following call diagram:

Everything that would have unparked a native runtime's thread wakes the host instead: a spawn, a task woken from another thread, a socket becoming readable, a timer expiring. Because a Waker is Send + Sync and carries no execution semantics, its implementation is just a notification to the host's event loop. This makes the contract safe by construction: a wake arriving from another thread, from a host callback, or even during a drive, queues work rather than re-entering the runtime. The drive that follows runs on the owning thread with nothing else on the stack.

The one thing you cannot do is wait. block_on still exists and runs a future as far as ready work carries it, but where a normal runtime would park, LocalEventLoop::block_on panics instead. This is because nothing could wake that future from inside the call since its wait belongs to the host.

With this design, the host event loop is never blocked, and interleaves its own work with Tokio's, one batch at a time. The same architecture embeds into a GTK main loop, a Win32 message pump, or a Cocoa run loop. In addition, any number of these runtimes can co-exist due to their cooperative execution semantics.

Supporting sockets and epoll on Emscripten

With both Tokio runtime integration approaches fleshed out for the Rust Emscripten target, we were then able to support most of the Tokio test suite, with one major subsystem still unsupported — the net feature. This includes its sockets APIs: TCP, UDP, and Unix sockets. The reason for this was that Emscripten only supported poll() and a WebSocket emulation layer, but not epoll_wait(), on which Tokio’s I/O driver is built (via mio).

For Cloudflare Workers, we wanted to be able to fully integrate with our TCP sockets APIs, including upcoming inbound TCP. To do this, we would need to build a bridge between Emscripten’s virtualization layer and our own sockets API layer.

Instead of having to build this bridge ourselves, we realized we already have one: the node:net API we support in our Node.js compatibility layer. Emscripten had an -sNODERAWFS mode to bridge natively into Node.js FS APIs, so the same approach could give us a sockets bridge without Emscripten or Cloudflare Workers needing to implement any custom APIs on either side.

We contributed this work in over 40 pull requests to Emscripten, which is now the –sNODERAWSOCKETS layer compilation option, enabling support for epoll, TCP, UDP, and Unix sockets for Emscripten applications in Node.js — and, because Workers implements the same node:net API, on Workers itself as well.

Under JSPI, Emscripten's epoll_wait() simply suspends the stack until readiness, so Tokio's I/O driver works as on native. For LocalEventLoop, readiness instead had to reach the Waker from a JS callback. To support this, we drafted an Emscripten proposal for a new emscripten_epoll_add_listener API to associate a callback on an epoll’s ready events. The next drive() then collects these events with a zero-timeout epoll_wait(), so Tokio's existing I/O driver can be used unchanged.

Running a Minecraft server on Workers

To test out this new target, the challenge was raised to see if a Minecraft server could be deployed to Workers. Dan Lapid then implemented a Pumpkin Minecraft server running on a Durable Object over a single weekend.

Pumpkin is a Rust Minecraft server built on Tokio and designed for multi-core machines. World generation runs on a dedicated thread pool, while the game tick loop and chunk scheduler each run on their own OS threads.

Inside a Durable Object there is exactly one thread, so getting Pumpkin to run there meant turning those threads into cooperative tasks on the event loop. Leaning on the Tokio integration, the tick loop and chunk scheduler became async tasks and each Rayon job became a Tokio task. World generation then runs on the event loop itself, taking one turn per chunk and interleaving with network I/O and game ticks rather than blocking them.

For persistence, Pumpkin writes its world through ordinary std::fs. Under Emscripten’s -sNODERAWFS option file system calls get forwarded to the node:fs Node.js compatibility layer for Workers. Since this bridge is just JavaScript, it is easy to swap out. Dan’s worker-fs-mount library enables mounting a node:fs-compatible file system with its durable-object-fs backend, which stores files as rows in the Durable Object's SQLite storage. With this integrated, every file Pumpkin saves becomes a row in the object's database, written synchronously and committed with the Durable Object's transaction. Pumpkin itself has no idea it isn't on disk, and a restarted object boots straight from the same world.

Networking needed no changes in Pumpkin either. Each player's connection arrives through Workers TCP ingress at the Worker's connect() handler, which forwards it into the Durable Object. There, handleAsNodeConnection() from cloudflare:node dispatches the socket to a net.Server listening on that port inside the object – a new TCP counterpart of handleAsNodeRequest(). Emscripten’s -sNODERAWSOCKETS backend implements TcpListener on top of net.Server, so the server accepts players exactly as it would on Linux, and the returned promise tells the object when the last player has left, so it can save and shut down.

Running a fully persistent Minecraft server inside a Durable Object with multiplayer support demonstrates the level of native compatibility that is possible with this new Emscripten target. We’re excited to see what native Rust applications you can bring to the platform.

Try it out

We’re making these full patchsets and workflows available today for experimental use. Improving wasm-bindgen’s support for modern WebAssembly standards for all users is part of our ongoing commitment to the Rust and Wasm ecosystem.

We welcome all contributions and feedback — find us on GitHub and in the #rust-on-workers channel on Cloudflare’s Discord.

EmDash 1.0: the stable CMS with a secure plugin registry

Post Syndicated from Scott Buscemi original https://blog.cloudflare.com/emdash-cms-plugin-registry/

When we introduced EmDash on April 1 as the “spiritual successor to WordPress”, the buzz was hard to ignore. Walking around a WordPress conference that month, we couldn’t walk far without hearing murmurs about EmDash from fellow attendees.

But alongside the excitement and curiosity has been a seed of doubt among some in the industry. Was this just an April Fools’ joke?

It was not. Today, we are releasing EmDash 1.0: a stable, free, and open source CMS built on Astro, ready to power a production website, your agency’s vibe-coding platform, or your hosting company’s site-building experience.

Developers build with Astro, editors manage content through the EmDash admin, and agents can work through the API, CLI, or built-in MCP server. EmDash 1.0 brings those pieces together with production-tested editorial, media, localization, migration, and deployment workflows.

We are also launching a decentralized plugin registry that lets developers publish without handing ownership of their identity or releases to a central marketplace, while site owners can discover and install their plugins directly from inside EmDash.

The road to 1.0

Since EmDash's first beta, developers have launched real websites with it. Even so, we kept hearing a reasonable response: “This looks interesting. Let me know when it is 1.0.” Before relying on EmDash for their sites, they wanted confidence that it was stable, secure, that upgrades would protect their content, and that we are fully committed to maintaining it.

EmDash 1.0 is our answer to that. For the past five months, we have worked with contributors and production users on the parts of the CMS that every site depends upon: data safety, database migrations, editorial workflows, localization, plugin security, performance, and the reliability of the admin, API, MCP, and media experiences.

Real deployments shaped much of that work, uncovering edge cases and identifying needs that only show up when a site is serving real traffic and being used by real teams of editors.

As one example of this, Avulux converted a custom microsite to EmDash after dealing with the maintenance burden of WordPress for too long. Since EmDash is an agent’s best friend, using the EmDash Agent Skills helped that transition take less than a day to complete.

“The site had to be fast to use and simple for our team to update,” says Greg Barbosa, Director of Innovation and Systems at Avulux. “WordPress had become the opposite of that. With EmDash, we now have a shared platform that developers can extend and marketers can edit content."

In August, we migrated the Cloudflare Blog to EmDash as part of our “Customer Zero” approach. Working alongside our content engineering team provided actionable insights around localization, media management, the admin editor experience, and scaling. Comfortably handling the traffic load for the Cloudflare Blog meant being ready for millions of pageviews per week, spikes up to 5,000 requests per second (RPS) of legitimate traffic, or sporadic DDoS attacks. The optional KV object caching, Hyperdrive database adapter, and Workers Cache compatibility were all features spawned from our migration project that are widely available to all customers now.

Built in public, open to everyone

A CMS sits at the heart of an organization’s web presence. It is trusted with its most important data, and is often used by dozens of editors every day. They need to be able to know they can rely on it, without fear of vendor lock-in or changing business priorities. For that reason, EmDash is completely free and open source, using the flexible and permissive MIT license.

EmDash 1.0 could not exist without its open-source development community. At the time of writing, more than 175 people have contributed to the project, across more than 1,800 commits. The rise of agentic coding tools has presented both challenges and opportunities to open-source projects, and we have deliberately built a project where the agents can help the human developers, rather than being overwhelmed by them.

We are particularly grateful to the core group of the most dedicated contributors, who between them have shipped hundreds of improvements to all areas of the project. They include @swissky, @danielmlr, @MA2153, @marcusbellamyshaw-cell, and dozens of others. Contributors have translated EmDash into 25 languages, from Arabic to Ukrainian.

Particular recognition is due to Noah Pham, who joined Cloudflare as an intern and became EmDash’s second maintainer alongside Matt. Noah contributed more than 80 changes, taking ownership of major parts of the media library, content editor, and admin interface. We have said that interns ship meaningful work at Cloudflare; Noah’s work is now at the heart of EmDash 1.0.

There is plenty more to build, and contributing does not have to mean writing code. If you want to help with code, translations, documentation, testing, design, issue triage, answering questions, or just welcoming new users, join over 800 others in the EmDash community on Discord.

A growing ecosystem

A content management system thrives when the ecosystem around it is healthy and supported. We’ve been excited by the theme companies, plugin shops, agencies, and platforms who are creating new services and products using EmDash.

  • Lexington Themes offers 44 Astro themes with EmDash variants, giving teams a beautiful and functional starting point for their site, complete with reusable components and built-in content collections.
  • Urumi is using their WooCommerce expertise to release EmDash’s first eCommerce plugin.
  • Empress is launching a platform for multi-brand entities, offering the flexibility of managing a fleet of sites with natural language or a conventional CMS admin panel.

“Empress's delightful multisite experience would not be possible without the foundations EmDash has laid: sites that are fast, safe to extend, and easy for people and agents to read and act on,” says Raj Makker, Empress’s founder. “EmDash unlocks powerful control over a website, and Empress builds on it to give you complete control over as many sites as you want. We're excited to be part of this journey.”

A plugin registry that does not own the ecosystem

With EmDash 1.0, developers can publish sandboxed plugins and site owners can discover, inspect, and install them from the EmDash plugin registry.

Traditional plugin registries usually combine three roles: they provide the publisher’s account, hold the authoritative package record, and operate the catalog where users discover it. That is convenient, but it also makes one company the gatekeeper for both identity and distribution. If an account is suspended, a listing is removed, the rules change, or the service shuts down, publishers cannot take the same identity and release history somewhere else.

EmDash separates the plugin from the catalog. Publishers retain control of their packages and release history, while EmDash provides a convenient place for people to find and install them. Other services can index the same publications, build their own catalogs, and apply their own policies without requiring developers to start again.

The EmDash catalog applies default content moderation to the names, descriptions, links, and images it displays. Moderation can hide harmful or inappropriate material from this catalog, but it does not rewrite a release, take ownership of the plugin, or erase the underlying publication.

The registry is built on AT Protocol (atproto), a decentralized network protocol that powers Bluesky and a growing ecosystem of applications. Plugin authors publish with an Atmosphere account, the portable identity also used across these applications. The package and release records are signed by the publisher and stored in the publisher's own account.

EmDash hosts the default registry services so publishers and site owners can use them without running infrastructure themselves. We hope others will build new catalogs, moderation systems, publishing tools, and other services we have not imagined yet. To help this, we have released all of our services as open-source software, including the aggregator that powers the plugin registry, the labeler service that uses Workers AI to moderate package descriptions, and an Astro live content loader to make it easy to include plugin listings in any Astro site. We are excited to see what the community will build with them!

There’s no centralized controller that can take over plugins or remove them on a whim. We believe the security of plugins and the registry should be baked into code, rather than trusting any central authority or code of conduct.

Decentralized publishing does not mean accepting unverified code. Atproto repositories use signed Merkle Search Trees, so an inclusion proof connects the exact release record to a signed commit from the publisher’s account. EmDash can therefore verify the record independently instead of trusting the catalog’s copy. It then checks the plugin’s checksum, package name, version, requested access, and any required build provenance, and confirms that the downloaded bundle matches the signed record.

The registry supports free plugins today, and we aim to add support for paid plugins in future, with the same decentralized model as now. We are watching the Atproto Spaces Alpha with interest. Anybody should be able to run a secure, paid plugin marketplace. We want publishers to be able to make money from their software.

For now, plugin authors can publish useful software without asking permission from EmDash — and site owners can understand exactly what that software is allowed to do before they run it.

Plugins with clear boundaries

Plugins are the core of a CMS ecosystem. They save site developers from rebuilding the same integrations and workflows, while giving experts a way to share their best practices — and potentially build a software business.

Agents make that reuse even more valuable. An agent can build a one-off integration, but it still has to understand the problem, generate and test the code, and maintain it afterward. A plugin captures that work. Another agent can install and configure a proven solution instead of starting again from scratch.

That convenience creates a serious security concern: plugins are code you did not write, operating alongside valuable content and customer data. With WordPress, plugins run inside the same PHP process as the rest of the application, with direct access to its database, filesystem, and network. A contact-form plugin can technically read unpublished posts, modify another plugin, or send data anywhere. Site owners must trust that it will not — and that a future update will not change its behavior.

Sandboxed EmDash plugins use a different model. Each plugin runs in an isolated runtime with access to its own private storage, but not to the site’s content, media, users, secrets, environment, filesystem, or network. It gains additional abilities only when they are declared by the plugin and approved by the site administrator.

Installing a sandboxed plugin therefore feels more like installing a mobile app than a traditional CMS plugin. EmDash shows what the plugin wants to do before it runs, and the runtime limits it to those approved abilities.

That lets plugins perform useful work without receiving unrelated access:

A plugin could…

It might need to…

It still cannot…

Index published articles for search

Read content and contact the search service

Edit articles or contact other hosts

Optimize uploaded images

Read and manage media

Read user content

Send publishing notifications

Observe publishing and send email

Change the content being published

Provide configurable webhooks

Read selected events and contact public destinations

Reach private networks or other plugins’ storage

The important part is that these abilities are independent. Giving a plugin access to media does not also expose users or unpublished content. Allowing it to contact one service does not open the rest of the network. These boundaries are enforced by the runtime, not left to the plugin author’s good intentions.

This isolation is not limited to Cloudflare deployments. On Cloudflare, EmDash runs each plugin as a Dynamic Worker through the Worker Loader. On Node.js, EmDash starts workerd — the open-source Workers runtime — as a separate process and runs each plugin as an isolated service inside it. Plugins use the same manifests and capability-gated APIs on either platform. See the plugin sandbox documentation for setup and runtime differences.

From a single site to a website platform

EmDash can power an individual Astro site, but it is also designed for companies building website creation and hosting products.

Workers for Platforms lets those companies run each customer’s site as a Worker on Cloudflare’s global network. They do not need to provision a server for every site or build the surrounding networking and deployment infrastructure themselves. That leaves them free to focus on the experience their customers use to create and manage a website.

EmDash provides the content layer for that experience. A platform can use the admin interface directly, build its own interface on the API and CLI, or put an agent in front of the built-in MCP server. A bakery owner could update opening hours by asking for the change in plain language; the platform’s agent would handle reading, updating, and saving the content.

The sandboxed plugin model also gives platforms a safer way to offer extensions across many customer sites. Plugins receive only approved access to content, media, users, email, or external services, rather than running with unrestricted access to the whole application. Platforms can run their own plugin marketplaces with access to the full registry — or curate a selection of pre-approved plugins.

We are always looking for additional hosting partners that want to join us in developing new agent-oriented CMS experiences. Reach out to us if you’d like to learn more about reinventing with EmDash.

EmDash Build: an open source AI site builder

We are also releasing and open-sourcing an alpha of EmDash Build, an AI site builder that hosting providers, website builders, and platforms can run themselves and integrate with their own systems. Try the demo today at build.emdashcms.com, or explore the code.

If you’re building a site today, for yourself or for a client, you won’t start in an IDE. You’re more likely going to start with a chat box, and describe to an agent what you want. But once you have something, you could find that changing simple things requires you to go back to that prompt box, and either roll the dice on the result, or burn credits for a one-line change.

EmDash Build creates an EmDash site instead, which brings the full stack: server-rendered Astro pages, a database, media storage, and an admin interface. The agent designs a content model from the brief, fills it in through EmDash's MCP server, and writes the pages that display it. After that, you or your customer can edit text right on the page, schedule posts, or let an agent do it for you.

In EmDash Build, each project gets its own Cloudflare Sandbox container, where the agent, built on the Agents SDK, verifies its own work. Artifacts tracks every change as a git commit. When the site is published, the content moves into a production EmDash site, and deploys to the host's Workers for Platforms namespace.

Get started and get involved

With the release of EmDash 1.0, now is a great time to migrate your company’s marketing site or have an agent spin up that side project you’ve been talking about. Try out the EmDash playground site here. 

To create a new EmDash site locally, via the CLI, run:

Or you can do the same via the Cloudflare dashboard below:

If you’re ready to develop an EmDash plugin, our documentation has a step-by-step guide for creating and publishing a plugin to the registry.

We also welcome you to join our growing community of contributors on Discord. You don’t have to be an engineer to get involved — we welcome translators, issue triage managers, user experience designers, marketers, and all others who are excited about the future of content management systems.

Introducing The Cold Start: pitch your startup live at Cloudflare Connect

Post Syndicated from Fatima Yusuf original https://blog.cloudflare.com/introducing-the-cold-start/

Sixteen years ago, Cloudflare was one of more than 1,000 startups hoping for a spot on stage at TechCrunch Disrupt.

We were not, on the face of it, an obvious choice. Cloudflare was infrastructure: we made websites faster and protected them from attack, which was not well understood by the general market at the time. Infrastructure is often invisible right up until the moment it becomes important.

But on September 27, 2010, Matthew Prince and Michelle Zatlyn got on the Startup Battlefield stage and launched Cloudflare to the public. During the presentation, people started signing up. Then more people started signing up. By the time the judges had finished asking questions, hundreds of websites had joined Cloudflare, putting our initial five data centers to the test in real time. In the seven days following, traffic through our network increased almost 10x and Cloudflare jumped from the 1,000th largest site online to one of the top 50.

Cloudflare didn’t win the main trophy that day. At the awards ceremony, TechCrunch founder Mike Arrington described what we did as something akin to "muffler repair for the Internet" and honestly, he had a point. But then he named us the Most Innovative Company. As Matthew wrote later: "You may not win the trophy, but you'll receive something else far more important."

There are moments in the life of a company when somebody gives you a room, a microphone, and a small amount of time to explain the thing you have spent months or years obsessing over. Most of the time nothing magical happens. Sometimes, though, the right people hear it at the right moment and suddenly an idea that has mostly existed between a handful of people starts moving through the world.

This October, as Cloudflare turns 16, we want to give five early-stage startups a stage of their own.

Introducing The Cold Start

The Cold Start is a live startup competition taking place next month at Cloudflare Connect in San Francisco. We’ll select five early-stage companies and give each of them five minutes on stage to explain what they’re building, why it needs to exist, and why they are the people who should build it.

We are less interested in perfect pitch decks than in interesting ideas clearly explained. You do not need thirty slides, a suspiciously precise TAM calculation, or a rehearsed story about how your childhood prepared you to disrupt accounts receivable. What we want is to understand the vision: what changed in the world that makes it possible, what you see that other people have missed, and why you cannot stop thinking about it.

The five finalists will make their case in front of the Cloudflare Connect audience and three people who've spent a fair amount of their lives thinking about companies, infrastructure, and the Internet:

  • Matthew Prince, co-founder and CEO of Cloudflare
  • Michelle Zatlyn, co-founder and President of Cloudflare
  • Dane Knecht, CTO of Cloudflare

The judges will select one startup to win $500,000 in Cloudflare credits, take over a billboard in San Francisco, and receive an invitation to our VIP speakers dinner that evening.

Five companies, five minutes each, and a room full of people paying attention.

What are we looking for?

The Cold Start is open to ambitious early-stage startups that have raised less than $10 million. Beyond that, we are deliberately keeping the definition broad because the most interesting companies rarely arrive neatly categorized.

We want to see ideas that seem obvious once somebody finally builds them, and ideas that initially sound slightly unreasonable. We want infrastructure that appears boring until you realize everyone is going to need it; products that could not have existed a few years ago; strange new interfaces; new ways of building software; things aimed at enormous existing markets and things aimed at markets nobody has bothered to name yet.

Most of all, we want to meet people who have noticed something about the world and decided to do something about it.

As part of the application, we’ll ask you to tell us who you are, give us your one-line pitch, explain what you’re building and why, tell us about your funding and revenue stage, show us how Cloudflare fits into your stack, and point us toward anything else that helps us understand you and your work.

The goal is simple: make us understand why the thing you’re building should exist.

Apply for The Cold Start.

Applications are open now and close Friday, October 2, 2026. 

Five minutes in San Francisco

The Cold Start will take place Monday, October 19, from 4:00-5:00 p.m. PDT at Moscone West in San Francisco, as part of Cloudflare Connect. Cloudflare will pay for the five finalists to fly to San Francisco for the competition.

Connect brings together people building and thinking about what comes next for the Internet. This year’s lineup includes AI pioneer Dr. Fei-Fei Li; Idealab founder Bill Gross; organizational psychologist and author Adam Grant; AMD CTO Mark Papermaster; Vue.js and Vite creator Evan You; Lovable co-founder and CTO Fabian Hedin; and OpenAI member of technical staff and creator of OpenClaw, Peter Steinberger.

For five young companies, we are reserving part of that stage. Each startup will have 5 minutes to pitch, followed by 3–5 minutes of questions by the judges.

There is something we like about that symmetry. Sixteen years ago, Cloudflare needed someone to take a chance on an infrastructure company with a difficult story to tell and give us a few minutes in front of the right room. Today, we are fortunate enough to have a stage of our own, and we want to pass that same opportunity on to companies that are just getting started.

Start small. Build something enormous.

There is a practical reason Cloudflare spends so much time working with startups: very small groups of people have an uncanny ability to attempt very large things.

The problem is that ambitious software increasingly depends on infrastructure that, not very long ago, only the largest technology companies in the world could afford to build for themselves. Global compute, storage, networking, security, real-time systems, AI inference, and the ability to survive the possibility that the thing you made suddenly becomes popular should not require a company to first become enormous.

We think you should be able to reach for those capabilities on day one.

That is part of the idea behind Cloudflare for Startups, through which eligible early-stage companies can receive up to $350,000 in Cloudflare credits for one year. It is also part of the reason we continue expanding Cloudflare’s developer platform: a tiny team should be able to build something on Tuesday and if the Internet decides it likes it on Wednesday, spend Thursday focused on the product rather than hastily becoming experts in global infrastructure.

Cloudflare began with its own slightly unreasonable premise: that the performance, security, and global compute available to the largest companies on the Internet should be available to everyone on day one. In 2010, we got a chance on a stage to explain why that matters.

Sixteen years later, we have a much larger network, a somewhat larger team, and significantly better circuit breakers.

Now we want to hear what you’re building.

Apply for The Cold Start.

Introducing Forge: the open source pipeline for generating SDKs, CLIs, docs, and more

Post Syndicated from Dimitri Mitropoulos original https://blog.cloudflare.com/forge-open-source-generation-pipeline/

Today we’re introducing Forge, a fresh approach to generating SDKs, CLIs, docs, and libraries. Forge is an open source, pluggable generation pipeline that anyone can deploy and run for free.

Forge is early in its life, but already generates the output required for the cf CLI, and over the next few months will power Cloudflare’s API documentation, SDKs, and much more.

We built Forge because we needed it ourselves in order to treat agents as our customers. Now, we’re open sourcing it because we think everyone should be able to generate all the surfaces that agents need. It used to be that only developer products needed CLIs, API SDKs, MCP servers, all with great corresponding docs. Now these are table stakes for every product.

Our API outgrew our generators

Cloudflare’s API has over 3,500 operations, and the hundreds of services that power these APIs are written in many languages, including Rust, Go, TypeScript, and Python. As we embarked on building a CLI for the entire Cloudflare API, including our SDKs and API docs, we needed a code generation pipeline that could handle this scale. That pipeline needs to be flexible enough to work across languages and the ways each of our engineering teams operate.

We needed a way to reduce coordination overhead between teams. When a Cloudflare product team makes an API change, they need to be able to use a preview build of the Cloudflare-wide CLI, SDK, and docs site that will be generated, before merging that change and shipping to customers. We needed a way to ensure they didn’t inadvertently break the generation pipeline. And we needed a system that we could extend to generate more than just an SDK, from Cap‘n Web to MCP and beyond.

We’ve tried several hosted products that attempt to solve this, and relied on some in production. None of them solved this problem for us, and some have shut down entirely. One team would merge a change that inadvertently would break the generation pipeline, another team would discover this at release time, and we spent too much time swimming upstream through hosted tools we couldn’t control, coordinating changes between teams and vendors.

That’s how we started building Forge.

Forge seeks to fix all these problems: it runs in CI, on each team’s API repos, just like our AI code reviewer and test pipelines. It lints every change, and then generates preview builds of the CLI, docs, and SDKs with just your changes highlighted that you can install to test. It’s the same premise as Workers Previews: a full preview build for every change, but applied to SDK generation at scale, including when the API surface is distributed across hundreds of services and repositories. That’s what Forge seeks to deliver.

Forge transformers can generate anything, including Cap’n Web

Cloudflare has more reasons than most to want a generator that can go well beyond the normal language targets. Cap’n Web is Cloudflare’s RPC system that lets TypeScript call a remote API as if it were calling a local method:

Forge makes it possible to take an OpenAPI spec and generate Cap’n Web directly. This opens the door to generating bindings from Workers to other APIs. After all, bindings in the Workers runtime are implemented as Workers that expose RPC methods.

This isn’t specific to Cap’n Web: other popular tools you may already rely on need the same thing. If you use TanStack Query, you’d ideally want to be able to generate TanStack Query bindings for your application, built directly from your API itself. Always up to date, always validated against your real API. The same is true of generating Zod or Valibot schemas, MCP servers, or anything else that makes it easier to consume your API.

This is possible because Forge code generators are flexible. They’re built for flowing information from one output to another.

Forge transformers can be chained: generate outputs from other outputs

We’ve designed Forge to be pluggable, and support many input and output types. Forge provides CLI, SDK and docs generators, but there’s nothing stopping you from adding a transformer that generates a library-specific package or even a full dashboard or application. Forge supports OpenAPI as an input type today, but we’ve designed it to allow AsyncAPI, GraphQL, Cap’n Proto, Protobuf, or other input formats in the future.

This is about more than just compatibility: it lets you chain targets, using one target output to produce others. This is common in other generators where the CLI and Terraform targets are produced from the Go SDK. But what’s missing, and what Forge provides, is a way for the user to control this chaining system themselves.

We needed a solution for this ourselves, because our own cf CLI is written in TypeScript, which other SDK generators don't generally chain from for CLIs. But our own situation made us recognize the deeper problem: why should any SDK generator tool make this decision for you? Maybe you’re a Python shop, and you want the CLI to be in Python.

If you’re thinking “Well, but who cares if it’s in Python or not? The code is automatically generated,” it’s because CLIs are different. CLIs often introduce local-only behaviors that wouldn’t make sense in an SDK. Behaviors that you write by hand since they’re inherently not something backed by any API call. For example, the cf CLI has commands like cf dev and cf build that are added on top of the rest of the generated output. These commands need to call TypeScript APIs from other packages like Vite. 

Now let’s add docs to the mix. If you’re generating your CLI and your docs purely from your OpenAPI spec, how do you feed those handwritten commands back into your docs, so they can be documented alongside the rest?

We couldn’t find an existing tool that does this today, and yet this is exactly what we need for cf. So we’re building it into Forge.

Change your API without breaking users

Forge is also setting us up for better API versioning. Cloudflare’s v4 API has been the one major version of our API for 10 years. Since then, it appears like we haven’t launched any new major versions, but by SemVer definitions we’ve made quite a few changes worthy of a new major version. At the same time, several operations across our API feature internal ‘v2’ tags or ‘beta’ identifiers that have long outlived that part of the product’s lifecycle.

After so many years of our v4 API, we’re keenly aware that a big new v5 would leave a lot of our customers behind. That’s why, with Forge releasing artifacts along the way, we’re working on an API versioning approach that allows us to release new major API versions without breaking old clients or SDKs.

We’ll have more on our SDKs very soon, including TypeScript, Rust, Python, Go, PHP, and Terraform. Especially Terraform. We know that upgrading any Terraform provider comes with its own set of rigor, and we’re going to put extra special care into the Terraform transition.

Critical tools should be open to all

We believe that building tools for APIs is a core part of the Internet, and you should be able to do that without needing a SaaS product. You should own your SDKs, CLIs, and docs. And if you generate them, then you should be able to do whatever you want, wherever you want, for free.

That’s why we’re making Forge available open source under the permissive Apache 2.0 license. We want people to join in on this journey with us, and contribute.

Or not? Maybe you want to keep everything to yourself. Go for it! You can run Forge on your own for any purpose, with custom modifications, for free, in private.

Acknowledgements: This project was also made possible by the design and implementation efforts of Dan Carter, Steven Chong, Krishna Paritala, and Shelley Jones.

Four months of VoidZero at Cloudflare: making the open-source JavaScript toolchain faster for all humans and agents

Post Syndicated from Evan You original https://blog.cloudflare.com/voidzero-update/

When VoidZero joined Cloudflare four months ago, we made a commitment to open source, promising that Vite, Vitest, Rolldown, Oxc, and Vite+ will stay open source, vendor-agnostic, and community-driven.  As part of Cloudflare’s Birthday Week, where Cloudflare gives back to the Internet, we thought it’d be a good time to check in on how we’ve been doing against this commitment.

In the four months since VoidZero joined Cloudflare, we’ve shipped more than 80 releases, closed over 1,200 issues, and landed some big performance improvements, including:

On top of this, Vite+, which unifies the entire toolchain with a set of great defaults, is now 1.0. And we’re making progress towards shipping “Bundled Dev” (f.k.a. Full Bundle Mode) — a new development mode that has been shaped by working with customers with massive web applications, including Cloudflare’s own dashboard. By being part of Cloudflare, our engineering team gets a much better view into scenarios that only occur in massive codebases.

VoidZero’s mission is to make the next generation of JavaScript developers more productive than ever before. And in 2026, that means making developers’ agents faster.

A faster developer experience for humans and agents

Developer experience and performance were always about improving the feedback loop while creating software. But now when we optimize for developer experience, we no longer just optimize for humans — we optimize for agents. And as inference gets faster, the speed of type-checking, linting, or building the code becomes the bottleneck again. The longer those processes take, the longer an agent has to wait before making progress and completing its goals.

VoidZero has always been focused on performance, but this new era of software has given us even more reason and motivation to make the entire toolchain faster for both agents and humans. And since the tools in the VoidZero toolchain are built on each other, from the compiler (Oxc) to the bundler (Rolldown) to the build tool (Vite), the linter (Oxlint) and the test runner (Vitest), any optimization at one layer automatically benefits everything built on top.

Oxc React Compiler — 10x faster compile times for React.js apps

We’ve recently shipped the Oxc React Compiler, a rewrite of the React Compiler, based on the React team’s rewrite in Rust. It is 10x faster than the original Babel implementation, uses less memory, and has a more complete implementation with better error handling. If you use Vite, you can enable the Oxc React Compiler by installing the oxc-transform-react package and enabling the compiler flag:

Vitest 5 — up to 50% faster than Vitest 4

Vitest 5, released in September, cuts test times by double-digit percentages across many common scenarios.

  • Faster test runs. Vitest shares transformed files across projects, caches modules on disk, and sends less data between its main process and workers.
  • vitest doctor. It breaks down setup, import, transform, and test time, then tests other configurations and recommends faster settings.
  • Trace View. It records browser interactions, assertions, and DOM snapshots, so you can replay failures or inspect them in an HTML report.
  • Conditional mocks with vi.when. Map arguments to responses without writing a manual mockImplementation.
  • Better benchmarks. Benchmarks now work like regular tests, with fixtures, hooks, retries, filters, and assertions.
  • Fewer false passes. Vitest fails unawaited async assertions, clears mock calls before each test, and adds the --repeats flag to help find intermittent failures.

tsgolint — now stable, up to 18x faster than ESLint in large codebases

tsgolint, the type-aware linting engine behind Oxlint, is now stable. It catches bugs that require TypeScript type data while running 12 to 18 times faster than ESLint with typescript-eslint.

Enable type-aware linting and TypeScript diagnostics in your Oxlint config:

Oxfmt — 7x faster than Prettier with formatters now written in Rust

Oxfmt brings the speed of the Oxc toolchain to formatting. We rewrote its JSON, CSS, SCSS, Less, GraphQL, and YAML formatters in Rust, making Oxfmt many times faster than Prettier, while keeping Prettier-compatible output and ergonomics.

Bundled Dev — faster dev server for larger apps

We’ve made progress towards shipping “Bundled Dev” (f.k.a. Full Bundle Mode), which uses Vite’s production bundler during development. This should lead to significant dev server speed-ups in larger applications and reduce network overhead when working with remote sandboxes.

We’re looking forward to bringing Bundled Dev out of experimental status soon. Cloudflare’s Dashboard already uses Bundled Dev for all internal developers.

Vite+ is now 1.0 — a unified toolchain

Speeding up tools is one way of making humans and agents ship software faster. Another way is by reducing decision fatigue (“Which linter shall I use?”) and providing great defaults. We are excited to announce that Vite+ is now 1.0. Vite+ bundles all of VoidZero’s tools together into a single unified and fast toolchain.

Vite+ ships with Vite 8, Vitest 5, Rolldown, Oxlint, Oxfmt and task caching built in, and comes with great defaults. Check out the Getting Started guide to try it out today.

Cloudflare’s Open Source Investment

VoidZero was born in open source, and we believe a more sustainable open-source ecosystem benefits everyone. VoidZero is a proud member of the Open Source Pledge and as part of Cloudflare, we have the opportunity to expand that impact.

Cloudflare committed $1M to a Vite ecosystem fund to support maintainers and contributors in the original announcement. Since then, Cloudflare has committed an additional $1M to open source. Together, we’re doubling down on open source.

The open-source toolchain for the entire JavaScript community

There is more to come! Features we plan to ship in the next few months include major improvements to Oxc’s parser with up to 3x potential speedup, a re-designed chunking algorithm in Rolldown, and an open-source, self-hostable version of Void, the Vite-native deployment platform built on top of Cloudflare. Cloudflare and VoidZero both recognize our responsibility that we have to developers, and we do not take the community’s trust in us for granted. We’re in this for the long haul.

We are excited to keep shipping faster tools, and will continue to make every decision with the community in mind, just like we did when raising venture capital, or when we open sourced Vite+, or when we joined Cloudflare. Thank you for continuing to trust us with your projects and apps, supporting us, and building with us.

How Cloudflare addressed a cross-tenant data exposure vulnerability in Containers

Post Syndicated from Rushil Mehra original https://blog.cloudflare.com/containers-cross-tenant-vulnerability/

On September 4, 2026, Oren Yomtov, a security researcher from Accomplish, responsibly reported a vulnerability affecting Cloudflare Containers and Cloudflare Sandboxes (which is built on Containers), through Cloudflare’s bug bounty program. Cloudflare has fully remediated the vulnerability, and we have no evidence that customer data has been compromised. 

This post was prepared in collaboration with Oren Yomtov and the Accomplish security research team, whose detailed report and controlled testing helped us validate the issue and respond quickly.

Cloudflare Containers run workloads on multi-tenant infrastructure and automatically assign them to eligible servers; customers cannot select the underlying host. The researchers demonstrated that a customer with a Workers Paid account could recover residual disk blocks previously used by Containers on the same host. The technique could not target a particular customer, workload, host, or data, and residual data was not guaranteed to be present.

Cloudflare applied a fix across the Containers fleet, with no customer-side configuration changes required. Within the historical disk-I/O telemetry available to us, we identified no evidence of malicious exploitation. Activity we could attribute to the reported technique came from the researchers and Cloudflare engineers conducting authorized validation.

Here, we explain the underlying storage behavior, its potential impact, our investigation, and the actions we took in response.

How container storage allocation works 

Cloudflare Containers use Linux device mapper thin provisioning (dm-thin) to provide each container with a writable root disk. Each container lives inside a dedicated virtual machine powered by the Firecracker virtual machine monitor. Firecracker presents this disk to the virtual machine as /dev/vdc.

Thin provisioning allocates physical storage only when a virtual disk writes to a previously unmapped region. The affected storage pools used a 64 KiB thin-block size. When the thin volume backing a container's root disk was deleted, its physical blocks were returned to a pool that served workloads belonging to multiple customer accounts.

The affected pool configuration included the following option:

skip_block_zeroing

With this option configured, dm-thin skips zeroing newly allocated blocks before making them accessible. Consequently, when a previously-used 64 KiB block was reassigned, a full-block write replaced its previous contents, but a smaller write changed only the written portion. The remainder could retain data from the block’s previous owner.

How the exploit worked

Reading an unmapped region of a new thin disk did not reveal residual data. For an unmapped region of the thin device, dm-thin returned zeroes without allocating a physical block.

The proof of concept identified 64 KiB-aligned regions corresponding to free space in the guest’s ext4 filesystem and wrote one aligned 4 KiB block into each region.

When such a write reached an unmapped thin block, dm-thin allocated a physical 64 KiB block from the shared pool. The 4 KiB write replaced only that portion of the block, and because block zeroing was disabled, the remaining 60 KiB could retain data from a previous container.

A subsequent raw-device read could therefore observe bytes that the new container had never written.

The proof of concept performed the following steps:

  1. Create a container using a Workers Paid account.
  2. Open the writable root disk at /dev/vdc.
  3. Read the disk and record a baseline.
  4. Write one 4 KiB block into each selected 64 KiB region corresponding to ext4 free space.
  5. Read the resulting blocks again.
  6. Examine only the portions not overwritten by the new container.

The submission included counts, block offsets, sizes, checksum results, and truncated hash prefixes. Although the researchers recovered raw blocks to validate the issue, the materials provided to Cloudflare contained no third-party filenames, identifiers, credentials, hostnames, addresses, or recovered content values. As described below, the researchers have also confirmed that they securely deleted the recovered data.

How the vulnerability was validated 

The researchers used ext4 directory block checksums to distinguish blocks belonging to their own test filesystem created for the proof of concept from blocks originating from other filesystems.

When ext4 uses the metadata_csum feature, directory block checksums incorporate values associated with the filesystem and inode. 

Across six production placements, the researchers reported:

  • All 5,614 testable directory blocks.
  • Zero of those blocks were attributed to the researchers’ filesystem. 
  • 2,700 distinct foreign directory inodes identified through checksum analysis.

To validate the method, the researchers tested it against blocks they had deliberately created and deleted in the controlled test filesystem used for the proof of concept. The method correctly attributed all 162 blocks to that filesystem.

The researchers ultimately observed residual material on 18 of 24 placements and 20 of 22 underlying nodes across four continents. The recovered block types included directory structures, database pages, and structurally complete SQLite databases. The researchers reported using scripts that output only aggregate counts and format checks, not recovered file contents. The materials submitted to Cloudflare contained no recovered content values or third-party identifiers. The researchers subsequently confirmed that recovered data under their control remained confidential and was securely deleted following submission, consistent with Cloudflare’s HackerOne disclosure policy.

Impact 

The vulnerability would potentially have allowed for a customer with a Workers Paid account to recover residual data from storage blocks previously used by other customers’ Containers on the same underlying host.

A successful exploitation would have crossed the tenant-isolation boundary and could disclose filesystem metadata, directory structures, database pages, and application data.

However, an attacker could not select a particular victim or access an actively attached disk. Exposure depended on Cloudflare’s workload placement and which previously released blocks dm-thin reassigned. Moreover, the researchers did not demonstrate modification of another customer’s active data or impact to workload availability.

How we mitigated the vulnerability

Our first mitigation was to remove skip_block_zeroing from the dm-thin pool configuration across the fleet. This restored dm-thin’s default behavior of clearing newly allocated blocks before exposing them to a container. It stopped the reported technique, in which a small write triggered allocation and a larger read recovered residual data from the remainder of the block. The researchers independently confirmed that their proof of concept no longer worked after this change.

Zeroing new allocations did not sanitize blocks already mapped into existing thin devices. These mappings existed in running container disks and in each host’s cache of prepared dm-thin snapshots for OCI image layers. A new container could inherit mappings from a cached layer without allocating those blocks again, allowing residual bytes in unused regions, including ext4 free space, to remain readable through raw reads of /dev/vdc.

We therefore also retired all running container disks and removed cached image snapshots created before the mitigation. We drained hosts during off-peak hours, restarted the VMs on each host, and cleared each host's image cache so that disks and cached layers were recreated using zeroed allocations. We have completed this cleanup across the Containers fleet.

No evidence of exploitation

As part of our response, we investigated whether other workloads showed activity consistent with the reported exploitation technique. We reviewed retained historical disk-I/O telemetry from our container infrastructure, using the researchers’ proof of concept and our internal reproduction as reference activity.

The proof of concept produced a characteristic relationship between writes and reads. When a 4 KiB write reached a previously unmapped region, it could trigger allocation of a reused 64 KiB storage block. With zeroing disabled, the remaining 60 KiB could retain data from a previous container. Subsequent reads could therefore recover substantially more data than the new container had overwritten.

Using these characteristics, we developed detection signatures and applied them to the historical telemetry available to us. We identified activity attributable to the researchers and Cloudflare engineers conducting authorized validation, and did not identify additional activity consistent with the reported technique.

We saw no evidence that this specific attack vector was exploited by anyone else.  

Cloudflare customers are protected

As we noted above, Cloudflare has patched this vulnerability and remediation does not require any further action by Cloudflare customers. In addition, we found no evidence of any malicious actor abusing this vulnerability.

Moving quickly with transparency 

We thank Oren Yomtov and the Accomplish security research team for their thorough research, responsible disclosure, and collaboration on this post. We encourage the Cloudflare community to submit any identified vulnerabilities to help us continually improve the security posture of our products and platform.

We also recognize that the trust you place in us is paramount to the success of your infrastructure on Cloudflare. We take these vulnerabilities very seriously and will continue to do everything in our power to mitigate impact. We deeply appreciate your continued support and trust in our platform, and remain committed not only to prioritizing security in all we do, but also acting swiftly and transparently whenever an issue arises.

Timeline

  • September 4, 15:26 UTC: Oren Yomtov from Accomplish reported the issue through HackerOne.
  • September 4, 18:45 UTC: Cloudflare opened a security incident and confirmed the production setup that caused the flaw.
  • September 4, 21:27 UTC: Cloudflare merged the runtime fix and its reuse test.
  • September 4, 22:03 UTC: Cloudflare merged the changes for new and live pools.
  • September 4, 23:15 UTC: Cloudflare started rolling out the changes.
  • September 7, 06:13 UTC: Cloudflare completed rolling out the changes and began clearing old pool data.
  • September 14, 10:50 UTC: The researchers reported that their proof of concept had stopped working.
  • September 14, 12:52 UTC: Cloudflare awarded the researcher a bounty.
  • September 19, 15:03 UTC: Cloudflare completed cleanup of all pre-mitigation cached snapshots across the affected fleet.