Tag Archives: Workers

Everything we launched during Birthday Week 2026

Post Syndicated from Carlos Armada original https://blog.cloudflare.com/birthday-week-2026-wrap-up/

We celebrated our 16th birthday last week by sharing how we’re building a better Internet for today’s world. As Matthew and Michelle reflected in this year’s Founders’ Letter, this year saw some of the most consequential changes in the history of the Internet.

For the first time, automated traffic surpassed human activity. AI is empowering people to build like never before, leading the Internet to grow massively in scale and unlocking more ambition and creativity. As we witnessed the influence that agent-driven recommendations have on consumer choices, we identified the need for a new approach that creates space for new businesses to succeed.

Each day of Birthday Week explored a different way we are helping to build the future of the Internet. We began on Monday by strengthening our commitment to open source. Tuesday focused on application security and the post-quantum transition. On Wednesday, we explored new economic models for the agentic Internet. Thursday, we expanded the Developer Platform with new tools for data analysis, storage, AI, and agent development. Finally, we closed out the week by launching features that make Cloudflare faster, easier to operate, and more accessible to everyone. As a special Birthday Week follow-up, we shared an update on our intern program, one year after announcing our goal to hire 1,111 interns. Interns directly contributed to many of the projects launched this week, including EmDash, post-quantum visibility, CryptoLabe, and Protected Quick Tunnels.

We shipped 46 announcements this week. In case you missed any, here’s the full list of everything we announced during Birthday Week 2026.

Monday, September 28 – Commitment to open source

With the announcement of our new CLI, which we released alongside the pipeline we use to generate it and our SDKs and docs, we shared how we’re building to support agents and developers as they use Cloudflare — and supporting the projects that you rely on, too.

What

In a sentence…

Introducing cf: the agentic CLI for the entire Cloudflare API

The new cf CLI mirrors the Cloudflare API, uses JSON-first output and typed configuration, and gives people and agents one consistent command-line interface.

Introducing Forge: the open source pipeline for generating SDKs, CLIs, docs, and more

Forge is a pluggable, open-source pipeline that runs in CI to generate SDKs, CLIs, documentation, and other interfaces directly from API definitions.

Introducing EmDash – the spiritual successor to WordPress that solves plugin security

EmDash is an open-source, Astro-based serverless CMS that runs plugins in isolated Worker sandboxes with explicitly approved capabilities.

Four months of VoidZero at Cloudflare: making the open-source JavaScript toolchain faster for all humans and agents

Since joining Cloudflare, VoidZero has delivered more than 80 releases across the Vite ecosystem, and its previously commercial Void platform will become fully open source.

Next.js applications, powered by Vite: introducing Vinext 1.0

Vinext 1.0 turns an AI-built experiment into a production-ready, portable way to run Next.js applications on Vite.

The road to the agentic browser: A Kitesurf update

Kitesurf, our Workers-based browser for agents, adds WebMCP support, faster DOM operations, broader web compatibility, and terminal-based rendering.

How fast is the web? Explore billions of real-user measurements with BEACON

BEACON makes billions of anonymized real-user performance measurements from 10,000 major websites available as a public BigQuery dataset.

Supporting native Rust in Workers with the new Emscripten target for wasm-bindgen

Experimental Emscripten target support lets developers bring more native Rust libraries and applications, including progress toward Tokio support, to Workers.

Introducing The Cold Start: pitch your startup live at Cloudflare Connect

The Cold Start gives five early-stage companies the opportunity to pitch live at Cloudflare Connect and compete for resources to help them grow.

Tuesday, September 29 – Helping secure the agentic Internet

Technological progress is rapidly changing how we think about application security. We announced our intention to become a certificate authority, as well as how we’re preparing foundational Internet cryptography for the post-quantum era and adapting application security to counter AI-driven attacks.

What

In a sentence…

Building a certificate authority for the whole Internet

Twelve years after launching Universal SSL, Cloudflare announced its intention to become a public certificate authority (CA) and add resilience to free, automated certificate issuance.

Building a post-quantum certificate authority with Merkle Tree Certificates

Our planned CA will issue free Merkle Tree Certificates designed to make post-quantum authentication practical without imposing large certificate and handshake costs.

Using AI to chart a course for our post-quantum migration

CryptoLabe uses AI to find and classify cryptography across our codebase as Cloudflare works toward completing its post-quantum migration by 2029.

Preventing quantum downgrade attacks against IPsec

Cloudflare helped develop an IETF extension that authenticates the full IKEv2 transcript and prevents attackers from downgrading post-quantum IPsec tunnels.

Is your domain using post-quantum encryption? Now you can see for yourself

HTTP Analytics, Log Explorer, and Logpush now show whether requests negotiated post-quantum key exchange, giving customers evidence they can inspect and report.

Enforce positive security with Cloudflare Application Profiles

Application Profiles learns the expected structure of HTTP requests so customers can identify deviations and enforce what valid application traffic should look like.

We tested our own WAF with frontier AI models. Here's what we found

An adaptive AI red-team system found WAF detection gaps across six attack categories, helping us improve normalization and managed rules for customers.

Introducing Threat Signals: agentic skills for open-source threat intelligence, free for every Cloudflare account

Threat Signals turns open-source reporting into structured indicators and connects the context to WAF rules, while the Threat Events Platform expands to every account.

Adaptive application security for the AI era: how Cloudflare connects code, traffic, and intelligence to stop attacks

Our application-security framework connects discovery, governance, runtime protection, investigation, and response in a continuous learning loop.

Wednesday, September 30 – Powering the agent economy

With our announcements of Pay Per Use and the release of our Monetization Gateway in beta, we shared how we’re building support for a new economic model that empowers creators to monetize their content and services.

What

In a sentence…

The Internet has a second audience

AI agent requests have grown rapidly, and our strategy helps creators see agents, set terms for access, and get paid when agents use their work.

Cloudflare Containers, rebuilt to scale agent sandboxes

Containers now has faster startup, flexible image and instance selection, new scheduling controls, and filesystem snapshots for persistent agent workspaces.

Monetization Gateway beta: charge AI agents for consumption with HTTP 402

Monetization Gateway lets sellers put a price on resources behind Cloudflare and collect agent payments using HTTP 402 and x402.

Pay Per Use: when AI uses your work, you should get paid

Pay Per Use gives enrolled publishers usage reports, billing, and payouts when verified AI buyers use their content.

Simplifying domains for people and agents

A new domain-search experience and expanded Registrar APIs make it easier for both people and agents to search, register, transfer, and manage domains.

Identify AI model overuse with User Insights

AI Gateway User Insights identifies tasks, model fit, and overuse, so teams can understand where a smaller or less expensive model may work.

Detect and send production issues straight to your agent

Issues groups Workers errors and sends the relevant stack traces, logs, and traces to coding agents or any webhook for faster investigation.

Cut your AI spend with AI Gateway's Auto Router

Auto Router classifies each request at the edge and sends it to a suitable model, reducing cost while preserving response quality.

Cloudflare Impact reaches $100 million in donations

Initiatives including Project Galileo, the Athenian Project, and Cloudflare for Campaigns have now delivered more than $100 million in donated services.

Thursday, October 1 – Bringing more of the developer stack to Cloudflare

We expanded what is possible to achieve on Cloudflare’s platform with the general availability launch of Cloudflare Basin, our data analytics platform, the launch of K2, a durable serverless event stream, and the announcement of our new contest — inviting developers to build a Git platform designed for agentic development.

What

In a sentence…

Introducing Cloudflare Basin: an open, serverless data platform, now generally available

Basin is now generally available, giving developers a serverless platform built on Apache Iceberg and R2 for ingesting, managing, and querying large datasets.

Support for modern cryptographic algorithms in Workers

Workers adds opt-in native Web Crypto support for ML-KEM and ML-DSA, giving developers post-quantum primitives without bundling their own implementations.

AI Search is now generally available

AI Search reaches general availability with visual search, OCR for scanned PDFs, larger files, and support for any chat model.

We want you to build the next Git platform on Cloudflare

Artifacts enters open beta and a new competition invites developers to build a Git platform designed for the era of AI agents.

Announcing Cloudflare K2: serverless event streams

K2 provides durable, ordered event streams on R2, separating producers and consumers without the operational overhead of managing broker clusters.

Cloudflare OS: your company's agent workspace, managed for you

Cloudflare OS provides an agent workspace connected to an organization’s data and systems, with a waitlist open for fully managed deployments.

Introducing Workers KV Instant – powered by Quicksilver

Workers KV Instant delivers sub-two-millisecond p99 reads and fast global replication across more than 300 locations using the familiar Workers KV API.

One year later: Sovereign AI and the fight for choice

We are expanding local open-source model choice and model-agnostic security tools, so nations can pursue AI sovereignty without isolation.

Introducing Clef: our open-source decision models, and new RL fine-tuning platform

Clef and Clef-flash are open-source decision models for fast classification and agent workflows, accompanied by a platform for reinforcement-learning fine-tuning.

Friday, October 2 – Delivering a faster, simpler Internet for everyone

We wrapped up the week with major updates to Cloudflare Observability, alongside adding Cloudflare Traces, network performance improvements that make Cloudflare faster, and an announcement on how we’re supporting civil society organizations.

What

In a sentence…

8 major updates to Cloudflare Observability

Eight updates bring logs, traces, analytics, alerts, dashboards, querying, and telemetry export into one observability platform with simpler pricing.

Introducing Cloudflare Traces: follow requests through our entire platform

Cloudflare Traces provides request-level visibility across security rules, transformations, cache, Workers, services, and origins without requiring an agent or SDK.

Updates on our pledge to make Cloudflare features accessible to everyone

One year after our pledge, Logpush, multi-account governance, higher platform limits, and other capabilities are available to more customers across plans.

Announcing Cloudflare OHTTP Gateway – expanding access to Cloudflare's privacy-preserving infrastructure

A self-serve OHTTP Gateway enters closed beta, while Privacy Gateway becomes Cloudflare OHTTP Relay to distinguish the two roles.

Follow the thread: a new dashboard to investigate account abuse

Account Abuse Protection uses stateful analysis and privacy-preserving Hashed User IDs to help teams investigate credential stuffing and fake-account creation.

Protected Quick Tunnels: simple accountless authentication for your next dev project

Quick Tunnels now support email authentication, letting developers share a local application with selected people or domains without requiring Cloudflare accounts.

Building for good: How civil society organizations are automating on Cloudflare

Civil society organizations are using Cloudflare’s developer platform to automate and scale work that protects human rights and the public interest.

2026 Birthday week: network performance update

Using an expanded real-user measurement methodology, Cloudflare now ranks as the fastest provider across 74% of the top 1,000 networks.

Introducing Web Search API via AI Gateway

AI Gateway’s Web Search API brings current web context from multiple providers into model calls through REST APIs, Workers bindings, or customer-managed keys.

Streamline: custom video pipelines with Cloudflare Stream and Workers

Streamline is an open-source example for building continuous video pipelines by combining Workers, Durable Objects, and a containerized media engine.

Building the Internet’s next chapter together

Across this week’s announcements, we kept returning to a consistent theme: the Internet should continue to open up more opportunities for people to create, contribute, and succeed. That means open tools developers can shape, security that keeps pace with new threats, a fairer exchange between agents and the people whose work they use, and infrastructure designed for the agentic Internet.

For 16 years, we have been building alongside developers, creators, researchers, customers, partners, and open-source communities. Your ideas, feedback, and willingness to challenge us have shaped Cloudflare, and that collaboration matters now more than ever.

Streamline: custom video pipelines with Cloudflare Stream and Workers

Post Syndicated from Willi Geiger original https://blog.cloudflare.com/streamline/

Cloudflare Stream is a powerful broadcasting platform that, for many of our customers, just works. But what if you wanted to render dynamic annotations on a livestream or create an alternate version of a hosted video with burned-in subtitles? You would need to run a custom video pipeline.

Today, we’re releasing a new developer playground, Streamline, that demonstrates how you can build a system to deliver these bespoke video experiences on Cloudflare’s Developer Platform. We’ll walk you through how Streamline leverages Workers, Containers, and several media protocols to modify video — and immediately publish that output as livestream or new hosted video. You’ll also have the opportunity to try it for your projects.

A processing pipeline needs a durable, long-running environment that can run specialized, compiled code with predictable memory and CPU capacity. Video streams can run for minutes or hours, so the media process needs a lifecycle independent of the request that started it. An application should be able to start a pipeline, send its input, inspect it, and stop it without needing to keep a single request open for the entire duration.

Cloudflare provides the primitives we need. Containers are long-lived runtimes suitable for media processing. Durable Objects help with orchestration. Finally, Workers are perfect for control signaling and monitoring.

For Streamline, we built a media engine running in a Container to handle media processing in real-time. The Container is controlled by a Worker exposing control, preview, and testing to an agent or user. Processing will continue even if the Worker disconnects. We've architected Streamline with modular components so that the media engine could be replaced with dedicated encoding products in the future.

Architecture

A Streamline deployment consists of two components: the Media Engine, which handles media input/output and processing, and a controlling Application, which creates, configures, observes, and stops media sessions.

Media Engine

The Media Engine has two components:

  • Controller. This is a control harness written in Go that implements an HTTP server, receives incoming requests, and translates them into operations that can be executed by the media engine.
  • Processor that performs the actual media processing. The current implementation uses FFmpeg, but that is an internal implementation detail rather than part of the user-facing API.

The Media Engine is hosted in a Container, and handles all media input/output as well as processing. It can pull RTMPS playback over the network from one Stream Live input and publish RTMPS output to another Stream Live input. It can pull a Cloudflare Stream HLS manifest and its segments to use hosted videos as input. It can accept video input from a source supplied by the controlling application, for example a webcam. It can publish preview video over an outbound WebSocket to a Durable Object relay. An application that needs preview can connect to that relay through its own WebSocket.

Application 

The application is built using Workers, and can be a full-stack browser application, an agent, or an embedded system. It consists of:

  • User interface (UI) including client logic, identity and access policy. This post uses a browser application as its concrete example, so it also includes a browser interface.
  • Orchestrator coordinates the session, the Container lifecycle, and preview relay. The orchestrator is implemented by a Durable Object.

It is possible to run the system locally during development, in which case the container is just a local Docker instance and the Durable Object is not used: there is a single user, the controlling application does not require authorization for local access, and the video preview can connect directly to a WebSocket on localhost.

When these components are deployed to Cloudflare, an authorized user or agent can visit the Worker to start a new session. This spins up a new Streamline container if needed, manages its lifecycle automatically, exposes an API to perform a number of video manipulation operations, and routes inputs from and outputs back to Cloudflare Stream.

Time for a technical deep dive on how the system works.

Container lifecycle and session management

The controlling Worker application initiates a long-running media processing session. After starting the session, the application can disconnect and reconnect safely, while the Container continues processing until the controlling application stops it. We also include a maximum duration to ensure a session is always eventually closed down and can’t run indefinitely, even without external control. While a media processing session is running, the container instance is unavailable for other applications to use.

A Cloudflare Container will automatically sleep if it has not received any incoming requests since a defined interval. However, in our case, once the pipeline is running, it must continue even if the controlling application disconnects and it receives no requests. We can implement this behavior by overriding the onActivityExpired() callback on the container. If the expiry time has not been reached, then we renew the activity, otherwise we destroy the container.

API

The HTTP server implemented by the Go harness and the Durable Object associated with the Container together define the low-level interface to the system. However, we wanted to provide an abstraction over this, so the system is as agnostic as possible to who or what is controlling the session and any unnecessary details of the backend implementation.

We implement this by exporting two packages from Streamline:

  • @cloudflare/streamline/client Defines a high-level, session-based API.
  • @cloudflare/streamline/ Exposes the Durable Object base class associated with the container. This routes the API requests, implements the preview relay server described below, and provides hooks for security and access policy.

In a remote deployment, the controlling Worker is expected to import @streamline/cloudflare and define a concrete subclass of the Durable Object exposed by the container that can be used for application-specific logic and storage.

In local mode, where there is no Durable Object, the frontend defines a thin adapter layer that maintains the session-based API, but connects directly to the local Docker instance with no access controls, etc.

The example below shows how the controlling application can use the API to access Streamline, prepare a session, and start a video processing pipeline.

config is a JSON object that defines the processing pipeline to be executed, described more in subsequent sections.

The table below shows the complete list of all API calls.

Client method

Function

createStreamline()

Creates a new Streamline instance.

streamline.sessions.create()

Creates a new processing session.

streamline.sessions.resume(id)

Reconnects to an existing session.

session.start(config)

Starts a new processing pipeline.

session.ingest(chunk)

Sends a chunk of video data in “webcam” mode.

session.annotation(png)

Updates the transparent annotation overlay.

session.metrics()

Receives metrics about the current session.

session.stop()

Stops the processing in the current session.

Defining and running a video processing pipeline

session.start() constructs and runs a processing pipeline. It takes a single argument which is a JSON configuration object defining the processing to be performed:

  • Input(s)
  • Operations
  • Output

The example below starts a pipeline that takes an RTMP (real-time messaging protocol) broadcast as input (for example, a feed of a Stream Live input receiving an inbound livestream), applies an overlay image with transparency, and sends the output to an RTMP destination (for example, to another Stream Live input for recording or broadcast). This allows the Worker application to create a modified version of a livestream in real time.

Video-on-demand input via HLS

Streamline can also ingest streaming video input via HLS (HTTP live streaming), for example a video hosted on Cloudflare Stream. The example below shows how a Worker application could run a pipeline that ingests a Stream video, reads the embedded closed caption subtitles and renders them as text on the video, and sends the output via RTMP, for example to a Stream Live Input for broadcasting or recording of the modified version.

Sending video to Streamline

It’s often useful to be able to quickly preview a processing pipeline by sending video data directly to Streamline, for example from a webcam. An agent or embedded device application may also want to use this capability, for example to send footage from factory cameras for AI analysis, or to combine multiple camera feeds into a composite view.

The example below creates a pipeline that expects input from the Worker application and produces a preview video output available over a WebSocket (we’ll talk more about the WebSocket preview video below). It applies two filters and an “annotation,” which is an overlay specified as a PNG image that can be updated while the processing is running, for example to implement an animated graphic.

The code snippet above just starts the pipeline. The controlling Worker is not sending any media to Streamline yet. We’ll discuss the openViewer() function below.

The Worker application sends video data to Streamline using the session.ingest() call. The example below shows how a web browser application might receive chunks from the webcam and forward them to Streamline.

Animated overlay

The annotation overlay can be updated using the session.annotation() call. The example below shows how the Worker application could snapshot a canvas and send it to Streamline. This could be done on an animation loop, although the update rate may be limited in practice by the size of the PNG overlay images, the available bandwidth, and processing power.

Receiving preview video from Streamline

Streamline can also produce preview video output, by specifying output: { mode: 'websocket' }.

Streamline uses WebSockets for low-latency preview video delivery back to the controlling application: the container publishes fMP4 fragments to the Durable Object, which forwards them to an output relay available over a WebSocket on the URL /relay/view, relative to the application origin. The application must connect a WebSocket to this URL, and will then receive video data pushed to it as it becomes available from Streamline. The code snippet below shows how a web browser application might display the preview video feed.

A production MediaSource player must queue fragments while SourceBuffer.updating is true. In local development, the browser or other controlling application simply opens a WebSocket connection directly on the local container.

Currently supported operations

In the configuration object passed to session.start() in the examples above, pipeline is an array of operations from the set supported by the underlying media engine. The operation order is currently fixed by the engine; the order specified in the array is not significant. The list of currently supported operations and the order in which they are applied is below.

Operation name

Function

filter

Applies filtering operations, e.g. blur, saturation.

overlay

Overlays an image referenced by URL or a binary PNG specified separately in a call to annotation().

subtitle

Burns in subtitles.

encode

Specifies output encoding parameters.

Security

This is Cloudflare, so it is important that security is part of the design rather than an addition at the end. We need to ensure that only authorized users can create a new session or take control of an existing one, and that sessions are isolated from each other. We must treat Stream RTMPS input/output keys as secrets that shouldn’t be leaked to the controlling application. We must ensure that resource use is bounded.

The owner deployment is kept private using Workers’ Access integration. The configured owner identity and other allowed users can edit the same shared profiles and start a session while the singleton is idle. The Worker verifies the Access session before accepting control requests and binds the active session to the verified principal. Only one session can run at a time, and a different principal cannot stop or replace the active session.

Stream Live Input keys are stored in Worker secrets or as write-only shared overrides in Durable Object storage. They are never returned by the settings API or placed in browser storage. The controlling application specifies RTMPS input and output by referring to a named profile. The Worker resolves the profile before contacting the container.

The preview video stream has two credentials with separate purposes. A Cloudflare Access service token authenticates the container workload to the publisher endpoint. A random per-session capability authorizes publishing only for the currently active relay. The service token is injected by the container's outbound Worker and never enters container memory. The initial deployment uses a temporary path-specific Access Bypass while the per-session capability remains enforced; after deployment and a successful smoke test, the rollout replaces Bypass with Service Auth.

The owner deployment is intentionally private and singleton-routed. It is not the security model for a public multi-user service.

Playground and open source

We want you to try out Streamline and start building! So together with this post, we are releasing the system as open source and deploying a public playground.

The Streamline container can be run locally or deployed on your account. It exports the Worker API for your control application to use.

There is also an example Worker application with an Astro web frontend that demonstrates Streamline functionality with a few common use cases, including overlays, subtitle decoding, filters and picture-in-picture. There is probe functionality that provides performance metrics and system tracing, and can be useful for debugging the system when developing new features. The example application can be run on a local Astro server, or is set up to be deployed behind Cloudflare Access, so you can control who has access to your Streamline instance.

Both repositories are available as open source on Cloudflare’s GitHub:

We have published a public playground deployment of the example application. This is also something a user can deploy if desired. It uses its own Access configuration, one container identity per verified user, one active session per user, global admission control, concurrency, media and session limits, and no ability for one user to replace another user's session.

You can try the public playground at:

Where we go from here

Streamline demonstrates one way to combine existing managed services, like Stream, with lower-level primitives to build highly customizable media pipelines. In this iteration, Streamline uses Container CPU for media processing, which introduces a bottleneck at higher qualities or frame-rates.

Moving forward, we’re excited to see how we and our developer community can extend this architecture to  build new support for computer vision pipelines, hardware-accelerated media processing, realtime experiences with next generation protocols like WebRTC and MoQ, and ultimately video encoding and decoding primitives natively in Workers.

Today, we invite you to check out our hosted demo of Streamline to see how powerful these tools can be. From there, check out the codebases we’ve open sourced to see how easy it is to deploy Streamline into your own account and use it to create your own experiences.

Introducing Cloudflare Traces: follow requests through our entire platform

Post Syndicated from Mar Witek original https://blog.cloudflare.com/cloudflare-tracing/

Today, we’re introducing Cloudflare Traces in open beta, extending automatic tracing beyond Workers to the rest of the request path. In one trace, you can see supported security rules, transformations, cache decisions, routing, Worker execution, and origin handling, then continue that trace through services running on Cloudflare, at your origin, or elsewhere in your stack. This is a long-term investment in OpenTelemetry and in making Cloudflare the most observable part of your stack.

You can now:

You can enable tracing in the Cloudflare dashboard on any domain or let your agent set up for you:

Giving you the visibility we use to debug Cloudflare

When our own teams investigate, we use our own internal traces, which often include thousands of spans for a single trace, generated by dozens of services and features. This lets us dig deep into every detail of a given request. We don’t think that visibility should stop at our internal systems.

Workers Tracing was our first step toward exposing what happens on our platform. Last year, we launched automatic instrumentation for Worker invocations, including outbound fetches and calls to KV, R2, D1, Durable Objects, and other Workers. It shows the work performed inside the Workers runtime without requiring tracing code for every operation.

The goal of Cloudflare Traces is to bring the same level of visibility to everyone using Cloudflare, whether you’re building on Cloudflare or just have Cloudflare in front of an origin. You get to see how your traffic moved through our platform, and connect the dots between how you’ve configured Cloudflare, and how this influences request processing time, routing decisions, and more. 

Follow one request end to end

A request’s path through Cloudflare can be complicated! It might pass through security rules, transformations, routing, caching, or proxied to another service entirely. Cloudflare Traces records each supported step as a span, including its timing, outcome, and relevant attributes. Instead of reconstructing the request from separate logs and configuration, you can see the request’s path through our system in one place.

You can answer questions like:

Why was the request blocked or challenged, and which security rule took action?

See when custom or managed rules evaluated the request, how long evaluation took, and the resulting action. Identify the rule responsible for a block or challenge through its span events.

Was the URL rewritten by a Transform Rule before it reached the application?

You can open the http_request_transform span to see each change, the request component it affected, and the rule responsible. You can also see where the transformation occurred relative to routing and origin handling.

Which Page Rules, Snippets, or Workers handled or changed the request?

The workers_routing span shows whether a route matched, which routing type was used, and the matching route pattern.

Was the response served from cache, and where was time spent between Cloudflare, the origin connection, and the application?

You can expand nested cache, upstream, and origin spans to see where the request spent its time. Here, you can see there was a cache miss that went to origin and spent 527ms of the 539ms getting a response.

Configure your tracing

There is no special instrumentation, config, or plugins required. Once tracing is enabled for a domain, Cloudflare generates these spans automatically. This lets you extend the trace through third-party services and back again by adhering to open standards. From there, you can control which requests are traced using a baseline sampling rate and Trace Rules.

Set a baseline sampling rate

You can enable tracing on any domain and set a baseline sampling rate to balance visibility, data volume, and cost. You might trace 1% of requests during normal operation, giving you a continuous view of request behavior without collecting a trace for every request.

Configure Trace Rules

Trace Rules let you keep a low baseline sampling rate while capturing complete traces for a specific investigation. If one customer reports a problem, you can trace 100% of traffic for their hostname, source IP, or identifying request header while leaving everyone else at 1%. Or during an investigation, you could trace 100% of requests carrying a temporary debug header, while leaving all other traffic at the baseline. This lets you reproduce an issue without increasing tracing across the entire domain.

Trace Rules use the same Cloudflare Rules language, so you can target paths, methods, headers, IP addresses, geographies, or combinations of those properties.

Accept and propagate trace context

One of the most common requests we hear is for true distributed tracing: a single trace that follows a request into Cloudflare, through our platform, and onward through the rest of your stack.

Cloudflare Traces can accept a W3C traceparent header from an incoming request, allowing Cloudflare spans to join a trace that began before the request reached our platform. An incoming propagation policy controls whether Cloudflare accepts that context.

Cloudflare can also forward a new traceparent header to your origin. Any other instrumented services can extract that context and continue the trace through APIs, databases, and services running on Cloudflare or elsewhere. To view everything as one connected trace, you can send both Cloudflare and application spans to the same OpenTelemetry-compatible backend.

Export traces to your observability platform

You can export Cloudflare spans over OTLP to a compatible observability platform, where they appear alongside telemetry from the rest of your stack. Configure an account-level destination, then choose which domains send traces to it. This is part of our commitment to OpenTelemetry: Cloudflare represents request activity as OpenTelemetry spans and delivers them using OTLP, keeping the data portable across observability tools.

Let your agent investigate Cloudflare Traces

When you ask a coding agent to debug a production issue, it might inspect your code and run tests, but it may not be able to see what happened to the request in production. With the Cloudflare Observability MCP server, your agent can leverage our SQL API to query your traces (and all of your observability data!), giving it access to your investigation production telemetry.

Let your agent find the right requests, comparing failed traces with successful ones, and identifying where their spans diverge. Since the agent can also inspect your repository, it can connect those findings to the relevant code, narrow down what needs to change, and help put up a fix for you to review.

Pricing

Cloudflare Traces will be a part of the unified Cloudflare Observability pricing model. Instead of charging by the number of spans/events, pricing is based on how much observability data you ingest and how long you retain it. New pricing will take effect across Cloudflare Tracing (and Workers Tracing!) starting December 1, 2026.

Plan

Included Usage

Retention

Additional Usage

Free

0.5 GB of ingestion per day

7 Days

Not available

Paid and Enterprise

50 GB of ingestion
10 GB-month of storage per billing cycle

Up to 1 year 
(coming soon)

$0.25 per GB ingested
$0.10 per GB-month stored

What's next

Following the open beta, we plan to launch:

  • Broader automatic instrumentation: Add more spans across both the HTTP request path (e.g. DDoS rules, Access) and the Workers execution path (e.g. Workflows, Queues, Pipelines).
  • Authenticated context propagation: Let trusted callers continue an existing trace without accepting context from every incoming request.
  • Ad hoc tracing: Capture a specific request on demand without changing the baseline sampling rate.
  • OpenTelemetry API support in Workers: Continue building out our OpenTelemetry APIs to enable adding attributes to existing spans or getting trace context.
  • Longer retention: Keep trace data available for up to 365 days for longer-running investigations.

Get started

Follow the Cloudflare Traces documentation to trace your first request and tune sampling with Trace Rules. Cloudflare Traces is available in open beta from the dashboard, through the API, or with Terraform, with support for exporting to an OTLP destination.

8 major updates to Cloudflare Observability

Post Syndicated from Nevi Shah original https://blog.cloudflare.com/one-observability-platform/

Today, we’re launching eight major updates that bring your logs, traces, analytics, alerts, dashboards, and exporting into one observability platform, with simpler and more predictable pricing.

Here's what's launching:

One observability platform for all of Cloudflare

Understanding an issue often requires data from more than one Cloudflare product. A spike in 5xx responses could come from a Worker, from your origin, or from Cloudflare failing to connect to your origin globally or regionally. But investigating it today requires knowing which product owns each signal and how to query it.

Observability should be a platform-wide capability: it should reflect how applications actually behave and give you the complete context needed to resolve an issue. Over the coming months, you’ll see more Cloudflare products, datasets, and workflows become part of this shared observability platform, with more consistent pricing, product experiences, and features. These eight updates are the first step into a more unified Observability problem.

1. Investigate all your logs in one place

The new Logs home combines Workers Observability (for debugging Workers applications and its connected resources) with Log Explorer (for searching across security logs). You can now choose from log datasets like HTTP events, firewall events, Workers, Containers, R2, and AI Gateway, and use the same investigative tools and capabilities for each.

Start with an increase in request latency, group it by hostname or data center, narrow the results to affected paths, and inspect individual requests by Ray ID. If the investigation leads to another Cloudflare product, switch datasets without leaving Logs. Support for querying across multiple datasets is coming soon, making it possible to connect related events across products in a single query.

You can query your logs with raw SQL or with built-in filters to narrow down on specific events. Create visualizations with natural language, and easily investigate and understand detected anomalies.

2. Trace requests through our entire platform — now in open beta

We’re launching Cloudflare Traces in open beta, giving you a request-level view of supported security rules, transformations, cache decisions, routing, Workers, and origin handling. You get to see how your traffic moved through our platform, and connect the dots between how you’ve configured Cloudflare, and how this influences request processing time, routing decisions, and more.

Set a baseline sampling rate for continuous visibility, then use Trace Rules to capture specific traffic at a higher rate during an investigation. Target hostnames, paths, IP addresses, or headers, search by Ray ID, and inspect the resulting spans directly in the Cloudflare dashboard.

You can export traces over OpenTelemetry, while W3C trace context propagation lets you accept incoming trace context and pass along context to your origin. Check out the full blog post to learn more about Cloudflare Tracing or give  this command to your agent to get started:

3. Have your agent query observability data with one unified SQL API

Agents also need a consistent way to sift through your observability data, investigate issues, correlate signals, and verify fixes. We’re launching a unified SQL API, now in beta, for querying telemetry across Cloudflare. Instead of integrating separately with Workers logs, Containers security events, HTTP request logs, and analytics data, people and agents can query them using one SQL dialect, authentication model, and API.

Your agent can use the new Cloudflare CLI, cf, to find and run queries from the command line or connect through Cloudflare’s Observability MCP server to investigate logs, traces and analytics. Dataset schemas, fields, and example queries are available to help both people and agents build queries.

Additionally, we’re also bringing the SQL interface directly into Workers with a native binding. Your Worker can now do things like query Analytics Engine data to meter customer usage and power billing workflows, build customer-facing analytics dashboards, generate health reports, or automate incident investigation without configuring a separate API client.

4. New pricing for all ingested and stored logs and traces

For all logs and traces ingested and stored on Cloudflare, we are moving to one unified Observability subscription and pricing. Beginning December 1, 2026, this pricing model will apply across all plans (effective upon renewal for all Enterprise customers) and cover existing Developer Platform logs, including Workers, Containers, AI Gateway, as well as all tracing data.

Because logs and traces can vary dramatically in size, the new model is based on the volume you ingest and store rather than an event-based count. This pricing adjustment will be. Check out our documentation for more details on pricing.

Plan

Included Usage

Retention

Additional usage

Free

0.5 GB of ingestion per day

7 days

Not available

Paid and Enterprise

50 GB of ingestion
10 GB-month of storage per billing cycle

Up to 1 year
(coming soon)

$0.25 per GB ingested
$0.10 per GB-month stored

5. Configure custom alerts on your observability data – now in beta

Notifications (now called “Alerts”) just got a major upgrade. You can now define custom alerts directly on anything supported by our new unified SQL API, including HTTP request logs, Workers events, Workers Analytics Engine datasets, analytics datasets, traces, and security events.

Choose a dataset in the dashboard or define the condition using custom SQL. Then select a threshold, anomaly, or SLO, set the evaluation window, and choose where the alert should go. You might alert when origin 5xx responses exceed a threshold for five minutes, a Container repeatedly fails, Worker errors increase after a deployment, or trace latency crosses an expected limit. 

You can send alerts right to tools your teams are already using, including incident management tools, chat platforms, and webhooks. Webhooks are now available on all plans, allowing you to route alerts to custom services or even your agent to begin investigating immediately. To get started check out our documentation or give this command to your agent:

6. See your domain analytics in one place — now with 30 days retention

Understanding what is happening on your domain has often meant piecing together metrics from different Cloudflare products. We’re bringing traffic, performance, security, cache, origin, and DNS data together so you can see how they relate. If latency increases, you can quickly see whether it is tied to a specific Cloudflare data center, hostname, or origin.

In addition, you now get 30 days of domain analytics on every plan. A full month of history gives you time to investigate issues after they happen, compare today with the same day in previous weeks, and tell the difference between a one-time spike and a longer trend.

7. Build custom dashboards

Prebuilt dashboards cover common use cases, but applications often use several parts of Cloudflare. With Custom Dashboards, you can bring together analytics from across Cloudflare, logs and traces from the Workers platform, and security events in one view. Track request volume, errors, latency, storage, and blocked traffic, then share the dashboard with your team. Instead of rebuilding queries during every investigation, you have one place to monitor the signals that matter to your application.

8. Logpush is now available on all self-serve plans

Logpush, previously available only to Enterprise, is now available on all self-serve plans, letting you export all Cloudflare logs to the tools and destinations you already use. Need to apply filters, perform redaction, enrich events or reshape output before delivery? Transformers is now generally available, letting you apply any SQL transformation without operating a separate ETL pipeline.

We’re introducing usage-based pricing for Logpush and Transformers. Each includes a free monthly allowance, with simple pricing for additional usage:

Export usage

Included each month

Additional usage

Exports to Cloudflare destinations

25 GB

$0.03 per GB

Exports to external destinations

25 GB

$0.10 per GB

Logpush Transformers

1 GB

$0.04 per GB

Visit the documentation to get started with Logpush and explore complete pricing details.

What's coming up:

  • Longer retention for your observability data: You’ll be able to retain logging and tracing data for up to one year, making it easier to investigate recurring issues, compare historical behavior, and analyze long-term trends.
  • OpenTelemetry API support in Workers: We’ll continue building out our OpenTelemetry APIs to enable adding attributes to existing spans or getting trace context.
  • Easier metrics export with OpenTelemetry: You’ll be able to send Cloudflare metrics to OpenTelemetry-compatible destinations and analyze them alongside telemetry from the rest of your stack.
  • New pricing takes effect December 1, 2026: If you ingest or store observability data on Cloudflare, the unified pricing plan will apply to your usage. We’ll notify you before the change takes effect.

Ready to start investigating?

We hear you when you say Cloudflare can feel like a black box. These updates are just the beginning of exposing what’s happening, making the underlying data accessible, and giving you the context that you need to act. That transparency matters even more as agents move from writing software to operating it. An agent can only close the loop between a change and its outcome if it can query what happened, identify the failure, and verify the fix.

By building around OpenTelemetry, W3C Trace Context, and SQL, we are committed to giving you and your agents standard, portable interfaces to that context. Check out our new Observability documentation home to learn more.

We want you to build the next Git platform on Cloudflare

Post Syndicated from Dina Kozlov original https://blog.cloudflare.com/next-git-platform-on-cloudflare/

GitHub was built for a world where humans write code, organize it into repositories, and collaborate through branches, commits, issues, and pull requests.

But the next generation of software is going to be built differently because it is going to be built by a different kind of developer: agents.

Agents are already writing more code than ever before — they’re fixing bugs, building features, writing tests, reviewing changes, updating dependencies, and doing the routine maintenance required to keep an application running.

So in this new world where you have hundreds, or even thousands, of agents working on the same codebase at the same time, what does the foundation look like?

How do agents know what other agents are working on? What happens when they make conflicting changes? How do you review everything they produce? How do you keep track of not just what changed, but why a change was made?

And so the burning question is: What does the next GitHub look like?

We want you to help us answer it, by building it out.

Earlier this year, we launched Artifacts, a versioned filesystem that speaks Git and can scale to millions of repositories. From the start, we designed Artifacts as a set of programmable primitives that developers could use to build their own products, workflows, and abstractions.

Artifacts provides the foundation: repositories that can be created and forked programmatically, versioned storage for code and agent context, and the Git operations agents already know how to use.

With that foundation in place, you can focus on the layer above it: how agents coordinate their work, how changes are reviewed and merged, and what the developer experience should look like when hundreds or thousands of agents are working on the same codebase.

That is the layer we want you to build.

Now that Artifacts is in open beta, we’re holding a competition to see who can build the next Git platform on Cloudflare using Workers and Artifacts.

Artifacts is in open beta. Here’s why you should build on it

When we launched Artifacts, our goal was to make it possible to create a repository for every agent, session, task, or user — and to do that at the scale agents require.

Since then, we’ve seen developers use Artifacts in a range of ways: Vibe-coding platforms are using it to store the projects their users create. Developers are using it to persist the code and context from agent sessions. Others are creating isolated repositories, so multiple agents can safely work from the same starting point and compare or merge the results later.

Here are some new capabilities we’ve added since the initial launch.

Deploy Artifacts repos to Workers

You can now connect an Artifacts repository to a Worker through Workers Builds. When you or an agent pushes code to the Artifacts repository, Cloudflare will build the project and, for the production branch, deploy the updated Worker. Pushes to other branches automatically create or update Workers Previews, giving you an isolated, shareable version of your Worker where you can test changes before they go live.

You can connect an existing Worker to an Artifacts repository or start a new project and automatically store it in Artifacts.

Manage Artifacts directly from Workers

You can interact with Artifacts repositories directly from a Worker using an Artifacts binding to create or fork repos, inspect files and commits, and issue repo-scoped Git tokens. This makes your Git workflow programmable. When a new task arrives, a Worker can fork the project for an agent, read the files it needs for context, and give it a repository to work in. When the agent pushes a change, your automation can inspect the result and start a review. You define those steps in code to fit how your agents work.

For example, here’s how to fork a project for a new agent task and read its AGENTS.md for instructions:

React to every change with event subscriptions

Artifacts publishes events whenever a repository is created, imported, forked, deleted, pushed to, cloned, or fetched. You can subscribe to these events to decide what happens next: run CI, kick off a code review agent, or deploy a change.

For example, you can subscribe to Artifacts push events and have a Worker start a code review workflow for each push. The Worker passes the repository, branch, and new commit to the Workflow, giving a review agent the context it needs to inspect the change:

Data jurisdiction for Artifacts repos

You can now choose where Artifacts stores and processes your repository data. Set a U.S. or EU jurisdiction when you create a namespace, and every repository created in that namespace will automatically follow the same restriction.

View Artifacts metrics

You can now see metrics for your Artifacts repositories in the Cloudflare dashboard. For each repository, you can now see total operations, pulls, pushes, errors, and error rate, helping you understand how the repository is being used and spot failures. You can also query Artifacts metrics directly to build your own dashboards or monitoring.

Pricing

Artifacts pricing is based on repository operations and the amount of data stored. We will begin billing for Artifacts usage on October 15, 2026.

Competition: Build the next Git platform on Cloudflare

We want you to build your vision for the Git platform of the agentic era using Cloudflare Workers and Artifacts.

You could rethink repositories, branches, pull requests, worktrees, code review, and merge conflicts — or build new ways to preserve agent context, compare multiple changes at the same time, and decide which one should ship.

We aren’t looking for GitHub as it exists today with agents added on top. At a minimum, we want to see multiple agents working on changes concurrently. Beyond that, we want you to get creative — what you think comes next.

How to enter

Submit:

  • A 5-10 minute video demonstrating what you built, what it enables agents and developers to do, and how it works
  • A link to the source code, which must be provided under a permissive open source license (MIT, Apache, BSD)
  • Instructions for running or trying the project

Deadline

Submissions are open until October 14, 2026.

Why should you participate?

We’ll select the top three projects and fly up to two members from each team to San Francisco to attend Cloudflare Connect and show what they built.

The first-place team will also receive $25,000 in Cloudflare credits, along with invitations to the VIP speaker dinner on Monday night at Connect.

Get started

Artifacts is available in open beta to customers on the Workers Paid plan.

Get started with your coding agent: copy the prompt below to set up your first Artifacts repository and start pushing code to it.

You can view or create the Artifacts repositories in the dashboard or if you’re looking to learn more, check out the documentation.

Announcing Cloudflare K2: serverless event streams

Post Syndicated from Micah Wylde original https://blog.cloudflare.com/cloudflare-k2-streams/

With traditional Remote Procedure Call (RPC) architectures, there exists a core challenge: producers and consumers must align in scale and in time. If your producers send too much data for your consumers to handle or if your consumers or downstream services become unavailable, events are dropped. This problem is compounded with multiple consumers that need to independently process the data. For example, an ecommerce backend may emit events when transactions are completed, which need to be read by an analytics system and a fraud detection service.

We can solve this by decoupling our producers and consumers — inserting a service in the middle that absorbs writes while allowing independent readers to consume at their own pace.

Today we are launching Cloudflare K2 in public beta to solve this problem. K2 is a durable event streaming primitive on the Developer Platform. You send events to a K2 stream, which stores them as an ordered log. Consumers can read them in a variety of ways, for example by splitting up reads across a set of consumers, or delivering all messages to all consumers. It's fully serverless, scales to vast quantities of data, and supports long-term retention, so even long periods of consumer downtime do not lose data.

Under the hood, K2 implements a partitioned, durable log on top of R2 object storage, which allows it to scale to huge volumes of storage.

If you’re ready to get started, you can create your first stream in seconds by following the guide here.

Streams on the edge

We first built K2 because we needed a durable buffer on the edge, initially to serve as the ingestion layer for Basin Pipelines. Pipelines is powered by a stream processing engine that operates on a pull-based model, which means some other system has to store events before they are read, transformed, and written to R2. And because we commit to never dropping events once they’re accepted into the Pipelines Stream, that storage has to be durable — meaning it can’t lose data — over potentially long periods of time.

This is where most companies would deploy Apache Kafka. However, Pipelines runs on the Cloudflare edge, which spans a huge number of servers across over 335 cities. Our unique architecture means we often cannot run traditional distributed systems software like Kafka, and need to rethink how these systems are built and operated.

For stateful services, in particular, Cloudflare’s global infrastructure presents some challenges: we get relatively small slices of machines, those machines are relatively ephemeral, and networking is often over the public Internet. But our infrastructure also has a few superpowers: it’s close to users wherever they are in the world and has an incredible capacity to scale horizontally.

In designing the durable buffering system that became K2, we decided to rely on the powerful state primitive we already have: R2. Object storage systems like R2 combine extremely durable storage (11 9s!) with strongly consistent APIs. Offloading replication and consensus to the storage layer allows us to make the application layer (K2 in this case) radically simpler, cheaper, and higher performance. A secondary benefit is that it separates compute and storage, meaning each can be scaled independently. This allows us to store vast quantities of historical data at low cost.

How do we build a log on top of object storage? An immediate issue is that R2 — like other object stores — does not support appends, the standard operation on a log. Instead, we must write complete files, or segments, that are large enough to overcome the cost of writing and reading each one. We do this by first accumulating writes in-memory on an edge service. After waiting a short period for data to arrive, we write all events as a segment file. We achieve ordering and strictly incrementing offsets using R2’s atomic operations without needing a separate coordination service.

While building on R2 has many advantages, there is one downside: higher produce latencies. Writing to object storage is slower than a local disk, and we have to wait for the local batch to accumulate before starting the write. In our initial release of K2, this adds up to about 1 second of produce latency at the 99th percentile of response times.

We will be sharing more details on the design of K2 in an upcoming technical deep dive.

Streams, Queues, or Pipelines?

Cloudflare has several existing asynchronous delivery primitives, including Queues and Basin Pipelines. When should you reach for K2 instead of these existing products?

There are some superficial similarities between Queues and K2 Streams: both receive events, durably store them, and deliver them to consumers. Queues are designed around tracking individual items of expensive or time consuming work that need to be asynchronously completed. For example, an image processing application may enqueue a user request to be handled by the actual image processing service. They support complex logic on the grain of a particular work item, like retries, delays, and dead-letter queues for failed attempts.

K2, by contrast, is designed for high-scale data movement, long-term retention, and fan-out consumption. Messages are produced and consumed as batches — enabling efficient processing at the expense of message-level retries. This batching also drives higher producer latency than for queues.

Basin Pipelines is a serverless ingestion service. You can send your Pipeline JSON events, which can be transformed and written to R2 or a Basin Catalog. We recommend Pipelines when the end result is writing your events to object storage or Iceberg tables, and K2 when doing custom processing or writing to other destinations.

Getting started

Using K2 involves first creating a stream. You can have many streams across your account for different use cases or types of events. Streams can be created via cf, Wrangler, the dashboard, or API.

Let's take the example of collecting and processing product analytics. First, we'll create a stream with cf:

Once we have a stream, we can start producing to it, via an HTTP API or Worker binding. For example, from a Worker:

K2 represents data as bytes, so you can use whatever format or encoding makes sense for your application.

Now that we have events in a stream, we can create a subscription. Subscriptions divide up work between consumers, enabling read parallelism — scaling out to multiple readers to handle more load than a single server can manage.

We can create a subscription via the HTTP API.

With the subscription created, we can then poll it from each of our consumers:

When a client calls consume, they receive a lease for that particular batch of events for 5 minutes. The client can do one of three things:

  • ack the batch, which marks it as processed and ensures it will not be redelivered
  • nack (negative ack) it, meaning we’ve failed to process it and would like it to be redelivered
  • extend its lease, in case it needs more time to complete processing

This is one way to consume from K2: splitting work amongst multiple consumers such that each consumer gets a portion of the data. Another way to read is with a separate subscription for each consumer — the pub/sub pattern — in which case each consumer sees all of the messages. Or you can mix-and-match between these two approaches, having multiple independent consumer pools.

See the K2 docs for full details on the APIs.

Pricing and Availability

K2 is available today in public beta for accounts with Workers Paid subscriptions, within these limits:

  • Maximum of 10GB of storage used
  • 30 MB/s produce per stream

If you need higher limits, please reach out to the team on Discord or fill out the limit increase form.

Usage of K2 will not be billed during the beta period. Once we begin billing, we anticipate this pricing:

Pricing

Data Produced

$0.04 / GB

Data Consumed

$0.04 / GB

Data Retained

$0.02 / GB / month

What’s next

We have an exciting roadmap for K2 over the coming months, including:

  • Higher write parallelism, up to multi-GB/s streams
  • Message keys and key-based ordering guarantees
  • Push-based worker consumers
  • Express tier with lower produce and end-to-end latencies
  • Drop-in support for Apache Kafka clients

We’re excited to see what you build on K2! Share your feedback on the Cloudflare Discord.

Cloudflare OS: your company’s agent workspace, managed for you

Post Syndicated from Phillip Jones original https://blog.cloudflare.com/managed-cloudflare-os/

Cloudflare OS gives everyone in your organization an agent workspace that knows how your company works and connects to its data and systems. Today, we're opening the waitlist for fully managed Cloudflare OS deployments.

If I asked you to prepare for an important customer meeting later today, what would you do? You might learn how your company typically runs customer meetings, review the account in your CRM, check recent support tickets and product usage, then turn it into a short presentation to review with the group. Now imagine doing that another 100 times this month.

Every team has work like this. With Cloudflare OS, you can ask your agent to handle the work for you, build a tool for your team, or move between the two as the work evolves.

Last month, we announced Cloudflare OS and shared the open source repository. Since then, thousands of organizations have started using it to work with company data, produce docs and slides, build tools for their teams, and automate work with agents.

With a few clicks in the Cloudflare dashboard, you’ll be able to launch your organization’s own agent workspace. Just tell us what custom domain you want to use, what Cloudflare Access policies apply, and which AI Gateway to connect. We’ll handle the rest.

Cloudflare OS, managed for you

Every company has its own terminology, procedures, systems, and requirements. We made Cloudflare OS open source so you can customize it around how your company works.

You can already deploy Cloudflare OS into your own Cloudflare account from the open-source repository. That gives you full control, but it also means someone has to configure the deployment, operate it, and keep it up to date.

With the fully managed option, you decide who can access Cloudflare OS, which organizational skills and context are available, and which systems it can reach. You can leave the rest to us.

If you want Cloudflare OS fully managed for your organization, join the waitlist and we’ll reach out.

What’s new in Cloudflare OS

We’ve also spent the last month expanding what people and agents can do in Cloudflare OS. Here are a few highlights.

Mount Git repos and work with code

When we launched Cloudflare OS, we focused first on work outside software development: creating documents and slides, automating tasks, and building collaborative tools. Agents could write code for an app, but they could not work with code in an existing Git repository.

You can now connect an existing GitHub repository to Cloudflare OS. Ask your agent to explore the codebase, fix a bug, add a feature, or open a pull request. It can search and edit files, review its changes, create commits, and push them to GitHub.

Work across Google Workspace

For many organizations, work starts and ends in Google Workspace. Decisions live in email threads, context lives in Google Drive, analysis happens in Sheets, and teams coordinate through Calendar. Agents need to do work across those systems too.

We’ve made significant improvements to the Google Workspace Gatekeeper (a service-specific Worker that sits between Cloudflare OS and an external service). Cloudflare OS can now read and research Gmail threads, create drafts, and send emails. You can also connect your entire Google Drive, a specific folder, or an individual doc or sheet.

Export work in the formats your team uses

Work often needs to move into the formats your team already uses. Finance may need an Excel spreadsheet, and a report may need to become a PDF before sending to a customer.

The built-in document, presentation, and spreadsheet experiences can now export work to familiar formats. Depending on what you create, you can export to Microsoft Excel (.xlsx), CSV, PDF, Markdown, or HTML. Microsoft Word (.docx) and PowerPoint (.pptx) export is coming soon.

Tools you build can also define their own export formats. Tell the agent what you need, like “let me download this schedule as a calendar file (.ics)”, and it’ll add the option to the tool’s export menu.

Sign up for the waitlist

Cloudflare OS is open source and available today. You can check out the source code or deploy it into your own Cloudflare account.

If you want Cloudflare OS fully managed for your organization, join the waitlist and we’ll reach out with more information.

Simplifying domains for people and agents

Post Syndicated from Ankit Shah original https://blog.cloudflare.com/simplifying-domains/

You just thought of your next great idea, and buying the right domain feels like the easiest way to make that first bit of progress. Naturally, you open a new tab in your browser, only to find yourself face-to-face with an experience that feels like a budget airline peppering you with add-ons at checkout: Want security? How about a website? Do you want email? You’re just a few minutes into building your next idea, and it doesn’t feel fun anymore.

Launched a decade ago, Cloudflare Registrar has always taken a simpler approach. Domains at cost, transparent pricing, and no unnecessary upsells. But simplicity shouldn’t begin at checkout. It should begin the moment you start looking for the right domain.

Today, we’re bringing that same simplicity to the entire experience of finding and buying a domain. Our new domain search shows every extension we support, responds as quickly as you type, and makes hundreds of possibilities easier to explore through sorting, filtering, and transparent pricing.

And you know what is particularly good at ignoring distractions and staying focused on the destination? An AI agent. We designed Cloudflare Registrar to work naturally with agents through the Registrar API, MCP, and our newly launched cf CLI. You can ask your favorite agent to find the right domain, buy it, or transfer one you already own.

The agentic registrar, expanded

In April, we launched the Registrar API beta, allowing developers and agents to search for, check, and register domains programmatically. We have expanded the API since then. The new sandbox lets you test registrar workflows without purchasing a domain or triggering a real transaction. Our extensions endpoint returns relevant information for each of the 420+ extensions we support, helping you account for the different requirements across registries. We also added transfers, so you can bring domains from another registrar into Cloudflare programmatically.

The Registrar API is available through Cloudflare MCP, giving agents access without requiring a separate integration. Earlier this week, we also announced the launch of cf CLI, bringing the same capabilities directly into your terminal. You can prompt your favorite agent to search for, register, or transfer a domain. These tasks already lend themselves naturally to a conversation:

  • “Is example.com available?”

  • “Buy example.com.”

  • “Transfer example.com from my current registrar.”

The way people interact with the Internet is changing. Cloudflare Registrar should feel natural whether you use it through an agent or in your browser. For many people, the browser is still where the search begins, and that experience was long overdue for an overhaul.

Search simplified

Previously, our search page showed around 20 available results from a subset of extensions, sometimes modifying your search term to suggest related options. You couldn’t see that exact name across every extension we support.

We decided to take a simpler approach: show you the exact term you searched across every supported extension. Results appear as you type and continue to load as you scroll, letting you explore hundreds of options without starting another search. We also include domains that have already been registered, giving you a more complete picture. If you only want domains available to buy, you can filter everything else out.

More results might not sound simpler, but these are the results you asked for. Sorting and filtering help you narrow them down. Whether you are logged in or logged out, on your phone or at your desk, the experience feels the same.

Making a complex question feel simple

Cloudflare Registrar supports 420+ extensions. A single search can thus create more than 420 separate availability questions. Each extension is operated by a registry that maintains its official registration records and provides the authoritative answer about a domain’s availability and price.

Asking every registry every question at once would be slow and wasteful. Registries respond at different speeds and impose request limits, and much of the work would be for results the person might never view. To solve this, our search gathers evidence from multiple sources: (1) prepared availability datasets (e.g. zone files) and cached answers; (2) DNS answers; (3) live registry lookups.

A hit against an availability dataset or DNS can tell us that a domain is already in use, but a miss cannot necessarily prove availability as a domain may be registered without being configured in DNS, or it may be on a blocked list. A recent registry-derived answer is stronger but becomes stale over time, while a live registry check provides the freshest authoritative answer but takes longer and draws on limited upstream capacity. Our new search progressively probes these sources while balancing speed, freshness, and certainty for each result. As better or more accurate information arrives, we update only the affected result dynamically.

How we built search for speed and scale

We built the new search on the same Cloudflare developer platform available to our customers. Workers run the public search entry point and the services that gather availability evidence. Durable Objects give each active search one coordinator, while Workers KV stores prepared data that can be reused across searches. Together, these primitives let the service scale across users and 420+ extensions while keeping operating costs low.

We prepare useful evidence before a search begins: a purpose-built pipeline converts registry zone files and other bulk sources into compact availability datasets in Workers KV. Large datasets are split into smaller pieces, so a lookup retrieves only the data required for that domain. These fast checks can answer many questions without making a new live registry request.

A search session is composed of a Durable Object that coordinates that specific search interaction. It establishes the result order from the query, sort, and filters without waiting for network lookups, remembers the best evidence received for each domain, and tracks which results are visible so lookup work follows the person's attention.

WebSockets provide the bidirectional connection, and we designed an application protocol on top of them to connect the browser to the resolution process. An initial snapshot establishes the ordered list. Subsequent delta messages contain only the fields that changed, letting the browser update one domain instead of downloading the full result set again. Before sending a delta, the Durable Object compares the new evidence with the current answer: stronger evidence can replace it but weaker evidence cannot.

A separate Worker gathers additional evidence. It can query DNS through Cloudflare's 1.1.1.1 resolver, make a live Registrar check, or reuse a recently cached answer. Reusing fresh answers avoids repeating upstream requests. Because an available domain can be registered at any moment, an available answer has a shorter useful cache life than evidence that a domain is already taken.

Together, these pieces turn hundreds of independent availability checks and all that coordination into one coherent search experience. Cloudflare’s Developer Platform gives us all the building blocks to hide that complexity and craft a domain search experience that feels simple and is among the fastest in the world.

Transparent pricing and price drops

Making search feel simple is not only about speed. It is also about knowing exactly what a domain will cost. Cloudflare Registrar has offered domains at cost since day one. Great prices are part of making domains simple, but so is knowing what you will pay. Our new search and our new pricing page show both the initial registration price and the renewal price for every domain. When a domain is discounted, we show the original at-cost registration price crossed out alongside the promotional price.

Beginning with Birthday Week (this week!), we’re offering first-year registration discounts on select extensions including .io, .dev, .app, and .tech. You can explore every discounted extension directly from the search page as well as our newly launched pricing page.

Whether you search in your browser or ask an agent, our goal is the same. Remove the friction between having an idea and making it real. Buying a domain for your next idea should be fun and feel like progress.

Find your next domain

Choose how you want to get started:

  • Search in your browser: Explore every available extension and find your next domain.
  • Ask your favorite agent: Install cf CLI, then prompt your agent to search for, register, or transfer a domain.
  • Build with the API: Use the Registrar API to bring domain search and registration into your own application or workflow.

However you choose to do it, finding your next domain should be the fun part.

Acknowledgements: This simplicity was a result of cross-team collaboration. Special thanks to Pedro Menezes, Shobhit Kuruvilla, Lucy Dryaeva, Fred Pinto, and the Registrar Team, the Design Engineering Team, and the Forge team.

Cloudflare Containers, rebuilt to scale agent sandboxes

Post Syndicated from Thomas Gauvin original https://blog.cloudflare.com/faster-agent-sandboxes/

Today, we’re making Cloudflare Containers more programmable and optimized for agent workloads. Agents don't deploy sandboxes ahead of time. They create sandboxes on demand, for each task, expect them to be ready immediately, and be able to pause and resume. So we rearchitected Containers to meet these requirements: your code can now choose each sandbox's image and instance type at runtime, Containers start 6x faster, and filesystem snapshots are available in public beta.

To make this possible, we’ve rethought the Containers infrastructure from the bottom up. A new scheduling policy moves control over each sandbox into application code, while a redesigned runtime provides a faster path to a running Container. In ComputeSDK’s independent benchmark, median startup fell from just over four seconds to 648 milliseconds, and, in our own preliminary tests, burst testing successfully created hundreds of thousands of containers in seconds.

All of this builds on what has always set Containers on Cloudflare apart: every Container gets its own Durable Object, a persistent, programmable controller running right next to it that manages its lifecycle, outbound traffic, and more. We are bringing more capabilities directly to the native ctx.container API, so the Durable Object can control its Container without a wrapper class in between, and we’re carrying this model into Sandbox SDK 1.0.

As we wrote earlier this year, your agent needs a computer. These changes make Containers a better complement to Workers, Dynamic Workers, and Durable Objects when agents need a full Linux workspace.

Rethinking Containers’ runtime for agents

Until now, Cloudflare Containers has been organized around application deployments. You choose an image and compute resources at deploy time, roll that configuration out across the application, and manage it centrally. The application is the unit of configuration and rollout.

An agent workspace is different: it’s created on demand, while the agent is working. The task determines its image, resources, tools, and starting filesystem. It might exist for a few minutes, sleep between requests, or be restored days later. Those decisions need to live with the application code handling the task, and the agent's sandbox needs to start up fast, because every second of startup is time your users spend waiting.

We’ve seen this pattern with Base44 on app-building workspaces and Kilo Code on cloud-agent sessions. It also appears in our integrations with Cursor Cloud Agents, Devin Outposts, the OpenAI Agents API, and Claude Managed Agents.

Each of these workloads needs something different from the workspace. Coding agents need repositories, package managers, compilers, test runners, and development servers. Evals need sandboxes that begin from a known state. Reinforcement learning systems need to create, grade, and reset large numbers of environments. Longer-running tasks need to preserve the files an agent produces, so work can continue later.

These requirements led us to fundamentally rethink how Cloudflare Containers are configured, scheduled, and saved. The result is a new way to provision and schedule Containers: the durable_object scheduling policy. It lets your code choose each sandbox’s image and compute resources at runtime, starts Containers more than 6x faster, and supports filesystem snapshots, so workspaces can be saved and restored.

“At Base44, we help anyone turn an idea into a working app. Cloudflare Containers gives each app an isolated development environment where our AI can execute commands, install dependencies, and bring changes to life in a live preview.”

Dolev Epshtein, Software Engineer, App Infrastructure at Base44

“At Kilo Code, every cloud-agent session needs its own workspace and environment, with the right repository, tools, and user configuration. Cloudflare Containers lets us create those isolated environments on demand, so our agents can start running commands quickly and get to work for our customers.”

Emilie Schario, Co-founder of Kilo Code and VP Engineering, AI Workspaces at Anaconda

Choose the sandbox your agent needs, from code

From the start, every Cloudflare Container instance has been attached to a Durable Object. The Durable Object gives the environment a stable identity and lets application code control when it starts, sleeps, and stops. This model has proven particularly well-suited to agent sandboxes.

Until now, though, the two decisions that matter most for an agent sandbox were locked in at deploy time: which image it runs and how much compute it gets. Each combination of image and instance type was its own Containers application, with its own Durable Object namespace, set up ahead of time with wrangler deploy. 

Say one agent needs a small Node.js sandbox and another needs a large Python sandbox for builds. With our previous approach, that required two applications, two namespaces, and routing logic in your Worker to send each task to the right one. Every new environment meant another deployment. 

The new durable_object scheduling policy removes that. The image and instance type are now arguments your code passes when the sandbox starts. To opt in, set the scheduling policy and declare the images your Durable Object can choose from:

Each image you declare is available as this.ctx.container.images.<name> within the Durable Object. When a task arrives, your code picks the image and instance type for that task:

This code makes two decisions after the task is known. It chooses the toolchain the workspace needs and how much compute to give it. What used to take a separate application and a separate wrangler deploy is now an if statement. One Durable Object class can start Node.js and Python sandboxes of different sizes side by side, and adding a new environment is a code change, not a new deployment. That’s the idea behind this whole update: infrastructure becomes code that runs at request time, right down to the environment itself.

Rollouts are now just code

Choosing the image at start time also fixes one of the most painful parts of running Containers: rollouts.

Before, updating an image meant updating the whole application. You set grace periods, so running instances could drain, defined percentage splits to move traffic gradually, and called our API to push the new configuration. Throughout that process, the platform decided which instances got replaced and when, whether an agent was in the middle of a task.

With the durable_object scheduling policy, there's no rollout configuration at all. A Container can keep running the image it started with until your code stops it. The next time that Durable Object starts a Container, it uses whatever image your code chooses. That means any rollout strategy you want is a few lines of code:

For example, you can:

  • Canary a new toolchain on 5% of new sandboxes by hashing the Durable Object ID.
  • Pin active projects to their current image, so an agent never has its environment swapped out mid-task.
  • Migrate a workspace at a natural checkpoint, like the next session or after a snapshot.
  • Roll back by changing which image future starts choose. No config push, no waiting for a drain.

The rollout policy lives in your Durable Object code, next to the rest of your logic, and it can be as simple or as sophisticated as you need.

Each of these improvements comes from leaning further into the Durable Object, which already owns the workspace's identity, state, and lifecycle. Letting it choose and control its Container gives you more control over every instance and its rollout. It also gives the scheduler a better place to start the Container: wherever the Durable Object is already running. That's what gets the agent to its first command faster.

Faster first commands

Previously, starting a Container required our global control plane to resolve the application configuration, find capacity, and coordinate placement. That model works well for application-wide fleets, but it put deployment machinery in the path of an agent’s first command.

With the durable_object scheduling policy, demand begins at the Durable Object. The Containers infrastructure serving it looks for capacity on the same machine first, then widens the search within the same location if it needs to. It also favors hosts that already have the Container’s image or snapshot in local storage, so the Container can start without downloading it first. 

We also cut work after the Container reaches a host. Instead of booting a new virtual machine from scratch, the new runtime restores a prepared virtual machine that isn’t yet assigned. It reuses networking and filesystem setup, batches repeated operations, and no longer waits on services the first command doesn’t need. 

Together, these changes substantially reduce the time it takes to go from creating a sandbox to running a command in it. On ComputeSDK’s independent Burst TTI Benchmark, which launches 100 sandboxes concurrently and measures time-to-interactive from the client:

Startup measurement

Previous scheduling path

New scheduling policy

Improvement

Median

4.049 seconds

648 milliseconds

6.2x faster

95th percentile

5.839 seconds

910 milliseconds

6.4x faster

99th percentile

6.717 seconds

1129 milliseconds

5.9x faster

The new path also holds up under burst load. In our preliminary burst test, a single account started 100,000 Containers in 5.387 seconds across six locations.

Start with a prepared system image

As the scheduling path gets faster, preparing the image becomes a larger part of the remaining wait. Before a Container can start, its image has to be on the host and unpacked into a filesystem. When that work happens after the request arrives, the agent is left waiting.

That’s why we are introducing cloudflare/debian-trixie: a ready-to-use system image for agents that can configure their environment at runtime, containing Debian Trixie Slim and Node.js 24.20.0 LTS:

This means your agent can start a Linux sandbox without first creating a Dockerfile, building an image, or pushing it to Cloudflare. Once the sandbox is running, your agent can use exec() to clone a repository, install packages, and configure the environment for its task.

And because Cloudflare controls this image, we can distribute and prepare it across eligible Containers hosts before requests arrive. Startups don’t have to download or unpack the base image while the user waits.

Save the workspace and return to it later

A fast start still leaves one more wait: setting up the workspace. Cloning a repository, installing dependencies, and configuring a toolchain can take much longer than starting the Container itself. As the agent works, it also produces files you want to keep. Repeating setup on every start wastes time, and losing the agent’s changes makes it difficult to continue a task.

That is why we’re adding native filesystem snapshots to Containers, in public beta. Snapshots let an agent begin a task in a prepared environment, save its workspace when the task pauses, and restore those files when the session resumes.

Snapshots enable two useful patterns.

First, one workspace can continue across many sessions. For a coding agent, that might mean saving the workspace when the user finishes working and restoring it when they return the next day. The repository, installed dependencies, build caches, configuration, and edits are available without rebuilding the environment.

Second, a snapshot can provide a shared checkpoint for many sandboxes. Since snapshots are immutable and reusable, multiple Containers can start independently of the same prepared environment and make their own changes from there.

Agent evaluations are a good example. An eval might run the same task across different system prompts, skills, models, or agent versions. To compare the results, everything else has to stay fixed: the repository, dependencies, tools, and input files. One snapshot can start many isolated environments from the same baseline, reducing setup time and preventing environment drift from affecting the results.

Snapshots also complement the new system image we introduced above. An agent can start from cloudflare/debian-trixie, set up its environment with exec(), and save the result as a snapshot. Future sandboxes then start from that snapshot with the repository, dependencies, and toolchain already in place.

With snapshots available through the new durable_object scheduling policy, Containers can act as persistent agent workspaces. Compute can stop when work pauses, and a new Container can start from the latest snapshot to pick up where it left off.

The Durable Object advantage for agent sandboxes

The faster scheduling path, runtime image selection, and filesystem snapshots all come from the same design choice: leaning further into the Durable Object as the controller for its attached Container.

Agent systems need a programmable, stateful environment outside the Container to keep state, hold credentials, and control the sandbox’s lifecycle. You can run the agent there, following the “decoupling the brain from the hands” pattern described by Anthropic: when the agent is separate from the sandbox where it works, the agent stays available while its sandboxes and tools can start, stop, fail, or be replaced independently. Or, if you run the agent inside the sandbox, the outside environment lets you supervise it and report progress back to the user.

This is where the Durable Object and Container architecture becomes uniquely well-suited. Every Container is attached to a stateful Durable Object with its own compute and storage running right next to it. You can run the agent in the Durable Object and use the Container as its workspace, or run the agent in the Container and use the Durable Object to supervise it. Add Dynamic Workers for lightweight isolated execution, and an application can choose the execution environment each task requires.

What’s new with the durable_object scheduling policy is that the Durable Object can now control its Container directly, with no wrapper class in between. exec() runs directly in the Workers runtime, and outbound request interception, runtime image and instance selection, and filesystem snapshots are all available on ctx.container. You can combine them with Durable Object storage, alarms, WebSockets, RPC and the rest of your code.

This makes the Container a compute extension of the Durable Object. The Container supplies the Linux environment, while your Durable Object retains the sandbox’s identity, state, policy, and lifecycle. That split is especially well-suited for several patterns:

An agent can remain available while its Linux workspace sleeps. The agent loop can run in the Durable Object, where it maintains session state, communicates with the user over WebSockets, and calls models. It can wake the Container when it needs a shell, compiler, or development server, then stop that compute while it waits for the user or model, paying nothing for idle Linux compute.

The Durable Object can program the security boundary around its Container. It can remember which services, repositories, and operations a user has authorized, then update the Container’s Outbound Request Handler to inject newly granted credentials, enforce new policies, or record additional activity. This resembles the Gatekeeper pattern used by Cloudflare OS, applied to each agent computer.

Evals and reinforcement learning systems can supervise each attempt from outside the environment being tested. A coordinator snapshots a base workspace and forks it into N attempts, each with its own Durable Object and a Container. Each Durable Object runs its attempt, monitors the run, and preserves the result, even if the Container crashes. The coordinator grades the attempts, snapshots the best one, and forks again from there. 

These patterns don’t fit one generic lifecycle. Native APIs let you combine the Container with the Durable Object primitives your application needs, while still using higher-level utilities where they help. You keep control over the Container’s lifecycle, policy, and state.

What this means for the Container class and Sandbox SDK

When we launched Containers, we deliberately hid the Durable Object behind the Container class. We wanted sandboxes to feel familiar and match what developers expected from other platforms. The Sandbox SDK was built on that class, and it filled real gaps: back then, the runtime had no native command execution, outbound request interception, or snapshots, so we built them in userspace.

Since then, agent workspaces have become one of the main workloads shaping Containers, and the cost of that abstraction has become clear. By hiding the Durable Object, we made it harder for you to see and combine the identity, state, and coordination it provides with the Container it controls. Nearly every team we worked with needed something slightly different from the generic lifecycle: their own sleep policy, their own credential handling, their own way of tracking eval runs.

These capabilities are now native, so we're making the Durable Object explicit in the developer experience:

  • New capabilities are native-only. The durable_object scheduling policy, faster startup, runtime image and instance selection, and filesystem snapshots are available only through ctx.container.
  • We'll maintain the Container class and legacy Sandbox class through December 31, 2026. Existing deployments keep running after that date, but the classes won't get updates. We recommend migrating to ctx.container.
  • Sandbox SDK 1.0 is a set of utilities, not a base class. Its helpers work inside your own Durable Object class, alongside ctx.container.
  • For a higher-level environment, @cloudflare/computer combines Dynamic Workers and Containers with a synchronized filesystem.

Migrating usually means changing extends Container to extends DurableObject and calling this.ctx.container directly. See the migration guide for details. Or, get started with your agents:

Get started

Today, most agents are measured by what they can accomplish in a single session. As agents take responsibility for projects that unfold across hours, days, and weeks, the environments where they work need to keep up.

We want every agent to be able to create the sandbox for the task at hand, release the compute when work pauses, and return to the same workspace when the project continues. Today’s changes bring us closer to sandboxes that are as programmable, persistent, and ready to work as the agents using them.

Try the new durable_object scheduling policy, available to all today in public beta, and see what faster startup, filesystem snapshots, and runtime configuration unlock for your agents:

Acknowledgements: This project was also made possible by the contributions of Greg Anders, Andrew Martinez, Nafeez Nazer, Kian Newman-Hazel, Sebastien Pahl, Naresh Ramesh, Cody Roseborough, Nikita Sharma, and Sarah Snell.

The road to the agentic browser: A Kitesurf update

Post Syndicated from Celso Martinho original https://blog.cloudflare.com/kitesurf-update/

In August, we introduced Kitesurf, a browser for the agentic age that runs entirely on Cloudflare Workers.  We built it around what agents need from the web, rather than carrying all the features and bloat of a browser designed for humans. If this is the first you’re hearing about it, we highly recommend you read the blog post where we introduced Kitesurf, for all the juicy technical details of how we did it.

Since then, we’ve put Kitesurf through increasingly realistic tasks and used internal and external feedback from customers to make it more capable and more efficient. Here’s what has changed, how you can try it today, and where we’re going next.

WebMCP support

Websites were not built for agents to use. Browsing today is a messy process of clicking pixels and hoping the right element loads. In a programmatic world this is slow and fragile. WebMCP helps by allowing developers to expose site functionality directly to agents, where they can call functions like searchFlights() instead of simulating clicks.

Cloudflare has been supporting WebMCP since its early days; just a few weeks ago we announced that site owners can now turn on WebMCP with one switch, so browser agents can discover and use tools on their sites without changing the site’s code, and Browser Run has been supporting WebMCP when using Chrome beta for some time now.

Today we are announcing that Kitesurf now supports WebMCP.

You can test this by going to our public Kitesurf playground, opening Cloudflare Radar, and navigating to WebMCP on the Application tab in the DevTools panel. As you can see, Radar exposes a list of WebMCP tools like navigate-to or set-location which allow clients to interact with the page and explore Radar programmatically.

If you target Kitesurf with your AI Agent:

You can see that the AI model can interact with the exposed WebMCP tools which you can use to complete tasks more reliably.

You can read more about how to use WebMCP with Kitesurf and Browser Run here.

New APIs, better WPT coverage

Since the initial announcement, we’ve been adding more browser standards so that agents can render more sophisticated pages. The list of the APIs that Kitesurf supports has grown, and now includes:

We added URL-based module resolution, JSON modules, and import map handling—important for sites that load JavaScript in chunks. Additionally we are using the new Cloudflare Workers’ module registry to support imports from URLs.

Iframe behavior has improved as well; now they load at the right time, stay better isolated, and display text correctly across more languages and encodings.

As we said at launch, running tests is how we keep the quality of both code and results under control without losing velocity while improving Kitesurf. Web Platform Tests (WPT) is a shared, open-source test suite that checks whether browsers implement web standards consistently.

We now pass 730,000+ WPT subtests and are growing. That’s 500,000 more subtests than when we launched. Here you can see the evolution over time, up to the latest version since we started the project:

Efficiency optimized for agents

For an AI agent, efficiency isn’t so much about loading pages fast, but about the latency of the agentic loop. To make Kitesurf truly agentic, we’ve aggressively optimized the browser engine’s internals so that every DOM traversal, timer, and font fetch is as lightweight as possible, ensuring the agent spends its compute cycles on reasoning, not waiting for the browser to catch up.

These optimizations include:

  • Improved JavaScript execution with less work crossing between Boa and the DOM, making the boundary more compatible with real Web frameworks. Common reads such as getAttribute, id, and parentNode can now be answered inside the Wasm DOM instead of making repeated Boa → JavaScript shim → Wasm trips.
  • Kitesurf does less repeated work when running timers and loading scripts, and releases memory from objects it no longer needs. It also handles objects and classes more consistently when code moves between its two JavaScript engines, resulting in running busy pages more efficiently.
  • Kitesurf now loads fonts when they’re needed, fetches fewer fonts a page won’t use, checks which characters appear on the page before fetching language-specific font files and renders synthetic italics more faithfully.

Together, these optimizations help keep Kitesurf efficient for agents. Despite adding support for more web standards—and bringing Kitesurf closer to the capabilities of full-featured browsers like Chrome—its wall-clock time and CPU usage remain roughly in line with our launch benchmarks, and in some cases they have improved.

Plays better with Browser Run

Browser Run is our developer platform product that lets you programmatically control and run headless browser instances. When you use this API, you can select from a list of browser flavors we support, including Kitesurf.

This means that we have to make sure that all of our browsers are supported across the API surface. Starting today, Kitesurf has full Browser Run API coverage. You can use Kitesurf with CDP, Playwright, Puppeteer, or MCP.

One of the most popular Browser Run features, Quick Actions, provide simple interfaces for common browser tasks like capturing screenshots, extracting HTML content, generating PDFs, and more. When we launched Kitesurf, you could use Quick Actions from our REST API. Now you can also use them from inside a Worker script using the env.BROWSER.quickAction() binding:

Kitesurf runs in the terminal now

As we detailed in the How we built it section of our announcement blog post, Kitesurf separates PageScript, the isolate that handles the page session and runs the page code, from PageRenderer, which is responsible for generating the actual pixels from the computed page objects.

This not only gives great isolation and flexibility, but it also allows us to decouple and move the rendering logic to outside Kitesurf (to the client, for example, or to another Worker), while keeping the security-critical parts server-side, running in our network.

If this model sounds familiar, it may be because Cloudflare has another SASE product called Cloudflare Browser Isolation, which runs all untrusted web code at the edge of our global network while it “streams” the rendering data back to the clients.

We can do something similar with Kitesurf. To prove it, we moved PageRenderer to our Playground Worker and patched this version so that instead of converting scene data into an image, it outputs to Kitty—a terminal graphics protocol supported by modern terminals like Kitty itself, Ghostty, WezTerm, and others. We even went a step further and added a pure ANSI text mode for environments where Kitty isn't available.

The result is that you can now quickly open and render a page using Kitesurf without leaving the comfort of your terminal application. This is super useful not only because you can now browse the modern Web at a glance without context-switching, but you can also use this tool to see how an agent using Kitesurf “sees” a page.

To install the terminal version of Kitesurf, do this:

From now on just type this in terminal:

Here’s a demo of it working.

The terminal also sends back scrolling and click events, so you use the keyboard, arrows, or the mouse normally, as if you were in a dedicated browser application window.

And here is our Silent Space Marine Doom demo running in Kitesurf inside the terminal:

Where we go from here

We continue to iterate rapidly on the road to the best agentic browser for our customers and developers. Expect ever-better performance benchmarks and for the list of supported Web standards and WPT test coverage to continue to rise quickly. In fact, we’ve decided to publish the results here and here, in the open, so that you track them as we move forward, in real time.

We are going to continue exploring scenarios where decoupling Kitesurf and moving PageRenderer away from PageScript is an advantage for agents, or where higher frame rates are important. We may or may not have a 30fps Doom version running in Kitesurf as we write this.

We also want to address the elephant in the room: While we are currently prioritizing rapid development, we remain committed to open-sourcing Kitesurf. This is coming soon, but we want to do this right, so we are set up to support it for the long term.

Kitesurf stays true to its initial design goal: It runs entirely on top of Workers just like any other customer application does; that means we only use our publicly available features and APIs and have no access to any special privileges. This is not only a great way to dogfood and prove our own platform, but also the only way to make Kitesurf very cheap and scale automatically across the Cloudflare global network.

Give Kitesurf a try in the refreshed kitesurf.dev playground, and use it in your projects via Browser Run. It’s available for free while in beta, behind per-account limits. Keep an eye on our changelog and come chat with the team on Discord. Share your experience and send us feedback—we’ll be listening.

Introducing Worker Previews: isolated preview environments for every change your agent makes

Post Syndicated from Yomna Shousha original https://blog.cloudflare.com/worker-previews/

Nothing is worse than testing out a change that works in staging, only to see it behave differently in production. That’s why we wanted to give you an environment that’s as close to production as possible — so you can battle-test your changes and make sure they behave exactly as you expect them to.

Agents are helping us push more lines of code than ever before, and larger changes mean more ground needs to be tested ahead of release. Ideally, that testing is done in a way that doesn’t slow agents down, but gives them the tools to take on more of the development lifecycle.

That’s why today we’re launching Worker Previews. Each Git branch gets a production-like place to run, with its own code, configuration, URL, observability, and state.

So now, for every change in your codebase, you can:

  • Deploy an isolated Preview with npx wrangler preview, using its own variables, secrets, and bindings, separate from production configuration and traffic.
  • Share a stable Preview URL for the branch so that every push updates the same running Preview where you can send requests, click through the UI, and test runtime responses.
  • Isolate Durable Objects and Containers per branch, keeping state changes, sessions, memory, migrations, and concurrent tests scoped to that Preview.
  • Inspect logs, errors, metrics, and traces for that Preview to confirm the change works, catch failures, push a fix, and verify it before production sees it.
  • Start from the Preview configuration you set, so each Preview begins with a copy of the variables, secrets, bindings, and settings you define — just like a code branch starts from main. We call this the base configuration.
  • Override a Preview’s configuration when needed, like pointing it at its own database or test API key for migrations — without changing production, the base, or other Previews’ configuration.
  • Serve Preview URLs on a custom domain so that auth providers, cookies, cross-origin resource sharing (CORS), and OAuth redirects work the same way they will in production.

The result is a pre-production feedback loop for every branch. Push your change to a branch, test behavior, inspect performance — before you merge to production.

This enables an Agent Development Lifecycle (ADLC) where each change is atomic, independently deployable, observable, and revisable. And it gives agents and humans the evidence they need to self-improve: catch what failed, push a fix, and verify the next deployment before it hits production.

Every Git branch gets its own environment

When you start work on a new feature, the first thing you do is branch off of main. You get your own copy of the code and make your changes without affecting anything in production.

Worker Previews extend that same model beyond code. Each branch gets its own isolated environment and URL. You can run hundreds of Previews at the same time — each operating independently without affecting other Previews or production.

Production and each Preview have their own configuration — served on their own URL.

When you run npx wrangler preview, the branch gets its own copy of your Previews configuration that you have defined, running on its own URL — all under the same Worker.

In the dashboard, this works like switching branches. Click the breadcrumb next to your Worker's name (it defaults to Production) to see all your Previews:

The dashboard brings every environment into one view. Production sits alongside as many Previews as you need, so contributors can work on separate changes without fighting over a shared staging site. Unlike Wrangler environments, where each environment requires deploying and managing a separate Worker, Previews keep that isolation in one dashboard view.

Each Preview runs as a real version of your Worker. Some changes can only be validated at runtime: an API endpoint has to handle a real request and return the right response. More subjective changes, like a UI update, a new onboarding step, or a different error state, need to be experienced in context before they reach production.

Every Preview has its own isolated and persistent state, with Durable Objects and Containers 

For isolation to extend across your application, stateful resources need special treatment. The reason for that is that Durable Objects run on a singleton model. One instance is responsible for a given object ID, and that instance owns its storage.

If a Preview shared the same DO namespace as production, you wouldn't just be reading stale data — you could modify the same instance serving live traffic in real time (scary!).

That is why every time you run npx wrangler preview, Cloudflare automatically creates a new Durable Object namespace and Container application for that Preview — so that a failed migration or a bad schema change stays contained to that branch and that branch only.

All you need to do is export the class, add its migration, and access it through ctx.exports:

In production, ctx.exports.Counter resolves to the production namespace, while in a Preview, it resolves to that Preview’s namespace.

You now have an entire playground to experiment with. Take Sandboxes, for example, where milliseconds of improvement to startup time can make or break the experience. If you have been trying to improve cold-start performance, you can run different configurations across branches at the same time, compare their cold and warm performance side by side, and find the best setup faster.

Test, observe, and revise each Preview (or have your agent do it)

Now that each branch runs at its own URL in an isolated environment with its own state, you can enter the feedback loop and start battle-testing every change before it reaches production.

You can send traffic to the Preview URL however you normally would — from your terminal, probe from CI, an agent, or by clicking through it yourself. Once that traffic starts flowing, every Workers Observability tool you’re already used to is available, scoped to each individual Preview.

As each request hits the Preview, Workers Observability traces its full lifecycle in a waterfall, including fetch calls, binding operations, and handler invocations. So when something fails, you can follow exactly what happened without sorting through production traffic or signals from other changes.

Observability for Previews looks just like you're already used to for production Workers. Select your Preview from the breadcrumb and open the Observability tab to see its events, errors, and traces:

To give your agents even more control, you can have them open the Preview URL in a headless browser, click through a login flow step by step, and capture a screenshot or record the entire session as replayable DOM events – with Browser Run. 

Below is an example where an agent opens the Preview, captures what was rendered, and connects a failed request to Workers Observability events from the same run.

A reviewer can watch the session in real time with Live View or step in with Human in the Loop when the automation needs judgment.

If something fails, you see it from both angles: what rendered and what happened at runtime. 

That gives the agent enough evidence to keep the pre-production loop running autonomously: deploy, open the URL with Playwright MCP, click through, query the traces through the Workers Observability MCP server, patch, redeploy, and verify. Every iteration stays scoped to the branch.

Configure a base configuration for Previews once, then override as needed

Just like you wouldn't reconfigure your code from scratch every time you branch, you shouldn't have to reconfigure your environment either. 

You set base configuration for Previews once, in a previews block in your Wrangler configuration file.

In the dashboard under Worker → Settings, you see this inlined as Production and Previews Base. Once the base is set, run npx wrangler preview from any branch to create a Preview. If your Worker is Git-connected through Workers Builds, it happens automatically on push.

You can override any setting for only one Preview — without affecting production, the base, or other Previews.

Preview URLs on your own custom domain, protected with Cloudflare Access

To bring the whole setup even closer to production, your preview URLs can be served from your own custom domain. If your app runs on example.com, a Preview for a login branch could run at feature-login.previews.example.com.

If you want to keep those URLs private, you can protect your Previews with Cloudflare Access and require visitors to sign in first.

Testing the whole system before production

We’ve already been dogfooding Worker Previews inside Cloudflare, most notably to build and test CloudflareOS, our open-source platform for safely connecting agents to company systems.

CloudflareOS lets agents work with services such as Google, GitHub, and Slack through Gatekeepers, which control what those agents can access and change. That makes Gatekeeper changes especially sensitive, because a bug could expose data or permit an action that should never have been allowed.

Some of these bugs only appear when OAuth callbacks, permissions, approval flows, and application state are running together. Because testing each component separately cannot show us how the complete system will behave, we deploy an isolated Preview of CloudflareOS and its Gatekeepers for every change under review. We then run the full workflow, fix what fails, and test it again before merging.

We’re seeing customers use Previews for the same basic reason: some problems only show themselves when the change is actually running.

"Previews gives us the ability to iterate earlier at the edge. For IKEA.com, custom domain support helps us avoid Content Security Policy and cookie issues. We’re especially excited for Service Binding support, which will enable communication between Previews and be a game changer for end-to-end testing across our Worker chain." — Santosh Kumar Dwivedi, Senior Software Engineer, IKEA

"At Supermemory, we use Cloudflare heavily, and Worker Previews are exactly the kind of developer experience improvement we wanted to see. For HTTP flows, we can preview Worker changes before they reach production, including routes backed by Durable Objects, and catch issues earlier without slowing down shipping." — Dhravya Shah, Founder, Supermemory

"Previews is amazing for Inspect [Ramp’s coding agent]. I used it to review and test an Inspect PR on my phone that is making reviewing and testing PRs with Inspect on phones responsive…with Inspect." — Dylan Garcia, Senior Staff Engineer, Ramp

What’s next?

You might be thinking: Didn't Workers already have preview URLs? It’s true, we did. We're now calling those Version URLs because they point to specific uploaded Worker versions. Unlike Worker Previews, they don't create an isolated environment for each branch and could only point to production resources. To learn more and compare the different workflows, check out our docs.

Worker Previews is a big improvement from what we offered before, but there's still more to come. Here's what we're working on next:

  • Preview multi-Worker applications. Today, a service binding from a Preview still calls the bound Worker's production deployment. We're working toward keeping the entire request path inside matching Previews.
  • Run Queue consumers and Workflows inside each Preview. Today, Previews can send messages to Queues but cannot consume them, while isolating Workflow executions requires separate configuration. We want the entire asynchronous flow scoped to the branch automatically.
  • Support long-lived Previews for staging and QA. We've heard from teams in the private beta that not every branch is short-lived — some maintain staging, QA, or per-developer environments that persist across sprints. We want to support these end-to-end, and we want to hear how you use them, so we can get it right.

Worker Previews are available now. Get started with the docs, and if you have a feature request or run into an issue, open an issue on GitHub or join the Cloudflare Developers community on Discord.

Acknowledgements: This project was made possible by the design and implementation efforts of Greg Brimble, Patrick O’Donnell, Matt Price, Korinne Alpers, Max Peterson, Cina Saffary, Josh Wheeler, Thomas Ankcorn, Matt Rothenberg, and Brandon Strittmatter, with leadership from Brendan Irvine-Broque and Dan Carter.

Python Workers are now generally available

Post Syndicated from Gyeongjae Choi original https://blog.cloudflare.com/python-workers-ga/

We introduced Python Workers two years ago, providing a way to run Python applications in the Cloudflare Workers runtime. Our goal was to make it as simple to write Workers in Python as it is in TypeScript, and to make the ecosystem of Python packages and frameworks “just work”.

Today, Python Workers are now generally available (GA).

What does GA mean? It means Python is now a first-class, fully supported language on the Cloudflare Developer Platform. You can bring the Python code, libraries, and design patterns you already know and connect them seamlessly to Workers AI, R2, D1, Hyperdrive, Durable Objects, Queues, Workflows, and the rest of the Cloudflare platform. You can also run popular Python frameworks like FastAPI, Django, and Flask inside Python Workers. You can even create a Python Worker inside another Worker using Dynamic Workers.

The journey behind Python Workers

Bringing Python to Cloudflare Workers was a natural choice. Because Workers has supported WebAssembly since 2018, it gave us the perfect environment to run a Wasm-compiled Python interpreter. By using Pyodide, we were able to quickly support a wide range of Python applications in Cloudflare Workers.

Our goal was to create the first platform for infinitely scalable Python apps, while making it as easy and performant as developing Python apps anywhere else.

The features we are highlighting today are the result of this multi-year effort. Many developers are already building applications within Python Workers; today, we are making these capabilities production-ready for everyone.

Python is now a first-class language in the Cloudflare Workers runtime

Python Workers now natively support Cloudflare Developer Platform bindings. Previously, using these Cloudflare bindings in Python Workers required converting Python objects into TypeScript objects explicitly at the RPC boundary. For example, sending a Python dictionary into a Cloudflare Queue required the following glue code to work:

This required Python developers to keep the JavaScript environment and code in mind while writing Python Workers, and it was a common source of error for both humans and AI agents. To address this, we have encapsulated the entire type conversion process within the Workers runtime and the Python SDK. This allows you to utilize all Cloudflare bindings in a Pythonic way without writing a single line of JavaScript code, making the following just work:

Web frameworks: FastAPI, Django, and Flask

You can now run your favorite Python framework, such as FastAPI, Django, or Flask, to build an API server in Python Workers. We implemented a built-in connector that you can use to easily connect your web application to Python Workers.

Let’s say you have a simple FastAPI web application:

In native environments, you would use a web server such as uvicorn to run this application.

In Python Workers, you can run the same application using the workers.asgi package we provide, just by adding this snippet to your code:

Similarly, you can use workers.wsgi package to run synchronous web applications such as Django.

So, what happens under the hood?

Python has a standard contract for how web applications should communicate with web servers, known as the Web Server Gateway Interface (WSGI), or its modern asynchronous counterpart, ASGI. This standard allows developers to build applications that are completely server-agnostic. In a traditional deployment, web servers like Uvicorn or Gunicorn are responsible for handling multiple concurrent client connections and threads to scale traffic, while web frameworks like FastAPI can focus purely on the application logic.

In Cloudflare Workers, the Workers platform itself serves as the web server. Since our global network already seamlessly handles load balancing and infinite scaling, we don't need to reinvent the wheel by running a server inside Python Workers.

Instead, our workers.asgi and workers.wsgi connectors act as a thin, optimized bridge. They translate the incoming native JavaScript request into the standard WSGI/ASGI structures that Python applications expect, and seamlessly pipe the response back out with minimal overhead. By doing this, Python developers get the best of both worlds: you can write and organize code using your favorite web frameworks, while letting the Cloudflare Workers platform instantly scale your API across the globe, without ever configuring a server.

These connectors can be used not only with FastAPI, Django, or Flask, but with any Python web framework that uses the WSGI or ASGI interface.

You can find more information about using each web framework in the Python Workers documentation.

Using PostgreSQL and MySQL with Hyperdrive

If you are building a Python application using relational databases such as PostgreSQL or MySQL, you can now integrate Hyperdrive into Python Workers.

Previously, Python Workers didn’t support TCP sockets, making database drivers unavailable. To understand why this was a blocker, you need to look at how WebAssembly operates. Python database drivers like aiomysql or asyncpg rely on the standard library's socket module to establish connections. In a standard environment, this module makes POSIX system calls to the underlying operating system. Inside a WebAssembly sandbox, those POSIX networking syscalls are normally stubs that always fail. Any attempt to open a standard socket would immediately fail. To solve this problem, we implemented socket system calls using the Workers connect API.

When a database driver attempts to open a TCP connection, it goes through our custom socket syscall implementation. It translates standard Python socket operations like opening a connection and reading bytes into the corresponding JavaScript calls used by the Workers runtime. Because this translation happens at the system call level, your database drivers don't have to know about the underlying implementation at all.

This socket bridge is what makes our Hyperdrive integration possible. To use Hyperdrive in Python Workers, first connect your database with Hyperdrive and set up the binding in the Wrangler config:

Then, connect to Hyperdrive using the database drivers you are familiar with:

You can refer to the Hyperdrive Python Workers documentation to find out how you can use Hyperdrive in Python Workers, and which packages are currently supported.

Expanding the WebAssembly package ecosystem

Because Python Workers run inside a WebAssembly sandbox, any packages with native C/C++/Rust extensions must be cross-compiled to WebAssembly to run in Python Workers. However, previously, there was no standard way to cross-compile any Python packages to WebAssembly. That meant our team had to manually compile and host custom WebAssembly packages. This greatly limited the number of packages you could actually use in Python Workers.

We wanted to fix this and allow users to use a wider variety of packages. However, we didn’t want to merely build packages usable only in Python Workers, which wouldn’t benefit the community. Since Python Workers are built on top of Pyodide, we wanted the ecosystem to evolve in a way that benefits Pyodide and the entire Python-on-WebAssembly community.

To this end, we proposed PEP 783, which standardizes a platform for running Python in the browser runtimes called PyEmscripten. After over a year of discussion and refinement, this proposal was accepted, enabling package maintainers to build and publish packages for the PyEmscripten platform and make them available across all environments that implement PyEmscripten.

We also stabilized the existing Pyodide build toolchain and evolved it into a form that is accessible to all package maintainers, enabling developers to easily build packages for the PyEmscripten platform. Furthermore, we added PyEmscripten platform support to cibuildwheel, to make it easier for others to adopt support for the PyEmscripten platform.

While the ecosystem is still adopting this standard, we hope every Python package will have a wheel that works with WebAssembly in the future. We are also actively working with major package maintainers to add PyEmscripten builds. If you encounter a package that isn’t supported yet, let us know on Discord or GitHub, and our team will work to get it built.

You can also check out our EuroPython 2026 talk: “Python Everywhere: The State of Python on WebAssembly” to see how we made this possible.

Building AI agents and pipelines in Python

The large ecosystem of data science and machine learning packages makes Python the natural choice for building intelligent agents and AI pipelines. But bringing these to Python Workers historically presented a challenge: libraries such as openai and langchain rely on HTTP clients like requests or httpx to communicate with external APIs. However, because of missing low-level socket operations support in Python Workers, these HTTP clients didn’t work properly.

To solve this, we contributed upstream to ensure these HTTP clients can route requests directly through the JavaScript fetch API in WebAssembly environments. Combined with our new support for low-level socket operations as explained in the previous section, this makes the entire networking stack work seamlessly inside Python Workers.

As a result, you can now run AI libraries like openai, langchain, and mcp natively in Python Workers. You can also combine them with Workers AI to run serverless inference on GPUs in Cloudflare’s network, or proxy requests through Cloudflare AI Gateway.

The example below shows a way to run Worker AI models in langchain, using the langchain-cloudflare package:

What you can build today

We have assembled a collection of production-ready patterns in our python-workers-examples repository. Here are some ways you can combine Python Workers with the Cloudflare ecosystem.

Asynchronous AI orchestration

Building a full-stack AI application often means connecting multiple services such as storage, queuing, and inference. This example shows how to build an AI-driven image-to-image generator purely in Python Workers. It accepts user requests, drops them into a Cloudflare Queue, and uses Workflows to orchestrate the image generation step via Workers AI, and stores the image to an R2 bucket.

Real-time stream processing with Bluesky Jetstream

Consuming a firehose of real-time events usually requires a dedicated server to maintain the connection. In this example, we use a Python Worker to connect to the ATProto/Bluesky Jetstream WebSocket. By backing this connection with a Durable Object, the Python Worker can maintain long-lived state, ensuring that the WebSocket connection stays alive.

More examples to explore

Model Context Protocol (MCP) Server

Build and deploy an MCP server using the official Python MCP package to give your AI assistants access to edge data.

Retrieval-Augmented Generation (RAG) system with Vectorize

Building a RAG system using Workers AI and Vectorize, Cloudflare’s vector database.

Python code examples across the Cloudflare developer docs

We’ve updated our docs across Cloudflare products to include Python example code. Nearly everywhere where there is a code example showing how to do something in TypeScript, there’s also a code example in Python. We’re committed to continuing to include Python examples across all of our products. You can toggle code snippets between JavaScript, TypeScript, and Python throughout our developer documentation.

What’s next?

Reaching GA is just the start. We have many plans to make Python Workers better, including making Python Workers more performant and memory efficient, as well as supporting more packages.

Keep telling us what you want to build on Python Workers, and we’ll keep pushing the bounds of what is possible. Check out Python Workers documentation and start building your first Python Worker!

Give every teammate and agent the right level of access to your Workers

Post Syndicated from Dina Kozlov original https://blog.cloudflare.com/workers-granular-authorization/

As more teams — and now agents — build applications on Cloudflare's Developer Platform, having the right access controls is crucial to allow you to ship safely. After all, the last thing you want is for an agent to make a change in production, just because it was granted more access than it needs.

Now, you can give a teammate or agent access to a specific Worker, so that they can only make changes to that application and no other resources in your account. Moreover, we’re giving you four new roles, so you can limit exactly what they can do: 

The new roles are available today, for all customers. You can assign them to a specific user, so when they log into the dashboard, they will only see the Worker you have given them access to. Or, you can create an API token with the scoped access, which you can give to your agent to ensure they only have access to that one application. 

Here’s an example of how to create an API token with permissions per Worker: 

Roles designed for how teams build

When defining these roles, we wanted to strike the right balance. Overly broad roles force you to grant more access than intended, undermining the principle of least privilege, while providing too many individual permissions makes it difficult to know which ones to grant. We landed on four roles that reflect the levels of access you may want to give a person or agent: enough to debug a resource without exposing its content, read the content without changing it, make changes without being able to delete the resource, or fully manage it.

We plan to use these same roles as we bring resource-level access controls to other Developer Platform products, including D1, R2, and KV. Each role can be applied at one of three scopes. For example, if you set the “metadata read-only” control, here’s what that would look like at different levels: 

  • Developer Platform level: Access to metadata for all Developer Platform resources.
  • Product level: Access to metadata for every resource of one product, such as every Worker.
  • Resource level: Access to metadata for one specific resource, such as one Worker.

The role and scope determine what someone can do and which resources they can do it to. Let’s take a look at how this would look in some common Workers workflows.

Debug without exposing source code

To debug an issue, an engineer or agent might need to look at a Worker’s settings, metrics, logs, and traces to understand what went wrong. But they do not need to see the Worker’s code or make changes to it.

Metadata Read-Only gives them access to that information without exposing the Worker’s source code. They can query analytics through the GraphQL API, access logs, and inspect traces and other observability data. Those requests only return data for the Workers they have access to. If an agent is scoped to one Worker, it can use the Cloudflare APIs to investigate an issue without seeing data from any other Worker in the account.

As we bring these roles to more Developer Platform products, we plan to preserve that separation. Someone could inspect settings and observability data for a D1 database or R2 bucket without being able to read the values in the database or the files in the bucket.

Review code without changing it 

A teammate or code review agent may need to read the code running in a Worker to understand how it works, investigate a bug, or review a proposed change. But that does not mean they should be able to deploy new code or update the Worker’s settings.

Content Read-Only provides that separation. It lets them retrieve and review the Worker’s code without being able to modify or deploy it. When scoped to an individual Worker, they can read only that Worker’s code, rather than the code for every Worker in the account.

Once supported for other Developer Platform products, Content Read-Only will work the same way: someone could read the data stored in a D1 database, KV namespace, or R2 bucket without being able to modify it.

Let CI deploy without giving it full control

A CI/CD workflow only needs access to the application it deploys. It should not be able to change another Worker or delete its own and take the application offline.

With Worker-level access controls, each workflow can have its own API token with the Editor role, scoped to one Worker. If the workflow is misconfigured or its token is exposed, the impact remains contained: it can deploy changes to that Worker, but it cannot delete it or touch any other application in your account.

Delete a Worker with Admin access

Admin is the highest level of access you can grant. It allows you to delete an application. You can still scope the role to an individual Worker, so that access does not extend to every Worker in the account.

Routes & Custom Domains 

You can add routes or Custom Domains to a Worker to specify which hostnames are routed to that application. For example, this configuration in your Wrangler file sends traffic for example.com to the Worker:

Because changing that route could redirect production traffic or take the application offline, access to the Worker alone is not enough. To add, change, or remove a route or Custom Domain, you need both Editor access to the Worker and Workers Routes permission for the zone.

Requiring Workers Routes permission, rather than broader access to the zone, means someone can manage how traffic reaches a Worker without being able to change unrelated settings for the domain.

However, once a route is configured, you can continue deploying new versions of the Worker without access to the connected zone or resource, as long as the deployment does not change that connection. This allows your CI/CD system to deploy the application without also giving it access to your domains, databases, or storage.

Workers permissions extend to Durable Objects

Durable Objects do not have their own roles or permissions. Instead, access to a Durable Object is determined by your access to the Worker that implements it. To give someone access to a Durable Object, grant them the appropriate role for that Worker.

Metadata Read-Only gives them access to Durable Object metrics, logs, and traces, but not the data stored in the object. Because Durable Objects Data Studio can query and modify that stored data directly, accessing it requires the Editor role.

Better errors that tell you and your agents which permissions you need

When you give someone narrowly scoped permissions, they may eventually try to perform an operation they do not have access to. When that happens, the error should tell them what permission they need, so they don’t get stuck.

Instead of returning only a generic 403 Forbidden response, our APIs now include a link to the relevant API documentation, where you can see exactly which permissions are required to make the request. This way, you and your agent can figure out exactly the right level of access that’s needed without granting broader permissions than necessary.

Available now

Worker-level access controls are available today for all customers. You can configure them in the Cloudflare dashboard, through the API, or with Terraform.

To give a team member access to a specific Worker, go to Manage Account > Members, select the member, and create a policy with the role and Worker scope they need.

Manage team access with user groups

If several people on the same team or project need the same access, you can create a User Group instead of assigning permissions to each person individually. Assign the policy to the group, then add the relevant members. Everyone in that group will automatically inherit that policy.

Replacing legacy permissions for Workers 

Previously, we used the following roles and permissions to manage access to Workers. Now that we are rolling out a consistent set of roles across the Developer Platform, we recommend using the new roles going forward.

There is no deprecation date for the legacy roles and permissions. Existing assignments will continue to work, and we will provide advance notice before any deprecation. That said, we recommend starting to move to the new roles, since they're the ones that support granular, resource-level access. 

What’s next? 

Worker-level access is the first step toward a more consistent authorization model across Cloudflare's Developer Platform.

Next, we are bringing the same resource-level access controls to more Developer Platform products, including resources like KV namespaces and D1 databases. Instead of granting someone access to every bucket or every database in an account, you will be able to scope access to the specific resource they need and pair that scope with the right role.

The same roles introduced for Workers will apply across these resources.

Check out our developer docs to get started.

Introducing context-aware vulnerability discovery and remediation with Cloudflare Managed Defense and OpenAI Daybreak models

Post Syndicated from Ken Sanderson original https://blog.cloudflare.com/vulnerability-discovery-remediation/

Your scanner just flagged 4,000 new vulnerabilities, 78 of them critical. Which one do you fix first?

To answer that question, Cloudflare is announcing early access to Vulnerability Discovery and Remediation, now part of Cloudflare Managed Defense. Vulnerability Discovery and Remediation is a new, invitation-only Cloudflare service that helps customers detect and mitigate vulnerabilities in their codebases.

Through the OpenAI Daybreak Defense Network, we use OpenAI Daybreak models, including GPT-5.6 Cyber, for reconnaissance, hunting, and validation against codebases that you authorize us to access. If we detect a vulnerability, we will then propose solutions to you, automatically checking each proposed patch and any accompanying proposed mitigation before presenting them for review. Importantly, you are in the driver’s seat: while we may propose code patches and other mitigations, you decide whether they are implemented.

Choosing what to fix first has always been hard. It's getting harder. Large language models can now surface weaknesses across a codebase in minutes, which means the number of findings keeps climbing. But the real problem is speed. Attackers can use AI to accelerate parts of vulnerability discovery and exploitation, giving security teams and developers less time to decide what matters and act on it.

Imagine that your scanner tells you there's a vulnerability in a handler. It doesn't tell you whether that code is deployed. It doesn't tell you whether anyone is actually hitting that route, what security activity surrounds it, or what controls you already have in place. You have to prioritize the finding without evidence of its production exposure or the protections already in place.

This is where we can help. With our global network, we can see which routes are active, how much traffic they carry, and what security events surround them. When customers enable Vulnerability Discovery and Remediation with Web Application Firewall (WAF), we can also see what rules are already applied and are actively blocking attacks. That context turns a generic finding into a specific priority: this vulnerability is in code that's live, on a route that's heavily used, with recent attack activity and no existing protection. And we can help you mitigate that vulnerability by proposing custom WAF mitigations and code patches tailored to your systems.

If this sounds familiar, it should. In “Build your own vulnerability harness”, we described the model-agnostic pipeline we use to scan Cloudflare's fleet, adversarially validate every finding, and turn raw model output into fixes engineers can trust. That internal system is one pillar of Vulnerability Discovery and Remediation. The harness gave us a way to find bugs at fleet scale. Vulnerability Discovery and Remediation brings that discovery process to the code the customer authorizes us to inspect, then connects the findings to production traffic, security events, and the edge controls that can act on them.

This diagram provides an overview of our process, which we explain in more detail below.

Adding context to a vulnerability harness

Our solution works across Cloudflare Workers and proxied applications. The process of detecting vulnerabilities begins with the collection of a traffic and security data snapshot from Web Assets and WAF. The snapshot shows which routes are active, how much traffic they receive, and whether recent security events are associated with them. For instance, a path exhibiting a high volume of detection triggers may also be considered critical for security context purposes. Web Assets and WAF itself serve as the first and second pillar of Vulnerability Discovery and Remediation respectively.

Next, we use source code vulnerability analysis to identify potential weaknesses in code. But that analysis does not show which routes reach it, how much traffic those routes receive, whether they receive suspicious requests, or which protections already apply. We treat routes carrying a high volume of requests as hot paths. Source code deployed to these routes undergoes stricter security profiling. Together, these signals provide evidence about how the API is used and where a vulnerability may be exposed.

For Workers, we retrieve the most recent source version of the Worker and its configured routes to identify the endpoints the Worker serves. Next, we match the Worker's routes to Web Assets and request metadata from Workers Observability, tying the exact source under review to the endpoints it handles in production. This collected network context stays available throughout the investigation, allowing agents to pull it when they need it. 

Our vulnerability harness then starts up. It begins by using the Reconnaissance agent to map request paths to the parts of the codebase that handle them. Reconnaissance uses that map to send hunter agents into specific sections of the customer-authorized code, where they look for vulnerabilities and pull in relevant network context as needed. That context can help the hunter agents pay more attention to code behind an active or recently targeted route, but it does not establish that a vulnerability exists. Every vulnerability finding has to be corroborated by evidence in the source code.

Once the hunters return their findings, the validation stage checks the proposed mitigations before assigning each vulnerability an initial risk rating based on source code. The network evidence we collect can raise that rating further when, for example, the affected endpoint carries significant traffic or shows signs of active probing.

The result is a prioritized list of findings, each with a recommended code patch and, when the evidence supports it, a Cloudflare WAF Custom rule that can reduce exposure while the code fix is reviewed. If you have authorized our VDR to defend your zone, we will deploy the rules, scoped conservatively around the method, path, and other request details needed to reach the vulnerable code. If a route pattern contains only variables and wildcards, we do not suggest a rule. We would rather miss a possible connection than claim one the evidence cannot support.

The HTTP method override bypass example above shows how these signals work together. The harness maps the source finding to the production route, uses traffic and security activity to prioritize it, and scopes a proposed WAF rule around the requests that can reach the vulnerable code. That rule can reduce exposure while engineering reviews and ships the code patch.

Where the model runs

When you authorize an investigation, Vulnerability Discovery and Remediation runs the harness on Cloudflare and sends model prompts from Workers through Cloudflare AI Gateway to OpenAI Daybreak models on OpenAI's servers. GPT-5.6 Cyber is used during reconnaissance, hunting, and validation, and its responses return to the harness so the workflow can continue on Cloudflare. No model inference runs at Cloudflare's edge, and the model cannot apply any patch or rule it proposes.

We keep each investigation narrow by limiting it to the source code and evidence the customer authorizes. Before that context reaches the model, Vulnerability Discovery and Remediation removes what the investigation does not need and applies the redaction controls configured for the engagement. The harness treats source code, logs, and request metadata as evidence to inspect, rather than instructions to follow.

Tool access follows the same boundary: each call is logged and checked against the investigation's access policy before it runs, and every patch or rule proposal must pass checks implemented outside the model. If one of those checks fails, the workflow stops before the proposal reaches customer review.

Nothing is presented for review until it has cleared the checks and our team validates the output. For an edge-defense suggestion, that means validating the rule syntax and running it against synthetic fixtures that represent expected requests, rather than against customer traffic. If a check fails or the result remains ambiguous, we hold the output back and route it for diagnosis.

Passing those checks still does not change your environment. After validation by our team, Vulnerability Discovery and Remediation prepares the source code patch and WAF rule.

Join early access

Vulnerability Discovery and Remediation is available to selected customers by invitation during early access through our Managed Defense team. Each engagement starts with one application whose codebase the customer authorizes us to investigate. To connect the findings to production, Vulnerability Discovery and Remediation uses authorized read access to the Web Assets operation inventory, the relevant WAF controls, and Workers Trace Events Logpush where available. The investigation is semi-automated, but you review every result before deciding whether to test or deploy a change.

If you're interested in learning more, talk to your Cloudflare account team.