Tag Archives: Developer Platform

Streamline: custom video pipelines with Cloudflare Stream and Workers

Post Syndicated from Willi Geiger original https://blog.cloudflare.com/streamline/

Cloudflare Stream is a powerful broadcasting platform that, for many of our customers, just works. But what if you wanted to render dynamic annotations on a livestream or create an alternate version of a hosted video with burned-in subtitles? You would need to run a custom video pipeline.

Today, we’re releasing a new developer playground, Streamline, that demonstrates how you can build a system to deliver these bespoke video experiences on Cloudflare’s Developer Platform. We’ll walk you through how Streamline leverages Workers, Containers, and several media protocols to modify video — and immediately publish that output as livestream or new hosted video. You’ll also have the opportunity to try it for your projects.

A processing pipeline needs a durable, long-running environment that can run specialized, compiled code with predictable memory and CPU capacity. Video streams can run for minutes or hours, so the media process needs a lifecycle independent of the request that started it. An application should be able to start a pipeline, send its input, inspect it, and stop it without needing to keep a single request open for the entire duration.

Cloudflare provides the primitives we need. Containers are long-lived runtimes suitable for media processing. Durable Objects help with orchestration. Finally, Workers are perfect for control signaling and monitoring.

For Streamline, we built a media engine running in a Container to handle media processing in real-time. The Container is controlled by a Worker exposing control, preview, and testing to an agent or user. Processing will continue even if the Worker disconnects. We've architected Streamline with modular components so that the media engine could be replaced with dedicated encoding products in the future.

Architecture

A Streamline deployment consists of two components: the Media Engine, which handles media input/output and processing, and a controlling Application, which creates, configures, observes, and stops media sessions.

Media Engine

The Media Engine has two components:

  • Controller. This is a control harness written in Go that implements an HTTP server, receives incoming requests, and translates them into operations that can be executed by the media engine.
  • Processor that performs the actual media processing. The current implementation uses FFmpeg, but that is an internal implementation detail rather than part of the user-facing API.

The Media Engine is hosted in a Container, and handles all media input/output as well as processing. It can pull RTMPS playback over the network from one Stream Live input and publish RTMPS output to another Stream Live input. It can pull a Cloudflare Stream HLS manifest and its segments to use hosted videos as input. It can accept video input from a source supplied by the controlling application, for example a webcam. It can publish preview video over an outbound WebSocket to a Durable Object relay. An application that needs preview can connect to that relay through its own WebSocket.

Application 

The application is built using Workers, and can be a full-stack browser application, an agent, or an embedded system. It consists of:

  • User interface (UI) including client logic, identity and access policy. This post uses a browser application as its concrete example, so it also includes a browser interface.
  • Orchestrator coordinates the session, the Container lifecycle, and preview relay. The orchestrator is implemented by a Durable Object.

It is possible to run the system locally during development, in which case the container is just a local Docker instance and the Durable Object is not used: there is a single user, the controlling application does not require authorization for local access, and the video preview can connect directly to a WebSocket on localhost.

When these components are deployed to Cloudflare, an authorized user or agent can visit the Worker to start a new session. This spins up a new Streamline container if needed, manages its lifecycle automatically, exposes an API to perform a number of video manipulation operations, and routes inputs from and outputs back to Cloudflare Stream.

Time for a technical deep dive on how the system works.

Container lifecycle and session management

The controlling Worker application initiates a long-running media processing session. After starting the session, the application can disconnect and reconnect safely, while the Container continues processing until the controlling application stops it. We also include a maximum duration to ensure a session is always eventually closed down and can’t run indefinitely, even without external control. While a media processing session is running, the container instance is unavailable for other applications to use.

A Cloudflare Container will automatically sleep if it has not received any incoming requests since a defined interval. However, in our case, once the pipeline is running, it must continue even if the controlling application disconnects and it receives no requests. We can implement this behavior by overriding the onActivityExpired() callback on the container. If the expiry time has not been reached, then we renew the activity, otherwise we destroy the container.

API

The HTTP server implemented by the Go harness and the Durable Object associated with the Container together define the low-level interface to the system. However, we wanted to provide an abstraction over this, so the system is as agnostic as possible to who or what is controlling the session and any unnecessary details of the backend implementation.

We implement this by exporting two packages from Streamline:

  • @cloudflare/streamline/client Defines a high-level, session-based API.
  • @cloudflare/streamline/ Exposes the Durable Object base class associated with the container. This routes the API requests, implements the preview relay server described below, and provides hooks for security and access policy.

In a remote deployment, the controlling Worker is expected to import @streamline/cloudflare and define a concrete subclass of the Durable Object exposed by the container that can be used for application-specific logic and storage.

In local mode, where there is no Durable Object, the frontend defines a thin adapter layer that maintains the session-based API, but connects directly to the local Docker instance with no access controls, etc.

The example below shows how the controlling application can use the API to access Streamline, prepare a session, and start a video processing pipeline.

config is a JSON object that defines the processing pipeline to be executed, described more in subsequent sections.

The table below shows the complete list of all API calls.

Client method

Function

createStreamline()

Creates a new Streamline instance.

streamline.sessions.create()

Creates a new processing session.

streamline.sessions.resume(id)

Reconnects to an existing session.

session.start(config)

Starts a new processing pipeline.

session.ingest(chunk)

Sends a chunk of video data in “webcam” mode.

session.annotation(png)

Updates the transparent annotation overlay.

session.metrics()

Receives metrics about the current session.

session.stop()

Stops the processing in the current session.

Defining and running a video processing pipeline

session.start() constructs and runs a processing pipeline. It takes a single argument which is a JSON configuration object defining the processing to be performed:

  • Input(s)
  • Operations
  • Output

The example below starts a pipeline that takes an RTMP (real-time messaging protocol) broadcast as input (for example, a feed of a Stream Live input receiving an inbound livestream), applies an overlay image with transparency, and sends the output to an RTMP destination (for example, to another Stream Live input for recording or broadcast). This allows the Worker application to create a modified version of a livestream in real time.

Video-on-demand input via HLS

Streamline can also ingest streaming video input via HLS (HTTP live streaming), for example a video hosted on Cloudflare Stream. The example below shows how a Worker application could run a pipeline that ingests a Stream video, reads the embedded closed caption subtitles and renders them as text on the video, and sends the output via RTMP, for example to a Stream Live Input for broadcasting or recording of the modified version.

Sending video to Streamline

It’s often useful to be able to quickly preview a processing pipeline by sending video data directly to Streamline, for example from a webcam. An agent or embedded device application may also want to use this capability, for example to send footage from factory cameras for AI analysis, or to combine multiple camera feeds into a composite view.

The example below creates a pipeline that expects input from the Worker application and produces a preview video output available over a WebSocket (we’ll talk more about the WebSocket preview video below). It applies two filters and an “annotation,” which is an overlay specified as a PNG image that can be updated while the processing is running, for example to implement an animated graphic.

The code snippet above just starts the pipeline. The controlling Worker is not sending any media to Streamline yet. We’ll discuss the openViewer() function below.

The Worker application sends video data to Streamline using the session.ingest() call. The example below shows how a web browser application might receive chunks from the webcam and forward them to Streamline.

Animated overlay

The annotation overlay can be updated using the session.annotation() call. The example below shows how the Worker application could snapshot a canvas and send it to Streamline. This could be done on an animation loop, although the update rate may be limited in practice by the size of the PNG overlay images, the available bandwidth, and processing power.

Receiving preview video from Streamline

Streamline can also produce preview video output, by specifying output: { mode: 'websocket' }.

Streamline uses WebSockets for low-latency preview video delivery back to the controlling application: the container publishes fMP4 fragments to the Durable Object, which forwards them to an output relay available over a WebSocket on the URL /relay/view, relative to the application origin. The application must connect a WebSocket to this URL, and will then receive video data pushed to it as it becomes available from Streamline. The code snippet below shows how a web browser application might display the preview video feed.

A production MediaSource player must queue fragments while SourceBuffer.updating is true. In local development, the browser or other controlling application simply opens a WebSocket connection directly on the local container.

Currently supported operations

In the configuration object passed to session.start() in the examples above, pipeline is an array of operations from the set supported by the underlying media engine. The operation order is currently fixed by the engine; the order specified in the array is not significant. The list of currently supported operations and the order in which they are applied is below.

Operation name

Function

filter

Applies filtering operations, e.g. blur, saturation.

overlay

Overlays an image referenced by URL or a binary PNG specified separately in a call to annotation().

subtitle

Burns in subtitles.

encode

Specifies output encoding parameters.

Security

This is Cloudflare, so it is important that security is part of the design rather than an addition at the end. We need to ensure that only authorized users can create a new session or take control of an existing one, and that sessions are isolated from each other. We must treat Stream RTMPS input/output keys as secrets that shouldn’t be leaked to the controlling application. We must ensure that resource use is bounded.

The owner deployment is kept private using Workers’ Access integration. The configured owner identity and other allowed users can edit the same shared profiles and start a session while the singleton is idle. The Worker verifies the Access session before accepting control requests and binds the active session to the verified principal. Only one session can run at a time, and a different principal cannot stop or replace the active session.

Stream Live Input keys are stored in Worker secrets or as write-only shared overrides in Durable Object storage. They are never returned by the settings API or placed in browser storage. The controlling application specifies RTMPS input and output by referring to a named profile. The Worker resolves the profile before contacting the container.

The preview video stream has two credentials with separate purposes. A Cloudflare Access service token authenticates the container workload to the publisher endpoint. A random per-session capability authorizes publishing only for the currently active relay. The service token is injected by the container's outbound Worker and never enters container memory. The initial deployment uses a temporary path-specific Access Bypass while the per-session capability remains enforced; after deployment and a successful smoke test, the rollout replaces Bypass with Service Auth.

The owner deployment is intentionally private and singleton-routed. It is not the security model for a public multi-user service.

Playground and open source

We want you to try out Streamline and start building! So together with this post, we are releasing the system as open source and deploying a public playground.

The Streamline container can be run locally or deployed on your account. It exports the Worker API for your control application to use.

There is also an example Worker application with an Astro web frontend that demonstrates Streamline functionality with a few common use cases, including overlays, subtitle decoding, filters and picture-in-picture. There is probe functionality that provides performance metrics and system tracing, and can be useful for debugging the system when developing new features. The example application can be run on a local Astro server, or is set up to be deployed behind Cloudflare Access, so you can control who has access to your Streamline instance.

Both repositories are available as open source on Cloudflare’s GitHub:

We have published a public playground deployment of the example application. This is also something a user can deploy if desired. It uses its own Access configuration, one container identity per verified user, one active session per user, global admission control, concurrency, media and session limits, and no ability for one user to replace another user's session.

You can try the public playground at:

Where we go from here

Streamline demonstrates one way to combine existing managed services, like Stream, with lower-level primitives to build highly customizable media pipelines. In this iteration, Streamline uses Container CPU for media processing, which introduces a bottleneck at higher qualities or frame-rates.

Moving forward, we’re excited to see how we and our developer community can extend this architecture to  build new support for computer vision pipelines, hardware-accelerated media processing, realtime experiences with next generation protocols like WebRTC and MoQ, and ultimately video encoding and decoding primitives natively in Workers.

Today, we invite you to check out our hosted demo of Streamline to see how powerful these tools can be. From there, check out the codebases we’ve open sourced to see how easy it is to deploy Streamline into your own account and use it to create your own experiences.

Introducing Web Search API via AI Gateway

Post Syndicated from Michelle Chen original https://blog.cloudflare.com/introducing-web-search-api/

Fun fact: when you use an agent and it needs to fetch a live web page, the agent usually just guesses the URL of the page and then makes a tool call to curl it. This is why you’ll sometimes see web fetches come back with a 404 Not Found, which happens if the agent incorrectly guesses the URL of that information. As you can imagine, it’s not super efficient to randomly guess URLs all the time.

There is a better way. What if your agent can actually browse the Internet, just like how humans start with a search engine query when we’re looking for information? This is what web search is designed to do — it enables agents to search for relevant data on the Internet and grounds an agent’s responses based on live information.

Today, we’re announcing Cloudflare’s partnership with web search providers to bring you grounded intelligence via AI Gateway. We’re kicking off this launch with our partners from Ceramic.ai, Exa, and Linkup.

What can I do with the Web Search API?

AI models are only as good as the context you feed them. Models are typically trained and then frozen at a point in time, operating only on information that existed before their knowledge cut off date. This makes it quite hard to engage with models about recent events, changing APIs, or fast-evolving news.

Integrating Web Search API directly into your inference pipeline equips your agents with a dynamic context layer. Your applications get fresh, structured snippets from the web injected straight into context, which gives your models access to live information.

For example, if your agent was building with Cloudflare developer tools, it might miss all the new products and features we’re releasing during this Birthday Week! With web search, you’ll be able to retrieve the latest and greatest documentation and releases, so you can build faster and smarter.

Elevating the industry standard for web search

At Cloudflare, we believe that crawlers should be honest, transparent, and respect all bot rules and preferences, and that site owners should have meaningful transparency and control over how their content is used. With our launch today, we’re excited to announce that our partners have committed to meeting Cloudflare’s bot crawling standards.

The crawler used by the web search provider must comply with Cloudflare’s publicly stated requirements for “Verified bots” as defined in our developer documentation, and web search responses must include a link to the location of crawled content. These rules are net-positive for a fair Internet, and allow creators to decide what they want to do with their data.

We’re extremely excited to be taking another step in setting the bar for what it means to be a good crawler on the Internet, and even more proud of the web search partners who have risen to the challenge to uphold these standards with us. When you use web search on Cloudflare, you choose to consume search knowledge from operators committed to providing the transparency, control, and visibility that helps build a better Internet, by identifying their crawlers, respecting robots.txt, and providing the source of search results.

Using Web Search API via AI Gateway

Cloudflare’s AI Gateway is the flagship integration point for our new Web Search API product. AI Gateway is designed to be the control plane for your applications, with observability, unified billing, security, and access controls all built-in. Naturally, we thought web search would fit right in: you can consume web search with your AI Gateway credits; collect logs and request data on web search calls; and control who gets access to what web search providers.

Requests show up in your normal AI Gateway observability logs, and web search queries draw down from your AI Gateway credit balance. We will identify partners supporting Zero Data Retention (ZDR), so that you know that your data is not retained. We also offer web search directly at list API pricing from our partners, without any additional markup. More details can be found on the web search developer docs.

We also support Bring-Your-Own-Key (BYOK) with web search providers, as we do with model inference providers. This way, you’re able to bring your existing organizational setup and get started with AI Gateway and web search in a few simple steps.

Direct REST API

If you’re making HTTP calls from an existing backend, mobile app, or external service, you can query web search directly through a standard REST endpoint. Simply pass your AI Gateway authentication token, specify your preferred provider in the payload, and it will return web search results.

Workers Bindings

For developers building directly on Cloudflare Workers, integration takes just a line of code. We also have a standalone worker binding that you can use to do web search for your agent:

Coming soon: Server tools

We are actively building native Server Tools directly into AI Gateway. Soon, you won’t need to define tools yourself: they will come built into our control plane so you can spend more time building rather than orchestrating the harness. Web search will be one of the first tools we incorporate into our stack of server tools, and we’re excited for you to try it.

However, if you’d like to orchestrate web search as a server tool yourself today, you can easily do so with the following Worker code snippet:

Try it out today!

We’re excited to bring web search to our platform and to be doing this with wonderful partners who are championing what it means to be a good steward of the Internet. Please give our new web search tools a try via AI Gateway, the standalone REST API, and other formats in the future. Get started with our developer docs today, or try it out on the AI Playground with your AI Gateway account and key.

Introducing Cloudflare Traces: follow requests through our entire platform

Post Syndicated from Mar Witek original https://blog.cloudflare.com/cloudflare-tracing/

Today, we’re introducing Cloudflare Traces in open beta, extending automatic tracing beyond Workers to the rest of the request path. In one trace, you can see supported security rules, transformations, cache decisions, routing, Worker execution, and origin handling, then continue that trace through services running on Cloudflare, at your origin, or elsewhere in your stack. This is a long-term investment in OpenTelemetry and in making Cloudflare the most observable part of your stack.

You can now:

You can enable tracing in the Cloudflare dashboard on any domain or let your agent set up for you:

Giving you the visibility we use to debug Cloudflare

When our own teams investigate, we use our own internal traces, which often include thousands of spans for a single trace, generated by dozens of services and features. This lets us dig deep into every detail of a given request. We don’t think that visibility should stop at our internal systems.

Workers Tracing was our first step toward exposing what happens on our platform. Last year, we launched automatic instrumentation for Worker invocations, including outbound fetches and calls to KV, R2, D1, Durable Objects, and other Workers. It shows the work performed inside the Workers runtime without requiring tracing code for every operation.

The goal of Cloudflare Traces is to bring the same level of visibility to everyone using Cloudflare, whether you’re building on Cloudflare or just have Cloudflare in front of an origin. You get to see how your traffic moved through our platform, and connect the dots between how you’ve configured Cloudflare, and how this influences request processing time, routing decisions, and more. 

Follow one request end to end

A request’s path through Cloudflare can be complicated! It might pass through security rules, transformations, routing, caching, or proxied to another service entirely. Cloudflare Traces records each supported step as a span, including its timing, outcome, and relevant attributes. Instead of reconstructing the request from separate logs and configuration, you can see the request’s path through our system in one place.

You can answer questions like:

Why was the request blocked or challenged, and which security rule took action?

See when custom or managed rules evaluated the request, how long evaluation took, and the resulting action. Identify the rule responsible for a block or challenge through its span events.

Was the URL rewritten by a Transform Rule before it reached the application?

You can open the http_request_transform span to see each change, the request component it affected, and the rule responsible. You can also see where the transformation occurred relative to routing and origin handling.

Which Page Rules, Snippets, or Workers handled or changed the request?

The workers_routing span shows whether a route matched, which routing type was used, and the matching route pattern.

Was the response served from cache, and where was time spent between Cloudflare, the origin connection, and the application?

You can expand nested cache, upstream, and origin spans to see where the request spent its time. Here, you can see there was a cache miss that went to origin and spent 527ms of the 539ms getting a response.

Configure your tracing

There is no special instrumentation, config, or plugins required. Once tracing is enabled for a domain, Cloudflare generates these spans automatically. This lets you extend the trace through third-party services and back again by adhering to open standards. From there, you can control which requests are traced using a baseline sampling rate and Trace Rules.

Set a baseline sampling rate

You can enable tracing on any domain and set a baseline sampling rate to balance visibility, data volume, and cost. You might trace 1% of requests during normal operation, giving you a continuous view of request behavior without collecting a trace for every request.

Configure Trace Rules

Trace Rules let you keep a low baseline sampling rate while capturing complete traces for a specific investigation. If one customer reports a problem, you can trace 100% of traffic for their hostname, source IP, or identifying request header while leaving everyone else at 1%. Or during an investigation, you could trace 100% of requests carrying a temporary debug header, while leaving all other traffic at the baseline. This lets you reproduce an issue without increasing tracing across the entire domain.

Trace Rules use the same Cloudflare Rules language, so you can target paths, methods, headers, IP addresses, geographies, or combinations of those properties.

Accept and propagate trace context

One of the most common requests we hear is for true distributed tracing: a single trace that follows a request into Cloudflare, through our platform, and onward through the rest of your stack.

Cloudflare Traces can accept a W3C traceparent header from an incoming request, allowing Cloudflare spans to join a trace that began before the request reached our platform. An incoming propagation policy controls whether Cloudflare accepts that context.

Cloudflare can also forward a new traceparent header to your origin. Any other instrumented services can extract that context and continue the trace through APIs, databases, and services running on Cloudflare or elsewhere. To view everything as one connected trace, you can send both Cloudflare and application spans to the same OpenTelemetry-compatible backend.

Export traces to your observability platform

You can export Cloudflare spans over OTLP to a compatible observability platform, where they appear alongside telemetry from the rest of your stack. Configure an account-level destination, then choose which domains send traces to it. This is part of our commitment to OpenTelemetry: Cloudflare represents request activity as OpenTelemetry spans and delivers them using OTLP, keeping the data portable across observability tools.

Let your agent investigate Cloudflare Traces

When you ask a coding agent to debug a production issue, it might inspect your code and run tests, but it may not be able to see what happened to the request in production. With the Cloudflare Observability MCP server, your agent can leverage our SQL API to query your traces (and all of your observability data!), giving it access to your investigation production telemetry.

Let your agent find the right requests, comparing failed traces with successful ones, and identifying where their spans diverge. Since the agent can also inspect your repository, it can connect those findings to the relevant code, narrow down what needs to change, and help put up a fix for you to review.

Pricing

Cloudflare Traces will be a part of the unified Cloudflare Observability pricing model. Instead of charging by the number of spans/events, pricing is based on how much observability data you ingest and how long you retain it. New pricing will take effect across Cloudflare Tracing (and Workers Tracing!) starting December 1, 2026.

Plan

Included Usage

Retention

Additional Usage

Free

0.5 GB of ingestion per day

7 Days

Not available

Paid and Enterprise

50 GB of ingestion
10 GB-month of storage per billing cycle

Up to 1 year 
(coming soon)

$0.25 per GB ingested
$0.10 per GB-month stored

What's next

Following the open beta, we plan to launch:

  • Broader automatic instrumentation: Add more spans across both the HTTP request path (e.g. DDoS rules, Access) and the Workers execution path (e.g. Workflows, Queues, Pipelines).
  • Authenticated context propagation: Let trusted callers continue an existing trace without accepting context from every incoming request.
  • Ad hoc tracing: Capture a specific request on demand without changing the baseline sampling rate.
  • OpenTelemetry API support in Workers: Continue building out our OpenTelemetry APIs to enable adding attributes to existing spans or getting trace context.
  • Longer retention: Keep trace data available for up to 365 days for longer-running investigations.

Get started

Follow the Cloudflare Traces documentation to trace your first request and tune sampling with Trace Rules. Cloudflare Traces is available in open beta from the dashboard, through the API, or with Terraform, with support for exporting to an OTLP destination.

AI Search is now generally available

Post Syndicated from Gabriel Massadas original https://blog.cloudflare.com/ai-search-ga/

Cloudflare’s AI Search combines Workers AI, Vectorize, R2, and Browser Run into a fully managed index and retrieval pipeline. Since we launched AI Search over a year ago, we’ve seen developers use it to power a wide range of search use cases, from searching internal documentation to powering search for their websites. We use AI Search ourselves to power search on our own blog and developer docs.

Starting today, AI Search is generally available. And as part of it, we've expanded and improved our support for multimodal formats beyond text, adding native image embeddings, optical character recognition (OCR) for PDFs, and support for larger files.

As part of general availability, we’ll start billing for AI Search on November 1, 2026, and continue to offer a generous free tier on all Workers plans.

New: multimodal embedding and retrieval

An image is more than the sentence used to describe it. Product texture, screenshot state, chart relationships, document layout, and fine visual detail can all disappear when pixels are compressed into a caption.

AI Search now preserves both signals: it embeds image pixels directly for visual retrieval while retaining captions for textual understanding. To keep these richer representations efficient, AI Search leverages Matryoshka Representation Learning (MRL), allowing smaller embeddings to retain useful information while keeping storage manageable and search fast.

Although we originally supported retrieval over images, the implementation was naive: we would perform object detection, generate a caption, and then embed that text. This made images searchable, but only through the details captured in the caption. Now, we do both — caption-based understanding and native image retrieval.

Native multimodal retrieval is available today with the Qwen3-VL-Embedding model. At query time, AI Search checks whether your instance’s embedding model supports images. If it does, a query image is embedded directly by that model, landing in the same vector space as your indexed images and text.

If your embedding model is text-only, you can still query with an image. AI Search converts the query image to text with ToMarkdown and searches using the resulting caption. This gives every model basic multimodal support, while models with native image support get the full visual signal.

Caption-based understanding

A small bird with black-and-white markings perched among golden fruit and green leaves

Native Image Retrieval

Can match details omitted from the caption: the geometry of the bird’s white eyebrow stripe, its yellow-green plumage, the mixture of smooth and weathered fruit, the leaves’ deeply ribbed texture, and the image’s warm palette and shallow-focus composition.

The caption is a compressed interpretation of the image. Capturing every potentially useful detail requires long or specialized captions, written with the eventual search query in mind. Native image embeddings preserve visual characteristics without requiring the caption to anticipate which details matter.

This enables searches that are difficult to express precisely with words. You can describe an image you want to find, provide another image to locate visually similar results, or combine both, such as “a bird with similar markings” or “a bird perched on a leafy branch with plums.” It is useful for product discovery, screenshot matching, charts, diagrams, scanned documents, and other collections where color, texture, composition, or spatial relationships matter.

Here’s how a query moves through AI Search. First, the query is optionally rewritten, then embedded (image queries are embedded directly by multimodal models, or captioned first by text-only models). Vector and keyword search run in parallel, and results are fused and optionally reranked. The top chunks are returned, or passed to a generation model to write an answer.

Bigger files and OCR for scanned documents

AI Search now accepts your text files (Markdown, HTML, CSV, JSON and similar) and PDFs up to 10 MiB, up from 4 MiB. Many PDFs are really scanned images with no extractable text. For those, turn on OCR and AI Search reads the text from each page before chunking and embedding it. OCR is available to every account and is billed under the new AI Search pricing as image processing ingestion tokens.

Now in GA: billing and pricing for AI Search

During our August 2026 Agents Week, we announced preview pricing for AI Search. With the product going GA today, we’re announcing billing for AI Search that is going live on November 1, 2026. We’ll send a reminder email before billing is enabled.

AI Search pricing is designed, so you can estimate your bill before you index a single file. You pay for three things: the content you ingest, the data you store, and the queries you run. The work in between (parsing, chunking, embedding with Workers AI models, keyword indexing, and reranking) is included. There are no instance hours, capacity units, or monthly minimums to size up front, and small projects fit inside the free monthly allotment.

Ingestion pricing is based on one rate per token with whichever Workers AI embedding model you pick, and tokens are counted the same way for every model. Switching from a text-only embedding model to a multimodal one doesn't change what you pay to ingest, unless you are also processing images (add-on fee). Storage is priced on the size of data in your indices. Querying is priced based on the type of query (semantic vs. full-text) and how many queries you send.

Estimating your cost comes down to how much content you index, how much you store, and how many queries you expect. Here's the pricing we announced in preview, with one tweak that we’re making: on the free monthly allotment, you will receive 1,000 semantic queries and 1,000 full-text queries (instead of a shared pool of 2,000 queries).

Pricing

Free monthly allotment (all Workers plans)

Ingestion

Base Ingestion

$0.75 / 1M tokens

5M tokens †

Image processing (add-on)

+$0.50 / 1M tokens

5M tokens †

Storage

Stored data

$2.00 / GB-month

10 GB

Query

Semantic (hybrid and vector search)

$0.75 / 1k queries

1,000 queries

Full-text

$0.10 / 1k queries

1,000 queries

Embedding and Reranking

Ingestion and query

Free with select Workers AI models; third-party billed separately

N/A

† A single pool of 5M ingestion tokens per month, covering any file type currently supported (e.g., text, images).

What’s next

Multimodal embedding support is just the first step; we’re building an ingestion pipeline to support full video and audio processing to allow our customers to search their rich media assets.

We’re also refactoring the keyword search engine so that it scales better with your content requirements, particularly for when you have big data stores where the current implementation has limits.

Finally, we’re developing better and simpler ways to enable AI Search and create indexes for websites already running on Cloudflare. This ensures AI agents can discover, explore, and consume content more easily and efficiently.

Stay tuned for these follow-up announcements and more.

AI Search is now generally available to enable and use today. Get started with our new multimodal embeddings, bigger files, OCR features, and our new managed instance pricing. Check out the AI Search developer docs for more information.

Announcing Cloudflare K2: serverless event streams

Post Syndicated from Micah Wylde original https://blog.cloudflare.com/cloudflare-k2-streams/

With traditional Remote Procedure Call (RPC) architectures, there exists a core challenge: producers and consumers must align in scale and in time. If your producers send too much data for your consumers to handle or if your consumers or downstream services become unavailable, events are dropped. This problem is compounded with multiple consumers that need to independently process the data. For example, an ecommerce backend may emit events when transactions are completed, which need to be read by an analytics system and a fraud detection service.

We can solve this by decoupling our producers and consumers — inserting a service in the middle that absorbs writes while allowing independent readers to consume at their own pace.

Today we are launching Cloudflare K2 in public beta to solve this problem. K2 is a durable event streaming primitive on the Developer Platform. You send events to a K2 stream, which stores them as an ordered log. Consumers can read them in a variety of ways, for example by splitting up reads across a set of consumers, or delivering all messages to all consumers. It's fully serverless, scales to vast quantities of data, and supports long-term retention, so even long periods of consumer downtime do not lose data.

Under the hood, K2 implements a partitioned, durable log on top of R2 object storage, which allows it to scale to huge volumes of storage.

If you’re ready to get started, you can create your first stream in seconds by following the guide here.

Streams on the edge

We first built K2 because we needed a durable buffer on the edge, initially to serve as the ingestion layer for Basin Pipelines. Pipelines is powered by a stream processing engine that operates on a pull-based model, which means some other system has to store events before they are read, transformed, and written to R2. And because we commit to never dropping events once they’re accepted into the Pipelines Stream, that storage has to be durable — meaning it can’t lose data — over potentially long periods of time.

This is where most companies would deploy Apache Kafka. However, Pipelines runs on the Cloudflare edge, which spans a huge number of servers across over 335 cities. Our unique architecture means we often cannot run traditional distributed systems software like Kafka, and need to rethink how these systems are built and operated.

For stateful services, in particular, Cloudflare’s global infrastructure presents some challenges: we get relatively small slices of machines, those machines are relatively ephemeral, and networking is often over the public Internet. But our infrastructure also has a few superpowers: it’s close to users wherever they are in the world and has an incredible capacity to scale horizontally.

In designing the durable buffering system that became K2, we decided to rely on the powerful state primitive we already have: R2. Object storage systems like R2 combine extremely durable storage (11 9s!) with strongly consistent APIs. Offloading replication and consensus to the storage layer allows us to make the application layer (K2 in this case) radically simpler, cheaper, and higher performance. A secondary benefit is that it separates compute and storage, meaning each can be scaled independently. This allows us to store vast quantities of historical data at low cost.

How do we build a log on top of object storage? An immediate issue is that R2 — like other object stores — does not support appends, the standard operation on a log. Instead, we must write complete files, or segments, that are large enough to overcome the cost of writing and reading each one. We do this by first accumulating writes in-memory on an edge service. After waiting a short period for data to arrive, we write all events as a segment file. We achieve ordering and strictly incrementing offsets using R2’s atomic operations without needing a separate coordination service.

While building on R2 has many advantages, there is one downside: higher produce latencies. Writing to object storage is slower than a local disk, and we have to wait for the local batch to accumulate before starting the write. In our initial release of K2, this adds up to about 1 second of produce latency at the 99th percentile of response times.

We will be sharing more details on the design of K2 in an upcoming technical deep dive.

Streams, Queues, or Pipelines?

Cloudflare has several existing asynchronous delivery primitives, including Queues and Basin Pipelines. When should you reach for K2 instead of these existing products?

There are some superficial similarities between Queues and K2 Streams: both receive events, durably store them, and deliver them to consumers. Queues are designed around tracking individual items of expensive or time consuming work that need to be asynchronously completed. For example, an image processing application may enqueue a user request to be handled by the actual image processing service. They support complex logic on the grain of a particular work item, like retries, delays, and dead-letter queues for failed attempts.

K2, by contrast, is designed for high-scale data movement, long-term retention, and fan-out consumption. Messages are produced and consumed as batches — enabling efficient processing at the expense of message-level retries. This batching also drives higher producer latency than for queues.

Basin Pipelines is a serverless ingestion service. You can send your Pipeline JSON events, which can be transformed and written to R2 or a Basin Catalog. We recommend Pipelines when the end result is writing your events to object storage or Iceberg tables, and K2 when doing custom processing or writing to other destinations.

Getting started

Using K2 involves first creating a stream. You can have many streams across your account for different use cases or types of events. Streams can be created via cf, Wrangler, the dashboard, or API.

Let's take the example of collecting and processing product analytics. First, we'll create a stream with cf:

Once we have a stream, we can start producing to it, via an HTTP API or Worker binding. For example, from a Worker:

K2 represents data as bytes, so you can use whatever format or encoding makes sense for your application.

Now that we have events in a stream, we can create a subscription. Subscriptions divide up work between consumers, enabling read parallelism — scaling out to multiple readers to handle more load than a single server can manage.

We can create a subscription via the HTTP API.

With the subscription created, we can then poll it from each of our consumers:

When a client calls consume, they receive a lease for that particular batch of events for 5 minutes. The client can do one of three things:

  • ack the batch, which marks it as processed and ensures it will not be redelivered
  • nack (negative ack) it, meaning we’ve failed to process it and would like it to be redelivered
  • extend its lease, in case it needs more time to complete processing

This is one way to consume from K2: splitting work amongst multiple consumers such that each consumer gets a portion of the data. Another way to read is with a separate subscription for each consumer — the pub/sub pattern — in which case each consumer sees all of the messages. Or you can mix-and-match between these two approaches, having multiple independent consumer pools.

See the K2 docs for full details on the APIs.

Pricing and Availability

K2 is available today in public beta for accounts with Workers Paid subscriptions, within these limits:

  • Maximum of 10GB of storage used
  • 30 MB/s produce per stream

If you need higher limits, please reach out to the team on Discord or fill out the limit increase form.

Usage of K2 will not be billed during the beta period. Once we begin billing, we anticipate this pricing:

Pricing

Data Produced

$0.04 / GB

Data Consumed

$0.04 / GB

Data Retained

$0.02 / GB / month

What’s next

We have an exciting roadmap for K2 over the coming months, including:

  • Higher write parallelism, up to multi-GB/s streams
  • Message keys and key-based ordering guarantees
  • Push-based worker consumers
  • Express tier with lower produce and end-to-end latencies
  • Drop-in support for Apache Kafka clients

We’re excited to see what you build on K2! Share your feedback on the Cloudflare Discord.

Identify AI model overuse with User Insights

Post Syndicated from Ayush Kumar original https://blog.cloudflare.com/ai-model-overuse-user-insights/

When we launched User Insights last month, we wanted to help teams answer a basic question: What are people actually doing with AI? User Insights gives teams a clearer view of their AI usage, showing which users, applications, tasks, and models are driving traffic. It also highlights user and agent anomalies, helping teams identify unexpected or out-of-control spending and usage before they become larger problems.

Our latest update adds something our users have been asking for: context. 

Since launch, we’ve heard from users that model names and request counts only tell part of the story. They show where traffic is going, but reveal little about the work behind it: is that request a code review, a research task, or an agent making several calls to complete a job? The same token count can represent very different kinds of work, and you can’t evaluate with model choice without understanding the task.

User Insights now shows when a model may be more capable than a task requires, which of your users and agents are driving that usage, and how the task, model, cost, and conversation patterns relate. Teams can use these insights to investigate and make targeted changes within their organization. These capabilities are available for free to AI Gateway users.

Why AI usage is hard to understand

Consider a team that has routed its internal AI traffic through AI Gateway. After a few weeks, spending is increasing and some requests feel slower than expected, a common challenge as organizations adopt AI at scale.

There could be several explanations. Developers may be using AI for increasingly complex coding work. Agents may be making too many follow-up calls to complete a task. Or a small group of users or agents may be responsible for a disproportionate share of the organization’s usage.

Tokens and request counts alone cannot show which pattern is driving the increase. Teams need to understand what the traffic represents before deciding whether a model, workflow, or routing rule should change.

Helping teams find where AI models are overkill 

The model overkill view helps teams identify conversations where the selected model appears to be more capable than the task requires. For example, a team might discover that users or agents are sending simple formatting or summarization requests to a high-capability reasoning model.

That gives the organization a place to start. They can see which users, agents, or applications are associated with the pattern, then investigate the tasks behind it. A team might find that a model is being used because it is the default, because users are unsure which model to choose, or because an agent has been configured to use the same model for every step.

The overkill view is not a leaderboard and does not automatically recommend a replacement model. It helps teams ask better questions:

  • Is this model appropriate for the task?
  • Is the extra capability improving the result?
  • Would a faster or less expensive model produce an equivalent outcome?
  • Is the issue limited to one workflow, user, or agent?

From there, teams can compare cost, latency, token usage, and conversation turns before deciding what to change.

These insights support both the new Potential Savings view and the Auto Router, which is launching in public beta alongside this release. The Potential Savings view helps teams identify requests that may be handled by a faster or less expensive model without compromising output quality. The Auto Router applies these task and model-fit signals automatically, helping reduce costs without requiring a separate routing rule for every workload.

The Overkill view is a starting point for evaluating model fit. Teams can compare latency, input and output tokens, conversation turns, and total cost for the same type of task. A difficult coding or research task may need a capable reasoning model, while a short summary or simple classification task may not. The goal is not to move every request to the least expensive model, but to understand whether the selected model is appropriate for the work.

Understand what people are using AI for

Task analysis groups conversations by the kind of work they represent. Initial categories include coding, research, writing, summarization, and data analysis.

This provides context that a list of model names cannot. An engineering team might use AI mostly for coding and debugging, while another team might use it for research and summarization. A team may also discover that a surprising amount of traffic comes from simple tasks, even though those tasks are being sent to a high-capability model.

The answers will vary by team. The category data provides a way to investigate those differences using traffic already passing through AI Gateway. Teams can determine whether a model is being used for the work it is best suited to handle, or whether a default model is being applied too broadly.

Understand the full cost of a task

Some tasks are finished in one exchange. Others take a few rounds of questions, corrections, and follow-ups. Turns analysis shows how much back-and-forth different tasks require. A long conversation is not necessarily a bad thing, especially for complex work. But if a simple task keeps taking several turns, it may be worth looking at the prompt, the model, or the workflow.

The first request is only part of the cost. Teams should also look at the time, tokens, and money spent before the task is finished. Comparing those numbers can show where a workflow is taking longer or costing more than expected.

Turn insights into auto routing

Once a team has identified an overkill pattern and confirmed it across task, cost, latency, and turn data, it can turn that insight into an automatic routing decision.

For example, the task view might show that much of the team’s AI usage is summarization and formatting. The model view could show that those requests are being sent to a large reasoning model, while the turns view shows that most conversations finish in a single turn. Together, these signals give the team a concrete workload to evaluate.

In addition to our updates to User Insights, the Auto Router is now available in closed beta. The Auto Router uses the conversation trajectory, task category, task complexity, and model-fit signals to automatically route requests to an appropriate model while taking cost into account.

Instead of creating a separate routing rule for every workload, customers in the beta can let the Auto Router select among the models available to their application. The router does not simply send every request to the least expensive model, but instead selects an appropriate model for the task at hand. Complex coding or research work may still need a more capable model, while simpler tasks may be handled by a faster or less expensive option.

To learn more about the Auto Router and sign up for the closed beta, read the blog post here.

The Auto Router uses the same task and conversation signals that power User Insights. The section below explains how those signals are produced.

How User Insights classifies traffic

Each conversation receives an analysis signal that can be grouped in User Insights. The signal is used for reporting and routing analysis, and is not intended to replace or expose the original request.

The categorization engine is a dedicated Cloudflare Worker that processes eligible AI Gateway logs. It examines the conversation trajectory, including user requests, assistant responses, tool calls, and tool results, and identifies the type of work being performed, such as coding, debugging, research, or summarization. It also returns a confidence score and evaluates dimensions such as task complexity, intent ambiguity, stakes, and context dependence.

The Worker returns a category that can be joined with the log metadata used by the dashboard. These signals can also be used to evaluate model fit by comparing how well candidate models suit the task against their cost. The current implementation focuses on a small set of categories that are easy to understand, rather than trying to infer every detail about a user’s work.

The pipeline follows the existing AI Gateway log architecture. Metadata is stored separately from log bodies, and the current implementation uses Durable Objects for metadata and R2 for log bodies. User Insights exposes derived categories and aggregate views. It does not turn the dashboard into a raw prompt browser. Retention of the underlying log bodies continues to follow the configured AI Gateway logging behavior, so teams should review those settings when deciding what to send through the classifier.

The classification is asynchronous, which means it happens after AI Gateway has handled the request rather than while the user is waiting for a response. AI Gateway writes the log to the existing storage path first, and the classification Worker processes it afterward. This keeps classification out of the request path and adds no latency to the user’s response.

The tradeoff is that User Insights is not a real-time view. Newly received conversations may not appear in the dashboard immediately, and analysis may trail incoming traffic by approximately one day as logs are processed and aggregated. Teams should use User Insights to identify usage patterns over time rather than monitor live request activity.

The flow looks like this:

Connect usage to users, teams, and tools

Task categories become more useful when they can be viewed by user, team, or application. AI Gateway is identity-aware, providing that context without requiring teams to build a separate reporting pipeline.

This works not only for applications that teams build themselves, but also for developer tools and agent harnesses such as Claude Code, Codex, and OpenCode. By putting AI Gateway behind Cloudflare Access, teams can connect authenticated users and sessions to their AI traffic, allowing User Insights to associate activity with the right person and conversation.

For custom applications, requests must include both a stable user_id and a session_id for User Insights analysis. The exact identity configuration and field names depend on how the application or tool is set up. The important part is to provide stable, non-sensitive user and session identifiers so usage can be grouped without putting identity data in the prompt itself.

For custom applications, the request metadata might look like this:

The request body contains the model and messages for the conversation.

With Access configured in front of AI Gateway, tools such as Claude Code, Codex, and OpenCode can inherit this identity context automatically. Cloudflare Access is available at no cost for teams with up to 50 users, making it an easy way to get started.

Get started with AI Gateway User Insights

AI usage is changing quickly. Models change, teams develop new workflows, and the right choice for one group may be the wrong choice for another.

User Insights lets teams start making smarter choices by identifying where certain models may be overkill. They can then see which users and agents are driving that usage, understand the tasks behind it, and compare the cost of completing the work.

Learn more with the AI Gateway User Insights documentation . Open AI Gateway in the Cloudflare dashboard, and use what you learn to make more targeted model and routing decisions.

Detect and send production issues straight to your agent

Post Syndicated from Thomas Ankcorn original https://blog.cloudflare.com/real-time-issue-detection/

As agents help us build more complex applications, both humans and agents need a better way to stay on top of what goes wrong in production. Coding agents can already query observability data, navigate a repository, change code, write tests, and open a pull request. What remains manual is connecting those steps: recognizing that repeated failures come from the same bug, gathering the relevant logs and traces, sending that context to an agent, and checking whether the fix worked. Without that structured handoff, the agent must search raw telemetry to reconstruct the scope and context of the failure before it can investigate.

Today, we are introducing Issues, built-in error monitoring for Cloudflare Workers (now in open beta!) to streamline this workflow. Issues can:

  • Group repeated exceptions, 5xx responses, and error logs into one issue.
  • Send the error, stack trace, logs, traces, and Worker version to a configured coding agent.
  • Trigger the agent’s configured workflow — from triaging an issue to querying more data to opening a pull request.

Fix your first issue with the following prompt for your agent with CF CLI or checkout the documentation to get started:

Catch failures automatically

With one line of configuration, you can start receiving Issues detected on your Worker with no additional instrumentation required. Issues are built into the Workers runtime, so there is no SDK to install or application wrapper to add. 

Once enabled, Issues records uncaught exceptions, failed invocations, HTTP 5xx responses, output from console.log() and console.error(), and logs that contain a stack trace. It also flags runaway alarm conditions and code that writes large volumes of logs inside loops.

Consider a Worker whose handler starts throwing errors after a deployment. Every failed request has a different request ID, but they all come from the same bug. Issues groups them together and shows when the error first appeared, how many times it has happened, and whether it is becoming more frequent.

When you open an issue you can see the error, a stack trace when available, the logs and traces leading up to it, the Worker version, request details and trend of the issue over time, as shown below:  

Contextualizing errors for your agent

Cloudflare can capture what happened inside the Worker, but it does not know which users, accounts, or sessions matter to your application. Use the Worker runtime's built-in OpenTelemetry API to add those identifiers, without installing another package:

Those identifiers appear with each occurrence. In this issue, you can now see whether failures are concentrated in one account or session before sending the issue to an agent.

Send detected issues to your agent

An issue no longer has to sit in a dashboard while someone copies a stack trace and pastes it into a prompt. Configure an automation once, and when an issue crosses an occurrence threshold or returns after a quiet period, Issues sends it straight to your agent through Automations. You can choose when the automation should run and where the issue should go.

This can be via:

  • Built-in coding agents: Connect Claude Code with a routine ID and token, Cursor with an automation webhook URL, or Devin with an API token and organization ID.
  • Generic webhooks: Send issue context to your own agent or HTTPS endpoint.
  • Chat and incident management: Notify your team through chat or an on-call workflow.

When the automation runs, Issues sends the failure summary and diagnostic context captured with the issue — the exception, error, source-mapped stack trace, leading and trailing logs and traces, Worker version, and the application context you added. For deeper investigation, you can connect the agent separately to Cloudflare MCP that lets the agent query the related logs and traces so that it can propose code and test changes and open a pull request.

You stay in control of what reaches production: review the pull request, deploy the fix, and mark the issue resolved.

How Issues uncovered and resolved two Workflows bugs in one day

Cloudflare Workflows, a primitive that powers long-running, multi-step applications, is built entirely on the Workers platform.

Behind the scenes, its services keep track of steps, retries, and saved state. This makes Workflows a useful place to test Issues on our own production systems. Within a day of turning it on, the team found two unusual problems hidden inside a large volume of traffic.

  • A migration stuck in a retry loop: A Workflows control plane migration repeatedly hit a SQLite foreign key error when attempting to apply migrations in an edge case. Issues allowed the Workflows team to identify the problem and fix it.
  • A deletion process that never completed: Workflows discovered that during deletion of Workflow instances, there was an edge case where they could exceed a Workers subrequest limit and not finish the deletion. Issues helped the Workflows team identify the issue and do a fix.

Instead of leaving the team to connect thousands of separate pieces of telemetry and user reports, their automation setup sent these issues directly to Cloudflare OS, which followed the errors into the Workflows code and proposed a fix for both issues.

Get started

Ready to see what Issues finds in your app? To get started:

  1. Set observability.issues.enabled to true in your wrangler.jsonc file
  2. Set up your first automation in the Cloudflare dashboard to send issues to your destination of choice whether that’s an agent, a webhook, incident management tool or chat platform.

If your agent is handling setup, it can also use the new cf CLI to inspect issues and create automations. Check out our documentation to learn more!

Cloudflare Containers, rebuilt to scale agent sandboxes

Post Syndicated from Thomas Gauvin original https://blog.cloudflare.com/faster-agent-sandboxes/

Today, we’re making Cloudflare Containers more programmable and optimized for agent workloads. Agents don't deploy sandboxes ahead of time. They create sandboxes on demand, for each task, expect them to be ready immediately, and be able to pause and resume. So we rearchitected Containers to meet these requirements: your code can now choose each sandbox's image and instance type at runtime, Containers start 6x faster, and filesystem snapshots are available in public beta.

To make this possible, we’ve rethought the Containers infrastructure from the bottom up. A new scheduling policy moves control over each sandbox into application code, while a redesigned runtime provides a faster path to a running Container. In ComputeSDK’s independent benchmark, median startup fell from just over four seconds to 648 milliseconds, and, in our own preliminary tests, burst testing successfully created hundreds of thousands of containers in seconds.

All of this builds on what has always set Containers on Cloudflare apart: every Container gets its own Durable Object, a persistent, programmable controller running right next to it that manages its lifecycle, outbound traffic, and more. We are bringing more capabilities directly to the native ctx.container API, so the Durable Object can control its Container without a wrapper class in between, and we’re carrying this model into Sandbox SDK 1.0.

As we wrote earlier this year, your agent needs a computer. These changes make Containers a better complement to Workers, Dynamic Workers, and Durable Objects when agents need a full Linux workspace.

Rethinking Containers’ runtime for agents

Until now, Cloudflare Containers has been organized around application deployments. You choose an image and compute resources at deploy time, roll that configuration out across the application, and manage it centrally. The application is the unit of configuration and rollout.

An agent workspace is different: it’s created on demand, while the agent is working. The task determines its image, resources, tools, and starting filesystem. It might exist for a few minutes, sleep between requests, or be restored days later. Those decisions need to live with the application code handling the task, and the agent's sandbox needs to start up fast, because every second of startup is time your users spend waiting.

We’ve seen this pattern with Base44 on app-building workspaces and Kilo Code on cloud-agent sessions. It also appears in our integrations with Cursor Cloud Agents, Devin Outposts, the OpenAI Agents API, and Claude Managed Agents.

Each of these workloads needs something different from the workspace. Coding agents need repositories, package managers, compilers, test runners, and development servers. Evals need sandboxes that begin from a known state. Reinforcement learning systems need to create, grade, and reset large numbers of environments. Longer-running tasks need to preserve the files an agent produces, so work can continue later.

These requirements led us to fundamentally rethink how Cloudflare Containers are configured, scheduled, and saved. The result is a new way to provision and schedule Containers: the durable_object scheduling policy. It lets your code choose each sandbox’s image and compute resources at runtime, starts Containers more than 6x faster, and supports filesystem snapshots, so workspaces can be saved and restored.

“At Base44, we help anyone turn an idea into a working app. Cloudflare Containers gives each app an isolated development environment where our AI can execute commands, install dependencies, and bring changes to life in a live preview.”

Dolev Epshtein, Software Engineer, App Infrastructure at Base44

“At Kilo Code, every cloud-agent session needs its own workspace and environment, with the right repository, tools, and user configuration. Cloudflare Containers lets us create those isolated environments on demand, so our agents can start running commands quickly and get to work for our customers.”

Emilie Schario, Co-founder of Kilo Code and VP Engineering, AI Workspaces at Anaconda

Choose the sandbox your agent needs, from code

From the start, every Cloudflare Container instance has been attached to a Durable Object. The Durable Object gives the environment a stable identity and lets application code control when it starts, sleeps, and stops. This model has proven particularly well-suited to agent sandboxes.

Until now, though, the two decisions that matter most for an agent sandbox were locked in at deploy time: which image it runs and how much compute it gets. Each combination of image and instance type was its own Containers application, with its own Durable Object namespace, set up ahead of time with wrangler deploy. 

Say one agent needs a small Node.js sandbox and another needs a large Python sandbox for builds. With our previous approach, that required two applications, two namespaces, and routing logic in your Worker to send each task to the right one. Every new environment meant another deployment. 

The new durable_object scheduling policy removes that. The image and instance type are now arguments your code passes when the sandbox starts. To opt in, set the scheduling policy and declare the images your Durable Object can choose from:

Each image you declare is available as this.ctx.container.images.<name> within the Durable Object. When a task arrives, your code picks the image and instance type for that task:

This code makes two decisions after the task is known. It chooses the toolchain the workspace needs and how much compute to give it. What used to take a separate application and a separate wrangler deploy is now an if statement. One Durable Object class can start Node.js and Python sandboxes of different sizes side by side, and adding a new environment is a code change, not a new deployment. That’s the idea behind this whole update: infrastructure becomes code that runs at request time, right down to the environment itself.

Rollouts are now just code

Choosing the image at start time also fixes one of the most painful parts of running Containers: rollouts.

Before, updating an image meant updating the whole application. You set grace periods, so running instances could drain, defined percentage splits to move traffic gradually, and called our API to push the new configuration. Throughout that process, the platform decided which instances got replaced and when, whether an agent was in the middle of a task.

With the durable_object scheduling policy, there's no rollout configuration at all. A Container can keep running the image it started with until your code stops it. The next time that Durable Object starts a Container, it uses whatever image your code chooses. That means any rollout strategy you want is a few lines of code:

For example, you can:

  • Canary a new toolchain on 5% of new sandboxes by hashing the Durable Object ID.
  • Pin active projects to their current image, so an agent never has its environment swapped out mid-task.
  • Migrate a workspace at a natural checkpoint, like the next session or after a snapshot.
  • Roll back by changing which image future starts choose. No config push, no waiting for a drain.

The rollout policy lives in your Durable Object code, next to the rest of your logic, and it can be as simple or as sophisticated as you need.

Each of these improvements comes from leaning further into the Durable Object, which already owns the workspace's identity, state, and lifecycle. Letting it choose and control its Container gives you more control over every instance and its rollout. It also gives the scheduler a better place to start the Container: wherever the Durable Object is already running. That's what gets the agent to its first command faster.

Faster first commands

Previously, starting a Container required our global control plane to resolve the application configuration, find capacity, and coordinate placement. That model works well for application-wide fleets, but it put deployment machinery in the path of an agent’s first command.

With the durable_object scheduling policy, demand begins at the Durable Object. The Containers infrastructure serving it looks for capacity on the same machine first, then widens the search within the same location if it needs to. It also favors hosts that already have the Container’s image or snapshot in local storage, so the Container can start without downloading it first. 

We also cut work after the Container reaches a host. Instead of booting a new virtual machine from scratch, the new runtime restores a prepared virtual machine that isn’t yet assigned. It reuses networking and filesystem setup, batches repeated operations, and no longer waits on services the first command doesn’t need. 

Together, these changes substantially reduce the time it takes to go from creating a sandbox to running a command in it. On ComputeSDK’s independent Burst TTI Benchmark, which launches 100 sandboxes concurrently and measures time-to-interactive from the client:

Startup measurement

Previous scheduling path

New scheduling policy

Improvement

Median

4.049 seconds

648 milliseconds

6.2x faster

95th percentile

5.839 seconds

910 milliseconds

6.4x faster

99th percentile

6.717 seconds

1129 milliseconds

5.9x faster

The new path also holds up under burst load. In our preliminary burst test, a single account started 100,000 Containers in 5.387 seconds across six locations.

Start with a prepared system image

As the scheduling path gets faster, preparing the image becomes a larger part of the remaining wait. Before a Container can start, its image has to be on the host and unpacked into a filesystem. When that work happens after the request arrives, the agent is left waiting.

That’s why we are introducing cloudflare/debian-trixie: a ready-to-use system image for agents that can configure their environment at runtime, containing Debian Trixie Slim and Node.js 24.20.0 LTS:

This means your agent can start a Linux sandbox without first creating a Dockerfile, building an image, or pushing it to Cloudflare. Once the sandbox is running, your agent can use exec() to clone a repository, install packages, and configure the environment for its task.

And because Cloudflare controls this image, we can distribute and prepare it across eligible Containers hosts before requests arrive. Startups don’t have to download or unpack the base image while the user waits.

Save the workspace and return to it later

A fast start still leaves one more wait: setting up the workspace. Cloning a repository, installing dependencies, and configuring a toolchain can take much longer than starting the Container itself. As the agent works, it also produces files you want to keep. Repeating setup on every start wastes time, and losing the agent’s changes makes it difficult to continue a task.

That is why we’re adding native filesystem snapshots to Containers, in public beta. Snapshots let an agent begin a task in a prepared environment, save its workspace when the task pauses, and restore those files when the session resumes.

Snapshots enable two useful patterns.

First, one workspace can continue across many sessions. For a coding agent, that might mean saving the workspace when the user finishes working and restoring it when they return the next day. The repository, installed dependencies, build caches, configuration, and edits are available without rebuilding the environment.

Second, a snapshot can provide a shared checkpoint for many sandboxes. Since snapshots are immutable and reusable, multiple Containers can start independently of the same prepared environment and make their own changes from there.

Agent evaluations are a good example. An eval might run the same task across different system prompts, skills, models, or agent versions. To compare the results, everything else has to stay fixed: the repository, dependencies, tools, and input files. One snapshot can start many isolated environments from the same baseline, reducing setup time and preventing environment drift from affecting the results.

Snapshots also complement the new system image we introduced above. An agent can start from cloudflare/debian-trixie, set up its environment with exec(), and save the result as a snapshot. Future sandboxes then start from that snapshot with the repository, dependencies, and toolchain already in place.

With snapshots available through the new durable_object scheduling policy, Containers can act as persistent agent workspaces. Compute can stop when work pauses, and a new Container can start from the latest snapshot to pick up where it left off.

The Durable Object advantage for agent sandboxes

The faster scheduling path, runtime image selection, and filesystem snapshots all come from the same design choice: leaning further into the Durable Object as the controller for its attached Container.

Agent systems need a programmable, stateful environment outside the Container to keep state, hold credentials, and control the sandbox’s lifecycle. You can run the agent there, following the “decoupling the brain from the hands” pattern described by Anthropic: when the agent is separate from the sandbox where it works, the agent stays available while its sandboxes and tools can start, stop, fail, or be replaced independently. Or, if you run the agent inside the sandbox, the outside environment lets you supervise it and report progress back to the user.

This is where the Durable Object and Container architecture becomes uniquely well-suited. Every Container is attached to a stateful Durable Object with its own compute and storage running right next to it. You can run the agent in the Durable Object and use the Container as its workspace, or run the agent in the Container and use the Durable Object to supervise it. Add Dynamic Workers for lightweight isolated execution, and an application can choose the execution environment each task requires.

What’s new with the durable_object scheduling policy is that the Durable Object can now control its Container directly, with no wrapper class in between. exec() runs directly in the Workers runtime, and outbound request interception, runtime image and instance selection, and filesystem snapshots are all available on ctx.container. You can combine them with Durable Object storage, alarms, WebSockets, RPC and the rest of your code.

This makes the Container a compute extension of the Durable Object. The Container supplies the Linux environment, while your Durable Object retains the sandbox’s identity, state, policy, and lifecycle. That split is especially well-suited for several patterns:

An agent can remain available while its Linux workspace sleeps. The agent loop can run in the Durable Object, where it maintains session state, communicates with the user over WebSockets, and calls models. It can wake the Container when it needs a shell, compiler, or development server, then stop that compute while it waits for the user or model, paying nothing for idle Linux compute.

The Durable Object can program the security boundary around its Container. It can remember which services, repositories, and operations a user has authorized, then update the Container’s Outbound Request Handler to inject newly granted credentials, enforce new policies, or record additional activity. This resembles the Gatekeeper pattern used by Cloudflare OS, applied to each agent computer.

Evals and reinforcement learning systems can supervise each attempt from outside the environment being tested. A coordinator snapshots a base workspace and forks it into N attempts, each with its own Durable Object and a Container. Each Durable Object runs its attempt, monitors the run, and preserves the result, even if the Container crashes. The coordinator grades the attempts, snapshots the best one, and forks again from there. 

These patterns don’t fit one generic lifecycle. Native APIs let you combine the Container with the Durable Object primitives your application needs, while still using higher-level utilities where they help. You keep control over the Container’s lifecycle, policy, and state.

What this means for the Container class and Sandbox SDK

When we launched Containers, we deliberately hid the Durable Object behind the Container class. We wanted sandboxes to feel familiar and match what developers expected from other platforms. The Sandbox SDK was built on that class, and it filled real gaps: back then, the runtime had no native command execution, outbound request interception, or snapshots, so we built them in userspace.

Since then, agent workspaces have become one of the main workloads shaping Containers, and the cost of that abstraction has become clear. By hiding the Durable Object, we made it harder for you to see and combine the identity, state, and coordination it provides with the Container it controls. Nearly every team we worked with needed something slightly different from the generic lifecycle: their own sleep policy, their own credential handling, their own way of tracking eval runs.

These capabilities are now native, so we're making the Durable Object explicit in the developer experience:

  • New capabilities are native-only. The durable_object scheduling policy, faster startup, runtime image and instance selection, and filesystem snapshots are available only through ctx.container.
  • We'll maintain the Container class and legacy Sandbox class through December 31, 2026. Existing deployments keep running after that date, but the classes won't get updates. We recommend migrating to ctx.container.
  • Sandbox SDK 1.0 is a set of utilities, not a base class. Its helpers work inside your own Durable Object class, alongside ctx.container.
  • For a higher-level environment, @cloudflare/computer combines Dynamic Workers and Containers with a synchronized filesystem.

Migrating usually means changing extends Container to extends DurableObject and calling this.ctx.container directly. See the migration guide for details. Or, get started with your agents:

Get started

Today, most agents are measured by what they can accomplish in a single session. As agents take responsibility for projects that unfold across hours, days, and weeks, the environments where they work need to keep up.

We want every agent to be able to create the sandbox for the task at hand, release the compute when work pauses, and return to the same workspace when the project continues. Today’s changes bring us closer to sandboxes that are as programmable, persistent, and ready to work as the agents using them.

Try the new durable_object scheduling policy, available to all today in public beta, and see what faster startup, filesystem snapshots, and runtime configuration unlock for your agents:

Acknowledgements: This project was also made possible by the contributions of Greg Anders, Andrew Martinez, Nafeez Nazer, Kian Newman-Hazel, Sebastien Pahl, Naresh Ramesh, Cody Roseborough, Nikita Sharma, and Sarah Snell.

Using AI to chart a course for our post-quantum migration

Post Syndicated from Sharon Goldberg original https://blog.cloudflare.com/ai-driven-cryptography-discovery/

As laboratories around the world race to build out a cryptographically relevant quantum computer, we at Cloudflare are racing towards a 2029 target deadline for full post-quantum readiness. While we’ve already transitioned many of our products to post-quantum encryption, we still have work to do to support post-quantum authentication and achieve full post-quantum readiness across our platform.

We’re taking a maximalist stance (“PQ everything!”), because as an infrastructure provider to the world, we want to give our customers the peace of mind that using Cloudflare ensures that their traffic is future-proofed against quantum adversaries.

But how does one accomplish such a massive migration at an organization of our size and scale? After all, cryptography is the base layer for almost all of the world’s digital systems, including the software services and the networking protocols that power our platform.

To drive our PQ migration, we have three key goals.

First, we want to help our product and engineering teams understand how cryptography is being used and how they should be upgrading it. This should cover both the upgrades to post-quantum encryption and to post-quantum authentication. Many of our products have already been upgraded to post-quantum encryption over TLS 1.3, but we still want to cover the long tail of TLS connections, as well as upgrade any other uses of public-key encryption. Meanwhile, it’s still early days for our deployment of post-quantum authentication.

Next, we want to provide progress metrics for the migration. These might include per-repository and per-product counts of the use of classical and post-quantum cryptography.

Finally, we want to surface prerequisites early. If our products or platform rely on protocols that don’t yet have a PQ migration plan (because PQ variants of the system have not yet been considered, because PQ standards do not exist or lack consensus, or because software libraries or other key ecosystem components do not yet have PQ support), then we need to know now. That way we can work with the relevant stakeholders, standards bodies and ecosystems to help drive their PQ migration plans, so that we can meet our own 2029 PQ migration timeline.

This post is the story of how we’re going about this. We explain how we turned to AI to help us solve some of our problems and how we’re developing an internal tool called CryptoLabe to help us. CryptoLabe is named after the mariner’s astrolabe, a navigation instrument refined by Portuguese navigators. Just as an astrolabe helped sailors determine where they were and chart a course, CryptoLabe helps us discover cryptography in our code, understand how it is used, and chart a path to post-quantum migration.

CryptoLabe is highly specialized to our internal systems (our repositories, our ticketing systems, and internal documentation processes) and still evolving as we continue its development, so we aren’t making it available to customers. Nevertheless, we are sharing our learnings so that other organizations can build upon our efforts as they work through their own PQ migration journey.

The scale of the problem

The software that powers most Cloudflare products lives inside our single centralized source control management platform. This means we can find most uses of cryptography across our platform by just looking through our codebase.

While the centralization of our codebase is a marked advantage for us, we still need to contend with three challenges that come with the scale of this problem. First, our code is spread across many repositories. Second, cryptography rarely announces itself plainly in the code. Instead, it hides in

  • shared libraries that a repository imports but may or may not actually call
  • upstream and protocol defaults, like a TLS 1.3 listener that is configured to negotiate a classical key exchange such as X25519 rather than post-quantum X25519MLKEM768
  • configuration files that select algorithms far away from the code that uses them, like a TLS responder whose key exchange protocols are pinned in a YAML file stored in a different repository
  • code paths that are dead, test-only, or on a path to being deprecated

Third, cryptography discovery is about more than just pattern matching. Grepping for certain algorithm names (e.g. “RSA” or “X25519”) overcounts, because it finds cryptography in unused code. Grepping also undercounts, because it misses defaults and indirect uses in dependencies and configuration. Most importantly, it can't tell you how the cryptography is used. A classical ECDSA signature could be part of a JWT, IPsec, TLS, or SSH, and each has a completely different migration path. Many uses also depend on the other side of the connection: a TLS server may support both post-quantum key exchange and classical key exchange; the one it chooses to use would depend on the client.

Turning to AI

It turns out that AI is pretty good at doing more than just grepping. A model can search a codebase, follow evidence across files, and return structured analysis. It can also enrich findings by pulling information from other sources, like our internal documentation and ticketing systems. In fact, AI can even explain how cryptography is being used and how it should be updated. We’ve been putting that idea to the test as we develop CryptoLabe.

As we said before, our first two goals are to (1) discover and understand the use of cryptography in our codebase, and also (2) to get metrics on the state of our PQ migration. Towards these goals, our current implementation of CryptoLabe performs scans in two stages, as shown in the figure below.

The first “discovery” stage starts by mapping the repository. It then searches for cryptography through source, configuration, manifests, lockfiles, scripts, tests, and documentation. Among other things, the scan looks for the use of cryptography like key agreement, signatures, asymmetric encryption, PKI, tokens, credentials, hardware security module integrations, and more. This discovery stage produces a set of "raw observations."

Each raw observation feeds a run of the second stage. This “analysis” stage first re-checks the observation against the source code. It then investigates how the cryptographic operation is used at runtime, what role the repository plays, and which internal or external parties it depends on. When necessary, it can inspect related code in other repositories to complete the analysis. Finally, it takes a pass over its own conclusions, searching for missing or conflicting evidence such as configuration overrides, test-only code, or incorrect assumptions about runtime behavior.

Next, the model assigns a classification to the finding. If there is not enough evidence to assign a classification, the model assigns More evidence needed, External dependency, or Unknown rather than guessing.

This is the current list of classifications used by CryptoLabe, containing catch-all classifiers which will likely be refined as we proceed through our migration. (As an example, we could refine our classifiers by splitting the “encryption” classifier into key agreement and HPKE; you get the idea.)

Classification

Examples

Classical encryption

This is a catch-all category that finds cases of elliptic-curve Diffie-Hellman key exchange (ECDHE) (e.g., X25519, P-256, P-384), RSA key agreement or other uses of public-key encryption (e.g., HPKE). These are broken by a quantum computer running Shor's algorithm, which puts them at risk of harvest-now-decrypt-later attacks.

Classical signature

This is a catch-all category that finds use of an RSA signature or elliptic-curve (ECDSA) signature in anything, for example a certificate, a TLS handshake, another protocol handshake. These signatures are broken by Shor's algorithm.

Classical token

We found a lot of RS256 or ES256 JWT tokens, so we created a special classification for them. These are JWTs that use classical RSA and ECDSA signatures; RFC 9964 defines a post-quantum replacement using ML-DSA.

PQ-ready hybrid key exchange

Finds hybrid post-quantum key exchange in TLS 1.3, i.e. X25519MLKEM768. This is the most prevalent use of PQ encryption in our codebase.

PQ-ready

Finds other uses of post-quantum cryptography that are not X25519MLKEM768 in TLS 1.3, like ML-DSA.

Finally, it generates a report that serves two audiences: (1) product managers who need to understand what the migration means for their product, and (2) engineers that need enough detail to execute the migration.  

Here’s a (cropped) view of one of our reports:

While we’ve been iteratively reviewing findings against the source code and with relevant engineers, we do not yet have a ground-truth dataset for reproducibly comparing different versions of the prompts we’ve tried for CryptoLabe.

Built on Cloudflare’s Developer Platform

We built CryptoLabe on Cloudflare's Developer Platform. Here’s the architecture:

CryptoLabe runs across two Cloudflare Workers. There’s a scanner Worker that runs the scans. And there’s an inventory Worker that serves the dashboard, exposes the API, and stores everything in a D1 database. The two communicate through Service Bindings. A scan starts when someone requests it from the dashboard, and the inventory Worker passes the request to the scanner.

Orchestrating a scan

We need a way to keep a scan alive and on track from start to finish, without building our own job orchestration system. We did this with Agents SDK. Each repository gets its own persistent coordinator built on a Durable Object (DO). A bounded queue in front of the coordinators limits how many scans run at once. When a scan's turn comes, the coordinator tracks its progress and handles cancellation, retries, and recovery.

The coordinator doesn't do the analysis itself. It hands the work to Cloudflare Workflows, so that they can persist progress and automatically retry failed steps. The coordinator moves each repository through four stages:

  1. discovery Workflow (the first scanning stage that produces raw observations)
  2. deep analysis Workflow (the second stage, run on each raw observation)
  3. merge Workflow (that builds a list of findings for a given repository, including combining repeated or similar finds)
  4. publish workflow (that hands results back to the inventory Worker)

The first two workflows need the model to have access to the repository's code. We want this access to be isolated, so we don’t risk damaging the codebase. That’s why CryptoLabe downloads the repository once, at an exact commit, at the start of each scan, and then stores that snapshot in R2. Each Workflow then restores the snapshot into a fresh, short-lived Cloudflare Sandbox, an isolated container. The model then works with the Sandbox through a small set of read-only tools on an immutable snapshot of the code, even if the codebase changes while the scan is still running.

Calling the model at scale

If we want to scan through all of our (many!) repositories, we have to worry about both cost and capacity.

For cost, the model loop sends its requests through AI Gateway to cost-effective open-weight models hosted on Workers AI. Putting the model behind AI Gateway also makes it easy to switch models as better or cheaper ones become available.  

Capacity became a problem once we scanned many repositories at once. Bursts of model requests began triggering HTTP 429 (rate limit) responses from AI Gateway, and scans retrying independently only made the bursts worse. We solved this with a single, global Durable Object that paces every model request across all scans, including retries. When any scan hits a rate limit, the cooldown is shared and all scans back off together, so concurrent scans share the available capacity instead of competing for it.

Prerequisites and hard cases

Let’s now get into our third goal: surfacing prerequisites and hard cases early.

A lot of ink has been spilled about ecosystem readiness for the PQ migration, and we are now going to spill some more. As everyone knows, a PQ migration cannot happen in a vacuum. For migration to succeed, post-quantum cryptography must be supported in relevant software libraries (e.g. BoringSSL) and across parties that participate in the ecosystem (e.g. clients, browsers, origins, cloud proxies, certificate authorities, etc.). Standards are also an important indicator of ecosystem support, although a standard that is still in “draft” state does not necessarily mean deployment cannot proceed. As an example, we deployed X25519MLKEM768 in TLS 1.3 back in 2022 when it was still a “draft” at the Internet Engineering Task Force (IETF) while it was only finalized as RFC 10024 in 2026.

Either way, our point is that in order to upgrade a system to PQ cryptography, we need to understand its dependencies and level of ecosystem support. 

That’s why CryptoLabe uses the concept of “prerequisites” to highlight findings that cannot be immediately remediated by an individual product team working alone.

A prerequisite can be something as straightforward as “we are currently blocked on migrating to post-quantum JWTs.” We say this is straightforward because there is already a standard (RFC 9964) for post-quantum JWTs. Nevertheless, if our software libraries don’t yet support validating post-quantum JWTs, or if we’re using a token issuer that does not yet issue post-quantum JWTs, we can’t go company-wide and ask each of our product teams to start PQ-ing their JWTs. This migration is blocked until we solve its core prerequisites. CryptoLabe lets us group together findings that (likely) have the same prerequisite, which also helps us decide how to prioritize resolving these prerequisites.

For example, the snapshot below shows the six findings from CryptoLabe that have post-quantum SAML as a prerequisite. (SAML is a protocol for single sign-on (SSO).)

On the other hand, there may be uses of cryptography that lack even a basic level of ecosystem support. We’ve been calling these “hard cases.” To find them, we wrote a separate prompt that ignores “vanilla” uses of cryptography (e.g. ordinary TLS between internal systems) and instead looks for custom cryptographic protocols, keys, or signatures used in size-constrained fields, cryptography built into hardware, specialized cryptographic constructions (like blind signatures), protocols without a PQ standard, and dependencies on external parties that do not yet support PQ cryptography.

This prompt is shorter and simpler than those used for CryptoLabe, since its only job is to find hard cases.  In our qualitative review, we found that it got better results when it ran in one fell swoop against all our repositories, while also taking in context from our internal ticketing and documentation system.  

Here’s an example of a “hard case” we found: a certificate carried in an HTTP header. Post-quantum certificates and signatures are larger than their classical counterparts, so if the header (or an intermediary, or the application processing the header) assumes a certificate has a certain size, changing the signature algorithm may break the system. Our next step is to determine whether this code will remain in use in the long term. If it will, we need to measure the relevant size limits and decide how to accommodate the larger certificate.

An important lesson here is that no single scan finds everything. Our repository-by-repository scans were effective at discovering common uses of cryptography. Meanwhile, this targeted scan worked better for “hard cases” because it ignored well-understood cryptography and had more context about each product and its dependencies.

The bottom line is that different approaches find different things, and every finding still needs to be checked by the engineers who understand how the system actually works.

Sharing our prompts

We’ve been messing around with the best way to write prompts for CryptoLabe for the last several months.  We don’t yet have a ground-truth dataset for comparing one prompt’s performance against another, and we are not convinced we have 100% coverage of all uses of cryptography in our codebase. Instead, we have iterated by running scans, reviewing findings with the engineers that maintain the repositories, investigating misses that came up during these reviews and revising the prompts.   Nevertheless, we decided to publish selected prompts, so other teams can learn from and adapt our approach. These prompts are starting points, not a standalone version of CryptoLabe, and the quality of their results will depend on the model, tools, context, and engineering review available.

Thinking through your own PQ migration

At Cloudflare, we’re taking a maximalist approach to our PQ migration because of our goal of acting as a provider of post-quantum cryptography for customers and the Internet at large. But most organizations do not need to start by finding every use of cryptography in every repository in every one of their products. In fact, most organizations should not be doing this, because at this time it's a waste of precious resources.

Before scanning a single repository, you can protect traffic in bulk wherever possible. If your websites run through Cloudflare, we protect your data in transit with post-quantum encryption already today; check this out with our new PQ visibility features. Our SASE platform, Cloudflare One, provides post-quantum encryption for private network traffic. Post-quantum encryption is provided at no additional cost and without requiring you to upgrade every origin server or private application on your enterprise network. This gives you a compensating control while you work through discovering and understanding the use of cryptography inside your own systems.

An exhaustive cryptographic inventory is not a prerequisite for action. Instead, organizations should first identify the systems whose compromise would matter most, discover their use of cryptography, and then PQ that cryptography in priority order. Here is one way to begin:

  1. Choose a repository for one important system. Start with something that handles sensitive or long-lived data, authenticates users or software, or is exposed to the public Internet.
  2. Run cryptography discovery against that repository. We hope our description of CryptoLabe will be helpful to this effort!
  3. Validate the results. Ask the team who owns the system to validate the results of cryptography discovery and confirm that the cryptography finding is needed long term and needs to be upgraded to PQ. It’s important to remember that it might not need to be immediately upgraded to PQ if there is another compensating control in place.
  4. Prioritize action. Figure out what upgrades you can make now and what upgrades are blocked. Record shared prerequisites that need help from a library, vendor, standards group, or another part of your organization. Prioritize your findings and make a plan for addressing the highest-impact systems and prerequisites first.

That gives you the beginning of a PQ transition plan, without requiring a complete map of every cryptographic operation in your organization. CryptoLabe is still ever-evolving, but its scans and results have been illuminating to us as we plan our migration. We hope these shared learnings will be useful as you continue to work through your own PQ migration.

Acknowledgements: Many people across Cloudflare provided feedback on and contributed to CryptoLabe, including Davide Marquês, Peter Wu, Phil Schmieder, JP Aumasson, Andrew Galloni, Christopher Patton, Luke Valenta, Mari Galicer, Vânia Gonçalves, and the Client, Tunnel and Gateway teams who reviewed reports produced by the tool.

Next.js applications, powered by Vite: introducing Vinext 1.0

Post Syndicated from James Anderson original https://blog.cloudflare.com/vinext-nextjs-on-vite/

When we launched Vinext in February, it was the result of an audacious week-long AI-driven experiment to see how far one engineer, and a stack of tokens, could get to replicating the NextJS framework backed by Vite.

In the seven months since that experiment, Vinext has grown into a framework that our customers trust and run in production for high-traffic, dynamic applications.

Today we are announcing the release of Vinext 1.0, the latest step on our journey to make it possible to deploy Next.js apps anywhere. Vinext lets you take any Next.js application, whether it was built for the Pages or App Router, and make it portable to be deployed to any web platform, including the Cloudflare Workers free plan, Netlify, or AWS Lambda.

Vinext 1.0 brings with it sweeping improvements to compatibility, stability, and caching behaviors, and sets the project up for the long term. There’s never been a better time to take your Next.js project and convert it to Vinext; just run npx vinext check and npx vinext init.

Graduation to 1.0

On release Vinext was promising, but it was incomplete. Since then, we’ve spent a lot of time both improving App Router compatibility and expanding that to Pages Router apps — which we’ve learned many customers are longtime fans of, with large applications that are complex to migrate. We didn’t want Vinext to be a tool that only worked for people using the latest App Router features.

Our focus has been on adopting both these routers, and watching our test compatibility closely, which for most important customer-requested features now surpasses 99%.

This improvement has been fueled through the community around our GitHub project. As soon as Vinext launched, that community threw it at a wide variety of applications to find the gaps. With their scrutiny, we found challenges not immediately obvious in the test coverage. Vinext needs to act exactly as Next.js behaves. It is not good enough to imitate functions with the same name. Building an alternative import { revalidatePath } is simple enough; the difficulty is in making sure it correctly affects the rendered pages, cache entry, and future requests.

Tracing requests through the application to make sure Vinext responds in the way expected — and replicating not just the API, but the behavior of this machine — was by far the more challenging aspect.

Once we’ve patched problems and brought new features forward, it’s important that we don’t regress, especially if Next.js makes a change. That’s why we’ve also built out our test suite: thousands of focused tests covering core framework behavior across both routers, the development and production server, and the deployment targets of Nodejs and Cloudflare Workers. We also run the Next.js end-to-end test suite against Vinext nightly, giving us a continually moving window on our compatibility, and making sure we immediately become aware of regressions coming from merged changes. Alongside the automated testing, we’ve been working directly with large customers that have Vinext in production to make sure they are not facing issues.

What’s in 1.0

The clearest messages we got from customers using Vinext is that certain Next.js features carry the framework and Vinext didn’t actually need to do everything that Next.js has launched in recent versions to be incredibly useful to them. So we focused on better support where you need it:

  • App Router, Pages Router, and Hybrid applications: We heard from customers that Pages Router was still important, and migrations are not a one-step process. Vinext therefore has support for both routing paths, including React Server Components, Server Actions, API routes, route handlers, middleware, and client-side navigation.
  • The complete page lifecycle: Pages can be rendered in many different ways: on the server, pre-rendered in the build, exported as static assets, or cached with page-level Incremental Static Regeneration (ISR). We’ve made sure that Background and on-demand revalidation work with any output.
  • Caching: Vinext has a shared set of caching functions across the App and Pages Router and the supported runtimes. We have further support for using Cloudflare’s Workers Cache.
  • Observability: Vinext provides Next.js-compatible tracing across both routers, so existing OpenTelemetry and Sentry setups continue to work. On Cloudflare Workers, traces also integrate with native Workers Observability.
  • Next.js ecosystem compatibility:  Vinext implements the public next/* surface and supports common Next patterns for use of authentication, MDX, image optimization, fonts, metadata, environment variables, and more.
  • First-class runtime support for Workers: While Vinext can run anywhere, server code can run in the Cloudflare workerd runtime during development and production, with direct access to bindings such as image optimization and hyperdrive. 

We’ve also made migration part of the framework: it takes two commands to verify that your Next.js install and any modification you have made is compatible, and set up the Vite and deployment configuration while keeping all your previous Next.js project structure.

When we talked to teams about what features were important for them, something stood out. Next.js 16 took a stance that Cache Components were an important part of the future of the framework, and yet most teams that we talked to were not using them and did not consider support a prerequisite to move. Therefore, Vinext today has limited support for the “use cache” directive that drives Cache Components, and though we will continue to improve compatibility there, we’re much more focused on the core priorities above.

Pre-rendering and cache warming

When we first announced Vinext, it supported Incremental Static Regeneration (ISR) after the first request, but it did not yet render pages during the build. Applications use generateStaticParams() and getStaticPaths() to identify pages that should be rendered when building, and they expect page-level ISR to connect those initial responses to background and on-demand revalidation.

Vinext 1.0 supports that lifecycle for both routers. It can prerender App Router and Pages Router routes during the build, serve those responses through page-level ISR, and invalidate them by path or tag. It also supports output: "export" when the result you want is a fully static site.

But this led us to question something: Why should this rendering happen during the build at all?

A site with tens or hundreds of thousands of possible URLs can spend a seriously long time rendering pages that receive little traffic. The build process cannot evaluate the long tail of traffic that most sites experience and therefore cannot focus compute time on the smaller number of more critical pages. Instead you waste hours of time waiting for sequential builds working their way through thousands of pages, long after the most important routes are done.

Cache warming is our solution to this, moving page prerendering from the build machine to Cloudflare’s network. Developers can continue to use Next.js primitives to identify the pages for prerendering, and Vinext can additionally identify high-traffic pages to add to this list. This happens in the background before your site is deployed to production, so that the moment it is, it is ready to serve rapid responses from the Cloudflare cache.

Inside the deployment process, this works by uploading a new Worker version and deploying it to 0% of production traffic, before then requesting pages specifically from that version. This allows the rendering pipeline to work before any real users hit the new deployment. Once the caches have been populated, the deployment can be promoted safely.

What we’re doing next

If the original experiment invented the one-off slopfork, the more consequential part has been how we can keep that process of self-improvement running indefinitely.

The project is now focused on keeping the framework up to date with everything happening upstream. Next.js canary receives new commits every day. Each morning, an agent reviews the changes, fetches diffs, and opens tracking issues for anything that could affect Vinext. Every night, the compatibility matrix is regenerated as we run the Next.js test suite against Vinext.

When one of these tests or issues reveals a gap, agents are now in the position where they can identify the change across both codebases, build a reproduction, port any relevant tests, and propose a fix.

This review has been catching missing cases, unsafe caching behaviors, and differences in the development vs. production servers.

Automation has helped us narrow the stream of activity into a focused set of changes that deserve attention, allowing the maintainers of the project to focus on only the issues that need context of how a process should map onto Vite from the Next.js implementation.

We’re building a software factory for open source at Cloudflare, and you can see what we’re up to on GitHub.

Try it out

Vinext is available for new applications, and existing Next.js projects.

Start a new application today:

Or migrate an existing application:

And then deploy it to Cloudflare Workers, with our cache warming:

Visit vinext.dev for documentation, examples, and the current compatibility matrix. Vinext is open source at github.com/cloudflare/vinext. Issues, pull requests, application reproductions, and feedback are welcome.

Supporting native Rust in Workers with the new Emscripten target for wasm-bindgen

Post Syndicated from Guy Bedford original https://blog.cloudflare.com/rust-workers-emscripten-target/

Today we’re announcing the first public experimental preview of a feature to better support native Rust code and even Tokio-based applications just running natively on Workers: first-class support for the Emscripten wasm32-unknown-emscripten Rust compiler target on the wasm-bindgen open source toolchain and Cloudflare’s Rust Workers.

wasm-bindgen is the open source toolchain powering Rust-based WebAssembly applications on our V8-based Workers Runtime. Enabling the Emscripten target for wasm-bindgen has been a long-term effort, first initiated by Google over a year ago, and then further reviewed and supported by the Cloudflare engineers maintaining wasm-bindgen.

While still in pre-release, we’re excited to share the new workflow possibilities this work enables in running native wasm-bindgen Rust applications with Emscripten on the web, Node.js, and on Cloudflare’s global Workers platform.

In testing we’ve been able to see significantly improved library compatibility for Rust Workers. To illustrate the sort of capabilities supported, we were able to get a Rust-native Minecraft server (Pumpkin) running inside of a Durable Object with TCP ingress, using real TCP sockets via Tokio. See the end of this post for a full description of this port.

Emscripten is an open-source WebAssembly compiler toolchain initially created by Mozilla and currently maintained by Google engineers, which makes it possible to run and bridge native code with the web platform, including supporting and virtualizing platform features such as timers, file system operations, sockets, and other native functionality.

Since Cloudflare Workers supports Web Platform APIs and Node.js compatibility, we are able to support Emscripten on Workers using its Node.js compilation flags, fully virtualizing native platform features such as timers, file system operations, and sockets on top of our existing Node.js APIs. And with our work on Tokio support, Rust code building on top of Tokio’s async runtime ecosystem can now also be fully integrated into JavaScript-based host environments with this Emscripten target.

We’ve made these current experimental patchsets available with example applications to try out today, including:

Supporting the wasm32-unknown-emscripten target in wasm-bindgen

Mitch Foley on Google’s Portable Toolchains team first encountered the need for Rust WebAssembly toolchains to interoperate with C++ when another internal team at Google was interested in using wasm-bindgen to interface with their JavaScript. The goal was for wasm-bindgen to drive the build and produce the companion JavaScript, while Emscripten’s linker (wasm-ld) would be able to link in any C++ dependencies needed.

Internally, Google uses Emscripten in C++ codebases to generate JavaScript along with the rest of their applications, allowing the C++ and JS to interact with one another. While Emscripten was capable of linking in Rust code as a dependency, the thing it couldn’t do was provide a robust binding system between Rust and JavaScript like wasm-bindgen does.

The problem was both tools assume they are in charge of loading and interacting with JavaScript and generating the final JS and Wasm output. Because of this conflict, Google internal users would have to commit to using one or the other toolchain and never both, and adding a second set of tooling would have doubled the support surface for Google’s Portable Toolchains team.

Mitch and his colleague Yifan Yang crafted a plan to have them work cooperatively: Emscripten would continue to drive the build, load the Wasm module, and provide the companion JS, while wasm-bindgen would produce a smaller portable version of its JavaScript bindings in a format that could be directly included in Emscripten’s library system.

This wasn’t just a technical problem — both Emscripten and the wasm-bindgen maintainers had to support this plan and be willing to maintain integration tests that depended on the other. With feedback from Google’s Portable Toolchains team, Google’s Wasm Tools team, and Cloudflare engineers, this finally resulted in the required changes landing in both projects.

As a result of these efforts, the wasm-bindgen and Emscripten toolchains now have seamless interoperability under the new -sWASM_BINDGEN configuration:

  1. C++ Emscripten code driven by Emscripten’s compiler can be built against static wasm-bindgen Rust code, fully supporting the wasm-bindgen bindings layer alongside the Emscripten bindings layer.
  2. Rust applications using wasm-bindgen and driven by Rust’s compiler can now be built for the Emscripten target, fully supporting Emscripten’s bindings layer alongside wasm-bindgen’s bindings layer.

See the wasm-bindgen Emscripten documentation page for more information about using this target.

Rust library support for Emscripten

After landing the wasm32-unknown-emscripten target support in wasm-bindgen, early prototypes by Cloudflare engineers demonstrated clear success for supporting this target on Cloudflare Workers. Many libraries worked out of the box, even including low-level systems libraries, since Emscripten already supports the target_family = unix in Rust.

Some low-level systems libraries that were unaware of Emscripten required patching, for example libc, socket2, and Mio. Even for these low-level libraries, patches primarily involved adding the Emscripten target to the existing platform gates, for example to explicitly allow target_os = “emscripten”, in place of Wasm platform gates.

Overall we posted these target support patches fairly infrequently, and they were mostly trivial when needed, with the Rust library maintainers very receptive to reviewing the support.

That said, one of the critical foundational libraries that needed significant support was Tokio.

Tokio support

Cloudflare Workers are single-threaded and hosted within a JS event loop, while Tokio async is designed around blocking operations being supported through threaded parking semantics. The two models are clearly not compatible with each other — a blocking operation such as a pending socket read or epoll wait cannot block the shared JS event loop.

To get around this, we had two options: WebAssembly JavaScript Promise Integration or to modify Tokio to support event loop runtime integration.

To support Tokio building for Emscripten, we have contributed full Tokio support patchsets that are being reviewed upstream, with the first target support patch for wasm32-unknown-emscripten already landed upstream in Tokio. With these patchsets, we’ve been able to fully support both approaches in Cloudflare Workers. In our Rust Workers Tokio examples, these patches are required directly pending further integration upstream.

Supporting JSPI

WebAssembly JavaScript Promise Integration (JSPI) maps directly onto Tokio’s existing parking semantics. This is because JSPI allows a blocking Wasm call to suspend the WebAssembly stack on a synchronous operation and return control to the JS event loop, exactly as one would expect of a park.

But when a new Wasm call is made while an earlier one is suspended, JSPI doesn’t break: it instead supports having a new WebAssembly stack being entered while arbitrary existing stacks are suspended at the same time.

The problem here, though, is that Rust itself isn’t aware that its stack is being switched out from under it. Tokio’s runtime context is tracked via a thread local, but a JSPI stack switch is not a thread switch, so the suspended and new stack still share the same thread-local runtime context. As a result, the Tokio runtime context still thinks it is in the parked context when a new Wasm call enters, and then panics because the runtime is already entered.

Supporting fully reentrant JSPI therefore requires careful thread local handling to split context between JSPI context switches. To support this, the thread-local context itself must be swapped on each JSPI enter, exit, suspend, and resume. In effect this is cooperative time-multiplexed threading, with each suspended stack carrying its own runtime context.

With our pre-release Tokio patchset, we were able to implement and verify this model. Finalizing the design and upstreaming it is ongoing in collaboration with the Tokio and Emscripten teams.

Adding an event loop runtime to Tokio

The other runtime approach is full event loop integration. Event loops are of course a common paradigm in native UI applications, so the idea of supporting an event loop runtime for Tokio was certainly not something completely unfamiliar to the maintainers in discussions we had around the Emscripten target support.

The question was rather how to design an event loop for Tokio in such a way that it could work across native applications (including Windows and macOS), as well as for WebAssembly embeddings in JavaScript hosts.

If we could design a general LocalEventLoop runtime for Tokio, we could solve this problem more generally.

A Tokio runtime does two things in a loop: (1) it polls tasks until nothing is ready, and then (2) it waits. The wait is what makes it a runtime rather than a library: the thread parks inside the I/O driver until a socket becomes readable, a timer expires, or another thread wakes it.

But when the host already has its own event loop, Tokio's loop cannot run inside it without blocking the host's. The way around this is to split Tokio's loop in two with both parts able to integrate with the host: (1) becomes an explicit drive() operation that runs one batch of ready tasks and returns, and (2) is replaced by a wake, so that instead of parking, the runtime tells the host it has work and the host calls drive() when it is ready to.

Our proposed LocalEventLoop design for Tokio is a LocalRuntime whose wait has been replaced by a wake. It is built with a standard std::task::Waker that the host owns, which instead of being used to signal that a future should be polled soon, is used by the Tokio runtime to signal that the runtime itself should be driven soon.

Consider the example of a socket read under a regular Tokio runtime:

This would correspond to the following call diagram:

Instead, with LocalEventLoop we can write:

Here, spawn_local returns immediately, with the host event loop taking responsibility for driving the spawned task to completion via el.drive() calls from the host. When stream data is unavailable, the runtime simply returns, handing back control to the host. Once the socket becomes readable, Tokio uses the host_waker to signal that the runtime needs driving.

This LocalEventLoop flow corresponds to the following call diagram:

Everything that would have unparked a native runtime's thread wakes the host instead: a spawn, a task woken from another thread, a socket becoming readable, a timer expiring. Because a Waker is Send + Sync and carries no execution semantics, its implementation is just a notification to the host's event loop. This makes the contract safe by construction: a wake arriving from another thread, from a host callback, or even during a drive, queues work rather than re-entering the runtime. The drive that follows runs on the owning thread with nothing else on the stack.

The one thing you cannot do is wait. block_on still exists and runs a future as far as ready work carries it, but where a normal runtime would park, LocalEventLoop::block_on panics instead. This is because nothing could wake that future from inside the call since its wait belongs to the host.

With this design, the host event loop is never blocked, and interleaves its own work with Tokio's, one batch at a time. The same architecture embeds into a GTK main loop, a Win32 message pump, or a Cocoa run loop. In addition, any number of these runtimes can co-exist due to their cooperative execution semantics.

Supporting sockets and epoll on Emscripten

With both Tokio runtime integration approaches fleshed out for the Rust Emscripten target, we were then able to support most of the Tokio test suite, with one major subsystem still unsupported — the net feature. This includes its sockets APIs: TCP, UDP, and Unix sockets. The reason for this was that Emscripten only supported poll() and a WebSocket emulation layer, but not epoll_wait(), on which Tokio’s I/O driver is built (via mio).

For Cloudflare Workers, we wanted to be able to fully integrate with our TCP sockets APIs, including upcoming inbound TCP. To do this, we would need to build a bridge between Emscripten’s virtualization layer and our own sockets API layer.

Instead of having to build this bridge ourselves, we realized we already have one: the node:net API we support in our Node.js compatibility layer. Emscripten had an -sNODERAWFS mode to bridge natively into Node.js FS APIs, so the same approach could give us a sockets bridge without Emscripten or Cloudflare Workers needing to implement any custom APIs on either side.

We contributed this work in over 40 pull requests to Emscripten, which is now the –sNODERAWSOCKETS layer compilation option, enabling support for epoll, TCP, UDP, and Unix sockets for Emscripten applications in Node.js — and, because Workers implements the same node:net API, on Workers itself as well.

Under JSPI, Emscripten's epoll_wait() simply suspends the stack until readiness, so Tokio's I/O driver works as on native. For LocalEventLoop, readiness instead had to reach the Waker from a JS callback. To support this, we drafted an Emscripten proposal for a new emscripten_epoll_add_listener API to associate a callback on an epoll’s ready events. The next drive() then collects these events with a zero-timeout epoll_wait(), so Tokio's existing I/O driver can be used unchanged.

Running a Minecraft server on Workers

To test out this new target, the challenge was raised to see if a Minecraft server could be deployed to Workers. Dan Lapid then implemented a Pumpkin Minecraft server running on a Durable Object over a single weekend.

Pumpkin is a Rust Minecraft server built on Tokio and designed for multi-core machines. World generation runs on a dedicated thread pool, while the game tick loop and chunk scheduler each run on their own OS threads.

Inside a Durable Object there is exactly one thread, so getting Pumpkin to run there meant turning those threads into cooperative tasks on the event loop. Leaning on the Tokio integration, the tick loop and chunk scheduler became async tasks and each Rayon job became a Tokio task. World generation then runs on the event loop itself, taking one turn per chunk and interleaving with network I/O and game ticks rather than blocking them.

For persistence, Pumpkin writes its world through ordinary std::fs. Under Emscripten’s -sNODERAWFS option file system calls get forwarded to the node:fs Node.js compatibility layer for Workers. Since this bridge is just JavaScript, it is easy to swap out. Dan’s worker-fs-mount library enables mounting a node:fs-compatible file system with its durable-object-fs backend, which stores files as rows in the Durable Object's SQLite storage. With this integrated, every file Pumpkin saves becomes a row in the object's database, written synchronously and committed with the Durable Object's transaction. Pumpkin itself has no idea it isn't on disk, and a restarted object boots straight from the same world.

Networking needed no changes in Pumpkin either. Each player's connection arrives through Workers TCP ingress at the Worker's connect() handler, which forwards it into the Durable Object. There, handleAsNodeConnection() from cloudflare:node dispatches the socket to a net.Server listening on that port inside the object – a new TCP counterpart of handleAsNodeRequest(). Emscripten’s -sNODERAWSOCKETS backend implements TcpListener on top of net.Server, so the server accepts players exactly as it would on Linux, and the returned promise tells the object when the last player has left, so it can save and shut down.

Running a fully persistent Minecraft server inside a Durable Object with multiplayer support demonstrates the level of native compatibility that is possible with this new Emscripten target. We’re excited to see what native Rust applications you can bring to the platform.

Try it out

We’re making these full patchsets and workflows available today for experimental use. Improving wasm-bindgen’s support for modern WebAssembly standards for all users is part of our ongoing commitment to the Rust and Wasm ecosystem.

We welcome all contributions and feedback — find us on GitHub and in the #rust-on-workers channel on Cloudflare’s Discord.

Cloudflare’s 2026 Annual Founders’ Letter

Post Syndicated from Matthew Prince original https://blog.cloudflare.com/cloudflares-2026-annual-founders-letter/

This week Cloudflare celebrates our 16th birthday. Like many 16-year-olds, we find ourselves looking around at the world we grew up in and feeling caught between the past and future. Also like many 16-year-olds, we look at all the change in the world and sometimes see risk. But, on balance, we come out incredibly optimistic for what lies ahead. What drives our optimism and fear is a recognition that change involves disruption. 

The Internet is changing more today than at any point since Cloudflare launched back on September 27, 2010. Some of that change seems undoubtedly good. Some of it is upending the way we think the Internet works.

One significant change is the rate of growth of the web itself. From 2012 through 2025, the web plateaued and by some measures even shrank. That changed in mid-2025 when there was an explosion of new websites. The popular narrative is that growth has been driven by AI "slop." And, while there is some of that, that's not the majority of what we see.

Instead, AI has unleashed a new cohort of creators. Individuals with ideas but without coding skills who, aided by so-called "vibe coding" tools, are able to bring new creations to life. We're proud that the majority of these tools have Cloudflare's developer platform as their preferred deployment target. Technology at its best allows more people to express their creativity. We have front row seats to watch students around the world building apps to solve real-world problems. Startups with new business ideas getting built in record time. Today more than 7 million developers are building the future on Cloudflare’s developer platform.

The last year has also changed in terms of who — and increasingly what — is using the Internet. We originally forecast that automated traffic would pass human traffic in the second half of 2027. The rise of agents and AI crawlers pulled that date forward to May of 2026. And, if current trends continue — which, if anything, seems conservative — automated traffic will be 1,000 times human traffic in just five years. Not because we believe human traffic will decline, but because traffic from agents is exploding.

For people using them, these agents are already astonishing. Ask one to find you a flight, a contractor or a cheaper phone plan, and it will read more pages in a minute than you could in an afternoon, then come back with an answer. It does the legwork so you don't have to.

But that legwork isn’t free. If you ask your AI agent to suggest where you should go to lunch, it may survey 1,000 restaurant menus in the area only to suggest one. The one restaurant may get your business, but the other 999 had to shoulder the load of serving the agent without getting anything in return. The risk here is a tragedy of the commons problem where the users benefiting from the agents and their traffic don't bear the costs of the load they put on the system and therefore don't exercise appropriate restraint.

AI has made it easier than ever for those students and startups to build something. Agents could make it much harder for anyone to find it. Today, small businesses win customers with emotion or convenience. You patronize a certain deli because the person working the counter remembers your name or because it’s part of how you define your identity. Or you shop at a local store even though you know it doesn't have the best selection or prices but it's on your drive home.

Your agent doesn't care who remembers your name, and it isn't driving home past the local store. It goes with whatever it has the most information about, and that's usually whoever has been around the longest. The risk then is that as agents handle more and more commerce, they'll make it harder for new entrants to break in. That, in turn, is likely to lead to consolidation and a less robust business environment.

We were a new entrant once. Sixteen years ago this week, we launched Cloudflare on stage at TechCrunch Disrupt while our engineers sat in the audience fixing bugs. There were eight left when we walked on. By the time we walked off, they were fixed and we were live in five data centers on three continents. Nobody's agent would have recommended us. People took a chance on us anyway. We want the next new entrants to get the same chance.

That’s what we’re playing for. Not a future with five AI companies, but one with 500,000, spread around the world. Not one where content creators wither away and die because they can't be compensated, but one where anyone can create, reach a global audience and get paid for their work. And not one where a few mega corporations win by default, but one where new entrants with better products can serve customers well, and win.

Last year we wrote about what AI was doing to publishers. This year the same shift is reaching the restaurant and the deli. What replaces the Internet's old business model is the most interesting question of the next five years. This week, we take another swing at answering it.

Like every birthday, we're celebrating by giving presents instead of getting them, and some of the presents are for the 999 restaurants. Agents crawl the web the way search engines always have: everything, over and over, whether it's changed or not. Our data suggests more than half of what good bots fetch hasn't changed since their last visit. We've been working so that crawlers see more of the web while fetching only what's new, which cuts the load on the sites they crawl. And we're giving anyone who puts content or applications online a way to get paid when agents use what they've built.

Cloudflare's mission isn't to build a better Internet, it's to help build a better Internet. That means we can't do it alone. So this week we'll also be announcing partnerships with companies and organizations that want the same future we do.

Like other 16-year-olds in this new era of AI, we get to ship our opinions. We're shipping them for an Internet that still has room for a teenager launching their first app, the deli where they remember your name, and whoever is building the next Cloudflare.

And we’ve never been more excited about it.

How Cloudflare addressed a cross-tenant data exposure vulnerability in Containers

Post Syndicated from Rushil Mehra original https://blog.cloudflare.com/containers-cross-tenant-vulnerability/

On September 4, 2026, Oren Yomtov, a security researcher from Accomplish, responsibly reported a vulnerability affecting Cloudflare Containers and Cloudflare Sandboxes (which is built on Containers), through Cloudflare’s bug bounty program. Cloudflare has fully remediated the vulnerability, and we have no evidence that customer data has been compromised. 

This post was prepared in collaboration with Oren Yomtov and the Accomplish security research team, whose detailed report and controlled testing helped us validate the issue and respond quickly.

Cloudflare Containers run workloads on multi-tenant infrastructure and automatically assign them to eligible servers; customers cannot select the underlying host. The researchers demonstrated that a customer with a Workers Paid account could recover residual disk blocks previously used by Containers on the same host. The technique could not target a particular customer, workload, host, or data, and residual data was not guaranteed to be present.

Cloudflare applied a fix across the Containers fleet, with no customer-side configuration changes required. Within the historical disk-I/O telemetry available to us, we identified no evidence of malicious exploitation. Activity we could attribute to the reported technique came from the researchers and Cloudflare engineers conducting authorized validation.

Here, we explain the underlying storage behavior, its potential impact, our investigation, and the actions we took in response.

How container storage allocation works 

Cloudflare Containers use Linux device mapper thin provisioning (dm-thin) to provide each container with a writable root disk. Each container lives inside a dedicated virtual machine powered by the Firecracker virtual machine monitor. Firecracker presents this disk to the virtual machine as /dev/vdc.

Thin provisioning allocates physical storage only when a virtual disk writes to a previously unmapped region. The affected storage pools used a 64 KiB thin-block size. When the thin volume backing a container's root disk was deleted, its physical blocks were returned to a pool that served workloads belonging to multiple customer accounts.

The affected pool configuration included the following option:

skip_block_zeroing

With this option configured, dm-thin skips zeroing newly allocated blocks before making them accessible. Consequently, when a previously-used 64 KiB block was reassigned, a full-block write replaced its previous contents, but a smaller write changed only the written portion. The remainder could retain data from the block’s previous owner.

How the exploit worked

Reading an unmapped region of a new thin disk did not reveal residual data. For an unmapped region of the thin device, dm-thin returned zeroes without allocating a physical block.

The proof of concept identified 64 KiB-aligned regions corresponding to free space in the guest’s ext4 filesystem and wrote one aligned 4 KiB block into each region.

When such a write reached an unmapped thin block, dm-thin allocated a physical 64 KiB block from the shared pool. The 4 KiB write replaced only that portion of the block, and because block zeroing was disabled, the remaining 60 KiB could retain data from a previous container.

A subsequent raw-device read could therefore observe bytes that the new container had never written.

The proof of concept performed the following steps:

  1. Create a container using a Workers Paid account.
  2. Open the writable root disk at /dev/vdc.
  3. Read the disk and record a baseline.
  4. Write one 4 KiB block into each selected 64 KiB region corresponding to ext4 free space.
  5. Read the resulting blocks again.
  6. Examine only the portions not overwritten by the new container.

The submission included counts, block offsets, sizes, checksum results, and truncated hash prefixes. Although the researchers recovered raw blocks to validate the issue, the materials provided to Cloudflare contained no third-party filenames, identifiers, credentials, hostnames, addresses, or recovered content values. As described below, the researchers have also confirmed that they securely deleted the recovered data.

How the vulnerability was validated 

The researchers used ext4 directory block checksums to distinguish blocks belonging to their own test filesystem created for the proof of concept from blocks originating from other filesystems.

When ext4 uses the metadata_csum feature, directory block checksums incorporate values associated with the filesystem and inode. 

Across six production placements, the researchers reported:

  • All 5,614 testable directory blocks.
  • Zero of those blocks were attributed to the researchers’ filesystem. 
  • 2,700 distinct foreign directory inodes identified through checksum analysis.

To validate the method, the researchers tested it against blocks they had deliberately created and deleted in the controlled test filesystem used for the proof of concept. The method correctly attributed all 162 blocks to that filesystem.

The researchers ultimately observed residual material on 18 of 24 placements and 20 of 22 underlying nodes across four continents. The recovered block types included directory structures, database pages, and structurally complete SQLite databases. The researchers reported using scripts that output only aggregate counts and format checks, not recovered file contents. The materials submitted to Cloudflare contained no recovered content values or third-party identifiers. The researchers subsequently confirmed that recovered data under their control remained confidential and was securely deleted following submission, consistent with Cloudflare’s HackerOne disclosure policy.

Impact 

The vulnerability would potentially have allowed for a customer with a Workers Paid account to recover residual data from storage blocks previously used by other customers’ Containers on the same underlying host.

A successful exploitation would have crossed the tenant-isolation boundary and could disclose filesystem metadata, directory structures, database pages, and application data.

However, an attacker could not select a particular victim or access an actively attached disk. Exposure depended on Cloudflare’s workload placement and which previously released blocks dm-thin reassigned. Moreover, the researchers did not demonstrate modification of another customer’s active data or impact to workload availability.

How we mitigated the vulnerability

Our first mitigation was to remove skip_block_zeroing from the dm-thin pool configuration across the fleet. This restored dm-thin’s default behavior of clearing newly allocated blocks before exposing them to a container. It stopped the reported technique, in which a small write triggered allocation and a larger read recovered residual data from the remainder of the block. The researchers independently confirmed that their proof of concept no longer worked after this change.

Zeroing new allocations did not sanitize blocks already mapped into existing thin devices. These mappings existed in running container disks and in each host’s cache of prepared dm-thin snapshots for OCI image layers. A new container could inherit mappings from a cached layer without allocating those blocks again, allowing residual bytes in unused regions, including ext4 free space, to remain readable through raw reads of /dev/vdc.

We therefore also retired all running container disks and removed cached image snapshots created before the mitigation. We drained hosts during off-peak hours, restarted the VMs on each host, and cleared each host's image cache so that disks and cached layers were recreated using zeroed allocations. We have completed this cleanup across the Containers fleet.

No evidence of exploitation

As part of our response, we investigated whether other workloads showed activity consistent with the reported exploitation technique. We reviewed retained historical disk-I/O telemetry from our container infrastructure, using the researchers’ proof of concept and our internal reproduction as reference activity.

The proof of concept produced a characteristic relationship between writes and reads. When a 4 KiB write reached a previously unmapped region, it could trigger allocation of a reused 64 KiB storage block. With zeroing disabled, the remaining 60 KiB could retain data from a previous container. Subsequent reads could therefore recover substantially more data than the new container had overwritten.

Using these characteristics, we developed detection signatures and applied them to the historical telemetry available to us. We identified activity attributable to the researchers and Cloudflare engineers conducting authorized validation, and did not identify additional activity consistent with the reported technique.

We saw no evidence that this specific attack vector was exploited by anyone else.  

Cloudflare customers are protected

As we noted above, Cloudflare has patched this vulnerability and remediation does not require any further action by Cloudflare customers. In addition, we found no evidence of any malicious actor abusing this vulnerability.

Moving quickly with transparency 

We thank Oren Yomtov and the Accomplish security research team for their thorough research, responsible disclosure, and collaboration on this post. We encourage the Cloudflare community to submit any identified vulnerabilities to help us continually improve the security posture of our products and platform.

We also recognize that the trust you place in us is paramount to the success of your infrastructure on Cloudflare. We take these vulnerabilities very seriously and will continue to do everything in our power to mitigate impact. We deeply appreciate your continued support and trust in our platform, and remain committed not only to prioritizing security in all we do, but also acting swiftly and transparently whenever an issue arises.

Timeline

  • September 4, 15:26 UTC: Oren Yomtov from Accomplish reported the issue through HackerOne.
  • September 4, 18:45 UTC: Cloudflare opened a security incident and confirmed the production setup that caused the flaw.
  • September 4, 21:27 UTC: Cloudflare merged the runtime fix and its reuse test.
  • September 4, 22:03 UTC: Cloudflare merged the changes for new and live pools.
  • September 4, 23:15 UTC: Cloudflare started rolling out the changes.
  • September 7, 06:13 UTC: Cloudflare completed rolling out the changes and began clearing old pool data.
  • September 14, 10:50 UTC: The researchers reported that their proof of concept had stopped working.
  • September 14, 12:52 UTC: Cloudflare awarded the researcher a bounty.
  • September 19, 15:03 UTC: Cloudflare completed cleanup of all pre-mitigation cached snapshots across the affected fleet.

Python Workers are now generally available

Post Syndicated from Gyeongjae Choi original https://blog.cloudflare.com/python-workers-ga/

We introduced Python Workers two years ago, providing a way to run Python applications in the Cloudflare Workers runtime. Our goal was to make it as simple to write Workers in Python as it is in TypeScript, and to make the ecosystem of Python packages and frameworks “just work”.

Today, Python Workers are now generally available (GA).

What does GA mean? It means Python is now a first-class, fully supported language on the Cloudflare Developer Platform. You can bring the Python code, libraries, and design patterns you already know and connect them seamlessly to Workers AI, R2, D1, Hyperdrive, Durable Objects, Queues, Workflows, and the rest of the Cloudflare platform. You can also run popular Python frameworks like FastAPI, Django, and Flask inside Python Workers. You can even create a Python Worker inside another Worker using Dynamic Workers.

The journey behind Python Workers

Bringing Python to Cloudflare Workers was a natural choice. Because Workers has supported WebAssembly since 2018, it gave us the perfect environment to run a Wasm-compiled Python interpreter. By using Pyodide, we were able to quickly support a wide range of Python applications in Cloudflare Workers.

Our goal was to create the first platform for infinitely scalable Python apps, while making it as easy and performant as developing Python apps anywhere else.

The features we are highlighting today are the result of this multi-year effort. Many developers are already building applications within Python Workers; today, we are making these capabilities production-ready for everyone.

Python is now a first-class language in the Cloudflare Workers runtime

Python Workers now natively support Cloudflare Developer Platform bindings. Previously, using these Cloudflare bindings in Python Workers required converting Python objects into TypeScript objects explicitly at the RPC boundary. For example, sending a Python dictionary into a Cloudflare Queue required the following glue code to work:

This required Python developers to keep the JavaScript environment and code in mind while writing Python Workers, and it was a common source of error for both humans and AI agents. To address this, we have encapsulated the entire type conversion process within the Workers runtime and the Python SDK. This allows you to utilize all Cloudflare bindings in a Pythonic way without writing a single line of JavaScript code, making the following just work:

Web frameworks: FastAPI, Django, and Flask

You can now run your favorite Python framework, such as FastAPI, Django, or Flask, to build an API server in Python Workers. We implemented a built-in connector that you can use to easily connect your web application to Python Workers.

Let’s say you have a simple FastAPI web application:

In native environments, you would use a web server such as uvicorn to run this application.

In Python Workers, you can run the same application using the workers.asgi package we provide, just by adding this snippet to your code:

Similarly, you can use workers.wsgi package to run synchronous web applications such as Django.

So, what happens under the hood?

Python has a standard contract for how web applications should communicate with web servers, known as the Web Server Gateway Interface (WSGI), or its modern asynchronous counterpart, ASGI. This standard allows developers to build applications that are completely server-agnostic. In a traditional deployment, web servers like Uvicorn or Gunicorn are responsible for handling multiple concurrent client connections and threads to scale traffic, while web frameworks like FastAPI can focus purely on the application logic.

In Cloudflare Workers, the Workers platform itself serves as the web server. Since our global network already seamlessly handles load balancing and infinite scaling, we don't need to reinvent the wheel by running a server inside Python Workers.

Instead, our workers.asgi and workers.wsgi connectors act as a thin, optimized bridge. They translate the incoming native JavaScript request into the standard WSGI/ASGI structures that Python applications expect, and seamlessly pipe the response back out with minimal overhead. By doing this, Python developers get the best of both worlds: you can write and organize code using your favorite web frameworks, while letting the Cloudflare Workers platform instantly scale your API across the globe, without ever configuring a server.

These connectors can be used not only with FastAPI, Django, or Flask, but with any Python web framework that uses the WSGI or ASGI interface.

You can find more information about using each web framework in the Python Workers documentation.

Using PostgreSQL and MySQL with Hyperdrive

If you are building a Python application using relational databases such as PostgreSQL or MySQL, you can now integrate Hyperdrive into Python Workers.

Previously, Python Workers didn’t support TCP sockets, making database drivers unavailable. To understand why this was a blocker, you need to look at how WebAssembly operates. Python database drivers like aiomysql or asyncpg rely on the standard library's socket module to establish connections. In a standard environment, this module makes POSIX system calls to the underlying operating system. Inside a WebAssembly sandbox, those POSIX networking syscalls are normally stubs that always fail. Any attempt to open a standard socket would immediately fail. To solve this problem, we implemented socket system calls using the Workers connect API.

When a database driver attempts to open a TCP connection, it goes through our custom socket syscall implementation. It translates standard Python socket operations like opening a connection and reading bytes into the corresponding JavaScript calls used by the Workers runtime. Because this translation happens at the system call level, your database drivers don't have to know about the underlying implementation at all.

This socket bridge is what makes our Hyperdrive integration possible. To use Hyperdrive in Python Workers, first connect your database with Hyperdrive and set up the binding in the Wrangler config:

Then, connect to Hyperdrive using the database drivers you are familiar with:

You can refer to the Hyperdrive Python Workers documentation to find out how you can use Hyperdrive in Python Workers, and which packages are currently supported.

Expanding the WebAssembly package ecosystem

Because Python Workers run inside a WebAssembly sandbox, any packages with native C/C++/Rust extensions must be cross-compiled to WebAssembly to run in Python Workers. However, previously, there was no standard way to cross-compile any Python packages to WebAssembly. That meant our team had to manually compile and host custom WebAssembly packages. This greatly limited the number of packages you could actually use in Python Workers.

We wanted to fix this and allow users to use a wider variety of packages. However, we didn’t want to merely build packages usable only in Python Workers, which wouldn’t benefit the community. Since Python Workers are built on top of Pyodide, we wanted the ecosystem to evolve in a way that benefits Pyodide and the entire Python-on-WebAssembly community.

To this end, we proposed PEP 783, which standardizes a platform for running Python in the browser runtimes called PyEmscripten. After over a year of discussion and refinement, this proposal was accepted, enabling package maintainers to build and publish packages for the PyEmscripten platform and make them available across all environments that implement PyEmscripten.

We also stabilized the existing Pyodide build toolchain and evolved it into a form that is accessible to all package maintainers, enabling developers to easily build packages for the PyEmscripten platform. Furthermore, we added PyEmscripten platform support to cibuildwheel, to make it easier for others to adopt support for the PyEmscripten platform.

While the ecosystem is still adopting this standard, we hope every Python package will have a wheel that works with WebAssembly in the future. We are also actively working with major package maintainers to add PyEmscripten builds. If you encounter a package that isn’t supported yet, let us know on Discord or GitHub, and our team will work to get it built.

You can also check out our EuroPython 2026 talk: “Python Everywhere: The State of Python on WebAssembly” to see how we made this possible.

Building AI agents and pipelines in Python

The large ecosystem of data science and machine learning packages makes Python the natural choice for building intelligent agents and AI pipelines. But bringing these to Python Workers historically presented a challenge: libraries such as openai and langchain rely on HTTP clients like requests or httpx to communicate with external APIs. However, because of missing low-level socket operations support in Python Workers, these HTTP clients didn’t work properly.

To solve this, we contributed upstream to ensure these HTTP clients can route requests directly through the JavaScript fetch API in WebAssembly environments. Combined with our new support for low-level socket operations as explained in the previous section, this makes the entire networking stack work seamlessly inside Python Workers.

As a result, you can now run AI libraries like openai, langchain, and mcp natively in Python Workers. You can also combine them with Workers AI to run serverless inference on GPUs in Cloudflare’s network, or proxy requests through Cloudflare AI Gateway.

The example below shows a way to run Worker AI models in langchain, using the langchain-cloudflare package:

What you can build today

We have assembled a collection of production-ready patterns in our python-workers-examples repository. Here are some ways you can combine Python Workers with the Cloudflare ecosystem.

Asynchronous AI orchestration

Building a full-stack AI application often means connecting multiple services such as storage, queuing, and inference. This example shows how to build an AI-driven image-to-image generator purely in Python Workers. It accepts user requests, drops them into a Cloudflare Queue, and uses Workflows to orchestrate the image generation step via Workers AI, and stores the image to an R2 bucket.

Real-time stream processing with Bluesky Jetstream

Consuming a firehose of real-time events usually requires a dedicated server to maintain the connection. In this example, we use a Python Worker to connect to the ATProto/Bluesky Jetstream WebSocket. By backing this connection with a Durable Object, the Python Worker can maintain long-lived state, ensuring that the WebSocket connection stays alive.

More examples to explore

Model Context Protocol (MCP) Server

Build and deploy an MCP server using the official Python MCP package to give your AI assistants access to edge data.

Retrieval-Augmented Generation (RAG) system with Vectorize

Building a RAG system using Workers AI and Vectorize, Cloudflare’s vector database.

Python code examples across the Cloudflare developer docs

We’ve updated our docs across Cloudflare products to include Python example code. Nearly everywhere where there is a code example showing how to do something in TypeScript, there’s also a code example in Python. We’re committed to continuing to include Python examples across all of our products. You can toggle code snippets between JavaScript, TypeScript, and Python throughout our developer documentation.

What’s next?

Reaching GA is just the start. We have many plans to make Python Workers better, including making Python Workers more performant and memory efficient, as well as supporting more packages.

Keep telling us what you want to build on Python Workers, and we’ll keep pushing the bounds of what is possible. Check out Python Workers documentation and start building your first Python Worker!

When scanners miss the attack: how Cloudflare Client-Side Security protects storefronts

Post Syndicated from Juan Miguel Cejuela original https://blog.cloudflare.com/client-side-security-finds-4-malicious-campaigns/

A modern storefront can look perfectly healthy while malicious JavaScript works underneath: siphoning affiliate revenue, hijacking searches and clicks, tampering with analytics, or asking a remote server what to execute next. Pages load, products appear, and checkout works — yet the browser may be quietly doing something the site owner never authorized.

That is the blind spot our Client-Side Security machine learning (ML) model is built to expose. This post follows four operations, spanning eight payloads, that our Page Shield ML uncovered in the wild. 

The detection of these malicious payloads was automated; humans verified each finding only after the system had flagged it. When we afterward reviewed the campaigns using security scanning tools, seven of the eight payloads were entirely absent from VirusTotal, and URLScan returned no malicious verdict for any of them. Page Shield ML, meanwhile, caught all eight in live traffic.

For instance, while security research documented the broader Lnkr family years earlier, one specific payload version sat indexed by URLScan for nearly two and a half years with “No classification,” including during a direct scan in January 2024. Only in this case had VirusTotal ingested the payload earlier: while it currently flags the script as malicious, public history does not reveal when that verdict was first assigned. Meanwhile, Page Shield ML independently surfaced those exact bytes live on an online retailer's storefront. More broadly, a hash can be known long before the code behind it is classified as malicious. If your defense waits for that label, you are already late. You need ML that can unravel the JavaScript itself and judge it at scale.

Indeed, seeing a file is not the same as understanding it. The tricky part was that the four operations shared no universal signature or common concealment technique. One remained dormant unless the device, country, time, referrer, or browser state matched what it was waiting for. Another concealed a clickless affiliate request within an invisible iframe. Others intercepted clicks, suppressed monitoring, or conditionally loaded additional code from remote servers. To catch them, you have to watch how those pieces work together: when the script wakes up, what it hides, what it intercepts, and what it fetches next. Checking the page once is not enough; as these cases show, such scripts are built to stay quiet until the right victim shows up. That is why ongoing browser visibility makes the difference between catching an attack and missing it entirely.

How we detect and label JavaScript at scale

The same GNN (graph neural network) that flagged the four operations in this post had already caught malicious npm packages and an in-the-wild Magecart payment skimmer. The GNN does not treat JavaScript as a flat chunk of text; it reasons through the code as a graph: a syntax tree connecting code symbols and exposing what calls what, what the attacker tried to bury, and what still phones home. That structure helps it recognize suspicious patterns across minification, renaming, and some obfuscation without relying on a known URL or byte signature. 

The few scripts that the GNN flags as malicious (under 0.3% of all analyzed traffic) go to a lightweight large language model (LLM) on Workers AI for a live second opinion. This further reduces false positives while keeping recall high. When the LLM corroborates the GNN, customers are alerted.

To investigate the most complex scripts at scale, we use a cohort of frontier models, which we call teachers (an ensemble of automated judges). The cohort draws leading models from around six different families, including open-weight models running on Workers AI. We spin up each as an agent to analyze the same suspicious script in its own fresh, independent session. When useful, their agentic tool access lets them use a restricted JavaScript evaluator to unpack small snippets and reveal concealed behavior. We will soon extend this workflow with Cloudflare Sandbox for deeper analysis in isolated environments.

The frontier models sometimes disagree, especially on the most intricate scripts. We treat that disagreement as signal, not noise. Each label becomes a vote, weighted by the model's score in the Artificial Analysis Intelligence Index, producing a probability distribution over four labels: benign, payment skimming (magecart), other malware, and cryptomining. Human reviewers therefore need only examine scripts flagged as malicious or lacking a clear two-thirds majority. We then feed those label distributions back into GNN training, helping it distinguish ever more nuanced cases. This feedback loop is still partly manual, though we are starting to automate it.

Four malicious JavaScript operations we caught

These four operations do very different things, from commission theft to stolen analytics on shoppers the store already paid to acquire. Stealing a commission is not like skimming a credit card; likewise, hijacking search is not like stealing a password. If an ML model only knows one of those tricks, it will sleep through the others. Instead, our Page Shield ML has to stay attuned to every kind of hostile behavior. 

Now, let’s dig deeper into each operation and how it worked.

Operation 1: The after-hours affiliate-commission hijacker

Picture a quiet Sunday afternoon: a shopper on a phone taps a product. Instead of following the tap normally, the script opens a product or campaign landing page from an attacker-preselected list in a new tab and sends the original tab through an affiliate route. The storefront still appears to work. If the shopper completes a purchase (either then or later), the detour hijacks the attribution, crediting the sale (and any resulting commission) to an account that did not earn the referral.

What the shop lost

The shop could pay an unearned commission to an account that did not bring the shopper. Worse, if a legitimate partner had made the referral, the forced request could misattribute it, diverting credit and a potential payout from the partner who did the work. The damage could outlast one commission: partners who stop trusting the attribution system may also stop trusting the retailer behind it.

Attack chain

Qualified mobile visitor → intercepted product tap → script-selected page opens in new tab + original tab follows attacker’s affiliate route

How it stayed hidden

We found five related script builds: two active and three paused when captured. Each active variant uses a different set of gates before it acts, checking things like the visitor’s device and local time, whether the trick has run recently, whether a product button has appeared, and whether someone actually clicks it. That maze of rules keeps the malicious behavior out of sight during a brief automated visit unless the variant’s specific conditions are met. The active scripts use a MutationObserver (a JavaScript API) to watch for product tiles and buttons that dynamically appear after the page is first loaded. This lets them intercept clicks on those late-arriving elements, while a crawler that loaded the HTML once and stopped there could miss the redirect path entirely.

In the active later variants, the script intercepts a qualifying click and writes a three-day cooldown to localStorage (staying dormant on that device for days). It then executes a dual-tab maneuver: popping an attacker-chosen product page into a fresh tab to keep the shopper engaged, while the original tab takes a quick, unnoticed round-trip through the attacker's affiliate tracking link and back to the shop, to plant the attacker’s attribution cookie in the background. Console masking and self-defending source checks make inspection harder, while the cooldowns and narrow schedules limit how often the malicious path can appear during otherwise normal shopping.

The following sanitized excerpt shows how the payload hooks dynamic product tiles and executes the dual-tab detour. We simplified identifiers, reformatted the code, and neutralized destination URLs for readability.

The paused builds showed how the campaign could go dark without removing the script. Their embedded configuration set status: "paused", so they exited before installing click handlers. These paused scripts carried different per-shopper cooldown configurations (3, 4, and 5 days). One of the paused scripts even recorded a version-history comment explicitly documenting that the campaign was paused after Black Friday.

To reach visitors in the first place, the operation leveraged the site's marketing supply chain: the third-party scripts and tag managers embedded by e-commerce sites to track ad campaigns and analytics. One confirmed delivery path ran through two otherwise ordinary tag managers: Google Tag Manager → another tag manager → malicious script. That is how the payload reached the browser, not proof that either tag manager was compromised.

The attacker even disguised the domain hosting the script to pass a quick marketing review. One delivery host hid in plain sight: adtargett[.]com differed by a single “t” from adtarget[.]com, an advertising domain registered in 1998. The lookalike was registered in 2025 and, when we checked, its homepage called itself “Adtarget.com – Performance Marketing Agency.” This is typosquatting: by mimicking a real ad agency, the host blended in with routine marketing tags, quietly serving the malicious payload that hijacked shopper clicks and redirected them through affiliate payout links.

Operation 2: The clickless affiliate theft

While the first scam still needed a click, this one requires even less. A shopper can open a booking page, linger over the product options, and never touch an ad. In the background, however, the script might have already sent an affiliate request that could make a later sale look as though someone else had referred the shopper. Indeed, when the script’s conditions are met, the payload sends that request through a hidden iframe or a link that clicks itself.

What the shop lost

For the affected tourism business, the attack could corrupt the economics of customer acquisition: a legitimate booking or purchase could be credited to an unearned affiliate account. The code proves covert, automated affiliate requests, but whether any specific request resulted in completed attribution, account crediting, or paid commission in practice remains unobserved.

Attack chain

Time-gated browser → covert affiliate request (off-screen iframe) → 1-hour throttle cookie → when blocked, automated hidden-link click fallback

How it stayed hidden

The script conceals the affiliate request in two layers: selective execution (a pre-flight network gate and hourly schedule), and stealth delivery (an off-screen iframe). The first layer is surprising because its country labels are disconnected from actual geography: neither the shopper’s nor the shop’s location drives the choice.

First, the script calls a public IP-based geolocation service but ignores everything it returns, including the shopper’s country. We could not determine why it required a successful response while ignoring the returned data; this may have been intended to confuse investigators or simply been a remnant of an earlier version. Interestingly, if the geolocation request fails, the script silently stops; its promise chain ends with .catch(() => {}). Although intent is unproven, this fail-closed behavior could help the script evade network-restricted sandboxes.

Next, instead of using the fetched geolocation data, the payload contains three TradeDoubler (an affiliate-marketing network) configuration objects labelled {AU, US, and UK}. These settings blocks are embedded in the code, and each contains an affiliate URL and start and end times. The script computes Asia/Kolkata time in JavaScript, checks those configured time windows, then applies fixed odd/even-hour rules to choose one of the three or else skip the affiliate request for that run. The choice is deterministic. 

Together, the schedule and browser-state checks create time-gated selective execution, a form of cloaking. When those conditions do not line up, the affiliate behavior stays dormant, so a one-off inspection can miss it.

Once the script chooses a configuration, it writes a local cookie named affiliateClicked_<market> as a one-hour retry throttle so it will not re-fire for that region right away (this is a client-side throttle to avoid noise, not an affiliate-network attribution cookie). Next, it loads that affiliate URL in an off-screen iframe with the referrer suppressed. The iframe is the primary delivery path, but it carries an aggressive fallback: if the iframe errors or fails to finish loading after one to two seconds, the script creates a hidden link (<a>) without a target attribute and clicks it programmatically, which could navigate the user's active tab. To the qualifying shopper, nothing seems out of place: they never see an ad, never have to click, and can close the tab as if nothing happened.

As for the script’s obfuscation, it is simple but effective: even property names are assembled one character at a time. The following sanitized excerpt shows the payload creating an invisible off-screen iframe. We renamed key identifiers and reformatted the code for readability. The destination has been removed.

Operation 3: The old search saboteur, now storefront backdoor

Years ago, the Lnkr malware family made the news by hiding inside shady browser extensions, intercepting Google and Bing searches to redirect results and pocket ad money. Now, attackers repurposed the codebase to plant a backdoor into an online retailer’s website.

Because the script was running on a shop rather than a search engine, its old redirect tricks stayed dormant. This time, the script was used to send telemetry back to the attacker. More dangerously, it gave the attacker a remote doorway to arbitrarily download and run fresh JavaScript in customers' browsers whenever they wanted, without touching a single file on the server. It even carried an old trick from its extension days: shutting itself off if someone typed words like “virus” or “popup” into Google. From the outside, the store kept selling without a hint that anything was wrong.

What the shop lost

The shop lost control over what code runs in its customers' browsers. Attackers were secretly tracking visitors' sessions and had a direct backdoor to push and run any JavaScript they wanted on the storefront at any time.

Attack chain

HTML-referenced script → analyst evasion gates → parallel host-gated branches (dormant search vs. live backdoor) → arbitrary remote JavaScript execution

How it stayed hidden

Unlike campaigns delivered through tag managers, this script was directly embedded into the merchant’s HTML. We could not determine the exact initial intrusion vector; in practice, direct HTML insertions usually happen through compromised store admin credentials, an unauthorized template edit, or an infected third-party theme or plugin.

Under the hood, the script is a modular toolkit carrying both active and dormant code. Its older modules (transparent click overlays, search-engine query interceptors, extension-store link rewriters, and redirects for typosquatted domains, like buking[.]com instead of booking[.]com) only wake up on specific target sites, so they stayed turned off on this storefront. Several embedded domains (sugabit[.]net, votetoda[.]com, cdnpps[.]us, and telemetry endpoint hanstrackr[.]com) sat inside these disabled modules.

On the shop, the active branches focused on evasion, telemetry, and remote control:

  • Playing dead for security researchers. An evasion trick inherited from its browser-extension days: the script monitored search inputs and URL queries for telltale adware terms. Searching one security keyword paused the script for that visit. Searching two or more wrote a persistent opt-out record to localStorage, permanently silencing the script on that analyst's machine so repeated tests would find nothing. While originally built to dodge analysts on search engines, as far as we could determine, this check was hard-coded specifically to Google search URLs and remained dormant on the merchant's storefront.
  • Dynamic remote code execution. The script didn't need to modify the storefront to change its behavior. While the hardcoded domain names (scrprime[.]com, youronlinesearches[.]com, jullyambery[.]net) remained identical to older captures, what those endpoints returned was entirely up to the attacker. The script could phone home visitor telemetry, ask those servers for new instructions, and pull down fresh JavaScript directly into the shopper's browser. Effectively, this gave attackers a live backdoor to run arbitrary code on the storefront. We could not determine what second-stage payloads were served in practice.

All in all, a static snapshot of the site showed only the normal storefront, while the underlying state checks, anti-analysis traps, and remote-loading branches exposed the backdoor.

Operation 4: The paid-mobile cloaker

The shop already paid to bring this visitor in from a mobile ad or marketing campaign. The malicious script lets that visit through, then cuts off the merchant's visibility. Analytics go dark, the live support chat vanishes, and a rogue observer starts recording telemetry on the very session the store just bought. 

Behind the scenes, the payload refuses to run unless that visit matches an elaborate set of conditions: the exact target storefront, a narrow mobile screen, and a campaign tag during the first two pages of the visit. It stays dormant on laptops, corporate networks, cloud providers, and VPNs, so the engineers most likely to debug the page never see it fire. The script also stays dormant across selected US cities and regions, backed by a handcrafted denylist of 325 IP strings to dodge automated scanners and security analysts. Only then does the script attempt to tear down the shop’s monitoring, substitute replacement advertising and analytics identities, and phone home. A second look from the wrong device or network will never trigger it. All the while, the storefront keeps selling.

What the shop lost

For a direct-to-consumer retailer, the malware specifically targeted high-value traffic the store had paid to acquire through paid-search and marketing campaigns (ppc, cpc, sms, paid). Those customers could still buy. Yet the shop faced three clear threats: diverted advertising attribution and unearned publisher payouts, the loss of critical session analytics across nine observability tools, and the suppression of the help chat and contact form (preventing shoppers from asking questions or reporting anomalies). Dynamic analysis in a sandboxed browser environment confirmed that the replacement analytics script loaded and fired a tracking beacon (an invisible network request sent to log visitor activity), but whether the attacker successfully captured session telemetry or diverted ad revenue in practice remains unproven.

Attack chain

Campaign-tagged mobile arrival → multi-tier cloaking & network gates → monitoring sabotaged → advertising, analytics, and support controls rewritten 

How it stayed hidden

To blend into the store's marketing supply chain, the attacker delivered the payload from sdk-amazonaws[.]com, a lookalike domain registered in 2024 and wholly unaffiliated with the official Amazon Web Services domain (amazonaws.com, registered in 2005). To compound the deception, the attacker prefixed the domain with a subdomain mimicking a popular e-commerce marketing platform too. This stacked, double-trusted-brand typosquat forged a convincing disguise, engineered to slip past quick tag reviews. Neither Amazon Web Services nor the impersonated marketing platform was involved in the attack or suffered any compromise.

Once loaded in the browser, the script executed an exceptionally dense gauntlet of cloaking gates before triggering its main payload:

  • Target host and browsing context. The script verified that window.location.hostname matched the specific merchant host it was built to target (exiting immediately anywhere else), ensured the current window was top-level (not an embedded iframe), and checked that the path did not contain /challenge. It also verified that tracking marker cookies (_cart_dr and logoalt) were not already present in the browser.
  • Device and campaign filtering. The visitor's viewport width had to be narrower than 477 pixels (a handheld smartphone). Furthermore, the visitor had to arrive via a first-touch (the visitor's initial referral) campaign tagged with one of six specific UTM mediums (Urchin Tracking Module, standard URL tags used to track marketing campaigns): ppc, cpc, sms, paid, flow, or campaign. It also had to be the first or second page load of their session. Curiously, while the code contained a nominal non-UTM path, it required the session page count to be simultaneously greater than -1 and less than -2 (a mathematical impossibility that left that branch completely unreachable). This could be yet another diversion technique or a code change leftover.
  • The "random" gate that always passed. The code featured what looked like a probabilistic throttle (Math.random() <= threshold) to make execution appear intermittent. However, when we solved the deobfuscated arithmetic, the threshold reduced to exactly 1. Because JavaScript's Math.random() always returns a value strictly below 1, this gate always evaluates to true. Like the unreachable non-campaign branch, this is a condition that never actually decides anything. It may be a leftover throttle, or purposeful camouflage for anyone reading the obfuscated math. Either way, the captured payload never used it to skip a qualifying visitor.
  • Third-party IP intelligence. The script queried an external third-party IP intelligence API using an embedded key. It required a U.S. mobile consumer connection, while immediately bailing out if the response flagged a business network, hosting facility, cloud provider, bogon, Tor exit node, VPN, proxy, relay, or general threat indicator.
  • Geographic exclusions. The payload did not activate for visitors from specific regions (US-NY, US-CA, and US-NHDD) or from cities named San Francisco, Plymouth, Compton, Hopkinton, and Lafayette.
  • The 325-entry IP substring trap. The script checked the visitor's IP against an embedded denylist containing 325 full IPv4 address strings. After deduplication, these represented 313 unique addresses across 249 distinct three-octet prefixes. Rather than performing structured CIDR (Classless Inter-Domain Routing) subnet matching, the author simply stripped the final octet from the visitor's IPv4 address and ran a raw substring search: !denylistString.includes(visitorPrefix).

In stripped-down pseudocode, the multi-layer primary activation funnel looks like this:

Sabotaging observability and hijacking identities:

Only after every primary gate passed did the script execute its payload:

  • Blinding monitoring tools. It searched the DOM and removed script tags for nine different observability and analytics services: Lucky Orange, Segment, Optimizely, New Relic, Bugsnag, LogRocket, Hotjar, Microsoft Clarity, and the store's Google Tag Manager container (GTM-<redacted>). In remaining inline scripts, it string-replaced references to these tools with undefined dummy identifiers (hji0) so calls to them failed silently, attempting to blind the store's error reporting and monitoring.
  • Suppressing customer support. It injected CSS and removed elements to hide its support-chat and contact-form containers, cutting off the customer's direct line to store support.
  • Replacing advertising and analytics identities. It purged Google Ads globals (google_ad_modifications, adsbygoogle), tore down existing ad slots (ca-pub-<original>), and loaded Google Ads under a replacement publisher ID (ca-pub-<replacement>). It then injected a new Microsoft Clarity session-replay script configured with a rogue, replacement project ID.

Simpler independent beacons and the 600-day marker:

In sharp contrast to the elaborate primary cloak, the payload also contained secondary beaconing branches (standalone routines that quietly ping an external server to confirm a visit) that completely bypassed the viewport, hostname, campaign, geography, and IP gates. If the visitor was on their second page or beyond, the script wrote a persistent cookie (_cart_dr=1) with an expiry of exactly 600 days (51,840,000,000 milliseconds) and fired an invisible zero-pixel image request to a remote telemetry endpoint on maper[.]info (a tracking beacon used to log that the browser reached this step).

A separate branch checked for an alternate marker (_logo_alt), which would trigger a second telemetry .png beacon (a cookie this script looked for, but never wrote itself; likely planted by a companion script). This gave the attacker a simple, persistent hit-counter to log basic traffic for all visitors (IP and User-Agent logged at the endpoint) across the entire store, while keeping their high-risk ad-hijacking routines strictly hidden behind the mobile cloak (high-value paid arrivals). It shows why analyzing only one visible effect does not reveal the full reach of a multi-purpose payload.

Indicators of Compromise (IOCs)

We are publishing these indicators to help security teams and researchers detect and hunt these campaigns across their own environments. All indicators are drawn directly from captured payloads and their network connections. Listed URLs are defanged. Some indicators have been withheld or generalized because publishing them could inadvertently divulge the identities of affected organizations. Listed domains reflect infrastructure observed participating in the delivery, redirection, or telemetry chain during these attacks; inclusion does not imply that a shared service or hosting provider is exclusively malicious.

Four lessons for defenders

Taken together, the operations tell one escalating story: attackers changed the objective, delivery path, and disguise, but the browser still had to execute their logic. Four lessons stand out.

Behavior beats signatures. These operations pursued different forms of monetization and manipulation, but every payload still had to act in the browser: observe events, inspect state, alter the page, schedule work, make network requests, or load another stage. That is what structural analysis looks for: the logic a hostile payload must carry, even as URLs, signatures, and objectives change.

Selective execution is part of the attack, not a footnote. Device, time, geography, referrer, session, network, and cooldown gates can all defeat a crawler that visits once and takes a static snapshot. Continuous visibility matters because an attack may appear only to one browser, in one state, at one moment. 

Obfuscation raised the cost of analysis, but in these cases it did not prevent detection. Self-defending loops, console suppression, debugger traps, rotated string tables, and dead branches complicated analysis. Page Shield ML still surfaced all four operations despite those barriers. Fast in-house models surface the suspicious code at scale, while frontier models investigate the hardest cases. Their disagreements highlight the trickiest obfuscation and logic, helping us narrow our focus.

Context completes the picture. Code that looks ordinary in isolation can reveal its malicious role once defenders link static analysis with dynamic context: how it arrived, which browser state activated it, what connections it opened, and what it actually did at runtime.

Continuous visibility into client-side execution

These four operations relied on different layers of misdirection, but they all shared one constraint: their JavaScript had to execute in the browser. Public scanners and static crawls can miss gated behavior. Continuous observation helps clarify what the code actually does when real visitors interact with the page

Cloudflare Client-Side Security provides that visibility across all plans. You can turn on Continuous script monitoring under Security settings to track first- and third-party scripts on your storefront, while automated malicious-script detection and alerting are available with Client-Side Security Advanced. You can review script activity and manage detections directly in the Cloudflare dashboard.

How we rebuilt Cloudflare Workers’ module registry for Node.js compatibility

Post Syndicated from Logan Gatlin original https://blog.cloudflare.com/workers-module-registry-nodejs/

We’ve rewritten the module registry in workerd, the core open-source component of the Workers runtime, to be faster, more standards-compliant, and more closely aligned with Node.js' module registry.

Over the past few years, we’ve been adding support for more and more Node.js runtime APIs. The Workers runtime now supports every stable API from Node.js that you might want to use in a serverless context, and these APIs are now enabled by default, letting you deploy even larger Node.js apps to Cloudflare (now up to 64 MiB on all plans — we’ve removed the limit on compressed bundle size).

But API compatibility alone is not enough: Node.js applications also depend on how the runtime resolves, loads, and caches modules. ESM, CommonJS, and WebAssembly are each types of modules that you can import in your Worker’s code. The system within the runtime that handles all of this is called the module registry.

You can start using it today by enabling the new_module_registry compatibility flag in your Worker.

When you enable the new_module_registry compatibility flag:

  • import.meta.url, import.meta.main, and import.meta.resolve() all work.
  • Module specifiers are parsed and resolved as real URLs, including query strings and fragments.
  • node: built-ins resolve to the same module instance no matter how you reach them.
  • Import attributes (with { type: 'json' }) are correctly validated.
  • require() on an ES module follows Node.js' require(esm) rules.
  • Errors use consistent classes and messages regardless of which loading path triggered them.
  • Modules compile lazily when first imported (statically or dynamically).
  • WebAssembly modules support source phase imports.

For the full deep-dive on how this new module registry interacts with V8’s module APIs, we’ve added reference docs to workerd that break down everything in detail. But for most people building on Workers, you want to understand how these changes improve compatibility and help you build. To do that, we’ll dive into each of these changes in the sections below.

How the Workers runtime loads the code you give it

When you deploy a Worker to Cloudflare, wrangler or Vite “bundles” all of your Worker’s code from many files and dependencies into one or many modules, which are then uploaded to Cloudflare when you run wrangler deploy.

By default, Wrangler bundles nearly all of this code into a single module script. It runs esbuild under the hood, which processes then inlines relative imports and require() calls for most npm dependencies into that one file. The import and require() statements are replaced with regular functions as part of the process. By the time that bundle reaches the Workers runtime (workerd), there usually isn't much of a module graph left for the Workers runtime to deal with. Most of the different modules are bundled into one file. We have seen these scripts grow to as many as multiple hundreds of thousands of lines long.

Why is it necessary to bundle many modules into a single file before uploading server-side code to Cloudflare? It has been technically possible to upload multiple modules, and even modules of different types, in the Workers runtime for many years now. However, the runtime has not resolved modules in a way that was consistent with all the other runtimes. If, for example, your code or dependencies used import.meta.resolve() to resolve the path to another module, that code would fail because import.meta.resolve() was not supported.

When you use the Cloudflare Vite plugin, Vite 8 bundles your code using Rolldown, instead of Wrangler bundling your code using esbuild. Rolldown resolves imports and npm dependencies, converts CommonJS to ESM where necessary, and emits an entry module plus any additional chunks created through code splitting, such as dynamic imports. As a result, the Workers runtime receives a smaller, build-generated module graph rather than the application’s original source graph.

The new module registry implementation in the Workers runtime opens the door to bundlers like Rolldown to perform fewer transformations, and to rely more on the runtime to handle module resolution.

When you import a Node.js API in your worker, by default you are importing a module that is built into workerd. It is not bundled into your code as a polyfill. Wasm, text, and binary modules are provided to the Workers runtime as separate files too. They are referenced by specifier instead of being inlined. And if you deploy with --no-bundle, or your tooling uploads a Worker as multiple modules directly, the full module graph shows up at runtime exactly as you wrote it.

In all of these cases, something has to take a specifier, work out what code it actually points to, compile it, and hand V8 a module object it can link and run. In workerd, that's the module registry's job.

Why a new implementation?

The original registry resolves specifiers as filesystem-style paths, not URLs. That sounds like a minor distinction, but it ruled out a bunch of things: there was no clean way to implement import.meta.url, relative imports didn't follow the same resolution rules as new URL(), and protocols like node: and cloudflare: were handled as special-cased string prefixes instead of, well, protocols.

It also compiles your entire Worker bundle up front, whether or not a given module ever gets imported, and it keeps a separate, private copy of everything per V8 isolate. Cloudflare runs multiple V8 isolate replicas of the same Worker to spread load across CPU cores, so in practice that meant compiling the exact same source more than once, with keeping multiple copies of the source in memory.

None of this is really a bug, but it made it difficult to evolve the implementation without breaking changes. The new registry starts from URLs as the specifier format and treats laziness and cache sharing as things to design in from day one. The existing registry implementation is not going anywhere. Currently, deployed Workers will continue to work as they always have.

import.meta

The import.meta API provides information about the module, such as the module's URL, and whether it is the main entry point module:

That prints something like file:///bundle/index.js, main: true. 

import.meta.main is true only for the module configured as your Worker's entrypoint; every other module gets false.

import.meta.resolve() resolves a specifier against the current module without importing it:

It's a pure string transform, same as in Node.js and in browsers: it doesn't check that the resolved URL corresponds to a real module, and it throws a TypeError for a specifier that can't be parsed as a URL at all, rather than returning null. One detail worth knowing if you ever look closely at the output: it normalizes percent-encoding the same way new URL() does, which means it collapses paths like ./a/../b.js, but it does not decode characters that were already percent-encoded. import.meta.resolve('%66oo.js') resolves to file:///bundle/%66oo.js, not file:///bundle/foo.js.

Specifiers are URLs

Relative imports now resolve the same way as new URL(specifier, base) would, because that's literally what's happening under the hood. Full URLs work as specifiers too, not just relative paths:

The more interesting consequence is what happens with query strings and fragments. Per the same module-identity rules browsers use, a specifier with a different query string or fragment is treated as a genuinely distinct module instance, even when it points at the same underlying source:

./counter.js?a and ./counter.js?b load the same source, but they're evaluated separately, each gets its own import.meta.url, and each gets its own copy of any top-level state. Importing the same specifier with the same query string again still gets you back the same instance, so this isn't a way to force re-evaluation on every import.

Import attributes are correctly validated

The original module registry implementation silently ignores the import attributes in violation of the spec. It is expected that implementations throw an exception when any import attribute it does not understand is used.

json is the only import attribute type enabled right now, since it's the only one of the relevant TC39 proposals that has reached Stage 4. text and bytes are recognized, because they track the Import Text and Import Bytes proposals, but they're rejected with a specific error instead of being silently ignored or treated as unsupported syntax:

Any attribute key other than type is now a hard error too, rather than being ignored:

And if the type you specify doesn't match what the module actually is:

require(esm) follows Node.js' rules

If you require() something that turns out to be an ES module, whether that's directly inside a CommonJS module or through require('node:module').createRequire(), the registry follows Node.js' require(esm) behavior:

  • If the module has a string-named export called 'module.exports', Node.js' actual mechanism for letting an ES module control what require() sees, that value is returned.
  • Otherwise, require() returns the module's namespace object.
  • The one exception is workerd's own node: built-ins. They're implemented as ES modules that wrap a CommonJS-style API in a default export, so requiring one returns that default export directly. require('node:buffer').Buffer behaves the way you'd expect; you don't get a namespace object with a .default you need to unwrap yourself.

There's a restriction that comes along with this: if the module you're requiring, or anything in its module graph, has a top-level await, require() throws instead of blocking or handing back something half-finished:

This matches Node.js' own ERR_REQUIRE_ASYNC_MODULE restriction: require() has to return synchronously, and there's no reasonable value to hand back for a module that hasn't finished evaluating yet. Use import() for anything async instead. The check holds regardless of import order too: a module doesn't become require()-able just because something already import()'d and fully evaluated it earlier.

If you're requiring output from a bundler that predates Node.js' require(esm) support and sets a truthy __cjsUnwrapDefault export as a marker, that takes priority over both rules above and returns the default export. That's purely there so existing prebuilt bundles keep working.

Errors are consistent, and use the right class

Regardless of whether resolution fails through a static import, a dynamic import(), or require(), you get the same class of error with the same message shape:

"Module not found" is a plain Error, since it's a failure to locate something rather than a problem with the value you passed in. A specifier that can't be parsed as a URL at all is a TypeError, matching Node.js' own ERR_INVALID_MODULE_SPECIFIER. A circular dependency that V8 can't unwind is also a plain Error, never a TypeError. This mostly matters if you're building something on top of dynamic import(), like your own loader or a retry wrapper, since you can now branch on the error class or message reliably no matter which loading path triggered it.

WebAssembly source phase imports

You can now import the compiled-but-not-instantiated form of a WebAssembly module directly, using source phase imports:

or dynamically:

Either way you get a WebAssembly.Module back directly, instead of importing the module normally and pulling it off the default export. As source phase imports are a new feature of the language, right now this only works for WebAssembly; trying it on any other module type throws a SyntaxError, matching the behavior of Node.js and other runtimes.

What's next

Try it out! Add the new_module_registry compatibility flag to your Worker:

It doesn't have a default on date yet, so it won't turn on automatically for your Worker, old or new, no matter what compatibility date it's using. You will need to add the flag explicitly.

We’d love your feedback. workerd is open source. If you run into behavior that looks like a regression rather than one of the changes described here, please file it against the workerd repository.

The Cloudflare Blog – Brought to you by EmDash

Post Syndicated from Kody Jackson original https://blog.cloudflare.com/cloudflare-blog-uses-emdash/

You likely noticed the recent redesign of the Cloudflare Blog. We added dark mode, modernized the look and feel, and made a lot of other small improvements along the way.

What you might not have noticed – well, except for those who are more terminally online – is that the redesign was part of a much bigger migration project. On Wednesday, August 12, we moved the blog to EmDash, a content management system (CMS) built especially to work on Astro and with Cloudflare.

We’ll take you into the migration story – what we learned and how EmDash got better – as well as into the benefits we’re already seeing from a new platform.

We are Customer Zero

At Cloudflare, Cloudflare itself is Customer Zero. This means that we use our products. And – in use – we make them better for ourselves and our customers.

This is a very real cultural value at Cloudflare. The burden of proof is on you if you want to use an external vendor. Why can’t that team support you, what gaps are there, why can’t those gaps be filled, and are those “gaps” true requirements?

This preference is even enshrined in our internal engineering standards, known as our Codex.

We don’t just build products for others; we build them to run Cloudflare itself. We are our own first, most demanding customer.

We validate scale, security, and usability on our own massive infrastructure before a paying customer ever touches the product. If a product breaks, it breaks us first. This forces us to fix issues immediately, ensuring that by the time a feature reaches the enterprise, it has already survived the harshest production environment on earth.

With the launch of EmDash and some limitations with our current CMS vendor, we knew that we’d likely be the Customer Zero for EmDash internally at Cloudflare.

Customer Zero in Action

When we began our initial migration conversations, we started with two main questions:

  • Does EmDash work for us?
  • Can EmDash scale?

Does the platform work?

Our first question was the most broad, does EmDash work for us? This is something you’d want to know broadly about any new platform, but especially one that’s pre-1.0. 

To answer this question, we ran through a bunch of common user flows, such as:

  • Publishing and unpublishing a post
  • Authoring a new post
  • Scheduling a post
  • Adding media items

By and large, EmDash held up pretty well to these usability tests. The gaps we found were generally related to:

The biggest oversight we found was around scheduled posts, which didn’t work until EmDash version 0.19.0. This gap was understandable given the early version of EmDash, but it was also definitely something we didn’t want to be finding out after the scheduled time for a post.

Can EmDash scale?

Our biggest concerns were whether our proposed EmDash setup could handle the traffic we saw on the Cloudflare Blog.

The traffic pattern to our blog is incredibly varied. Normal load sits in the neighborhood of 75 requests per second (RPS), but also spikes up to over 5,000 RPS. Some of these spikes line up with the publishing times of new posts, meaning those posts went viral and attracted a lot of attention. Others happen during all points of the day and night, which likely means folks are sending some extra traffic our way, just to see what happens. 

Performance also matters for our systems (and our readers). Cloudflare is a web performance company, after all, so the speed at which a page loads becomes incredibly important.

With those two concerns in mind, we built out some scenarios using k6, an open-source performance testing tool:

  • Ramp: Where we gradually increase requests up to triple the prod baseline and then cool down.
  • Breakpoint: Where we ramp from 0 to 100 RPS over 10 minutes, stopping when something breaks.
  • Burst: Where we throw an immediate traffic load of 7,000 RPS and see what happens.

For each of those scenarios, we evaluated:

  • Availability: Failure when more than 0.01% of HTTP requests lead to 5xx errors, meaning the application couldn’t handle the traffic.
  • Latency:
    • P95 latency: Failure when more than 5% of responses exceed 500ms.
    • P99 latency: Failure when more than 1% of responses exceed 1000ms.

Armed with these tests – and a lot of internal discussion and data points – we came to our production architecture:

The multiple layers of caching we put in place play a key role in making the blog both fast and resilient. In the diagram below, they are ordered from top to bottom by proximity to the user:

With this setup, we’re typically serving 99.5% of static files from a cache and 70% of requests from a cache, improving frontend performance and decreasing load on the database.

Once we had that architecture in place, we could start thinking about the frontend redesign as well.

Frontend redesign

Beyond updating the backend architecture, the migration offered us the perfect opportunity to bring the blog's interface into alignment with Cloudflare’s updated visual language. We rebuilt the frontend experience using patterns established by the Kumo design system, creating visual and structural consistency between the Cloudflare homepage, dashboard, and marketing sites. The result is a cohesive reading experience that feels like a natural extension of the broader Cloudflare ecosystem.

A major priority for this redesign, and a long-overdue request from our readers, was native support for light and dark modes. We implemented theme switching tied directly to system preferences, alongside an explicit toggle, and ensured that accessibility guidelines were strictly met across both themes. Regardless of preference, the updated palette and code syntax highlighting adapt seamlessly without sacrificing legibility.

We also took the opportunity to solve a few long-standing user experience quirks, starting with our email subscription form. Previously, the subscription box lived in the top right corner of the page. Because of its placement, readers frequently mistook it for a search bar and typed their search queries directly into the input field.

To fix this, we moved the email sign-up into a dedicated call-to-action block at the bottom of posts. 

Now, once a reader finishes an article and wants to stay updated, the prompt to subscribe appears naturally at the end of a post.

Finally, we introduced two dedicated sidebar features on interior post pages to improve navigation and community engagement. On the right, an "On this page" table of contents tracks your progress and lets you jump directly to specific sections of longer technical posts. On the left, a new "Discuss Online" section makes it effortless to share articles and engage in conversations across social platforms and developer communities.

Rollout strategy

As we got nearer to our migration, we started focusing on the broader question of “how do we make this change safely?” Ensuring zero downtime for our readers was a non-negotiable requirement, alongside guaranteeing a seamless fallback mechanism if something went wrong at the last minute.

To achieve this, we deployed a proxy Worker to intelligently route traffic between the legacy blog and the new EmDash-powered site. This Worker set a version cookie on requests, which then let us route incoming traffic to the new or legacy experience accordingly. Additionally, this strategy allowed us to fall back to the legacy blog if the new site experienced any 500 errors. Thanks to the flexibility of Cloudflare Workers, this proxy was relatively simple to create and scaled without any issues. The ability to configure a direct worker-to-worker connection through the NEW_BLOG service binding was particularly useful here, as it reduced latency for any end user going through the proxy. This service binding let the proxy Worker dispatch incoming requests directly to the new blog Worker instead of sending them through a public hostname, DNS, TLS, and an outbound HTTP connection.

On launch day, we initiated a gradual rollout, starting at just 1% of total traffic, then incrementally stepping up to 5%, 15%, and beyond as we validated system health. This phased approach allowed us to observe how the platform handled real-world production load while catching a few last-minute edge cases without impacting the vast majority of our audience. By the end of the day, we had comfortably shifted 100% of traffic over to the new platform.

Results

Measurable performance gains

One of our primary objectives for this migration was to deliver a faster, more reliable site to our readers, and the early data shows we accomplished exactly that.

Comparing p95 response latencies between the old architecture (green line) and the new EmDash setup (yellow line) revealed a stark difference. Where the previous platform experienced periodic latency spikes under load, the new system maintains a remarkably flat, consistent response profile. By running EmDash on Cloudflare Workers alongside our new caching layers, we’ve delivered a significantly faster and more performant reading experience across the board.

We’ve seen all these performance gains – and minimal errors – while serving up to 850 RPS.

MCP servers

With this change, the blog also got more accessible for agents, in two distinct ways.

The first is that we released a new Model Context Protocol (MCP) server for the Cloudflare Blog.

An MCP server bundles up a bunch of specific tools that your agent can then use to interact with an external resource, almost like an API for agents.

Using that MCP, you can now use the following tools with your agents:

  • search_posts
  • list_posts
  • get_post
  • list_tags

With the new, intuitive EmDash APIs and AI search endpoints exposed by our Worker, creating this new MCP took just a few hours of work.

The second is that – for our blog authors – EmDash has an MCP server for EmDash itself, meaning that they can browse, create, and edit content, publish and schedule posts, remove files, and more.

Though this sort of agentic tooling is becoming more standardized in the CMS industry, what’s not standard is that it’s available without any additional cost. The MCP is just another part of the platform, reflecting a growing trend of designing for agents, as well as humans.

The first test: Agents Week

At Cloudflare, we run multiple innovation “weeks” a year, where we set ambitious goals for internal teams around specific themes. These weeks push our products forward, as well as help customers digest the changes that are constantly happening at Cloudflare.

The latest of these, Agents Week, was quite a test for the new blog. We launched 18 new posts over 9 days. And those posts got a lot of traffic, close to 3 million pageviews.

On the frontend, our new blog Worker did very well, serving up to 450 RPS without any noticeable issues. Thanks to Cloudflare’s built-in DDoS protection, we also absorbed a 28,000 RPS DDoS attack on August 10th, also without any noticeable issues.

On the editing side, we continued to find some issues. Most of these involved small quirks of the editing experience, though we also found some bugs specifically around scheduled posts. We’ve since raised these to the EmDash team and are confident that they’ll be fixed before Birthday Week.

Give EmDash a try

We want to give a heartfelt thank you to the EmDash team, who made this migration about as smooth as possible and were incredibly receptive to our feedback. This is how Customer Zero is supposed to work, and it’s incredibly gratifying to share an inside look into that process with all of our readers as well.

If you’re in the market for a new CMS, try out EmDash today. It’s pretty amazing and – with the upcoming launch to v1 – it’ll be getting even better soon.

From all-or-nothing to task-based OAuth consent

Post Syndicated from Miller Vargas original https://blog.cloudflare.com/task-based-oauth-consent/

Since June, developers have created thousands of third-party OAuth apps on Cloudflare, with more than a million authorizations since.  

OAuth makes delegated access possible. It lets applications act on a user’s behalf without asking them to handle long-lived credentials or hand over a password. That model works well when an application can describe its access needs with a small set of scopes. 

Developers use OAuth for SaaS integrations, internal tools, CLIs, and agents. Our permission model has become more granular over time to support better scoping of these different workflows. That is great for security, but it makes a purely all-or-nothing consent screen hard to justify.

Cloudflare OAuth already allows clients to request a subset of their configured scopes. But once the client made that request, the user could not narrow it any further on the consent screen. For the user on the consent screen, the experience was still an all-or-nothing one. If an application requested more access than a user was comfortable granting, their only options were to approve the full request, or deny outright. 

MCP servers are a good example of this. An MCP server might request a broad set of permissions, because in theory an agent could use all of them. But most users would not want an agent to have that much access. Before this feature, the only way to handle this was for the app developer to build a custom scope selection screen before sending the user to our consent flow.

Today, we’re introducing OAuth scope customization. Client owners can mark specific scopes as optional when configuring an OAuth client, giving users the ability to grant a narrower subset of an application’s requested access at authorization time.

The OAuth spec already allows authorization servers to grant a narrower set of scopes than what was requested. We built on top of that flexibility to make this work cleanly for every existing app.

More control, without overwhelming users

Our goal with introducing scope selection is to give security conscious users more flexibility to make the right choices for their use case, without turning the consent screen into a long scope checklist. 

With scope customization: 

  • Developers can mark specific scopes on an OAuth client as required or optional
  • At authorization time, users can deselect optional scopes from the requested set
  • Required and optional scopes are evaluated against the scopes requested for that authorization flow
  • If no optional scopes are requested, the consent experience stays the same
  • By default, the consent screen still grants the full requested scope set.  

Scoping to the authorization request

One important detail is that required and optional scopes are evaluated only against the scopes requested in a specific authorization flow, not every scope configured on the client. That matters because OAuth clients do not always request their full configured scope set.

For example, a client might be configured with user-details.read, workers-scripts.write, workers-kv-storage.write, and zone.read, while marking workers-kv-storage.write and zone.read as optional. If that client starts an authorization flow requesting all four scopes, the consent screen will evaluate all four. In that case, user-details.read and workers-scripts.write remain required, while the user can choose whether to grant workers-kv-storage.write and zone.read.

But if the client later requests only workers-scripts.write and zone.read, then only those two scopes are considered for that authorization flow. user-details.read and workers-kv-storage.write would not be shown or enforced, because they were not requested.

This keeps the consent screen focused on the task at hand, rather than every capability the application could request. It also means existing OAuth clients keep their current behavior by default: if a client does not opt into optional scopes, the consent flow remains unchanged.

Configuring an OAuth client to use optional scopes

Developers can opt into scope customization when configuring an OAuth client. Scopes continue to be configured as they are today, and clients can now additionally specify which of those scopes are optional: 

In the example above, the client can request all four scopes, but the user may only opt out of the workers-kv-storage.write and zone.read scopes during consent. user-details:read and workers-scripts.write remain required if they are included in the authorization request. 

If the client later requests only workers-scripts.write and zone.read, then only those two scopes are considered for that authorization flow. user-details.read and workers-kv-storage.write would not be shown or enforced because it was not requested.

Building with partial grants in mind

When a user deselects any optional scopes and completes the authorization flow, the generated access token will only contain the scopes they consented to. For developers, this means you need to check the granted scope set after exchanging the authorization code, rather than assuming the full requested set of scopes was approved.

An app that handles a narrower grant gracefully, for example an agent that operates within whatever subset of permissions it receives, is one that users feel comfortable authorizing. Requesting only the permissions needed and marking the rest as optional is a good sign to users that your app respects their access decisions.

Scopes for every Product

Over the next few weeks, we will be expanding our account & zone-level role surface to cover nearly every Cloudflare product. That means more API token roles, account membership options, and OAuth scopes, giving customers the tools to secure workloads with the right level of access. 

Build with Optional Scopes

Allowing developers and users to better restrict access through optional OAuth scopes is an important step toward a more flexible and trustworthy consent experience on Cloudflare. With optional scopes, developers can build more nuanced authorization flows, and users gain more control over what they approve. 

To get started with Third Party OAuth, take a look at our documentation or jump straight to the OAuth apps page in the dashboard and create your first OAuth app. 

Thank you to our amazing interns

This feature is one of the many that we built with the help of our 1,111 interns. Congratulations to Miller Vargas and José Enrique Rodriguez on your high impact contributions here. Miller is a senior at the University of Texas – Austin studying computer science and math; and José is a senior at Universidad Panamericana studying engineering, data intelligence, and cybersecurity.

Secure all your internal vibe-coded applications — in one click

Post Syndicated from Chythra Malapati original https://blog.cloudflare.com/workers-protected-by-access/

AI has enabled employees across every team to build applications faster than ever before.

But that speed is also what's keeping every CISO up at night: any employee can build an application, deploy it to the public Internet, and accidentally expose internal work or company data.

Today, we're launching new tools to make it easy to keep your applications hosted on Workers private. You can now apply Cloudflare Access directly to a Worker or to every Worker in your account, so that your applications are behind your company login by default, without relying on each developer to set that up themselves.

You can now:

  • Set a policy at the account level to ensure that all preview and production deployments are behind your company login by default.
  • Set a policy on a single application to ensure authentication is enforced on every domain associated with it, no matter how it's deployed.
  • See exactly who visits your application. Get every authenticated user’s email, name, and groups directly in your code — no JWT (JSON Web Token) validation required. 
  • Deploy an internal platform where every deployment is private by default. We've open-sourced an example: an internal static site platform where every Worker deployed is private.

Access on Workers: how it works

When you enable Access on a Worker, Cloudflare enforces authentication before any request reaches your application code. It doesn't matter how the request gets to your Worker, whether it's through a custom domain, a route, a workers.dev subdomain, or a preview URL. If Access is on, the user has to authenticate first.

Previously, you had to configure this at the hostname level, which meant setting up Access policies on each domain your Worker was reachable on. If you wanted to add a new custom domain to your Worker, you needed to update the Access policy first or that hostname would be reachable without authentication.
Now the policy is attached to the Worker itself, so any domain or URL associated with that Worker is automatically protected. You can choose what to protect: just preview URLs, or all hostnames. 

If you set it to previews only, every preview URL created for that application, whether it's a workers.dev preview URL or a custom domain you use for previews, will require authentication whenever you deploy a new version. If you set it to all hostnames, every domain associated with that Worker is protected — custom domains, routes, workers.dev subdomains, and preview URLs.

Access gives you control over how users authenticate. You can connect your existing identity provider, so employees sign in with the credentials they already use, or restrict access to specific email addresses, email domains, or groups. For agents, you can grant access through service tokens.

Read more in the Cloudflare Access for Workers documentation here.

Keep every Worker in your account private by default

If you have developers across your organization deploying Workers, you don't want to rely on each one to remember to enable Access. You want the default to be private.

You can set an Access policy once at the account level, and every Worker in your account, current and future, is private from the moment it's created.

You choose what the policy covers: only preview URL traffic, all production traffic, or both. Preview-only is useful if your production Workers are intentionally public, but you never want an in-progress deployment exposed.

Need a Worker to be public? Bypass the account-wide policy on that one Worker.

Protect a specific Worker

If you don't need an account-wide default and just want to lock down one specific Worker, you can apply Access to that Worker directly.

The new Access tab in the Worker view shows exactly which policies apply to that application. If you have multiple, the most specific one takes priority: hostname policies first, then Worker policies, then account policies.

See who is accessing your application

When Access is protecting your Worker, you can get information about who is making each request — their email, name, and groups — so you can personalize what they see, enforce permissions, or log activity per user.

This works through your Worker's context object (ctx). Every request to your Worker carries a ctx with metadata about that request. When Access is enabled, we attach the authenticated user's identity to it as ctx.access. From there, call ctx.access.getIdentity() to get back the user's email, name, and more.

Before, this meant validating a JWT yourself — parsing the token, verifying the signature, and extracting the claims. Now, when Access is enabled on your Worker, every authenticated request includes ctx.access.

Here's all you need to get the user's identity:

Test locally before you deploy

We showed how you can use ctx.access.getIdentity() to give your Worker information about who is making a request — their email, name, and groups. 

You can use this when developing locally with wrangler dev. Add an access block to your wrangler.jsonc to simulate an authenticated user:

Your Worker picks it up through ctx.access.getIdentity() — returning an identity object shaped like what you'd get in production. Swap the email in your config to test as a different user.

This means you can verify that the right content shows up for the right user without having to deploy and sign in through Access every time you make a change.

Deploy an internal platform where every application is private by default

If you manage an internal platform where employees can prototype and deploy applications, you need every application to be private without configuring access controls on each one.

Workers for Platforms lets you deploy Workers at scale. Every Worker lives inside a namespace, and all traffic to that namespace goes through a single entry point: the dispatch Worker.

Set an Access policy on your dispatch Worker, and every Worker deployed through it is private by default.

We also have an open-source example where you can deploy your own internal drag-and-drop deployment platform — configure access on the dispatcher worker once and every site deployed through it is private by default.

Click the button below to deploy it yourself!

For the full architecture, see our Workers for Platforms reference architecture.

Built on solid foundations

This feature was made possible by FL2, the new Rust-based modular proxy that powers Cloudflare's edge. Access is the front gate to your applications, and as such, it traditionally ran before all Workers logic in the request pipeline. But in order for Access applications to target individual Workers themselves instead of their hostnames, Access needs to know which Worker a given request is destined to reach. Therefore, we needed to split Workers routing from Workers execution, and move the routing logic, so it could run before Access.

In our old FL1 system based on NGINX and modules written in Lua, this change would have been complex and risky. Interactions between products can be subtle, and moving logic to an earlier phase of the request pipeline can be unsafe if it depends on shared state that is modified by another product.

FL2 made it easy. Its strict module system separates logic into well-defined, consistently ordered phases that statically declare their inputs and outputs. We were able to lean on the compiler to surface any broken interactions between phases, and gradually roll out this refactor with confidence.

Try it today

This is now available to everyone. Try it out in the dashboard or read the Cloudflare Access for Workers documentation to get started.

Acknowledgments

Thank you to Jesse Li, Brandon Strittmatter, Kyle Hiller, Kenny Johnson, Matt "TK" Taylor, Brendan Irvine-Broque, Yomna Shousha, and Mike Aizatsky for the engineering and design work that made this possible!

Everything we launched during Agents Week

Post Syndicated from Shelley Jones original https://blog.cloudflare.com/agents-week-review-august-2026/

At the beginning of Agents Week, Rita shared that agents represent the next evolution of computing: not only as a new application of AI but also as a new class of software that’s shaping how people interact with technology, and how software interacts with the Internet. Over the last year or so, we set out to explore what this shift means for developers and customers building AI-native apps and the infrastructure needed to support them. As agents become more capable and autonomous, the challenges extend beyond the models themselves — to identity, communication, orchestration, memory, observability, and security.

Over the past week we’ve shared how we’re bringing those pieces together across the Cloudflare platform to serve an Agentic Internet. Each day we presented new tools, products, and ideas toward building for an Internet where humans and agents cooperate instead of collide.

Monday, August 3

Monday focused on the foundations for building and running intelligent, autonomous apps — the runtime and infrastructure agents rely on.

Tuesday, August 4

Tuesday introduced the Agent Development Lifecycle (ADLC) and the primitives that take agentic software from prototype to production.

Wednesday, August 5

Wednesday extended Zero Trust from users and devices to agents themselves — and we shared how we’re running it internally at Cloudflare. 

Thursday, August 6

Thursday defined the Agentic Internet, and how website owners, publishers, and agents can all contribute to an Internet that works for people and agents alike.

Friday, August 7

Friday put a lens on what’s actually happening: what agents are really doing on the web, where AI is running in your apps, who’s contributing to the ecosystems, and new tools for analyzing Internet data.

Agents Week is done, but we aren’t

Five days on, the answer to Rita’s question of “What does your agent need from an Agent Cloud?” is starting to take shape. It needs an execution layer and primitives to run on, a development lifecycle that increasingly writes itself, secure access for the people and agents doing the work, an Agentic Internet, and the humans and communities keeping all of it grounded. There's plenty still to come, but the shape of what’s next is becoming clearer: an Internet that natively supports the humans it was built for and the agents now acting on their behalf.

Our work doesn’t stop here. Keep an eye on our changelog for the latest updates. And if you’re building any part of this with us, we’d love to hear from you! Come find us on X or Discord.