Tag Archives: Birthday Week

How fast is the web? Explore billions of real-user measurements with BEACON

Post Syndicated from Ryan Townsend original https://blog.cloudflare.com/how-fast-is-the-web/

If you work in technology, you’re probably reading this on a powerful laptop or flagship mobile on robust, lightning-fast Wi-Fi. This is fantastic for building software, but often is far removed from the reality facing many who are using that software.

End users might be nursing a four-year-old budget phone, running low on battery, on a data plan that throttles after 2 GB, living with under-invested public infrastructure, or even just walking into that corner of the gym where the Wi-Fi never seems to work. Multiply this by billions of people around the world and the gap between “works on my machine” and “works for everyone” starts to widen, distorting critical decisions regarding our technology choices and priorities.

Closing the perception gap with objective data is central to our mission of helping build a better Internet, one that's fast and accessible to everyone, not just those of us using the best hardware.

That’s why today, we’re sharing a view that offers insight on how real people experience the web, by publishing the Cloudflare BEACON dataset — Browser Experience Across Cloudflare's Observed Network. Cloudflare has collected this kind of telemetry for years on behalf of our customers, giving them a customer-specific, detailed understanding of how real users experience their sites. Today, we're opening that view up to everyone.

BEACON is an anonymized dataset built from billions of real-world performance measurements across 10,000 of the largest websites on our network. It covers every major browser engine, is updated daily in Google BigQuery, and uses standards defined by the community-led RUM Archive, a publicly available Real User Monitoring (RUM) database. By expanding the footprint of that project 100-fold, BEACON gives researchers an unprecedented view of how the web performs across browsers, devices, and countries.

What BEACON reveals

The Core Web Vitals have long been the de facto standard for measuring perceived performance on the web, and BEACON reports all three:

  • Largest Contentful Paint (LCP): load time
  • Cumulative Layout Shift (CLS): visual stability
  • Interaction to Next Paint (INP): interaction responsiveness

Because we’re publishing these as full histograms rather than single averages, you can derive any percentile you like. Instead of stopping at P75 (the 75th percentile), you can examine the long tail and see where the industry still struggles to deliver fast experiences for everyone.

Who experiences a slower web?

WebKit, currently the only browser engine on iOS, performs best on these metrics overall, but that advantage is not universal. In 46 countries where WebKit represents more than 10% of traffic, its LCP, INP, or both are at least 10% worse than those of Blink-based browsers such as Chrome, Edge, and Opera. In Cambodia, for example, WebKit accounts for 17.5% of page views, but its LCP is 50% worse than Blink’s.

BEACON also includes domain industry classifications from our Intel API. Government and Politics, Health, and Safe for Kids stand out as high-performing categories, while Ads, Religion, and Weather typically perform worst:

The additional percentiles expose differences hidden by a single P75 result. In several industries, the slowest experiences fall sharply in the long tail, particularly for visual stability as measured by Cumulative Layout Shift.

What makes pages feel slow?

BEACON extends the RUM Archive standard with LCP and INP sub-parts that separate the stages of loading and responding to an interaction. We’ll also add these metrics to our Real User Monitoring (RUM) dashboard in the coming weeks. Aggregating them into the suggested ‘Good’, ‘Needs Improvement’, and ‘Poor’ thresholds helps narrow down what needs to be optimized:

LCP Sub-part

Document TTFB
(Time to First Byte)

Nothing can be rendered until we have our HTML document.

Load Delay

Is JavaScript dependence slowing discovery of our LCP candidates?

Load Duration

Is bandwidth an issue, with the LCP image/video/webfont taking a long time to download?

Render Delay

When all is ready, is there something blocking the LCP from rendering?

Good

598ms

76ms

119ms

157ms

Needs Improvement

1015ms

1049ms

199ms

437ms

Poor

1891ms

1485ms

119ms

2002ms

The query for the table above and others for every data is stored in BigQuery alongside the dataset as ‘Global LCP Sub-parts’ so you can customize it as you see fit.

The results challenge a common assumption: downloading the resource itself, such as an image, font, or video, typically contributes the least to perceived loading time. For most page views that breach the ‘Good’ threshold, the larger opportunities are discovering the LCP (load time) candidate and unblocking its render. Cloudflare customers can address some resource-discovery delays with Smart Hints.

The same workflow breaks down Interaction to Next Paint (INP), our measure of responsiveness, into input delay, processing time, and presentation delay:

INP Sub-part

Input Delay

Are our interactions waiting on the main thread becoming available?

Processing Time

Does processing the interactions themselves block the main thread?

Presentation Delay

How long does it take to paint any update to the screen?

Good

18ms

55ms

56ms

Needs Improvement

32ms

112ms

111ms

Poor

84ms

284ms

217ms

Query for above table stored as ‘Global INP Sub-parts’ in BigQuery

For the slowest interactions, JavaScript execution time covers the longest period but we also see a significant rise in presentation time, which is typically dominated by complex CSS layout recalculations. Cloudflare customers can use tools such as Zaraz to reduce the performance impact of third-party JavaScript.

How application architecture changes the picture

Speaking of JavaScript, we recently added support for Google Chrome’s new Soft Navigations API, which provides accurate LCP reporting for client-side navigations commonly used in single-page applications built with frameworks such as React, Vue, Angular, and Svelte.

LCP Percentile

P50

P75

P90

P95

Hard Navigations

791ms

1,421ms

2,636ms

4,122ms

Soft Navigations

274ms

582ms

1,169ms

1,816ms

Query for above table stored as ‘Blink – Hard vs Soft Navigations’ in BigQuery

Soft navigations render two to three times faster than hard navigations at every percentile. But they do not eliminate the cost of the initial landing page, which is often considerably heavier:

LCP Percentile

P50

P75

P90

P95

Landing Page

1,370ms

2,681ms

5,397ms

8,940ms

Query for above table stored as ‘Blink – Landing Pages’ in BigQuery

For teams choosing this architecture, it’s important to be mindful of tradeoffs: faster subsequent navigations must offset a slower first experience. If users rarely progress beyond the landing page, a heavier initial load may never pay for itself.

What can researchers discover by combining data?

Because BEACON is an open dataset, its value is not limited to the fields it contains. Researchers can join it with other sources to explore new questions. For example, combining BEACON with the World Bank Group’s measure of GDP per capita reveals how economic conditions correlate with web performance by country:

Cloudflare Radar, our public data insights and visualizations platform, will start using this approach in a new Web Performance section on Radar, featuring correlations of their Internet Quality Index (IQI) data with BEACON data. IQI is an aggregation of the performance benchmarking data that powers our bi-annual network performance updates.

Pairing BEACON and IQI is particularly compelling because it splits the user experience into its two most influential components: the performance of the website a user is visiting, and the performance of the eyeball network getting them there. Both affect how quickly the page will load, and together they determine whether a visit to a website feels painful or seamless.

Early analysis shows the two tend to move together: good web performance usually coincides with good network quality, and vice versa. In the above graphs we see that the higher the bandwidth, the faster the perceived load time (LCP). For example in Europe, the bandwidth is higher relative to other continents, while the LCP is higher overall with the lower end.

The relationship between bandwidth and LCP was expected, but when comparing IQI to Transfer Size we saw something surprising. We would expect that transfer size would be uniform across continents — after all, the content of the sites are typically the same. However, see that Africa has a noticeably smaller transfer size, suggesting less content being downloaded by these users. Although we can’t be sure of the cause, we can observe that in the IQI data Africa has lower bandwidth, which suggests that businesses on the continent are adapting their websites to optimize towards the constraints of network conditions. High-quality eyeball networks are also potentially more forgiving to poorly-optimized websites, while a slow one exacerbates bottlenecks. Each of these hypotheses are potential directions for future analysis.

The new Web Performance section will explore how common these patterns are, and in particular how often one half of the experience cancels out the gains made by the other. Follow our progress on Radar.

We’ll continue to introduce more metrics and more dimensions over time, and we welcome requests for data you’d like to see next.

How we process all this data and make it useful

Anonymization

Publicizing a real-world dataset at this scale brings with it responsibility for privacy. Our Real User Monitoring (RUM) product is already built to be privacy-first, and for BEACON, we also strip out any potential customer website identifiers such as the domain name and URL paths. Ultimately, the community gains valuable insight without compromising the trust of the people and businesses behind it.

Normalization

For the primary table, including all websites from our data would lead to one of two problems:

  1. The largest sites would dominate the data by traffic volume, skewing performance metrics towards their architecture, visitor profiles etc., or:
  2. If we instead normalize every site down to the traffic volume smallest website, that would drastically reduce the overall number of records in the data.

These two extremes made it necessary to scope the dataset to the greatest number of the largest websites on our network to provide maximum diversity across architectures, technologies, geography, and more, all while retaining the total beacon count after normalizing. We found 10,000 to be a good balance: a globally representative sample with enough volume in the 10,000th that when we normalize the data down to their level, the overall dataset still represents billions of daily records collectively.

Aggregation

Finally, we aggregate records together where they share dimensions such as country, operating system, browser, and connection protocol, and discard any records with fewer than five data points to further guarantee no individual or specific site can be identified.

How to get access and contribute

BEACON is publicly available on Google BigQuery. We’ve included queries for all the data included in this post as examples you can adapt for your own analysis, and the RUM Archive website includes further documentation too.

We’d love to hear what you discover. Share your findings in our community forum or on Discord.

BEACON makes it possible to study web performance at a scale and level of geographic and browser diversity that has not previously been publicly available. We hope researchers, browser vendors, developers, and standards groups use it to identify where the web falls short and help make fast experiences available to everyone.

A special shout out to Cloudflare’s 1,111 intern project. This couldn’t have happened without the hard work of two of our wonderful summer interns, taking the initial idea through to what you see today. Their contributions were instrumental in launching this project. Thanks Chisara Duru and Tong Zhou!

Supporting native Rust in Workers with the new Emscripten target for wasm-bindgen

Post Syndicated from Guy Bedford original https://blog.cloudflare.com/rust-workers-emscripten-target/

Today we’re announcing the first public experimental preview of a feature to better support native Rust code and even Tokio-based applications just running natively on Workers: first-class support for the Emscripten wasm32-unknown-emscripten Rust compiler target on the wasm-bindgen open source toolchain and Cloudflare’s Rust Workers.

wasm-bindgen is the open source toolchain powering Rust-based WebAssembly applications on our V8-based Workers Runtime. Enabling the Emscripten target for wasm-bindgen has been a long-term effort, first initiated by Google over a year ago, and then further reviewed and supported by the Cloudflare engineers maintaining wasm-bindgen.

While still in pre-release, we’re excited to share the new workflow possibilities this work enables in running native wasm-bindgen Rust applications with Emscripten on the web, Node.js, and on Cloudflare’s global Workers platform.

In testing we’ve been able to see significantly improved library compatibility for Rust Workers. To illustrate the sort of capabilities supported, we were able to get a Rust-native Minecraft server (Pumpkin) running inside of a Durable Object with TCP ingress, using real TCP sockets via Tokio. See the end of this post for a full description of this port.

Emscripten is an open-source WebAssembly compiler toolchain initially created by Mozilla and currently maintained by Google engineers, which makes it possible to run and bridge native code with the web platform, including supporting and virtualizing platform features such as timers, file system operations, sockets, and other native functionality.

Since Cloudflare Workers supports Web Platform APIs and Node.js compatibility, we are able to support Emscripten on Workers using its Node.js compilation flags, fully virtualizing native platform features such as timers, file system operations, and sockets on top of our existing Node.js APIs. And with our work on Tokio support, Rust code building on top of Tokio’s async runtime ecosystem can now also be fully integrated into JavaScript-based host environments with this Emscripten target.

We’ve made these current experimental patchsets available with example applications to try out today, including:

Supporting the wasm32-unknown-emscripten target in wasm-bindgen

Mitch Foley on Google’s Portable Toolchains team first encountered the need for Rust WebAssembly toolchains to interoperate with C++ when another internal team at Google was interested in using wasm-bindgen to interface with their JavaScript. The goal was for wasm-bindgen to drive the build and produce the companion JavaScript, while Emscripten’s linker (wasm-ld) would be able to link in any C++ dependencies needed.

Internally, Google uses Emscripten in C++ codebases to generate JavaScript along with the rest of their applications, allowing the C++ and JS to interact with one another. While Emscripten was capable of linking in Rust code as a dependency, the thing it couldn’t do was provide a robust binding system between Rust and JavaScript like wasm-bindgen does.

The problem was both tools assume they are in charge of loading and interacting with JavaScript and generating the final JS and Wasm output. Because of this conflict, Google internal users would have to commit to using one or the other toolchain and never both, and adding a second set of tooling would have doubled the support surface for Google’s Portable Toolchains team.

Mitch and his colleague Yifan Yang crafted a plan to have them work cooperatively: Emscripten would continue to drive the build, load the Wasm module, and provide the companion JS, while wasm-bindgen would produce a smaller portable version of its JavaScript bindings in a format that could be directly included in Emscripten’s library system.

This wasn’t just a technical problem — both Emscripten and the wasm-bindgen maintainers had to support this plan and be willing to maintain integration tests that depended on the other. With feedback from Google’s Portable Toolchains team, Google’s Wasm Tools team, and Cloudflare engineers, this finally resulted in the required changes landing in both projects.

As a result of these efforts, the wasm-bindgen and Emscripten toolchains now have seamless interoperability under the new -sWASM_BINDGEN configuration:

  1. C++ Emscripten code driven by Emscripten’s compiler can be built against static wasm-bindgen Rust code, fully supporting the wasm-bindgen bindings layer alongside the Emscripten bindings layer.
  2. Rust applications using wasm-bindgen and driven by Rust’s compiler can now be built for the Emscripten target, fully supporting Emscripten’s bindings layer alongside wasm-bindgen’s bindings layer.

See the wasm-bindgen Emscripten documentation page for more information about using this target.

Rust library support for Emscripten

After landing the wasm32-unknown-emscripten target support in wasm-bindgen, early prototypes by Cloudflare engineers demonstrated clear success for supporting this target on Cloudflare Workers. Many libraries worked out of the box, even including low-level systems libraries, since Emscripten already supports the target_family = unix in Rust.

Some low-level systems libraries that were unaware of Emscripten required patching, for example libc, socket2, and Mio. Even for these low-level libraries, patches primarily involved adding the Emscripten target to the existing platform gates, for example to explicitly allow target_os = “emscripten”, in place of Wasm platform gates.

Overall we posted these target support patches fairly infrequently, and they were mostly trivial when needed, with the Rust library maintainers very receptive to reviewing the support.

That said, one of the critical foundational libraries that needed significant support was Tokio.

Tokio support

Cloudflare Workers are single-threaded and hosted within a JS event loop, while Tokio async is designed around blocking operations being supported through threaded parking semantics. The two models are clearly not compatible with each other — a blocking operation such as a pending socket read or epoll wait cannot block the shared JS event loop.

To get around this, we had two options: WebAssembly JavaScript Promise Integration or to modify Tokio to support event loop runtime integration.

To support Tokio building for Emscripten, we have contributed full Tokio support patchsets that are being reviewed upstream, with the first target support patch for wasm32-unknown-emscripten already landed upstream in Tokio. With these patchsets, we’ve been able to fully support both approaches in Cloudflare Workers. In our Rust Workers Tokio examples, these patches are required directly pending further integration upstream.

Supporting JSPI

WebAssembly JavaScript Promise Integration (JSPI) maps directly onto Tokio’s existing parking semantics. This is because JSPI allows a blocking Wasm call to suspend the WebAssembly stack on a synchronous operation and return control to the JS event loop, exactly as one would expect of a park.

But when a new Wasm call is made while an earlier one is suspended, JSPI doesn’t break: it instead supports having a new WebAssembly stack being entered while arbitrary existing stacks are suspended at the same time.

The problem here, though, is that Rust itself isn’t aware that its stack is being switched out from under it. Tokio’s runtime context is tracked via a thread local, but a JSPI stack switch is not a thread switch, so the suspended and new stack still share the same thread-local runtime context. As a result, the Tokio runtime context still thinks it is in the parked context when a new Wasm call enters, and then panics because the runtime is already entered.

Supporting fully reentrant JSPI therefore requires careful thread local handling to split context between JSPI context switches. To support this, the thread-local context itself must be swapped on each JSPI enter, exit, suspend, and resume. In effect this is cooperative time-multiplexed threading, with each suspended stack carrying its own runtime context.

With our pre-release Tokio patchset, we were able to implement and verify this model. Finalizing the design and upstreaming it is ongoing in collaboration with the Tokio and Emscripten teams.

Adding an event loop runtime to Tokio

The other runtime approach is full event loop integration. Event loops are of course a common paradigm in native UI applications, so the idea of supporting an event loop runtime for Tokio was certainly not something completely unfamiliar to the maintainers in discussions we had around the Emscripten target support.

The question was rather how to design an event loop for Tokio in such a way that it could work across native applications (including Windows and macOS), as well as for WebAssembly embeddings in JavaScript hosts.

If we could design a general LocalEventLoop runtime for Tokio, we could solve this problem more generally.

A Tokio runtime does two things in a loop: (1) it polls tasks until nothing is ready, and then (2) it waits. The wait is what makes it a runtime rather than a library: the thread parks inside the I/O driver until a socket becomes readable, a timer expires, or another thread wakes it.

But when the host already has its own event loop, Tokio's loop cannot run inside it without blocking the host's. The way around this is to split Tokio's loop in two with both parts able to integrate with the host: (1) becomes an explicit drive() operation that runs one batch of ready tasks and returns, and (2) is replaced by a wake, so that instead of parking, the runtime tells the host it has work and the host calls drive() when it is ready to.

Our proposed LocalEventLoop design for Tokio is a LocalRuntime whose wait has been replaced by a wake. It is built with a standard std::task::Waker that the host owns, which instead of being used to signal that a future should be polled soon, is used by the Tokio runtime to signal that the runtime itself should be driven soon.

Consider the example of a socket read under a regular Tokio runtime:

This would correspond to the following call diagram:

Instead, with LocalEventLoop we can write:

Here, spawn_local returns immediately, with the host event loop taking responsibility for driving the spawned task to completion via el.drive() calls from the host. When stream data is unavailable, the runtime simply returns, handing back control to the host. Once the socket becomes readable, Tokio uses the host_waker to signal that the runtime needs driving.

This LocalEventLoop flow corresponds to the following call diagram:

Everything that would have unparked a native runtime's thread wakes the host instead: a spawn, a task woken from another thread, a socket becoming readable, a timer expiring. Because a Waker is Send + Sync and carries no execution semantics, its implementation is just a notification to the host's event loop. This makes the contract safe by construction: a wake arriving from another thread, from a host callback, or even during a drive, queues work rather than re-entering the runtime. The drive that follows runs on the owning thread with nothing else on the stack.

The one thing you cannot do is wait. block_on still exists and runs a future as far as ready work carries it, but where a normal runtime would park, LocalEventLoop::block_on panics instead. This is because nothing could wake that future from inside the call since its wait belongs to the host.

With this design, the host event loop is never blocked, and interleaves its own work with Tokio's, one batch at a time. The same architecture embeds into a GTK main loop, a Win32 message pump, or a Cocoa run loop. In addition, any number of these runtimes can co-exist due to their cooperative execution semantics.

Supporting sockets and epoll on Emscripten

With both Tokio runtime integration approaches fleshed out for the Rust Emscripten target, we were then able to support most of the Tokio test suite, with one major subsystem still unsupported — the net feature. This includes its sockets APIs: TCP, UDP, and Unix sockets. The reason for this was that Emscripten only supported poll() and a WebSocket emulation layer, but not epoll_wait(), on which Tokio’s I/O driver is built (via mio).

For Cloudflare Workers, we wanted to be able to fully integrate with our TCP sockets APIs, including upcoming inbound TCP. To do this, we would need to build a bridge between Emscripten’s virtualization layer and our own sockets API layer.

Instead of having to build this bridge ourselves, we realized we already have one: the node:net API we support in our Node.js compatibility layer. Emscripten had an -sNODERAWFS mode to bridge natively into Node.js FS APIs, so the same approach could give us a sockets bridge without Emscripten or Cloudflare Workers needing to implement any custom APIs on either side.

We contributed this work in over 40 pull requests to Emscripten, which is now the –sNODERAWSOCKETS layer compilation option, enabling support for epoll, TCP, UDP, and Unix sockets for Emscripten applications in Node.js — and, because Workers implements the same node:net API, on Workers itself as well.

Under JSPI, Emscripten's epoll_wait() simply suspends the stack until readiness, so Tokio's I/O driver works as on native. For LocalEventLoop, readiness instead had to reach the Waker from a JS callback. To support this, we drafted an Emscripten proposal for a new emscripten_epoll_add_listener API to associate a callback on an epoll’s ready events. The next drive() then collects these events with a zero-timeout epoll_wait(), so Tokio's existing I/O driver can be used unchanged.

Running a Minecraft server on Workers

To test out this new target, the challenge was raised to see if a Minecraft server could be deployed to Workers. Dan Lapid then implemented a Pumpkin Minecraft server running on a Durable Object over a single weekend.

Pumpkin is a Rust Minecraft server built on Tokio and designed for multi-core machines. World generation runs on a dedicated thread pool, while the game tick loop and chunk scheduler each run on their own OS threads.

Inside a Durable Object there is exactly one thread, so getting Pumpkin to run there meant turning those threads into cooperative tasks on the event loop. Leaning on the Tokio integration, the tick loop and chunk scheduler became async tasks and each Rayon job became a Tokio task. World generation then runs on the event loop itself, taking one turn per chunk and interleaving with network I/O and game ticks rather than blocking them.

For persistence, Pumpkin writes its world through ordinary std::fs. Under Emscripten’s -sNODERAWFS option file system calls get forwarded to the node:fs Node.js compatibility layer for Workers. Since this bridge is just JavaScript, it is easy to swap out. Dan’s worker-fs-mount library enables mounting a node:fs-compatible file system with its durable-object-fs backend, which stores files as rows in the Durable Object's SQLite storage. With this integrated, every file Pumpkin saves becomes a row in the object's database, written synchronously and committed with the Durable Object's transaction. Pumpkin itself has no idea it isn't on disk, and a restarted object boots straight from the same world.

Networking needed no changes in Pumpkin either. Each player's connection arrives through Workers TCP ingress at the Worker's connect() handler, which forwards it into the Durable Object. There, handleAsNodeConnection() from cloudflare:node dispatches the socket to a net.Server listening on that port inside the object – a new TCP counterpart of handleAsNodeRequest(). Emscripten’s -sNODERAWSOCKETS backend implements TcpListener on top of net.Server, so the server accepts players exactly as it would on Linux, and the returned promise tells the object when the last player has left, so it can save and shut down.

Running a fully persistent Minecraft server inside a Durable Object with multiplayer support demonstrates the level of native compatibility that is possible with this new Emscripten target. We’re excited to see what native Rust applications you can bring to the platform.

Try it out

We’re making these full patchsets and workflows available today for experimental use. Improving wasm-bindgen’s support for modern WebAssembly standards for all users is part of our ongoing commitment to the Rust and Wasm ecosystem.

We welcome all contributions and feedback — find us on GitHub and in the #rust-on-workers channel on Cloudflare’s Discord.

EmDash 1.0: the stable CMS with a secure plugin registry

Post Syndicated from Scott Buscemi original https://blog.cloudflare.com/emdash-cms-plugin-registry/

When we introduced EmDash on April 1 as the “spiritual successor to WordPress”, the buzz was hard to ignore. Walking around a WordPress conference that month, we couldn’t walk far without hearing murmurs about EmDash from fellow attendees.

But alongside the excitement and curiosity has been a seed of doubt among some in the industry. Was this just an April Fools’ joke?

It was not. Today, we are releasing EmDash 1.0: a stable, free, and open source CMS built on Astro, ready to power a production website, your agency’s vibe-coding platform, or your hosting company’s site-building experience.

Developers build with Astro, editors manage content through the EmDash admin, and agents can work through the API, CLI, or built-in MCP server. EmDash 1.0 brings those pieces together with production-tested editorial, media, localization, migration, and deployment workflows.

We are also launching a decentralized plugin registry that lets developers publish without handing ownership of their identity or releases to a central marketplace, while site owners can discover and install their plugins directly from inside EmDash.

The road to 1.0

Since EmDash's first beta, developers have launched real websites with it. Even so, we kept hearing a reasonable response: “This looks interesting. Let me know when it is 1.0.” Before relying on EmDash for their sites, they wanted confidence that it was stable, secure, that upgrades would protect their content, and that we are fully committed to maintaining it.

EmDash 1.0 is our answer to that. For the past five months, we have worked with contributors and production users on the parts of the CMS that every site depends upon: data safety, database migrations, editorial workflows, localization, plugin security, performance, and the reliability of the admin, API, MCP, and media experiences.

Real deployments shaped much of that work, uncovering edge cases and identifying needs that only show up when a site is serving real traffic and being used by real teams of editors.

As one example of this, Avulux converted a custom microsite to EmDash after dealing with the maintenance burden of WordPress for too long. Since EmDash is an agent’s best friend, using the EmDash Agent Skills helped that transition take less than a day to complete.

“The site had to be fast to use and simple for our team to update,” says Greg Barbosa, Director of Innovation and Systems at Avulux. “WordPress had become the opposite of that. With EmDash, we now have a shared platform that developers can extend and marketers can edit content."

In August, we migrated the Cloudflare Blog to EmDash as part of our “Customer Zero” approach. Working alongside our content engineering team provided actionable insights around localization, media management, the admin editor experience, and scaling. Comfortably handling the traffic load for the Cloudflare Blog meant being ready for millions of pageviews per week, spikes up to 5,000 requests per second (RPS) of legitimate traffic, or sporadic DDoS attacks. The optional KV object caching, Hyperdrive database adapter, and Workers Cache compatibility were all features spawned from our migration project that are widely available to all customers now.

Built in public, open to everyone

A CMS sits at the heart of an organization’s web presence. It is trusted with its most important data, and is often used by dozens of editors every day. They need to be able to know they can rely on it, without fear of vendor lock-in or changing business priorities. For that reason, EmDash is completely free and open source, using the flexible and permissive MIT license.

EmDash 1.0 could not exist without its open-source development community. At the time of writing, more than 175 people have contributed to the project, across more than 1,800 commits. The rise of agentic coding tools has presented both challenges and opportunities to open-source projects, and we have deliberately built a project where the agents can help the human developers, rather than being overwhelmed by them.

We are particularly grateful to the core group of the most dedicated contributors, who between them have shipped hundreds of improvements to all areas of the project. They include @swissky, @danielmlr, @MA2153, @marcusbellamyshaw-cell, and dozens of others. Contributors have translated EmDash into 25 languages, from Arabic to Ukrainian.

Particular recognition is due to Noah Pham, who joined Cloudflare as an intern and became EmDash’s second maintainer alongside Matt. Noah contributed more than 80 changes, taking ownership of major parts of the media library, content editor, and admin interface. We have said that interns ship meaningful work at Cloudflare; Noah’s work is now at the heart of EmDash 1.0.

There is plenty more to build, and contributing does not have to mean writing code. If you want to help with code, translations, documentation, testing, design, issue triage, answering questions, or just welcoming new users, join over 800 others in the EmDash community on Discord.

A growing ecosystem

A content management system thrives when the ecosystem around it is healthy and supported. We’ve been excited by the theme companies, plugin shops, agencies, and platforms who are creating new services and products using EmDash.

  • Lexington Themes offers 44 Astro themes with EmDash variants, giving teams a beautiful and functional starting point for their site, complete with reusable components and built-in content collections.
  • Urumi is using their WooCommerce expertise to release EmDash’s first eCommerce plugin.
  • Empress is launching a platform for multi-brand entities, offering the flexibility of managing a fleet of sites with natural language or a conventional CMS admin panel.

“Empress's delightful multisite experience would not be possible without the foundations EmDash has laid: sites that are fast, safe to extend, and easy for people and agents to read and act on,” says Raj Makker, Empress’s founder. “EmDash unlocks powerful control over a website, and Empress builds on it to give you complete control over as many sites as you want. We're excited to be part of this journey.”

A plugin registry that does not own the ecosystem

With EmDash 1.0, developers can publish sandboxed plugins and site owners can discover, inspect, and install them from the EmDash plugin registry.

Traditional plugin registries usually combine three roles: they provide the publisher’s account, hold the authoritative package record, and operate the catalog where users discover it. That is convenient, but it also makes one company the gatekeeper for both identity and distribution. If an account is suspended, a listing is removed, the rules change, or the service shuts down, publishers cannot take the same identity and release history somewhere else.

EmDash separates the plugin from the catalog. Publishers retain control of their packages and release history, while EmDash provides a convenient place for people to find and install them. Other services can index the same publications, build their own catalogs, and apply their own policies without requiring developers to start again.

The EmDash catalog applies default content moderation to the names, descriptions, links, and images it displays. Moderation can hide harmful or inappropriate material from this catalog, but it does not rewrite a release, take ownership of the plugin, or erase the underlying publication.

The registry is built on AT Protocol (atproto), a decentralized network protocol that powers Bluesky and a growing ecosystem of applications. Plugin authors publish with an Atmosphere account, the portable identity also used across these applications. The package and release records are signed by the publisher and stored in the publisher's own account.

EmDash hosts the default registry services so publishers and site owners can use them without running infrastructure themselves. We hope others will build new catalogs, moderation systems, publishing tools, and other services we have not imagined yet. To help this, we have released all of our services as open-source software, including the aggregator that powers the plugin registry, the labeler service that uses Workers AI to moderate package descriptions, and an Astro live content loader to make it easy to include plugin listings in any Astro site. We are excited to see what the community will build with them!

There’s no centralized controller that can take over plugins or remove them on a whim. We believe the security of plugins and the registry should be baked into code, rather than trusting any central authority or code of conduct.

Decentralized publishing does not mean accepting unverified code. Atproto repositories use signed Merkle Search Trees, so an inclusion proof connects the exact release record to a signed commit from the publisher’s account. EmDash can therefore verify the record independently instead of trusting the catalog’s copy. It then checks the plugin’s checksum, package name, version, requested access, and any required build provenance, and confirms that the downloaded bundle matches the signed record.

The registry supports free plugins today, and we aim to add support for paid plugins in future, with the same decentralized model as now. We are watching the Atproto Spaces Alpha with interest. Anybody should be able to run a secure, paid plugin marketplace. We want publishers to be able to make money from their software.

For now, plugin authors can publish useful software without asking permission from EmDash — and site owners can understand exactly what that software is allowed to do before they run it.

Plugins with clear boundaries

Plugins are the core of a CMS ecosystem. They save site developers from rebuilding the same integrations and workflows, while giving experts a way to share their best practices — and potentially build a software business.

Agents make that reuse even more valuable. An agent can build a one-off integration, but it still has to understand the problem, generate and test the code, and maintain it afterward. A plugin captures that work. Another agent can install and configure a proven solution instead of starting again from scratch.

That convenience creates a serious security concern: plugins are code you did not write, operating alongside valuable content and customer data. With WordPress, plugins run inside the same PHP process as the rest of the application, with direct access to its database, filesystem, and network. A contact-form plugin can technically read unpublished posts, modify another plugin, or send data anywhere. Site owners must trust that it will not — and that a future update will not change its behavior.

Sandboxed EmDash plugins use a different model. Each plugin runs in an isolated runtime with access to its own private storage, but not to the site’s content, media, users, secrets, environment, filesystem, or network. It gains additional abilities only when they are declared by the plugin and approved by the site administrator.

Installing a sandboxed plugin therefore feels more like installing a mobile app than a traditional CMS plugin. EmDash shows what the plugin wants to do before it runs, and the runtime limits it to those approved abilities.

That lets plugins perform useful work without receiving unrelated access:

A plugin could…

It might need to…

It still cannot…

Index published articles for search

Read content and contact the search service

Edit articles or contact other hosts

Optimize uploaded images

Read and manage media

Read user content

Send publishing notifications

Observe publishing and send email

Change the content being published

Provide configurable webhooks

Read selected events and contact public destinations

Reach private networks or other plugins’ storage

The important part is that these abilities are independent. Giving a plugin access to media does not also expose users or unpublished content. Allowing it to contact one service does not open the rest of the network. These boundaries are enforced by the runtime, not left to the plugin author’s good intentions.

This isolation is not limited to Cloudflare deployments. On Cloudflare, EmDash runs each plugin as a Dynamic Worker through the Worker Loader. On Node.js, EmDash starts workerd — the open-source Workers runtime — as a separate process and runs each plugin as an isolated service inside it. Plugins use the same manifests and capability-gated APIs on either platform. See the plugin sandbox documentation for setup and runtime differences.

From a single site to a website platform

EmDash can power an individual Astro site, but it is also designed for companies building website creation and hosting products.

Workers for Platforms lets those companies run each customer’s site as a Worker on Cloudflare’s global network. They do not need to provision a server for every site or build the surrounding networking and deployment infrastructure themselves. That leaves them free to focus on the experience their customers use to create and manage a website.

EmDash provides the content layer for that experience. A platform can use the admin interface directly, build its own interface on the API and CLI, or put an agent in front of the built-in MCP server. A bakery owner could update opening hours by asking for the change in plain language; the platform’s agent would handle reading, updating, and saving the content.

The sandboxed plugin model also gives platforms a safer way to offer extensions across many customer sites. Plugins receive only approved access to content, media, users, email, or external services, rather than running with unrestricted access to the whole application. Platforms can run their own plugin marketplaces with access to the full registry — or curate a selection of pre-approved plugins.

We are always looking for additional hosting partners that want to join us in developing new agent-oriented CMS experiences. Reach out to us if you’d like to learn more about reinventing with EmDash.

EmDash Build: an open source AI site builder

We are also releasing and open-sourcing an alpha of EmDash Build, an AI site builder that hosting providers, website builders, and platforms can run themselves and integrate with their own systems. Try the demo today at build.emdashcms.com, or explore the code.

If you’re building a site today, for yourself or for a client, you won’t start in an IDE. You’re more likely going to start with a chat box, and describe to an agent what you want. But once you have something, you could find that changing simple things requires you to go back to that prompt box, and either roll the dice on the result, or burn credits for a one-line change.

EmDash Build creates an EmDash site instead, which brings the full stack: server-rendered Astro pages, a database, media storage, and an admin interface. The agent designs a content model from the brief, fills it in through EmDash's MCP server, and writes the pages that display it. After that, you or your customer can edit text right on the page, schedule posts, or let an agent do it for you.

In EmDash Build, each project gets its own Cloudflare Sandbox container, where the agent, built on the Agents SDK, verifies its own work. Artifacts tracks every change as a git commit. When the site is published, the content moves into a production EmDash site, and deploys to the host's Workers for Platforms namespace.

Get started and get involved

With the release of EmDash 1.0, now is a great time to migrate your company’s marketing site or have an agent spin up that side project you’ve been talking about. Try out the EmDash playground site here. 

To create a new EmDash site locally, via the CLI, run:

Or you can do the same via the Cloudflare dashboard below:

If you’re ready to develop an EmDash plugin, our documentation has a step-by-step guide for creating and publishing a plugin to the registry.

We also welcome you to join our growing community of contributors on Discord. You don’t have to be an engineer to get involved — we welcome translators, issue triage managers, user experience designers, marketers, and all others who are excited about the future of content management systems.

Introducing The Cold Start: pitch your startup live at Cloudflare Connect

Post Syndicated from Fatima Yusuf original https://blog.cloudflare.com/introducing-the-cold-start/

Sixteen years ago, Cloudflare was one of more than 1,000 startups hoping for a spot on stage at TechCrunch Disrupt.

We were not, on the face of it, an obvious choice. Cloudflare was infrastructure: we made websites faster and protected them from attack, which was not well understood by the general market at the time. Infrastructure is often invisible right up until the moment it becomes important.

But on September 27, 2010, Matthew Prince and Michelle Zatlyn got on the Startup Battlefield stage and launched Cloudflare to the public. During the presentation, people started signing up. Then more people started signing up. By the time the judges had finished asking questions, hundreds of websites had joined Cloudflare, putting our initial five data centers to the test in real time. In the seven days following, traffic through our network increased almost 10x and Cloudflare jumped from the 1,000th largest site online to one of the top 50.

Cloudflare didn’t win the main trophy that day. At the awards ceremony, TechCrunch founder Mike Arrington described what we did as something akin to "muffler repair for the Internet" and honestly, he had a point. But then he named us the Most Innovative Company. As Matthew wrote later: "You may not win the trophy, but you'll receive something else far more important."

There are moments in the life of a company when somebody gives you a room, a microphone, and a small amount of time to explain the thing you have spent months or years obsessing over. Most of the time nothing magical happens. Sometimes, though, the right people hear it at the right moment and suddenly an idea that has mostly existed between a handful of people starts moving through the world.

This October, as Cloudflare turns 16, we want to give five early-stage startups a stage of their own.

Introducing The Cold Start

The Cold Start is a live startup competition taking place next month at Cloudflare Connect in San Francisco. We’ll select five early-stage companies and give each of them five minutes on stage to explain what they’re building, why it needs to exist, and why they are the people who should build it.

We are less interested in perfect pitch decks than in interesting ideas clearly explained. You do not need thirty slides, a suspiciously precise TAM calculation, or a rehearsed story about how your childhood prepared you to disrupt accounts receivable. What we want is to understand the vision: what changed in the world that makes it possible, what you see that other people have missed, and why you cannot stop thinking about it.

The five finalists will make their case in front of the Cloudflare Connect audience and three people who've spent a fair amount of their lives thinking about companies, infrastructure, and the Internet:

  • Matthew Prince, co-founder and CEO of Cloudflare
  • Michelle Zatlyn, co-founder and President of Cloudflare
  • Dane Knecht, CTO of Cloudflare

The judges will select one startup to win $500,000 in Cloudflare credits, take over a billboard in San Francisco, and receive an invitation to our VIP speakers dinner that evening.

Five companies, five minutes each, and a room full of people paying attention.

What are we looking for?

The Cold Start is open to ambitious early-stage startups that have raised less than $10 million. Beyond that, we are deliberately keeping the definition broad because the most interesting companies rarely arrive neatly categorized.

We want to see ideas that seem obvious once somebody finally builds them, and ideas that initially sound slightly unreasonable. We want infrastructure that appears boring until you realize everyone is going to need it; products that could not have existed a few years ago; strange new interfaces; new ways of building software; things aimed at enormous existing markets and things aimed at markets nobody has bothered to name yet.

Most of all, we want to meet people who have noticed something about the world and decided to do something about it.

As part of the application, we’ll ask you to tell us who you are, give us your one-line pitch, explain what you’re building and why, tell us about your funding and revenue stage, show us how Cloudflare fits into your stack, and point us toward anything else that helps us understand you and your work.

The goal is simple: make us understand why the thing you’re building should exist.

Apply for The Cold Start.

Applications are open now and close Friday, October 2, 2026. 

Five minutes in San Francisco

The Cold Start will take place Monday, October 19, from 4:00-5:00 p.m. PDT at Moscone West in San Francisco, as part of Cloudflare Connect. Cloudflare will pay for the five finalists to fly to San Francisco for the competition.

Connect brings together people building and thinking about what comes next for the Internet. This year’s lineup includes AI pioneer Dr. Fei-Fei Li; Idealab founder Bill Gross; organizational psychologist and author Adam Grant; AMD CTO Mark Papermaster; Vue.js and Vite creator Evan You; Lovable co-founder and CTO Fabian Hedin; and OpenAI member of technical staff and creator of OpenClaw, Peter Steinberger.

For five young companies, we are reserving part of that stage. Each startup will have 5 minutes to pitch, followed by 3–5 minutes of questions by the judges.

There is something we like about that symmetry. Sixteen years ago, Cloudflare needed someone to take a chance on an infrastructure company with a difficult story to tell and give us a few minutes in front of the right room. Today, we are fortunate enough to have a stage of our own, and we want to pass that same opportunity on to companies that are just getting started.

Start small. Build something enormous.

There is a practical reason Cloudflare spends so much time working with startups: very small groups of people have an uncanny ability to attempt very large things.

The problem is that ambitious software increasingly depends on infrastructure that, not very long ago, only the largest technology companies in the world could afford to build for themselves. Global compute, storage, networking, security, real-time systems, AI inference, and the ability to survive the possibility that the thing you made suddenly becomes popular should not require a company to first become enormous.

We think you should be able to reach for those capabilities on day one.

That is part of the idea behind Cloudflare for Startups, through which eligible early-stage companies can receive up to $350,000 in Cloudflare credits for one year. It is also part of the reason we continue expanding Cloudflare’s developer platform: a tiny team should be able to build something on Tuesday and if the Internet decides it likes it on Wednesday, spend Thursday focused on the product rather than hastily becoming experts in global infrastructure.

Cloudflare began with its own slightly unreasonable premise: that the performance, security, and global compute available to the largest companies on the Internet should be available to everyone on day one. In 2010, we got a chance on a stage to explain why that matters.

Sixteen years later, we have a much larger network, a somewhat larger team, and significantly better circuit breakers.

Now we want to hear what you’re building.

Apply for The Cold Start.

The road to the agentic browser: A Kitesurf update

Post Syndicated from Celso Martinho original https://blog.cloudflare.com/kitesurf-update/

In August, we introduced Kitesurf, a browser for the agentic age that runs entirely on Cloudflare Workers.  We built it around what agents need from the web, rather than carrying all the features and bloat of a browser designed for humans. If this is the first you’re hearing about it, we highly recommend you read the blog post where we introduced Kitesurf, for all the juicy technical details of how we did it.

Since then, we’ve put Kitesurf through increasingly realistic tasks and used internal and external feedback from customers to make it more capable and more efficient. Here’s what has changed, how you can try it today, and where we’re going next.

WebMCP support

Websites were not built for agents to use. Browsing today is a messy process of clicking pixels and hoping the right element loads. In a programmatic world this is slow and fragile. WebMCP helps by allowing developers to expose site functionality directly to agents, where they can call functions like searchFlights() instead of simulating clicks.

Cloudflare has been supporting WebMCP since its early days; just a few weeks ago we announced that site owners can now turn on WebMCP with one switch, so browser agents can discover and use tools on their sites without changing the site’s code, and Browser Run has been supporting WebMCP when using Chrome beta for some time now.

Today we are announcing that Kitesurf now supports WebMCP.

You can test this by going to our public Kitesurf playground, opening Cloudflare Radar, and navigating to WebMCP on the Application tab in the DevTools panel. As you can see, Radar exposes a list of WebMCP tools like navigate-to or set-location which allow clients to interact with the page and explore Radar programmatically.

If you target Kitesurf with your AI Agent:

You can see that the AI model can interact with the exposed WebMCP tools which you can use to complete tasks more reliably.

You can read more about how to use WebMCP with Kitesurf and Browser Run here.

New APIs, better WPT coverage

Since the initial announcement, we’ve been adding more browser standards so that agents can render more sophisticated pages. The list of the APIs that Kitesurf supports has grown, and now includes:

We added URL-based module resolution, JSON modules, and import map handling—important for sites that load JavaScript in chunks. Additionally we are using the new Cloudflare Workers’ module registry to support imports from URLs.

Iframe behavior has improved as well; now they load at the right time, stay better isolated, and display text correctly across more languages and encodings.

As we said at launch, running tests is how we keep the quality of both code and results under control without losing velocity while improving Kitesurf. Web Platform Tests (WPT) is a shared, open-source test suite that checks whether browsers implement web standards consistently.

We now pass 730,000+ WPT subtests and are growing. That’s 500,000 more subtests than when we launched. Here you can see the evolution over time, up to the latest version since we started the project:

Efficiency optimized for agents

For an AI agent, efficiency isn’t so much about loading pages fast, but about the latency of the agentic loop. To make Kitesurf truly agentic, we’ve aggressively optimized the browser engine’s internals so that every DOM traversal, timer, and font fetch is as lightweight as possible, ensuring the agent spends its compute cycles on reasoning, not waiting for the browser to catch up.

These optimizations include:

  • Improved JavaScript execution with less work crossing between Boa and the DOM, making the boundary more compatible with real Web frameworks. Common reads such as getAttribute, id, and parentNode can now be answered inside the Wasm DOM instead of making repeated Boa → JavaScript shim → Wasm trips.
  • Kitesurf does less repeated work when running timers and loading scripts, and releases memory from objects it no longer needs. It also handles objects and classes more consistently when code moves between its two JavaScript engines, resulting in running busy pages more efficiently.
  • Kitesurf now loads fonts when they’re needed, fetches fewer fonts a page won’t use, checks which characters appear on the page before fetching language-specific font files and renders synthetic italics more faithfully.

Together, these optimizations help keep Kitesurf efficient for agents. Despite adding support for more web standards—and bringing Kitesurf closer to the capabilities of full-featured browsers like Chrome—its wall-clock time and CPU usage remain roughly in line with our launch benchmarks, and in some cases they have improved.

Plays better with Browser Run

Browser Run is our developer platform product that lets you programmatically control and run headless browser instances. When you use this API, you can select from a list of browser flavors we support, including Kitesurf.

This means that we have to make sure that all of our browsers are supported across the API surface. Starting today, Kitesurf has full Browser Run API coverage. You can use Kitesurf with CDP, Playwright, Puppeteer, or MCP.

One of the most popular Browser Run features, Quick Actions, provide simple interfaces for common browser tasks like capturing screenshots, extracting HTML content, generating PDFs, and more. When we launched Kitesurf, you could use Quick Actions from our REST API. Now you can also use them from inside a Worker script using the env.BROWSER.quickAction() binding:

Kitesurf runs in the terminal now

As we detailed in the How we built it section of our announcement blog post, Kitesurf separates PageScript, the isolate that handles the page session and runs the page code, from PageRenderer, which is responsible for generating the actual pixels from the computed page objects.

This not only gives great isolation and flexibility, but it also allows us to decouple and move the rendering logic to outside Kitesurf (to the client, for example, or to another Worker), while keeping the security-critical parts server-side, running in our network.

If this model sounds familiar, it may be because Cloudflare has another SASE product called Cloudflare Browser Isolation, which runs all untrusted web code at the edge of our global network while it “streams” the rendering data back to the clients.

We can do something similar with Kitesurf. To prove it, we moved PageRenderer to our Playground Worker and patched this version so that instead of converting scene data into an image, it outputs to Kitty—a terminal graphics protocol supported by modern terminals like Kitty itself, Ghostty, WezTerm, and others. We even went a step further and added a pure ANSI text mode for environments where Kitty isn't available.

The result is that you can now quickly open and render a page using Kitesurf without leaving the comfort of your terminal application. This is super useful not only because you can now browse the modern Web at a glance without context-switching, but you can also use this tool to see how an agent using Kitesurf “sees” a page.

To install the terminal version of Kitesurf, do this:

From now on just type this in terminal:

Here’s a demo of it working.

The terminal also sends back scrolling and click events, so you use the keyboard, arrows, or the mouse normally, as if you were in a dedicated browser application window.

And here is our Silent Space Marine Doom demo running in Kitesurf inside the terminal:

Where we go from here

We continue to iterate rapidly on the road to the best agentic browser for our customers and developers. Expect ever-better performance benchmarks and for the list of supported Web standards and WPT test coverage to continue to rise quickly. In fact, we’ve decided to publish the results here and here, in the open, so that you track them as we move forward, in real time.

We are going to continue exploring scenarios where decoupling Kitesurf and moving PageRenderer away from PageScript is an advantage for agents, or where higher frame rates are important. We may or may not have a 30fps Doom version running in Kitesurf as we write this.

We also want to address the elephant in the room: While we are currently prioritizing rapid development, we remain committed to open-sourcing Kitesurf. This is coming soon, but we want to do this right, so we are set up to support it for the long term.

Kitesurf stays true to its initial design goal: It runs entirely on top of Workers just like any other customer application does; that means we only use our publicly available features and APIs and have no access to any special privileges. This is not only a great way to dogfood and prove our own platform, but also the only way to make Kitesurf very cheap and scale automatically across the Cloudflare global network.

Give Kitesurf a try in the refreshed kitesurf.dev playground, and use it in your projects via Browser Run. It’s available for free while in beta, behind per-account limits. Keep an eye on our changelog and come chat with the team on Discord. Share your experience and send us feedback—we’ll be listening.

Introducing Forge: the open source pipeline for generating SDKs, CLIs, docs, and more

Post Syndicated from Dimitri Mitropoulos original https://blog.cloudflare.com/forge-open-source-generation-pipeline/

Today we’re introducing Forge, a fresh approach to generating SDKs, CLIs, docs, and libraries. Forge is an open source, pluggable generation pipeline that anyone can deploy and run for free.

Forge is early in its life, but already generates the output required for the cf CLI, and over the next few months will power Cloudflare’s API documentation, SDKs, and much more.

We built Forge because we needed it ourselves in order to treat agents as our customers. Now, we’re open sourcing it because we think everyone should be able to generate all the surfaces that agents need. It used to be that only developer products needed CLIs, API SDKs, MCP servers, all with great corresponding docs. Now these are table stakes for every product.

Our API outgrew our generators

Cloudflare’s API has over 3,500 operations, and the hundreds of services that power these APIs are written in many languages, including Rust, Go, TypeScript, and Python. As we embarked on building a CLI for the entire Cloudflare API, including our SDKs and API docs, we needed a code generation pipeline that could handle this scale. That pipeline needs to be flexible enough to work across languages and the ways each of our engineering teams operate.

We needed a way to reduce coordination overhead between teams. When a Cloudflare product team makes an API change, they need to be able to use a preview build of the Cloudflare-wide CLI, SDK, and docs site that will be generated, before merging that change and shipping to customers. We needed a way to ensure they didn’t inadvertently break the generation pipeline. And we needed a system that we could extend to generate more than just an SDK, from Cap‘n Web to MCP and beyond.

We’ve tried several hosted products that attempt to solve this, and relied on some in production. None of them solved this problem for us, and some have shut down entirely. One team would merge a change that inadvertently would break the generation pipeline, another team would discover this at release time, and we spent too much time swimming upstream through hosted tools we couldn’t control, coordinating changes between teams and vendors.

That’s how we started building Forge.

Forge seeks to fix all these problems: it runs in CI, on each team’s API repos, just like our AI code reviewer and test pipelines. It lints every change, and then generates preview builds of the CLI, docs, and SDKs with just your changes highlighted that you can install to test. It’s the same premise as Workers Previews: a full preview build for every change, but applied to SDK generation at scale, including when the API surface is distributed across hundreds of services and repositories. That’s what Forge seeks to deliver.

Forge transformers can generate anything, including Cap’n Web

Cloudflare has more reasons than most to want a generator that can go well beyond the normal language targets. Cap’n Web is Cloudflare’s RPC system that lets TypeScript call a remote API as if it were calling a local method:

Forge makes it possible to take an OpenAPI spec and generate Cap’n Web directly. This opens the door to generating bindings from Workers to other APIs. After all, bindings in the Workers runtime are implemented as Workers that expose RPC methods.

This isn’t specific to Cap’n Web: other popular tools you may already rely on need the same thing. If you use TanStack Query, you’d ideally want to be able to generate TanStack Query bindings for your application, built directly from your API itself. Always up to date, always validated against your real API. The same is true of generating Zod or Valibot schemas, MCP servers, or anything else that makes it easier to consume your API.

This is possible because Forge code generators are flexible. They’re built for flowing information from one output to another.

Forge transformers can be chained: generate outputs from other outputs

We’ve designed Forge to be pluggable, and support many input and output types. Forge provides CLI, SDK and docs generators, but there’s nothing stopping you from adding a transformer that generates a library-specific package or even a full dashboard or application. Forge supports OpenAPI as an input type today, but we’ve designed it to allow AsyncAPI, GraphQL, Cap’n Proto, Protobuf, or other input formats in the future.

This is about more than just compatibility: it lets you chain targets, using one target output to produce others. This is common in other generators where the CLI and Terraform targets are produced from the Go SDK. But what’s missing, and what Forge provides, is a way for the user to control this chaining system themselves.

We needed a solution for this ourselves, because our own cf CLI is written in TypeScript, which other SDK generators don't generally chain from for CLIs. But our own situation made us recognize the deeper problem: why should any SDK generator tool make this decision for you? Maybe you’re a Python shop, and you want the CLI to be in Python.

If you’re thinking “Well, but who cares if it’s in Python or not? The code is automatically generated,” it’s because CLIs are different. CLIs often introduce local-only behaviors that wouldn’t make sense in an SDK. Behaviors that you write by hand since they’re inherently not something backed by any API call. For example, the cf CLI has commands like cf dev and cf build that are added on top of the rest of the generated output. These commands need to call TypeScript APIs from other packages like Vite. 

Now let’s add docs to the mix. If you’re generating your CLI and your docs purely from your OpenAPI spec, how do you feed those handwritten commands back into your docs, so they can be documented alongside the rest?

We couldn’t find an existing tool that does this today, and yet this is exactly what we need for cf. So we’re building it into Forge.

Change your API without breaking users

Forge is also setting us up for better API versioning. Cloudflare’s v4 API has been the one major version of our API for 10 years. Since then, it appears like we haven’t launched any new major versions, but by SemVer definitions we’ve made quite a few changes worthy of a new major version. At the same time, several operations across our API feature internal ‘v2’ tags or ‘beta’ identifiers that have long outlived that part of the product’s lifecycle.

After so many years of our v4 API, we’re keenly aware that a big new v5 would leave a lot of our customers behind. That’s why, with Forge releasing artifacts along the way, we’re working on an API versioning approach that allows us to release new major API versions without breaking old clients or SDKs.

We’ll have more on our SDKs very soon, including TypeScript, Rust, Python, Go, PHP, and Terraform. Especially Terraform. We know that upgrading any Terraform provider comes with its own set of rigor, and we’re going to put extra special care into the Terraform transition.

Critical tools should be open to all

We believe that building tools for APIs is a core part of the Internet, and you should be able to do that without needing a SaaS product. You should own your SDKs, CLIs, and docs. And if you generate them, then you should be able to do whatever you want, wherever you want, for free.

That’s why we’re making Forge available open source under the permissive Apache 2.0 license. We want people to join in on this journey with us, and contribute.

Or not? Maybe you want to keep everything to yourself. Go for it! You can run Forge on your own for any purpose, with custom modifications, for free, in private.

Acknowledgements: This project was also made possible by the design and implementation efforts of Dan Carter, Steven Chong, Krishna Paritala, and Shelley Jones.

Four months of VoidZero at Cloudflare: making the open-source JavaScript toolchain faster for all humans and agents

Post Syndicated from Evan You original https://blog.cloudflare.com/voidzero-update/

When VoidZero joined Cloudflare four months ago, we made a commitment to open source, promising that Vite, Vitest, Rolldown, Oxc, and Vite+ will stay open source, vendor-agnostic, and community-driven.  As part of Cloudflare’s Birthday Week, where Cloudflare gives back to the Internet, we thought it’d be a good time to check in on how we’ve been doing against this commitment.

In the four months since VoidZero joined Cloudflare, we’ve shipped more than 80 releases, closed over 1,200 issues, and landed some big performance improvements, including:

On top of this, Vite+, which unifies the entire toolchain with a set of great defaults, is now 1.0. And we’re making progress towards shipping “Bundled Dev” (f.k.a. Full Bundle Mode) — a new development mode that has been shaped by working with customers with massive web applications, including Cloudflare’s own dashboard. By being part of Cloudflare, our engineering team gets a much better view into scenarios that only occur in massive codebases.

VoidZero’s mission is to make the next generation of JavaScript developers more productive than ever before. And in 2026, that means making developers’ agents faster.

A faster developer experience for humans and agents

Developer experience and performance were always about improving the feedback loop while creating software. But now when we optimize for developer experience, we no longer just optimize for humans — we optimize for agents. And as inference gets faster, the speed of type-checking, linting, or building the code becomes the bottleneck again. The longer those processes take, the longer an agent has to wait before making progress and completing its goals.

VoidZero has always been focused on performance, but this new era of software has given us even more reason and motivation to make the entire toolchain faster for both agents and humans. And since the tools in the VoidZero toolchain are built on each other, from the compiler (Oxc) to the bundler (Rolldown) to the build tool (Vite), the linter (Oxlint) and the test runner (Vitest), any optimization at one layer automatically benefits everything built on top.

Oxc React Compiler — 10x faster compile times for React.js apps

We’ve recently shipped the Oxc React Compiler, a rewrite of the React Compiler, based on the React team’s rewrite in Rust. It is 10x faster than the original Babel implementation, uses less memory, and has a more complete implementation with better error handling. If you use Vite, you can enable the Oxc React Compiler by installing the oxc-transform-react package and enabling the compiler flag:

Vitest 5 — up to 50% faster than Vitest 4

Vitest 5, released in September, cuts test times by double-digit percentages across many common scenarios.

  • Faster test runs. Vitest shares transformed files across projects, caches modules on disk, and sends less data between its main process and workers.
  • vitest doctor. It breaks down setup, import, transform, and test time, then tests other configurations and recommends faster settings.
  • Trace View. It records browser interactions, assertions, and DOM snapshots, so you can replay failures or inspect them in an HTML report.
  • Conditional mocks with vi.when. Map arguments to responses without writing a manual mockImplementation.
  • Better benchmarks. Benchmarks now work like regular tests, with fixtures, hooks, retries, filters, and assertions.
  • Fewer false passes. Vitest fails unawaited async assertions, clears mock calls before each test, and adds the --repeats flag to help find intermittent failures.

tsgolint — now stable, up to 18x faster than ESLint in large codebases

tsgolint, the type-aware linting engine behind Oxlint, is now stable. It catches bugs that require TypeScript type data while running 12 to 18 times faster than ESLint with typescript-eslint.

Enable type-aware linting and TypeScript diagnostics in your Oxlint config:

Oxfmt — 7x faster than Prettier with formatters now written in Rust

Oxfmt brings the speed of the Oxc toolchain to formatting. We rewrote its JSON, CSS, SCSS, Less, GraphQL, and YAML formatters in Rust, making Oxfmt many times faster than Prettier, while keeping Prettier-compatible output and ergonomics.

Bundled Dev — faster dev server for larger apps

We’ve made progress towards shipping “Bundled Dev” (f.k.a. Full Bundle Mode), which uses Vite’s production bundler during development. This should lead to significant dev server speed-ups in larger applications and reduce network overhead when working with remote sandboxes.

We’re looking forward to bringing Bundled Dev out of experimental status soon. Cloudflare’s Dashboard already uses Bundled Dev for all internal developers.

Vite+ is now 1.0 — a unified toolchain

Speeding up tools is one way of making humans and agents ship software faster. Another way is by reducing decision fatigue (“Which linter shall I use?”) and providing great defaults. We are excited to announce that Vite+ is now 1.0. Vite+ bundles all of VoidZero’s tools together into a single unified and fast toolchain.

Vite+ ships with Vite 8, Vitest 5, Rolldown, Oxlint, Oxfmt and task caching built in, and comes with great defaults. Check out the Getting Started guide to try it out today.

Cloudflare’s Open Source Investment

VoidZero was born in open source, and we believe a more sustainable open-source ecosystem benefits everyone. VoidZero is a proud member of the Open Source Pledge and as part of Cloudflare, we have the opportunity to expand that impact.

Cloudflare committed $1M to a Vite ecosystem fund to support maintainers and contributors in the original announcement. Since then, Cloudflare has committed an additional $1M to open source. Together, we’re doubling down on open source.

The open-source toolchain for the entire JavaScript community

There is more to come! Features we plan to ship in the next few months include major improvements to Oxc’s parser with up to 3x potential speedup, a re-designed chunking algorithm in Rolldown, and an open-source, self-hostable version of Void, the Vite-native deployment platform built on top of Cloudflare. Cloudflare and VoidZero both recognize our responsibility that we have to developers, and we do not take the community’s trust in us for granted. We’re in this for the long haul.

We are excited to keep shipping faster tools, and will continue to make every decision with the community in mind, just like we did when raising venture capital, or when we open sourced Vite+, or when we joined Cloudflare. Thank you for continuing to trust us with your projects and apps, supporting us, and building with us.

Cloudflare’s 2026 Annual Founders’ Letter

Post Syndicated from Matthew Prince original https://blog.cloudflare.com/cloudflares-2026-annual-founders-letter/

This week Cloudflare celebrates our 16th birthday. Like many 16-year-olds, we find ourselves looking around at the world we grew up in and feeling caught between the past and future. Also like many 16-year-olds, we look at all the change in the world and sometimes see risk. But, on balance, we come out incredibly optimistic for what lies ahead. What drives our optimism and fear is a recognition that change involves disruption. 

The Internet is changing more today than at any point since Cloudflare launched back on September 27, 2010. Some of that change seems undoubtedly good. Some of it is upending the way we think the Internet works.

One significant change is the rate of growth of the web itself. From 2012 through 2025, the web plateaued and by some measures even shrank. That changed in mid-2025 when there was an explosion of new websites. The popular narrative is that growth has been driven by AI "slop." And, while there is some of that, that's not the majority of what we see.

Instead, AI has unleashed a new cohort of creators. Individuals with ideas but without coding skills who, aided by so-called "vibe coding" tools, are able to bring new creations to life. We're proud that the majority of these tools have Cloudflare's developer platform as their preferred deployment target. Technology at its best allows more people to express their creativity. We have front row seats to watch students around the world building apps to solve real-world problems. Startups with new business ideas getting built in record time. Today more than 7 million developers are building the future on Cloudflare’s developer platform.

The last year has also changed in terms of who — and increasingly what — is using the Internet. We originally forecast that automated traffic would pass human traffic in the second half of 2027. The rise of agents and AI crawlers pulled that date forward to May of 2026. And, if current trends continue — which, if anything, seems conservative — automated traffic will be 1,000 times human traffic in just five years. Not because we believe human traffic will decline, but because traffic from agents is exploding.

For people using them, these agents are already astonishing. Ask one to find you a flight, a contractor or a cheaper phone plan, and it will read more pages in a minute than you could in an afternoon, then come back with an answer. It does the legwork so you don't have to.

But that legwork isn’t free. If you ask your AI agent to suggest where you should go to lunch, it may survey 1,000 restaurant menus in the area only to suggest one. The one restaurant may get your business, but the other 999 had to shoulder the load of serving the agent without getting anything in return. The risk here is a tragedy of the commons problem where the users benefiting from the agents and their traffic don't bear the costs of the load they put on the system and therefore don't exercise appropriate restraint.

AI has made it easier than ever for those students and startups to build something. Agents could make it much harder for anyone to find it. Today, small businesses win customers with emotion or convenience. You patronize a certain deli because the person working the counter remembers your name or because it’s part of how you define your identity. Or you shop at a local store even though you know it doesn't have the best selection or prices but it's on your drive home.

Your agent doesn't care who remembers your name, and it isn't driving home past the local store. It goes with whatever it has the most information about, and that's usually whoever has been around the longest. The risk then is that as agents handle more and more commerce, they'll make it harder for new entrants to break in. That, in turn, is likely to lead to consolidation and a less robust business environment.

We were a new entrant once. Sixteen years ago this week, we launched Cloudflare on stage at TechCrunch Disrupt while our engineers sat in the audience fixing bugs. There were eight left when we walked on. By the time we walked off, they were fixed and we were live in five data centers on three continents. Nobody's agent would have recommended us. People took a chance on us anyway. We want the next new entrants to get the same chance.

That’s what we’re playing for. Not a future with five AI companies, but one with 500,000, spread around the world. Not one where content creators wither away and die because they can't be compensated, but one where anyone can create, reach a global audience and get paid for their work. And not one where a few mega corporations win by default, but one where new entrants with better products can serve customers well, and win.

Last year we wrote about what AI was doing to publishers. This year the same shift is reaching the restaurant and the deli. What replaces the Internet's old business model is the most interesting question of the next five years. This week, we take another swing at answering it.

Like every birthday, we're celebrating by giving presents instead of getting them, and some of the presents are for the 999 restaurants. Agents crawl the web the way search engines always have: everything, over and over, whether it's changed or not. Our data suggests more than half of what good bots fetch hasn't changed since their last visit. We've been working so that crawlers see more of the web while fetching only what's new, which cuts the load on the sites they crawl. And we're giving anyone who puts content or applications online a way to get paid when agents use what they've built.

Cloudflare's mission isn't to build a better Internet, it's to help build a better Internet. That means we can't do it alone. So this week we'll also be announcing partnerships with companies and organizations that want the same future we do.

Like other 16-year-olds in this new era of AI, we get to ship our opinions. We're shipping them for an Internet that still has room for a teenager launching their first app, the deli where they remember your name, and whoever is building the next Cloudflare.

And we’ve never been more excited about it.

15 years of helping build a better Internet: a look back at Birthday Week 2025

Post Syndicated from Nikita Cano original https://blog.cloudflare.com/birthday-week-2025-wrap-up/

Cloudflare launched fifteen years ago with a mission to help build a better Internet. Over that time the Internet has changed and so has what it needs from teams like ours.  In this year’s Founder’s Letter, Matthew and Michelle discussed the role we have played in the evolution of the Internet, from helping encryption grow from 10% to 95% of Internet traffic to more recent challenges like how people consume content. 

We spend Birthday Week every year releasing the products and capabilities we believe the Internet needs at this moment and around the corner. Previous Birthday Weeks saw the launch of IPv6 gateway in 2011,  Universal SSL in 2014, Cloudflare Workers and unmetered DDoS protection in 2017, Cloudflare Radar in 2020, R2 Object Storage with zero egress fees in 2021,  post-quantum upgrades for Cloudflare Tunnel in 2022, Workers AI and Encrypted Client Hello in 2023. And those are just a sample of the launches.

This year’s themes focused on helping prepare the Internet for a new model of monetization that encourages great content to be published, fostering more opportunities to build community both inside and outside of Cloudflare, and evergreen missions like making more features available to everyone and constantly improving the speed and security of what we offer.

We shipped a lot of new things this year. In case you missed the dozens of blog posts, here is a breakdown of everything we announced during Birthday Week 2025. 

Monday, September 22

What

In a sentence …

Help build the future: announcing Cloudflare’s goal to hire 1,111 interns in 2026

To invest in the next generation of builders, we announced our most ambitious intern program yet with a goal to hire 1,111 interns in 2026.

Supporting the future of the open web: Cloudflare is sponsoring Ladybird and Omarchy

To support a diverse and open Internet, we are now sponsoring Ladybird (an independent browser) and Omarchy (an open-source Linux distribution and developer environment).

Come build with us: Cloudflare’s new hubs for startups

We are opening our office doors in four major cities (San Francisco, Austin, London, and Lisbon) as free hubs for startups to collaborate and connect with the builder community.

Free access to Cloudflare developer services for non-profit and civil society organizations

We extended our Cloudflare for Startups program to non-profits and public-interest organizations, offering free credits for our developer tools.

Introducing free access to Cloudflare developer features for students

We are removing cost as a barrier for the next generation by giving students with .edu emails 12 months of free access to our paid developer platform features.

Cap’n Web: a new RPC system for browsers and web servers

We open-sourced Cap’n Web, a new JavaScript-native RPC protocol that simplifies powerful, schema-free communication for web applications.

A lookback at Workers Launchpad and a warm welcome to Cohort #6

We announced Cohort #6 of the Workers Launchpad, our accelerator program for startups building on Cloudflare.

Tuesday, September 23

What

In a sentence …

Building unique, per-customer defenses against advanced bot threats in the AI era

New anomaly detection system that uses machine learning trained on each zone to build defenses against AI-driven bot attacks. 

Why Cloudflare, Netlify, and Webflow are collaborating to support Open Source tools

To support the open web, we joined forces with Webflow to sponsor Astro, and with Netlify to sponsor TanStack.

Launching the x402 Foundation with Coinbase, and support for x402 transactions

We are partnering with Coinbase to create the x402 Foundation, encouraging the adoption of the x402 protocol to allow clients and services to exchange value on the web using a common language

Helping protect journalists and local news from AI crawlers with Project Galileo

We are extending our free Bot Management and AI Crawl Control services to journalists and news organizations through Project Galileo.

Cloudflare Confidence Scorecards – making AI safer for the Internet

Automated evaluation of AI and SaaS tools, helping organizations to embrace AI without compromising security.

Wednesday, September 24

What

In a sentence …

Automatically Secure: how we upgraded 6,000,000 domains by default

Our Automatic SSL/TLS system has upgraded over 6 million domains to more secure encryption modes by default and will soon automatically enable post-quantum connections.

Giving users choice with Cloudflare’s new Content Signals Policy

The Content Signals Policy is a new standard for robots.txt that lets creators express clear preferences for how AI can use their content.

To build a better Internet in the age of AI, we need responsible AI bot principles

A proposed set of responsible AI bot principles to start a conversation around transparency and respect for content creators’ preferences.

Securing data in SaaS to SaaS applications

New security tools to give companies visibility and control over data flowing between SaaS applications.

Securing today for the quantum future: WARP client now supports post-quantum cryptography (PQC)

Cloudflare’s WARP client now supports post-quantum cryptography, providing quantum-resistant encryption for traffic. 

A simpler path to a safer Internet: an update to our CSAM scanning tool

We made our CSAM Scanning Tool easier to adopt by removing the need to create and provide unique credentials, helping more site owners protect their platforms.

Thursday, September 25

What

In a sentence …

Every Cloudflare feature, available to everyone

We are making every Cloudflare feature, starting with Single Sign On (SSO), available for anyone to purchase on any plan. 

Cloudflare’s developer platform keeps getting better, faster, and more powerful

Updates across Workers and beyond for a more powerful developer platform – such as support for larger and more concurrent Container images, support for external models from OpenAI and Anthropic in AI Search (previously AutoRAG), and more. 

Partnering to make full-stack fast: deploy PlanetScale databases directly from Workers

You can now connect Cloudflare Workers to PlanetScale databases directly, with connections automatically optimized by Hyperdrive.

Announcing the Cloudflare Data Platform

A complete solution for ingesting, storing, and querying analytical data tables using open standards like Apache Iceberg. 

R2 SQL: a deep dive into our new distributed query engine

A technical deep dive on R2 SQL, a serverless query engine for petabyte-scale datasets in R2.

Safe in the sandbox: security hardening for Cloudflare Workers

A deep-dive into how we’ve hardened the Workers runtime with new defense-in-depth security measures, including V8 sandboxes and hardware-assisted memory protection keys.

Choice: the path to AI sovereignty

To champion AI sovereignty, we’ve added locally-developed open-source models from India, Japan, and Southeast Asia to our Workers AI platform.

Announcing Cloudflare Email Service’s private beta

We announced the Cloudflare Email Service private beta, allowing developers to reliably send and receive transactional emails directly from Cloudflare Workers.

A year of improving Node.js compatibility in Cloudflare Workers

There are hundreds of new Node.js APIs now available that make it easier to run existing Node.js code on our platform. 

Friday, September 26

What

In a sentence …

Cloudflare just got faster and more secure, powered by Rust

We have re-engineered our core proxy with a new modular, Rust-based architecture, cutting median response time by 10ms for millions. 

Introducing Observatory and Smart Shield

New monitoring tools in the Cloudflare dashboard that provide actionable recommendations and one-click fixes for performance issues.

Monitoring AS-SETs and why they matter

Cloudflare Radar now includes Internet Routing Registry (IRR) data, allowing network operators to monitor AS-SETs to help prevent route leaks.

An AI Index for all our customers

We announced the private beta of AI Index, a new service that creates an AI-optimized search index for your domain that you control and can monetize.

Introducing new regional Internet traffic and Certificate Transparency insights on Cloudflare Radar

Sub-national traffic insights and Certificate Transparency dashboards for TLS monitoring.

Eliminating Cold Starts 2: shard and conquer

We have reduced Workers cold starts by 10x by implementing a new “worker sharding” system that routes requests to already-loaded Workers.

Network performance update: Birthday Week 2025

The TCP Connection Time (Trimean) graph shows that we are the fastest TCP connection time in 40% of measured ISPs – and the fastest across the top networks.

How Cloudflare uses performance data to make the world’s fastest global network even faster

We are using our network’s vast performance data to tune congestion control algorithms, improving speeds by an average of 10% for QUIC traffic.

Come build with us!

Helping build a better Internet has always been about more than just technology. Like the announcements about interns or working together in our offices, the community of people behind helping build a better Internet matters to its future. This week, we rolled out our most ambitious set of initiatives ever to support the builders, founders, and students who are creating the future.

For founders and startups, we are thrilled to welcome Cohort #6 to the Workers Launchpad, our accelerator program that gives early-stage companies the resources they need to scale. But we’re not stopping there. We’re opening our doors, literally, by launching new physical hubs for startups in our San Francisco, Austin, London, and Lisbon offices. These spaces will provide access to mentorship, resources, and a community of fellow builders.

We’re also investing in the next generation of talent. We announced free access to the Cloudflare developer platform for all students, giving them the tools to learn and experiment without limits. To provide a path from the classroom to the industry, we also announced our goal to hire 1,111 interns in 2026 — our biggest commitment yet to fostering future tech leaders.

And because a better Internet is for everyone, we’re extending our support to non-profits and public-interest organizations, offering them free access to our production-grade developer tools, so they can focus on their missions.

Whether you’re a founder with a big idea, a student just getting started, or a team working for a cause you believe in, we want to help you succeed.

Until next year

Thank you to our customers, our community, and the millions of developers who trust us to help them build, secure, and accelerate the Internet. Your curiosity and feedback drive our innovation.

It’s been an incredible 15 years. And as always, we’re just getting started!

An AI Index for all our customers

Post Syndicated from Celso Martinho original https://blog.cloudflare.com/an-ai-index-for-all-our-customers/

Today, we’re announcing the private beta of AI Index for domains on Cloudflare, a new type of web index that gives content creators the tools to make their data discoverable by AI, and gives AI builders access to better data for fair compensation.

With AI Index enabled on your domain, we will automatically create an AI-optimized search index for your website, and expose a set of ready-to-use standard APIs and tools including an MCP server, LLMs.txt, and a search API. Our customers will own and control that index and how it’s used, and you will have the ability to monetize access through Pay per crawl and the new x402 integrations. You will be able to use it to build modern search experiences on your own site, and more importantly, interact with external AI and Agentic providers to make your content more discoverable while being fairly compensated.

For AI builders—whether developers creating agentic applications, or AI platform companies providing foundational LLM models—Cloudflare will offer a new way to discover and retrieve web content: direct pub/sub connections to individual websites with AI Index. Instead of indiscriminate crawling, builders will be able to subscribe to specific sites that have opted in for discovery, receive structured updates as soon as content changes, and pay fairly for each access. Access is always at the discretion of the site owner.

From the individual indexes, Cloudflare will also build an aggregated layer, the Open Index, that bundles together participating sites. Builders get a single place to search across collections or the broader web, while every site still retains control and can earn from participation. 

Why build an AI Index?

AI platforms are quickly becoming one of the main ways people discover information online. Whether asking a chatbot to summarize a news article or find a product recommendation, the path to that answer almost always starts with crawling original content and indexing or using that data for training. However, today, that process is largely controlled by platforms: what gets crawled, how often, and whether the site owner has any input in the matter.

Although Cloudflare now offers to monitor and control how AI services respect your access policies and how they access your content, it’s still challenging to make new content visible. Content creators have no efficient way to signal to AI builders when a page is published or updated. On the other hand, for AI builders, crawling and recrawling unstructured content is costly, wastes resources, especially when you don’t know the quality and cost in advance.

We need a fairer and healthier ecosystem for content discovery and usage that bridges the gap between content creators and AI builders.

How AI Index will work

When you onboard a domain to Cloudflare, or if you have an existing domain on Cloudflare, you will have the choice to enable an AI Index. If enabled, we will automatically create an AI-optimized search index for your domain that you own and control.


As your site updates and grows, the index will evolve with it. New or updated pages will be processed in real-time using the same technology that powers Cloudflare AI Search (formerly AutoRAG) and its Website as a data source. Best of all, we will manage everything; you won’t have to worry about each individual component of compute, storage resources, databases, embeddings, chunking, or AI models. Everything will happen behind the scenes, automatically.

Importantly, you will have control over what content to include or exclude from your website’s index, and who can get access to your content via AI Crawl Control, ensuring that only the data you want to expose is made searchable and accessible. You also will be able to opt out of the AI Index completely; it will all be up to you.

When your AI Index is set up, you will get a set of ready-to-use APIs:                                                                                                                                                   

  • An MCP Server: Agentic applications will be able to connect directly to your site using the Model Context Protocol (MCP), making your content discoverable to agents in a standardized way. This includes support for NLWeb tools, an open project developed by Microsoft that defines a standard protocol for natural language queries on websites.

  • A flexible search API: This endpoint will return relevant results in structured JSON. 

  • LLMs.txt and LLMs-full.txt: Standard files that provide LLMs with a machine-readable map of your site, following emerging open standards. These will help models understand how to use your site’s content at inference time. An example of llms.txt exists in the Cloudflare Developer Documentation.

  • A bulk data API: An endpoint for transferring large amounts of content efficiently, available under the rules you set. Instead of querying for every document, AI providers will be able to ingest in one shot.

  • Pub-sub subscriptions: AI platforms will be able to subscribe to your site’s index and receive events and content updates directly from Cloudflare in a structured format in real-time, making it easy for them to stay current without re-crawling.

  • Discoverability directives: In robots.txt and well-known URIs to allow AI agents and crawlers visiting your site to discover and use the available API automatically.


The index will integrate directly with AI Crawl Control, so you will be able to see who’s accessing your content, set rules, and manage permissions. And with Pay per crawl and x402 integrations, you can choose to directly monetize access to your content. 

A feed of the web for AI builders

As an AI builder, you will be able to discover and subscribe to high-quality, permissioned web data through individual site’s AI indexes. Instead of sending crawlers blindly across the open Internet, you will connect via a pub/sub model: participating websites will expose structured updates whenever their content changes, and you will be able to subscribe to receive those updates in real-time. With this model, your new workflow may look something like this:

  1. Discover websites that have opted in: Browse and filter through a directory of websites that make their indexes available through Cloudflare.

  2. Evaluate content with metadata and metrics: Get content metadata information on various metrics (e.g., uniqueness, depth, contextual relevance, popularity) before accessing it.

  3. Pay fairly for access: When content is valuable, platforms can compensate creators directly through Pay per crawl. These payments not only enable access but also support the continued creation of original content, helping to sustain a healthier ecosystem for discovery.

  4. Subscribe to updates: Use pub-sub subscriptions to receive events about changes made by the website, so you know when to retrieve or crawl for new content without wasting resources on constant re-crawling. 

By shifting from blind crawling to a permissioned pub/sub system for the web, AI builders save time, cut costs, and gain access to cleaner, high-quality data while content creators remain in control and are fairly compensated.

The aggregated Open Index

Individual indexes provide AI platforms with the ability to access data directly from specific sites, allowing them to subscribe for updates, evaluate value, and pay for full content access on a per-site basis. But when builders need to work at a larger scale, managing dozens or hundreds of separate subscriptions can become complex. The Open Index will provide an additional option: a bundled, opt-in collection of those indexes, featuring sophisticated features such as quality, uniqueness, originality, and depth of content filters, all accessible in one place.


The Open Index is designed to make content discovery at scale easier:

  • Get unified access: Query and retrieve data across many participating sites simultaneously. This reduces integration overhead and enables builders to plug into a curated collection of data, or use it as a ready-made web search layer that can be accessed at query time.

  • Discover broader scopes: Work with topic-specific bundles (e.g., news, documentation, scientific research) or a general discovery index covering the broader web. This makes it simple to explore new content sources you may not have identified individually.

  • Bottom-up monetization: Results still originate from an individual site’s AI index, with monetization flowing back to that site through Pay per crawl, helping preserve fairness and sustainability at scale.

Together, per-site AI indexes and the Open Index will provide flexibility and precise control when you want full content from individual sites (i.e., for training, AI agents, or search experiences), and broad search coverage when you need a unified search across the web.

How you can participate in the shift

With AI Index and the Cloudflare Open Index, we’re creating a model where websites decide how their content is accessed, and AI builders receive structured, reliable data at scale to build a fairer and healthier ecosystem for content discovery and usage on the Internet.

We’re starting with a private beta. If you want to enroll your website into the AI Index or access the pub/sub web feed as an AI builder, you can sign up today.

Introducing Observatory and Smart Shield — see how the world sees your website, and make it faster in one click

Post Syndicated from Tim Kadlec original https://blog.cloudflare.com/introducing-observatory-and-smart-shield/

Modern users expect instant, reliable web experiences. When your application is slow, they don’t just complain — they leave. Even delays as small as 100 ms have been shown to have a measurable impact on revenue, conversions, bounce rate, engagement and more. 

If you’re responsible for delivering on these expectations to the users of your product, you know there are many monitoring tools that show you how visitors experience your website, and can let you know when things are slow or causing issues. This is essential, but we believe understanding the condition is only half the story. The real value comes from integrating monitoring and remedies in the same view, giving customers the ability to quickly identify and resolve issues.

That’s why today, we’re excited to launch the new and improved Observatory, now in open beta. This monitoring and observability tool goes beyond charts and graphs, by also telling you exactly how to improve your application’s performance and resilience, and immediately showing you the impact of those changes. And we’re releasing it to all subscription tiers (including Free!), available today.


But wait, there’s more! To make your users’ experience in Cloudflare even faster, we’re launching Smart Shield, available today for all subscription tiers. Using Observatory, you can pinpoint performance bottlenecks, and for many of the most common issues, you can now apply the fix in just a few clicks with our Smart Shield product. Double the fun!

Our unique perspective: leveraging data from 20% of the web

Every day, Cloudflare handles traffic for over 20% of the web, giving us a unique vantage point into what makes websites faster and more resilient. We built Observatory to take advantage of this position, uniting data that is normally scattered across different tools — including real-user data, synthetic testing, error rates, and backend telemetry — into a single platform. This gives you a complete, cohesive picture of your application’s health end-to-end, in one spot, and enables you to easily identify and resolve performance issues.

For this launch, we’re bringing together:

  • Real-user data: See how your application performs for real people, in the real world.

  • Back-end telemetry: Break down the lifecycle of a request to pinpoint areas for improvement.

  • Error rates: Understand the stability of your application at both the edge and origin.

  • Cache hit ratios: Ensure you’re maximizing the performance of your configuration.

  • Synthetic testing: Proactively test and monitor key endpoints with powerful, accurate simulations.

Let’s take a quick look at each data set to see how we use them in Observatory.

Real-user data

There are two primary forms of data collection: real-user data and synthetic data. Real-user data are performance metrics collected from real traffic, from real visitors, to your application. It’s how users are actually seeing your application perform in the real world. It’s unpredictable, and covers every scenario.

Synthetic data is data collected using some sort of simulated test (loading a site in a headless browser, making network requests from a testing system to an endpoint, etc.). Tests are run under a predefined set of characteristics — location, network speed, etc. — to provide a consistent baseline.

Both forms of data have their uses, and companies with a strongly established culture of operational excellence tend to use both.

The first data you’ll see when you visit Observatory is real-user data collected with Real User Monitoring (RUM), with a particular focus on the Core Web Vital metrics.


This is very intentional.

Real-user data should be the source of truth when it comes to measuring performance and resiliency of your application. Even the best of synthetic data sources are always going to be an approximation. They cannot cover every possible scenario, and because they are being run from a lab environment, they will not always reveal issues that may be more sporadic and unpredictable.

They’re also the best representation of what your users are experiencing when they access your site and, at the end of the day, that’s why we focus on improving performance, resiliency,  and security for our users.

We believe so strongly in the importance of every company having access to accurate, detailed RUM data that we are providing it for free, to all accounts. In fact, we’re about to make our privacy-first analytics — which doesn’t track individual users for analytics — available by default for all free zones (excluding data from EU or UK visitors), no setup necessary. We believe the right thing is arming everyone with detailed, actionable, real-user data, and we want to make it easy.

Backend telemetry

Front-end performance metrics are our best proxy for understanding the actual user experience of an application and as a result, they work great as key performance indicators (KPI’s).

But they’re not enough. Every primary metric should have some level of supporting diagnostic metrics that help us understand why our user metrics are performing like they are — so that we can quickly identify issues, bottlenecks, and areas of improvement.


While the industry has largely, and rightfully, moved on from Time to First Byte (TTFB) as a primary metric of focus, it still has value as a diagnostic metric. In fact, we analyzed our RUM data and found a very strong connection between Time to First Byte and Largest Contentful Paint.

Google’s recommended thresholds for Time to First Byte are:

  • Good: <= 800ms

  • Needs Improvement: > 800ms and <= 1800ms

  • Poor: > 1800ms

Similarly, their official thresholds for Largest Contentful Paint are:

  • Good: <= 2500ms

  • Needs Improvement > 2500ms and <= 4000ms

  • Poor: > 4000ms

Looking across over 9 billion events, we found that when compared to the average site, sites with a “poor” (>1800ms) TTFB are:

  • 70.1 percentage points less likely to have a “good” LCP

  • 21.9 percentage points more likely to have a “needs improvement” LCP

  • 48.2 percentage points more likely to have a “poor” LCP

TTFB is an ill-defined blackbox, so we’re making a point to break that down into its various subparts so you can quickly pinpoint if the issue is with the connection establishment, the server response time, the network itself, and more. We’ll be working to break this down even further in the coming months as we expose the complete lifecycle of a request so you’re able to pinpoint exactly where the bottlenecks lie.

Errors & cache ratios

Degradation in stability and performance are frequently directly connected to configuration changes or an increase in errors. Clear visibility into these characteristics can often cut right to the heart of the issue at hand, as well as point to opportunities for improvement of the overall efficiency and effectiveness of your application.


Observatory prominently surfaces cache hit ratio and error rates for both the edge and origin. This compliments the backend telemetry nicely, and helps to further breakdown the backend metrics you are seeing to help pinpoint areas of improvement.

Take cache hit ratio for example. Intuitively, we know that when content is served from cache on an edge server, it should be faster than when the request has to go all the way back to the origin server. Based on our data, again, that’s exactly what we see.

If we consider our Time To First Byte thresholds again (good is <= 800ms; needs improvement is > 800ms and less than 1800ms; poor is anything over 1800ms), when looking across 9 billion data points as collected by our RUM solution, we see that a whopping 91.7% of all pages served from Cloudflare’s cache have a “good” TTFB compared to 79.7% when the request has to be served from the origin server.

In other words, optimizing origin performance (more on that in a bit) and moving more content to the edge are sure-fire ways to give you a much stronger performance baseline.

Accurate and detailed synthetic testing

While real-user data is our source of truth, synthetic testing and monitoring is important as well. Because tests are run in a more controlled environment (test from this location, at this time, with this criteria, etc.), the resulting data is a lot less noisy and variable. In addition, because there is not a user involved and we don’t have to worry about any observer effect, synthetic tests are able to grab a lot more information about the request and page lifecycle.

As a result, synthetic data tends to work very well for arming engineers with debugging information, as well as providing a cleaner set of data for comparing and contrasting results across different platforms, releases, and other situations.

Observatory provides two different types of synthetic tests.

The first synthetic test is a browser test. A browser test will load the requested page in a headless browser, run Google’s Lighthouse on it to report on key performance metrics, and provide some light suggestions for improvement. 


The second type of synthetic test Observatory provides is a network test. This is a brand new test type in Cloudflare, and is focused on giving you a better breakdown of the network and back-end performance of an endpoint.

Each network test will hit the provided endpoint for the test and record the wait time, server response time, connect time, SSL negotiation time, and total load time for the endpoint response. Because these tests are much more targeted, a single test in itself is not as valuable and can be prone to variation. That variation isn’t necessarily a bad thing—in fact, variability in these results can actually give you a better understanding of the breadth of results when real users hit that same endpoint.

For that reason, network tests trigger a series of individual runs against the provided endpoint spread out over a short period of time. The data for each response is recorded, and then presented as a histogram on the test results page, letting you see not just a single datapoint, but the long and short-tail of each metric. This gives you a much more accurate representation of reality than what a single test run can provide.


You are also able to compare network tests in Observatory, by selecting two network tests that have been completed. Again, all the data points for each test will be provided in a histogram, where you can easily compare the results of the two.


We are working on improving both synthetic test types in Q4 2025, focusing on making them more powerful and diagnostic.

As we mentioned before, even at its best, synthetic data is an approximation of what is actually happening. Accuracy is critical. Inaccurate data can distract teams with variability and faulty measurements.

It’s important that these tools are as accurate and true to the real world as possible. It’s also important to us that we give back to the community, both because it’s the right thing to do, and because we believe the best way to have the highest level of confidence in the measurement tools and frameworks we’re using is the rigor and scrutiny that open-source provides.

For those reasons, we’ll be working on open-sourcing many of the testing agents we’re using to power Observatory. We’ll share more on that soon, as well as more details about how we’ve built each different testing tool, and why.

Doing something about it: Smart Suggestions

People don’t measure for the sake of having data and pretty charts. They measure because they want to be able to stay on top of the health of their application and find ways to improve it. Data is easy. Understanding what to do about the data you’re presented is both the hardest, and most important, part.

Monitoring without action is useless.

We’re building Observatory to have a relentless focus on actionability. Before any new metric is presented, we take some time to explore why that metric matters, when it’s something worth addressing, and what actions you should take if those metrics need improvement.

All of that leads us to our new Smart Suggestions. Wherever possible, we want to pair each metric with a set of opinionated, data-driven suggestions for how to make things better. We want to avoid vague hand-wavy advice and instead be prescriptive and specific and precise.

For example, let’s look at one particular recommendation we provide around improving Largest Contentful Paint.

Largest Contentful Paint is a core web vital metric that measures when the largest piece of content is displayed on the screen. That piece of content could be an image, video or text.

Much like TTFB, Largest Contentful Paint is a bit of a black box by itself. While it tells us how long it takes for that content to get on screen, there are a large number of potential bottlenecks that could be causing the delay. Perhaps the server response time was very slow. Or maybe there was something blocking the content from being displayed on the page. If the object was an image or video, perhaps the filesize was large and the resulting download was slow. LCP by itself doesn’t give us that level of granularity, so it’s hard to give more than hand wavy guidance on how to address it.

Thankfully, just like we can break TTFB into subparts, we can break LCP into its subparts as well. Specifically we can look at:

  • Time to First Byte: how quickly the server responds to the request for HTML

  • Resource Load Delay: How long it takes after TTFB for the browser to discover the LCP resource

  • Resource Load Duration: How long it takes for the browser to download the LCP resource

  • Render Delay: How long it takes the browser to render the content, after it has the resource in hand.

Breaking it down into these subparts, we can be much more diagnostic about what to do.


In the example above, our recommendation engine analyzes the site’s real-user data and notices that Resource Load Delay accounts for over 10% of total LCP time. As a result, there’s a high likelihood that the resource triggering LCP is large and could potentially be compressed to reduce file size. So we make a recommendation to enable compression using Polish.

We’re very excited about the impact these suggestions will have on helping everyone quickly zero in on meaningful solutions for improving performance and resiliency, without having to wade through mountains of data to get there. As we analyze data, we’ll find more and more patterns of problems and the solutions they can map to. Expanding on our Smart Suggestions will be a constant and ongoing focus as we move forward, and we are working on adding much more content about those patterns and what we find in Q4.

Fixing the biggest pain point: Smart Shield

Observatory gives you unprecedented insight into your application’s health, but insights are only half the battle. The next challenge is acting on them, which brings us to another layer of complexity: protecting your origin. For many of our customers, proper management of origin routes and connections is one of the largest drivers of aggregate overall performance. As we mentioned before, we see a clear negative impact on user-facing performance metrics when we have to go back to the origin, and we want to make it as easy as possible for our customers to improve those experiences. Achieving this requires protecting against unnecessary load while ensuring only trusted traffic reaches your servers.

Today’s customers have powerful tools to protect their origins, but achieving basic use cases remains frustratingly complex:

  • Making applications faster

  • Reducing origin load

  • Understanding origin health issues

  • Restricting IP address access to origin servers

These fundamental needs currently require navigating multiple APIs and dashboard settings. You shouldn’t need to become an expert in each feature — we should analyze your traffic patterns and provide clear, actionable solutions.

Smart Shield: the future of origin shielding

Smart Shield transforms origin protection from a complex, multi-tool challenge into a streamlined, intelligent solution that works on your behalf. Our unified API and UI combines all origin protection essentials — dynamic traffic acceleration, intelligent caching, health monitoring, and dedicated egress IPs — into one place that enables single-click configuration.

But we didn’t stop at simplification. Smart Shield integrates with Observatory to provide both the “what” — identifying performance bottlenecks and health issues — and the “how” — delivering capabilities that increase performance, availability, and security.

This creates a continuous feedback loop: Observatory identifies problems, Smart Shield provides solutions, and real-time analytics verify the impact.



But what does this mean for you? 

  • Reduce total cost of ownership (TCO)

  • Reduce the time-to-value (TTV) for performance, availability, and security issues pertaining to customer origins

  • Enable new features without guesswork and validate effectiveness in the data

Your time stays focused on building incredible user experiences, not becoming a configuration expert. We are excited to give you back time for your customers and your engineers, while paving the way for how you make sure your origin infrastructure is easily optimized to delight your customers. 

Protecting and accelerating origins with smart Connection Reuse

Keeping your origins fast and stable is a big part of what we do at Cloudflare. When you experience a traffic surge, the last thing you want is for a flood of TLS handshakes to knock your origin down, or for those new connections to stall your requests, leaving your users to wait for slow pages to load.

This is why we’ve made significant changes to how Cloudflare’s network talks to your origins to dramatically improve the performance of our origin connections. 

When Cloudflare makes a request to your origins, we make them from a subset of the available machines in every Cloudflare data center so that we can improve your connection reuse. Until now, this pool would be sized the same by default for every application within a data center, and changes to the sizing of the pool for a particular customer would need to be made manually. This often led to suboptimal connection reuse for our customers, as we might be making requests from way more machines than were actually needed, resulting in fewer warm connection pools than we otherwise could have had. This also caused issues at our data centers from time to time, as larger applications might have more traffic than the default pool size was capable of serving, resulting in production incidents where engineers are paged and had to manually increase the fanout factor for specific customers.

Now, these pool sizes are determined automatically and dynamically. By tracking domain-level traffic volume within a datacenter, we can automatically scale up and scale down the number of machines that serve traffic destined for customer origin servers for any particular customer, improving both the performance of customer websites and the reliability of our network. A massive, high-volume website with a considerable amount of API traffic will no longer be processed by the same number of machines as a smaller and more typical website. Our systems can respond to changes in customer traffic patterns within seconds, allowing us to quickly ramp up and respond to surges in origin traffic.

Thanks to these improvements, Cloudflare now uses over 30% fewer connections across the board to talk to origins. To put this into a more understandable perspective, this translates to saving approximately 402 years of handshake time every day across our global traffic, or 12,060 years of handshake time saved per month! This means just by proxying your traffic through Cloudflare, you’ll see a 30% on average reduction in the amount of connections to your origin, keeping it more available while serving the same traffic volume and in turn lowering your egress fees. But, in many cases, the results observed can be far greater than 30%. For example, in one data center which is particularly heavy in API traffic, we saw a reduction in origin connections of ~60%! 

Many don’t realize that making more connections to an origin requires more compute and time for systems to create TCP and SSL handshakes. This takes time away from serving content requested by your end-users and can act as a hidden tax on your performance and overall to your application. We are proud to reduce the Internet’s hidden tax by finding intelligent, innovative ways to reduce the amount of connections needed while supporting the same traffic volume.

Watch out for more updates to Smart Shield at the start of 2026 — we’re working on adding self-serve support for dedicated CDN egress IP addresses, along with significant performance, reliability, and resilience improvements!

Charting the course: next steps for Observatory & Smart Shield

We’re really excited to share these two products with everyone today. Smart Shield and Observatory combine to provide a powerful one-two punch of insight and easy remediation.

As we navigate the beta launch of Observatory, we know this is just the start.

Our vision for Observatory is to be the single source of truth for your application’s health. We know that making the right decisions requires robust, accurate data, and we want to arm our customers with the most comprehensive picture available.

In the coming months, we plan to continue driving forward with our goal of providing comprehensive data, backed by a clear path to action.

  • Deeper, more diagnostic data. We’ll continue to break down data silos, bringing in more metrics to make sure you have a truly comprehensive view of your application’s health. We’ll be focused on going deeper and being more diagnostic, breaking down every aspect of both the request and page lifecycle to give you more granular data.

  • More paths to solutions. People don’t measure for the sake of looking at data, they measure to solve problems. We’re going to continue to expand our suggestions, arming you with more precise, data-driven solutions to a wider range of issues, letting you fix problems with a single click through Smart Shield and bringing a tighter feedback loop to validate the impact of your configuration updates.

  • Benchmarking against other products. Some of our customers split traffic between different CDNs due to regulatory or compliance requirements. Naturally, this brings up a whole series of questions about comparing the performance of the split traffic. In Observatory, you can compare these today, but we have a lot of things planned to make this even easier.

Try out Observatory and Smart Shield yourself today. And if you have ideas or suggestions for making Observatory and Smart Shield better, we’re all ears and would love to talk!

Monitoring AS-SETs and why they matter

Post Syndicated from Mingwei Zhang original https://blog.cloudflare.com/monitoring-as-sets-and-why-they-matter/

Introduction to AS-SETs

An AS-SET, not to be confused with the recently deprecated BGP AS_SET, is an Internet Routing Registry (IRR) object that allows network operators to group related networks together. AS-SETs have been used historically for multiple purposes such as grouping together a list of downstream customers of a particular network provider. For example, Cloudflare uses the AS13335:AS-CLOUDFLARE AS-SET to group together our list of our own Autonomous System Numbers (ASNs) and our downstream Bring-Your-Own-IP (BYOIP) customer networks, so we can ultimately communicate to other networks whose prefixes they should accept from us. 

In other words, an AS-SET is currently the way on the Internet that allows someone to attest the networks for which they are the provider. This system of provider authorization is completely trust-based, meaning it’s not reliable at all, and is best-effort. The future of an RPKI-based provider authorization system is coming in the form of ASPA (Autonomous System Provider Authorization), but it will take time for standardization and adoption. Until then, we are left with AS-SETs.

Because AS-SETs are so critical for BGP routing on the Internet, network operators need to be able to monitor valid and invalid AS-SET memberships for their networks. Cloudflare Radar now introduces a transparent, public listing to help network operators in our routing page per ASN.

AS-SETs and building BGP route filters

AS-SETs are a critical component of BGP policies, and often paired with the expressive Routing Policy Specification Language (RPSL) that describes how a particular BGP ASN accepts and propagates routes to other networks. Most often, networks use AS-SET to express what other networks should accept from them, in terms of downstream customers. 

Back to the AS13335:AS-CLOUDFLARE example AS-SET, this is published clearly on PeeringDB for other peering networks to reference and build filters against. 


When turning up a new transit provider service, we also ask the provider networks to build their route filters using the same AS-SET. Because BGP prefixes are also created in IRR registries using the route or route6 objects, peers and providers now know what BGP prefixes they should accept from us and deny the rest. A popular tool for building prefix-lists based on AS-SETs and IRR databases is bgpq4, and it’s one you can easily try out yourself. 


For example, to generate a Juniper router’s IPv4 prefix-list containing prefixes that AS13335 could propagate for Cloudflare and its customers, you may use: 

% bgpq4 -4Jl CLOUDFLARE-PREFIXES -m24 AS13335:AS-CLOUDFLARE | head -n 10
policy-options {
replace:
 prefix-list CLOUDFLARE-PREFIXES {
    1.0.0.0/24;
    1.0.4.0/22;
    1.1.1.0/24;
    1.1.2.0/24;
    1.178.32.0/19;
    1.178.32.0/20;
    1.178.48.0/20;

Restricted to 10 lines, actual output of prefix-list would be much greater

This prefix list would be applied within an eBGP import policy by our providers and peers to make sure AS13335 is only able to propagate announcements for ourselves and our customers.

How accurate AS-SETs prevent route leaks

Let’s see how accurate AS-SETs can help prevent route leaks with a simple example. In this example, AS64502 has two providers – AS64501 and AS64503. AS64502 has accidentally messed up their BGP export policy configuration toward the AS64503 neighbor, and is exporting all routes, including those it receives from their AS64501 provider. This is a typical Type 1 Hairpin route leak.


Fortunately, AS64503 has implemented an import policy that they generated using IRR data including AS-SETs and route objects. By doing so, they will only accept the prefixes that originate from the AS Cone of AS64502, since they are their customer. Instead of having a major reachability or latency impact for many prefixes on the Internet because of this route leak propagating, it is stopped in its tracks thanks to the responsible filtering by the AS64503 provider network. Again it is worth keeping in mind the success of this strategy is dependent upon data accuracy for the fictional AS64502:AS-CUSTOMERS AS-SET.

Monitoring AS-SET misuse

Besides using AS-SETs to group together one’s downstream customers, AS-SETs can also represent other types of relationships, such as peers, transits, or IXP participations.

For example, there are 76 AS-SETs that directly include one of the Tier-1 networks, Telecom Italia / Sparkle (AS6762). Judging from the names of the AS-SETs, most of them are representing peers and transits of certain ASNs, which includes AS6762. You can view this output yourself at https://radar.cloudflare.com/routing/as6762#irr-as-sets


There is nothing wrong with defining AS-SETs that contain one’s peers or upstreams as long as those AS-SETs are not submitted upstream for customer->provider BGP session filtering. In fact, an AS-SET for upstreams or peer-to-peer relationships can be useful for defining a network’s policies in RPSL.

However, some AS-SETs in the AS6762 membership list such as AS-10099 look to attest customer relationships. 

% whois -h rr.ntt.net AS-10099 | grep "descr"
descr:          CUHK Customer

We know AS6762 is transit free and this customer membership must be invalid, so it is a prime example of AS-SET misuse that would ideally be cleaned up. Many Internet Service Providers and network operators are more than happy to correct an invalid AS-SET entry when asked to. It is reasonable to look at each AS-SET membership like this as a potential risk of having higher route leak propagation to major networks and the Internet when they happen.

AS-SET information on Cloudflare Radar

Cloudflare Radar is a hub that showcases global Internet traffic, attack, and technology trends and insights. Today, we are adding IRR AS-SET information to Radar’s routing section, freely available to the public via both website and API access. To view all AS-SETs an AS is a member of, directly or indirectly via other AS-SETs, a user can visit the corresponding AS’s routing page. For example, the AS-SETs list for Cloudflare (AS13335) is available at https://radar.cloudflare.com/routing/as13335#irr-as-sets

The AS-SET data on IRR contains only limited information like the AS members and AS-SET members. Here at Radar, we also enhance the AS-SET table with additional useful information as follows.

  • Inferred ASN shows the AS number that is inferred to be the creator of the AS-SET. We use PeeringDB AS-SET information match if available. Otherwise, we parse the AS-SET name to infer the creator.

  • IRR Sources shows which IRR databases we see the corresponding AS-SET. We are currently using the following databases: AFRINIC, APNIC, ARIN, LACNIC, RIPE, RADB, ALTDB, NTTCOM, and TC.

  • AS Members and AS-SET members show the count of the corresponding types of members.

  • AS Cone is the count of the unique ASNs that are included by the AS-SET directly or indirectly.

  • Upstreams is the count of unique AS-SETs that includes the corresponding AS-SET.

Users can further filter the table by searching for a specific AS-SET name or ASN. A toggle to show only direct or indirect AS-SETs is also available.


In addition to listing AS-SETs, we also provide a tree-view to display how an AS-SET includes a given ASN. For example, the following screenshot shows how as-delta indirectly includes AS6762 through 7 additional other AS-SETs. Users can copy or download this tree-view content in the text format, making it easy to share with others.


We built this Radar feature using our publicly available API, the same way other Radar websites are built. We have also experimented using this API to build additional features like a full AS-SET tree visualization. We encourage developers to give this API (and other Radar APIs) a try, and tell us what you think!


Looking ahead

We know AS-SETs are hard to keep clean of error or misuse, and even though Radar is making them easier to monitor, the mistakes and misuse will continue. Because of this, we as a community need to push forth adoption of RFC9234 and implementations of it from the major vendors. RFC9234 embeds roles and an Only-To-Customer (OTC) attribute directly into the BGP protocol itself, helping to detect and prevent route leaks in-line. In addition to BGP misconfiguration protection with RFC9234, Autonomous System Provider Authorization (ASPA) is still making its way through the IETF and will eventually help offer an authoritative means of attesting who the actual providers are per BGP Autonomous System (AS).

If you are a network operator and manage an AS-SET, you should seriously consider moving to hierarchical AS-SETs if you have not already. A hierarchical AS-SET looks like AS13335:AS-CLOUDFLARE instead of AS-CLOUDFLARE, but the difference is very important. Only a proper maintainer of the AS13335 ASN can create AS13335:AS-CLOUDFLARE, whereas anyone could create AS-CLOUDFLARE in an IRR database if they wanted to. In other words, using hierarchical AS-SETs helps guarantee ownership and prevent the malicious poisoning of routing information.

While keeping track of AS-SET memberships seems like a chore, it can have significant payoffs in preventing BGP-related incidents such as route leaks. We encourage all network operators to do their part in making sure the AS-SETs you submit to your providers and peers to communicate your downstream customer cone are accurate. Every small adjustment or clean-up effort in AS-SETs could help lessen the impact of a BGP incident later.

Visit Cloudflare Radar for additional insights around (Internet disruptions, routing issues, Internet traffic trends, attacks, Internet quality, etc.). Follow us on social media at @CloudflareRadar (X), https://noc.social/@cloudflareradar (Mastodon), and radar.cloudflare.com (Bluesky), or contact us via e-mail.

Cloudflare just got faster and more secure, powered by Rust

Post Syndicated from Richard Boulton original https://blog.cloudflare.com/20-percent-internet-upgrade/

Cloudflare is relentless about building and running the world’s fastest network. We have been tracking and reporting on our network performance since 2021: you can see the latest update here.

Building the fastest network requires work in many areas. We invest a lot of time in our hardware, to have efficient and fast machines. We invest in peering arrangements, to make sure we can talk to every part of the Internet with minimal delay. On top of this, we also have to invest in the software we run our network on, especially as each new product can otherwise add more processing delay.

No matter how fast messages arrive, we introduce a bottleneck if that software takes too long to think about how to process and respond to requests. Today we are excited to share a significant upgrade to our software that cuts the median time we take to respond by 10ms and delivers a 25% performance boost, as measured by third-party CDN performance tests.

We’ve spent the last year rebuilding major components of our system, and we’ve just slashed the latency of traffic passing through our network for millions of our customers. At the same time, we’ve made our system more secure, and we’ve reduced the time it takes for us to build and release new products. 

Where did we start?

Every request that hits Cloudflare starts a journey through our network. It might come from a browser loading a webpage, a mobile app calling an API, or automated traffic from another service. These requests first terminate at our HTTP and TLS layer, then pass into a system we call FL, and finally through Pingora, which performs cache lookups or fetches data from the origin if needed.


FL is the brain of Cloudflare. Once a request reaches FL, we then run the various security and performance features in our network. It applies each customer’s unique configuration and settings, from enforcing WAF rules and DDoS protection to routing traffic to the Developer Platform and R2. 

Built more than 15 years ago, FL has been at the core of Cloudflare’s network. It enables us to deliver a broad range of features, but over time that flexibility became a challenge. As we added more products, FL grew harder to maintain, slower to process requests, and more difficult to extend. Each new feature required careful checks across existing logic, and every addition introduced a little more latency, making it increasingly difficult to sustain the performance we wanted.

You can see how FL is key to our system — we’ve often called it the “brain” of Cloudflare. It’s also one of the oldest parts of our system: the first commit to the codebase was made by one of our founders, Lee Holloway, well before our initial launch. We’re celebrating our 15th Birthday this week – this system started 9 months before that!

commit 39c72e5edc1f05ae4c04929eda4e4d125f86c5ce
Author: Lee Holloway <q@t60.(none)>
Date:   Wed Jan 6 09:57:55 2010 -0800

    nginx-fl initial configuration

As the commit implies, the first version of FL was implemented based on the NGINX webserver, with product logic implemented in PHP.  After 3 years, the system became too complex to manage effectively, and too slow to respond, and an almost complete rewrite of the running system was performed. This led to another significant commit, this time made by Dane Knecht, who is now our CTO.

commit bedf6e7080391683e46ab698aacdfa9b3126a75f
Author: Dane Knecht
Date:   Thu Sep 19 19:31:15 2013 -0700

    remove PHP.

From this point on, FL was implemented using NGINX, the OpenResty framework, and LuaJIT.  While this was great for a long time, over the last few years it started to show its age. We had to spend increasing amounts of time fixing or working around obscure bugs in LuaJIT. The highly dynamic and unstructured nature of our Lua code, which was a blessing when first trying to implement logic quickly, became a source of errors and delay when trying to integrate large amounts of complex product logic. Each time a new product was introduced, we had to go through all the other existing products to check if they might be affected by the new logic.

It was clear that we needed a rethink. So, in July 2024, we cut an initial commit for a brand new, and radically different, implementation. To save time agreeing on a new name for this, we just called it “FL2”, and started, of course, referring to the original FL as “FL1”.

commit a72698fc7404a353a09a3b20ab92797ab4744ea8
Author: Maciej Lechowski
Date:   Wed Jul 10 15:19:28 2024 +0100

    Create fl2 project

Rust and rigid modularization

We weren’t starting from scratch. We’ve previously blogged about how we replaced another one of our legacy systems with Pingora, which is built in the Rust programming language, using the Tokio runtime. We’ve also blogged about Oxy, our internal framework for building proxies in Rust. We write a lot of Rust, and we’ve gotten pretty good at it.

We built FL2 in Rust, on Oxy, and built a strict module framework to structure all the logic in FL2.

Why Oxy?

When we set out to build FL2, we knew we weren’t just replacing an old system; we were rebuilding the foundations of Cloudflare. That meant we needed more than just a proxy; we needed a framework that could evolve with us, handle the immense scale of our network, and let teams move quickly without sacrificing safety or performance. 

Oxy gives us a powerful combination of performance, safety, and flexibility. Built in Rust, it eliminates entire classes of bugs that plagued our Nginx/LuaJIT-based FL1, like memory safety issues and data races, while delivering C-level performance. At Cloudflare’s scale, those guarantees aren’t nice-to-haves, they’re essential. Every microsecond saved per request translates into tangible improvements in user experience, and every crash or edge case avoided keeps the Internet running smoothly. Rust’s strict compile-time guarantees also pair perfectly with FL2’s modular architecture, where we enforce clear contracts between product modules and their inputs and outputs.

But the choice wasn’t just about language. Oxy is the culmination of years of experience building high-performance proxies. It already powers several major Cloudflare services, from our Zero Trust Gateway to Apple’s iCloud Private Relay, so we knew it could handle the diverse traffic patterns and protocol combinations that FL2 would see. Its extensibility model lets us intercept, analyze, and manipulate traffic from layer 3 up to layer 7, and even decapsulate and reprocess traffic at different layers. That flexibility is key to FL2’s design because it means we can treat everything from HTTP to raw IP traffic consistently and evolve the platform to support new protocols and features without rewriting fundamental pieces.

Oxy also comes with a rich set of built-in capabilities that previously required large amounts of bespoke code. Things like monitoring, soft reloads, dynamic configuration loading and swapping are all part of the framework. That lets product teams focus on the unique business logic of their module rather than reinventing the plumbing every time. This solid foundation means we can make changes with confidence, ship them quickly, and trust they’ll behave as expected once deployed.

Smooth restarts – keeping the Internet flowing

One of the most impactful improvements Oxy brings is handling of restarts. Any software under continuous development and improvement will eventually need to be updated. In desktop software, this is easy: you close the program, install the update, and reopen it. On the web, things are much harder. Our software is in constant use and cannot simply stop. A dropped HTTP request can cause a page to fail to load, and a broken connection can kick you out of a video call. Reliability is not optional.

In FL1, upgrades meant restarts of the proxy process. Restarting a proxy meant terminating the process entirely, which immediately broke any active connections. That was particularly painful for long-lived connections such as WebSockets, streaming sessions, and real-time APIs. Even planned upgrades could cause user-visible interruptions, and unplanned restarts during incidents could be even worse.

Oxy changes that. It includes a built-in mechanism for graceful restarts that lets us roll out new versions without dropping connections whenever possible. When a new instance of an Oxy-based service starts up, the old one stops accepting new connections but continues to serve existing ones, allowing those sessions to continue uninterrupted until they end naturally.

This means that if you have an ongoing WebSocket session when we deploy a new version, that session can continue uninterrupted until it ends naturally, rather than being torn down by the restart. Across Cloudflare’s fleet, deployments are orchestrated over several hours, so the aggregate rollout is smooth and nearly invisible to end users.

We take this a step further by using systemd socket activation. Instead of letting each proxy manage its own sockets, we let systemd create and own them. This decouples the lifetime of sockets from the lifetime of the Oxy application itself. If an Oxy process restarts or crashes, the sockets remain open and ready to accept new connections, which will be served as soon as the new process is running. That eliminates the “connection refused” errors that could happen during restarts in FL1 and improves overall availability during upgrades.

We also built our own coordination mechanisms in Rust to replace Go libraries like tableflip with shellflip. This uses a restart coordination socket that validates configuration, spawns new instances, and ensures the new version is healthy before the old one shuts down. This improves feedback loops and lets our automation tools detect and react to failures immediately, rather than relying on blind signal-based restarts.

Composing FL2 from Modules

To avoid the problems we had in FL1, we wanted a design where all interactions between product logic were explicit and easy to understand. 

So, on top of the foundations provided by Oxy, we built a platform which separates all the logic built for our products into well-defined modules. After some experimentation and research, we designed a module system which enforces some strict rules:

  • No IO (input or output) can be performed by the module.

  • The module provides a list of phases.

  • Phases are evaluated in a strictly defined order, which is the same for every request.

  • Each phase defines a set of inputs which the platform provides to it, and a set of outputs which it may emit.

Here’s an example of what a module phase definition looks like:

Phase {
    name: phases::SERVE_ERROR_PAGE,
    request_types_enabled: PHASE_ENABLED_FOR_REQUEST_TYPE,
    inputs: vec![
        InputKind::IPInfo,
        InputKind::ModuleValue(
            MODULE_VALUE_CUSTOM_ERRORS_FETCH_WORKER_RESPONSE.as_str(),
        ),
        InputKind::ModuleValue(MODULE_VALUE_ORIGINAL_SERVE_RESPONSE.as_str()),
        InputKind::ModuleValue(MODULE_VALUE_RULESETS_CUSTOM_ERRORS_OUTPUT.as_str()),
        InputKind::ModuleValue(MODULE_VALUE_RULESETS_UPSTREAM_ERROR_DETAILS.as_str()),
        InputKind::RayId,
        InputKind::StatusCode,
        InputKind::Visitor,
    ],
    outputs: vec![OutputValue::ServeResponse],
    filters: vec![],
    func: phase_serve_error_page::callback,
}

This phase is for our custom error page product.  It takes a few things as input — information about the IP of the visitor, some header and other HTTP information, and some “module values.” Module values allow one module to pass information to another, and they’re key to making the strict properties of the module system workable. For example, this module needs some information that is produced by the output of our rulesets-based custom errors product (the “MODULE_VALUE_RULESETS_CUSTOM_ERRORS_OUTPUT” input). These input and output definitions are enforced at compile time.

While these rules are strict, we’ve found that we can implement all our product logic within this framework. The benefit of doing so is that we can immediately tell which other products might affect each other.

How to replace a running system

Building a framework is one thing. Building all the product logic and getting it right, so that customers don’t notice anything other than a performance improvement, is another.

The FL code base supports 15 years of Cloudflare products, and it’s changing all the time. We couldn’t stop development. So, one of our first tasks was to find ways to make the migration easier and safer.

Step 1 – Rust modules in OpenResty

It’s a big enough distraction from shipping products to customers to rebuild product logic in Rust. Asking all our teams to maintain two versions of their product logic, and reimplement every change a second time until we finished our migration was too much.

So, we implemented a layer in our old NGINX and OpenResty based FL which allowed the new modules to be run. Instead of maintaining a parallel implementation, teams could implement their logic in Rust, and replace their old Lua logic with that, without waiting for the full replacement of the old system.

For example, here’s part of the implementation for the custom error page module phase defined earlier (we’ve cut out some of the more boring details, so this doesn’t quite compile as-written):

pub(crate) fn callback(_services: &mut Services, input: &Input<'_>) -> Output {
    // Rulesets produced a response to serve - this can either come from a special
    // Cloudflare worker for serving custom errors, or be directly embedded in the rule.
    if let Some(rulesets_params) = input
        .get_module_value(MODULE_VALUE_RULESETS_CUSTOM_ERRORS_OUTPUT)
        .cloned()
    {
        // Select either the result from the special worker, or the parameters embedded
        // in the rule.
        let body = input
            .get_module_value(MODULE_VALUE_CUSTOM_ERRORS_FETCH_WORKER_RESPONSE)
            .and_then(|response| {
                handle_custom_errors_fetch_response("rulesets", response.to_owned())
            })
            .or(rulesets_params.body);

        // If we were able to load a body, serve it, otherwise let the next bit of logic
        // handle the response
        if let Some(body) = body {
            let final_body = replace_custom_error_tokens(input, &body);

            // Increment a metric recording number of custom error pages served
            custom_pages::pages_served("rulesets").inc();

            // Return a phase output with one final action, causing an HTTP response to be served.
            return Output::from(TerminalAction::ServeResponse(ResponseAction::OriginError {
                rulesets_params.status,
                source: "rulesets http_custom_errors",
                headers: rulesets_params.headers,
                body: Some(Bytes::from(final_body)),
            }));
        }
    }
}

The internal logic in each module is quite cleanly separated from the handling of data, with very clear and explicit error handling encouraged by the design of the Rust language.

Many of our most actively developed modules were handled this way, allowing the teams to maintain their change velocity during our migration.

Step 2 – Testing and automated rollouts


It’s essential to have a seriously powerful test framework to cover such a migration.  We built a system, internally named Flamingo, which allows us to run thousands of full end-to-end test requests concurrently against our production and pre-production systems. The same tests run against FL1 and FL2, giving us confidence that we’re not changing behaviours.

Whenever we deploy a change, that change is rolled out gradually across many stages, with increasing amounts of traffic. Each stage is automatically evaluated, and only passes when the full set of tests have been successfully run against it – as well as overall performance and resource usage metrics being within acceptable bounds. This system is fully automated, and pauses or rolls back changes if the tests fail.

The benefit is that we’re able to build and ship new product features in FL2 within 48 hours – where it would have taken weeks in FL1. In fact, at least one of the announcements this week involved such a change!

Step 3 – Fallbacks

Over 100 engineers have worked on FL2, and we have over 130 modules. And we’re not quite done yet. We’re still putting the final touches on the system, to make sure it replicates all the behaviours of FL1.

So how do we send traffic to FL2 without it being able to handle everything? If FL2 receives a request, or a piece of configuration for a request, that it doesn’t know how to handle, it gives up and does what we’ve called a fallback – it passes the whole thing over to FL1. It does this at the network level – it just passes the bytes on to FL1.

As well as making it possible for us to send traffic to FL2 without it being fully complete, this has another massive benefit. When we have implemented a piece of new functionality in FL2, but want to double check that it is working the same as in FL1, we can evaluate the functionality in FL2, and then trigger a fallback. We are able to compare the behaviour of the two systems, allowing us to get a high confidence that our implementation was correct.

Step 4 – Rollout

We started running customer traffic through FL2 early in 2025, and have been progressively increasing the amount of traffic served throughout the year. Essentially, we’ve been watching two graphs: one with the proportion of traffic routed to FL2 going up, and another with the proportion of traffic failing to be served by FL2 and falling back to FL1 going down.

We started this process by passing traffic for our free customers through the system. We were able to prove that the system worked correctly, and drive the fallback rates down for our major modules. Our Cloudflare Community MVPs acted as an early warning system, smoke testing and flagging when they suspected the new platform might be the cause of a new reported problem. Crucially their support allowed our team to investigate quickly, apply targeted fixes, or confirm the move to FL2 was not to blame.

We then advanced to our paying customers, gradually increasing the amount of customers using the system. We also worked closely with some of our largest customers, who wanted the performance benefits of FL2, and onboarded them early in exchange for lots of feedback on the system.

Right now, most of our customers are using FL2. We still have a few features to complete, and are not quite ready to onboard everyone, but our target is to turn off FL1 within a few more months.

Impact of FL2

As we described at the start of this post, FL2 is substantially faster than FL1. The biggest reason for this is simply that FL2 performs less work. You might have noticed in the module definition example a line

    filters: vec![],

Every module is able to provide a set of filters, which control whether they run or not. This means that we don’t run logic for every product for every request — we can very easily select just the required set of modules. The incremental cost for each new product we develop has gone away.

Another huge reason for better performance is that FL2 is a single codebase, implemented in a performance focussed language. In comparison, FL1 was based on NGINX (which is written in C), combined with LuaJIT (Lua, and C interface layers), and also contained plenty of Rust modules.  In FL1, we spent a lot of time and memory converting data from the representation needed by one language, to the representation needed by another.

As a result, our internal measures show that FL2 uses less than half the CPU of FL1, and much less than half the memory. That’s a huge bonus — we can spend the CPU on delivering more and more features for our customers!

How do we measure if we are getting better?

Using our own tools and independent benchmarks like CDNPerf, we measured the impact of FL2 as we rolled it out across the network. The results are clear: websites are responding 10 ms faster at the median, a 25% performance boost.


Security

FL2 is also more secure by design than FL1. No software system is perfect, but the Rust language brings us huge benefits over LuaJIT. Rust has strong compile-time memory checks and a type system that avoids large classes of errors. Combine that with our rigid module system, and we can make most changes with high confidence.

Of course, no system is secure if used badly. It’s easy to write code in Rust, which causes memory corruption. To reduce risk, we maintain strong compile time linting and checking, together with strict coding standards, testing and review processes.

We have long followed a policy that any unexplained crash of our systems needs to be investigated as a high priority. We won’t be relaxing that policy, though the main cause of novel crashes in FL2 so far has been due to hardware failure. The massively reduced rates of such crashes will give us time to do a good job of such investigations.

What’s next?

We’re spending the rest of 2025 completing the migration from FL1 to FL2, and will turn off FL1 in early 2026. We’re already seeing the benefits in terms of customer performance and speed of development, and we’re looking forward to giving these to all our customers.

We have one last service to completely migrate. The “HTTP & TLS Termination” box from the diagram way back at the top is also an NGINX service, and we’re midway through a rewrite in Rust. We’re making good progress on this migration, and expect to complete it early next year.

After that, when everything is modular, in Rust and tested and scaled, we can really start to optimize! We’ll reorganize and simplify how the modules connect to each other, expand support for non-HTTP traffic like RPC and streams, and much more. 

If you’re interested in being part of this journey, check out our careers page for open roles – we’re always looking for new talent to help us to help build a better Internet.

Network performance update: Birthday Week 2025

Post Syndicated from Lai Yi Ohlsen original https://blog.cloudflare.com/network-performance-update-birthday-week-2025/

We are committed to being the fastest network in the world because improvements in our performance translate to improvements for the own end users of your application. We are excited to share that Cloudflare continues to be the fastest network for the most peered networks in the world.

We relentlessly measure our own performance and our performance against peers. We publish those results routinely, starting with our first update in June 2021 and most recently with our last post in September 2024.

Today’s update breaks down where we have improved since our update last year and what our priorities are going into the next year. While we are excited to be the fastest in the greatest number of last-mile ISPs, we are never done improving and have more work to do.

How do we measure this metric, and what are the results?

We measure network performance by attempting to capture what the experience is like for Internet users across the globe. To do that we need to simulate what their connection is like from their last-mile ISP to our networks.

We start by taking the 1,000 largest networks in the world based on estimated population. We use that to give ourselves a representation of real users in nearly every geography.

We then measure performance itself with TCP connection time. TCP connection time is the time it takes for an end user to connect to the website or endpoint they are trying to reach. We chose this metric because we believe this most closely approximates what users perceive to be Internet speed, as opposed to other metrics which are either too scientific (ignoring real world challenges like congestion or distance) or too broad.

We take the trimean measurement of TCP connection times to calculate our metric. The trimean is a weighted average of three statistical values: the first quartile, the median, and the third quartile. This approach allows us to reduce some of the noise and outliers and get a comprehensive picture of quality.

For this year’s update, we examined the trimean of TCP connection times measured from August 6 to September 4, Cloudflare is the #1 provider in 40% of the top 1000 networks. In our September 2024 update, we shared that we were the #1 provider in 44% of the top 1000 networks.


The TCP Connection Time (Trimean) graph shows that we are the fastest TCP connection time in 383 networks, but that would make us the fastest in 38% of the top 1,000. We exclude networks that aren’t last-mile ISPs, such as transit networks, since they don’t reflect the end user experience, which brings the number of measured networks to 964 and makes Cloudflare the fastest in 40% of measured ISPs and the fastest across the top networks.

How do we capture this data? 

A Cloudflare-branded error page does more than just display an error; it kicks off a real-world speed test. Behind the scenes, on a selection of our error pages, we use Real User Measurements (RUM), which involves a browser retrieving a small file from multiple networks, including Cloudflare, Amazon CloudFront, Google, Fastly and Akamai.

Running these tests lets us gather performance data directly from the user’s perspective, providing a genuine comparison of different network speeds. We do this to understand where our network is fastest and, more importantly, where we can make further improvements. For a deeper dive into the technical details, the Speed Week blog post covers the full methodology.

By using RUM data, we track key metrics like TCP Connection Time, Time to First Byte (TTFB), and Time to Last Byte (TTLB). These are widely recognized, industry-standard metrics that allow us to objectively measure how quickly and efficiently a website loads for actual users. By monitoring these benchmarks, we can objectively compare our performance against other networks.

We specifically chose the top 1000 networks by estimated population from APNIC, excluding those that aren’t last-mile ISPs. Consistency is key: by analyzing the same group of networks in every cycle, we ensure our measurements and reporting remain reliable and directly comparable over time.

How do the results compare across countries?

The map below shows the fastest providers per country and Cloudflare is fastest in dozens of countries. 


The color coding is generated by grouping all the measurements we generate by which country the measurement originates from. Then we look at the trimean measurements for each provider to identify who is the fastest… Akamai was measured as well, but providers are only represented in the map if they ranked first in a country which Akamai does not anywhere in the world.

These slim margins mean that the fastest provider in a country is often determined by latency differences so small that the fastest provider is often only faster by less than 5%. As an example, let’s look at India, a country where we are currently the second-fastest provider.

India (IN)

Rank

Entity 

Connect Time (Trimean)

#1 Diff

#1

CloudFront

107 ms

–

#2

Cloudflare

113 ms

+4.81% (+5.16 ms)

#3

Google

117 ms

+8.74% (+9.39 ms)

#4 

Fastly

133 ms

+24% (+26 ms)

#5

Akamai

144 ms

+34% (+37 ms)

In India, Cloudflare is 5ms behind Cloudfront, the #1 provider (To put milliseconds into perspective, the average human eye blink lasts between 100ms and 400ms). The competition for the number one spot in many countries is fierce and often shifts day by day. For example, in Mexico on Tuesday, August 5th, Cloudflare was the second-fastest provider by 0.73 ms but then on Tuesday, August 12th, Cloudflare was the fastest provider by 3.72 ms. 

Mexico (MX)

Date

Rank

Entity 

Connect Time (Trimean)

#1 Diff

August 5, 2025

#1

CloudFront

116 ms

–

#2

Cloudflare

116 ms

+0.63% (+0.73 ms)

August 12, 2025

#1

Cloudflare

106 ms

–

#2

CloudFront

109 ms

+3.52% (+3.72 ms)

Because ranking reorderings are common, we also review country and network level rankings to evaluate and benchmark our performance. 

Focusing on where we are not the fastest yet

As mentioned above, in September 2024, Cloudflare was fastest in 44% of measured ISPs. These values can shift as providers constantly make improvements to their networks. One way we focus in on how we are prioritizing improving is to not just observe where we are not the fastest but to measure how far we are from the leader.

In these locations we tend to pace extremely close to the fastest provider, giving us an opportunity to capture the spot as we relentlessly improve. In networks where Cloudflare is 2nd, over 50% of those networks have a less than 5% difference (10ms or less) with the top provider.

Country

ASN

#1

Cloudflare Rank

#1 Diff (ms)

#1 Diff (%)

US

AS36352

Google

2

25 ms

32%

US

AS46475

Google

2

35 ms

29%

US

AS29802

Google

2

8.03 ms

21%

US

AS20473

Google

2

15 ms

13%

US

AS7018

CloudFront

2

23 ms

13%

US

AS4181

CloudFront

2

8.19 ms

11%

US

AS62240

Google

2

18 ms

9.77%

US

AS22773

CloudFront

2

12 ms

9.48%

US

AS6167

CloudFront

2

13 ms

7.55%

US

AS11427

Google

2

9.33 ms

5.27%

US

AS6614

CloudFront

2

6.68 ms

4.12%

US

AS4922

Google

2

3.38 ms

3.86%

US

AS11492

Fastly

2

3.73 ms

3.33%

US

AS11351

Google

2

5.14 ms

3.04%

US

AS396356

Google

2

4.12 ms

2.23%

US

AS212238

Google

2

3.42 ms

1.35%

US

AS20055

Fastly

2

1.22 ms

1.33%

US

AS40021

CloudFront

2

2.06 ms

0.91%

US

AS12271

Fastly

2

1.26 ms

0.89%

US

AS141039

CloudFront

2

1.26 ms

0.88%

In networks where Cloudflare is 3rd, 50% of those networks are less than a 10% difference with the top provider (10ms or less). Margins are small and suggest that in instances where Cloudflare isn’t number one across networks, we’re extremely close to our competitors and the top networks change day over day. 

Country

ASN

#1

Cloudflare Rank

#1 Diff (ms)

#1 Diff (%)

US

AS6461

Google

3

33 ms

39%

US

AS81

Fastly

3

43 ms

35%

US

AS14615

Google

3

24 ms

24%

US

AS13977

CloudFront

3

21 ms

19%

US

AS33363

Google

3

29 ms

18%

US

AS63949

Google

3

9.56 ms

14%

US

AS14593

Fastly

3

17 ms

13%

US

AS23089

CloudFront

3

7.4 ms

11%

US

AS16509

Fastly

3

10 ms

9.48%

US

AS209

CloudFront

3

9.69 ms

6.87%

US

AS27364

CloudFront

3

8.76 ms

6.61%

US

AS11404

CloudFront

3

6.11 ms

6.16%

US

AS46690

CloudFront

3

5.91 ms

5.43%

US

AS136787

CloudFront

3

8.23 ms

5.18%

US

AS6079

Fastly

3

5.45 ms

4.49%

US

AS5650

Google

3

3.91 ms

3.35%

Countries with an abundance of networks, like the United States, have a lot of noise we need to calibrate against. For example, the graph below represents the performance of all providers for a major ISP like AS701 (Verizon Business).

AS701 (Verizon Business) Connect Time (P95) between 2025-08-09 and 2025-09-09


In this chart, the “P95” value, or 95th percentile, refers to one point of a percentile distribution. The P95 shows the value below which 95% of the data points fall and is specifically good at helping identify the slowest or worst-case user experiences, such as those on poor networks or older devices. Additionally, we review the other numbers lower on the percentile chain in the table below, which tell us how performance varies across the full range of data. When we do so, the picture becomes more nuanced.

AS701 (Verizon Business) Provider Rankings for Connect Time at P95, P75 and P50

Rank

Entity 

Connect Time (P95)

Connect Time (P75)

Connect Time (P50)

#1

Fastly

128 ms

66 ms

48 ms

#2

Google

134 ms

72 ms

54 ms

#3

CloudFront

139 ms

67 ms

47 ms

#4 

Cloudflare

141 ms

68 ms

49 ms

#5

Akamai

160 ms

84 ms

61 ms

At the 95th percentile for AS701, Cloudflare ranks 4th but at the 75th and 50th, Cloudflare is only 2 milliseconds slower than the fastest provider. In other words, when reviewing more than one point along the distribution at the network level, Cloudflare is keeping up with the top providers for the less extreme samples. To capture these details, it’s important to look at the range of outcomes, not just one percentile.

To better reflect the full spectrum of user experiences, we started using the trimean in July 2025 to rank providers. This metric combines values from across the distribution of data – specifically the 75th, 50th and 25th percentiles – which gives a more balanced representation of overall performance, rather than only focusing on the extremes. Summarizing user experience with a single number is always challenging, but the trimean helps us compare providers in a way that better reflects how users actually experience the Internet.

Cloudflare is the fastest provider in 40% of networks in the majority of real-world conditions, not just in worst-case scenarios. Still, the 95th percentile remains key to understanding how performance holds up in challenging conditions and where other providers might fall behind in performance. When we review the 95th percentile across the same date range for all the networks, not just AS701, Cloudflare is fastest across roughly the same amount of networks but by 103 more networks than the next fastest provider. Being faster in such a wide margin of networks tells us that Cloudflare is particularly strong in the challenging, long-tail cases that other providers struggle with.


Our performance data shows that even when we are not the top-ranked provider, we remain exceptionally competitive, often trailing the leader by a mere handful of percentage points. Our strength at the 95th percentile also highlights our superior performance in the most challenging scenarios. Cloudflare’s ability to outperform other providers, in the worst-case, is a testament to the resilience and efficiency of our network.

Moving forward, we’ll continue to share multiple metrics and continue to make improvements to our network —and we’ll use this data to do it! Let’s talk about how. 

How does Cloudflare use this data to improve?

Cloudflare applies this data to identify regions and networks that need prioritization. If we are consistently slower than other providers in a network, we want to know why, so we can fix it.

For example, the graph below shows the 95th percentile of Connect Time for AS8966. Prior to June 13, 2025, our performance was suffering, and we were the slowest provider for the network. By referencing our own measurement data, we prioritized partner data centers in the region and almost immediately performance improved for users connecting through AS8966.

Cloudflare’s partner data centers consist of collaborations with local service providers who host Cloudflare’s equipment within their own facilities. This allows us to expand our network to new locations and get closer to users more quickly. In the case of AS8966, adding a new partner data center took us from being ranked last to ranked first and improved latency by roughly 150ms in one day. By using a data-driven approach, we made our network faster and most importantly, improved the end user experience.

TCP Connect Time (P95) for AS8966


What’s next?

We are always working to build a faster network and will continue sharing our process as we go. Our approach is straightforward: identify performance bottlenecks, implement fixes, and report the results. We believe in being transparent about our methods and are committed to a continuous cycle of improvement to achieve the best possible performance. Follow our blog for the latest performance updates as we continue to optimize our network and share our progress.

Eliminating Cold Starts 2: shard and conquer

Post Syndicated from Harris Hancock original https://blog.cloudflare.com/eliminating-cold-starts-2-shard-and-conquer/

Five years ago, we announced that we were Eliminating Cold Starts with Cloudflare Workers. In that episode, we introduced a technique to pre-warm Workers during the TLS handshake of their first request. That technique takes advantage of the fact that the TLS Server Name Indication (SNI) is sent in the very first message of the TLS handshake. Armed with that SNI, we often have enough information to pre-warm the request’s target Worker.

Eliminating cold starts by pre-warming Workers during TLS handshakes was a huge step forward for us, but “eliminate” is a strong word. Back then, Workers were still relatively small, and had cold starts constrained by limits explained later in this post. We’ve relaxed those limits, and users routinely deploy complex applications on Workers, often replacing origin servers. Simultaneously, TLS handshakes haven’t gotten any slower. In fact, TLS 1.3 only requires a single round trip for a handshake – compared to three round trips for TLS 1.2 – and is more widely used than it was in 2021.

Earlier this month, we finished deploying a new technique intended to keep pushing the boundary on cold start reduction. The new technique (or old, depending on your perspective) uses a consistent hash ring to take advantage of our global network. We call this mechanism “Worker sharding”.


What’s in a cold start?

A Worker is the basic unit of compute in our serverless computing platform. It has a simple lifecycle. We instantiate it from source code (typically JavaScript), make it serve a bunch of requests (often HTTP, but not always), and eventually shut it down some time after it stops receiving traffic, to re-use its resources for other Workers. We call that shutdown process “eviction”.

The most expensive part of the Worker’s lifecycle is the initial instantiation and first request invocation. We call this part a “cold start”. Cold starts have several phases: fetching the script source code, compiling the source code, performing a top-level execution of the resulting JavaScript module, and finally, performing the initial invocation to serve the incoming HTTP request that triggered the whole sequence of events in the first place.


Cold starts have become longer than TLS handshakes

Fundamentally, our TLS handshake technique depends on the handshake lasting longer than the cold start. This is because the duration of the TLS handshake is time that the visitor must spend waiting, regardless, so it’s beneficial to everyone if we do as much work during that time as possible. If we can run the Worker’s cold start in the background while the handshake is still taking place, and if that cold start finishes before the handshake, then the request will ultimately see zero cold start delay. If, on the other hand, the cold start takes longer than the TLS handshake, then the request will see some part of the cold start delay – though the technique still helps reduce that visible delay.


In the early days, TLS handshakes lasting longer than Worker cold starts was a safe bet, and cold starts typically won the race. One of our early blog posts explaining how our platform works mentions 5 millisecond cold start times – and that was correct, at the time!

For every limit we have, our users have challenged us to relax them. Cold start times are no different. 

There are two crucial limits which affect cold start time: Worker script size and the startup CPU time limit. While we didn’t make big announcements at the time, we have quietly raised both of those limits since our last Eliminating Cold Starts blog post:

We relaxed these limits because our users wanted to deploy increasingly complex applications to our platform. And deploy they did! But the increases have a cost:

  • Increasing script size increases the amount of data we must transfer from script storage to the Workers runtime.

  • Increasing script size also increases the time complexity of the script compilation phase.

  • Increasing the startup CPU time limit increases the maximum top-level execution time.

Taken together, cold starts for complex applications began to lose the TLS handshake race.

Routing requests to an existing Worker

With relaxed script size and startup time limits, optimizing cold start time directly was a losing battle. Instead, we needed to figure out how to reduce the absolute number of cold starts, so that requests are simply less likely to incur one.

One option is to route requests to existing Worker instances, where before we might have chosen to start a new instance.

Previously, we weren’t particularly good at routing requests to existing Worker instances. We could trivially coalesce requests to a single Worker instance if they happened to land on a machine which already hosted a Worker, because in that case it’s not a distributed systems problem. But what if a Worker already existed in our data center on a different server, and some other server received a request for the Worker? We would always choose to cold start a new Worker on the machine which received the request, rather than forward the request to the machine with the already-existing Worker, even though forwarding the request would avoid the cold start.


To drive the point home: Imagine a visitor sends one request per minute to a data center with 300 servers, and that the traffic is load balanced evenly across all servers. On average, each server will receive one request every five hours. In particularly busy data centers, this span of time could be long enough that we need to evict the Worker to re-use its resources, resulting in a 100% cold start rate. That’s a terrible experience for the visitor.

Consequently, we found ourselves explaining to users, who saw high latency while prototyping their applications, that their latency would counterintuitively decrease once they put sufficient traffic on our network. This highlighted the inefficiency in our original, simple design.

If, instead, those requests were all coalesced onto one single server, we would notice multiple benefits. The Worker would receive one request per minute, which is short enough to virtually guarantee that it won’t be evicted. This would mean the visitor may experience a single cold start, and then have a 100% “warm request rate.” We would also use 99.7% (299 / 300) less memory serving this traffic. This makes room for other Workers, decreasing their eviction rate, and increasing their warm request rates, too – a virtuous cycle!

There’s a cost to coalescing requests to a single instance, though, right? After all, we’re adding latency to requests if we have to proxy them around the data center to a different server.

In practice, the added time-to-first-byte is less than one millisecond, and is the subject of continual optimization by our IPC and performance teams. One millisecond is far less than a typical cold start, meaning it’s always better, in every measurable way, to proxy a request to a warm Worker than it is to cold start a new one.

The consistent hash ring

A solution to this very problem lies at the heart of many of our products, including one of our oldest: the HTTP cache in our Content Delivery Network.

When a visitor requests a cacheable web asset through Cloudflare, the request gets routed through a pipeline of proxies. One of those proxies is a caching proxy, which stores the asset for later, so we can serve it to future requests without having to request it from the origin again.

A Worker cold start is analogous to an HTTP cache miss, in that a request to a warm Worker is like an HTTP cache hit.

When our standard HTTP proxy pipeline routes requests to the caching layer, it chooses a cache server based on the request’s cache key to optimize the HTTP cache hit rate. The cache key is the request’s URL, plus some other details. This technique is often called “sharding”. The servers are considered to be individual shards of a larger, logical system – in this case a data center’s HTTP cache. So, we can say things like, “Each data center contains one logical HTTP cache, and that cache is sharded across every server in the data center.”

Until recently, we could not make the same claim about the set of Workers in a data center. Instead, each server contained its own standalone set of Workers, and they could easily duplicate effort.

We borrow the cache’s trick to solve that. In fact, we even use the same type of data structure used by our HTTP cache to choose servers: a consistent hash ring. A naive sharding implementation might use a classic hash table mapping Worker script IDs to server addresses. That would work fine for a set of servers which never changes. But servers are actually ephemeral and have their own lifecycle. They can crash, get rebooted, taken out for maintenance, or decommissioned. New ones can come online. When these events occur, the size of the hash table would change, necessitating a re-hashing of the whole table. Every Worker’s home server would change, and all sharded Workers would be cold started again!

A consistent hash ring improves this scenario significantly. Instead of establishing a direct correspondence between script IDs and server addresses, we map them both to a number line whose end wraps around to its beginning, also known as a ring. To look up the home server of a Worker, first we hash its script, and then we find where it lies on the ring. Next, we take the server address which comes directly on or after that position on the ring, and consider that the Worker’s home.


If a new server appears for some reason, all the Workers that lie before it on the ring get re-homed, but none of the other Workers are disturbed. Similarly, if a server disappears, all the Workers which lay before it on the ring get re-homed.


We refer to the Worker’s home server as the “shard server”. In request flows involving sharding, there is also a “shard client”. It’s also a server! The shard client initially receives a request, and, using its consistent hash ring, looks up which shard server it should send the request to. I’ll be using these two terms – shard client and shard server – in the rest of this post.

Handling overload

The nature of HTTP assets lend themselves well to sharding. If they are cacheable, they are static, at least for their cache Time to Live (TTL) duration. So, serving them requires time and space complexity which scales linearly with their size.

But Workers aren’t JPEGs. They are live units of compute which can use up to five minutes of CPU time per request. Their time and space complexity do not necessarily scale with their input size, and can vastly outstrip the amount of computing power we must dedicate to serving even a huge file from cache.

This means that individual Workers can easily get overloaded when given sufficient traffic. So, no matter what we do, we need to keep in mind that we must be able to scale back up to infinity. We will never be able to guarantee that a data center has only one instance of a Worker, and we must always be able to horizontally scale at the drop of a hat to support burst traffic. Ideally this is all done without producing any errors.

This means that a shard server must have the ability to refuse requests to invoke Workers on it, and shard clients must always gracefully handle this scenario.

Two load shedding options

I am aware of two general solutions to shedding load gracefully, without serving errors.

In the first solution, the client asks politely if it may issue the request. It then sends the request if it  receives a positive response. If it instead receives a “go away” response, it handles the request differently, like serving it locally. In HTTP, this pattern can be found in Expect: 100-continue semantics. The main downside is that this introduces one round-trip of latency to set the expectation of success before the request can be sent. (Note that a common naive solution is to just retry requests. This works for some kinds of requests, but is not a general solution, as requests may carry arbitrarily large bodies.)


The second general solution is to send the request without confirming that it can be handled by the server, then count on the server to forward the request elsewhere if it needs to. This could even be back to the client. This avoids the round-trip of latency that the first solution incurs, but there is a tradeoff: It puts the shard server in the request path, pumping bytes back to the client. Fortunately, we have a trick to minimize the amount of bytes we actually have to send back in this fashion, which I’ll describe in the next section.


Optimistically sending sharded requests

There are a couple of reasons why we chose to optimistically send sharded requests without waiting for permission.

The first reason of note is that we expect to see very few of these refused requests in practice. The reason is simple: If a shard client receives a refusal for a Worker, then it must cold start the Worker locally. As a consequence, it can serve all future requests locally without incurring another cold start. So, after a single refusal, the shard client won’t shard that Worker any more (until traffic for the Worker tapers off enough for an eviction, at least).

Generally, this means we expect that if a request gets sharded to a different server, the shard server will most likely accept the request for invocation. Since we expect success, it makes a lot more sense to optimistically send the entire request to the shard server than it does to incur a round-trip penalty to establish permission first.

The second reason is that we have a trick to avoid paying too high a cost for proxying the request back to the client, as I mentioned above.

We implement our cross-instance communication in the Workers runtime using Cap’n Proto RPC, whose distributed object model enables some incredible features, like JavaScript-native RPC. It is also the elder, spiritual sibling to the just-released Cap’n Web.

In the case of sharding, Cap’n Proto makes it very easy to implement an optimal request refusal mechanism. When the shard client assembles the sharded request, it includes a handle (called a capability in Cap’n Proto) to a lazily-loaded local instance of the Worker. This lazily-loaded instance has the same exact interface as any other Worker exposed over RPC. The difference is just that it’s lazy – it doesn’t get cold started until invoked. In the event the shard server decides it must refuse the request, it does not return a “go away” response, but instead returns the shard client’s own lazy capability!

The shard client’s application code only sees that it received a capability from the shard server. It doesn’t know where that capability is actually implemented. But the shard client’s RPC system does know where the capability lives! Specifically, it recognizes that the returned capability is actually a local capability – the same one that it passed to the shard server. Once it realizes this, it also realizes that any request bytes it continues to send to the shard server will just come looping back. So, it stops sending more request bytes, waits to receive back from the shard server all the bytes it already sent, and shortens the request path as soon as possible. This takes the shard server entirely out of the loop, preventing a “trombone effect.”

Workers invoking Workers

With load shedding behavior figured out, we thought the hard part was over.

But, of course, Workers may invoke other Workers. There are many ways this could occur, most obviously via Service Bindings. Less obviously, many of our favorite features, such as Workers KV, are actually cross-Worker invocations. But there is one product, in particular, that stands out for its powerful ability to invoke other Workers: Workers for Platforms.

Workers for Platforms allows you to run your own functions-as-a-service on Cloudflare infrastructure. To use the product, you deploy three special types of Workers:

  • a dynamic dispatch Worker

  • any number of user Workers

  • an optional, parameterized outbound Worker

A typical request flow for Workers for Platforms goes like so: First, we invoke the dynamic dispatch Worker. The dynamic dispatch Worker chooses and invokes a user Worker. Then, the user Worker invokes the outbound Worker to intercept its subrequests. The dynamic dispatch Worker chose the outbound Worker’s arguments prior to invoking the user Worker.

To really amp up the fun, the dynamic dispatch Worker could have a tail Worker attached to it. This tail Worker would need to be invoked with traces related to all the preceding invocations. Importantly, it should be invoked one single time with all events related to the request flow, not invoked multiple times for different fragments of the request flow.

You might further ask, can you nest Workers for Platforms? I don’t know the official answer, but I can tell you that the code paths do exist, and they do get exercised.

To support this nesting doll of Workers, we keep a context stack during invocations. This context includes things like ownership overrides, resource limit overrides, trust levels, tail Worker configurations, outbound Worker configurations, feature flags, and so on. This context stack was manageable-ish when everything was executed on a single thread. For sharding to be truly useful, though, we needed to be able to move this context stack around to other machines.

Our choice of Cap’n Proto RPC as our primary communications medium helped us make sense of it all. To shard Workers deep within a stack of invocations, we serialize the context stack into a Cap’n Proto data structure and send it to the shard server. The shard server deserializes it into native objects, and continues the execution where things left off.

As with load shedding, Cap’n Proto’s distributed object model provides us simple answers to otherwise difficult questions. Take the tail Worker question – how do we coalesce tracing data from invocations which got fanned out across any number of other servers back to one single place? Easy: create a capability (a live Cap’n Proto object) for a reportTraces() callback on the dynamic dispatch Worker’s home server, and put that in the serialized context stack. Now, that context stack can be passed around at will. That context stack will end up in multiple places: At a minimum, it will end up on the user Worker’s shard server and the outbound Worker’s shard server. It may also find its way to other shard servers if any of those Workers invoked service bindings! Each of those shard servers can call the reportTraces() callback, and be confident that the data will make its way back to the right place: the dynamic dispatch Worker’s home server. None of those shard servers need to actually know where that home server is. Phew!

Eviction rates down, warm request rates up

Features like this are always satisfying to roll out, because they produce graphs showing huge efficiency gains.

Once fully rolled out, only about 4% of total requests from enterprise traffic ended up being sharded. To put that another way, 96% of all enterprise requests are to Workers which are sufficiently loaded that we must run multiple instances of them in a data center.


Despite that low total rate of sharding, we reduced our global Worker eviction rate by 10x. 


Our eviction rate is a measure of memory pressure within our system. You can think of it like garbage collection at a macro level, and it has the same implications. Fewer evictions means our system uses memory more efficiently. This has the happy consequence of using less CPU to clean up our memory. More relevant to Workers users, the increased efficiency means we can keep Workers in memory for an order of magnitude longer, improving their warm request rate and reducing their latency.

The high leverage shown – sharding just 4% of our traffic to improve memory efficiency by 10x – is a consequence of the power-law distribution of Internet traffic.

A power law distribution is a phenomenon which occurs across many fields of science, including linguistics, sociology, physics, and, of course, computer science. Events which follow power law distributions typically see a huge amount clustered in some small number of “buckets”, and the rest spread out across a large number of those “buckets”. Word frequency is a classic example: A small handful of words like “the”, “and”, and “it” occur in texts with extremely high frequency, while other words like “eviction” or “trombone” might occur only once or twice in a text.

In our case, the majority of Workers requests goes to a small handful of high-traffic Workers, while a very long tail goes to a huge number of low-traffic Workers. The 4% of requests which were sharded are all to low-traffic Workers, which are the ones that benefit the most from sharding.

So did we eliminate cold starts? Or will there be an Eliminating Cold Starts 3 in our future?


For enterprise traffic, our warm request rate increased from 99.9% to 99.99% – that’s three 9’s to four 9’s. Conversely, this means that the cold start rate went from 0.1% to 0.01% of requests, a 10x decrease. A moment’s thought, and you’ll realize that this is coherent with the eviction rate graph I shared above: A 10x decrease in the number of Workers we destroy over time must imply we’re creating 10x fewer to begin with.

Simultaneously, our warm request rate became less volatile throughout the course of the day.

Hmm.

I hate to admit this to you, but I still notice a little bit of space at the top of the graph. 😟

Can you help us get to five 9’s?

Code Mode: the better way to use MCP

Post Syndicated from Kenton Varda original https://blog.cloudflare.com/code-mode/

It turns out we’ve all been using MCP wrong.

Most agents today use MCP by directly exposing the “tools” to the LLM.

We tried something different: Convert the MCP tools into a TypeScript API, and then ask an LLM to write code that calls that API.

The results are striking:

  1. We found agents are able to handle many more tools, and more complex tools, when those tools are presented as a TypeScript API rather than directly. Perhaps this is because LLMs have an enormous amount of real-world TypeScript in their training set, but only a small set of contrived examples of tool calls.

  2. The approach really shines when an agent needs to string together multiple calls. With the traditional approach, the output of each tool call must feed into the LLM’s neural network, just to be copied over to the inputs of the next call, wasting time, energy, and tokens. When the LLM can write code, it can skip all that, and only read back the final results it needs.

In short, LLMs are better at writing code to call MCP, than at calling MCP directly.

What’s MCP?

For those that aren’t familiar: Model Context Protocol is a standard protocol for giving AI agents access to external tools, so that they can directly perform work, rather than just chat with you.

Seen another way, MCP is a uniform way to:

  • expose an API for doing something,

  • along with documentation needed for an LLM to understand it,

  • with authorization handled out-of-band.

MCP has been making waves throughout 2025 as it has suddenly greatly expanded the capabilities of AI agents.

The “API” exposed by an MCP server is expressed as a set of “tools”. Each tool is essentially a remote procedure call (RPC) function – it is called with some parameters and returns a response. Most modern LLMs have the capability to use “tools” (sometimes called “function calling”), meaning they are trained to output text in a certain format when they want to invoke a tool. The program invoking the LLM sees this format and invokes the tool as specified, then feeds the results back into the LLM as input.

Anatomy of a tool call

Under the hood, an LLM generates a stream of “tokens” representing its output. A token might represent a word, a syllable, some sort of punctuation, or some other component of text.

A tool call, though, involves a token that does not have any textual equivalent. The LLM is trained (or, more often, fine-tuned) to understand a special token that it can output that means “the following should be interpreted as a tool call,” and another special token that means “this is the end of the tool call.” Between these two tokens, the LLM will typically write tokens corresponding to some sort of JSON message that describes the call.

For instance, imagine you have connected an agent to an MCP server that provides weather info, and you then ask the agent what the weather is like in Austin, TX. Under the hood, the LLM might generate output like the following. Note that here we’ve used words in <| and |> to represent our special tokens, but in fact, these tokens do not represent text at all; this is just for illustration.

I will use the Weather MCP server to find out the weather in Austin, TX.

I will use the Weather MCP server to find out the weather in Austin, TX.

<|tool_call|>
{
  "name": "get_current_weather",
  "arguments": {
    "location": "Austin, TX, USA"
  }
}
<|end_tool_call|>

Upon seeing these special tokens in the output, the LLM’s harness will interpret the sequence as a tool call. After seeing the end token, the harness pauses execution of the LLM. It parses the JSON message and returns it as a separate component of the structured API result. The agent calling the LLM API sees the tool call, invokes the relevant MCP server, and then sends the results back to the LLM API. The LLM’s harness will then use another set of special tokens to feed the result back into the LLM:

<|tool_result|>
{
  "location": "Austin, TX, USA",
  "temperature": 93,
  "unit": "fahrenheit",
  "conditions": "sunny"
}
<|end_tool_result|>

The LLM reads these tokens in exactly the same way it would read input from the user – except that the user cannot produce these special tokens, so the LLM knows it is the result of the tool call. The LLM then continues generating output like normal.

Different LLMs may use different formats for tool calling, but this is the basic idea.

What’s wrong with this?

The special tokens used in tool calls are things LLMs have never seen in the wild. They must be specially trained to use tools, based on synthetic training data. They aren’t always that good at it. If you present an LLM with too many tools, or overly complex tools, it may struggle to choose the right one or to use it correctly. As a result, MCP server designers are encouraged to present greatly simplified APIs as compared to the more traditional API they might expose to developers.

Meanwhile, LLMs are getting really good at writing code. In fact, LLMs asked to write code against the full, complex APIs normally exposed to developers don’t seem to have too much trouble with it. Why, then, do MCP interfaces have to “dumb it down”? Writing code and calling tools are almost the same thing, but it seems like LLMs can do one much better than the other?

The answer is simple: LLMs have seen a lot of code. They have not seen a lot of “tool calls”. In fact, the tool calls they have seen are probably limited to a contrived training set constructed by the LLM’s own developers, in order to try to train it. Whereas they have seen real-world code from millions of open source projects.

Making an LLM perform tasks with tool calling is like putting Shakespeare through a month-long class in Mandarin and then asking him to write a play in it. It’s just not going to be his best work.

But MCP is still useful, because it is uniform

MCP is designed for tool-calling, but it doesn’t actually have to be used that way.

The “tools” that an MCP server exposes are really just an RPC interface with attached documentation. We don’t really have to present them as tools. We can take the tools, and turn them into a programming language API instead.

But why would we do that, when the programming language APIs already exist independently? Almost every MCP server is just a wrapper around an existing traditional API – why not expose those APIs?

Well, it turns out MCP does something else that’s really useful: It provides a uniform way to connect to and learn about an API.

An AI agent can use an MCP server even if the agent’s developers never heard of the particular MCP server, and the MCP server’s developers never heard of the particular agent. This has rarely been true of traditional APIs in the past. Usually, the client developer always knows exactly what API they are coding for. As a result, every API is able to do things like basic connectivity, authorization, and documentation a little bit differently.

This uniformity is useful even when the AI agent is writing code. We’d like the AI agent to run in a sandbox such that it can only access the tools we give it. MCP makes it possible for the agentic framework to implement this, by handling connectivity and authorization in a standard way, independent of the AI code. We also don’t want the AI to have to search the Internet for documentation; MCP provides it directly in the protocol.

OK, how does it work?

We have already extended the Cloudflare Agents SDK to support this new model!

For example, say you have an app built with ai-sdk that looks like this:

const stream = streamText({
  model: openai("gpt-5"),
  system: "You are a helpful assistant",
  messages: [
    { role: "user", content: "Write a function that adds two numbers" }
  ],
  tools: {
    // tool definitions 
  }
})

You can wrap the tools and prompt with the codemode helper, and use them in your app: 

import { codemode } from "agents/codemode/ai";

const {system, tools} = codemode({
  system: "You are a helpful assistant",
  tools: {
    // tool definitions 
  },
  // ...config
})

const stream = streamText({
  model: openai("gpt-5"),
  system,
  tools,
  messages: [
    { role: "user", content: "Write a function that adds two numbers" }
  ]
})

With this change, your app will now start generating and running code that itself will make calls to the tools you defined, MCP servers included. We will introduce variants for other libraries in the very near future. Read the docs for more details and examples. 

Converting MCP to TypeScript

When you connect to an MCP server in “code mode”, the Agents SDK will fetch the MCP server’s schema, and then convert it into a TypeScript API, complete with doc comments based on the schema.

For example, connecting to the MCP server at https://gitmcp.io/cloudflare/agents, will generate a TypeScript definition like this:

interface FetchAgentsDocumentationInput {
  [k: string]: unknown;
}
interface FetchAgentsDocumentationOutput {
  [key: string]: any;
}

interface SearchAgentsDocumentationInput {
  /**
   * The search query to find relevant documentation
   */
  query: string;
}
interface SearchAgentsDocumentationOutput {
  [key: string]: any;
}

interface SearchAgentsCodeInput {
  /**
   * The search query to find relevant code files
   */
  query: string;
  /**
   * Page number to retrieve (starting from 1). Each page contains 30
   * results.
   */
  page?: number;
}
interface SearchAgentsCodeOutput {
  [key: string]: any;
}

interface FetchGenericUrlContentInput {
  /**
   * The URL of the document or page to fetch
   */
  url: string;
}
interface FetchGenericUrlContentOutput {
  [key: string]: any;
}

declare const codemode: {
  /**
   * Fetch entire documentation file from GitHub repository:
   * cloudflare/agents. Useful for general questions. Always call
   * this tool first if asked about cloudflare/agents.
   */
  fetch_agents_documentation: (
    input: FetchAgentsDocumentationInput
  ) => Promise<FetchAgentsDocumentationOutput>;

  /**
   * Semantically search within the fetched documentation from
   * GitHub repository: cloudflare/agents. Useful for specific queries.
   */
  search_agents_documentation: (
    input: SearchAgentsDocumentationInput
  ) => Promise<SearchAgentsDocumentationOutput>;

  /**
   * Search for code within the GitHub repository: "cloudflare/agents"
   * using the GitHub Search API (exact match). Returns matching files
   * for you to query further if relevant.
   */
  search_agents_code: (
    input: SearchAgentsCodeInput
  ) => Promise<SearchAgentsCodeOutput>;

  /**
   * Generic tool to fetch content from any absolute URL, respecting
   * robots.txt rules. Use this to retrieve referenced urls (absolute
   * urls) that were mentioned in previously fetched documentation.
   */
  fetch_generic_url_content: (
    input: FetchGenericUrlContentInput
  ) => Promise<FetchGenericUrlContentOutput>;
};

This TypeScript is then loaded into the agent’s context. Currently, the entire API is loaded, but future improvements could allow an agent to search and browse the API more dynamically – much like an agentic coding assistant would.

Running code in a sandbox

Instead of being presented with all the tools of all the connected MCP servers, our agent is presented with just one tool, which simply executes some TypeScript code.

The code is then executed in a secure sandbox. The sandbox is totally isolated from the Internet. Its only access to the outside world is through the TypeScript APIs representing its connected MCP servers.

These APIs are backed by RPC invocation which calls back to the agent loop. There, the Agents SDK dispatches the call to the appropriate MCP server.

The sandboxed code returns results to the agent in the obvious way: by invoking console.log(). When the script finishes, all the output logs are passed back to the agent.


Dynamic Worker loading: no containers here

This new approach requires access to a secure sandbox where arbitrary code can run. So where do we find one? Do we have to run containers? Is that expensive?

No. There are no containers. We have something much better: isolates.

The Cloudflare Workers platform has always been based on V8 isolates, that is, isolated JavaScript runtimes powered by the V8 JavaScript engine.

Isolates are far more lightweight than containers. An isolate can start in a handful of milliseconds using only a few megabytes of memory.

Isolates are so fast that we can just create a new one for every piece of code the agent runs. There’s no need to reuse them. There’s no need to prewarm them. Just create it, on demand, run the code, and throw it away. It all happens so fast that the overhead is negligible; it’s almost as if you were just eval()ing the code directly. But with security.

The Worker Loader API

Until now, though, there was no way for a Worker to directly load an isolate containing arbitrary code. All Worker code instead had to be uploaded via the Cloudflare API, which would then deploy it globally, so that it could run anywhere. That’s not what we want for Agents! We want the code to just run right where the agent is.

To that end, we’ve added a new API to the Workers platform: the Worker Loader API. With it, you can load Worker code on-demand. Here’s what it looks like:

// Gets the Worker with the given ID, creating it if no such Worker exists yet.
let worker = env.LOADER.get(id, async () => {
  // If the Worker does not already exist, this callback is invoked to fetch
  // its code.

  return {
    compatibilityDate: "2025-06-01",

    // Specify the worker's code (module files).
    mainModule: "foo.js",
    modules: {
      "foo.js":
        "export default {\n" +
        "  fetch(req, env, ctx) { return new Response('Hello'); }\n" +
        "}\n",
    },

    // Specify the dynamic Worker's environment (`env`).
    env: {
      // It can contain basic serializable data types...
      SOME_NUMBER: 123,

      // ... and bindings back to the parent worker's exported RPC
      // interfaces, using the new `ctx.exports` loopback bindings API.
      SOME_RPC_BINDING: ctx.exports.MyBindingImpl({props})
    },

    // Redirect the Worker's `fetch()` and `connect()` to proxy through
    // the parent worker, to monitor or filter all Internet access. You
    // can also block Internet access completely by passing `null`.
    globalOutbound: ctx.exports.OutboundProxy({props}),
  };
});

// Now you can get the Worker's entrypoint and send requests to it.
let defaultEntrypoint = worker.getEntrypoint();
await defaultEntrypoint.fetch("http://example.com");

// You can get non-default entrypoints as well, and specify the
// `ctx.props` value to be delivered to the entrypoint.
let someEntrypoint = worker.getEntrypoint("SomeEntrypointClass", {
  props: {someProp: 123}
});

You can start playing with this API right now when running workerd locally with Wrangler (check out the docs), and you can sign up for beta access to use it in production.

Workers are better sandboxes

The design of Workers makes it unusually good at sandboxing, especially for this use case, for a few reasons:

Faster, cheaper, disposable sandboxes

The Workers platform uses isolates instead of containers. Isolates are much lighter-weight and faster to start up. It takes mere milliseconds to start a fresh isolate, and it’s so cheap we can just create a new one for every single code snippet the agent generates. There’s no need to worry about pooling isolates for reuse, prewarming, etc.

We have not yet finalized pricing for the Worker Loader API, but because it is based on isolates, we will be able to offer it at a significantly lower cost than container-based solutions.

Isolated by default, but connected with bindings

Workers are just better at handling isolation.

In Code Mode, we prohibit the sandboxed worker from talking to the Internet. The global fetch() and connect() functions throw errors.

But on most platforms, this would be a problem. On most platforms, the way you get access to private resources is, you start with general network access. Then, using that network access, you send requests to specific services, passing them some sort of API key to authorize private access.

But Workers has always had a better answer. In Workers, the “environment” (env object) doesn’t just contain strings, it contains live objects, also known as “bindings”. These objects can provide direct access to private resources without involving generic network requests.

In Code Mode, we give the sandbox access to bindings representing the MCP servers it is connected to. Thus, the agent can specifically access those MCP servers without having network access in general.

Limiting access via bindings is much cleaner than doing it via, say, network-level filtering or HTTP proxies. Filtering is hard on both the LLM and the supervisor, because the boundaries are often unclear: the supervisor may have a hard time identifying exactly what traffic is legitimately necessary to talk to an API. Meanwhile, the LLM may have difficulty guessing what kinds of requests will be blocked. With the bindings approach, it’s well-defined: the binding provides a JavaScript interface, and that interface is allowed to be used. It’s just better this way.

No API keys to leak

An additional benefit of bindings is that they hide API keys. The binding itself provides an already-authorized client interface to the MCP server. All calls made on it go to the agent supervisor first, which holds the access tokens and adds them into requests sent on to MCP.

This means that the AI cannot possibly write code that leaks any keys, solving a common security problem seen in AI-authored code today.

Try it now!

Sign up for the production beta

The Dynamic Worker Loader API is in closed beta. To use it in production, sign up today.

Or try it locally

If you just want to play around, though, Dynamic Worker Loading is fully available today when developing locally with Wrangler and workerd – check out the docs for Dynamic Worker Loading and code mode in the Agents SDK to get started.

Introducing new regional Internet traffic and Certificate Transparency insights on Cloudflare Radar

Post Syndicated from David Belson original https://blog.cloudflare.com/new-regional-internet-traffic-and-certificate-transparency-insights-on-radar/

Since launching during Birthday Week in 2020, Radar has announced significant new capabilities and data sets during subsequent Birthday Weeks. We continue that tradition this year with a two-part launch, adding more dimensions to Radar’s ability to slice and dice the Internet.

First, we’re adding regional traffic insights. Regional traffic insights bring a more localized perspective to the traffic trends shown on Radar.

Second, we’re adding detailed Certificate Transparency (CT) data, too. The new CT data builds on the work that Cloudflare has been doing around CT since 2018, including Merkle Town, our initial CT dashboard.

Both features extend Radar’s mission of providing deeper, more granular visibility into the health and security of the Internet. Below, we dig into these new capabilities and data sets.

Introducing regional Internet traffic insights on Radar

Cloudflare Radar initially launched with visibility into Internet traffic trends at a national level: want to see how that Internet shutdown impacted traffic in Iraq, or what IPv6 adoption looks like in India? It’s visible on Radar. Just a year and a half later, in March 2022, we launched Autonomous System (ASN) pages on Radar. This has enabled us to bring more granular visibility to many of our metrics: What’s network performance like on AS701 (Verizon Fios)? How thoroughly has AS812 (Rogers Communications) implemented routing security? Did AS58322 (Halasat) just go offline? It’s all visible on Radar.

However, sometimes Internet usage shifts on a more local level — maybe a sporting event in a particular region drives people online to find out more information. Or maybe a storm or other natural disaster causes infrastructure damage and power outages in a given state, impacting Internet traffic.

For the last few years, the Radar team relied on internal data sets and Jupyter notebooks to visualize these “sub-national” traffic shifts. But today, we are bringing that insight to Cloudflare Radar, and to you, with the launch of regional traffic insights. With this new capability, you’ll be able to see traffic trends at a more local level, including bytes and requests, as well as breakouts of desktop/mobile device and bot/human traffic shares. And for even more granular visibility, within the Data Explorer, you’ll also be able to select an autonomous system to join with the regional selection — for example, looking at AS7922 (Comcast) in Massachusetts (United States).

Geographic guidance

In line with common industry practice, the region names displayed on Radar are sourced in data from GeoNames (geonames.org), a crowdsourced geographical database. Specifically, we are using the “first-order administrative divisions” listed for each country — for example, the states of America, the departments of Honduras, or the provinces of Canada. Those geographical names reflect data provided by GeoNames; for more information, please refer to their About page.

Requests logged by Cloudflare’s services include the IP address of the device making the request. The address range (“prefix”) that includes this address is associated with a GeoNames ID within our IP address geolocation data, and we then match that GeoNames ID with the associated country and “first order administrative division” found in the GeoNames dataset. (For example: 155.246.1.142 → 155.246.0.0/16 → GeoNames ID 5101760 → United States > New Jersey) 


Drilling down into Radar traffic data

Within Cloudflare Radar, there are several ways to get to this regional data. If you know the name of the region of interest, you can type it into the search bar at the top of the page, and select it from the results. For example, beginning to type Massachusetts returns the U.S. state, linked to its regional traffic page. Typing the region name into the Traffic in dropdown at the top of a Traffic page will also return the same set of results.


Radar’s country-level pages now have a new Traffic characteristics by region card that includes both summary and time series views of regional traffic. The summary view is presented as a map and table, similar to the Traffic characteristics card in the Worldwide traffic view. After selecting a metric from the dropdown at the top right of the card, the table and map are updated to reflect the relevant summary values for the chosen time period. Within the paginated table, the region names are linked, and clicking one will take you to the relevant page. Within the map, the summary values are represented by circles placed in the centroid of each region, sized in relation to their value. Clicking a circle will take you to the relevant page.


Below the summary map and table, the card also includes a time series graph of traffic at a regional level for the top five highest traffic regions within the country. These graphs can reveal interesting regional differences in traffic patterns. For example, the Traffic volume by region in Iraq graph for HTTP request traffic shown below highlights the differing Internet shutdown schedules (Kurdistan Region, central and southern Iraq) across the different governorates. On days when the schedules do not overlap, such as September 2 and 7, traffic from the Erbil and Sulaymaniyah governorates, which are located in the Kurdistan Region, does not drop concurrent with the loss in traffic observed in Baghdad and Basra.


Mobile vs. desktop device traffic trends

Over the past several years, a number of Radar blog posts have explored how human activity impacts Internet traffic, including holiday celebrations, elections, and the Paris 2024 Summer Olympics. With the new regional views, this impact now becomes even clearer at a more local level. For instance, mobile devices account for, on average, just over half of the request traffic seen from Nairobi Country in Kenya. A clear diurnal pattern is seen on weekdays, where mobile device usage drops during workday hours, and then rises again in the evening. However, during the weekends, mobile traffic remains elevated, presumably due to fewer people using desktop computers in office environments, as well as fewer desktop computers in use at home, in line with Kenya’s mobile-first culture.


Bot vs human traffic trends

Similar to how the mobile vs. desktop view exposes shifts in human activity, bot vs. human traffic insights do as well. One interpretation of the graph below is that overnight bot activity from Lisbon increased significantly during the first few days of September. However, since the graph shows traffic shares, and given the timing of the apparent increases, the more likely cause is increasingly larger drops in human-driven traffic – users in Lisbon appear to begin logging off around 23:00 UTC (midnight local time), and start getting back online around 05:00 UTC (06:00 local time). The shares and shifts will obviously vary by country and region, but they can provide a perspective on the nocturnal habits of users in a region.


Customize regional analysis with Radar’s Data Explorer

Within the Data Explorer, you can use the breakdown options and filters to customize your analysis of regional traffic data.

At a country level, choosing to breakdown by regions generates a stacked area graph that shows the relative traffic shares of the top 20 regions in the selected country, along with a bar graph showing summary share values. For example, the graph below shows that in aggregate, Virginia and California are responsible for just over a quarter of the HTTP request volume in the United States.



You can also use Data Explorer to drill down on traffic at a network (ASN) level in a given region, in both summary and timeseries views. For example, looking at HTTP request traffic for Massachusetts by ASN, we can see that AS7922 (Comcast), accounts for a third, followed by AS701 (Verizon Fios, 15%), AS21928 (T-Mobile, 8.8%), AS6167 (Verizon Wireless, 5.1%), AS7018 (AT&T, 4.7%), and AS20115 (Charter/Spectrum, 4.5%). Over 70% of the request traffic is concentrated in these six providers, with nearly half of that from one provider.



Going a level deeper, you can also look at traffic trends over time for an ASN within a given region, and even compare it with another time period. The graph below shows traffic for AS7922 (Comcast) in Massachusetts over a seven-day period, compared with the prior week. While the traffic volumes on most days were largely in line with the previous week, Saturday and Sunday were noticeably higher. These differences may reflect a shift in human activity, as September 6 & 7 were quite rainy in Massachusetts, so people may have spent more time indoors and online. (The prior weekend was Labor Day weekend, but those Saturday and Sunday traffic levels were in line with the preceding weekend.) You can also add another ASN to the traffic trends comparison. Selecting Massachusetts (Location) and AS701 (ASN) (Verizon Fios) in the Compare section finds that traffic on that network was higher on Saturday and Sunday as well, lending credence to the rainy weekend theory.



Regional comparisons, whether within the same country or across different countries, are also possible in Data Explorer. For instance, if the Kansas City Chiefs and Philadelphia Eagles were to meet yet again in the Super Bowl, the configuration below could be used to compare traffic patterns in the teams’ respective home states, as well as comparing the trends with the previous week, showing how human activity impacted it over the course of the game.


As always, the data powering the visualizations described above are also available through the Radar API. The timeseries_groups and summary methods for the NetFlows and HTTP endpoints now have an ADM1 dimension, allowing traffic to be broken down by first-order administrative divisions. In addition, the new geoId filter for the NetFlows and HTTP endpoints allows you to filter the results by a specific geolocation, using its GeoNames ID. And finally, there are new get and list endpoints for fetching geolocation details.

A note regarding data quantity and quality

As you’d expect, the more traffic we see from a given geography, the better the “signal”, and the clearer the associated graph is — this is generally the case when traffic is aggregated at a country level. However, for some smaller or less populous regions, especially in developing countries or countries with poor Internet connectivity, lower traffic will likely cause the signal to be weaker, resulting in graphs that appear spiky or incomplete. (Note that this will also be true for region+ASN views.) An illustrative example is shown below, for Northern Darfur State in Sudan. Traffic is observed somewhat inconsistently, resulting in the spikes seen in the graph. Similarly, the “Previous 7 days” line is largely incomplete, indicating a lack of traffic data for that period. In these cases, it will be hard to draw definitive conclusions from such graphs.


Although the Internet arguably transcends geographical boundaries, the reality is that usage patterns can vary by location, with traffic trends that reflect more localized human activity. The new regional insights on Cloudflare Radar traffic pages, and in the Data Explorer, provide a perspective at a sub-national level. We are exploring the potential to go a level deeper in the future, providing traffic data for “second-order administrative divisions” (such as counties, cities, etc.).

If you share our regional traffic graphs on social media, be sure to tag us: @CloudflareRadar (X), noc.social/@cloudflareradar (Mastodon), and radar.cloudflare.com (Bluesky). If you have questions or comments, you can reach out to us on social media, or contact us via email.

Introducing Certificate Transparency insights on Radar

Just as we’re bringing more granular detail to traffic patterns, we’re also shedding more light on the very foundation of trust on the Internet: TLS certificates. Certificate Authorities (CAs) serve as trusted gatekeepers for the Internet: any website that wants to prove its identity to clients must present a certificate issued by a CA that the client trusts. But how do we know that CAs themselves are trustworthy and only issue certificates they are authorized to issue?

That’s where Certificate Transparency (CT) comes in. Clients that enforce CT (most major browsers) will only trust a website certificate if it is both signed by a trusted CA and has proof that the certificate has been added to a public, append-only CT log, so that it can be publicly audited. Only recently, CT played a key role in detecting the unauthorized issuance of certificates for 1.1.1.1, a public DNS resolver service that Cloudflare operates.

In addition to its role as a vital safety mechanism for the Internet, CT has proven to be invaluable in other ways, as it provides publicly-accessible lists of all website certificates used on the Internet. This dataset is a treasure trove of intelligence for researchers measuring the Internet, security teams detecting malicious activity like phishing campaigns, or penetration testers mapping a target’s external attack surface.

The sheer amount of data (multiple terabytes) available in CT makes it difficult for regular Internet users to download and explore themselves. Instead, services like crt.sh, Censys, and Merklemap provide easy search interfaces to allow discoverability for specific domain names and certificates. We launched Merkle Town in 2018 to share broad insights into the CT ecosystem using data from our own CT monitoring service.

Certificate Transparency on Cloudflare Radar is the next evolution of Merkle Town, providing integration with security and domain information already on Radar and more interactive ways to explore and analyze CT data. (For long-time Merkle Town users, we’re keeping it around until we’ve reached full feature parity.)

In the sections below, we’ll walk you through the features available in the new dashboard.

Certificate volume and characteristics

The CT page leads with a view of how many certificates are being issued and logged over time. Because the same certificate can appear multiple times within a single log or be submitted to several logs, the total count can be inflated. To address this, two distinct lines are shown: one for total entries and another for unique entries. Uniqueness, however, is calculated only within the selected time range — for example, if certificate C is added to log A in one period and to log B in another, it will appear in the unique count for both periods. It is also important to note that the CT charts and date filters use the log timestamp, which is the time a certificate was added to a CT log. Additionally, the data displayed on the page was collected from the logs monitored by Cloudflare — delays, backlogs, or other inconsistencies may exist, so please report any issues or discrepancies.

Alongside this chart is a comparison between certificates and pre-certificates. A pre-certificate is a special type of certificate used in CT that allows a CA to publicly log a certificate before it is officially issued. CAs are not required to log full certificates if corresponding pre-certificates have already been logged (although many CAs do anyway), so typically there are more pre-certificates logged than full certificates, as seen in the chart.


While certificate issuance trends are interesting on their own, analyzing the characteristics of issued certificates provides deeper insight into the state of the web’s trust infrastructure. Starting with the public key algorithm, which defines how secure connections are established between clients and servers, we found that more than 65% of certificates still use RSA, while the remainder use ECDSA. RSA remains dominant due to its long-standing compatibility with a wide range of clients, while ECDSA is increasingly adopted for its efficiency and smaller key sizes, which can improve performance and reduce computational overhead. In the coming years, we expect post-quantum signature algorithms like ML-DSA to appear when public CAs begin to offer support.

Next, a breakdown of certificates by signature algorithm reveals how Certificate Authorities (CAs) sign the certificates they issue. Most certificates (over 65%) use RSA with SHA-256, followed by ECDSA with SHA-384 at 19%, ECDSA with SHA-256 at 12%, and a small fraction using other algorithms. The choice of signature algorithm reflects a balance between widespread support, security, and performance, with stronger algorithms like ECDSA gradually gaining traction for modern deployments.

Certificates are also categorized by validation level, which reflects the degree to which the CA has verified the identity of the certificate requester. The main validation types are Domain Validation (DV), Organization Validation (OV), and Extended Validation (EV). DV certificates verify only control of the domain, OV certificates verify both domain control and the organization behind it, and EV certificates involve more rigorous checks and display additional identity information in browsers. The industry trend is toward simpler, automated issuance, with DV certificates now making up almost 98% of issued certificates, while EV issuance has become largely obsolete.


Finally, the chart on certificate duration shows the difference between the NotBefore and NotAfter dates embedded in each certificate, which define the period during which the certificate is valid. Currently, the majority (92%) of issued certificates have durations between 47 and 100 days. Shorter certificate lifetimes improve security by limiting exposure if a certificate is compromised, and the industry is moving toward even shorter durations, driven by browser policies and automated renewal systems.


Certificate issuance

Certificate issuance is the process by which CAs generate certificates for domain owners. Many CAs are operated by larger organizations that manage multiple subordinate CAs under a single corporate umbrella. The CT page highlights the distribution of certificate issuance across the top CA owners. At the moment, the Internet Security Research Group (ISRG), also known as Let’s Encrypt, issues more than 66% of all certificates, followed by other widely used CA owners including Google Trust Services, Sectigo, and GoDaddy.


The impact of events like the July 21-22 Let’s Encrypt API outage due to internal DNS failures that significantly reduced certificate issuance rates are visible in this visualization, as issuance rates dropped significantly during the two-day period.


In addition to CA owners, the page provides a breakdown of certificate issuance by individual CA certificates. Among the top five CAs, Let’s Encrypt’s four intermediate CAs — R12, R13, E7, and E8 — represent the bulk of its issuance. The bar chart can also be filtered by CA owner to display only the certificates associated with a specified organization.


The CT section also offers dedicated CA-specific pages. By searching for a CA name or fingerprint in the top search bar, you can reach a page showing all insights and trends available on the main CT page, filtered by the selected CA. The page also includes an additional CA information card, which provides details such as the CA’s owner, revocation status, parent certificate, validity period, country, inclusion in public root stores, and a list of all CAs operated by the same owner. All of this information is derived from the Common CA Database (CCADB).


Certificate Transparency logs

Next on the CT page is a section focused on CT logs. This section shows the distribution of certificates across CT log operators, identifying the organizations that manage the infrastructure behind the logs. Over the last three months, Sectigo operated the logs containing the largest number of certificates (2.8 billion), followed by Google (2.5 billion), Cloudflare (1.6 billion), and Let’s Encrypt (1.4 billion). Note that the same certificate can be logged multiple times across CT logs, so organizations that operate multiple CT logs with overlapping acceptance criteria may log certificates at an elevated rate. As such, the relative rank of the operators in this graph should not be construed as a measure of how load-bearing the logs are within the ecosystem.


Below this, a bar chart displays the distribution of certificates across individual CT logs. Among the top five logs are Google’s xenon2025h1 and argon2025h2, Cloudflare’s nimbus2025, and Let’s Encrypt’s oak2025h2. This chart can also be filtered by operator to show only the logs associated with a specific owner. Next to the chart, another view shows the distribution of certificates by log API, distinguishing between logs following the original RFC 6962 API versus those compatible with the newer and more efficient static CT API.


Similar to the dedicated CA pages, the CT section also provides log-specific pages. By searching for a log name in the top search bar, you can access a page showing all insights and trends available on the main CT page, filtered by the selected log. Two additional cards are included: one showing information about the log, derived from Google Chrome’s log list, including details such as the operator, API type, documentation, and a list of other logs operated by the same organization; and another displaying performance metrics with two radar charts tracking uptime and response time over the past 90 days, as observed by Cloudflare’s CT monitor. These metrics are useful to determine if logs are meeting the ongoing requirements for inclusion in CT programs like Google’s.


Certificate coverage

Last but not least, the CT page includes a section on certificate coverage. Certificates can cover multiple top-level domains (TLDs), include wildcard entries, and support IP addresses in Subject Alternative Names (SANs).

The distribution of pre-certificates across the top 10 TLDs highlights the domains most commonly covered. .com leads with 45% of certificates, followed by other popular TLDs such as .dev and .net.

Next to this view, two half-donut charts provide further insights into certificate coverage: one shows the share of certificates that include wildcard entries — almost 25% of certificates use wildcards to cover multiple subdomains — while the other shows certificates that include IP addresses, revealing that the vast majority of certificates do not contain IPs in their SAN fields


Expanded domain certificate data 

The domain information page has also been updated to provide richer details about certificates. The certificates table, which displays certificates recorded in active CT logs for the specified domain, now includes expandable rows. Expanding a row reveals further information, including the certificate’s SHA-256 fingerprint, subject and issuer details — Common Name (CN), Organization (O), and Country (C) — the validity period (NotBefore and NotAfter), and the CT log where the certificate was found.


While the charts above highlight key insights in the CT ecosystem, all underlying data is accessible via the API and can be explored interactively across time periods, CAs, logs, and additional filters and dimensions using Radar’s Data Explorer. And as always, Radar charts and graphs can be downloaded for sharing or embedded directly into blogs, websites, and dashboards for further analysis. Don’t hesitate to reach out to us with feedback, suggestions, and feature requests — we’re already working through a list of early feedback from the CT community!

How Cloudflare uses the world’s greatest collection of performance data to make the world’s fastest global network even faster

Post Syndicated from Steve Goldsmith original https://blog.cloudflare.com/how-cloudflare-uses-the-worlds-greatest-collection-of-performance-data/

Cloudflare operates the fastest network on the planet. We’ve shared an update today about how we are overhauling the software technology that accelerates every server in our fleet, improving speed globally.

That is not where the work stops, though. To improve speed even further, we have to also make sure that our network swiftly handles the Internet-scale congestion that hits it every day, routing traffic to our now-faster servers.

We have invested in congestion control for years. Today, we are excited to share how we are applying a superpower of our network, our massive Free Plan user base, to optimize performance and find the best way to route traffic across our network for all our customers globally.

Early results have seen performance increases that average 10% faster than the prior baseline. We achieved this by applying different algorithmic methods to improve performance based on the data we observe about the Internet each day. We are excited to begin rolling out these improvements to all customers.

How does traffic arrive in our network?

The Internet is a massive collection of interconnected networks, each composed of many machines (“nodes”). Data is transmitted by breaking it up into small packets, and passing them from one machine to another (over a “link”). Each one of these machines is linked to many others, and each link has limited capacity.


When we send a packet over the Internet, it will travel in a series of “hops” over the links from A to B.  At any given time, there will be one link (one “hop”) with the least available capacity for that path. It doesn’t matter where in the connection this hop is — it will be the bottleneck.

But there’s a challenge — when you’re sending data over the Internet, you don’t know what route it’s going to take. In fact, each node decides for itself which route to send the traffic through, and different packets going from A to B can take entirely different routes. The dynamic and decentralized nature of the system is what makes the Internet so effective, but it also makes it very hard to work out how much data can be sent.  So — how can a sender know where the bottleneck is, and how fast to send data?

Between Cloudflare nodes, our Argo Smart Routing product takes advantage of our visibility into the global network to speed up communication. Similarly, when we initiate connections to customer origins, we can leverage Argo and other insights to optimize them. However, the speed of a connection from your phone or laptop (the Client below) to the nearest Cloudflare datacenter will depend on the capacity of the bottleneck hop in the chain from you to Cloudflare, which happens outside our network.


What happens when too much data arrives at once?

If too much data arrives at any one node in a network in the path of a request being processed, the requestor will experience delays due to congestion. The data will either be queued for a while (risking bufferbloat), or some of it will simply get dropped. Protocols like TCP and QUIC respond to packets being dropped by retransmitting the data, but this introduces a delay, and can even make the problem worse by further overloading the limited capacity.

If cloud infrastructure providers like Cloudflare don’t manage congestion carefully, we risk overloading the system, slowing down the rate of data getting through. This actually happened in the early days of the Internet. To avoid this, the Internet infrastructure community has developed systems for controlling congestion, which give everyone a turn to send their data, without overloading the network. This is an evolving challenge, as the network grows ever more complicated, and the best method to implement congestion control is a constant pursuit. Many different algorithms have been developed, which take different sources of information and signals, optimize in a particular method, and respond to congestion in different ways.

Congestion control algorithms use a number of signals to estimate the right rate to send traffic, without knowing how the network is set up. One important signal has been loss. When a packet is received, the receiver sends an “ACK,” telling the sender the packet got through. If it’s dropped somewhere along the way, the sender never gets the receipt, and after a timeout will treat the packet as having been lost.

More recent algorithms have used additional data. For example, a popular algorithm called BBR (Bottleneck Bandwidth and Round-trip propagation time), which we have been using for much of our traffic, attempts to build a model during each connection of the maximum amount of data that can be transmitted in a given time period, using estimates of the round trip time as well as loss information.

The best algorithm to use often depends on the workload. For example, for interactive traffic like a video call, an algorithm that biases towards sending too much traffic can cause queues to build up, leading to high latency and poor video experience. If one were to optimize solely for that use case though, and avoid that by sending less traffic, the network will not make the best use of the connection for clients doing bulk downloads. The performance optimization outcome varies, depending on a lot of different factors.  But – we have visibility into many of them!

BBR was an exciting development in congestion control approach, moving from reactive loss-based approaches to proactive model-based optimization, resulting in significantly better performance for modern networks. Our data gives us an opportunity to go further, applying different algorithmic methods to improve performance. 

How can we do better?

All the existing algorithms are constrained to use only information gathered during the lifetime of the current connection. Thankfully, we know far more about the Internet at any given moment than this!  With Cloudflare’s perspective on traffic, we see much more than any one customer or ISP might see at any given time.

Every day, we see traffic from essentially every major network on the planet. When a request comes into our system, we know what client device we’re talking to, what type of network is enabling the connection, and whether we’re talking to consumer ISPs or cloud infrastructure providers.

We know about the patterns of load across the global Internet, and the locations where we believe systems are overloaded, within our network, or externally. We know about the networks that have stable properties, which have high packet loss due to cellular data connections, and the ones that traverse low earth orbit satellite links and radically change their routes every 15 seconds. 

How does this work?

We have been in the process of migrating our network technology stack to use a new platform, powered by Rust, that provides more flexibility to experiment with varying the parameters in the algorithms used to handle congestion control. Then we needed data.

The data powering these experiments needs to reflect the measure we’re trying to optimize, which is the user experience. It’s not just enough that we’re sending data to nearly all the networks on the planet; we have to be able to see what is the experience that customers have. So how do we do that, at our scale?

First, we have detailed “passive” logs of the rate at which data is able to be sent from our network, and how long it takes for the destination to acknowledge receipt. This covers all our traffic, and gives us an idea of how quickly the data was received by the client, but doesn’t guarantee to tell us about the user experience.

Next, we have a system for gathering Real User Measurement (RUM) data, which records information in supported web browsers about metrics such as Page Load Time (PLT). Any Cloudflare customer can enable this and will receive detailed insights in their dashboard. In addition, we use this metadata in aggregate across all our customers and networks to understand what customers are really experiencing. 

However, RUM data is only going to be present for a small proportion of connections across our network. So, we’ve been working to find a way to predict the RUM measures by extrapolating from the data we see only in passive logs. For example, here are the results of an experiment we performed comparing two different algorithms against the cubic baseline.


Now, here’s the same timescale, observed through the prediction based on our passive logs. The curves are very similar – but even more importantly, the ratio between the curves is very similar. This is huge! We can use a relatively small amount of RUM data to validate our findings, but optimize our network in a much more fine-grained way by using the full firehose of our passive logs.


Extrapolating too far becomes unreliable, so we’re also working with some of our largest customers to improve our visibility of the behaviour of the network from their clients’ point of view, which allows us to extend this predictive model even further. In return, we’ll be able to give our customers insights into the true experience of their clients, in a way that no other platform can offer.

What is next?

We’re currently running our experiments and improved algorithms for congestion control on all of our free tier QUIC traffic. As we learn more, verify on more complex customers, and expand to TCP traffic, we’ll gradually roll this out to all our customers, for all traffic, over 2026 and beyond. The results have led to as much as a 10% improvement as compared to the baseline!

We’re working with a select group of enterprises to test this in an early access program. If you’re interested in learning more, contact us! 

Choice: the path to AI sovereignty

Post Syndicated from Carly Ramsey original https://blog.cloudflare.com/sovereign-ai-and-choice/

Every government is laser-focused on the potential for national transformation by AI. Many view AI as an unparalleled opportunity to solve complex national challenges, drive economic growth, and improve the lives of their citizens. Others are concerned about the risks AI can bring to its society and economy. Some sit somewhere between these two perspectives. But as plans are drawn up by governments around the world to address the question of AI development and adoption, all are grappling with the critical question of sovereignty — how much of this technology, mostly centered in the United States and China, needs to be in their direct control? 

Each nation has their own response to that question — some seek ‘self-sufficiency’ and total authority. Others, particularly those that do not have the capacity to build the full AI technology stack, are approaching it layer-by-layer, seeking to build on the capacities their country does have and then forming strategic partnerships to fill the gaps. 

We believe AI sovereignty at its core is about choice. Each nation should have the ability to select the right tools for the task, to control its own data, and to deploy applications at will, all without being locked into a single provider or a single way of doing things. It’s about autonomy and options, realized through a diversified, resilient digital supply chain. 

Cloudflare’s mission is to help build a better Internet. We make tools for developers around the world to build Internet and AI applications that are widely, and in many cases, freely, available. We work on standards to improve interoperability and prevent vendor lock-in. And we are global — our network spans 330 cities in over 125 countries. By supporting local developers to build and deploy AI tools and services right where they are, Cloudflare can help each nation on their path to greater AI sovereignty. 

Creating a future that enables many AI options

Many nations recognize the practical challenge of realizing a robust AI-driven future that incorporates sovereignty — the significant cost and complexity of the infrastructure needed to set AI in action. Cloudflare believes that countries can achieve their objectives by creating vibrant marketplaces that allow multiple options, and we are creating a path for governments that provides maximum choice:

Infrastructure accessibility: Countries often focus on building large data centers that have the compute capacity to train general purpose AI models, neglecting the infrastructure needed to effectively deploy AI. Because of their proximity to end users, distributed edge networks are critical to ensuring that consumers can actually use AI technologies at scale. Although some AI technologies will be designed to work on-device, many will need more power to run AI inference, the tasks that users ask an AI engine to complete.  Distributed networks are equipped to run AI workloads at the edge, to help deliver the low latency and high performance needed for advanced technologies. Cloudflare’s distributed network gives developers a path to rapidly deploy their apps globally without massive upfront investments. 

Inclusivity: Nations want their entire economies, from the small businesses, to research institutions, to non-profits and enterprises, to benefit from AI transformation. Serverless models like Cloudflare’s make it easy to get started. Developers pay only for what they use, rather than being locked into paying for expensive and unnecessary compute, dramatically lowering the barrier to entry. Our free tier allows developers to experiment, build, and even launch applications without any cost, while our pay-as-you-go model for increased usage removes the significant financial barriers that might otherwise keep advanced AI out of reach. 

Control over data: An important part of sovereignty is the ability to control your own data. We believe countries should avoid equating this type of control with data locality, focusing instead on integrating security tools that provide visibility and the ability to restrict access to data.  Cloudflare’s global, distributed network ensures that developers can experiment, build, and deploy AI-powered applications right where they are, setting rules and controls at the Internet edge.

Multi-modal, dynamic markets: Building new applications with closed AI models can make it challenging to switch models later, and can make developers dependent on particular providers. AI strategies must embrace diversity — developers should have access to a wide variety of both open source and closed AI models. Cloudflare’s Workers AI platform, with over 50 open source models, is model agnostic, helping to create a competitive, dynamic environment where developers can swap models in and out as better, cheaper, or more specialized options become available. Cloudflare’s AI Gateway allows our customers to connect and control all their AI models, regardless of vendor, in a single, unified, interoperable platform. 

Underpinning all of this is the importance of open standards that encourage interoperability. Open standards and protocols throughout the AI technology stack help prevent dependency, create dynamic and competitive markets, and create choice for governments and their developers.

Championing regional AI innovation  

Many countries have started to put their own mark on how to spur innovation in their markets, starting with large language models (LLM).  AI development to date has mostly centered around LLMs  trained on English-centric data, and increasingly, Chinese-centric data, leaving behind those who can’t fully access this technology in these two languages. Recognizing this gap, these nations are building and freely offering AI models trained on local language datasets that are fine-tuned to the nuances of their own cultures and languages. This approach lowers the barrier to entry for local businesses, organizations, and governments to create customized AI solutions for their specific markets. Open-sourcing these LLMs is to recognize that AI sovereignty is a means to an end. The goal is innovation, economic growth, and the ability to solve meaningful problems.

Cloudflare is now supporting these sovereign AI initiatives in India, Japan, and Southeast Asia. We are bringing these locally-developed, open-source AI models to developers around the world through our serverless inference platform, Workers AI.

India: India’s national vision is “AI for All”, which focuses on AI driving inclusive growth and social empowerment. India will host the momentous global AI Impact Summit in 2026, and a key element of showcasing empowering technological advancements that are accessible to the Global South. With its immense linguistic diversity, India is at the forefront of creating models that serve its hundreds of millions of Internet users in their native tongues. A cornerstone in this endeavor is the Government of India’s Bhashini, a digital public good platform that enables all Indian citizens to access the Internet and digital services in 22 official languages.

Cloudflare is now offering AI4Bharat’s IndicTrans2 model, a key open source language model that is also part of the Bhashini initiative. The model is able to translate text across 22 Indic languages, including Bengali, Gujarati, Hindi, Tamil, Sanskrit and even traditionally low-resourced languages like Kashmiri, Manipuri and Sindhi. 

You can use the @cf/ai4bharat/indictrans2-en-indic-1B model on Workers AI as follows:

curl --request POST \
  --url https://api.cloudflare.com/client/v4/accounts/ACCOUNT_ID/ai/run/@cf/ai4bharat/indictrans2-en-indic-1B \
  --header 'Authorization: Bearer TOKEN' \
  --header 'Content-Type: application/json' \
  --data '{
    "text": ["What is your favourite food?", "I like pizza"],
    "target_language": "guj_Gujr"
}'

Japan: Japan has a very clear and expansive vision of AI development. Concerned about Japan’s slow AI uptake, the Japanese government aims to make the country “the world’s most friendly AI nation” by creating the ideal conditions for AI growth, both at home and abroad. A major initiative for Japan’s government is supporting AI that deeply understands the complexities and cultural context of the Japanese language. 

Cloudflare is offering Preferred Networks, Inc.(PFN)  PLaMo-Embedding-1B, a home-grown Japanese text embedding model, made freely and openly available. The Japanese government supported PFN through its Generative AI Accelerator Challenge (GENIAC) program, which supports local LLM development through subsidized access to compute resources for training. The PLaMo Embedding model enables users to generate high-quality embeddings for Japanese text, which is helpful for building RAG-powered applications and semantic search use cases.

You can use the @cf/pfnet/plamo-embedding-1b model on Workers AI as follows:

curl --request POST \
  --url https://api.cloudflare.com/client/v4/accounts/ACCOUNT_ID/ai/run/@cf/pfnet/plamo-embedding-1b \
  --header 'Authorization: Bearer TOKEN' \
  --header 'Content-Type: application/json' \
  --data '{
  	"text": [
            "PLaMo-Embedding-1Bは、Preferred Networks, Inc. によって開発された日本語テキスト埋め込みモデルです。",
            "最近は随分と暖かくなりましたね。"
        ]
}'

Southeast Asia: As Chair of the Association of Southeast Asian Nations (ASEAN) Working Group on AI Governance, Singapore’s ambitious National AI Strategy 2.0 aims to ensure that AI is a public good, both for Southeast Asia and the world. As a cornerstone of this strategy, Singapore is championing the development and adoption of SEA-LION, a family of open-source LLMs designed for Southeast Asia’s diverse languages and cultures. The initiative aims to establish the nation as an inclusive global AI leader, ensuring the technology is both accessible and regionally relevant to its multilingual and multicultural populaces. The models are adept in numerous regional languages, including Bahasa Indonesia, Bahasa Malaysia, Thai, Vietnamese, and Tamil, unlocking AI technologies for a significant portion of the Asian and global population.

SEA-LION model v4-27B is now available on the Workers AI platform. SEA-LION v4 stands out on the Singapore government’s leaderboard as its most powerful, efficient, multimodal and multilingual model yet. 

You can use the @cf/aisingapore/gemma-sea-lion-v4-27b-it model on Workers AI as follows:

curl --request POST \
  --url https://api.cloudflare.com/client/v4/accounts/ACCOUNT_ID/ai/run/@cf/aisingapore/gemma-sea-lion-v4-27b-it \
  --header 'Authorization: Bearer TOKEN' \
  --header 'Content-Type: application/json' \
  --data '{
  "messages": [
    {
      "role": "user",
      "content": "แล้วทำผัดไทยอย่างไร"
    }
  ]
}'

Bringing AI models to the world

Singapore, India and Japan have all chosen to open-source many of their local language models, a strategy that champions an expansive vision of AI sovereignty. This approach demonstrates a crucial understanding: true AI sovereignty is ensuring you have choices. 

Supporting local language open source models is more than just supporting technology; this is a shared commitment to fostering an open, interoperable, and competitive AI ecosystem by empowering governments and developers to solve local problems, create economic opportunities, and preserve their digital and cultural heritages.

We are honored to support the initiatives of the governments of India, Japan, and Singapore on this journey. We believe that by putting their sovereign AI models into the hands of developers in their economies, we can help unlock a powerful wave of innovation that is more diverse, equitable, and representative of the world we live in. The future of AI is being built today, and we are proud to ensure that AI developers everywhere are at the forefront. 

Choice is the foundation of AI sovereignty. We’re starting with the models from India, Japan, and Singapore on our serverless inference platform, but it’s only the start. Come build with us! Take the first step for free on Workers AI.

Every Cloudflare feature, available to everyone

Post Syndicated from Dane Knecht original https://blog.cloudflare.com/enterprise-grade-features-for-all/

Over the next year Cloudflare will make nearly every feature we offer available to any customer who wants to buy and use it regardless of whether they are an enterprise account. No need to pick up a phone and talk to a sales team member. No requirement to find time with a solutions engineer in our team to turn on a feature. No contract necessary. We believe that if you want to use something we offer, you should just be able to buy it.

Today’s launch starts by bringing Single Sign-On (SSO) into our dashboard out of our enterprise plan and making it available to any user. That capability is the first of many. We will be sharing updates over the next few months as more and more features become available for purchase on any plan.

We are also making a commitment to ensuring that all future releases will follow this model. The goal is not to restrict new tools to the enterprise tier for some amount of time before making them widely available. We believe helping build a better Internet means making sure the best tools are available to anyone who needs them.

Enterprise grade for everyone

It’s not enough to build the best tools on the web. At Cloudflare our mission is to help build a better Internet and that means making the tools we build accessible. We believe the best way to make the Internet faster and more secure is to put powerful features into the hands of as many people as possible.

We first launched an Enterprise tier years ago when larger customers came to us looking to scale their usage of Cloudflare in new ways. They needed procurement options beyond a credit card, like invoices, custom contracts, and dedicated support. This offering was a necessary and important step to bring the benefits of our network and tools to large organizations with complex needs.

This created an unintended side effect in how we shipped products. Some of our most powerful and innovative features were launched within an enterprise-only tier. This created a gap, a two-tiered system where some of the most advanced features were reserved only for the largest companies.

It also created a divergence in our product development. Features built for our self-service customers had to be incredibly simple and intuitive from day-one. Features designated “enterprise-only” didn’t always face that same pressure to scale – we could instead rely on our solutions teams or partners to help set up and support.

It’s time to fix that. Starting today, we are doing away with the concept of “enterprise-only” features. Over the coming months and quarters, we will make many of our most advanced capabilities available to all of our customers.

The change will help build a more secure Internet by removing barriers to the adoption of the most advanced tools available. The change improves the experience for all customers. Smaller teams on our self-service plans will have access to the most powerful configuration options we offer. Existing enterprise teams will have easier pathways to adopt new tools without calling their account manager. And our own Product teams have even more reason to continue to make all features we ship easy to use.

Today we are beginning with dashboard SSO with instructions on how to begin setting that up right now below. It is the first of many though and capabilities like apex proxying and expanded upload limits, along with many others of our most requested enterprise features, will follow.

Starting with how you sign in to Cloudflare

One example of a feature we launched only to enterprise customers because of the complexity in setting it up is SSO. Enterprise teams maintain their own identity provider where they can manage internal employee accounts and how their team members log into different services.

They integrate these identity providers with the tools their employees need so that team members do not need to create and remember a username and password for each and every service. More importantly, the management of identity in a single place gives enterprises the ability to control authentication policies, onboard and offboard users, and hand out licenses for tools.

We first launched our own SSO support way back in 2018. In the last seven years we have been helping thousands of enterprise customers manually set this up, but we know that teams of all sizes rely on the security and convenience of an identity provider. As part of this announcement, the first enterprise feature we are making available to everyone is dashboard SSO.

The functionality is available immediately to anyone on any plan. To get started, follow the instructions here to integrate your identity provider with Cloudflare and to then connect your domain with your account. By setting up your identity provider for dashboard SSO you will also be able to begin using the vast majority of our Zero Trust security features, as well, which are available at no cost for up to 50 users.

We also know that some teams are too early or distributed to have a full-fledged identity provider but want the convenience and security of managing logins in one place. To that end, we are also excited to launch support for GitHub as a social login provider to the Cloudflare dashboard as part of today’s announcement.


And extending to almost everything else over the next year

We prioritized dashboard SSO because just about every team that uses Cloudflare wants it. This one change helps make nearly every customer safer by allowing them to centrally manage team access. As we burn down the list of previously enterprise-only features, we will continue targeting those that have similar broad impact.

Some capabilities, like Magic Transit, have less broad appeal. The organizations that maintain their own networks and want to deploy Magic Transit tend to already want to be enterprise customers for account management reasons. That said, we still can improve their experience by making tools like Magic Transit available to all plans because we will have to remove some of the friction in the setup that we have historically just solved with people hours from our solution engineers and partners.

We also realize that the way some of these features are priced only made sense with an invoice or enterprise license agreement model. To make this work, we need to revisit how some of our usage metering and billing functions. That will continue to be a priority for us, and we are excited about how this will push us to continue making our packaging and billing even simpler for all customers.

There are some features that we can’t make available to everyone because of non-technical reasons. For example, using our China Network has complicated legal requirements in China that are impossible for us to manage for millions of customers. 

Self-service by default going forward

One thing we are not announcing today is a strategy to continue to release “enterprise-only” features for a while before they eventually make it to the self-service plans. Going forward, to launch something at Cloudflare the team will need to make sure that any customer can buy it off the shelf without talking to someone.

We expect that requirement to improve how all products are built here, not just the more advanced capabilities. We also consider it mission-critical. We have a long history of making the kinds of tools that only the largest businesses could buy available to anyone, from universal SSL over a decade ago to newer features this week that were available for self-service plans immediately like per-customer bot detection IDs and security of data in transit between SaaS applications. We are excited to continue this tradition.

What’s next?

You can get started right now setting up dashboard SSO in your Cloudflare account using the documentation available here. We will continue to share updates as previously enterprise-only features are made available to any plan.