Tag Archives: open source

Introducing Clef: our open-source decision models, and new RL fine-tuning platform

Post Syndicated from Michelle Chen original https://blog.cloudflare.com/clef-decision-models/

Over the last few weeks, there has been lots of buzz around decision models such as Typesafe AI’s Jev System One model. While classifier models have been around for some time, Jev introduces a new decision model concept into the world of AI — a model that produces bounded structured outputs cheaply, quickly and consistently that can be added into a workflow when a decision is required. These models are capable enough to work over any set of inputs without constantly retraining the model to incorporate new classification categories. This contrasts with the world of Large Language Models (LLMs), which are largely non-deterministic, but are open-ended enough to reason and generate text and tool calls for agentic workloads. 

Today, we’re releasing two Cloudflare-trained decision models, Clef and Clef-flash, hosted on Workers AI. Clef is currently the leader when evaluated against the Jev Decision Index, you can view full results on the live benchmark demo site. These models are smarter, faster, and fully Jev-API compatible, so you can experiment with these hosted models easily. We’re fully open-sourcing these models on Hugging Face under an Apache 2.0 license for you to run locally and experiment with yourselves. 

Lastly, we’re excited to debut our new reinforcement learning (RL) product, which allows customers to fine-tune Clef to suit their use cases as well.

What is a decision model?

A decision model makes classifications to help agents decide how to act, based on certain probabilities. For example, you can pass in a customer support message (inputs) and ask if it is urgent and which team should handle it. A decision model will return typed answers with probabilities (outputs), which your code can use to route the ticket, trigger an escalation, or defer to a human. This means that a human does not necessarily need to be in the loop for agentic decisions anymore — agents can programmatically gather context, make decisions, and take actions on tasks, or defer to a human when needed.

Specifically at Cloudflare, we’ve been testing our new Clef model on our Threat Intelligence team to help us classify website domains. By giving a domain to Clef (with Browser Run) it can quickly identify categories that the domain falls under — for example, it might classify a domain with a 95% chance it is a fashion website, 85% ecommerce, <1% phishing, etc. This classification took our Clef model 2.2s to fetch, render, and classify the website. In contrast, our fastest general LLM gpt-oss-120b took 4.7s in the same workflow, and only returned two classifications. As a user, you can imagine how a 2x savings in latency and results can help us improve our threat intelligence workflows and be faster in identifying malicious or legitimate domains. Generalize this to any use case where you need to make quick programmatic decisions, and you unlock powerful agentic workflows that are able to autonomously decide, reason, and execute.

In music theory, a clef is a symbol placed at the beginning of a musical staff that assigns specific pitch names to the lines and spaces. A decision model is analogous to a music clef because it helps define the domain of the context and the subsequent notes (actions) that follow it. We chose Clef as the name of our family of decision models, as it serves similar purposes, and the CF hearkens to Cloudflare.

How is Clef different from other decision models?

Although the market is getting increasingly saturated with decision models, Clef has some unique properties that make us excited to release it to the public. First, it has a vision encoder so it’s able to take in images and classify visual content. This is different from Jev, which only does text classification today. Secondly, our model has a 64k context window (compared to Jev’s 32k), which allows users to squeeze more input state for the model to classify against.

Third, our model is accurate and powerful, scoring competitively against other decision models on the market across various quality benchmarks. We shortlisted some evaluations below that are important for decision-making as defined by the Jev Decision Index and scored some of the more popular models on the market for it. Check out the table below for benchmarks, or view the scores on our live decision index demo site:

Benchmark

Clef

Clef-flash

Jev

DiffusionGemma Jev

Kev 9B

Laya

BFCL · case exact

98.47

98.76

95.75

96.52

94.51

38.13

ToolRet · nDCG@10

69.19

66.43

65.28

61.21

64.26

12.69

API-Bank · accuracy

91.93

93.11

88.19

83.66

56.30

11.41

Home appliances · case exact

82.95

97.73

52.27

42.05

25.00

0.00

When2Call · accuracy

72.37

65.58

80.97

75.44

49.62

11.94

BANKING77 · macro-F1

94.20

90.93

79.74

74.28

84.83

14.29

CLINC150+OOS · macro-F1

97.43

66.77

89.27

83.49

79.03

3.19

BRIGHT · nDCG@10

45.91

39.26

47.52

42.94

38.53

19.90

Amazon ESCI · macro-F1

57.48

57.39

55.21

53.37

49.22

24.40

PhishNChips · accuracy

79.60

75.05

62.55

85.35

50.75

50.15

We also ran benchmarks across Typesafe’s own eval suite and our Clef models fared well, beating Jev in 3 out of 4 areas. Notably, our Clef-flash performs exceptionally well, given how much faster it is.

Workflow

Clef

Clef-flash

Jev

Invoice processing

64.7

57.1

61.8

Customer service

76.3

77

76.0

Security incidents

62.9

61.7

61.7

Agent trace observability

68.5

69.8

71.6

Across the 43 eval benchmarks that we ran, our Clef models beat the decision models on latency (except for Laya which is very fast but trades off quality in the benchmarks above):

Benchmark

Clef

Clef-flash

Jev

DiffusionGemma Jev

Kev-9B

Laya

Median latency  · ms

209.3

38.8

524.1

84.4

51.4

5.8

p95 latency · ms

238.6

122.4

536.0

211.2

187.9

222.5

On top of the latency benefits from the model itself, our Clef models are hosted on Workers AI. Because they are hosted on Cloudflare’s infrastructure, we’re able to take advantage of our GPUs at the edge, leading to low network latency and faster decisions. This means that you could put Clef into the hot path for agents to make decisions and combine that with one of our LLMs on Workers AI to take action. 

Clef also produces strictly typed outputs similar to Jev and is fully API-compatible, so you can make the swap extremely easily. The larger Clef model is your more powerful precision model, while the Clef-Flash model is great for latency-critical decisions. The models are enterprise-ready with our guarantee that we don’t read, store, or train on your requests or responses (unless you want to use our fine-tuning product, which we go into below). You can get started with the Clef models today, starting with our developer documentation or play around with the open-source model on the Hugging Face repo.

If you’d like help tuning Clef for a specific workload, we are also offering fine-tuning services — first as a hands-on partner with our forward-deployed engineer (FDE) team, and then later as a self-serve fine-tuning platform for customers to train and redeploy the model onto Cloudflare.

How we trained Clef

In the same week that Jev came out, we posted about some experiments we had with our own homegrown decision model. Our demo goes into how we adapted the DiffusionGemma model to output deterministic probabilities by exposing the logprobs that are generated by a large language model. Our initial approach built upon independent research by Matt Mastracci, who has been active in the machine learning (ML) community with sharing new ideas and pull requests to vLLM inference engine to make DiffusionGemma support stronger.

Clef builds upon this concept, but uses a different base model as the backbone. We currently use Qwen as the base model and post-trained it to suit decision model use cases. During inference, Clef uses Qwen for a prefill-only pass, then scores the valid schema choices in parallel. The decision step is non-autoregressive, so there’s no intermediate text to generate token by token, making Clef significantly faster than autoregressive LLMs. Rather than generating intermediate text to produce structured answers, Clef and Clef-flash derive schema choices directly from internal backbone representations. This approach relies on a specialized two-stage attention routing process: every valid choice extracts context relevant to the prompt, allowing individual field parameters to cross-attend with other fields and back to the original payload prior to scoring. By leveraging a lexical prior, the model preserves semantic intent across options. Ultimately, the architecture unites option-specific evidence routing, joint cross-field attention, and schema-bound scoring.

By freezing Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, we jointly optimized the routing head alongside rank-256 low-rank adapters. Our post-training utilizes label-smoothed cross-entropy for valid schema outputs paired with a Brier loss to refine probability calibration. This training leverages our own internal synthetic datasets permutating field orders, prompts, and schema structures. We also developed Reinforcement Learning for Calibrated Decisions (RLCD) to serve as a secondary optimization target, granting partial credit to adjacent ordinal choices, rewarding fully precise record outputs, and applying a reference penalty to prevent distribution shift, giving us better accuracy and generalization.

This means that we were able to achieve a few novel things with Clef: we improved accuracy of the model in classification, constrained it to output only probabilities instead of text generation, and made it faster than Jev and the base Qwen models.

How fine-tuning can extend the capabilities of Clef

We heard a lot of internal use cases that required fine-tuning our Clef model to be built into our agentic workflows at Cloudflare. For example, internal teams want a classifier model to be able to evaluate Trust & Safety submissions, help us triage Cloudflare Support requests, or even to be built-in to our Bot products to decide if a crawler is a good bot or bad bot.

These use cases are incredibly specific and we have had many years of labelled decisions that we could use to train a specific classifier. When you fine-tune a model, you may give up some general purpose performance in exchange for higher accuracy in a specific domain.. Because Cloudflare has more than 15 years of network data across different domains, we can fine-tune a model to fit these specific use cases which is more accurate and faster than our generic Clef model. We’re working with internal teams already to figure out how we can post-train Clef to create powerful ML models that boost our impact and improve workflows across Cloudflare. These internal teams and use cases are the next remit of our new FDE fine-tuning team and basis for our reinforcement learning (RL) product.

Our new RL service

We are offering a service to help customers fine-tune Clef to suit their workloads with our hands-on FDE team. From that, we’ll learn from our hands-on experiences to build a self-serve platform that customers can use to capture data, fine-tune, and redeploy the model, all on Cloudflare.

This has actually been a long time coming — we’ve been building our AI platform to have the right primitives where we could be building a custom RL product. The interest in Jev shows the need for a fast, small, specific, classifier model, and we chose this to be our niche to start experimenting with RL environments.

To do this, we leverage the primitives that we already have built on our Cloudflare platform:

  • Cloudflare AI Gateway – pass all your AI traffic through AI Gateway and automatically create a dataset of requests for your use case
  • Cloudflare Workers AI – generate rollouts against the base Clef model
  • Cloudflare Containers – RL sandbox for scoring and replaying agent actions
  • [NEW] Trainer – update weights of fine-tuned Clef model
  • Cloudflare Workers AI + BYO Model – redeploy the fine-tuned model on Workers AI

This combines a few work-in-progress pieces of the AI Platform that we’ve been working on, including AI Gateway that captures your AI traffic so you can leverage your own request/response data, Containers for RL Sandboxes, and Workers AI’s Bring Your Own Model (Cog) work that has been progressing since our acquisition of Replicate. 

Try it out today

We’re excited to launch our first Cloudflare-trained ML model from the Workers AI team today. We’re still early here and have a lot more improvements in store, but it is a wonderful first showcase of the hard work we’ve been doing on the AI Platform team. We believe that Clef has the ability to disrupt the way we use agents, which fits naturally into Cloudflare’s mission of being the agent cloud.

If you have specific use cases and are already customers of these products — we’d love to chat with you and be design partners as we experiment in this space.

Try out the Clef models hosted on Workers AI, download the weights on Hugging Face if you’d like to explore for yourself, and reach out if you have fine-tuning use cases you’d like us to help with.

Our ML team has been growing in impact, from model optimizations to model training research. If you’re interested in joining our mission, check out our open roles. 

One year later: Sovereign AI and the fight for choice

Post Syndicated from Carly Ramsey original https://blog.cloudflare.com/sovereign-ai-choice-one-year-later/

It's Birthday Week, when we traditionally ship presents to the Internet. This year, two of them come from Europe: EuroLLM, which covers all 24 official EU languages, and Apertus, Switzerland's fully open model, trained on more than 1,500 languages. Both were built by public universities and research institutions. Both are coming to Workers AI, and you can request access today.

We're also launching hands-on workshops that help government cyber agencies and critical infrastructure operators build AI defenses that work with any model. The first runs in Singapore in October.

Today's announcements follow from an argument we made a year ago, when questions about AI access and sovereignty were swirling in national capitals. Our answer was choice: the freedom to pick the right tools for the job, and to switch when you need to.

Since then, those conversations have hardened. Attackers have used frontier models to run cyber attacks. Access to some frontier models now depends on where you are. Calls to restrict open models are getting louder. Put it all together and it's easy to conclude that AI sovereignty is zero-sum: every model another country controls is one you can't count on, so the safe move is to build walls.

We think the past year can point the other way. India, Japan and Singapore focused on open-sourced models, and people across Asia-Pacific built tools on them for rural citizens, elderly patients and the nurses who care for them. Our own security team built AI defenses that work with any model, so losing access to one doesn't mean losing your defenses.

Helping build a better Internet has always meant more options, not fewer. That's why we work on open standards that prevent vendor lock-in, why so much of what we build is free to start with, and why our network runs in more than 335 cities across 125+ countries, with GPUs for AI inference in more than 230 of them. We don't think any country should have to depend on one company for its AI. That includes us.

Two European open models on Workers AI

In February, Matthew Prince told the India AI Impact Summit 2026 in New Delhi that decentralized, affordable access to AI is a matter of national resilience. The models from India, Japan and Singapore we'd added a few months earlier were our first proof. The summit series moves to Geneva in June 2027, with a mission of "prosperity and progress for all."

EuroLLM supports 35 languages, including all 24 official EU languages, many of which are underserved by existing open models. It was developed with support from Horizon Europe, the European Research Council and EuroHPC by a consortium that includes Instituto Superior Técnico, the University of Edinburgh, Instituto de Telecomunicações, Université Paris-Saclay, Unbabel, Sorbonne University, Naver Labs and the University of Amsterdam. It was trained on the MareNostrum 5 supercomputer and, according to the consortium, outperforms similar-sized models on EU multilingual benchmarks and machine translation.

You can request access to EuroLLM on Workers AI here.

Apertus (Latin for "open") is Switzerland's first large-scale, fully open, multilingual language model. It was trained on more than 15 trillion tokens across more than 1,500 languages, with 40% of training data in languages other than English. It was developed by ETH Zurich, EPFL and the Swiss National Supercomputing Centre (CSCS) as part of the Swiss AI Initiative: built by public institutions, for the public good. Its architecture, weights, training data and methods are all published. It was designed with Swiss and European rules such as the EU AI Act and GDPR in mind, which means respecting training opt-outs, removing personal data and preventing memorization. It was trained on CSCS's Alps supercomputer (more than 10,000 GH200 GPUs), and its developers report that it significantly outperforms leading closed and open models on rare and regional languages, from Romansh and Swiss German to low-resource languages across Asia and Africa.

You can request access to Apertus on Workers AI here.

What people built last year

Last year's national models didn't sit on a shelf. Since we added them to Workers AI, hundreds of students, startups, small businesses and public servants across Asia-Pacific have built on them, many at buildathons we ran with local partners. Three of them:

  • Government forms, no reading required – Form Mitra (India): Benefit forms written in dense English shut out many of the rural, low-literacy and visually impaired citizens they're meant for. Students at the Indian Institute of Technology Delhi built Form Mitra at a buildathon we ran with CyberPeace: a voice-guided assistant that walks people through the form in any of 22 Indian languages, using AI4Bharat's IndicTrans2 model.
  • Care in the patient's own dialect – MedBridge (Singapore): Many of the nurses caring for Singapore's elderly patients come from across Southeast Asia and don't speak the local dialects, so critical clinical information can get lost in translation. MedBridge guides patients through health conversations in 14 languages, including Hokkien and Cantonese.
  • The right public service, in one tap – Anshin Concierge (Japan): For elderly residents, people with disabilities and anyone less comfortable with digital tools, working out which public service to call in a moment of need can be overwhelming. Built as a hackathon prototype, Anshin Concierge lets people describe the problem in their own words ("my knees hurt", "a strange screen appeared on my phone") and connects them to the right Tokyo Metropolitan Government support desk, by phone or web page, in one tap. 

AI defenses that don't depend on one model

Governments want to use AI to defend essential services and national infrastructure. The frontier models that can find vulnerabilities at scale can find them for defenders too. But a defense built on one model is only as dependable as your access to that model, and governments have watched access to critical models get constrained with little warning.

We had the same problem, and in security we are always our own first customer. Over the past year, our Security team, working with teams across the company, set out to build AI defenses that don't depend on any one model. We built a harness: an orchestration layer that coordinates multiple AI models working in parallel to hunt for vulnerabilities, verify findings and prioritize threats. We published what we learned about using frontier and open models together, open-sourced the harness so any organization can run it with the models of its choice, and laid out the layered architecture we use to stop attackers armed with frontier models from finding vulnerabilities in the first place.

Because the harness works with any model, closed or open, losing access to one provider doesn't switch your defenses off. When we walked governments through it, the most common reaction was relief. Then came the practical questions: how to stand it up in their own environments, under their own rules. Briefings quickly turned into requests for hands-on training.

So today we're launching a program of hands-on workshops for government cybersecurity agencies and critical infrastructure operators. Participants build their own AI security harness and layered defenses, and leave knowing how to adapt both to their organization. The modules are plug-and-play, designed to slot into national AI skilling and cyber resilience programs. The first workshop runs in Singapore this October, at Singapore International Cyber Week.

Come build with us

None of this happened alone. Partners like CyberPeace in India and Code for Japan helped turn open models into working tools. If you run a national AI program, a cyber agency or critical infrastructure, and you'd like more options than you have today, write to us at [email protected]. Request access to EuroLLM and Apertus, or start with the harness.

A year ago, we said choice is the path to AI sovereignty. This year showed it's the path to AI security, too.

Cloudflare OS: your company’s agent workspace, managed for you

Post Syndicated from Phillip Jones original https://blog.cloudflare.com/managed-cloudflare-os/

Cloudflare OS gives everyone in your organization an agent workspace that knows how your company works and connects to its data and systems. Today, we're opening the waitlist for fully managed Cloudflare OS deployments.

If I asked you to prepare for an important customer meeting later today, what would you do? You might learn how your company typically runs customer meetings, review the account in your CRM, check recent support tickets and product usage, then turn it into a short presentation to review with the group. Now imagine doing that another 100 times this month.

Every team has work like this. With Cloudflare OS, you can ask your agent to handle the work for you, build a tool for your team, or move between the two as the work evolves.

Last month, we announced Cloudflare OS and shared the open source repository. Since then, thousands of organizations have started using it to work with company data, produce docs and slides, build tools for their teams, and automate work with agents.

With a few clicks in the Cloudflare dashboard, you’ll be able to launch your organization’s own agent workspace. Just tell us what custom domain you want to use, what Cloudflare Access policies apply, and which AI Gateway to connect. We’ll handle the rest.

Cloudflare OS, managed for you

Every company has its own terminology, procedures, systems, and requirements. We made Cloudflare OS open source so you can customize it around how your company works.

You can already deploy Cloudflare OS into your own Cloudflare account from the open-source repository. That gives you full control, but it also means someone has to configure the deployment, operate it, and keep it up to date.

With the fully managed option, you decide who can access Cloudflare OS, which organizational skills and context are available, and which systems it can reach. You can leave the rest to us.

If you want Cloudflare OS fully managed for your organization, join the waitlist and we’ll reach out.

What’s new in Cloudflare OS

We’ve also spent the last month expanding what people and agents can do in Cloudflare OS. Here are a few highlights.

Mount Git repos and work with code

When we launched Cloudflare OS, we focused first on work outside software development: creating documents and slides, automating tasks, and building collaborative tools. Agents could write code for an app, but they could not work with code in an existing Git repository.

You can now connect an existing GitHub repository to Cloudflare OS. Ask your agent to explore the codebase, fix a bug, add a feature, or open a pull request. It can search and edit files, review its changes, create commits, and push them to GitHub.

Work across Google Workspace

For many organizations, work starts and ends in Google Workspace. Decisions live in email threads, context lives in Google Drive, analysis happens in Sheets, and teams coordinate through Calendar. Agents need to do work across those systems too.

We’ve made significant improvements to the Google Workspace Gatekeeper (a service-specific Worker that sits between Cloudflare OS and an external service). Cloudflare OS can now read and research Gmail threads, create drafts, and send emails. You can also connect your entire Google Drive, a specific folder, or an individual doc or sheet.

Export work in the formats your team uses

Work often needs to move into the formats your team already uses. Finance may need an Excel spreadsheet, and a report may need to become a PDF before sending to a customer.

The built-in document, presentation, and spreadsheet experiences can now export work to familiar formats. Depending on what you create, you can export to Microsoft Excel (.xlsx), CSV, PDF, Markdown, or HTML. Microsoft Word (.docx) and PowerPoint (.pptx) export is coming soon.

Tools you build can also define their own export formats. Tell the agent what you need, like “let me download this schedule as a calendar file (.ics)”, and it’ll add the option to the tool’s export menu.

Sign up for the waitlist

Cloudflare OS is open source and available today. You can check out the source code or deploy it into your own Cloudflare account.

If you want Cloudflare OS fully managed for your organization, join the waitlist and we’ll reach out with more information.

Next.js applications, powered by Vite: introducing Vinext 1.0

Post Syndicated from James Anderson original https://blog.cloudflare.com/vinext-nextjs-on-vite/

When we launched Vinext in February, it was the result of an audacious week-long AI-driven experiment to see how far one engineer, and a stack of tokens, could get to replicating the NextJS framework backed by Vite.

In the seven months since that experiment, Vinext has grown into a framework that our customers trust and run in production for high-traffic, dynamic applications.

Today we are announcing the release of Vinext 1.0, the latest step on our journey to make it possible to deploy Next.js apps anywhere. Vinext lets you take any Next.js application, whether it was built for the Pages or App Router, and make it portable to be deployed to any web platform, including the Cloudflare Workers free plan, Netlify, or AWS Lambda.

Vinext 1.0 brings with it sweeping improvements to compatibility, stability, and caching behaviors, and sets the project up for the long term. There’s never been a better time to take your Next.js project and convert it to Vinext; just run npx vinext check and npx vinext init.

Graduation to 1.0

On release Vinext was promising, but it was incomplete. Since then, we’ve spent a lot of time both improving App Router compatibility and expanding that to Pages Router apps — which we’ve learned many customers are longtime fans of, with large applications that are complex to migrate. We didn’t want Vinext to be a tool that only worked for people using the latest App Router features.

Our focus has been on adopting both these routers, and watching our test compatibility closely, which for most important customer-requested features now surpasses 99%.

This improvement has been fueled through the community around our GitHub project. As soon as Vinext launched, that community threw it at a wide variety of applications to find the gaps. With their scrutiny, we found challenges not immediately obvious in the test coverage. Vinext needs to act exactly as Next.js behaves. It is not good enough to imitate functions with the same name. Building an alternative import { revalidatePath } is simple enough; the difficulty is in making sure it correctly affects the rendered pages, cache entry, and future requests.

Tracing requests through the application to make sure Vinext responds in the way expected — and replicating not just the API, but the behavior of this machine — was by far the more challenging aspect.

Once we’ve patched problems and brought new features forward, it’s important that we don’t regress, especially if Next.js makes a change. That’s why we’ve also built out our test suite: thousands of focused tests covering core framework behavior across both routers, the development and production server, and the deployment targets of Nodejs and Cloudflare Workers. We also run the Next.js end-to-end test suite against Vinext nightly, giving us a continually moving window on our compatibility, and making sure we immediately become aware of regressions coming from merged changes. Alongside the automated testing, we’ve been working directly with large customers that have Vinext in production to make sure they are not facing issues.

What’s in 1.0

The clearest messages we got from customers using Vinext is that certain Next.js features carry the framework and Vinext didn’t actually need to do everything that Next.js has launched in recent versions to be incredibly useful to them. So we focused on better support where you need it:

  • App Router, Pages Router, and Hybrid applications: We heard from customers that Pages Router was still important, and migrations are not a one-step process. Vinext therefore has support for both routing paths, including React Server Components, Server Actions, API routes, route handlers, middleware, and client-side navigation.
  • The complete page lifecycle: Pages can be rendered in many different ways: on the server, pre-rendered in the build, exported as static assets, or cached with page-level Incremental Static Regeneration (ISR). We’ve made sure that Background and on-demand revalidation work with any output.
  • Caching: Vinext has a shared set of caching functions across the App and Pages Router and the supported runtimes. We have further support for using Cloudflare’s Workers Cache.
  • Observability: Vinext provides Next.js-compatible tracing across both routers, so existing OpenTelemetry and Sentry setups continue to work. On Cloudflare Workers, traces also integrate with native Workers Observability.
  • Next.js ecosystem compatibility:  Vinext implements the public next/* surface and supports common Next patterns for use of authentication, MDX, image optimization, fonts, metadata, environment variables, and more.
  • First-class runtime support for Workers: While Vinext can run anywhere, server code can run in the Cloudflare workerd runtime during development and production, with direct access to bindings such as image optimization and hyperdrive. 

We’ve also made migration part of the framework: it takes two commands to verify that your Next.js install and any modification you have made is compatible, and set up the Vite and deployment configuration while keeping all your previous Next.js project structure.

When we talked to teams about what features were important for them, something stood out. Next.js 16 took a stance that Cache Components were an important part of the future of the framework, and yet most teams that we talked to were not using them and did not consider support a prerequisite to move. Therefore, Vinext today has limited support for the “use cache” directive that drives Cache Components, and though we will continue to improve compatibility there, we’re much more focused on the core priorities above.

Pre-rendering and cache warming

When we first announced Vinext, it supported Incremental Static Regeneration (ISR) after the first request, but it did not yet render pages during the build. Applications use generateStaticParams() and getStaticPaths() to identify pages that should be rendered when building, and they expect page-level ISR to connect those initial responses to background and on-demand revalidation.

Vinext 1.0 supports that lifecycle for both routers. It can prerender App Router and Pages Router routes during the build, serve those responses through page-level ISR, and invalidate them by path or tag. It also supports output: "export" when the result you want is a fully static site.

But this led us to question something: Why should this rendering happen during the build at all?

A site with tens or hundreds of thousands of possible URLs can spend a seriously long time rendering pages that receive little traffic. The build process cannot evaluate the long tail of traffic that most sites experience and therefore cannot focus compute time on the smaller number of more critical pages. Instead you waste hours of time waiting for sequential builds working their way through thousands of pages, long after the most important routes are done.

Cache warming is our solution to this, moving page prerendering from the build machine to Cloudflare’s network. Developers can continue to use Next.js primitives to identify the pages for prerendering, and Vinext can additionally identify high-traffic pages to add to this list. This happens in the background before your site is deployed to production, so that the moment it is, it is ready to serve rapid responses from the Cloudflare cache.

Inside the deployment process, this works by uploading a new Worker version and deploying it to 0% of production traffic, before then requesting pages specifically from that version. This allows the rendering pipeline to work before any real users hit the new deployment. Once the caches have been populated, the deployment can be promoted safely.

What we’re doing next

If the original experiment invented the one-off slopfork, the more consequential part has been how we can keep that process of self-improvement running indefinitely.

The project is now focused on keeping the framework up to date with everything happening upstream. Next.js canary receives new commits every day. Each morning, an agent reviews the changes, fetches diffs, and opens tracking issues for anything that could affect Vinext. Every night, the compatibility matrix is regenerated as we run the Next.js test suite against Vinext.

When one of these tests or issues reveals a gap, agents are now in the position where they can identify the change across both codebases, build a reproduction, port any relevant tests, and propose a fix.

This review has been catching missing cases, unsafe caching behaviors, and differences in the development vs. production servers.

Automation has helped us narrow the stream of activity into a focused set of changes that deserve attention, allowing the maintainers of the project to focus on only the issues that need context of how a process should map onto Vite from the Next.js implementation.

We’re building a software factory for open source at Cloudflare, and you can see what we’re up to on GitHub.

Try it out

Vinext is available for new applications, and existing Next.js projects.

Start a new application today:

Or migrate an existing application:

And then deploy it to Cloudflare Workers, with our cache warming:

Visit vinext.dev for documentation, examples, and the current compatibility matrix. Vinext is open source at github.com/cloudflare/vinext. Issues, pull requests, application reproductions, and feedback are welcome.

Supporting native Rust in Workers with the new Emscripten target for wasm-bindgen

Post Syndicated from Guy Bedford original https://blog.cloudflare.com/rust-workers-emscripten-target/

Today we’re announcing the first public experimental preview of a feature to better support native Rust code and even Tokio-based applications just running natively on Workers: first-class support for the Emscripten wasm32-unknown-emscripten Rust compiler target on the wasm-bindgen open source toolchain and Cloudflare’s Rust Workers.

wasm-bindgen is the open source toolchain powering Rust-based WebAssembly applications on our V8-based Workers Runtime. Enabling the Emscripten target for wasm-bindgen has been a long-term effort, first initiated by Google over a year ago, and then further reviewed and supported by the Cloudflare engineers maintaining wasm-bindgen.

While still in pre-release, we’re excited to share the new workflow possibilities this work enables in running native wasm-bindgen Rust applications with Emscripten on the web, Node.js, and on Cloudflare’s global Workers platform.

In testing we’ve been able to see significantly improved library compatibility for Rust Workers. To illustrate the sort of capabilities supported, we were able to get a Rust-native Minecraft server (Pumpkin) running inside of a Durable Object with TCP ingress, using real TCP sockets via Tokio. See the end of this post for a full description of this port.

Emscripten is an open-source WebAssembly compiler toolchain initially created by Mozilla and currently maintained by Google engineers, which makes it possible to run and bridge native code with the web platform, including supporting and virtualizing platform features such as timers, file system operations, sockets, and other native functionality.

Since Cloudflare Workers supports Web Platform APIs and Node.js compatibility, we are able to support Emscripten on Workers using its Node.js compilation flags, fully virtualizing native platform features such as timers, file system operations, and sockets on top of our existing Node.js APIs. And with our work on Tokio support, Rust code building on top of Tokio’s async runtime ecosystem can now also be fully integrated into JavaScript-based host environments with this Emscripten target.

We’ve made these current experimental patchsets available with example applications to try out today, including:

Supporting the wasm32-unknown-emscripten target in wasm-bindgen

Mitch Foley on Google’s Portable Toolchains team first encountered the need for Rust WebAssembly toolchains to interoperate with C++ when another internal team at Google was interested in using wasm-bindgen to interface with their JavaScript. The goal was for wasm-bindgen to drive the build and produce the companion JavaScript, while Emscripten’s linker (wasm-ld) would be able to link in any C++ dependencies needed.

Internally, Google uses Emscripten in C++ codebases to generate JavaScript along with the rest of their applications, allowing the C++ and JS to interact with one another. While Emscripten was capable of linking in Rust code as a dependency, the thing it couldn’t do was provide a robust binding system between Rust and JavaScript like wasm-bindgen does.

The problem was both tools assume they are in charge of loading and interacting with JavaScript and generating the final JS and Wasm output. Because of this conflict, Google internal users would have to commit to using one or the other toolchain and never both, and adding a second set of tooling would have doubled the support surface for Google’s Portable Toolchains team.

Mitch and his colleague Yifan Yang crafted a plan to have them work cooperatively: Emscripten would continue to drive the build, load the Wasm module, and provide the companion JS, while wasm-bindgen would produce a smaller portable version of its JavaScript bindings in a format that could be directly included in Emscripten’s library system.

This wasn’t just a technical problem — both Emscripten and the wasm-bindgen maintainers had to support this plan and be willing to maintain integration tests that depended on the other. With feedback from Google’s Portable Toolchains team, Google’s Wasm Tools team, and Cloudflare engineers, this finally resulted in the required changes landing in both projects.

As a result of these efforts, the wasm-bindgen and Emscripten toolchains now have seamless interoperability under the new -sWASM_BINDGEN configuration:

  1. C++ Emscripten code driven by Emscripten’s compiler can be built against static wasm-bindgen Rust code, fully supporting the wasm-bindgen bindings layer alongside the Emscripten bindings layer.
  2. Rust applications using wasm-bindgen and driven by Rust’s compiler can now be built for the Emscripten target, fully supporting Emscripten’s bindings layer alongside wasm-bindgen’s bindings layer.

See the wasm-bindgen Emscripten documentation page for more information about using this target.

Rust library support for Emscripten

After landing the wasm32-unknown-emscripten target support in wasm-bindgen, early prototypes by Cloudflare engineers demonstrated clear success for supporting this target on Cloudflare Workers. Many libraries worked out of the box, even including low-level systems libraries, since Emscripten already supports the target_family = unix in Rust.

Some low-level systems libraries that were unaware of Emscripten required patching, for example libc, socket2, and Mio. Even for these low-level libraries, patches primarily involved adding the Emscripten target to the existing platform gates, for example to explicitly allow target_os = “emscripten”, in place of Wasm platform gates.

Overall we posted these target support patches fairly infrequently, and they were mostly trivial when needed, with the Rust library maintainers very receptive to reviewing the support.

That said, one of the critical foundational libraries that needed significant support was Tokio.

Tokio support

Cloudflare Workers are single-threaded and hosted within a JS event loop, while Tokio async is designed around blocking operations being supported through threaded parking semantics. The two models are clearly not compatible with each other — a blocking operation such as a pending socket read or epoll wait cannot block the shared JS event loop.

To get around this, we had two options: WebAssembly JavaScript Promise Integration or to modify Tokio to support event loop runtime integration.

To support Tokio building for Emscripten, we have contributed full Tokio support patchsets that are being reviewed upstream, with the first target support patch for wasm32-unknown-emscripten already landed upstream in Tokio. With these patchsets, we’ve been able to fully support both approaches in Cloudflare Workers. In our Rust Workers Tokio examples, these patches are required directly pending further integration upstream.

Supporting JSPI

WebAssembly JavaScript Promise Integration (JSPI) maps directly onto Tokio’s existing parking semantics. This is because JSPI allows a blocking Wasm call to suspend the WebAssembly stack on a synchronous operation and return control to the JS event loop, exactly as one would expect of a park.

But when a new Wasm call is made while an earlier one is suspended, JSPI doesn’t break: it instead supports having a new WebAssembly stack being entered while arbitrary existing stacks are suspended at the same time.

The problem here, though, is that Rust itself isn’t aware that its stack is being switched out from under it. Tokio’s runtime context is tracked via a thread local, but a JSPI stack switch is not a thread switch, so the suspended and new stack still share the same thread-local runtime context. As a result, the Tokio runtime context still thinks it is in the parked context when a new Wasm call enters, and then panics because the runtime is already entered.

Supporting fully reentrant JSPI therefore requires careful thread local handling to split context between JSPI context switches. To support this, the thread-local context itself must be swapped on each JSPI enter, exit, suspend, and resume. In effect this is cooperative time-multiplexed threading, with each suspended stack carrying its own runtime context.

With our pre-release Tokio patchset, we were able to implement and verify this model. Finalizing the design and upstreaming it is ongoing in collaboration with the Tokio and Emscripten teams.

Adding an event loop runtime to Tokio

The other runtime approach is full event loop integration. Event loops are of course a common paradigm in native UI applications, so the idea of supporting an event loop runtime for Tokio was certainly not something completely unfamiliar to the maintainers in discussions we had around the Emscripten target support.

The question was rather how to design an event loop for Tokio in such a way that it could work across native applications (including Windows and macOS), as well as for WebAssembly embeddings in JavaScript hosts.

If we could design a general LocalEventLoop runtime for Tokio, we could solve this problem more generally.

A Tokio runtime does two things in a loop: (1) it polls tasks until nothing is ready, and then (2) it waits. The wait is what makes it a runtime rather than a library: the thread parks inside the I/O driver until a socket becomes readable, a timer expires, or another thread wakes it.

But when the host already has its own event loop, Tokio's loop cannot run inside it without blocking the host's. The way around this is to split Tokio's loop in two with both parts able to integrate with the host: (1) becomes an explicit drive() operation that runs one batch of ready tasks and returns, and (2) is replaced by a wake, so that instead of parking, the runtime tells the host it has work and the host calls drive() when it is ready to.

Our proposed LocalEventLoop design for Tokio is a LocalRuntime whose wait has been replaced by a wake. It is built with a standard std::task::Waker that the host owns, which instead of being used to signal that a future should be polled soon, is used by the Tokio runtime to signal that the runtime itself should be driven soon.

Consider the example of a socket read under a regular Tokio runtime:

This would correspond to the following call diagram:

Instead, with LocalEventLoop we can write:

Here, spawn_local returns immediately, with the host event loop taking responsibility for driving the spawned task to completion via el.drive() calls from the host. When stream data is unavailable, the runtime simply returns, handing back control to the host. Once the socket becomes readable, Tokio uses the host_waker to signal that the runtime needs driving.

This LocalEventLoop flow corresponds to the following call diagram:

Everything that would have unparked a native runtime's thread wakes the host instead: a spawn, a task woken from another thread, a socket becoming readable, a timer expiring. Because a Waker is Send + Sync and carries no execution semantics, its implementation is just a notification to the host's event loop. This makes the contract safe by construction: a wake arriving from another thread, from a host callback, or even during a drive, queues work rather than re-entering the runtime. The drive that follows runs on the owning thread with nothing else on the stack.

The one thing you cannot do is wait. block_on still exists and runs a future as far as ready work carries it, but where a normal runtime would park, LocalEventLoop::block_on panics instead. This is because nothing could wake that future from inside the call since its wait belongs to the host.

With this design, the host event loop is never blocked, and interleaves its own work with Tokio's, one batch at a time. The same architecture embeds into a GTK main loop, a Win32 message pump, or a Cocoa run loop. In addition, any number of these runtimes can co-exist due to their cooperative execution semantics.

Supporting sockets and epoll on Emscripten

With both Tokio runtime integration approaches fleshed out for the Rust Emscripten target, we were then able to support most of the Tokio test suite, with one major subsystem still unsupported — the net feature. This includes its sockets APIs: TCP, UDP, and Unix sockets. The reason for this was that Emscripten only supported poll() and a WebSocket emulation layer, but not epoll_wait(), on which Tokio’s I/O driver is built (via mio).

For Cloudflare Workers, we wanted to be able to fully integrate with our TCP sockets APIs, including upcoming inbound TCP. To do this, we would need to build a bridge between Emscripten’s virtualization layer and our own sockets API layer.

Instead of having to build this bridge ourselves, we realized we already have one: the node:net API we support in our Node.js compatibility layer. Emscripten had an -sNODERAWFS mode to bridge natively into Node.js FS APIs, so the same approach could give us a sockets bridge without Emscripten or Cloudflare Workers needing to implement any custom APIs on either side.

We contributed this work in over 40 pull requests to Emscripten, which is now the –sNODERAWSOCKETS layer compilation option, enabling support for epoll, TCP, UDP, and Unix sockets for Emscripten applications in Node.js — and, because Workers implements the same node:net API, on Workers itself as well.

Under JSPI, Emscripten's epoll_wait() simply suspends the stack until readiness, so Tokio's I/O driver works as on native. For LocalEventLoop, readiness instead had to reach the Waker from a JS callback. To support this, we drafted an Emscripten proposal for a new emscripten_epoll_add_listener API to associate a callback on an epoll’s ready events. The next drive() then collects these events with a zero-timeout epoll_wait(), so Tokio's existing I/O driver can be used unchanged.

Running a Minecraft server on Workers

To test out this new target, the challenge was raised to see if a Minecraft server could be deployed to Workers. Dan Lapid then implemented a Pumpkin Minecraft server running on a Durable Object over a single weekend.

Pumpkin is a Rust Minecraft server built on Tokio and designed for multi-core machines. World generation runs on a dedicated thread pool, while the game tick loop and chunk scheduler each run on their own OS threads.

Inside a Durable Object there is exactly one thread, so getting Pumpkin to run there meant turning those threads into cooperative tasks on the event loop. Leaning on the Tokio integration, the tick loop and chunk scheduler became async tasks and each Rayon job became a Tokio task. World generation then runs on the event loop itself, taking one turn per chunk and interleaving with network I/O and game ticks rather than blocking them.

For persistence, Pumpkin writes its world through ordinary std::fs. Under Emscripten’s -sNODERAWFS option file system calls get forwarded to the node:fs Node.js compatibility layer for Workers. Since this bridge is just JavaScript, it is easy to swap out. Dan’s worker-fs-mount library enables mounting a node:fs-compatible file system with its durable-object-fs backend, which stores files as rows in the Durable Object's SQLite storage. With this integrated, every file Pumpkin saves becomes a row in the object's database, written synchronously and committed with the Durable Object's transaction. Pumpkin itself has no idea it isn't on disk, and a restarted object boots straight from the same world.

Networking needed no changes in Pumpkin either. Each player's connection arrives through Workers TCP ingress at the Worker's connect() handler, which forwards it into the Durable Object. There, handleAsNodeConnection() from cloudflare:node dispatches the socket to a net.Server listening on that port inside the object – a new TCP counterpart of handleAsNodeRequest(). Emscripten’s -sNODERAWSOCKETS backend implements TcpListener on top of net.Server, so the server accepts players exactly as it would on Linux, and the returned promise tells the object when the last player has left, so it can save and shut down.

Running a fully persistent Minecraft server inside a Durable Object with multiplayer support demonstrates the level of native compatibility that is possible with this new Emscripten target. We’re excited to see what native Rust applications you can bring to the platform.

Try it out

We’re making these full patchsets and workflows available today for experimental use. Improving wasm-bindgen’s support for modern WebAssembly standards for all users is part of our ongoing commitment to the Rust and Wasm ecosystem.

We welcome all contributions and feedback — find us on GitHub and in the #rust-on-workers channel on Cloudflare’s Discord.

EmDash 1.0: the stable CMS with a secure plugin registry

Post Syndicated from Scott Buscemi original https://blog.cloudflare.com/emdash-cms-plugin-registry/

When we introduced EmDash on April 1 as the “spiritual successor to WordPress”, the buzz was hard to ignore. Walking around a WordPress conference that month, we couldn’t walk far without hearing murmurs about EmDash from fellow attendees.

But alongside the excitement and curiosity has been a seed of doubt among some in the industry. Was this just an April Fools’ joke?

It was not. Today, we are releasing EmDash 1.0: a stable, free, and open source CMS built on Astro, ready to power a production website, your agency’s vibe-coding platform, or your hosting company’s site-building experience.

Developers build with Astro, editors manage content through the EmDash admin, and agents can work through the API, CLI, or built-in MCP server. EmDash 1.0 brings those pieces together with production-tested editorial, media, localization, migration, and deployment workflows.

We are also launching a decentralized plugin registry that lets developers publish without handing ownership of their identity or releases to a central marketplace, while site owners can discover and install their plugins directly from inside EmDash.

The road to 1.0

Since EmDash's first beta, developers have launched real websites with it. Even so, we kept hearing a reasonable response: “This looks interesting. Let me know when it is 1.0.” Before relying on EmDash for their sites, they wanted confidence that it was stable, secure, that upgrades would protect their content, and that we are fully committed to maintaining it.

EmDash 1.0 is our answer to that. For the past five months, we have worked with contributors and production users on the parts of the CMS that every site depends upon: data safety, database migrations, editorial workflows, localization, plugin security, performance, and the reliability of the admin, API, MCP, and media experiences.

Real deployments shaped much of that work, uncovering edge cases and identifying needs that only show up when a site is serving real traffic and being used by real teams of editors.

As one example of this, Avulux converted a custom microsite to EmDash after dealing with the maintenance burden of WordPress for too long. Since EmDash is an agent’s best friend, using the EmDash Agent Skills helped that transition take less than a day to complete.

“The site had to be fast to use and simple for our team to update,” says Greg Barbosa, Director of Innovation and Systems at Avulux. “WordPress had become the opposite of that. With EmDash, we now have a shared platform that developers can extend and marketers can edit content."

In August, we migrated the Cloudflare Blog to EmDash as part of our “Customer Zero” approach. Working alongside our content engineering team provided actionable insights around localization, media management, the admin editor experience, and scaling. Comfortably handling the traffic load for the Cloudflare Blog meant being ready for millions of pageviews per week, spikes up to 5,000 requests per second (RPS) of legitimate traffic, or sporadic DDoS attacks. The optional KV object caching, Hyperdrive database adapter, and Workers Cache compatibility were all features spawned from our migration project that are widely available to all customers now.

Built in public, open to everyone

A CMS sits at the heart of an organization’s web presence. It is trusted with its most important data, and is often used by dozens of editors every day. They need to be able to know they can rely on it, without fear of vendor lock-in or changing business priorities. For that reason, EmDash is completely free and open source, using the flexible and permissive MIT license.

EmDash 1.0 could not exist without its open-source development community. At the time of writing, more than 175 people have contributed to the project, across more than 1,800 commits. The rise of agentic coding tools has presented both challenges and opportunities to open-source projects, and we have deliberately built a project where the agents can help the human developers, rather than being overwhelmed by them.

We are particularly grateful to the core group of the most dedicated contributors, who between them have shipped hundreds of improvements to all areas of the project. They include @swissky, @danielmlr, @MA2153, @marcusbellamyshaw-cell, and dozens of others. Contributors have translated EmDash into 25 languages, from Arabic to Ukrainian.

Particular recognition is due to Noah Pham, who joined Cloudflare as an intern and became EmDash’s second maintainer alongside Matt. Noah contributed more than 80 changes, taking ownership of major parts of the media library, content editor, and admin interface. We have said that interns ship meaningful work at Cloudflare; Noah’s work is now at the heart of EmDash 1.0.

There is plenty more to build, and contributing does not have to mean writing code. If you want to help with code, translations, documentation, testing, design, issue triage, answering questions, or just welcoming new users, join over 800 others in the EmDash community on Discord.

A growing ecosystem

A content management system thrives when the ecosystem around it is healthy and supported. We’ve been excited by the theme companies, plugin shops, agencies, and platforms who are creating new services and products using EmDash.

  • Lexington Themes offers 44 Astro themes with EmDash variants, giving teams a beautiful and functional starting point for their site, complete with reusable components and built-in content collections.
  • Urumi is using their WooCommerce expertise to release EmDash’s first eCommerce plugin.
  • Empress is launching a platform for multi-brand entities, offering the flexibility of managing a fleet of sites with natural language or a conventional CMS admin panel.

“Empress's delightful multisite experience would not be possible without the foundations EmDash has laid: sites that are fast, safe to extend, and easy for people and agents to read and act on,” says Raj Makker, Empress’s founder. “EmDash unlocks powerful control over a website, and Empress builds on it to give you complete control over as many sites as you want. We're excited to be part of this journey.”

A plugin registry that does not own the ecosystem

With EmDash 1.0, developers can publish sandboxed plugins and site owners can discover, inspect, and install them from the EmDash plugin registry.

Traditional plugin registries usually combine three roles: they provide the publisher’s account, hold the authoritative package record, and operate the catalog where users discover it. That is convenient, but it also makes one company the gatekeeper for both identity and distribution. If an account is suspended, a listing is removed, the rules change, or the service shuts down, publishers cannot take the same identity and release history somewhere else.

EmDash separates the plugin from the catalog. Publishers retain control of their packages and release history, while EmDash provides a convenient place for people to find and install them. Other services can index the same publications, build their own catalogs, and apply their own policies without requiring developers to start again.

The EmDash catalog applies default content moderation to the names, descriptions, links, and images it displays. Moderation can hide harmful or inappropriate material from this catalog, but it does not rewrite a release, take ownership of the plugin, or erase the underlying publication.

The registry is built on AT Protocol (atproto), a decentralized network protocol that powers Bluesky and a growing ecosystem of applications. Plugin authors publish with an Atmosphere account, the portable identity also used across these applications. The package and release records are signed by the publisher and stored in the publisher's own account.

EmDash hosts the default registry services so publishers and site owners can use them without running infrastructure themselves. We hope others will build new catalogs, moderation systems, publishing tools, and other services we have not imagined yet. To help this, we have released all of our services as open-source software, including the aggregator that powers the plugin registry, the labeler service that uses Workers AI to moderate package descriptions, and an Astro live content loader to make it easy to include plugin listings in any Astro site. We are excited to see what the community will build with them!

There’s no centralized controller that can take over plugins or remove them on a whim. We believe the security of plugins and the registry should be baked into code, rather than trusting any central authority or code of conduct.

Decentralized publishing does not mean accepting unverified code. Atproto repositories use signed Merkle Search Trees, so an inclusion proof connects the exact release record to a signed commit from the publisher’s account. EmDash can therefore verify the record independently instead of trusting the catalog’s copy. It then checks the plugin’s checksum, package name, version, requested access, and any required build provenance, and confirms that the downloaded bundle matches the signed record.

The registry supports free plugins today, and we aim to add support for paid plugins in future, with the same decentralized model as now. We are watching the Atproto Spaces Alpha with interest. Anybody should be able to run a secure, paid plugin marketplace. We want publishers to be able to make money from their software.

For now, plugin authors can publish useful software without asking permission from EmDash — and site owners can understand exactly what that software is allowed to do before they run it.

Plugins with clear boundaries

Plugins are the core of a CMS ecosystem. They save site developers from rebuilding the same integrations and workflows, while giving experts a way to share their best practices — and potentially build a software business.

Agents make that reuse even more valuable. An agent can build a one-off integration, but it still has to understand the problem, generate and test the code, and maintain it afterward. A plugin captures that work. Another agent can install and configure a proven solution instead of starting again from scratch.

That convenience creates a serious security concern: plugins are code you did not write, operating alongside valuable content and customer data. With WordPress, plugins run inside the same PHP process as the rest of the application, with direct access to its database, filesystem, and network. A contact-form plugin can technically read unpublished posts, modify another plugin, or send data anywhere. Site owners must trust that it will not — and that a future update will not change its behavior.

Sandboxed EmDash plugins use a different model. Each plugin runs in an isolated runtime with access to its own private storage, but not to the site’s content, media, users, secrets, environment, filesystem, or network. It gains additional abilities only when they are declared by the plugin and approved by the site administrator.

Installing a sandboxed plugin therefore feels more like installing a mobile app than a traditional CMS plugin. EmDash shows what the plugin wants to do before it runs, and the runtime limits it to those approved abilities.

That lets plugins perform useful work without receiving unrelated access:

A plugin could…

It might need to…

It still cannot…

Index published articles for search

Read content and contact the search service

Edit articles or contact other hosts

Optimize uploaded images

Read and manage media

Read user content

Send publishing notifications

Observe publishing and send email

Change the content being published

Provide configurable webhooks

Read selected events and contact public destinations

Reach private networks or other plugins’ storage

The important part is that these abilities are independent. Giving a plugin access to media does not also expose users or unpublished content. Allowing it to contact one service does not open the rest of the network. These boundaries are enforced by the runtime, not left to the plugin author’s good intentions.

This isolation is not limited to Cloudflare deployments. On Cloudflare, EmDash runs each plugin as a Dynamic Worker through the Worker Loader. On Node.js, EmDash starts workerd — the open-source Workers runtime — as a separate process and runs each plugin as an isolated service inside it. Plugins use the same manifests and capability-gated APIs on either platform. See the plugin sandbox documentation for setup and runtime differences.

From a single site to a website platform

EmDash can power an individual Astro site, but it is also designed for companies building website creation and hosting products.

Workers for Platforms lets those companies run each customer’s site as a Worker on Cloudflare’s global network. They do not need to provision a server for every site or build the surrounding networking and deployment infrastructure themselves. That leaves them free to focus on the experience their customers use to create and manage a website.

EmDash provides the content layer for that experience. A platform can use the admin interface directly, build its own interface on the API and CLI, or put an agent in front of the built-in MCP server. A bakery owner could update opening hours by asking for the change in plain language; the platform’s agent would handle reading, updating, and saving the content.

The sandboxed plugin model also gives platforms a safer way to offer extensions across many customer sites. Plugins receive only approved access to content, media, users, email, or external services, rather than running with unrestricted access to the whole application. Platforms can run their own plugin marketplaces with access to the full registry — or curate a selection of pre-approved plugins.

We are always looking for additional hosting partners that want to join us in developing new agent-oriented CMS experiences. Reach out to us if you’d like to learn more about reinventing with EmDash.

EmDash Build: an open source AI site builder

We are also releasing and open-sourcing an alpha of EmDash Build, an AI site builder that hosting providers, website builders, and platforms can run themselves and integrate with their own systems. Try the demo today at build.emdashcms.com, or explore the code.

If you’re building a site today, for yourself or for a client, you won’t start in an IDE. You’re more likely going to start with a chat box, and describe to an agent what you want. But once you have something, you could find that changing simple things requires you to go back to that prompt box, and either roll the dice on the result, or burn credits for a one-line change.

EmDash Build creates an EmDash site instead, which brings the full stack: server-rendered Astro pages, a database, media storage, and an admin interface. The agent designs a content model from the brief, fills it in through EmDash's MCP server, and writes the pages that display it. After that, you or your customer can edit text right on the page, schedule posts, or let an agent do it for you.

In EmDash Build, each project gets its own Cloudflare Sandbox container, where the agent, built on the Agents SDK, verifies its own work. Artifacts tracks every change as a git commit. When the site is published, the content moves into a production EmDash site, and deploys to the host's Workers for Platforms namespace.

Get started and get involved

With the release of EmDash 1.0, now is a great time to migrate your company’s marketing site or have an agent spin up that side project you’ve been talking about. Try out the EmDash playground site here. 

To create a new EmDash site locally, via the CLI, run:

Or you can do the same via the Cloudflare dashboard below:

If you’re ready to develop an EmDash plugin, our documentation has a step-by-step guide for creating and publishing a plugin to the registry.

We also welcome you to join our growing community of contributors on Discord. You don’t have to be an engineer to get involved — we welcome translators, issue triage managers, user experience designers, marketers, and all others who are excited about the future of content management systems.

Four months of VoidZero at Cloudflare: making the open-source JavaScript toolchain faster for all humans and agents

Post Syndicated from Evan You original https://blog.cloudflare.com/voidzero-update/

When VoidZero joined Cloudflare four months ago, we made a commitment to open source, promising that Vite, Vitest, Rolldown, Oxc, and Vite+ will stay open source, vendor-agnostic, and community-driven.  As part of Cloudflare’s Birthday Week, where Cloudflare gives back to the Internet, we thought it’d be a good time to check in on how we’ve been doing against this commitment.

In the four months since VoidZero joined Cloudflare, we’ve shipped more than 80 releases, closed over 1,200 issues, and landed some big performance improvements, including:

On top of this, Vite+, which unifies the entire toolchain with a set of great defaults, is now 1.0. And we’re making progress towards shipping “Bundled Dev” (f.k.a. Full Bundle Mode) — a new development mode that has been shaped by working with customers with massive web applications, including Cloudflare’s own dashboard. By being part of Cloudflare, our engineering team gets a much better view into scenarios that only occur in massive codebases.

VoidZero’s mission is to make the next generation of JavaScript developers more productive than ever before. And in 2026, that means making developers’ agents faster.

A faster developer experience for humans and agents

Developer experience and performance were always about improving the feedback loop while creating software. But now when we optimize for developer experience, we no longer just optimize for humans — we optimize for agents. And as inference gets faster, the speed of type-checking, linting, or building the code becomes the bottleneck again. The longer those processes take, the longer an agent has to wait before making progress and completing its goals.

VoidZero has always been focused on performance, but this new era of software has given us even more reason and motivation to make the entire toolchain faster for both agents and humans. And since the tools in the VoidZero toolchain are built on each other, from the compiler (Oxc) to the bundler (Rolldown) to the build tool (Vite), the linter (Oxlint) and the test runner (Vitest), any optimization at one layer automatically benefits everything built on top.

Oxc React Compiler — 10x faster compile times for React.js apps

We’ve recently shipped the Oxc React Compiler, a rewrite of the React Compiler, based on the React team’s rewrite in Rust. It is 10x faster than the original Babel implementation, uses less memory, and has a more complete implementation with better error handling. If you use Vite, you can enable the Oxc React Compiler by installing the oxc-transform-react package and enabling the compiler flag:

Vitest 5 — up to 50% faster than Vitest 4

Vitest 5, released in September, cuts test times by double-digit percentages across many common scenarios.

  • Faster test runs. Vitest shares transformed files across projects, caches modules on disk, and sends less data between its main process and workers.
  • vitest doctor. It breaks down setup, import, transform, and test time, then tests other configurations and recommends faster settings.
  • Trace View. It records browser interactions, assertions, and DOM snapshots, so you can replay failures or inspect them in an HTML report.
  • Conditional mocks with vi.when. Map arguments to responses without writing a manual mockImplementation.
  • Better benchmarks. Benchmarks now work like regular tests, with fixtures, hooks, retries, filters, and assertions.
  • Fewer false passes. Vitest fails unawaited async assertions, clears mock calls before each test, and adds the --repeats flag to help find intermittent failures.

tsgolint — now stable, up to 18x faster than ESLint in large codebases

tsgolint, the type-aware linting engine behind Oxlint, is now stable. It catches bugs that require TypeScript type data while running 12 to 18 times faster than ESLint with typescript-eslint.

Enable type-aware linting and TypeScript diagnostics in your Oxlint config:

Oxfmt — 7x faster than Prettier with formatters now written in Rust

Oxfmt brings the speed of the Oxc toolchain to formatting. We rewrote its JSON, CSS, SCSS, Less, GraphQL, and YAML formatters in Rust, making Oxfmt many times faster than Prettier, while keeping Prettier-compatible output and ergonomics.

Bundled Dev — faster dev server for larger apps

We’ve made progress towards shipping “Bundled Dev” (f.k.a. Full Bundle Mode), which uses Vite’s production bundler during development. This should lead to significant dev server speed-ups in larger applications and reduce network overhead when working with remote sandboxes.

We’re looking forward to bringing Bundled Dev out of experimental status soon. Cloudflare’s Dashboard already uses Bundled Dev for all internal developers.

Vite+ is now 1.0 — a unified toolchain

Speeding up tools is one way of making humans and agents ship software faster. Another way is by reducing decision fatigue (“Which linter shall I use?”) and providing great defaults. We are excited to announce that Vite+ is now 1.0. Vite+ bundles all of VoidZero’s tools together into a single unified and fast toolchain.

Vite+ ships with Vite 8, Vitest 5, Rolldown, Oxlint, Oxfmt and task caching built in, and comes with great defaults. Check out the Getting Started guide to try it out today.

Cloudflare’s Open Source Investment

VoidZero was born in open source, and we believe a more sustainable open-source ecosystem benefits everyone. VoidZero is a proud member of the Open Source Pledge and as part of Cloudflare, we have the opportunity to expand that impact.

Cloudflare committed $1M to a Vite ecosystem fund to support maintainers and contributors in the original announcement. Since then, Cloudflare has committed an additional $1M to open source. Together, we’re doubling down on open source.

The open-source toolchain for the entire JavaScript community

There is more to come! Features we plan to ship in the next few months include major improvements to Oxc’s parser with up to 3x potential speedup, a re-designed chunking algorithm in Rolldown, and an open-source, self-hostable version of Void, the Vite-native deployment platform built on top of Cloudflare. Cloudflare and VoidZero both recognize our responsibility that we have to developers, and we do not take the community’s trust in us for granted. We’re in this for the long haul.

We are excited to keep shipping faster tools, and will continue to make every decision with the community in mind, just like we did when raising venture capital, or when we open sourced Vite+, or when we joined Cloudflare. Thank you for continuing to trust us with your projects and apps, supporting us, and building with us.

Saving another 100TB of RAM with math (and Rust)

Post Syndicated from Kevin Guthrie original https://blog.cloudflare.com/saving-100-tb-of-ram-with-math/

Cloudflare operates at a scale so big that even after working here for years, it doesn’t seem real. We have thousands of servers all over the world with petabytes of RAM and millions of CPU cores, and all of it is pushed to the max. As vast as those resources feel, they are still finite, and when you need every service to run on every node, it doesn’t leave room for wasted space.

At this scale, small improvements are greatly magnified, so even 1%-at-a-time improvements are worth celebrating. And some tweaks add up to a lot more: in this post, we’ll look at how small changes to a single algorithm reduced the memory footprint of one of our Pingora-based services significantly. That allowed us to reclaim more than 100TB of RAM globally, on top of the 100TB of memory the DNS team was able to shed last month.

Waste not

Maintaining equitable resource sharing between teams is not easy, especially in large organizations. One of the ways Cloudflare ensures the balance is kept is through the tireless efforts of the wonderful Performance team. 

This story starts with a ticket filed by Ivan who found: Excessive memory usage from pingora-ketama in Pingora Backend Router. The finding was that our internal load-balancing service, Pingora Backend Router (yes, PBR), was using significantly more memory than expected — specifically in structures associated with pingora-ketama, which is our open-source library for handling consistent hashing.

In order to talk about how we addressed this seeming overuse of memory, we need to talk about what consistent hashing even is, why we are using it in PBR, and how it became so memory hungry. Along the way, we’ll learn some Rust and even a little math.

Consistent hashing

Consistent hashing is a widely used method for distributing tasks across multiple servers in a way that does not require large changes when servers are added or removed. Internally we use it to route cacheable requests to servers by URL. This allows us to keep only one copy of a file stored per data center and gives a stable way to find the location of each file. We have mentioned this system before, but let’s take the time to walk through how and why this algorithm is used and how it works.

The key concept of consistent hashing is that while hash functions can accept any kind of input, their output is limited to a single unsigned integer (32, 64, or 128-bit integers depending on which hash function). This allows us to relate tasks and servers to each other in a consistent way. Most discussions of consistent hashing have you think of that output space as a continuous, circular ring that wraps around from its max value to zero. This depiction makes for some nice visualizations, but it can also make the simple concept of integer ranges seem more complicated than it needs to be. For our discussion, we’ll represent the 32-bit output of our hash function as a number line.

Now, let’s say we have a set of servers, A, B, & C, and a set of tasks t-z. We can map each onto the number line based on the hash of their representative values, so something like IP addresses for servers and cache keys for tasks.

Assigning tasks to servers is now just a matter of finding the first server to the left of each task. We can represent this visually by coloring in the region of hashes that will be associated with each server. Notice that the range covered by server C wraps around to the beginning, hence the idea that hashes exist in a ring.

And that’s it. At a base level, consistent hashing is this simple — but it doesn’t take long to see that there is room for improvement. Notice that the range covered by server A in our example is significantly larger than that of either B or C. This is a problem because the fraction of the requests a server handles is going to be proportional to the size of its range on the number line. Ideally we would like to guarantee each server will have an equal size, but because hashes are essentially random numbers, we have to talk about the size of the regions in terms of statistics. 😨

Math and consequences

First: don’t panic. I promise I'm not about to lie to you and that we will stay safely within the bounds of a day-one probability lesson. When we talk about statistical distributions, there are two big factors that help us quantify uncertainty in helpful ways: expected value and standard deviation. In (over-)simplified terms, expected value gives us a point where measurements based on a distribution will be centered, and standard deviation tells how close to that central point most measurements are likely to be.

For consistent hashing, we can calculate these factors for the fractional size of the range associated with one of N servers. (Details on where this formula comes from later).

In terms of concrete numbers, let’s say we have 100 servers. The formulas above give:

That tells us that we can expect that the range each server handles will be centered around 0.99% of the total and most of the lengths to fall within 1% of what's expected. This sounds good until we realize that that’s 0.99% of the total length. We need to scale the standard deviation by the expected value to see how big the error is as a fraction of the target size. This value is called the coefficient of variation.

What if we add hashes?

The simplicity of consistent hashing is a double-edged sword. It’s easy to understand and implement because everything is turned into easily-relatable hashes on the same numberline, but any improvements to the system will also need to be relatable to that numberline. That means the solution to any consistent hashing problem can only be more hashes. It’s less like a golden hammer (a tool with which all problems look like nails) and more like a golden nail in that it turns all tools into hammers.

To solve the problem of imbalanced workloads, we can add multiple hashes to represent each server instead of just one. We’ll get to the math behind this momentarily, but it should make some intuitive sense that while each individual range has a large standard deviation, adding a bunch together should make their total size even out. If we take our three-server example from the above diagrams and add two more hashes at random for each server, we see that it helps even out each server’s workload. 

This is an admittedly contrived example. The random nature of the system means there’s no guarantee how much improvement you will get from adding 2 additional hashes per server, but it should make some intuitive sense that combining more of these hash segments together produces a more even distribution. Each segment in the sum has a chance of balancing another. Maybe one is too short; maybe one is too long. This is essentially what the law of large numbers tells us should happen… The obvious problem is it only works for large numbers. In NGINX, the baseline number of hashes per server is hardcoded to 160, and Pingora uses the same value as the default. I’ll spare you the math for now, but if we go back to our 100-server example, if we use 160 points per server instead of just one, the coefficient of variation (which we can think of like an error margin) drops from about 99% to about 8%, a significant improvement.

What if we add more hashes?

We saw above that increasing the number of hashes per server by a constant amount allows us to improve how evenly workloads are distributed per server, but what if we don’t want to distribute the work evenly? In Cloudflare’s case, we have some servers that have more storage space than others, so it would be better to have the number of requests allotted to a server be proportional to its disk space. One way to accomplish this is with the ketama algorithm. The naming is a little funny because the algorithm is named after the library where it was first implemented, and the library was named … well you can google it 😶‍🌫️.

For us, since we want workload to be scaled based on storage, we can use the disk space as the weight, which is exactly what the Pingora team has been doing for years. Elsewhere in the company where workloads are more compute-intensive, weights might be based on CPU or GPU count.

What if we add even more hashes???

The last problem we need to address is that so far we are working under the assumption that any server can handle any request, but in practice that is not the case. Things like compliance requirements or enabled caching features mean only a subset of servers can handle any particular request. Unfortunately, unlike before, we can’t solve this problem by adding more hashes to the same ring. We have to add completely new rings, and not only that — every combination of features potentially needs its own specific ring!

Storage improvements

One big improvement came from Zaidoon, who had an insight about our struct for storing hashes in PBR. That struct looks like this:

Unfortunately, Rust doesn’t make it that easy. Changing the size of the index as we did above does nothing to reduce the memory footprint. This is because Rust has alignment rules that require the size of a structure in memory to be a multiple of its largest (or “most aligned”) field. In this case, the hash is the largest with four bytes, so when stored in memory, a Point is required to have size $mN \times 4m$, so the minimum size is eight bytes.

Luckily there are well-known ways around this. You (meaning me) might be tempted to use #[repr(packed)], but that is controversial for good reasons. A safer but less readable solution is to store the hash and index as raw byte array and access them with getters. Both methods compile to the same thing.

This simple (if wordy) change reduces the amount of memory used for consistent hashing by a whopping 25%! In order to do better than that, we’ll need to jump back into the math, so everybody hang on to something; this is the home stretch.

What if we tried fewer hashes?

You may have noticed that we gave the formula for the standard deviation for the case where there is only one hash per server. Deriving the formula for the case where there are $m k m$ hashes per server is not easy, and most sources only give you an approximation or an asymptotic limit, but not us. I might not be a statistician, but I grew up with a calculus teacher (Hi, Mom!), and I wanted to know the actual value. The full derivation is in a supplemental post, but here is the payoff.

To see how increasing the hash count improves the accuracy, we need to look again at the coefficient of variation.

The predictions from my beautiful math only work if we think about hashes in a continuous ring, but in practice we use 32-bit numbers for the hashes that have the potential for collisions, and the probability of collisions goes up surprisingly quickly as the number of hashes increases (see the birthday paradox). Collisions matter because in the ideal case, every hash contributes to the volume and distribution of requests handled by the associated server, but a collision means some contributions are randomly dropped, introducing unpredictable error. If we compare some simulated results with 32-bit hashes with the predicted error rate, we can see that for data centers with 2048 servers, the error rate increases: between 10,000 and 100,000 hashes per server.

Ultimately, even though this realization feels kind of bad, it’s great news for our plan to reclaim some RAM! Now that we have some math to back it up, we determined that we could decrease the number of hashes we were generating for each server by 90% without incurring any appreciable error, so that is what we set out to do.

Migrating without melting origins

There was one more problem: changing the hash ring changes where some cacheable requests go. Even if the new ring is better, switching the whole network at once would effectively invalidate almost all cached content. It would turn a memory optimization into an apocalyptic increase in origin traffic.

So we did not make this a single global flip. For a while, PBR carried both versions of the cacheable load balancer in memory: the old ketama ring and the new smaller one. Each request used our normal migration framework to decide which ring should select the backend. That meant the rollout decision was stable per request hash, and it also gave us a clean rollback path. If anything looked wrong, we could send new requests back through the old ring without redeploying PBR.

We then rolled the migration out in layers. We started with small validation locations, moved through progressively larger groups of data centers, and only then continued toward the rest of the world. 

The important part was that we controlled two dimensions independently: how much traffic used the new ring, and where that traffic was allowed to move. A plain global percentage rollout would have spread cache churn everywhere at once. Data-center-scoped rollout kept the blast radius small and made it much easier to tell whether a change was actually safe.

During the migration, we watched backend-selection traces, ring-version counters, PBR connection errors, process memory, startup time, cache behavior, and origin traffic. Once the migration reached 100%, we removed the temporary old-ring path, and voila!

The chart above shows the comparison of the memory used by PBR the week of the change compared with data from a few weeks before, as well as the result of subtracting one from the other. The sharp drop is the day where the version of PBR with the large (now unused) hash rings was decommissioned forever. Looking at the difference, we get the satisfying result that our changes dropped the used memory by 100TB!

Try it yourself

All the changes we talked about in this post are available now in the pingora-ketama crate in the form of a (for now) unadvertised cargo feature. The v2 ring has the compacted storage format, a faster sorting method, and the ability to scale the base number of hashes per node. Our focus in making these changes had to be on stability and control, so the v1 ring is identical to what pingora ketama has always used, and the library makes it possible to run both simultaneously and decide on a request-by-request basis which to use and when. 

Beyond trying our literal consistent hashing changes, I would like you to take away from this some inspiration to dig into your own systems to see what “simple” or “obvious” decisions are hiding potential wins, if you’re willing to get into the numbers. You might not be able to solve all your problems with Rust, but math is universal.

AIs Compress Exploit Timeline

Post Syndicated from Bruce Schneier original https://www.schneier.com/blog/archives/2026/09/ais-compress-exploit-timeline.html

Give an AI agent a mere rumor of an exploit, and it’s enough for them to find it.

What’s worse, I found I could use my own agents to find the exploit just by knowing roughly what it was about and so could have been exploiting it well before the public patch was available! Given that just the rumour of a security issue seems enough to give attackers enough info to find new exploits, we’re going to need to change the way we deal with security responses in open source.

Simon Willison comments:

Anil points out that this rate of discovery appears incompatible with existing open source embargo practices for new issues. If an issue can become an exploit this fast, we need to figure out new processes for keeping our communities safe.

The state of AI for security: Measuring what matters most for building trust

Post Syndicated from Anshumali Shrivastava original https://aws.amazon.com/blogs/security/the-state-of-ai-for-security-measuring-what-matters-most-for-building-trust/

Security teams are starting to actively use AI for security work, including vulnerability triage, penetration testing, threat modeling, incident response, and code review. The promise is speed, but a security tool that moves fast and raises too many false alarms doesn’t save time. Engineers spend time on false alarms, on-call is noisier, and teams distrust findings that matter.

Today, we’re releasing Deception Benchmark, the first benchmark designed to measure that trust problem directly. It tests whether a model can distinguish real vulnerabilities from code that looks risky but is actually safe. The benchmark includes 14,822 samples across 16 languages and more than 70 Common Weakness Enumeration (CWE) categories. We evaluated 12 models from five providers and are releasing the dataset and whitepaper to the community. Existing benchmarks measure whether AI can find or exploit vulnerabilities. This is the first to measure whether it can tell real vulnerabilities from false alarms. Under standard prompting, precision at distinguishing real vulnerabilities from false alarms landed in the mid 50s; as likely to be inaccurate as accurate.

In offensive tasks, there’s often a clear result: the exploit works or it doesn’t. Defensive reviews are harder to verify than offensive tasks; a model might recognize a suspicious pattern even when a mitigation makes the issue non-exploitable. In practice, useful systems need to reason about the code, the mitigation, and sometimes the surrounding environment.

The measurement gap

The community has made progress on security evaluations. CyberGym tests agents on more than 1,500 realistic tasks. Meta’s CyberSecEval and CyberSecEval 2 measure exploit generation. CYBENCH evaluates capture the flag (CTF) challenges. SEC-Bench and VulnBench push toward authentic security workflows.

Recent work reinforces both the progress and the gap. ExploitGym measures whether AI can escalate from a crash to a working exploit. Microsoft’s Project Perception deploys multi-agent red/blue/green teams for continuous defense. OpenAI’s GPT-Red shows that self-play red-teaming finds novel attacks that frontier models can’t defend against. Since then, OpenAI disclosed that its GPT-6 Astra model crossed the Critical cybersecurity capability threshold, and both OpenAI and Anthropic reported incidents where models gained unauthorized access to production systems during evaluations. The offensive side is moving fast. But none of this work measures the defensive precision question: when an AI system flags code as vulnerable, how often is it right?

Introducing Deception Benchmark

14,822 samples, 16 languages, and more than 70 CWE categories. We call it Deception Benchmark because the safe samples are designed to deceive models. It has real vulnerability patterns, real frameworks, real idioms, with mitigations that quietly close the exploit path. The goal is to classify code as vulnerable or safe, with no hints.

Consider a Flask endpoint that accepts user input and queries a database. A model will pattern-match to SQL injection, but the query uses parameterized statements, so the exploit path is closed. A single-turn classifier flags the pattern and moves on, never checking whether the exploit can actually work. Production tools rely on multi-step loops and agentic workflows to compensate, but that scaffolding masks whether the model itself understands the code. This benchmark strips the scaffolding away and asks the model to make the call in a single pass, so what it measures is understanding, not how many tries a harness takes to get there.

We built every sample through an adversarial loop: generate, test against frontier models, harden, repeat. If a model gets it right easily, the sample doesn’t survive. The result is a benchmark calibrated to the frontier, not below it. Building it this way is expensive. Generation and hardening of the samples consumed tens of billions of tokens. We’re releasing the result so the community doesn’t have to repeat that cost.

This benchmark generates two challenge types. Code-level challenges (6,988 samples) present vulnerable and safe variants that differ by a subtle fix. Both look suspicious, only one is exploitable. Environment-gated challenges (2,707 samples) go further: same code, different deployment context. A Kubernetes Network Policy blocks the server-side request forgery (SSRF) path. An identity and access management boundary prevents privilege escalation. The pattern is visible in the source. The infrastructure makes it unexploitable. The model has to figure out which scenario applies.

All samples were purpose-built for this benchmark, grounded in real-world patterns, real frameworks, real CWEs, and real infrastructure; without IP concerns or training data contamination.

Large-scale quality data with LLMs and humans in the loop

Generating reliable labels at this scale is difficult: a single pass—by people or by models—leaves errors that skew scores. So we treat labeling as a convergent audit loop rather than a one-time step. Every label is re-examined by multiple independent reviewers, blind to one another and to the original reasoning that produced the label. Disagreements escalate to direct adjudication, where the original reasoning is evaluated against the challenge. Unresolved cases go to human review. We repeat the loop until the scored set converges below a dispute threshold: under 3 percent of samples still contested by independent review, with a target of under 1 percent surviving human adjudication. One choice makes this defensible: we never relabel a disputed sample. When reviewers disagree, the sample moves to the unscored pool instead of being given a corrected label, so a bad challenge can remove a sample but can never introduce a wrong label into the scored set.

A human review of 100 randomly drawn scored samples found no label errors. We describe the full process in the whitepaper.

The results

The benchmark is roughly balanced: half vulnerable, half safe, so a random classifier scores 50 percent. We report two error rates separately, because they fail in opposite directions. The false positive rate (FPR) is how often the model flags safe code as vulnerable. These are the false alarms that waste an engineer’s time. The false negative rate (FNR) is how often it misses a real vulnerability and calls it safe. Accuracy alone hides this: a model that labels everything vulnerable catches every bug (0 percent FNR) but flags all safe code (100 percent FPR) and still scores about 50 percent. We consider FPR below 10 percent and FNR below 10 percent the minimum bar for production use.

Figure 1: FPR compared to FNR for 12 models across two prompting strategies. No model reaches the generous bar

Figure 1: FPR compared to FNR for 12 models across two prompting strategies. No model reaches the generous bar.

Model

Prompt

Accuracy

FPR

FNR

GPT-5.6 Sol Direct 54.9% 92.5% 0.9%
GPT-5.6 Sol PoE 58.9% 58.6% 23.1%
GPT-5.5 Direct 56.9% 87.8% 1.3%
GPT-5.5 PoE 62.9% 63.6% 12.4%
GPT-5.4 Direct 60.2% 81.0% 1.5%
GPT-5.4 PoE 77.7% 10.1% 33.6%
Llama 3.3 70B Direct 58.8% 84.2% 1.1%
Llama 3.3 70B PoE 72.2% 10.2% 44.2%
Claude Haiku 4.5 Direct 55.6% 92.1% 0.0%
Claude Haiku 4.5 PoE 75.6% 22.4% 26.3%
Claude Opus 4.6 Direct 55.9% 91.3% 0.1%
Claude Opus 4.6 PoE 75.8% 42.7% 7.0%
Claude Opus 4.7 Direct 58.3% 85.5% 0.9%
Claude Opus 4.7 PoE 75.9% 32.0% 16.8%
Claude Opus 4.8 Direct 53.8% 95.7% 0.2%
Claude Opus 4.8 PoE 75.8% 32.5% 16.4%
Claude Opus 5 Direct 77.3% 41.5% 5.2%
Claude Opus 5 PoE 79.3% 24.9% 16.8%
Claude Sonnet 5 Direct 62.9% 74.7% 2.2%
Claude Sonnet 5 PoE 74.7% 31.8% 19.2%
Amazon Nova 2 Lite Direct 56.3% 89.2% 1.2%
Amazon Nova 2 Lite PoE 70.1% 45.2% 15.5%
Mistral Large Direct 52.2% 99.0% 0.0%
Mistral Large PoE 65.5% 49.3% 20.6%

Among the general-purpose frontier models tested, no configuration achieves both FPR and FNR less than 10 percent on this benchmark.

Every model has the same failure mode. With direct prompting, they catch up to 95 percent of real vulnerabilities but also flag 41–99 percent of safe code. Precision runs from 52 percent to 71 percent, clustered in the mid-50s; effectively as likely to be inaccurate as accurate. The models see a vulnerability pattern and stop reasoning. Proof-of-exploit prompting cuts false positives by 17–74 points but misses 7–44 percent of real vulnerabilities. The environment-gated challenges are worse: models flag the code and ignore the Kubernetes Network Policy next to it. No tested configuration keeps both false positives and false negatives below 10 percent.

These results reflect general-purpose models in single-turn prompting. Purpose-built systems with multi-step validation and tool use are a different operating point that we didn’t measure, and if a harness can close the gap between pattern recognition and genuine understanding, this benchmark is the place to demonstrate it. Two cautions before assuming it already does. Agentic verification is proven mostly on offensive tasks, where success can be confirmed: the exploit fires or it doesn’t. Judging that code is safe has no such oracle. Extra iterations re-sample the same judgment rather than confirm a negative, and a harness still inherits the base model’s understanding. If the model can’t separate an effective mitigation from an ineffective one in a single pass, more passes won’t add the missing knowledge. That’s what this benchmark measures: the model’s intrinsic ability to understand code, tested at the single-turn baseline where no scaffolding can mask the gap.

For security teams evaluating AI tools today: ask your vendors how their system performs on tasks like this, not just whether it finds vulnerabilities, but how often it’s wrong. Pair any AI-assisted review with human verification on high-risk code paths, and use Deception Benchmark to hold your tools accountable.

Availability

We built Deception Benchmark to simplify measuring this problem in a reproducible way. The public release includes the samples and evaluation workflow. We don’t release the labels, so submissions can be scored consistently over time without turning the benchmark into a memorization exercise.

Of the 14,822 samples, 9,695 are scored; the remaining 5,127 are held out and unscored, mixed in with the rest of the benchmark. The goal is straightforward: make it more difficult to optimize the benchmark compared to improving the underlying system. We describe that design in more detail in the whitepaper.

Deception Benchmark is available on GitHub, along with the whitepaper and submission instructions for verified scoring. If you’re building security tooling, you can download the dataset, run your system against the benchmark, and submit predictions for scored evaluation.

If you have feedback about this post, submit comments in the Comments section below.


Anshumali-Shrivastava

Anshumali Shrivastava

Anshumali is an Amazon Scholar and Full Professor of Computer Science at Rice University. His research on dynamic sparsity, sketching, and hashing pioneered techniques now central to efficient LLM training and inference. A two-time founder — ThirdAI (acquired by ServiceNow) and XMAD.ai (acquired by Workato) — he bridges theoretical computer science and practical AI systems at scale.

Neha Rungta

Neha Rungta

Neha is a scientist and builder who has spent her career making machines reason about complex systems at scale. Her work spans automated reasoning, formal verification, security, and AI, shaping systems including Cedar, IAM Access Analyzer, and Continuum. Today, she is forging the next generation of machine reasoning, combining LLMs, formal methods, and agentic systems.

How we rebuilt Cloudflare Workers’ module registry for Node.js compatibility

Post Syndicated from Logan Gatlin original https://blog.cloudflare.com/workers-module-registry-nodejs/

We’ve rewritten the module registry in workerd, the core open-source component of the Workers runtime, to be faster, more standards-compliant, and more closely aligned with Node.js' module registry.

Over the past few years, we’ve been adding support for more and more Node.js runtime APIs. The Workers runtime now supports every stable API from Node.js that you might want to use in a serverless context, and these APIs are now enabled by default, letting you deploy even larger Node.js apps to Cloudflare (now up to 64 MiB on all plans — we’ve removed the limit on compressed bundle size).

But API compatibility alone is not enough: Node.js applications also depend on how the runtime resolves, loads, and caches modules. ESM, CommonJS, and WebAssembly are each types of modules that you can import in your Worker’s code. The system within the runtime that handles all of this is called the module registry.

You can start using it today by enabling the new_module_registry compatibility flag in your Worker.

When you enable the new_module_registry compatibility flag:

  • import.meta.url, import.meta.main, and import.meta.resolve() all work.
  • Module specifiers are parsed and resolved as real URLs, including query strings and fragments.
  • node: built-ins resolve to the same module instance no matter how you reach them.
  • Import attributes (with { type: 'json' }) are correctly validated.
  • require() on an ES module follows Node.js' require(esm) rules.
  • Errors use consistent classes and messages regardless of which loading path triggered them.
  • Modules compile lazily when first imported (statically or dynamically).
  • WebAssembly modules support source phase imports.

For the full deep-dive on how this new module registry interacts with V8’s module APIs, we’ve added reference docs to workerd that break down everything in detail. But for most people building on Workers, you want to understand how these changes improve compatibility and help you build. To do that, we’ll dive into each of these changes in the sections below.

How the Workers runtime loads the code you give it

When you deploy a Worker to Cloudflare, wrangler or Vite “bundles” all of your Worker’s code from many files and dependencies into one or many modules, which are then uploaded to Cloudflare when you run wrangler deploy.

By default, Wrangler bundles nearly all of this code into a single module script. It runs esbuild under the hood, which processes then inlines relative imports and require() calls for most npm dependencies into that one file. The import and require() statements are replaced with regular functions as part of the process. By the time that bundle reaches the Workers runtime (workerd), there usually isn't much of a module graph left for the Workers runtime to deal with. Most of the different modules are bundled into one file. We have seen these scripts grow to as many as multiple hundreds of thousands of lines long.

Why is it necessary to bundle many modules into a single file before uploading server-side code to Cloudflare? It has been technically possible to upload multiple modules, and even modules of different types, in the Workers runtime for many years now. However, the runtime has not resolved modules in a way that was consistent with all the other runtimes. If, for example, your code or dependencies used import.meta.resolve() to resolve the path to another module, that code would fail because import.meta.resolve() was not supported.

When you use the Cloudflare Vite plugin, Vite 8 bundles your code using Rolldown, instead of Wrangler bundling your code using esbuild. Rolldown resolves imports and npm dependencies, converts CommonJS to ESM where necessary, and emits an entry module plus any additional chunks created through code splitting, such as dynamic imports. As a result, the Workers runtime receives a smaller, build-generated module graph rather than the application’s original source graph.

The new module registry implementation in the Workers runtime opens the door to bundlers like Rolldown to perform fewer transformations, and to rely more on the runtime to handle module resolution.

When you import a Node.js API in your worker, by default you are importing a module that is built into workerd. It is not bundled into your code as a polyfill. Wasm, text, and binary modules are provided to the Workers runtime as separate files too. They are referenced by specifier instead of being inlined. And if you deploy with --no-bundle, or your tooling uploads a Worker as multiple modules directly, the full module graph shows up at runtime exactly as you wrote it.

In all of these cases, something has to take a specifier, work out what code it actually points to, compile it, and hand V8 a module object it can link and run. In workerd, that's the module registry's job.

Why a new implementation?

The original registry resolves specifiers as filesystem-style paths, not URLs. That sounds like a minor distinction, but it ruled out a bunch of things: there was no clean way to implement import.meta.url, relative imports didn't follow the same resolution rules as new URL(), and protocols like node: and cloudflare: were handled as special-cased string prefixes instead of, well, protocols.

It also compiles your entire Worker bundle up front, whether or not a given module ever gets imported, and it keeps a separate, private copy of everything per V8 isolate. Cloudflare runs multiple V8 isolate replicas of the same Worker to spread load across CPU cores, so in practice that meant compiling the exact same source more than once, with keeping multiple copies of the source in memory.

None of this is really a bug, but it made it difficult to evolve the implementation without breaking changes. The new registry starts from URLs as the specifier format and treats laziness and cache sharing as things to design in from day one. The existing registry implementation is not going anywhere. Currently, deployed Workers will continue to work as they always have.

import.meta

The import.meta API provides information about the module, such as the module's URL, and whether it is the main entry point module:

That prints something like file:///bundle/index.js, main: true. 

import.meta.main is true only for the module configured as your Worker's entrypoint; every other module gets false.

import.meta.resolve() resolves a specifier against the current module without importing it:

It's a pure string transform, same as in Node.js and in browsers: it doesn't check that the resolved URL corresponds to a real module, and it throws a TypeError for a specifier that can't be parsed as a URL at all, rather than returning null. One detail worth knowing if you ever look closely at the output: it normalizes percent-encoding the same way new URL() does, which means it collapses paths like ./a/../b.js, but it does not decode characters that were already percent-encoded. import.meta.resolve('%66oo.js') resolves to file:///bundle/%66oo.js, not file:///bundle/foo.js.

Specifiers are URLs

Relative imports now resolve the same way as new URL(specifier, base) would, because that's literally what's happening under the hood. Full URLs work as specifiers too, not just relative paths:

The more interesting consequence is what happens with query strings and fragments. Per the same module-identity rules browsers use, a specifier with a different query string or fragment is treated as a genuinely distinct module instance, even when it points at the same underlying source:

./counter.js?a and ./counter.js?b load the same source, but they're evaluated separately, each gets its own import.meta.url, and each gets its own copy of any top-level state. Importing the same specifier with the same query string again still gets you back the same instance, so this isn't a way to force re-evaluation on every import.

Import attributes are correctly validated

The original module registry implementation silently ignores the import attributes in violation of the spec. It is expected that implementations throw an exception when any import attribute it does not understand is used.

json is the only import attribute type enabled right now, since it's the only one of the relevant TC39 proposals that has reached Stage 4. text and bytes are recognized, because they track the Import Text and Import Bytes proposals, but they're rejected with a specific error instead of being silently ignored or treated as unsupported syntax:

Any attribute key other than type is now a hard error too, rather than being ignored:

And if the type you specify doesn't match what the module actually is:

require(esm) follows Node.js' rules

If you require() something that turns out to be an ES module, whether that's directly inside a CommonJS module or through require('node:module').createRequire(), the registry follows Node.js' require(esm) behavior:

  • If the module has a string-named export called 'module.exports', Node.js' actual mechanism for letting an ES module control what require() sees, that value is returned.
  • Otherwise, require() returns the module's namespace object.
  • The one exception is workerd's own node: built-ins. They're implemented as ES modules that wrap a CommonJS-style API in a default export, so requiring one returns that default export directly. require('node:buffer').Buffer behaves the way you'd expect; you don't get a namespace object with a .default you need to unwrap yourself.

There's a restriction that comes along with this: if the module you're requiring, or anything in its module graph, has a top-level await, require() throws instead of blocking or handing back something half-finished:

This matches Node.js' own ERR_REQUIRE_ASYNC_MODULE restriction: require() has to return synchronously, and there's no reasonable value to hand back for a module that hasn't finished evaluating yet. Use import() for anything async instead. The check holds regardless of import order too: a module doesn't become require()-able just because something already import()'d and fully evaluated it earlier.

If you're requiring output from a bundler that predates Node.js' require(esm) support and sets a truthy __cjsUnwrapDefault export as a marker, that takes priority over both rules above and returns the default export. That's purely there so existing prebuilt bundles keep working.

Errors are consistent, and use the right class

Regardless of whether resolution fails through a static import, a dynamic import(), or require(), you get the same class of error with the same message shape:

"Module not found" is a plain Error, since it's a failure to locate something rather than a problem with the value you passed in. A specifier that can't be parsed as a URL at all is a TypeError, matching Node.js' own ERR_INVALID_MODULE_SPECIFIER. A circular dependency that V8 can't unwind is also a plain Error, never a TypeError. This mostly matters if you're building something on top of dynamic import(), like your own loader or a retry wrapper, since you can now branch on the error class or message reliably no matter which loading path triggered it.

WebAssembly source phase imports

You can now import the compiled-but-not-instantiated form of a WebAssembly module directly, using source phase imports:

or dynamically:

Either way you get a WebAssembly.Module back directly, instead of importing the module normally and pulling it off the default export. As source phase imports are a new feature of the language, right now this only works for WebAssembly; trying it on any other module type throws a SyntaxError, matching the behavior of Node.js and other runtimes.

What's next

Try it out! Add the new_module_registry compatibility flag to your Worker:

It doesn't have a default on date yet, so it won't turn on automatically for your Worker, old or new, no matter what compatibility date it's using. You will need to add the flag explicitly.

We’d love your feedback. workerd is open source. If you run into behavior that looks like a regression rather than one of the changes described here, please file it against the workerd repository.

AWS Weekly Roundup: EC2 application status checks, IAM role manager, OpenAI Daybreak on Bedrock, and more (August 17, 2026)

Post Syndicated from Channy Yun (윤석찬) original https://aws.amazon.com/blogs/aws/aws-weekly-roundup-ec2-application-status-checks-iam-role-manager-openai-daybreak-on-bedrock-and-more-august-17-2026/

Last week, the OpenSearch and Valkey teams visited Seoul to meet open source developers and contributors in the Open Source Summit Korea 2026 and MCP DevSummit Seoul 2026. At the four-day event, community leaders and users of open source projects and emerging agent AI gathered to share knowledge, collaborate on solutions, and push the projects forward.

Leaders of the Korean OpenSearch communities volunteered to participate in the booth, and also had time to network and interact in the user group meetup.

OpenSearch is an open source, enterprise-grade search and observability suite that brings order to unstructured data at scale. On June 9, 2026, OpenSearch 3.7 introduced new tools designed to query, alert, and track SLOs across logs, traces, and metrics through a single interface and retrieve vectors up to 5.5x faster for improved search performance. Since July 30, 2026, you can run OpenSearch version 3.7 on Amazon OpenSearch Service for improvements in vector search performance, search relevance, and Query Insights.

Valkey is an open source high-performance key/value datastore that supports a variety of workloads such as caching, message queues, and it can act as a primary database. On May 19, 2026, Valkey 9.1 introduced a redesigned I/O threading model that improves throughput by up to 17% and reduces memory usage for strings under 128 bytes by up to 20%. Since June 23, 2026, you can run Valkey 9.1 in Amazon ElastiCache for node-based clusters, delivering higher throughput, improved memory efficiency, and stronger access control for multi-tenant workloads.

You can meet our open source teams at upcoming OpenSearch and Valkey events.

Last week’s launches
Here are some launches that got my attention:

For a full list of AWS announcements, be sure to keep an eye on the What’s New with AWS page.

Other AWS news
Here are some additional projects and news items you may find interesting:

  • The deprecation of email validation in AWS Certificate Manager: ACM will discontinue support for email-validated public certificates by September 30, 2027. If you use email validation for your ACM public certificates, you need to migrate to DNS validation before that date. For Amazon CloudFront distributions, HTTP validation is also available.
  • The next-generation AWS VPN Client with CLI support and admin controls: You can use a new AWS VPN Client built on OpenVPN3. With the new client, you get full backward compatibility with existing AWS Client VPN endpoints while delivering the automation capabilities and security posture that enterprise networking teams have been asking for.
  • Oracle Exadata on Exascale for Oracle AI Database@AWS: ExaDB-XS brings Exadata-class performance and availability through a consumption-based model. With ExaDB-XS, you can scale compute and storage independently in small increments and pay only for what you consume.

For a full list of AWS blog posts, be sure to keep an eye on the AWS Blogs page.

Learn more about AWS, browse and join upcoming AWS-led in-person and virtual events, startup events, and developer-focused events including AWS Summits and AWS Community Days. Join the AWS Builder Center to connect with builders, share solutions, and access content that supports your development.

That is all for this week. Check back next Monday for another Weekly Roundup!

— Channy

Python Now Has a Post-Quantum Encryption Library

Post Syndicated from Bruce Schneier original https://www.schneier.com/blog/archives/2026/08/python-now-has-a-post-quantum-encryption-library.html

This is good:

Post-quantum cryptography is now one pip-install away for the entire Python ecosystem. With funding from the Sovereign Tech Agency, we implemented support for ML-KEM, the NIST-standard key-establishment primitive, and ML-DSA, the NIST-standard digital-signature primitive, in pyca/cryptography.

Remember, the reason to do this now is because there’s no emergency. And because you will make your systems crypto agile, which is always a good idea.

Announcing Cloudflare Ambassadors, Community Engineers, and another $1M in open-source funding

Post Syndicated from Kristian Freeman original https://blog.cloudflare.com/community-program-refresh/

As a platform for helping build a better Internet, Cloudflare helps turn ideas into real products and experiences around the world. Across communities and backgrounds, developers build with Cloudflare using the tools they love, shaping what comes next for the Internet while inspiring, collaborating with, and teaching others.

The community is where some of Cloudflare’s best moments happen. Students show their friends how to deploy Workers for the first time. Discord users answer questions from other developers via working code samples, instead of links to documentation. Open-source contributors build novel solutions to solve their own problems, then share them with the world. Organizers host events that give builders from all backgrounds the space to start building their dream project.

All of these represent a community at its best: people helping other people build.

This spirit of community is an exciting and vital part of helping to build the Internet. Those who step up to educate and support others, or to invent, build, or maintain tools shared across the ecosystem, make lasting contributions to the health and potential of the Internet.

We want to have their backs.

That's why today we’re announcing an improved community program, designed to better support, recognize, and empower the people getting involved, while working with them to shape what comes next.
The program has two main tracks:

  1. Cloudflare Ambassadors: Bringing Cloudflare to their own communities.
  2. Cloudflare Community Engineers: Contributing to open-source projects that improve the Internet.

We’re launching a new home for the program where you can learn more and get involved: cloudflare.com/community.

Cloudflare Ambassadors

Cloudflare Ambassadors are people who bring Cloudflare into their own communities. You can probably think of people in the communities you value who share a genuine passion for a product or technology. It’s inspiring and we love to see it. When that enthusiasm includes the tools we’re building here at Cloudflare, it’s especially exciting for us.

Following our annual application process (more below), we’ll announce the year’s Cloudflare Ambassadors cohort. Selected Ambassadors will receive support, resources, and benefits to help their community thrive and bring their ideas to life. Ambassadors can serve for up to two years, giving them meaningful time to build momentum while helping us support more communities over time.

What Ambassadors do and what we provide

Being an Ambassador might mean organizing a local event, leading a student group, creating spaces where builders can learn together, publishing tutorials or sharing content online, or being the person others turn to when they want to understand what’s possible with Cloudflare. 

Ambassadors will take the lead on events in their communities, whether on campus, through local organizations, or across their city. When hosting meetups, hackathons, workshops, or talks, they will be able to apply for support in the form of credits, marketing assets, technical resources, and more.

We’ll also give them a visible role in Cloudflare’s online community spaces, including Discord, so that other developers know who they are, and that they’re here to help.

Applications are open now, and will be accepted through September 6. Those selected as Ambassadors will be informed of their selection by October 5.
Apply to become a Cloudflare Ambassador

A great example of the enthusiasm we’re looking for comes from Sruthi Pereddy, a Computer Science major at University of Michigan and a current intern on Cloudflare’s Recruiting Ops team. Sruthi’s work within Cloudflare has created a drive to share and explore more with others:

“Whether it’s hackathons, startup venture funds, or coursework, I want to show my peers that Cloudflare is a go-to developer platform for whatever they’re building,” Pereddy says. “Students are ready to build, but often feel constrained by resources. I’m excited to bridge that gap and make sure they have the infrastructure to turn their ideas into reality from day one.”

Cloudflare Community Engineers

Some community work happens in person, but a great deal of community work also happens in code. Much of Cloudflare’s Developer Platform is built on open-source work, or is open-source, like workerd and quiche. Open-source contributors, especially maintainers, do wonderful work and embody so much passion and determination. We’re eager to support them, especially since their work can sometimes feel thankless. So we’re doubling down on our efforts to build stronger incentives and directly support the maintainers doing this important work.

Last year, we announced our sponsorship of the web framework TanStack. TanStack creator Tanner Linsley says that sponsorship has had a major impact.

“Cloudflare’s sponsorship has given us room to keep investing in foundational open-source work that’s hard to tie to a single product or launch, maintaining the core libraries, improving docs and tooling, supporting contributors, and putting real time into bigger bets like TanStack Router and Start,” Linsley says. “It’s also helped us make sure TanStack apps have a really solid path onto Cloudflare’s platform. More than anything, that support buys stability, which is kind of everything when you’re building open source for the long haul.”

Today, we’re expanding on our previous open-source investments by introducing Cloudflare Community Engineers. Earlier this year, we announced a $1M fund as part of our acquisition of VoidZero to support the Vite community. We’re committing an additional $1M in funding to sponsor and support open-source projects over the next two years, with eligible Community Engineers receiving grants from the fund to support their continuing work in open source.

The Community Engineer program does not have a maximum term. Open source work doesn’t neatly fit into annual cycles. Some projects require maintenance for years, while other times, contributors do the work that is needed at exactly the right moment. This program is intended to support that.

To begin, we’ll focus on developers working on things in the orbit of our own open-source projects — projects like Astro, Agents SDK, EmDash, Hono, and Vinext. We’ll also grant our Community Engineers a special designation in Cloudflare’s Discord server and other online spaces.

Applications for Community Engineer grants will open at a later date.

Making our Discord better as it grows

Since we launched Cloudflare’s Discord server in 2020, almost 100,000 Cloudflare users have joined. Our Discord server has become one of the main places where developers ask questions, share projects, and provide valuable feedback. But of course, the more a Discord community grows, the more effort is required to keep it healthy and approachable.

To address this, a new Discord committee will help to maintain and grow our Discord community, with Cloudflare Ambassadors joining Cloudflare staff on the committee.

This is not about being on hand to perform moderation and admin tasks. We’ve been building tools and automations to help us do that with far less human intervention. Our new automated protections against spam and malicious links are starting to relieve this burden, allowing our Developer Relations team to help manage things where some human insight is needed.

In fact, we’ll be open-sourcing and sharing those tools soon because we think every Discord server could benefit from less spam and malicious content.

The committee will help provide a useful connection to those building and managing products at Cloudflare. They’ll be able to steer people and conversations to domain experts and convene conversations and sessions with internal teams and makers around the community. They’ll be much more focused on content and opportunities than on the type of Discord administrivia that can otherwise swallow so much time and energy.

We want our Discord to be easier to use, contribute to, and trust. It should be a place where builders find each other, help each other, and shape the future of the platform together. We believe this is the way.

Ready, set, go!

To learn more about the community program, and to apply for a role, visit the new community site at cloudflare.com/community.

Applications to join the 2026-27 Cloudflare Ambassadors cohort have now officially opened. Be sure to apply by September 6.

And don’t forget to join the conversation in the Cloudflare Discord.

Cloudflare OS: an open platform for agents, apps, and work

Post Syndicated from Phillip Jones original https://blog.cloudflare.com/cloudflare-os/

Every organization has a mission, a reason for being. Organizations pass that mission — along with their terminology, procedures, systems, standards, and ways of working — to their people. People, in turn, take this context together with their own experience and work towards the mission.

Work can take many forms, from code, to documents and slides, to relationships, to outcomes in the physical world.

Some of these are straightforward: code either runs or it doesn’t. Agents have been using this feedback loop to produce code that “works” for developers over the last couple of years. But what about the rest of us?

Bringing the same leverage to the rest of the organization is a harder problem. Agents need to understand the context of the company and be able to reach the systems people use to do their jobs. They need to turn that context and access into work that moves the organization towards its mission.

That’s why we created Cloudflare OS. It gives every person an agent and workspace built around their company: how it works, what it knows, and the systems it relies on.

In May of this year, we gave every person at Cloudflare access to the first version of Cloudflare OS. Thousands of people across every function, many of them outside of engineering, use it every day to create documents and slides, automate repeatable tasks, and build small apps to visualize data and help them do their work.

Cloudflare OS also gave everyone a shared library of context and skills built by teams at Cloudflare. It captures our terminology, procedures, and best-known ways of doing recurring work as instructions an agent can follow. When one person figures out a better way to do something, everyone else can use it.

Today, we are open sourcing a new version of Cloudflare OS. Any organization can deploy it, connect it to internal systems, and make it their own.

What we learned from the first version

The Cloudflare OS we are open sourcing today is based on what we learned from running the first version internally, a journey our CIO, Sam Rhea, covers in his blog post.

The first version centered on individuals working with agents through private workspaces. Apps were static rather than live software connected to internal systems, and mostly deterministic jobs still required running an agent skill again and consuming more model tokens.

Collaboration exposed a more fundamental challenge. Access to an MCP server told us which tools an agent could call, but not which underlying resources the agent had observed. Once people began sharing workspaces, apps, and outputs, we needed to ensure that collaboration could not expose information someone was not permitted to see.

We rebuilt Cloudflare OS on a new foundation to solve these problems. Security had to be part of the platform, not something every person building an app or using an agent has to implement correctly.

The result is a platform designed to belong to the company running it. You can customize the interfaces, connect your tools, and add the skills and context that capture how your organization works.

Introducing Cloudflare OS

Cloudflare OS starts with a conversation in your browser, like many other AI tools. What makes it different is that each conversation is grounded in the context and skills your organization has curated. Give your workspace a goal, and it can draw on that knowledge and work with the tools and data your organization already uses to achieve it.

Cloudflare OS combines three parts:

  • An agent workspace grounded in context and skills your company curates, with an isolated runtime where agents can write and run code.
  • A new security and governance framework for safe access to internal data and services.
  • A platform for personal, modifiable apps that people can build, share, and continue changing.

What begins as a conversation can become a doc, an app, or a workflow that continues doing the work.

An agent workspace for everyone in your company

Agent workspaces were designed for everyone in your organization to use. You interact with them in your browser, so you don’t have to be a developer or know how to use a terminal. 

A workspace combines agent sessions, persistent state, outputs and files, resource access, and an isolated runtime where the agent can write and run code.

They come loaded with the curated context and skills your team or company has collected. No more reinventing the wheel for every task — if someone on your team has figured out the best way to do something, everyone benefits. People no longer have to explain the same process, terminology, and best practices to a model every time they start a task.

A few things you can do:

Research and ask questions

Ask a workspace to research a topic using company context and the resources you make available to it. The agent can write code to search, filter, join, and analyze information instead of pulling an entire dataset into the model’s context window.

Create docs, slides, and spreadsheets

A workspace can turn its research into a document, presentation, or spreadsheet that you can continue editing. These outputs do not have to be static files. They can remain connected to live data, be updated as their sources change, and still be exported to familiar formats or services such as Google Drive.

Create collaborative, connected apps for your team

When a document or spreadsheet is not enough, the agent can build an app with its own interface, logic, and state. The app can use connected company resources and support multiple people working together.

Run deterministic workflows 

Not every job needs a full agent session. Many are a known sequence of steps with one or two places where judgment is useful. A workspace can turn those jobs into mostly deterministic workflows, using code for the predictable steps and a model only where it adds value. Workflows can run on demand, on a schedule, or when an event occurs in a connected system.

Cloudflare OS gives agents and apps governed access to systems of record through Gatekeepers (more on this in the security section below). It also supports existing Model Context Protocol (MCP) servers your organization already uses via MCP Server Portals.

A new security and governance framework for safe access to internal data and services

As people begin experimenting with AI at work, one of their first requests is often for API keys to company systems. This makes sense: AI isn’t much use at work if it doesn’t have access to the systems people use to do their jobs.

But handing over API keys to people and agents is dangerous and does not scale. Keys often provide broad, long-lived access that is difficult to constrain, share safely, and audit.

MCP gives agents a better way to use these systems. An MCP server can hold the credential and expose a defined set of tools instead of handing the key directly to the agent. But controlling which tools an agent can call is only the first step. MCP alone does not tell us which underlying resources an agent has observed. The agent can combine information across systems, send it somewhere less restricted, or expose it through apps and outputs to people who may not be allowed to see the original resources. Authorization has to account for where the data can go next.

Agents start with no access

Cloudflare Access controls who can enter Cloudflare OS. Inside, every agent and app starts with access to nothing. An agent can ask for access to a specific resource, which you can grant or deny. Generated code receives that resource as a typed binding:

env.PROJECT is a capability representing permission to use a specific resource under a specific policy. The credential remains completely isolated from the agent and any generated code.

Server code runs in a Dynamic Worker with global outbound networking disabled. Client code runs in a sandboxed frame in the browser. Neither can reach the Internet except through capabilities you explicitly provide.

Gatekeepers govern resources and actions

A Gatekeeper is a service-specific Worker that sits between Cloudflare OS and an external service. It understands the service’s API, its resources, and the operations that can be performed on them.

Giving an agent access to your entire GitHub account is likely too broad. A Gatekeeper can give it access to a single repository, allow it to read issues but not source code, mask particular fields, apply rate limits, and require approval before merging a pull request.

The agent and its apps see a small TypeScript API. The Gatekeeper handles OAuth, holds the credential, enforces policy, records what was read, and mediates anything with an externally visible side effect.

Policy follows what the agent has seen

Controlling the initial read is not enough. Take, for example, the case where an agent reads a sensitive table in a data warehouse and uses it to produce a live dashboard. Sharing the dashboard must not become a way to share the table with people who could not access it directly.

Cloudflare OS records every resource agents observe. These observations remain attached to the agent and its work. When another person tries to open the workspace, interact with the agent, or view what it produced, Gatekeepers verify that person's access to the observed resources.

The same observation log is used to inform policies that determine when agents can make external requests. A read of sensitive data can prevent the agent from writing data to certain sources, inviting new collaborators, handing work to another agent, or making an outbound request.

People using agents or building apps do not have to worry about making these mistakes. The platform can now be used to handle this.

A platform for building and sharing personal, modifiable apps

Most productivity suites give you a fixed set of applications: documents, spreadsheets, and presentations. In Cloudflare OS, each “file” can be its own application, written by an agent for one person, one project, or one team.

These are not prototypes that you have to export and deploy somewhere else. Each one is a full-stack application with client code, server code, an API, and durable state. Apps are private by default, but can be shared like documents.

Every app is a Worker

When you ask your workspace to build an app, the agent writes two parts:

  • Client code that renders the app’s UI in the browser
  • Server code that stores state and implements the app’s behavior

The server is loaded on demand as a Dynamic Worker and instantiated as a Durable Object Facet (both are features we built for this project). The facet gives the app its own SQLite database, separate from the Cloudflare OS runtime managing it. Dynamic Workers use lightweight V8 isolates, so every app can have its own isolated runtime without needing a dedicated server or container sitting around.

The browser client talks to the server using Cap’n Web, Cloudflare’s open source object-capability Remote Procedure Call (RPC) system. A server method can be called from the client like a normal JavaScript function:

The special part is that the agent can also call the same method.

So if you can build a tool to do a job yourself, agents can use your tool to do the job when you’re not there.

Share the app, or share how it was built

When you build an app in Cloudflare OS, you have two ways to share them:

  • Sharing your app itself lets other people collaborate in real time using the same state.
  • Sharing a blueprint of your app lets other people create their own copy of your app.

An app instantiated from a blueprint contains the original app’s code. But it does not contain its SQLite data, conversation history, credentials, or connected resources. Each new app starts with independent state and resources.

This means when you share apps with your team, they can modify them themselves with AI instead of filing a feature request and assigning you.

Use any model, and control what it costs

Cloudflare OS can be used with any model. Every inference call runs through Cloudflare AI Gateway, giving your organization one place to decide which models are available and which model should handle each job.

Not every task needs the most expensive model. You may not want to run the most expensive frontier model to summarize your unread emails every morning. AI Gateway gives you the control needed to make sure expensive models are only being used for the hardest work.

Every request is attributed to the person, team, or workspace that made it. Administrators can see where inference spend is going, set budgets and rate limits, and decide what happens when a limit is reached. 

Open source, so you can make it yours

Cloudflare OS is available today and is open source. Check out the cloudflare-os GitHub repository. You can deploy it into your own Cloudflare account and use your own Access policies, AI Gateway configuration, data, and integrations.

Our internal deployment reflects Cloudflare’s systems, terminology, policies, and ways of working. Yours should reflect your organization.

Cloudflare OS is designed so you can customize the interface, add internal Gatekeepers, and build organization-specific features without changing the core product.

We are releasing two repositories: the Cloudflare OS core and an example deployment based on how we run it internally at Cloudflare. The deployment repository consumes the core without patching it, providing a place for configuration, custom UI, internal integrations, analytics, and deployment pipelines.

Delivered together with our partners

The source code is only the starting point. The context, skills, workflows, internal systems, and policies are what make Cloudflare OS even more useful for your organization.

Cloudflare’s strategic partners, Presidio and Happy Cog, will work with you to customize Cloudflare OS around how your organization operates and roll it out across your workforce.

Partners can help you curate shared skills and institutional context, build custom interfaces, connect internal systems through Gatekeepers and MCP Server Portals, and configure security, model, and cost controls.

You get your own branded Cloudflare OS, connected to your systems, running on Cloudflare, and shaped around how your people actually work.

Get started

Cloudflare OS is available today on GitHub. You can explore the source code, try the demo, or deploy it into your own Cloudflare account in a few minutes using our starter repository.

We’re just getting started. We’re working on bringing Cloudflare OS to the Cloudflare dashboard as a fully managed product, adding containers for development workflows, and bringing workspaces into Slack and other chat tools.

If you’re interested in talking with our team, we would love to chat. Use this form to reach out!

How we built a software factory to drive Astro’s GitHub issue count to zero

Post Syndicated from Matthew Phillips original https://blog.cloudflare.com/astro-issue-triage/

Everyone is talking about software factories: the idea that AI agents can be assembled into a pipeline that produces working software on their own, the way a factory turns raw materials into finished goods. There’s endless debate over whether that’s actually possible, how far the automation can really go, and whether the “loops” people are demoing count for anything. Some have already written them off as a failure.

Running alongside that is a quieter, more worried conversation: open source maintainers are burning out. The AI boom has made it nearly free to generate issues, pull requests, and security reports, and enormously expensive for a maintainer to read through them all. The old ways of keeping a project healthy are buckling under the volume.

Everyone has a hot take on both topics. We think we have something rarer to offer: real results. For the past several months we’ve run an automated triage pipeline on the Astro repository. It reads incoming bug reports, reproduces them in sandboxes, diagnoses the root cause, and ships preview releases for the reporter to verify. The engine underneath it grew into Flue, an open framework for building this kind of agent automation, and it’s the same tool you could use to build your own.

It wasn’t an instant success. But through a lot of iteration, we’ve used it to bring our open issues down from over 200 to about 30, and we expect to hit zero sometime in the next month. That would be the first time this repository has seen zero open issues in its 5+ year history. 

We didn’t get there by declaring "issue bankruptcy," auto-closing cold tickets, or ignoring reports. We did it by automating issue triage with a team of isolated AI subagents running right inside GitHub Actions. Here’s the story of how we got there, and what you might take back to your own projects.

Starting with an agent skill

At the start of the year, we focused on automating one specific area of development: issue triage. As an open source project, manual issue triage can be one of the more time-consuming, least-rewarding parts of the job. A single issue can sometimes take hours just to reproduce, let alone fix. It was a natural (yet often overlooked) place for us to start our automation journey.

We began by developing an agent skill. This allowed us to develop and test the automation locally as maintainers, running a coding harness on our own machines. We could then run that same harness in a GitHub Action on our repo, and get total reuse of that exact same triage workflow skill.

The triage skill mirrors the exact steps we take during manual issue resolution:

  1. Reproduce: Clone the provided reproduction repository to verify the reported issue.
  2. Diagnose: Instrument the codebase and introduce logging to pinpoint the root cause of the bug.
  3. Verify: Review relevant test suites, code comments, and documentation to determine if the behavior is genuinely a bug or intended functionality.
  4. Fix: Convert the reproduction into failing unit tests, identify the appropriate solution via the architecture guide, and deploy the fix.

To prevent the frequent LLM bias toward forcing a solution when a bug might not actually exist, each phase is executed by an isolated subagent. These subagents pass information forward sequentially by compiling their discoveries into a report.md file.

Turning the skill into an automation

Following initial internal testing of the triage skill, our focus shifted toward building a fully automated pipeline. We specifically wanted to integrate this logic directly into a GitHub workflow, ensuring complete transparency so that anyone could easily audit the agent's sequential reasoning and operational steps.

As we wired it up, we realized the whole pipeline was really just a state machine driven by issue labels. Every new submission starts with the label triage needed, and once a user confirms a fix it moves to fix verified. Beyond those label transitions the pipeline holds no state of its own; it simply reads back through the issue’s existing comments to work out where a given issue is and what should happen next.

From there the flow runs on its own. When the agents land on a fix, the pipeline spins up a preview release with pkg.pr.new and posts everything back to the issue: a summary of what it found, the full logs, and instructions for installing the preview. The original reporter can then try the patch against their own project, and if they confirm it works, the automation opens a pull request linked to the issue.

From triage to a framework

As we built this out, we kept noticing that nothing about it was really specific to GitHub. Reacting to an event, running a sequence of isolated subagents, and separating their reasoning from the actions they’re allowed to take — it’s all just a workflow. One that could run just as well from a Slack message, a cron job, or a webhook as from a GitHub issue. Generalizing that realization into a runtime that works the same way regardless of where it’s deployed, or which model it’s driving, is what became Flue: an open, platform-agnostic framework for building durable agents and workflows.

Benefits of agent automation

When we first launched this automated system, we had shared concerns about its efficacy and the potential negative impacts it might have on our developer community. There was a valid fear that relying on automated bot responses might feel impersonal and create just one more disconnect between us as maintainers and our user base.

That did not happen. If anything, we talk to users more now, just in more useful places:

  • Engaging directly with our community members within Discord.
  • Actively participating in RFC discussions and addressing new feature requests.
  • Collaborating closely with contributors to help integrate their ideas into the framework.

Regarding the quality of automated patches, our core philosophy is that our AI agents should successfully resolve the vast majority of incoming issues. When an agent fails to identify a correct solution, we interpret that failure as an indicator of an underlying architectural or documentation issue within the codebase, pointing to one of three areas:

  • Opaque Abstractions: If an agent cannot interpret the boundaries between components, human developers likely struggle with the code structure as well.
  • Missing Documentation: Critical code segments lack explicit comments explaining the rationale behind their implementation.
  • Insufficient Testing: The repository suffers from a lack of comprehensive test coverage, particularly unit tests.

A clear example occurred with a series of related Hot Module Replacement (HMR) bugs. The triage bot repeatedly attempted to modify a specific if condition to resolve the issue. While this change fixed the targeted bug, it introduced regressions elsewhere due to a lack of test coverage for that specific condition. Once we added a descriptive comment explaining the exact logic governing that statement, the bot adapted and stopped attempting incorrect modifications in that area.

Every time we chase down one of these failures and add the missing comment, test, or clearer boundary, the bot gets noticeably better at that part of the codebase, and so does the next human who works on it.

Turning the workflow into a GitHub Action

Initially, our triage logic lived directly within the Astro monorepo. This coupling made iteration difficult; upgrading Flue or modifying the workflow felt like performing surgery on live infrastructure without a safety net. To solve this, we decoupled the logic into a standalone, testable repository: triagebot-action. This isolation allowed us to introduce automated testing and ensure stability before ever touching our primary codebase.

Today, this action powers issue management in Astro, and it has spread from there. Several other teams have picked it up, some using it directly, and others forking it to build their own automated "factories" tailored to their projects. That second path is really the point: triagebot-action is young and still actively evolving, so we’re sharing it less as a finished product and more as a working reference you can read, learn from, and adapt. 

The wiring for the action itself looks like this:

Or point your own agent at the repository and have it read through the setup, including adding the labels the state machine relies on.

Whichever route you take, the underlying idea matters more than our specific implementation: a sustainable feedback loop that frees maintainers to focus on the framework itself instead of administering a backlog. The code is open. Fork it, strip it down, or just borrow the parts that fit your project.

Want to build something like this? Dig into the code of the triagebot-action to see how it works, or fork it as a starting point for your own repository’s automation. And if you’re building agent-based infrastructure more seriously, that’s exactly what Flue is for: dive into the Flue framework to build your own. We’d love to see what you build. Come share your "factory" stories in the Astro Discord.

Don’t stop early: Case-folding source code at memory speed

Post Syndicated from Alexander Neubeck original https://github.blog/engineering/architecture-optimization/dont-stop-early-case-folding-source-code-at-memory-speed/


Suppose a user searches for café and your corpus contains CAFÉ, or they type straße and you’ve stored STRASSE. To make these count as matches, you need a canonical form that erases case distinctions, so that two strings which differ only in case compare equal. That form is case folding, and it shows up wherever text is matched rather than displayed: search engines, regex (?i) flags, case-insensitive usernames and hostnames.

It’s a basic operation, but at GitHub we run it a lot. Blackbird, GitHub’s code search engine, indexes over 180 million repositories—more than 480TB of source code. Every byte is case-folded before we extract ngrams and build the index, and for every potential query result, another (implicit or explicit) case folding operation is needed to locate matches. At that scale, the speed of even a basic operation starts to matter.

This post is about how we made it fast, and it starts somewhere counterintuitive: the biggest win in the ASCII fast path came from removing an optimization, not adding one. It turns out to be faster to sweep the whole buffer with no branches than to stop early at the first non-ASCII byte. We open-sourced the result as a Rust crate called casefold.

Folding is not lowercasing

It is tempting to reach for str::to_lowercase, but lowercasing and folding are different operations with different goals:

Lowercasing is for display, and it’s locale- and context-sensitive: Greek final sigma lowercases to ς at the end of a word and σ elsewhere, and Turkish I lowercases differently than English I. Case folding is for comparison, and it’s deliberately context-free and locale-independent. The point is a relation that stays stable and symmetric, so that if A folds to match B, B folds to match A in any locale. The Unicode Character Database ships an explicit CaseFolding.txt for exactly that.

The two operations diverge on real characters—ß, İ, final sigma—which is why lowercasing as a stand-in silently produces wrong matches. This crate implements only the simple (1-to-1) folds—statuses C and S in CaseFolding.txt—and not the multi-character “full” folds (ß → ss) or Turkic locale folds (the dotted İ). This isn’t an unusual choice: common tools and regex engines like ripgrep make the same restriction, and being consistent across tools is important.

The counterintuitive core: Don’t stop early

We deal mostly with source code, so the text we fold is overwhelmingly ASCII and making it run at memory speed is the single most important thing we can do. Everything else just has to keep the rare non-ASCII path from spoiling it.

The fold of an ASCII letter is trivial—A..=Z map to a..=z, everything else is unchanged—so the ASCII pass is really just “sweep the buffer, lowercase in place.” Ask any LLM for it and you might get something like this:

let bytes = s.as_bytes_mut(); 
for (i, b) in bytes.iter_mut().enumerate() { 
    if *b >= 0x80 { 
        break; // non-ASCII at index i: hand the rest to the Unicode path 
    } 
    if b.is_ascii_uppercase() { 
        *b += 32; // 'A'..='Z' → 'a'..='z' 
    } 
}

It looks ideal: do the cheap byte work, and the instant you hit a non-ASCII byte, break and let the “real” Unicode path take over: “only do the cheap work until you have to.” On an Apple M4 this runs at about 3 GiB/s. That sounds fine in isolation, but it is more than 15× short of “optimal” because of the if branches.

Let’s delete every branch, line by line:

  • if b >= 0x80 { break } → don’t stop at all. ORevery byte into an accumulator and test it once, after the loop: high_bit_acc |= *b. Same information (was there any non-ASCII byte?), zero branches in the body.
  • The A..=Z range test → make it arithmetic. b.wrapping_sub(b'A') < 26 is true exactly for A..=Z (any other byte wraps to ≥ 26), yielding a 0/1 mask with no branch.
  • The conditional write → fold the mask into the store.| (is_upper << 5)sets bit 5—turning an upper-case letter lower-case and being a no-op on everything else—the byte is always written, never branched on.

What’s left has no branch in its body and no early exit:

let mut high_bit_acc: u8 = 0; 
for b in &mut bytes { 
    high_bit_acc |= *b; // detect any non-ASCII byte 
    let is_upper = b.wrapping_sub(b'A') < 26; // branchless A..=Z test 
    *b |= u8::from(is_upper) << 5; // set bit 5 → lowercase, else no-op 
} 
if high_bit_acc & 0x80 == 0 { 
    return bytes; // pure ASCII: already folded in place, no second buffer 
}

A loop with no data-dependent control flow is trivially vectorizable: LLVM emits 16-byte-at-a-time NEON and the whole thing runs at > 45 GiB/s—essentially memory bandwidth. And we come out of the pass already knowing, from high_bit_acc, whether there’s any non-ASCII work left to do.

How much did each step matter? Measuring the cumulative ladder on pure ASCII (Apple M4, 5.7 KB buffer):

Version  Throughput  Vectorized? 
naive (break + branch test)  3.1 GiB/s  no (0 vector instrs) 
→ branchless test/write, keep break  2.6 GiB/s  no (0 vector instrs) 
→ drop the early-exit break  7.6 GiB/s  partially (25 vector instrs) 
→ branchless test + write (the loop)  >45 GiB/s  fully (41 vector instrs) 

The early-exit is what gates vectorization: keep the break but make the body perfectly branch-free and you still get zero vector instructions (~2.6 GiB/s); a data-dependent loop exit is enough on its own to keep the loop scalar. Only once the break is gone can the compiler vectorize. The final step—making the upper-case fold branchless—then turns a partially vectorized loop (which still compiles the conditional store to a compare-blend-masked-store, ~7.6 GiB/s) into the straight-line arithmetic that hits memory bandwidth.

Note: Branchless is a pessimization in scalar code. Look again at the table: making the body branchless while keeping the break (2.6 GiB/s) is actually slower than the naive branchy loop (3.1 GiB/s). The asm explains why. The branchy version only stores a byte when it actually changes one; its conditional strbis skipped for every lowercase letter, digit and space (the vast majority of real text), and the well-predicted branch that guards it is nearly free. The branchless version replaces that rarely taken store with an unconditional strbevery iteration, writing back all ~5,700 bytes instead of just the handful of upper-case ones. Extra write traffic for no benefit. Branchless-write only wins once the loop vectorizes, because then the store becomes a single 16-byte vector write regardless of content, and the per-byte cost disappears. The lesson: a branchless body is worth it only as the enabler for vectorization. On its own, in scalar code, it can cost you.

There’s also a middle ground, and it’s what standard libraries use. Instead of testing one byte at a time, [u8]::is_ascii scans a machine word at a time—on a 64-bit target it tests 16 bytes per iteration by OR-ing two u64 lanes and checking all their high bits with a single & 0x8080_8080_8080_8080 mask. You can build the ASCII fast path on top of that: chunk-scan to find the ASCII prefix, then run the branchless (vectorizable) convert over it. That keeps the early-exit ability—it still bails on the first non-ASCII block—while letting both halves go fast. The catch is that it reads the data twice (once to scan, once to convert), landing at about 23 GiB/s—roughly half of the single-pass branchless sweep, and ~7× the naive break loop. A solid, general-purpose default; just not the absolute ceiling when you control the whole loop and can fold detection and conversion into one branch-free pass.

Wouldn’t fusing the two passes be faster? It’s the obvious next thought: keep the chunked early-exit but convert each 16-byte block right after you’ve confirmed it’s ASCII, reading the data only once. Measured, it’s ~2.6× slower—8.7 GiB/s versus the two-pass 23. The inner block convert still vectorizes to a single 16-byte op, but now there’s a data-dependent early-exit branch every 16 bytes, and that branch pins the loop to one block at a time: the compiler doesn’t unroll or software-pipeline across blocks, and each iteration pays the full load→test→branch→convert→store latency with nothing to hide it behind. Split into two passes, each one is clean: the scan is a branch-light, store-free word scan that races through memory, and the convert is the fully-vectorized branch-free sweep at >45 GiB/s. Two fast, branch-free passes beat one branchy fused pass—even though the fused version touches the data half as many times. It’s the same lesson one more time: in the hot loop, the branch is the enemy.

Avoiding the heap

Forty-Five GiB/s also means doing zero unnecessary allocation. simple_fold takes the input String by value, owning the heap buffer it can mutate and return it. If the OR-accumulator’s high bit was clear, the input was pure ASCII already folded in place. We hand the same allocation straight back, no second buffer and no copy. Otherwise, we memchrto the first non-ASCII byte and scan the tail from there, leaving the output buffer unallocated (a null write cursor) until we hit a character that folds to different bytes. Text whose multibyte content never folds—CJK, Hangul, Kana, Arabic, Hebrew, symbols—also returns the original allocation untouched, never copying a byte.

Why a second buffer rather than rewriting in place like the ASCII pass? Because folding can make the string longer: almost every fold preserves the UTF-8 length or shrinks it, but two outliers grow—U+023A (Ⱥ) and U+023E (Ɀ) are 2 bytes each yet fold to 3-byte characters (ⱥ, ɀ). Once one appears, the output no longer fits in the input’s bytes, and we need somewhere new to write.

We allocate that buffer once, sized for the worst case, rather than growing it as more folds appear. Incremental reserve calls would mean re-checking capacity, occasionally reallocating, copying everything written so far, and juggling extra length/capacity bookkeeping; a single up-front allocation lets a raw write cursor run straight to the end with none of that. And since the cursor is nulluntil that first growing/changing fold, it doubles as the “have we allocated the extra buffer yet?” flag.

Sizing it needs a bound on growth, and those same two outliers give it: every 2 input bytes yield at most 3 output bytes, capping the output at 1.5× the input—exactly the capacity we reserve:

out = Vec::with_capacity(bytes.len() + bytes.len() / 2 + 4); 

After that the loop writes through a raw pointer with no capacity checks and calls set_len once at the end. Two more details keep it branch-light. The run of unchanged bytes between two folds is moved with a single copy_nonoverlapping rather than byte by byte. And each fold unconditionally writes all 4 bytes of a little-endian word before bumping the cursor by only the folded length (1–4)—dropping a branch on the output length from the hot path, with the + 4 in the reservation as the headroom that makes the final character’s over-store safe.

Making Unicode cheap too

When a character does fold, we still don’t want to fall off a cliff—decode UTF-8, hash lookup, re-encode. Unicode 16.0 has 1484 simple-fold mappings, but they’re a very sparse and very structured relation. Four observations shrink them to 1776 bytes and let the fold run without ever decoding a full character.

Even on the non-ASCII path, the overwhelming majority of characters do not fold. The hot operation isn’t really “fold this character,” it’s “does this character fold?” Almost always no. The table has to make that negative test as cheap as possible; the actual folding is the rare case on an already-rare path. That priority is what shapes the layout below—the page bitmap exists precisely so a non-folding character is rejected in a single bit test, straight from its leading UTF-8 bytes, without decoding or scanning anything.

This is exactly why a HashMap<u32, u32> is the wrong shape for the job, not just a bigger one. A hash map is optimized for the hit: it finds a present key in roughly one probe, and only spends extra work (more probes, full key comparison) when load factor or collisions bite. But our workload is dominated by misses—characters that aren’t in the table at all—and a miss is a hash map’s least favorite query: it still has to hash the key, jump to a bucket, and walk the probe sequence far enough to prove absence.

Foldable code points cluster into 64-code-point “pages”

Foldable code points bunch together. Slice the code space into 64-code-point “pages” and the ~1484 folds touch just 59 of ~1960 possible pages. A one-bit-per-page presence bitmap answers the negative test on its own: a clear bit is a definitive “no fold”—copy through, done—which is what makes fold-free scripts cheap. Only on a set bit do we consult a second structure, a cumulative-popcount side table that ranks the page (how many populated pages precede it) to find its slice of entries, storing nothing for the ~1900 empty pages.

let (word_idx, bit_idx, c_len) = if lead < 0xE0 { 
    (0usize, lead & 0x1F, 2usize) // 2-byte: word 0 
} else if lead < 0xF0 { 
    ((lead & 0x0F) as usize, bytes[read + 1] & 0x3F, 3) // 3-byte: word = nibble 
 
} else { 
    ( 
        (((lead & 0x07) as usize) << 6) | (bytes[read + 1] & 0x3F) as usize, 
        bytes[read + 2] & 0x3F, 
        4usize, 
    ) // 4-byte: merge 2 bytes 
}; 
// reject without decoding: clear bit ⇒ no fold 
if word_idx >= PAGE_BITMAP.len() || (PAGE_BITMAP[word_idx] >> bit_idx) & 1 == 0 { 
    read += c_len; 
    continue; 
} 

Because word_idxdepends only on the lead byte (and, for four-byte sequences, the first continuation byte), the bitmap load can be issued early.

Within a page, folds come in runs

A set page bit tells us something on this page folds, but not which code points or to what. The obvious encoding is one entry per foldable code point—but that is both bulky and slow to search: a page can hold dozens of folds, and we’d have to scan them all to find the one matching the current code point. The structure of the data rescues us again. Adjacent code points overwhelmingly share the same delta to their fold: A–Z all map +32, and Latin Extended is full of alternating runs like 0x0100, 0x0102, 0x0104, … where every second code point folds. Instead of per-code-point entries we store runs—start, end, stride, delta—and a 1-bit stride flag covers both the contiguous and the every-other case. This interval compression collapses the ~1484 individual folds into just 238 runs across the 59 pages (≈four per page), leaving the within-page search only a handful of entries to look at instead of dozens. This range-with-delta encoding (including the stride trick) is borrowed from Go’s unicode package, whose CaseRange records store a Lo/Hi range plus per-case deltas, with an UpperLower sentinel marking the alternating blocks. Runs are split at the page boundaries so a run never straddles two pages.

A run record is two clean bytes

With both endpoints inside one page they fit in 6 bits, split across two arrays: RUN_END_LOW[``i``] = end & 0x3F (the scan key) and RUN_START_STRIDE[``i``] = (start & 0x3F) | ((stride − 1) << 6) (read only on a hit). Because each key is one clean byte, the within-page search can go wide: rather than comparing cp & 0x3F against the runs one at a time, we load 8 end_low bytes into a single u64 and test all of them at once with one branchless SWAR step—(chunk | 0x80…80) − broadcast(low) & 0x80…80 sets the top bit of every lane whose key is ≥ cp & 0x3F. A single bit-scan of that mask (the keys are sorted, so the first set lane is the run we want) finds the slot. A page holds ~4 runs on average; that one 8-wide compare almost always resolves the entire search in a single step. One unlucky page does hold 30 runs, which puts the compare inside a short loop that strides eight keys at a time—but that loop trips at most a handful of times on exactly one page in all of Unicode, and never on the common ones. Either way: no per-run branch, and no code-point reconstruction anywhere.

/// Offset of the first run with `end_low >= low_v` in a page of `n` runs, 
/// or `n` if none. Scans 8 `end_low` bytes at a time via SWAR. 
#[inline] 
fn scan_end_low(lo: usize, n: usize, low_v: u8) -> usize { 
    const HIGH: u64 = 0x8080_8080_8080_8080; 
    const ONES: u64 = 0x0101_0101_0101_0101; 
    let bcast = (low_v as u64).wrapping_mul(ONES); 
    let mut base = 0; 
    while base < n { 
        // RUN_END_LOW is padded by 8 bytes so this read is always in bounds. 
        let chunk = u64::from_le_bytes( 
            RUN_END_LOW[lo + base..lo + base + 8] 
                .try_into() 
                .expect("8-byte slice"), 
        ); 
        // `(b | 0x80) - low_v` keeps its high bit iff `b >= low_v` (no 
        // cross-lane borrow). The first set lane is the first run `>= low_v`. 
        let ge = (chunk | HIGH).wrapping_sub(bcast) & HIGH; 
        if ge != 0 { 
            let j = base + (ge.trailing_zeros() / 8) as usize; 
            return if j < n { j } else { n }; 
        } 
        base += 8; 
    } 
    n 
} 

Folding is a little-endian byte addition

On a little-endian machine the folded character’s UTF-8 bytes, read as a u32, equal the source bytes (as a u32) plus a per-run constant. A parallel BYTE_DELTA[i] table then turns the whole fold into a masked load, one wrapping_add, and a 4-byte store:

let word = u32::from_le_bytes(next_four_bytes) & length_mask; // keep this char's bytes 
let folded = word.wrapping_add(BYTE_DELTA[i]); // the fold, as one byte add 
write_u32_le(dst, folded); // store all 4 bytes... 
dst += utf8_len(folded); // ...advance by the folded length

Both lengths in that snippet—the length_mask for the source character and the advance by the folded length for the destination—come from one more tiny trick. A UTF-8 sequence’s length is fixed by the top four bits of its lead byte, letting the 16 possible lengths pack one nibble each into a single 64-bit constant (0x4322_1111_1111_1111); the length is then a shift and a mask, (LEN_BITS >> (4 * (lead >> 4))) & 0xF—no if chain, no table memory, nothing for the predictor to get wrong. (A count leading ones—(!lead).leading_zeros()—would also work, since a lead byte carries one leading 1-bit per byte of the sequence.)

/// Number of bytes in the UTF-8 sequence whose lead byte is `lead`. 
#[inline] 
pub fn utf8_len(lead: u8) -> usize { 
    const UTF8_LEN_BY_LEAD: u64 = 0x4322_1111_1111_1111; 
    ((UTF8_LEN_BY_LEAD >> (4 * (lead >> 4))) & 0xF) as usize 
}

Because we advance by the folded length, this even handles length-changing folds—U+212A KELVIN SIGN (3 bytes) → k (1 byte), or U+023A Ⱥ (2 bytes) → U+2C65 ⱥ (3 bytes)—by writing fewer or more bytes than were read. That’s the part we believe is genuinely new: every other folder we looked at—ICU, Go’s unicode, Rust’s regex, CPython, glibc—decodes UTF-8 to a code point, applies the fold there, and re-encodes (even SIMD folders decode first). Doing the arithmetic in byte space skips both the decode and the encode, which is exactly why this path can outrun a hash map that already has the answer tabulated—the hash map still has to decode its key and encode its result. The byte-space arithmetic assumes the input is well-formed, shortest-form UTF-8—every code point encoded with the minimal number of bytes. Reading the source bytes as a u32and adding a per-run delta only lands on the correct folded encoding when the source is in canonical form; an overlong encoding (a code point padded into more bytes than necessary, e.g. / as 0xC0 0xAF) has a different byte pattern and would break thelength_mask and the delta arithmetic. This is not a real restriction in Rust—&str/String are guaranteed to hold valid UTF-8, which by definition rejects overlong sequences—but a caller feeding raw bytes from elsewhere must validate (or otherwise normalize) them first.

The ASCII shortcut in the tail loop

One more shortcut rounds out the tail loop. Remember the first pass already lowercased every ASCII byte, so when the scan meets an ASCII byte in the tail it advances a single byte and moves on—no page probe, no table touch at all. And it doesn’t copy that byte either: unmodified bytes (ASCII and non-folding multibyte alike) aren’t moved one at a time. The scan just keeps walking until it reaches a character that actually folds, then flushes the whole unchanged run between the last fold and this one with a single copy_nonoverlapping. Mixed text—CJK with ASCII spaces and punctuation, or code with the occasional accented identifier—therefore races through the ASCII filler and only consults the bitmap for genuine multibyte characters, copying in bulk rather than byte by byte.

Putting it together: the whole table

Component  Bytes 
PAGE_BITMAP (1 bit per 64-cp page)  248 
POPCNT_SAMPLES (cumulative popcount)  32 
PAGE_OFFSET (per populated page)  60 
RUN_END_LOW (scan key, end & 0x3F, +8 pad)  246 
RUN_START_STRIDE (start & 0x3F | stride)  238 
BYTE_DELTA (little-endian fold delta per run)  952 
Total  1776 

That’s 9.6 bits per fold entry, over half of it the BYTE_DELTA side table we trade for the decode-free path; the index + run records alone are ~4.4 bits/entry.

Next to the obvious alternatives, that 1776 bytes is an order of magnitude or more smaller—and unlike most of them, it never decodes a character:

Representation  Size
Naïve [(u32, u32); 1484]  ~11.6 KB 
regex-syntax’s case_folding_simple table  ~70 KB 
Go’s unicode.SimpleFold (orbit + ASCII + ranges)  ~7.3 KB 
A runtime HashMap<u32, u32>  ~17 KB 
This crate (paged bitmap + packed runs)  1776 B 

Where it lands against the alternatives

On the common case, ASCII, folding runs at memory bandwidth (>45 GiB/s), more than an order of magnitude ahead of other real folders and more than 50% faster than the (non-equivalent) str::to_lowercase function. To get a rough “upper bound” for the non-ASCII case, we measured the optimized Utf8 decoding + encoding round trip without performing any actual case folding using the simdutf crate. This experiment achieves consistently about 2GB/sec and is only about twice as fast than our solution for the worst case all-folding input. A naive hash map trails everything on all workloads.

The three columns are real case folders that produce identical output: simple_fold (this crate), simd_normalizer (the simd-normalizer crate), and HashMap (naive CaseFolding.txt lookup). The workload rows are chosen to simulate different scenarios from typical to worst case:

Workload (input size)  simple_fold  simd_normalizer  HashMap (byte path) 
Pure ASCII (5.7 KB)  >45 GiB/s  1.21 GiB/s  213 MiB/s 
Chinese/Japanese/Korean, no folds (8.1 KB)  2.95 GiB/s  1.97 GiB/s  558 MiB/s 
Symbols / Myanmar, no folds (9.0 KB)  2.96 GiB/s  1.56 GiB/s  410 MiB/s 
Worst case: Latin/Greek/Cyrillic (Unicode U+0000–U+FFFF), all folding (8.8 KB)  869 MiB/s  922 MiB/s  334 MiB/s 
Length-changing folds (1.7 KB)  1.26 GiB/s  716 MiB/s  233 MiB/s 

Treat the absolute figures as illustrative, not portable: the whole design leans on auto-vectorization, SWAR, and little-endian byte arithmetic, so the numbers—and even the ratios between rows—can shift substantially on a different microarchitecture (a wider or narrower vector unit, different memory bandwidth, a big-endian target, x86 vs ARM).

More details can be found in the performance section of the README.

Take this with you

Case folding is about as basic as text operations get, which is exactly why it was worth the effort: we run it across every byte we index. The wins came from two ideas that both cut against instinct—sweep the whole buffer branch-free instead of stopping early, and do the fold as byte-space arithmetic instead of decoding to a code point. Together they let the common case run at memory bandwidth and the rare fold run without a decode, in a table small enough (1776 bytes) to stay resident. The decode-free byte-space fold is the piece we believe is genuinely new; it’s why this path can beat a hash map that already has the answer.

There’s surely more to find here, and we’d like to see it. The crate is casefold; the generated table and full design notes live alongside the source.

The post Don’t stop early: Case-folding source code at memory speed appeared first on The GitHub Blog.

Dogfooding at scale: migrating cdnjs to Cloudflare’s Developer Platform

Post Syndicated from Simona Badoiu original https://blog.cloudflare.com/cdnjs-dev-platform-migration/

As of June 23, 2026, cdnjs, one of the Internet's busiest open-source CDNs, is running exclusively on Cloudflare’s Developer Platform. Along the way, cdnjs surfaced limits in the platform, and the platform grew to meet them.

cdnjs is a free, open-source content delivery network for JavaScript and CSS libraries. Instead of using a bundler or self-hosting jQuery, Bootstrap, or Lodash, you drop a <script> tag pointing to cdnjs.cloudflare.com and the library loads from Cloudflare's edge, instantly, anywhere in the world, with no signup, no API keys, and no rate limits. It's the infrastructure behind a significant portion of “intro to JavaScript” tutorials, CodePen demos, and Stack Overflow answers.

Community-driven, cdnjs is used on roughly 12% of all websites, a 48.3% share of the JavaScript CDN market. It serves an average of 108,000 requests per second, 9 billion per day, across more than 330 Cloudflare data centers, with a 98.6% cache hit rate. Pretty cool, Internet!

In 2011, when bundlers were exotic, npm was barely a year old, and "just drop a <script> tag" was how the web shipped JavaScript, Ryan Kirkman and Thomas Davis built cdnjs as a free, community-run mirror of every popular open-source library. 

Cloudflare stepped in to host it free of charge months later, and took over project maintenance in 2019. Back then, Cloudflare didn't have a mature Developer Platform that could fully sustain the entire cdnjs ecosystem. Fifteen years and a lot of building blocks later, the platform is mature enough to run cdnjs end to end, on Workers, Workflows, D1, Queues, Workers Cache, R2, KV, and Containers.

Why cdnjs has 9 billion requests a day 

The web has changed beyond recognition from those days. We have ES Modules (ESM), the standardized import / export syntax browsers understand natively. We have import maps, Vite, Bun, Turbopack. We have AI assistants that scaffold entire apps in seconds. Bundlers are everywhere. So why does a CDN for <script> tags still serve 9 billion requests a day?

One reason: LLMs love cdnjs. When ChatGPT, Claude, or Cursor scaffold a quick HTML demo, they reach for cdnjs because their training data is full of it. There have been 15 years of blog posts, GitHub READMEs, tutorial sites, and Q&A threads pointing to cdnjs.cloudflare.com. The URL pattern is consistent and versions are immutable — exactly the kind of dependency a model can produce reliably without hallucinating.

Every file on cdnjs has an SRI hash (we're still working on ensuring all the existing stored hashes match reality due to bugs in the old system), mirrors are auditable, and the whole project is open source. In a world increasingly worried about supply-chain attacks, an immutable, hash-verified mirror of well-known libraries is indispensable.

And it's free, forever, for everyone. No API keys. No rate limits. No "sign up to continue." That's a rare thing on today's Internet, and it's worth protecting.

Why we migrated

We didn't migrate because cdnjs was slow. We migrated because we want to keep improving it.

The previous architecture served users well: 98% cache hit, billions of requests, no outages. But internally, shipping anything new or fixing existing issues in how packages were processed was getting harder. Making a change meant coordinating deployments across GCP Functions, a VM, and Cloudflare. Observability was painful too.

The pain points  

In 2020, we migrated cdnjs to serverless, moving file serving onto Cloudflare Workers and KV, with a bare-metal origin as fallback. That change dramatically improved resilience and scalability, and let us pre-compress every asset with Brotli and gzip for smaller, faster responses, but only on the serving side.

The publishing side — the pipeline that watches npm and GitHub for new library versions, downloads them, processes them, and writes the results so cdnjs can serve them — stayed on Google Cloud Platform (GCP). At the time, Cloudflare Workers were designed for fast, short-lived HTTP requests; they didn't yet have the building blocks for a long-running, multi-step pipeline that fetches large tarballs, runs CPU-heavy compression, and orchestrates work over hours. Workflows, Queues, Durable Objects, R2, and Containers didn't exist yet.

So we built the publishing bot on what was available: a chain of GCP Functions, a VM running git-sync, and a GitHub repository as the source of truth. It worked, but six years later, that architecture was showing its age. Here's a diagram of the previous architecture:

 

The architecture had five pain points. The one that hurt most was observability: debugging meant stitching logs together by hand. We'll start there.

  1. No shared trace
    A single package update could pass through Cloud Functions, Google Cloud Storage (GCS) object events, Pub/Sub topics, a git-sync VM, and Workers KV before a file reached a user. None of those systems shared a correlation ID. GCP Logging held one half of the story, Cloudflare Logpush held the other, and the two had no common key to join on.

    The problem wasn't outright failure, it was partial success. A version that processed cleanly, wrote to KV, and then silently failed to land in the GitHub repo would serve fine for weeks until someone noticed the two stores had diverged. There was no alert for that. There couldn't be, because nothing in the system knew the full pipeline state.

  2. Split-brain storage
    Files lived in two places at once: Workers KV at the edge (with a bare-metal origin as fallback) and a GitHub repository as the source of truth. The ingestion pipeline wrote to both at the end of every run. Neither was authoritative, and when they drifted, there was no clean way to reconcile them.
  3. Pipeline glued together with object events
    The ingestion pipeline was a chain of small GCP Cloud Functions, each doing one step and handing off to the next through shared storage. One function fetched the package's release archive from npm and dropped it in a bucket. The bucket firing a "new file" event triggered the next function, which unpacked it and wrote the results somewhere else, triggering the next, and so on. Storage was doing double duty as a message queue, with no dead-letter queue, no backlog visibility, and no clean replay when a step failed.
  4. 26 functions for 26 letters
    Just checking npm for updates required 26 Cloud Functions, one per letter of the alphabet. Each shard had its own deployment and its own logs, and the only way to know if the fleet was healthy was to check all 26.
  5. The GitHub repo GitHub couldn't serve
    A separate VM ran git-sync, mirroring every processed file into cdnjs/cdnjs. Years of releases pushed it past 1.1TB of packed storage, large enough that GitHub's own archive service refused to generate tarballs or zip downloads for it. Forking became impractical, clones were slow, and the .gitignore had grown to 274 hand-curated entries blocking broken or weirdly-versioned releases. It was a documented graveyard of everything the pipeline couldn't reasonably reject upstream.

    A genuine thank you to the GitHub team for hosting this giant for over a decade. They bore with us through years of storage growth, and the project wouldn't have survived without them.

    A quieter benefit that came with the migration is having fewer moving parts to secure. Cloud Functions, a git-sync VM, container images, GCS buckets, service-account keys — every one of those was a thing to secure, patch, and audit. Retiring the pipeline closed all of the recently opened cdnjs vulnerabilities.

How we re-built it

The new cdnjs architecture runs entirely on Cloudflare’s Developer Platform.

R2 is the single source of truth for file content. It has no practical size limit, so the files that couldn't fit in KV before, like source maps, big bundles, and font packs, now live alongside everything else. As a bonus, the S3 API makes the entire cdnjs catalog accessible to any S3 client. Maintain a mirror? Open an issue on the cdnjs repository and we'll set you up with read-only credentials.

KV stores only metadata now: package info, version lists, SRI hashes. KV is built for high read volume with infrequent writes, which is exactly the shape of metadata access.

In front of the Worker sits Workers Cache, a tiered cache Cloudflare launched this year. Before, we relied on a separate internal caching layer between the edge and the Worker. That layer is gone now, replaced by one owned by the Developer Platform, the same platform that runs the rest of cdnjs. One less moving part!

The new architecture also extends a long-standing partnership. DigitalOcean has hosted the cdnjs website for years as a sponsor; now it hosts the storage too. Every file published to R2 is mirrored to DigitalOcean Spaces: architecturally a disaster-recovery copy, operationally also a live fallback. The serving worker reads through to it whenever R2 can’t return a file. The chain is cache → R2 → DigitalOcean, so R2 having a bad day doesn't take cdnjs down. A Cloudflare-hosted origin still sits in the chain during the transition, but it will retire once the GitHub backfill lands in R2.

The ingestion pipeline is built on Cloudflare Workflows. Every ten minutes, a cron job triggers PackageUpdatesWorkflow, which checks npm and GitHub for new versions. For each new version found, it spawns a DownloadPackageWorkflow that fetches the tarball into R2, then a ProcessingWorkflow per file that extracts, minifies, and compresses. Finally, PublishingWorkflow writes the results to R2 and KV and updates the Algolia search index.

Because Workflows provides durable execution, the state of each step is preserved. If anything fails — a network timeout, a compression error — the workflow resumes from the last successful step.

The trickier piece is how we glue Workflows to the external compression container. We pre-compress text-based files to streamline the delivery process. But compression is too CPU-intensive for a Worker, so we hand it off to Cloudflare Containers, wait for compression to complete, and then pick up where we left off.

The pipeline has two kinds of waiting:

Per file: Each ProcessingWorkflow writes the uncompressed file to an R2 bucket, sends a job to a Queue, and hibernates. A Rust compression service running in the container picks it up, compresses it, and writes the result to another bucket. An R2 event notification wakes the workflow up so it can continue.

Per package: The parent workflow needs to wait for all its file children before moving on to publish. A package with thousands of files means thousands of children running in parallel. We use a small Durable Object as a counter: a parent increments on each child it spawns, children decrement when they finish. The parent wakes up when the counter reaches zero.

An overview of the new architecture, with R2 as the source of truth and Workflows running the pipeline:

Pushing the limits

Designing the new architecture was one challenge. Migrating the existing catalog into it, without disturbing a single file already in the wild, was another.

We'd actually tried this once before and had to roll back. The plan, back then, was to re-process old packages and write the results directly to R2, but the regenerated files didn't byte-match what KV had been serving. Minifiers and compressors aren't fully deterministic across versions, so the new outputs were correct but had different SRI hashes. For a CDN where users pin those hashes in their HTML, that's a serving break. So we rolled back, and now, we migrated the existing content from KV to R2 as-is instead of regenerating it.

That decision shifted the problem from "re-process millions of files" to "copy millions of files between accounts, without missing any." And that's where we ran into the Workers subrequest limit, capped at 1,000 per invocation on paid plans. A package with thousands of files would burn through it in one go. Parallelizing didn't help, since every Worker hits the same ceiling. So we sharded the migration by package name and fanned the work out across many invocations via Queues, whose at-least-once delivery guarantee meant no package could silently fall out of the migration.

We hit two platform limits during the migration: 1,000 subrequests per Worker invocation and 1,024 steps per Workflow. Instead of just working around them, we asked the Workers and Workflows teams to raise them — which they did. Subrequests now go up to 10 million on paid plans; Workflows now default to 10,000 steps, configurable to 25,000.

The cdnjs pipeline runs on the same building blocks anyone can use: Workers, Workflows, R2, KV, Queues, Containers, and Durable Objects. The limits we hit are limits we lifted for everyone. If the Cloudflare Developer Platform can serve 9 billion requests a day and publish packages with hundreds of thousands of compressed, minified files, it can probably run whatever you're building.

What’s next

There's an obvious next question hiding in all of this: could cdnjs also serve modern, browser-native ES modules? The same packages, transformed on publish, ready to import without a bundler. The architecture doesn't rule it out. The Workflows-plus-Containers pattern that pre-compresses files today would work just as well for transforming them. We're not committing to it, but it's the kind of thing that's now possible to consider, which wasn't true a year ago.

git commit -m "with love" --author="cdnjs team"

We follow every open issue on GitHub and we want your feedback. Don't hesitate to contribute and help make the Internet better for everyone.

Tame Dependabot: Group your updates, slow the cadence, keep security fast

Post Syndicated from Bruno Borges original https://github.blog/security/supply-chain-security/tame-dependabot-group-your-updates-slow-the-cadence-keep-security-fast/


If you maintain an active repository, you know the feeling. You open your notifications on a Monday morning and there they are: five, 10, sometimes a dozen Dependabot pull requests, each bumping a single dependency by a single patch version. Individually, every one of them is helpful. Collectively, they’re noise. And noise is how important updates get ignored.

We looked at Microsoft’s GCToolkit, an open source Java library for analyzing garbage collection logs. As of July 2026, a git log of the repository showed that 92 of its 578 commits, roughly one in six, were Dependabot version bumps, with 61 in the previous 12 months alone, sometimes several in a single day. That’s a lot of review, merge, and CI cycles spent on routine maintenance.

The good news: Dependabot already ships with the features to fix this. In a recent pull request, the project changed its dependabot.yml in three small but meaningful ways, turning a daily drip of single-dependency pull requests into a predictable, grouped, monthly batch per ecosystem. Here’s what changed, why it works, and how to apply the same pattern to your own repositories, following the GCToolkit example.

The problem: Good defaults, wrong cadence

Here’s what GCToolkit’s configuration looked like before:

version: 2
updates:
- package-ecosystem: github-actions
  directory: "/"
  schedule:
    interval: daily
  open-pull-requests-limit: 10

This is a common starting point, but the daily interval here was a deliberate choice, not a default: schedule.interval is required, and GitHub’s suggested starter template uses weekly. Two things make this configuration noisy:

  • interval: daily tells Dependabot to check for updates every weekday (Monday through Friday). For a repository that references a handful of GitHub Actions, that can mean new pull requests landing on any weekday.
  • No grouping means every dependency gets its own pull request. Ten available updates equals 10 pull requests, 10 CI runs, and 10 review notifications.

The open-pull-requests-limit: 10 line is a symptom, not a cure: it caps the flood at 10 open pull requests, but it doesn’t stop the flood.

The fix: Three changes that compound

Here’s the configuration after the change:

version: 2
updates:
  - package-ecosystem: "github-actions"
    directory: "/"
    schedule:
      interval: "monthly"
    groups:
      monthly-batch:
        patterns:
          - "*"

  - package-ecosystem: "maven"
    directory: "/"
    schedule:
      interval: "monthly"
    groups:
      monthly-batch:
        patterns:
          - "*"

Three things are happening here, and they build on each other.

1. Group everything into a single pull request

The groups block is the heart of this change:

groups:
  monthly-batch:
    patterns:
      - "*"

A Dependabot group bundles multiple dependency updates into one pull request. The name (monthly-batch) is yours to choose. It shows up in the pull request title and branch name. The patterns list decides which dependencies belong to the group, and "*" is a wildcard that matches all of them.

So instead of 10 pull requests, you get one pull request titled something like “Bump the monthly-batch group with 10 updates.” One branch. One CI run. One review. If the whole batch is green, you merge once and you’re done. If something breaks, it’s contained in a single, reviewable place.

For larger projects, you don’t have to lump everything together. You can define multiple named groups with more specific patterns. For example, you could keep all your testing libraries in one group and your production dependencies in another, so related updates travel together and unrelated ones stay separate.

Grouping keeps getting more capable, too. In a February 2026 update, Dependabot gained the ability to group updates for the same dependency across multiple directories into a single pull request. That’s aimed squarely at monorepos: if one library is pinned in a dozen services, a single bump used to open a dozen near-identical pull requests, one per directory. Now you can point the directories key (note the plural) at a list of paths, or a glob like /apps/*, and let your group collapse all of them into one:

- package-ecosystem: "npm"
  directories:
    - "/apps/*"
  schedule:
    interval: "monthly"
  groups:
    monthly-batch:
      group-by: dependency-name
      patterns:Expand comment
        - "*"

That’s the same monthly-batch group as before, now spanning every service in the repository instead of a single directory. For the full set of options, see the Dependabot options reference.

2. Slow the cadence from daily to monthly

schedule:
  interval: "monthly"

Switching from daily to monthly changes the rhythm from “whenever anything changes” to “once, on a schedule you can plan around.” Combined with grouping, this is the real noise reduction: Dependabot now opens one batched pull request per ecosystem, per month, instead of a steady trickle all month long.

Monthly is the right call for a mature library where dependencies are stable and updates are rarely urgent. If you want something in between, weekly is also available, and you can pin the exact day and time with schedule.day and schedule.time.

3. Cover every ecosystem you actually use

The original config only requested version updates for github-actions. But GCToolkit is a Java project built with Maven, so its application dependencies weren’t receiving Dependabot version updates. The updated config adds a second updates entry:

- package-ecosystem: "maven"
  directory: "/"

This is an easy one to miss. Reducing noise is only half the win; the other half is making sure Dependabot is watching the dependencies that matter most. Each ecosystem gets its own schedule and its own group, so your Actions updates and your Maven updates arrive as two clean, separate batches.

But what about security updates?

This is the question every maintainer should ask before slowing anything down, and it’s where the design really shines: by default, the groups and schedule you set here shape your version updates, not your security fixes.

Dependabot security updates are raised as soon as a vulnerability with a fix is disclosed, independent of your schedule and separate from your version-update groups. So a monthly batch cadence for routine bumps doesn’t delay a critical patch. (You can batch security fixes on purpose with a group scoped to applies-to: security-updates, but even then they’re triggered by disclosures, not by your version-update schedule.)

One caveat: this safety net only exists if Dependabot security updates are actually turned on for the repository, which also requires the dependency graph and Dependabot alerts to be enabled. Confirm those are on before you rely on a slower version-update cadence. Do that, and you get the best of both worlds: quiet, predictable maintenance for the routine stuff, and immediate action when a real vulnerability lands.

That separation is what makes “slow down Dependabot” a safe recommendation rather than a risky one.

A new safety net: default package cooldown

There’s one more piece of noise reduction that landed recently, and it happens automatically. Dependabot now waits until a new release has been on its registry for at least three days before opening a version-update pull request. This cooldown is the default and requires no configuration.

Why wait? A brand-new release is one of the most common entry points for a supply chain attack. A compromised or simply broken version can reach your dependency updates before maintainers and the wider community have caught the problem. A short delay gives that signal time to surface, so you’re far less likely to merge a bad release the moment it ships.

Two things worth knowing:

  • It only applies to version updates. Security updates still open immediately, so critical fixes are never held back by the cooldown.
  • You stay in control. Use the cooldown option in your .github/dependabot.yml to widen or shorten the window, tune it per semantic-versioning level, or opt out entirely:
- package-ecosystem: "maven"
  directory: "/"
  schedule:
    interval: "monthly"
  cooldown:
    default-days: 7
  groups:
    monthly-batch:
      patterns:
        - "*"

Pair cooldown with grouping and a monthly cadence and the effect compounds: fewer pull requests, and the ones you do get have had a few days to prove they’re safe to merge.

How to apply this to your own repositories

You can adopt this pattern in a few minutes:

  1. Open (or create) .github/dependabot.yml in the default branch of your repository.
  2. For each package-ecosystem you depend on, set schedule.interval to weekly or monthly.
  3. Add a groups block with a single wildcard group (patterns: ["*"]) to batch updates into one pull request per ecosystem.
  4. Make sure every ecosystem you actually ship with is listed: not just github-actions, but maven, npm, pip, gomod, docker, and so on.
  5. Commit, and let the next scheduled run produce a single, grouped pull request.

A few tips as you tune it:

  • Start broad, then split. A single wildcard group is the simplest starting point. If you later find you want, say, patch-level and major-version updates handled differently, break the wildcard into more targeted named groups.
  • Don’t fold security fixes into this cadence. Dependabot security updates are triggered by vulnerability disclosures, not your version-update schedule, so a monthly cadence never delays them. You can even group them with applies-to: security-updates without slowing them down.
  • Lean on cooldown. The three-day default already shields you from brand-new bad releases; bump cooldown.default-days higher if you want an even wider safety margin on version updates.
  • Right-size the interval. Fast-moving apps may prefer weekly; stable libraries do fine on monthly.
  • Consolidate monorepo directories. If the same dependency lives in many directories, list them under directories and set group-by: dependency-name in the group so a single bump produces one pull request instead of one per directory.

The takeaway

Dependency updates are one of those chores that’s easy to automate and then easy to start ignoring, which defeats the purpose. The fix isn’t to turn Dependabot off or to merge pull requests without looking. It’s to shape its output so that the routine work is quiet and batched, and the urgent work still cuts through.

GCToolkit did it with about a dozen lines of YAML: group everything, slow the cadence to monthly, and make sure every ecosystem is covered. Add the new default cooldown on top, and even that monthly batch has had a few days to prove itself before it reaches you. The result is fewer pull requests, fewer CI runs, and, most importantly, a review queue where the updates that matter don’t get lost in the ones that don’t.

Further reading: once the routine pull request noise is under control, the harder question is which security alerts to fix first. Our earlier post, Cutting through the noise: How to prioritize Dependabot alerts, walks through using EPSS scores and repository properties to turn an overwhelming alert list into a clear, risk-ranked queue.

Configure your own Dependabot updates >

The post Tame Dependabot: Group your updates, slow the cadence, keep security fast appeared first on The GitHub Blog.

We’re open sourcing our privacy proxy CLI

Post Syndicated from Hannah Wang original https://blog.cloudflare.com/open-sourcing-our-privacy-proxy-cli/

Debugging privacy-preserving protocols is hard. Oblivious HTTP has several different steps across four different parties, not to mention binary HTTP encoding and details spread across many draft RFCs. We've taken what we've learned operating protocols like Oblivious HTTP at the scale of millions of requests per second, and wrapped it up in a nice, clean CLI tool — that we are open sourcing today.

We call it our privacy-client, or pvcli. We’re releasing it under the Apache-2.0 License, and it is open for contributions.

Here’s a single line of code that executes a full Oblivious HTTP request with a relay, gateway and origin. Don't worry if you don't know what that means, we'll cover it below.

We’ll explain why we built this tool, and show just how handy it can be.

Why privacy protocols can be hard to debug

Let’s take a closer look at our motivation for creating pvcli. Over time, the Privacy team’s product suite and customer base grew. We added products like Privacy Proxy and Privacy Gateway, which power Apple’s Private Relay, Microsoft’s Edge Secure Network VPN, Flo Health’s Anonymous Mode, and more. With it came an increasing amount of special customer requirements, domain knowledge, and complexity. As a result, we saw increased friction in development and incident response.

To see this in action, let’s look at how one of our products implements Oblivious HTTP, also known as OHTTP. First, a quick primer. OHTTP provides users with a privacy guarantee: no one can know both who made a request and what they’re requesting. To achieve this, OHTTP requires two servers, a relay and a gateway, operated by two non-colluding parties. 

Below is a sequence diagram of OHTTP, where our customer owns the relay and Cloudflare owns the gateway. At a high level, OHTTP can be broken down into these steps:

  1. Client gets public key from the gateway.
  2. Client encrypts the request and sends it to the relay.
  3. Relay removes “who” the client is from the encrypted request, and sends it to the gateway.
  4. Gateway decrypts the request, and sends it to the target.
  5. Target processes the request, and sends a response to the gateway.
  6. Gateway encrypts the response, and sends it to the relay.
  7. Relay sends the encrypted response to the client.
  8. Client decrypts, and gets the plaintext response.

It involves quite a bit of back-and-forth, as you can see: 

Each step is a potential point of failure that we have to consider while debugging!

In particular, we saw certain kinds of problems when debugging OHTTP.

  • Customers asked for ways to test the live system from their end, and we often wrote one-off, custom clients for our customers specific deployments.
  • Figuring out which step caused an issue was time-consuming. Was the root cause a bug in our system or our customer’s system?
  • Examining raw bits was tedious and highly prone to human error. OHTTP builds on binary HTTP, which is a binary encoded HTTP request. Anytime we needed to check the binary encoding, we were painstakingly going through raw bits.

As a result, we decided to place all of our privacy protocols in one tool. It has a clean interface that’s already familiar, displays every single step of the protocol in order, and is flexible enough to support new protocols and architectures.

To see the difference this makes, let’s see an OHTTP debugging scenario — before and after pvcli.

Debugging without pvcli

Say we operate an OHTTP relay that sits in front of a customer's gateway. The customer has asked us to do an end-to-end test with a request:

Recall the OHTTP steps from earlier. The first step is to fetch the public key from the gateway. We use curl to fetch it from the customer gateway and get this back:

That’s a big binary string in hex. To make sense of it, we look at OHTTP RFC 9458 §3 and parse it manually:

  • 0029 is 41 in decimal, telling us this public key entry has 41 bytes associated with it.
  • 55 is the public key ID.
  • 0020 identifies the asymmetric encryption method we can use. In this case, DHKEM(X25519, HKDF-SHA256).
  • b9bb667e2230dc01c6d6cc047f94a1083beb185c63e50ec09f7692a5a0832540 is the public key.
  • 0004 tells us there are 4 bytes of symmetric cryptographic IDs that follow.
  • 0001 and 0001 identify the symmetric encryption methods we can use: HKDF-SHA256 and AES-128-GCM.

We repeat this process for however many public keys are in the binary string.

Next, we convert our original HTTP request into binary HTTP, referencing RFC 9292. We manually craft the binary with the help of some bespoke scripts:

We verify each field:

  • 02 means it's an indeterminate-length request
  • 04504f5354 is POST
  • 056874747073 is https
  • 117461726765742e6f687474702e696e666f is target.ohttp.info
  • and so on

Finally, we form a wrapper HTTP request that will hold our OHTTP request. To do so, we spend some more time writing another makeshift script that encrypts the binary HTTP request in the manner OHTTP specifies, using the public key from earlier. We create a header, which is the concatenation of public key ID, asymmetric encryption method ID, and symmetric encryption method IDs. Then, we concatenate header and encrypted binary HTTP request, resulting in:

We put those bytes into the body of our wrapper HTTP request, and send it to our relay. We get back a response.

What does that mean? We reach out to the customer to ask if they can share logs from their gateway. In the meantime, we double-check the bits we've crafted. The decoded public keys look fine. The binary HTTP request… Oh! We see:

BHTTP is length prefixed. That means we specify a length (0x0a is 10 in decimal), and then 10 bytes follow. But here, 11 bytes follow. There is an extra 20 before the 00. 20 represents a space character, so we must have accidentally added that when building the body. We remove the extra character, resend, and it works!

Debugging with pvcli

With pvcli, all of that is now a single command:

It handles all the binary parsing and encrypting for us, and prints logs in case we want to dive deeper:

What used to be a fragile process — involving manipulating bits, gluing together scripts, and referencing long RFCs — is now one command.

What pvcli can do

To install:

pvcli takes a lot of inspiration from curl. We designed it with the “principle of least surprise” in mind. As a result, a lot of the arguments are the same as curl’s! Try a quick GET request to our cdn-cgi endpoint:

If you’re curious about what is happening under the hood, you can use -v to get detailed logs:

Now, about that OHTTP command from earlier: you use –ohttp to tell pvcli to construct an OHTTP request. You pass in the relay as the –first-hop and the gateway as the –proxy. The target will be an echo server, so you can see what the target would see. In this command, we filled in the arguments with a relay, gateway, and target from ohttp.info.

Try running the command yourself!

We’ve encountered many cases where we wanted to pass headers to the relay, rather than the target. You are able to do that with --first-hop-header:

Similarly, we’ve also had cases where we wanted to authenticate to the relay with mTLS, to ensure that the correct client is talking with the correct relay. To do that, you can use –first-hop-client and --first-hop-key.

And it just works. Need to test a full Oblivious HTTP request with a relay, a gateway, arbitrary headers, and mTLS? Or perhaps only request through a gateway? Or maybe you just want to see the OHTTP key configuration? pvcli can do it with a single command, debugging included.

Why build our own tool?

There are some great tools for OHTTP that already exist. Martin Thomson’s Rust implementation and Chris Wood’s Go implementation were incredibly helpful when we built out our original OHTTP implementation a few years ago. But pvcli is not only focused on OHTTP. We’re looking to add as many privacy-preserving protocols as we can to the tool. So while there are other OSS tools out there for debugging OHTTP, nothing combines OHTTP, CONNECT proxying, MASQUE and Privacy Pass (coming soon) all in one place.

Contribute to pvcli

Oblivious HTTP is an amazing protocol, and we would love to see you use it. We hope that this tool helps people debug OHTTP and write their own OHTTP implementations. 

We are accepting contributions! To get started, clone the repo at https://github.com/cloudflareresearch/pvcli, and submit a pull request. 

If you're looking for ways to contribute, here are some things on our to-do list. For MASQUE, we plan to add support for proxying TCP over HTTP/3, and UDP and/or IP over HTTP/2 and HTTP/3. For OHTTP, we plan to support post-quantum cryptography, add timing/latency information, support Chunked OHTTP, and improve logging.

Contact us if you are interested in using Cloudflare’s OHTTP Relays and Gateways.