Building a certificate authority for the whole Internet

Post Syndicated from Steve Goldsmith original https://blog.cloudflare.com/cloudflare-certificate-authority/

Twelve years ago, during Birthday Week 2014, we turned on Universal SSL and nearly doubled the number of encrypted sites on the web overnight, giving free TLS to every site behind Cloudflare, including the ones that never paid us a cent. Encryption stopped being an expensive, time-intensive undertaking and instead became the default.

For Birthday Week this year, we are taking the next step on that path. For more than a decade we have been one of the largest consumers of publicly trusted certificates on the Internet, and have never issued a single one ourselves. That is changing. Cloudflare is announcing our intent to become a public certificate authority (CA).

Today we are announcing the first concrete milestones in that effort: We have applied for inclusion in the Chrome, Apple, Microsoft, and Mozilla root programs, and we have signed a definitive agreement to acquire an established, broadly trusted root from GlobalSign, so that we can offer certificates with the widest possible device reach the day we begin issuing. We’re also announcing our plans to be one of the first CAs to serve post-quantum certificates, targeting Chrome’s recently announced Quantum-resistant Root Program.

We are not issuing certificates yet, and it will be a little while before we do. What we are doing is committing to the work in public, sharing the milestones as they land, and telling you exactly what we are building while working with the root programs and other members of the WebPKI community to achieve this.

Two paths to trust

A brand-new root is not widely useful for years. Even after a root program accepts it, that root has to propagate out into the world's operating systems, browsers, and devices, and it never reaches the large set of devices that have stopped receiving updates, or never received them in the first place. That long tail of older clients is where a great deal of the world’s Internet traffic originates, and where a correspondingly large set of avoidable breakage lives. We believe that all clients deserve the highest level of security possible, regardless of their manufacturer, operating system, or time since last update.

Acquiring an existing root with a high degree of trust store coverage across a diverse set of clients solves that on day one. The existing GlobalSign root has been trusted across browsers, operating systems, and devices since 2012, and it reaches older clients that a fresh root never will. The new root that we will be submitting for inclusion in root key programs is built for where the ecosystem is heading, including the programs that are starting to cap how old a trusted root may be. The established root gives us reach across the devices of the past. The new roots give us standing under the policies of the future. We want both to ensure certificates issued by our CA provide the widest set of customer compatibility possible.

A new source of free certificates

The free-of-charge, automated certificate model now carries most of the encrypted web, and much of it runs through one remarkable operator. Let's Encrypt issues on the order of ten million certificates a day, serves more than 500 million sites, and passed four billion active certificates in 2025. It is one of the best things to happen to the Internet in twenty years, and we say that as one of its largest users.

That success comes with some systemic risk: if the dominant free certificate authority had a bad week, much of the web would have no comparable free, automated alternative ready to take the load. At the certificate pack level, we have spent years building exactly this kind of redundancy for our own customers. Every Cloudflare Universal SSL certificate already ships with a backup certificate, wrapped with a separate key and issued from a different authority, ready to deploy automatically if the primary is ever revoked or compromised. A public CA is that same idea, but at the scale of the whole Internet.

To make it easy to adopt, we will be Automated Certificate Management Environment (ACME)-first, an open standard protocol that is widely accepted. Automated issuance and renewal through ACME will be the way you get a certificate from us, which means anyone already pointed at any existing free CA can move to us by changing a directory URL, with no new tooling and nothing to re-architect.

Certificate growth projections are huge

Cloudflare sits in front of more than 20 percent of global Internet request traffic and terminates TLS for millions of domains, relying on millions of certificates per year to do so. We provision those certificates through multiple CAs, with primary and backup paths so customer services stay up through CA outages and revocation events.

That has taught us not just how the WebPKI ecosystem works, but also that it occasionally fails, from the consuming side, the hard way. We have dealt with rate limits, validation edge cases, revocation latency, chain building, and root distribution lag. We have lived through the CA churn of recent years and felt it through our customers. We know what reliable issuance has to look like from the outside, because our customers' uptime has depended on us being resilient and responsive when an issuer has a bad day.

And as certificate maximum validity period decreases over the next few years, agentic activity increases, and PQ certs go mainstream, we expect the raw number of certificates we rely on annually on to continue to grow, quickly — and we are not alone. We want to not just solve this problem for ourselves, but be part of providing this utility to the Internet, and ensure that the certificate supply chain for our customers has even more providers.

Designing for resilience: transparency and fail small

In taking on this new responsibility of being our own CA, we're committed to making the most reliable and resilient CA possible. We intend to build a certificate authority whose reliability depends not just on avoiding mistakes, but as with the rest of Cloudflare’s products, to “fail small” and limit the impact of any one issue.

That means instituting processes to design and test recovery before any incident occurs. As an example, we will make renewal automation a condition of issuance. We will only issue to clients that support ACME Renewal Information (ARI), standardized in RFC 9773. Subscribers must maintain automation that polls our renewal endpoint, acts on the renewal windows we publish, and identifies the certificate it is replacing.

We're also learning from what we've observed over the past 16 years. We have seen certificate authorities caught between timely revocation and keeping subscribers’ sites online because too many subscribers could not replace their certificates quickly enough. When certificates need to be retired, whether for a compliance issue or a security incident, we can bring forward renewal windows for the affected certificates, spread replacements across the available time, and track replacement issuance.

This is just one of the many ways we intend to build. We will be transparent with our issuance stack and operations, publish reproducible builds of the software that signs certificates, attest the hardware security modules that hold our keys, and run a public dashboard for issuance health and incidents. Audits are point-in-time and tell you a CA passed, not how it runs on an ordinary Tuesday. We want root programs, researchers, and ordinary site owners to watch how a modern CA actually operates between audits.

A certificate authority for the post-quantum Internet

We also intend to lead on where certificates are going, not just where they are. We plan to be one of the first CAs to issue production Merkle Tree Certificates (MTCs), with the first certificates issued in the first quarter of 2027.

MTCs are a new and far more compact way to deliver publicly trusted certificates, designed for a post-quantum world where traditional certificate chains grow large enough to strain TLS handshakes. We have been championing the standards-based proposal for MTCs at the IETF, and earlier this year, Chrome named MTCs as the preferred path for post-quantum authentication. Issuing them in production allows us to protect Cloudflare customers as well as the wider Internet against the post-quantum threat, with real volume behind a transition the whole web has to make. We’ve shared much more about MTCs and what this new Web Public Key Infrastructure (PKI) will look like in a blog post on the topic.

We do not expect that transition to be sudden. Much of the Internet will continue to rely on classic certificates and existing WebPKI for many more years. But across that window we expect MTCs to take a steadily growing share of issuance, and that is why we are building one service that does both. By carrying classic certificates and Merkle Tree Certificates under one CA, with one lifecycle and one set of guarantees, customers can adopt at the pace that suits them and help the web make the crossing without a hard cutover. Customers should not have to pick a side of a multi-decade migration, run two systems, or rebuild when the balance shifts.

As always, Cloudflare will be Customer Zero

In addition to providing certificate packs via Universal SSL for our customers, Cloudflare consumes certificates from many different CAs to run our systems and internal operations. Just like our other products, we will be Customer Zero for the new CA and its certificates (both WebPKI and MTC), ensuring that all aspects of the new systems and processes meet our high internal standards, and that our CA’s infrastructure is exercised at Cloudflare scale.

What happens next

We are working through the application and approval process with each of the core web root key programs. These processes happen in the open, and we’ll share more updates as they proceed, through to the first Merkle Tree Certificates in early 2027. If you want to follow this work or be one of the first to use a Cloudflare CA certificate in the future, you can register for updates.

As we build out this new capability, we will continue to work closely with the network of partner public CAs we have relied on for many years — 16 in fact! — as we all work together to ensure a trusted and open Internet.

When we launched Universal SSL, the argument was simple: every byte that flows encrypted across the Internet makes it harder to intercept, throttle, or censor, and the open web is something we all build together. A public, redundant, transparent certificate authority is that same argument carried one layer down, to the trust that makes the encrypted web possible in the first place. We have been working toward this for a long time, and we are glad to finally be on the road.

Happy Birthday Week!

Adaptive application security for the AI era: how Cloudflare connects code, traffic, and intelligence to stop attacks

Post Syndicated from Daniele Molteni original https://blog.cloudflare.com/ai-era-framework/

In July, AI agents testing new cybersecurity models compromised parts of OpenAI’s infrastructure and Hugging Face’s production environment.

We've all just witnessed one of the first AI-driven successful cyber attacks. When given a task, the agents ignored existing guardrails and autonomously discovered previously unknown vulnerabilities, recovered exposed credentials, moved between cloud environments and coordinated their work through communication channels they created themselves.

The speed of the final compromise was incredible. In under 13 hours, the agents went from executing code on a Hugging Face worker to gaining admin-level access across multiple clusters. But the incident had been brewing for much longer. Responders found clues of activity tracing back to May (agents created an unauthorized message board), to June (internal network scanning) and early July. The relationship between these events was understood only on July 20.

The lesson here is not that AI agents exploit vulnerabilities. That’s not news; human attackers already do that. The change is that agents can work persistently, test multiple paths simultaneously, share discoveries, and chain vulnerabilities, credentials, and permissions into sophisticated attacks.

The incident also shows why application security cannot depend on single tools. For example, network restrictions were bypassed by services connected to the Internet; valid credentials were used to perform unauthorized actions. Rebuilding Artifactory removed one attack path, but agents found another. The key insight is that individual alerts identified pieces of the activity without revealing the complete campaign. OpenAI reached a similar conclusion in its report: organizations need overlapping and independent controls across prevention, detection, and mitigation, continuous validation of security boundaries, and faster mechanisms to correlate and contain suspicious behavior.

We address this challenge by connecting application security across four activities that are too often separated: discovering which risks matter, governing what humans and agents may do, protecting applications at runtime, and turning every investigation into stronger protection. Cloudflare can deliver this framework because of its broad security portfolio and visibility across a vast share of Internet traffic.

Alongside the framework, we connect existing Cloudflare solutions with new capabilities across each stage. These include: using Large Language Models (LLMs) to conduct a penetration test of our Web Application Firewall (WAF), expanding threat intelligence to all customers, and a new feature to automate deploying positive security.

What has changed

The security landscape is shifting. These are the emerging trends we see:

  • The way we build software has fundamentally changed. AI-assisted development allows engineers to produce and deploy software faster outside traditional engineering workflows. That speed creates both more code and more opportunities for vulnerabilities to reach production.
  • Software composition risk is still a risk: applications depend on large chains of open-source libraries, packages, and operating-system components that are intrinsically trusted and are difficult for any team to inspect. What’s new is that AI is now importing libraries that we might not be aware of.
  • Techniques and tactics are changing. LLMs can chain vulnerabilities and use feedback in real time to mutate payloads, evade defenses, and make decisions autonomously. They can operate continuously and at machine speed. Patching faster remains important, but patching alone cannot close the gap. Attackers are always going to be faster than you can update your systems.
  • Agentic traffic. In the past, automation was a synonym for malicious activity. Today, a request generated by an agent or bot may be malicious automation, a search crawler, or an agent purchasing a product on behalf of a customer.
  • Compromised servers, residential proxies, IoT devices, and cloud resources allow attacks to move quickly across infrastructure and identities. A coordinated attack can leverage a number of devices, making it difficult to be identified as a unique campaign.

Application Security’s goal is also expanding. It now needs to address three connected problems: protecting conventional applications from AI-enabled attackers, governing legitimate and malicious agentic clients, and securing applications that contain models, agents, tools, and data.

A connected application security framework

Application security in the AI era must operate as a continuous system rather than a collection of controls that teams update after each new vulnerability. To protect applications in the era of AI, you need to work on multiple activities, which we have organized around four stages:

  1. Discover and prioritize risks
  2. Govern access and agent behavior
  3. Protect applications at runtime
  4. Investigate, respond and learn

None of these activities is new in isolation. What changes is connecting them so discoveries, runtime signals, and investigation outcomes continually improve the controls that follow.

More than 20% of the web sits behind Cloudflare’s network, which gives us visibility into attack infrastructure, payload mutations, emerging techniques, and coordinated campaigns at a scale that few organizations can match. Patterns that look isolated from the perspective of one application can become clear across our network. This combination of global threat intelligence, local application context, and inline enforcement powers every stage of the framework. That local context includes which code is deployed, which endpoints are exposed, what legitimate traffic looks like, which identities are acting, and which controls are already active. Because Cloudflare is inline, we can turn those insights into protections immediately.

Cloudflare is the adaptive security control plane for applications, APIs, and agents. Here is what we are launching today to advance every stage of the security journey.

Discover and prioritize risks

Security teams do not suffer from a shortage of findings. They struggle to determine which findings represent an immediate risk. A useful discovery system must connect vulnerabilities to the production reality, including whether a vulnerability buried in your stack is actually reachable in the first place. We see three main areas you should look into: software composition risk, proprietary code, and runtime penetration testing (pentesting).

Understand software composition risk

Applications inherit risk from open source libraries, packages, operating-system components, and the services on which they depend. This represents the Supply Chain of your application. A package vulnerability alone does not tell a team whether the affected component is deployed, reachable, or exposed to hostile traffic. Open-source software is the top priority when it comes to supply chain risk, and Cloudflare is part of Chainguard Athena, an industry coalition aiming at protecting open-source software from AI attacks.

Scan proprietary code

When it comes to code scanning, you have two options: getting a managed service or developing in-house expertise to run it yourself.

Cloudflare recently announced early access to Vulnerability Discovery and Remediation, a service that uses frontier models to identify application-specific vulnerabilities and deploy WAF mitigations to block targeted exploits while engineers fix the code. The important step is prioritization. Cloudflare connects source-code findings to production traffic and security signals. We can identify whether the affected route is active, and how much traffic it receives.

Pentest your application at runtime

Defenders can also use the same capabilities as attackers. Customers can build their own LLM-based pentesting harness to search for weaknesses, validate findings, and test whether their applications are vulnerable. Discovery becomes continuous rather than a periodic exercise. We have done this internally at Cloudflare since Anthropic’s Claude Mythos was released, and we shared our learnings.

A vulnerability buried deep inside your code is harder to exploit if it can’t be reached from the outside. Our Security Analyst team has already used LLM-based red teaming to test customer applications and our own runtime detections, turning the findings into improved detections for all customers. We are now developing Adaptive Security, a self-service capability that will periodically pentest selected URLs behind Cloudflare, using LLM-powered agents to identify vulnerabilities that are reachable and exploitable before attackers find them.

Govern access and agent behavior

Agentic traffic operates in the space between automation and human: tasks delegated by people, executed by software. This changes how access decisions must be made. Detecting automation is no longer enough. For every interaction, application owners need to answer two questions: Is this entity who it claims to be, and can this interaction be trusted?

These questions can be hard to answer. A recognized agent with a long history of legitimate activity may have high trust, but an unusual action can still create immediate risk. An unknown agent may simply be new; a lack of history does not necessarily mean malicious intent. Cloudflare’s approach is to keep trust and risk signals separate, thereby giving application owners more control than a single bot score or allow-or-block decision, and providing more powerful tools to quickly adapt to change.

Establish identity and trust

Trust accumulates over time, while risk is evaluated for each interaction.

Botbase provides a directory of known automated entities that have registered with Cloudflare. Registration gives legitimate bots and agents a way to declare who they are, while application owners retain control over whether and how those agents may access their sites. Cloudflare is also making registration more accessible to smaller and custom agents, building a verified identity layer across all agentic traffic, not just the major platforms.

Identity alone does not establish trust. Cloudflare can evaluate whether an entity has been seen before, whether its historical behavior was legitimate, and whether its current activity is consistent with that history. This makes it possible to distinguish a recognized agent behaving normally from the same agent suddenly changing established request patterns, location, identity, or transaction behavior.

Understand agentic behavior across the journey

Precursor adds client-side and session-level signals to distinguish human from automated behavior, such as typing cadence, mouse movement, navigation patterns, and sequences of actions. An agent that navigates a checkout flow in two seconds, skipping the browsing and comparison steps a human would take, reveals its nature through the session. These signals help identify whether behavior across a session is consistent with human interaction or automation, giving application owners a clearer picture of the traffic they're managing.

Manage access and adapt

Application owners can block traffic from AI crawlers and decide what activity is allowed on their asset (e.g. search, training, etc.). Adaptive Intelligence combines network, client-side, historical, and behavioral validation signals in a probabilistic model that can be updated as attackers change their techniques. Customer outcomes, including chargebacks and successful legitimate transactions, can feed back into the system to improve future decisions.

The result is a continuously updated assessment of every entity and interaction. Application owners can encourage known, useful automation while applying stronger controls when identity, history, and current behavior indicate greater risk.

Protect applications at runtime

Cloudflare’s reverse proxy protects applications at runtime by filtering traffic before it reaches the origin. In the AI era, a new layered approach is emerging to best filter traffic from malicious requests:

  1. Enforce positive security
  2. Detect attacks and identify LLM tactics and techniques
  3. Protect business logic
  4. Deploy real-time threat intelligence

Enforce a positive security model

You can dramatically reduce the attack surface by learning what legitimate traffic looks like, allowing conforming requests and blocking everything else. Today we are announcing Application Profiles, which automatically learns the structure of your web or API application and detects non-conforming requests. Application Profiles automates the learning process and adds a layer of interpretation. Based on the learned profile, we can understand the business logic of different endpoints and request parameters and help you prioritize what endpoints require more scrutiny and attention.

Detect attacks and identify LLM tactics and techniques

Traditional WAFs are designed to run highly crafted rules to detect Common Vulnerabilities and Exposures (CVEs) and malicious payloads. Before AI, the time to disclose new vulnerabilities was measured in months and days. Not anymore: now we see vulnerabilities being exploited before they are disclosed, so the time to patch is nearing zero.

Different tools can be deployed to detect known exploits. These tools include:

  • Managed Rules hardened with frontier models. We’ve partnered with major model providers to use frontier models for adversarial validation. We used the latest models to pentest the WAF to uncover bypasses and vulnerabilities. All customers benefit automatically from ongoing improvements.
  • Machine Learning detection. While signatures are great for high-precision attack detections, machine learning can stop attacks before they are discovered and disclosed. Attack Score detects attack mutations and evading techniques that are often used by LLMs. Attack Score is available to all Cloudflare Customers
  • AI Security for Applications. Chatbots and Internet-facing LLMs are subject to a new class of attacks, such as prompt injection and sensitive data exposure. You can protect generative AI traffic by deploying guardrails and security detections designed to stop these attacks.

Protect business logic

Attackers can still craft legitimate requests and abuse business logic to gain advantage on the application. For example, an attacker uses a valid password-reset flow repeatedly to take over accounts. Fraud detection tools, including account takeover and leaked credential detections, help prevent abuse in which the request appears legitimate, but the intent is malicious.

Real-time threat intelligence detection

Back in June, we launched always-on detection based on our threat intelligence feeds. Cloudforce One customers can deploy protections to block requests originating from compromised infrastructure. We are now also expanding access to Cloudforce One’s Threat Events Platform, our core threat intelligence offering, to all Cloudflare accounts for free.

Investigate, respond and learn

The OpenAI Hugging Face incident did not begin with the final 13-hour compromise. The activity stretched from May to July, with signals including an unauthorized message board, internal network scanning, and movement across environments. Viewed separately, each event revealed only part of the activity. Together, they showed the behavior of a developing breach. Security operations must therefore identify sequences of behavior that lead to compromise, not simply evaluate alerts in isolation.

This is difficult for security teams that already protect large attack surfaces with limited resources. Alerts arrive from different tools and datasets, leaving analysts to determine which events are connected, collect the evidence, and identify whether the activity is escalating.

Cloudflare is building a platform to automate security operations. Deterministic workflows establish the customer and investigation context using trigger history, traffic baselines, enforcement outcomes, and network observations. A detection agent searches authorized datasets for anomalies and correlations. When it finds suspicious activity, specialist agents review the evidence alongside customer history and threat intelligence, helping analysts connect isolated events to broader campaigns. The system can then recommend mitigations, such as rate limiting, WAF, or DDoS protection changes, for human approval.

We are developing these capabilities with Cloudflare’s Managed Defense team, whose analysts are helping us test how evidence is collected, correlated, and turned into recommendations. We plan to make them available more broadly over time and will share more as this work progresses.

Cloudflare’s combination of reverse proxy and forward proxy services makes this correlation especially powerful. Application Security signals can reveal attempts to exploit a public-facing application, while Cloudflare One can surface subsequent activity across corporate traffic. Connecting these datasets can link an external attack with unusual access, internal scanning, or potential lateral movement, turning separate alerts into a timeline of compromise and helping analysts intervene before the breach progresses.

Looking ahead

AI is changing how software is built, how attacks unfold, and who interacts with applications. Security teams can no longer manage discovery, access, runtime protection, and response as separate activities.

Cloudflare is bringing these capabilities together in a closed-loop system powered by global intelligence, local application context, and inline enforcement. A vulnerability finding can strengthen runtime protection, runtime activity can guide an investigation, and each analyst decision can improve future detections and controls.

No organization can anticipate every new technique. The goal is to build a security system that learns from each attempt, responds faster, and becomes more effective over time. The capabilities announced today are the next step toward that adaptive model of application security.

Building a post-quantum certificate authority with Merkle Tree Certificates

Post Syndicated from Mari Galicer original https://blog.cloudflare.com/pq-ca-with-mtcs/

When you type in an address into a browser, how do you know you’re connecting to the right website? The Web Public Key Infrastructure (Web PKI) is the complex and distributed ecosystem of policies, protocols, and infrastructure operators that helps you trust that you’re not being misdirected to an incorrect or malicious website. In the past few decades, this ecosystem has undergone significant changes. One is the addition of transparency: the now-mandatory requirement that all certificates be logged in public certificate transparency logs. Now it faces another challenge: the imminent arrival of a quantum computer, which has prompted us to upgrade to post-quantum (PQ) cryptography by 2029.

This transition is not straightforward: simply swapping post-quantum cryptography into certificates at Internet scale would lead to unacceptable performance degradation. This moment calls for a new approach to the Web PKI, one that allows us to treat transparency as a first-party property rather than an add-on, and design a new system that scales post-quantum signatures efficiently.

After gaining broad support across the industry, Merkle Tree Certificates (MTCs) have emerged as the path forward. This year, after a successful experimental deployment with Chrome, Cloudflare is full steam ahead on MTCs.

Following today’s announcement that Cloudflare is becoming a certificate authority (CA), we’re excited to share that this CA will support MTC issuance, targeting early 2027 for inclusion in Chrome’s newly launched Quantum-resistant Root Store. As part of our mission to help build a better Internet, and following in Cloudflare tradition of offering the strongest available cryptography for free, we will provide standard MTC issuance at no cost. Having a CA that supports both classical certificate and MTC issuance allows us to default to the most secure authentication method available, providing a painless and performant PQ upgrade path for a large swath of the Internet.

The current trust ecosystem

To understand how MTCs are changing the game, let's start with some background on how trust works on the web today.

On the client side, browsers — in this case, “TLS clients” — maintain root programs, which specify a set of policies that CAs must follow to be trusted. On the server side, CAs are the trusted gatekeepers: they operate certificate issuance infrastructure where they validate domain ownership and attest to the binding of a domain name and a public key that shows ownership of that domain.

But how do we check that CAs are following the rules? Enter certificate transparency (CT), which makes certificate issuance publicly auditable. When a CA issues a certificate, it must also submit that certificate to at least two public logs. Cloudflare has operated the Nimbus family of CT logs since 2016, and is launching Raio, a new family of static CT logs, going forward.

While the CT ecosystem makes certificates publicly viewable, it doesn't mean they are correctly issued or safe to use. Monitoring helps with this by comparing those log records with what domain owners expected and reporting suspicious activity. Cloudflare launched Certificate Transparency Monitoring in 2019 and recently made it generally available. We also publish large-scale measurements about certificates on the Certificate Transparency page in Radar (formerly known as Merkle Town).

As organizations begin upgrading their servers to use PQ authentication, certificate transparency monitoring will take on an even more important role in detecting potential post-quantum downgrades. Domain owners who have upgraded their domains to post-quantum authentication should monitor CT logs for unexpectedly issued legacy certificates to prevent clients from falling back on a malicious downgrade path.

Part of the problem with this current system is that transparency was an add-on, causing it to run into scaling issues. Certificates are frequently logged multiple times, in different forms, across multiple logs, requiring monitors to download and process every log to avoid missing an issuance. This can be expensive — making it difficult to encourage a diverse set of log operators at Internet scale. According to our estimates, PQ signatures will balloon the amount of data that CT logs need to store by 40x. This scaling challenge, and subsequent incentive misalignment, is at the heart of the post-quantum scaling problem.

The post-quantum scaling problem

We've written extensively about the challenges of scaling post-quantum cryptography, but in short: to support server authentication at Internet scale, the WebPKI must authenticate roughly a billion TLS servers without preloading every server’s public key into every client. Traditionally, CAs addressed this problem by using certificate chains as a trust-distribution mechanism. But over time, additions like key revocation checks and certificate transparency have added more public keys and signatures — five signatures and two keys in a typical TLS handshake. PQ signatures are roughly 40 times larger than classical ones, creating larger overheads that would be expensive for clients, CAs, logs, and monitors to handle at scale.

Enter Merkle Tree Certificates (MTCs), a draft specification from the IETF PLANTS working group that describes an architecture for compact, efficient, post-quantum certificates. MTCs batch certificates into an append-only Merkle tree, allowing a CA to sign the root of that tree instead of many individual certificates. This allows browsers or other clients to verify a certificate using a compact inclusion proof — a sequence of cryptographic hashes — against a signed tree head rather than validating each certificate individually. A key idea behind MTCs is "don't log what you issue, issue by logging." By coupling issuance and logging, transparency becomes a requirement for operation, rather than an add-on.

The role of a certificate authority in a redesigned PKI

We’re building out our capability to issue MTCs as an integral part of our creation of a Cloudflare CA. That means keeping track of new PQ Root Program requirements, and writing an issuance and mirroring software stack at the same time we’re building the facilities, operations, and compliance functions of the traditional  CA — no small feat!

The upside is that we get to prioritize the requirements and architecture for this new, post-quantum PKI from day one, building our setup in a way that feels right for Cloudflare's values and global network — aiming to be as transparent as possible as we embark on this new journey.

Let’s take a look at the architecture updated for MTC:

If you compare this to the traditional CA ecosystem, you'll notice that the responsibilities of a CA stay mostly the same: to validate control of a domain, bind it to a public key, and issue certificates. The main difference is that in the MTC ecosystem, instead of signing certificates directly and then logging them, the CA now maintains a transparency log backed by a Merkle tree, where an inclusion proof that the certificate is indeed in the tree serves as the trust anchor. CAs will also operate Mirroring cosigners that store a copy of issuance logs, verifying their append-only consistency and ensuring the transparency and availability of these logs for the broader ecosystem.  

Issuing MTCs

MTCs come in two forms, both of which can be encoded in the X.509 certificate format that client software recognizes today — just with a “funny” signature algorithm. In standalone form, the certificate’s signature value contains a cosigned tree head of an issuance log and an inclusion proof (a sequence of hashes) demonstrating that the certificate is contained in that log. If clients are able to obtain the cosigned tree heads out of band (e.g., via a browser update mechanism), the certificate can instead be served in landmark-relative form, where the signature value consists of the lightweight inclusion proof with no heavyweight post-quantum signatures at all.

For simplicity’s sake, let’s take a look at an example of standalone certificate issuance. When a website wants a certificate for their domain, they can request it from a CA via the Automatic Certificate Management Environment (ACME) protocol, which handles certificate requests, domain-control validation, and issuance workflows. Cloudflare's ACME infrastructure will be a fork of Boulder, the widely deployed and well-tested ACME software that powers Let's Encrypt. Let's Encrypt is actively developing MTC support in Boulder, and we plan to maintain our own fork that incorporates these upstream changes along with Cloudflare-specific modifications, contributing back upstream where possible.

When the MTC CA receives a certificate issuance request, the CA's ACME server checks that the server actually controls the domain. If those checks pass, the CA serializes that data and adds it to an append-only log.

After adding the MTC entry into its issuance log, the CA computes the updated state of the log, and then signs a checkpoint over that state. This checkpoint attests that the CA issued every entry included in the log’s Merkle tree up until that point in time.

The CA then sends its updated log state and new checkpoint to a trusted cosigner, which durably stores a copy of the CA's issuance log and checks that each new state is append-only, consistent with the previous tree, and correctly formed. This additional cosignature gives clients and monitors confidence that another trusted party has observed the same log state and verified that the CA is not presenting different views of issuance to different parts of the ecosystem. It also ensures that the issued certificates will be available for monitoring even if the CA issuance log is unavailable.

Chrome’s Quantum-resistant Root Program draft policy mandates at least two cosignatures: one from a Chrome-recognized Mirroring Cosigner operated by a distinct organization, and one from the issuing MTC CA itself. As such, we'll operate mirrors for other pilot CAs — and require at least one independent cosignature on our own issued certificates.

Cloudflare will implement our mirroring cosigner in Azul, our open-source Rust-based transparency log, and for maximal interoperability, it will implement c2sp's tlog mirror protocol.

Finally, after successfully receiving a cosignature from a mirroring cosigner, the CA constructs an MTC with the cosignatures, server's public key, and an inclusion proof. It then sends that MTC to the server, which can then use it for TLS moving forward!

Delivering PQ signatures efficiently: the landmark optimization

While standalone certificates are functional, they still send large PQ signatures over the TLS handshake, limiting their efficiency. The real performance improvements provided by the MTC design are landmark-relative certificates.

Instead of sending cosignatures in every certificate, CAs can designate a sequence of subtrees that cover all active certificates in the log as a landmark, and distribute those subtrees (along with data to authenticate them) to clients via an out-of-band update service. During a TLS handshake, the actual authentication to the server happens by the browser checking that the server's certificate data — including its domain name and public key — appears in a trusted subtree of the CA’s log. If the inclusion proof connects that certificate to a cosigned landmark, and the public key then proves possession during the TLS handshake, the client knows it is talking to the right server.

Periodically transmitting these signatures and tree metadata to TLS clients out of band, a small set of MTC batch signatures can efficiently cover billions of certificates issued by a given CA. While landmarks are more efficient at scale, they do not eliminate the need for standalone MTCs — clients may be newly installed, offline, or missing the relevant landmark update. That’s why it’s important that servers retain a standalone certificate fallback.

MTCs in the wild: results of our experiment with Chrome

This year, we ran an experiment with Chrome to test the feasibility of MTCs between a client and server. We operated a "bootstrap CA" (a fake CA that stubbed the issuance pipeline) that issued MTCs backed by a traditional certificate chain for a selection of Cloudflare domains on Cloudflare's "free" plan and served them to 50% of Chrome Beta 146. Over the course of the experiment we successfully served billions of MTCs.

For TLS, we found that the common case is fairly efficient: with a landmark-relative certificate, the handshake only needs to transmit one public key, one signature, and one inclusion proof of less than 1kB. In the experiment, we fell back to the traditional certificate chain instead of serving a standalone certificate in cases where we were unable to negotiate a landmark-relative certificate with the client. On the CT side, MTCs also change the scaling properties of transparency: the log only needs to carry hashes of public keys; there are no per-entry signatures, and the signature on the tree head covers the whole log. This prevents certificate explosion because the CA issuance log is the source of truth for all certificates the CA issues, and log consumers only need to fetch a single copy of each certificate.

The result: MTCs really work! At median, using a MTC is 9% faster using landmark MTCs over a classical signature chain (admittedly, most of this performance benefit is due to intermediate elision). And because we tested MTCs with classical signatures, we expect an even greater improvement with post-quantum signatures. Satisfied with these results, and with the level of cross-industry collaboration with MTCs at the PLANTS WG at the IETF, we began winding down the experiment last month (August 2026).

The road ahead for MTCs

We’re excited that our experiment with Chrome showed that MTCs can work in practice, and are especially excited to be able to issue certificates as a real CA.

However, there are still broader questions that we can only answer by running this great experiment with the full PKI ecosystem. Can independent monitors consume and verify MTC issuance logs at production volume? Will multiple CAs and cosigners emerge so that the system has the diversity needed for resilience? How should browsers balance the performance benefits of compact landmark MTCs with the fallback paths needed for clients without fresh landmarks? MTCs have emerged as the authoritative design for post-quantum authentication, but proving it out at production Internet scale will require participation from a diverse set of root programs, browser vendors, CAs, mirrors, monitors, and the wider community.

We see the opportunity to participate in this next phase of the Web PKI as an honor, and we take the responsibility of operating CA infrastructure seriously. CAs occupy a privileged position in the trust ecosystem — browsers, domain owners, and everyday people rely on them to validate identities correctly, protect signing keys, follow policy, and operate reliably. Before Cloudflare's CA can be trusted by browsers to issue MTCs, we will need to apply to Chrome's Quantum Resistant root store and undergo a rigorous evaluation process. We welcome that scrutiny, and we expect to hold ourselves to the same high bar as any other CA trusted with helping secure the Internet. We hope other CAs will emerge to support MTC adoption, and we're excited to work with any browser that wants to deploy MTCs.

We tested our own WAF with frontier AI models. Here’s what we found

Post Syndicated from Vikram Grover original https://blog.cloudflare.com/adaptive-ai-waf-testing/

“Is your WAF ready for frontier AI models?” We keep hearing this question from our customers, so we decided to find out.

When it comes to exploiting applications, what LLMs are really good at is iterating and mutating attack payloads faster than any human hacker could do. LLMs can use real-time responses to iterate and change their techniques by, for example, testing different encodings, sending the payload in a different part of the HTTP request, or moving to the next vulnerability to test.

Even before LLMs were around, security engineers used two common approaches to test applications: static and dynamic application security testing. The former analyzes code without executing it to identify vulnerabilities, while the latter probes running applications to find runtime flaws. There are plenty of works scanning code with frontier AI models, including details on how to build your own harness.

For the project described in this blog post, we took a dynamic approach: making the LLM act as if it was a hacker to evaluate whether a WAF is doing its job. The LLM had no visibility into source code, no view of the WAF's rules, and could only see selected HTTP response data.

We built a WAF tester that starts from known exploits and then iterates by changing how it is encoded or delivered, sends it again, and uses the response to choose the next variation. A request that was not blocked became a lead for human review, not a confirmed exploit.

We ran the tester against an authorized customer staging environment across six attack categories and recorded 1,107 attempts. After reviewing the non-blocked requests and removing malformed, benign, duplicate, and out-of-scope observations, the vast majority of the attacks were blocked by the Cloudflare WAF. The requests that got through helped us create new detections to harden our security to benefit all Cloudflare customers.

Here we will explain how we set up the system, the types of attacks we tested, which attack vectors bypassed the WAF more easily, and how we fixed it. Most importantly, we share what we learned from this process and how this exercise is becoming a foundational building block of our WAF development lifecycle.

Finally, we offer guidance to help you correctly deploy your WAF in front of your application and, most importantly, patch your software. A payload that bypasses the WAF still needs an exploitable application to succeed, so keeping your stack up-to-date remains one of the strongest defenses against attackers.

How the adaptive loop works

To test our WAF with frontier models, we built a system that iterates over multiple scenarios. A scenario means choosing one attack category, placing the input in a specific part of the request, starting with a version the WAF already blocked, and giving the tester a fixed number of attempts to try other variations. The loop runs LLM models twice: the first is the proposal call, the second is the review call.

The first call receives the starting request, the context, a short history of earlier results, and suggests the next variation, then the code builds and sends the request. The review call receives the request context, response status, selected headers, and the response body. The loop stops when mutations stop producing useful variations or when a hard coded attempt limit has been reached.

Both model calls work without access to WAF internal information. Neither receives rule expressions, rule IDs, WAF Attack Score details, or the identity of the security layer that acted. We implemented the system in Python rather than wrapping an existing penetration-testing tool. It handles HTTP replay, scenario orchestration, state tracking, and result collection.

In the current implementation, the models do not send requests directly — code controls what happens at each step. Before each request, it checks the target hostname against an allowlist, disables redirects, records the attempt, and enforces the attempt limit. After each request, it records the response and uses the model's review to choose the next predefined step. Response text may appear in a later prompt, so the tester treats it as untrusted input. Neither model call can deploy a rule nor change enforcement.

The system records structured evidence for each attempt.

Six attack categories against one WAF configuration

The main run targeted an authorized customer staging environment protected by Cloudflare’s WAF. We used an allowlisted test User-Agent so the customer’s automated-traffic controls would not stop the test before requests reached the WAF.

We ran 45 scenarios. For each, we looked for ways to deliver the same attack differently: different encoding, different part of the request, or the same destination written another way. Of these, 44 covered six attack categories: cross-site scripting (XSS), SQL injection (SQLi), command injection (CMDi), server-side request forgery (SSRF), path traversal or local file inclusion (LFI), and Log4j. The remaining scenario covered log injection, reported separately.

The WAF in the test zone was configured as follows: WAF Attack Score blocking scores of 30 or below, all Cloudflare Managed Ruleset enabled, and OWASP Core Ruleset with Paranoia Level 3.

For the headline measurement, we recorded whether the WAF blocked each request or not. The results describe the configured WAF boundary as a whole, not the performance of any individual rule or detection mechanism.

What adaptation looked like in one recorded session

Here is an example of how the LLM adapts a Server-Side Request Forgery (SSRF) attack during the test.

Cloud metadata services can expose temporary credentials to workloads. An SSRF vulnerability can let an application fetch that data on an attacker's behalf. A WAF can help stop the malicious request before it reaches the application, but it is only one layer of protection.

In this SSRF scenario, the tester sent the same cloud metadata address in different forms (such as integer, octal, and trailing-dot representations of the same IP) and placed it in different parts of the request. The WAF blocked all of them except one. At attempt 18, the model kept the same request structure as the previous blocked attempt and switched to the trailing-dot form. The client encountered a redirect rather than a WAF block.

The table below shows selected moments from the session. The hypothesis column summarizes what the model said it was trying before each move. It is not a verbatim transcript, and it is not proof that the explanation was correct.

Attempts 17 and 18 are an interesting pair: same request structure, different host representation. One was blocked, one was not. That gave us a specific question: does the trailing dot change how the WAF reads the destination? It was a lead to investigate, but not proof that metadata was accessed.

This was one selected trajectory among 45 scenarios. The next section shows how we counted and triaged the full run.

What we found

Our tester generated 1,107 attempts and the overall result was strong with XSS, LFI, SQLi, and Log4j having near full coverage. While the run produced useful findings, it also produced noise. After human review, we were left with 49 findings worth investigating, 48 of them belonging to CMDi and SSRF. 

Here is how they break down:

Metric

Value

What it means

Recorded mutation attempts

1,107

Model iterations across 45 active scenarios; not all produced a usable result

Post-triage result set

607

The 558 blocked requests plus 49 documented WAF-relevant findings

Blocked requests

558

The WAF stopped these before they reached the application

WAF-relevant findings

49

Documented for remediation analysis after human review

The rest did not produce a result worth counting as the model failed to generate a usable HTTP request, some failed before reaching the target, or the payload generated was benign.

When a request was not blocked, we worked through five questions before counting it as a finding:

Question

Why it matters

Did the tester actually send a valid request?

If the model failed or the request never reached the target, the result tells us nothing about the WAF.

Was the request clearly not blocked?

An ambiguous response is not enough to count.

Was the request still malicious?

Changing a request to get it past the WAF can also make it harmless.

Did the behavior belong to the WAF?

Some attacks only work through DNS or network paths the WAF cannot stop at request time.

Could engineers reproduce it safely?

A fix needs a stable test case with a clear expected result.

We removed anything that failed those checks and combined duplicate cases. What remained became the input for rule, normalization, and mitigation work.

Findings became detections

Not every finding needed a new rule. Some pointed to gaps in existing Managed Rules coverage. Others pointed to how the WAF normalized the request or belonged to another security control. We replayed each case and decided where the change should happen.

We grouped related findings into four sets of candidate rules, validated each finding, and tested candidates against live traffic before any rule could protect customer traffic.

Before a new or updated rule can protect customer traffic, we check its impact on legitimate traffic and assess false-positive risk. Some of the issues we find when evaluating a new rule candidate include:

Issue

Next step

Missing or narrow detection

Review whether existing rules cover the finding

Equivalent inputs interpreted differently

Engine or normalization review

False-positive risk is too high

Revise or reject the candidate

This work contributed to three changes in Cloudflare's Managed Ruleset: new detections for SSRF – Obfuscated Host and SSRF – Restricted Protocol in the July 21 release, and improvement of the existing SSRF – Cloud rule. The SSRF – Obfuscated Host detection came directly from requests that encoded internal addresses in non-standard numeric forms.

What we learned

The model was only one part of the test. We ran the same scenarios with two versions of the same model family. They produced different variations – and the same underlying issues appeared in both. Because request replay and evidence capture stayed consistent, we could compare the runs without treating either model's output as ground truth.

More attempts within one scenario did not always find more. Some scenarios started repeating earlier ideas near the end of the 25-attempt limit. We got broader coverage by testing more starting requests, attack categories, and input locations instead of extending one sequence.

The model generated requests. We decided which ones mattered. A request that was not blocked still needed replay and human review before it could become a finding, a mitigation, or a regression test. Without that review, there were no findings.

What customers can do now

WAF is just one layer of detections you can deploy. When you deploy all available protections you increase the effectiveness of your overall stack. 

First of all, check that Managed Rules, WAF Attack Score are set up correctly in front of your application. Other tools you can deploy include API Security, Bots and Fraud detection, and Threat Intelligence to strengthen your posture even further. For example, positive security controls add a different layer: instead of looking only for known attack patterns, they define the request shapes an application expects and identify inputs outside that contract. This drastically reduces your attack surface area. 

Customers do not need to reproduce this experiment. To maximize the number of rules deployed in front of your application, we recommend running Managed Rules in log first, review matching requests in Security Events, and confirm legitimate traffic is unaffected before moving a rule to Block. Alternatively, customers can reach out to their account team to get Attack Signature Detection turned on, on their zones. This new feature simplifies how to review matched traffic and how to deploy signature detections. If you already perform application security testing, run those tests against a staging hostname protected by the same Cloudflare controls as production.

Next steps

By combining adaptive AI-driven testing with human triage and validation, we found detection gaps that fixed tests might miss and turned those findings into stronger WAF protections, improving our block rate. In a future post, we will share results from further testing using a white-box approach, where the model knows both the application’s vulnerabilities and the WAF rules protecting it.

Is your domain using post-quantum encryption? Now you can see for yourself

Post Syndicated from Andrew Depke original https://blog.cloudflare.com/post-quantum-visibility/

Today, we are introducing additional post-quantum (PQ) cryptography visibility tools into Cloudflare's Application Security and Logs products. You can now inspect and graph the adoption of post-quantum TLS 1.3 encryption for live traffic from directly within Logpush, Log Explorer, and the HTTP Traffic Analytics dashboard. By surfacing the key exchange algorithm negotiated on every incoming request from visitors to our platform, Cloudflare gives customers granular, per-connection telemetry to audit their post-quantum posture, assess compliance, and identify cryptographic gaps across their domains.

Cloudflare is targeting 2029 for full post-quantum security, and executing a cryptographic transition at scale requires detailed telemetry. We’ve already deployed post-quantum encryption across many of our products, including in our cloud-proxy platform and on every on-ramp and off-ramp of our SASE platform.   As many of our customers work towards quantum-readiness deadlines around 2030, we’re helping ease the transition by making post-quantum encryption the default in many of our products, sharing learnings from our internal cryptography discovery tool, and launching the new post-quantum visibility features for TLS that we’ll cover in this blog.

Bringing post-quantum visibility to the domain level

When it comes to post-quantum visibility, we already have macro-level visibility into Internet-wide post-quantum adoption in TLS through Cloudflare Radar. On Radar, we track global post-quantum encryption statistics, both when Cloudflare proxies HTTP requests from visitors (the visitor-to-Cloudflare connection) and when Cloudflare connects to origin servers (the Cloudflare-to-origin connections), as shown in this figure.

From Radar we can see that about 70% of browser-generated traffic hitting Cloudflare's network (on the visitor-to-Cloudflare connection) is protected with post-quantum encryption using hybrid ML-KEM (FIPS 203).  Meanwhile, we can see that today, just about 15% of origins that Cloudflare connects to use hybrid ML-KEM. These are aggregate numbers; the first number is aggregated across all the browser-generated traffic we see, and the second number is aggregated across all the origins we connect to.

We’ve also recently launched Automatic Key Exchange for the Cloudflare-to-origin connection, which reveals which cryptographic algorithms are supported by a given origin. This is useful because outdated configurations can cause an origin to connect to Cloudflare using classical cryptography, even if it does support a post-quantum encryption. 

While Radar and Automatic Key Exchange both provide valuable macro-level views of Internet-wide readiness, our customers have asked us to be able to go beyond aggregate numbers and dive into the behavior of individual domains.

We have long provided visibility into the TLS version used at individual domains (TLS 1.3, TLS 1.2, etc.).

But until now we have not exposed information about the cryptographic algorithms used with the TLS version used at the domain level. This means customers could not answer questions like “What fraction of traffic to my domain www.example.com is using post-quantum encryption?” This information is helpful when aiming to comply with regulatory frameworks, troubleshooting a migration to post-quantum encryption, or seeking to understand which fraction of traffic that is exposed to future quantum adversaries. Now, these questions can be answered.

Post-quantum cryptography in TLS

Before we get into the new product features, let’s do a quick review of post-quantum cryptography in TLS, so we can understand the information that the feature surfaces.

In 2024, the National Institute of Standards and Technology (NIST) stated that RSA and Elliptic Curve Cryptography (ECC) should be deprecated by 2030, and many governments and regulators have since gotten behind that deadline. That’s why today, many of our products are protected with post-quantum encryption using a cryptographic key agreement algorithm called hybrid ML-KEM. Post-quantum encryption is needed right now to stop harvest-now-decrypt-later attacks, where an adversary harvests data today and then decrypts it in the future once powerful quantum computers come online. Organizations that have data that are valuable even if decrypted in 3–10 years (public sector, defense, finance, telecom, healthcare, and others), should consider immediately protecting their traffic with post-quantum encryption.  

 In TLS 1.3, the key exchange group X25519MLKEM768 is the only recommended algorithm for post-quantum encryption. It is now the algorithm preferred by most major browsers. (Note: post-quantum encryption is not available in TLS 1.2 or any earlier version of TLS.)   If you are using Chrome, you can check the key agreement algorithm used by this webpage (or any other) by right-clicking “Inspect”, going to the “Security” tab and looking for the below:

With X25519MLKEM768 in TLS 1.3, the client and server execute both:

  • the Elliptic Curve Diffie-Hellman Key Exchange (ECDHE) over curve X25519 and
  • the post-quantum Module Lattice Key Encapsulation Mechanism (ML-KEM)

X25519 and MLKEM768 each produce a shared secret. TLS then combines those two secrets and uses the result to encrypt TLS traffic. This hybrid approach provides belt-and-suspenders security; as long as one of the two key exchanges is secure, the resulting shared secret is also secure. TLS 1.3 also supports other key exchange groups, including X25519, P-256 and P-384, all of which are just classical ECDHE over different elliptic curves; these algorithms are still used all over the web. In earlier versions of TLS you can also find key agreement based on the RSA algorithm, which is quantum-vulnerable and thankfully much less popular these days due to many known classical security problems.

But post-quantum encryption is only the first part of the story; the second part is post-quantum authentication. Once powerful quantum computers exist, we need to worry about upgrading the certificates and signatures used in TLS 1.3 away from RSA and ECC and towards post-quantum algorithms like ML-DSA. We’re actively making progress towards that goal. In fact, we recently announced that origins can use ML-DSA-44 certificates over TLS 1.3 to connect to Cloudflare, and today we announced that we’re launching a certificate authority that will support post-quantum Merkle Tree Certificates. Nevertheless, for now it remains true that post-quantum encryption with hybrid MLKEM is more broadly deployed than post-quantum authentication.

Bringing post-quantum visibility to the visitor-to-Cloudflare connection

Today we’re making it possible to see the extent to which post-quantum key agreement is used on the visitor-to-Cloudflare connection for any domain in HTTP Traffic Analytics dashboard, Logpush, and Log Explorer.

To view the TLS key exchange data on your domains, go to the Cloudflare Dashboard, and navigate to HTTP Traffic under the Analytics tab. Here you’ll get in-depth statistics about the kinds of traffic visiting your domains, now including a dedicated card for TLS Key Exchange groups on the visitor-to-Cloudflare connection. (Scroll down to find it!) Here’s a look at a TLS Key Exchange card for one of our test domains:

As you can see, the majority of the traffic to this domain uses post-quantum X25519MLKEM768 (in TLS 1.3).  We see some traffic using classical ECDHE over curve X25519 or P-256 (in TLS 1.3 or below).  The traffic labeled “None” is using either RSA key agreement (in TLS 1.2 or below) or no TLS at all. And finally we have a small number of visitors using the now-deprecated X25519Kyber768Draft00 algorithm with TLS 1.3, which we implemented back before X25519MLKEM768 was fully standardized by the Internet Engineering Task Force (IETF). We’ve waited to remove support for X25519Kyber768Draft00 until observed connections are diminishingly small, to avoid regressing clients for which this is their only way to support PQ encryption.

While we’re here, we’ll just drop a few tips about PQ-ing your traffic. If you look at your domain and find no use of X25519MLKEM768 at all, you should confirm that TLS 1.3 is enabled. In the Cloudflare dashboard, select your domain, go to SSL/TLS > Edge Certificates, and then scroll until you find the TLS 1.3 switch; switch TLS 1.3 to On. (There is no separate post-quantum setting: when TLS 1.3 is enabled and a visitor supports X25519MLKEM768, Cloudflare negotiates it automatically.) Also, if the vast majority of your traffic is over classical X25519, P-256, P-384, or None, it might be because most visitors to that domain are non-browser clients that lack support for X25519MLKEM768 and/or TLS 1.3. (Again, most major browsers do prefer to negotiate a TLS 1.3 connection with X25519MLKEM768.)

The key exchange group can now also be a filtering term in the HTTP Traffic dash. Here’s how to take a look at the traffic that is not using post-quantum encryption with X25519MLKEM768:

Analytics are great for aggregate investigations, but being able to see this information in individual log lines can be even more powerful. You can enable the new ClientTLSKeyExchangeGroup field, under the TLS category in the HTTP Requests dataset, to gain visibility into individual post-quantum key exchange in your Log Explorer and Logpush connection logs.

With this new field enabled, you’ll see it start appearing in your Logpush HTTP Request logs, like so:

Visibility to origins and more

The release of the key exchange group stats represents the first major milestone in our broader cryptographic visibility initiative. Designed for scalability, our underlying telemetry pipeline is built to ingest additional cryptographic parameters from TLS handshakes.

That’s why we’ve also surfaced the key exchange group from the Cloudflare-to-origin connection and to provide end-to-end visibility from eyeball to origin in Logpush as OriginTLSKeyExchangeGroup. (This group will be the same for all visitor connections made to that domain, which is why it's not shown in the HTTP Traffic Analytics dashboard).

And for customers that use legacy origin servers that are unlikely to support modern post-quantum cryptography, don’t despair. You can put the origin server behind a Cloudflare Tunnel, to tunnel traffic from the origin server to Cloudflare over TLS 1.3 with X25519MLKEM768, without need to upgrade the legacy origin server itself. This is what the network configuration would look like if you put your origin server behind a Cloudflare Tunnel:

Eventually we’ll be able to also surface post-quantum authentication (namely the algorithm used for certificates and signatures in TLS, including Merkle Tree Certificates) once we start to see a broader-based deployment of that technology.

Your domain has started its post-quantum journey

If your domain is behind Cloudflare, its post-quantum journey is already underway. Check HTTP Traffic Analytics dash and your logs to see the percentage of visitor connections to your domain that already use TLS 1.3 with post-quantum encryption (X25519MLKEM768).  You can also check logs to see if you’re using post-quantum encryption on the Cloudflare-to-origin connection. If your origin server is too ossified to support post-quantum cryptography, then just put it behind Cloudflare Tunnel. With the right settings and visibility, you can protect more of your traffic on Cloudflare from harvest-now-decrypt-later attacks today.

We thank Luke Valenta, Ollie Hsieh and Alex Krivit for contributions to this work.

Introducing Threat Signals: agentic skills for open-source threat intelligence, free for every Cloudflare account

Post Syndicated from Emilia Yoffie original https://blog.cloudflare.com/threat-signals/

Organizations can now scale threat intelligence expertise the way they scale infrastructure. Threat intelligence analysts and network defenders have long automated the ingestion of structured threat feeds to help enrich their SIEM or WAF. The harder work has always been unstructured reporting: turning a research post into indicators your tools can use, without losing the context that explains why they matter. AI skills make that work possible to automate. A skill is a set of rich, detailed instructions that captures how an experienced analyst handles one part of the job, and it runs the same way on every report. 

Threat Signals puts that process into practice at scale. It’s launching today, and we made it available to every Cloudflare account. 

Threat Signals turns open-source reporting that you choose into intelligence you can act on. Its agentic skills summarize reports, surface key context, extract and normalize indicators of compromise, and apply tags — all within a private, account-scoped dataset. The end result is a contextualized indicator stored in your account’s private Threat Intelligence dataset as a Threat Event that can instantly be applied in your WAF policy.

Starting today, we are also expanding access to Cloudforce One’s Threat Events Platform, our core threat intelligence offering, to all Cloudflare accounts for free. With this expansion, each account gets:

  • API and dashboard access to Threat Signals and the ability to select one RSS feed
  • A private dataset built from the RSS feed in Threat Signals, tailored to your reporting requirements and stored for up to 30 days
  • API and dashboard access to Threat Events Platform to investigate events, indicators, and tags related to your private dataset

Essentials, Advantage, and Elite enterprise customers can extend this offering to include an expanded number of RSS feeds, access to Cloudforce One’s proprietary threat intelligence datasets, the ability to generate custom agentic skills, higher storage options for Threat Signals’ derived open-source reporting, and the ability to create custom WAF rules on open-source and proprietary threat events.

Discovery is only the beginning

We started with open-source intelligence because it is the most obvious place to prove the power of agentic workflows. We also heard from customers that their existing platforms cannot scale beyond polling 100 RSS feeds. Recognizing the critical impact open-source reporting plays in understanding the threat landscape, we sought to build an infinitely scalable platform (more on that later).

Researchers regularly publish detailed findings on vulnerabilities, malicious infrastructure, phishing campaigns, malware families, and threat actors. While RSS feed readers make it easier to discover new reporting, discovery is only the beginning. Harnessing data into a usable workflow with consistent expertise is the key to building actionable defense.

Expertise has never been something organizations can replicate at scale. A report explains how a campaign works and identifies the infrastructure behind it, but before an analyst can use that information, they need to:

  • Read and summarize the report
  • Identify relevant indicators
  • Convert indicator values into a consistent format
  • Classify the report using an internal taxonomy for tagging
  • Populate the indicators into a threat intelligence platform (TIP)
  • Preserve a link to the original source
  • Share the intelligence with the rest of the security team

Repeating that process across dozens of sources takes time; moreover, almost every step is entirely about human judgment. As a result, context is lost. Indicators inserted into your TIP are separated from the context that explains why they matter and helps assess the risk later in the remediation cycle. It's not surprising that weeks later, a domain is pushed to a blocklist and nobody understands why. 

How Threat Signals works

Threat Signals uses RSS to monitor the open-source reporting that matters to your organization. You can add an RSS feed, give it a recognizable name and category, and configure how frequently Threat Signals checks for new content. All three feed specifications (RSS 2.0, Atom, and RSS 1.0/RDF) are supported.

Each feed you select enters a Workflow that periodically polls for new articles. It uses Browser Run’s Markdown quick action to fetch and clean the article text into a readable markdown format, which is then stored in R2. The text is passed into an indicator of compromise extractor and a set of default Cloudforce One-defined skills to summarize the content, apply tags based on your account configuration, and add indicator contextualization at the IOC level.

The output is a concise summary and key points that help an analyst quickly understand what happened, who was affected, and why the report matters. All of it is searchable and tagged, so you can find the articles you care about across the platform.

Lastly, each indicator extracted is backed by a threat event within the account's own private Threat Signals dataset. The event, its indicators and tags, and the original report stay connected, so an analyst can always trace where the intelligence came from and why it is there. These indicators can then be used to create WAF rules from threat events to protect your applications and infrastructure.

What we learned

It’s not hard to write a script that pulls an RSS feed and regexes IP addresses out of it. The first version of Threat Signals was a one-week internal prototype, built by a threat analyst who wanted more out of the reports she was already reading. Turning that into something every account can rely on was harder, and most of what slowed us down had nothing to do with parsing. The hard work was in making the output something analysts would trust and actually use. 

We were tempted to let the system invent whatever tags seemed useful. The teams we talked to pushed back: intelligence labeled in an unfamiliar vocabulary is harder to use, because now there are two vocabularies to reconcile. So we limited AI tagging to each account's existing tag catalog. 

Recording whether a tag was applied automatically or by an analyst sounds like a minor piece of metadata, but it turned out to be essential. In our experience, analysts were far more willing to trust automatic tagging when they could see exactly which tags it applied.

Summaries are useful, and they are what users notice first. But what analysts kept returning to in early testing was the link between an event and the report it came from. As investigations progressed, we discovered that link consistently helped them keep track of indicators and understand why each one mattered in the first place. 

What’s next

Open-source reporting isn’t limited to RSS feeds. Analysts need to be able to quickly consume threat intelligence in various formats and pipelines. Now that we’ve laid out the building blocks for ingesting indicators from data feeds into our platform, the natural next step is to add more consumers. Be on the lookout for more data ingestion pipelines that we will support so that you can bring more actionable intelligence onto the platform to protect your organization.

Open the Cloudflare dashboard and set up your feed today

The best investigations begin with trusted context, and Threat Signals helps keep that context close from the first lead onward. Threat Signals is now generally available for every Cloudflare account via API and the dashboard. Open the Cloudflare dashboard, navigate to Application Security → Threat Intelligence → Threat Signals, and add your RSS feed. The documentation is here. 

You can also read threat intelligence research from our team, and talk to your account team about putting Threat Events to work in your enterprise environment.

Preventing quantum downgrade attacks against IPsec

Post Syndicated from Christopher Patton original https://blog.cloudflare.com/ipsec-downgrade-protection/

For Birthday Week, Cloudflare is helping one of the Internet’s core security protocols develop stronger protections against quantum downgrade attacks. To protect our customers and the Internet at large, we worked with the IETF to develop a mitigation against downgrade attacks on IPsec, which we’ve implemented and made available in beta across our IPsec products.

The world is racing to build the first generation of quantum computers. These new machines hold great promise, but they also create a new threat: early quantum computers will be capable of cracking cryptography we've relied on for secure communication. To address this, it is necessary to migrate to post-quantum (PQ) cryptography: cryptography we believe even quantum computers cannot break. Diffie-Hellman key agreement will have to be replaced by PQ key agreement mechanisms such as ML-KEM; classical signature schemes, like ECDSA and RSA, will have to be replaced by PQ schemes such as ML-DSA; and so on.

The PQ migration is well underway, and we’re helping the migration along by making post-quantum encryption the default in our products, open-sourcing part of our internal cryptography discovery tool, launching new post-quantum visibility features, and leading the way in the web’s migration to post-quantum certificates. Still, it will take years before all clients and servers on the Internet have been upgraded to post-quantum cryptography. In the meantime, it will be necessary for modern devices to maintain support for classical cryptography in order to connect with today’s endpoints.

The need for backwards compatibility creates its own risk. In a downgrade attack, an on-path attacker between a client and server tricks the endpoints into using weaker crypto than they support. It does so by manipulating the messages sent between client and server, making it appear to one party that its peer does not support PQ at all. In other words, a downgrade attack eliminates the protection provided by PQ cryptography by downgrading the victims back to classical, so it can be attacked by a quantum computer.

What this means is that merely adding support for the cryptographic primitives themselves is not sufficient to head off the quantum threat. The next frontier in the PQ migration is to prevent active attackers from bypassing PQ by downgrading the connection.

In this post, we focus on the IPsec protocol, a central component of a variety of Cloudflare products, namely Cloudflare IPsec, Cloudflare WAN, and Magic Transit. Like all secure channel protocols, including TLS, IPsec is vulnerable to the following simple downgrade attack as long as both classical and post-quantum authentication are supported. An attacker can impersonate a party by cracking its classical credentials and can pretend the party doesn't support PQ. However, several months ago, we discovered — or rather rediscovered, as we'll explain — a design flaw in IPsec that admits a more sophisticated attack that works regardless of which authentication method is used.

The vulnerability allows a quantum attacker to decrypt all traffic between PQ-capable endpoints. The attack is relatively hard to pull off, as it requires a quantum computation to be carried out in real time during the protocol handshake. (This is different from a harvest-now, decrypt-later attack, where the quantum computation is entirely offline.) We don't yet know if and when this attack will be feasible, but recent trends give us ample reason to be cautious: at the time of writing, resource estimates for quantum attacks on public key cryptography have decreased dramatically, leading Cloudflare to move up our transition deadline to 2029.

To inoculate IPsec to this threat, we helped the IETF develop an extension that adds a downgrade protection mechanism to IPsec. Both parties must support this extension for it to be effective: for our part, Cloudflare has rolled out beta support in Cloudflare WAN and Magic Transit, which customers can now enable by requesting the account managers to turn on the ipsec_downgrade_protection flag for their accounts. We hope to see the rest of the IPsec ecosystem follow suit in short order.

IPsec's place on the Internet

Frequent readers of the Cloudflare blog are likely already familiar with the TLS and QUIC protocols. Between them, TLS/QUIC secure virtually all the web traffic transiting the Internet today. Both operate at the transport layer of the network stack: TLS runs over TCP, while QUIC runs over UDP. 

IPsec serves a similar function, but operates at the IP layer. Because IPsec operates at an even lower layer of the network stack than TLS and QUIC, it is deeply rooted in modern network infrastructure. Cloudflare IPsec allows organizations to extend their IPsec connections over Cloudflare’s global anycast network without expensive multiprotocol label switching (MPLS) connections. IPsec is also part of Cloudflare’s Magic Transit product. With Magic Transit, Cloudflare’s global anycast network sits in front of an organization’s IP range to shield it from attacks and threats like Distributed Denial of Service (DDoS) attacks, and then hands the scrubbed traffic back to the organization via IPsec tunnels.

Despite being so deeply rooted in today's Internet infrastructure, the IPsec protocol continues to evolve. It has seen many important upgrades in the past several years, including the addition of PQ key agreement. IPsec is also on track to adopt PQ authentication on about the same timeline as TLS/QUIC. (In fact, IPsec is actually further along, depending on how it's configured. A pre-shared key is frequently used for authentication in IPsec, and this is already fully PQ!) This suggests that the IPsec ecosystem is more than capable of adapting to shifting threats.

Background on IPsec

Let's now take a peek into the protocol details that are relevant to the downgrade attack. "IPsec" refers to the mechanism used to encrypt IP packets. Before encryption can begin, the endpoints must first perform an authenticated key agreement. They do so using the IKEv2 protocol.

IKEv2 typically has two phases, called exchanges. In the initial exchange, the initiator advertises the parameters it supports and sends a Diffie-Hellman key share. The responder completes the initial exchange by telling the initiator which parameters it selected and sending its own key share.

After the initial exchange, the initiator and responder derive an encryption key from the key shares and encrypt all subsequent exchanges. The key shares are not yet authenticated, meaning each endpoint has no way of knowing where the key share came from. This is accomplished in the authentication exchange, in which the initiator identifies itself to its peer and sends a signature of its key share and advertised parameters. The responder uses the identity to resolve the initiator's credentials and verifies the signature before accepting the new connection. The responder does the same in the authentication message it sends in reply.

One crucial detail to point out here: each party only signs its outbound messages, rather than the entire handshake transcript, as in more modern protocols like TLS 1.3. This means the authenticating party never confirms to the relying party that they've observed the same sequence of messages. This will be crucial for the attack.

Encrypting handshake messages has two purposes. First, it hides the identity of the endpoints from the network. (TLS/QUIC don't have this feature by default, but can enable it using the Encrypted Client Hello extension.) Second, it allows the endpoints to begin using IPsec's packet fragmentation mechanism, making transmission of long messages over multiple packets more reliable. (This is especially relevant to handling large ML-KEM key exchange messages.)

This protocol relies on classical Diffie-Hellman key exchange, meaning a quantum attacker will eventually be able to derive the encryption key from the exchanged key shares. To mitigate this threat, IKEv2 includes an option to run an intermediate exchange following the initial exchange using ML-KEM as the key exchange algorithm:

Backwards compatibility. Crucially, this exchange is only performed if the initiator advertises support for it in the initial exchange and the responder agrees to use it. This allows for backwards compatibility with endpoints that don't yet support PQ. In particular, if the responder selects a classical-only key agreement, then the initiator will assume the responder doesn't support PQ and fall back to classical-only. Likewise, if the initiator doesn't advertise support for PQ key agreement, then the responder will assume the initiator doesn't support it.

Hello my name is Mallory

Let's think about how to exploit this parameter negotiation behavior. We'll start with a simple idea that doesn't quite work, and see what it takes to make it work.

Suppose there's an attacker between the endpoints — let's call them Mallory — who has a quantum computer. Mallory can make it appear to the responder that the initiator doesn't support PQ by intercepting the initiator's initial key exchange message, rewriting it to advertise classical-only, and forwarding the modified message to the responder.

This would cause the authentication exchange to fail. The initiator signs the message it sent, but the responder verifies the message it received. Since the message received is different from the message sent, verification of the signature would fail, unless the attacker also manages to forge a signature that the responder would accept.

That's not all, however: in IKEv2, the authentication messages are encrypted, which means Mallory also needs to compute the encryption key. But this is precisely what the downgrade attack enables: Mallory has already convinced the endpoints to fall back to classical-only, and they can use their quantum computer to recover the encryption key from the Diffie-Hellman key shares.

Still, there's no obvious way to forge a signature from the honest initiator, unless Mallory has compromised the initiator's authentication key. A paper from 2016 observes the following: because the responder only signs its own outbound messages, it doesn't actually confirm to its peer which initiator identity it accepted. This means the responder will accept an authentication message from any initiator it trusts, not just the initiator of the connection.

Suppose Mallory themself is an initiator whose credentials the responder will accept. In this case, Mallory can produce a valid signature using their own credentials. The responder will complete the connection, believing it's talking to Mallory, who is identified by IDm in the figure below. Meanwhile, the initiator (IDi) will complete the connection, believing it's talking to the responder (IDr):

This is a kind of identity-misbinding attack: the endpoints have both accepted an encryption key known to the attacker, but one endpoint has authenticated the wrong entity.

More variants of this attack are possible. For example, in a key-compromise impersonation attack, Mallory would just steal the initiator's credentials and impersonate the initiator directly, allowing them to eavesdrop until the responder has revoked the stolen credentials; this kind of attack does not require identity misbinding. These attacks are also not PQ-specific: Mallory can force the endpoints to use the weakest key agreement method they both support.

Does this attack actually matter?

The main difficulty with the quantum variant of this attack is that the quantum computation is online, meaning it must be carried out during the attack before the handshake completes. This is in contrast to other quantum threats to the Internet, where the computation is offline (harvest-now, decrypt-later attacks, cracking a TLS certificate, etc.). This gives us a little breathing room: downgrade attacks are unlikely to be the first target of cryptographically relevant quantum computers, given there is much, much more low-hanging fruit.

On the other hand, there's a non-negligible chance that Q-day will arrive before we've had time to disable classical-only across the IPsec ecosystem. We don't yet know precisely how long it will take to crack a Diffie-Hellman key agreement, but it's a safe bet that the capabilities of quantum computers will ramp up quickly once they arrive. It's best to get ahead of the threat while we're in the midst of other PQ upgrades for IPsec, especially given how long it takes for these upgrades to get deployed across the ecosystem.

Protecting IPsec

The simplest way to mitigate this attack is to disable classical-only key agreement (i.e., IKEv2 configurations with an initial Diffie-Hellman exchange but with no PQ key exchange following it). This is easier said than done, however: the reason parameter negotiation exists in TLS and IPsec at all is because the initiator doesn't always know the capabilities of the responder before attempting to connect (and vice versa).

In some cases, an HSTS-like mechanism is possible. With HSTS (HTTP Strict Transport Security), a client remembers which of its peers and servers have supported PQ in an earlier connection, and then rejects classical-only in all future connections to those peers. This works as long as you know who is trying to connect, i.e., when your peer identifies themselves. But in IKEv2, negotiation happens in the initial exchange; the peer doesn't identify themselves until the authentication exchange, by which time it's too late.

In any case, this solution fails to address the fundamental problem. Remember that each endpoint signs its outbound messages only, and doesn't sign the messages sent by its peer. This allows an attacker to create a "split view" of the protocol's execution: the initiator sees one sequence of messages, and the responder sees another. Downgrade attacks wouldn't be possible had the initiator and responder confirmed they had a matching conversation. In modern handshake protocols, like TLS 1.3, each authenticating party signs the entire handshake transcript, including the messages they received from the relying party. This allows the relying party to confirm it had the same conversation, thereby preventing the split view exploited by the downgrade attack. We prefer this more principled approach.

Introducing the full transcript authentication extension of IKEv2

We worked with the IPsec Maintenance (IPSECME) Working Group at IETF to develop an extension for IKEv2 (soon to be an RFC!) called IKE_SA_INIT_FULL_TRANSCRIPT_AUTH that endows the protocol with full transcript authentication. For backwards compatibility, use of this extension is negotiated just like any other feature. This means the extension itself is subject to downgrade attack, but the extension uses a clever trick to prevent this.

The extension is very simple:

  • Support for the extension is signaled by a notify message sent in the initial key exchange. The notification is sent unconditionally: the initiator always notifies; and the responder notifies even if the initiator didn't. This is different from TLS 1.3 extensions, where the server is only supposed to reply to an extension if requested by the client.
  • If the peer notifies support for the extension, then an IKEv2 endpoint opts into updated authentication logic. In particular, instead of signing only its outbound messages, it signs the entire transcript. Likewise, it expects its peer to sign the entire transcript.

The trick that prevents downgrades is unconditional notification. Let's say Mallory modifies the initial exchange by dropping the IKE_SA_INIT_FULL_TRANSCRIPT_AUTH notification from the initiator's message, but allows the responder's notification to go through. In this case, the responder falls back to the old authentication logic, but the initiator opts in to the new logic. The responder will end up signing a different byte sequence than the initiator verifies, causing the authentication exchange to fail and resulting in an AUTHENTICATION_FAILURE notification. A similar thing happens if Mallory drops the responder's notification but lets the initiator's through.

Now consider what happens if Mallory drops the notification from both messages. This would cause both parties to fall back to the old authentication logic, allowing Mallory to downgrade the connection and compute the encryption key. But to pull off the attack, Mallory would need to forge a signature not just from the initiator, but the responder as well.

When attempting identity misbinding, Mallory would need to present an identity for a different responder than the initiator wanted to connect to. It's as if the initiator attempted to connect to example.com, but got a certificate for cloudflare.com. Unless the initiator is severely misconfigured, this will cause the authentication step to fail.

If Mallory manages to compromise the credentials of both the initiator and responder, then they can indeed pull off the key compromise impersonation variant of this attack. However, in this case Mallory has much simpler attacks at their disposal. For IKE negotiations, Cloudflare simply acts as a responder. 

How to enable full transcript authentication

This feature is gated under a feature flag scoped to each customer account. Any customer interested in trying it out can request this flag to be enabled on their behalf by reaching out to their account team. 

Here’s what happens at the protocol level, for accounts that enable this flag.  The IKE_SA_INIT_FULL_TRANSCRIPT_AUTH notification will be sent during the IKE_SA_INIT response. We will enable this flag for all customer accounts after sufficient beta testing. The feature gate is created to account for the unlikely scenario that the customer's IKEv2 initiator incorrectly handles the new notification.

Looking forward

As of this writing, this feature is on its way to RFC status. Much of the credit goes to our co-author Valery Smyslov, who did much of the heavy lifting of shepherding the document. He also spotted the trick that makes the extension downgrade-resistant.

The PQ migration is full of surprises. Ideally these surprises are few and far between. The design flaw in IPsec that allows downgrade attacks has been known for some time, at least 10 years as of this writing. There are perhaps many cryptographic protocols in use today with latent bugs that have renewed relevance in the quantum era.

Cloudflare has implemented the full transcript authentication extension and made it available on an opt-in basis. We encourage customers to reach out to their account manager to implement and begin testing the extension, and the rest of the IPsec ecosystem to consider implementing it as the draft continues to advance through the IETF.

Enforce positive security with Cloudflare Application Profiles

Post Syndicated from Daniele Molteni original https://blog.cloudflare.com/application-profiles/

Today, we are launching Application Profiles, a seamless way to enforce a positive security policy. By analyzing the structure and format of HTTP requests and identifying deviations, Cloudflare can help you significantly reduce the attack surface area.

Every customer we speak to wants to know how we can protect them from attacks that use frontier AI models. This has become the number one priority for anyone working in security. Large language models (LLMs) allow even non-technical people to launch attacks with a single prompt. LLMs can generate malicious payloads, test known techniques, and probe applications autonomously by mutating their tactics based on the feedback from the application or the Web Application Firewall (WAF). 

Our tools have changed to stay a step ahead of the attackers. Managed WAF rules and machine learning-based detections remain essential for detecting techniques such as SQL injection, cross-site scripting, remote code execution, and new CVEs, including many variations of those attacks. The answer can’t simply be “patch faster”: this is not sustainable, and it doesn’t work if you haven’t completely mapped your vulnerabilities.

What if you could learn what good requests look like by analyzing your traffic structure? Instead of looking only for requests that resemble known attacks, we could allow only requests that conform with what we expect. By doing this, we’d dramatically reduce the attack surface area. For example, if the search field in your query doesn’t expect special characters, we can only accept alphanumeric strings. This would already prevent a vast library of known attacks.

But we don’t stop here. Once we have learned the structure and format of your HTTP requests, we can infer the goal of each operation and then understand what the application ultimately does. With this information, we can identify and prioritize the most critical and vulnerable operations and fields you should take care of first.

Cloudflare already supports positive security for APIs through Schema Learning and Schema Validation. We are now extending this protection to web applications through Application Schema Profiles. You onboard an application, we learn its profile, and then we start to deploy an always-on detection that identifies non-conformity. All automated and enriched by powerful analytics.

We are opening a closed beta to invited Enterprise customers without API Security; customers with API Security already have access.

Validating requests based on learned profiles

Schema Profiles periodically learn the expected request structure from observed traffic. After a profile is available, an always-on validation layer is automatically deployed on live traffic. For every request, the detection evaluates whether it conforms or not with the profile, and it adds the result as metadata, augmenting the information already associated with the request. The signal does not take action by itself: customers can analyze past traffic in Security Analytics and decide where enforcement is appropriate and create Security Rules to block non-conforming requests. Requests to operations without a profile are not classified by this feature.

Unlike Managed Rules, failing validation does not require a request to match a known attack signature. A value outside an expected range, an unknown enum value, an invalid universally unique identifier (UUID), or unexpected characters — all can be identified because they differ from the learned profile.

For example, consider the following operation: 

www.example.com/shop/2dbda2e7-cfc9-448d-9465-799d2e6ff363/inventory?product_id=938062541

Below we describe the learning process, which evaluates only the structure and format of the request. When enough traffic has been observed, we learn that the path expects a UUID variable and that product_id is an integer and what its boundaries are. When product_id contains a string, it will be flagged as a violation. Similarly, Cloudflare can identify malformed UUID values and, when the customer enables enforcement, prevent non-UUID input from reaching the corresponding handler. These simple filters reduce the range of inputs an attacker can send, preventing the vast majority of typical attack vectors, such as SQL injection, cross-site scripting, remote code execution and more. 

Non-conforming does not always mean malicious. An application release, a new client, or an unusual but valid request may also introduce a difference. We recommend starting in observation mode, so customers can review a profile's effect before enforcement.

Learn the expected structure of requests

To determine the anticipated request structure for a web or API application, Schema Profiles routinely analyze observed traffic. Each profile may include the following, depending on the application traffic:

  • Path variables 
  • Query parameters
  • Headers and cookies
  • Body structure (JSON body or form-encoded)

For each field, the system learns its data type (integer, string, boolean, arrays, UUID or enum) and constraints such as numeric ranges, short enumerations, string lengths, and character classes.

Learning applies to operations that customers select for profiling. In Web Assets, an operation is Cloudflare's term for an operation identified by its HTTP method, hostname pattern, and path pattern. Web Assets continuously discovers operations and lists them under Web Assets > Operations. Customers can also add operations manually. Profiling doesn’t automatically start for discovered operations, while manually created operations do trigger profiling when created. For discovered operations, the customer must intentionally select Learn profile from the operation's overflow menu. 

After profiling is enabled, Cloudflare collects qualifying traffic and runs learning automatically once a week for each zone, using the most recent successful traffic. An operation needs at least 1,000 requests that received a 2xx response in the previous seven days to learn fields, and at least 10,000 to learn data boundaries. Successful requests can include bots and scanners, so customers should review a learned profile before enforcing it. Our roadmap includes allowing customers to trigger learning on demand and excluding automated traffic.

Once learned, profiles can be reviewed by selecting View details of the operation and finding the learned schema in the Security overview panel. If a learned schema is not shown, Cloudflare is still collecting data for the profile. Customers can also export the profile as an OpenAPI v3 schema file.

Learned profiles update each week as application traffic changes. New fields are added and fields that are no longer observed are removed, so validation tracks how the application changes. Customers can pin and save the learned schema by downloading the learned schema and uploading it to Schema Validation.

Review before you block

Security Analytics now includes a new Profile Analysis tab. Customers can select a validation profile and see traffic trends, including how many requests did not conform to the learned profile during the previous seven days. 

Customers can review the conforming and non-conforming traffic. They can drill into violations and review sampled logs to see where the violation occurred, which field was affected, and why it failed validation. Violations are classified into ten reasons, including type mismatches, values outside a learned range, and invalid formats.

Once a team understands the effect, it can use Security Rules to act on the signal. A rule can cover an entire application or be limited to selected paths, operations, or fields. Teams control where to monitor and where to block.

Positive security for web and API traffic

Traditional WAF learning modes can build detailed positive-security policies, but they often require operators to review suggestions, stage changes, and maintain policy entities. Cloudflare Schema Profiles expose validation as a request field cf.schema_validation.learned.violated, allowing customers to combine it with request properties, Bot Score, Attack Score, and other signals in a single Security Rule. By creating simple rules, teams can combine detections and define precisely when Cloudflare should take action.

Two other classes of fields are available to create more targeted rules. First, there are fields that collect where the violation occurred. For example, based on our initial example, if the value of product_id query parameter does not conform with the profile, the following field will be populated cf.schema_validation.uploaded.query.violated_parameters = ["product_id"]. This allows customers to create rules that enforce positive security only on specific fields or exclude them from the enforcement.

The second class of field collects new parameters that are not present in the profile. This is useful when you want to handle requests with new parameters (e.g. when deploying a new version of your application), or restrict your posture even further by blocking any parameters that were not detected or defined in the past.

Use case

Field

Location values

Example

Identify where in the request the violation occurred

Array up to 20 items

cf.schema_validation.learned.[location].violated_parameters

query,path,headers,cookies,body

cf.schema_validation.learned.query.violated_parameters = ["product_id"]

Identify whether an undeclared parameter is seen in the request 

Array up to 20 items

cf.schema_validation.learned.[location].undeclared_parameters

query

cf.schema_validation.learned.query.undeclared_parameters = ["adminMode", "utm"]

Coming up: critical field analysis, how we help you roll out positive security

Even with a flexible enforcement design, customers tell us that deploying a positive security policy is operationally complex. A large application can have thousands of operations with tens of thousands of fields. But not all operations and fields carry the same risk. Contextualization and prioritization helps security teams roll out positive security in a controlled and confident manner.

LLMs can help contextualize learned profiles to provide additional insight. For web applications, paths and field names are usually self-explanatory, thus semantic. For example, we piloted running a model hosted on Workers AI across four random applications’ learned profiles. The model successfully identified the link between clientId and account_number across two applications of a system, as well as the common dependency of using One-Time Password (OTP) for enhanced authentication. Highlighting this context enables security teams to prioritize actions such as configuring Rate Limiting Rules to defend against account-focused brute force attacks.

These LLM-powered insights will be accessible directly within the dashboard alongside each operation in Web Assets. Before executing a one-click deployment, teams can evaluate rule recommendations designed to secure these key fields, backed by mitigation simulation using past traffic to gain confidence.

Beyond contextualizing operations with semantic insights and risk indicators, we are developing additional metrics to order operations using historical request trends and signals. This enables security teams to focus mitigation efforts on the highest-priority operations first, including:

  • Data loss: upward trend of unusual increased data transfer
  • Reconnaissance activity: high count of unknown parameters
  • Business criticality: total volume of traffic correlated with the unique session IDs served

What’s available today

Customers with API Security already have access, given that this is an extension of Schema Learning and Schema Validation. We are opening a closed beta to customers without API Security who can test Schema Profiles on production web application traffic, meet with the product team, and provide detailed feedback on profile accuracy, analytics, and enforcement controls. Access is by invitation and does not imply future plan availability. If you are not an API Security customer and want to get access, contact your account team.

The feature supports paths, query parameters, headers, cookies, JSON request bodies, and form-encoded request bodies. Profiles can validate integers, strings, UUIDs, arrays, and enums containing up to three values. Multipart forms, GraphQL, and XML are not supported at this time.

Schema Profiles validate every value when a parameter name is repeated, but they do not enforce parameter uniqueness. They also do not learn and enforce required parameters or block a request solely because it includes a new parameter.

Get ahead of zero-days

Our idea for Application Profiles does not stop at validating request structure. The same workflow can learn other characteristics of what an application expects (such as ASNs or JA4s), explain when traffic deviates from them, and give security teams confidence in defining what “good” looks like. With a Proactive Security workflow, we help security teams get ahead of zero-days!

Using Device Linking to Eavesdrop on WhatsApp and Signal

Post Syndicated from Bruce Schneier original https://www.schneier.com/blog/archives/2026/09/using-device-linking-to-eavesdrop-on-whatsapp-and-signal.html

Modern messaging apps allow users to link their phone accounts to their computer desktop. Eavesdroppers are taking advantage of this capability:

Apps such as WhatsApp Web and Signal Desktop allow people to use their accounts on other devices, such as laptops or desktop computers.

Germany’s Customs Office has been using these features to connect a police-controlled computer to a suspect’s account.

Once connected, messages can be delivered to that computer without the police having to crack the encryption protecting them.

Netzpoltik details that police are able to gain access in this way either through physical access to someone’s phone or by intercepting verification codes via a state-sanctioned phishing attack or intercepting SMS messages via telephone surveillance.

That last paragraph is important. Making this work requires user consent.

What we want is a feature that displays connected devices, so users could notice if a new device gets connected to their account.

Колъм Тойбин: Без уютни събития, без лесни преживявания, без неоспорима развръзка

Post Syndicated from Антония Апостолова original https://www.toest.bg/kolm-toybin-bez-uyutni-subitiya-bez-lesni-prezhivyavaniya-bez-neosporima-razvruzka/

Колъм Тойбин: Без уютни събития, без лесни преживявания, без неоспорима развръзка

Световноизвестният ирландски писател Колъм Тойбин идва за пръв път у нас по покана на ICU – българското издателство на книгите му. В навечерието на неговото гостуване излезе и романът му от 2017 г. „Дом на имена“. Срещата на автора с читатели във формат „въпроси и отговори“ и подписване на книги ще се състои на 6 октомври от 16 ч. в книжарница Umberto & Co. На 7 октомври ще се проведе и галавечер с водеща Надежда Московска и с участието на преводачките Бистра Андреева и Елка Виденова. Събитието ще започне в 19 ч. на голямата сцена на Младежкия театър.

Когато не пишете за велики писатели или митологични персонажи, историите Ви следват живота на съвсем обикновени хора, на които се случват съвсем обикновени неща (по Нортръп Фрай). В какво се състои литературното обаяние на обикновеността за Вас? 

Знам какво имате предвид под „обикновено“, само че понятието „обикновен“ не означава кой знае какво за мен, когато работя. Често пиша за света на детството и за семейството. Но дори когато не е така, се опитвам да драматизирам сложността и двусмислието, а за това е необходимо да съзра някакво вътрешно богатство в героя си, независимо от житейските и другите обстоятелства около него. Написал съм около дузина романи и без да съм го планирал, се очерта следният модел: един роман е за известен автор или митичен персонаж, а следващият – за някой „най-обикновен“ човек от моя роден град или семейство. Изпитвам облекчение, когато преминавам от едното към другото и обратно. 

Двусмислената (да се захвана за тази Ваша дума) концепция за дома бележи голяма част от творчеството Ви. Самият Вие сте живял в редица държави. Стигнахте ли до надеждна дефиниция за „дом“?

За човек на моята възраст домът е там, където са компактдисковете му! Признавам си, че колкото и да е странно, не разсъждавам много върху това, върху концепцията за дом. Когато обаче хвана самолетен полет от някой американски град за Дъблин, знам много добре, че се прибирам у дома. Знам къде и какво е домът. Това е Ирландия. Това е Уексфорд. Това е мястото, откъдето съм. 

В „Празното семейство“ пишете: „Всяка събота ходех до Пойнт Рейъс, за да усетя болката по дома.“ Коя емоция за Вас най-автентично изразява принадлежността? 

На някои езици – на испански например – е трудно да се преведе самата дума miss [в оригиналния текст Тойбин пише буквално to miss home, „за да ми липсва домът“ – б.а.]. Предполагам, че в изречението, което цитирате, се опитвам да внуша идеята, че това „да ти липсва домът“ е вид фантазия, представление, нещо може би реално, но също така вероятно и изкуствено. 

Централен персонаж в много от историите Ви е преобърнатият архетип на майката – понякога проблемно, отчуждено, дестабилизиращо присъствие. Или по-скоро отсъствие – нещо, с което често работите. Сред най-мощните въплъщения на последното е quest-ът, търсенето на майката, в разказа „Дълга зима“…

Пристъпвам към историите една по една. Не следвам някаква определена теория. Не се опитвам да доказвам нищо. Художествената литература се нуждае от разрив, от нарушаване на баланса, от герои, които не се държат по очакван и обичаен за образа си начин. Една любяща майка няма да ми свърши особена работа. Няма какво да я правя. Ами ако майката не е майчински настроена? Ами ако присъствието ѝ е вредно и пагубно? Ами ако именно отсъствието ѝ е онова, което има значение и въздействие? Няма ли да бъде по-интересно? Иначе, що се отнася до „Дълга зима“, написах разказа малко след като майка ми и брат ми починаха и цялата мъка, привнесена в тази история, беше все още съвсем жива и оголена за мен. Не бях я планирал като почти автобиографична, но се получи тъкмо такава. 

Подобно на „Одисея“ изобразявате завръщането у дома като много по-голямото изпитание, отколкото напускането му. У Вас и двете могат да се четат, понякога едновременно, като акт на окончателно пристигане, на бягство, на спасение, на поражение…

Обожавам края на „Одисея“. Няма щастливо завръщане. Шеги, ирония, увъртания. Обичам да се завръщам в Ирландия. Но това простичко чувство не трае дълго. То е фалшиво усещане. Изобщо, много ми харесва идеята за измамното, невярното усещане в литературата – то ми допада повече от автентичността или искреността. Предполагам, това, което в крайна сметка се опитвам да кажа, е, че белетристиката трябва да бъде чисто и просто интересна. Без уютни събития, без лесни преживявания. Без неоспорима развръзка. 

В какво се състои себепознанието у героите Ви? Тяхната идентичност често е диалектична – нещо, което градят по необходимост след криза, посттравматично, компромисно.

Този въпрос е лесен, или поне отговорът му е такъв. В моите романи героите ми правят нещо, мислят, помнят. Не ме занимава идеята за някаква всеобхватна идентичност, нито дори проблемите на себепознанието. Нямам теория за човешкия характер. Нямам дарба за абстрактно мислене. За мен съществуват единствено и само следващият образ, следващото изречение, следващата сцена. Работата ми се състои в това да създам нещо интересно и истинско. Така че оставям персонажите си да живеят колкото могат. Много често те мълчат за важните неща, а това придава на вътрешния им свят суров и нелицеприятен вид, прави ги неспособни да общуват лесно и директно. Разликата между това, което чувстват, и това, което разкриват, дава огромна енергия на повествованието. Интересувам се от тази раздалеченост. Интересувам се от един герой, най-много двама, на едно място. Интересувам се от личния интимен живот – какъв е, как се усеща; от вътрешния свят.

Заговаряйки за интимността, със сигурност сте казал достатъчно за това какво е да се пише за гей сексуалността в един, поне доскоро, репресивен социално-културен контекст като ирландския (и българския, уви). И по-важното – за автоцензурата, преодолявана по пътя към тези Ваши сурови, неподправени, натуралистични описания на секса между мъже. 

Понякога е важно да не се пишат графични сексуални сцени. Няма закон, който да казва, че трябва. Но понякога начинът, по който героите правят секс, е съществен за историята. Предполагам, че има няколко правила: без метафори, без сравнения, без завоалиран или натруфен стил. Просто кажете какво са направили персонажите, не как са се чувствали. Впрочем точно днес получих имейл от хетеросексуален приятел, който реагира на интимните гей сцени в новия ми роман (The Bridge). Та той пише, че се е почувствал възбуден от описанията. Това ми се стори хубаво. Една от секс сцените в тази книга е в затвор, та имах добро основание да я направя графична – важен беше начинът на правене на любов, конкретните физически действия. 

В творчеството Ви любовта често се явява форма на задължение – преплетена с неизбежност, с наложителност, с премълчаване. Има ли изобщо нещо лесно и освобождаващо в нея? 

Може би, но то не върши работа в един роман. Иначе, в разказите се опитвам да работя по ръба на това, което може да бъде изречено, и онова, което трябва да остане неизказано. Кимването е важно, смръщването, въздишката, полуизказаното, необлечената в слово мисъл, внезапното изтърсване на нещо.

В последния си издаден на български роман „Дом на имена“ ни давате гледните точки на Клитемнестра и децата ѝ Електра и Орест, но не и на Агамемнон. (Интересно, че той бе лишен от лице и почти от глас и в „Одисея“ на Нолан). Вашият специфичен поглед върху тази архетипна история? 

Първоначално исках да работя с това, което бих нарекъл „стакато в първо лице“: гласовете на жените – Клитемнестра и нейната дъщеря. Но впоследствие бях очарован от историята на Орест, от неговата срамежливост, от неговото отсъствие, от неговата сдържаност. Нямах никакъв интерес към гласа на Агамемнон [който бива убит от съпругата си, след като принася в жертва дъщеря им Ифигения – б.а.], към мотивите му или към неговата версия за случилото се. Накрая поставих именно Орест в центъра на историята. 

Всеки акт на насилие там води не до развръзка, а до отварянето на нов цикъл от болка. 

Да, всяко убийство следваше като вид възмездие. В един момент, когато пишех за насилието в Северна Ирландия, забелязах тази спирала – убийства тип „око за око“, убийства за отмъщение. Именно това беше в съзнанието ми. 

Намираме се на прага на настъпващата ера на изкуствения интелект. Когато един ден ИИ овладее писането, кой недостатък на създадената от човека литература би Ви липсвал най-много?

Знанието, че си се провалил. Самата концепция за провал.

 

Building event-driven applications at scale with Amazon EventBridge

Post Syndicated from Nahid Karimaghalou original https://aws.amazon.com/blogs/compute/building-event-driven-applications-at-scale-with-amazon-eventbridge/

Event-driven applications on Amazon EventBridge usually start small and then spread. One team creates a Custom event bus, adds a few rules, and ships. Another team needs some of those events, so a rule forwards them to a bus in a second account. A third team needs a subset of what the second team receives, so another rule forwards again. A year later the organization runs dozens of Custom event buses joined by forwarding rules, and that topology has become a thing to operate in its own right.

That shape has a price, and the smallest part of it is the bill. Every forwarding hop is a separate ingestion, so cost tracks the topology rather than the number of consumers that needed the event. The harder problem is that nobody can see the whole picture. Governance spreads across the accounts it was meant to cover. Answering who publishes to a bus, who consumes a given event type, or what breaks when a team stops publishing means visiting each account and reading its rule configuration. Tracing one event is harder still: its path crosses several buses in several accounts, each with its own metrics and logs, and no single view follows it from publication to the consumer that never received it.

Application teams also wait. Publishing to a bus in another account, or consuming from one, needs a resource policy, a role, and a forwarding rule owned by a central team. The team that wants to build opens a ticket, and the platform team becomes a queue. Both the missing visibility and the waiting grow with every team onboarded.

Amazon EventBridge recently relaunched the Custom event bus, which tackles these challenges directly. A platform team creates one bus, shares it across the organization, and keeps control of who can publish and who can subscribe. Every consumer of those events is listed on the one bus rather than inferred from configuration spread across accounts. Application teams create their own Subscribers in their own accounts. The bus stores events for a retention period you choose, preserves order within a key the publisher sets, accepts Avro and Protocol Buffers (Protobuf) alongside JSON (including CloudEvents), and delivers to targets without a function in the path to translate a call. It runs alongside the Custom event bus – classic, so adoption is incremental.

In this post, you see how a platform team stands up a shared bus and governs access to it, how application teams onboard themselves with a single Subscriber resource, and how retention, ordering, open formats, transformation, and direct target integrations change what one bus can carry.

One bus, shared with the organization

The platform team’s job on a shared bus is narrower than it was on a fleet of them. It owns the bus and sets the boundaries: which principals can publish and what their events can declare, which principals can subscribe, and, where it matters, what those principals are allowed to filter on. Application teams then manage their own configuration within those boundaries, such as filters, targets, delivery roles, retry policies, and failure destinations, none of which the platform team needs to write or review. That division is the point of the design. The platform team keeps governance of the bus and stops owning everyone else’s configuration, which is what takes it out of the provisioning path without giving up control of who is on the bus.

Creating the bus is a single call in a platform account.

BUS_ARN=$(aws eventsv2 create-event-bus \
    --name company-events \
    --storage-configuration '{"RetentionPeriodInDays":7}' \
    --query EventBusArn --output text)

Retention is the one setting worth deciding deliberately here rather than revisiting after an incident. It runs from 1 to 365 days and can be modified later, but a change only applies going forward. Raising it widens the window for events published from that point on, and does not make older events readable again. Seven days covers a working week of history, which is usually enough to onboard a consumer or reprocess after a bug without paying to store a year of events nobody will read.

Sharing the bus is the second decision. AWS Resource Access Manager is the route to reach for first: it associates automatically for accounts in the same organization and reaches accounts outside it by invitation the consumer accepts. A resource policy written on the bus directly is the alternative, and can also name accounts inside or outside the organization.

Access is granted per principal, and publishing and subscribing are separate permissions. A team that produces order events gains no ability to read payment events from the same bus. One grant is not enough for a cross-account caller, as usual on AWS: the role that publishes or subscribes also needs its own IAM policy allowing those actions. The platform team decides which accounts can reach the bus, and each consuming team decides which of its own principals can use that access.

Taken together, those decisions produce the architecture in the following diagram. One bus lives in a platform account, and application teams publish to it and subscribe from their own accounts. An AWS Lambda function in Team A’s account calls PutRawEvents to publish events onto the Amazon EventBridge bus in the platform account. Team B and Team C each attach their own Subscriber: Team B’s delivers to a Lambda function, Team C’s to an Amazon DynamoDB table.

Architecture diagram of one Custom event bus in a platform account. A Lambda function in Team A’s account calls PutRawEvents to publish events onto the Amazon EventBridge bus in the platform account. Team B and Team C each attach their own Subscriber in their own accounts: Team B’s Subscriber delivers to a Lambda function, and Team C’s Subscriber delivers to an Amazon DynamoDB table.

Figure 1: Multi-account sharing

Cost follows team boundaries because charges separate ingestion from delivery. The account that publishes an event pays to put it on the bus, and the account that owns a Subscriber pays for what that Subscriber consumes. Each team’s usage appears on its own bill, which is what makes a shared bus something a platform team can charge back rather than a shared cost center nobody can decompose. Removing the forwarding hops also removes the duplicated ingestion and delivery those hops created: the same event reaching the same three consumers is ingested once instead of three times.

Publishing in the format teams already use

Not every producer speaks JSON. Teams that standardize event exchange across an organization often register schemas and publish compact binary payloads, because the schema is the contract between teams that deploy on their own timetables. Accepting the formats those producers already emit is simpler than changing each one to convert to JSON first.

With the new Custom event bus, application teams can publish events in Avro, Protobuf, and CloudEvents (JSON) formats. For the binary formats, a schema registry named on the request is used to deserialize the events.

There are two publish APIs, and the payload decides which one to call. PutEvents takes structured JSON with the familiar Detail, Source, and DetailType fields. PutRawEvents takes a binary payload plus metadata you define, and is the one to use for Avro, Protobuf, CloudEvents, or bytes the bus should not interpret.

import boto3

events = boto3.client("eventbridgev2")
events.put_raw_events(
    EventBusArn=BUS_ARN,
    SchemaRegistryConfiguration={"RegistryUri": GLUE_REGISTRY_ARN},
    Entries=[
        {
            "Data": avro_encoded_order,  # bytes, straight from your existing producer
            "SystemMetadata": {"ContentType": "application/avro"},
            "Metadata": {"eventType": "OrderPlaced"},
        }
    ],
)

The schema registry can be either the AWS Glue Schema Registry or the Confluent Cloud Schema Registry.

Because the bus decodes the event before filters and transformations run, a consumer subscribing to Avro events written by another team needs no schema, no decoder, and no access to the registry. It writes the same filter it would write against JSON. Producers and consumers stay decoupled, and no deserialization code has to be repeated in each consuming team.

Publishers get one more setting on the same request: deduplication. A retry that already succeeded would otherwise leave a duplicate for every consumer to handle. It works one of two ways: the bus hashes the content of each event, or it uses a deduplication ID you supply. Content-based hashing suits producers with no natural key, since two identical events hash the same. A deduplication ID fits when you already have one, such as an order ID combined with a state transition. It keeps matching even when parts of the payload differ in ways that should not count as a new event.

Self-service onboarding for application teams

The new Custom event bus introduces a new resource called a Subscriber. Application teams create and configure their own Subscribers in their own accounts, provided they have been granted subscribe access to the bus. A Subscriber is the one place a consumer’s behavior is defined: which events it receives, where they are delivered, how delivery is retried, and where events go when delivery does not succeed. Reviewing or changing a consumer is one thing to read and one thing to update.

SUBSCRIBER_ARN=$(aws eventsv2 create-subscriber \
    --name orders-to-fulfilment \
    --event-bus-arn "$BUS_ARN" \
    --filter-configuration '{"Filters":[{"Scope":"METADATA","Pattern":"{\"eventType\":[\"OrderPlaced\"]}"}]}' \
    --invoke-configuration '{"TargetArn":"'"$QUEUE_ARN"'","RoleArn":"'"$ROLE_ARN"'"}' \
    --retry-policy '{"MaxRetryAttempts":10,"MaxEventAgeInSeconds":3600}' \
    --on-failure-configuration '{"Arn":"'"$DLQ_ARN"'"}' \
    --query SubscriberArn --output text)

A filter’s scope decides which part of the event the pattern is matched against. DATA matches the payload, METADATA matches the key-value pairs the publisher attached to the event, and SYSTEM_METADATA matches the event’s system fields: the content type and ordering key a publisher declares, plus the fields Amazon EventBridge adds itself. Because Avro and Protobuf payloads are decoded as they are published, a DATA filter reads their fields directly, the same as it would for JSON.

The retry policy says how the bus should behave when a target is failing. MaxRetryAttempts sets how many times a delivery is retried, and MaxEventAgeInSeconds sets how long an event stays eligible for retry, measured from when it was published. Retries stop as soon as either limit is reached, so both bound the same delivery.

When deliveries do fail, the reason shows up in the Subscriber’s own logs, which application teams can turn on themselves. They record the error from each delivery attempt alongside the exact input sent to the target, which makes a problem quick to place. Seeing what the target actually received separates a transformation that produced the wrong shape from a target that rejected a correct one.

Screenshot of the Amazon EventBridge console showing the Create subscriber form, with fields for the subscriber name, event bus, filter configuration, target (invoke configuration), retry policy, and on-failure destination.

History for consumers that did not exist yet

A Subscriber sometimes needs events that were published before it existed. For example, a new analytics service needs hydrating with recent history, or a target processed a window of events incorrectly and needs that window replayed. Because the bus retains events for the period configured on it, a Subscriber can be created with a starting position in the past, so it reads history, catches up, and continues with live traffic:

aws eventsv2 create-subscriber \
    --name analytics-backfill \
    --event-bus-arn "$BUS_ARN" \
    --starting-position POINT_IN_TIME \
    --point-in-time-configuration '{"PointType":"TIMESTAMP","StartingPoint":"2026-09-14T06:00:00Z"}' \
    --filter-configuration '{"Filters":[{"Scope":"METADATA","Pattern":"{\"eventType\":[\"OrderPlaced\"]}"}]}' \
    --invoke-configuration '{"TargetArn":"'"$ANALYTICS_ARN"'","RoleArn":"'"$ROLE_ARN"'"}'

A starting position is either LATEST or POINT_IN_TIME. Choosing POINT_IN_TIME then needs a point-in-time configuration: a PointType of TIMESTAMP with a starting point, or HORIZON to begin at the earliest event still retained. An optional end point stops the read at a chosen time, which is what you want when reprocessing a known-bad window rather than catching up to live traffic.

Two things to keep in mind. The starting position is fixed when the Subscriber is created, so reading a different window means a new Subscriber. Treat the starting position as part of a Subscriber’s identity rather than a dial to turn later. And retention cannot reach back beyond the retention window, so the read starts at the earliest retained event however far back the timestamp asks for.

Order, where order matters

In event-driven architectures, where components are built to work asynchronously, the order events arrive in usually does not matter. There are still use cases where a consumer relies on ordered delivery, and the new Custom event bus offers it as an option on individual Subscribers.

Ordering is scoped by a key the publisher sets. A publisher includes an event group ID (a customer ID, an order ID, a driver ID), and a Subscriber created with FIFO delivery type receives the events for each group in the order they were published. A FIFO Subscriber reading events published without a group ID has nothing to sequence by, so the two sides work together. Creating one takes the same call as an unordered Subscriber, with the delivery type set to FIFO:

aws eventsv2 create-subscriber \
    --name inventory-ordered \
    --event-bus-arn "$BUS_ARN" \
    --type FIFO \
    --filter-configuration '{"Filters":[{"Scope":"METADATA","Pattern":"{\"eventType\":[\"OrderPlaced\"]}"}]}' \
    --invoke-configuration '{"TargetArn":"'"$FIFO_QUEUE_ARN"'","RoleArn":"'"$ROLE_ARN"'","SqsParameters":{"MessageGroupId":"{% $events.SystemMetadata.EventGroupId %}","MessageDeduplicationId":"{% $events.SystemMetadata.DeduplicationId %}"}}'

Ordering is per group, so throughput scales with the number of groups. If an event cannot be delivered, it holds up the rest of its own group while other groups keep moving. Choosing the key therefore matters: one that maps to a business entity, such as an order or a customer, gives sequencing where it is needed and independence everywhere else. A key so broad that most events share it puts them all in a single sequence, and a key so specific that every event has its own leaves nothing to order.

Because ordering is set on each Subscriber, consumers of the same events do not need to agree on it. An inventory service can receive a group’s events in sequence while an analytics service subscribing to those same events takes them as they arrive.

Reshaping events, and delivering directly to a target

A consumer’s business logic expects events in a particular shape, and the events on the bus are not always in that shape. Where the two get reconciled is an ownership decision: inside the consumer, where it becomes part of that team’s code, or on the Subscriber, ahead of it.

The first case is reformatting. A downstream system, often owned by another domain or outside the organization entirely, expects a different structure from the one the publisher emits. A JSONata transformer on the Subscriber produces that structure before delivery, so the consumer receives what it already expects. The business logic stays where it belongs, and when the published shape changes upstream, or another event type needs deriving into the same input, it is the transformer that changes rather than the consumer:

--transformer '{
    "Type":"JSONATA",
    "JsonataConfiguration":{
        "Expression":"{% {\"orderRef\": $events.Data.detail.orderId, \"total\": $events.Data.detail.amount} %}"
    }
}'

The transformer type determines the shape of what gets delivered. RAW delivers the event payload as is and is the default, so a Subscriber with no transformer configuration receives only the payload. WITH_METADATA adds the event envelope alongside it, and JSONATA reshapes it with an expression wrapped in {% %}.

The transformation reshapes events only for the Subscriber that owns it and does not affect what other Subscribers of the same bus receive. That also makes it a data minimization control: a partner can receive only the fields it needs rather than a whole internal event. Defining it at the Subscriber means it holds for every event without anyone remembering to strip fields.

The second case is calling an AWS service API. A Subscriber delivers directly to targets including Amazon Simple Queue Service (Amazon SQS), Amazon Simple Notification Service (Amazon SNS), AWS Lambda, and Amazon Kinesis Data Streams. For other services it has been common practice to add a proxy step whose only job is to make the call. With universal targets, the new Custom event bus can call a supported AWS service API directly, with the request body built by a JSONata expression.

TargetArn: arn:aws:events:::aws-sdk:dynamodb:putItem
UniversalTargetParameters.Input:
{% { "TableName": "orders", "Item": { "pk": { "S": $events.Data.detail.orderId } } } %}

Note that a universal target shapes its input through that parameter rather than through the preceding transformer, and setting a transformer on one is rejected when the Subscriber is created. The two mechanisms do the same kind of work on different targets.

That removes the proxy processing that existed only to make the call. The delivery role still needs the action the target requires and getting that wrong is the most common cause of a Subscriber that looks healthy and delivers nothing.

Conclusion

Running an event-driven application across many accounts no longer means running many event buses and the forwarding between them. A platform team creates one new Custom event bus, shares it across the organization through AWS Resource Access Manager or a resource policy on the bus, and keeps one place to decide who publishes and who consumes. Application teams create and own their Subscribers without waiting for provisioning. Ingestion and delivery are charged separately, so each team’s usage appears on its own bill, and the duplicated ingestion that forwarding hops created disappears with the hops.

The capabilities that used to send individual teams elsewhere now sit on the same bus. Ordering is per Subscriber and scoped by a publisher-supplied key, so one team’s sequencing requirement no longer fragments an architecture. Retention makes it possible to onboard a consumer that needs history it was never subscribed to. Avro and Protobuf are decoded by the bus, so producers keep their binary contracts. Transformation and universal targets keep business logic where it belongs, removing the proxy steps that existed only to reshape an event or make an API call.

Because the new Custom event bus runs alongside the Custom event bus – classic, adoption is incremental. Point one new consumer at a shared bus or forward a slice of an existing bus into it and move the rest as teams are ready.

Next steps. Create a bus, add a Subscriber, and publish an event, starting from the Amazon EventBridge documentation for the resource model and the AWS Command Line Interface (AWS CLI) reference. If you already run Custom event buses, the migration guidance covers routing existing events into a new Custom event bus without changing producers. From there, look at the Subscriber logging and metrics options for tracing an event from publication to delivery, and at AWS Resource Access Manager for how sharing and permissions work across an organization. If you have questions or feedback about the new Custom event bus, leave a comment on this post. We’d like to hear how you’re using it.

Improving Lambda function latency with scalable network bandwidth

Post Syndicated from Rahul Shandilya original https://aws.amazon.com/blogs/compute/improving-lambda-function-latency-with-scalable-network-bandwidth/

AWS Lambda now supports scalable network bandwidth for functions configured with 2,048 MB of memory or more, running outside of a virtual private cloud (VPC). Previously, sustained network throughput was capped at 625 Mbps regardless of your function’s memory configuration. Now, sustained throughput scales proportionally from 625 Mbps at configurations below 2,048 MB up to 3,000 Mbps at 10,240 MB, increasing the rate at which data moves to and from your execution environment.

In this post, you learn how to apply this new capability to latency-sensitive data processing workloads, helping reduce function execution times and per-invocation costs while improving the end-user experience through reduced latency. You also walk through a deployable implementation that demonstrates the performance improvements this capability unlocks.

Latency-sensitive data processing

Latency-sensitive data processing applications are data processing workloads that must be completed in a defined period of time. They often experience bursty, ad hoc traffic patterns while being required to download gigabytes or even terabytes of data from a data store, process it in a compute environment, and return a result to a waiting end user.

Latency-sensitive data processing is often highly parallelizable. Data can be divided into smaller pieces with each piece being individually processed before combining them together to obtain a result.

These workloads can be found in multiple industries and verticals. Examples include:

  • Log querying engines – An end user initiates an on-demand search across terabytes of log data and expects results within seconds.
  • Insurance underwriting – A prospective customer submits an application, triggering real-time evaluation of historical claims and risk data. The underwriting process determines what coverage and premiums to offer to the prospective customer.
  • Financial ETL pipelines – An economic announcement triggers an unexpected burst of market data that must be ingested, transformed, and made available to downstream trading systems before the next market tick.
  • Genomics platforms – A clinician orders a diagnostic test, requiring gigabytes of DNA or RNA sequencing data to pass through a bioinformatics pipeline and be compared against a reference genome while the patient awaits results.

These workloads are challenging to build on traditional compute clusters. Their spiky and unpredictable nature forces you to choose between under-provisioning compute to optimize costs (and risk missing your SLA) or over-provisioning and paying for idle capacity.

Why Lambda fits latency-sensitive data processing

Lambda eliminates this tradeoff. Instead of pre-provisioning a compute cluster, Lambda scales compute capacity in response to incoming requests, matching processing power to unpredictable traffic patterns. Because latency-sensitive data processing is highly parallelizable, the ability of Lambda to rapidly scale out execution environments makes it a natural fit. You can fan out across thousands of concurrent functions to process data in parallel, paying only for the compute you use.

However, as data volume and performance requirements grow, network bandwidth to and from the compute environment can become the limiting factor in minimizing workload latency.

Scalable network bandwidth directly addresses this limitation by raising the per-environment network throughput ceiling, improving the rate at which data can be transferred to and from the execution environment. Each execution environment can now drive up to 3,000 Mbps of sustained throughput when configured with 10,240 MB of memory, a 4.8x increase from the previous ceiling of 625 Mbps. Combined with the ability of Lambda to scale out at a rate of 1,000 execution environments every 10 seconds, you can download more than 3 TB of data in under 10 seconds.

New network throughput behavior for Lambda functions

Scalable network bandwidth applies to both data ingress to and egress from an execution environment for functions outside of a VPC. For functions configured with 2 GB of memory or more, network bandwidth scales by approximately 280 Mbps increments for every 1 GB of additional memory allocated.

The following table shows the maximum sustained bandwidth available to each execution environment at each memory configuration.

Memory Configuration Max Sustained Bandwidth
Less than 2,048 MB 625 Mbps
2,048 MB 765 Mbps
3,072 MB 1,044 Mbps
4,096 MB 1,324 Mbps
5,120 MB 1,603 Mbps
6,144 MB 1,883 Mbps
7,168 MB 2,162 Mbps
8,192 MB 2,441 Mbps
9,216 MB 2,721 Mbps
10,240 MB 3,000 Mbps (4.8x increase)

Table 1. Lambda sustained network bandwidth by memory configuration. Bandwidth scales at ~280 Mbps per additional GB of memory above 2 GB.

In the following section, you learn how scalable network bandwidth improves end-user latency by building an ETL pipeline that demonstrates it. You can find the source code in the GitHub repository.

Solution overview

Consider a SaaS analytics platform where users submit ad hoc queries against a data store. The application must extract the relevant data, apply a filter or transformation, and return an aggregate result while the user waits. In this example, the result needs to be returned in 8 seconds or less.

The following diagram illustrates the architecture of the solution.

ETL fan-out architecture: a client calls an orchestrator Lambda function, which fans out to multiple worker Lambda functions that read data in parallel from Amazon S3, with bandwidth scaling callouts for each memory tier.

Figure 1. ETL fan-out pattern: an orchestrator Lambda function distributes work to multiple worker Lambda functions that read from Amazon S3 in parallel, with bandwidth scaling callouts per memory tier.

A client initiates an ad hoc query by calling the orchestrator Lambda function through the Lambda API. The orchestrator function determines how to split the work. To process the data in parallel, the orchestrator function uses a ThreadPoolExecutor to issue synchronous invoke requests to the Lambda worker function, fanning out the worker across multiple execution environments at the same time.

Each Lambda worker function is configured with 10,240 MB of memory, so it has access to up to 3,000 Mbps of sustained network throughput. After the data is processed, the aggregated result is returned to the client.

Prerequisites

Before you start the deployment process, make sure that you have completed the following steps:

  1. Install the AWS SAM CLI on your computer and confirm that you are running Python 3.12 or later.
  2. Have your AWS account credentials ready.
  3. Submit a request to AWS Service Quotas to turn on scalable network bandwidth for your Lambda functions. This quota is listed under Network bandwidth per execution environment.

Clone the source code from the GitHub repo and deploy the application within your AWS account. Creating the 10 GB test dataset and running the benchmark can incur charges to your AWS account.

git clone https://github.com/aws-samples/sample-lambda-enhanced-bandwidth
cd sample-lambda-enhanced-bandwidth
sam build
sam deploy --guided

After the AWS CloudFormation stack is deployed, record the DataBucketName and orchestrator function name from the stack outputs to use in subsequent commands.

To simulate data for the end user to query, the GitHub repo has a script that creates 10 GB of synthetic data and uploads it to your S3 bucket.

Mode 1: Processing pre-partitioned data

In Mode 1, the 10 GB of synthetic data is pre-partitioned. Pre-partitioned data is typically produced incrementally by many sources over a period of time, which can be the case with IoT data or access logs. The following command creates 10 GB of data divided into 20 partitions that are 512 MB each.

python scripts/generate_data.py \
    --bucket <DATA_BUCKET_NAME> \
    --total-gb 10 \
    --chunk-mb 512

Turning on scalable network bandwidth does not, on its own, make your downloads faster. A single download request only opens one connection to Amazon S3, and one connection does not move data fast enough to fill all the bandwidth now available to your Lambda function. To actually use your full allotment of network bandwidth, the execution environment has to pull the data over several connections at once. It does this by preferring the AWS Common Runtime (CRT) transfer client, a high-performance download engine built into Boto3. When the worker calls download_fileobj, the CRT client automatically breaks the 512 MB object into smaller parts and downloads them in parallel across multiple Amazon S3 requests. Those parallel downloads are what let a single Lambda worker take advantage of its full network bandwidth.

The following command runs the benchmark on the pre-partitioned data.

python scripts/run_fanout_benchmark.py \
    --orchestrator-name <STACK_NAME>-orchestrator \
    --bucket <DATA_BUCKET_NAME> \
    --iterations 5

Mode 2: Processing single large objects

Mode 2 generates 10 GB of data in one large object. This arrangement is more common when data is produced or delivered as one complete unit, such as database backups or genomic datasets. The following command creates 10 GB of data in a single large object.

python scripts/generate_data.py \
    --bucket <DATA_BUCKET_NAME> \
    --single-object-gb 10 \
    --key large-object/large-file.bin

In Mode 1, the CRT preference applies to Boto3 managed transfer methods such as download_file and download_fileobj. Mode 2 takes a different approach. Each worker reads a specific byte range of a single large object using get_object. The CRT preference setting has no effect on these calls. Instead, you can control concurrency by explicitly tuning the number of Lambda workers and using a bounded ThreadPoolExecutor to issue multiple byte-range requests at the same time.

When you run the following command, the orchestrator takes the single large object and divides it into consecutive byte ranges of 512 MB each. Each of the individual ranges is then processed by a Lambda worker execution environment in parallel.

python scripts/run_fanout_benchmark.py \
    --orchestrator-name <STACK_NAME>-orchestrator \
    --bucket <DATA_BUCKET_NAME> \
    --key large-object/large-file.bin \
    --slice-size-mb 512 \
    --iterations 5

The benchmark reports wall-clock duration, client-observed duration, aggregate throughput across workers, worker completion counts, and target compliance. When comparing memory configurations, keep the code, dataset, AWS Region, partition count, warm-up policy, and measurement count identical. You should run the benchmark multiple times in your account because placement, cold starts, concurrency, S3 behavior, and execution-environment reuse could affect results.

Results

To compare results, we ran the benchmark using a baseline configuration where the worker Lambda function is configured with only 1,024 MB of memory, well below the 2,048 MB threshold required for scalable network bandwidth to take effect. The 1,024 MB configuration limits network throughput to the previous sustained ceiling of 625 Mbps.

In our baseline test run, a worker downloaded and processed a single 512 MB partition with a 6.61-second download time at a 649.8 Mbps throughput (at p50). The 649.8 Mbps throughput exceeds the 625 Mbps ceiling because Lambda is capable of bursts in network throughput over a short period of time. Across twenty measured fan-out queries, the complete 10 GB query was completed with a 7.113-second wall-clock at p50. This fits within the 8-second SLA but leaves very little headroom.

To run our scalable network bandwidth benchmark, we re-deployed our worker Lambda function with a 10,240 MB memory configuration and re-ran the application. At a 10,240 MB memory configuration, each execution environment can now access up to 3,000 Mbps in sustained throughput. Direct 512 MB downloads achieved a 1.70-second download time and 2,521.3 Mbps throughput (both at p50). The complete 10 GB query was completed with a 2.640-second wall-clock at p50. That is 2.69 times faster, or 62.9% lower median latency, than the 1,024 MB configuration.

Table 2 summarizes the direct worker and end-to-end fan-out measurements for the same 10 GB dataset and 20 × 512 MB orchestration pattern. Aggregate throughput is the total data transfer rate across all twenty execution environments spun up to run the benchmark.

Memory Configuration Single 512 MB partition download time and throughput (p50) 10 GB fan-out wall time (p50) Aggregate throughput p50
1,024 MB baseline tier (sustained 625 Mbps) 6.61s / 649.8 Mbps 7.113s 12.08 Gbps
10,240 MB scalable tier (up to 3,000 Mbps) 1.70s / 2,521.3 Mbps 2.640s 32.54 Gbps

Table 2. Measured 1,024 MB baseline tier and 10,240 MB scalable bandwidth performance for a 10 GB fan-out ETL query.

Using scalable network bandwidth, the customer’s SLA headroom has improved by nearly 5 seconds. The Lambda function can now handle larger partitions within the same SLA window, reducing costs while still remaining comfortably within the customer’s SLA.

Clean up

To clean up the resources you created for the benchmark test, run the following commands:

aws s3 rm s3://<DATA_BUCKET_NAME> --recursive
sam delete --stack-name <STACK_NAME>

Best practices

After scalable network bandwidth is turned on for your AWS account, the following practices help you get the most out of it.

Profiling and planning

  • Test before you tune. Not every function is network-bound. Before increasing memory, profile your function to confirm that network I/O is the primary contributor to invocation duration and not CPU or application logic. Use Amazon CloudWatch Lambda Insights to inspect rx_bytes, tx_bytes, and duration. Functions where network I/O dominates invocation time are prime candidates for tuning.
  • Design for parallelism. Break your data into parallelizable chunks that can be processed independently in a fan-out pattern across multiple execution environments. You can use Amazon S3 byte-range reads to split large files into independently downloadable partitions. For implementation details, see Downloading an object with part numbers in the Amazon S3 User Guide.
  • Run AWS Lambda Power Tuning. Lambda Power Tuning is a state machine that helps you optimize your Lambda functions for cost and performance. Use Power Tuning to sweep memory configurations from 1,024 MB to 10,240 MB and identify the optimal cost-vs-latency point for your workload.

Implementation

  • Check upstream and downstream limits. Check the throughput limits of your data sources. For example, a Lambda function running at 3,000 Mbps can exceed the throughput capacity of a single S3 prefix, which supports up to 5,500 GET requests per second. When this happens, you will see HTTP 503 (Slow Down) errors in your application logs. Distribute your S3 objects across multiple prefixes to parallelize reads and avoid per-prefix throttling.
  • Balance bandwidth and CPU. Lambda allocates CPU proportionally to memory. For example, at a 1.7 GB memory configuration you are allocated 1 vCPU while a 10 GB memory configuration is allocated up to 6 vCPU. If your function processes data in parallel threads, the higher memory tiers give you both more network bandwidth and more CPU to process it. Use the concurrent.futures module in Python or worker_threads in Node.js to process data across parallel threads and maximize both CPU and network utilization.
  • Turn on Amazon S3 CRT for Boto3. If your function uses the Python runtime, initialize your Amazon S3 client with preferred_transfer_client: 'crt' to maximize single-connection throughput. The AWS Common Runtime automatically parallelizes requests across multiple TCP connections, which matters because individual TCP connections have a throughput ceiling.
  • Use SnapStart for JVM workloads. If you use Lambda SnapStart for Java functions, scalable network bandwidth reduces afterRestore hook latency. Network activity that occurs during function restore, such as pre-warming connections or pre-fetching configuration data, can complete faster.

Conclusion

Scalable network bandwidth raises the per-environment sustained throughput ceiling of AWS Lambda from 625 Mbps to 3,000 Mbps, directly reducing end-to-end latency for data-intensive workloads. Combined with the Lambda scaling rate, you can now move terabytes of data in seconds, without provisioning or managing infrastructure.

To get started, request the Network bandwidth per execution environment quota increase through AWS Service Quotas and deploy the sample application from the GitHub repository to see the improvement firsthand.

How Property Finder automated incident management with AWS DevOps Agent

Post Syndicated from Nada Tlohi original https://aws.amazon.com/blogs/devops/how-property-finder-automated-incident-management-with-aws-devops-agent/

When a production service starts saturating the CPU at 1 AM, every minute counts for incident management. For Property Finder, a production incident could mean failed searches, frustrated users, and direct revenue impact. Property Finder is the leading property portal in the Middle East and North Africa (MENA), serving millions of property seekers across five markets.

Before adopting AWS DevOps Agent, incident response followed a familiar pattern: an alert fires, an on-call engineer wakes up, spends 20–40 minutes correlating metrics across tools, manually documents findings, and opens a fix. Mean Time to Resolution stretched to 2–3 days for non-critical issues.

Today, that entire workflow runs autonomously. From alert to root cause analysis, Slack notification, Jira ticket, on-call phone call with context, and auto-remediation pull request (PR), the full lifecycle completes in 14 minutes. This post walks through the implementation and shows how a separate custom agent that automatically generates code fixes is the key differentiator.

The business problem

Property Finder runs a distributed microservices architecture on Amazon Elastic Container Service (Amazon ECS) fronted by Application Load Balancers (ALBs). When infrastructure issues occur, the impact is immediate: users see failed searches, agents cannot update listings, and revenue is directly impacted during peak hours.

The traditional workflow had three gaps:

  1. Detection lag. Non-critical anomalies could go undetected for days.
  2. Context switching. Engineers bounced between five or more tools per incident.
  3. Knowledge silos. Runbooks lived in people’s heads, not automation.

Solution architecture

Property Finder’s implementation connects AWS DevOps Agent at the center of a three-tier pipeline: Detection and Trigger, Autonomous Investigation, and Event-Driven Output.

Three-tier incident pipeline from a CloudWatch alarm through AWS DevOps Agent investigation to Slack, Jira, and GitHub outputs

Figure 1: End-to-end autonomous incident management architecture

The numbered steps correspond to the data flow in Figure 1:

  1. ECS CPU spike triggers an Amazon CloudWatch Alarm. CloudWatch Metrics Insights monitors service health across all ECS clusters. When sustained CPU exceeds 98%, the alarm transitions to ALARM state.
  2. AWS Lambda formats and HMAC-signs the payload. Triggered directly by the CloudWatch alarm action (which fires only on ALARM state transitions), AWS Lambda enriches the payload with service metadata, signs it with HMAC-SHA256 using credentials from AWS Secrets Manager, and POSTs to the webhook.
  3. The agent begins autonomous investigation. Parallel subagents query ECS metrics, AWS CloudTrail, ALB traffic patterns, and Grafana telemetry (Prometheus, Loki, Pyroscope). The agent reads relevant source code from GitHub for correlation.
  4. Findings post to Slack in real time. The native Slack integration posts investigation progress to #incidents. The full root cause analysis, impact assessment, and mitigation plan appear at the end of the thread.
  5. Investigation Completed event fires to Amazon EventBridge. Amazon EventBridge triggers an orchestrator Lambda that fans out to three independent targets simultaneously.
  6. Lambda creates a Jira ticket with the full root cause analysis. The Lambda retrieves the investigation summary from journal records and creates a prioritized ticket with root cause, severity, and affected service.
  7. Grafana IRM pages the on-call engineer by phone. A Lambda posts a Grafana Alerting-compatible payload to the IRM webhook. The escalation chain calls the engineer with full investigation context: what broke, why, and the recommended fix.
  8. The remediation agent opens a GitHub PR with the auto-fix. It receives the root cause, generates a Terraform or code fix, and opens a Draft PR through a GitHub Model Context Protocol (MCP) server. Engineers review before merging.

A real incident

The example-service, Property Finder’s core property search microservice serving millions of queries per day across five MENA markets, experienced CPU saturation at 99.11%. The pipeline resolved it end-to-end in 14 minutes.

1:21 AM │ Alarm fires (ECS CPU > 98%)

1:22 AM │ Investigation starts + Slack posted

1:22 AM │ 4 parallel subagents launched

1:32 AM │ Root cause identified

1:33 AM │ Jira ticket [redacted] created

1:34 AM │ On-call paged via phone call

1:35 AM │ GitHub PR [redacted] opened with fix

The detection Lambda handles three tasks: (1) retrieves the webhook secret from AWS Secrets Manager, (2) enriches the CloudWatch alarm event with ECS service metadata (cluster name, service name, task count), and (3) HMAC-signs the payload before POSTing to the webhook. The key authentication pattern:

# HMAC-SHA256 signing for webhook authentication
ts = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%S.000Z")
body = json.dumps(incident)
sig = hmac.new(webhook_secret.encode("utf-8"),
               f"{ts}:{body}".encode(), hashlib.sha256).digest()
http.request("POST", webhook_url, body=body,
             headers={"x-amzn-event-timestamp": ts,
                      "x-amzn-event-signature": base64.b64encode(sig).decode()})
Four parallel subagents querying ECS, CloudTrail, ALB, and Grafana data sources during the investigation

Figure 2: Four parallel subagents investigating ECS, CloudTrail, ALB, and Grafana data sources simultaneously

Root cause: Conflicting CPU and memory target-tracking autoscaling policies combined with an insufficient capacity floor. The service had both a CPU policy (target 70%) and a memory policy (target 75%). Actual memory usage sat at 3–8%, creating a persistent conflict between the two policies.

With MinCapacity set too low, the service could not sustain the task count needed to absorb CPU load. The resulting instability (22+ scaling flips observed) prevented stable scale-out, leaving the service effectively pinned at two tasks with no CPU headroom.

This is a common organizational issue: teams configure both scaling dimensions without realizing the interaction, especially when the capacity floor is not sized for baseline traffic. The agent identified the pattern in 10 minutes, a task that typically requires senior engineers with deep scaling expertise and hours of CloudWatch metric correlation.

Investigation output naming conflicting CPU and memory autoscaling policies as the root cause

Figure 3: Root cause analysis identifying the conflicting autoscaling policy

Slack incidents channel message linking to the running investigation at 1:22 AM

Figure 4: Slack notification with investigation link posted at 1:22 AM

Auto-created Jira ticket showing priority, root cause, and affected service

Figure 5: Jira ticket [redacted] auto-created with priority, root cause, and affected service

At 1:34 AM, the on-call engineer received a phone call through Grafana IRM with the complete investigation context. No need to wake up and hunt for root cause across dashboards.

Grafana IRM escalation chain routing the alert to the on-call engineer

Figure 6: Grafana IRM escalation chain routing the alert and calling the on-call engineer

Incoming on-call phone call at 1:34 AM carrying the investigation context

Figure 7: Incoming phone call at 1:34 AM with investigation context

Mitigation plan generated: (1) Remove the memory-based scaling policy, (2) raise MinCapacity to handle baseline traffic, (3) implement CPU-only target tracking at 70%. This plan was passed to a separate custom agent for remediation.

Remediation

Remediation is the key differentiator in this pipeline. It is a dedicated remediation agent (pr-creation-agent) invoked only after investigation completes. AWS DevOps Agent enforces read-only access to infrastructure through a per-session permission guardrail. Effective permissions are the intersection of the execution role’s IAM policy and the guardrail, and write actions are excluded.

The split separates concerns: investigation stays within that read-only envelope, whereas the remediation agent is scoped to a GitHub MCP server as its only external integration. Safety at the remediation layer does not rely on the agent’s built-in directed actions approval mechanism. Instead, two controls enforce the boundary. First, the remediation agent is a separate, narrowly scoped agent with access limited to GitHub MCP. Second, every output is a Draft pull request that requires human review and merge before taking effect. The GitHub MCP connection is authenticated with a fine-grained personal access token scoped to the specific infrastructure repositories, with an expiration and rotation policy. No elevated IAM role or additional agent permissions are required.

How it works

When the “Investigation Completed” Amazon EventBridge event fires, a Lambda orchestrator invokes the remediation agent with the investigation ID. The agent then:

  1. Reads findings from journal records to understand the root cause and recommended fix.
  2. Maps the AWS account to the correct repository. Property Finder has six infrastructure repos for different teams (B2B, B2C, core-platform, growth, data-engineering, shared-infra). The agent extracts the account ID from resource ARNs and routes to the right repo. This mapping is validated through automated tests and updated as new accounts or repositories are onboarded.
  3. Checks for duplicate PRs by searching existing PR titles and bodies for the investigation ID. If a matching PR exists, it reports the URL and exits without creating a duplicate.
  4. Reads the relevant Terraform files through GitHub MCP (GITHUB-MCP_get_file_contents), identifies the exact changes required, and plans the fix.
  5. Creates a feature branch (fix/{investigation_id}), commits the changes, and opens a Draft PR with a structured template including problem summary, root cause, changes made, and a testing checklist.

AWS also supports remediation through Kiro CLI with AWS CodeBuild or Kiro-ready prompts. Property Finder chose an approach that fits their multi-team repository structure: the remediation agent runs entirely within the Agent Space (the managed environment where custom agents execute), uses GitHub MCP for repository access, and maps multiple repositories to different teams automatically.

The orchestrator Lambda is triggered by the “Investigation Completed” Amazon EventBridge event. It first retrieves the investigation findings from journal records, then fans out to three targets simultaneously. Target one creates a Jira ticket with the full root cause analysis, severity, and affected service. Target two posts a Grafana Alerting-compatible payload to the Grafana IRM webhook to trigger phone call escalation. Target three invokes the remediation agent through the CreateChat and SendMessage API, passing the investigation ID and root cause context so it can generate the appropriate code fix.

Draft GitHub pull request with a problem summary, root cause, and changes template

Figure 8: GitHub PR [redacted] generated by the remediation agent with a structured problem, root cause, and changes template

Terraform diff replacing the memory scaling policy with a CPU-only target-tracking policy

Figure 9: Terraform diff showing the new CPU-only scaling policy replacing the conflicting memory configuration

The PR is always opened as Draft. Engineers review, run terraform plan, validate in staging, and merge. The agent never auto-merges.

Results

Metric Before After Improvement
End-to-end time Hours to days 14 minutes >88% reduction
Investigation 20 to 40 min (manual) 10 min (autonomous) 50–75% reduction
Documentation Manual, incomplete Auto-generated root cause analysis + Jira 100% documented
Remediation Manual PR by engineer Auto-fix PR + review Minutes to code fix

Cost considerations: Each incident invokes two agent sessions (investigation + remediation) with up to four parallel subagents. Billing is based on agent minutes. For detailed pricing, see the AWS DevOps Agent pricing page. We recommend reviewing pricing for all services used in this architecture.

“We now rely fully on AWS DevOps Agent to identify infrastructure-related issues. It has helped us identify multiple complex issues without even opening a support ticket. Even if we had raised tickets, it would likely have taken support engineers hours to find the root cause, whereas we resolved these issues in minutes.”

— Yasitha Bogamuwa, Cloud Engineering Manager, Property Finder

Getting started

Prerequisites:

  1. An Agent Space configured in your account.
  2. Amazon CloudWatch and AWS CloudTrail enabled for observability.
  3. Slack, Grafana, and GitHub connected as capabilities.
  4. Infrastructure resources tagged for topology mapping.

Step 1: Configure the webhook trigger. Set up CloudWatch Alarm action to invoke a Lambda function. The Lambda enriches the payload, HMAC-signs it, and POSTs to your Agent Space webhook endpoint.

Step 2: Set up event-driven outputs. Create an Amazon EventBridge rule for “Investigation Completed” events (source: aws.aidevops). Add Lambda targets for Jira, Grafana IRM, and optionally a remediation custom agent.

Step 3: Test end-to-end. Trigger a test alarm and verify the full pipeline: investigation starts, Slack posts, Jira ticket created, on-call paged, and PR opened.

For a similar integration pattern with Salesforce, see Automating Incident Investigation with AWS DevOps Agent and Salesforce MCP Server on the AWS DevOps Blog.

Clean up

This post describes an architecture pattern implemented by Property Finder. If you deployed test resources while following along, remember to delete any CloudWatch Alarms, Lambda functions, Amazon EventBridge rules, and Agent Space configurations to avoid ongoing charges. For a full list of resources and associated costs, review the pricing pages for each AWS service used in this architecture.

Conclusion

Property Finder’s implementation shows that autonomous incident management works in production today, with their pipeline running since early 2026. The agent never auto-merges. Human review remains in the loop by design: the agent accelerates, the engineer decides. The on-call engineer wakes up to a phone call with the root cause already identified, a Jira ticket filed, and a PR ready for review.

Explore the AWS DevOps Agent documentation to get started with your own autonomous pipeline.

  1. Getting Started with AWS DevOps Agent.
  2. Automating Incident Investigation with Salesforce MCP.
  3. Building an End-to-End Agentic SRE.
  4. Amazon EventBridge User Guide.
  5. Grafana IRM Documentation.

About the authors

Nada Tlohi

Nada Tlohi

Nada is a Technical Account Manager at AWS based in Dubai, UAE. She helps strategic enterprise customers across the MENA region transform their cloud operations and improve system reliability by adopting AIOps, incident automation, and DevOps best practices.

Conor Manton

Conor Manton

Conor is a Principal Technical Account Manager at AWS, based in San Francisco. He works with strategic enterprise customers to accelerate their cloud journey, with a focus to operationalize AI-powered workflows to drive business outcomes.

Jaydeep Singh

Jaydeep Singh

Jaydeep is a Senior DevOps Engineer at Property Finder. He specializes in designing and operating scalable cloud infrastructure, containerized platforms, and Kubernetes ecosystems. He leads platform reliability, infrastructure automation, and continuous integration and continuous delivery (CI/CD) initiatives, so engineering teams can build and deploy applications securely, efficiently, and at scale.

Git v2.56.0 released

Post Syndicated from jake original https://lwn.net/Articles/1097213/

Version 2.56 of the Git distributed
version-control system has been released. It has 748 non-merge commits
since Git 2.55 was released back in
June; those commits came from 104 developers, 39 of whom are first-time
contributors. New features include a safer workflow for conflict
resolution, smaller path-walk repacks, a new git history drop
sub-command, and much more. LWN looked at Git
2.56
recently and the GitHub blog has a lengthy
look at 2.56
as well.

AWS European Sovereign Cloud: Demonstrating an independent operation

Post Syndicated from Stéphane Israël original https://aws.amazon.com/blogs/security/aws-european-sovereign-cloud-demonstrating-an-independent-operation/

On Saturday, October 24, 2026 we will conduct an exercise, demonstrating that the AWS European Sovereign Cloud can operate without depending on any infrastructure outside of the European Union (EU).

For several hours, the AWS European Sovereign Cloud will operate without a connection to the AWS Global Network backbone. The backbone is the private network that moves authorized AWS operational data between AWS locations without using the public internet. During the exercise, this traffic will securely reroute over the public internet.

The exercise will not affect service availability within the AWS European Sovereign Cloud, other AWS Regions, or private connectivity through AWS Direct Connect. Customers may experience brief connectivity disruptions as traffic moves onto a separate network route at the beginning or the end of the exercise, after which normal connectivity resumes.

The operational team, composed entirely of EU residents within the EU, will execute the exercise using only the hardware and software resources of the AWS European Sovereign Cloud. The AWS European Sovereign Cloud Managing Directors called for this exercise to showcase its operational independence.

An independent cloud for Europe

The AWS European Sovereign Cloud is a new, independent cloud for Europe. Located in Brandenburg, Germany, its data centers are physically and logically separate from other AWS Regions, with a local in-EU copy of the source code. All customer content and customer-created metadata stay in the EU. It has no critical dependencies on non-EU infrastructure and is operated exclusively by EU residents.

In standard operations, the AWS European Sovereign Cloud uses two global systems. The first is the AWS Global Network backbone. The second is a dedicated system that the local EU team controls and supervises to securely manage limited, controlled transfers of operational AWS data.

Neither is operation-critical, and neither affects the sovereignty assurance of the AWS European Sovereign Cloud. The AWS European Sovereign Cloud can operate independently at any time without a connection to these global systems, and on October 24 that’s what the team will demonstrate.

Built to meet regulatory standards

This exercise will produce verifiable technical and operational evidence that the AWS European Sovereign Cloud can operate independently within the EU. The exercise is designed to be consistent with the objectives of the European Commission’s EU Cloud Sovereignty Framework (CSF) and the criteria of the C3A framework from Germany’s Federal Office for Information Security (BSI). These frameworks set out objectives and criteria for assessing whether cloud services can be provided independently and autonomously.

We designed the AWS European Sovereign Cloud for regulated customers and the public sector across the EU. The AWS European Sovereign Cloud: Sovereign Reference Framework (ESC-SRF) gives our customers and partners a comprehensive set of evidence points, maps to controls, artifacts, and other elements regulators and compliance authorities need to accelerate their adoption of the AWS European Sovereign Cloud. The results of this exercise will provide additional evidence for their compliance and assurance packages.

Standalone and fully secure

The AWS European Sovereign Cloud runs connected to the AWS Global Network backbone because it delivers superior performance, capacity, reliability, security, and cost savings to customers. That includes always-on encryption and distributed denial of service (DDoS) defenses; and the backbone can’t decrypt or see the encrypted data that AWS European Sovereign Cloud customers send and receive. While the backbone delivers these benefits day-to-day, the AWS European Sovereign Cloud can continue to operate independently, with the appropriate security controls in place.

During the exercise, instead of using the AWS Global Network backbone, the AWS European Sovereign Cloud will exclusively use its dedicated internet connectivity from European internet service providers. This will provide connectivity to the worldwide internet.

Whenever traffic moves between internet links, there’s a small window of limited disruption called convergence, a short time when other non-AWS networks change their routing information to reflect the change. This could happen at the beginning of the exercise, when traffic moves to dedicated AWS European Sovereign Cloud internet providers, and at the end of the exercise, when traffic moves back to the AWS Global Network backbone.

Customer data stays in the EU

AWS has committed to not moving AWS European Sovereign Cloud customer content and customer-created metadata outside of the EU. Only certain data, which is neither customer content nor customer-created metadata, such as AWS operational data, leaves the EU. We use a dedicated system to securely manage these limited, controlled transfers under the control and supervision of the local EU team. We’re rigorous about what the system transfers. It accepts vetted source code mirroring and software updates, and transfers out very limited and approved routine information. During the exercise, the AWS European Sovereign Cloud team will disable the system entirely, confirming that the AWS European Sovereign Cloud continues to operate independently without it.

Learn more

AWS will share an update after the exercise with regulators and customers. To learn more about the AWS European Sovereign Cloud’s design and digital sovereignty controls, visit aws.eu. If you have questions about this exercise or would like to discuss how it may impact your workloads, reach out to AWS Support or contact your AWS Account team.

Stephane Israel

Stéphane Israël

Stéphane is the leader and Managing Director of the AWS European Sovereign Cloud. He is responsible for the management and operations of the AWS European Sovereign Cloud, including infrastructure, technology, and services, in addition to broader digital sovereignty efforts at AWS. Prior to AWS, he was the CEO of Arianespace, where he oversaw numerous successful space missions, including the launch of the James Webb Space Telescope.

AWS Weekly Roundup: GPT-6 Sol and Luna, Claude Opus 5.5 on Amazon Bedrock, Strands harness, and more (September 28, 2026)

Post Syndicated from Daniel Abib original https://aws.amazon.com/blogs/aws/aws-weekly-roundup-gpt-6-sol-and-luna-claude-opus-5-5-on-amazon-bedrock-strands-harness-and-more-september-28-2026/

If there’s one theme that defined last week, it’s choice. The frontier models keep arriving, and the interesting question is no longer just “how smart is it?” but “which model fits this step, at this cost, at this latency?” That’s exactly what landed on Amazon Bedrock over the past few days: GPT-6 Sol and GPT-6 Luna from OpenAI, giving you two new points on the intelligence-versus-efficiency curve, and Claude Opus 5.5 from Anthropic, the first of the Claude 5.5 family.

GPT-6 Sol is built for the demanding, recurring work of development and operations, while GPT-6 Luna makes focused, repeatable tasks practical at high volume, and both ship at significantly lower pricing than their GPT-5.6 predecessors. Claude Opus 5.5, meanwhile, does more with fewer tokens than Opus 5 and is tuned for agentic coding and long-running tasks. What I like about all three is that they push toward the same idea: match the model to the job instead of reaching for the biggest one every time. The other thread was observability catching up to this agentic world, including a launch I had the pleasure of writing about myself.

Now, let’s get into this week’s AWS news…

Last week’s launches

Here are some launches and updates from this past week that caught my attention:

  • Introducing Amazon CloudWatch Omni – You can now observe your applications and AI agents together in a single, collaborative experience. Amazon CloudWatch Omni is built on OpenTelemetry, so your existing telemetry shows up with nothing to reconfigure, and your whole team reaches it through one URL with enterprise SSO — no console access required. It auto-discovers your services, maps dependencies, and brings AWS DevOps Agent into investigation sessions to correlate signals and trace root causes. There’s a companion post on the agent-observability side, a deeper dive on the AWS Cloud Operations blog on what observability for the AI era looks like, and the announcement on What’s New with the specifics. If you want the bigger picture, Matt Wood’s Wrong, not broken is a great read on why correctness now has to be measured at the level of the run.
  • Enhanced custom event buses in Amazon EventBridge – Amazon EventBridge now offers an enhanced custom event bus purpose-built for organizations scaling event-driven applications across teams and accounts. You can now deploy a single centralized bus shared across every account in your organization through AWS RAM, with optional event ordering, a simplified Subscriber resource that bundles filtering, targets, and retries, content-based deduplication, and synchronous invocation for targets like AWS Lambda. A new ingress/egress pricing model replaces the compounding cross-account routing charges of multi-bus setups, and your existing buses keep working unchanged as “classic.”
  • Amazon SageMaker HyperPod Inference Gateway – You can now front your LLM inference on Amazon SageMaker HyperPod with a Kubernetes-native, GPU-aware routing layer that deploys as a single Amazon EKS managed add-on with zero application changes. Instead of round-robin load balancing, it routes on real-time inference signals — KV cache utilization, queue depth, prefix cache hits, predicted latency, and more — cutting first-token latency by up to 82% in mixed-hardware and bursty scenarios. It works with any OpenAI-compatible model server, including vLLM and SGLang.
  • AI agent skills for AWS End User Messaging and Amazon SES – You can now build and send messages by asking your AI coding agent in plain language. Amazon SES and AWS End User Messaging publish AI agent skills for the AWS MCP Server, giving your agent step-by-step, validated guidance for tasks like verifying a sending identity, sending a production email, or building a branded RCS agent with cards and buttons. The skills work with Claude Code, Codex, Cursor, and Kiro, so you can complete messaging workflows without hopping between docs and console screens.

For a full list of AWS announcements, be sure to keep an eye on the What’s New with AWS page.

Other AWS news

Here are some additional posts and resources that you might find interesting:

  • Introducing Strands harness – The Strands Agents team released Strands harness, a fully assembled, general-purpose agent harness you can run locally or deploy anywhere, under Apache 2.0. It takes one line of Python or TypeScript to wire up your model of choice across Amazon Bedrock, Anthropic, OpenAI, Google, or a local Ollama model, and it ships with sensible defaults for prompt caching and context management (truncating bulky tool results, compacting when the context window fills up, and keeping memory across runs). The team reports it costs about 28% less than comparable harnesses on the same models while holding accuracy steady.
  • Announcing the new AWS Reimagine report on AI – The AWS Executive in Residence team spent nine months interviewing 154 leaders across 27 countries about what separates organizations that turn AI into value from those that don’t. The report is candid (including where AI hasn’t worked at Amazon), and the recurring insight is that once building gets fast, the bottleneck moves to deciding, funding, and governing the work. Well worth a read if you’re thinking about how your teams adopt AI in practice.

For a full list of AWS blog posts, be sure to keep an eye on the AWS Blogs page.

Upcoming AWS events

Check your calendar and sign up for upcoming AWS events:

  • AWS re:Invent – AWS re:Invent returns to Las Vegas from November 30 to December 4, and session times, locations, and speakers are live. Reserved seating opens October 6, so register now and be ready to claim your spot in chalk talks, workshops, and builders’ sessions.
  • AWS Summits – With re:Invent on the horizon, the Summits are wrapping up for the year. The last stop is Dubai (September 30) at the Dubai World Trade Center, with 60+ sessions, an AWS Village, and hands-on workshops.

Join the AWS Builder Center to connect with builders, share solutions, and access content that supports your development. Browse here for upcoming AWS-led in-person and virtual events and developer-focused events. That’s all for this week. Check back next Monday for another Weekly Roundup!

— Daniel Abib

Audit trails for autonomous agents with AWS DevOps Agent

Post Syndicated from Ben Peterson original https://aws.amazon.com/blogs/devops/audit-trails-for-autonomous-agents-with-aws-devops-agent/

Autonomous agents need audit trails. AWS DevOps Agent (DevOps Agent) investigates production incidents and proposes or applies fixes on your behalf. Every operation and security review then raises the same two questions: what did the agent do, and how do you understand its impact?

AWS DevOps Agent maintains an immutable, step-by-step record of its own reasoning and actions. This post shows how to capture the agent’s full operational trail using the agent journal, recommendations, Amazon EventBridge lifecycle events, and AWS CloudTrail. We then wire them into an audit pipeline built on Amazon EventBridge, AWS Lambda, and Amazon Simple Storage Service (Amazon S3).

By the end, you will have deployable audit patterns that show, for any investigation the agent runs, what it concluded, what it recommended, when it ran, and whether the fix landed.

Why auditing an autonomous agent is different

CloudTrail records the API calls made in your account, but an autonomous agent adds reasoning that CloudTrail doesn’t capture. “The agent ran a metric query” is far less valuable than “the agent concluded the Lambda was timing out because its security group blocks egress to the database.” The latter is a decision, and that’s what an agent audit needs to capture.

The four surfaces

AWS DevOps Agent exposes four surfaces. Two capture the agent’s output, what it found and what it advises, and two capture context: when it ran, and who configured the agent and its permissions.

The agent journal

The agent journal (API) is the heart of the audit trail. For every execution, AWS DevOps Agent records an ordered, immutable log of its reasoning step, sub-agent it dispatches, observations, findings, and root-cause summary. Journal entries cannot be modified once written, making them resistant to prompt injection and trustworthy as an audit record.

aws devops-agent list-journal-records \
  --agent-space-id <id> --execution-id <execution-id>

Each record carries a recordType: symptom, observation, finding, and investigation_summary / investigation_summary_md are what matters for audit. This is the surface you archive per investigation.

Recommendations polling

Recommendations (API) are cross-incident preventative advice. The agent generates these on a schedule through a goal, and each recommendation carries a status and a version. Each evaluation run writes new records rather than updating the previous run’s, so advice that persists week over week appears as a series of records. The superseded ones remain at whatever status they last held. “The agent recommended X, the same failure recurred Y weeks later, and here is every version of that advice in between” is something you reconstruct from the archived snapshots, because the API returns current and superseded records together. Recommendations have no Amazon EventBridge event. You capture them by polling on a schedule.

aws devops-agent list-recommendations --agent-space-id <id>

Amazon EventBridge lifecycle events

Amazon EventBridge is how you capture lifecycle transitions in real time. A successful investigation produces Created, In Progress, and Completed events. Each carries the execution_id you need to fetch the journal and a summary_record_id pointing at the root-cause summary. Investigations can also end as Failed, Timed Out, or Canceled, and mitigations emit their own parallel set.

AWS CloudTrail

CloudTrail records API calls made to the AWS DevOps Agent service and stamps agent-initiated service calls: invokedBy: aidevops.amazonaws.com. It doesn’t capture the agent’s investigation reads, the metric and log queries it runs while diagnosing an incident in your account’s trail. Use CloudTrail for control-plane accountability, and the journal for behavioral audit.

IAM: Action boundary

As with anything in AWS, the agent can only do what its AWS Identity and Access Management (IAM) role permits. During an investigation, AWS DevOps Agent assumes an Agent Space role. That role’s policies are the hard ceiling on its capabilities. You can inspect it directly:

aws iam list-attached-role-policies --role-name DevOpsAgentRole-AgentSpace-<suffix>

The AWS-managed AIOpsAssistantPolicy is attached to the default role. As of policy version 15, 848 of its actions are reads except 6 read-oriented query lifecycle operations. The only actions that change anything come from a companion policy: support:CreateCase and a service-linked-role creation scoped to the Amazon Resource Name (ARN) of a single role.

Keep that role least-privilege, and your audit surface stays small by construction. If you enable agent actions, a later section covers the write path which uses a separate actions role.

The reference architecture

The agent produces output that arrives two different ways, and this shapes how you capture each:

Agent output Delivery How you capture it Latency
Investigation lifecycle Push: Amazon EventBridge events React to events (rules + targets) Seconds
Recommendations Pull: no event emitted. Generated on goal cadence Poll list-recommendations on a schedule depends on your poll frequency

The journal itself has no dedicated event, but the terminal lifecycle event carries the execution_id you need to fetch it. The journal is push-triggered, pull-retrieved: the event tells you when to look, and the API gives you what to archive.

Five capture layers inside the Agent Space Region: lifecycle events to CloudWatch Logs, terminal events to a Lambda that archives journals to Amazon S3 with a dead-letter queue, a scheduled poll for recommendations, control-plane mutations to an SNS topic through Amazon EventBridge, and a Glue/Athena query layer. Two operator CLIs read the archive: correlate.py joins findings to AWS Config and CloudTrail, and correlate_agent.py joins agent actions to their approvals.

Figure 1: Reference architecture for auditing AWS DevOps Agent across five capture layers

Layer 1: Lifecycle capture. One Amazon EventBridge rule matching {"source":["aws.aidevops"]}, targeting an Amazon CloudWatch Logs (CloudWatch Logs) group directly. This durably records every lifecycle transition. Start here for operational visibility. If your primary goal is behavioral audit rather than operational visibility, deploy layer 2 alongside it.

Layer 2: Behavior capture. A second rule matches only terminal events and invokes a Lambda function. The function reads the execution_id from the event, calls list-journal-records, and writes the journal to Amazon S3. Subscribe to each terminal investigation and mitigation type. This is the layer that captures the agent’s decisions for the long term, including agent-based mitigations.

Layer 3: Recommendations snapshot. Because recommendations are generated on a schedule and have no event, capture them with an Amazon EventBridge Scheduler rule that invokes a Lambda function on a cadence (start daily). The function calls list-recommendations and writes each to Amazon S3, keyed on recommendation ID and version. It also calls list-goals in the same invocation, because a recommendation carries no field saying whether it is still current and the owning goal is the only thing that does. The journal captures what the agent found, and this layer captures what it advised and what you did about it.

Layer 4: Control-plane alerting. On your existing organization trail, alert on mutating aidevops.amazonaws.com events including UpdateApprovalAction, which is produced on elevated actions. This is your tripwire for changes to the agent itself.

Layer 5: Query. AWS Glue Data Catalog tables and an Amazon Athena (Athena) workgroup over the archived journals, recommendations, and goals.

Querying the archive: AWS Glue and Athena

The sample implementation overlays an AWS Glue Data Catalog and an Athena workgroup on the Amazon S3 archive. Three external tables cover the full archive. The journals table uses Athena partition projection, and Hive-partitioned by agent space and date:

s3://<amzn-s3-demo-archive-bucket>/journals/space=<agent-space-id>/dt=2026-07-28/<execution-id>.json

Volume of recommendations is low (tens to hundreds of objects), so a flat external table over the recommendations/ prefix is sufficient. Athena recurses subdirectories by default, picking up every versioned snapshot.

The result bucket has Amazon S3 Object Lock but Object Lock prevents Athena from managing its own query-result objects. The query layer deploys a dedicated results bucket with a seven-day lifecycle rule for ephemeral query outputs.

Access control

Use IAM to control access. Investigation journals contain the agent’s full reasoning about your infrastructure. Scope your IAM permissions on the Athena workgroup, AWS Glue database, and on the archive bucket itself since bucket read access bypasses Athena entirely. Scope all three to your audit and operations teams.

To find all findings from the past 7 days for a specific resource:

SELECT
  execution_id,
  event_time,
  task.title,
  record.content
FROM devops_agent_audit.journals
CROSS JOIN UNNEST(journal_records) AS t(record)
WHERE dt >= date_format(current_date - interval '7' day, '%Y-%m-%d')
  AND record.recordType IN ('finding', 'investigation_result')
  AND record.content LIKE '%sg-0123456789abcdef0%'
ORDER BY event_time DESC;

The Athena workgroup integrates with Amazon Quick or any business intelligence tool that speaks JDBC/ODBC. Additional examples are available in the sample repository.

Closing the loop: Correlating findings to actual changes

The capture layers record what the agent found and what it recommended. But did the recommended fix actually land? This requires connecting the agent’s output to the real infrastructure change that followed.

The sample implementation includes correlate.py, an on-demand operator CLI that takes an archived finding or recommendation, resolves the resource it references, and reports what changed, when, and who did it. The correlation is heuristic by looking at resource identity and a tight time window in minutes to produce reliable attribution. This is why the sample implementation pairs it with a deterministic engine for agent-initiated actions.

It works by pivoting through two services:

  1. AWS Config resolves the resource identity by using select-resource-config, then pulls its configuration timeline from get-resource-config-history. This shows the before/after state of the resource around the time of the agent’s finding.
  2. CloudTrail looks up the write event that caused the change: who called what API, from where, and when. This attributes the change to a principal.

The output is a correlated record: the agent found X, the resource changed from state A to state B, and that change was made by principal Y at time T.

Because CloudTrail indexes resources by different identifiers depending on the service, you require a strategy registry. Examples are in the following table:

Resource type How CloudTrail indexes it Lookup strategy
S3 bucket Bucket name By name
Lambda function Function name By name
Amazon Relational Database Service (Amazon RDS) instance/cluster Full ARN (not the DB ID) Build ARN from template
Amazon Elastic Compute Cloud (Amazon EC2) security group Group ID as ResourceName By name, with a resource-type scan as fallback

A naive “look up by resource name” works for Amazon S3 and Lambda but returns zero results for Amazon RDS (RDS). The strategy registry encodes the right ID per resource type.

Correlating agent actions

When an operator approves an elevated action, the service stamps the approval ID into the credential it mints, so the executed call carries that ID inside its own principal ARN (op.system.apr.<approvalId>). The sample implementation includes correlate_agent.py that uses this. Because the ID is present on both sides, the correlation is a join. The engine checks the executed call against the argumentPins the operator was shown at approval time, so you can prove the agent’s behavior.

. correlate.py correlate_agent.py
Pivots on A resource the agent named Agent’s approval ID
Correlation heuristic deterministic
Answers Who changed? Who approved, and did it match?
Dependency CloudTrail and AWS Config CloudTrail

Production considerations

Understand the data volume. Journal size scales with investigation complexity. As an example:

Scenario Journal size API calls (pagination) Notes
Minimal (single-service, shallow investigation) ~65 KB 2–3 pages Quick symptom to finding arc
Typical (multi-signal, 1–2 findings) 250–340 KB 65–106 calls Typical investigations
Exhaustive (account-wide, high-priority) ~428 KB 150+ calls Full cross-service correlation

At 100 investigations/month at 300 KB average, you are storing roughly 30 MB/month of journal data.

Concurrency per agent space. By default, you can run three concurrent investigations per agent space. Additional requests queue as PENDING_START and start when a slot opens. The archival pipeline is unaffected because each terminal event triggers its own Lambda invocation. Refer to the AWS DevOps Agent Quotas page for future updates.

Paginate the journal. The journal API is server-paginated: pass limit, follow nextToken until it’s empty. A real incident’s journal can span several pages. Always loop.

Design for at-least-once delivery. Amazon EventBridge can deliver an event more than once. Key the Amazon S3 object on execution_id so a redelivery overwrites rather than duplicates, and attach an Amazon Simple Queue Service (Amazon SQS) dead-letter queue (DLQ) so a dropped terminal event is not lost.

Deploy per AWS Region and per account. Events land on the default bus in each Agent Space’s hosting account and Region. If you run agent spaces in multiple accounts, you must aggregate events to a central monitoring account for unified visibility. Refer to Amazon EventBridge cross-account document for further details.

Make the archive immutable. Enable Amazon S3 Object Lock and versioning. The sample implementation defaults to GOVERNANCE mode but for stronger compliance posture, use COMPLIANCE mode.

Warning: COMPLIANCE mode is irreversible. After it’s set, no principal (including the account root user) can delete or modify locked objects before their retention period expires. The only way out is closing the AWS account, and Object Lock itself can’t be disabled once enabled. Choose COMPLIANCE mode deliberately. If you use GOVERNANCE mode, enable CloudTrail data events on the bucket.

Encrypt your data. The sample implementation uses SSE-S3. If your compliance framework requires you to control and audit decryption events, use SSE-KMS with customer managed key.

The full loop

Here’s what a complete audit trail looks like for a single incident through resolution.

Step 1: Investigation. The agent investigates a failing Lambda function, concludes its security group restricts necessary egress, and writes the finding to the journal. Layer 2 archives the journal to Amazon S3.

Step 2: Recommendation. On its goal cadence, the agent generates a recommendation: “Update the security group egress rules to allow…” Layer 3 polls and captures it as recommendations/rec-a1b2c3.../v1.json with status PROPOSED. A later poll captures v2.json as the status changes.

Step 3: Engineer applies the fix. An engineer runs the suggested command. AWS Config records the new configuration item, and CloudTrail records the API call with principal, source IP, and timestamp.

Step 4: Correlation.

$ python correlate.py --archive-bucket $BUCKET \
    --recommendation rec-a1b2c3d4-5e6f-7890-abcd-ef1234567890 \
    --window-hours 24

Recommendation: rec-a1b2c3d4-5e6f-7890-abcd-ef1234567890
Title:          Update the Lambda security group egress rules to allow API access
Status:         PROPOSED → UPDATE_IN_PROGRESS (v2)

AWS Config change detected:

  Resource:     AWS::EC2::SecurityGroup / sg-0123456789abcdef0
  Changed:      2026-07-23 08:45:54.105000-04:00
CloudTrail attribution:
  Event:        AuthorizeSecurityGroupEgress
  Principal:    arn:aws:iam::111122223333:user/jsmith
  Source IP:    203.0.113.10
  Time:         2026-07-23 08:44:33-04:00

Correlation:    OK Recommendation → AWS Config change → CloudTrail event aligned

The agent found the problem, recommended the fix, and you can prove who applied it and when.

Step 5: A new investigation. A later investigation examines the same Lambda function, still erroring. The agent compares new advice against advice it has already given, and that comparison is semantic. But it compares against the recommendations currently attached to the goal, not against everything it has ever advised, and when the comparison is uncertain it keeps the two separate. Older advice drops out of that comparison set over time. Because you archived every recommendation and every finding with their resource identifiers, you can now compare across the full history:

$ python correlate.py --archive-bucket $BUCKET \
    --finding  exe-ops1-0f1e2d3c-4b5a-6978-8796-a5b4c3d2e1f0 \
    --check-prior-recommendations

Resource:       AWS::EC2::SecurityGroup /  sg-0123456789abcdef0

Prior recommendations referencing this resource:
   rec-a1b2c3d4-5e6f-7890-abcd-ef1234567890 (v2, UPDATE_IN_PROGRESS):
    "Update the security group egress rules..."
   rec-b2c3d4e5-6f70-8901-bcde-f01234567890 (v1, PROPOSED):
    "Update the security group egress rules..."

! This finding may be a consequence of recommendation(s): rec-a1b2c3d4-5e6f-7890-abcd-ef1234567890, rec-b2c3d4e5-6f70-8901-bcde-f01234567890
Last change to this resource (CloudTrail): Event: AuthorizeSecurityGroupEgress Principal: arn:aws:iam::111122223333:user/jsmith Source IP: 203.0.113.10 Time: 2026-07-23 08:44:33-04:00

The archive diagnosed the cause of the cause. Two recommendations, raised separately, on one resource, in one view. The agent’s own comparison covers the advice currently attached to the goal. The archive covers all of it. That is the feedback loop the audit trail adds.

Agent Actions changes Step 3’s actor, and the agent applies the fix directly. In the recommendation path, the human runs the command. In the elevated-action path, the human approves a specific call, and the agent executes it under a single-use session. correlate_agent.py uses a single-use session named for the approval (op.system.apr.<approvalId>), with invokedBy: aidevops.amazonaws.com rather than a time-window heuristic.

Operating the pipeline: Common failures

Always design for failure. Here are some common failures and how to detect and recover.

Failure Symptom Detection Recovery
Lambda timeout No archive in Amazon S3. Event in DLQ DLQ ApproximateNumberOfMessagesVisible alarm Increase timeout above the 2-minute default. Replay DLQ message which is idempotent on the execution_id key
Missed recommendation poll Gap in recommendations/ prefix with a version number skipped Periodic reconciliation: compare Amazon S3 keys against list-recommendations response Re-run poll Lambda manually (idempotent)
Amazon EventBridge delivery failure Missing lifecycle event in Layer 1 logs Layer 2 archive exists without matching Layer 1 log entry No data loss since journal already archived. Gap is in lifecycle visibility only
Amazon S3 write failure Lambda errors spike. DLQ grows Lambda error rate metric and DLQ alarm Fix IAM/bucket policy. Replay DLQ (all messages are idempotent)
AWS Config recorder stopped correlate.py returns no configuration history AWS Config recorder status alarm Re-enable recorder. Note: historical gap is permanent for the stopped period
Journal API throttled Partial archive. Lambda retries exhaust timeout Lambda error logs showing throttling exceptions Implement exponential backoff in the pagination loop. Increase timeout
Approval recorded but not executing Approval exists in CloudTrail with no corresponding write Join approvals to execution on the approval ID None needed

The highest value alarm is on the DLQ message count. A non-empty DLQ means a terminal event triggered, but the journal was not archived. Terminal events aren’t re-emitted, and the DLQ retains messages for 14 days. After that, the record is lost. The sample implementation ships this alarm at a threshold of 1, wired to an Amazon Simple Notification Service topic.

Run a reconciliation check weekly or monthly. Compare the execution_id values in the Layer 1 lifecycle log against the set of keys in the Amazon S3 journals/ prefix. Any ID in the logs but not in Amazon S3 represents a missed archive.

Limitations

Automated correlation – The current design requires a human to run correlate.py. Extend to a Lambda function that triggers on each new journal archive, cross-references the finding’s resource identifiers against the recommendations table, and alerts when a new finding touches a resource that was the subject of a prior recommendation.

Schema evolution – The Athena table definitions depend on the journal’s recordType values and content structure. If new record types appear, queries can return incomplete results without raising an error. Monitor for unknown recordType values. A query that returns zero findings for a week of active investigations is a signal that the schema moved.

Conclusion

Adopting an autonomous agent is a trust decision, and trust needs evidence. AWS DevOps Agent gives you the raw material: a journal of its reasoning, a real-time lifecycle event stream, a control-plane audit in CloudTrail, and an action boundary you can read straight from IAM. The pattern in this post assembles those into a durable, low-maintenance audit trail using services you already run.

The archive is more than compliance paperwork. With a persistent record of every finding and every recommendation, you can correlate across investigations and recommendations the agent no longer has in view, and against the present state of your infrastructure. That feedback loop is the difference between trusting the agent and understanding it.

Start with Layer 1. A single Amazon EventBridge rule to a log group gives you visibility into every investigation within minutes. Add the journal-archiving Lambda when you are ready to retain the agent’s decisions for the long term. Add the correlation layer when you want to prove that recommendations were acted on and catch the ones that created new problems.

Clone the sample repo to get started. It covers prerequisites, deploy steps, codebases, and teardown instruction. If you want the agent’s mitigations to become code, Automated incident remediation with AWS DevOps Agent and Kiro CLI builds a pipeline.


About the authors

Ben Peterson

Ben Peterson

Ben is a Senior Solutions Architect at AWS, focused on the developer experience and helping ISV customers modernize on AWS. He provides strategic guidance on using the AWS suite of services to modernize legacy systems, optimize performance, and unlock new capabilities. Connect with Ben on LinkedIn.

Jake Izumi

Jake Izumi

Jake is a Senior Solutions Architect supporting the NAMER ISV customers at AWS. Using his previous experience supporting corporate growth strategies, Jake works with business and technology leaders to innovate and grow on top of AWS. Connect with Jake on LinkedIn.

Sean Falconer

Sean Falconer

Sean is a Senior Solutions Architect at AWS, focused on agentic AI and event-driven architectures for ISV customers. His current work centers on the trust and governance patterns that let teams adopt autonomous agents in production. Connect with Sean on LinkedIn.

The collective thoughts of the interwebz