Frozen package management for air-gapped RHEL-family AMIs

Post Syndicated from Anand Krishna Varanasi original https://aws.amazon.com/blogs/compute/frozen-package-management-for-air-gapped-rhel-family-amis/

If you run a regulated, air-gapped compute fleet on RHEL-family instances, you have probably felt three requirements pulling against each other. Your organization must configure the network to remove internet access from the instances. Your team must review and approve new packages or version upgrades before you adopt them. Your team removes public repository definitions, restricts network paths, and configures instances to use only the internal repository your team has approved. Teams in chip design, finance, healthcare, defense, and the public sector often face this combination while still needing operating system updates.

This post describes a two-account pattern that separates the connected package-ingestion path from the air-gapped fleet. You create an authorized initial baseline and approve later changes to form a versioned package snapshot in Amazon Simple Storage Service (Amazon S3). EC2 Image Builder uses the frozen snapshot to build Amazon Machine Images (AMIs). Your team configures AWS Systems Manager Patch Manager to patch the instances your organization runs from the same internal package source.

The accompanying reference implementation demonstrates the pattern for RPM-based RHEL-family systems (AlmaLinux for example). It is a reference, not a substitute for distribution of licensing, vulnerability analysis, testing, or an organization’s change-management process.

The challenge: Getting packages into an air-gapped approval-gated fleet

Common delivery models each assume something an air-gapped fleet might not provide:

  • Red Hat Update Infrastructure (RHUI) expects each instance to reach the service. A fleet with no internet egress needs a different content path.
  • Red Hat Satellite supports disconnected content management, but it is a separate product and operational footprint. Teams that need a custom package-level approval workflow must integrate that workflow with their content-management process.
  • The Red Hat CDN requires a connected, entitled content-management path. Centralizing that path changes the network architecture, not the customer’s Red Hat subscription obligations.

The objective is not to replace these products universally. It is to show a serverless AWS pattern for teams that need an authorized repository baseline, explicit approval for later package changes, and a fleet with no public package source.

How the pattern works

The pattern combines three controls:

  1. A frozen package repository on Amazon S3: The pattern stores a deployment-authorized baseline and subsequent approved package changes in versioned, per-OS repository prefixes. The repository manifest records the package inventory for each state.
  2. EC2 Image Builder Orchestration builds AMIs from that repository: The build helps remove upstream repository definitions and configures the internal frozen mirror to be used for all package operations.
  3. The launched fleet has no internet egress: The dnf operations are configured to resolve the internal mirror. Patch Manager uses the same repository source, so image builds and in-place patching draw from one frozen snapshot.

Choosing the upstream source

Choose one package lineage end to end. The parent AMI, repository content, and trusted signing keys must belong to that same lineage.

The reference implementation defaults to AlmaLinux vault content plus EPEL and an AlmaLinux parent AMI. The AlmaLinux OS Foundation states that AlmaLinux aims for binary and application binary interface (ABI) compatibility with RHEL. This is an AlmaLinux compatibility goal, not a Red Hat certification, and it does not make repository mixing a supported practice.

For genuine RHEL systems, use a Red Hat parent AMI, entitled Red Hat repositories, and Red Hat signing keys. A connected content-management host can retrieve content for the isolated environment. This centralizes the network path but does not reduce or change the customer’s Red Hat subscription obligations. Confirm those obligations against the applicable Red Hat agreement.

Do not pair AlmaLinux repositories with genuine RHEL hosts, or Red Hat repositories with AlmaLinux hosts. Mixed-vendor package lineages can create support, stability, and maintainability problems even when the RPMs appear mechanically compatible.

Architecture and workflow

The account boundary provides a primary security boundary for this architecture. The following diagram shows the connected Distribution account, the read-only Workload account, and an example cross-Region layout.

Two-account architecture showing the connected Distribution account with the control plane and internet path, and the air-gapped read-only Workload account, spanning two Regions

Figure 1: Two-account, cross-Region architecture separating the connected Distribution account from the air-gapped Workload account

The Distribution account owns the writable control plane and the only internet path. It runs Amazon EventBridge, three AWS Lambda functions, Amazon DynamoDB, Amazon Simple Notification Service (Amazon SNS), and the AWS Fargate sync task. It also owns the frozen S3 repository and its AWS Key Management Service (AWS KMS) key.

The Workload account is air-gapped and read-only with respect to the repository. It runs the internal HTTPS mirror, EC2 Image Builder, Patch Manager, and the compute fleet. Its mirror task role can read and decrypt frozen content but cannot write it.

The sample repository places the Distribution control plane in US East (N. Virginia), the frozen store in US West (Oregon), and the Workload resources in US West (Oregon) to demonstrate API-only cross-account and cross-Region operation. This Region split is not required. In most deployments, place the Distribution control plane and frozen store in the same Region unless data residency, disaster recovery, or an existing regional footprint justifies the additional latency, transfer cost, and KMS policy complexity.

The two accounts do not need Amazon Virtual Private Cloud (VPC) peering or a transit gateway. Cross-account access uses S3, KMS, and IAM policies. The VPC address ranges can overlap because no VPC-to-VPC route is required.

Package baseline and scheduled upgrade workflow

Before the scheduled workflow begins, your organization must authorize and run a full sync to establish the initial repository baseline. This bootstrap does not provide package-by-package approval. If your organization requires individual approval for every initial RPM, your team should generate and review the baseline manifest before promotion instead of relying solely on deployment authorization.

After the baseline, the detector runs on a customer-defined schedule. The reference implementation defaults to monthly. The following diagram shows the bootstrap distinction and the selective approval flow.

Workflow diagram distinguishing the initial baseline bootstrap sync from the recurring detect, request approval, review, record, selective sync, and manifest update steps

Figure 2: Package baseline bootstrap and the scheduled selective approval workflow

  1. Detect. Amazon EventBridge invokes the detector Lambda function on the configured schedule. The detector compares upstream repository metadata with manifest.json, which records the current frozen inventory. It classifies a newer version as an upgrade and an absent package as new.
  2. Request approval. The detector writes candidates to S3, creates a KMS-protected review token carrying the request ID and expiry, and is designed to send a review link through SNS. The detector can use kms:Encrypt but not kms:Decrypt.
  3. Review. A human opens the review page through an Amazon API Gateway HTTP API, reviews the proposed package versions, and chooses which changes to approve. The approver can use kms:Decrypt but not kms:Encrypt.
  4. Record and start. A conditional DynamoDB update changes a request from pending to approved only once. The approver then starts the Fargate sync task and passes the request ID.
  5. Selective sync. The task reads the approved package list, downloads those package versions, is designed to perform verification checks, and regenerates repository metadata.
  6. Update the manifest. When the task stops, Amazon EventBridge invokes the manifest-updater Lambda function. It archives the outgoing manifest and records the resulting repository inventory.

The approval decision controls adoption. It does not prove that package code is safe. Advisory review, vulnerability scanning, testing, and staged rollout remain in separate controls.

Evidence from the approval workflow

The token ties a review action to a specific request and expiry. Separating kms:Encrypt from kms:Decrypt prevents either Lambda function from performing both token roles. The conditional DynamoDB write makes the approval transition single-use.

DynamoDB records request state, AWS CloudTrail records control-plane API activity, and manifest history records repository inventory changes. These service records can feed the organization’s existing audit and evidence-management workflow. Object-level S3 access auditing requires CloudTrail S3 data events. KMS activity alone is not a substitute for those events.

The frozen package repository on Amazon S3

The following diagram shows the per-OS, per-component prefix layout, and manifest objects.

Amazon S3 prefix layout with one prefix per operating system version, each holding BaseOS, AppStream, and EPEL components with Packages and repodata trees plus manifest objects

Figure 3: Per-OS, per-component prefix layout of the frozen repository on Amazon S3

Each pinned operating system version receives its own prefix. Repository components such as BaseOS, AppStream, and EPEL contain Packages/ and repodata/ trees. manifest.json records the active inventory, and archived manifests preserve historical evidence and comparison points.

Your organization configures the bucket with versioning and SSE-KMS. Public RPM content does not require a customer-managed KMS key for confidentiality, so your organization could instead configure SSE-S3 for encryption at rest. However, SSE-S3 would remove the separate cross-account authorization control provided by the customer-managed KMS key policy. The customer managed key is used here for explicit cross-account key-policy control and revocation, and the manifests reveal the fleet’s exact software inventory. S3 Bucket Keys reduce KMS request volume. If object-level access evidence is required, enable CloudTrail S3 data events.

A rollback must restore a coherent repository state, including metadata and any required object versions. Restoring only manifest.json does not roll back repository contents.

Building, patching, and running the fleet

At AMI build time, an Image Builder component installs the configured repository keys, moves existing repository definitions aside, and writes one frozen repository definition per component. It locks the package manager to the frozen repository directory, fetches metadata through the internal mirror, and fails the build if the mirror validation step fails. An optional curated package list demonstrates that the AMI can install real packages through the frozen path.

Patch Manager uses the same mirror for the running fleet. A host created from an older AMI and a newly built host are therefore patched toward the same frozen snapshot. Instances run without an Amazon VPC NAT gateway, public IP, or an Amazon VPC internet gateway route in the Workload VPC, and their repository configuration contains no public fallback.

The intended verification model is defense in depth: the sync task helps verify a vendor’s signature before content enters the trusted repository, and the system verifies it again at installation through dnf. The ingestion gate helps reject digest-only results and can be configured to help confirm that only valid package signatures from a trusted lineage key are accepted.

The package mirror

Nginx fronts aws-sigv4-proxy, which signs cross-account S3 GET requests using the mirror task role. To a client, the service appears as a standard HTTPS package repository behind an internal Application Load Balancer and private DNS name.

Use the latest version of aws-sigv4-proxy (current latest is v1.12). This version 1.12 contains the fix for signing S3 paths (or the OS package names) with special characters such as +. Earlier versions can return SignatureDoesNotMatch. The reference implementation pins the reviewed v1.12 release commit immutably. Keep it current through dependency-update reviews.

Security boundaries and limits

The design provides the following controls:

  • No automatic public-repository adoption: A new upstream version enters the selective path only after an explicit, recorded decision by the user.
  • Repository ingestion and installation checks: You configure strict sync-time signature validation to help validate content before it enters the trusted store. dnf verifies again during installation.
  • No package-channel egress: Workload instances are configured to prevent access to public package sources.
  • A read-only workload boundary: A Workload-account principal cannot modify the frozen repository.

Human approval is not a malware detection. A reviewer cannot reliably identify a backdoor in a legitimately signed package merely by seeing its name, version, or changelog. Use vulnerability intelligence, scanning, pre-production tests, and staged deployment as additional controls.

The approval state, manifests, and CloudTrail records can help support evidence for control frameworks such as SOC 2 change management, ISO 27001 patch-management controls, and FDA 21 CFR Part 11 electronic records. Applicability depends on the organization’s environment, audit scope, and assessor. Confirm it with the compliance team under the AWS shared responsibility model.

Cost and operations

Cost depends on the amount of repository content and the chosen networking and availability design. Components can include S3 storage and requests, KMS requests, Lambda invocations, DynamoDB, SNS, Fargate tasks, the internal load balancer, Distribution-account internet egress, Amazon VPC endpoints, and AMI snapshots. A three-task always-on mirror costs more than an S3 bucket alone. Estimate the target topology with current AWS pricing rather than applying a fixed monthly figure from the sample.

Run detection and review at an interval defined by patch policy and risk tolerance. The supplied default is monthly, but the Terraform input is configurable. If a package change must be reversed, restore a tested, coherent repository version and rebuild or patch affected hosts as appropriate.

Prerequisites

To set up the reference implementation, work through these in order:

  1. Two AWS accounts: a connected Distribution account and an air-gapped Workload account.
  2. Deployment tools: Terraform 1.5 or later, Terragrunt, Finch or Docker, and Python with pip.
  3. AWS Command Line Interface (AWS CLI): one named profile per account.
  4. Distribution networking: private subnets with internet egress that works without public IPs, security-group egress on port 443, and DNS resolution for public names.
  5. Workload networking: VPC interface endpoints for ssm, ssmmessages, ec2messages, logs, kms, and imagebuilder, plus an S3 gateway endpoint.
  6. Internal mirror identity: an AWS Certificate Manager (ACM) certificate and a private hosted zone. If the parent AMI does not trust the issuing CA, configure the CA file so the build installs the trust anchor.
  7. Package lineage: a parent AMI, repositories, and signing keys from the same distribution lineage. A RHEL subscription is required when retrieving genuine entitled Red Hat content.
  8. Optional deployment roles: otherwise, the stack uses each profile’s credentials.

The deployment creates state backend and Amazon Elastic Container Registry (ECR) repositories. Do not create those ECR repositories separately before applying their own Terraform units.

Reference implementation

The companion repository provides Terraform modules, Lambda handlers, two container images, a Terragrunt two-account layout, and Makefile targets for deployment and verification. The shipped alma810 example defaults to the AlmaLinux lineage (RHEL family).

Choose one distribution lineage before deployment:

  • AlmaLinux Parent Image default: use an AlmaLinux parent AMI, AlmaLinux vault repositories, EPEL, and the included AlmaLinux and EPEL signing keys. No Red Hat subscription is required.
  • Genuine RHEL Parent Image: use a Red Hat parent AMI, an entitled Red Hat content source, and Red Hat signing keys. The customer supplies the Red Hat subscription and content-access integration.

Note: The reference implementation only provides AlmaLinux lineage setup, not genuine RHEL. If you choose to use a genuine RHEL parent image lineage, only the reference implementation code needs to be updated to fetch the Red Hat credentials or subscription access, and the rest of the workflow remains the same.

Please follow the repository README for detailed setup instructions:

  1. Fill in environments/config.hcl and both account files with the two accounts, networking, mirror certificate, package lineage, and parent AMI.
  2. Create the Terraform state backend with make bootstrap DIST_PROFILE=<dist> WORK_PROFILE=<work>.
  3. Review both account plans with make plan DIST_PROFILE=<dist> WORK_PROFILE=<work>.
  4. Deploy in dependency order with make all BASELINE_APPROVED=true DIST_PROFILE=<dist> WORK_PROFILE=<work>. The baseline sync can take several hours.

Validate the internal package management workflow

Validation begins by running the AlmaLinux EC2 Image Builder pipeline. A successful build produces private, encrypted AMIs in the configured AWS Regions. To validate the configuration, launch a test instance from the generated AMI in a no-egress Workload subnet and verify that the instance is connected to the internal frozen repository for all package management workflow.

Terminal output listing only the internal frozen-baseos, frozen-appstream, and frozen-epel repositories enabled, with dnf makecache downloading metadata through the private mirror

Figure 4: Instance showing only the internal frozen repositories enabled, with no public repository configured

The output confirms that only the internal frozen AlmaLinux repositories are enabled: frozen-baseos, frozen-appstream, and frozen-epel. The dnf makecache command successfully downloads metadata for all three repositories through the private mirror. No public repository is configured or used.

Now try installing, upgrading, or installing a new package to test that package operations are served by the internal frozen repository.

Terminal output showing a package install and upgrade completing successfully through the internal frozen repository mirror

Figure 5: Package install and upgrade served by the internal frozen repository

  • Fail closed on approval data. If the approved package list cannot be retrieved, stop the sync task and avoid substituting a full sync.
  • Govern the initial baseline. Record who authorized the bootstrap full sync, or require explicit review of its manifest before promotion.
  • Require vendor signatures at ingestion. Do not treat a valid package digest as equivalent to a trusted signature.
  • Keep package lineages consistent. Parent AMI, repository content, and signing keys must come from the same distribution lineage.
  • Use a customer-defined review cadence. Monthly is only the sample default.
  • Use immutable dependency references with active updates. Require aws-sigv4-proxy v1.12 or later, pin the reviewed artifact by digest or full SHA, and automate update proposals.
  • Test rollback as a repository operation. Restore metadata and objects together, then validate the mirror before using the restored state.

Clean up

The walkthrough deploys billable resources in both accounts. Please follow the repository README for detailed setup cleanup instructions.

Conclusion

This pattern separates connected package ingestion from an air-gapped fleet, provides an authorized package baseline and an explicit decision point for later package changes or upgrades, and keeps image builds and running hosts on one frozen repository snapshot. To learn more, visit the EC2 Image Builder service page, the EC2 Image Builder documentation, the Patch Manager documentation, and the Amazon S3 user guide. The reference implementation is available in aws-samples.

Build adaptive AI interfaces with the AG-UI protocol, agent swarms, and Nova Act on AWS

Post Syndicated from Anand Bilgaiyan original https://aws.amazon.com/blogs/architecture/build-adaptive-ai-interfaces-with-the-ag-ui-protocol-agent-swarms-and-nova-act-on-aws/

Your generative AI applications produce different results each time they run. One medical scan shows a single fracture, and another reveals twenty ambiguous regions that require expert review. Static interfaces can’t adapt to this variability.

In this post, we show you how to build interfaces that automatically adapt to your AI’s variable outputs. This approach reduces interface development time and removes the need for custom integration code by using the AG-UI protocol, the Strands Agents Software Development Kit (SDK), and Amazon Nova Act. You learn to build adaptive interfaces using the agent-to-UI (AG-UI) protocol for dynamic UI generation, the Strands Agents SDK Swarm pattern for multi-agent collaboration, and Amazon Nova Act for legacy system integration. After reading this post, you understand when to use adaptive interfaces and how to deploy them in your applications.

The problem with static interfaces

You design screens with fixed layouts, predetermined controls, and static data binding. This works well when your application’s output space is known and consistent. An ecommerce checkout page needs the same fields for every transaction. A dashboard displays the same metrics regardless of the data. Static interfaces excel at these predictable scenarios.

AI-driven applications introduce new requirements. Consider two scenarios when you review bone X-rays. In the first scenario, one image shows a single obvious fracture requiring minimal interface controls. In the second scenario, another image reveals three subtle regions where multiple AI agents must debate findings, track confidence progression, and reach consensus before presenting results. A static interface optimized for the first scenario lacks the controls needed for the second. An interface built for the second scenario presents unnecessary complexity in the first scenario with empty panels and unused controls.

This problem appears in multiple domains. Fraud detection systems encounter variable evidence chains. Legal document review surfaces unpredictable numbers of relevant clauses. Security event response reveals different threat patterns requiring different analysis tools. Domains where AI discovers things dynamically rather than classifying into predetermined categories face this architectural challenge.

The core challenge is that AI agents discover and reason about the world dynamically, while traditional interface design assumes static, predetermined outputs.

Business impact

This mismatch costs development teams significant time and creates poor user experiences. You spend weeks building interface variations to handle different scenarios, then maintain multiple code paths as your AI models evolve. Your users face either overwhelming complexity when AI finds simple results, or insufficient controls when AI discovers complex patterns requiring deeper analysis. The development cost compounds as you add new AI capabilities. Each new agent or model requires rethinking your entire interface architecture.

Solution overview

Three AWS technologies address this challenge, so you can build interfaces that adapt to what AI agents discover.

The AG-UI protocol offers standardized streaming for agent-to-user interface (UI) communication. Before AG-UI, connecting agents to interfaces required custom code for each framework. You built custom WebSocket formats, polling mechanisms, and bespoke integration code every time you switched agent frameworks or added new capabilities. AG-UI removes this work by providing a standard format of typed events that stream over Server-Sent Events (SSE). Agent frameworks emit AG-UI events, and frontends consume them, creating a universal contract so you can swap frameworks without rewriting integration code.

The Strands Agents SDK offers the Swarm pattern for peer-to-peer multi-agent collaboration. In swarm patterns, your agents operate as peers that share hypotheses and iteratively refine findings until reaching consensus. This debate process, visible to you in real time, builds trust and catches errors that single-agent systems miss. For medical imaging, the swarm pattern mirrors how radiologists consult specialists, with multiple expert perspectives converging on accurate diagnoses.

Amazon Nova Act offers browser-based automation using natural language commands, so your AI agents can interact with legacy systems through their web interfaces. Your agent navigates login screens, searches for related records, fills form fields, and captures confirmation numbers, while streaming actions back to the primary interface so you can observe the process.

These three technologies work together naturally: AG-UI adapts your interface to swarm findings, the swarm produces explainable multi-agent analysis, and Nova Act bridges the gap with legacy systems that lack API access.

Prerequisites

Before starting, verify you have:

Required AWS Resources:

  • AWS account with Amazon Bedrock access in a supported region (us-east-1, us-west-2, or eu-west-1)
  • Your Identity and Access Management (IAM) user or role needs these specific permissions:
    • bedrock:InvokeModel – For calling foundation models.
    • bedrock:CreateAgent and bedrock:CreateAgentActionGroup – For agent deployment.
    • lambda:CreateFunction and lambda:InvokeFunction – For serverless compute.
    • s3:PutObject and s3:GetObject – For file storage.
    • dynamodb:PutItem and dynamodb:GetItem – For state management.
    • secretsmanager:GetSecretValue – For credential retrieval.
    • logs:CreateLogGroup and logs:PutLogEvents – For Amazon CloudWatch logging.

Development Environment:

  • Python 3.9+ with pip installed.
  • Node.js 16+ and React 18+ installed.
  • AWS Command Line Interface (CLI) configured with your credentials.

Technical Skills:

  • Intermediate Python programming experience.
  • Familiarity with event-driven architectures (your interface listens for messages from agents and updates in real time, similar to how chat applications work).
  • Basic understanding of Representational State Transfer (REST) APIs and Server-Sent Events (SSE).

Data protection and HIPAA compliance

This solution processes Protected Health Information (PHI) including medical images, patient MRN, and clinical findings. Apply the following safeguards before deploying to any environment handling real patient data.

Encryption at rest — Configure SSE-KMS with a customer managed key on all Amazon S3 buckets storing medical images. Enable encryption with a customer managed AWS KMS key on all DynamoDB tables storing session state, conversation history, and analysis findings.

Encryption in transit — Enforce TLS 1.2+ on all connections. Attach a bucket policy denying all S3 actions when aws:SecureTransport is false. Do not override DynamoDB SDK endpoints to HTTP. Do not set ignore_https_errors=True on Nova Act workflows. If the legacy system uses self-signed certificates, add its CA to your runtime trust store.

S3 Block Public Access — Enable Block Public Access at the account level and on every bucket in this solution. Medical images must never be exposed through public bucket policies or ACLs.

HIPAA-eligible services and BAA — All AWS services in this architecture (Amazon Bedrock, Amazon S3, DynamoDB, Lambda, API Gateway, CloudFront, Cognito, Secrets Manager, CloudWatch) are HIPAA-eligible. Before processing PHI, execute a Business Associate Agreement (BAA) with AWS covering these services.

PHI minimization — Never write MRN, patient name, or clinical findings to plaintext logs. Enable CloudWatch Logs data protection policies to detect and mask PHI patterns automatically. Suppress Nova Act trajectory logging during steps that display patient data. Verify Cognito JWT on the SSE endpoint before emitting any PHI-bearing event.

Important: Code samples in this post are for educational purposes. Review all configurations against your organization’s HIPAA Security Rule implementation before production deployment.

Architecture

You implement an orchestrator pattern where your frontend interacts with a single entry point that internally coordinates specialized sub-agents. The solution follows this pattern: your React app sends requests to Amazon API Gateway, which triggers AWS Lambda functions that coordinate AI agents through Amazon Bedrock, then streams results back over Server-Sent Events. We use Amazon Bedrock AgentCore, a platform to build, connect, and optimize agents at scale with any framework or model.

Logical view of the radiology portal: the React AG-UI client streams over SSE to a backend orchestrator that coordinates a Strands agent swarm, Nova Act browser automation, a data store, the legacy RIS/EMR portal, and Amazon Bedrock

Figure 1: Logical view of the adaptive interface, showing the AG-UI client, the backend orchestrator, the Strands agent swarm, and integrations with legacy systems and Amazon Bedrock

Detailed AWS architecture with 16 numbered components spanning Amazon Cognito and AWS STS, Amazon CloudFront and Amazon S3, Amazon API Gateway, AWS Lambda, Amazon DynamoDB, Amazon Bedrock AgentCore, the Strands agent swarm, Amazon Bedrock, Amazon OpenSearch Serverless, Amazon Nova Act, and AWS Secrets Manager

Figure 2: End-to-end AWS architecture showing the 16 components that deliver authentication, content delivery, agent orchestration, foundation model inference, and legacy system automation

Architecture components

The architecture consists of 16 integrated components working together:

① User authentication You authenticate through Amazon Cognito, which provides secure identity management and generates JWT tokens for accessing the radiology portal application.

② Content delivery Amazon CloudFront serves the React application and static assets from Amazon S3, providing global low-latency access and caching for optimal performance.

③ API Gateway Amazon API Gateway handles both REST API requests and Server-Sent Events (SSE) connections, providing the entry point for client-server communication.

④ Authorization of user API request with token validation

⑤ Backend processing AWS Lambda functions act as the AG-UI handler, managing authentication, invoking agents through Amazon Bedrock AgentCore Gateway, and formatting responses as SSE streams.

⑥ Image storage Amazon S3 stores medical images with SSE-KMS encryption using a customer managed key. S3 Block Public Access is enabled at the bucket level. Lambda generates short-lived, tightly scoped pre-signed URLs for secure direct uploads, pinned to the PUT method, scoped to a per-user key prefix, restricted to the application/dicom content type, and set to expire in 300 seconds. A bucket policy enforces maximum upload size through the s3:content-length-range condition and denies requests where aws:SecureTransport is false.

⑦ Session management Amazon DynamoDB maintains session state, conversation history, analysis findings, and agent registry data for stateful multi-turn interactions.

⑧ AgentCore Gateway Amazon Bedrock AgentCore Gateway serves as the orchestration layer, routing agent requests, managing sessions, coordinating multi-agent workflows, and load balancing across runtime instances.

⑨ AG-UI handler A specialized component within AgentCore that manages AG-UI protocol events, formatting agent responses as standardized events for dynamic UI rendering.

⑩ AgentCore runtime AgentCore Runtime provides the execution environment for agent instances, running the Strands Agent Swarm with isolated runtime instances for each agent type.

⑪ Agent swarm execution The Strands Agent Swarm consists of three specialized agents (Image Analysis, Clinical Reasoning, Reporting) that collaborate through iterative debate rounds to reach consensus.

⑫ Foundation model inference Amazon Bedrock provides access to the Claude Sonnet foundation model for reasoning, interpretation, and natural language generation across agents using a large language model (LLM).

⑬ Knowledge base retrieval Amazon OpenSearch Serverless stores medical knowledge base with vector embeddings, providing semantic search for relevant medical literature and clinical guidelines.

⑭ Legacy system automation Amazon Nova Act performs browser-based automation to submit validated findings to a legacy hospital’s Radiology Information System (RIS) and Electronic Medical Record (EMR) systems, with each action streamed back to your UI.

⑮ Credential management AWS Secrets Manager securely stores and rotates credentials for legacy system access, providing Nova Act with authentication details at runtime.

⑯ Observability and monitoring Amazon CloudWatch captures logs and metrics, with data protection policies enabled to detect and mask PHI. AWS Distro for OpenTelemetry (ADOT) provides distributed tracing with AWS X-Ray-compatible trace export. AWS Security Token Service (AWS STS) manages temporary security credentials.

Key architectural decisions

Your frontend sees one agent (radiology-assistant), not three, through the orchestrator pattern. This simplifies integration and encapsulates workflow complexity. The orchestrator internally coordinates the swarm based on analysis stage.

Rather than generating arbitrary HTML, agents select from themed, accessible components (ROICard, DebatePanel, ConfidenceMeter). This balances flexibility with design consistency and security through the predefined component library approach.

SSE provides real-time updates as agents work through the streaming protocol. You see agent contributions character by character, creating transparency into the reasoning process.

Your frontend exposes state (current findings, validation decisions) to agents through bidirectional state synchronization. Agents update state through actions. This synchronization supports human-in-the-loop workflows where agents pause for validation before proceeding.

Solution walkthrough

The following steps walk through the solution, from authentication and swarm configuration to the AG-UI endpoint, legacy system integration, and deployment.

Step 0: Configure authentication

Add JWT validation to your API endpoint to enforce authentication before processing requests containing PHI.

from fastapi import Depends, HTTPException, Request
import os

COGNITO_USER_POOL_ID = os.environ["COGNITO_USER_POOL_ID"]
COGNITO_APP_CLIENT_ID = os.environ["COGNITO_APP_CLIENT_ID"]

async def verify_token(request: Request):
    """Validate Cognito JWT from Authorization header."""
    token = request.headers.get("Authorization", "").replace("Bearer ", "")
    if not token:
        raise HTTPException(status_code=401, detail="Missing authorization")
    # Validate token against Cognito JWKS endpoint
    # See: https://docs.aws.amazon.com/cognito/latest/developerguide/amazon-cognito-user-pools-using-tokens-verifying-a-jwt.html
    return validate_jwt(token, COGNITO_USER_POOL_ID, COGNITO_APP_CLIENT_ID)

@app.post("/api/agui")
async def agui_endpoint(request: Request, user=Depends(verify_token)):
    ...

Full Amazon Cognito user pool setup (pool creation, hosted UI, and token exchange) is covered in the Amazon Cognito Developer Guide. This post focuses on the agent architecture layer.

Step 1: Install the Strands Agents SDK

Install required packages by running the following command:

pip install strands-agents ag-ui-strands fastapi uvicorn

This command installs the required packages for building your multi-agent swarm, including the Strands framework, AG-UI integration, and web server components.

Step 2: Configure your swarm agents

Create specialized agents for your swarm by adding the following code to your project:

from strands import Agent
from strands.models import BedrockModel
from typing import Dict, List, AsyncGenerator
import os

region = os.environ.get("AWS_REGION", "us-east-1")

class RadiologySwarm:
    """Coordinates multi-agent analysis with visible debate."""

    def __init__(
        self,
        consensus_threshold: float = 0.80,
        max_rounds: int = 5,
    ):
        self.consensus_threshold = consensus_threshold
        self.max_rounds = max_rounds
        self.agents = self._initialize_agents()
        self.hypotheses = []

    def _initialize_agents(self) -> List[Agent]:
        """Create specialized agents for swarm."""
        model = BedrockModel(
            model_id=os.environ.get("BEDROCK_MODEL_ID", "us.anthropic.claude-sonnet-4-20250514-v1:0")
        )
        return [
            Agent(
                name="image_analysis",
                system_prompt="Detect abnormal patterns in medical images...",
                model=model,
            ),
            Agent(
                name="clinical_reasoning",
                system_prompt="Validate findings with clinical context...",
                model=model,
            ),
            Agent(
                name="reporting",
                system_prompt="Structure findings in clinical format...",
                model=model,
            ),
        ]

    async def analyze_with_debate_stream(
        self,
        image_data: bytes,
        patient_context: Dict,
    ) -> AsyncGenerator[Dict, None]:
        """Stream swarm analysis events for real-time UI updates."""
        # Implementation continues in next steps...
        pass

Note: The agents use a foundation model accessed through Amazon Bedrock. The example defaults to Claude Sonnet 4, but you should select a currently active model from the Amazon Bedrock model lifecycle page for your AWS Region. Read the model ID from an environment variable so you can update it without code changes as newer models become available. Amazon Bedrock retires models on a published lifecycle schedule, so hardcoding a model ID risks failure when that model reaches end of life. Check the Amazon Bedrock model lifecycle page and supported Regions, and use an environment variable or AWS Systems Manager Parameter Store to manage the model ID externally.

This code creates three specialized agents (Image Analysis, Clinical Reasoning, Reporting) that work together as peers in a swarm pattern, sharing hypotheses and refining findings through iterative debate until reaching consensus.

Step 3: Implement consensus mechanism

Configure your swarm to continue debate rounds until agents reach consensus or exhaust maximum iterations:

from typing import Dict, List

class RadiologySwarm:
    """Consensus mechanism implementation."""

    def __init__(
        self,
        consensus_threshold: float = 0.80,
        max_rounds: int = 5,
    ):
        self.consensus_threshold = consensus_threshold
        self.max_rounds = max_rounds
        self.hypotheses = []

    async def analyze_with_debate_stream(
        self,
        image_data: bytes,
        patient_context: Dict,
    ):
        """Stream swarm analysis with consensus checking."""
        yield {"type": "swarm_start", "max_rounds": self.max_rounds}
        for round_num in range(1, self.max_rounds + 1):
            # Each agent contributes based on current hypotheses
            for agent in self.agents:
                try:
                    context = {
                        "image_data": image_data,
                        "patient_context": patient_context,
                        "round": round_num,
                    }
                    async for chunk in agent.stream_response(context):
                        yield {
                            "type": "agent_contribution_chunk",
                            "agent": agent.name,
                            "round": round_num,
                            "text": chunk,
                        }
                except Exception as e:
                    yield {
                        "type": "agent_error",
                        "agent": agent.name,
                        "round": round_num,
                        "error": str(e),
                    }
            # Check if consensus reached after all agents contribute
            if self._check_consensus():
                yield {"type": "consensus_reached", "round": round_num}
                break
        yield {"type": "swarm_complete", "findings": self.hypotheses}

    def _check_consensus(self) -> bool:
        """Verify hypotheses meet confidence threshold."""
        if not self.hypotheses:
            return False
        # Check if all hypotheses have confidence above threshold
        for hypothesis in self.hypotheses:
            if hypothesis.get("confidence", 0) < self.consensus_threshold:
                return False
        return True

Your swarm reaches consensus when proposed findings achieve confidence scores above your configured threshold (80% for this example). Confidence progresses as agents validate or debate each other.

In the first example of subtle fracture detection, Round 1 shows the Image Agent proposing a Region of Interest (ROI) with confidence above 80%. Round 2 shows the Clinical Agent validating with anatomical context, maintaining confidence above the threshold. Round 3 shows agents agreeing on clinical significance, reaching consensus when the three agents reach confidence scores above the configured threshold.

In the second example of false positive challenge, Round 1 shows the Image Agent proposing an ROI with confidence above the threshold. Round 2 shows the Clinical Agent challenging it as an artifact, with confidence dropping below 80%. Round 3 shows the Image Agent acknowledging the challenge, with confidence dropping further. Round 4 shows agents continuing debate without consensus. Round 5 shows maximum rounds reached, with the finding marked as disputed.

This iterative refinement, visible to you in real time, provides explainability that black-box AI systems can’t match.

Step 4: Set up the AG-UI protocol endpoint

Create a streaming endpoint by adding the following code:

from fastapi import FastAPI, Request
from fastapi.responses import StreamingResponse
import json

app = FastAPI()

@app.post("/api/agui")
async def agui_endpoint(request: Request):
    """AG-UI streaming endpoint emitting typed events."""
    body = await request.json()

    async def event_stream():
        """Stream standardized AG-UI events."""
        swarm = RadiologySwarm()
        async for event in swarm.analyze_with_debate_stream(
            image_data=body["image_data"],
            patient_context=body["patient_context"],
        ):
            yield format_sse_event(event)

    return StreamingResponse(
        event_stream(),
        media_type="text/event-stream",
    )

def format_sse_event(event: Dict) -> str:
    """Format as Server-Sent Event."""
    return f"data: {json.dumps(event)}\n\n"

This code creates a streaming endpoint that emits standardized AG-UI events, so your frontend receives real-time updates as agents work without requiring custom integration code for each agent framework.

The AG-UI protocol defines event types that cover agent-to-UI communication needs. Instead of custom formats, agents emit structured events over SSE, and frontends subscribe to this stream and react to each event type. TEXT_MESSAGE_CONTENT streams agent reasoning or responses token by token for real-time visibility into agent thinking. STATE_DELTA provides incremental state updates for bidirectional synchronization between agent and interface. TOOL_CALL_START and TOOL_CALL_END show tool execution when agents invoke external functions or APIs. With UI_COMPONENT_SPEC, agents control which UI components appear and how they’re configured through interface element specifications.

Configure your frontend to subscribe to this stream and handle each event type:

// Frontend event handler
eventSource.onmessage = (event) => {
  const data = JSON.parse(event.data);
  switch (data.type) {
    case "text_message_content":
      appendAgentText(data.agent, data.text);
      break;
    case "state_delta":
      updateApplicationState(data.path, data.value);
      break;
    case "ui_component_spec":
      renderComponent(data.component, data.props);
      break;
  }
};

Step 5: Implement real-time agent debate visualization

Create interface components that make multi-agent collaboration visible to you by implementing the following visualization features:

A key benefit of the swarm pattern combined with AG-UI is making multi-agent collaboration visible. Rather than presenting you with final results from a black-box system, the interface shows agents debating findings in real time.

Your interface includes specialized components for visualizing agent collaboration. The agent contribution display shows each agent’s reasoning streaming character by character with a typing effect, indicating which agent is currently “thinking.” Color-coding distinguishes agents (blue for Image Analysis, purple for Clinical Reasoning, cyan for Reporting). The confidence timeline uses a line graph to show how confidence evolves across rounds. Increasing confidence (72% → 78% → 85%) indicates agents converging on consensus. Decreasing confidence (68% → 52% → 45%) shows successful challenge of a false positive. Evidence cards display each region of interest with a status badge (Challenged, Consensus, Disputed) based on the debate outcome, an AI-generated summary of key reasoning points, and an expandable section showing complete agent contributions for radiologists who want detailed analysis.

This visibility serves multiple purposes. For explainability, you see why agents reached conclusions, not only what they concluded. For error detection, visible debate helps you spot flawed reasoning. For confidence calibration, watching agents debate each other helps you assess reliability. For educational value, radiologists learn from agent reasoning, improving their own analysis.

Step 6: Integrate Amazon Nova Act for legacy systems

Implement browser automation to submit findings to legacy systems by adding the following code:

import json
import boto3
from typing import Dict

class LegacyRISSubmission:
    """Browser automation for legacy RIS submission."""

    def __init__(
        self, secret_name: str = "ris-credentials", region_name: str = "us-east-1"
    ):
        """Initialize with AWS Secrets Manager configuration."""
        self.secret_name = secret_name
        self.secretsmanager = boto3.client("secretsmanager", region_name=region_name)

    def _get_credentials(self) -> Dict:
        """Retrieve credentials from AWS Secrets Manager."""
        response = self.secretsmanager.get_secret_value(SecretId=self.secret_name)
        return json.loads(response["SecretString"])

    def submit_report(self, report_data: Dict, ris_url: str) -> Dict:
        """Submit validated findings to legacy RIS."""
        from nova_act import NovaAct
        from nova_act.types.workflow import Workflow
        creds = self._get_credentials()
        with Workflow(
            model_id="us.amazon.nova-act-v1:0",
            workflow_definition_name="radiology-submission"
        ) as workflow:
            with NovaAct(
                starting_page=f"{ris_url}/login",
                headless=True,
                workflow=workflow,
            ) as nova:
                # Authentication
                nova.act("Click on the username input field")
                nova.type_text(creds["username"], sensitive=True)
                nova.act("Click on the password input field")
                nova.type_text(creds["password"], sensitive=True)
                nova.act("Click Sign In button")
                # Patient lookup
                mrn = report_data["patient_mrn"]
                nova.act(f"Search for patient MRN '{mrn}'")
                nova.act("Click View Record button")
                # Report submission
                nova.act("Click New Report button")
                findings_text = report_data["findings_summary"]
                nova.act(f"Fill findings textarea with: {findings_text}")
                nova.act("Upload annotated image file")
                nova.act("Click Submit Report button")
                # Capture confirmation
                result = nova.act("Find and return the RPT- confirmation number")
        return {"status": "success", "confirmation": result, "mrn": mrn}

Security note: Nova Act captures prompts and screenshots as trajectory data. Never interpolate credentials or PHI into act() commands. Use type_text with sensitive=True to prevent credential capture in logs and trajectories. In production, suppress trajectory capture entirely for authentication steps, or route trajectory storage to a KMS-encrypted, access-controlled bucket subject to your BAA

Many enterprise systems, particularly in healthcare, lack modern APIs. Hospital RIS and EMR systems often run on decades-old technology stacks that can’t be modified without significant effort. Amazon Nova Act offers a pragmatic solution: browser-based automation using natural language commands.

This code implements browser automation that interacts with legacy systems through their existing web interfaces, retrieving credentials securely from AWS Secrets Manager and streaming each action back to your interface for transparency.

Each Nova Act command streams back to the radiology portal, so you can observe the submission process. Your interface displays a live action log showing authentication, navigation, form filling, and confirmation capture. This transparency helps you understand what the automation is doing and intervene if issues arise.

Browser automation requires careful credential management. The implementation stores credentials in AWS Secrets Manager, retrieves them at runtime, and never exposes them to the frontend. Audit logging captures Nova Act actions for compliance and troubleshooting. For production deployment, additional controls include IP allowlisting, session timeout enforcement, and multi-factor authentication where supported by legacy systems.

Step 7: Deploy to AWS

Configure DynamoDB tables with customer-managed KMS encryption:

import boto3

dynamodb = boto3.client("dynamodb")

dynamodb.create_table(
    TableName="radiology-sessions",
    KeySchema=[{"AttributeName": "session_id", "KeyType": "HASH"}],
    AttributeDefinitions=[{"AttributeName": "session_id", "AttributeType": "S"}],
    BillingMode="PAY_PER_REQUEST",
    SSESpecification={
        "Enabled": True,
        "SSEType": "KMS",
        "KMSMasterKeyId": "arn:aws:kms:us-east-1:ACCOUNT:key/YOUR-KEY-ID"
    },
)

Deploy your solution using Amazon Bedrock AgentCore, which offers managed runtime for agents with built-in identity, memory, and observability. AgentCore supports long-running tasks (up to 8 hours), asynchronous tool execution, and native CloudWatch integration, making it ideal for production deployments requiring minimal operational overhead.

Pattern selection guidance

Choosing the right architecture pattern depends on problem characteristics and requirements.

When to use dynamic agent-generated UIs

Characteristic Dynamic UI (AG-UI) Static UI (Traditional)
Output Variability High – unpredictable number/structure of results Low – consistent data structure
Multi-Agent Value Visible collaboration builds trust Single agent or no collaboration to show
Interface Complexity Varies per case Consistent across cases
Explainability Needs Important – users must understand reasoning Less important – results speak for themselves
Development Effort Higher – protocol integration, component library Lower – standard REST API

Swarm compared to other multi-agent patterns

Use swarm when multiple perspectives improve accuracy (peer review, consensus building), debate process offers value (explainability, error detection), no clear hierarchy exists (agents are peers, not supervisor/worker), and iterative refinement is beneficial (findings improve through rounds).

Use agents as tools when clear task decomposition exists (supervisor delegates to specialists), subtasks are independent (parallel execution possible), hierarchy is natural (manager coordinating experts), and no need for peer debate exists (each agent’s output stands alone).

Use sequential workflow when strict ordering is required (step B needs step A’s output), checkpoints are needed (validate before proceeding), process is well-defined (known stages, dependencies), and no benefit from parallel exploration exists (linear pipeline).

Use cases beyond medical imaging

This architectural pattern applies to domains with high output variability and multi-agent value. For fraud detection, you encounter variable evidence chains (1-20 suspicious transactions), multiple analyst perspectives (financial, behavioral, network analysis), and legacy banking systems requiring browser automation. For legal document review, you face unpredictable numbers of relevant clauses, expert opinions from different legal domains (contract law, regulatory compliance, risk assessment), and integration with legacy case management systems. For security event response, you deal with variable threat indicators, team collaboration (network analysis, malware analysis, threat intelligence), and legacy Security Information and Event Management (SIEM) systems without modern APIs.

Anti-patterns

Avoid this approach when you have simple classification (predetermined categories with fixed confidence scores), deterministic calculations (no uncertainty or debate needed), low-latency requirements (multi-round debate adds latency), cost-sensitive scenarios with predictable outputs (swarm pattern increases token usage), or minimal explainability needs (users trust results without seeing reasoning).

Clean up

To avoid incurring future charges, delete the resources you created while following this walkthrough. Delete them in the following order. The sequence removes the frontend and compute layers first, then the data stores, and finally the encryption keys and IAM roles, so no deletion fails because another resource still depends on it. Because this solution can process PHI, this teardown also removes any stored medical images, patient identifiers, and findings.

  1. Amazon CloudFront — Disable the distribution, let it propagate, then delete it. This stops traffic and releases the S3 app bucket.
  2. Amazon API Gateway — Delete the REST API and SSE endpoint.
  3. Amazon Bedrock AgentCore — Delete the Gateway, then the Runtime instances (and their identity, memory, and session resources). This stops all agent activity against the downstream services.
  4. AWS Lambda — Delete the AG-UI handler and any other functions.
  5. Amazon Bedrock model access — Confirm nothing is still invoking the models. Revoke access to Claude Sonnet or Amazon Nova Act if enabled only for this walkthrough.
  6. Amazon OpenSearch Serverless — Delete the knowledge base collection and its access, network, and encryption policies. OCUs bill hourly.
  7. Amazon DynamoDB — Delete the radiology-sessions table and any other tables you created.
  8. Amazon S3 — Empty (including all object versions) and delete the medical-image and app-hosting buckets.
  9. AWS Secrets Manager — Delete the legacy RIS/EMR credential secrets. A 7 to 30 day recovery window applies unless you force deletion.
  10. Amazon Cognito — Delete the user pool and app client.
  11. AWS KMS — Only now, schedule deletion of the customer managed keys used for Amazon S3 and DynamoDB. Deleting them earlier can make encrypted data unrecoverable.
  12. AWS IAM — Delete the roles and policies created for Lambda and AgentCore.
  13. Amazon CloudWatch — Delete the log groups, dashboards, and alarms (kept until last for troubleshooting).

Finally, review the AWS Billing and Cost Management console to confirm no unexpected charges remain.

Conclusion

In this post, we showed you how to build adaptive AI interfaces using three AWS technologies. You learned to use the AG-UI protocol for dynamic UI generation, apply Strands swarm patterns for multi-agent collaboration, and integrate Amazon Nova Act for legacy systems. Through a complete radiology assistant implementation, you saw how these technologies work together to handle variable AI outputs while maintaining explainability and practical enterprise integration.

The AG-UI protocol creates interfaces that adapt to what your agents discover, not what you anticipated. Use it when output variability is high and you can’t predetermine interface requirements. The Strands swarm pattern creates explainability through visible multi-agent debate. Use when multiple perspectives improve accuracy and you benefit from seeing reasoning processes. Amazon Nova Act bridges the gap with legacy systems lacking APIs. Use when browser automation is the only viable integration path and action transparency matters.

These technologies work together naturally because they address different aspects of the same problem: building AI systems that are flexible, explainable, and practical for real-world enterprise environments.

Next steps

Read related AWS blogs:

  • Multi-Agent Collaboration Patterns with Strands Agents and Amazon Nova.
  • Strands Agents SDK: A Technical Deep Dive into Agent Architectures and Observability.
  • Build a Drug Discovery Research Assistant using Strands Agents and Amazon Bedrock.

 


About the authors

Using AI to chart a course for our post-quantum migration

Post Syndicated from Sharon Goldberg original https://blog.cloudflare.com/ai-driven-cryptography-discovery/

As laboratories around the world race to build out a cryptographically relevant quantum computer, we at Cloudflare are racing towards a 2029 target deadline for full post-quantum readiness. While we’ve already transitioned many of our products to post-quantum encryption, we still have work to do to support post-quantum authentication and achieve full post-quantum readiness across our platform.

We’re taking a maximalist stance (“PQ everything!”), because as an infrastructure provider to the world, we want to give our customers the peace of mind that using Cloudflare ensures that their traffic is future-proofed against quantum adversaries.

But how does one accomplish such a massive migration at an organization of our size and scale? After all, cryptography is the base layer for almost all of the world’s digital systems, including the software services and the networking protocols that power our platform.

To drive our PQ migration, we have three key goals.

First, we want to help our product and engineering teams understand how cryptography is being used and how they should be upgrading it. This should cover both the upgrades to post-quantum encryption and to post-quantum authentication. Many of our products have already been upgraded to post-quantum encryption over TLS 1.3, but we still want to cover the long tail of TLS connections, as well as upgrade any other uses of public-key encryption. Meanwhile, it’s still early days for our deployment of post-quantum authentication.

Next, we want to provide progress metrics for the migration. These might include per-repository and per-product counts of the use of classical and post-quantum cryptography.

Finally, we want to surface prerequisites early. If our products or platform rely on protocols that don’t yet have a PQ migration plan (because PQ variants of the system have not yet been considered, because PQ standards do not exist or lack consensus, or because software libraries or other key ecosystem components do not yet have PQ support), then we need to know now. That way we can work with the relevant stakeholders, standards bodies and ecosystems to help drive their PQ migration plans, so that we can meet our own 2029 PQ migration timeline.

This post is the story of how we’re going about this. We explain how we turned to AI to help us solve some of our problems and how we’re developing an internal tool called CryptoLabe to help us. CryptoLabe is named after the mariner’s astrolabe, a navigation instrument refined by Portuguese navigators. Just as an astrolabe helped sailors determine where they were and chart a course, CryptoLabe helps us discover cryptography in our code, understand how it is used, and chart a path to post-quantum migration.

CryptoLabe is highly specialized to our internal systems (our repositories, our ticketing systems, and internal documentation processes) and still evolving as we continue its development, so we aren’t making it available to customers. Nevertheless, we are sharing our learnings so that other organizations can build upon our efforts as they work through their own PQ migration journey.

The scale of the problem

The software that powers most Cloudflare products lives inside our single centralized source control management platform. This means we can find most uses of cryptography across our platform by just looking through our codebase.

While the centralization of our codebase is a marked advantage for us, we still need to contend with three challenges that come with the scale of this problem. First, our code is spread across many repositories. Second, cryptography rarely announces itself plainly in the code. Instead, it hides in

  • shared libraries that a repository imports but may or may not actually call
  • upstream and protocol defaults, like a TLS 1.3 listener that is configured to negotiate a classical key exchange such as X25519 rather than post-quantum X25519MLKEM768
  • configuration files that select algorithms far away from the code that uses them, like a TLS responder whose key exchange protocols are pinned in a YAML file stored in a different repository
  • code paths that are dead, test-only, or on a path to being deprecated

Third, cryptography discovery is about more than just pattern matching. Grepping for certain algorithm names (e.g. “RSA” or “X25519”) overcounts, because it finds cryptography in unused code. Grepping also undercounts, because it misses defaults and indirect uses in dependencies and configuration. Most importantly, it can't tell you how the cryptography is used. A classical ECDSA signature could be part of a JWT, IPsec, TLS, or SSH, and each has a completely different migration path. Many uses also depend on the other side of the connection: a TLS server may support both post-quantum key exchange and classical key exchange; the one it chooses to use would depend on the client.

Turning to AI

It turns out that AI is pretty good at doing more than just grepping. A model can search a codebase, follow evidence across files, and return structured analysis. It can also enrich findings by pulling information from other sources, like our internal documentation and ticketing systems. In fact, AI can even explain how cryptography is being used and how it should be updated. We’ve been putting that idea to the test as we develop CryptoLabe.

As we said before, our first two goals are to (1) discover and understand the use of cryptography in our codebase, and also (2) to get metrics on the state of our PQ migration. Towards these goals, our current implementation of CryptoLabe performs scans in two stages, as shown in the figure below.

The first “discovery” stage starts by mapping the repository. It then searches for cryptography through source, configuration, manifests, lockfiles, scripts, tests, and documentation. Among other things, the scan looks for the use of cryptography like key agreement, signatures, asymmetric encryption, PKI, tokens, credentials, hardware security module integrations, and more. This discovery stage produces a set of "raw observations."

Each raw observation feeds a run of the second stage. This “analysis” stage first re-checks the observation against the source code. It then investigates how the cryptographic operation is used at runtime, what role the repository plays, and which internal or external parties it depends on. When necessary, it can inspect related code in other repositories to complete the analysis. Finally, it takes a pass over its own conclusions, searching for missing or conflicting evidence such as configuration overrides, test-only code, or incorrect assumptions about runtime behavior.

Next, the model assigns a classification to the finding. If there is not enough evidence to assign a classification, the model assigns More evidence needed, External dependency, or Unknown rather than guessing.

This is the current list of classifications used by CryptoLabe, containing catch-all classifiers which will likely be refined as we proceed through our migration. (As an example, we could refine our classifiers by splitting the “encryption” classifier into key agreement and HPKE; you get the idea.)

Classification

Examples

Classical encryption

This is a catch-all category that finds cases of elliptic-curve Diffie-Hellman key exchange (ECDHE) (e.g., X25519, P-256, P-384), RSA key agreement or other uses of public-key encryption (e.g., HPKE). These are broken by a quantum computer running Shor's algorithm, which puts them at risk of harvest-now-decrypt-later attacks.

Classical signature

This is a catch-all category that finds use of an RSA signature or elliptic-curve (ECDSA) signature in anything, for example a certificate, a TLS handshake, another protocol handshake. These signatures are broken by Shor's algorithm.

Classical token

We found a lot of RS256 or ES256 JWT tokens, so we created a special classification for them. These are JWTs that use classical RSA and ECDSA signatures; RFC 9964 defines a post-quantum replacement using ML-DSA.

PQ-ready hybrid key exchange

Finds hybrid post-quantum key exchange in TLS 1.3, i.e. X25519MLKEM768. This is the most prevalent use of PQ encryption in our codebase.

PQ-ready

Finds other uses of post-quantum cryptography that are not X25519MLKEM768 in TLS 1.3, like ML-DSA.

Finally, it generates a report that serves two audiences: (1) product managers who need to understand what the migration means for their product, and (2) engineers that need enough detail to execute the migration.  

Here’s a (cropped) view of one of our reports:

While we’ve been iteratively reviewing findings against the source code and with relevant engineers, we do not yet have a ground-truth dataset for reproducibly comparing different versions of the prompts we’ve tried for CryptoLabe.

Built on Cloudflare’s Developer Platform

We built CryptoLabe on Cloudflare's Developer Platform. Here’s the architecture:

CryptoLabe runs across two Cloudflare Workers. There’s a scanner Worker that runs the scans. And there’s an inventory Worker that serves the dashboard, exposes the API, and stores everything in a D1 database. The two communicate through Service Bindings. A scan starts when someone requests it from the dashboard, and the inventory Worker passes the request to the scanner.

Orchestrating a scan

We need a way to keep a scan alive and on track from start to finish, without building our own job orchestration system. We did this with Agents SDK. Each repository gets its own persistent coordinator built on a Durable Object (DO). A bounded queue in front of the coordinators limits how many scans run at once. When a scan's turn comes, the coordinator tracks its progress and handles cancellation, retries, and recovery.

The coordinator doesn't do the analysis itself. It hands the work to Cloudflare Workflows, so that they can persist progress and automatically retry failed steps. The coordinator moves each repository through four stages:

  1. discovery Workflow (the first scanning stage that produces raw observations)
  2. deep analysis Workflow (the second stage, run on each raw observation)
  3. merge Workflow (that builds a list of findings for a given repository, including combining repeated or similar finds)
  4. publish workflow (that hands results back to the inventory Worker)

The first two workflows need the model to have access to the repository's code. We want this access to be isolated, so we don’t risk damaging the codebase. That’s why CryptoLabe downloads the repository once, at an exact commit, at the start of each scan, and then stores that snapshot in R2. Each Workflow then restores the snapshot into a fresh, short-lived Cloudflare Sandbox, an isolated container. The model then works with the Sandbox through a small set of read-only tools on an immutable snapshot of the code, even if the codebase changes while the scan is still running.

Calling the model at scale

If we want to scan through all of our (many!) repositories, we have to worry about both cost and capacity.

For cost, the model loop sends its requests through AI Gateway to cost-effective open-weight models hosted on Workers AI. Putting the model behind AI Gateway also makes it easy to switch models as better or cheaper ones become available.  

Capacity became a problem once we scanned many repositories at once. Bursts of model requests began triggering HTTP 429 (rate limit) responses from AI Gateway, and scans retrying independently only made the bursts worse. We solved this with a single, global Durable Object that paces every model request across all scans, including retries. When any scan hits a rate limit, the cooldown is shared and all scans back off together, so concurrent scans share the available capacity instead of competing for it.

Prerequisites and hard cases

Let’s now get into our third goal: surfacing prerequisites and hard cases early.

A lot of ink has been spilled about ecosystem readiness for the PQ migration, and we are now going to spill some more. As everyone knows, a PQ migration cannot happen in a vacuum. For migration to succeed, post-quantum cryptography must be supported in relevant software libraries (e.g. BoringSSL) and across parties that participate in the ecosystem (e.g. clients, browsers, origins, cloud proxies, certificate authorities, etc.). Standards are also an important indicator of ecosystem support, although a standard that is still in “draft” state does not necessarily mean deployment cannot proceed. As an example, we deployed X25519MLKEM768 in TLS 1.3 back in 2022 when it was still a “draft” at the Internet Engineering Task Force (IETF) while it was only finalized as RFC 10024 in 2026.

Either way, our point is that in order to upgrade a system to PQ cryptography, we need to understand its dependencies and level of ecosystem support. 

That’s why CryptoLabe uses the concept of “prerequisites” to highlight findings that cannot be immediately remediated by an individual product team working alone.

A prerequisite can be something as straightforward as “we are currently blocked on migrating to post-quantum JWTs.” We say this is straightforward because there is already a standard (RFC 9964) for post-quantum JWTs. Nevertheless, if our software libraries don’t yet support validating post-quantum JWTs, or if we’re using a token issuer that does not yet issue post-quantum JWTs, we can’t go company-wide and ask each of our product teams to start PQ-ing their JWTs. This migration is blocked until we solve its core prerequisites. CryptoLabe lets us group together findings that (likely) have the same prerequisite, which also helps us decide how to prioritize resolving these prerequisites.

For example, the snapshot below shows the six findings from CryptoLabe that have post-quantum SAML as a prerequisite. (SAML is a protocol for single sign-on (SSO).)

On the other hand, there may be uses of cryptography that lack even a basic level of ecosystem support. We’ve been calling these “hard cases.” To find them, we wrote a separate prompt that ignores “vanilla” uses of cryptography (e.g. ordinary TLS between internal systems) and instead looks for custom cryptographic protocols, keys, or signatures used in size-constrained fields, cryptography built into hardware, specialized cryptographic constructions (like blind signatures), protocols without a PQ standard, and dependencies on external parties that do not yet support PQ cryptography.

This prompt is shorter and simpler than those used for CryptoLabe, since its only job is to find hard cases.  In our qualitative review, we found that it got better results when it ran in one fell swoop against all our repositories, while also taking in context from our internal ticketing and documentation system.  

Here’s an example of a “hard case” we found: a certificate carried in an HTTP header. Post-quantum certificates and signatures are larger than their classical counterparts, so if the header (or an intermediary, or the application processing the header) assumes a certificate has a certain size, changing the signature algorithm may break the system. Our next step is to determine whether this code will remain in use in the long term. If it will, we need to measure the relevant size limits and decide how to accommodate the larger certificate.

An important lesson here is that no single scan finds everything. Our repository-by-repository scans were effective at discovering common uses of cryptography. Meanwhile, this targeted scan worked better for “hard cases” because it ignored well-understood cryptography and had more context about each product and its dependencies.

The bottom line is that different approaches find different things, and every finding still needs to be checked by the engineers who understand how the system actually works.

Sharing our prompts

We’ve been messing around with the best way to write prompts for CryptoLabe for the last several months.  We don’t yet have a ground-truth dataset for comparing one prompt’s performance against another, and we are not convinced we have 100% coverage of all uses of cryptography in our codebase. Instead, we have iterated by running scans, reviewing findings with the engineers that maintain the repositories, investigating misses that came up during these reviews and revising the prompts.   Nevertheless, we decided to publish selected prompts, so other teams can learn from and adapt our approach. These prompts are starting points, not a standalone version of CryptoLabe, and the quality of their results will depend on the model, tools, context, and engineering review available.

Thinking through your own PQ migration

At Cloudflare, we’re taking a maximalist approach to our PQ migration because of our goal of acting as a provider of post-quantum cryptography for customers and the Internet at large. But most organizations do not need to start by finding every use of cryptography in every repository in every one of their products. In fact, most organizations should not be doing this, because at this time it's a waste of precious resources.

Before scanning a single repository, you can protect traffic in bulk wherever possible. If your websites run through Cloudflare, we protect your data in transit with post-quantum encryption already today; check this out with our new PQ visibility features. Our SASE platform, Cloudflare One, provides post-quantum encryption for private network traffic. Post-quantum encryption is provided at no additional cost and without requiring you to upgrade every origin server or private application on your enterprise network. This gives you a compensating control while you work through discovering and understanding the use of cryptography inside your own systems.

An exhaustive cryptographic inventory is not a prerequisite for action. Instead, organizations should first identify the systems whose compromise would matter most, discover their use of cryptography, and then PQ that cryptography in priority order. Here is one way to begin:

  1. Choose a repository for one important system. Start with something that handles sensitive or long-lived data, authenticates users or software, or is exposed to the public Internet.
  2. Run cryptography discovery against that repository. We hope our description of CryptoLabe will be helpful to this effort!
  3. Validate the results. Ask the team who owns the system to validate the results of cryptography discovery and confirm that the cryptography finding is needed long term and needs to be upgraded to PQ. It’s important to remember that it might not need to be immediately upgraded to PQ if there is another compensating control in place.
  4. Prioritize action. Figure out what upgrades you can make now and what upgrades are blocked. Record shared prerequisites that need help from a library, vendor, standards group, or another part of your organization. Prioritize your findings and make a plan for addressing the highest-impact systems and prerequisites first.

That gives you the beginning of a PQ transition plan, without requiring a complete map of every cryptographic operation in your organization. CryptoLabe is still ever-evolving, but its scans and results have been illuminating to us as we plan our migration. We hope these shared learnings will be useful as you continue to work through your own PQ migration.

Acknowledgements: Many people across Cloudflare provided feedback on and contributed to CryptoLabe, including Davide Marquês, Peter Wu, Phil Schmieder, JP Aumasson, Andrew Galloni, Christopher Patton, Luke Valenta, Mari Galicer, Vânia Gonçalves, and the Client, Tunnel and Gateway teams who reviewed reports produced by the tool.

Building a certificate authority for the whole Internet

Post Syndicated from Steve Goldsmith original https://blog.cloudflare.com/cloudflare-certificate-authority/

Twelve years ago, during Birthday Week 2014, we turned on Universal SSL and nearly doubled the number of encrypted sites on the web overnight, giving free TLS to every site behind Cloudflare, including the ones that never paid us a cent. Encryption stopped being an expensive, time-intensive undertaking and instead became the default.

For Birthday Week this year, we are taking the next step on that path. For more than a decade we have been one of the largest consumers of publicly trusted certificates on the Internet, and have never issued a single one ourselves. That is changing. Cloudflare is announcing our intent to become a public certificate authority (CA).

Today we are announcing the first concrete milestones in that effort: We have applied for inclusion in the Chrome, Apple, Microsoft, and Mozilla root programs, and we have signed a definitive agreement to acquire an established, broadly trusted root from GlobalSign, so that we can offer certificates with the widest possible device reach the day we begin issuing. We’re also announcing our plans to be one of the first CAs to serve post-quantum certificates, targeting Chrome’s recently announced Quantum-resistant Root Program.

We are not issuing certificates yet, and it will be a little while before we do. What we are doing is committing to the work in public, sharing the milestones as they land, and telling you exactly what we are building while working with the root programs and other members of the WebPKI community to achieve this.

Two paths to trust

A brand-new root is not widely useful for years. Even after a root program accepts it, that root has to propagate out into the world's operating systems, browsers, and devices, and it never reaches the large set of devices that have stopped receiving updates, or never received them in the first place. That long tail of older clients is where a great deal of the world’s Internet traffic originates, and where a correspondingly large set of avoidable breakage lives. We believe that all clients deserve the highest level of security possible, regardless of their manufacturer, operating system, or time since last update.

Acquiring an existing root with a high degree of trust store coverage across a diverse set of clients solves that on day one. The existing GlobalSign root has been trusted across browsers, operating systems, and devices since 2012, and it reaches older clients that a fresh root never will. The new root that we will be submitting for inclusion in root key programs is built for where the ecosystem is heading, including the programs that are starting to cap how old a trusted root may be. The established root gives us reach across the devices of the past. The new roots give us standing under the policies of the future. We want both to ensure certificates issued by our CA provide the widest set of customer compatibility possible.

A new source of free certificates

The free-of-charge, automated certificate model now carries most of the encrypted web, and much of it runs through one remarkable operator. Let's Encrypt issues on the order of ten million certificates a day, serves more than 500 million sites, and passed four billion active certificates in 2025. It is one of the best things to happen to the Internet in twenty years, and we say that as one of its largest users.

That success comes with some systemic risk: if the dominant free certificate authority had a bad week, much of the web would have no comparable free, automated alternative ready to take the load. At the certificate pack level, we have spent years building exactly this kind of redundancy for our own customers. Every Cloudflare Universal SSL certificate already ships with a backup certificate, wrapped with a separate key and issued from a different authority, ready to deploy automatically if the primary is ever revoked or compromised. A public CA is that same idea, but at the scale of the whole Internet.

To make it easy to adopt, we will be Automated Certificate Management Environment (ACME)-first, an open standard protocol that is widely accepted. Automated issuance and renewal through ACME will be the way you get a certificate from us, which means anyone already pointed at any existing free CA can move to us by changing a directory URL, with no new tooling and nothing to re-architect.

Certificate growth projections are huge

Cloudflare sits in front of more than 20 percent of global Internet request traffic and terminates TLS for millions of domains, relying on millions of certificates per year to do so. We provision those certificates through multiple CAs, with primary and backup paths so customer services stay up through CA outages and revocation events.

That has taught us not just how the WebPKI ecosystem works, but also that it occasionally fails, from the consuming side, the hard way. We have dealt with rate limits, validation edge cases, revocation latency, chain building, and root distribution lag. We have lived through the CA churn of recent years and felt it through our customers. We know what reliable issuance has to look like from the outside, because our customers' uptime has depended on us being resilient and responsive when an issuer has a bad day.

And as certificate maximum validity period decreases over the next few years, agentic activity increases, and PQ certs go mainstream, we expect the raw number of certificates we rely on annually on to continue to grow, quickly — and we are not alone. We want to not just solve this problem for ourselves, but be part of providing this utility to the Internet, and ensure that the certificate supply chain for our customers has even more providers.

Designing for resilience: transparency and fail small

In taking on this new responsibility of being our own CA, we're committed to making the most reliable and resilient CA possible. We intend to build a certificate authority whose reliability depends not just on avoiding mistakes, but as with the rest of Cloudflare’s products, to “fail small” and limit the impact of any one issue.

That means instituting processes to design and test recovery before any incident occurs. As an example, we will make renewal automation a condition of issuance. We will only issue to clients that support ACME Renewal Information (ARI), standardized in RFC 9773. Subscribers must maintain automation that polls our renewal endpoint, acts on the renewal windows we publish, and identifies the certificate it is replacing.

We're also learning from what we've observed over the past 16 years. We have seen certificate authorities caught between timely revocation and keeping subscribers’ sites online because too many subscribers could not replace their certificates quickly enough. When certificates need to be retired, whether for a compliance issue or a security incident, we can bring forward renewal windows for the affected certificates, spread replacements across the available time, and track replacement issuance.

This is just one of the many ways we intend to build. We will be transparent with our issuance stack and operations, publish reproducible builds of the software that signs certificates, attest the hardware security modules that hold our keys, and run a public dashboard for issuance health and incidents. Audits are point-in-time and tell you a CA passed, not how it runs on an ordinary Tuesday. We want root programs, researchers, and ordinary site owners to watch how a modern CA actually operates between audits.

A certificate authority for the post-quantum Internet

We also intend to lead on where certificates are going, not just where they are. We plan to be one of the first CAs to issue production Merkle Tree Certificates (MTCs), with the first certificates issued in the first quarter of 2027.

MTCs are a new and far more compact way to deliver publicly trusted certificates, designed for a post-quantum world where traditional certificate chains grow large enough to strain TLS handshakes. We have been championing the standards-based proposal for MTCs at the IETF, and earlier this year, Chrome named MTCs as the preferred path for post-quantum authentication. Issuing them in production allows us to protect Cloudflare customers as well as the wider Internet against the post-quantum threat, with real volume behind a transition the whole web has to make. We’ve shared much more about MTCs and what this new Web Public Key Infrastructure (PKI) will look like in a blog post on the topic.

We do not expect that transition to be sudden. Much of the Internet will continue to rely on classic certificates and existing WebPKI for many more years. But across that window we expect MTCs to take a steadily growing share of issuance, and that is why we are building one service that does both. By carrying classic certificates and Merkle Tree Certificates under one CA, with one lifecycle and one set of guarantees, customers can adopt at the pace that suits them and help the web make the crossing without a hard cutover. Customers should not have to pick a side of a multi-decade migration, run two systems, or rebuild when the balance shifts.

As always, Cloudflare will be Customer Zero

In addition to providing certificate packs via Universal SSL for our customers, Cloudflare consumes certificates from many different CAs to run our systems and internal operations. Just like our other products, we will be Customer Zero for the new CA and its certificates (both WebPKI and MTC), ensuring that all aspects of the new systems and processes meet our high internal standards, and that our CA’s infrastructure is exercised at Cloudflare scale.

What happens next

We are working through the application and approval process with each of the core web root key programs. These processes happen in the open, and we’ll share more updates as they proceed, through to the first Merkle Tree Certificates in early 2027. If you want to follow this work or be one of the first to use a Cloudflare CA certificate in the future, you can register for updates.

As we build out this new capability, we will continue to work closely with the network of partner public CAs we have relied on for many years — 16 in fact! — as we all work together to ensure a trusted and open Internet.

When we launched Universal SSL, the argument was simple: every byte that flows encrypted across the Internet makes it harder to intercept, throttle, or censor, and the open web is something we all build together. A public, redundant, transparent certificate authority is that same argument carried one layer down, to the trust that makes the encrypted web possible in the first place. We have been working toward this for a long time, and we are glad to finally be on the road.

Happy Birthday Week!

Adaptive application security for the AI era: how Cloudflare connects code, traffic, and intelligence to stop attacks

Post Syndicated from Daniele Molteni original https://blog.cloudflare.com/ai-era-framework/

In July, AI agents testing new cybersecurity models compromised parts of OpenAI’s infrastructure and Hugging Face’s production environment.

We've all just witnessed one of the first AI-driven successful cyber attacks. When given a task, the agents ignored existing guardrails and autonomously discovered previously unknown vulnerabilities, recovered exposed credentials, moved between cloud environments and coordinated their work through communication channels they created themselves.

The speed of the final compromise was incredible. In under 13 hours, the agents went from executing code on a Hugging Face worker to gaining admin-level access across multiple clusters. But the incident had been brewing for much longer. Responders found clues of activity tracing back to May (agents created an unauthorized message board), to June (internal network scanning) and early July. The relationship between these events was understood only on July 20.

The lesson here is not that AI agents exploit vulnerabilities. That’s not news; human attackers already do that. The change is that agents can work persistently, test multiple paths simultaneously, share discoveries, and chain vulnerabilities, credentials, and permissions into sophisticated attacks.

The incident also shows why application security cannot depend on single tools. For example, network restrictions were bypassed by services connected to the Internet; valid credentials were used to perform unauthorized actions. Rebuilding Artifactory removed one attack path, but agents found another. The key insight is that individual alerts identified pieces of the activity without revealing the complete campaign. OpenAI reached a similar conclusion in its report: organizations need overlapping and independent controls across prevention, detection, and mitigation, continuous validation of security boundaries, and faster mechanisms to correlate and contain suspicious behavior.

We address this challenge by connecting application security across four activities that are too often separated: discovering which risks matter, governing what humans and agents may do, protecting applications at runtime, and turning every investigation into stronger protection. Cloudflare can deliver this framework because of its broad security portfolio and visibility across a vast share of Internet traffic.

Alongside the framework, we connect existing Cloudflare solutions with new capabilities across each stage. These include: using Large Language Models (LLMs) to conduct a penetration test of our Web Application Firewall (WAF), expanding threat intelligence to all customers, and a new feature to automate deploying positive security.

What has changed

The security landscape is shifting. These are the emerging trends we see:

  • The way we build software has fundamentally changed. AI-assisted development allows engineers to produce and deploy software faster outside traditional engineering workflows. That speed creates both more code and more opportunities for vulnerabilities to reach production.
  • Software composition risk is still a risk: applications depend on large chains of open-source libraries, packages, and operating-system components that are intrinsically trusted and are difficult for any team to inspect. What’s new is that AI is now importing libraries that we might not be aware of.
  • Techniques and tactics are changing. LLMs can chain vulnerabilities and use feedback in real time to mutate payloads, evade defenses, and make decisions autonomously. They can operate continuously and at machine speed. Patching faster remains important, but patching alone cannot close the gap. Attackers are always going to be faster than you can update your systems.
  • Agentic traffic. In the past, automation was a synonym for malicious activity. Today, a request generated by an agent or bot may be malicious automation, a search crawler, or an agent purchasing a product on behalf of a customer.
  • Compromised servers, residential proxies, IoT devices, and cloud resources allow attacks to move quickly across infrastructure and identities. A coordinated attack can leverage a number of devices, making it difficult to be identified as a unique campaign.

Application Security’s goal is also expanding. It now needs to address three connected problems: protecting conventional applications from AI-enabled attackers, governing legitimate and malicious agentic clients, and securing applications that contain models, agents, tools, and data.

A connected application security framework

Application security in the AI era must operate as a continuous system rather than a collection of controls that teams update after each new vulnerability. To protect applications in the era of AI, you need to work on multiple activities, which we have organized around four stages:

  1. Discover and prioritize risks
  2. Govern access and agent behavior
  3. Protect applications at runtime
  4. Investigate, respond and learn

None of these activities is new in isolation. What changes is connecting them so discoveries, runtime signals, and investigation outcomes continually improve the controls that follow.

More than 20% of the web sits behind Cloudflare’s network, which gives us visibility into attack infrastructure, payload mutations, emerging techniques, and coordinated campaigns at a scale that few organizations can match. Patterns that look isolated from the perspective of one application can become clear across our network. This combination of global threat intelligence, local application context, and inline enforcement powers every stage of the framework. That local context includes which code is deployed, which endpoints are exposed, what legitimate traffic looks like, which identities are acting, and which controls are already active. Because Cloudflare is inline, we can turn those insights into protections immediately.

Cloudflare is the adaptive security control plane for applications, APIs, and agents. Here is what we are launching today to advance every stage of the security journey.

Discover and prioritize risks

Security teams do not suffer from a shortage of findings. They struggle to determine which findings represent an immediate risk. A useful discovery system must connect vulnerabilities to the production reality, including whether a vulnerability buried in your stack is actually reachable in the first place. We see three main areas you should look into: software composition risk, proprietary code, and runtime penetration testing (pentesting).

Understand software composition risk

Applications inherit risk from open source libraries, packages, operating-system components, and the services on which they depend. This represents the Supply Chain of your application. A package vulnerability alone does not tell a team whether the affected component is deployed, reachable, or exposed to hostile traffic. Open-source software is the top priority when it comes to supply chain risk, and Cloudflare is part of Chainguard Athena, an industry coalition aiming at protecting open-source software from AI attacks.

Scan proprietary code

When it comes to code scanning, you have two options: getting a managed service or developing in-house expertise to run it yourself.

Cloudflare recently announced early access to Vulnerability Discovery and Remediation, a service that uses frontier models to identify application-specific vulnerabilities and deploy WAF mitigations to block targeted exploits while engineers fix the code. The important step is prioritization. Cloudflare connects source-code findings to production traffic and security signals. We can identify whether the affected route is active, and how much traffic it receives.

Pentest your application at runtime

Defenders can also use the same capabilities as attackers. Customers can build their own LLM-based pentesting harness to search for weaknesses, validate findings, and test whether their applications are vulnerable. Discovery becomes continuous rather than a periodic exercise. We have done this internally at Cloudflare since Anthropic’s Claude Mythos was released, and we shared our learnings.

A vulnerability buried deep inside your code is harder to exploit if it can’t be reached from the outside. Our Security Analyst team has already used LLM-based red teaming to test customer applications and our own runtime detections, turning the findings into improved detections for all customers. We are now developing Adaptive Security, a self-service capability that will periodically pentest selected URLs behind Cloudflare, using LLM-powered agents to identify vulnerabilities that are reachable and exploitable before attackers find them.

Govern access and agent behavior

Agentic traffic operates in the space between automation and human: tasks delegated by people, executed by software. This changes how access decisions must be made. Detecting automation is no longer enough. For every interaction, application owners need to answer two questions: Is this entity who it claims to be, and can this interaction be trusted?

These questions can be hard to answer. A recognized agent with a long history of legitimate activity may have high trust, but an unusual action can still create immediate risk. An unknown agent may simply be new; a lack of history does not necessarily mean malicious intent. Cloudflare’s approach is to keep trust and risk signals separate, thereby giving application owners more control than a single bot score or allow-or-block decision, and providing more powerful tools to quickly adapt to change.

Establish identity and trust

Trust accumulates over time, while risk is evaluated for each interaction.

Botbase provides a directory of known automated entities that have registered with Cloudflare. Registration gives legitimate bots and agents a way to declare who they are, while application owners retain control over whether and how those agents may access their sites. Cloudflare is also making registration more accessible to smaller and custom agents, building a verified identity layer across all agentic traffic, not just the major platforms.

Identity alone does not establish trust. Cloudflare can evaluate whether an entity has been seen before, whether its historical behavior was legitimate, and whether its current activity is consistent with that history. This makes it possible to distinguish a recognized agent behaving normally from the same agent suddenly changing established request patterns, location, identity, or transaction behavior.

Understand agentic behavior across the journey

Precursor adds client-side and session-level signals to distinguish human from automated behavior, such as typing cadence, mouse movement, navigation patterns, and sequences of actions. An agent that navigates a checkout flow in two seconds, skipping the browsing and comparison steps a human would take, reveals its nature through the session. These signals help identify whether behavior across a session is consistent with human interaction or automation, giving application owners a clearer picture of the traffic they're managing.

Manage access and adapt

Application owners can block traffic from AI crawlers and decide what activity is allowed on their asset (e.g. search, training, etc.). Adaptive Intelligence combines network, client-side, historical, and behavioral validation signals in a probabilistic model that can be updated as attackers change their techniques. Customer outcomes, including chargebacks and successful legitimate transactions, can feed back into the system to improve future decisions.

The result is a continuously updated assessment of every entity and interaction. Application owners can encourage known, useful automation while applying stronger controls when identity, history, and current behavior indicate greater risk.

Protect applications at runtime

Cloudflare’s reverse proxy protects applications at runtime by filtering traffic before it reaches the origin. In the AI era, a new layered approach is emerging to best filter traffic from malicious requests:

  1. Enforce positive security
  2. Detect attacks and identify LLM tactics and techniques
  3. Protect business logic
  4. Deploy real-time threat intelligence

Enforce a positive security model

You can dramatically reduce the attack surface by learning what legitimate traffic looks like, allowing conforming requests and blocking everything else. Today we are announcing Application Profiles, which automatically learns the structure of your web or API application and detects non-conforming requests. Application Profiles automates the learning process and adds a layer of interpretation. Based on the learned profile, we can understand the business logic of different endpoints and request parameters and help you prioritize what endpoints require more scrutiny and attention.

Detect attacks and identify LLM tactics and techniques

Traditional WAFs are designed to run highly crafted rules to detect Common Vulnerabilities and Exposures (CVEs) and malicious payloads. Before AI, the time to disclose new vulnerabilities was measured in months and days. Not anymore: now we see vulnerabilities being exploited before they are disclosed, so the time to patch is nearing zero.

Different tools can be deployed to detect known exploits. These tools include:

  • Managed Rules hardened with frontier models. We’ve partnered with major model providers to use frontier models for adversarial validation. We used the latest models to pentest the WAF to uncover bypasses and vulnerabilities. All customers benefit automatically from ongoing improvements.
  • Machine Learning detection. While signatures are great for high-precision attack detections, machine learning can stop attacks before they are discovered and disclosed. Attack Score detects attack mutations and evading techniques that are often used by LLMs. Attack Score is available to all Cloudflare Customers
  • AI Security for Applications. Chatbots and Internet-facing LLMs are subject to a new class of attacks, such as prompt injection and sensitive data exposure. You can protect generative AI traffic by deploying guardrails and security detections designed to stop these attacks.

Protect business logic

Attackers can still craft legitimate requests and abuse business logic to gain advantage on the application. For example, an attacker uses a valid password-reset flow repeatedly to take over accounts. Fraud detection tools, including account takeover and leaked credential detections, help prevent abuse in which the request appears legitimate, but the intent is malicious.

Real-time threat intelligence detection

Back in June, we launched always-on detection based on our threat intelligence feeds. Cloudforce One customers can deploy protections to block requests originating from compromised infrastructure. We are now also expanding access to Cloudforce One’s Threat Events Platform, our core threat intelligence offering, to all Cloudflare accounts for free.

Investigate, respond and learn

The OpenAI Hugging Face incident did not begin with the final 13-hour compromise. The activity stretched from May to July, with signals including an unauthorized message board, internal network scanning, and movement across environments. Viewed separately, each event revealed only part of the activity. Together, they showed the behavior of a developing breach. Security operations must therefore identify sequences of behavior that lead to compromise, not simply evaluate alerts in isolation.

This is difficult for security teams that already protect large attack surfaces with limited resources. Alerts arrive from different tools and datasets, leaving analysts to determine which events are connected, collect the evidence, and identify whether the activity is escalating.

Cloudflare is building a platform to automate security operations. Deterministic workflows establish the customer and investigation context using trigger history, traffic baselines, enforcement outcomes, and network observations. A detection agent searches authorized datasets for anomalies and correlations. When it finds suspicious activity, specialist agents review the evidence alongside customer history and threat intelligence, helping analysts connect isolated events to broader campaigns. The system can then recommend mitigations, such as rate limiting, WAF, or DDoS protection changes, for human approval.

We are developing these capabilities with Cloudflare’s Managed Defense team, whose analysts are helping us test how evidence is collected, correlated, and turned into recommendations. We plan to make them available more broadly over time and will share more as this work progresses.

Cloudflare’s combination of reverse proxy and forward proxy services makes this correlation especially powerful. Application Security signals can reveal attempts to exploit a public-facing application, while Cloudflare One can surface subsequent activity across corporate traffic. Connecting these datasets can link an external attack with unusual access, internal scanning, or potential lateral movement, turning separate alerts into a timeline of compromise and helping analysts intervene before the breach progresses.

Looking ahead

AI is changing how software is built, how attacks unfold, and who interacts with applications. Security teams can no longer manage discovery, access, runtime protection, and response as separate activities.

Cloudflare is bringing these capabilities together in a closed-loop system powered by global intelligence, local application context, and inline enforcement. A vulnerability finding can strengthen runtime protection, runtime activity can guide an investigation, and each analyst decision can improve future detections and controls.

No organization can anticipate every new technique. The goal is to build a security system that learns from each attempt, responds faster, and becomes more effective over time. The capabilities announced today are the next step toward that adaptive model of application security.

Building a post-quantum certificate authority with Merkle Tree Certificates

Post Syndicated from Mari Galicer original https://blog.cloudflare.com/pq-ca-with-mtcs/

When you type in an address into a browser, how do you know you’re connecting to the right website? The Web Public Key Infrastructure (Web PKI) is the complex and distributed ecosystem of policies, protocols, and infrastructure operators that helps you trust that you’re not being misdirected to an incorrect or malicious website. In the past few decades, this ecosystem has undergone significant changes. One is the addition of transparency: the now-mandatory requirement that all certificates be logged in public certificate transparency logs. Now it faces another challenge: the imminent arrival of a quantum computer, which has prompted us to upgrade to post-quantum (PQ) cryptography by 2029.

This transition is not straightforward: simply swapping post-quantum cryptography into certificates at Internet scale would lead to unacceptable performance degradation. This moment calls for a new approach to the Web PKI, one that allows us to treat transparency as a first-party property rather than an add-on, and design a new system that scales post-quantum signatures efficiently.

After gaining broad support across the industry, Merkle Tree Certificates (MTCs) have emerged as the path forward. This year, after a successful experimental deployment with Chrome, Cloudflare is full steam ahead on MTCs.

Following today’s announcement that Cloudflare is becoming a certificate authority (CA), we’re excited to share that this CA will support MTC issuance, targeting early 2027 for inclusion in Chrome’s newly launched Quantum-resistant Root Store. As part of our mission to help build a better Internet, and following in Cloudflare tradition of offering the strongest available cryptography for free, we will provide standard MTC issuance at no cost. Having a CA that supports both classical certificate and MTC issuance allows us to default to the most secure authentication method available, providing a painless and performant PQ upgrade path for a large swath of the Internet.

The current trust ecosystem

To understand how MTCs are changing the game, let's start with some background on how trust works on the web today.

On the client side, browsers — in this case, “TLS clients” — maintain root programs, which specify a set of policies that CAs must follow to be trusted. On the server side, CAs are the trusted gatekeepers: they operate certificate issuance infrastructure where they validate domain ownership and attest to the binding of a domain name and a public key that shows ownership of that domain.

But how do we check that CAs are following the rules? Enter certificate transparency (CT), which makes certificate issuance publicly auditable. When a CA issues a certificate, it must also submit that certificate to at least two public logs. Cloudflare has operated the Nimbus family of CT logs since 2016, and is launching Raio, a new family of static CT logs, going forward.

While the CT ecosystem makes certificates publicly viewable, it doesn't mean they are correctly issued or safe to use. Monitoring helps with this by comparing those log records with what domain owners expected and reporting suspicious activity. Cloudflare launched Certificate Transparency Monitoring in 2019 and recently made it generally available. We also publish large-scale measurements about certificates on the Certificate Transparency page in Radar (formerly known as Merkle Town).

As organizations begin upgrading their servers to use PQ authentication, certificate transparency monitoring will take on an even more important role in detecting potential post-quantum downgrades. Domain owners who have upgraded their domains to post-quantum authentication should monitor CT logs for unexpectedly issued legacy certificates to prevent clients from falling back on a malicious downgrade path.

Part of the problem with this current system is that transparency was an add-on, causing it to run into scaling issues. Certificates are frequently logged multiple times, in different forms, across multiple logs, requiring monitors to download and process every log to avoid missing an issuance. This can be expensive — making it difficult to encourage a diverse set of log operators at Internet scale. According to our estimates, PQ signatures will balloon the amount of data that CT logs need to store by 40x. This scaling challenge, and subsequent incentive misalignment, is at the heart of the post-quantum scaling problem.

The post-quantum scaling problem

We've written extensively about the challenges of scaling post-quantum cryptography, but in short: to support server authentication at Internet scale, the WebPKI must authenticate roughly a billion TLS servers without preloading every server’s public key into every client. Traditionally, CAs addressed this problem by using certificate chains as a trust-distribution mechanism. But over time, additions like key revocation checks and certificate transparency have added more public keys and signatures — five signatures and two keys in a typical TLS handshake. PQ signatures are roughly 40 times larger than classical ones, creating larger overheads that would be expensive for clients, CAs, logs, and monitors to handle at scale.

Enter Merkle Tree Certificates (MTCs), a draft specification from the IETF PLANTS working group that describes an architecture for compact, efficient, post-quantum certificates. MTCs batch certificates into an append-only Merkle tree, allowing a CA to sign the root of that tree instead of many individual certificates. This allows browsers or other clients to verify a certificate using a compact inclusion proof — a sequence of cryptographic hashes — against a signed tree head rather than validating each certificate individually. A key idea behind MTCs is "don't log what you issue, issue by logging." By coupling issuance and logging, transparency becomes a requirement for operation, rather than an add-on.

The role of a certificate authority in a redesigned PKI

We’re building out our capability to issue MTCs as an integral part of our creation of a Cloudflare CA. That means keeping track of new PQ Root Program requirements, and writing an issuance and mirroring software stack at the same time we’re building the facilities, operations, and compliance functions of the traditional  CA — no small feat!

The upside is that we get to prioritize the requirements and architecture for this new, post-quantum PKI from day one, building our setup in a way that feels right for Cloudflare's values and global network — aiming to be as transparent as possible as we embark on this new journey.

Let’s take a look at the architecture updated for MTC:

If you compare this to the traditional CA ecosystem, you'll notice that the responsibilities of a CA stay mostly the same: to validate control of a domain, bind it to a public key, and issue certificates. The main difference is that in the MTC ecosystem, instead of signing certificates directly and then logging them, the CA now maintains a transparency log backed by a Merkle tree, where an inclusion proof that the certificate is indeed in the tree serves as the trust anchor. CAs will also operate Mirroring cosigners that store a copy of issuance logs, verifying their append-only consistency and ensuring the transparency and availability of these logs for the broader ecosystem.  

Issuing MTCs

MTCs come in two forms, both of which can be encoded in the X.509 certificate format that client software recognizes today — just with a “funny” signature algorithm. In standalone form, the certificate’s signature value contains a cosigned tree head of an issuance log and an inclusion proof (a sequence of hashes) demonstrating that the certificate is contained in that log. If clients are able to obtain the cosigned tree heads out of band (e.g., via a browser update mechanism), the certificate can instead be served in landmark-relative form, where the signature value consists of the lightweight inclusion proof with no heavyweight post-quantum signatures at all.

For simplicity’s sake, let’s take a look at an example of standalone certificate issuance. When a website wants a certificate for their domain, they can request it from a CA via the Automatic Certificate Management Environment (ACME) protocol, which handles certificate requests, domain-control validation, and issuance workflows. Cloudflare's ACME infrastructure will be a fork of Boulder, the widely deployed and well-tested ACME software that powers Let's Encrypt. Let's Encrypt is actively developing MTC support in Boulder, and we plan to maintain our own fork that incorporates these upstream changes along with Cloudflare-specific modifications, contributing back upstream where possible.

When the MTC CA receives a certificate issuance request, the CA's ACME server checks that the server actually controls the domain. If those checks pass, the CA serializes that data and adds it to an append-only log.

After adding the MTC entry into its issuance log, the CA computes the updated state of the log, and then signs a checkpoint over that state. This checkpoint attests that the CA issued every entry included in the log’s Merkle tree up until that point in time.

The CA then sends its updated log state and new checkpoint to a trusted cosigner, which durably stores a copy of the CA's issuance log and checks that each new state is append-only, consistent with the previous tree, and correctly formed. This additional cosignature gives clients and monitors confidence that another trusted party has observed the same log state and verified that the CA is not presenting different views of issuance to different parts of the ecosystem. It also ensures that the issued certificates will be available for monitoring even if the CA issuance log is unavailable.

Chrome’s Quantum-resistant Root Program draft policy mandates at least two cosignatures: one from a Chrome-recognized Mirroring Cosigner operated by a distinct organization, and one from the issuing MTC CA itself. As such, we'll operate mirrors for other pilot CAs — and require at least one independent cosignature on our own issued certificates.

Cloudflare will implement our mirroring cosigner in Azul, our open-source Rust-based transparency log, and for maximal interoperability, it will implement c2sp's tlog mirror protocol.

Finally, after successfully receiving a cosignature from a mirroring cosigner, the CA constructs an MTC with the cosignatures, server's public key, and an inclusion proof. It then sends that MTC to the server, which can then use it for TLS moving forward!

Delivering PQ signatures efficiently: the landmark optimization

While standalone certificates are functional, they still send large PQ signatures over the TLS handshake, limiting their efficiency. The real performance improvements provided by the MTC design are landmark-relative certificates.

Instead of sending cosignatures in every certificate, CAs can designate a sequence of subtrees that cover all active certificates in the log as a landmark, and distribute those subtrees (along with data to authenticate them) to clients via an out-of-band update service. During a TLS handshake, the actual authentication to the server happens by the browser checking that the server's certificate data — including its domain name and public key — appears in a trusted subtree of the CA’s log. If the inclusion proof connects that certificate to a cosigned landmark, and the public key then proves possession during the TLS handshake, the client knows it is talking to the right server.

Periodically transmitting these signatures and tree metadata to TLS clients out of band, a small set of MTC batch signatures can efficiently cover billions of certificates issued by a given CA. While landmarks are more efficient at scale, they do not eliminate the need for standalone MTCs — clients may be newly installed, offline, or missing the relevant landmark update. That’s why it’s important that servers retain a standalone certificate fallback.

MTCs in the wild: results of our experiment with Chrome

This year, we ran an experiment with Chrome to test the feasibility of MTCs between a client and server. We operated a "bootstrap CA" (a fake CA that stubbed the issuance pipeline) that issued MTCs backed by a traditional certificate chain for a selection of Cloudflare domains on Cloudflare's "free" plan and served them to 50% of Chrome Beta 146. Over the course of the experiment we successfully served billions of MTCs.

For TLS, we found that the common case is fairly efficient: with a landmark-relative certificate, the handshake only needs to transmit one public key, one signature, and one inclusion proof of less than 1kB. In the experiment, we fell back to the traditional certificate chain instead of serving a standalone certificate in cases where we were unable to negotiate a landmark-relative certificate with the client. On the CT side, MTCs also change the scaling properties of transparency: the log only needs to carry hashes of public keys; there are no per-entry signatures, and the signature on the tree head covers the whole log. This prevents certificate explosion because the CA issuance log is the source of truth for all certificates the CA issues, and log consumers only need to fetch a single copy of each certificate.

The result: MTCs really work! At median, using a MTC is 9% faster using landmark MTCs over a classical signature chain (admittedly, most of this performance benefit is due to intermediate elision). And because we tested MTCs with classical signatures, we expect an even greater improvement with post-quantum signatures. Satisfied with these results, and with the level of cross-industry collaboration with MTCs at the PLANTS WG at the IETF, we began winding down the experiment last month (August 2026).

The road ahead for MTCs

We’re excited that our experiment with Chrome showed that MTCs can work in practice, and are especially excited to be able to issue certificates as a real CA.

However, there are still broader questions that we can only answer by running this great experiment with the full PKI ecosystem. Can independent monitors consume and verify MTC issuance logs at production volume? Will multiple CAs and cosigners emerge so that the system has the diversity needed for resilience? How should browsers balance the performance benefits of compact landmark MTCs with the fallback paths needed for clients without fresh landmarks? MTCs have emerged as the authoritative design for post-quantum authentication, but proving it out at production Internet scale will require participation from a diverse set of root programs, browser vendors, CAs, mirrors, monitors, and the wider community.

We see the opportunity to participate in this next phase of the Web PKI as an honor, and we take the responsibility of operating CA infrastructure seriously. CAs occupy a privileged position in the trust ecosystem — browsers, domain owners, and everyday people rely on them to validate identities correctly, protect signing keys, follow policy, and operate reliably. Before Cloudflare's CA can be trusted by browsers to issue MTCs, we will need to apply to Chrome's Quantum Resistant root store and undergo a rigorous evaluation process. We welcome that scrutiny, and we expect to hold ourselves to the same high bar as any other CA trusted with helping secure the Internet. We hope other CAs will emerge to support MTC adoption, and we're excited to work with any browser that wants to deploy MTCs.

We tested our own WAF with frontier AI models. Here’s what we found

Post Syndicated from Vikram Grover original https://blog.cloudflare.com/adaptive-ai-waf-testing/

“Is your WAF ready for frontier AI models?” We keep hearing this question from our customers, so we decided to find out.

When it comes to exploiting applications, what LLMs are really good at is iterating and mutating attack payloads faster than any human hacker could do. LLMs can use real-time responses to iterate and change their techniques by, for example, testing different encodings, sending the payload in a different part of the HTTP request, or moving to the next vulnerability to test.

Even before LLMs were around, security engineers used two common approaches to test applications: static and dynamic application security testing. The former analyzes code without executing it to identify vulnerabilities, while the latter probes running applications to find runtime flaws. There are plenty of works scanning code with frontier AI models, including details on how to build your own harness.

For the project described in this blog post, we took a dynamic approach: making the LLM act as if it was a hacker to evaluate whether a WAF is doing its job. The LLM had no visibility into source code, no view of the WAF's rules, and could only see selected HTTP response data.

We built a WAF tester that starts from known exploits and then iterates by changing how it is encoded or delivered, sends it again, and uses the response to choose the next variation. A request that was not blocked became a lead for human review, not a confirmed exploit.

We ran the tester against an authorized customer staging environment across six attack categories and recorded 1,107 attempts. After reviewing the non-blocked requests and removing malformed, benign, duplicate, and out-of-scope observations, the vast majority of the attacks were blocked by the Cloudflare WAF. The requests that got through helped us create new detections to harden our security to benefit all Cloudflare customers.

Here we will explain how we set up the system, the types of attacks we tested, which attack vectors bypassed the WAF more easily, and how we fixed it. Most importantly, we share what we learned from this process and how this exercise is becoming a foundational building block of our WAF development lifecycle.

Finally, we offer guidance to help you correctly deploy your WAF in front of your application and, most importantly, patch your software. A payload that bypasses the WAF still needs an exploitable application to succeed, so keeping your stack up-to-date remains one of the strongest defenses against attackers.

How the adaptive loop works

To test our WAF with frontier models, we built a system that iterates over multiple scenarios. A scenario means choosing one attack category, placing the input in a specific part of the request, starting with a version the WAF already blocked, and giving the tester a fixed number of attempts to try other variations. The loop runs LLM models twice: the first is the proposal call, the second is the review call.

The first call receives the starting request, the context, a short history of earlier results, and suggests the next variation, then the code builds and sends the request. The review call receives the request context, response status, selected headers, and the response body. The loop stops when mutations stop producing useful variations or when a hard coded attempt limit has been reached.

Both model calls work without access to WAF internal information. Neither receives rule expressions, rule IDs, WAF Attack Score details, or the identity of the security layer that acted. We implemented the system in Python rather than wrapping an existing penetration-testing tool. It handles HTTP replay, scenario orchestration, state tracking, and result collection.

In the current implementation, the models do not send requests directly — code controls what happens at each step. Before each request, it checks the target hostname against an allowlist, disables redirects, records the attempt, and enforces the attempt limit. After each request, it records the response and uses the model's review to choose the next predefined step. Response text may appear in a later prompt, so the tester treats it as untrusted input. Neither model call can deploy a rule nor change enforcement.

The system records structured evidence for each attempt.

Six attack categories against one WAF configuration

The main run targeted an authorized customer staging environment protected by Cloudflare’s WAF. We used an allowlisted test User-Agent so the customer’s automated-traffic controls would not stop the test before requests reached the WAF.

We ran 45 scenarios. For each, we looked for ways to deliver the same attack differently: different encoding, different part of the request, or the same destination written another way. Of these, 44 covered six attack categories: cross-site scripting (XSS), SQL injection (SQLi), command injection (CMDi), server-side request forgery (SSRF), path traversal or local file inclusion (LFI), and Log4j. The remaining scenario covered log injection, reported separately.

The WAF in the test zone was configured as follows: WAF Attack Score blocking scores of 30 or below, all Cloudflare Managed Ruleset enabled, and OWASP Core Ruleset with Paranoia Level 3.

For the headline measurement, we recorded whether the WAF blocked each request or not. The results describe the configured WAF boundary as a whole, not the performance of any individual rule or detection mechanism.

What adaptation looked like in one recorded session

Here is an example of how the LLM adapts a Server-Side Request Forgery (SSRF) attack during the test.

Cloud metadata services can expose temporary credentials to workloads. An SSRF vulnerability can let an application fetch that data on an attacker's behalf. A WAF can help stop the malicious request before it reaches the application, but it is only one layer of protection.

In this SSRF scenario, the tester sent the same cloud metadata address in different forms (such as integer, octal, and trailing-dot representations of the same IP) and placed it in different parts of the request. The WAF blocked all of them except one. At attempt 18, the model kept the same request structure as the previous blocked attempt and switched to the trailing-dot form. The client encountered a redirect rather than a WAF block.

The table below shows selected moments from the session. The hypothesis column summarizes what the model said it was trying before each move. It is not a verbatim transcript, and it is not proof that the explanation was correct.

Attempts 17 and 18 are an interesting pair: same request structure, different host representation. One was blocked, one was not. That gave us a specific question: does the trailing dot change how the WAF reads the destination? It was a lead to investigate, but not proof that metadata was accessed.

This was one selected trajectory among 45 scenarios. The next section shows how we counted and triaged the full run.

What we found

Our tester generated 1,107 attempts and the overall result was strong with XSS, LFI, SQLi, and Log4j having near full coverage. While the run produced useful findings, it also produced noise. After human review, we were left with 49 findings worth investigating, 48 of them belonging to CMDi and SSRF. 

Here is how they break down:

Metric

Value

What it means

Recorded mutation attempts

1,107

Model iterations across 45 active scenarios; not all produced a usable result

Post-triage result set

607

The 558 blocked requests plus 49 documented WAF-relevant findings

Blocked requests

558

The WAF stopped these before they reached the application

WAF-relevant findings

49

Documented for remediation analysis after human review

The rest did not produce a result worth counting as the model failed to generate a usable HTTP request, some failed before reaching the target, or the payload generated was benign.

When a request was not blocked, we worked through five questions before counting it as a finding:

Question

Why it matters

Did the tester actually send a valid request?

If the model failed or the request never reached the target, the result tells us nothing about the WAF.

Was the request clearly not blocked?

An ambiguous response is not enough to count.

Was the request still malicious?

Changing a request to get it past the WAF can also make it harmless.

Did the behavior belong to the WAF?

Some attacks only work through DNS or network paths the WAF cannot stop at request time.

Could engineers reproduce it safely?

A fix needs a stable test case with a clear expected result.

We removed anything that failed those checks and combined duplicate cases. What remained became the input for rule, normalization, and mitigation work.

Findings became detections

Not every finding needed a new rule. Some pointed to gaps in existing Managed Rules coverage. Others pointed to how the WAF normalized the request or belonged to another security control. We replayed each case and decided where the change should happen.

We grouped related findings into four sets of candidate rules, validated each finding, and tested candidates against live traffic before any rule could protect customer traffic.

Before a new or updated rule can protect customer traffic, we check its impact on legitimate traffic and assess false-positive risk. Some of the issues we find when evaluating a new rule candidate include:

Issue

Next step

Missing or narrow detection

Review whether existing rules cover the finding

Equivalent inputs interpreted differently

Engine or normalization review

False-positive risk is too high

Revise or reject the candidate

This work contributed to three changes in Cloudflare's Managed Ruleset: new detections for SSRF – Obfuscated Host and SSRF – Restricted Protocol in the July 21 release, and improvement of the existing SSRF – Cloud rule. The SSRF – Obfuscated Host detection came directly from requests that encoded internal addresses in non-standard numeric forms.

What we learned

The model was only one part of the test. We ran the same scenarios with two versions of the same model family. They produced different variations – and the same underlying issues appeared in both. Because request replay and evidence capture stayed consistent, we could compare the runs without treating either model's output as ground truth.

More attempts within one scenario did not always find more. Some scenarios started repeating earlier ideas near the end of the 25-attempt limit. We got broader coverage by testing more starting requests, attack categories, and input locations instead of extending one sequence.

The model generated requests. We decided which ones mattered. A request that was not blocked still needed replay and human review before it could become a finding, a mitigation, or a regression test. Without that review, there were no findings.

What customers can do now

WAF is just one layer of detections you can deploy. When you deploy all available protections you increase the effectiveness of your overall stack. 

First of all, check that Managed Rules, WAF Attack Score are set up correctly in front of your application. Other tools you can deploy include API Security, Bots and Fraud detection, and Threat Intelligence to strengthen your posture even further. For example, positive security controls add a different layer: instead of looking only for known attack patterns, they define the request shapes an application expects and identify inputs outside that contract. This drastically reduces your attack surface area. 

Customers do not need to reproduce this experiment. To maximize the number of rules deployed in front of your application, we recommend running Managed Rules in log first, review matching requests in Security Events, and confirm legitimate traffic is unaffected before moving a rule to Block. Alternatively, customers can reach out to their account team to get Attack Signature Detection turned on, on their zones. This new feature simplifies how to review matched traffic and how to deploy signature detections. If you already perform application security testing, run those tests against a staging hostname protected by the same Cloudflare controls as production.

Next steps

By combining adaptive AI-driven testing with human triage and validation, we found detection gaps that fixed tests might miss and turned those findings into stronger WAF protections, improving our block rate. In a future post, we will share results from further testing using a white-box approach, where the model knows both the application’s vulnerabilities and the WAF rules protecting it.

Is your domain using post-quantum encryption? Now you can see for yourself

Post Syndicated from Andrew Depke original https://blog.cloudflare.com/post-quantum-visibility/

Today, we are introducing additional post-quantum (PQ) cryptography visibility tools into Cloudflare's Application Security and Logs products. You can now inspect and graph the adoption of post-quantum TLS 1.3 encryption for live traffic from directly within Logpush, Log Explorer, and the HTTP Traffic Analytics dashboard. By surfacing the key exchange algorithm negotiated on every incoming request from visitors to our platform, Cloudflare gives customers granular, per-connection telemetry to audit their post-quantum posture, assess compliance, and identify cryptographic gaps across their domains.

Cloudflare is targeting 2029 for full post-quantum security, and executing a cryptographic transition at scale requires detailed telemetry. We’ve already deployed post-quantum encryption across many of our products, including in our cloud-proxy platform and on every on-ramp and off-ramp of our SASE platform.   As many of our customers work towards quantum-readiness deadlines around 2030, we’re helping ease the transition by making post-quantum encryption the default in many of our products, sharing learnings from our internal cryptography discovery tool, and launching the new post-quantum visibility features for TLS that we’ll cover in this blog.

Bringing post-quantum visibility to the domain level

When it comes to post-quantum visibility, we already have macro-level visibility into Internet-wide post-quantum adoption in TLS through Cloudflare Radar. On Radar, we track global post-quantum encryption statistics, both when Cloudflare proxies HTTP requests from visitors (the visitor-to-Cloudflare connection) and when Cloudflare connects to origin servers (the Cloudflare-to-origin connections), as shown in this figure.

From Radar we can see that about 70% of browser-generated traffic hitting Cloudflare's network (on the visitor-to-Cloudflare connection) is protected with post-quantum encryption using hybrid ML-KEM (FIPS 203).  Meanwhile, we can see that today, just about 15% of origins that Cloudflare connects to use hybrid ML-KEM. These are aggregate numbers; the first number is aggregated across all the browser-generated traffic we see, and the second number is aggregated across all the origins we connect to.

We’ve also recently launched Automatic Key Exchange for the Cloudflare-to-origin connection, which reveals which cryptographic algorithms are supported by a given origin. This is useful because outdated configurations can cause an origin to connect to Cloudflare using classical cryptography, even if it does support a post-quantum encryption. 

While Radar and Automatic Key Exchange both provide valuable macro-level views of Internet-wide readiness, our customers have asked us to be able to go beyond aggregate numbers and dive into the behavior of individual domains.

We have long provided visibility into the TLS version used at individual domains (TLS 1.3, TLS 1.2, etc.).

But until now we have not exposed information about the cryptographic algorithms used with the TLS version used at the domain level. This means customers could not answer questions like “What fraction of traffic to my domain www.example.com is using post-quantum encryption?” This information is helpful when aiming to comply with regulatory frameworks, troubleshooting a migration to post-quantum encryption, or seeking to understand which fraction of traffic that is exposed to future quantum adversaries. Now, these questions can be answered.

Post-quantum cryptography in TLS

Before we get into the new product features, let’s do a quick review of post-quantum cryptography in TLS, so we can understand the information that the feature surfaces.

In 2024, the National Institute of Standards and Technology (NIST) stated that RSA and Elliptic Curve Cryptography (ECC) should be deprecated by 2030, and many governments and regulators have since gotten behind that deadline. That’s why today, many of our products are protected with post-quantum encryption using a cryptographic key agreement algorithm called hybrid ML-KEM. Post-quantum encryption is needed right now to stop harvest-now-decrypt-later attacks, where an adversary harvests data today and then decrypts it in the future once powerful quantum computers come online. Organizations that have data that are valuable even if decrypted in 3–10 years (public sector, defense, finance, telecom, healthcare, and others), should consider immediately protecting their traffic with post-quantum encryption.  

 In TLS 1.3, the key exchange group X25519MLKEM768 is the only recommended algorithm for post-quantum encryption. It is now the algorithm preferred by most major browsers. (Note: post-quantum encryption is not available in TLS 1.2 or any earlier version of TLS.)   If you are using Chrome, you can check the key agreement algorithm used by this webpage (or any other) by right-clicking “Inspect”, going to the “Security” tab and looking for the below:

With X25519MLKEM768 in TLS 1.3, the client and server execute both:

  • the Elliptic Curve Diffie-Hellman Key Exchange (ECDHE) over curve X25519 and
  • the post-quantum Module Lattice Key Encapsulation Mechanism (ML-KEM)

X25519 and MLKEM768 each produce a shared secret. TLS then combines those two secrets and uses the result to encrypt TLS traffic. This hybrid approach provides belt-and-suspenders security; as long as one of the two key exchanges is secure, the resulting shared secret is also secure. TLS 1.3 also supports other key exchange groups, including X25519, P-256 and P-384, all of which are just classical ECDHE over different elliptic curves; these algorithms are still used all over the web. In earlier versions of TLS you can also find key agreement based on the RSA algorithm, which is quantum-vulnerable and thankfully much less popular these days due to many known classical security problems.

But post-quantum encryption is only the first part of the story; the second part is post-quantum authentication. Once powerful quantum computers exist, we need to worry about upgrading the certificates and signatures used in TLS 1.3 away from RSA and ECC and towards post-quantum algorithms like ML-DSA. We’re actively making progress towards that goal. In fact, we recently announced that origins can use ML-DSA-44 certificates over TLS 1.3 to connect to Cloudflare, and today we announced that we’re launching a certificate authority that will support post-quantum Merkle Tree Certificates. Nevertheless, for now it remains true that post-quantum encryption with hybrid MLKEM is more broadly deployed than post-quantum authentication.

Bringing post-quantum visibility to the visitor-to-Cloudflare connection

Today we’re making it possible to see the extent to which post-quantum key agreement is used on the visitor-to-Cloudflare connection for any domain in HTTP Traffic Analytics dashboard, Logpush, and Log Explorer.

To view the TLS key exchange data on your domains, go to the Cloudflare Dashboard, and navigate to HTTP Traffic under the Analytics tab. Here you’ll get in-depth statistics about the kinds of traffic visiting your domains, now including a dedicated card for TLS Key Exchange groups on the visitor-to-Cloudflare connection. (Scroll down to find it!) Here’s a look at a TLS Key Exchange card for one of our test domains:

As you can see, the majority of the traffic to this domain uses post-quantum X25519MLKEM768 (in TLS 1.3).  We see some traffic using classical ECDHE over curve X25519 or P-256 (in TLS 1.3 or below).  The traffic labeled “None” is using either RSA key agreement (in TLS 1.2 or below) or no TLS at all. And finally we have a small number of visitors using the now-deprecated X25519Kyber768Draft00 algorithm with TLS 1.3, which we implemented back before X25519MLKEM768 was fully standardized by the Internet Engineering Task Force (IETF). We’ve waited to remove support for X25519Kyber768Draft00 until observed connections are diminishingly small, to avoid regressing clients for which this is their only way to support PQ encryption.

While we’re here, we’ll just drop a few tips about PQ-ing your traffic. If you look at your domain and find no use of X25519MLKEM768 at all, you should confirm that TLS 1.3 is enabled. In the Cloudflare dashboard, select your domain, go to SSL/TLS > Edge Certificates, and then scroll until you find the TLS 1.3 switch; switch TLS 1.3 to On. (There is no separate post-quantum setting: when TLS 1.3 is enabled and a visitor supports X25519MLKEM768, Cloudflare negotiates it automatically.) Also, if the vast majority of your traffic is over classical X25519, P-256, P-384, or None, it might be because most visitors to that domain are non-browser clients that lack support for X25519MLKEM768 and/or TLS 1.3. (Again, most major browsers do prefer to negotiate a TLS 1.3 connection with X25519MLKEM768.)

The key exchange group can now also be a filtering term in the HTTP Traffic dash. Here’s how to take a look at the traffic that is not using post-quantum encryption with X25519MLKEM768:

Analytics are great for aggregate investigations, but being able to see this information in individual log lines can be even more powerful. You can enable the new ClientTLSKeyExchangeGroup field, under the TLS category in the HTTP Requests dataset, to gain visibility into individual post-quantum key exchange in your Log Explorer and Logpush connection logs.

With this new field enabled, you’ll see it start appearing in your Logpush HTTP Request logs, like so:

Visibility to origins and more

The release of the key exchange group stats represents the first major milestone in our broader cryptographic visibility initiative. Designed for scalability, our underlying telemetry pipeline is built to ingest additional cryptographic parameters from TLS handshakes.

That’s why we’ve also surfaced the key exchange group from the Cloudflare-to-origin connection and to provide end-to-end visibility from eyeball to origin in Logpush as OriginTLSKeyExchangeGroup. (This group will be the same for all visitor connections made to that domain, which is why it's not shown in the HTTP Traffic Analytics dashboard).

And for customers that use legacy origin servers that are unlikely to support modern post-quantum cryptography, don’t despair. You can put the origin server behind a Cloudflare Tunnel, to tunnel traffic from the origin server to Cloudflare over TLS 1.3 with X25519MLKEM768, without need to upgrade the legacy origin server itself. This is what the network configuration would look like if you put your origin server behind a Cloudflare Tunnel:

Eventually we’ll be able to also surface post-quantum authentication (namely the algorithm used for certificates and signatures in TLS, including Merkle Tree Certificates) once we start to see a broader-based deployment of that technology.

Your domain has started its post-quantum journey

If your domain is behind Cloudflare, its post-quantum journey is already underway. Check HTTP Traffic Analytics dash and your logs to see the percentage of visitor connections to your domain that already use TLS 1.3 with post-quantum encryption (X25519MLKEM768).  You can also check logs to see if you’re using post-quantum encryption on the Cloudflare-to-origin connection. If your origin server is too ossified to support post-quantum cryptography, then just put it behind Cloudflare Tunnel. With the right settings and visibility, you can protect more of your traffic on Cloudflare from harvest-now-decrypt-later attacks today.

We thank Luke Valenta, Ollie Hsieh and Alex Krivit for contributions to this work.

Introducing Threat Signals: agentic skills for open-source threat intelligence, free for every Cloudflare account

Post Syndicated from Emilia Yoffie original https://blog.cloudflare.com/threat-signals/

Organizations can now scale threat intelligence expertise the way they scale infrastructure. Threat intelligence analysts and network defenders have long automated the ingestion of structured threat feeds to help enrich their SIEM or WAF. The harder work has always been unstructured reporting: turning a research post into indicators your tools can use, without losing the context that explains why they matter. AI skills make that work possible to automate. A skill is a set of rich, detailed instructions that captures how an experienced analyst handles one part of the job, and it runs the same way on every report. 

Threat Signals puts that process into practice at scale. It’s launching today, and we made it available to every Cloudflare account. 

Threat Signals turns open-source reporting that you choose into intelligence you can act on. Its agentic skills summarize reports, surface key context, extract and normalize indicators of compromise, and apply tags — all within a private, account-scoped dataset. The end result is a contextualized indicator stored in your account’s private Threat Intelligence dataset as a Threat Event that can instantly be applied in your WAF policy.

Starting today, we are also expanding access to Cloudforce One’s Threat Events Platform, our core threat intelligence offering, to all Cloudflare accounts for free. With this expansion, each account gets:

  • API and dashboard access to Threat Signals and the ability to select one RSS feed
  • A private dataset built from the RSS feed in Threat Signals, tailored to your reporting requirements and stored for up to 30 days
  • API and dashboard access to Threat Events Platform to investigate events, indicators, and tags related to your private dataset

Essentials, Advantage, and Elite enterprise customers can extend this offering to include an expanded number of RSS feeds, access to Cloudforce One’s proprietary threat intelligence datasets, the ability to generate custom agentic skills, higher storage options for Threat Signals’ derived open-source reporting, and the ability to create custom WAF rules on open-source and proprietary threat events.

Discovery is only the beginning

We started with open-source intelligence because it is the most obvious place to prove the power of agentic workflows. We also heard from customers that their existing platforms cannot scale beyond polling 100 RSS feeds. Recognizing the critical impact open-source reporting plays in understanding the threat landscape, we sought to build an infinitely scalable platform (more on that later).

Researchers regularly publish detailed findings on vulnerabilities, malicious infrastructure, phishing campaigns, malware families, and threat actors. While RSS feed readers make it easier to discover new reporting, discovery is only the beginning. Harnessing data into a usable workflow with consistent expertise is the key to building actionable defense.

Expertise has never been something organizations can replicate at scale. A report explains how a campaign works and identifies the infrastructure behind it, but before an analyst can use that information, they need to:

  • Read and summarize the report
  • Identify relevant indicators
  • Convert indicator values into a consistent format
  • Classify the report using an internal taxonomy for tagging
  • Populate the indicators into a threat intelligence platform (TIP)
  • Preserve a link to the original source
  • Share the intelligence with the rest of the security team

Repeating that process across dozens of sources takes time; moreover, almost every step is entirely about human judgment. As a result, context is lost. Indicators inserted into your TIP are separated from the context that explains why they matter and helps assess the risk later in the remediation cycle. It's not surprising that weeks later, a domain is pushed to a blocklist and nobody understands why. 

How Threat Signals works

Threat Signals uses RSS to monitor the open-source reporting that matters to your organization. You can add an RSS feed, give it a recognizable name and category, and configure how frequently Threat Signals checks for new content. All three feed specifications (RSS 2.0, Atom, and RSS 1.0/RDF) are supported.

Each feed you select enters a Workflow that periodically polls for new articles. It uses Browser Run’s Markdown quick action to fetch and clean the article text into a readable markdown format, which is then stored in R2. The text is passed into an indicator of compromise extractor and a set of default Cloudforce One-defined skills to summarize the content, apply tags based on your account configuration, and add indicator contextualization at the IOC level.

The output is a concise summary and key points that help an analyst quickly understand what happened, who was affected, and why the report matters. All of it is searchable and tagged, so you can find the articles you care about across the platform.

Lastly, each indicator extracted is backed by a threat event within the account's own private Threat Signals dataset. The event, its indicators and tags, and the original report stay connected, so an analyst can always trace where the intelligence came from and why it is there. These indicators can then be used to create WAF rules from threat events to protect your applications and infrastructure.

What we learned

It’s not hard to write a script that pulls an RSS feed and regexes IP addresses out of it. The first version of Threat Signals was a one-week internal prototype, built by a threat analyst who wanted more out of the reports she was already reading. Turning that into something every account can rely on was harder, and most of what slowed us down had nothing to do with parsing. The hard work was in making the output something analysts would trust and actually use. 

We were tempted to let the system invent whatever tags seemed useful. The teams we talked to pushed back: intelligence labeled in an unfamiliar vocabulary is harder to use, because now there are two vocabularies to reconcile. So we limited AI tagging to each account's existing tag catalog. 

Recording whether a tag was applied automatically or by an analyst sounds like a minor piece of metadata, but it turned out to be essential. In our experience, analysts were far more willing to trust automatic tagging when they could see exactly which tags it applied.

Summaries are useful, and they are what users notice first. But what analysts kept returning to in early testing was the link between an event and the report it came from. As investigations progressed, we discovered that link consistently helped them keep track of indicators and understand why each one mattered in the first place. 

What’s next

Open-source reporting isn’t limited to RSS feeds. Analysts need to be able to quickly consume threat intelligence in various formats and pipelines. Now that we’ve laid out the building blocks for ingesting indicators from data feeds into our platform, the natural next step is to add more consumers. Be on the lookout for more data ingestion pipelines that we will support so that you can bring more actionable intelligence onto the platform to protect your organization.

Open the Cloudflare dashboard and set up your feed today

The best investigations begin with trusted context, and Threat Signals helps keep that context close from the first lead onward. Threat Signals is now generally available for every Cloudflare account via API and the dashboard. Open the Cloudflare dashboard, navigate to Application Security → Threat Intelligence → Threat Signals, and add your RSS feed. The documentation is here. 

You can also read threat intelligence research from our team, and talk to your account team about putting Threat Events to work in your enterprise environment.

Preventing quantum downgrade attacks against IPsec

Post Syndicated from Christopher Patton original https://blog.cloudflare.com/ipsec-downgrade-protection/

For Birthday Week, Cloudflare is helping one of the Internet’s core security protocols develop stronger protections against quantum downgrade attacks. To protect our customers and the Internet at large, we worked with the IETF to develop a mitigation against downgrade attacks on IPsec, which we’ve implemented and made available in beta across our IPsec products.

The world is racing to build the first generation of quantum computers. These new machines hold great promise, but they also create a new threat: early quantum computers will be capable of cracking cryptography we've relied on for secure communication. To address this, it is necessary to migrate to post-quantum (PQ) cryptography: cryptography we believe even quantum computers cannot break. Diffie-Hellman key agreement will have to be replaced by PQ key agreement mechanisms such as ML-KEM; classical signature schemes, like ECDSA and RSA, will have to be replaced by PQ schemes such as ML-DSA; and so on.

The PQ migration is well underway, and we’re helping the migration along by making post-quantum encryption the default in our products, open-sourcing part of our internal cryptography discovery tool, launching new post-quantum visibility features, and leading the way in the web’s migration to post-quantum certificates. Still, it will take years before all clients and servers on the Internet have been upgraded to post-quantum cryptography. In the meantime, it will be necessary for modern devices to maintain support for classical cryptography in order to connect with today’s endpoints.

The need for backwards compatibility creates its own risk. In a downgrade attack, an on-path attacker between a client and server tricks the endpoints into using weaker crypto than they support. It does so by manipulating the messages sent between client and server, making it appear to one party that its peer does not support PQ at all. In other words, a downgrade attack eliminates the protection provided by PQ cryptography by downgrading the victims back to classical, so it can be attacked by a quantum computer.

What this means is that merely adding support for the cryptographic primitives themselves is not sufficient to head off the quantum threat. The next frontier in the PQ migration is to prevent active attackers from bypassing PQ by downgrading the connection.

In this post, we focus on the IPsec protocol, a central component of a variety of Cloudflare products, namely Cloudflare IPsec, Cloudflare WAN, and Magic Transit. Like all secure channel protocols, including TLS, IPsec is vulnerable to the following simple downgrade attack as long as both classical and post-quantum authentication are supported. An attacker can impersonate a party by cracking its classical credentials and can pretend the party doesn't support PQ. However, several months ago, we discovered — or rather rediscovered, as we'll explain — a design flaw in IPsec that admits a more sophisticated attack that works regardless of which authentication method is used.

The vulnerability allows a quantum attacker to decrypt all traffic between PQ-capable endpoints. The attack is relatively hard to pull off, as it requires a quantum computation to be carried out in real time during the protocol handshake. (This is different from a harvest-now, decrypt-later attack, where the quantum computation is entirely offline.) We don't yet know if and when this attack will be feasible, but recent trends give us ample reason to be cautious: at the time of writing, resource estimates for quantum attacks on public key cryptography have decreased dramatically, leading Cloudflare to move up our transition deadline to 2029.

To inoculate IPsec to this threat, we helped the IETF develop an extension that adds a downgrade protection mechanism to IPsec. Both parties must support this extension for it to be effective: for our part, Cloudflare has rolled out beta support in Cloudflare WAN and Magic Transit, which customers can now enable by requesting the account managers to turn on the ipsec_downgrade_protection flag for their accounts. We hope to see the rest of the IPsec ecosystem follow suit in short order.

IPsec's place on the Internet

Frequent readers of the Cloudflare blog are likely already familiar with the TLS and QUIC protocols. Between them, TLS/QUIC secure virtually all the web traffic transiting the Internet today. Both operate at the transport layer of the network stack: TLS runs over TCP, while QUIC runs over UDP. 

IPsec serves a similar function, but operates at the IP layer. Because IPsec operates at an even lower layer of the network stack than TLS and QUIC, it is deeply rooted in modern network infrastructure. Cloudflare IPsec allows organizations to extend their IPsec connections over Cloudflare’s global anycast network without expensive multiprotocol label switching (MPLS) connections. IPsec is also part of Cloudflare’s Magic Transit product. With Magic Transit, Cloudflare’s global anycast network sits in front of an organization’s IP range to shield it from attacks and threats like Distributed Denial of Service (DDoS) attacks, and then hands the scrubbed traffic back to the organization via IPsec tunnels.

Despite being so deeply rooted in today's Internet infrastructure, the IPsec protocol continues to evolve. It has seen many important upgrades in the past several years, including the addition of PQ key agreement. IPsec is also on track to adopt PQ authentication on about the same timeline as TLS/QUIC. (In fact, IPsec is actually further along, depending on how it's configured. A pre-shared key is frequently used for authentication in IPsec, and this is already fully PQ!) This suggests that the IPsec ecosystem is more than capable of adapting to shifting threats.

Background on IPsec

Let's now take a peek into the protocol details that are relevant to the downgrade attack. "IPsec" refers to the mechanism used to encrypt IP packets. Before encryption can begin, the endpoints must first perform an authenticated key agreement. They do so using the IKEv2 protocol.

IKEv2 typically has two phases, called exchanges. In the initial exchange, the initiator advertises the parameters it supports and sends a Diffie-Hellman key share. The responder completes the initial exchange by telling the initiator which parameters it selected and sending its own key share.

After the initial exchange, the initiator and responder derive an encryption key from the key shares and encrypt all subsequent exchanges. The key shares are not yet authenticated, meaning each endpoint has no way of knowing where the key share came from. This is accomplished in the authentication exchange, in which the initiator identifies itself to its peer and sends a signature of its key share and advertised parameters. The responder uses the identity to resolve the initiator's credentials and verifies the signature before accepting the new connection. The responder does the same in the authentication message it sends in reply.

One crucial detail to point out here: each party only signs its outbound messages, rather than the entire handshake transcript, as in more modern protocols like TLS 1.3. This means the authenticating party never confirms to the relying party that they've observed the same sequence of messages. This will be crucial for the attack.

Encrypting handshake messages has two purposes. First, it hides the identity of the endpoints from the network. (TLS/QUIC don't have this feature by default, but can enable it using the Encrypted Client Hello extension.) Second, it allows the endpoints to begin using IPsec's packet fragmentation mechanism, making transmission of long messages over multiple packets more reliable. (This is especially relevant to handling large ML-KEM key exchange messages.)

This protocol relies on classical Diffie-Hellman key exchange, meaning a quantum attacker will eventually be able to derive the encryption key from the exchanged key shares. To mitigate this threat, IKEv2 includes an option to run an intermediate exchange following the initial exchange using ML-KEM as the key exchange algorithm:

Backwards compatibility. Crucially, this exchange is only performed if the initiator advertises support for it in the initial exchange and the responder agrees to use it. This allows for backwards compatibility with endpoints that don't yet support PQ. In particular, if the responder selects a classical-only key agreement, then the initiator will assume the responder doesn't support PQ and fall back to classical-only. Likewise, if the initiator doesn't advertise support for PQ key agreement, then the responder will assume the initiator doesn't support it.

Hello my name is Mallory

Let's think about how to exploit this parameter negotiation behavior. We'll start with a simple idea that doesn't quite work, and see what it takes to make it work.

Suppose there's an attacker between the endpoints — let's call them Mallory — who has a quantum computer. Mallory can make it appear to the responder that the initiator doesn't support PQ by intercepting the initiator's initial key exchange message, rewriting it to advertise classical-only, and forwarding the modified message to the responder.

This would cause the authentication exchange to fail. The initiator signs the message it sent, but the responder verifies the message it received. Since the message received is different from the message sent, verification of the signature would fail, unless the attacker also manages to forge a signature that the responder would accept.

That's not all, however: in IKEv2, the authentication messages are encrypted, which means Mallory also needs to compute the encryption key. But this is precisely what the downgrade attack enables: Mallory has already convinced the endpoints to fall back to classical-only, and they can use their quantum computer to recover the encryption key from the Diffie-Hellman key shares.

Still, there's no obvious way to forge a signature from the honest initiator, unless Mallory has compromised the initiator's authentication key. A paper from 2016 observes the following: because the responder only signs its own outbound messages, it doesn't actually confirm to its peer which initiator identity it accepted. This means the responder will accept an authentication message from any initiator it trusts, not just the initiator of the connection.

Suppose Mallory themself is an initiator whose credentials the responder will accept. In this case, Mallory can produce a valid signature using their own credentials. The responder will complete the connection, believing it's talking to Mallory, who is identified by IDm in the figure below. Meanwhile, the initiator (IDi) will complete the connection, believing it's talking to the responder (IDr):

This is a kind of identity-misbinding attack: the endpoints have both accepted an encryption key known to the attacker, but one endpoint has authenticated the wrong entity.

More variants of this attack are possible. For example, in a key-compromise impersonation attack, Mallory would just steal the initiator's credentials and impersonate the initiator directly, allowing them to eavesdrop until the responder has revoked the stolen credentials; this kind of attack does not require identity misbinding. These attacks are also not PQ-specific: Mallory can force the endpoints to use the weakest key agreement method they both support.

Does this attack actually matter?

The main difficulty with the quantum variant of this attack is that the quantum computation is online, meaning it must be carried out during the attack before the handshake completes. This is in contrast to other quantum threats to the Internet, where the computation is offline (harvest-now, decrypt-later attacks, cracking a TLS certificate, etc.). This gives us a little breathing room: downgrade attacks are unlikely to be the first target of cryptographically relevant quantum computers, given there is much, much more low-hanging fruit.

On the other hand, there's a non-negligible chance that Q-day will arrive before we've had time to disable classical-only across the IPsec ecosystem. We don't yet know precisely how long it will take to crack a Diffie-Hellman key agreement, but it's a safe bet that the capabilities of quantum computers will ramp up quickly once they arrive. It's best to get ahead of the threat while we're in the midst of other PQ upgrades for IPsec, especially given how long it takes for these upgrades to get deployed across the ecosystem.

Protecting IPsec

The simplest way to mitigate this attack is to disable classical-only key agreement (i.e., IKEv2 configurations with an initial Diffie-Hellman exchange but with no PQ key exchange following it). This is easier said than done, however: the reason parameter negotiation exists in TLS and IPsec at all is because the initiator doesn't always know the capabilities of the responder before attempting to connect (and vice versa).

In some cases, an HSTS-like mechanism is possible. With HSTS (HTTP Strict Transport Security), a client remembers which of its peers and servers have supported PQ in an earlier connection, and then rejects classical-only in all future connections to those peers. This works as long as you know who is trying to connect, i.e., when your peer identifies themselves. But in IKEv2, negotiation happens in the initial exchange; the peer doesn't identify themselves until the authentication exchange, by which time it's too late.

In any case, this solution fails to address the fundamental problem. Remember that each endpoint signs its outbound messages only, and doesn't sign the messages sent by its peer. This allows an attacker to create a "split view" of the protocol's execution: the initiator sees one sequence of messages, and the responder sees another. Downgrade attacks wouldn't be possible had the initiator and responder confirmed they had a matching conversation. In modern handshake protocols, like TLS 1.3, each authenticating party signs the entire handshake transcript, including the messages they received from the relying party. This allows the relying party to confirm it had the same conversation, thereby preventing the split view exploited by the downgrade attack. We prefer this more principled approach.

Introducing the full transcript authentication extension of IKEv2

We worked with the IPsec Maintenance (IPSECME) Working Group at IETF to develop an extension for IKEv2 (soon to be an RFC!) called IKE_SA_INIT_FULL_TRANSCRIPT_AUTH that endows the protocol with full transcript authentication. For backwards compatibility, use of this extension is negotiated just like any other feature. This means the extension itself is subject to downgrade attack, but the extension uses a clever trick to prevent this.

The extension is very simple:

  • Support for the extension is signaled by a notify message sent in the initial key exchange. The notification is sent unconditionally: the initiator always notifies; and the responder notifies even if the initiator didn't. This is different from TLS 1.3 extensions, where the server is only supposed to reply to an extension if requested by the client.
  • If the peer notifies support for the extension, then an IKEv2 endpoint opts into updated authentication logic. In particular, instead of signing only its outbound messages, it signs the entire transcript. Likewise, it expects its peer to sign the entire transcript.

The trick that prevents downgrades is unconditional notification. Let's say Mallory modifies the initial exchange by dropping the IKE_SA_INIT_FULL_TRANSCRIPT_AUTH notification from the initiator's message, but allows the responder's notification to go through. In this case, the responder falls back to the old authentication logic, but the initiator opts in to the new logic. The responder will end up signing a different byte sequence than the initiator verifies, causing the authentication exchange to fail and resulting in an AUTHENTICATION_FAILURE notification. A similar thing happens if Mallory drops the responder's notification but lets the initiator's through.

Now consider what happens if Mallory drops the notification from both messages. This would cause both parties to fall back to the old authentication logic, allowing Mallory to downgrade the connection and compute the encryption key. But to pull off the attack, Mallory would need to forge a signature not just from the initiator, but the responder as well.

When attempting identity misbinding, Mallory would need to present an identity for a different responder than the initiator wanted to connect to. It's as if the initiator attempted to connect to example.com, but got a certificate for cloudflare.com. Unless the initiator is severely misconfigured, this will cause the authentication step to fail.

If Mallory manages to compromise the credentials of both the initiator and responder, then they can indeed pull off the key compromise impersonation variant of this attack. However, in this case Mallory has much simpler attacks at their disposal. For IKE negotiations, Cloudflare simply acts as a responder. 

How to enable full transcript authentication

This feature is gated under a feature flag scoped to each customer account. Any customer interested in trying it out can request this flag to be enabled on their behalf by reaching out to their account team. 

Here’s what happens at the protocol level, for accounts that enable this flag.  The IKE_SA_INIT_FULL_TRANSCRIPT_AUTH notification will be sent during the IKE_SA_INIT response. We will enable this flag for all customer accounts after sufficient beta testing. The feature gate is created to account for the unlikely scenario that the customer's IKEv2 initiator incorrectly handles the new notification.

Looking forward

As of this writing, this feature is on its way to RFC status. Much of the credit goes to our co-author Valery Smyslov, who did much of the heavy lifting of shepherding the document. He also spotted the trick that makes the extension downgrade-resistant.

The PQ migration is full of surprises. Ideally these surprises are few and far between. The design flaw in IPsec that allows downgrade attacks has been known for some time, at least 10 years as of this writing. There are perhaps many cryptographic protocols in use today with latent bugs that have renewed relevance in the quantum era.

Cloudflare has implemented the full transcript authentication extension and made it available on an opt-in basis. We encourage customers to reach out to their account manager to implement and begin testing the extension, and the rest of the IPsec ecosystem to consider implementing it as the draft continues to advance through the IETF.

Enforce positive security with Cloudflare Application Profiles

Post Syndicated from Daniele Molteni original https://blog.cloudflare.com/application-profiles/

Today, we are launching Application Profiles, a seamless way to enforce a positive security policy. By analyzing the structure and format of HTTP requests and identifying deviations, Cloudflare can help you significantly reduce the attack surface area.

Every customer we speak to wants to know how we can protect them from attacks that use frontier AI models. This has become the number one priority for anyone working in security. Large language models (LLMs) allow even non-technical people to launch attacks with a single prompt. LLMs can generate malicious payloads, test known techniques, and probe applications autonomously by mutating their tactics based on the feedback from the application or the Web Application Firewall (WAF). 

Our tools have changed to stay a step ahead of the attackers. Managed WAF rules and machine learning-based detections remain essential for detecting techniques such as SQL injection, cross-site scripting, remote code execution, and new CVEs, including many variations of those attacks. The answer can’t simply be “patch faster”: this is not sustainable, and it doesn’t work if you haven’t completely mapped your vulnerabilities.

What if you could learn what good requests look like by analyzing your traffic structure? Instead of looking only for requests that resemble known attacks, we could allow only requests that conform with what we expect. By doing this, we’d dramatically reduce the attack surface area. For example, if the search field in your query doesn’t expect special characters, we can only accept alphanumeric strings. This would already prevent a vast library of known attacks.

But we don’t stop here. Once we have learned the structure and format of your HTTP requests, we can infer the goal of each operation and then understand what the application ultimately does. With this information, we can identify and prioritize the most critical and vulnerable operations and fields you should take care of first.

Cloudflare already supports positive security for APIs through Schema Learning and Schema Validation. We are now extending this protection to web applications through Application Schema Profiles. You onboard an application, we learn its profile, and then we start to deploy an always-on detection that identifies non-conformity. All automated and enriched by powerful analytics.

We are opening a closed beta to invited Enterprise customers without API Security; customers with API Security already have access.

Validating requests based on learned profiles

Schema Profiles periodically learn the expected request structure from observed traffic. After a profile is available, an always-on validation layer is automatically deployed on live traffic. For every request, the detection evaluates whether it conforms or not with the profile, and it adds the result as metadata, augmenting the information already associated with the request. The signal does not take action by itself: customers can analyze past traffic in Security Analytics and decide where enforcement is appropriate and create Security Rules to block non-conforming requests. Requests to operations without a profile are not classified by this feature.

Unlike Managed Rules, failing validation does not require a request to match a known attack signature. A value outside an expected range, an unknown enum value, an invalid universally unique identifier (UUID), or unexpected characters — all can be identified because they differ from the learned profile.

For example, consider the following operation: 

www.example.com/shop/2dbda2e7-cfc9-448d-9465-799d2e6ff363/inventory?product_id=938062541

Below we describe the learning process, which evaluates only the structure and format of the request. When enough traffic has been observed, we learn that the path expects a UUID variable and that product_id is an integer and what its boundaries are. When product_id contains a string, it will be flagged as a violation. Similarly, Cloudflare can identify malformed UUID values and, when the customer enables enforcement, prevent non-UUID input from reaching the corresponding handler. These simple filters reduce the range of inputs an attacker can send, preventing the vast majority of typical attack vectors, such as SQL injection, cross-site scripting, remote code execution and more. 

Non-conforming does not always mean malicious. An application release, a new client, or an unusual but valid request may also introduce a difference. We recommend starting in observation mode, so customers can review a profile's effect before enforcement.

Learn the expected structure of requests

To determine the anticipated request structure for a web or API application, Schema Profiles routinely analyze observed traffic. Each profile may include the following, depending on the application traffic:

  • Path variables 
  • Query parameters
  • Headers and cookies
  • Body structure (JSON body or form-encoded)

For each field, the system learns its data type (integer, string, boolean, arrays, UUID or enum) and constraints such as numeric ranges, short enumerations, string lengths, and character classes.

Learning applies to operations that customers select for profiling. In Web Assets, an operation is Cloudflare's term for an operation identified by its HTTP method, hostname pattern, and path pattern. Web Assets continuously discovers operations and lists them under Web Assets > Operations. Customers can also add operations manually. Profiling doesn’t automatically start for discovered operations, while manually created operations do trigger profiling when created. For discovered operations, the customer must intentionally select Learn profile from the operation's overflow menu. 

After profiling is enabled, Cloudflare collects qualifying traffic and runs learning automatically once a week for each zone, using the most recent successful traffic. An operation needs at least 1,000 requests that received a 2xx response in the previous seven days to learn fields, and at least 10,000 to learn data boundaries. Successful requests can include bots and scanners, so customers should review a learned profile before enforcing it. Our roadmap includes allowing customers to trigger learning on demand and excluding automated traffic.

Once learned, profiles can be reviewed by selecting View details of the operation and finding the learned schema in the Security overview panel. If a learned schema is not shown, Cloudflare is still collecting data for the profile. Customers can also export the profile as an OpenAPI v3 schema file.

Learned profiles update each week as application traffic changes. New fields are added and fields that are no longer observed are removed, so validation tracks how the application changes. Customers can pin and save the learned schema by downloading the learned schema and uploading it to Schema Validation.

Review before you block

Security Analytics now includes a new Profile Analysis tab. Customers can select a validation profile and see traffic trends, including how many requests did not conform to the learned profile during the previous seven days. 

Customers can review the conforming and non-conforming traffic. They can drill into violations and review sampled logs to see where the violation occurred, which field was affected, and why it failed validation. Violations are classified into ten reasons, including type mismatches, values outside a learned range, and invalid formats.

Once a team understands the effect, it can use Security Rules to act on the signal. A rule can cover an entire application or be limited to selected paths, operations, or fields. Teams control where to monitor and where to block.

Positive security for web and API traffic

Traditional WAF learning modes can build detailed positive-security policies, but they often require operators to review suggestions, stage changes, and maintain policy entities. Cloudflare Schema Profiles expose validation as a request field cf.schema_validation.learned.violated, allowing customers to combine it with request properties, Bot Score, Attack Score, and other signals in a single Security Rule. By creating simple rules, teams can combine detections and define precisely when Cloudflare should take action.

Two other classes of fields are available to create more targeted rules. First, there are fields that collect where the violation occurred. For example, based on our initial example, if the value of product_id query parameter does not conform with the profile, the following field will be populated cf.schema_validation.uploaded.query.violated_parameters = ["product_id"]. This allows customers to create rules that enforce positive security only on specific fields or exclude them from the enforcement.

The second class of field collects new parameters that are not present in the profile. This is useful when you want to handle requests with new parameters (e.g. when deploying a new version of your application), or restrict your posture even further by blocking any parameters that were not detected or defined in the past.

Use case

Field

Location values

Example

Identify where in the request the violation occurred

Array up to 20 items

cf.schema_validation.learned.[location].violated_parameters

query,path,headers,cookies,body

cf.schema_validation.learned.query.violated_parameters = ["product_id"]

Identify whether an undeclared parameter is seen in the request 

Array up to 20 items

cf.schema_validation.learned.[location].undeclared_parameters

query

cf.schema_validation.learned.query.undeclared_parameters = ["adminMode", "utm"]

Coming up: critical field analysis, how we help you roll out positive security

Even with a flexible enforcement design, customers tell us that deploying a positive security policy is operationally complex. A large application can have thousands of operations with tens of thousands of fields. But not all operations and fields carry the same risk. Contextualization and prioritization helps security teams roll out positive security in a controlled and confident manner.

LLMs can help contextualize learned profiles to provide additional insight. For web applications, paths and field names are usually self-explanatory, thus semantic. For example, we piloted running a model hosted on Workers AI across four random applications’ learned profiles. The model successfully identified the link between clientId and account_number across two applications of a system, as well as the common dependency of using One-Time Password (OTP) for enhanced authentication. Highlighting this context enables security teams to prioritize actions such as configuring Rate Limiting Rules to defend against account-focused brute force attacks.

These LLM-powered insights will be accessible directly within the dashboard alongside each operation in Web Assets. Before executing a one-click deployment, teams can evaluate rule recommendations designed to secure these key fields, backed by mitigation simulation using past traffic to gain confidence.

Beyond contextualizing operations with semantic insights and risk indicators, we are developing additional metrics to order operations using historical request trends and signals. This enables security teams to focus mitigation efforts on the highest-priority operations first, including:

  • Data loss: upward trend of unusual increased data transfer
  • Reconnaissance activity: high count of unknown parameters
  • Business criticality: total volume of traffic correlated with the unique session IDs served

What’s available today

Customers with API Security already have access, given that this is an extension of Schema Learning and Schema Validation. We are opening a closed beta to customers without API Security who can test Schema Profiles on production web application traffic, meet with the product team, and provide detailed feedback on profile accuracy, analytics, and enforcement controls. Access is by invitation and does not imply future plan availability. If you are not an API Security customer and want to get access, contact your account team.

The feature supports paths, query parameters, headers, cookies, JSON request bodies, and form-encoded request bodies. Profiles can validate integers, strings, UUIDs, arrays, and enums containing up to three values. Multipart forms, GraphQL, and XML are not supported at this time.

Schema Profiles validate every value when a parameter name is repeated, but they do not enforce parameter uniqueness. They also do not learn and enforce required parameters or block a request solely because it includes a new parameter.

Get ahead of zero-days

Our idea for Application Profiles does not stop at validating request structure. The same workflow can learn other characteristics of what an application expects (such as ASNs or JA4s), explain when traffic deviates from them, and give security teams confidence in defining what “good” looks like. With a Proactive Security workflow, we help security teams get ahead of zero-days!

Using Device Linking to Eavesdrop on WhatsApp and Signal

Post Syndicated from Bruce Schneier original https://www.schneier.com/blog/archives/2026/09/using-device-linking-to-eavesdrop-on-whatsapp-and-signal.html

Modern messaging apps allow users to link their phone accounts to their computer desktop. Eavesdroppers are taking advantage of this capability:

Apps such as WhatsApp Web and Signal Desktop allow people to use their accounts on other devices, such as laptops or desktop computers.

Germany’s Customs Office has been using these features to connect a police-controlled computer to a suspect’s account.

Once connected, messages can be delivered to that computer without the police having to crack the encryption protecting them.

Netzpoltik details that police are able to gain access in this way either through physical access to someone’s phone or by intercepting verification codes via a state-sanctioned phishing attack or intercepting SMS messages via telephone surveillance.

That last paragraph is important. Making this work requires user consent.

What we want is a feature that displays connected devices, so users could notice if a new device gets connected to their account.

Колъм Тойбин: Без уютни събития, без лесни преживявания, без неоспорима развръзка

Post Syndicated from Антония Апостолова original https://www.toest.bg/kolm-toybin-bez-uyutni-subitiya-bez-lesni-prezhivyavaniya-bez-neosporima-razvruzka/

Колъм Тойбин: Без уютни събития, без лесни преживявания, без неоспорима развръзка

Световноизвестният ирландски писател Колъм Тойбин идва за пръв път у нас по покана на ICU – българското издателство на книгите му. В навечерието на неговото гостуване излезе и романът му от 2017 г. „Дом на имена“. Срещата на автора с читатели във формат „въпроси и отговори“ и подписване на книги ще се състои на 6 октомври от 16 ч. в книжарница Umberto & Co. На 7 октомври ще се проведе и галавечер с водеща Надежда Московска и с участието на преводачките Бистра Андреева и Елка Виденова. Събитието ще започне в 19 ч. на голямата сцена на Младежкия театър.

Когато не пишете за велики писатели или митологични персонажи, историите Ви следват живота на съвсем обикновени хора, на които се случват съвсем обикновени неща (по Нортръп Фрай). В какво се състои литературното обаяние на обикновеността за Вас? 

Знам какво имате предвид под „обикновено“, само че понятието „обикновен“ не означава кой знае какво за мен, когато работя. Често пиша за света на детството и за семейството. Но дори когато не е така, се опитвам да драматизирам сложността и двусмислието, а за това е необходимо да съзра някакво вътрешно богатство в героя си, независимо от житейските и другите обстоятелства около него. Написал съм около дузина романи и без да съм го планирал, се очерта следният модел: един роман е за известен автор или митичен персонаж, а следващият – за някой „най-обикновен“ човек от моя роден град или семейство. Изпитвам облекчение, когато преминавам от едното към другото и обратно. 

Двусмислената (да се захвана за тази Ваша дума) концепция за дома бележи голяма част от творчеството Ви. Самият Вие сте живял в редица държави. Стигнахте ли до надеждна дефиниция за „дом“?

За човек на моята възраст домът е там, където са компактдисковете му! Признавам си, че колкото и да е странно, не разсъждавам много върху това, върху концепцията за дом. Когато обаче хвана самолетен полет от някой американски град за Дъблин, знам много добре, че се прибирам у дома. Знам къде и какво е домът. Това е Ирландия. Това е Уексфорд. Това е мястото, откъдето съм. 

В „Празното семейство“ пишете: „Всяка събота ходех до Пойнт Рейъс, за да усетя болката по дома.“ Коя емоция за Вас най-автентично изразява принадлежността? 

На някои езици – на испански например – е трудно да се преведе самата дума miss [в оригиналния текст Тойбин пише буквално to miss home, „за да ми липсва домът“ – б.а.]. Предполагам, че в изречението, което цитирате, се опитвам да внуша идеята, че това „да ти липсва домът“ е вид фантазия, представление, нещо може би реално, но също така вероятно и изкуствено. 

Централен персонаж в много от историите Ви е преобърнатият архетип на майката – понякога проблемно, отчуждено, дестабилизиращо присъствие. Или по-скоро отсъствие – нещо, с което често работите. Сред най-мощните въплъщения на последното е quest-ът, търсенето на майката, в разказа „Дълга зима“…

Пристъпвам към историите една по една. Не следвам някаква определена теория. Не се опитвам да доказвам нищо. Художествената литература се нуждае от разрив, от нарушаване на баланса, от герои, които не се държат по очакван и обичаен за образа си начин. Една любяща майка няма да ми свърши особена работа. Няма какво да я правя. Ами ако майката не е майчински настроена? Ами ако присъствието ѝ е вредно и пагубно? Ами ако именно отсъствието ѝ е онова, което има значение и въздействие? Няма ли да бъде по-интересно? Иначе, що се отнася до „Дълга зима“, написах разказа малко след като майка ми и брат ми починаха и цялата мъка, привнесена в тази история, беше все още съвсем жива и оголена за мен. Не бях я планирал като почти автобиографична, но се получи тъкмо такава. 

Подобно на „Одисея“ изобразявате завръщането у дома като много по-голямото изпитание, отколкото напускането му. У Вас и двете могат да се четат, понякога едновременно, като акт на окончателно пристигане, на бягство, на спасение, на поражение…

Обожавам края на „Одисея“. Няма щастливо завръщане. Шеги, ирония, увъртания. Обичам да се завръщам в Ирландия. Но това простичко чувство не трае дълго. То е фалшиво усещане. Изобщо, много ми харесва идеята за измамното, невярното усещане в литературата – то ми допада повече от автентичността или искреността. Предполагам, това, което в крайна сметка се опитвам да кажа, е, че белетристиката трябва да бъде чисто и просто интересна. Без уютни събития, без лесни преживявания. Без неоспорима развръзка. 

В какво се състои себепознанието у героите Ви? Тяхната идентичност често е диалектична – нещо, което градят по необходимост след криза, посттравматично, компромисно.

Този въпрос е лесен, или поне отговорът му е такъв. В моите романи героите ми правят нещо, мислят, помнят. Не ме занимава идеята за някаква всеобхватна идентичност, нито дори проблемите на себепознанието. Нямам теория за човешкия характер. Нямам дарба за абстрактно мислене. За мен съществуват единствено и само следващият образ, следващото изречение, следващата сцена. Работата ми се състои в това да създам нещо интересно и истинско. Така че оставям персонажите си да живеят колкото могат. Много често те мълчат за важните неща, а това придава на вътрешния им свят суров и нелицеприятен вид, прави ги неспособни да общуват лесно и директно. Разликата между това, което чувстват, и това, което разкриват, дава огромна енергия на повествованието. Интересувам се от тази раздалеченост. Интересувам се от един герой, най-много двама, на едно място. Интересувам се от личния интимен живот – какъв е, как се усеща; от вътрешния свят.

Заговаряйки за интимността, със сигурност сте казал достатъчно за това какво е да се пише за гей сексуалността в един, поне доскоро, репресивен социално-културен контекст като ирландския (и българския, уви). И по-важното – за автоцензурата, преодолявана по пътя към тези Ваши сурови, неподправени, натуралистични описания на секса между мъже. 

Понякога е важно да не се пишат графични сексуални сцени. Няма закон, който да казва, че трябва. Но понякога начинът, по който героите правят секс, е съществен за историята. Предполагам, че има няколко правила: без метафори, без сравнения, без завоалиран или натруфен стил. Просто кажете какво са направили персонажите, не как са се чувствали. Впрочем точно днес получих имейл от хетеросексуален приятел, който реагира на интимните гей сцени в новия ми роман (The Bridge). Та той пише, че се е почувствал възбуден от описанията. Това ми се стори хубаво. Една от секс сцените в тази книга е в затвор, та имах добро основание да я направя графична – важен беше начинът на правене на любов, конкретните физически действия. 

В творчеството Ви любовта често се явява форма на задължение – преплетена с неизбежност, с наложителност, с премълчаване. Има ли изобщо нещо лесно и освобождаващо в нея? 

Може би, но то не върши работа в един роман. Иначе, в разказите се опитвам да работя по ръба на това, което може да бъде изречено, и онова, което трябва да остане неизказано. Кимването е важно, смръщването, въздишката, полуизказаното, необлечената в слово мисъл, внезапното изтърсване на нещо.

В последния си издаден на български роман „Дом на имена“ ни давате гледните точки на Клитемнестра и децата ѝ Електра и Орест, но не и на Агамемнон. (Интересно, че той бе лишен от лице и почти от глас и в „Одисея“ на Нолан). Вашият специфичен поглед върху тази архетипна история? 

Първоначално исках да работя с това, което бих нарекъл „стакато в първо лице“: гласовете на жените – Клитемнестра и нейната дъщеря. Но впоследствие бях очарован от историята на Орест, от неговата срамежливост, от неговото отсъствие, от неговата сдържаност. Нямах никакъв интерес към гласа на Агамемнон [който бива убит от съпругата си, след като принася в жертва дъщеря им Ифигения – б.а.], към мотивите му или към неговата версия за случилото се. Накрая поставих именно Орест в центъра на историята. 

Всеки акт на насилие там води не до развръзка, а до отварянето на нов цикъл от болка. 

Да, всяко убийство следваше като вид възмездие. В един момент, когато пишех за насилието в Северна Ирландия, забелязах тази спирала – убийства тип „око за око“, убийства за отмъщение. Именно това беше в съзнанието ми. 

Намираме се на прага на настъпващата ера на изкуствения интелект. Когато един ден ИИ овладее писането, кой недостатък на създадената от човека литература би Ви липсвал най-много?

Знанието, че си се провалил. Самата концепция за провал.

 

Building event-driven applications at scale with Amazon EventBridge

Post Syndicated from Nahid Karimaghalou original https://aws.amazon.com/blogs/compute/building-event-driven-applications-at-scale-with-amazon-eventbridge/

Event-driven applications on Amazon EventBridge usually start small and then spread. One team creates a Custom event bus, adds a few rules, and ships. Another team needs some of those events, so a rule forwards them to a bus in a second account. A third team needs a subset of what the second team receives, so another rule forwards again. A year later the organization runs dozens of Custom event buses joined by forwarding rules, and that topology has become a thing to operate in its own right.

That shape has a price, and the smallest part of it is the bill. Every forwarding hop is a separate ingestion, so cost tracks the topology rather than the number of consumers that needed the event. The harder problem is that nobody can see the whole picture. Governance spreads across the accounts it was meant to cover. Answering who publishes to a bus, who consumes a given event type, or what breaks when a team stops publishing means visiting each account and reading its rule configuration. Tracing one event is harder still: its path crosses several buses in several accounts, each with its own metrics and logs, and no single view follows it from publication to the consumer that never received it.

Application teams also wait. Publishing to a bus in another account, or consuming from one, needs a resource policy, a role, and a forwarding rule owned by a central team. The team that wants to build opens a ticket, and the platform team becomes a queue. Both the missing visibility and the waiting grow with every team onboarded.

Amazon EventBridge recently relaunched the Custom event bus, which tackles these challenges directly. A platform team creates one bus, shares it across the organization, and keeps control of who can publish and who can subscribe. Every consumer of those events is listed on the one bus rather than inferred from configuration spread across accounts. Application teams create their own Subscribers in their own accounts. The bus stores events for a retention period you choose, preserves order within a key the publisher sets, accepts Avro and Protocol Buffers (Protobuf) alongside JSON (including CloudEvents), and delivers to targets without a function in the path to translate a call. It runs alongside the Custom event bus – classic, so adoption is incremental.

In this post, you see how a platform team stands up a shared bus and governs access to it, how application teams onboard themselves with a single Subscriber resource, and how retention, ordering, open formats, transformation, and direct target integrations change what one bus can carry.

One bus, shared with the organization

The platform team’s job on a shared bus is narrower than it was on a fleet of them. It owns the bus and sets the boundaries: which principals can publish and what their events can declare, which principals can subscribe, and, where it matters, what those principals are allowed to filter on. Application teams then manage their own configuration within those boundaries, such as filters, targets, delivery roles, retry policies, and failure destinations, none of which the platform team needs to write or review. That division is the point of the design. The platform team keeps governance of the bus and stops owning everyone else’s configuration, which is what takes it out of the provisioning path without giving up control of who is on the bus.

Creating the bus is a single call in a platform account.

BUS_ARN=$(aws eventsv2 create-event-bus \
    --name company-events \
    --storage-configuration '{"RetentionPeriodInDays":7}' \
    --query EventBusArn --output text)

Retention is the one setting worth deciding deliberately here rather than revisiting after an incident. It runs from 1 to 365 days and can be modified later, but a change only applies going forward. Raising it widens the window for events published from that point on, and does not make older events readable again. Seven days covers a working week of history, which is usually enough to onboard a consumer or reprocess after a bug without paying to store a year of events nobody will read.

Sharing the bus is the second decision. AWS Resource Access Manager is the route to reach for first: it associates automatically for accounts in the same organization and reaches accounts outside it by invitation the consumer accepts. A resource policy written on the bus directly is the alternative, and can also name accounts inside or outside the organization.

Access is granted per principal, and publishing and subscribing are separate permissions. A team that produces order events gains no ability to read payment events from the same bus. One grant is not enough for a cross-account caller, as usual on AWS: the role that publishes or subscribes also needs its own IAM policy allowing those actions. The platform team decides which accounts can reach the bus, and each consuming team decides which of its own principals can use that access.

Taken together, those decisions produce the architecture in the following diagram. One bus lives in a platform account, and application teams publish to it and subscribe from their own accounts. An AWS Lambda function in Team A’s account calls PutRawEvents to publish events onto the Amazon EventBridge bus in the platform account. Team B and Team C each attach their own Subscriber: Team B’s delivers to a Lambda function, Team C’s to an Amazon DynamoDB table.

Architecture diagram of one Custom event bus in a platform account. A Lambda function in Team A’s account calls PutRawEvents to publish events onto the Amazon EventBridge bus in the platform account. Team B and Team C each attach their own Subscriber in their own accounts: Team B’s Subscriber delivers to a Lambda function, and Team C’s Subscriber delivers to an Amazon DynamoDB table.

Figure 1: Multi-account sharing

Cost follows team boundaries because charges separate ingestion from delivery. The account that publishes an event pays to put it on the bus, and the account that owns a Subscriber pays for what that Subscriber consumes. Each team’s usage appears on its own bill, which is what makes a shared bus something a platform team can charge back rather than a shared cost center nobody can decompose. Removing the forwarding hops also removes the duplicated ingestion and delivery those hops created: the same event reaching the same three consumers is ingested once instead of three times.

Publishing in the format teams already use

Not every producer speaks JSON. Teams that standardize event exchange across an organization often register schemas and publish compact binary payloads, because the schema is the contract between teams that deploy on their own timetables. Accepting the formats those producers already emit is simpler than changing each one to convert to JSON first.

With the new Custom event bus, application teams can publish events in Avro, Protobuf, and CloudEvents (JSON) formats. For the binary formats, a schema registry named on the request is used to deserialize the events.

There are two publish APIs, and the payload decides which one to call. PutEvents takes structured JSON with the familiar Detail, Source, and DetailType fields. PutRawEvents takes a binary payload plus metadata you define, and is the one to use for Avro, Protobuf, CloudEvents, or bytes the bus should not interpret.

import boto3

events = boto3.client("eventbridgev2")
events.put_raw_events(
    EventBusArn=BUS_ARN,
    SchemaRegistryConfiguration={"RegistryUri": GLUE_REGISTRY_ARN},
    Entries=[
        {
            "Data": avro_encoded_order,  # bytes, straight from your existing producer
            "SystemMetadata": {"ContentType": "application/avro"},
            "Metadata": {"eventType": "OrderPlaced"},
        }
    ],
)

The schema registry can be either the AWS Glue Schema Registry or the Confluent Cloud Schema Registry.

Because the bus decodes the event before filters and transformations run, a consumer subscribing to Avro events written by another team needs no schema, no decoder, and no access to the registry. It writes the same filter it would write against JSON. Producers and consumers stay decoupled, and no deserialization code has to be repeated in each consuming team.

Publishers get one more setting on the same request: deduplication. A retry that already succeeded would otherwise leave a duplicate for every consumer to handle. It works one of two ways: the bus hashes the content of each event, or it uses a deduplication ID you supply. Content-based hashing suits producers with no natural key, since two identical events hash the same. A deduplication ID fits when you already have one, such as an order ID combined with a state transition. It keeps matching even when parts of the payload differ in ways that should not count as a new event.

Self-service onboarding for application teams

The new Custom event bus introduces a new resource called a Subscriber. Application teams create and configure their own Subscribers in their own accounts, provided they have been granted subscribe access to the bus. A Subscriber is the one place a consumer’s behavior is defined: which events it receives, where they are delivered, how delivery is retried, and where events go when delivery does not succeed. Reviewing or changing a consumer is one thing to read and one thing to update.

SUBSCRIBER_ARN=$(aws eventsv2 create-subscriber \
    --name orders-to-fulfilment \
    --event-bus-arn "$BUS_ARN" \
    --filter-configuration '{"Filters":[{"Scope":"METADATA","Pattern":"{\"eventType\":[\"OrderPlaced\"]}"}]}' \
    --invoke-configuration '{"TargetArn":"'"$QUEUE_ARN"'","RoleArn":"'"$ROLE_ARN"'"}' \
    --retry-policy '{"MaxRetryAttempts":10,"MaxEventAgeInSeconds":3600}' \
    --on-failure-configuration '{"Arn":"'"$DLQ_ARN"'"}' \
    --query SubscriberArn --output text)

A filter’s scope decides which part of the event the pattern is matched against. DATA matches the payload, METADATA matches the key-value pairs the publisher attached to the event, and SYSTEM_METADATA matches the event’s system fields: the content type and ordering key a publisher declares, plus the fields Amazon EventBridge adds itself. Because Avro and Protobuf payloads are decoded as they are published, a DATA filter reads their fields directly, the same as it would for JSON.

The retry policy says how the bus should behave when a target is failing. MaxRetryAttempts sets how many times a delivery is retried, and MaxEventAgeInSeconds sets how long an event stays eligible for retry, measured from when it was published. Retries stop as soon as either limit is reached, so both bound the same delivery.

When deliveries do fail, the reason shows up in the Subscriber’s own logs, which application teams can turn on themselves. They record the error from each delivery attempt alongside the exact input sent to the target, which makes a problem quick to place. Seeing what the target actually received separates a transformation that produced the wrong shape from a target that rejected a correct one.

Screenshot of the Amazon EventBridge console showing the Create subscriber form, with fields for the subscriber name, event bus, filter configuration, target (invoke configuration), retry policy, and on-failure destination.

History for consumers that did not exist yet

A Subscriber sometimes needs events that were published before it existed. For example, a new analytics service needs hydrating with recent history, or a target processed a window of events incorrectly and needs that window replayed. Because the bus retains events for the period configured on it, a Subscriber can be created with a starting position in the past, so it reads history, catches up, and continues with live traffic:

aws eventsv2 create-subscriber \
    --name analytics-backfill \
    --event-bus-arn "$BUS_ARN" \
    --starting-position POINT_IN_TIME \
    --point-in-time-configuration '{"PointType":"TIMESTAMP","StartingPoint":"2026-09-14T06:00:00Z"}' \
    --filter-configuration '{"Filters":[{"Scope":"METADATA","Pattern":"{\"eventType\":[\"OrderPlaced\"]}"}]}' \
    --invoke-configuration '{"TargetArn":"'"$ANALYTICS_ARN"'","RoleArn":"'"$ROLE_ARN"'"}'

A starting position is either LATEST or POINT_IN_TIME. Choosing POINT_IN_TIME then needs a point-in-time configuration: a PointType of TIMESTAMP with a starting point, or HORIZON to begin at the earliest event still retained. An optional end point stops the read at a chosen time, which is what you want when reprocessing a known-bad window rather than catching up to live traffic.

Two things to keep in mind. The starting position is fixed when the Subscriber is created, so reading a different window means a new Subscriber. Treat the starting position as part of a Subscriber’s identity rather than a dial to turn later. And retention cannot reach back beyond the retention window, so the read starts at the earliest retained event however far back the timestamp asks for.

Order, where order matters

In event-driven architectures, where components are built to work asynchronously, the order events arrive in usually does not matter. There are still use cases where a consumer relies on ordered delivery, and the new Custom event bus offers it as an option on individual Subscribers.

Ordering is scoped by a key the publisher sets. A publisher includes an event group ID (a customer ID, an order ID, a driver ID), and a Subscriber created with FIFO delivery type receives the events for each group in the order they were published. A FIFO Subscriber reading events published without a group ID has nothing to sequence by, so the two sides work together. Creating one takes the same call as an unordered Subscriber, with the delivery type set to FIFO:

aws eventsv2 create-subscriber \
    --name inventory-ordered \
    --event-bus-arn "$BUS_ARN" \
    --type FIFO \
    --filter-configuration '{"Filters":[{"Scope":"METADATA","Pattern":"{\"eventType\":[\"OrderPlaced\"]}"}]}' \
    --invoke-configuration '{"TargetArn":"'"$FIFO_QUEUE_ARN"'","RoleArn":"'"$ROLE_ARN"'","SqsParameters":{"MessageGroupId":"{% $events.SystemMetadata.EventGroupId %}","MessageDeduplicationId":"{% $events.SystemMetadata.DeduplicationId %}"}}'

Ordering is per group, so throughput scales with the number of groups. If an event cannot be delivered, it holds up the rest of its own group while other groups keep moving. Choosing the key therefore matters: one that maps to a business entity, such as an order or a customer, gives sequencing where it is needed and independence everywhere else. A key so broad that most events share it puts them all in a single sequence, and a key so specific that every event has its own leaves nothing to order.

Because ordering is set on each Subscriber, consumers of the same events do not need to agree on it. An inventory service can receive a group’s events in sequence while an analytics service subscribing to those same events takes them as they arrive.

Reshaping events, and delivering directly to a target

A consumer’s business logic expects events in a particular shape, and the events on the bus are not always in that shape. Where the two get reconciled is an ownership decision: inside the consumer, where it becomes part of that team’s code, or on the Subscriber, ahead of it.

The first case is reformatting. A downstream system, often owned by another domain or outside the organization entirely, expects a different structure from the one the publisher emits. A JSONata transformer on the Subscriber produces that structure before delivery, so the consumer receives what it already expects. The business logic stays where it belongs, and when the published shape changes upstream, or another event type needs deriving into the same input, it is the transformer that changes rather than the consumer:

--transformer '{
    "Type":"JSONATA",
    "JsonataConfiguration":{
        "Expression":"{% {\"orderRef\": $events.Data.detail.orderId, \"total\": $events.Data.detail.amount} %}"
    }
}'

The transformer type determines the shape of what gets delivered. RAW delivers the event payload as is and is the default, so a Subscriber with no transformer configuration receives only the payload. WITH_METADATA adds the event envelope alongside it, and JSONATA reshapes it with an expression wrapped in {% %}.

The transformation reshapes events only for the Subscriber that owns it and does not affect what other Subscribers of the same bus receive. That also makes it a data minimization control: a partner can receive only the fields it needs rather than a whole internal event. Defining it at the Subscriber means it holds for every event without anyone remembering to strip fields.

The second case is calling an AWS service API. A Subscriber delivers directly to targets including Amazon Simple Queue Service (Amazon SQS), Amazon Simple Notification Service (Amazon SNS), AWS Lambda, and Amazon Kinesis Data Streams. For other services it has been common practice to add a proxy step whose only job is to make the call. With universal targets, the new Custom event bus can call a supported AWS service API directly, with the request body built by a JSONata expression.

TargetArn: arn:aws:events:::aws-sdk:dynamodb:putItem
UniversalTargetParameters.Input:
{% { "TableName": "orders", "Item": { "pk": { "S": $events.Data.detail.orderId } } } %}

Note that a universal target shapes its input through that parameter rather than through the preceding transformer, and setting a transformer on one is rejected when the Subscriber is created. The two mechanisms do the same kind of work on different targets.

That removes the proxy processing that existed only to make the call. The delivery role still needs the action the target requires and getting that wrong is the most common cause of a Subscriber that looks healthy and delivers nothing.

Conclusion

Running an event-driven application across many accounts no longer means running many event buses and the forwarding between them. A platform team creates one new Custom event bus, shares it across the organization through AWS Resource Access Manager or a resource policy on the bus, and keeps one place to decide who publishes and who consumes. Application teams create and own their Subscribers without waiting for provisioning. Ingestion and delivery are charged separately, so each team’s usage appears on its own bill, and the duplicated ingestion that forwarding hops created disappears with the hops.

The capabilities that used to send individual teams elsewhere now sit on the same bus. Ordering is per Subscriber and scoped by a publisher-supplied key, so one team’s sequencing requirement no longer fragments an architecture. Retention makes it possible to onboard a consumer that needs history it was never subscribed to. Avro and Protobuf are decoded by the bus, so producers keep their binary contracts. Transformation and universal targets keep business logic where it belongs, removing the proxy steps that existed only to reshape an event or make an API call.

Because the new Custom event bus runs alongside the Custom event bus – classic, adoption is incremental. Point one new consumer at a shared bus or forward a slice of an existing bus into it and move the rest as teams are ready.

Next steps. Create a bus, add a Subscriber, and publish an event, starting from the Amazon EventBridge documentation for the resource model and the AWS Command Line Interface (AWS CLI) reference. If you already run Custom event buses, the migration guidance covers routing existing events into a new Custom event bus without changing producers. From there, look at the Subscriber logging and metrics options for tracing an event from publication to delivery, and at AWS Resource Access Manager for how sharing and permissions work across an organization. If you have questions or feedback about the new Custom event bus, leave a comment on this post. We’d like to hear how you’re using it.

Improving Lambda function latency with scalable network bandwidth

Post Syndicated from Rahul Shandilya original https://aws.amazon.com/blogs/compute/improving-lambda-function-latency-with-scalable-network-bandwidth/

AWS Lambda now supports scalable network bandwidth for functions configured with 2,048 MB of memory or more, running outside of a virtual private cloud (VPC). Previously, sustained network throughput was capped at 625 Mbps regardless of your function’s memory configuration. Now, sustained throughput scales proportionally from 625 Mbps at configurations below 2,048 MB up to 3,000 Mbps at 10,240 MB, increasing the rate at which data moves to and from your execution environment.

In this post, you learn how to apply this new capability to latency-sensitive data processing workloads, helping reduce function execution times and per-invocation costs while improving the end-user experience through reduced latency. You also walk through a deployable implementation that demonstrates the performance improvements this capability unlocks.

Latency-sensitive data processing

Latency-sensitive data processing applications are data processing workloads that must be completed in a defined period of time. They often experience bursty, ad hoc traffic patterns while being required to download gigabytes or even terabytes of data from a data store, process it in a compute environment, and return a result to a waiting end user.

Latency-sensitive data processing is often highly parallelizable. Data can be divided into smaller pieces with each piece being individually processed before combining them together to obtain a result.

These workloads can be found in multiple industries and verticals. Examples include:

  • Log querying engines – An end user initiates an on-demand search across terabytes of log data and expects results within seconds.
  • Insurance underwriting – A prospective customer submits an application, triggering real-time evaluation of historical claims and risk data. The underwriting process determines what coverage and premiums to offer to the prospective customer.
  • Financial ETL pipelines – An economic announcement triggers an unexpected burst of market data that must be ingested, transformed, and made available to downstream trading systems before the next market tick.
  • Genomics platforms – A clinician orders a diagnostic test, requiring gigabytes of DNA or RNA sequencing data to pass through a bioinformatics pipeline and be compared against a reference genome while the patient awaits results.

These workloads are challenging to build on traditional compute clusters. Their spiky and unpredictable nature forces you to choose between under-provisioning compute to optimize costs (and risk missing your SLA) or over-provisioning and paying for idle capacity.

Why Lambda fits latency-sensitive data processing

Lambda eliminates this tradeoff. Instead of pre-provisioning a compute cluster, Lambda scales compute capacity in response to incoming requests, matching processing power to unpredictable traffic patterns. Because latency-sensitive data processing is highly parallelizable, the ability of Lambda to rapidly scale out execution environments makes it a natural fit. You can fan out across thousands of concurrent functions to process data in parallel, paying only for the compute you use.

However, as data volume and performance requirements grow, network bandwidth to and from the compute environment can become the limiting factor in minimizing workload latency.

Scalable network bandwidth directly addresses this limitation by raising the per-environment network throughput ceiling, improving the rate at which data can be transferred to and from the execution environment. Each execution environment can now drive up to 3,000 Mbps of sustained throughput when configured with 10,240 MB of memory, a 4.8x increase from the previous ceiling of 625 Mbps. Combined with the ability of Lambda to scale out at a rate of 1,000 execution environments every 10 seconds, you can download more than 3 TB of data in under 10 seconds.

New network throughput behavior for Lambda functions

Scalable network bandwidth applies to both data ingress to and egress from an execution environment for functions outside of a VPC. For functions configured with 2 GB of memory or more, network bandwidth scales by approximately 280 Mbps increments for every 1 GB of additional memory allocated.

The following table shows the maximum sustained bandwidth available to each execution environment at each memory configuration.

Memory Configuration Max Sustained Bandwidth
Less than 2,048 MB 625 Mbps
2,048 MB 765 Mbps
3,072 MB 1,044 Mbps
4,096 MB 1,324 Mbps
5,120 MB 1,603 Mbps
6,144 MB 1,883 Mbps
7,168 MB 2,162 Mbps
8,192 MB 2,441 Mbps
9,216 MB 2,721 Mbps
10,240 MB 3,000 Mbps (4.8x increase)

Table 1. Lambda sustained network bandwidth by memory configuration. Bandwidth scales at ~280 Mbps per additional GB of memory above 2 GB.

In the following section, you learn how scalable network bandwidth improves end-user latency by building an ETL pipeline that demonstrates it. You can find the source code in the GitHub repository.

Solution overview

Consider a SaaS analytics platform where users submit ad hoc queries against a data store. The application must extract the relevant data, apply a filter or transformation, and return an aggregate result while the user waits. In this example, the result needs to be returned in 8 seconds or less.

The following diagram illustrates the architecture of the solution.

ETL fan-out architecture: a client calls an orchestrator Lambda function, which fans out to multiple worker Lambda functions that read data in parallel from Amazon S3, with bandwidth scaling callouts for each memory tier.

Figure 1. ETL fan-out pattern: an orchestrator Lambda function distributes work to multiple worker Lambda functions that read from Amazon S3 in parallel, with bandwidth scaling callouts per memory tier.

A client initiates an ad hoc query by calling the orchestrator Lambda function through the Lambda API. The orchestrator function determines how to split the work. To process the data in parallel, the orchestrator function uses a ThreadPoolExecutor to issue synchronous invoke requests to the Lambda worker function, fanning out the worker across multiple execution environments at the same time.

Each Lambda worker function is configured with 10,240 MB of memory, so it has access to up to 3,000 Mbps of sustained network throughput. After the data is processed, the aggregated result is returned to the client.

Prerequisites

Before you start the deployment process, make sure that you have completed the following steps:

  1. Install the AWS SAM CLI on your computer and confirm that you are running Python 3.12 or later.
  2. Have your AWS account credentials ready.
  3. Submit a request to AWS Service Quotas to turn on scalable network bandwidth for your Lambda functions. This quota is listed under Network bandwidth per execution environment.

Clone the source code from the GitHub repo and deploy the application within your AWS account. Creating the 10 GB test dataset and running the benchmark can incur charges to your AWS account.

git clone https://github.com/aws-samples/sample-lambda-enhanced-bandwidth
cd sample-lambda-enhanced-bandwidth
sam build
sam deploy --guided

After the AWS CloudFormation stack is deployed, record the DataBucketName and orchestrator function name from the stack outputs to use in subsequent commands.

To simulate data for the end user to query, the GitHub repo has a script that creates 10 GB of synthetic data and uploads it to your S3 bucket.

Mode 1: Processing pre-partitioned data

In Mode 1, the 10 GB of synthetic data is pre-partitioned. Pre-partitioned data is typically produced incrementally by many sources over a period of time, which can be the case with IoT data or access logs. The following command creates 10 GB of data divided into 20 partitions that are 512 MB each.

python scripts/generate_data.py \
    --bucket <DATA_BUCKET_NAME> \
    --total-gb 10 \
    --chunk-mb 512

Turning on scalable network bandwidth does not, on its own, make your downloads faster. A single download request only opens one connection to Amazon S3, and one connection does not move data fast enough to fill all the bandwidth now available to your Lambda function. To actually use your full allotment of network bandwidth, the execution environment has to pull the data over several connections at once. It does this by preferring the AWS Common Runtime (CRT) transfer client, a high-performance download engine built into Boto3. When the worker calls download_fileobj, the CRT client automatically breaks the 512 MB object into smaller parts and downloads them in parallel across multiple Amazon S3 requests. Those parallel downloads are what let a single Lambda worker take advantage of its full network bandwidth.

The following command runs the benchmark on the pre-partitioned data.

python scripts/run_fanout_benchmark.py \
    --orchestrator-name <STACK_NAME>-orchestrator \
    --bucket <DATA_BUCKET_NAME> \
    --iterations 5

Mode 2: Processing single large objects

Mode 2 generates 10 GB of data in one large object. This arrangement is more common when data is produced or delivered as one complete unit, such as database backups or genomic datasets. The following command creates 10 GB of data in a single large object.

python scripts/generate_data.py \
    --bucket <DATA_BUCKET_NAME> \
    --single-object-gb 10 \
    --key large-object/large-file.bin

In Mode 1, the CRT preference applies to Boto3 managed transfer methods such as download_file and download_fileobj. Mode 2 takes a different approach. Each worker reads a specific byte range of a single large object using get_object. The CRT preference setting has no effect on these calls. Instead, you can control concurrency by explicitly tuning the number of Lambda workers and using a bounded ThreadPoolExecutor to issue multiple byte-range requests at the same time.

When you run the following command, the orchestrator takes the single large object and divides it into consecutive byte ranges of 512 MB each. Each of the individual ranges is then processed by a Lambda worker execution environment in parallel.

python scripts/run_fanout_benchmark.py \
    --orchestrator-name <STACK_NAME>-orchestrator \
    --bucket <DATA_BUCKET_NAME> \
    --key large-object/large-file.bin \
    --slice-size-mb 512 \
    --iterations 5

The benchmark reports wall-clock duration, client-observed duration, aggregate throughput across workers, worker completion counts, and target compliance. When comparing memory configurations, keep the code, dataset, AWS Region, partition count, warm-up policy, and measurement count identical. You should run the benchmark multiple times in your account because placement, cold starts, concurrency, S3 behavior, and execution-environment reuse could affect results.

Results

To compare results, we ran the benchmark using a baseline configuration where the worker Lambda function is configured with only 1,024 MB of memory, well below the 2,048 MB threshold required for scalable network bandwidth to take effect. The 1,024 MB configuration limits network throughput to the previous sustained ceiling of 625 Mbps.

In our baseline test run, a worker downloaded and processed a single 512 MB partition with a 6.61-second download time at a 649.8 Mbps throughput (at p50). The 649.8 Mbps throughput exceeds the 625 Mbps ceiling because Lambda is capable of bursts in network throughput over a short period of time. Across twenty measured fan-out queries, the complete 10 GB query was completed with a 7.113-second wall-clock at p50. This fits within the 8-second SLA but leaves very little headroom.

To run our scalable network bandwidth benchmark, we re-deployed our worker Lambda function with a 10,240 MB memory configuration and re-ran the application. At a 10,240 MB memory configuration, each execution environment can now access up to 3,000 Mbps in sustained throughput. Direct 512 MB downloads achieved a 1.70-second download time and 2,521.3 Mbps throughput (both at p50). The complete 10 GB query was completed with a 2.640-second wall-clock at p50. That is 2.69 times faster, or 62.9% lower median latency, than the 1,024 MB configuration.

Table 2 summarizes the direct worker and end-to-end fan-out measurements for the same 10 GB dataset and 20 × 512 MB orchestration pattern. Aggregate throughput is the total data transfer rate across all twenty execution environments spun up to run the benchmark.

Memory Configuration Single 512 MB partition download time and throughput (p50) 10 GB fan-out wall time (p50) Aggregate throughput p50
1,024 MB baseline tier (sustained 625 Mbps) 6.61s / 649.8 Mbps 7.113s 12.08 Gbps
10,240 MB scalable tier (up to 3,000 Mbps) 1.70s / 2,521.3 Mbps 2.640s 32.54 Gbps

Table 2. Measured 1,024 MB baseline tier and 10,240 MB scalable bandwidth performance for a 10 GB fan-out ETL query.

Using scalable network bandwidth, the customer’s SLA headroom has improved by nearly 5 seconds. The Lambda function can now handle larger partitions within the same SLA window, reducing costs while still remaining comfortably within the customer’s SLA.

Clean up

To clean up the resources you created for the benchmark test, run the following commands:

aws s3 rm s3://<DATA_BUCKET_NAME> --recursive
sam delete --stack-name <STACK_NAME>

Best practices

After scalable network bandwidth is turned on for your AWS account, the following practices help you get the most out of it.

Profiling and planning

  • Test before you tune. Not every function is network-bound. Before increasing memory, profile your function to confirm that network I/O is the primary contributor to invocation duration and not CPU or application logic. Use Amazon CloudWatch Lambda Insights to inspect rx_bytes, tx_bytes, and duration. Functions where network I/O dominates invocation time are prime candidates for tuning.
  • Design for parallelism. Break your data into parallelizable chunks that can be processed independently in a fan-out pattern across multiple execution environments. You can use Amazon S3 byte-range reads to split large files into independently downloadable partitions. For implementation details, see Downloading an object with part numbers in the Amazon S3 User Guide.
  • Run AWS Lambda Power Tuning. Lambda Power Tuning is a state machine that helps you optimize your Lambda functions for cost and performance. Use Power Tuning to sweep memory configurations from 1,024 MB to 10,240 MB and identify the optimal cost-vs-latency point for your workload.

Implementation

  • Check upstream and downstream limits. Check the throughput limits of your data sources. For example, a Lambda function running at 3,000 Mbps can exceed the throughput capacity of a single S3 prefix, which supports up to 5,500 GET requests per second. When this happens, you will see HTTP 503 (Slow Down) errors in your application logs. Distribute your S3 objects across multiple prefixes to parallelize reads and avoid per-prefix throttling.
  • Balance bandwidth and CPU. Lambda allocates CPU proportionally to memory. For example, at a 1.7 GB memory configuration you are allocated 1 vCPU while a 10 GB memory configuration is allocated up to 6 vCPU. If your function processes data in parallel threads, the higher memory tiers give you both more network bandwidth and more CPU to process it. Use the concurrent.futures module in Python or worker_threads in Node.js to process data across parallel threads and maximize both CPU and network utilization.
  • Turn on Amazon S3 CRT for Boto3. If your function uses the Python runtime, initialize your Amazon S3 client with preferred_transfer_client: 'crt' to maximize single-connection throughput. The AWS Common Runtime automatically parallelizes requests across multiple TCP connections, which matters because individual TCP connections have a throughput ceiling.
  • Use SnapStart for JVM workloads. If you use Lambda SnapStart for Java functions, scalable network bandwidth reduces afterRestore hook latency. Network activity that occurs during function restore, such as pre-warming connections or pre-fetching configuration data, can complete faster.

Conclusion

Scalable network bandwidth raises the per-environment sustained throughput ceiling of AWS Lambda from 625 Mbps to 3,000 Mbps, directly reducing end-to-end latency for data-intensive workloads. Combined with the Lambda scaling rate, you can now move terabytes of data in seconds, without provisioning or managing infrastructure.

To get started, request the Network bandwidth per execution environment quota increase through AWS Service Quotas and deploy the sample application from the GitHub repository to see the improvement firsthand.

How Property Finder automated incident management with AWS DevOps Agent

Post Syndicated from Nada Tlohi original https://aws.amazon.com/blogs/devops/how-property-finder-automated-incident-management-with-aws-devops-agent/

When a production service starts saturating the CPU at 1 AM, every minute counts for incident management. For Property Finder, a production incident could mean failed searches, frustrated users, and direct revenue impact. Property Finder is the leading property portal in the Middle East and North Africa (MENA), serving millions of property seekers across five markets.

Before adopting AWS DevOps Agent, incident response followed a familiar pattern: an alert fires, an on-call engineer wakes up, spends 20–40 minutes correlating metrics across tools, manually documents findings, and opens a fix. Mean Time to Resolution stretched to 2–3 days for non-critical issues.

Today, that entire workflow runs autonomously. From alert to root cause analysis, Slack notification, Jira ticket, on-call phone call with context, and auto-remediation pull request (PR), the full lifecycle completes in 14 minutes. This post walks through the implementation and shows how a separate custom agent that automatically generates code fixes is the key differentiator.

The business problem

Property Finder runs a distributed microservices architecture on Amazon Elastic Container Service (Amazon ECS) fronted by Application Load Balancers (ALBs). When infrastructure issues occur, the impact is immediate: users see failed searches, agents cannot update listings, and revenue is directly impacted during peak hours.

The traditional workflow had three gaps:

  1. Detection lag. Non-critical anomalies could go undetected for days.
  2. Context switching. Engineers bounced between five or more tools per incident.
  3. Knowledge silos. Runbooks lived in people’s heads, not automation.

Solution architecture

Property Finder’s implementation connects AWS DevOps Agent at the center of a three-tier pipeline: Detection and Trigger, Autonomous Investigation, and Event-Driven Output.

Three-tier incident pipeline from a CloudWatch alarm through AWS DevOps Agent investigation to Slack, Jira, and GitHub outputs

Figure 1: End-to-end autonomous incident management architecture

The numbered steps correspond to the data flow in Figure 1:

  1. ECS CPU spike triggers an Amazon CloudWatch Alarm. CloudWatch Metrics Insights monitors service health across all ECS clusters. When sustained CPU exceeds 98%, the alarm transitions to ALARM state.
  2. AWS Lambda formats and HMAC-signs the payload. Triggered directly by the CloudWatch alarm action (which fires only on ALARM state transitions), AWS Lambda enriches the payload with service metadata, signs it with HMAC-SHA256 using credentials from AWS Secrets Manager, and POSTs to the webhook.
  3. The agent begins autonomous investigation. Parallel subagents query ECS metrics, AWS CloudTrail, ALB traffic patterns, and Grafana telemetry (Prometheus, Loki, Pyroscope). The agent reads relevant source code from GitHub for correlation.
  4. Findings post to Slack in real time. The native Slack integration posts investigation progress to #incidents. The full root cause analysis, impact assessment, and mitigation plan appear at the end of the thread.
  5. Investigation Completed event fires to Amazon EventBridge. Amazon EventBridge triggers an orchestrator Lambda that fans out to three independent targets simultaneously.
  6. Lambda creates a Jira ticket with the full root cause analysis. The Lambda retrieves the investigation summary from journal records and creates a prioritized ticket with root cause, severity, and affected service.
  7. Grafana IRM pages the on-call engineer by phone. A Lambda posts a Grafana Alerting-compatible payload to the IRM webhook. The escalation chain calls the engineer with full investigation context: what broke, why, and the recommended fix.
  8. The remediation agent opens a GitHub PR with the auto-fix. It receives the root cause, generates a Terraform or code fix, and opens a Draft PR through a GitHub Model Context Protocol (MCP) server. Engineers review before merging.

A real incident

The example-service, Property Finder’s core property search microservice serving millions of queries per day across five MENA markets, experienced CPU saturation at 99.11%. The pipeline resolved it end-to-end in 14 minutes.

1:21 AM │ Alarm fires (ECS CPU > 98%)

1:22 AM │ Investigation starts + Slack posted

1:22 AM │ 4 parallel subagents launched

1:32 AM │ Root cause identified

1:33 AM │ Jira ticket [redacted] created

1:34 AM │ On-call paged via phone call

1:35 AM │ GitHub PR [redacted] opened with fix

The detection Lambda handles three tasks: (1) retrieves the webhook secret from AWS Secrets Manager, (2) enriches the CloudWatch alarm event with ECS service metadata (cluster name, service name, task count), and (3) HMAC-signs the payload before POSTing to the webhook. The key authentication pattern:

# HMAC-SHA256 signing for webhook authentication
ts = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%S.000Z")
body = json.dumps(incident)
sig = hmac.new(webhook_secret.encode("utf-8"),
               f"{ts}:{body}".encode(), hashlib.sha256).digest()
http.request("POST", webhook_url, body=body,
             headers={"x-amzn-event-timestamp": ts,
                      "x-amzn-event-signature": base64.b64encode(sig).decode()})
Four parallel subagents querying ECS, CloudTrail, ALB, and Grafana data sources during the investigation

Figure 2: Four parallel subagents investigating ECS, CloudTrail, ALB, and Grafana data sources simultaneously

Root cause: Conflicting CPU and memory target-tracking autoscaling policies combined with an insufficient capacity floor. The service had both a CPU policy (target 70%) and a memory policy (target 75%). Actual memory usage sat at 3–8%, creating a persistent conflict between the two policies.

With MinCapacity set too low, the service could not sustain the task count needed to absorb CPU load. The resulting instability (22+ scaling flips observed) prevented stable scale-out, leaving the service effectively pinned at two tasks with no CPU headroom.

This is a common organizational issue: teams configure both scaling dimensions without realizing the interaction, especially when the capacity floor is not sized for baseline traffic. The agent identified the pattern in 10 minutes, a task that typically requires senior engineers with deep scaling expertise and hours of CloudWatch metric correlation.

Investigation output naming conflicting CPU and memory autoscaling policies as the root cause

Figure 3: Root cause analysis identifying the conflicting autoscaling policy

Slack incidents channel message linking to the running investigation at 1:22 AM

Figure 4: Slack notification with investigation link posted at 1:22 AM

Auto-created Jira ticket showing priority, root cause, and affected service

Figure 5: Jira ticket [redacted] auto-created with priority, root cause, and affected service

At 1:34 AM, the on-call engineer received a phone call through Grafana IRM with the complete investigation context. No need to wake up and hunt for root cause across dashboards.

Grafana IRM escalation chain routing the alert to the on-call engineer

Figure 6: Grafana IRM escalation chain routing the alert and calling the on-call engineer

Incoming on-call phone call at 1:34 AM carrying the investigation context

Figure 7: Incoming phone call at 1:34 AM with investigation context

Mitigation plan generated: (1) Remove the memory-based scaling policy, (2) raise MinCapacity to handle baseline traffic, (3) implement CPU-only target tracking at 70%. This plan was passed to a separate custom agent for remediation.

Remediation

Remediation is the key differentiator in this pipeline. It is a dedicated remediation agent (pr-creation-agent) invoked only after investigation completes. AWS DevOps Agent enforces read-only access to infrastructure through a per-session permission guardrail. Effective permissions are the intersection of the execution role’s IAM policy and the guardrail, and write actions are excluded.

The split separates concerns: investigation stays within that read-only envelope, whereas the remediation agent is scoped to a GitHub MCP server as its only external integration. Safety at the remediation layer does not rely on the agent’s built-in directed actions approval mechanism. Instead, two controls enforce the boundary. First, the remediation agent is a separate, narrowly scoped agent with access limited to GitHub MCP. Second, every output is a Draft pull request that requires human review and merge before taking effect. The GitHub MCP connection is authenticated with a fine-grained personal access token scoped to the specific infrastructure repositories, with an expiration and rotation policy. No elevated IAM role or additional agent permissions are required.

How it works

When the “Investigation Completed” Amazon EventBridge event fires, a Lambda orchestrator invokes the remediation agent with the investigation ID. The agent then:

  1. Reads findings from journal records to understand the root cause and recommended fix.
  2. Maps the AWS account to the correct repository. Property Finder has six infrastructure repos for different teams (B2B, B2C, core-platform, growth, data-engineering, shared-infra). The agent extracts the account ID from resource ARNs and routes to the right repo. This mapping is validated through automated tests and updated as new accounts or repositories are onboarded.
  3. Checks for duplicate PRs by searching existing PR titles and bodies for the investigation ID. If a matching PR exists, it reports the URL and exits without creating a duplicate.
  4. Reads the relevant Terraform files through GitHub MCP (GITHUB-MCP_get_file_contents), identifies the exact changes required, and plans the fix.
  5. Creates a feature branch (fix/{investigation_id}), commits the changes, and opens a Draft PR with a structured template including problem summary, root cause, changes made, and a testing checklist.

AWS also supports remediation through Kiro CLI with AWS CodeBuild or Kiro-ready prompts. Property Finder chose an approach that fits their multi-team repository structure: the remediation agent runs entirely within the Agent Space (the managed environment where custom agents execute), uses GitHub MCP for repository access, and maps multiple repositories to different teams automatically.

The orchestrator Lambda is triggered by the “Investigation Completed” Amazon EventBridge event. It first retrieves the investigation findings from journal records, then fans out to three targets simultaneously. Target one creates a Jira ticket with the full root cause analysis, severity, and affected service. Target two posts a Grafana Alerting-compatible payload to the Grafana IRM webhook to trigger phone call escalation. Target three invokes the remediation agent through the CreateChat and SendMessage API, passing the investigation ID and root cause context so it can generate the appropriate code fix.

Draft GitHub pull request with a problem summary, root cause, and changes template

Figure 8: GitHub PR [redacted] generated by the remediation agent with a structured problem, root cause, and changes template

Terraform diff replacing the memory scaling policy with a CPU-only target-tracking policy

Figure 9: Terraform diff showing the new CPU-only scaling policy replacing the conflicting memory configuration

The PR is always opened as Draft. Engineers review, run terraform plan, validate in staging, and merge. The agent never auto-merges.

Results

Metric Before After Improvement
End-to-end time Hours to days 14 minutes >88% reduction
Investigation 20 to 40 min (manual) 10 min (autonomous) 50–75% reduction
Documentation Manual, incomplete Auto-generated root cause analysis + Jira 100% documented
Remediation Manual PR by engineer Auto-fix PR + review Minutes to code fix

Cost considerations: Each incident invokes two agent sessions (investigation + remediation) with up to four parallel subagents. Billing is based on agent minutes. For detailed pricing, see the AWS DevOps Agent pricing page. We recommend reviewing pricing for all services used in this architecture.

“We now rely fully on AWS DevOps Agent to identify infrastructure-related issues. It has helped us identify multiple complex issues without even opening a support ticket. Even if we had raised tickets, it would likely have taken support engineers hours to find the root cause, whereas we resolved these issues in minutes.”

— Yasitha Bogamuwa, Cloud Engineering Manager, Property Finder

Getting started

Prerequisites:

  1. An Agent Space configured in your account.
  2. Amazon CloudWatch and AWS CloudTrail enabled for observability.
  3. Slack, Grafana, and GitHub connected as capabilities.
  4. Infrastructure resources tagged for topology mapping.

Step 1: Configure the webhook trigger. Set up CloudWatch Alarm action to invoke a Lambda function. The Lambda enriches the payload, HMAC-signs it, and POSTs to your Agent Space webhook endpoint.

Step 2: Set up event-driven outputs. Create an Amazon EventBridge rule for “Investigation Completed” events (source: aws.aidevops). Add Lambda targets for Jira, Grafana IRM, and optionally a remediation custom agent.

Step 3: Test end-to-end. Trigger a test alarm and verify the full pipeline: investigation starts, Slack posts, Jira ticket created, on-call paged, and PR opened.

For a similar integration pattern with Salesforce, see Automating Incident Investigation with AWS DevOps Agent and Salesforce MCP Server on the AWS DevOps Blog.

Clean up

This post describes an architecture pattern implemented by Property Finder. If you deployed test resources while following along, remember to delete any CloudWatch Alarms, Lambda functions, Amazon EventBridge rules, and Agent Space configurations to avoid ongoing charges. For a full list of resources and associated costs, review the pricing pages for each AWS service used in this architecture.

Conclusion

Property Finder’s implementation shows that autonomous incident management works in production today, with their pipeline running since early 2026. The agent never auto-merges. Human review remains in the loop by design: the agent accelerates, the engineer decides. The on-call engineer wakes up to a phone call with the root cause already identified, a Jira ticket filed, and a PR ready for review.

Explore the AWS DevOps Agent documentation to get started with your own autonomous pipeline.

  1. Getting Started with AWS DevOps Agent.
  2. Automating Incident Investigation with Salesforce MCP.
  3. Building an End-to-End Agentic SRE.
  4. Amazon EventBridge User Guide.
  5. Grafana IRM Documentation.

About the authors

Nada Tlohi

Nada Tlohi

Nada is a Technical Account Manager at AWS based in Dubai, UAE. She helps strategic enterprise customers across the MENA region transform their cloud operations and improve system reliability by adopting AIOps, incident automation, and DevOps best practices.

Conor Manton

Conor Manton

Conor is a Principal Technical Account Manager at AWS, based in San Francisco. He works with strategic enterprise customers to accelerate their cloud journey, with a focus to operationalize AI-powered workflows to drive business outcomes.

Jaydeep Singh

Jaydeep Singh

Jaydeep is a Senior DevOps Engineer at Property Finder. He specializes in designing and operating scalable cloud infrastructure, containerized platforms, and Kubernetes ecosystems. He leads platform reliability, infrastructure automation, and continuous integration and continuous delivery (CI/CD) initiatives, so engineering teams can build and deploy applications securely, efficiently, and at scale.

Git v2.56.0 released

Post Syndicated from jake original https://lwn.net/Articles/1097213/

Version 2.56 of the Git distributed
version-control system has been released. It has 748 non-merge commits
since Git 2.55 was released back in
June; those commits came from 104 developers, 39 of whom are first-time
contributors. New features include a safer workflow for conflict
resolution, smaller path-walk repacks, a new git history drop
sub-command, and much more. LWN looked at Git
2.56
recently and the GitHub blog has a lengthy
look at 2.56
as well.

The collective thoughts of the interwebz