This post provides step-by-step procedures to deploy Oracle Database on Amazon Elastic VMware Service (Amazon EVS) with Amazon FSx for NetApp ONTAP as NFS datastore storage. You will provision storage volumes, mount NFS datastores, install Oracle, and configure SnapMirror replication for cross-region disaster recovery.
This post picks up where the architecture left off. We provide step-by-step procedures to deploy the entire stack — from provisioning your first Oracle VM and creating FSx for ONTAP volumes, through mounting NFS datastores, installing Oracle 19c, and configuring SnapMirror and SnapCenter. We also cover four migration paths for moving existing on-premises Oracle workloads to EVS and day-2 operations including snapshot backup, point-in-time recovery, and database cloning.
This section walks through the end-to-end deployment workflow: provisioning the Oracle VM, creating FSx for ONTAP storage volumes, mounting NFS datastores in vSphere, installing Oracle 19c, and configuring SnapMirror for cross-region DR. Complete the prerequisites first, then follow Steps 1 through 8 in order.
Prerequisites
Amazon EVS environment deployed (VCF 9.1, ESXi 9.1. With VCF 9.x, Amazon EVS provisions the bare metal infrastructure and VLANs and you deploy VCF using Broadcom’s VCF Installer).
Minimum 4 hosts per cluster. Dedicated DB cluster recommended for Oracle.
Log in to the vSphere Client connected to your EVS vCenter.
Create a new Virtual Machine in the Production DB Cluster:
Guest OS: Red Hat Enterprise Linux 8 (64-bit) or Oracle Linux 8.
vCPU: Size per Oracle workload (8–32 vCPU typical)
Memory: 32–256 GiB based on SGA/PGA requirements.
Disk: 100 GiB on vSAN datastore (OS + Oracle Home + swap + temp tablespace)
Network: Attach to DB segment on the prod-trusted Tier-1 gateway.
Power on the VM and configure the guest OS:
# Set hostname
hostnamectl set-hostname ora-db1
# Create swap on vSAN-backed disk (local NVMe, single-digit ms latency)
# The vSAN datastore is already available as the VM's primary disk
# Allocate a dedicated partition or LV for swap
lvcreate -L 16G -n swap vgos
mkswap /dev/vgos/swap
swapon /dev/vgos/swap
echo "/dev/vgos/swap swap swap defaults 0 0" >> /etc/fstab
# Install Oracle prerequisites
sudo yum install -y oracle-database-preinstall-19c python3
For HA/DR, consider replicating the entire Oracle VM rather than maintaining a separate licensed instance in the DR cluster (see DR Licensing Consideration in the architecture post).
Oracle licensing consideration: Instance type selection affects Oracle license cost, which is an important factor to take into consideration. We recommend requesting an AWS Optimization and Licensing Assessment (AWS OLA) for further guidelines.
Step 2: Provision FSx for NetApp ONTAP
Open the Amazon FSx console, select Create file system, then select Amazon FSx for NetApp ONTAP.
Select Standard create and configure:
Setting
Value
Deployment type
Single-AZ (required for EVS)
SSD storage capacity
Size for Oracle data + logs + 20% headroom
Throughput capacity
512–2,048 MB/s (size for write workload. See asymmetry note)
IOPS
Automatic (3/GiB) or user-provisioned (up to 80,000)
VPC
Same VPC as EVS environment
Subnet
EVS service access subnet, same AZ as DB cluster
Security group
Allow NFS (TCP 2049, 111, 635) from EVS management VLAN
Set the fsxadmin password (required for ONTAP CLI automation).
Create an SVM (Storage Virtual Machine) with vsadmin password.
Disable automatic daily backups. Use SnapCenter for Oracle-aware scheduling instead.
After creation, select SVM, select Endpoints, and copy the NFS DNS name.
Step 3: Create Oracle database volumes
Connect to the FSx for ONTAP cluster using SSH (ssh fsxadmin@management-endpoint) and create volumes:
Set -tiering-policy none to pin all data to SSD tier. Size the log volume for 24 hours of archive logs.
Step 4: Mount FSx for ONTAP as NFS datastore in vSphere
In vSphere Client, select the DB Cluster, then select Configure > Storage > New Datastore.
Select NFS, then select NFS 3.
Enter:
Server:svm-id.fs-id.fsx.region.amazonaws.com
Folder:/oradb1data
Datastore name:fsx-ora-db1-data
Repeat for binary (/oradb1bin) and log (/oradb1log) volumes.
Verify all three datastores show correct capacity in the cluster storage view.
Step 5: Create Oracle VMDKs on FSx for ONTAP datastores
With the NFS datastores mounted at the ESXi host level (Step 4), create virtual disks for Oracle on these datastores:
In vSphere Client, select the Oracle VM, select Edit Settings, then select Add New Device > Hard Disk.
Create these VMDKs:
VMDK
Datastore
Size
Guest Mount
Purpose
Hard Disk 2
fsx-ora-db1-data
500 GiB
/u02
Oracle data files
Hard Disk 3
fsx-ora-db1-log
250 GiB
/u03
Oracle redo + archive logs
Hard Disk 4
fsx-ora-db1-bin
50 GiB
/u01
Oracle Home binaries
Select Thick Provision, Eager Zeroed for data and log VMDKs (best Oracle performance).
Oracle Database can be created on Oracle ASM or Filesystem (local/NFS). This installation is based on creating the Oracle database on local XFS filesystem.
Inside the Oracle VM guest OS, partition and mount the new disks:
# Identify new disks
lsblk
# Create filesystem on each disk (example: /dev/sdb for data)
mkfs.xfs /dev/sdb
mkfs.xfs /dev/sdc
mkfs.xfs /dev/sdd
# Create mount points
mkdir -p /u01 /u02 /u03
# Mount
mount /dev/sdd /u01 # Oracle Home (binaries)
mount /dev/sdb /u02 # Oracle data files
mount /dev/sdc /u03 # Oracle redo + archive logs
# Persist in /etc/fstab
cat >> /etc/fstab <<EOF
/dev/sdb /u02 xfs defaults,noatime 0 0
/dev/sdc /u03 xfs defaults,noatime 0 0
/dev/sdd /u01 xfs defaults,noatime 0 0
EOF
# Set ownership
chown -R oracle:oinstall /u01 /u02 /u03
Key insight: The Oracle VM accesses /u02 and /u03 as local XFS block devices. It has no awareness that the underlying storage is an NFS datastore backed by FSx for ONTAP. All NFS communication happens at the ESXi host level, where each host uses its own network path to FSx for ONTAP.
Step 6: Install and configure Oracle 19c
# As oracle user
export ORACLE_HOME=/u01/app/oracle/product/19.0.0/dbhome_1
cd $ORACLE_HOME
./runInstaller -silent -responseFile /path/to/db_install.rsp
Create the database with data on /u02 and logs on /u03:
CREATE DATABASE orcl
DATAFILE '/u02/oradata/orcl/system01.dbf' SIZE 1G
LOGFILE
GROUP 1 '/u03/oralogs/orcl/redo01.log' SIZE 512M,
GROUP 2 '/u03/oralogs/orcl/redo02.log' SIZE 512M,
GROUP 3 '/u03/oralogs/orcl/redo03.log' SIZE 512M;
Oracle accesses /u02 and /u03 as local XFS filesystems. Standard Oracle ASM or filesystem-based storage management applies. No NFS-specific Oracle configuration is needed because the NFS layer is abstracted by the ESXi hypervisor.
Important: Consider Oracle licensing requirements when planning your DR strategy. An alternative is to replicate the Oracle VM through NetApp SnapMirror from Production to DR. Keep the DR replicated volumes as data-protection (DP) volumes that are NOT mounted as NFS datastores on DR hosts until a failover event is declared to avoid Oracle double licensing. Only then break the SnapMirror, mount the NFS datastore on the DR Host, and power on the VM. Pre-mounting the SnapMirror volume as a datastore — even with no VM powered on — means Oracle binaries are accessible on those hosts, which Oracle may consider an “installation” requiring licenses across the entire DR cluster. We recommend requesting an AWS Optimization and Licensing Assessment (AWS OLA) for further guidelines.
Step 8: Configure SnapCenter backup
Deploy SnapCenter Server (or use SnapCenter SaaS).
Add FSx for ONTAP storage system using the cluster management IP.
Install SnapCenter Plugin for Oracle on each Oracle VM.
Create backup policies:
Policy
Scope
Frequency
SnapMirror Update
Full DB Backup
Data + Control + Archive
Every 4–6 hours
Yes
Archive Log
Archive logs only
Every 10–15 minutes
Yes
Create resource groups, assign policies, and schedule.
Database migration from on-premises VMware to EVS
Option 1: VMware HCX live migration
For enterprises with existing VMware on-premises:
Deploy HCX Connector on-premises, HCX Cloud Manager on EVS.
Create site pairing and network extensions (L2 stretch).
Migrate Oracle VMs using HCX vMotion (zero downtime) or Bulk Migration.
Post-migration: storage vMotion Oracle VMDKs from vSAN to FSx for ONTAP NFS datastores for snapshot/replication capabilities.
Option 2: SnapMirror ONTAP-to-ONTAP
If on-premises Oracle already uses NetApp ONTAP storage:
Establish SnapMirror between on-premises ONTAP and AWS FSx for ONTAP.
Incrementally replicate until cutover.
At switchover: quiesce Oracle, flush archive logs, final SnapMirror sync, break mirror.
Mount FSx for ONTAP volumes on EVS Oracle VM, recover database, open for service.
Option 3: Oracle PDB relocation (multitenant)
For Oracle databases already in PDB/CDB multitenant model:
Create target CDB on EVS with FSx for ONTAP storage.
Use PDB hot clone to relocate PDBs from on-premises CDB to AWS CDB.
Minimal service interruption. Only final switchover requires brief outage.
Provision Oracle VM on EVS, mount FSx for ONTAP volumes.
Restore from RMAN backup, apply archive logs.
Open database and redirect applications.
Day-2 operations
Snapshot backup
SnapCenter manages full database snapshots as storage-layer operations, providing efficient backup capabilities.
Point-in-time recovery
In SnapCenter, select the SCN or timestamp, mount the log snapshot, restore the data snapshot, apply archive logs, and open with RESETLOGS.
Database cloning
SnapCenter FlexClone creates space-efficient database copies. Clones share unchanged blocks with the source and consume storage only for deltas. Use for dev/test, patch validation, and reporting.
HA failover procedure
Break SnapMirror on DR volumes.
Mount SnapMirror volumes as NFS datastores on DR ESXi hosts, then power on the Oracle VM.
Recover to last available archive log.
Open database. Update DNS/connection strings.
Clean up
To stop incurring charges after testing this deployment, remove the following resources in this order:
Oracle VMs — Power off and delete Oracle database VMs from the vSphere inventory.
NFS datastores — Unmount FSx for ONTAP datastores from ESXi hosts in vSphere.
SnapMirror relationships — Delete SnapMirror relationships and DP volumes on the DR FSx for ONTAP file system.
FSx for ONTAP file systems — Delete both production and DR file systems from the Amazon FSx console. This action deletes all volumes and data on those file systems.
Amazon EVS environment — Delete the EVS environment from the Amazon EVS console. This terminates the underlying EC2 bare metal instances.
Networking — Remove Transit Gateway attachments, VPC Route Server configurations, and Direct Connect connections if they were created solely for this deployment.
Important: Deleting an FSx for ONTAP file system permanently removes all data. Confirm that you have backed up any data you need before proceeding.
Conclusion
In this post, we walked through deploying Oracle Database on Amazon Elastic VMware Service (Amazon EVS) with Amazon FSx for NetApp ONTAP as NFS datastore storage. You provisioned the Oracle VM, created FSx for ONTAP volumes, mounted NFS datastores in vSphere, installed Oracle 19c, and configured SnapMirror for cross-region disaster recovery and SnapCenter for Oracle-aware backup. We also covered four migration paths for existing on-premises Oracle workloads and day-2 operations for backup, point-in-time recovery, and cloning.
Enterprises with existing VMware Cloud Foundation (VCF) investments want to migrate their Oracle databases to AWS without rearchitecting applications or retraining operations teams. Oracle on Amazon Elastic VMware Service (Amazon EVS) can take advantage of sub-millisecond storage latency, snapshot-based backup, cross-region disaster recovery (DR), and independent storage scaling, all while preserving existing VMware operational workflows.
In this post, we show you how to architect a complete Oracle Database environment on Amazon EVS with Amazon FSx for NetApp ONTAP. You learn how the storage, compute, networking, and disaster recovery layers work together to deliver high availability, cross-region DR, and sub-millisecond storage latency while maintaining your familiar VMware operational tooling.
In this post, you learn how to:
Design Oracle Database architecture on Amazon EVS running VMware Cloud Foundation 9.1.
Select optimal EC2 bare metal instance types and VM sizing for Oracle workloads.
Architect storage using Amazon FSx for NetApp ONTAP as NFS datastores for Oracle data and log volumes.
Plan SnapMirror replication for cross-region disaster recovery.
Evaluate migration options for moving existing Oracle workloads from on-premises VMware to EVS.
Amazon EVS directly runs VMware Cloud Foundation (VCF) environments on Amazon Elastic Compute Cloud (Amazon EC2) bare metal instances within an Amazon Virtual Private Cloud (Amazon VPC). With VCF 9.x, Amazon EVS provisions the bare metal infrastructure and VLAN subnets, and you then deploy VCF using Broadcom’s VCF Installer (the self-deployed model). VCF 9.x also supports evaluation mode, so you can validate the design before applying license keys, and the Solutions for Amazon EVS GitHub repository provides CloudFormation and Terraform templates to automate the phased VCF 9 deployment. Amazon FSx for NetApp ONTAP provides managed ONTAP storage with NFS, SMB, iSCSI, and NVMe over TCP access. FSx for NetApp ONTAP can deliver sub-millisecond response times, multiple GBps of throughput, and up to 80,000 IOPS per file system (see FSx for ONTAP performance).
This section describes a highly available Oracle Database deployment on Amazon EVS with Amazon FSx for NetApp ONTAP storage across two AWS Regions. Figure 1 illustrates the numbered data flow from on-premises through hybrid connectivity into the production EVS environment and cross-region DR.
Figure 1: Oracle Database on Amazon EVS with FSx for NetApp ONTAP and cross-region SnapMirror DR
The architecture consists of eight functional components, described in the following section.
Architecture components
These numbered items correspond to the data flow shown in the diagram.
Hybrid link — On-premises data center connects to AWS through AWS Direct Connect for dedicated, low-latency bandwidth.
Transit routing — AWS Transit Gateway routes traffic between the production VPC, DR region, and on-premises networks.
Workload landing — Traffic reaches the EVS Database Cluster running on i7i.metal-24xl Amazon EC2 bare metal instances with VCF 9.1.
Live migration — VMware HCX (Hybrid Cloud Extension) migrates Oracle VMs from on-premises VMware to EVS with near-zero downtime using Replication Assisted vMotion or bulk migration.
NFS data path — Each ESXi host mounts Amazon FSx for NetApp ONTAP as an NFS datastore. Oracle VMs access database volumes as block devices (VMDKs on the NFS datastore) with sub-millisecond latency and up to 80,000 IOPS. Each host has its own independent NFS path to FSx for ONTAP, distributing bandwidth across the cluster.
Cross-region DR — SnapMirror asynchronously replicates FSx for ONTAP volumes (data, logs, binaries) to the DR region with configurable Recovery Point Objective (RPO).
Dynamic routing — NSX Tier-0 gateway peers through BGP with Amazon VPC Route Server. This is one-way BGP. Route Server listens for routes advertised by NSX and writes them to the VPC route table, but does not advertise VPC routes back to NSX.
DR failover — On failover, Transit Gateway routes to the DR region where pre-provisioned standby Oracle VMs mount the SnapMirror replica volumes.
Factors to consider for Oracle Database deployment on EVS
Before you begin deployment, evaluate your Oracle workload requirements against the available infrastructure options. The decisions you make for instance types, VM sizing, and storage architecture directly affect database performance, cost, and operational complexity.
Oracle licensing consideration: Instance type selection affects Oracle license cost, which is an important factor to take into consideration. We recommend requesting an AWS Optimization and Licensing Assessment (AWS OLA) for further guidelines.
EC2 instance type selection for ESXi hosts
Amazon EVS currently supports three bare metal instance types for ESXi hosts. The choice depends on whether you prioritize per-core compute speed (i7i) or per-host memory and storage density (i4i), and on how much you want to scale a single host vertically. For most new Oracle deployments, i7i is the better choice because Oracle query performance is sensitive to CPU instruction throughput and storage IO latency.
i7i.metal-24xl (recommended default): 96 vCPUs (48 cores), 768 GiB RAM. 5th Gen Intel Xeon delivers improved compute performance, critical for Oracle CPU-bound queries. 3rd Gen AWS Nitro SSDs provide enhanced real-time storage performance and lower IO latency for vSAN.
i7i.metal-48xl: Same 5th Gen Intel Xeon as the 24xl but double the capacity per host (192 vCPUs, 96 cores, 1,536 GiB RAM, up to 100 Gbps network and 60 Gbps Amazon Elastic Block Store (Amazon EBS) bandwidth). Choose this when you want to scale a single Oracle host vertically, such as for large SGAs, high core counts, or fewer and denser hosts to reduce VMware per-host licensing. It keeps the same per-core performance profile as the 24xl.
VCF compatibility: All three instance types support VCF 9.1 with ESXi 9.1 (build 9.1.0.0100.25433460). VCF 5.2.2 (ESXi 8.0U3g) remains available but is on a path to end of support, so new deployments should default to 9.x. An EVS environment supports 4–32 hosts per cluster. You can mix instance types across clusters within the same SDDC.
VM sizing for Oracle Database guests
Size Oracle Database VMs based on these workload characteristics:
Size memory for SGA + PGA + OS overhead (typically 75–85% of allocated VM memory for SGA)
Configure VM swap on vSAN datastore (not NFS). vSAN uses local NVMe with single-digit millisecond latency.
Use NUMA-aware VM placement for VMs exceeding single-socket core count.
Place OS swap and Oracle temp tablespace on the vSAN datastore for single digit ms latency at no additional cost.
Storage architecture: vSAN + FSx for NetApp ONTAP
The recommended design splits storage responsibilities between two tiers:
vSAN (backed by local NVMe drives on each ESXi host) handles low-latency, non-replicated workloads: VM boot disks, OS swap, and Oracle temp tablespace.
FSx for NetApp ONTAP handles Oracle data files and redo logs that require snapshot-based backup and cross-region replication. ESXi hosts mount FSx for ONTAP volumes as NFS datastores, and Oracle VMs access standard VMDKs on those datastores.
This separation gives local NVMe speed for transient IO while adding snapshot, clone, and SnapMirror capabilities for persistent database files.
Why NFS datastore (host-level) instead of in-guest NFS (dNFS)? With in-guest NFS, all Oracle NFS traffic routes through the NSX overlay and an NSX Edge node before reaching FSx for ONTAP. This creates a single Edge chokepoint. With NFS datastores, each ESXi host talks NFS directly to FSx for ONTAP using its own network bandwidth. There is no Edge bottleneck and no extra latency hop.
FSx for NetApp ONTAP sizing considerations
Important: Use 100% SSD for Oracle. We recommend against using capacity pool tiering for Oracle database volumes. Keep tiering policy set to none for all Oracle volumes.
Important: SSD capacity planning. If the SSD tier fills to capacity, FSx for ONTAP blocks writes. Monitor SSD utilization and provision headroom (minimum 20% free).
Important: Read/write throughput asymmetry. On a 6 GB/s filesystem, read throughput can reach 6 GB/s, but write throughput is limited to approximately 1 GB/s. Size throughput capacity based on Oracle write workload requirements. This write ceiling applies per high-availability (HA) pair. To scale write throughput beyond a single HA pair, deploy a file system with more than one HA pair and distribute Oracle volumes across the additional aggregates so writes are spread across pairs. This placement is not automatic, so plan the layout up front to keep write-heavy datasets balanced.
Network architecture
Amazon EVS uses VLAN subnets (defined at environment creation and unchangeable later) to segment traffic.
VPC Route Server replaces static routes within the VPC. NSX Tier-0 gateways peer through BGP with Route Server endpoints. This is one-way BGP: Route Server listens for routes from NSX and programs them into the VPC route table but does not advertise VPC routes back to NSX. Beyond the VPC, routes remain static at Transit Gateway.
NSX-T segmentation for Oracle
Dedicated Tier-1 gateway for production database segments (DB subnets)
Separate Tier-1 for application tier (App subnets) and perimeter network.
Micro-segmentation between Oracle instances prevents lateral movement.
Important: Security group rules are not enforced on VLAN subnet interfaces. Use network ACLs and NSX distributed firewall for traffic control.
High availability and disaster recovery
SnapMirror replication frequency determines RPO. Configure based on business requirements.
Pre-provision standby Oracle VMs in the DR cluster to reduce Recovery Time Objective (RTO).
Replicate binary volumes so that Oracle installation is not required during recovery.
Automate failover with Ansible/SnapCenter to reduce human error.
Oracle licensing consideration for DR: There are license impacts based on how DR replication is implemented. If you have licensing questions, we recommend requesting an AWS Optimization and Licensing Assessment (AWS OLA).
To comply with the Oracle licensing rules, an alternative is to replicate the Oracle VM through NetApp SnapMirror from Production to DR. Keep the DR replicated volumes as data-protection (DP) volumes that are NOT mounted as NFS datastores on DR hosts until a failover event is declared to avoid Oracle double licensing. Only then break the SnapMirror, mount the NFS datastore on the DR Host, and power on the VM. Pre-mounting the SnapMirror volume as a datastore, even with no VM powered on, means Oracle binaries are accessible on those hosts, which Oracle may consider an “installation” requiring licenses across the entire DR cluster.
Database migration from on-premises VMware to EVS
The following table compares migration options from on-premises VMware to EVS.
Security for Oracle on Amazon EVS spans multiple layers from network isolation to database-level encryption. The following table summarizes the security controls across each layer.
VPC encryption for NFS traffic, and IPsec for SnapMirror cross-region
Database encryption
Oracle TDE (Transparent Data Encryption) for additional protection
Administrative access
Zero-trust access (for example, Banyan or Zscaler) for VMware admin consoles
VLAN subnet security
Network ACLs (security groups not enforced on VLAN interfaces)
Cost optimization
Cost optimization for Oracle on Amazon EVS focuses on matching infrastructure capacity to workload demands and using AWS pricing models. The following table summarizes key strategies across compute, storage, networking, and Oracle licensing. We recommend requesting an AWS Optimization and Licensing Assessment (AWS OLA) for further guidelines.
i7i.metal-24xl delivers ~10% price-performance over i4i.metal
FSx for ONTAP throughput
Right-size for the write workload, and adjust on the fly
FSx for ONTAP storage
100% SSD for Oracle, and storage efficiency for non-production
Data transfer
Place FSx for ONTAP in same AZ as EVS cluster
SnapMirror replication
Schedule frequency based on RPO (less frequent = lower cost)
Oracle DR licensing
Replicate the Oracle VMs through NetApp SnapMirror from Production to DR, keeping the DR replicated volumes as data-protection (DP) volumes that are NOT mounted as NFS datastores on DR hosts until a failover event is declared to avoid Oracle double licensing. When a failover event is declared, then break the SnapMirror, mount the NFS datastore on the DR Host, and power on the VM. We recommend requesting an AWS Optimization and Licensing Assessment (AWS OLA) for further guidelines.
Non-production
Use fewer hosts, and apply tiering for dev/test data
Summary
Deploying Oracle databases on Amazon EVS with Amazon FSx for NetApp ONTAP provides high availability, cross-region DR, and sub-millisecond storage latency while combining VMware operational consistency with AWS cloud economics:
Performance: i7i.metal-24xl delivers up to 23% better compute and 50% lower IO latency. FSx for ONTAP delivers sub-millisecond latency with up to 80,000 IOPS. For details, see Amazon EC2 i7i instances in the AWS GovCloud (US) Regions.
Availability: vSphere HA, SnapMirror cross-region replication, and optional Oracle Data Guard.
Manageability: SnapCenter for backup, clone, and recovery in seconds regardless of database size.
Security: NSX micro-segmentation, AWS KMS encryption, and zero-trust access.
Cost efficiency: Improved price performance with i7i, on-the-fly throughput adjustment, and AWS Savings Plans.
This architecture provides you with high availability, cross-region DR, and storage-based backup and cloning similar to Oracle RAC and Data Guard functions while maintaining your familiar VMware operational tooling and procedures.
Next steps
To get started with this deployment:
Provision an Amazon EVS environment in your target Region. See the Amazon EVS User Guide for setup instructions.
Deploy an Amazon FSx for NetApp ONTAP file system in the same VPC and Availability Zone as your EVS cluster.
Follow the step-by-step procedures in our companion post to mount NFS datastores, create Oracle VMDKs, and configure SnapMirror DR.
Test in a non-production environment first, then migrate production Oracle workloads.
When you build multi-Region architectures on AWS, one question comes up early: “What services are available in each AWS Region?” The answer shapes architecture decisions, from which AWS Regions to expand into, to how you design for resilience. For teams navigating data residency requirements and compliance reporting, getting the answer wrong has real consequences: deployment failures, compliance gaps, and delayed launches.
AWS publishes Regional availability data covering services, features, APIs, and AWS CloudFormation resource types across all AWS Regions. You can explore this data on the AWS Capabilities by Region page, access it through Amazon Simple Storage Service (Amazon S3) for pipeline integration, or query it using the AWS Knowledge MCP server. These options work well for exploration and automation. However, teams told us they need this data deployed as infrastructure they own, refreshing on their schedule, inside their network, filtered to their workload. That means running availability data the same way you run the rest of your stack: in your Amazon Virtual Private Cloud (Amazon VPC), under your governance.
In this post, we introduce two open source solutions that deliver that ownership. Capability Insights for AWS deploys a Regional availability dashboard into your VPC that auto-refreshes every 24 hours. The dashboard, its API, and availability data all run in your own account, and the scheduled refresh makes the only call that leaves your account when it reads the dataset that AWS publishes. Workload Analysis scans your AWS CloudTrail logs and CloudFormation stacks, then narrows 200+ services to the 20–30 your account actually runs, significantly reducing the scope of a Regional gap analysis.
Together with the Capabilities by Region S3 Access Point, you now have three levels of control: consume from Amazon S3, deploy into your VPC, or filter to your workload. We walk through how to deploy each solution, run a workload analysis, and view personalized results during your Regional expansion planning. Whether you’re building multi-Region recovery strategies, standardizing compliance reporting, or accelerating expansion timelines, these solutions put the data and the decisions it drives inside your perimeter.
Figure 1 shows how Capability Insights for AWS integrates with the VPC and subnets in your existing AWS account. A client in the public subnet accesses the dashboard through an S3 gateway endpoint and calls the private API through an API Gateway VPC endpoint.
Figure 1: Capability Insights for AWS architecture, including private dashboard access and scheduled Regional availability data refreshes.
An Amazon EventBridge schedule invokes the data-fetch Lambda function every 24 hours. The data-fetch function runs outside your VPC, so it reads the S3 access point published by AWS over the AWS network. The function pulls Regional availability data from the Capabilities by Region S3 bucket published by AWS and writes it to the website bucket within your account. The dashboard application accesses the data within the website through the S3 gateway endpoint from your VPC. The API Lambda can also invoke the data-fetch function on demand through the Lambda VPC endpoint. The deployment assets bucket supplies the Lambda code during deployment only.
Prerequisites
To follow along, you need:
An AWS account with permissions to deploy the CloudFormation stacks, including permission to create named AWS Identity and Access Management (IAM) roles. At runtime, the roles created by the stacks use scoped permissions for Amazon S3 object read and write access, AWS CloudFormation read operations, Amazon Athena queries, AWS Glue and AWS Lake Formation catalog access, AWS Lambda invocation, and AWS Step Functions execution. For starting-point deployment policies, see the documentation folder in the repository.
A VPC with DNS resolution and DNS hostnames enabled, one public subnet (routed to an internet gateway) for dashboard and API access, and one private subnet (no internet route) for the in-VPC Lambda function. Both subnets need a route to Amazon S3 through a gateway VPC endpoint.
Node.js v24.18.0 (includes npm and npx) for automated deployment.
An S3 bucket for deployment assets.
An active CloudTrail configuration where you store logs in an S3 bucket (required for Part 2).
Part 1: Deploy a self-hosted dashboard (Capability Insights for AWS)
Capability Insights for AWS deploys a searchable Regional availability dashboard into your own AWS account. The solution pulls data from the AWS Capabilities by Region S3 bucket, stores it inside your VPC, and serves it through a static website backed by Amazon API Gateway, AWS Lambda, and Amazon EventBridge.
If your organization has access to additional data sources beyond the public dataset, the solution incorporates those as well, giving you a unified view across all AWS partitions you have access to. To learn more about accessing additional data, work with your AWS representative.
The dashboard covers:
Services and features: availability status, expected launch dates, and expansion plans per Region.
API operations: individual API action availability per Region for each AWS service.
CloudFormation resource types: which resource types each Region supports.
You provide your own VPC, subnets, and S3 bucket so the solution integrates with your existing infrastructure and security controls.
Clone and install
git clone https://github.com/aws/capability-insights-for-aws.git
cd capability-insights-for-aws
npm install
Create a deployment assets bucket
Create an S3 bucket with public access blocked to store the Lambda code package during deployment. We recommend naming it capability-insights-assets-<ACCOUNT_ID>-<REGION>.
Deploy the stack
Automated deployment:
npm run deploy
The script builds all assets, prompts for parameters, deploys the CloudFormation stack, uploads the website, and triggers an initial data sync. You will be prompted for SourceFolders, a comma-separated list of data sources to pull from. The default is public.
To skip interactive prompts, pass all parameters as flags:
VPC ID with DNS resolution and DNS hostnames enabled
–backend-subnet-id
Private subnet with no internet route. Needs a route to Amazon S3 through a gateway VPC endpoint.
–api-access-subnet-id
Public subnet routed to an internet gateway (user access, API Gateway VPC endpoint)
–deployment-assets-bucket-name
S3 bucket for deployment assets
–source-access-point-arn
Public Capabilities by Region S3 access point ARN published by AWS. Account ID 686591367145 identifies the account managed by AWS that hosts the public dataset. Use the ARN as shown.
EC2 instance with SOCKS proxy: SSH into an instance in the VPC and proxy browser traffic through it.
Explore the dashboard
Once connected, the dashboard provides:
Search and filter: find services by name across all Regions.
Expandable service details: select any service to see individual feature availability per Region.
Status indicators: Available, Planning, Not Expanding, and projected dates (for example, “2026 Q3”)
Export: download the current view as JSON or CSV for sharing or further analysis.
Settings: view last sync time and trigger manual data refreshes.
Figure 2 shows the Capability Insights for AWS dashboard and its controls for comparing service availability across Regions.
Figure 2: Dashboard overview showing service availability across Regions with status indicators and export options.
Summary counts show coverage for services and features, API operations, CloudFormation resources, and Regions. Tabs switch between catalog views, while filtering, Region columns, status values, and the expand and export controls help you inspect and download availability data.
At this point you have a working dashboard inside your VPC, refreshing daily, with no external dependencies during normal operation.
Part 2 adds personalization by filtering this catalog to the specific services your account runs. You can view the workload analysis results through the dashboard or access them through dedicated API endpoints.
Part 2: Personalize the catalog with Workload Analysis
The catalog covers 200+ services across 35+ Regions. When you plan expansion to a new Region, you don’t need to evaluate all of them. You need to evaluate the 20–30 your account runs. Without that filter, a gap analysis is a multi-week project. With it, it’s a 30-minute review.
Workload Analysis deploys as an additive CloudFormation stack alongside the Capability Insights stack. The pipeline runs as an AWS Step Functions state machine with three stages:
Parallel analyzers (two branches run in parallel):
CloudTrail Analyzer: Creates an AWS Glue Data Catalog table pointing at your CloudTrail log bucket, then runs an Amazon Athena query that extracts distinct service/API/Region/account combinations from the last N days (configurable, default 30). This identifies services your account has called.
CloudFormation Analyzer: Calls ListStacks and GetTemplate for every active stack in the account. Extracts AWS:: resource types and scalar property values (strings, numbers, booleans), maps them to service names, and records which stack contributed each resource. This identifies what your account deploys, not only what it calls.
Usage Decorator runs after both analyzers complete. It reads the primary capability catalogs (products.json, apis.json, cfn_resources.json) from the website bucket, intersects them with the analyzer outputs, and writes personalized files back to the bucket for the dashboard to consume.
Scheduled execution through an Amazon EventBridge rule triggers the state machine daily (configurable schedule expression), keeping personalized data fresh without manual intervention.
Figure 3 shows how the Workload Analysis components coordinate scheduled and on-demand analysis.
An Amazon EventBridge schedule or an on-demand API request starts the AWS Step Functions state machine. The workflow runs two branches in parallel: the CloudTrail analyzer uses Amazon Athena with AWS Glue Data Catalog and AWS Lake Formation catalog access to query activity, while the CloudFormation analyzer inspects active stacks. After both branches complete, the Usage Decorator combines their results with the primary capability catalogs and writes personalized data to the Capability Insights website bucket for the dashboard and API to consume.
Deploy Workload Analysis
Verify that CloudTrail is active and your logs land in an S3 bucket. Then deploy with the usage analysis flag:
This deploys both the core Insights stack and the Usage Analysis stack, then wires them together so the API Lambda can trigger analysis and serve personalized results.
Additional flag
Description
--enable-usage-analysis
Enables the Workload Analysis pipeline
--cloudtrail-bucket
S3 bucket containing your CloudTrail logs
A typical first run completes in 2–5 minutes depending on account size and the number of active CloudFormation stacks.
Retrieve API endpoint
Use the following command to retrieve the API base URL from the API configuration file on the dashboard bucket:
Use apiBaseUrl as the base URL for the dedicated API endpoints. You must run API requests from within the VPC because the API Gateway endpoint is private.
Run an analysis
Start an analysis by calling POST /analysis through the API Gateway endpoint:
The Amazon EventBridge rule triggers subsequent runs daily. You can also trigger ad-hoc runs through the dashboard’s Settings page.
View personalized results
After an analysis completes, the dashboard toggles between the full catalog and your personalized My Stuff view, showing only services and resources your account uses.
Through the API:
GET /capabilities?usageFilter=combined&scope=account
The response contains filtered products, APIs, and CloudFormation resources with usage attribution, including which stacks deploy each resource type and which property configurations you use:
Consider an account running a typical web application. Without Workload Analysis, the dashboard shows all 200+ services across 35 Regions. With the combined filter applied:
Before: 200+ services, thousands of features, hundreds of CloudFormation resource types.
After: 28 services your account uses, with per-stack attribution showing which CloudFormation stacks deploy each resource type.
The CloudFormation resource view goes deeper. For each resource type your stacks deploy, you can see which property configurations are in use and which stacks contributed them. For example, your AWS::EC2::Instance resources might show InstanceType: t3.medium from your application stack and InstanceType: m5.xlarge from your data processing stack, each traced back to its source.
This targeted view means your Regional expansion gap analysis focuses on the 28 services you care about rather than the full catalog. For example, if you plan to deploy into the Europe (Zurich) Region (eu-central-2), the dashboard shows which 28 services are available there, which have planned launch dates, and which have no roadmap entry yet. These results highlight service-availability gaps to consider during Regional expansion. From here, you can evaluate your options: wait for planned launches, architect around unavailable features, or evaluate a different target Region. You’re not evaluating 200+ services against a Region matrix. You’re scanning a personalized list where every row is something your stacks deploy.
Clean up
To remove the resources created in this walkthrough:
Workload Analysis (if deployed): Delete the Usage Analysis stack first:
Optionally delete the deployment assets bucket you created during setup. This solution uses standard AWS service pricing for Lambda, S3, API Gateway, Athena, and Step Functions. There is no additional charge for the solution itself.
Automated cleanup: Run npm run teardown to remove both stacks and empty the website bucket in one step.
Conclusion
Knowing which services are available in your target Region is the first step in designing for resilience. You can’t build a multi-Region recovery strategy for services that aren’t there yet. With Capability Insights for AWS and Workload Analysis, you own that data inside your VPC, filtered to what your account runs, refreshing on your schedule.
Cloudflare Stream is a powerful broadcasting platform that, for many of our customers, just works. But what if you wanted to render dynamic annotations on a livestream or create an alternate version of a hosted video with burned-in subtitles? You would need to run a custom video pipeline.
Today, we’re releasing a new developer playground, Streamline, that demonstrates how you can build a system to deliver these bespoke video experiences on Cloudflare’s Developer Platform. We’ll walk you through how Streamline leverages Workers, Containers, and several media protocols to modify video — and immediately publish that output as livestream or new hosted video. You’ll also have the opportunity to try it for your projects.
A processing pipeline needs a durable, long-running environment that can run specialized, compiled code with predictable memory and CPU capacity. Video streams can run for minutes or hours, so the media process needs a lifecycle independent of the request that started it. An application should be able to start a pipeline, send its input, inspect it, and stop it without needing to keep a single request open for the entire duration.
Cloudflare provides the primitives we need. Containers are long-lived runtimes suitable for media processing. Durable Objects help with orchestration. Finally, Workers are perfect for control signaling and monitoring.
For Streamline, we built a media engine running in a Container to handle media processing in real-time. The Container is controlled by a Worker exposing control, preview, and testing to an agent or user. Processing will continue even if the Worker disconnects. We've architected Streamline with modular components so that the media engine could be replaced with dedicated encoding products in the future.
Architecture
A Streamline deployment consists of two components: the Media Engine, which handles media input/output and processing, and a controlling Application, which creates, configures, observes, and stops media sessions.
Media Engine
The Media Engine has two components:
Controller. This is a control harness written in Go that implements an HTTP server, receives incoming requests, and translates them into operations that can be executed by the media engine.
Processor that performs the actual media processing. The current implementation uses FFmpeg, but that is an internal implementation detail rather than part of the user-facing API.
The Media Engine is hosted in a Container, and handles all media input/output as well as processing. It can pull RTMPS playback over the network from one Stream Live input and publish RTMPS output to another Stream Live input. It can pull a Cloudflare Stream HLS manifest and its segments to use hosted videos as input. It can accept video input from a source supplied by the controlling application, for example a webcam. It can publish preview video over an outbound WebSocket to a Durable Object relay. An application that needs preview can connect to that relay through its own WebSocket.
Application
The application is built using Workers, and can be a full-stack browser application, an agent, or an embedded system. It consists of:
User interface (UI) including client logic, identity and access policy. This post uses a browser application as its concrete example, so it also includes a browser interface.
Orchestrator coordinates the session, the Container lifecycle, and preview relay. The orchestrator is implemented by a Durable Object.
It is possible to run the system locally during development, in which case the container is just a local Docker instance and the Durable Object is not used: there is a single user, the controlling application does not require authorization for local access, and the video preview can connect directly to a WebSocket on localhost.
When these components are deployed to Cloudflare, an authorized user or agent can visit the Worker to start a new session. This spins up a new Streamline container if needed, manages its lifecycle automatically, exposes an API to perform a number of video manipulation operations, and routes inputs from and outputs back to Cloudflare Stream.
Time for a technical deep dive on how the system works.
Container lifecycle and session management
The controlling Worker application initiates a long-running media processing session. After starting the session, the application can disconnect and reconnect safely, while the Container continues processing until the controlling application stops it. We also include a maximum duration to ensure a session is always eventually closed down and can’t run indefinitely, even without external control. While a media processing session is running, the container instance is unavailable for other applications to use.
A Cloudflare Container will automatically sleep if it has not received any incoming requests since a defined interval. However, in our case, once the pipeline is running, it must continue even if the controlling application disconnects and it receives no requests. We can implement this behavior by overriding the onActivityExpired() callback on the container. If the expiry time has not been reached, then we renew the activity, otherwise we destroy the container.
API
The HTTP server implemented by the Go harness and the Durable Object associated with the Container together define the low-level interface to the system. However, we wanted to provide an abstraction over this, so the system is as agnostic as possible to who or what is controlling the session and any unnecessary details of the backend implementation.
We implement this by exporting two packages from Streamline:
@cloudflare/streamline/client Defines a high-level, session-based API.
@cloudflare/streamline/ Exposes the Durable Object base class associated with the container. This routes the API requests, implements the preview relay server described below, and provides hooks for security and access policy.
In a remote deployment, the controlling Worker is expected to import @streamline/cloudflare and define a concrete subclass of the Durable Object exposed by the container that can be used for application-specific logic and storage.
In local mode, where there is no Durable Object, the frontend defines a thin adapter layer that maintains the session-based API, but connects directly to the local Docker instance with no access controls, etc.
The example below shows how the controlling application can use the API to access Streamline, prepare a session, and start a video processing pipeline.
config is a JSON object that defines the processing pipeline to be executed, described more in subsequent sections.
The table below shows the complete list of all API calls.
Client method
Function
createStreamline()
Creates a new Streamline instance.
streamline.sessions.create()
Creates a new processing session.
streamline.sessions.resume(id)
Reconnects to an existing session.
session.start(config)
Starts a new processing pipeline.
session.ingest(chunk)
Sends a chunk of video data in “webcam” mode.
session.annotation(png)
Updates the transparent annotation overlay.
session.metrics()
Receives metrics about the current session.
session.stop()
Stops the processing in the current session.
Defining and running a video processing pipeline
session.start() constructs and runs a processing pipeline. It takes a single argument which is a JSON configuration object defining the processing to be performed:
Input(s)
Operations
Output
The example below starts a pipeline that takes an RTMP (real-time messaging protocol) broadcast as input (for example, a feed of a Stream Live input receiving an inbound livestream), applies an overlay image with transparency, and sends the output to an RTMP destination (for example, to another Stream Live input for recording or broadcast). This allows the Worker application to create a modified version of a livestream in real time.
Video-on-demand input via HLS
Streamline can also ingest streaming video input via HLS (HTTP live streaming), for example a video hosted on Cloudflare Stream. The example below shows how a Worker application could run a pipeline that ingests a Stream video, reads the embedded closed caption subtitles and renders them as text on the video, and sends the output via RTMP, for example to a Stream Live Input for broadcasting or recording of the modified version.
Sending video to Streamline
It’s often useful to be able to quickly preview a processing pipeline by sending video data directly to Streamline, for example from a webcam. An agent or embedded device application may also want to use this capability, for example to send footage from factory cameras for AI analysis, or to combine multiple camera feeds into a composite view.
The example below creates a pipeline that expects input from the Worker application and produces a preview video output available over a WebSocket (we’ll talk more about the WebSocket preview video below). It applies two filters and an “annotation,” which is an overlay specified as a PNG image that can be updated while the processing is running, for example to implement an animated graphic.
The code snippet above just starts the pipeline. The controlling Worker is not sending any media to Streamline yet. We’ll discuss the openViewer() function below.
The Worker application sends video data to Streamline using the session.ingest() call. The example below shows how a web browser application might receive chunks from the webcam and forward them to Streamline.
Animated overlay
The annotation overlay can be updated using the session.annotation() call. The example below shows how the Worker application could snapshot a canvas and send it to Streamline. This could be done on an animation loop, although the update rate may be limited in practice by the size of the PNG overlay images, the available bandwidth, and processing power.
Receiving preview video from Streamline
Streamline can also produce preview video output, by specifying output: { mode: 'websocket' }.
Streamline uses WebSockets for low-latency preview video delivery back to the controlling application: the container publishes fMP4 fragments to the Durable Object, which forwards them to an output relay available over a WebSocket on the URL /relay/view, relative to the application origin. The application must connect a WebSocket to this URL, and will then receive video data pushed to it as it becomes available from Streamline. The code snippet below shows how a web browser application might display the preview video feed.
A production MediaSource player must queue fragments while SourceBuffer.updating is true. In local development, the browser or other controlling application simply opens a WebSocket connection directly on the local container.
Currently supported operations
In the configuration object passed to session.start() in the examples above, pipeline is an array of operations from the set supported by the underlying media engine. The operation order is currently fixed by the engine; the order specified in the array is not significant. The list of currently supported operations and the order in which they are applied is below.
Operation name
Function
filter
Applies filtering operations, e.g. blur, saturation.
overlay
Overlays an image referenced by URL or a binary PNG specified separately in a call to annotation().
subtitle
Burns in subtitles.
encode
Specifies output encoding parameters.
Security
This is Cloudflare, so it is important that security is part of the design rather than an addition at the end. We need to ensure that only authorized users can create a new session or take control of an existing one, and that sessions are isolated from each other. We must treat Stream RTMPS input/output keys as secrets that shouldn’t be leaked to the controlling application. We must ensure that resource use is bounded.
The owner deployment is kept private using Workers’ Access integration. The configured owner identity and other allowed users can edit the same shared profiles and start a session while the singleton is idle. The Worker verifies the Access session before accepting control requests and binds the active session to the verified principal. Only one session can run at a time, and a different principal cannot stop or replace the active session.
Stream Live Input keys are stored in Worker secrets or as write-only shared overrides in Durable Object storage. They are never returned by the settings API or placed in browser storage. The controlling application specifies RTMPS input and output by referring to a named profile. The Worker resolves the profile before contacting the container.
The preview video stream has two credentials with separate purposes. A Cloudflare Access service token authenticates the container workload to the publisher endpoint. A random per-session capability authorizes publishing only for the currently active relay. The service token is injected by the container's outbound Worker and never enters container memory. The initial deployment uses a temporary path-specific Access Bypass while the per-session capability remains enforced; after deployment and a successful smoke test, the rollout replaces Bypass with Service Auth.
The owner deployment is intentionally private and singleton-routed. It is not the security model for a public multi-user service.
Playground and open source
We want you to try out Streamline and start building! So together with this post, we are releasing the system as open source and deploying a public playground.
The Streamline container can be run locally or deployed on your account. It exports the Worker API for your control application to use.
There is also an example Worker application with an Astro web frontend that demonstrates Streamline functionality with a few common use cases, including overlays, subtitle decoding, filters and picture-in-picture. There is probe functionality that provides performance metrics and system tracing, and can be useful for debugging the system when developing new features. The example application can be run on a local Astro server, or is set up to be deployed behind Cloudflare Access, so you can control who has access to your Streamline instance.
Both repositories are available as open source on Cloudflare’s GitHub:
We have published a public playground deployment of the example application. This is also something a user can deploy if desired. It uses its own Access configuration, one container identity per verified user, one active session per user, global admission control, concurrency, media and session limits, and no ability for one user to replace another user's session.
Streamline demonstrates one way to combine existing managed services, like Stream, with lower-level primitives to build highly customizable media pipelines. In this iteration, Streamline uses Container CPU for media processing, which introduces a bottleneck at higher qualities or frame-rates.
Moving forward, we’re excited to see how we and our developer community can extend this architecture to build new support for computer vision pipelines, hardware-accelerated media processing, realtime experiences with next generation protocols like WebRTC and MoQ, and ultimately video encoding and decoding primitives natively in Workers.
Today, we invite you to check out our hosted demo of Streamline to see how powerful these tools can be. From there, check out the codebases we’ve open sourced to see how easy it is to deploy Streamline into your own account and use it to create your own experiences.
AWS DevOps Agent investigates operational issues and proposes likely root causes. Many teams, though, want to follow an investigation from where they already work: a ticket in Jira or ServiceNow, or a notification in PagerDuty. When an investigation stays inside the AWS DevOps Agent console, engineers move between tools, and the history of the work is spread across them. This makes the work harder to piece together later.
In this post, I use the AWS Cloud Development Kit (AWS CDK) to build a solution for this. It receives AWS DevOps Agent investigation events through Amazon EventBridge and processes them with AWS Lambda. The solution then creates and updates issues in Jira Cloud. I use Jira as the example, but the same pattern applies to other tools with an API.
Solution overview
This solution is an event-driven integration that starts from the investigation events that AWS DevOps Agent emits. When AWS DevOps Agent creates an investigation, it sends an event with the source aws.aidevops to Amazon EventBridge. An Amazon EventBridge rule matches detail-types that begin with Investigation (a prefix match) and invokes an AWS Lambda function. The Lambda function calls the Jira Cloud REST API: on Investigation Created it creates a new issue, and on every other investigation event it adds a comment to the existing issue.
To link an issue to an investigation, the solution uses Amazon DynamoDB. On Investigation Created, it stores the mapping between the investigation task_id and the Jira issue key in DynamoDB. For later events, it looks up the issue key from that mapping and appends a comment to the same issue. Jira connection details (base URL, user, API token, and project key) are stored in AWS Secrets Manager, so they stay out of the Lambda function’s environment variables and code.
With this design, you can add the integration without changing AWS DevOps Agent itself. Amazon EventBridge handles event delivery, so you add a rule and a target when you want another destination. The processing lives in Lambda, so replacing Jira with another tool keeps the change inside the function code.
Architecture diagram
The following diagram shows the path an investigation event takes from AWS DevOps Agent to Jira Cloud. The flow runs from the event to the created or updated issue without a manual step.
Figure 1: Event flow from AWS DevOps Agent through Amazon EventBridge and AWS Lambda to Jira Cloud
The flow works as follows. AWS DevOps Agent emits an investigation event, and an Amazon EventBridge rule (prefix match on Investigation) captures it and invokes the AWS Lambda function (Jira Issue Creator). The Lambda function reads the Jira credentials from AWS Secrets Manager and stores or reads the mapping between the task_id and the issue key in Amazon DynamoDB. It then calls the Jira Cloud REST API v3 to create an issue or add a comment.
Before you start, prepare the following. First, an AWS account with permissions to create Lambda, Amazon EventBridge, DynamoDB, Secrets Manager, AWS Identity and Access Management (IAM), and AWS CloudFormation resources. Second, a local development environment with the AWS Command Line Interface (AWS CLI) with configured credentials, the AWS CDK, and Node.js 18 or later. You also need an Agent Space in AWS DevOps Agent and a target Jira Cloud project for issue creation.
On the Jira side, create one API token. Create the token from your Atlassian account settings, and note the Jira base URL, the email address tied to the token, and the project key where issues are created. You store these values in Secrets Manager in the next step.
Step 1: Store the Jira credentials in Secrets Manager
First, store the Jira connection details in AWS Secrets Manager. The Lambda function reads the credentials from here, so no secret stays in the code or in a parameter. The following command stores the four values as a single secret.
When the command succeeds, it returns the Amazon Resource Name (ARN) of the secret. Note this ARN. You use it in the next step.
Step 2: Deploy with the AWS CDK
Get the repository and deploy the stack with the CDK. If this is the first time you use the CDK in this Region, run cdk bootstrap first. When you deploy, pass the ARN of the secret from Step 1 as a parameter.
cd cdk
npm install
npm run build
cdk deploy --parameters SecretArn=arn:aws:secretsmanager:ap-northeast-1:123456789012:secret:devops-agent-jira-credentials-AbCdEf
This stack creates the Amazon EventBridge rule, the Lambda function, the DynamoDB table, and the related IAM roles. The Lambda function receives only the permission to read the specified secret and to read from and write to the DynamoDB table.
Step 3: Understand the investigation event structure
Knowing what the Lambda function receives makes it more straightforward to adapt the solution to other tools. AWS DevOps Agent sends an event each time the state of an investigation changes. The state moves from PENDING_START to IN_PROGRESS to COMPLETED, and arrives with the detail-types Investigation Created, Investigation In Progress, and Investigation Completed. Alongside these three, you handle Investigation Failed, Investigation Timed Out, Investigation Cancelled, and Investigation Priority Updated through the same mechanism.
The following is part of an Investigation Created event. The detail.metadata.task_id value uniquely identifies the investigation, and the solution uses it as the DynamoDB key.
The Amazon EventBridge event does not include the investigation title or description. To fill in the summary and body of the issue, the Lambda function calls the AWS DevOps Agent GetBacklogTask API. The Investigation Completed event includes data.summary_record_id, which you use to retrieve the investigation summary.
Step 4: Confirm the behavior
When you start an investigation in AWS DevOps Agent, it emits an Investigation Created event, and the Lambda function creates an issue in Jira. The following screen shows an investigation starting in the AWS DevOps Agent console. The Investigation timeline shows the first event, where an Amazon CloudWatch alarm entered the ALARM state.
Figure 2: An investigation starting in the AWS DevOps Agent console
At this point, a new issue is created on the Jira side. The following screen shows the Jira board, where a single issue created by AWS DevOps Agent (its summary begins with [DevOps Agent]) appears in the To Do column.
Figure 3: The new Jira issue in the To Do column
When you open the issue, the detail looks like the following. The description holds the investigation title and the alarm details, along with metadata such as task_id, execution_id, agent_space_id, and the status. The Lambda function assembles these from the Amazon EventBridge event and the GetBacklogTask API.
Figure 4: The Jira issue detail with investigation metadata
As the investigation progresses and the state changes, the Lambda function looks up the issue key in DynamoDB and adds a comment to the same issue. When the investigation completes, AWS DevOps Agent presents a root cause. The following screen shows the cause summarized as an intentionally failing Lambda function.
Figure 5: The completed investigation and its root cause in the AWS DevOps Agent console
At the same time, a comment with the investigation result is appended to the Jira issue. The comment in the following screen includes the Investigation Completed status and an Investigation Summary that covers the symptoms, findings, and root cause.
Figure 6: The investigation result appended as a Jira comment
Through this flow, an engineer follows the investigation from start to finish by looking at Jira alone, without switching between the AWS DevOps Agent console and Jira.
Adapting to other tools
The same pattern applies to other tools with an API. What you change is mainly the body of the Lambda function. For ServiceNow, you replace the issue-creation call with the incident-creation API. For PagerDuty, you call its notification API instead of creating or updating an issue. The Amazon EventBridge rule and the DynamoDB mapping stay the same, so you don’t rebuild them for each destination. Storing credentials in Secrets Manager is also common across tools.
Clean up
When you finish testing, delete the resources you no longer need to avoid future charges. First, delete the CDK stack.
cd cdk
cdk destroy
Next, delete the Jira credentials stored in Secrets Manager.
The DynamoDB table and the Lambda function are created as part of the CDK stack, so deleting the stack removes them as well.
Conclusion
In this post, I used the AWS CDK to build a solution for connecting AWS DevOps Agent to Jira Cloud. It receives investigation events through Amazon EventBridge, processes them with AWS Lambda, and creates and updates issues in Jira. Because you follow the investigation from the Jira side, engineers keep working inside the tools they already use. By storing credentials in AWS Secrets Manager and mapping investigations to issues in Amazon DynamoDB, you keep each state change on the same issue. The same design applies to other tools with an API, such as ServiceNow and PagerDuty.
The gap in most observability setups isn’t the data, it’s closing the loop between incident detection and response. Your Amazon OpenSearch Service domain already stores structured logs and distributed traces, detects anomalies through alerting monitors, and fires notifications reliably. But when the system sends that alert lands at 2 AM, your engineers must still investigate manually. The investigating engineer opens Dashboards, crafts domain-specific language (DSL) queries, hunts for correlated trace IDs, pivots to AWS CloudTrail, and pieces together a root cause. That manual loop can take anywhere from minutes for straightforward issues to hours when failures cascade across microservices.
What if you closed the loop by connecting an AI agent directly to the indices that triggered the alert?
This post shows you how to connect AWS DevOps Agent to your OpenSearch observability data using the Model Context Protocol (MCP). The same alert that would page a human instead triggers the agent to query the logs and traces automatically, correlate them with AWS CloudTrail and Amazon CloudWatch, and deliver a root cause analysis.
Note: AWS DevOps Agent access to logs and traces in OpenSearch is controlled by the fine-grained access control (FGAC) role it is given.
We present three hosting paths for the MCP server so you can pick by AWS Region and operational preference: self-managed on Amazon Elastic Container Service (Amazon ECS) (AWS Fargate behind an internal Network Load Balancer (NLB), reachable through Amazon VPC Lattice, which works everywhere today), Amazon Bedrock AgentCore (a one-click AWS CloudFormation template, where available), or the built-in MCP endpoint on OpenSearch 3.3+ (no separate server). All paths use the official opensearch-mcp-server-py package.
By the end, you can deploy the MCP server (any of the three paths), register it as a Capability Provider, configure the AWS Identity and Access Management (IAM)-to-FGAC role mapping, route your OpenSearch alerts to the agent’s Event Channel, and verify the closed loop with a controlled failure.
Solution overview
The architecture creates a closed loop: your OpenSearch domain handles both detection and investigation, with AWS DevOps Agent orchestrating the response.
Figure 1: Architecture for Amazon OpenSearch alert to AWS DevOps Agent MCP investigation
Closed-loop architecture: OpenSearch Alerting triggers Amazon Simple Notification Service (Amazon SNS), which flows through a webhook forwarder to the Event Channel of AWS DevOps Agent. The agent investigates by querying OpenSearch indices using MCP tools, correlates with CloudTrail and CloudWatch, and delivers a root cause analysis.
The loop runs in six stages:
Applications emit logs and traces to OpenSearch.
Alerting monitors detect anomalies and publish to Amazon SNS.
Amazon SNS triggers the webhook forwarder AWS Lambda function.
The webhook forwarder Lambda transforms each notification into a hash-based message authentication code (HMAC)-signed payload for the agent’s Event Channel.
AWS DevOps Agent queries the OpenSearch indices through MCP and correlates them with CloudTrail and CloudWatch.
AWS DevOps Agent delivers a root cause analysis.
The critical insight: AWS DevOps Agent consumes remote MCP servers registered as Capability Providers. AWS DevOps Agent doesn’t connect to local MCP servers running on developer workstations (those are used by OpenSearch MCP apps for IDEs). The agent needs a network-accessible endpoint: either self-managed on Amazon ECS, hosted on Bedrock AgentCore, or built into the OpenSearch domain itself (3.3+).
Choosing a hosting path
This post walks through the self-managed path in detail (works everywhere today) and calls out the AgentCore and built-in 3.3+ alternatives at each step.
Prerequisites
Confirm the following before starting:
An Amazon OpenSearch Service managed domain (2.x+) with FGAC enabled, application logs and traces already indexed, and Alerting monitors publishing to an SNS topic.
AWS DevOps Agent enabled in your account.
AWS Cloud Development Kit (AWS CDK) (npm install -g aws-cdk, Node.js 18+) and AWS Command Line Interface (AWS CLI) v2 configured with admin access to your OpenSearch domain.
Note: This walkthrough assumes you already have an application emitting logs and traces to OpenSearch. The infrastructure in this post is shown as inline CDK snippets you can drop into your own CDK app and adapt to your environment.
Step 1: Deploy the OpenSearch MCP server
Choose the path that fits your Region. The suggested options host the official opensearch-mcp-server-py and expose the same MCP tools (SearchIndexTool, ListIndexTool, and the broader observability tool set) to AWS DevOps Agent.
Option A: Self-managed on Amazon ECS and NLB (deploy anywhere)
This path runs the MCP server on ECS Fargate behind an internal Network Load Balancer and exposes it to AWS DevOps Agent through a VPC Lattice private connection. It works in every Region today.
Figure 2: Architecture for self-managed OpenSearch MCP deployed on AWS Fargate
1. Generate a Transport Layer Security (TLS) certificate for the MCP server
AWS DevOps Agent requires HTTPS endpoints. Generate a self-signed certificate whose subject alternative name (SAN) matches your NLB DNS name, and import it to AWS Certificate Manager (ACM):
# Generate a self-signed cert whose SAN matches the NLB DNS, then import to ACM
openssl req -x509 -nodes -days 365 -newkey rsa:2048 \
-keyout /tmp/mcp-key.pem -out /tmp/mcp-cert.pem \
-subj "/CN=mcp-server" -addext "subjectAltName=DNS:<your-nlb-dns>"
aws acm import-certificate --certificate fileb:///tmp/mcp-cert.pem \
--private-key fileb:///tmp/mcp-key.pem --region <region> \
--query "CertificateArn" --output text
Save the returned certificate Amazon Resource Name (ARN) for the next step.
Why the SAN matters: VPC Lattice validates the certificate against the host address you configure for the private connection. If the SAN doesn’t match the NLB DNS, TLS validation fails and the connection never reaches Completed.
2. Define the MCP server in your CDK app
Run opensearch-mcp-server-py as an Amazon ECS Fargate service behind an internal NLB. The following snippet shows the essential wiring. Adapt it to your existing CDK app:
// Representative wiring — adapt into your CDK app (full construct in the linked Guidance).
// Fargate task runs the MCP server; grant it read on the domain (FGAC handles index auth).
taskDef.addContainer('mcp', {
image: ecs.ContainerImage.fromRegistry('python:3.12-slim'),
command: ['sh','-c', 'pip install opensearch-mcp-server-py --quiet && '
+ 'opensearch-mcp-server-py --transport stream --host 0.0.0.0 --port 8080'],
environment: { OPENSEARCH_URL: https://${props.openSearchDomain.domainEndpoint},
OPENSEARCH_USE_SSL: 'true' }, portMappings: [{ containerPort: 8080 }] });
// Internal NLB — MUST have a security group so VPC Lattice can reach it; TLS listener
// terminates with your ACM cert and forwards TCP:8080 to the service.
// Output the NLB DNS and task role ARN for Steps 2 and 3.
Deploy it (cdk deploy --require-approval broadening --region <region>) and note the NLB DNS and task role ARN from the stack outputs. You’ll need the stack outputs for the private connection (Step 2) and the FGAC mapping (Step 3).
The server starts in streamable-HTTP mode (–transport stream). Installing the package at container startup adds approximately 30 seconds to the first task boot up time. The higher task sizing (1024 MB/512 CPU) helps pip install complete quickly, and the health check’s unhealthyThresholdCount: 5 gives the service approximately 2.5 minutes to stabilize. For production, bake the package into a prebuilt image to avoid startup latency entirely.
3. Critical: Use a Network Load Balancer (NLB), not an Application Load Balancer (ALB)
The OpenSearch MCP servers use the streamable-HTTP transport, which delivers responses as Server-Sent Events (SSE) with chunked transfer encoding. Application Load Balancers (ALBs) operate at Layer 7 and can strip the Transfer-Encoding: chunked header, breaking the SSE stream. Network Load Balancers operate at Layer 4 (TCP) and pass HTTP framing untouched after TLS termination. Therefore deploy a streamable-HTTP MCP server behind an NLB with TLS termination and not an ALB.
Important: The NLB must have a security group so VPC Lattice resource gateway ENIs can reach it, and a security group can only be attached to an NLB at creation time (it cannot be added later). The CDK stack attaches one. If you build your own NLB, specify the security group when you create it.
Option B: Amazon Bedrock AgentCore (managed, where available)
If your Region supports the integration, the OpenSearch console provides a one-click CloudFormation template that deploys opensearch-mcp-server-py on AgentCore.
Figure 3: Architecture for OpenSearch MCP deployed on Amazon Bedrock AgentCore
Open the Amazon OpenSearch Service console, select your domain, and go to Integrations.
Locate the MCP server template and choose Launch stack.
Provide the parameters:
Parameter
Description
Example
OpenSearchEndpoint
Your domain’s endpoint
https://my-domain.<region>.es.amazonaws.com
AWSRegion
Region where the domain runs
<region>
AgentName
Logical name for this MCP server
opensearch-observability-mcp
After the stack reaches CREATE_COMPLETE, note from the Outputs tab: AgentCoreEndpoint, CognitoClientId, CognitoClientSecret, and McpServerRoleArn (needed for FGAC in Step 3).
Verify the tools are registered:
# Fetch an OAuth token, then confirm the tools list. Expect SearchIndexTool, ListIndexTool.
curl -s -X POST "https://<AgentCoreEndpoint>" -H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}' \
| jq '.result.tools[].name'
Option C: OpenSearch 3.3+ built-in MCP
Figure 4: MCP deployment architecture for domains running OpenSearch version 3.3+
This is the simplified path for domains running OpenSearch 3.3+. These versions have a built-in MCP endpoint, so no separate server deployment is needed. Enable the endpoint with two API calls:
Your MCP endpoint is live at https://<domain-endpoint>/_plugins/_ml/mcp. In Step 2, register this URL directly with AWS DevOps Agent using SigV4 authentication (service=es).
Step 2: Register with AWS DevOps Agent
With the MCP server running, register it as a Capability Provider so AWS DevOps Agent can invoke its tools during investigations.
The private connection in section 2.1 applies only to the self-managed path (Option A). For the AgentCore and built-in paths, skip to section 2.2.
2.1 Create a private connection
A private connection creates a secure network path between AWS DevOps Agent and your NLB using Amazon VPC Lattice.
Open the AWS DevOps Agent console.
Navigate to Capability Providers then Private connections and choose Create a new connection.
Configure the connection with these settings:
For Name, enter opensearch-mcp-connection.
Select your virtual private cloud (VPC) and private subnets (one per Availability Zone (AZ)).
Attach a security group that allows inbound TCP 443.
For Host address, enter your NLB DNS name (the Step 1 output), and set TCP port to 443.
For Certificate public key, paste the contents of /tmp/mcp-cert.pem.
Choose Create Connection and wait for status Completed (~5–10 minutes).
2.2 Register the MCP server
Then register the Capability Provider. The fields differ by path:
Field
Self-managed (ECS)
AgentCore
Built-in (3.3+)
Name
opensearch-observability
opensearch-observability
opensearch-observability
Endpoint URL
https://<nlb-dns>/mcp
https://<AgentCoreEndpoint>
https://<domain-endpoint>/_plugins/_ml/mcp
Private connection
opensearch-mcp-connection
—
—
Authentication
API key (x-api-key)
OAuth
SigV4
OAuth client ID / secret
—
from stack outputs
—
SigV4 service
—
—
es
2.3 Configure allowed tools
Classify the MCP tools as read-only so the agent can query but not mutate your domain: SearchIndexTool, ListIndexTool (and, if present, GetMappingsTool and GetShardsTool).
2.4 Verify the registration
In the AWS DevOps Agent console test interface, ask:
“List the indices in my OpenSearch domain that match application-logs-*”
The agent should invoke the list tool and return your index names. A 403 Forbidden or security_exception means you haven’t configured the FGAC mapping yet. Continue to Step 3. A connection timeout on the self-managed path points to the private connection status or NLB security group.
Step 3: Configure FGAC role mapping
OpenSearch managed domains with FGAC enforce a strict separation: IAM authenticates the caller, but the internal security plugin of OpenSearch authorizes access to the indices. Without an explicit mapping between the IAM role and an OpenSearch backend role, authenticated requests still receive 403 Forbidden.
Identify the role to map
Path
IAM Role
Where to find it
Self-managed (ECS)
ECS task role
CDK output McpTaskRoleArn (from the Step 1 snippet)
AgentCore
MCP server role (trust: agentcore.bedrock.amazonaws.com)
CloudFormation Outputs → McpServerRoleArn
Built-in endpoint
DevOps Agent role (trust: aidevops.amazonaws.com)
DevOps Agent console → Agent Space → IAM Configuration
Apply the role mapping (recommended: scoped to your indices)
With the role identified, apply the mapping. We recommend that you scope it to your observability indices rather than granting cluster-wide read, so you can control which data the agent can query. The first call creates a read-only role restricted to those indices. The second maps your IAM role to it:
# Create a read-only role scoped to your observability indices
curl -XPUT "https://<domain-endpoint>/_plugins/_security/api/roles/devops_agent_readonly" \
--aws-sigv4 "aws:amz:<region>:es" -H "Content-Type: application/json" -d '{
"cluster_permissions":["cluster_composite_ops_ro"],
"index_permissions":[{"index_patterns":["application-logs-*","otel-traces-*"],
"allowed_actions":["read","search"]}] }'
# Map the IAM role to it
curl -XPUT "https://<domain-endpoint>/_plugins/_security/api/rolesmapping/devops_agent_readonly" \
--aws-sigv4 "aws:amz:<region>:es" -H "Content-Type: application/json" -d '{
"backend_roles":["arn:aws:iam::<ACCOUNT_ID>:role/<YourMcpTaskRole-or-DevOpsAgentRole>"] }'
Your OpenSearch Alerting monitors already fire to SNS. The remaining connection is a lightweight Lambda that transforms those SNS notifications into the format the AWS DevOps Agent Event Channel expects, with HMAC signing for payload integrity. This is the piece that closes the loop: investigations trigger automatically, without a human forwarding the alert.
4.1 The webhook forwarder
The forwarder performs two operations: payload transformation and HMAC-SHA256 signing.
# HMAC-SHA256 signing contract. DevOps Agent expects two headers:
# x-amzn-event-signature = base64(HMAC-SHA256(secret, f"{timestamp}:{body}"))
# x-amzn-event-timestamp = %Y-%m-%dT%H:%M:%S.000Z (UTC)
# transform_alert() maps the OpenSearch Alerting payload to a DevOps Agent event
# (severity 1/2->HIGH, 3->MEDIUM, 4->LOW), deriving title/service/incidentId/metadata
# from monitor_name, trigger_name, and results.
The complete implementation adds retry logic (three attempts with exponential backoff at 1, 2, and 4 seconds), AWS Secrets Manager integration for HMAC secret caching, and structured JSON error logging.
4.2 Deploy the forwarder
Define the forwarder as a Lambda subscribed to your alerting SNS topic. Package the preceding transformation and signing logic as the handler, wire it up in AWS CDK, and deploy with cdk deploy --require-approval broadening --region <region>:
With the forwarder deployed, create the webhook in the AWS DevOps Agent console and wire its URL and signing secret back into the Lambda.
In the AWS DevOps Agent console, go to Agent Space, Webhooks, Agent Space Webhook, and then choose Add webhook.
Complete the setup steps: verify the data schema, configure HMAC authentication, and generate the URL and credentials.
Note the HMAC signing secret and store it in the AWS Secrets Manager secret referenced by the forwarder (devops-agent-webhook-secret). See Create an AWS Secrets Manager secret in the AWS Secrets Manager User Guide.
Set the generated Webhook URL as the WEBHOOK_URL environment variable on the forwarder Lambda (the eventChannelUrl prop in the preceding snippet).
If you don’t have a monitor yet, create an SNS notification channel (Notifications plugin) and a query-level alerting monitor over application-logs-* that triggers when error_count > 5 and posts to that channel. Full request bodies are in the OpenSearch Alerting docs. The action’s message_template must emit monitor_name, trigger_name, severity, period_start, period_end, and results to match the forwarder’s schema.
The IAM role (role_arn) needs sns:Publish permission on your topic and a trust policy allowing es.amazonaws.com to assume it. The destination_id in the action must match the config_id from the notification channel.
Note:401 Unauthorized means the HMAC secret in Secrets Manager doesn’t match the Event Channel secret. Connection refused means the Event Channel URL is wrong or the Lambda lacks outbound access.
Step 5: Verify the closed loop
With all four components connected (MCP server, Capability Provider, FGAC mapping, and alert routing), trigger a controlled failure to verify the full loop.
5.1 Inject failure
Set reserved concurrency to zero on one of your application’s Lambda functions. This causes all invocations to be throttled:
After you inject the failure and generate traffic, events should unfold roughly as follows. Use this timeline to confirm each stage of the loop is firing:
Elapsed
Event
T+0s
put-function-concurrency executed
T+30s
Throttle errors appear in application-logs-* index
T+~120s
Alerting monitor evaluates and triggers
T+~130s
SNS → Forwarder Lambda → Event Channel delivery
T+~135s
DevOps Agent begins investigation
T+~300s
Root cause analysis delivered
What AWS DevOps Agent produces
ROOT CAUSE ANALYSIS - error-rate-monitor / high-error-rate
1. OpenSearch logs (SearchIndexTool): 47 ERROR entries, "TooManyRequestsException" on service=OrderProcessor
2. Traces (SearchIndexTool): matching spans show status.code=429, function never executed
3. CloudTrail: PutFunctionConcurrency set ReservedConcurrentExecutions=0 two minutes before first error
4. CloudWatch: Throttles=47, Invocations=0
ROOT CAUSE: reserved concurrency set to 0 on OrderProcessorFunction, blocking all invocations.
REMEDIATION: aws lambda delete-function-concurrency --function-name OrderProcessorFunction
The agent identified the root cause by querying the same data that triggered the alert, which completes the closed loop.
5.4 Revert the failure
After the investigation completes, restore normal capacity by removing the reserved concurrency limit you set earlier:
Note: If the agent doesn’t begin investigation within 3 minutes, check: (1) the forwarder Lambda executed (CloudWatch Logs), (2) the Event Channel shows the received event, and (3) the Capability Provider is registered and healthy.
Cost considerations
This walkthrough adds roughly $70/month (as of September 2026, and varies by AWS Region and usage): ECS Fargate MCP task approximately $15, network address translation (NAT) gateway approximately $35, NLB approximately $18, Secrets Manager approximately $0.40, and the Lambda forwarder under $1. The AgentCore path removes the NLB and ECS costs but adds AgentCore hosted-endpoint charges. Your existing OpenSearch domain and application aren’t included.
Security considerations
The design keeps everything inside the VPC: OpenSearch and the MCP server run in private subnets with nothing public-facing, and AWS DevOps Agent reaches the MCP server over a VPC Lattice private connection. NLB terminates TLS (OpenSearch enforces HTTPS, TLS 1.2 minimum), and encryption at rest is enabled. IAM is least-privilege: the MCP server’s role maps to a read-only OpenSearch backend role scoped to your observability indices.
For production, also consider Security Assertion Markup Language (SAML)/IAM FGAC, request validation in front of the MCP server, and VPC endpoints for Amazon Elastic Container Registry (Amazon ECR), CloudWatch, and Secrets Manager.
Cleanup
# Destroy the CDK stacks (MCP server on ECS, webhook forwarder)
cdk destroy --all --region <region>
# AgentCore path: delete its CloudFormation stack
aws cloudformation delete-stack --stack-name opensearch-mcp-agentcore
# In the DevOps Agent console: remove the Capability Provider and the private connection.
# Remove the FGAC role mapping
curl -XDELETE "https://<domain-endpoint>/_plugins/_security/api/rolesmapping/devops_agent_readonly" \
--aws-sigv4 "aws:amz:<region>:es"
# 3.3+ path: disable the built-in endpoint. Delete the imported ACM certificate.
Verify the resources are removed: ECS tasks, NLB, NAT Gateway, AgentCore hosted endpoint (if used), and the Secrets Manager secret.
Conclusion
You connected AWS DevOps Agent to your OpenSearch observability data through MCP, with three hosting paths so Region availability doesn’t block you: self-managed ECS (everywhere today), AgentCore (low-ops, where available), or the built-in 3.3+ endpoint. The loop is now closed: the same domain that stores your data and fires alerts becomes the investigation source, and the agent can help determine why an alert fired automatically.
Start with one alert that fires frequently and costs your team time to investigate manually. Connect it through this pipeline, watch the agent produce its first root cause analysis, and iterate from there.
Rust has a number of kinds of smart pointers, both in the standard
library and defined by users. Still, some operations that are possible with
built-in references are not possible to perform with user-defined smart
pointers. Tyler Mandry, lead of the Rust project’s
language team, spoke at
RustConf 2026 about the lengthy effort to change that, and make smart pointers
just as flexible as built-in references.
St. Lucie County in Florida discovered (alt link) a dozen Flock cameras whose ownership it can’t identify, and that the county government had not permitted.
I am reminded of the decade-old story of StingRay cell phone surveillance devices in Washington, DC, whose operators were also unknown.
My guess is that in the StingRay case, the devices were operated by foreign actors. This Flock case is more likely some local government entity that didn’t bother getting approval. Were I a foreign actor, I would rather hack the existing Flock network—like Israel did with Tehran’s surveillance cameras—than risk installing my own.
Regardless, once we normalize a surveillance infrastructure, both friends and foes will take advantage of it.
The Linux Test
Project has announced its
latest stable release for September 2026. There have been 382 patches
from 41 authors since the May 2026
release. See the announcement for a list of new tests, changes, and
more.
Fun fact: when you use an agent and it needs to fetch a live web page, the agent usually just guesses the URL of the page and then makes a tool call to curl it. This is why you’ll sometimes see web fetches come back with a 404 Not Found, which happens if the agent incorrectly guesses the URL of that information. As you can imagine, it’s not super efficient to randomly guess URLs all the time.
There is a better way. What if your agent can actually browse the Internet, just like how humans start with a search engine query when we’re looking for information? This is what web search is designed to do — it enables agents to search for relevant data on the Internet and grounds an agent’s responses based on live information.
Today, we’re announcing Cloudflare’s partnership with web search providers to bring you grounded intelligence via AI Gateway. We’re kicking off this launch with our partners from Ceramic.ai, Exa, and Linkup.
What can I do with the Web Search API?
AI models are only as good as the context you feed them. Models are typically trained and then frozen at a point in time, operating only on information that existed before their knowledge cut off date. This makes it quite hard to engage with models about recent events, changing APIs, or fast-evolving news.
Integrating Web Search API directly into your inference pipeline equips your agents with a dynamic context layer. Your applications get fresh, structured snippets from the web injected straight into context, which gives your models access to live information.
For example, if your agent was building with Cloudflare developer tools, it might miss all the new products and features we’re releasing during this Birthday Week! With web search, you’ll be able to retrieve the latest and greatest documentation and releases, so you can build faster and smarter.
Elevating the industry standard for web search
At Cloudflare, we believe that crawlers should be honest, transparent, and respect all bot rules and preferences, and that site owners should have meaningful transparency and control over how their content is used. With our launch today, we’re excited to announce that our partners have committed to meeting Cloudflare’s bot crawling standards.
The crawler used by the web search provider must comply with Cloudflare’s publicly stated requirements for “Verified bots” as defined in our developer documentation, and web search responses must include a link to the location of crawled content. These rules are net-positive for a fair Internet, and allow creators to decide what they want to do with their data.
We’re extremely excited to be taking another step in setting the bar for what it means to be a good crawler on the Internet, and even more proud of the web search partners who have risen to the challenge to uphold these standards with us. When you use web search on Cloudflare, you choose to consume search knowledge from operators committed to providing the transparency, control, and visibility that helps build a better Internet, by identifying their crawlers, respecting robots.txt, and providing the source of search results.
Using Web Search API via AI Gateway
Cloudflare’s AI Gateway is the flagship integration point for our new Web Search API product. AI Gateway is designed to be the control plane for your applications, with observability, unified billing, security, and access controls all built-in. Naturally, we thought web search would fit right in: you can consume web search with your AI Gateway credits; collect logs and request data on web search calls; and control who gets access to what web search providers.
Requests show up in your normal AI Gateway observability logs, and web search queries draw down from your AI Gateway credit balance. We will identify partners supporting Zero Data Retention (ZDR), so that you know that your data is not retained. We also offer web search directly at list API pricing from our partners, without any additional markup. More details can be found on the web search developer docs.
We also support Bring-Your-Own-Key (BYOK) with web search providers, as we do with model inference providers. This way, you’re able to bring your existing organizational setup and get started with AI Gateway and web search in a few simple steps.
Direct REST API
If you’re making HTTP calls from an existing backend, mobile app, or external service, you can query web search directly through a standard REST endpoint. Simply pass your AI Gateway authentication token, specify your preferred provider in the payload, and it will return web search results.
Workers Bindings
For developers building directly on Cloudflare Workers, integration takes just a line of code. We also have a standalone worker binding that you can use to do web search for your agent:
Coming soon: Server tools
We are actively building native Server Tools directly into AI Gateway. Soon, you won’t need to define tools yourself: they will come built into our control plane so you can spend more time building rather than orchestrating the harness. Web search will be one of the first tools we incorporate into our stack of server tools, and we’re excited for you to try it.
However, if you’d like to orchestrate web search as a server tool yourself today, you can easily do so with the following Worker code snippet:
Try it out today!
We’re excited to bring web search to our platform and to be doing this with wonderful partners who are championing what it means to be a good steward of the Internet. Please give our new web search tools a try via AI Gateway, the standalone REST API, and other formats in the future. Get started with our developer docs today, or try it out on the AI Playground with your AI Gateway account and key.
Rapid7 tracked a set of Linux samples that blend into the software and device conventions of the telecom environments they target. The set spans a newly observed BPFDoor variant, a BPF Rekoobe build seen against South Korean targets, a dropper, and six builds of a Linux implant we track as AVERAT, deployed against Taiwanese appliances. Additionally, we provide source code details of the Rapid7 BPFDoor controller introduced in our April 2026 blog, Stealthy BPFDoor Variants are a Needle That Looks Like Hay.
The chain uses two binaries. A dropper writes a shell script to the appliance’s storage mount and executes it. The script stages both payloads into /sbin under the names ntpdate and udevds, launches them, and deletes each file ten seconds later while the processes continue running. One of those payloads is the dropper itself, re-executing as a resident watchdog, leaving both processes running without an on-disk image.
The dropper derives its encryption key from the string ShareTech and lives in the appliance’s own add-on package directory. The BPFDoor variants seen against South Korean systems impersonate the PID file of SpamSniper, a Korean anti-spam product, and rotate through ten Linux daemon names. Across the samples, each component adopts names and conventions designed to look unremarkable in the environment it targets.
The common thread is regionalized disguise: each sample is aware of the vendor’s software running on the targeted systems and implements process spoofing accordingly. Passive BPF implants avoid conventional port scans; while outbound beacons hide inside ordinary DNS, TCP, and traffic, the threat-actor(s) are leveraging SMTP to stay under the radar. Telecommunications and network-edge operators are most affected, including embedded devices such as CCTV and DVR systems that can sit close to the network core. Readers will learn how each component works, what binds the six AVERAT builds to one another, and which behaviors and indicators to hunt for.
Technical analysis
Rapid7 BPFDoor controller
Figure 1: Overview of BPFDoor HTTP-tunneled trigger flow through edge proxy
⠀
Following our introduction of the Rapid7 BPFDoor controller, published in March, this section examines new features from the reconstructed source code.
Earlier BPFDoor variants relied on raw “magic bytes” (like 0x7255 or 0x5293) sitting in the TCP or UDP headers. Once security vendors wrote static network signatures (Suricata/Snort) to detect these Layer 4 anomalies, the operators began targeting the edge proxies. By wrapping the magic packet in standard HTTPS POST requests and relying on SSL offloading common in telecom environments, the trigger can be delivered to the BPFDoor-infected node in a way that may evade conventional deep packet inspection.
Because proxies alter HTTP headers (adding X-Forwarded-For and changing User-Agent lengths), the malware can no longer rely on static byte offsets to find its payload. To solve this, the new controller sends fake, benign-looking web requests (e.g., POST /admin/login.aspx?id=99990) that are mathematically padded. This guarantees that the string “9999” lands at exactly offset 26 of the TCP payload consistently.
The backdoor uses this “9999” as a reference point, dynamically scans for the \r\n\r\n terminator, and extracts the hex-encoded command payload from the HTTP body.
The dogetlogin function contains the hardcoded paths blending in with legitimate requests:
Figure 2: Hardcoded web login paths used by the dogetlogin function
⠀
When running, the controller spoofs the identity of /usr/sbin/abrtd via set_proc_name and PR_SET_NAME. The #ifndef SOLARIS compiles safely across different operating systems, applying the abrtd disguise only where the Linux-specific prctl function is supported.
Figure 3: Process name spoofing logic applying the abrtd disguise on non-Solaris systems
⠀
The table below lists the Rapid7 controller flags, with new features identified relative to the TrendAI analysis marked accordingly.
Switch
Variable/Action
Description
-h
destip
Specifies the target host (the infected machine’s IP address) to control.
-d
dport
Sets the destination port on the infected host to send the trigger packet to.
-l
lhost
Sets the remote IP address that the infected machine will connect back to (Reverse Shell).
-s
lport
Sets the destination port to listen for incoming connections on the attacker’s machine.
-m
self = 1
Sets the attacker’s local IP address as the remote host, automatically setting up the local listener (overwrites -l).
-b
bport
Instructs the controller to bind to a specified TCP port locally (Bind Shell mode).
-n
nopass = 1
Sends the packet without prompting for a password (sends an empty/hashed password). Often used just to check if the backdoor is alive.
-i
raw = 2
ICMP mode. Embeds the magic packet into an ICMP Echo Request.
-u
raw = 3
UDP mode. Sends the magic packet via a UDP datagram.
-w
raw = 1
TCP mode. Sends the magic packet via a raw TCP SYN packet.
-f
magic_flag
Allows the operator to manually define a custom magic byte sequence (integer value).
-o
magic_flag = 0x5571
Quick-sets the magic bytes/flag to 0x5571.
-H
hdestip
[NEW] Specifies a secondary “hidden” IP address to embed inside the newly added hip field used to relay the magic packet.
-g
gethost
[NEW] Activates the HTTPS POST tunneling mode (dogetlogin).
-D
dir
[NEW] Customizes the URI directory path to blend into specific web server logs when using the -g (HTTPS POST) mode.
-v
debug = 1
[NEW] Enables verbose/debug mode, which is particularly useful for printing out the crafted HTTP requests and responses.
-t
tmout
[NEW] Sets a custom timeout value.
-c
break;
[DEPRECATED] Parses the flag but takes no actions
Table 1: Rapid7 BPFDoor Controller Flags and Descriptions
A new BPFDoor variant tied to the South Korean cluster
The BPFDoor variants create a raw PF_PACKET socket, attaching a classic BPF filter matching Rapid7 Variant F and using magic bytes 0x6693 (UDP), 0x4274 (TCP) and 0x7820 (ICMP). On a match, the implant extracts the source address and connects back to the sender if the password is gZbpx0, opens a bind shell if the password is sT21xf, and otherwise defaults to a UDP knock.
Strings are hidden with a rotating substitution alphabet. Decoding reveals a direct product-spoofing artifact and a set of service-name disguises. The SpamSniper /var/run/spamsniper.pid mutex, together with the sample provenance, ties this build to the South Korean cluster.
SpamSniper is antispam software used mainly in South Korea, so this masquerade is consistent with targeting a Korean mail or telecom environment. The variants a37ea9897221d4495b538de72b74f2aa1d2ff09b7b6dcedd395aee58931adbf3 and 7e667ba5f9df912e02275d3cfe3809d16f822fe776f4035c84b118ebd925b1b5 share the same filter, packet parser, callback, and command paths.
The data plane variant
The BPFDoor sample (a6f3b7f932761fb1fd5e74123f2482e36c65dd13e769af2ce08c65da195bfa7a) attaches a SOCK_RAW 16-BPF instructions parsing IP/TCP offsets and gates on a 14-byte payload (2B 76 C0 63 83 E9 5F E1 EE 69 3F 32 CD 94, unique per sample).
Figure 4: BPF filtering for abc00922 TCP magic bytes
⠀
It spoofs its process name to ora_ppmond, mimicking the naming convention of Oracle-backed telecom subscriber and provisioning platforms (HSS, OSS/BSS), a disguise that only reads as legitimate on hosts actually running that class of infrastructure. Once triggered, it opens a stock Tiny Shell session and dispatches single-byte ‘S’/’U’/’D’ commands — interactive shell, upload, download — the same switch-case and iptables NAT-redirect staging/teardown logic found byte-for-byte in a second sample 4435fcd6862921092614dbeaa880e4192352984686ebcd98f0ba13ee8e226ef9 (the latter spoofing /sniper/snipe/bin/dtnpd and /sniper/bin/ofgmd). These samples show BPFDoor operating as a modular framework that adapts to the telecom layer it targets, integrating Tiny Shell and Rekoobe logic to support exfiltration.
Figure 5: Tinyshell logic integrated into BPFDoor
⠀
BPF Rekoobe
The sample 652508a9cf40bee883dc0e5e219dfeba71fe7dac591d01c89f74c21f73b4963f is a Rekoobe-based backdoor. It attaches a 26 BPF instruction filter, sniffing for TCP/UDP/SCTP IPv4 and UDP IPv6 traffic with source and destination ports equal 25. Strings are protected with a repeating-key XOR routine (uvTIgh47,@#R), which decodes internal markers and command tokens. The magic packet is authenticated against a 32-byte sequence: 5C A3 1E F9 72 84 DB 40 26 9F C8 35 E1 7D 0A B2 4D 68 93 0F E7 5A B4 21 8C D6 39 F2 47 1B 60 CE. C2 interactions begin by sending the following 12-byte handshake: 50 01 13 3F 08 5C 73 7B 1A 72 53 78 (decrypting to “%wGvo4GL62p*” using the XOR key above).
Process names are drawn from an encrypted table and set through argv rewriting (Table 2).
On a SpamSniper appliance, SMTP server-to-server relay traffic is the primary legitimate traffic type the appliance is designed to handle. A firewall in front of the appliance will commonly allow rules such as:
A magic packet withsrc=25, dst=25 would match the first rule and reach the raw socket before any stateful inspection. The implant authors understood exactly what traffic profile would be invisible on this specific class of host.
The command interface relies on the same cryptography (HMAC-SHA1, AES-CBC) and opcodes as the standard Tinyshell/Rekoobe.
By setting variables like VIMINIT=”set viminfo=”, HISTFILE=/dev/null, HISTSIZE=0, andHISTFILESIZE=0, the malware ensures that the attacker’s commands are not logged to bash history nor to vim logs. The reverse shell spoofs “/sniper/autorun/rblsmtpd –start -n 9“ and connects on the attacker’s port 25. Rblsmtpd is a standard daemon used by mail servers (like qmail) to block mail from IPs listed in Real-time Blackhole Lists (RBLs).
A dropper likely built for ShareTech appliances
The dropper (update:2bedc26d4b29b435c21962beed7db21188a0219a0d28334bba8b4fb1656d7b15) is an x86-64 ELF with a minimal import table — fopen, fwrite, fputs, fclose, chmod, system, strlen, sleep, access, memcpy, exit — and no networking. It is a local installer, run after access is already established.
It carries four AES-128-ECB blobs keyed on the first sixteen bytes of SHA1(“ShareTech”), or 6C CA D5 17 0E 3D B8 17 B3 DF 52 E0 D9 71 B1 48. The operators seeded their own key derivation with the target vendor’s name.
Figure 6: AES key derivation seeding routine from SHA1(“ShareTech”)
Execution is gated on one precondition: the install branch fires only when /tmp/flag already exists on the filesystem and the .php script does not. When the gate passes, the dropper writes the script to the appliance’s bulk-storage mount, chmods it 0777, hands it to system(), and exits. Because system() runs sh -c against a file carrying a real shebang, /HDD must be both writable and exec-capable. A noexec mount returns EACCES, for which the shell offers no interpreter fallback.
The script then copies two malicious binaries from /addpkg into /sbin as ntpdate and udevds, runs each, and unlinks them ten seconds later. Both keep running with no on-disk image: /proc/<pid>/exe resolves to (deleted), so there is nothing to hash, quarantine or submit, and a responder grepping /sbin finds nothing at all.
The elegant part is that /sbin/ntpdate is the dropper re-executing itself; the second instance finds the .php already present, fails the gate, and drops into a two-second watchdog that recreates execProcEnd and re-writes the script whenever either disappears. That also explains the absence of any persistence code: /addpkg/sbin/ is the appliance’s own add-on package directory, so the firmware’s package startup very likely relaunches it at boot.
AVERAT: A modular implant reaching into ORB
The dropper copies AVERAT into /sbin/udevds, marks it executable, launches it, sleeps ten seconds, and removes it.
Command and control
The implant connects outbound to port 25 and speaks SMTP: it issues EHLO, requests STARTTLS, and only then begins its own encrypted session. On a mail security gateway, outbound SMTP to arbitrary mail exchangers is the device’s core function, so the traffic is indistinguishable from legitimate work in flow records.
The TLS layer is hand-built rather than linked from a standard cryptographic library. A fixed ClientHello template is compiled into the binary, including a 40-entry cipher suite list and a fixed extension ordering. Peer authentication is deferred entirely to the application layer via a shared-secret handshake carrying the magic value 1571 (0x0623).
Check-ins occur every 600 to 699 seconds. Each reports hostname, current user, OS version, network interfaces, and logged-in users. The interval is stored in a hidden file at /var/lib/.db and can be changed by the operator, persisting across restarts.
Figure 7: AVERAT beacon configuration details and persistent state file path structure
⠀
AVERAT takes its name from its only disk artifact: var, which reads as AVE when XOR-encoded (Figure 7).
Configuration
All operational values are held in a 276-byte encrypted blob. The key is derived from the blob’s own first 16 bytes, folded with a further byte and a reverse XOR cascade; that key seeds RC4’s key-scheduling algorithm, and the resulting S-box is used directly as a keystream. Each field starts at its own keystream offset, and the parity of that offset selects whether bytes are bitwise-inverted or nibble-rotated before the XOR.
Decrypted, the blob yields three 16-byte keys (transport, authentication, and a secondary handshake secret), the host table, the port table, a three-byte build tag, the timing state, and the .db path.
Figure 8: Key derivation and keystream offset mapping
⠀
The schema provides three host and port pairs; this build populates one.
Figure 9: Decrypted AVERAT host and port table configuration slots
⠀
Rapid7 developed an extractor for the AVERAT family. Figure 10 shows the results for the samples identified at the time of writing.
Figure 10: AVERAT extractor results and sample hashes identified at the time of writing
⠀
The MAC key and aux key are byte-identical in all six, while only one stream key is shared between a pair (Figure 10).
Command set
Command codes are uint16 values grouped into bands by subsystem. The most operationally significant are below.
Code
Capability
20
Enumerate directory contents
21
Download a file from the host, with resume support
22
Upload a file to the host in chunks, appending on resume
25
Recursively delete a file or directory tree
30
Recursively walk a directory tree, resolving file ownership
629
Enumerate running processes with command lines
632
Terminate a process (SIGTERM)
842
Overwrite the C2 host and port tables at runtime
912
Open an interactive shell session — up to ten concurrently
914
Write a command into an open shell session
916
Reboot the appliance, flushing buffers to disk beforehand
1010
Load or unload a shared-object module, extending the implant
1576
Set the callback interval and persist it to .db
1618
Open a proxy or port-forward channel through the appliance
unknown
Close the socket and terminate the process immediately
Table 4: AVERAT Command Codes and Capabilities
Infrastructure
The three IP addresses recovered from the configs represent compromised CPE belonging to third-party victims rather than intentional operator assets. Scan data shows all three sitting in Chunghwa Telecom’s HiNet address space (AS3462), in three separate Taiwanese cities — Tainan, Banqiao and Taoyuan — each representing a distinct class of neglected, internet-facing consumer or SMB appliance.
59.125.211.65 is a Synology NAS belonging to a Taiwanese fuel-retail business, still serving a Laravel-based “cloud management system” on 81/82 behind a Let’s Encrypt certificate that expired in October 2021, alongside an exposed MariaDB 5.5.62 instance that reached end of life in 2020. 122.116.138.33 is an embedded Taiwanese ADSL/FTTH SMB network appliance — gSOAP 2.8 on 8000 with a recording-management interface, HTTP Basic realms named SMB on 8081, 8082 and 10443, and a self-signed NetKlass Technology certificate valid from 2004 to 2014, MD5-signed with a 1024-bit key. 1.34.200.85 is a Dahua DH-XVR5116HS-I3 recorder. These are victim hosts repurposed as operational relays, selected on consistent criteria: reachable, unpatched, unmonitored, and unlikely to be audited.
The detail that binds them is PPTP on 1723, present on all three, returning a byte-identical banner (Firmware: 1, Hostname: local, Vendor: linux, fingerprint 261189147). A Dahua XVR ships no PPTP server. Synology’s DSM offers one only as an optional package, disabled by default and deprecated in current releases. We assess this as operator-installed, which makes each node dual-purpose: an outbound relay terminating SMTP-disguised implant traffic, and an inbound routed VPN foothold into the host’s own LAN. Combined with the RAT’s own proxy commands, the campaign has relay capability at both ends of the connection.
The campaign is running two distinct C2 addressing strategies:
Strategy
Samples
Trade-off
Attacker-registered domain
bf8135f4, 2fe2dd40, a65048eb
Survives IP churn, re-pointable via DNS — leaves a seizable, sinkholable, monitorable name
Hardcoded consumer-broadband IP
4925bcca, a4379e11, 925c0418
No DNS artifact at all, brittle against address rotation
Table 5: C2 Addressing Strategies Comparison
The relay layer is composed of consumer and small-business broadband CPE: a fuel retailer’s NAS, an obsolete NetKlass appliance, and a CCTV recorder, each sitting on a domestic-grade line.
AVERAT’s command-and-control infrastructure matches the device-class profile that CISA, NCSC-UK, and partner agencies described in their April 2026 joint advisory (AA26-113A) as the standard building blocks of China-nexus covert/ORB networks: end-of-life NAS, edge appliances, and DVRs chosen because their vulnerabilities will never be patched. We found no infrastructure or indicator overlap with any specific named ORB network (LapDogs/UAT-7810, SPACEHOP, or FLORAHOX), whose documented targeting instead centers on SOHO routers; AVERAT’s infrastructure is consistent with the broader ORB device-class pattern rather than confirmed membership in a known network.
Operator tradecraft
Three design decisions indicate operational maturity.
The dispatcher terminates the process on any unrecognized command code. This is inexpensive to implement and raises the cost of interactive probing or automated scanning against a live implant.
Support for ten concurrent shell sessions is consistent with provisioning for parallel operator access rather than a single interactive session..
The reboot handler calls sync() before forcing a restart through the kernel rather than through init. Flushing filesystem buffers before destroying volatile state is the behavior of an operator who knows precisely which artifacts persist across a restart and which do not — and in this chain everything happens in memory while the persistence needed to re-establish access is on disk.
Detectionguidance
File-based detection on the appliance is unlikely to succeed, because no payload persists in /sbin. We recommend prioritizing the following.
On the host, hunt for processes whose executable has been unlinked, which on Linux presents as a (deleted) suffix on the /proc/<pid>/exe target. To catch fileless and unlinked process execution, monitor process descriptors for instances where /proc/<pid>/exe points to an unlinked path, and inspect memory maps for executable pages lacking backing file paths on disk.
Additionally, alert on the presence of .db state files and dropper artifacts: the directory /HDD/ms6x2xTo64/, a shell script carrying a .php extension whose first bytes are #!/bin/sh, and a marker file execProcEnd containing the literal string end. A process-tree sequence of sh -c on a .php path, followed by cp and chmod into /sbin and an rm -rf of the same path within roughly ten seconds, provides high-fidelity detection of the staging sequence.
On the network, the fixed ClientHello template means the implant’s TLS fingerprint does not vary between infections. Fingerprinting it is more durable than watching the port, because the port is configurable at runtime through command 842 while the template is compiled in. Outbound SMTP from an appliance to mail-role hostnames that resolve to consumer-grade or embedded devices warrants investigation on its own.
Mitigation
Investigate unexpected raw packet sockets and classic BPF filters on Linux systems that do not require packet capture. Review outbound TCP port-25 callbacks from processes that are not mail services, particularly when the process renames itself to a common daemon or creates hidden PID and socket markers. Preserve short-lived staged binaries and collect process arguments, open file descriptors, socket metadata, and historical DNS records. Restrict management access to routers, DVRs, and other edge appliances, and monitor NFS or SMB mounts that could let an adjacent host write executables onto an embedded device.
YARA rules and more IoCs are available in Rapid7’s Intelligence Hub along with ongoing intelligence on the latest campaigns.
What defenders should take away from these campaigns
The components form a modular access ecosystem. The dropper executes AVERAT, which beacons to changeable infrastructure, while BPFDoor and Rekoobe samples wait for a magic packet before opening interactive access.
Across both campaigns, the network edge is a consistent focus. Each targets mail-security appliances that sit inline in front of the mail server, giving an implant positioned there visibility into an organization’s inbound and outbound traffic. Both also use port 25 to blend into expected SMTP activity, although the mechanism differs between the campaigns. In either case, command-and-control traffic can hide within a protocol that is normal for the device and may therefore attract less scrutiny.
This fits the broader BPFDoor pattern, where compromised IoT and SMB devices, including NAS units and DVRs, can act as operational relays that obscure the true source of magic-packet traffic before it reaches the passive backdoor. Newer samples also show how the malware continues to adapt to the environments it targets, including a second userland-level magic-packet check layered on top of the kernel BPF gate.
Rapid7 Variant G provides another example of that adaptation. It uses three BPF filters to preserve operational resilience on high-traffic edge nodes, with the filters left unoptimized because the libpcap version running on its end-of-life targets does not support filter optimization.
For defenders, the most useful detection opportunities remain raw packet sockets, BPF filters, port-25 callbacks from unexpected processes, process masquerading, and appliance-specific staging paths. Specific attribution should remain an ongoing assessment as new samples and infrastructure emerge.
Today, we’re launching eight major updates that bring your logs, traces, analytics, alerts, dashboards, and exporting into one observability platform, with simpler and more predictable pricing.
Understanding an issue often requires data from more than one Cloudflare product. A spike in 5xx responses could come from a Worker, from your origin, or from Cloudflare failing to connect to your origin globally or regionally. But investigating it today requires knowing which product owns each signal and how to query it.
Observability should be a platform-wide capability: it should reflect how applications actually behave and give you the complete context needed to resolve an issue. Over the coming months, you’ll see more Cloudflare products, datasets, and workflows become part of this shared observability platform, with more consistent pricing, product experiences, and features. These eight updates are the first step into a more unified Observability problem.
1. Investigate all your logs in one place
The new Logs home combines Workers Observability (for debugging Workers applications and its connected resources) with Log Explorer (for searching across security logs). You can now choose from log datasets like HTTP events, firewall events, Workers, Containers, R2, and AI Gateway, and use the same investigative tools and capabilities for each.
Start with an increase in request latency, group it by hostname or data center, narrow the results to affected paths, and inspect individual requests by Ray ID. If the investigation leads to another Cloudflare product, switch datasets without leaving Logs. Support for querying across multiple datasets is coming soon, making it possible to connect related events across products in a single query.
You can query your logs with raw SQL or with built-in filters to narrow down on specific events. Create visualizations with natural language, and easily investigate and understand detected anomalies.
2. Trace requests through our entire platform — now in open beta
We’re launching Cloudflare Traces in open beta, giving you a request-level view of supported security rules, transformations, cache decisions, routing, Workers, and origin handling. You get to see how your traffic moved through our platform, and connect the dots between how you’ve configured Cloudflare, and how this influences request processing time, routing decisions, and more.
Set a baseline sampling rate for continuous visibility, then use Trace Rules to capture specific traffic at a higher rate during an investigation. Target hostnames, paths, IP addresses, or headers, search by Ray ID, and inspect the resulting spans directly in the Cloudflare dashboard.
3. Have your agent query observability data with one unified SQL API
Agents also need a consistent way to sift through your observability data, investigate issues, correlate signals, and verify fixes. We’re launching a unified SQL API, now in beta, for querying telemetry across Cloudflare. Instead of integrating separately with Workers logs, Containers security events, HTTP request logs, and analytics data, people and agents can query them using one SQL dialect, authentication model, and API.
Additionally, we’re also bringing the SQL interface directly into Workers with a native binding. Your Worker can now do things like query Analytics Engine data to meter customer usage and power billing workflows, build customer-facing analytics dashboards, generate health reports, or automate incident investigation without configuring a separate API client.
4. New pricing for all ingested and stored logs and traces
For all logs and traces ingested and stored on Cloudflare, we are moving to one unified Observability subscription and pricing. Beginning December 1, 2026, this pricing model will apply across all plans (effective upon renewal for all Enterprise customers) and cover existing Developer Platform logs, including Workers, Containers, AI Gateway, as well as all tracing data.
Because logs and traces can vary dramatically in size, the new model is based on the volume you ingest and store rather than an event-based count. This pricing adjustment will be. Check out our documentation for more details on pricing.
Plan
Included Usage
Retention
Additional usage
Free
0.5 GB of ingestion per day
7 days
Not available
Paid and Enterprise
50 GB of ingestion 10 GB-month of storage per billing cycle
Up to 1 year (coming soon)
$0.25 per GB ingested $0.10 per GB-month stored
5. Configure custom alerts on your observability data – now in beta
Notifications (now called “Alerts”) just got a major upgrade. You can now define custom alerts directly on anything supported by our new unified SQL API, including HTTP request logs, Workers events, Workers Analytics Engine datasets, analytics datasets, traces, and security events.
Choose a dataset in the dashboard or define the condition using custom SQL. Then select a threshold, anomaly, or SLO, set the evaluation window, and choose where the alert should go. You might alert when origin 5xx responses exceed a threshold for five minutes, a Container repeatedly fails, Worker errors increase after a deployment, or trace latency crosses an expected limit.
You can send alerts right to tools your teams are already using, including incident management tools, chat platforms, and webhooks. Webhooks are now available on all plans, allowing you to route alerts to custom services or even your agent to begin investigating immediately. To get started check out our documentation or give this command to your agent:
6. See your domain analytics in one place — now with 30 days retention
Understanding what is happening on your domain has often meant piecing together metrics from different Cloudflare products. We’re bringing traffic, performance, security, cache, origin, and DNS data together so you can see how they relate. If latency increases, you can quickly see whether it is tied to a specific Cloudflare data center, hostname, or origin.
In addition, you now get 30 days of domain analytics on every plan. A full month of history gives you time to investigate issues after they happen, compare today with the same day in previous weeks, and tell the difference between a one-time spike and a longer trend.
7. Build custom dashboards
Prebuilt dashboards cover common use cases, but applications often use several parts of Cloudflare. With Custom Dashboards, you can bring together analytics from across Cloudflare, logs and traces from the Workers platform, and security events in one view. Track request volume, errors, latency, storage, and blocked traffic, then share the dashboard with your team. Instead of rebuilding queries during every investigation, you have one place to monitor the signals that matter to your application.
8. Logpush is now available on all self-serve plans
Logpush, previously available only to Enterprise, is now available on all self-serve plans, letting you export all Cloudflare logs to the tools and destinations you already use. Need to apply filters, perform redaction, enrich events or reshape output before delivery? Transformers is now generally available, letting you apply any SQL transformation without operating a separate ETL pipeline.
We’re introducing usage-based pricing for Logpush and Transformers. Each includes a free monthly allowance, with simple pricing for additional usage:
Export usage
Included each month
Additional usage
Exports to Cloudflare destinations
25 GB
$0.03 per GB
Exports to external destinations
25 GB
$0.10 per GB
Logpush Transformers
1 GB
$0.04 per GB
Visit the documentation to get started with Logpush and explore complete pricing details.
What's coming up:
Longer retention for your observability data: You’ll be able to retain logging and tracing data for up to one year, making it easier to investigate recurring issues, compare historical behavior, and analyze long-term trends.
OpenTelemetry API support in Workers: We’ll continue building out our OpenTelemetry APIs to enable adding attributes to existing spans or getting trace context.
Easier metrics export with OpenTelemetry: You’ll be able to send Cloudflare metrics to OpenTelemetry-compatible destinations and analyze them alongside telemetry from the rest of your stack.
New pricing takes effect December 1, 2026: If you ingest or store observability data on Cloudflare, the unified pricing plan will apply to your usage. We’ll notify you before the change takes effect.
Ready to start investigating?
We hear you when you say Cloudflare can feel like a black box. These updates are just the beginning of exposing what’s happening, making the underlying data accessible, and giving you the context that you need to act. That transparency matters even more as agents move from writing software to operating it. An agent can only close the loop between a change and its outcome if it can query what happened, identify the failure, and verify the fix.
A year ago, Cloudflare CTO Dane Knecht announced our intention to make every Cloudflare feature available to everyone. Cloudflare launched an Enterprise tier years ago when larger customers came to us looking for procurement options beyond a credit card, like invoices, custom contracts, and dedicated support. Those offerings met a customer need but over time, a two-tier system developed where some of our most advanced and powerful features were only available to Enterprise customers. Our goal was to close that gap.
Today, teams of every size use Cloudflare, from Fortune 100 enterprises to small businesses, open-source projects, and individuals. Across the platform, we’re committed to ensuring that every user or team can make use of all of Cloudflare’s capabilities in a way that helps their organization thrive.
The underlying philosophy is that Cloudflare should offer products suitable for our most demanding customers — and make those capabilities available to everyone. Large or small, every customer would prefer not to have to call support. Building products that are easy to buy, configure, and consume means more of our products in use and a step closer to a better Internet for everybody.
Every generally available (GA) feature we launched this week that is available on an Enterprise plan is also available to Pay-as-you-go customers, and most are available on the free tier. Where our plans differ, it's in how much you can use, not what you can use. While we haven’t yet met our goal that every feature be available to everyone, in the year since Dane’s announcement, we’ve made great progress.
Here are a few products and features making the transition today from Enterprise to everyone.
Logpush and Logpush Transformers now available to all plans
Flexibility on pushing logs to third parties and how logs are formatted expanded this week from Enterprise-only to all customers.
Logpush delivers Cloudflare logs to storage, security, and analytics destinations, helping customers monitor traffic, investigate issues, and analyze their data using existing tools. Previously available only to Enterprise customers, Logpush is now available to Free, Pro, and Business customers through self-service, pay-as-you-go pricing. Datasets available to Logpush have been expanding as well. We’ve recently added account-scoped firewall events, WebSocket analytics and per-zone post-quantum visibility.
Transformers is also becoming generally available to all customers. With Transformers, customers can use SQL to filter unnecessary records, redact sensitive information, enrich events, and reformat logs before delivery without operating a separate extraction, transformation and loading (ETL) pipeline. Together, Logpush and Transformers give every customer greater control over how their Cloudflare data is prepared and delivered.
In addition, Custom Dashboards which let customers create personalized views highlighting the metrics most critical to them, is now available to all customers.
New tools for managing Cloudflare at scale
Expanding RBAC
Over the last year, we’ve dramatically expanded the availability of Role-Based Access Control (RBAC) across all Cloudflare products and for all customers. Today, nearly all products have RBAC roles available at the account and zone level. Recently, Workers joined R2 and Access in having RBAC roles available at the individual resource level as well, so Administrators can decide who on their team gets specific access to individual Workers.
Multiple Accounts
While fine-grained RBAC lets customers manage subsets of an account, this setup still relies on a small number of super administrators making choices about who gets access to what. Centralized authority works great when your problem space is small, but as the number of teams and projects being managed on Cloudflare grows, it can turn into an organizational bottleneck.
The single account model is excellent in its simplicity, but it can start to feel a little crowded for customers maintaining hundreds or thousands of zones, workers, and storage products. That’s why we’ve been expanding our capabilities around managing multiple accounts.
New Account button
Last month, we quietly launched the New Account button on the dashboard that, for the first time, lets users create additional accounts directly. The response has been overwhelmingly positive, and we’re seeing thousands of customers branching out into additional accounts every week. When you use this button, it creates a new, free, Cloudflare account that you can use to segment your open source projects, or segment the work of multiple teams in your organization. Each of these accounts is independently billed, so you can segment spending across multiple cost-centers directly. Safeguards are in place to prevent fraud and abuse.
New Accounts for Enterprises
While the New Account button is for everyone, for the time being, we recommend that Enterprise customers reach out to their account team to get new accounts provisioned instead. This lets you reuse your existing enterprise agreement and subscriptions across all of your accounts. There is no preset limit on how many accounts an enterprise can request. We will be adding additional features in the future that make this process self-serve for enterprises too.
Organizations
Once you’ve created multiple accounts, how do you organize and track them all? Organizations allow customers to group accounts together with a single analytics and shared configuration surface. It’s in beta for Enterprise customers now, will be GA in October, and will be rolling out to free accounts in early 2027. Adding your multiple accounts to a single organization makes managing them easier by providing a unified surface for visibility and management. Organizations provide shared administrators with unified analytics and audit logging as well as shared WAF, Gateway, and Access IdP configurations.
Enterprises are eligible for exactly one organization. We limit enterprises to a single organization, so there’s a single pane of glass that shows all the company’s assets in one place. This makes life easier, so you can invite the CISO, CTO, or other executive stakeholders and give them unified visibility. If you’re an Enterprise customer and haven’t tried organizations yet, you can set one up directly as long as you are a super administrator of at least one account and nobody else has already created the organization. If the organization has already been started, talk to the other Cloudflare administrators in your company to get your accounts added to it. This process ensures that there’s never an elevation of privilege as we layer on this new management plane.
Terraform and Tags
Once a customer has created multiple accounts, an organization to manage them, and set RBAC rules for the products and resources they contain, they need to be able to manage them in a way that’s auditable and repeatable. Terraform lets customers use Infrastructure as Code to manage everything using version-controlled code rather than clicking on the dashboard in a way that may not be repeatable. In the last year, Cloudflare has made dramatic progress creating a Terraform provider that is built programmatically, so it’s always up-to-date with the latest version of the Cloudflare API. Terraform, like the other features mentioned in this post, is available to all customers, Enterprise and not.
Additionally, Resource Tagging lets customers apply key value tags to a very broad set of resources within the Accounts and Organizations. Today tags can be produced interactively or via API and are useful for organizing resources in the dashboard. In the future we intend to make tags useful in billing and access control scenarios and to be manageable via Terraform.
How we use it all at Cloudflare
With the increasing menu of enterprise-ready options for everyone, one of the top questions we get is “What does Cloudflare do internally?” Within Cloudflare, we create accounts per team, or per service, depending on the nature of the team. We then use Terraform to manage account access and production configuration, giving teams a peer-reviewed, auditable path for changes. Because the scope of each account is narrow, we can grant broader permissions to the engineers responsible for that account while keeping the blast radius contained. This lets teams grow their accounts organically without bottlenecking on a small number of central administrators, and it makes operational work like on-call response faster and safer.
Every account at Cloudflare lives within Cloudflare’s organization, which provides our security team with administrative access to every account within the organization, as well as analytics, policy management, and shared configurations. This makes it easier to align every account in the organization to our security standards. Our teams have the right blend of autonomy and centralized control to go fast.
Enabling teams to quickly sort, organize, and filter their resources is critical in our production environments. While it’s still early, Resource Tagging is enabled internally and teams have begun to roll out tags to make finding the WAF rule, R2 bucket, etc. that they need to interact with easier.
More features for everyone
We launched support for the Authentik identity provider (IdP), SCIM Audit logging, and SCIM 2.0 Group Sync. MCP Server Portals moved into general availability. All these features were once in some way Enterprise-only. Even network management is going self-serve: the Network Overview page and Unified Routing both recently became available for all.
Starting with free
Solving big problems starts with first ensuring they aren’t getting any larger. This year, as part of Code Orange: Fail Small, we announced a commitment to rolling new code out by traffic cohort, starting with our free customers. As a result, today we are committed to introducing no new Enterprise-only features. Naturally there will be some carve-outs for things like Cloudflare for Government that are inherently Enterprise-oriented in nature.
Other progress for free and pay-as-you-go customers
Beyond making previously enterprise-only features available to everyone, we’ve also done a lot of work to make Cloudflare more powerful and accessible for everyone
Billable Usage Dashboard and API
In August, we introduced the billable usage dashboard and API which lets non-Enterprise customers see how much they’ve spent and download their consumption data to use offline directly or through third-party tools like Vantage. We also introduced budget alerts, which are on by default to prevent unpleasant billing surprises. We're prototyping hard spending caps now, with early availability in Q4 2026. Because Enterprise customers have dramatically more variation on contract terms and how they pay, this experience is not yet available to Enterprise customers, but we are hard at work and expect to have an announcement in 2027.
Higher limits available to all customers
Over the past year we’ve increased limits across Cloudflare products. We’re constantly working to increase these defaults, and keep our front door as open as possible to people building the next big thing.
Between exposing formerly enterprise-only features to everyone and increasing the power of features that were already available to everyone, Cloudflare is committed to building the most powerful and accessible platform for customers large and small without the need for a contract. We still have much work to do on Dane’s pledge from a year ago, but we are committed to getting there and are delighted to be able to highlight our progress over the last year.
Take advantage of these new offerings
Create additional accounts to partition the concerns of your organization.
Use RBAC to define security policies at the zone and account level.
If you’re an Enterprise customer, create an Organization and onboard these accounts. For other customers, we’ll see you in early 2027.
Use Terraform to manage the state across your whole organization.
Attend Cloudflare Connect next month to learn more about everything discussed here and meet the team that built it.
Today, end users carry too much of the burden of online privacy. To avoid third-party trackers or targeted ads, users are instructed to use a VPN, disable cookies, or install adblockers. Meanwhile, some app developers end up knowing more about their users than they’d care to: a typical client-server exchange creates a trail of user data, like the client’s IP address or TLS fingerprint. This level of visibility can be a burden.
That’s why Cloudflare builds infrastructure that helps developers bake privacy into their apps. Oblivious HTTP (OHTTP) is an IETF standard designed to enable app backends to receive HTTP requests without seeing user IP addresses.
This fall, we’re launching the Cloudflare OHTTP Gateway. Customers will be able to enable our new OHTTP Gateway as a paid add-on to their zone and start receiving OHTTP traffic with just a few clicks. Register through our form to join our waitlist. Read on to learn more.
Expanding our OHTTP product suite
With OHTTP, requests travel through two independently-operated hops: a relay and a gateway. An OHTTP relay blindly forwards encrypted requests in order to hide client identifiers from app servers. An OHTTP gateway performs the cryptographic work of decapsulating encrypted requests and encapsulating responses such that app servers can handle OHTTP requests as if they were plain HTTP. The separation of trust between relay and gateway is critical: it ensures that no single party sees both client identifiers and request contents.
In 2022, we launched an OHTTP relay product, Privacy Gateway. Privacy Gateway enables our customers to offer more privacy-preserving experiences to their users. For example, Flo Health uses OHTTP for their app’s Anonymous Mode, and Apple’s Private Cloud Compute uses OHTTP to disassociate AI inference requests from user identities. But customers who are already protecting their servers behind Cloudflare can’t also use a Cloudflare-operated relay — they need an OHTTP gateway instead.
In our experience running OHTTP relays, we’ve seen how difficult it can be to build and operate a secure, performant OHTTP gateway at scale. Today, we’re launching the closed beta for our self-serve Cloudflare OHTTP Gateway. We’re also renaming our “Privacy Gateway” to “Cloudflare OHTTP Relay” to better distinguish the two products.
Now, customers who want an OHTTP architecture with the necessary separation of trust have two options:
Use Cloudflare’s OHTTP Relay (formerly Cloudflare Privacy Gateway) and run your gateway yourself. This is best if your application servers are hosted off Cloudflare, and you’re able to run your own OHTTP gateway.
Use Cloudflare’s new OHTTP Gateway with a third-party relay. This is best if your app servers are already behind Cloudflare (on our CDN or Workers, for example), if you’re accepting OHTTP requests from a third party (like Apple’s LiveCallerID), or if you want a managed gateway to minimize latency and operational overhead.
We’re working to raise the bar for privacy across the Internet, and we believe that protocols like OHTTP can help — if we make them easy enough to adopt. It’s always been our goal to expand our OHTTP product suite and make our trusted privacy infrastructure accessible to a broader swath of the Internet.
Why we built the Cloudflare OHTTP Gateway
Since we launched our OHTTP Relay product, we’ve observed a few things.
First, we’ve seen that there's a growing appetite among developers for accessible, usable privacy infrastructure. Developers of privacy-oriented apps want to bake network privacy into their applications by default, but doing so remains harder than it should be.
Second, we’ve learned that building and operating an OHTTP gateway can be tough for customers. Any proxying architecture introduces some latency because requests must travel an extra hop or two around the Internet. Combine that with the cost to decrypt requests and encrypt responses, and the latency hit of a homegrown OHTTP setup can be significant. We’re well-positioned to solve this problem: the same building blocks that enable us to operate fast, reliable privacy infrastructure for products like 1.1.1.1 and iCloud Private Relay make us a good home for an OHTTP gateway. Because of Cloudflare’s anycast approach, our OHTTP Gateway will run on every server on Cloudflare’s global edge network, minimizing latency in relay-to-gateway hops. If you use our CDN, user requests can be decrypted by our Gateway and resolved by your app servers on the same Cloudflare metals, saving gateway-to-origin latency.
Finally, recall that OHTTP’s privacy model requires that the relay and app server be operated by separate, non-colluding parties. We want to provide our customers with the best possible range of options for their privacy infrastructure. Before, developers who protected their app servers behind Cloudflare weren’t able to use our OHTTP Relay, because Cloudflare would see both client metadata and the decrypted contents of requests, breaking OHTTP’s privacy model. Now, developers can choose whether a Cloudflare OHTTP Relay or Gateway is a better fit for their architecture.
A primer on OHTTP
A typical interaction between a client and application server reveals information about the client. When a client and app server talk to one another, the app server learns the client’s IP address because each packet in which data is sent is labeled with a source IP — similar to the “from” label on an envelope. App servers can also “fingerprint” a client based on attributes like supported TLS versions or cipher suites. These signals make it possible for app servers to link multiple requests back to the same user.
But what if I wanted to build an app that really doesn’t know much about my users? For example: Flo Health wanted to build an Anonymous Mode to enable users to access personal health data without it being linkable to possible user identifiers.
OHTTP introduces a proxy, called a “relay,” that forwards requests and responses between client and app server to obfuscate the client’s identity from the app server. The relay sees client identifiers like IP address and TLS fingerprint, but strips them before forwarding on requests. This prevents app servers from linking multiple requests back to the same user, and means that request contents can’t be associated with the user’s IP address.
For example, a regular client-server exchange might reveal the following information about a client:
A request first sent through an OHTTP relay would reveal only the relay’s information to the app server receiving the request:
This means that for each request, the app server doesn’t learn the location and TLS fingerprint of the end user. Plus, if many different users are sending requests through the relay, the app server won’t be able to distinguish which requests are coming from whom, limiting their ability to trace app activity back to a single end user. This creates a strong privacy boundary.
What really differentiates OHTTP from a basic forwarding proxy, however, is the encryption of data between client and app server. Requests and responses are encapsulated using Hybrid Public Key Encryption (HPKE) such that only the client and app server can see plaintext, and the relay sees only a jumble of ciphertext. A “gateway” sits between the relay and app server to handle all of this cryptography — decapsulating requests, encapsulating responses — and the app server handles only plain HTTP.
This creates a “double-blind” privacy model: the relay sees only client identifiers; the gateway and app server see only request contents; no party sees both.
How we built the OHTTP Gateway
In building our OHTTP gateway-as-a-service, our goal is to bring our secure, performant privacy infrastructure to a broader swath of the Internet. Performance and easy onboarding are critical. So, we built our Gateway as a flexible service deployed across our global network. With just a couple of clicks, you can enable the Gateway on your zone and start sending OHTTP to https://your-zone.com/.well-known/ohttp-gateway. We’ll scale the service up and down automatically, so you don’t need to worry about capacity.
We had a few other user needs in mind, informed by the pain points we’d seen OHTTP Relay customers run into when operating their own OHTTP gateways.
First: We wanted to abstract away as much of the complexity of OHTTP as possible for your app servers. We wanted developers to be able to start receiving OHTTP while continuing to accept regular HTTP traffic if they chose. So, we designed the Gateway as a feature of your zone, where clients send well-formatted OHTTP requests to a /.well-known/ohttp-gateway endpoint on your zone. We support both standard and chunked OHTTP — and we recommend using chunked OHTTP for better performance, because it enables us to process requests incrementally (in “chunks”).
Our Gateway service will intercept each request, decrypt it, issue a subrequest to your app server, and return an encrypted response to the client. All non-OHTTP requests will travel to your server without invoking the Gateway.
Binding your Gateway to your zone also enables us to protect your Gateway from abuse. A client sending requests to your zone `example.com` may send to `foo.example.com` or `bar.example.com`, but not wikipedia.com. Without you needing to worry about it, this prevents unauthorized clients from using your zone as a way to target other domains.
Second: Seamless key management is critical. Gateways need to maintain a public HPKE key configuration to enable clients to encrypt requests, but managing keys securely is a challenge. So, we designed the Gateway to fully manage all keys for customers, and to serve public keys as responses to GET requests to /.well-known/ohttp-gateway. For stronger privacy, clients can download keys over a different IP than they request the gateway.
Third: Gateways need to be able to authenticate relays. Because the Gateway (by design) knows very little about the client sending a given request, it places trust in the relay to authenticate clients and forward traffic responsibly. But how do you ensure that only trusted relays can send traffic to your gateway?
We designed the Gateway such that Cloudflare Access, Cloudflare’s zero trust network access product, runs before requests are decrypted, enabling you to use any standard Access policies to authenticate incoming traffic and protect your Gateway from abuse. Options include mutual TLS, static service credentials, and custom external logic.
Finally: Mistakes happen, and we anticipated that customers might accidentally break OHTTP’s privacy model by running both their relay and gateway on Cloudflare. So, to preserve OHTTP’s separation of trust and ensure that Cloudflare never sees both client identities and decrypted inner requests, our Gateway will refuseto decrypt requests sent from Cloudflare Workers or from proxied hosts on Cloudflare.
When is the OHTTP Gateway a better fit than the OHTTP Relay?
If you want to use Cloudflare’s OHTTP product suite, but you’re wondering why you’d pick Cloudflare’s OHTTP Gateway instead of the OHTTP Relay, here are a couple of considerations.
First, do you want your app servers on Cloudflare – behind our CDN or built on Workers, for example? If so, the OHTTP Gateway is a better fit to ensure adherence to OHTTP’s privacy model.
Second, what’s your use case? If you want to receive OHTTP requests from a third-party client and relay — to use Apple’s LiveCallerID SDK, for example — then the OHTTP Gateway is likely the better solution for you.
Getting started
If you have a feature request or would like to register for our waitlist, so we can notify you when the product launches, sign up here.
Then, you’ll need to implement an OHTTP client. See ohttp.info or our sample client library for some examples to help you get started. One flag as you build the client: OHTTP provides privacy at the network level, and doesn’t touch the inner request body. So, to preserve user privacy, it’s up to you not to send identifying information (e.g. a user’s email address or username) in the request body.
Next, you’ll need to bring your own relay. Relays can run on any infrastructure provider, and they’re simple: here’s some sample code. The challenge and the reason you might want a dedicated OHTTP relay provider, is to verifiably promise to your users that you won’t inspect logs with client identifiers. Otherwise, you’d be able to correlate clients at the relay with decrypted requests at your app servers.
Finally, once your OHTTP deployment is live, check out our pvcli client to help with testing and debugging.
We’re excited to bring accessible privacy infrastructure to developers everywhere. Reach out to us if you’d like to try out the new OHTTP Gateway and raise the bar for privacy online.
The collective thoughts of the interwebz
Manage Consent
To provide the best experiences, we use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us to process data such as browsing behavior or unique IDs on this site. Not consenting or withdrawing consent, may adversely affect certain features and functions.
Functional
Always active
The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.
Preferences
The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.
Statistics
The technical storage or access that is used exclusively for statistical purposes.The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
Marketing
The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.