Tag Archives: Advanced (300)

Deploy Oracle Database step by step on Amazon EVS with FSx for ONTAP

Post Syndicated from Satish Bhoi original https://aws.amazon.com/blogs/architecture/deploy-oracle-database-step-by-step-on-amazon-evs-with-fsx-for-ontap/


This post provides step-by-step procedures to deploy Oracle Database on Amazon Elastic VMware Service (Amazon EVS) with Amazon FSx for NetApp ONTAP as NFS datastore storage. You will provision storage volumes, mount NFS datastores, install Oracle, and configure SnapMirror replication for cross-region disaster recovery.

Enterprises running Oracle databases on VMware want a path to AWS that preserves their existing operational workflows with no rearchitecting and retraining. In our related post, Architect highly available Oracle Database on Amazon EVS and FSx for ONTAP, we explained how to design that environment: selecting Amazon Elastic Compute Cloud (Amazon EC2) bare metal instances, sizing VMs, splitting storage between vSAN and FSx for NetApp ONTAP, and planning SnapMirror replication for cross-region DR.

This post picks up where the architecture left off. We provide step-by-step procedures to deploy the entire stack — from provisioning your first Oracle VM and creating FSx for ONTAP volumes, through mounting NFS datastores, installing Oracle 19c, and configuring SnapMirror and SnapCenter. We also cover four migration paths for moving existing on-premises Oracle workloads to EVS and day-2 operations including snapshot backup, point-in-time recovery, and database cloning.

For architecture decisions, instance type selection, storage design rationale, and high availability planning, see our related post: Architect highly available Oracle Database on Amazon EVS and FSx for ONTAP.


Step-by-step deployment procedures

This section walks through the end-to-end deployment workflow: provisioning the Oracle VM, creating FSx for ONTAP storage volumes, mounting NFS datastores in vSphere, installing Oracle 19c, and configuring SnapMirror for cross-region DR. Complete the prerequisites first, then follow Steps 1 through 8 in order.

Prerequisites


Step 1: Deploy Oracle VM on EVS

  1. Log in to the vSphere Client connected to your EVS vCenter.
  2. Create a new Virtual Machine in the Production DB Cluster:
    • Guest OS: Red Hat Enterprise Linux 8 (64-bit) or Oracle Linux 8.
    • vCPU: Size per Oracle workload (8–32 vCPU typical)
    • Memory: 32–256 GiB based on SGA/PGA requirements.
    • Disk: 100 GiB on vSAN datastore (OS + Oracle Home + swap + temp tablespace)
    • Network: Attach to DB segment on the prod-trusted Tier-1 gateway.
  3. Power on the VM and configure the guest OS:
# Set hostname
hostnamectl set-hostname ora-db1

# Create swap on vSAN-backed disk (local NVMe, single-digit ms latency)
# The vSAN datastore is already available as the VM's primary disk
# Allocate a dedicated partition or LV for swap
lvcreate -L 16G -n swap vgos
mkswap /dev/vgos/swap
swapon /dev/vgos/swap
echo "/dev/vgos/swap swap swap defaults 0 0" >> /etc/fstab

# Install Oracle prerequisites
sudo yum install -y oracle-database-preinstall-19c python3

For HA/DR, consider replicating the entire Oracle VM rather than maintaining a separate licensed instance in the DR cluster (see DR Licensing Consideration in the architecture post).

Oracle licensing consideration: Instance type selection affects Oracle license cost, which is an important factor to take into consideration. We recommend requesting an AWS Optimization and Licensing Assessment (AWS OLA) for further guidelines.


Step 2: Provision FSx for NetApp ONTAP

  1. Open the Amazon FSx console, select Create file system, then select Amazon FSx for NetApp ONTAP.
  2. Select Standard create and configure:
Setting Value
Deployment type Single-AZ (required for EVS)
SSD storage capacity Size for Oracle data + logs + 20% headroom
Throughput capacity 512–2,048 MB/s (size for write workload. See asymmetry note)
IOPS Automatic (3/GiB) or user-provisioned (up to 80,000)
VPC Same VPC as EVS environment
Subnet EVS service access subnet, same AZ as DB cluster
Security group Allow NFS (TCP 2049, 111, 635) from EVS management VLAN
  1. Set the fsxadmin password (required for ONTAP CLI automation).
  2. Create an SVM (Storage Virtual Machine) with vsadmin password.
  3. Disable automatic daily backups. Use SnapCenter for Oracle-aware scheduling instead.
  4. After creation, select SVM, select Endpoints, and copy the NFS DNS name.

Step 3: Create Oracle database volumes

Connect to the FSx for ONTAP cluster using SSH (ssh fsxadmin@management-endpoint) and create volumes:

# Oracle binary volume
vol create -volume oradb1bin -aggregate aggr1 -size 50G \
  -state online -policy default -tiering-policy none \
  -junction-path /oradb1bin

# Oracle data volume
vol create -volume oradb1data -aggregate aggr1 -size 500G \
  -state online -policy default -tiering-policy none \
  -junction-path /oradb1data

# Oracle log volume (redo + archive)
vol create -volume oradb1log -aggregate aggr1 -size 250G \
  -state online -policy default -tiering-policy none \
  -junction-path /oradb1log

Set -tiering-policy none to pin all data to SSD tier. Size the log volume for 24 hours of archive logs.


Step 4: Mount FSx for ONTAP as NFS datastore in vSphere

  1. In vSphere Client, select the DB Cluster, then select Configure > Storage > New Datastore.
  2. Select NFS, then select NFS 3.
  3. Enter:
    • Server: svm-id.fs-id.fsx.region.amazonaws.com
    • Folder: /oradb1data
    • Datastore name: fsx-ora-db1-data
  4. Repeat for binary (/oradb1bin) and log (/oradb1log) volumes.
  5. Verify all three datastores show correct capacity in the cluster storage view.

Step 5: Create Oracle VMDKs on FSx for ONTAP datastores

With the NFS datastores mounted at the ESXi host level (Step 4), create virtual disks for Oracle on these datastores:

  1. In vSphere Client, select the Oracle VM, select Edit Settings, then select Add New Device > Hard Disk.
  2. Create these VMDKs:
VMDK Datastore Size Guest Mount Purpose
Hard Disk 2 fsx-ora-db1-data 500 GiB /u02 Oracle data files
Hard Disk 3 fsx-ora-db1-log 250 GiB /u03 Oracle redo + archive logs
Hard Disk 4 fsx-ora-db1-bin 50 GiB /u01 Oracle Home binaries
  1. Select Thick Provision, Eager Zeroed for data and log VMDKs (best Oracle performance).

Oracle Database can be created on Oracle ASM or Filesystem (local/NFS). This installation is based on creating the Oracle database on local XFS filesystem.

Inside the Oracle VM guest OS, partition and mount the new disks:

# Identify new disks
lsblk

# Create filesystem on each disk (example: /dev/sdb for data)
mkfs.xfs /dev/sdb
mkfs.xfs /dev/sdc
mkfs.xfs /dev/sdd

# Create mount points
mkdir -p /u01 /u02 /u03

# Mount
mount /dev/sdd /u01   # Oracle Home (binaries)
mount /dev/sdb /u02   # Oracle data files
mount /dev/sdc /u03   # Oracle redo + archive logs

# Persist in /etc/fstab
cat >> /etc/fstab <<EOF
/dev/sdb /u02 xfs defaults,noatime 0 0
/dev/sdc /u03 xfs defaults,noatime 0 0
/dev/sdd /u01 xfs defaults,noatime 0 0
EOF

# Set ownership
chown -R oracle:oinstall /u01 /u02 /u03

Key insight: The Oracle VM accesses /u02 and /u03 as local XFS block devices. It has no awareness that the underlying storage is an NFS datastore backed by FSx for ONTAP. All NFS communication happens at the ESXi host level, where each host uses its own network path to FSx for ONTAP.


Step 6: Install and configure Oracle 19c

# As oracle user
export ORACLE_HOME=/u01/app/oracle/product/19.0.0/dbhome_1
cd $ORACLE_HOME
./runInstaller -silent -responseFile /path/to/db_install.rsp

Create the database with data on /u02 and logs on /u03:

CREATE DATABASE orcl
  DATAFILE '/u02/oradata/orcl/system01.dbf' SIZE 1G
  LOGFILE
    GROUP 1 '/u03/oralogs/orcl/redo01.log' SIZE 512M,
    GROUP 2 '/u03/oralogs/orcl/redo02.log' SIZE 512M,
    GROUP 3 '/u03/oralogs/orcl/redo03.log' SIZE 512M;

Oracle accesses /u02 and /u03 as local XFS filesystems. Standard Oracle ASM or filesystem-based storage management applies. No NFS-specific Oracle configuration is needed because the NFS layer is abstracted by the ESXi hypervisor.


Step 7: Set up SnapMirror for cross-region DR

Peer clusters (production → DR):

cluster peer create -peer-addrs <dr-cluster-intercluster-ip> \
  -username fsxadmin -initial-allowed-vserver-peers *

Peer SVMs:

vserver peer create -vserver svm-prod -peer-vserver svm-dr \
  -peer-cluster FSxDR -applications snapmirror

Create DP volumes on DR FSx for ONTAP:

vol create -volume oradb1bin -aggregate aggr1 -size 50G -state online -type DP
vol create -volume oradb1data -aggregate aggr1 -size 500G -state online -type DP
vol create -volume oradb1log -aggregate aggr1 -size 250G -state online -type DP

Create and initialize SnapMirror:

snapmirror create -source-path svm-prod:oradb1data \
  -destination-path svm-dr:oradb1data -throttle unlimited \
  -policy MirrorAllSnapshots -type DP

snapmirror create -source-path svm-prod:oradb1log \
  -destination-path svm-dr:oradb1log -throttle unlimited \
  -policy MirrorAllSnapshots -type DP

snapmirror create -source-path svm-prod:oradb1bin \
  -destination-path svm-dr:oradb1bin -throttle unlimited \
  -policy MirrorAllSnapshots -type DP

# Initialize
snapmirror initialize -destination-path svm-dr:oradb1data
snapmirror initialize -destination-path svm-dr:oradb1log
snapmirror initialize -destination-path svm-dr:oradb1bin

Important: Consider Oracle licensing requirements when planning your DR strategy. An alternative is to replicate the Oracle VM through NetApp SnapMirror from Production to DR. Keep the DR replicated volumes as data-protection (DP) volumes that are NOT mounted as NFS datastores on DR hosts until a failover event is declared to avoid Oracle double licensing. Only then break the SnapMirror, mount the NFS datastore on the DR Host, and power on the VM. Pre-mounting the SnapMirror volume as a datastore — even with no VM powered on — means Oracle binaries are accessible on those hosts, which Oracle may consider an “installation” requiring licenses across the entire DR cluster. We recommend requesting an AWS Optimization and Licensing Assessment (AWS OLA) for further guidelines.


Step 8: Configure SnapCenter backup

  1. Deploy SnapCenter Server (or use SnapCenter SaaS).
  2. Add FSx for ONTAP storage system using the cluster management IP.
  3. Install SnapCenter Plugin for Oracle on each Oracle VM.
  4. Create backup policies:
Policy Scope Frequency SnapMirror Update
Full DB Backup Data + Control + Archive Every 4–6 hours Yes
Archive Log Archive logs only Every 10–15 minutes Yes
  1. Create resource groups, assign policies, and schedule.

Database migration from on-premises VMware to EVS

Option 1: VMware HCX live migration

For enterprises with existing VMware on-premises:

  1. Deploy HCX Connector on-premises, HCX Cloud Manager on EVS.
  2. Create site pairing and network extensions (L2 stretch).
  3. Migrate Oracle VMs using HCX vMotion (zero downtime) or Bulk Migration.
  4. Post-migration: storage vMotion Oracle VMDKs from vSAN to FSx for ONTAP NFS datastores for snapshot/replication capabilities.

Option 2: SnapMirror ONTAP-to-ONTAP

If on-premises Oracle already uses NetApp ONTAP storage:

  1. Establish SnapMirror between on-premises ONTAP and AWS FSx for ONTAP.
  2. Incrementally replicate until cutover.
  3. At switchover: quiesce Oracle, flush archive logs, final SnapMirror sync, break mirror.
  4. Mount FSx for ONTAP volumes on EVS Oracle VM, recover database, open for service.

Option 3: Oracle PDB relocation (multitenant)

For Oracle databases already in PDB/CDB multitenant model:

  1. Create target CDB on EVS with FSx for ONTAP storage.
  2. Use PDB hot clone to relocate PDBs from on-premises CDB to AWS CDB.
  3. Minimal service interruption. Only final switchover requires brief outage.

Option 4: RMAN backup/restore (non-ONTAP on-premises)

If Oracle runs on non-ONTAP storage on-premises:

  1. Create RMAN backup, stage to Amazon Simple Storage Service (Amazon S3) using AWS DataSync or AWS Direct Connect.
  2. Provision Oracle VM on EVS, mount FSx for ONTAP volumes.
  3. Restore from RMAN backup, apply archive logs.
  4. Open database and redirect applications.

Day-2 operations

Snapshot backup

SnapCenter manages full database snapshots as storage-layer operations, providing efficient backup capabilities.

Point-in-time recovery

In SnapCenter, select the SCN or timestamp, mount the log snapshot, restore the data snapshot, apply archive logs, and open with RESETLOGS.

Database cloning

SnapCenter FlexClone creates space-efficient database copies. Clones share unchanged blocks with the source and consume storage only for deltas. Use for dev/test, patch validation, and reporting.

HA failover procedure

  1. Break SnapMirror on DR volumes.
  2. Mount SnapMirror volumes as NFS datastores on DR ESXi hosts, then power on the Oracle VM.
  3. Recover to last available archive log.
  4. Open database. Update DNS/connection strings.

Clean up

To stop incurring charges after testing this deployment, remove the following resources in this order:

  1. Oracle VMs — Power off and delete Oracle database VMs from the vSphere inventory.
  2. NFS datastores — Unmount FSx for ONTAP datastores from ESXi hosts in vSphere.
  3. SnapMirror relationships — Delete SnapMirror relationships and DP volumes on the DR FSx for ONTAP file system.
  4. FSx for ONTAP file systems — Delete both production and DR file systems from the Amazon FSx console. This action deletes all volumes and data on those file systems.
  5. Amazon EVS environment — Delete the EVS environment from the Amazon EVS console. This terminates the underlying EC2 bare metal instances.
  6. Networking — Remove Transit Gateway attachments, VPC Route Server configurations, and Direct Connect connections if they were created solely for this deployment.

Important: Deleting an FSx for ONTAP file system permanently removes all data. Confirm that you have backed up any data you need before proceeding.


Conclusion

In this post, we walked through deploying Oracle Database on Amazon Elastic VMware Service (Amazon EVS) with Amazon FSx for NetApp ONTAP as NFS datastore storage. You provisioned the Oracle VM, created FSx for ONTAP volumes, mounted NFS datastores in vSphere, installed Oracle 19c, and configured SnapMirror for cross-region disaster recovery and SnapCenter for Oracle-aware backup. We also covered four migration paths for existing on-premises Oracle workloads and day-2 operations for backup, point-in-time recovery, and cloning.

To get started, review the Amazon EVS User Guide and Configure FSx for ONTAP as NFS Datastore for EVS, then deploy your first Oracle VM on Amazon EVS. For the architecture decisions behind this deployment, see our related post, Architect highly available Oracle Database on Amazon EVS and FSx for ONTAP. Share your feedback and questions in the comments.


Additional resources


About the authors

Architect highly available Oracle Database on Amazon EVS and FSx for ONTAP

Post Syndicated from Satish Bhoi original https://aws.amazon.com/blogs/architecture/architect-highly-available-oracle-database-on-amazon-evs-and-fsx-for-ontap/


Enterprises with existing VMware Cloud Foundation (VCF) investments want to migrate their Oracle databases to AWS without rearchitecting applications or retraining operations teams. Oracle on Amazon Elastic VMware Service (Amazon EVS) can take advantage of sub-millisecond storage latency, snapshot-based backup, cross-region disaster recovery (DR), and independent storage scaling, all while preserving existing VMware operational workflows.

In this post, we show you how to architect a complete Oracle Database environment on Amazon EVS with Amazon FSx for NetApp ONTAP. You learn how the storage, compute, networking, and disaster recovery layers work together to deliver high availability, cross-region DR, and sub-millisecond storage latency while maintaining your familiar VMware operational tooling.

In this post, you learn how to:

  • Design Oracle Database architecture on Amazon EVS running VMware Cloud Foundation 9.1.
  • Select optimal EC2 bare metal instance types and VM sizing for Oracle workloads.
  • Architect storage using Amazon FSx for NetApp ONTAP as NFS datastores for Oracle data and log volumes.
  • Plan SnapMirror replication for cross-region disaster recovery.
  • Evaluate migration options for moving existing Oracle workloads from on-premises VMware to EVS.

Amazon EVS directly runs VMware Cloud Foundation (VCF) environments on Amazon Elastic Compute Cloud (Amazon EC2) bare metal instances within an Amazon Virtual Private Cloud (Amazon VPC). With VCF 9.x, Amazon EVS provisions the bare metal infrastructure and VLAN subnets, and you then deploy VCF using Broadcom’s VCF Installer (the self-deployed model). VCF 9.x also supports evaluation mode, so you can validate the design before applying license keys, and the Solutions for Amazon EVS GitHub repository provides CloudFormation and Terraform templates to automate the phased VCF 9 deployment. Amazon FSx for NetApp ONTAP provides managed ONTAP storage with NFS, SMB, iSCSI, and NVMe over TCP access. FSx for NetApp ONTAP can deliver sub-millisecond response times, multiple GBps of throughput, and up to 80,000 IOPS per file system (see FSx for ONTAP performance).

For more information, see FSx for NetApp ONTAP features.

For step-by-step deployment procedures including provisioning, storage configuration, Oracle installation, and SnapMirror setup, see our companion post: Deploy Oracle Database step by step on Amazon EVS with FSx for ONTAP.


Solution architecture

This section describes a highly available Oracle Database deployment on Amazon EVS with Amazon FSx for NetApp ONTAP storage across two AWS Regions. Figure 1 illustrates the numbered data flow from on-premises through hybrid connectivity into the production EVS environment and cross-region DR.

Oracle on Amazon EVS reference architecture showing on-premises to AWS connectivity, the production EVS cluster on FSx for ONTAP, and cross-region SnapMirror DR

Figure 1: Oracle Database on Amazon EVS with FSx for NetApp ONTAP and cross-region SnapMirror DR

The architecture consists of eight functional components, described in the following section.

Architecture components

These numbered items correspond to the data flow shown in the diagram.

  1. Hybrid link — On-premises data center connects to AWS through AWS Direct Connect for dedicated, low-latency bandwidth.
  2. Transit routing — AWS Transit Gateway routes traffic between the production VPC, DR region, and on-premises networks.
  3. Workload landing — Traffic reaches the EVS Database Cluster running on i7i.metal-24xl Amazon EC2 bare metal instances with VCF 9.1.
  4. Live migration — VMware HCX (Hybrid Cloud Extension) migrates Oracle VMs from on-premises VMware to EVS with near-zero downtime using Replication Assisted vMotion or bulk migration.
  5. NFS data path — Each ESXi host mounts Amazon FSx for NetApp ONTAP as an NFS datastore. Oracle VMs access database volumes as block devices (VMDKs on the NFS datastore) with sub-millisecond latency and up to 80,000 IOPS. Each host has its own independent NFS path to FSx for ONTAP, distributing bandwidth across the cluster.
  6. Cross-region DR — SnapMirror asynchronously replicates FSx for ONTAP volumes (data, logs, binaries) to the DR region with configurable Recovery Point Objective (RPO).
  7. Dynamic routing — NSX Tier-0 gateway peers through BGP with Amazon VPC Route Server. This is one-way BGP. Route Server listens for routes advertised by NSX and writes them to the VPC route table, but does not advertise VPC routes back to NSX.
  8. DR failover — On failover, Transit Gateway routes to the DR region where pre-provisioned standby Oracle VMs mount the SnapMirror replica volumes.

Factors to consider for Oracle Database deployment on EVS

Before you begin deployment, evaluate your Oracle workload requirements against the available infrastructure options. The decisions you make for instance types, VM sizing, and storage architecture directly affect database performance, cost, and operational complexity.

Oracle licensing consideration: Instance type selection affects Oracle license cost, which is an important factor to take into consideration. We recommend requesting an AWS Optimization and Licensing Assessment (AWS OLA) for further guidelines.

EC2 instance type selection for ESXi hosts

Amazon EVS currently supports three bare metal instance types for ESXi hosts. The choice depends on whether you prioritize per-core compute speed (i7i) or per-host memory and storage density (i4i), and on how much you want to scale a single host vertically. For most new Oracle deployments, i7i is the better choice because Oracle query performance is sensitive to CPU instruction throughput and storage IO latency.

  • i7i.metal-24xl (recommended default): 96 vCPUs (48 cores), 768 GiB RAM. 5th Gen Intel Xeon delivers improved compute performance, critical for Oracle CPU-bound queries. 3rd Gen AWS Nitro SSDs provide enhanced real-time storage performance and lower IO latency for vSAN.
  • i7i.metal-48xl: Same 5th Gen Intel Xeon as the 24xl but double the capacity per host (192 vCPUs, 96 cores, 1,536 GiB RAM, up to 100 Gbps network and 60 Gbps Amazon Elastic Block Store (Amazon EBS) bandwidth). Choose this when you want to scale a single Oracle host vertically, such as for large SGAs, high core counts, or fewer and denser hosts to reduce VMware per-host licensing. It keeps the same per-core performance profile as the 24xl.
  • i4i.metal: Higher per-host density (128 vCPUs, 1,024 GiB RAM, 30 TB NVMe) suits environments requiring fewer, larger hosts to reduce VMware licensing costs.

VCF compatibility: All three instance types support VCF 9.1 with ESXi 9.1 (build 9.1.0.0100.25433460). VCF 5.2.2 (ESXi 8.0U3g) remains available but is on a path to end of support, so new deployments should default to 9.x. An EVS environment supports 4–32 hosts per cluster. You can mix instance types across clusters within the same SDDC.

VM sizing for Oracle Database guests

Size Oracle Database VMs based on these workload characteristics:

  • Allocate vCPU count matching Oracle CPU_COUNT parameter.
  • Size memory for SGA + PGA + OS overhead (typically 75–85% of allocated VM memory for SGA)
  • Configure VM swap on vSAN datastore (not NFS). vSAN uses local NVMe with single-digit millisecond latency.
  • Use NUMA-aware VM placement for VMs exceeding single-socket core count.
  • Place OS swap and Oracle temp tablespace on the vSAN datastore for single digit ms latency at no additional cost.

Storage architecture: vSAN + FSx for NetApp ONTAP

The recommended design splits storage responsibilities between two tiers:

  • vSAN (backed by local NVMe drives on each ESXi host) handles low-latency, non-replicated workloads: VM boot disks, OS swap, and Oracle temp tablespace.
  • FSx for NetApp ONTAP handles Oracle data files and redo logs that require snapshot-based backup and cross-region replication. ESXi hosts mount FSx for ONTAP volumes as NFS datastores, and Oracle VMs access standard VMDKs on those datastores.

This separation gives local NVMe speed for transient IO while adding snapshot, clone, and SnapMirror capabilities for persistent database files.

Why NFS datastore (host-level) instead of in-guest NFS (dNFS)? With in-guest NFS, all Oracle NFS traffic routes through the NSX overlay and an NSX Edge node before reaching FSx for ONTAP. This creates a single Edge chokepoint. With NFS datastores, each ESXi host talks NFS directly to FSx for ONTAP using its own network bandwidth. There is no Edge bottleneck and no extra latency hop.

FSx for NetApp ONTAP sizing considerations

Important: Use 100% SSD for Oracle. We recommend against using capacity pool tiering for Oracle database volumes. Keep tiering policy set to none for all Oracle volumes.

Important: SSD capacity planning. If the SSD tier fills to capacity, FSx for ONTAP blocks writes. Monitor SSD utilization and provision headroom (minimum 20% free).

Important: Read/write throughput asymmetry. On a 6 GB/s filesystem, read throughput can reach 6 GB/s, but write throughput is limited to approximately 1 GB/s. Size throughput capacity based on Oracle write workload requirements. This write ceiling applies per high-availability (HA) pair. To scale write throughput beyond a single HA pair, deploy a file system with more than one HA pair and distribute Oracle volumes across the additional aggregates so writes are spread across pairs. This placement is not automatic, so plan the layout up front to keep write-heavy datasets balanced.


Network architecture

Amazon EVS uses VLAN subnets (defined at environment creation and unchangeable later) to segment traffic.

VPC Route Server replaces static routes within the VPC. NSX Tier-0 gateways peer through BGP with Route Server endpoints. This is one-way BGP: Route Server listens for routes from NSX and programs them into the VPC route table but does not advertise VPC routes back to NSX. Beyond the VPC, routes remain static at Transit Gateway.

NSX-T segmentation for Oracle

  • Dedicated Tier-1 gateway for production database segments (DB subnets)
  • Separate Tier-1 for application tier (App subnets) and perimeter network.
  • Distributed firewall rules restrict Oracle listener access (TCP 1521) to authorized application segments only.
  • Micro-segmentation between Oracle instances prevents lateral movement.

Important: Security group rules are not enforced on VLAN subnet interfaces. Use network ACLs and NSX distributed firewall for traffic control.


High availability and disaster recovery

  • SnapMirror replication frequency determines RPO. Configure based on business requirements.
  • Pre-provision standby Oracle VMs in the DR cluster to reduce Recovery Time Objective (RTO).
  • Replicate binary volumes so that Oracle installation is not required during recovery.
  • Automate failover with Ansible/SnapCenter to reduce human error.

Oracle licensing consideration for DR: There are license impacts based on how DR replication is implemented. If you have licensing questions, we recommend requesting an AWS Optimization and Licensing Assessment (AWS OLA).

To comply with the Oracle licensing rules, an alternative is to replicate the Oracle VM through NetApp SnapMirror from Production to DR. Keep the DR replicated volumes as data-protection (DP) volumes that are NOT mounted as NFS datastores on DR hosts until a failover event is declared to avoid Oracle double licensing. Only then break the SnapMirror, mount the NFS datastore on the DR Host, and power on the VM. Pre-mounting the SnapMirror volume as a datastore, even with no VM powered on, means Oracle binaries are accessible on those hosts, which Oracle may consider an “installation” requiring licenses across the entire DR cluster.


Database migration from on-premises VMware to EVS

The following table compares migration options from on-premises VMware to EVS.

Option Method Best for Downtime
1 VMware HCX Live Migration Existing VMware on-prem Near-zero (vMotion) or planned bulk
2 SnapMirror ONTAP-to-ONTAP On-prem Oracle on NetApp ONTAP Minutes (final sync switchover)
3 Oracle PDB Relocation PDB/CDB multitenant model Brief (final switchover only)
4 RMAN Backup/Restore Non-ONTAP on-prem (universal) Hours (backup + restore + apply)

For detailed procedures on each migration option, see our companion post: Deploy Oracle Database step by step on Amazon EVS with FSx for ONTAP.


Security

Security for Oracle on Amazon EVS spans multiple layers from network isolation to database-level encryption. The following table summarizes the security controls across each layer.

Layer Control
Network segmentation NSX-T Tier-1 gateways isolate DB/App/perimeter network segments
East-west traffic NSX Distributed Firewall: restrict TCP 1521 to authorized app segments
North-south traffic FortiGate or equivalent inspection VPC for ingress/egress filtering
Encryption at rest FSx for ONTAP volumes encrypted with AWS Key Management Service (AWS KMS)
Encryption in transit VPC encryption for NFS traffic, and IPsec for SnapMirror cross-region
Database encryption Oracle TDE (Transparent Data Encryption) for additional protection
Administrative access Zero-trust access (for example, Banyan or Zscaler) for VMware admin consoles
VLAN subnet security Network ACLs (security groups not enforced on VLAN interfaces)

Cost optimization

Cost optimization for Oracle on Amazon EVS focuses on matching infrastructure capacity to workload demands and using AWS pricing models. The following table summarizes key strategies across compute, storage, networking, and Oracle licensing. We recommend requesting an AWS Optimization and Licensing Assessment (AWS OLA) for further guidelines.

Component Strategy
EC2 bare metal hosts Compute Savings Plans or Reserved Instances (up to 54% savings)
Instance type selection i7i.metal-24xl delivers ~10% price-performance over i4i.metal
FSx for ONTAP throughput Right-size for the write workload, and adjust on the fly
FSx for ONTAP storage 100% SSD for Oracle, and storage efficiency for non-production
Data transfer Place FSx for ONTAP in same AZ as EVS cluster
SnapMirror replication Schedule frequency based on RPO (less frequent = lower cost)
Oracle DR licensing Replicate the Oracle VMs through NetApp SnapMirror from Production to DR, keeping the DR replicated volumes as data-protection (DP) volumes that are NOT mounted as NFS datastores on DR hosts until a failover event is declared to avoid Oracle double licensing. When a failover event is declared, then break the SnapMirror, mount the NFS datastore on the DR Host, and power on the VM. We recommend requesting an AWS Optimization and Licensing Assessment (AWS OLA) for further guidelines.
Non-production Use fewer hosts, and apply tiering for dev/test data

Summary

Deploying Oracle databases on Amazon EVS with Amazon FSx for NetApp ONTAP provides high availability, cross-region DR, and sub-millisecond storage latency while combining VMware operational consistency with AWS cloud economics:

  • Performance: i7i.metal-24xl delivers up to 23% better compute and 50% lower IO latency. FSx for ONTAP delivers sub-millisecond latency with up to 80,000 IOPS. For details, see Amazon EC2 i7i instances in the AWS GovCloud (US) Regions.
  • Availability: vSphere HA, SnapMirror cross-region replication, and optional Oracle Data Guard.
  • Manageability: SnapCenter for backup, clone, and recovery in seconds regardless of database size.
  • Migration flexibility: HCX (live), SnapMirror (ONTAP-to-ONTAP), PDB Relocation (multi-tenant), and RMAN (universal).
  • Security: NSX micro-segmentation, AWS KMS encryption, and zero-trust access.
  • Cost efficiency: Improved price performance with i7i, on-the-fly throughput adjustment, and AWS Savings Plans.

This architecture provides you with high availability, cross-region DR, and storage-based backup and cloning similar to Oracle RAC and Data Guard functions while maintaining your familiar VMware operational tooling and procedures.


Next steps

To get started with this deployment:

  1. Provision an Amazon EVS environment in your target Region. See the Amazon EVS User Guide for setup instructions.
  2. Deploy an Amazon FSx for NetApp ONTAP file system in the same VPC and Availability Zone as your EVS cluster.
  3. Follow the step-by-step procedures in our companion post to mount NFS datastores, create Oracle VMDKs, and configure SnapMirror DR.
  4. Test in a non-production environment first, then migrate production Oracle workloads.

Additional resources


About the authors

Closed-loop incident response: connect AWS DevOps Agent to OpenSearch

Post Syndicated from Sitaraman Vijay Krishna original https://aws.amazon.com/blogs/devops/closed-loop-incident-response-connect-aws-devops-agent-to-opensearch/

The gap in most observability setups isn’t the data, it’s closing the loop between incident detection and response. Your Amazon OpenSearch Service domain already stores structured logs and distributed traces, detects anomalies through alerting monitors, and fires notifications reliably. But when the system sends that alert lands at 2 AM, your engineers must still investigate manually. The investigating engineer opens Dashboards, crafts domain-specific language (DSL) queries, hunts for correlated trace IDs, pivots to AWS CloudTrail, and pieces together a root cause. That manual loop can take anywhere from minutes for straightforward issues to hours when failures cascade across microservices.

What if you closed the loop by connecting an AI agent directly to the indices that triggered the alert?

This post shows you how to connect AWS DevOps Agent to your OpenSearch observability data using the Model Context Protocol (MCP). The same alert that would page a human instead triggers the agent to query the logs and traces automatically, correlate them with AWS CloudTrail and Amazon CloudWatch, and deliver a root cause analysis.

Note: AWS DevOps Agent access to logs and traces in OpenSearch is controlled by the fine-grained access control (FGAC) role it is given.

We present three hosting paths for the MCP server so you can pick by AWS Region and operational preference: self-managed on Amazon Elastic Container Service (Amazon ECS) (AWS Fargate behind an internal Network Load Balancer (NLB), reachable through Amazon VPC Lattice, which works everywhere today), Amazon Bedrock AgentCore (a one-click AWS CloudFormation template, where available), or the built-in MCP endpoint on OpenSearch 3.3+ (no separate server). All paths use the official opensearch-mcp-server-py package.

By the end, you can deploy the MCP server (any of the three paths), register it as a Capability Provider, configure the AWS Identity and Access Management (IAM)-to-FGAC role mapping, route your OpenSearch alerts to the agent’s Event Channel, and verify the closed loop with a controlled failure.

Solution overview

The architecture creates a closed loop: your OpenSearch domain handles both detection and investigation, with AWS DevOps Agent orchestrating the response.

Closed-loop flow from OpenSearch alerting through Amazon SNS and a webhook forwarder to AWS DevOps Agent investigation

Figure 1: Architecture for Amazon OpenSearch alert to AWS DevOps Agent MCP investigation

Closed-loop architecture: OpenSearch Alerting triggers Amazon Simple Notification Service (Amazon SNS), which flows through a webhook forwarder to the Event Channel of AWS DevOps Agent. The agent investigates by querying OpenSearch indices using MCP tools, correlates with CloudTrail and CloudWatch, and delivers a root cause analysis.

The loop runs in six stages:

  1. Applications emit logs and traces to OpenSearch.
  2. Alerting monitors detect anomalies and publish to Amazon SNS.
  3. Amazon SNS triggers the webhook forwarder AWS Lambda function.
  4. The webhook forwarder Lambda transforms each notification into a hash-based message authentication code (HMAC)-signed payload for the agent’s Event Channel.
  5. AWS DevOps Agent queries the OpenSearch indices through MCP and correlates them with CloudTrail and CloudWatch.
  6. AWS DevOps Agent delivers a root cause analysis.

The critical insight: AWS DevOps Agent consumes remote MCP servers registered as Capability Providers. AWS DevOps Agent doesn’t connect to local MCP servers running on developer workstations (those are used by OpenSearch MCP apps for IDEs). The agent needs a network-accessible endpoint: either self-managed on Amazon ECS, hosted on Bedrock AgentCore, or built into the OpenSearch domain itself (3.3+).

Choosing a hosting path

This post walks through the self-managed path in detail (works everywhere today) and calls out the AgentCore and built-in 3.3+ alternatives at each step.

Prerequisites

Confirm the following before starting:

  • An Amazon OpenSearch Service managed domain (2.x+) with FGAC enabled, application logs and traces already indexed, and Alerting monitors publishing to an SNS topic.
  • AWS DevOps Agent enabled in your account.
  • AWS Cloud Development Kit (AWS CDK) (npm install -g aws-cdk, Node.js 18+) and AWS Command Line Interface (AWS CLI) v2 configured with admin access to your OpenSearch domain.

Note: This walkthrough assumes you already have an application emitting logs and traces to OpenSearch. The infrastructure in this post is shown as inline CDK snippets you can drop into your own CDK app and adapt to your environment.

Step 1: Deploy the OpenSearch MCP server

Choose the path that fits your Region. The suggested options host the official opensearch-mcp-server-py and expose the same MCP tools (SearchIndexTool, ListIndexTool, and the broader observability tool set) to AWS DevOps Agent.

Option A: Self-managed on Amazon ECS and NLB (deploy anywhere)

This path runs the MCP server on ECS Fargate behind an internal Network Load Balancer and exposes it to AWS DevOps Agent through a VPC Lattice private connection. It works in every Region today.

Self-managed MCP server on Amazon ECS Fargate behind an internal NLB, reached by AWS DevOps Agent over VPC Lattice

Figure 2: Architecture for self-managed OpenSearch MCP deployed on AWS Fargate

1. Generate a Transport Layer Security (TLS) certificate for the MCP server

AWS DevOps Agent requires HTTPS endpoints. Generate a self-signed certificate whose subject alternative name (SAN) matches your NLB DNS name, and import it to AWS Certificate Manager (ACM):

# Generate a self-signed cert whose SAN matches the NLB DNS, then import to ACM
openssl req -x509 -nodes -days 365 -newkey rsa:2048 \
  -keyout /tmp/mcp-key.pem -out /tmp/mcp-cert.pem \
  -subj "/CN=mcp-server" -addext "subjectAltName=DNS:<your-nlb-dns>"
aws acm import-certificate --certificate fileb:///tmp/mcp-cert.pem \
  --private-key fileb:///tmp/mcp-key.pem --region <region> \
  --query "CertificateArn" --output text

Save the returned certificate Amazon Resource Name (ARN) for the next step.

Why the SAN matters: VPC Lattice validates the certificate against the host address you configure for the private connection. If the SAN doesn’t match the NLB DNS, TLS validation fails and the connection never reaches Completed.

2. Define the MCP server in your CDK app

Run opensearch-mcp-server-py as an Amazon ECS Fargate service behind an internal NLB. The following snippet shows the essential wiring. Adapt it to your existing CDK app:

// Representative wiring — adapt into your CDK app (full construct in the linked Guidance).
// Fargate task runs the MCP server; grant it read on the domain (FGAC handles index auth).
taskDef.addContainer('mcp', {
  image: ecs.ContainerImage.fromRegistry('python:3.12-slim'),
  command: ['sh','-c', 'pip install opensearch-mcp-server-py --quiet && '
    + 'opensearch-mcp-server-py --transport stream --host 0.0.0.0 --port 8080'],
  environment: { OPENSEARCH_URL: https://${props.openSearchDomain.domainEndpoint},
    OPENSEARCH_USE_SSL: 'true' }, portMappings: [{ containerPort: 8080 }] });
// Internal NLB — MUST have a security group so VPC Lattice can reach it; TLS listener
// terminates with your ACM cert and forwards TCP:8080 to the service.
// Output the NLB DNS and task role ARN for Steps 2 and 3.

Deploy it (cdk deploy --require-approval broadening --region <region>) and note the NLB DNS and task role ARN from the stack outputs. You’ll need the stack outputs for the private connection (Step 2) and the FGAC mapping (Step 3).

The server starts in streamable-HTTP mode (–transport stream). Installing the package at container startup adds approximately 30 seconds to the first task boot up time. The higher task sizing (1024 MB/512 CPU) helps pip install complete quickly, and the health check’s unhealthyThresholdCount: 5 gives the service approximately 2.5 minutes to stabilize. For production, bake the package into a prebuilt image to avoid startup latency entirely.

3. Critical: Use a Network Load Balancer (NLB), not an Application Load Balancer (ALB)

The OpenSearch MCP servers use the streamable-HTTP transport, which delivers responses as Server-Sent Events (SSE) with chunked transfer encoding. Application Load Balancers (ALBs) operate at Layer 7 and can strip the Transfer-Encoding: chunked header, breaking the SSE stream. Network Load Balancers operate at Layer 4 (TCP) and pass HTTP framing untouched after TLS termination. Therefore deploy a streamable-HTTP MCP server behind an NLB with TLS termination and not an ALB.

Important: The NLB must have a security group so VPC Lattice resource gateway ENIs can reach it, and a security group can only be attached to an NLB at creation time (it cannot be added later). The CDK stack attaches one. If you build your own NLB, specify the security group when you create it.

Option B: Amazon Bedrock AgentCore (managed, where available)

If your Region supports the integration, the OpenSearch console provides a one-click CloudFormation template that deploys opensearch-mcp-server-py on AgentCore.

OpenSearch MCP server hosted on Amazon Bedrock AgentCore and registered with AWS DevOps Agent

Figure 3: Architecture for OpenSearch MCP deployed on Amazon Bedrock AgentCore

  1. Open the Amazon OpenSearch Service console, select your domain, and go to Integrations.
  2. Locate the MCP server template and choose Launch stack.
  3. Provide the parameters:
Parameter Description Example
OpenSearchEndpoint Your domain’s endpoint https://my-domain.<region>.es.amazonaws.com
AWSRegion Region where the domain runs <region>
AgentName Logical name for this MCP server opensearch-observability-mcp

After the stack reaches CREATE_COMPLETE, note from the Outputs tab: AgentCoreEndpoint, CognitoClientId, CognitoClientSecret, and McpServerRoleArn (needed for FGAC in Step 3).

Verify the tools are registered:

# Fetch an OAuth token, then confirm the tools list. Expect SearchIndexTool, ListIndexTool.
curl -s -X POST "https://<AgentCoreEndpoint>" -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}' \
  | jq '.result.tools[].name'

Option C: OpenSearch 3.3+ built-in MCP

OpenSearch 3.3+ domain exposing a built-in MCP endpoint registered directly with AWS DevOps Agent

Figure 4: MCP deployment architecture for domains running OpenSearch version 3.3+

This is the simplified path for domains running OpenSearch 3.3+. These versions have a built-in MCP endpoint, so no separate server deployment is needed. Enable the endpoint with two API calls:

curl -XPUT "https://<domain-endpoint>/_cluster/settings" \
  -H "Content-Type: application/json" \
  --aws-sigv4 "aws:amz:<region>:es" \
  -d '{"persistent": {"plugins.ml_commons.mcp_server_enabled": "true"}}'

curl -XPOST "https://<domain-endpoint>/_plugins/_ml/mcp/tools/_register" \
  -H "Content-Type: application/json" \
  --aws-sigv4 "aws:amz:<region>:es" \
  -d '{"tools": ["SearchIndexTool", "ListIndexTool"]}'

Your MCP endpoint is live at https://<domain-endpoint>/_plugins/_ml/mcp. In Step 2, register this URL directly with AWS DevOps Agent using SigV4 authentication (service=es).

Step 2: Register with AWS DevOps Agent

With the MCP server running, register it as a Capability Provider so AWS DevOps Agent can invoke its tools during investigations.

The private connection in section 2.1 applies only to the self-managed path (Option A). For the AgentCore and built-in paths, skip to section 2.2.

2.1 Create a private connection

A private connection creates a secure network path between AWS DevOps Agent and your NLB using Amazon VPC Lattice.

  1. Open the AWS DevOps Agent console.
  2. Navigate to Capability Providers then Private connections and choose Create a new connection.
    • Configure the connection with these settings:
      • For Name, enter opensearch-mcp-connection.
      • Select your virtual private cloud (VPC) and private subnets (one per Availability Zone (AZ)).
      • Attach a security group that allows inbound TCP 443.
      • For Host address, enter your NLB DNS name (the Step 1 output), and set TCP port to 443.
      • For Certificate public key, paste the contents of /tmp/mcp-cert.pem.
  3. Choose Create Connection and wait for status Completed (~5–10 minutes).

2.2 Register the MCP server

Then register the Capability Provider. The fields differ by path:

Field Self-managed (ECS) AgentCore Built-in (3.3+)
Name opensearch-observability opensearch-observability opensearch-observability
Endpoint URL https://<nlb-dns>/mcp https://<AgentCoreEndpoint> https://<domain-endpoint>/_plugins/_ml/mcp
Private connection opensearch-mcp-connection — —
Authentication API key (x-api-key) OAuth SigV4
OAuth client ID / secret — from stack outputs —
SigV4 service — — es

2.3 Configure allowed tools

Classify the MCP tools as read-only so the agent can query but not mutate your domain: SearchIndexTool, ListIndexTool (and, if present, GetMappingsTool and GetShardsTool).

2.4 Verify the registration

In the AWS DevOps Agent console test interface, ask:

“List the indices in my OpenSearch domain that match application-logs-*”

The agent should invoke the list tool and return your index names. A 403 Forbidden or security_exception means you haven’t configured the FGAC mapping yet. Continue to Step 3. A connection timeout on the self-managed path points to the private connection status or NLB security group.

Step 3: Configure FGAC role mapping

OpenSearch managed domains with FGAC enforce a strict separation: IAM authenticates the caller, but the internal security plugin of OpenSearch authorizes access to the indices. Without an explicit mapping between the IAM role and an OpenSearch backend role, authenticated requests still receive 403 Forbidden.

Identify the role to map

Path IAM Role Where to find it
Self-managed (ECS) ECS task role CDK output McpTaskRoleArn (from the Step 1 snippet)
AgentCore MCP server role (trust: agentcore.bedrock.amazonaws.com) CloudFormation Outputs → McpServerRoleArn
Built-in endpoint DevOps Agent role (trust: aidevops.amazonaws.com) DevOps Agent console → Agent Space → IAM Configuration

With the role identified, apply the mapping. We recommend that you scope it to your observability indices rather than granting cluster-wide read, so you can control which data the agent can query. The first call creates a read-only role restricted to those indices. The second maps your IAM role to it:

# Create a read-only role scoped to your observability indices
curl -XPUT "https://<domain-endpoint>/_plugins/_security/api/roles/devops_agent_readonly" \
  --aws-sigv4 "aws:amz:<region>:es" -H "Content-Type: application/json" -d '{
    "cluster_permissions":["cluster_composite_ops_ro"],
    "index_permissions":[{"index_patterns":["application-logs-*","otel-traces-*"],
      "allowed_actions":["read","search"]}] }'
# Map the IAM role to it
curl -XPUT "https://<domain-endpoint>/_plugins/_security/api/rolesmapping/devops_agent_readonly" \
  --aws-sigv4 "aws:amz:<region>:es" -H "Content-Type: application/json" -d '{
    "backend_roles":["arn:aws:iam::<ACCOUNT_ID>:role/<YourMcpTaskRole-or-DevOpsAgentRole>"] }'

Verify:

curl -s "https://<domain-endpoint>/_plugins/_security/api/rolesmapping/devops_agent_readonly" \
  --aws-sigv4 "aws:amz:<region>:es" | jq '.devops_agent_readonly.backend_roles'

Step 4: Wire up alert routing

Your OpenSearch Alerting monitors already fire to SNS. The remaining connection is a lightweight Lambda that transforms those SNS notifications into the format the AWS DevOps Agent Event Channel expects, with HMAC signing for payload integrity. This is the piece that closes the loop: investigations trigger automatically, without a human forwarding the alert.

4.1 The webhook forwarder

The forwarder performs two operations: payload transformation and HMAC-SHA256 signing.

# HMAC-SHA256 signing contract. DevOps Agent expects two headers:
#   x-amzn-event-signature = base64(HMAC-SHA256(secret, f"{timestamp}:{body}"))
#   x-amzn-event-timestamp = %Y-%m-%dT%H:%M:%S.000Z (UTC)
# transform_alert() maps the OpenSearch Alerting payload to a DevOps Agent event
# (severity 1/2->HIGH, 3->MEDIUM, 4->LOW), deriving title/service/incidentId/metadata
# from monitor_name, trigger_name, and results.

The complete implementation adds retry logic (three attempts with exponential backoff at 1, 2, and 4 seconds), AWS Secrets Manager integration for HMAC secret caching, and structured JSON error logging.

4.2 Deploy the forwarder

Define the forwarder as a Lambda subscribed to your alerting SNS topic. Package the preceding transformation and signing logic as the handler, wire it up in AWS CDK, and deploy with cdk deploy --require-approval broadening --region <region>:

// Representative wiring — the forwarder Lambda subscribes to the alerting SNS topic.
const forwarder = new lambda.Function(this, 'WebhookForwarder', {
  runtime: lambda.Runtime.PYTHON_3_12, handler: 'forwarder.handler',
  code: lambda.Code.fromAsset('lambda/webhook-forwarder'), timeout: Duration.seconds(60),
  environment: { WEBHOOK_URL: props.eventChannelUrl,        // Event Channel URL (below)
                 WEBHOOK_SECRET_NAME: 'devops-agent-webhook-secret' } });  // HMAC secret
secret.grantRead(forwarder);
alertsTopic.addSubscription(new subs.LambdaSubscription(forwarder));

4.3 Configure the Event Channel

With the forwarder deployed, create the webhook in the AWS DevOps Agent console and wire its URL and signing secret back into the Lambda.

  1. In the AWS DevOps Agent console, go to Agent Space, Webhooks, Agent Space Webhook, and then choose Add webhook.
  2. Complete the setup steps: verify the data schema, configure HMAC authentication, and generate the URL and credentials.
  3. Note the HMAC signing secret and store it in the AWS Secrets Manager secret referenced by the forwarder (devops-agent-webhook-secret). See Create an AWS Secrets Manager secret in the AWS Secrets Manager User Guide.
  4. Set the generated Webhook URL as the WEBHOOK_URL environment variable on the forwarder Lambda (the eventChannelUrl prop in the preceding snippet).

Note: Creating an OpenSearch Alerting monitor (if you don’t have one yet).

If you don’t have a monitor yet, create an SNS notification channel (Notifications plugin) and a query-level alerting monitor over application-logs-* that triggers when error_count > 5 and posts to that channel. Full request bodies are in the OpenSearch Alerting docs. The action’s message_template must emit monitor_name, trigger_name, severity, period_start, period_end, and results to match the forwarder’s schema.

The IAM role (role_arn) needs sns:Publish permission on your topic and a trust policy allowing es.amazonaws.com to assume it. The destination_id in the action must match the config_id from the notification channel.

4.4 Verify delivery

Publish a test alert to your SNS topic:

aws sns publish --topic-arn <AlertSnsTopicArn> \
  --message '{"monitor_name":"test-connectivity","trigger_name":"manual-test",
    "severity":"3","period_start":"2026-07-15T10:00:00Z",
    "period_end":"2026-07-15T10:05:00Z",
    "results":[{"index":"application-logs-2026.07.15","doc_count":1}]}'

Check the forwarder’s CloudWatch Logs for:

{"level": "INFO", "message": "Delivered successfully", "status_code": 200, "incident_id": "test-connectivity-1783166700"}

Note: 401 Unauthorized means the HMAC secret in Secrets Manager doesn’t match the Event Channel secret. Connection refused means the Event Channel URL is wrong or the Lambda lacks outbound access.

Step 5: Verify the closed loop

With all four components connected (MCP server, Capability Provider, FGAC mapping, and alert routing), trigger a controlled failure to verify the full loop.

5.1 Inject failure

Set reserved concurrency to zero on one of your application’s Lambda functions. This causes all invocations to be throttled:

aws lambda put-function-concurrency \
  --function-name <YourApplicationLambdaFunction> \
  --reserved-concurrent-executions 0

5.2 Generate traffic

Send requests to trigger errors that will appear in your OpenSearch logs:

for i in $(seq 1 20); do
  curl -s -o /dev/null -w "HTTP %{http_code}\n" \
    https://<your-api-endpoint>/orders
  sleep 2
done

5.3 Expected timeline

After you inject the failure and generate traffic, events should unfold roughly as follows. Use this timeline to confirm each stage of the loop is firing:

Elapsed Event
T+0s put-function-concurrency executed
T+30s Throttle errors appear in application-logs-* index
T+~120s Alerting monitor evaluates and triggers
T+~130s SNS → Forwarder Lambda → Event Channel delivery
T+~135s DevOps Agent begins investigation
T+~300s Root cause analysis delivered

What AWS DevOps Agent produces

ROOT CAUSE ANALYSIS - error-rate-monitor / high-error-rate
1. OpenSearch logs (SearchIndexTool): 47 ERROR entries, "TooManyRequestsException" on service=OrderProcessor
2. Traces (SearchIndexTool): matching spans show status.code=429, function never executed
3. CloudTrail: PutFunctionConcurrency set ReservedConcurrentExecutions=0 two minutes before first error
4. CloudWatch: Throttles=47, Invocations=0
ROOT CAUSE: reserved concurrency set to 0 on OrderProcessorFunction, blocking all invocations.
REMEDIATION: aws lambda delete-function-concurrency --function-name OrderProcessorFunction

The agent identified the root cause by querying the same data that triggered the alert, which completes the closed loop.

5.4 Revert the failure

After the investigation completes, restore normal capacity by removing the reserved concurrency limit you set earlier:

aws lambda delete-function-concurrency \
  --function-name <YourApplicationLambdaFunction>

Note: If the agent doesn’t begin investigation within 3 minutes, check: (1) the forwarder Lambda executed (CloudWatch Logs), (2) the Event Channel shows the received event, and (3) the Capability Provider is registered and healthy.

Cost considerations

This walkthrough adds roughly $70/month (as of September 2026, and varies by AWS Region and usage): ECS Fargate MCP task approximately $15, network address translation (NAT) gateway approximately $35, NLB approximately $18, Secrets Manager approximately $0.40, and the Lambda forwarder under $1. The AgentCore path removes the NLB and ECS costs but adds AgentCore hosted-endpoint charges. Your existing OpenSearch domain and application aren’t included.

Security considerations

The design keeps everything inside the VPC: OpenSearch and the MCP server run in private subnets with nothing public-facing, and AWS DevOps Agent reaches the MCP server over a VPC Lattice private connection. NLB terminates TLS (OpenSearch enforces HTTPS, TLS 1.2 minimum), and encryption at rest is enabled. IAM is least-privilege: the MCP server’s role maps to a read-only OpenSearch backend role scoped to your observability indices.

For production, also consider Security Assertion Markup Language (SAML)/IAM FGAC, request validation in front of the MCP server, and VPC endpoints for Amazon Elastic Container Registry (Amazon ECR), CloudWatch, and Secrets Manager.

Cleanup

# Destroy the CDK stacks (MCP server on ECS, webhook forwarder)
cdk destroy --all --region <region>
# AgentCore path: delete its CloudFormation stack
aws cloudformation delete-stack --stack-name opensearch-mcp-agentcore
# In the DevOps Agent console: remove the Capability Provider and the private connection.
# Remove the FGAC role mapping
curl -XDELETE "https://<domain-endpoint>/_plugins/_security/api/rolesmapping/devops_agent_readonly" \
  --aws-sigv4 "aws:amz:<region>:es"
# 3.3+ path: disable the built-in endpoint. Delete the imported ACM certificate.

Verify the resources are removed: ECS tasks, NLB, NAT Gateway, AgentCore hosted endpoint (if used), and the Secrets Manager secret.

Conclusion

You connected AWS DevOps Agent to your OpenSearch observability data through MCP, with three hosting paths so Region availability doesn’t block you: self-managed ECS (everywhere today), AgentCore (low-ops, where available), or the built-in 3.3+ endpoint. The loop is now closed: the same domain that stores your data and fires alerts becomes the investigation source, and the agent can help determine why an alert fired automatically.

Start with one alert that fires frequently and costs your team time to investigate manually. Connect it through this pipeline, watch the agent produce its first root cause analysis, and iterate from there.


About the authors

Sitaraman Vijay Krishna

Sitaraman Vijay Krishna

Sitaraman is a Senior Technical Account Manager at AWS, where he works with customers on Generative AI, Agentic AI, and AI observability, including hands-on adoption of the AWS DevOps Agent. Outside work, he’s a lifelong sports fan who’s as happy on the field as watching from the stands.

Prateek Sethi

Prateek Sethi

Prateek is a Senior Technical Account Manager who excels in architecting and implementing complex distributed systems, particularly transforming operations for global manufacturing and retail organizations. His passion for customer success drives him to nurture long-term partnerships, guiding organizations through their digital transformation journeys while ensuring optimal outcomes. When not solving technical challenges, Prateek enjoys exploring European cities on his motorcycle.

Optimize consumer rebalancing on Amazon MSK with next generation protocol

Post Syndicated from Yashika Jain original https://aws.amazon.com/blogs/big-data/optimize-consumer-rebalancing-on-amazon-msk-with-next-generation-protocol/

If you run large consumer groups on Apache Kafka and Amazon Managed Streaming for Apache Kafka (Amazon MSK), you’ve likely experienced the pain of slow rebalances: processing stalls across all consumers, “rebalance storms” triggered by routine scaling events, and prolonged recovery times that impact downstream applications. With the classic rebalance protocol, even a single consumer joining or leaving the group forces a global synchronization barrier, pausing every consumer regardless of whether its partition assignments changed.

The KIP-848 consumer protocol, introduced in Apache Kafka 4.0, fundamentally redesigns how consumer group rebalancing works. Also referred to as “the Next Generation Consumer Rebalance Protocol”, KIP-848 shifts coordination logic from the client to the broker-side group coordinator. This supports fully incremental, server-driven rebalancing that significantly improves performance for large consumer groups. You can use the consumer protocol on Amazon MSK on all 4.x Apache Kafka versions on both MSK Standard and Express brokers.

In this post, we explain how the consumer protocol works, how to enable it on Amazon MSK, and how to diagnose and resolve slow rebalancing issues to help improve performance.

The classic protocol compared to the consumer protocol

The classic protocol relied on client-side rebalance logic with a global synchronization barrier. Every rebalance caused all consumers in the group to pause processing simultaneously regardless of whether their partition assignments were changing. This led to “rebalance storms” in large consumer groups where cascading rebalances could take minutes to resolve. The CooperativeStickyAssignor is a client-side partition assignment strategy that supports incremental, cooperative rebalancing. It significantly improves rebalance performance and minimizes disruption to groups during rebalance events. However, it still suffers from bottlenecks as consumer group size and partition count increase. For large workloads, client-side rebalancing behavior can result in longer rebalancing times and require significant client tuning and monitoring during rebalances.

The consumer protocol addresses these limitations by moving all rebalancing logic to the server. The broker handles coordination using a continuous heartbeat mechanism and server-driven reconciliation process. Only affected partitions move during a rebalance, and consumers with unchanged assignments continue processing uninterrupted. This results in faster recovery compared to the classic protocol, and improved scalability as workloads grow.

The following table compares the classic and next generation protocols across key dimensions.

Aspect Classic Protocol Next Generation Protocol (KIP-848)
Rebalance logic Client-side Fully server-driven
Consumer impact Depends on assignor, all consumers pause, or rebalance is limited by group size Only affected consumers impacted, scales effectively as groups grow
Mechanism Client-side algorithm and cross-group coordination Incremental, async reconciliation
Commit processing Paused during rebalance Able to progress during rebalance
Scalability Complex, fragile at scale Resilient, broker-driven
Rebalance storms Common in large groups Eliminated

Server-side configuration

With the consumer protocol, key parameters are now configured on the server rather than the client:

  • group.consumer.heartbeat.interval.ms – Controls the consumer heartbeat interval (server-side).
  • group.consumer.session.timeout.ms – Controls the session timeout (server-side).
  • group.consumer.assignors – Specifies available assignors (uniform and range by default).

In Amazon MSK Express brokers, these configurations are read-only and cannot be modified. In Amazon MSK Standard brokers, these configurations are managed with broker configurations. To update these configurations in Amazon MSK Standard brokers, refer to Update the configuration of an Amazon MSK cluster.

When to use the consumer protocol

The consumer protocol provides the most benefit to workloads with the following requirements:

  • Large consumer groups: Groups with many consumers and partitions see the most significant improvements because of the elimination of global synchronization barriers.
  • High-availability applications: Applications that cannot afford processing interruptions benefit from continuous message processing during rebalances. Financial services, real-time analytics, and fraud detection systems are ideal candidates.
  • Frequently rebalancing environments: Automatic scaling deployments, Kubernetes with frequent pod restarts, or continuous integration and continuous delivery (CI/CD) environments experience significantly less disruption.
  • Dynamic partition scaling: Workloads that regularly add partitions or topics benefit from the incremental, server-driven approach.

Prerequisites

Before you begin, make sure that you have the following:

  • An Amazon MSK cluster running Apache Kafka version 4.0 or later (both MSK Standard and Express brokers are supported).
  • A Kafka client library that supports the KIP-848 consumer protocol (see Step 4 for supported versions).
  • Basic familiarity with Apache Kafka consumer groups and partition assignment.
  • An AWS account with appropriate permissions to manage your MSK cluster.

Enabling the consumer protocol on Amazon MSK

The following steps walk you through verifying your cluster version, configuring your consumer client, removing deprecated configurations, and confirming client library support.

Step 1: Verify cluster version

The consumer protocol requires Apache Kafka 4.0 or later. To use the consumer protocol on Amazon MSK, verify that your cluster is running Apache Kafka version 4.0.x or later. You can verify your cluster’s Apache Kafka version using the AWS Management Console, AWS Command Line Interface (AWS CLI), or AWS SDKs:

aws kafka describe-cluster-v2 --cluster-arn <your-cluster-arn> \
    --query "ClusterInfo.Provisioned.CurrentBrokerSoftwareInfo.KafkaVersion"

If your cluster is running Apache Kafka 4.0.x or later, the consumer protocol is automatically enabled on the server and ready to use. No additional server-side feature flag verification is needed.

Step 2: Configure consumer client

Set group.protocol=consumer in your consumer configuration. The protocol is not enabled by default:

# confluent-kafka-python example
config = {
    'bootstrap.servers': bootstrap_servers,
    'group.id': group_id,
    'group.protocol': 'consumer',  # Required — defaults to 'classic' if omitted
    'auto.offset.reset': 'earliest'
}

The consumer protocol can be changed in-place for existing consumer groups. When you update the group.protocol, perform a rolling restart of your consumers. The broker-side group coordinator automatically handles the upgrade to the consumer protocol and handles classic protocol requests from old clients alongside the upgraded clients.

Step 3: Remove deprecated client configurations

When the consumer protocol is enabled, the following client-side configurations are no longer supported because they are controlled by the brokers:

  • heartbeat.interval.ms.
  • session.timeout.ms.
  • partition.assignment.strategy.

Step 4: Verify client library support

Verify that your Kafka client version supports the consumer protocol:

  • Java clients: Generally available (GA) in Apache Kafka 4.0+.
  • confluent-kafka-python: Version 2.12.0+ (GA support for KIP-848). See the release notes.
  • librdkafka-based clients (Go, .NET, C/C++): Based on librdkafka 2.12.0+.

Note: For other Kafka client libraries, verify your client library’s documentation for group.protocol=consumer support before enabling the next generation protocol. If your client doesn’t support KIP-848, it will continue to use the classic protocol.

Diagnosing slow consumer group rebalancing with the consumer protocol

Even after enabling the consumer protocol, you may encounter situations where consumer group rebalancing takes longer than expected. The following sections help you diagnose and resolve these issues.

Common symptoms

  • Consumer group rebalancing takes longer than expected despite setting group.protocol=consumer.
  • Consumers pause processing during rebalances.
  • Broker logs show “member session expired” or “fenced” messages.
  • Frequent rebalances triggered during rolling deployments or pod restarts.

Step 1: Confirm the consumer protocol is active using broker logs

Before troubleshooting performance, verify which protocol your consumers are actually using. Check broker logs in Amazon CloudWatch Logs Insights. The log patterns differ significantly between protocols.

Consumer protocol expected logs:

Key indicators: “consumer protocol”, “epoch” terminology, “target assignment” with server-side assignor, “fenced” for member removal.

[GroupCoordinator id=X] [GroupId <group-id>] Member <member-id> joins the consumer group using the consumer protocol.
[GroupCoordinator id=X] [GroupId <group-id>] Bumped group epoch to 309 with metadata hash 4064309670987706693.
[GroupCoordinator id=X] [GroupId <group-id>] Computed a new target assignment for epoch 309 with 'uniform' assignor in 0ms.
[GroupCoordinator id=X] [GroupId <group-id>] Member <member-id> fenced from the group because the member session expired.

Classic protocol expected logs:

Key indicators: “PreparingRebalance” state, “old generation” terminology, “Assignment received from leader”.

If you see classic protocol logs, the consumer protocol is not active. Proceed to Step 2 to troubleshoot why.

[GroupCoordinator id=X] Preparing to rebalance group <group-id> in state PreparingRebalance with old generation X
[GroupCoordinator id=X] Stabilized group <group-id> with X members
[GroupCoordinator id=X] Assignment received from leader for group <group-id>

Step 2: Troubleshoot why the consumer protocol is not active

Verify that your client configuration, client library versions, and cluster versions support the consumer protocol, as described in the preceding Step 1 through Step 4.

Step 3: Resolve slow rebalancing when KIP-848 is active

After you verify the consumer protocol is active but rebalancing is still slow, investigate the following causes:

A. Consumer session timeout causing premature member removal

With the consumer protocol, session timeout is server-controlled through group.consumer.session.timeout.ms (default: 45 seconds). The diagnostic path depends on whether you are using static group membership. The following table outlines the diagnostic path and recommended actions for each scenario.

Scenario Symptom Root cause Recommended action
With static group membership (group.instance.id configured) Slow rebalancing when a static member terminates without calling consumer.close() The coordinator waits for the full session timeout before reassigning partitions. This is the most common cause of slow rebalancing in containerized environments. MSK Standard: Implement graceful shutdown to trigger an immediate leave-group request, or increase the session timeout: group.consumer.session.timeout.ms=60000 (default is 45000). MSK Express: This configuration is not editable in Amazon MSK Express clusters. For Amazon MSK Express, optimize your client’s cold starts to allow members to restart within the 45 second consumer session timeout.
Without static group membership Session timeouts expiring during normal operations Your consumer is freezing or becoming unresponsive, which prevents heartbeats from reaching the coordinator.

Investigate long-running message processing, garbage collection pauses, network connectivity issues, or resource exhaustion on the consumer host. Look for this in broker logs:

[GroupCoordinator id=X] [GroupId <group-id>] Member <member-id> has timed out

B. Missing graceful shutdown handling

When consumers terminate without calling consumer.close(), the coordinator waits for the full session timeout before removing the member. This is the most common cause of slow rebalancing in containerized environments.

Resolution: Implement proper SIGTERM handling to trigger an immediate leave-group:

import signal
import sys
from confluent_kafka import Consumer

class GracefulKafkaConsumer:
    def __init__(self, config):
        self.running = True
        self.consumer = Consumer(config)
        signal.signal(signal.SIGTERM, self.shutdown_handler)
        signal.signal(signal.SIGINT, self.shutdown_handler)

    def shutdown_handler(self, signum, frame):
        print(f"Received signal {signum}, initiating graceful shutdown...")
        self.running = False

    def consume_loop(self):
        self.consumer.subscribe(['your-topic'])
        while self.running:
            msg = self.consumer.poll(timeout=1.0)
            if msg is None:
                continue
            # Process message

        print("Closing consumer gracefully...")
        self.consumer.close()  # Sends LeaveGroup — triggers immediate rebalance
        sys.exit(0)

For Kubernetes, verify that terminationGracePeriodSeconds allows time for consumer.close() to complete:

spec:
  terminationGracePeriodSeconds: 60
  containers:
    - name: kafka-consumer

C. Frequent rebalances from unstable consumers

If consumers repeatedly join and leave (out-of-memory (OOM) kills, CrashLoopBackOff, or short-lived tasks), each event triggers a new rebalance epoch.

Resolution: Use static membership by assigning a unique group.instance.id:

config = {
    'bootstrap.servers': bootstrap_servers,
    'group.id': 'my-group',
    'group.protocol': 'consumer',
    'group.instance.id': f'consumer-{unique_identifier}'  # Unique per consumer
}

With static membership:

  • Short restarts within the session timeout don’t trigger rebalances.
  • The consumer rejoins with the same partition assignment.
  • Scaling up (adding new consumers) still works. New group.instance.id values trigger assignment of unassigned partitions only.

Monitoring and validation

After applying changes, confirm the improvement:

  • Check broker logs: Confirm that “member session expired” messages no longer appear during normal operations or deployments.
  • Monitor consumer lag: Use the SumOffsetLag and EstimatedMaxTimeLag Amazon CloudWatch metrics to verify that lag returns to zero quickly after a rebalance.
  • Describe consumer group: Use kafka-consumer-groups.sh --describe to verify that all members are active and stable.

Conclusion

After implementing the consumer protocol, you should observe the following behavior for consumer group rebalances:

  • Consistently faster rebalance times compared to the classic protocol.
  • Fewer session timeout-related rebalances.
  • More stable consumer group membership.
  • Smooth scaling operations without disrupting existing consumers.
  • Fewer unnecessary rebalances during consumer restarts when using static membership.
  • Clean consumer departures without waiting for timeout expiration when using graceful shutdown.

To get started, try the consumer protocol in your non-production workloads and observe the rebalance improvements as you scale your workload up and down.

To learn more about Amazon MSK and the consumer rebalance protocol, see the following resources:

 


About the authors

Yashika Jain

Yashika Jain

Yashika is a Senior Cloud Analytics Engineer at AWS, specializing in real-time analytics and event-driven architectures. She is committed to helping customers by providing deep technical guidance, driving best practices across real-time data platforms and solving complex issues related to their streaming data architectures.

Vinayaka Gangadhar

Vinayaka Gangadhar

Vinayaka is an Analytics Specialist at Amazon Web Services (AWS), where he helps customers build and troubleshoot scalable data platforms and derive meaningful insights through AWS analytics services, with deep expertise in Amazon Redshift and Amazon OpenSearch. When not solving complex analytics challenges, he enjoys exploring new technologies and spending quality time with his family.

Kalyan Janaki

Kalyan Janaki

Kalyan is Senior Big Data & Analytics Specialist with Amazon Web Services. He helps customers architect and build highly scalable, performant, and secure cloud-based solutions on AWS.

Running multi-day AZ evacuation drills with ARC Zonal Shift

Post Syndicated from Antoine Boucherie original https://aws.amazon.com/blogs/architecture/running-multi-day-az-evacuation-drills-with-arc-zonal-shift/

Running a multi-day Availability Zone evacuation drill with ARC Zonal Shift is an effective way to prove your application can withstand a sustained impairment. Multi-Availability Zone deployment is an architectural best practice for building resilient applications on AWS, but there is a gap between deploying multi-AZ and proving it works under sustained stress. Traditional disaster recovery (DR) tests validate the failover mechanism. A typical test shifts traffic, confirms targets respond, and rolls back within minutes. These short exercises don’t surface the issues that only appear over hours or days.

A multi-day AZ evacuation forces time-dependent behaviors to play out completely, exposing failure modes that brief tests miss:

  • Auto Scaling policies that aren’t tuned for sustained N-1 operation over a full day.
  • Deployment pipelines that don’t validate AZ health before placing new workloads.
  • Stale DNS or cached database endpoints.
  • Time-based routine operational processes tested against an N-1 architecture (certificate rotations, credential and secret rotations, maintenance windows, log rotation, backup automation, and cron jobs).
  • Long-lived database connections pinned to a specific AZ that are only used infrequently.
  • Recovery after a sustained multi-day shift, which is a different operational procedure than rolling back to a warm, nearly identical AZ within minutes.

By shifting all traffic away from a single AZ for 48–72 hours using Amazon Application Recovery Controller (ARC) Zonal Shift, you force your architecture to sustain full production load on N-1 zones. This proves capacity sufficiency, database stability, and client reconnection behavior, but most importantly, that your teams can operate normally for days on N-1 capacity.

This post shows you how to plan and run a multi-day AZ evacuation drill across Amazon Elastic Container Service (Amazon ECS), Amazon Elastic Kubernetes Service (Amazon EKS), Amazon Relational Database Service (Amazon RDS) for PostgreSQL, and Amazon Aurora PostgreSQL, with step-by-step CLI commands, a prerequisites section, observability metrics, and restore procedures.

Why financial services institutions are doing this already

Financial services organizations face unique regulatory pressure to demonstrate, not just document, their disaster recovery capabilities. Across the world, there is an increasing focus on building and demonstrating operational resilience within regulated entities. This is shifting the mindset from “show us your runbook” to “show us the evidence”. This uplift in control environments is driving financial services companies to conduct live DR testing under realistic conditions and produce auditable proof of recovery within defined timeframes. Some insurers and banks are now running periodic AZ evacuation drills as part of their operational resilience programs, shifting from “we have multi-AZ” to “we have proven multi-AZ”. Using ARC Zonal Shift, you can shift traffic at the infrastructure layer without changing application code. It works natively across Application Load Balancer, Network Load Balancer, Amazon Elastic Compute Cloud (Amazon EC2) Auto Scaling groups, and Amazon EKS clusters.

Solution overview

In this walkthrough, we demonstrate how to evacuate an Availability Zone for a multi-tier digital platform running on AWS.

We deliberately include both ECS and EKS, and two database engines, to show the evacuation procedure for each major service type readers are likely to run. The architecture is illustrative only.

The following table outlines the architecture:

Layer Components Multi-AZ Configuration
Traffic ingress Application Load Balancer (ALB) fronting ECS Deployed across 3 AZs, cross-zone load balancing activated
Compute (containers) Amazon ECS (Fargate) Stateless tasks distributed across 3 AZ subnets
Traffic ingress Network Load Balancer (NLB) fronting EKS Deployed across 3 AZs, cross-zone load balancing activated
Compute (Kubernetes) Amazon EKS or EKS Auto Mode Stateless services with topology spread constraints across 3 AZs
Database Amazon RDS for PostgreSQL Multi-AZ: primary in AZ A, standby in AZ B
Database Amazon Aurora PostgreSQL Writer in AZ A, reader in AZ B. Storage replicated across all 3 AZs

Note that the Aurora storage layer differs from standard RDS Multi-AZ. Aurora synchronously replicates data to six storage nodes across Availability Zones independently of compute instances, so only the writer or reader instance needs to be failed over as the storage remains fully available throughout the drill.

In this walkthrough, we evacuate AZ A, the zone hosting the RDS primary and Aurora writer.

Multi-tier architecture spanning three Availability Zones: an Application Load Balancer fronting Amazon ECS and a Network Load Balancer fronting Amazon EKS, with Amazon RDS for PostgreSQL and Amazon Aurora PostgreSQL databases, before evacuating AZ A.

How ARC Zonal Shift works

When you initiate a zonal shift, ARC takes two coordinated actions for Amazon Route 53 and Elastic Load Balancers:

  1. DNS removal: The load balancer’s IP address in the affected AZ is removed from DNS. New client queries don’t resolve to that endpoint.
  2. Cross-zone traffic blocking: Load balancer nodes in the remaining AZs stop routing requests to targets in the shifted AZ, even when cross-zone load balancing is enabled.

For Amazon EKS clusters with zonal shift enabled, ARC goes further. It performs the following actions:

  • Cordons all nodes in the impacted AZ, preventing new pod scheduling.
  • Removes pod endpoints in the impacted AZ from EndpointSlice resources, redirecting east-west service-to-service traffic to healthy AZs.
  • Suspends AZ rebalancing for managed node groups and updates ASGs to launch instances only in healthy AZs.
  • Preserves nodes and pods in the shifted AZ (they are not terminated), keeping full capacity available for when the shift ends.

Combined with service-specific procedures for ECS task redistribution and database failover, this creates a complete AZ evacuation across all three traffic dimensions: north-south ingress, east-west service communication, and outbound database connections.

ARC Zonal Shift is a data plane operation by design. Because it works independently of the AWS control plane, it remains available even during an AZ impairment. The other steps in this walkthrough (ECS service updates, manual RDS failovers, manual Aurora failovers, subnet group modifications) are control plane operations. For a planned drill, this distinction has no practical impact because the control plane is healthy. During a real AZ impairment, prioritize the data plane action (start the zonal shift first to stop traffic immediately) and perform control plane operations only after traffic has already been shifted.

Prerequisites

Configure these prerequisites well in advance of your first shift. These are foundational settings that verify that your architecture is shift-ready at all times. For this walkthrough, you should have the following:

  • An AWS account
  • A multi-tier application deployed across 3 Availability Zones with ALB/NLB, ECS or EKS workloads, and RDS or Aurora databases.
  • IAM permissions to manage ARC Zonal Shift, ECS, EKS, and RDS resources.
  • AWS Command Line Interface (AWS CLI) v2 installed and configured.
  • Familiarity with ARC Zonal Shift concepts.
  • Amazon CloudWatch dashboards with per-AZ metric breakdowns (fault rate, latency, and target health).
  • Auto Scaling policies validated for sustained N-1 AZ operation.

Specifically for Elastic Load Balancing (ELB):

  • ALB/NLB deregistration delay set to 60 seconds, which allows existing connections to drain quickly after a shift instead of the default 300 seconds.
  • target_group_health.dns_failover.minimum_healthy_targets.count configured on each target group.

Specifically, for EKS:

  • kubectl installed and configured for your EKS cluster.
  • Turn on Topology Aware Routing on EKS services (or configure Istio locality-aware load balancing).
  • Zonal shift activated on your EKS cluster (one-time setup).

Specifically, for ECS:

  • ECS stopTimeout set to 55 seconds in task definitions, slightly below the ALB deregistration delay so tasks finish in-flight requests before being force-stopped, avoiding 502 errors during the drain window.

Specifically, for RDS:

  • Verify that your RDS primary and standby are provisioned in different Availability Zones.

What changes for a multi-day shift

The mechanics of starting a zonal shift are the same whether you run it for 1 hour or 72 hours. What changes is the operational surface area:

  • Expiry management: Zonal shifts have a maximum duration. You must monitor and extend them before they expire, or traffic returns to the shifted AZ unexpectedly.
  • Scaling drift: Over days, Auto Scaling events in healthy AZs may create capacity imbalances. Monitor and cap scaling so recovery doesn’t overload the returning AZ.
  • Connection pool cycling: After 24+ hours, most client connections will have recycled. This validates DNS TTL compliance across your entire client fleet, something a 1-hour test won’t fully exercise.
  • Operational confidence: Teams will learn to deploy, patch, and troubleshoot in a reduced AZ environment. A multi-day drill forces this to happen naturally rather than in a controlled window.
  • Safe recovery: After days at N-1 capacity, restoring the shifted AZ requires careful ordering. Verify health, scale back gradually, and reintroduce traffic incrementally rather than all at once.

Solution details

Each section below walks through the zonal shift procedure for one layer of the architecture, starting with the compute tier and finishing at the database layer.

Amazon ECS — Zonal Shift with task redistribution

For ECS services behind an ALB, initiating a zonal shift at the load balancer layer stops new traffic from reaching targets in the evacuated AZ. Existing ECS tasks in that AZ remain running but stop receiving requests. To perform a complete evacuation, follow these steps:

Step 0. Before starting, verify ARC Zonal Shift is enabled on the load balancer (disabled by default)

aws elbv2 modify-load-balancer-attributes \
    --load-balancer-arn $ALB_ARN \
    --attributes Key=zonal_shift.config.enabled,Value=true

Step 1. Initiate the zonal shift on the load balancer.

aws arc-zonal-shift start-zonal-shift \
    --resource-identifier $ALB_ARN \
    --away-from $AZ_ID_TO_EVACUATE \
    --expires-in "72h" \
    --comment "Multi-day AZ evacuation drill"

Zonal shifts expire after the duration set in --expires-in. If the shift expires before you cancel it, traffic automatically returns to the shifted AZ. For a multi-day drill, monitor the remaining time and extend before expiry using:

aws arc-zonal-shift update-zonal-shift \
    --zonal-shift-id $SHIFT_ID \
    --resource-identifier $RESOURCE_ARN \
    --expires-in "24h" \
    --comment "Extending drill"

When cross-zone load balancing is enabled (the default for ALB), the shift instructs load balancer nodes in healthy AZs not to route requests to targets in the impaired AZ. Targets are fully isolated regardless of your cross-zone configuration.

Step 2. Restrict new task placement to healthy AZs.

Update the ECS service’s network configuration to exclude subnets in the evacuated AZ. This prevents new tasks from launching in the shifted zone:

aws ecs update-service \
    --cluster $CLUSTER_NAME \
    --service $SERVICE_NAME \
    --network-configuration "awsvpcConfiguration={subnets=[$AZB_SUBNET,$AZC_SUBNET],securityGroups=[$SG_ID],assignPublicIp=DISABLED}"

Step 3. If needed, scale to verify N-1 AZ capacity.

aws ecs update-service \
    --cluster $CLUSTER_NAME \
    --service $SERVICE_NAME \
    --desired-count $N_MINUS_1_COUNT

Note: updating the ECS service network configuration and count that you want are control plane operations. Perform these changes before the drill starts, not during a real AZ impairment when control plane availability may be degraded.

We recommend that you pre-scale your services to handle the loss of an AZ’s worth of capacity before the drill. Your architecture should tolerate AZ loss without needing to scale reactively. See static stability in the Amazon Builders’ library.

In a 3 AZ environment, pre-scaling for N-1 capacity means running approximately 50% more compute than your baseline peak requires. If the cost isn’t justifiable for all workloads, consider scheduled scaling policies that increase capacity during planned drill windows, Auto Scaling with aggressive scale-out thresholds, or load shedding mechanisms. Keep in mind that during an unplanned impairment, you won’t have time to scale reactively. Workloads that aren’t pre-scaled will operate in a degraded state until scaling catches up, which can take minutes under load.

Step 4. Monitor task distribution.

aws ecs describe-tasks \
    --cluster $CLUSTER_NAME \
    --tasks $(aws ecs list-tasks --cluster $CLUSTER_NAME --service-name $SERVICE_NAME --query 'taskArns' --output text) \
    --query 'tasks[].[taskArn,availabilityZone]' --output table

Restore: Revert the network configuration to include all three AZ subnets, then cancel the zonal shift. Tasks will gradually rebalance across all AZs during subsequent deployments.

Important: zonal shift won’t work for single-AZ target groups as the ALB will refuse the shift if healthy targets only exist in one Availability Zone. Verify each target group has targets registered in at least two AZs before proceeding. For more details, refer to Application Load Balancers in the ARC documentation.

Amazon EKS — Zonal Shift with EndpointSlice isolation

Amazon EKS natively supports ARC zonal shift. When you turn on this capability and trigger a shift, ARC handles both the infrastructure and Kubernetes networking layers automatically.

What ARC does when you shift an EKS cluster:

  • Nodes in the impacted AZ are cordoned (no new pod scheduling).
  • The built-in Kubernetes EndpointSlice controller removes pod endpoints in the impacted AZ, so east-west service traffic is automatically redirected to pods in healthy AZs.
  • For managed node groups, AZ rebalancing is suspended and ASGs are updated to only launch in healthy AZs.
  • Nodes and pods in the shifted AZ are not terminated, ensuring full capacity is immediately available when the shift ends.

Step 1. Activate zonal shift for your EKS cluster (one-time setup):

aws eks update-cluster-config \
    --name $CLUSTER_NAME \
    --zonal-shift-config enabled=true

Step 2. Start the zonal shift on both the load balancer and EKS cluster:

# Shift north-south traffic at the load balancer
aws arc-zonal-shift start-zonal-shift \
    --resource-identifier $NLB_ARN \
    --away-from $AZ_ID_TO_EVACUATE \
    --expires-in "72h" \
    --comment "Multi-day AZ evacuation drill - north-south traffic"

# Shift east-west traffic within the EKS cluster
aws arc-zonal-shift start-zonal-shift \
    --resource-identifier $EKS_CLUSTER_ARN \
    --away-from $AZ_ID_TO_EVACUATE \
    --expires-in "72h" \
    --comment "Multi-day AZ evacuation drill - east-west traffic"

Step 3. Verify EndpointSlice update: confirm pods in the shifted AZ are no longer receiving traffic:

# Endpoints should only show AZ B/AZ C
kubectl get endpointslices -l kubernetes.io/service-name=$SERVICE_NAME -o yaml | \
    grep -A2 "zone:"

Step 4. Verify node and pod status:

# Nodes in evacuated AZ should show SchedulingDisabled
kubectl get nodes -l topology.kubernetes.io/zone=$AZ_NAME_TO_EVACUATE

# Confirm traffic distribution across healthy AZs
kubectl get pods -o wide -l app=$APP_LABEL

Verify your pods use topologySpreadConstraints with maxSkew: 1 on topology.kubernetes.io/zone and are pre-scaled to handle N-1 AZ load. The zonal shift doesn’t evict pods or trigger autoscaling by itself.

Note that ARC zonal shift doesn’t control outbound connections from pods to external dependencies like Amazon RDS. If your pods connect to AZ-specific database endpoints, consider using Istio with locality-aware routing. For implementation details, refer to End-to-end recovery from AZ impairments in EKS using Zonal Shift and Istio.

For Aurora, the cluster endpoint automatically routes to the current writer regardless of AZ, so no Istio configuration is needed for writer traffic. However, if you use AZ-specific reader instance endpoints, configure Istio ServiceEntry resources for each endpoint and apply a DestinationRule with localityLbSetting to prefer healthy AZs. This directs outbound database traffic to follow the same shift pattern as your north-south and east-west traffic.

Restore: Cancel both zonal shifts. ARC automatically uncordons nodes, adds pod endpoints back to EndpointSlices, and restores AZ rebalancing. Traffic returns to all three AZs with full capacity already in place.

aws arc-zonal-shift cancel-zonal-shift \
    --zonal-shift-id $EKS_SHIFT_ID \
    --resource-identifier $EKS_CLUSTER_ARN

aws arc-zonal-shift cancel-zonal-shift \
    --zonal-shift-id $NLB_SHIFT_ID \
    --resource-identifier $NLB_ARN

Amazon RDS for PostgreSQL — multi-AZ failover

Regarding Amazon RDS for PostgreSQL in a Multi-AZ deployment, if the primary instance resides in the AZ being evacuated, you must trigger a failover to the synchronous standby. RDS handles this through a reboot with failover.

Prerequisite: verify your RDS primary and standby are provisioned in different Availability Zones. If both are in the same zone, the following steps wouldn’t evacuate the zone as intended.

Step 1. Check current primary location:

aws rds describe-db-instances \
    --db-instance-identifier $RDS_INSTANCE \
    --query 'DBInstances[0].[DBInstanceIdentifier,AvailabilityZone,MultiAZ,SecondaryAvailabilityZone]' \
    --output table

Step 2. If the primary is in the evacuated AZ, manually force failover:

aws rds reboot-db-instance \
    --db-instance-identifier $RDS_INSTANCE \
    --force-failover

Step 3. Wait for availability and verify the new primary AZ:

aws rds wait db-instance-available \
    --db-instance-identifier $RDS_INSTANCE

aws rds describe-db-instances \
    --db-instance-identifier $RDS_INSTANCE \
    --query 'DBInstances[0].AvailabilityZone'

After failover, RDS recreates the standby in the evacuated AZ automatically. This is acceptable for a sustained drill as the standby receives no client traffic. Monitor ReplicaLag to confirm replication health.

Step 4. (Optional) Remove the standby from the evacuated AZ.

If you want zero RDS presence in the evacuated Availability Zone, you can relocate the standby by modifying the DB subnet group:

  1. Create a manual snapshot as a safety net.
  2. Disable Multi-AZ on the instance.
  3. Modify the DB subnet group to include only healthy AZ subnets, removing the evacuated AZ subnet.
  4. Re-enable Multi-AZ so the new standby is created in one of the remaining healthy AZs.

This approach works the same way for Amazon RDS for PostgreSQL as it does for any RDS engine using Multi-AZ deployments. Note that RDS Multi-AZ modifications (disabling/re-enabling Multi-AZ, subnet group changes) can take several minutes to complete. Plan for this during the drill window.

Amazon Aurora PostgreSQL — writer failover & reader management

Aurora provides more control over AZ placement than standard RDS Multi-AZ. You can explicitly choose which reader to promote and manage reader placement across AZs using failover priority tiers.

Step 1. Identify the cluster topology:

aws rds describe-db-clusters \
    --db-cluster-identifier $CLUSTER_ID \
    --query 'DBClusters[0].DBClusterMembers[].{Instance:DBInstanceIdentifier,IsWriter:IsClusterWriter}'

aws rds describe-db-instances \
    --filters Name=db-cluster-id,Values=$CLUSTER_ID \
    --query 'DBInstances[].[DBInstanceIdentifier,AvailabilityZone,DBInstanceStatus]' \
    --output table

Step 2. If the writer is in the evacuated AZ, failover to a reader in a healthy AZ:

aws rds failover-db-cluster \
    --db-cluster-identifier $CLUSTER_ID \
    --target-db-instance-identifier $READER_IN_HEALTHY_AZ

Step 3. Wait for the cluster to stabilize:

aws rds wait db-cluster-available \
    --db-cluster-identifier $CLUSTER_ID

If your Aurora cluster has no pre-existing reader in a healthy AZ, writer promotion requires creating a new instance, which typically takes less than 10 minutes. Pre-provisioning a reader in a separate AZ reduces failover time, often to less than 30 seconds.

Step 4. (Optional) Remove the reader in the evacuated AZ and create one in a healthy AZ.

For a full AZ evacuation where you want zero database presence in the shifted zone:

# Delete the reader instance in the evacuated AZ
aws rds delete-db-instance \
    --db-instance-identifier $INSTANCE_IN_EVACUATED_AZ \
    --skip-final-snapshot

# Create a new reader in a healthy AZ
aws rds create-db-instance \
    --db-instance-identifier ${CLUSTER_ID}-reader-${TARGET_AZ} \
    --db-cluster-identifier $CLUSTER_ID \
    --db-instance-class $INSTANCE_CLASS \
    --engine aurora-postgresql \
    --availability-zone $TARGET_AZ

Step 5. Monitor replication and performance throughout the drill:

aws cloudwatch get-metric-statistics \
    --namespace AWS/RDS \
    --metric-name AuroraReplicaLag \
    --dimensions Name=DBInstanceIdentifier,Value=$READER_INSTANCE \
    --start-time $TIMESTAMP_5MIN_AGO \
    --end-time $TIMESTAMP \
    --period 60 --statistics Average

Monitoring the drill with CloudWatch

A multi-day drill is only as valuable as the evidence it produces. Unlike a brief failover test where you visually confirm targets respond, a 48–72-hour evacuation requires continuous, automated observation, capturing capacity trends, replication health, and latency shifts that only surface under sustained N-1 AZ load.

Before starting the drill, verify you have CloudWatch dashboards with per-AZ metric breakdowns for each layer of your architecture. During the drill, these metrics serve two purposes: real-time operational awareness and post-drill evidence for stakeholders.

Key metrics by layer

The following metrics give you real-time visibility into each layer of the architecture during the drill.

Application Load Balancer / Network Load Balancer

Metric Dimension What to watch
HealthyHostCount Per target group, per AZ Should drop to 0 in evacuated AZ. Stable in healthy AZs
UnHealthyHostCount Per target group, per AZ Targets in evacuated AZ may show unhealthy (expected)
RequestCount Per AZ Zero traffic in shifted AZ. Even distribution in remaining AZs
TargetResponseTime Per AZ Watch for latency increases in healthy AZs under concentrated load
HTTPCode_Target_5XX_Count Per target group Sustained increase signals capacity pressure

Amazon ECS

Metric Dimension What to watch
CPUUtilization Per service Should not exceed 70–80% sustained (indicates capacity headroom)
MemoryUtilization Per service Memory pressure under concentrated load
RunningTaskCount Per service Confirms tasks running only in healthy AZs
DesiredTaskCount vs RunningTaskCount Per service Gap indicates placement failures (check subnet/capacity)

Amazon EKS (using Container Insights)

Metric Dimension What to watch
node_cpu_utilization Per node, filtered by AZ Nodes in healthy AZs absorbing shifted load
pod_cpu_utilization Per pod/namespace Hotspot detection under N-1 operation
node_status_condition Per node Nodes in evacuated AZ should show SchedulingDisabled
pod_number_of_container_restarts Per pod Restart loops may indicate resource pressure

Amazon RDS for PostgreSQL

Metric Dimension What to watch
CPUUtilization Per instance Primary under higher load post-failover
DatabaseConnections Per instance Connection spike after failover (watch for pool exhaustion)
ReadIOPS / WriteIOPS Per instance I/O patterns shift when primary moves AZs
ReplicaLag Per standby Should stabilize within seconds after failover
FreeableMemory Per instance Memory pressure under full client reconnection

Amazon Aurora PostgreSQL

Metric Dimension What to watch
AuroraReplicaLag Per reader instance Establish your cluster baseline during normal operation. Sustained increases from baseline indicate storage pressure. Aurora Replicas typically lag 100 ms or less
CommitLatency Per writer Increased commit latency indicates write contention
BufferCacheHitRatio Per instance Drop below 99% may indicate working set doesn’t fit in memory
DatabaseConnections Per instance Client reconnection behavior after writer promotion
VolumeBytesUsed Per cluster Aurora storage is AZ-independent (should be unaffected)

Export your per-AZ CloudWatch dashboards as snapshots before, during, and after the drill. Combine these with the ARC zonal shift event history (available through list-zonal-shifts) to create an auditable evidence package.

Cleaning up

After completing the drill, restore services carefully. The order matters, especially if Auto Scaling has increased capacity in healthy AZs:

  1. Verify the evacuated AZ is healthy: confirm targets are registered, pods are running, and database instances are available.
  2. Cancel the EKS cluster zonal shift first (east-west traffic resumes). Monitor for errors as internal traffic rebalances.
  3. Cancel the load balancer zonal shift (north-south traffic resumes). Traffic returns gradually as DNS propagates.
  4. If Auto Scaling added capacity in the remaining AZs, scale back gradually over 15 to 30 minutes. Don’t remove capacity before traffic has redistributed evenly.
  5. If cross-zone load balancing is disabled, verify target_group_health.dns_failover.minimum_healthy_targets.count is configured. This allows Route 53 to only route traffic to an AZ once it has enough healthy targets registered, preventing the restored zone from receiving traffic before it’s ready to handle it.
  6. Monitor per-AZ metrics for 30 minutes after restoring to confirm even distribution and no error spikes.

No additional AWS resources are created by ARC Zonal Shift that incur ongoing charges. The zonal shift itself is available at no additional cost.

Conclusion

In this post, you learned how to run a sustained AZ evacuation drill using ARC Zonal Shift across Amazon Elastic Container Service (Amazon ECS), Amazon Elastic Kubernetes Service (Amazon EKS), Amazon RDS for PostgreSQL, and Amazon Aurora. By operating on N-1 Availability Zones for 48–72 hours, rather than a brief failover test, you produce evidence that your multi-AZ architecture delivers genuine, sustained resilience. This is particularly valuable for financial services organizations facing regulatory mandates that require demonstrated recovery capabilities.

To get started, use the prerequisites checklist and step-by-step procedures in this post. Begin in non-production, progress to production during low-traffic windows, and build toward sustained operation under shift. As confidence grows, activate zonal autoshift so AWS can shift traffic automatically when internal telemetry detects an impairment.

You can also use AWS Resilience Hub to assess your application’s resilience posture before and after the drill. It validates that your architecture meets your defined Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets.

Amazon Application Recovery Controller – Zonal Shift

Best practices for zonal shifts in ARC

Using cross-zone load balancing with zonal shift

New AWS Fault Injection Service recovery action for zonal autoshift

End-to-end recovery from AZ impairments in Amazon EKS using EKS Zonal Shift and Istio

Amazon EKS now supports Amazon Application Recovery Controller


About the authors

Implementing customer managed keys for AWS Lambda durable functions with Terraform

Post Syndicated from Rajdeep Banerjee original https://aws.amazon.com/blogs/compute/implementing-customer-managed-keys-for-aws-lambda-durable-functions-with-terraform/

If you run regulated workloads, you must control how persisted data is encrypted and who can access it. You need to manage encryption key rotation schedules, restrict decryption to authorized principals, and produce audit evidence that proves encryption controls are operating as designed.

AWS Lambda durable functions build resilient, multi-step workflows that survive failures through automatic checkpointing. The checkpoint mechanism persists execution state, including step results, payloads, and callback responses, to durable storage. For payment processing workloads, this persisted data is sensitive. AWS Lambda durable functions support customer managed keys from AWS Key Management Service (AWS KMS). A customer managed key gives you three controls: you set the key rotation schedule, you restrict decryption access through the key policy, and you generate per-function audit trails in AWS CloudTrail. A durable execution uses the same encryption key it started with for its entire lifetime. Changing or removing the key affects only executions that start after the change.

Updating the customer managed key policy to remove decrypt permissions, or disabling the key, stops the Lambda service from accessing previously checkpointed state. Customer managed key deletion is a permanent action, and all durable executions encrypted with that key become unrecoverable because the Lambda service has no mechanism to restore the data. Before scheduling key deletion, use the AWS KMS waiting period (7 to 30 days) and monitor AWS CloudTrail for Decrypt calls to confirm that the key is no longer in active use.

In this post, you learn to configure a customer managed key to encrypt durable execution data in an event-driven payment processing workflow. You create a symmetric encryption key in AWS KMS and define a key policy that grants the Lambda service, the function’s execution role, the function author, and durable execution operators only the AWS KMS actions each principal requires. You then configure the function to use the key for durable execution encryption and verify encryption operations through AWS CloudTrail logs. By the end, you have a deployable reference architecture you can adapt for regulated workloads running on Lambda durable functions.

To learn more about how AWS Lambda encrypts durable execution data, see Encrypting AWS Lambda durable execution data in the AWS Lambda Developer Guide.

Solution overview

The sample application implements an event-driven payment processing pipeline using Amazon DynamoDB, Amazon EventBridge, Amazon EventBridge Pipes, AWS Lambda, and Amazon SQS. The pipeline receives authorized payment transactions, validates and enriches them. A Lambda durable function applies business rules to the enriched transactions. The approved transactions are sent to a downstream settlement system for posting.

The following section covers the key architectural steps.

Architecture steps

  1. The upstream authorization system writes authorized payment records to a DynamoDB table.
  2. DynamoDB Streams captures each new record as an ordered change event.
  3. Amazon EventBridge Pipes polls the record from the DynamoDB stream. The pipe triggers a Lambda function as part of enrichment step for duplicate checking.
  4. The deduplication Lambda uses a DynamoDB table with conditional writes to identify duplicate inbound transactions based on transaction properties and time window.
  5. When the deduplication is successful, the pipe publishes an event to the Amazon EventBridge custom event bus.
  6. An Amazon EventBridge rule invokes a Lambda function for matching events. The function adds business context such as account type, bank routing details, and merchant category codes. The function publishes a new enriched event to the custom event bus.
  7. Another Amazon EventBridge rule matches the enriched events to a Lambda durable function. The durable function applies business rules to the incoming event. When the event passes all business rules, the function publishes a new event to the event bus.
  8. An Amazon EventBridge rule routes the approved event to an Amazon SQS queue preserving ordering for settlement and buffering against downstream throughput limits.
  9. The Posting Lambda function reads from the Amazon SQS and invokes the downstream posting subsystem to post the transaction. Finally, the function publishes a completion event to the event bus completing the transaction lifecycle.

With customer managed keys configured on DynamoDB, Amazon EventBridge, SQS, and the AWS Lambda durable function, every piece of persisted data in this pipeline is encrypted with keys you own and control. The walkthrough that follows shows you how to deploy this configuration with Terraform.

Figure 1 shows the reference architecture for this solution.

Reference architecture

Reference architecture for payment processing using Lambda durable functions

Figure 1: Payment processing using Lambda durable functions

Prerequisites

To deploy this solution, you need the following prerequisites:

  1. AWS account and CLI: An active AWS account with the AWS CLI installed and configured with appropriate credentials.
  2. Terraform: Terraform installed (version 1.0 or later) for infrastructure provisioning.
  3. Python environment: Python 3.11 or later, with pytest for running unit tests. The aws-durable-execution-sdk-python package requires Python 3.11 or later.
  4. AWS Identity and Access Management (IAM) permissions: The IAM permissions to create the resources. Follow the sample repository for the sample policy.
  5. Basic understanding and familiarity with AWS Serverless services.

Solution walkthrough

The following is a step-by-step guide to deploy and test the payment processing solution.

Step 1: Clone the repository

git clone https://github.com/aws-samples/sample-payment-processing-with-lambda-durable-functions.git
cd sample-payment-processing-with-lambda-durable-functions/source

Step 2: Run unit tests

Validate the payment processing logic locally before deploying:

cd lambda-src/business_rules
pip3 install -r requirements-test.txt
pytest test_app.py -v

This runs unit tests that cover transaction validation, business rule checks (foreign transaction detection, currency conversion, merchant type), event schema validation, and misconfiguration handling. The tests use the AWS Durable Execution Testing SDK to run the handler locally without deploying AWS resources.

Figure 2 shows an example of test results running locally.

Test run results of the business rules

Figure 2: Test run results of the business rules

Step 3: Inspect the Lambda durable functions construct

Open the payments-business-rules Lambda function in source/lambda-src/business_rules/business-rules-app.py for a sample Lambda durable function. Refer to Figure 3 for the code walkthrough.

Key features used

  1. @durable_execution decorator: Transforms a standard Lambda handler into a durable function handler. The durable execution SDK manages checkpointing automatically. No infrastructure changes are required.
  2. context.step("validate-transaction"): Validates that the transaction has a non-empty issuingCountryCode. The durable execution checkpoints the result (True or False) to durable storage. The durable execution restores checkpoint results instead of re-executing steps during the replay phase. This phase occurs whenever the function is re-invoked after an interruption such as a wait period completing, a failure, or a suspension. This checkpointed result is part of the durable execution data encrypted by your customer managed key.
  3. context.step("publish-posting-failure"): Publishes the full Amazon EventBridge envelope to Amazon SNS when validation fails. This step only runs on the failure path. The runtime checkpoints the Amazon SNS publish response to durable storage.
  4. context.parallel("run-business-rules"): Runs three independent rule checks concurrently: foreign transaction detection, currency conversion, and merchant type validation. Each branch checkpoints independently. If one branch fails, the others are not replayed on resume. Each branch result is persisted to durable storage and encrypted by the customer managed key.
  5. ctx.step("trigger-foreign-transaction-rule") (inside parallel): Compares billingAmount against transactionAmount. If they differ, it emits a ForeignTransactionFound event to Amazon EventBridge. This step is checkpointed independently within the parallel group.
  6. ctx.step("trigger-conversion-rate-rule") (inside parallel): Checks whether conversionRate equals 1. If so, it emits a CurrencyConversionTransactionFound event to Amazon EventBridge. This step is checkpointed independently within the parallel group.
  7. ctx.step("trigger-merchant-rule") (inside parallel): Checks whether merchantType equals AAFF. If so, it emits a WarningMerchantTypeTransactionFound event to Amazon EventBridge. This step is checkpointed independently within the parallel group.
  8. context.step("post-transaction-processed"): Emits the final TransactionPostingApproved event to Amazon EventBridge. This step is only reached when validation passes and all business rules complete. The runtime checkpoints the Amazon EventBridge response. On replay, if this step already succeeded, the event is not re-published, which guarantees exactly-once approval semantics.
  9. context.logger: Provides replay-aware logging throughout the handler. During replay of previously completed steps, log statements are suppressed to prevent duplicate log entries in Amazon CloudWatch.
Lambda durable functions code walkthrough

Figure 3: Sample Lambda durable functions code

Step 4: Deploy infrastructure with Terraform

Terraform currently doesn’t support attaching a customer managed key directly to the durable function. You create the symmetric key in Terraform and then associate the key with the durable function on the AWS Management Console. Refer to source/durable_kms.tf for the key configuration.

Initialize and deploy the AWS resources that make up the solution:

cd ../../
terraform init
terraform plan -var="region=us-east-2"

Review the plan output, then apply:

terraform apply -var="region=us-east-2" --auto-approve

Note: Replace us-east-2 with your preferred AWS Region.

On successful completion, Terraform outputs the AWS KMS key alias, key ARN, and DynamoDB Streams ARN used by the event-driven pipeline:

Apply complete! Resources: N added, 0 changed, 0 destroyed.

Outputs:

durable_kms_key_alias = "durable-function-encryption"
durable_kms_key_arn = "arn:aws:kms:us-east-2:xxxxxxxxxxxx:key/4e87d4c2-1190-4db4-8b97-46657f83ee00"
stream_arn = "arn:aws:dynamodb:us-east-2:xxxxxxxxxxxx:table/visa/stream/2026-03-16T14:25:47.847"

Step 5: Verify Lambda durable functions configuration

In the AWS Lambda console, navigate to the payments-business-rules function. Confirm that the function Type displays Durable, which indicates that the checkpoint-and-replay mechanism is active. Figure 4 shows the expected function configuration.

The durable function in the AWS Lambda console

Figure 4: The Lambda durable function in the AWS Lambda console

Step 6: Add the AWS KMS key to the Lambda durable function

  1. The durable function is not encrypted with a customer managed key. Figure 5 shows the function’s encryption configuration as empty.

    The Lambda durable function missing an AWS KMS key in the AWS Lambda console

    Figure 5: The Lambda durable function missing a customer managed key in the AWS Lambda console

  2. Choose Edit, then turn on Customize encryption settings as shown in Figure 6.

    The Lambda durable function encryption settings in the AWS Lambda console

    Figure 6: The Lambda durable function check encryption in the AWS Lambda console

  3. Select the AWS KMS key ARN created for the durable function. The key ARN is available in the Terraform output from Step 4. Figure 7 shows the key selection.

    Selecting the AWS KMS key for the durable function in the AWS Lambda console

    Figure 7: Select the AWS KMS key ARN for the durable function in the AWS Lambda console

  4. Choose Save and confirm that the durable function is now encrypted with a customer managed key, as shown in Figure 8.

    The Lambda durable function with the AWS KMS key in the AWS Lambda console

    Figure 8: AWS Lambda durable function with the customer managed key in the AWS Lambda console

Step 7: Execute a test payment

Invoke the payments-visa-mock Lambda function to simulate an end-to-end authorization flow. The mock function reads sample Visa authorization messages from a CSV file and writes them to DynamoDB, which triggers the event-driven pipeline. Figure 9 shows a sample test invocation.

Invoking the payments-visa-mock function to trigger the workflow

Figure 9: Invoke the payments-visa-mock function to trigger workflow

Figure 10 shows a sample response after invocation.

Test results from the payments-visa-mock function

Figure 10: Test results from the payments-visa-mock function to trigger workflow

The mock Lambda invocation creates records that follow the process described in the preceding architecture steps.

Step 8: Verify results

Open Amazon CloudWatch Logs and inspect the log group /aws/lambda/payments-business_rules. This log group belongs to the Lambda durable function for this use case. Figure 11 shows the CloudWatch log group on the console.

Search in CloudWatch Logs for the durable function log group

Figure 11: Search in CloudWatch

You see the complete business rules lifecycle for each transaction, as shown in Figure 12. The highlighted sections show all the business rules performed by the durable function. Each step is checkpointed by the runtime and encrypted by the customer managed key.

Search Results in lambda durable functions console

Figure 12: Search Results in lambda durable functions console

You can also check the other log groups to trace the full pipeline:

  • /aws/lambda/payments-enrich: Transaction enrichment logs.
  • /aws/lambda/payments-posting: Settlement posting logs.

Step 9: Verify the customer managed key configuration

You can verify the key configuration by using the AWS CLI:

aws lambda get-function-configuration \
    --function-name payments-business-rules \
    --query "DurableConfig" \
    --region us-east-2

Expected response:

{
    "KMSKeyArn": "arn:aws:kms:us-east-2:xxxxxxxxxxxx:key/4e87d4c2-1190-4db4-8b97-46657f83ee00",
    "RetentionPeriodInDays": 7,
    "ExecutionTimeout": 180
}

You can search in AWS CloudTrail to track the AWS KMS calls. When you configure or update the customer managed key on a durable function, Lambda validates the key policy with dry-run GenerateDataKey and Decrypt calls. These appear in CloudTrail with a DryRunOperationException error code, which confirms that the key policy permissions are correct and does not indicate an actual error. For more details, see Encrypting AWS Lambda durable execution data.

Clean up

To avoid ongoing charges, destroy all deployed resources using the following command:

terraform destroy -var="region=us-east-2" --auto-approve

Expected output:

Destroy complete! Resources: N destroyed.

Conclusion

In this post, you configured a customer managed key to encrypt durable execution data in a Lambda durable function. With a customer managed key, you control the key rotation schedule, restrict decryption access through the key policy, and generate per-function audit trails in AWS CloudTrail. You can revoke access to durable execution data at any time by updating the key policy, giving you full control over who can read execution state. In-flight executions stop at the next checkpoint call and new executions must be started after restoring access. For details, see When the customer managed key is unavailable.

For payment processors and financial institutions, encrypting durable execution data with a customer managed key satisfies compliance obligations for data-at-rest encryption, key governance, and access auditability across multi-step transaction workflows.

To get started, clone the sample repository and follow the preceding walkthrough. To learn more about Lambda durable functions, see the AWS Lambda Developer Guide.

Building an LLM-powered DAG failure analysis plugin for Amazon MWAA

Post Syndicated from Sushant Samantaray original https://aws.amazon.com/blogs/big-data/building-an-llm-powered-dag-failure-analysis-plugin-for-amazon-mwaa/

Apache Airflow has become the orchestration backbone for data pipelines across industries. But as those pipelines grow to hundreds of directed acyclic graphs (DAGs) spanning services like AWS Glue, Amazon EMR, Amazon Athena, and Amazon Redshift, debugging a single task failure turns into a significant operational challenge. When a task fails, data engineers sift through logs, cross-reference DAG configurations, and analyze error messages to find the root cause, delaying pipeline service level agreements (SLAs) and impacting team productivity.

In this post, we show you how to build a custom Apache Airflow plugin that integrates with Amazon Bedrock to automatically analyze DAG task failures and provide actionable diagnostic insights. The plugin deploys to Amazon Managed Workflows for Apache Airflow (Amazon MWAA) and provides AI-powered root cause analysis on demand.

The complete source code for this solution is available in the sample-aws-mwaa-llm-powered-plugin GitHub repository. Clone the repository and follow along as we explain the design decisions throughout this post.

Solution overview

Apache Airflow is a widely adopted open source platform for programmatically authoring, scheduling, and monitoring complex data pipelines. Teams use Airflow to orchestrate extract, transform, and load (ETL) processes, machine learning workflows, and data lake management across industries.

Amazon MWAA is a managed service that makes it straightforward to run Apache Airflow on AWS without the operational burden of managing the underlying infrastructure. With Amazon MWAA, you can focus on authoring workflows and business logic while AWS handles provisioning, patching, scaling, and securing your Airflow environments.

The solution uses the following AWS services:

The plugin adds an analysis view directly into your Airflow UI. At a high level, when a task fails and you trigger an analysis, the plugin automatically does the following:

  1. Retrieves the failed task instance metadata from the Airflow metadata database.
  2. Collects comprehensive context including task logs, DAG source code, and operator-specific scripts.
  3. Sends the enriched context to Amazon Bedrock for analysis.
  4. Returns a structured diagnostic report with root cause identification, step-by-step resolution, and prevention recommendations.

How it works

The preceding four steps happen behind a single Analyze Task action. The following diagram and pipeline show the high-level architecture and how the plugin carries them out.

Architecture of the task analyzer plugin connecting the Airflow UI on Amazon MWAA to Amazon Bedrock and Amazon S3

Figure 1: High-level architecture of the LLM-powered task analyzer plugin on Amazon MWAA

The plugin follows a multi-step analysis pipeline:

  1. User triggers analysis – From the Airflow UI, you select a failed task and choose Analyze Task.
  2. Context collection – The plugin retrieves task metadata, execution logs, and DAG source code from the Airflow metadata database and Amazon S3.
  3. Operator-aware enrichment – Based on the operator type, the plugin fetches the actual code or query that failed (for example, a PySpark script from AWS Glue or a SQL query from Amazon Athena).
  4. Foundation model analysis – The enriched context is sent to Amazon Bedrock, which returns a structured diagnostic report.
  5. Results presentation – The analysis displays in the Airflow UI with actionable recommendations.

All AWS API calls (Amazon Bedrock, Amazon S3, and AWS Glue) are authenticated through the aws_default Airflow connection. By default on Amazon MWAA, this connection has no static credentials, so boto3 falls back to the environment’s execution role. This means there are no keys to manage or rotate. If you need to call Amazon Bedrock or fetch scripts using a different identity, you can supply those credentials in the aws_default connection. This can be a dedicated IAM role or a cross-account principal, used instead of the execution role.

Operator-aware context collection

A key differentiator of this solution is its ability to understand different Airflow operator types and automatically fetch the associated code or queries. Unlike generic log analyzers, the plugin retrieves the actual code that failed, not just the error message.

The following table summarizes what the plugin fetches for each operator type:

Operator type What the plugin fetches Source
GlueJobOperator PySpark or Python script Amazon S3 (from the AWS Glue job definition)
EmrAddStepsOperator Spark or Python script Amazon S3 (from step arguments)
EmrServerlessStartJobOperator Spark script Amazon S3 (from job driver)
AthenaOperator SQL query Inline (from operator parameters)
RedshiftDataOperator SQL query Inline (from operator parameters)
BashOperator Bash command Inline (from operator parameters)
PythonOperator Python function DAG source code

This approach means the foundation model can analyze the actual logic that failed, correlating error messages with specific lines in your code for precise root cause identification.

Prerequisites

Before you begin, make sure that you have the following:

  • An Amazon MWAA environment running Apache Airflow 3.x (this walkthrough uses Airflow 3.2). The plugin registers its UI through the FastAPI-based plugin interface (fastapi_apps) introduced in Airflow 3.x. For setup instructions, see Get started with Amazon MWAA.
  • Access to Amazon Bedrock with the Anthropic Claude model family enabled in your AWS Region. This walkthrough uses Anthropic Claude, but you can adapt the plugin to work with Amazon Nova or other foundation models by modifying the prompt payload format in prompts.py. See Model access.
  • An AWS Identity and Access Management (IAM) execution role for Amazon MWAA with bedrock:InvokeModel and s3:GetObject permissions.
  • An Amazon S3 bucket backing your Amazon MWAA environment with bucket versioning enabled. See Create an Amazon S3 bucket for Amazon MWAA.
  • Python 3.10 or later installed locally.
  • The AWS Command Line Interface (AWS CLI) configured with appropriate permissions.

Note: In most Regions, you invoke Claude through an inference profile ID (for example, us.anthropic.claude-sonnet-4-5-20250929-v1:0) rather than a bare on-demand model ID. Run aws bedrock list-inference-profiles to confirm a model is ACTIVE before configuring it.

Plugin design

In this section, we explain the plugin design and its key components. The next section walks through deploying it to your Amazon MWAA environment.

Plugin structure

The plugin follows the standard Apache Airflow plugin architecture. The repository is organized as follows:

plugins/
├── task_analyzer_plugin.py    # Main plugin: FastAPI app, endpoints, registration
└── task_analyzer/
    ├── __init__.py
    ├── prompts.py             # Bedrock model configuration and prompt templates
    ├── script_utils.py        # Operator-specific script fetching logic
    ├── templates/
    │   └── index.html
    └── static/
        ├── css/
        │   └── styles.css
        └── js/
            ├── app.jsx
            ├── components.jsx
            ├── config.js
            ├── template.jsx
            └── utils.jsx

The repository also includes example DAGs that simulate various failure scenarios across different operator types.

Plugin registration

In Apache Airflow 3.x, the web component of a plugin is registered as a FastAPI application through the fastapi_apps attribute. In task_analyzer_plugin.py, the TaskAnalyzerPlugin class registers the FastAPI app under /task-analyzer and adds a view to the task instance page:

class TaskAnalyzerPlugin(AirflowPlugin):
    name = "task_analyzer_plugin"

    fastapi_apps = [
        {
            "app": app,
            "url_prefix": "/task-analyzer",
            "name": "Task Analyzer",
        }
    ]

    external_views = [
        {
            "name": "Analyze Task",
            "href": "/task-analyzer/",
            "url_route": "task_analyzer_view",
            "destination": "task_instance",
        }
    ]

Airflow automatically discovers any AirflowPlugin subclass in the plugins folder. No registration call or configuration change is needed. On Amazon MWAA, the file is delivered inside plugins.zip and extracted to /usr/local/airflow/plugins/.

Analysis engine

The analysis engine is the POST /api/analyze-task endpoint in task_analyzer_plugin.py. When you trigger an analysis, the endpoint performs the following steps:

  1. Retrieves AWS credentials from the aws_default Airflow connection. To override, edit the aws_default connection in the Airflow UI (Admin > Connections).
  2. Assembles a context dictionary from the request (task metadata, logs, DAG source).
  3. Enriches the context with an operator-specific script through fetch_and_add_operator_script.
  4. Builds the prompt using the template in prompts.py.
  5. Invokes Amazon Bedrock and returns the structured analysis.

Operator script fetching

The process_operator_script function in script_utils.py routes script retrieval based on operator type:

  • External scripts (AWS Glue, Amazon EMR) – The plugin calls the AWS Glue API to look up the job definition, then reads the PySpark script from Amazon S3. Amazon EMR handlers follow the same pattern, extracting the script path from the step configuration or job driver.
  • Inline scripts (Amazon Athena, Amazon Redshift, BashOperator, PythonOperator, DBTOperator) – The plugin reads the query or command directly from the task’s rendered template fields with no external API call.

The plugin implements smart fetching: for external scripts, it only makes the Amazon S3 API call when the error message contains code-relevant patterns (such as SyntaxError, TypeError, or data type mismatch). Infrastructure errors like timeouts skip the script fetch entirely, minimizing unnecessary API calls.

Prompt engineering

The prompt template in prompts.py provides the foundation model with:

  • Task metadata (DAG ID, task ID, run ID, state).
  • Error message and execution logs.
  • DAG source code.
  • Operator-specific script (when available).

The model produces a structured diagnostic report with root cause identification, step-by-step resolution, and prevention recommendations. Model IDs are configurable through Airflow Variables, so you can switch between Claude Sonnet and Claude Opus without redeploying the plugin.

Security measures

Before sending content to Amazon Bedrock, the plugin applies the following safeguards:

  • Credential redaction – The sanitize_script function removes sensitive patterns (passwords, tokens, access keys) from scripts and logs.
  • Content truncation – The truncate_script function caps content size to stay within model context windows.
  • Path traversal prevention – The read_allowlisted_file function resolves canonical paths and verifies they reside within allowed base directories before reading any file.

For the full implementation, see script_utils.py.

Optional: PII detection and redaction. The built-in sanitize_script function targets credential patterns. If your logs or scripts might contain personally identifiable information (PII), consider adding a detection pass with Amazon Comprehend before invoking Amazon Bedrock. The DetectPiiEntities API returns the entity types (such as names, email addresses, or account numbers) and their character offsets. You can use these offsets to mask or obfuscate the spans before the context leaves your environment. This adds one API call and cost per analysis, so add it where your compliance requirements call for it. For guidance, see Detecting PII entities.

Deploy the plugin

Follow these steps to deploy the plugin to your Amazon MWAA environment.

Step 1: Clone the repository

git clone https://github.com/aws-samples/sample-aws-mwaa-llm-powered-plugin.git
cd sample-aws-mwaa-llm-powered-plugin

Step 2: Package and upload to Amazon S3

Create the plugins.zip archive from the plugins/ directory and upload it to your Amazon MWAA S3 bucket:

cd plugins
zip -r ../plugins.zip .
cd ..

aws s3 cp plugins.zip s3://<amzn-s3-demo-bucket>/plugins.zip

aws s3api head-object \
  --bucket <amzn-s3-demo-bucket> \
  --key plugins.zip \
  --query VersionId --output text

Note the VersionId returned. You need it in the next step.

Note: This plugin requires only fastapi and Boto3, both pre-installed on Amazon MWAA for Airflow 3.x. You don’t need a requirements.txt file. Skipping the requirements file avoids package resolution conflicts that are a common cause of failed Amazon MWAA environment updates.

Step 3: Update the Amazon MWAA environment

Update your environment to use the new plugin archive:

aws mwaa update-environment \
  --name <your-environment-name> \
  --plugins-s3-path plugins.zip \
  --plugins-s3-object-version <version-id-from-step-2>

The environment restarts automatically. This process typically takes 10–30 minutes. Monitor the status with:

aws mwaa get-environment \
  --name <your-environment-name> \
  --query "Environment.{Status:Status,Plugins:PluginsS3Path}" --output json

Step 4: Configure the Amazon Bedrock connection

On Amazon MWAA, the aws_default connection exists by default and resolves to your environment’s execution role. In most cases, no action is needed.

To override the Region, edit the aws_default connection in the Airflow UI (Admin > Connections) and set the Extra field to:

{"region_name": "us-east-1"}

Leave login and password empty so the execution role is used.

Step 5: Verify the deployment

After the environment finishes updating, navigate to Admin > Plugins in the Airflow UI. Verify that task_analyzer_plugin appears in the list. The Analyze Task entry is now available from any task instance view.

Test the solution

The repository includes example DAGs that simulate failure scenarios across different operator types. To validate the deployment:

  1. Copy the dags/ directory contents to your Amazon MWAA S3 bucket’s DAGs folder:
    aws s3 cp dags/ s3://<amzn-s3-demo-bucket>/dags/ --recursive

  2. Wait for Amazon MWAA to sync the DAGs (typically 1–2 minutes).
  3. In the Airflow UI, trigger one of the test DAGs (for example, test_aws_sql_operators) and let the intentional failure occur.
  4. Navigate to the failed task instance.
  5. Choose Analyze Task in the task instance view.
  6. Review the generated analysis, which includes:
    • Root cause identification with file and line references.
    • Step-by-step resolution with code examples.
    • Prevention recommendations and monitoring suggestions.

The analysis typically completes within 5–10 seconds.

Cost considerations

The primary cost driver for this solution is Amazon Bedrock inference, which is billed by the number of input and output tokens each analysis consumes. Input tokens come from the task logs, DAG source, and operator script sent to the model. Output tokens come from the diagnostic report the model returns. Larger logs and scripts increase input tokens, and the model you select affects the per-token rate. For current per-model rates, see Amazon Bedrock pricing.

To help control cost, the plugin includes a caching mechanism that stores results keyed by a hash of the error context. Repeated analyses of the same failure pattern return cached results without invoking Amazon Bedrock again.

Best practices

When you deploy this solution in production, consider the following:

  • IAM least privilege – Grant only bedrock:InvokeModel for your chosen model IDs and scope s3:GetObject to specific bucket paths where your operator scripts reside. For guidance, see Amazon MWAA execution role.
  • Data sanitization – The plugin redacts credentials and truncates content before sending data to Amazon Bedrock. Store configuration values in AWS Secrets Manager rather than hardcoding them in DAG source files.
  • Access control – The plugin’s endpoints are protected by Airflow’s built-in authentication. For DAG-level access management at scale, see Automated tag-based DAG permission management in Amazon MWAA.
  • Operational resilience – Add retry logic and circuit breaker patterns around the Amazon Bedrock API call. Use Amazon CloudWatch to monitor plugin performance and set alarms on failure rates.

Extending the solution

You can extend this solution in the following ways:

  • Proactive notifications – Integrate with Amazon Simple Notification Service (Amazon SNS) or Slack to deliver analyses automatically when failures occur.
  • Knowledge base integration – Build a knowledge base of past analyses using Amazon Bedrock Knowledge Bases for Retrieval Augmented Generation (RAG) powered recommendations that learn from your organization’s historical failures.
  • Additional operator support – Add handlers for custom operators specific to your organization, such as proprietary data connectors or internal platform integrations.
  • Automated remediation – For well-understood failure patterns, trigger automated fixes such as restarting tasks with adjusted resource configurations.

Clean up

To remove the plugin from your environment:

  1. Delete the plugin archive from Amazon S3:
    aws s3 rm s3://<amzn-s3-demo-bucket>/plugins.zip

  2. Update your Amazon MWAA environment to remove the plugin reference, then wait for the environment to restart.
  3. Optionally, remove the Amazon Bedrock permissions from your execution role if they are no longer needed.

Conclusion

In this post, we showed you how to deploy an LLM-powered DAG failure analysis plugin for Amazon MWAA using Amazon Bedrock. The operator-aware context collection differentiates this approach from generic log analyzers. By fetching the actual code from AWS Glue, Amazon EMR, and other services, the foundation model provides precise, actionable recommendations with specific line references.

To get started, clone the sample-aws-mwaa-llm-powered-plugin repository, deploy it to a development Amazon MWAA environment, and test with the included example DAGs. As your team builds confidence in the analysis quality, roll it out to production environments where it serves as the first line of investigation for any pipeline failure.


About the authors

Sushant Samantaray

Sushant Samantaray

Sushant is a Sr. Delivery Consultant at AWS, bringing 19 years of industry experience with a focus on Data Analytics and Generative AI/Agentic AI solutions. He works closely with enterprise customers to design and deliver innovative solutions across Big Data, Generative AI, and Agentic AI, leveraging AWS native services, partner offerings, and open-source technologies. A passionate technologist and problem solver at heart, he balances his professional life with watching and playing sports and spending quality time with family.

Parameswara Reddy Gajjela

Parameswara Reddy Gajjela

Parameswara is a Delivery Consultant at AWS with 11+ years of experience in Data Analytics. He works closely with enterprise customers to architect innovative, end-to-end Big Data and Agentic AI solutions powered by AWS native services, partner ecosystems, and open-source technologies. His areas of expertise include modern DataLake and Data Warehouse migration and implementations on AWS.

Kamen Sharlandjiev

Kamen Sharlandjiev

Kamen is a Pr. Big Data and ETL Solutions Architect, MWAA and AWS Glue ETL expert. He’s on a mission to make life easier for customers who are facing complex data integration and orchestration challenges. His secret weapon? Fully managed AWS services that can get the job done with minimal effort. Follow Kamen on LinkedIn to keep up to date with the latest MWAA and AWS Glue features and news!

Announcing Spark Connect on Amazon EMR on EKS: Interactive PySpark development, anywhere

Post Syndicated from Amit Maindola original https://aws.amazon.com/blogs/big-data/announcing-spark-connect-on-amazon-emr-on-eks/

Today, we’re announcing support for Spark Connect on Amazon EMR on EKS, starting from EMR release 7.14 (Apache Spark 3.5.8) and emr-spark-8.1 (Apache Spark 4.1.1). You can now build, test, and debug Spark applications from your preferred tools, such as VS Code, PyCharm, Jupyter notebooks, Amazon SageMaker Unified Studio. At the same time, your full-scale Spark operations run on Amazon Elastic Kubernetes Service (Amazon EKS).

Deploying Spark applications from a local development environment to a remote Amazon EKS cluster often means dealing with environment differences, dependency conflicts, and performance gaps at scale. Spark Connect removes this friction. It separates your application client from the Spark server, so you develop and debug locally while Spark Connect routes your operations to a scalable Spark cluster running on Amazon EKS.

This client-server architecture supports a range of use cases, including interactive development from notebooks and IDEs, embedded Spark in web services, and continuous integration and continuous delivery (CI/CD) data-quality tests. All of these run on your existing EKS infrastructure. Each Spark Connect session uses its own AWS Identity and Access Management (IAM) execution role, custom tags, and cost tracking. For more information, see the Amazon EMR on EKS documentation.

Here are two demonstrations of using Spark Connect in Amazon SageMaker Unified Studio Notebooks and in a VS Code local IDE:

Amazon SageMaker Unified Studio Notebooks demo:

Local IDE demo:

For a runnable end-to-end example in an IDE, try the Spark Connect sample notebook in the aws-emr-utilities repository. It includes a client wrapper solution, built by AWS architects, for simplified connectivity:

How Spark Connect works on Amazon EMR on EKS

Spark Connect uses a client-server architecture that separates application code from the Spark engine:

  1. Client – A lightweight PySpark library running in your environment (such as an IDE or notebook). It doesn’t need Spark installed, direct access to data, or resources sized for the workload.
  2. Connection (EMR managed endpoint) – The client sends Spark operations over a secure gRPC/TLS channel to the Spark Connect server.
  3. Server – Runs Spark pods in your Amazon EMR on EKS namespace, starting from a minimum of two executors (adjustable) with autoscaling. The server performs Spark operations using the EKS compute resources and accesses data stores, such as an Amazon Simple Storage Service (Amazon S3) bucket, through job execution roles.
  4. Results – The server streams query results back to the client through gRPC as Apache Arrow-encoded row batches.
Client-server data flow from a local PySpark client through a gRPC channel to Spark pods on Amazon EMR on EKS

Figure 1: Spark Connect’s client-server architecture

On endpoint creation, Amazon EMR on EKS launches the Spark Connect server as pods on EKS and returns an Elastic Load Balancing (ELB)-backed endpoint and a short-lived token. You don’t need to provision any server or networking manually. Because the Spark Connect server runs on the EKS cluster you already operate, it inherits the node types, container images, and Spark configurations. What you see while developing Spark applications on the client side is what runs in the EKS environment at scale.

To provide a secure, simplified experience, Amazon EMR on EKS provisions two additional components on first use of Spark Connect on the EKS cluster:

Shared Envoy authentication-proxy router and Secret Agent service on the EKS cluster

Figure 2: Shared Envoy router and Secret Agent service on the EKS cluster

  • Managed authentication-proxy router – a shared Envoy router with three replicas by default (adjustable), fronted by a Network Load Balancer (NLB). It routes client traffic to the correct server pods, terminates TLS, and validates the session token. One router serves Spark Connect endpoints on the EKS cluster.
  • Secret Agent service – a lightweight, long-running pod that manages the short-lived credentials for session authentication. One service per EMR security configuration.

These components are long-running and shared across endpoints. Amazon EMR on EKS creates them automatically with the first endpoint on the cluster. Because the router is cluster-scoped and Secret Agent is namespace-scoped, deleting a managed endpoint doesn’t remove them. They keep running so that new endpoints can start within a minute. The router’s replica count is tunable. Scale down for non-production environments to reduce cost or scale up for higher throughput.

To fully remove these components:

  • Terminate all active managed endpoints and their virtual cluster that reference the Secret Agent’s security configuration, then delete the security configuration.
  • Once the last session-enabled virtual cluster is deleted, the authentication-proxy router and its underly resources, including the NLB and VPC endpoint, are removed automatically.
  • Alternatively, delete the EKS cluster to remove all in-cluster components at once.

Why use Spark Connect on Amazon EMR on EKS

With Amazon EMR on EKS, teams can run Spark alongside other applications on shared Kubernetes clusters with existing infrastructure, operational tooling, and system expertise. Spark Connect extends that value to interactive, embedded, and self-service Spark workloads. Your client stays lightweight while Spark code runs in governed, scalable server pods on EKS.

Interactive development on shared Kubernetes clusters

Data engineers and scientists iterate on Spark code cell-by-cell in notebooks or local IDEs. The Spark engine runs remotely on EKS, so validation runs on the same engine as your batch workloads. After validation on the Spark Connect client, the same Spark code deploys as a batch StartJobRun with no changes.

Spark Connect sessions run as pods on your existing cluster. They reuse your EKS RBAC, network policies, node autoscaling, and observability stack (Prometheus, Grafana, Amazon CloudWatch Container Insights). There are no separate compute and monitoring layers to operate.

Embedded Spark in applications and services

The Spark Connect client is a compact PySpark library. Teams can embed Spark operations directly into Python applications such as web services, dashboards, automation scripts, or backend APIs. The heavy processing runs on EKS while the application stays lightweight.

Teams can also expose Spark Connect as a self-service capability on their internal application. Business users submit Spark SQL scripts from a web UI. The compute runs on Spark Connect server on EKS, so the team manages capacity, security, and upgrades centrally.

Multi-tenant data exploration with governance

Each Spark Connect session uses the data user’s IAM permissions that you configure, limiting their access to authorized AWS services, data lake tables, and S3 paths. Every session carries tags with user, project, endpoint and virtual cluster IDs, feeding directly into billing and compliance reports. Meanwhile, data producers maintain guardrails on source data without blocking self-service exploration.

To manage resource consumption across teams, Amazon EMR on EKS virtual clusters provide namespace-level isolation. Each tenant binds their Spark Connect endpoints to a virtual cluster (a namespace) with independent IAM roles. Using resource quotas and limit ranges on EKS, you can protect each virtual cluster by controlling the compute resources that Spark Connect sessions can consume. Importantly, activating EKS split-cost allocation tags helps with chargeback reporting in a multi-tenant environment.

Reusable container images and scalable deployment

Teams often maintain custom container images with proprietary libraries, including internal feature stores, compliance toolkits, UDFs, or machine learning (ML) frameworks. With Spark Connect on Amazon EMR on EKS, teams reuse those same images as the Spark runtime for interactive sessions. No separate dependency lists needed. The same image works for both batch jobs and Spark Connect sessions.

Beyond the image itself, you can control Spark pod scheduling in Amazon EMR on EKS through pod templates and managed endpoint APIs, scaling across your environment. For example, you can:

  • Pin server pods to specific node types through pod templates. For example, Spot for cost savings.
  • Apply Spark Dynamic Resource allocation (DRA) to right-size each interactive session.
  • Use GPU node pools for accelerated Spark RAPIDS or ML.
cat > /tmp/spark-connect-endpoint.json << EOF
{
  "name": "spark-connect-custom-config",
  "virtualClusterId": "$VC_ID",
  "type": "SPARK_CONNECT",
  "releaseLabel": "emr-7.14.0-latest",
  "executionRoleArn": "$ROLE_ARN",
  "configurationOverrides": {
    "applicationConfiguration": [{
      "classification": "spark-defaults",
      "properties": {
        "spark.kubernetes.container.image": "${CUSTOM_IMAGE_URI}",
        "spark.kubernetes.executor.podTemplateFile": "s3://$S3BUCKET/exec-pod-template.yaml",
        "spark.kubernetes.node.selector.karpenter.sh/nodepool": "gpu-pool",
        "spark.dynamicAllocation.enabled": "true",
        "spark.dynamicAllocation.minExecutors": "0"
      }
    }]
  }
}
EOF

aws emr-containers create-managed-endpoint \
--cli-input-json file:///tmp/spark-connect-endpoint.json

Multi-cluster, multi-Region, and hybrid architectures

Enterprises running EKS clusters across multiple AWS accounts, AWS Regions, or hybrid environments with on-premises Kubernetes can use Spark Connect to query data wherever it’s processed. The lightweight client only needs to reach the Spark Connect endpoint, not the underlying S3 buckets or AWS Glue data catalogs. This means no VPC peering or direct network paths to every data store.

The client-server split is the core architectural advantage of Spark Connect on Amazon EMR on EKS. A developer on a laptop behind a VPN, a CI/CD deployment pipeline in a centralized service account, or an Airflow DAG orchestrating across Regions can all connect to a remote Spark server on EKS. This works regardless of where the client itself runs. This decoupling simplifies cross-Region or cross-account analytics without duplicating data or requiring direct access to each data store.

Getting started

To create a Spark Connect endpoint on Amazon EMR on EKS, complete the following steps:

  1. Create EMR namespaces on EKS.
  2. Create an EMR security configuration.
  3. Create a virtual cluster with the security configuration.
  4. Create a Spark Connect managed endpoint.
  5. Obtain a session token.
  6. Connect from your application.

Prerequisites

To proceed with this post, make sure you have the following:

Step 1: Create EMR namespaces

# set environment variables
export EKS_CLUSTER_NAME=my-eks-cluster
export USER_NAMESPACE=spark-connect-1
export SYS_NAMESPACE=spark-connect-1-system
export AWS_REGION=us-west-2
# connect to your EKS cluster
aws eks update-kubeconfig --name $EKS_CLUSTER_NAME --region $AWS_REGION
kubectl create namespace $USER_NAMESPACE
kubectl create namespace $SYS_NAMESPACE

Step 2: Create a security configuration

cat > /tmp/sec-config.json << EOF
{
  "name": "spark-connect-1-sc",
  "securityConfigurationData": {
    "authenticationConfiguration": {
      "identityCenterConfiguration": { "enableIdentityCenter": false }
    }
  },
  "containerProvider": {
    "type": "EKS",
    "id": "$EKS_CLUSTER_NAME",
    "info": { "eksInfo": { "namespace": "$SYS_NAMESPACE" } }
  }
}
EOF

SEC_CONFIG_ID=$(aws emr-containers create-security-configuration \
--region $AWS_REGION \
--cli-input-json file:///tmp/sec-config.json \
--query id \
--output text)
echo "Security Configuration ID: $SEC_CONFIG_ID"

Step 3: Create a virtual cluster with the security configuration

cat > /tmp/vc.json << EOF
{
  "name": "spark-connect-demo",
  "containerProvider": {
    "id": "$EKS_CLUSTER_NAME",
    "type": "EKS",
    "info": {"eksInfo": {"namespace": "$USER_NAMESPACE"}}
  },
  "securityConfigurationId": "$SEC_CONFIG_ID",
  "sessionEnabled": true
}
EOF

VC_ID=$(aws emr-containers create-virtual-cluster \
--region $AWS_REGION \
--cli-input-json file:///tmp/vc.json \
--query 'id' \
--output text)
# validate the virtual cluster
echo "Virtual Cluster ID: $VC_ID"
aws emr-containers describe-virtual-cluster --region $AWS_REGION --id $VC_ID

Step 4: Create a Spark Connect managed endpoint

Start an interactive session on your virtual cluster. Provide a job execution role that grants the session access to your data sources.

# reuse an existing execution role
ROLE_ARN="arn:aws:iam::YOUR_ACCOUNT_ID:role/EMRonEKSExecutionRole"
cat > /tmp/spark-connect-endpoint.json << EOF
{
  "name": "spark-connect-demo",
  "virtualClusterId": "$VC_ID",
  "type": "SPARK_CONNECT",
  "releaseLabel": "emr-7.14.0-latest",
  "executionRoleArn": "$ROLE_ARN",
  "sessionIdleTimeoutInMinutes": 1440,
  "configurationOverrides": {
    "applicationConfiguration": [{
      "classification": "spark-defaults",
      "properties": {
        "spark.dynamicAllocation.enabled": "true",
        "spark.dynamicAllocation.minExecutors": "0",
        "spark.dynamicAllocation.maxExecutors": "2"
      }
    }]
  }
}
EOF

EP_ID=$(aws emr-containers create-managed-endpoint \
--region $AWS_REGION \
--cli-input-json file:///tmp/spark-connect-endpoint.json \
--query 'id' \
--output text)
# Validate
echo "Endpoint ID: $EP_ID"
export EP_URL=$(aws emr-containers describe-managed-endpoint \
--region $AWS_REGION \
--virtual-cluster-id $VC_ID \
--id $EP_ID \
--query 'endpoint.authProxyUrl' \
--output text)
echo "Endpoint URL: $EP_URL"
Managed endpoint creation output showing the endpoint ID and endpoint URL

Figure 3: Managed endpoint creation output with the endpoint ID and URL

You can optionally pass some custom configuration overrides and tags:

aws emr-containers create-managed-endpoint \
--type SPARK_CONNECT \
--virtual-cluster-id $VC_ID \
--name more-endpoint \
--execution-role-arn $ROLE_ARN \
--release-label emr-7.14.0-latest \
--configuration-overrides '{
"applicationConfiguration": [{
"classification": "spark-defaults",
"properties": {
"spark.executor.instances": "1",
"spark.executor.memory": "4g",
"spark.executor.cores": "1",
"spark.sql.extensions": "org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions"
}
}]
}' \
--tags '{
"team": "data-engineering",
"project": "customer-analytics"
}'

Step 5: Obtain a session token

Request a session token after the managed endpoint is active:

# get a session token with a 12-hour expiry (adjustable)
export TOKEN=$(aws emr-containers get-managed-endpoint-session-credentials \
--region $AWS_REGION \
--virtual-cluster-identifier $VC_ID \
--endpoint-identifier $EP_ID \
--execution-role-arn $ROLE_ARN \
--credential-type TOKEN \
--duration-in-seconds 43200 \
--query 'credentials.token' \
--output text)
echo "Session Token: $TOKEN"

Security note: Communication between your environment and the Spark Connect server is encrypted using TLS. The authentication token is time-limited (15 minutes by default). For long-running sessions, refresh the token periodically by calling get-managed-endpoint-session-credentials again. Consider using AWS Secrets Manager to store and retrieve tokens programmatically.

Step 6: Connect from your application

Use the returned endpoint URL and token to connect from a PySpark-compatible environment. The following Python code shows how to establish a Spark Connect session:

import os
from pyspark.sql import SparkSession
session_endpoint = os.environ["EP_URL"]
auth_token = os.environ["TOKEN"]
spark_conn_url = (f"{session_endpoint};use_ssl=true;x-aws-proxy-auth={auth_token}")
spark = SparkSession.builder
.remote(spark_conn_url)
.getOrCreate()
# verify the connection
print(f"Connected remotely! Spark version: {spark.version}")
# query data through the AWS Glue Data Catalog
df = spark.sql("SELECT * FROM my_catalog.my_database.my_table LIMIT 10")
df.show()

After you’re connected, you can:

  • Debug interactively – Set breakpoints, inspect DataFrames, and step through Spark code in your IDE or notebook while the operations run remotely on EKS.
  • Combine local and remote processing – Pull query results back to the client as a pandas or PyArrow DataFrame for local analysis, visualization, or ML (scikit-learn, notebook widgets), then push further Spark operations back to the server in the same session. Heavy processing stays on Amazon EMR on EKS. Only the results you request cross the wire.
  • Reconnect without losing state – A managed endpoint runs independently of single clients for a configurable idle timeout (default: 60 minutes). Your Spark session, cached data, and temporary views are preserved on the server between connections. When a session token expires (default: 15 minutes, configurable up to 12 hours), request a new token and reconnect to the same endpoint to resume where you left off.
  • Reuse across workload types – The same client connection pattern works everywhere Python runs: notebooks, IDEs, batch scripts, Airflow operators, or web services. One endpoint, one connection pattern, many workload types.

Validation

After you create the endpoint, verify that the Spark Connect server is running and reachable through Amazon EMR on EKS API and standard Kubernetes tooling:

# get endpoint status
aws emr-containers describe-managed-endpoint --virtual-cluster-id $VC_ID --id $EP_ID
# inspect the server pods (driver + executors) in your namespace
kubectl get pods -n $USER_NAMESPACE -l "emr-containers.amazonaws.com/managed-endpoint-id=$EP_ID"
# View driver logs
kubectl logs -n $USER_NAMESPACE <driver-pod-name> -c spark-kubernetes-driver
Terminal output showing endpoint status and the running driver and executor pods

Figure 4: Endpoint status and the running driver and executor pods

# to view the live Spark UI, port-forward your driver pod:
DRIVER_POD=$(kubectl get pods -n $USER_NAMESPACE \
-l "emr-containers.amazonaws.com/managed-endpoint-id=$EP_ID,emr-containers.amazonaws.com/component=driver" \
-o name)
kubectl port-forward -n $USER_NAMESPACE "$DRIVER_POD" 4040:4040
# Open http://localhost:4040 in your browser
Live Spark UI for the Spark Connect session viewed in a browser through port forwarding

Figure 5: Live Spark UI for the Spark Connect session

Spark Connect endpoints run as pods on your EKS cluster. The existing Kubernetes observability stack, such as CloudWatch Container Insights, Prometheus, and Grafana, captures Spark Connect endpoint metrics alongside other cluster workloads.

Clean up resources

Terminate your session when you’re done to avoid ongoing costs:

# (OPTIONAL) Endpoints are auto-deleted after the idle timeout (default: 60 minutes).
aws emr-containers delete-managed-endpoint \
--virtual-cluster-id $VC_ID \
--id $EP_ID
# Delete the virtual cluster only when no active endpoints remain
aws emr-containers delete-virtual-cluster --id $VC_ID
# Delete Security Configuration
aws emr-containers delete-security-configuration --id $SEC_CONFIG_ID
# remove the remaining EKS namespaces
kubectl delete namespace $USER_NAMESPACE $SYS_NAMESPACE spark-connect-router

Deleting or timing out a managed endpoint automatically removes its corresponding driver and executor pods. The Envoy router and Secret Agent service are shared across endpoints on the EKS cluster and remain running when individual endpoints are terminated. To fully remove these shared components, delete the virtual cluster to remove its corresponding Secret Agent service. Before doing so, ensure that no managed endpoints in the virtual cluster are active. Terminating the last session-enabled virtual cluster automatically removes the Envoy router from the EKS cluster.

Availability and pricing

Spark Connect on Amazon EMR on EKS is available with EMR release 7.14 (Apache Spark 3.5) and emr-spark-8.1 (Apache Spark 4.1), in all AWS Regions where Amazon EMR on EKS is available, except the AWS GovCloud (US) Regions and the China Regions. The Amazon SageMaker Unified Studio experience is available in supported Regions.

There is no additional charge for Spark Connect managed endpoints beyond the standard Amazon EMR on EKS pricing. You pay for underlying Amazon EKS resources such as EC2 and ELB. For timed-out or terminated managed endpoints, EMR automatically removes their Spark pods from the EKS cluster.

Recommendations for cost efficiency:

  • Use Karpenter (or Cluster Autoscaler) to right-size cluster capacity to session workload demand. This provisions nodes when endpoints need them and removes them when idle, which keeps cost aligned to actual usage.
  • Schedule interactive session pods on On-Demand instances for persistent compute.
  • Use AWS Graviton processors for better performance on Spark workloads.
  • Activate Amazon EMR on EKS Cost Allocation tags to track per-team and per-project spending at granular level.
  • Keep a single, shared Envoy router and NLB serving all Spark Connect endpoints (the default) on the cluster. Right-size the router replica count (three by default) for your availability requirements.

Considerations and limitations

Before you build on Spark Connect for Amazon EMR on EKS, review the Considerations and limitations in the Amazon EMR on EKS documentation.

Conclusion

In this post, we showed how, with Spark Connect on Amazon EMR on EKS, you can build, test, and debug Spark applications from the tools you already use: IDEs, notebooks, Amazon SageMaker Unified Studio or Airflow. Your workloads run at scale on your existing Kubernetes clusters, with no application code changes.

For teams already running Amazon EMR on EKS, Spark Connect extends your virtual clusters to interactive and embedded workloads. The same virtual cluster that runs your batch StartJobRun jobs now also serves Spark Connect sessions. Each session runs as pods on your EKS cluster, inheriting your node groups, container images, and Spark configurations. Each session also carries its own IAM execution role and cost tags. This extends the security, multi-tenancy, and observability of your Amazon EMR on EKS investment to a broader set of users and use cases.

To get started, visit the Spark Connect on Amazon EMR on EKS documentation, try the Amazon SageMaker Unified Studio Getting Started guide, and review the Amazon EMR on EKS release notes for EMR 7.14.


About the authors

Amit Maindola

Amit Maindola

Amit is a Senior Data & AI Architect with AWS ProServe team focused on data engineering, analytics, and AI/ML at Amazon Web Services. He helps customers in their digital transformation journey and enables them to build highly scalable, robust, and secure cloud-based analytical solutions on AWS to gain timely insights and make critical business decisions.

Melody Yang

Melody Yang

Melody Yang is a Principal Analytics Specialist Solution Architect at AWS with expertise in Big Data technologies. She is an experienced analytics leader working with AWS customers to provide best practice guidance and technical advice in order to assist their success in data transformation. Her areas of interests are open-source frameworks and automation, data engineering and DataOps.

Al MS

Al MS

Al is a product manager for Amazon EMR at AWS.

Aurora PostgreSQL zero-ETL integration with Amazon SageMaker

Post Syndicated from Apurwa Pawar original https://aws.amazon.com/blogs/big-data/aurora-postgresql-zero-etl-integration-with-amazon-sagemaker/

When you need quick insights from your Amazon Aurora PostgreSQL operational data, traditional analytics approaches force you to build complex extract, transform, and load (ETL) pipelines. These pipelines introduce latency, operational overhead, and data silos, which slow down decision making and increase cost. AWS introduced the support for Amazon Aurora PostgreSQL zero-ETL integration with Amazon SageMaker, providing near real-time data availability for analytics workloads.

The zero-ETL integration automatically replicates the data from your Amazon Aurora PostgreSQL database into a target AWS Glue managed catalog, where it’s available as Apache Iceberg tables. You can then analyze this data through Amazon SageMaker alongside data from other sources using your preferred analytics and machine learning (ML) tools. The data is compatible with Apache Iceberg open standards, so you can use SQL, Apache Spark, business intelligence, and artificial intelligence and machine learning (AI/ML) tools.

In this post, you explore the benefits of this integration, the architectural concepts, and the underlying change data capture (CDC) mechanics. You also go through the setup process and learn how to query your Aurora PostgreSQL data in Amazon SageMaker AI.

Zero-ETL in the lakehouse architecture

The lakehouse architecture of Amazon SageMaker AI brings together data across Amazon Simple Storage Service (Amazon S3) data lakes and Amazon Redshift data warehouses. Because it’s built on open standards, you can build analytics and AI/ML applications on a single copy of data, without moving it between systems.

Amazon SageMaker AI uses AWS Glue Data Catalog and AWS Lake Formation to provide integrated access controls across S3 data lakes and Amazon Redshift data warehouses from a single governance plane.

Understanding change data capture mechanics

At its core, Aurora PostgreSQL zero-ETL integration is powered by CDC. CDC continuously monitors the database transaction log and streams every insert, update, and delete to a downstream target in near real time.

Aurora PostgreSQL uses enhanced logical replication as its CDC engine. Standard PostgreSQL logical replication publishes row-level changes from the write-ahead log (WAL). The enhanced logical replication in Aurora offers added capabilities that make it well-suited for zero-ETL integrations, including automatic DDL propagation and continuous streaming of transactional changes.

Solution overview

With Amazon Aurora PostgreSQL zero-ETL integration with Amazon SageMaker AI, you can:

  • Remove ETL complexity – Automatically replicate data without building custom ETL pipelines.
  • Near real-time analytics – Access operational data in Amazon SageMaker AI within seconds of changes in Aurora PostgreSQL.
  • Unify data analysis – Combine Aurora PostgreSQL data with data from other sources in a single lakehouse architecture.
  • Reduce costs – Minimize operational overhead and infrastructure costs associated with maintaining ETL pipelines.
  • Accelerate insights – Query data using familiar SQL tools and integrate with ML workflows in Amazon SageMaker AI.

The following diagram illustrates the architecture of this solution:

Aurora PostgreSQL zero-ETL integration replicating data into an AWS Glue managed catalog queried through Amazon SageMaker

Figure 1: Architecture of the Aurora PostgreSQL zero-ETL integration with Amazon SageMaker

The workflow includes the following steps:

  1. Your application writes data to an Amazon Aurora PostgreSQL database cluster.
  2. The zero-ETL integration automatically captures changes from the Aurora PostgreSQL database.
  3. Data is replicated to the target AWS Glue managed catalog in near real time.
  4. You can query and analyze the data using Amazon Athena, Amazon Redshift, or other analytics tools integrated with Amazon SageMaker AI.
  5. Data scientists can build and train ML models using Amazon SageMaker AI with direct access to the Apache Iceberg tables in the target AWS Glue managed catalog.

Prerequisites

Before setting up the zero-ETL integration, verify that you have the following:

Configure the source PostgreSQL database for zero-ETL integration

When you have all the prerequisites in place, you can configure the source PostgreSQL database for zero-ETL integration.

Create a custom Aurora PostgreSQL cluster parameter group

Your Aurora PostgreSQL database needs to have parameters configured for real-time replication. In this section, you will create the DB cluster parameter group and configure parameters. For more information, see Getting started with Aurora zero-ETL integrations.

Use the following AWS CLI command to create an Aurora PostgreSQL cluster parameter group:

aws rds create-db-cluster-parameter-group \
    --db-cluster-parameter-group-name aurora-pgsql-zetl-cluster-pg \
    --db-parameter-group-family aurora-postgresql16 \
    --description "Aurora PostgreSQL with enhanced logical replication" \
    --region us-east-1 --output json

Now set the parameters by modifying the parameter group:

aws rds modify-db-cluster-parameter-group --db-cluster-parameter-group-name <aurora-pgsql-zetl-cluster-pg> \
    --parameters \
    ParameterName=rds.logical_replication,ParameterValue=1,ApplyMethod=pending-reboot \
    ParameterName=aurora.enhanced_logical_replication,ParameterValue=1,ApplyMethod=pending-reboot \
    ParameterName=aurora.logical_replication_backup,ParameterValue=0,ApplyMethod=pending-reboot \
    ParameterName=aurora.logical_replication_globaldb,ParameterValue=0,ApplyMethod=pending-reboot \
    --region us-east-1 --output json

The parameter group is now fully configured and ready to be applied to your Aurora PostgreSQL cluster.

Select or create a source Aurora PostgreSQL cluster

If you already have an Aurora PostgreSQL cluster, you can use it, or you can create a new Aurora PostgreSQL cluster.

Note: Your source DB cluster must be running a supported version of Aurora PostgreSQL. For a list of supported versions, see Regions and database engines supported for Aurora zero-ETL integrations.

While creating an Aurora PostgreSQL cluster, use the parameter group (aurora-pgsql-zetl-cluster-pg) you created earlier:

Note: Throughout this post, make sure to replace the with your own information.

aws rds create-db-cluster \
    --db-cluster-identifier <aurora-pgsql-zetl> \
    --engine aurora-postgresql \
    --engine-version 16 \
    --master-username <admin> \
    --master-user-password <password> \
    --database-name <my_db> \
    --db-cluster-parameter-group-name <aurora-pgsql-zetl-cluster-pg> \
    --storage-encrypted \
    --kms-key-id alias/aws/rds \
    --backup-retention-period 7 \
    --db-subnet-group-name <dbsubnet> \
    --vpc-security-group-ids <sg-c14219ba> \
    --region <us-east-1> \
    --output json
aws rds create-db-instance \
    --db-instance-identifier <aurora-pgsql-zetl-instance-1> \
    --db-instance-class db.r5.large \
    --engine aurora-postgresql \
    --db-cluster-identifier <aurora-pgsql-zetl> \
    --region us-east-1 \
    --output json

If you’re creating a new Aurora PostgreSQL cluster, wait for your DB instance(s) to be in an “Available” status. You can verify DB instance status by using the describe-db-instances API call:

aws rds describe-db-instances --filters 'Name=db-cluster-id,Values=<aurora-pgsql-zetl>' --output json | grep -o '"DBInstanceStatus": "[^"]*"'
"DBInstanceStatus": "available"

Reboot the cluster to apply parameter changes

A cluster reboot is needed before zero-ETL integration can function correctly:

aws rds reboot-db-instance \
    --db-instance-identifier <aurora-pgsql-zetl-instance-1> \
    --region <us-east-1>

Wait until the cluster and the primary instance are back in Available status. For more information, see reboot-db-instance.

Create a target AWS Glue managed catalog

With your source PostgreSQL database configured for enhanced logical replication, the next step is setting up your target Amazon SageMaker AI. Zero-ETL integration uses AWS Glue Data Catalog backed by Amazon Redshift managed storage as its target. To have this functionality, you need to create a managed catalog, configure IAM permissions for Amazon SageMaker AI to access and query the managed catalog, and set up authorization for incoming integration requests from your source database.

Create an AWS Glue managed catalog

You must create a new catalog (if it doesn’t exist already) managed by AWS Glue to store table metadata and serve as the landing zone for your replicated datasets. Zero-ETL integration streams the data into Amazon Redshift managed storage, and AWS Glue keeps track of table definitions so that tools such as SageMaker AI, Athena, and Amazon Redshift Spectrum can query the data.

Create an IAM role for AWS Glue and Amazon Redshift to access the AWS Glue managed catalog

Now, use the following command to create an IAM role so that AWS Glue and Amazon Redshift can interact with the catalog. This role serves two key functions: It allows AWS Glue and Amazon Redshift to perform catalog operations, and it authorizes incoming integration requests from your source database.

aws iam create-role \
    --role-name <GlueDataCatalogDataTransferRole> \
    --assume-role-policy-document '{
        "Version": "2012-10-17",
        "Statement": [
            {
                "Effect": "Allow",
                "Principal": {
                    "Service": [
                        "glue.amazonaws.com",
                        "redshift.amazonaws.com"
                    ]
                },
                "Action": "sts:AssumeRole"
            }
        ]
    }'

Next, attach a policy to this IAM role that provides the minimum required permissions for AWS Glue and Amazon Redshift. This policy should also include the necessary permissions for encryption key actions to help maintain secure data handling throughout the integration process:

aws iam put-role-policy \
    --role-name <GlueDataCatalogDataTransferRole> \
    --policy-name <GlueDataTransferPolicy> \
    --policy-document '{
        "Version": "2012-10-17",
        "Statement": [
            {
                "Sid": "DataTransferRolePolicy",
                "Effect": "Allow",
                "Action": [
                    "kms:GenerateDataKey",
                    "kms:Decrypt",
                    "glue:GetDatabase",
                    "glue:GetCatalog"
                ],
                "Resource": ["*"]
            }
        ]
    }'

Set up AWS Lake Formation access

Before using the managed catalog for zero-ETL integration, you must configure data lake administrators in AWS Lake Formation who have administrative or read-only permissions on the managed resources. Additionally, you need to grant ReadOnlyAdmin permissions to the Amazon Redshift service-linked role, AWSServiceRoleForRedshift, in your account. If this role doesn’t exist in your account or you need to verify its permissions, see Using service-linked roles for Amazon Redshift.

aws lakeformation put-data-lake-settings \
    --region <us-east-1> \
    --cli-input-json '{
        "DataLakeSettings": {
            "DataLakeAdmins": [
                {
                    "DataLakePrincipalIdentifier": "<arn:aws:iam::111122223333:role/Admin>"
                }
            ],
            "ReadOnlyAdmins": [
                {
                    "DataLakePrincipalIdentifier": "<arn:aws:iam::111122223333:role/aws-service-role/redshift.amazonaws.com/AWSServiceRoleForRedshift>"
                }
            ],
            "CreateDatabaseDefaultPermissions": [],
            "CreateTableDefaultPermissions": [],
            "Parameters": {
                "CROSS_ACCOUNT_VERSION": "4",
                "SET_CONTEXT": "TRUE"
            },
            "AllowExternalDataFiltering": false,
            "ExternalDataFilteringAllowList": []
        }
    }'

Create the AWS Glue managed catalog backed by Amazon Redshift managed storage

Because you have configured IAM permissions and Lake Formation settings, you can now create the AWS Glue managed catalog.

aws glue create-catalog \
    --region <us-east-1> \
    --cli-input-json '{
        "Name": "<zetl-catalog>",
        "CatalogInput": {
            "Description": "A Glue Data Catalog backed by Redshift Managed Storage",
            "CreateDatabaseDefaultPermissions": [],
            "CreateTableDefaultPermissions": [],
            "CatalogProperties": {
                "DataLakeAccessProperties": {
                    "DataLakeAccess": true,
                    "DataTransferRole": "<arn:aws:iam::111122223333:role/GlueDataCatalogDataTransferRole>",
                    "CatalogType": "aws:redshift"
                }
            }
        }
    }'

Register the catalog as a zero-ETL integration target

To prepare your target AWS Glue managed catalog for zero-ETL integration, use the create-integration-resource-property command with these required parameters:

  • The –resource-arn parameter specifies the Amazon Resource Name (ARN) of your AWS Glue managed catalog that will serve as the integration target.
  • The –target-processing-properties parameter requires the ARN of an IAM role that has describe permissions on the target AWS Glue managed catalog.

You can use the GlueDataCatalogDataTransferRole created in the earlier step because it already includes the minimal describe permissions needed for this integration. Alternatively, you can create a new IAM role specifically for this purpose and attach the necessary minimal permissions to meet your company’s security requirements.

aws glue create-integration-resource-property \
    --region <us-east-1> \
    <arn:aws:glue:us-east-1:111122223333:catalog/zetl-catalog> \
    '{"RoleArn": "<arn:aws:iam::111122223333:role/GlueDataCatalogDataTransferRole>"}'

Example output:

{
    "ResourceArn": "arn:aws:glue:us-east-1:111122223333:catalog/zetl-catalog",
    "TargetProcessingProperties": {
        "RoleArn": "arn:aws:iam::111122223333:role/GlueDataCatalogDataTransferRole"
    }
}

Configure authorization for inbound integration requests

The last step in creating a target managed catalog is to define a resource-based access policy that authorizes zero-ETL integration to push data into your catalog. This policy grants AWS Glue the necessary permissions to create and authorize incoming integration requests from your source database. Apply this resource policy by using the AWS Glue put-resource-policy API call to complete the catalog configuration for your zero-ETL integration:

aws glue put-resource-policy \
    --region <us-east-1> \
    --policy-in-json '{
        "Version": "2012-10-17",
        "Statement": [
            {
                "Principal": {
                    "AWS": [
                        "111122223333"
                    ]
                },
                "Effect": "Allow",
                "Action": [
                    "glue:CreateInboundIntegration"
                ],
                "Resource": [
                    "arn:aws:glue:us-east-1:111122223333:catalog/zetl-catalog"
                ],
                "Condition": {
                    "StringEquals": {
                        "aws:SourceArn": "arn:aws:rds:us-east-1:111122223333:cluster:aurora-pgsql-zetl"
                    }
                }
            },
            {
                "Principal": {
                    "Service": [
                        "glue.amazonaws.com"
                    ]
                },
                "Effect": "Allow",
                "Action": [
                    "glue:AuthorizeInboundIntegration"
                ],
                "Resource": [
                    "arn:aws:glue:us-east-1:111122223333:catalog/zetl-catalog"
                ],
                "Condition": {
                    "StringEquals": {
                        "aws:SourceArn": "arn:aws:rds:us-east-1:111122223333:cluster:aurora-pgsql-zetl"
                    }
                }
            }
        ]
    }'

Your AWS Glue managed catalog is now ready to receive data from the zero-ETL integration.

Load data in the source Aurora PostgreSQL database

Now that your Aurora PostgreSQL database is configured and ready, you must populate it with sample data that serves as the historical baseline for your zero-ETL integration. This first dataset provides the foundation for testing and demonstrating the integration capabilities. After you set up the zero-ETL integration, subsequent database changes stream automatically in near real time to your target AWS Glue managed catalog.

Connect to the source Aurora PostgreSQL cluster

Use the following commands to create a connection to your source Aurora PostgreSQL cluster:

psql --host aurora-pgsql-zetl-xxxxx.us-east-1.rds.amazonaws.com --username admin --port 5432 --dbname my_db --password

Create a database and table

Create a table named products to store product information:

CREATE TABLE products (product_id SERIAL PRIMARY KEY,product_name VARCHAR(100) NOT NULL, description TEXT,category VARCHAR(50),price NUMERIC(10,2) NOT NULL,stock_quantity INTEGER DEFAULT 0,is_active BOOLEAN DEFAULT TRUE,created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,updated_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP );

Insert historical data

Use the following code to insert a row:

INSERT INTO products (product_name, description, category, price, stock_quantity) VALUES ('Laptop', 'High-performance laptop with 16GB RAM and 512GB SSD', 'Electronics', 1299.99, 50);

This table serves as a representative dataset to demonstrate the data capture and streaming capabilities of the zero-ETL integration. After your zero-ETL integration is active, all database changes, including inserts, updates, and deletes, are automatically captured and streamed to your AWS Glue managed catalog. This creates a data pipeline from your Aurora PostgreSQL database to your Amazon SageMaker for real-time analytics on your operational data.

Create a zero-ETL integration

Because your Aurora PostgreSQL database is now populated with historical data, you can set up the zero-ETL integration that continuously streams database changes to your AWS Glue managed catalog backed by Amazon Redshift managed storage.

Create the integration

Create the integration between your source PostgreSQL database and target AWS Glue catalog by using the aws rds create-integration AWS CLI command. You can customize the integration by specifying added configurations, such as data filters, to control which data gets replicated to your target environment:

aws rds create-integration \
    --source-arn <arn:aws:rds:us-east-1:111122223333:cluster:aurora-pgsql-zetl> \
    --target-arn <arn:aws:glue:us-east-1:111122223333:catalog/zetl-catalog> \
    --integration-name <zetl-test-integration> \
    --data-filter "include: *.*" \
    --region us-east-1

When you run the command, the zero-ETL integration begins provisioning and enters a ‘creating’ state. The AWS CLI response provides key details about the integration configuration.

Example CLI output:

{
    "SourceArn": "<arn:aws:rds:us-east-1:111122223333:cluster:aurora-pgsql-zetl>",
    "TargetArn": "<arn:aws:glue:us-east-1:111122223333:catalog/zetl-catalog>",
    "IntegrationName": "<zetl-test-integration>",
    "IntegrationArn": "<arn:aws:rds:us-east-1:111122223333:integration:4c4d81b9-xxxx>",
    "KMSKeyId": "<arn:aws:kms:us-east-1:111122223333:key/b9130ae1-exxxx>",
    "Status": "creating",
    "Tags": [],
    "CreateTime": "2025-07-04T05:42:56.841000+00:00",
    "DataFilter": "include: *.*"
}

When the integration status changes to “active”, your zero-ETL integration pipeline is fully operational.

Monitor the integration

Before generating new live data, verify that the integration has reached an “active” state by running the describe-integrations AWS CLI command. This monitoring step is important to confirm that changes from your source Aurora cluster are successfully streaming to the AWS Glue managed catalog without errors:

aws rds describe-integrations
{
    "Integrations": [
        {
            "SourceArn": "<arn:aws:rds:us-east-1:111122223333:cluster:aurora-pgsql-zetl>",
            "TargetArn": "<arn:aws:glue:us-east-1:111122223333:catalog/zetl-catalog>",
            "IntegrationName": "<zetl-test-integration>",
            "IntegrationArn": "<arn:aws:rds:us-east-1:111122223333:integration:4c4d81b9-xxxx>",
            "KMSKeyId": "<arn:aws:kms:us-east-1:111122223333:key/b9130ae1-xxxx>",
            "Status": "active",
            "Tags": [],
            "CreateTime": "2025-07-04T05:42:56.841000+00:00",
            "DataFilter": "include: *.*"
        }
    ]
}

Verify the zero-ETL integration

Now that your historical data is loaded and the zero-ETL integration is “active”, you must confirm that the data has been successfully replicated.

Grant Lake Formation permissions

Before you can query the AWS Glue managed catalog by using the Amazon Redshift Data API, you must make sure the IAM user or role has the right permissions to create and manage tables within the catalog. Use the Lake Formation grant-permissions API to provide these necessary permissions so that Amazon Redshift can access your AWS Glue managed catalog for the zero-ETL integration. For more information, see Creating an Amazon Redshift managed catalog in the AWS Glue Data Catalog.

aws lakeformation grant-permissions \
    --region <us-east-1> \
    --cli-input-json '{
        "Principal": {
            "DataLakePrincipalIdentifier": "<arn:aws:iam::111122223333:role/Admin>"
        },
        "Resource": {
            "Table": {
                "DatabaseName": "my_db",
                "CatalogId": "<111122223333:zetl-catalog/zetl_0dff6d97-xxxx>",
            },
            "Permissions": [
                "CREATE_CATALOG",
                "DESCRIBE",
                "CREATE_DATABASE",
                "DROP",
                "ALTER"
            ],
            "PermissionsWithGrantOption": [
                "CREATE_CATALOG",
                "DESCRIBE",
                "CREATE_DATABASE",
                "DROP",
                "ALTER"
            ]
        }'

These permissions allow for query execution and metadata inspection on the managed catalog.

Query historical data by using the Amazon Redshift Data API

With the necessary permissions in place, you can now verify your historical data by querying the AWS Glue managed catalog through the Amazon Redshift execute-statement Data API. Begin this verification process by running a SELECT statement against the catalog:

aws redshift-data execute-statement --sql 'SELECT * FROM "zetl_0dff6d97-xxxx@zetl-catalog"."my_db"."products" LIMIT 10;' --database "<arn:aws:glue:us-east-1:111122223333:catalog/pg-zetl-catalog>"

The following command returns a unique query ID that you can use to monitor the execution status and retrieve results from your query:

{
    "CreatedAt": "2025-07-15T00:31:48.778000+00:00",
    "Database": "arn:aws:glue:us-east-1:111122223333:catalog/zetl-catalog",
    "DbUser": "IAMR:Admin",
    "Id": "ce1ff0el-xxxx",
}

Monitor your query’s progress by using the describe-statement API with the query ID. Continue checking until the status shows that your query has completed successfully:

//use the Id to make the describe-statement API call to verify execution status is Started
aws redshift-data describe-statement --id <ce1ff0el-xxxx>
{
    "CreatedAt": "2025-07-15T00:31:48.778000+00:00",
    "Database": "arn:aws:glue:us-east-1:111122223333:catalog/zetl-catalog",
    "DbUser": "IAMR:Admin",
    "Duration": 6238051060,
    "HasResultSet": true,
    "Id": "2ca8fedf-xxxx",
    "QueryString": "SELECT * FROM \"zetl_0dff6d97-xxxx_zeroetl@pg-zetl-catalog\".\"zetl_default\".\"products\" LIMIT 10;",
    "RedshiftPid": 1073791309,
    "RedshiftQueryId": 1018598,
    "ResultFormat": "json",
    "ResultRows": 1,
    "ResultSize": 149,
    "Status": "FINISHED",
    "UpdatedAt": "2025-07-15T00:31:55.491000+00:00"
}

To complete the verification process and view your historical data now available in Amazon SageMaker AI, retrieve the query results by using the get-statement-result API call:

aws redshift-data get-statement-result --id <ce1ff0el-xxxx>
{
    "Records": [
        [
            {
                "longValue": 1
            },
            {
                "stringValue": "Laptop"
            },
            {
                "stringValue": "High-performance laptop with 16GB RAM and 512GB SSD"
            },
            {
                "stringValue": "Electronics"
            },
            {
                "stringValue": "1299.99"
            },
            {
                "longValue": 50
            },
            {
                "booleanValue": true
            },
            {
                "stringValue": "2026-02-27 13:32:02.697117"
            },
            {
                "stringValue": "2026-02-27 13:32:02.697117"
            }
        ]
    ],
    "ColumnMetadata": [
        ....
        //Skipping metadata
    ],
    "TotalNumRows": 1
}

With your zero-ETL integration now active, you can demonstrate real-time data streaming by adding new data to your source Aurora PostgreSQL instance. Run the following INSERT query to add a new row, which shows how changes are automatically replicated in near real time:

INSERT INTO products (product_name, description, category, price, stock_quantity)
VALUES ('Wireless Mouse', 'Ergonomic wireless mouse with USB receiver and long battery life', 'Electronics', 29.99, 150);

You can verify that the recent changes from your source database have been replicated to the target environment within seconds. Use the same Amazon Redshift Data API workflow you used earlier to confirm the real-time replication:

aws redshift-data execute-statement --sql 'SELECT * FROM "zetl_0dff6d97-xxxx@zetl-catalog"."my_db"."products" LIMIT 10;' --database "<arn:aws:glue:us-east-1:111122223333:catalog/pg-zetl-catalog>"
{
    "CreatedAt": "2025-07-15T00:31:48.778000+00:00",
    "Database": "arn:aws:glue:us-east-1:111122223333:catalog/zetl-catalog",
    "DbUser": "IAMR:Admin",
    "Id": "2ca8fedf-a604-4c87-a183-3a553d62354c",
}

Use the describe-statement API call to monitor the query execution and confirm that the status shows ‘FINISHED’ before proceeding to retrieve the results:

aws redshift-data describe-statement --id <ce1ff0ef-xxxx>
{
    "CreatedAt": "2025-07-15T00:31:48.778000+00:00",
    "Database": "arn:aws:glue:us-east-1:111122223333:catalog/zetl-catalog",
    "DbUser": "IAMR:Admin",
    "Duration": 6238051060,
    "HasResultSet": true,
    "Id": "2ca8fedf-xxxx",
    "QueryString": "SELECT * FROM \"zetl_0dff6d97-xxxx_zeroetl@pg-zetl-catalog\".\"zetl_default\".\"products\" LIMIT 10;",
    "RedshiftPid": 1073791309,
    "RedshiftQueryId": 1018598,
    "ResultFormat": "json",
    "ResultRows": 2,
    "ResultSize": 317,
    "Status": "FINISHED",
    "UpdatedAt": "2025-07-15T00:31:55.491000+00:00"
}

Finally, retrieve the query results by using the get-statement-result API call:

aws redshift-data get-statement-result --id <ce1ff0ef-xxxx>
{
    "Records": [
        [
            {
                "longValue": 1
            },
            {
                "stringValue": "Laptop"
            },
            {
                "stringValue": "High-performance laptop with 16GB RAM and 512GB SSD"
            },
            {
                "stringValue": "Electronics"
            },
            {
                "stringValue": "1299.99"
            },
            {
                "longValue": 50
            },
            {
                "booleanValue": true
            },
            {
                "stringValue": "2026-02-27 13:32:02.697117"
            },
            {
                "stringValue": "2026-02-27 13:32:02.697117"
            }
        ],
        [
            {
                "longValue": 2
            },
            {
                "stringValue": "Wireless Mouse"
            },
            {
                "stringValue": "Ergonomic wireless mouse with USB receiver and long battery life"
            },
            {
                "stringValue": "Electronics"
            },
            {
                "stringValue": "29.99"
            },
            {
                "longValue": 150
            },
            {
                "booleanValue": true
            },
            {
                "stringValue": "2026-02-27 15:41:00.206273"
            },
            {
                "stringValue": "2026-02-27 15:41:00.206273"
            }
        ]
    ],
    "ColumnMetadata": [
        ....
        //Skipping metadata
    ],
    "TotalNumRows": 2
}

This verification process confirms that your zero-ETL integration from Aurora PostgreSQL to Amazon SageMaker AI is working and continuously replicating both historical and real-time data. Although zero-ETL integration significantly simplifies data replication, it’s important to understand certain limitations on supported data types, schema change handling, and data filtering capabilities. For more details about these considerations and best practices, see Aurora zero-ETL integrations and Amazon RDS zero-ETL integrations.

Clean up

This section guides you through the cleanup process to remove the resources and components you created during this walkthrough. When you delete a zero-ETL integration, Amazon Aurora removes it from the source Aurora DB cluster. Your transactional data isn’t removed from Amazon Aurora or the analytics destination, but Aurora doesn’t send new data to Amazon SageMaker AI.

Delete the zero-ETL integration: Begin the cleanup process by removing the integration between your source Amazon Relational Database Service (Amazon RDS) database and the AWS Glue managed catalog. Run the following command to delete the integration:

aws rds delete-integration --integration-identifier <arn:aws:rds:us-east-1:111122223333:integration:4c4d81b9-af2a-4b09-b922-007636ba7f66>

Delete the AWS Glue managed catalog: After you successfully delete the integration, delete the AWS Glue managed catalog that served as your zero-ETL target destination. Use the following command to remove the catalog:

aws glue delete-catalog --catalog-id <111122223333:zetl-catalog>

This permanently removes all associated table metadata and Amazon Redshift managed storage references.

Delete the Aurora DB cluster: If you created the source Aurora DB cluster for this demonstration and you no longer need it, you can complete the cleanup by deleting the entire DB cluster. By skipping the final snapshot option, you avoid retaining any test data and confirm complete resource removal:

aws rds delete-db-instance --db-instance-identifier <aurora-pgsql-zetl-instance-1> --skip-final-snapshot --region <us-east-1>

aws rds delete-db-cluster --db-cluster-identifier <aurora-pgsql-zetl> --skip-final-snapshot --region <us-east-1>

aws rds delete-db-cluster-parameter-group --db-cluster-parameter-group-name <aurora-zetl-cluster-pg> --region <us-east-1>

Conclusion

In this post, you learned how to configure zero-ETL integration between Aurora PostgreSQL and your Amazon SageMaker AI using AWS CLI. This integration automatically replicates your PostgreSQL data to a lakehouse in near real time, removing the need for custom ETL pipelines.

As you move forward, consider expanding this zero-ETL approach to more supported data sources, such as Amazon RDS for MySQL and Amazon DynamoDB. This creates a centralized data access strategy across your company. You can also explore advanced analytics scenarios by combining zero-ETL integrations with Amazon Redshift capabilities. These include large-scale SQL analytics, Amazon Redshift ML for in-database ML, and federated queries that span multiple data lakes and warehouses. These integrations provide the foundation for building a near real-time data platform that scales with your business needs.

To get started, see the AWS zero-ETL documentation for setup guidance, supported configurations, troubleshooting integrations, and architectural best practices.

Related posts and references:


About the authors

Apurwa Pawar

Apurwa Pawar

Apurwa is a Solutions Architect at AWS and a Data Analytics and AI enthusiast. She helps customers build their modern data strategy and cloud-native innovative solutions on AWS. She works with enterprise organizations across industries including healthcare, life sciences, financial services, and hospitality, partnering with engineering and business leadership to turn data into insight and action.

Sarika Subramaniam

Sarika Subramaniam

Sarika is a Solutions Architect at AWS, specializing in analytics and data platforms. She helps customers design scalable, secure, and cloud-based modern data architectures on AWS. She works with enterprise customers across industries, partnering with engineering teams to build innovative data solutions and drive business outcomes.

Frozen package management for air-gapped RHEL-family AMIs

Post Syndicated from Anand Krishna Varanasi original https://aws.amazon.com/blogs/compute/frozen-package-management-for-air-gapped-rhel-family-amis/

If you run a regulated, air-gapped compute fleet on RHEL-family instances, you have probably felt three requirements pulling against each other. Your organization must configure the network to remove internet access from the instances. Your team must review and approve new packages or version upgrades before you adopt them. Your team removes public repository definitions, restricts network paths, and configures instances to use only the internal repository your team has approved. Teams in chip design, finance, healthcare, defense, and the public sector often face this combination while still needing operating system updates.

This post describes a two-account pattern that separates the connected package-ingestion path from the air-gapped fleet. You create an authorized initial baseline and approve later changes to form a versioned package snapshot in Amazon Simple Storage Service (Amazon S3). EC2 Image Builder uses the frozen snapshot to build Amazon Machine Images (AMIs). Your team configures AWS Systems Manager Patch Manager to patch the instances your organization runs from the same internal package source.

The accompanying reference implementation demonstrates the pattern for RPM-based RHEL-family systems (AlmaLinux for example). It is a reference, not a substitute for distribution of licensing, vulnerability analysis, testing, or an organization’s change-management process.

The challenge: Getting packages into an air-gapped approval-gated fleet

Common delivery models each assume something an air-gapped fleet might not provide:

  • Red Hat Update Infrastructure (RHUI) expects each instance to reach the service. A fleet with no internet egress needs a different content path.
  • Red Hat Satellite supports disconnected content management, but it is a separate product and operational footprint. Teams that need a custom package-level approval workflow must integrate that workflow with their content-management process.
  • The Red Hat CDN requires a connected, entitled content-management path. Centralizing that path changes the network architecture, not the customer’s Red Hat subscription obligations.

The objective is not to replace these products universally. It is to show a serverless AWS pattern for teams that need an authorized repository baseline, explicit approval for later package changes, and a fleet with no public package source.

How the pattern works

The pattern combines three controls:

  1. A frozen package repository on Amazon S3: The pattern stores a deployment-authorized baseline and subsequent approved package changes in versioned, per-OS repository prefixes. The repository manifest records the package inventory for each state.
  2. EC2 Image Builder Orchestration builds AMIs from that repository: The build helps remove upstream repository definitions and configures the internal frozen mirror to be used for all package operations.
  3. The launched fleet has no internet egress: The dnf operations are configured to resolve the internal mirror. Patch Manager uses the same repository source, so image builds and in-place patching draw from one frozen snapshot.

Choosing the upstream source

Choose one package lineage end to end. The parent AMI, repository content, and trusted signing keys must belong to that same lineage.

The reference implementation defaults to AlmaLinux vault content plus EPEL and an AlmaLinux parent AMI. The AlmaLinux OS Foundation states that AlmaLinux aims for binary and application binary interface (ABI) compatibility with RHEL. This is an AlmaLinux compatibility goal, not a Red Hat certification, and it does not make repository mixing a supported practice.

For genuine RHEL systems, use a Red Hat parent AMI, entitled Red Hat repositories, and Red Hat signing keys. A connected content-management host can retrieve content for the isolated environment. This centralizes the network path but does not reduce or change the customer’s Red Hat subscription obligations. Confirm those obligations against the applicable Red Hat agreement.

Do not pair AlmaLinux repositories with genuine RHEL hosts, or Red Hat repositories with AlmaLinux hosts. Mixed-vendor package lineages can create support, stability, and maintainability problems even when the RPMs appear mechanically compatible.

Architecture and workflow

The account boundary provides a primary security boundary for this architecture. The following diagram shows the connected Distribution account, the read-only Workload account, and an example cross-Region layout.

Two-account architecture showing the connected Distribution account with the control plane and internet path, and the air-gapped read-only Workload account, spanning two Regions

Figure 1: Two-account, cross-Region architecture separating the connected Distribution account from the air-gapped Workload account

The Distribution account owns the writable control plane and the only internet path. It runs Amazon EventBridge, three AWS Lambda functions, Amazon DynamoDB, Amazon Simple Notification Service (Amazon SNS), and the AWS Fargate sync task. It also owns the frozen S3 repository and its AWS Key Management Service (AWS KMS) key.

The Workload account is air-gapped and read-only with respect to the repository. It runs the internal HTTPS mirror, EC2 Image Builder, Patch Manager, and the compute fleet. Its mirror task role can read and decrypt frozen content but cannot write it.

The sample repository places the Distribution control plane in US East (N. Virginia), the frozen store in US West (Oregon), and the Workload resources in US West (Oregon) to demonstrate API-only cross-account and cross-Region operation. This Region split is not required. In most deployments, place the Distribution control plane and frozen store in the same Region unless data residency, disaster recovery, or an existing regional footprint justifies the additional latency, transfer cost, and KMS policy complexity.

The two accounts do not need Amazon Virtual Private Cloud (VPC) peering or a transit gateway. Cross-account access uses S3, KMS, and IAM policies. The VPC address ranges can overlap because no VPC-to-VPC route is required.

Package baseline and scheduled upgrade workflow

Before the scheduled workflow begins, your organization must authorize and run a full sync to establish the initial repository baseline. This bootstrap does not provide package-by-package approval. If your organization requires individual approval for every initial RPM, your team should generate and review the baseline manifest before promotion instead of relying solely on deployment authorization.

After the baseline, the detector runs on a customer-defined schedule. The reference implementation defaults to monthly. The following diagram shows the bootstrap distinction and the selective approval flow.

Workflow diagram distinguishing the initial baseline bootstrap sync from the recurring detect, request approval, review, record, selective sync, and manifest update steps

Figure 2: Package baseline bootstrap and the scheduled selective approval workflow

  1. Detect. Amazon EventBridge invokes the detector Lambda function on the configured schedule. The detector compares upstream repository metadata with manifest.json, which records the current frozen inventory. It classifies a newer version as an upgrade and an absent package as new.
  2. Request approval. The detector writes candidates to S3, creates a KMS-protected review token carrying the request ID and expiry, and is designed to send a review link through SNS. The detector can use kms:Encrypt but not kms:Decrypt.
  3. Review. A human opens the review page through an Amazon API Gateway HTTP API, reviews the proposed package versions, and chooses which changes to approve. The approver can use kms:Decrypt but not kms:Encrypt.
  4. Record and start. A conditional DynamoDB update changes a request from pending to approved only once. The approver then starts the Fargate sync task and passes the request ID.
  5. Selective sync. The task reads the approved package list, downloads those package versions, is designed to perform verification checks, and regenerates repository metadata.
  6. Update the manifest. When the task stops, Amazon EventBridge invokes the manifest-updater Lambda function. It archives the outgoing manifest and records the resulting repository inventory.

The approval decision controls adoption. It does not prove that package code is safe. Advisory review, vulnerability scanning, testing, and staged rollout remain in separate controls.

Evidence from the approval workflow

The token ties a review action to a specific request and expiry. Separating kms:Encrypt from kms:Decrypt prevents either Lambda function from performing both token roles. The conditional DynamoDB write makes the approval transition single-use.

DynamoDB records request state, AWS CloudTrail records control-plane API activity, and manifest history records repository inventory changes. These service records can feed the organization’s existing audit and evidence-management workflow. Object-level S3 access auditing requires CloudTrail S3 data events. KMS activity alone is not a substitute for those events.

The frozen package repository on Amazon S3

The following diagram shows the per-OS, per-component prefix layout, and manifest objects.

Amazon S3 prefix layout with one prefix per operating system version, each holding BaseOS, AppStream, and EPEL components with Packages and repodata trees plus manifest objects

Figure 3: Per-OS, per-component prefix layout of the frozen repository on Amazon S3

Each pinned operating system version receives its own prefix. Repository components such as BaseOS, AppStream, and EPEL contain Packages/ and repodata/ trees. manifest.json records the active inventory, and archived manifests preserve historical evidence and comparison points.

Your organization configures the bucket with versioning and SSE-KMS. Public RPM content does not require a customer-managed KMS key for confidentiality, so your organization could instead configure SSE-S3 for encryption at rest. However, SSE-S3 would remove the separate cross-account authorization control provided by the customer-managed KMS key policy. The customer managed key is used here for explicit cross-account key-policy control and revocation, and the manifests reveal the fleet’s exact software inventory. S3 Bucket Keys reduce KMS request volume. If object-level access evidence is required, enable CloudTrail S3 data events.

A rollback must restore a coherent repository state, including metadata and any required object versions. Restoring only manifest.json does not roll back repository contents.

Building, patching, and running the fleet

At AMI build time, an Image Builder component installs the configured repository keys, moves existing repository definitions aside, and writes one frozen repository definition per component. It locks the package manager to the frozen repository directory, fetches metadata through the internal mirror, and fails the build if the mirror validation step fails. An optional curated package list demonstrates that the AMI can install real packages through the frozen path.

Patch Manager uses the same mirror for the running fleet. A host created from an older AMI and a newly built host are therefore patched toward the same frozen snapshot. Instances run without an Amazon VPC NAT gateway, public IP, or an Amazon VPC internet gateway route in the Workload VPC, and their repository configuration contains no public fallback.

The intended verification model is defense in depth: the sync task helps verify a vendor’s signature before content enters the trusted repository, and the system verifies it again at installation through dnf. The ingestion gate helps reject digest-only results and can be configured to help confirm that only valid package signatures from a trusted lineage key are accepted.

The package mirror

Nginx fronts aws-sigv4-proxy, which signs cross-account S3 GET requests using the mirror task role. To a client, the service appears as a standard HTTPS package repository behind an internal Application Load Balancer and private DNS name.

Use the latest version of aws-sigv4-proxy (current latest is v1.12). This version 1.12 contains the fix for signing S3 paths (or the OS package names) with special characters such as +. Earlier versions can return SignatureDoesNotMatch. The reference implementation pins the reviewed v1.12 release commit immutably. Keep it current through dependency-update reviews.

Security boundaries and limits

The design provides the following controls:

  • No automatic public-repository adoption: A new upstream version enters the selective path only after an explicit, recorded decision by the user.
  • Repository ingestion and installation checks: You configure strict sync-time signature validation to help validate content before it enters the trusted store. dnf verifies again during installation.
  • No package-channel egress: Workload instances are configured to prevent access to public package sources.
  • A read-only workload boundary: A Workload-account principal cannot modify the frozen repository.

Human approval is not a malware detection. A reviewer cannot reliably identify a backdoor in a legitimately signed package merely by seeing its name, version, or changelog. Use vulnerability intelligence, scanning, pre-production tests, and staged deployment as additional controls.

The approval state, manifests, and CloudTrail records can help support evidence for control frameworks such as SOC 2 change management, ISO 27001 patch-management controls, and FDA 21 CFR Part 11 electronic records. Applicability depends on the organization’s environment, audit scope, and assessor. Confirm it with the compliance team under the AWS shared responsibility model.

Cost and operations

Cost depends on the amount of repository content and the chosen networking and availability design. Components can include S3 storage and requests, KMS requests, Lambda invocations, DynamoDB, SNS, Fargate tasks, the internal load balancer, Distribution-account internet egress, Amazon VPC endpoints, and AMI snapshots. A three-task always-on mirror costs more than an S3 bucket alone. Estimate the target topology with current AWS pricing rather than applying a fixed monthly figure from the sample.

Run detection and review at an interval defined by patch policy and risk tolerance. The supplied default is monthly, but the Terraform input is configurable. If a package change must be reversed, restore a tested, coherent repository version and rebuild or patch affected hosts as appropriate.

Prerequisites

To set up the reference implementation, work through these in order:

  1. Two AWS accounts: a connected Distribution account and an air-gapped Workload account.
  2. Deployment tools: Terraform 1.5 or later, Terragrunt, Finch or Docker, and Python with pip.
  3. AWS Command Line Interface (AWS CLI): one named profile per account.
  4. Distribution networking: private subnets with internet egress that works without public IPs, security-group egress on port 443, and DNS resolution for public names.
  5. Workload networking: VPC interface endpoints for ssm, ssmmessages, ec2messages, logs, kms, and imagebuilder, plus an S3 gateway endpoint.
  6. Internal mirror identity: an AWS Certificate Manager (ACM) certificate and a private hosted zone. If the parent AMI does not trust the issuing CA, configure the CA file so the build installs the trust anchor.
  7. Package lineage: a parent AMI, repositories, and signing keys from the same distribution lineage. A RHEL subscription is required when retrieving genuine entitled Red Hat content.
  8. Optional deployment roles: otherwise, the stack uses each profile’s credentials.

The deployment creates state backend and Amazon Elastic Container Registry (ECR) repositories. Do not create those ECR repositories separately before applying their own Terraform units.

Reference implementation

The companion repository provides Terraform modules, Lambda handlers, two container images, a Terragrunt two-account layout, and Makefile targets for deployment and verification. The shipped alma810 example defaults to the AlmaLinux lineage (RHEL family).

Choose one distribution lineage before deployment:

  • AlmaLinux Parent Image default: use an AlmaLinux parent AMI, AlmaLinux vault repositories, EPEL, and the included AlmaLinux and EPEL signing keys. No Red Hat subscription is required.
  • Genuine RHEL Parent Image: use a Red Hat parent AMI, an entitled Red Hat content source, and Red Hat signing keys. The customer supplies the Red Hat subscription and content-access integration.

Note: The reference implementation only provides AlmaLinux lineage setup, not genuine RHEL. If you choose to use a genuine RHEL parent image lineage, only the reference implementation code needs to be updated to fetch the Red Hat credentials or subscription access, and the rest of the workflow remains the same.

Please follow the repository README for detailed setup instructions:

  1. Fill in environments/config.hcl and both account files with the two accounts, networking, mirror certificate, package lineage, and parent AMI.
  2. Create the Terraform state backend with make bootstrap DIST_PROFILE=<dist> WORK_PROFILE=<work>.
  3. Review both account plans with make plan DIST_PROFILE=<dist> WORK_PROFILE=<work>.
  4. Deploy in dependency order with make all BASELINE_APPROVED=true DIST_PROFILE=<dist> WORK_PROFILE=<work>. The baseline sync can take several hours.

Validate the internal package management workflow

Validation begins by running the AlmaLinux EC2 Image Builder pipeline. A successful build produces private, encrypted AMIs in the configured AWS Regions. To validate the configuration, launch a test instance from the generated AMI in a no-egress Workload subnet and verify that the instance is connected to the internal frozen repository for all package management workflow.

Terminal output listing only the internal frozen-baseos, frozen-appstream, and frozen-epel repositories enabled, with dnf makecache downloading metadata through the private mirror

Figure 4: Instance showing only the internal frozen repositories enabled, with no public repository configured

The output confirms that only the internal frozen AlmaLinux repositories are enabled: frozen-baseos, frozen-appstream, and frozen-epel. The dnf makecache command successfully downloads metadata for all three repositories through the private mirror. No public repository is configured or used.

Now try installing, upgrading, or installing a new package to test that package operations are served by the internal frozen repository.

Terminal output showing a package install and upgrade completing successfully through the internal frozen repository mirror

Figure 5: Package install and upgrade served by the internal frozen repository

  • Fail closed on approval data. If the approved package list cannot be retrieved, stop the sync task and avoid substituting a full sync.
  • Govern the initial baseline. Record who authorized the bootstrap full sync, or require explicit review of its manifest before promotion.
  • Require vendor signatures at ingestion. Do not treat a valid package digest as equivalent to a trusted signature.
  • Keep package lineages consistent. Parent AMI, repository content, and signing keys must come from the same distribution lineage.
  • Use a customer-defined review cadence. Monthly is only the sample default.
  • Use immutable dependency references with active updates. Require aws-sigv4-proxy v1.12 or later, pin the reviewed artifact by digest or full SHA, and automate update proposals.
  • Test rollback as a repository operation. Restore metadata and objects together, then validate the mirror before using the restored state.

Clean up

The walkthrough deploys billable resources in both accounts. Please follow the repository README for detailed setup cleanup instructions.

Conclusion

This pattern separates connected package ingestion from an air-gapped fleet, provides an authorized package baseline and an explicit decision point for later package changes or upgrades, and keeps image builds and running hosts on one frozen repository snapshot. To learn more, visit the EC2 Image Builder service page, the EC2 Image Builder documentation, the Patch Manager documentation, and the Amazon S3 user guide. The reference implementation is available in aws-samples.

Building event-driven applications at scale with Amazon EventBridge

Post Syndicated from Nahid Karimaghalou original https://aws.amazon.com/blogs/compute/building-event-driven-applications-at-scale-with-amazon-eventbridge/

Event-driven applications on Amazon EventBridge usually start small and then spread. One team creates a Custom event bus, adds a few rules, and ships. Another team needs some of those events, so a rule forwards them to a bus in a second account. A third team needs a subset of what the second team receives, so another rule forwards again. A year later the organization runs dozens of Custom event buses joined by forwarding rules, and that topology has become a thing to operate in its own right.

That shape has a price, and the smallest part of it is the bill. Every forwarding hop is a separate ingestion, so cost tracks the topology rather than the number of consumers that needed the event. The harder problem is that nobody can see the whole picture. Governance spreads across the accounts it was meant to cover. Answering who publishes to a bus, who consumes a given event type, or what breaks when a team stops publishing means visiting each account and reading its rule configuration. Tracing one event is harder still: its path crosses several buses in several accounts, each with its own metrics and logs, and no single view follows it from publication to the consumer that never received it.

Application teams also wait. Publishing to a bus in another account, or consuming from one, needs a resource policy, a role, and a forwarding rule owned by a central team. The team that wants to build opens a ticket, and the platform team becomes a queue. Both the missing visibility and the waiting grow with every team onboarded.

Amazon EventBridge recently relaunched the Custom event bus, which tackles these challenges directly. A platform team creates one bus, shares it across the organization, and keeps control of who can publish and who can subscribe. Every consumer of those events is listed on the one bus rather than inferred from configuration spread across accounts. Application teams create their own Subscribers in their own accounts. The bus stores events for a retention period you choose, preserves order within a key the publisher sets, accepts Avro and Protocol Buffers (Protobuf) alongside JSON (including CloudEvents), and delivers to targets without a function in the path to translate a call. It runs alongside the Custom event bus – classic, so adoption is incremental.

In this post, you see how a platform team stands up a shared bus and governs access to it, how application teams onboard themselves with a single Subscriber resource, and how retention, ordering, open formats, transformation, and direct target integrations change what one bus can carry.

One bus, shared with the organization

The platform team’s job on a shared bus is narrower than it was on a fleet of them. It owns the bus and sets the boundaries: which principals can publish and what their events can declare, which principals can subscribe, and, where it matters, what those principals are allowed to filter on. Application teams then manage their own configuration within those boundaries, such as filters, targets, delivery roles, retry policies, and failure destinations, none of which the platform team needs to write or review. That division is the point of the design. The platform team keeps governance of the bus and stops owning everyone else’s configuration, which is what takes it out of the provisioning path without giving up control of who is on the bus.

Creating the bus is a single call in a platform account.

BUS_ARN=$(aws eventsv2 create-event-bus \
    --name company-events \
    --storage-configuration '{"RetentionPeriodInDays":7}' \
    --query EventBusArn --output text)

Retention is the one setting worth deciding deliberately here rather than revisiting after an incident. It runs from 1 to 365 days and can be modified later, but a change only applies going forward. Raising it widens the window for events published from that point on, and does not make older events readable again. Seven days covers a working week of history, which is usually enough to onboard a consumer or reprocess after a bug without paying to store a year of events nobody will read.

Sharing the bus is the second decision. AWS Resource Access Manager is the route to reach for first: it associates automatically for accounts in the same organization and reaches accounts outside it by invitation the consumer accepts. A resource policy written on the bus directly is the alternative, and can also name accounts inside or outside the organization.

Access is granted per principal, and publishing and subscribing are separate permissions. A team that produces order events gains no ability to read payment events from the same bus. One grant is not enough for a cross-account caller, as usual on AWS: the role that publishes or subscribes also needs its own IAM policy allowing those actions. The platform team decides which accounts can reach the bus, and each consuming team decides which of its own principals can use that access.

Taken together, those decisions produce the architecture in the following diagram. One bus lives in a platform account, and application teams publish to it and subscribe from their own accounts. An AWS Lambda function in Team A’s account calls PutRawEvents to publish events onto the Amazon EventBridge bus in the platform account. Team B and Team C each attach their own Subscriber: Team B’s delivers to a Lambda function, Team C’s to an Amazon DynamoDB table.

Architecture diagram of one Custom event bus in a platform account. A Lambda function in Team A’s account calls PutRawEvents to publish events onto the Amazon EventBridge bus in the platform account. Team B and Team C each attach their own Subscriber in their own accounts: Team B’s Subscriber delivers to a Lambda function, and Team C’s Subscriber delivers to an Amazon DynamoDB table.

Figure 1: Multi-account sharing

Cost follows team boundaries because charges separate ingestion from delivery. The account that publishes an event pays to put it on the bus, and the account that owns a Subscriber pays for what that Subscriber consumes. Each team’s usage appears on its own bill, which is what makes a shared bus something a platform team can charge back rather than a shared cost center nobody can decompose. Removing the forwarding hops also removes the duplicated ingestion and delivery those hops created: the same event reaching the same three consumers is ingested once instead of three times.

Publishing in the format teams already use

Not every producer speaks JSON. Teams that standardize event exchange across an organization often register schemas and publish compact binary payloads, because the schema is the contract between teams that deploy on their own timetables. Accepting the formats those producers already emit is simpler than changing each one to convert to JSON first.

With the new Custom event bus, application teams can publish events in Avro, Protobuf, and CloudEvents (JSON) formats. For the binary formats, a schema registry named on the request is used to deserialize the events.

There are two publish APIs, and the payload decides which one to call. PutEvents takes structured JSON with the familiar Detail, Source, and DetailType fields. PutRawEvents takes a binary payload plus metadata you define, and is the one to use for Avro, Protobuf, CloudEvents, or bytes the bus should not interpret.

import boto3

events = boto3.client("eventbridgev2")
events.put_raw_events(
    EventBusArn=BUS_ARN,
    SchemaRegistryConfiguration={"RegistryUri": GLUE_REGISTRY_ARN},
    Entries=[
        {
            "Data": avro_encoded_order,  # bytes, straight from your existing producer
            "SystemMetadata": {"ContentType": "application/avro"},
            "Metadata": {"eventType": "OrderPlaced"},
        }
    ],
)

The schema registry can be either the AWS Glue Schema Registry or the Confluent Cloud Schema Registry.

Because the bus decodes the event before filters and transformations run, a consumer subscribing to Avro events written by another team needs no schema, no decoder, and no access to the registry. It writes the same filter it would write against JSON. Producers and consumers stay decoupled, and no deserialization code has to be repeated in each consuming team.

Publishers get one more setting on the same request: deduplication. A retry that already succeeded would otherwise leave a duplicate for every consumer to handle. It works one of two ways: the bus hashes the content of each event, or it uses a deduplication ID you supply. Content-based hashing suits producers with no natural key, since two identical events hash the same. A deduplication ID fits when you already have one, such as an order ID combined with a state transition. It keeps matching even when parts of the payload differ in ways that should not count as a new event.

Self-service onboarding for application teams

The new Custom event bus introduces a new resource called a Subscriber. Application teams create and configure their own Subscribers in their own accounts, provided they have been granted subscribe access to the bus. A Subscriber is the one place a consumer’s behavior is defined: which events it receives, where they are delivered, how delivery is retried, and where events go when delivery does not succeed. Reviewing or changing a consumer is one thing to read and one thing to update.

SUBSCRIBER_ARN=$(aws eventsv2 create-subscriber \
    --name orders-to-fulfilment \
    --event-bus-arn "$BUS_ARN" \
    --filter-configuration '{"Filters":[{"Scope":"METADATA","Pattern":"{\"eventType\":[\"OrderPlaced\"]}"}]}' \
    --invoke-configuration '{"TargetArn":"'"$QUEUE_ARN"'","RoleArn":"'"$ROLE_ARN"'"}' \
    --retry-policy '{"MaxRetryAttempts":10,"MaxEventAgeInSeconds":3600}' \
    --on-failure-configuration '{"Arn":"'"$DLQ_ARN"'"}' \
    --query SubscriberArn --output text)

A filter’s scope decides which part of the event the pattern is matched against. DATA matches the payload, METADATA matches the key-value pairs the publisher attached to the event, and SYSTEM_METADATA matches the event’s system fields: the content type and ordering key a publisher declares, plus the fields Amazon EventBridge adds itself. Because Avro and Protobuf payloads are decoded as they are published, a DATA filter reads their fields directly, the same as it would for JSON.

The retry policy says how the bus should behave when a target is failing. MaxRetryAttempts sets how many times a delivery is retried, and MaxEventAgeInSeconds sets how long an event stays eligible for retry, measured from when it was published. Retries stop as soon as either limit is reached, so both bound the same delivery.

When deliveries do fail, the reason shows up in the Subscriber’s own logs, which application teams can turn on themselves. They record the error from each delivery attempt alongside the exact input sent to the target, which makes a problem quick to place. Seeing what the target actually received separates a transformation that produced the wrong shape from a target that rejected a correct one.

Screenshot of the Amazon EventBridge console showing the Create subscriber form, with fields for the subscriber name, event bus, filter configuration, target (invoke configuration), retry policy, and on-failure destination.

History for consumers that did not exist yet

A Subscriber sometimes needs events that were published before it existed. For example, a new analytics service needs hydrating with recent history, or a target processed a window of events incorrectly and needs that window replayed. Because the bus retains events for the period configured on it, a Subscriber can be created with a starting position in the past, so it reads history, catches up, and continues with live traffic:

aws eventsv2 create-subscriber \
    --name analytics-backfill \
    --event-bus-arn "$BUS_ARN" \
    --starting-position POINT_IN_TIME \
    --point-in-time-configuration '{"PointType":"TIMESTAMP","StartingPoint":"2026-09-14T06:00:00Z"}' \
    --filter-configuration '{"Filters":[{"Scope":"METADATA","Pattern":"{\"eventType\":[\"OrderPlaced\"]}"}]}' \
    --invoke-configuration '{"TargetArn":"'"$ANALYTICS_ARN"'","RoleArn":"'"$ROLE_ARN"'"}'

A starting position is either LATEST or POINT_IN_TIME. Choosing POINT_IN_TIME then needs a point-in-time configuration: a PointType of TIMESTAMP with a starting point, or HORIZON to begin at the earliest event still retained. An optional end point stops the read at a chosen time, which is what you want when reprocessing a known-bad window rather than catching up to live traffic.

Two things to keep in mind. The starting position is fixed when the Subscriber is created, so reading a different window means a new Subscriber. Treat the starting position as part of a Subscriber’s identity rather than a dial to turn later. And retention cannot reach back beyond the retention window, so the read starts at the earliest retained event however far back the timestamp asks for.

Order, where order matters

In event-driven architectures, where components are built to work asynchronously, the order events arrive in usually does not matter. There are still use cases where a consumer relies on ordered delivery, and the new Custom event bus offers it as an option on individual Subscribers.

Ordering is scoped by a key the publisher sets. A publisher includes an event group ID (a customer ID, an order ID, a driver ID), and a Subscriber created with FIFO delivery type receives the events for each group in the order they were published. A FIFO Subscriber reading events published without a group ID has nothing to sequence by, so the two sides work together. Creating one takes the same call as an unordered Subscriber, with the delivery type set to FIFO:

aws eventsv2 create-subscriber \
    --name inventory-ordered \
    --event-bus-arn "$BUS_ARN" \
    --type FIFO \
    --filter-configuration '{"Filters":[{"Scope":"METADATA","Pattern":"{\"eventType\":[\"OrderPlaced\"]}"}]}' \
    --invoke-configuration '{"TargetArn":"'"$FIFO_QUEUE_ARN"'","RoleArn":"'"$ROLE_ARN"'","SqsParameters":{"MessageGroupId":"{% $events.SystemMetadata.EventGroupId %}","MessageDeduplicationId":"{% $events.SystemMetadata.DeduplicationId %}"}}'

Ordering is per group, so throughput scales with the number of groups. If an event cannot be delivered, it holds up the rest of its own group while other groups keep moving. Choosing the key therefore matters: one that maps to a business entity, such as an order or a customer, gives sequencing where it is needed and independence everywhere else. A key so broad that most events share it puts them all in a single sequence, and a key so specific that every event has its own leaves nothing to order.

Because ordering is set on each Subscriber, consumers of the same events do not need to agree on it. An inventory service can receive a group’s events in sequence while an analytics service subscribing to those same events takes them as they arrive.

Reshaping events, and delivering directly to a target

A consumer’s business logic expects events in a particular shape, and the events on the bus are not always in that shape. Where the two get reconciled is an ownership decision: inside the consumer, where it becomes part of that team’s code, or on the Subscriber, ahead of it.

The first case is reformatting. A downstream system, often owned by another domain or outside the organization entirely, expects a different structure from the one the publisher emits. A JSONata transformer on the Subscriber produces that structure before delivery, so the consumer receives what it already expects. The business logic stays where it belongs, and when the published shape changes upstream, or another event type needs deriving into the same input, it is the transformer that changes rather than the consumer:

--transformer '{
    "Type":"JSONATA",
    "JsonataConfiguration":{
        "Expression":"{% {\"orderRef\": $events.Data.detail.orderId, \"total\": $events.Data.detail.amount} %}"
    }
}'

The transformer type determines the shape of what gets delivered. RAW delivers the event payload as is and is the default, so a Subscriber with no transformer configuration receives only the payload. WITH_METADATA adds the event envelope alongside it, and JSONATA reshapes it with an expression wrapped in {% %}.

The transformation reshapes events only for the Subscriber that owns it and does not affect what other Subscribers of the same bus receive. That also makes it a data minimization control: a partner can receive only the fields it needs rather than a whole internal event. Defining it at the Subscriber means it holds for every event without anyone remembering to strip fields.

The second case is calling an AWS service API. A Subscriber delivers directly to targets including Amazon Simple Queue Service (Amazon SQS), Amazon Simple Notification Service (Amazon SNS), AWS Lambda, and Amazon Kinesis Data Streams. For other services it has been common practice to add a proxy step whose only job is to make the call. With universal targets, the new Custom event bus can call a supported AWS service API directly, with the request body built by a JSONata expression.

TargetArn: arn:aws:events:::aws-sdk:dynamodb:putItem
UniversalTargetParameters.Input:
{% { "TableName": "orders", "Item": { "pk": { "S": $events.Data.detail.orderId } } } %}

Note that a universal target shapes its input through that parameter rather than through the preceding transformer, and setting a transformer on one is rejected when the Subscriber is created. The two mechanisms do the same kind of work on different targets.

That removes the proxy processing that existed only to make the call. The delivery role still needs the action the target requires and getting that wrong is the most common cause of a Subscriber that looks healthy and delivers nothing.

Conclusion

Running an event-driven application across many accounts no longer means running many event buses and the forwarding between them. A platform team creates one new Custom event bus, shares it across the organization through AWS Resource Access Manager or a resource policy on the bus, and keeps one place to decide who publishes and who consumes. Application teams create and own their Subscribers without waiting for provisioning. Ingestion and delivery are charged separately, so each team’s usage appears on its own bill, and the duplicated ingestion that forwarding hops created disappears with the hops.

The capabilities that used to send individual teams elsewhere now sit on the same bus. Ordering is per Subscriber and scoped by a publisher-supplied key, so one team’s sequencing requirement no longer fragments an architecture. Retention makes it possible to onboard a consumer that needs history it was never subscribed to. Avro and Protobuf are decoded by the bus, so producers keep their binary contracts. Transformation and universal targets keep business logic where it belongs, removing the proxy steps that existed only to reshape an event or make an API call.

Because the new Custom event bus runs alongside the Custom event bus – classic, adoption is incremental. Point one new consumer at a shared bus or forward a slice of an existing bus into it and move the rest as teams are ready.

Next steps. Create a bus, add a Subscriber, and publish an event, starting from the Amazon EventBridge documentation for the resource model and the AWS Command Line Interface (AWS CLI) reference. If you already run Custom event buses, the migration guidance covers routing existing events into a new Custom event bus without changing producers. From there, look at the Subscriber logging and metrics options for tracing an event from publication to delivery, and at AWS Resource Access Manager for how sharing and permissions work across an organization. If you have questions or feedback about the new Custom event bus, leave a comment on this post. We’d like to hear how you’re using it.

How Property Finder automated incident management with AWS DevOps Agent

Post Syndicated from Nada Tlohi original https://aws.amazon.com/blogs/devops/how-property-finder-automated-incident-management-with-aws-devops-agent/

When a production service starts saturating the CPU at 1 AM, every minute counts for incident management. For Property Finder, a production incident could mean failed searches, frustrated users, and direct revenue impact. Property Finder is the leading property portal in the Middle East and North Africa (MENA), serving millions of property seekers across five markets.

Before adopting AWS DevOps Agent, incident response followed a familiar pattern: an alert fires, an on-call engineer wakes up, spends 20–40 minutes correlating metrics across tools, manually documents findings, and opens a fix. Mean Time to Resolution stretched to 2–3 days for non-critical issues.

Today, that entire workflow runs autonomously. From alert to root cause analysis, Slack notification, Jira ticket, on-call phone call with context, and auto-remediation pull request (PR), the full lifecycle completes in 14 minutes. This post walks through the implementation and shows how a separate custom agent that automatically generates code fixes is the key differentiator.

The business problem

Property Finder runs a distributed microservices architecture on Amazon Elastic Container Service (Amazon ECS) fronted by Application Load Balancers (ALBs). When infrastructure issues occur, the impact is immediate: users see failed searches, agents cannot update listings, and revenue is directly impacted during peak hours.

The traditional workflow had three gaps:

  1. Detection lag. Non-critical anomalies could go undetected for days.
  2. Context switching. Engineers bounced between five or more tools per incident.
  3. Knowledge silos. Runbooks lived in people’s heads, not automation.

Solution architecture

Property Finder’s implementation connects AWS DevOps Agent at the center of a three-tier pipeline: Detection and Trigger, Autonomous Investigation, and Event-Driven Output.

Three-tier incident pipeline from a CloudWatch alarm through AWS DevOps Agent investigation to Slack, Jira, and GitHub outputs

Figure 1: End-to-end autonomous incident management architecture

The numbered steps correspond to the data flow in Figure 1:

  1. ECS CPU spike triggers an Amazon CloudWatch Alarm. CloudWatch Metrics Insights monitors service health across all ECS clusters. When sustained CPU exceeds 98%, the alarm transitions to ALARM state.
  2. AWS Lambda formats and HMAC-signs the payload. Triggered directly by the CloudWatch alarm action (which fires only on ALARM state transitions), AWS Lambda enriches the payload with service metadata, signs it with HMAC-SHA256 using credentials from AWS Secrets Manager, and POSTs to the webhook.
  3. The agent begins autonomous investigation. Parallel subagents query ECS metrics, AWS CloudTrail, ALB traffic patterns, and Grafana telemetry (Prometheus, Loki, Pyroscope). The agent reads relevant source code from GitHub for correlation.
  4. Findings post to Slack in real time. The native Slack integration posts investigation progress to #incidents. The full root cause analysis, impact assessment, and mitigation plan appear at the end of the thread.
  5. Investigation Completed event fires to Amazon EventBridge. Amazon EventBridge triggers an orchestrator Lambda that fans out to three independent targets simultaneously.
  6. Lambda creates a Jira ticket with the full root cause analysis. The Lambda retrieves the investigation summary from journal records and creates a prioritized ticket with root cause, severity, and affected service.
  7. Grafana IRM pages the on-call engineer by phone. A Lambda posts a Grafana Alerting-compatible payload to the IRM webhook. The escalation chain calls the engineer with full investigation context: what broke, why, and the recommended fix.
  8. The remediation agent opens a GitHub PR with the auto-fix. It receives the root cause, generates a Terraform or code fix, and opens a Draft PR through a GitHub Model Context Protocol (MCP) server. Engineers review before merging.

A real incident

The example-service, Property Finder’s core property search microservice serving millions of queries per day across five MENA markets, experienced CPU saturation at 99.11%. The pipeline resolved it end-to-end in 14 minutes.

1:21 AM │ Alarm fires (ECS CPU > 98%)

1:22 AM │ Investigation starts + Slack posted

1:22 AM │ 4 parallel subagents launched

1:32 AM │ Root cause identified

1:33 AM │ Jira ticket [redacted] created

1:34 AM │ On-call paged via phone call

1:35 AM │ GitHub PR [redacted] opened with fix

The detection Lambda handles three tasks: (1) retrieves the webhook secret from AWS Secrets Manager, (2) enriches the CloudWatch alarm event with ECS service metadata (cluster name, service name, task count), and (3) HMAC-signs the payload before POSTing to the webhook. The key authentication pattern:

# HMAC-SHA256 signing for webhook authentication
ts = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%S.000Z")
body = json.dumps(incident)
sig = hmac.new(webhook_secret.encode("utf-8"),
               f"{ts}:{body}".encode(), hashlib.sha256).digest()
http.request("POST", webhook_url, body=body,
             headers={"x-amzn-event-timestamp": ts,
                      "x-amzn-event-signature": base64.b64encode(sig).decode()})
Four parallel subagents querying ECS, CloudTrail, ALB, and Grafana data sources during the investigation

Figure 2: Four parallel subagents investigating ECS, CloudTrail, ALB, and Grafana data sources simultaneously

Root cause: Conflicting CPU and memory target-tracking autoscaling policies combined with an insufficient capacity floor. The service had both a CPU policy (target 70%) and a memory policy (target 75%). Actual memory usage sat at 3–8%, creating a persistent conflict between the two policies.

With MinCapacity set too low, the service could not sustain the task count needed to absorb CPU load. The resulting instability (22+ scaling flips observed) prevented stable scale-out, leaving the service effectively pinned at two tasks with no CPU headroom.

This is a common organizational issue: teams configure both scaling dimensions without realizing the interaction, especially when the capacity floor is not sized for baseline traffic. The agent identified the pattern in 10 minutes, a task that typically requires senior engineers with deep scaling expertise and hours of CloudWatch metric correlation.

Investigation output naming conflicting CPU and memory autoscaling policies as the root cause

Figure 3: Root cause analysis identifying the conflicting autoscaling policy

Slack incidents channel message linking to the running investigation at 1:22 AM

Figure 4: Slack notification with investigation link posted at 1:22 AM

Auto-created Jira ticket showing priority, root cause, and affected service

Figure 5: Jira ticket [redacted] auto-created with priority, root cause, and affected service

At 1:34 AM, the on-call engineer received a phone call through Grafana IRM with the complete investigation context. No need to wake up and hunt for root cause across dashboards.

Grafana IRM escalation chain routing the alert to the on-call engineer

Figure 6: Grafana IRM escalation chain routing the alert and calling the on-call engineer

Incoming on-call phone call at 1:34 AM carrying the investigation context

Figure 7: Incoming phone call at 1:34 AM with investigation context

Mitigation plan generated: (1) Remove the memory-based scaling policy, (2) raise MinCapacity to handle baseline traffic, (3) implement CPU-only target tracking at 70%. This plan was passed to a separate custom agent for remediation.

Remediation

Remediation is the key differentiator in this pipeline. It is a dedicated remediation agent (pr-creation-agent) invoked only after investigation completes. AWS DevOps Agent enforces read-only access to infrastructure through a per-session permission guardrail. Effective permissions are the intersection of the execution role’s IAM policy and the guardrail, and write actions are excluded.

The split separates concerns: investigation stays within that read-only envelope, whereas the remediation agent is scoped to a GitHub MCP server as its only external integration. Safety at the remediation layer does not rely on the agent’s built-in directed actions approval mechanism. Instead, two controls enforce the boundary. First, the remediation agent is a separate, narrowly scoped agent with access limited to GitHub MCP. Second, every output is a Draft pull request that requires human review and merge before taking effect. The GitHub MCP connection is authenticated with a fine-grained personal access token scoped to the specific infrastructure repositories, with an expiration and rotation policy. No elevated IAM role or additional agent permissions are required.

How it works

When the “Investigation Completed” Amazon EventBridge event fires, a Lambda orchestrator invokes the remediation agent with the investigation ID. The agent then:

  1. Reads findings from journal records to understand the root cause and recommended fix.
  2. Maps the AWS account to the correct repository. Property Finder has six infrastructure repos for different teams (B2B, B2C, core-platform, growth, data-engineering, shared-infra). The agent extracts the account ID from resource ARNs and routes to the right repo. This mapping is validated through automated tests and updated as new accounts or repositories are onboarded.
  3. Checks for duplicate PRs by searching existing PR titles and bodies for the investigation ID. If a matching PR exists, it reports the URL and exits without creating a duplicate.
  4. Reads the relevant Terraform files through GitHub MCP (GITHUB-MCP_get_file_contents), identifies the exact changes required, and plans the fix.
  5. Creates a feature branch (fix/{investigation_id}), commits the changes, and opens a Draft PR with a structured template including problem summary, root cause, changes made, and a testing checklist.

AWS also supports remediation through Kiro CLI with AWS CodeBuild or Kiro-ready prompts. Property Finder chose an approach that fits their multi-team repository structure: the remediation agent runs entirely within the Agent Space (the managed environment where custom agents execute), uses GitHub MCP for repository access, and maps multiple repositories to different teams automatically.

The orchestrator Lambda is triggered by the “Investigation Completed” Amazon EventBridge event. It first retrieves the investigation findings from journal records, then fans out to three targets simultaneously. Target one creates a Jira ticket with the full root cause analysis, severity, and affected service. Target two posts a Grafana Alerting-compatible payload to the Grafana IRM webhook to trigger phone call escalation. Target three invokes the remediation agent through the CreateChat and SendMessage API, passing the investigation ID and root cause context so it can generate the appropriate code fix.

Draft GitHub pull request with a problem summary, root cause, and changes template

Figure 8: GitHub PR [redacted] generated by the remediation agent with a structured problem, root cause, and changes template

Terraform diff replacing the memory scaling policy with a CPU-only target-tracking policy

Figure 9: Terraform diff showing the new CPU-only scaling policy replacing the conflicting memory configuration

The PR is always opened as Draft. Engineers review, run terraform plan, validate in staging, and merge. The agent never auto-merges.

Results

Metric Before After Improvement
End-to-end time Hours to days 14 minutes >88% reduction
Investigation 20 to 40 min (manual) 10 min (autonomous) 50–75% reduction
Documentation Manual, incomplete Auto-generated root cause analysis + Jira 100% documented
Remediation Manual PR by engineer Auto-fix PR + review Minutes to code fix

Cost considerations: Each incident invokes two agent sessions (investigation + remediation) with up to four parallel subagents. Billing is based on agent minutes. For detailed pricing, see the AWS DevOps Agent pricing page. We recommend reviewing pricing for all services used in this architecture.

“We now rely fully on AWS DevOps Agent to identify infrastructure-related issues. It has helped us identify multiple complex issues without even opening a support ticket. Even if we had raised tickets, it would likely have taken support engineers hours to find the root cause, whereas we resolved these issues in minutes.”

— Yasitha Bogamuwa, Cloud Engineering Manager, Property Finder

Getting started

Prerequisites:

  1. An Agent Space configured in your account.
  2. Amazon CloudWatch and AWS CloudTrail enabled for observability.
  3. Slack, Grafana, and GitHub connected as capabilities.
  4. Infrastructure resources tagged for topology mapping.

Step 1: Configure the webhook trigger. Set up CloudWatch Alarm action to invoke a Lambda function. The Lambda enriches the payload, HMAC-signs it, and POSTs to your Agent Space webhook endpoint.

Step 2: Set up event-driven outputs. Create an Amazon EventBridge rule for “Investigation Completed” events (source: aws.aidevops). Add Lambda targets for Jira, Grafana IRM, and optionally a remediation custom agent.

Step 3: Test end-to-end. Trigger a test alarm and verify the full pipeline: investigation starts, Slack posts, Jira ticket created, on-call paged, and PR opened.

For a similar integration pattern with Salesforce, see Automating Incident Investigation with AWS DevOps Agent and Salesforce MCP Server on the AWS DevOps Blog.

Clean up

This post describes an architecture pattern implemented by Property Finder. If you deployed test resources while following along, remember to delete any CloudWatch Alarms, Lambda functions, Amazon EventBridge rules, and Agent Space configurations to avoid ongoing charges. For a full list of resources and associated costs, review the pricing pages for each AWS service used in this architecture.

Conclusion

Property Finder’s implementation shows that autonomous incident management works in production today, with their pipeline running since early 2026. The agent never auto-merges. Human review remains in the loop by design: the agent accelerates, the engineer decides. The on-call engineer wakes up to a phone call with the root cause already identified, a Jira ticket filed, and a PR ready for review.

Explore the AWS DevOps Agent documentation to get started with your own autonomous pipeline.

  1. Getting Started with AWS DevOps Agent.
  2. Automating Incident Investigation with Salesforce MCP.
  3. Building an End-to-End Agentic SRE.
  4. Amazon EventBridge User Guide.
  5. Grafana IRM Documentation.

About the authors

Nada Tlohi

Nada Tlohi

Nada is a Technical Account Manager at AWS based in Dubai, UAE. She helps strategic enterprise customers across the MENA region transform their cloud operations and improve system reliability by adopting AIOps, incident automation, and DevOps best practices.

Conor Manton

Conor Manton

Conor is a Principal Technical Account Manager at AWS, based in San Francisco. He works with strategic enterprise customers to accelerate their cloud journey, with a focus to operationalize AI-powered workflows to drive business outcomes.

Jaydeep Singh

Jaydeep Singh

Jaydeep is a Senior DevOps Engineer at Property Finder. He specializes in designing and operating scalable cloud infrastructure, containerized platforms, and Kubernetes ecosystems. He leads platform reliability, infrastructure automation, and continuous integration and continuous delivery (CI/CD) initiatives, so engineering teams can build and deploy applications securely, efficiently, and at scale.

Isolate email reputation in Amazon SES Mail Manager with tenant management

Post Syndicated from Abilashkumar P C original https://aws.amazon.com/blogs/messaging-and-targeting/isolate-email-reputation-in-amazon-ses-mail-manager-with-tenant-management/

When AnyCompany’s new IT outsource team misconfigured the email settings on 200 of the company’s multifunction printer/scanners, it had two bad outcomes. First, nobody received their scanned documents in their inboxes. Somewhat predictably, many users rescanned the same documents multiple times before creating support tickets. Second, the misconfiguration along with the multiple failed attempts resulted in a “bounce storm” that was quickly reported by a major email service provider, but unfortunately ignored by the IT team.

Within 48 hours, the bounce rate crossed the provider’s threshold. The company’s entire Amazon Simple Email Service (Amazon SES) account lost its sending reputation. Password resets, order confirmations, and service notifications from every business unit on the account started landing in spam or failing to deliver. The damage spread because every sender on the account, from the mission-critical billing system to the misconfigured printers, shared the same reputation score.

This is a preventable problem. With Amazon SES tenant management you can isolate email reputation per tenant inside a single account so one misbehaving sender cannot affect the rest. If you use Amazon SES Mail Manager for Simple Mail Transfer Protocol (SMTP) filtering, routing, archiving, or relay, you can activate tenant isolation. To do so, tag each message with the X-SES-TENANT header in your Mail Manager rule set. In this post, you will compare five architectural patterns for applying the X-SES-TENANT header, from static per-tenant endpoints to AWS Lambda driven runtime resolution.

This post complements Isolate email suppression per tenant with Amazon SES. That post explains how tenant-level suppression lists prevent cross-tenant bounce and complaint contamination, which is the “what happens after the message is tagged” story. This post focuses on the upstream problem: how to get the X-SES-TENANT tag onto messages when your senders are legacy appliances, printers, or applications that can’t set custom MIME headers. Together, the two posts cover the full tenant isolation pipeline, from tagging through delivery and suppression.

This post provides architectural guidance. For step-by-step implementation, refer to the Amazon SES documentation.

How SES tenant isolation works

Amazon SES tenant management isolates reputation per tenant inside a single Amazon SES account. Each tenant acts as a container organized around sending identities, configuration sets, and the resulting reputation metrics. Amazon SES attributes bounces, complaints, and Trust and Safety signals to the tenant, not the account, so a deliverability issue in one tenant doesn’t affect the others.

A critical benefit of tenant isolation: when one tenant’s reputation degrades beyond a threshold, Amazon SES can pause sending for that tenant only. Other tenants continue delivering normally. Without tenant isolation, a reputation issue affects the entire account. This pause-and-contain mechanism is one of the strongest reasons to adopt tenant management, especially for accounts with diverse sender types.

You associate a message with a tenant by passing the TenantName parameter on the Amazon SES API v2 SendEmail operation, or by adding an X-SES-TENANT Multipurpose Internet Mail Extensions (MIME) header to an SMTP message. For a detailed walkthrough of tenant management concepts, including identity ownership, the ses:TenantName AWS Identity and Access Management (IAM) condition key, and tenant-level suppression lists, see Improve email deliverability with tenant management in Amazon SES.

How Mail Manager works

Mail Manager processes inbound and outbound SMTP traffic through a pipeline of three components:

  1. Ingress endpoint: an authenticated SMTP endpoint that accepts connections from your senders. Mail Manager ingress endpoints handle SMTP only, not the Amazon SES API.
  2. Traffic policy: filters connections based on sender attributes (IP, TLS version, authentication) before messages reach rule processing.
  3. Rule set: an ordered list of rules. Each rule has conditions (match on envelope sender, recipient, source IP, or header values) and actions (Add header, Write to S3, Invoke Lambda, Send to internet, SMTP relay, Drop).

The “Add header” rule action is what makes tenant isolation possible for legacy senders: it injects the X-SES-TENANT SMTP header before the “Send to internet” action hands the message to Amazon SES for delivery.

With the “Add header” rule action inserted before the “Send to internet” action in the same rule, Mail Manager effectively tags the message with the SMTP header that defines the tenant. When Amazon SES processes the send, it reads the X-SES-TENANT header and attributes the message to the corresponding tenant.

Amazon SES performs tenant attribution only during send processing. A Send to internet action, or a Lambda function that calls SendEmail with the TenantName parameter or X-SES-TENANT header, activates tenant management. An SMTP relay action forwards to a third-party SMTP server (Google Workspace, Microsoft 365, or on-premises mail), so Amazon SES doesn’t process the send and tenant attribution doesn’t apply. Write to S3 and Drop don’t hand messages to Amazon SES, so they don’t activate tenant management either. This post describes flows that include a Send to internet action or a Lambda function calling SendEmail.

Understanding the outbound email flow

An outbound message flows from the SMTP client to the Mail Manager ingress endpoint, passes through the traffic policy and rule set, then routes through Amazon SES to the internet.

Figure 1: Outbound email flow from an SMTP client through Mail Manager to Amazon SES

Compare the patterns

Before diving into each pattern, use this table to identify which one fits your workload. You can then read only the pattern section that applies, or read all five for the full picture.

Consideration Pattern 1 Pattern 2 Pattern 3 Pattern 4 Pattern 5
Works for legacy and appliance senders — Yes Yes Yes Yes
Retrieve tenant from static value Yes Yes Yes Yes Yes
Retrieve tenant from source IP or sender condition — Yes Yes Yes Yes
Retrieve tenant from runtime lookup or body inspection — — — Yes Yes
Records Send in Mail Manager log Yes Yes Yes — —
Tenants per Region Up to 10,000 ~50 400 (per-tenant Send) or 1,560 (chained) Up to 10,000 Up to 10,000

One difference cuts across the patterns: where the tenant mapping lives determines what it takes to change it. Patterns 2 and 3 hold the mapping in rule-set configuration, so adding or removing a tenant is a rule-set edit and deployment (a control-plane change, not a data change). Patterns 4 and 5 resolve the tenant from a runtime source such as a database, so onboarding or offboarding a tenant is a data update that takes effect without a deployment. In Pattern 1, the sender supplies the tenant, so there’s no mapping to maintain in Mail Manager at all.

Pattern 1: The SMTP sender sets the header before Mail Manager

Pattern 1, where the SMTP sender sets the X-SES-TENANT header before the message reaches the Mail Manager ingress endpoint

Figure 2: Pattern 1, where the SMTP sender sets the tenant header before Mail Manager

If the SMTP sender (a backend service, internal tool, or any application that can add a custom MIME header) sets X-SES-TENANT on the message before connecting to the Mail Manager ingress endpoint, the message arrives pre-tagged. The rule set only needs a Send to internet action.

Pattern 1 fits customers who already use Mail Manager for filtering, archiving, or compliance and whose sending applications can add one header at send time. You keep Mail Manager gateway capabilities without adding Add header or conditional logic to the rule set.

Pattern 2: Mail Manager adds a static header with Add header

Pattern 2, where a Mail Manager rule adds a static X-SES-TENANT header and then sends the message to the internet

Figure 3: Pattern 2, where a Mail Manager rule adds a static tenant header

A rule with Add header followed by Send to internet attaches a fixed tenant value to each message. This pattern fits a one-tenant-per-endpoint model: provision one authenticated ingress endpoint per tenant, give each tenant its own SMTP credentials, and attach a rule set that injects the tenant value.

For example, an enterprise provisions one endpoint for facilities-printer notifications and a second for corporate alerts. The Send to internet action’s IAM role grants permission only to that tenant’s Amazon SES identities, preventing a misrouted client from sending as another tenant.

You can group tenants behind one endpoint when they share a sending configuration. The header value and IAM scope live in the rule-set configuration, and no code runs at send time.

Pattern 3: Mail Manager derives the header from rule conditions

If multiple tenants share an endpoint but have stable distinguishing attributes (like source IP), one rule set handles each of them. Rule conditions match on envelope properties, and matching rules run an Add header action that sets X-SES-TENANT to the correct value.

Pattern 3, where Mail Manager derives the X-SES-TENANT header value from rule conditions before sending to the internet

Figure 4: Pattern 3, where Mail Manager derives the tenant header from rule conditions

Mail Manager rule sets allow 40 rules with up to 10 conditions and 10 actions per rule, but caps Send to internet and SMTP relay actions at 10 per rule set (counting every occurrence). One Send to internet per tenant rule tops out at 10 tenants.

To support more tenants, separate header-setting from delivery:

  • Rules 1 to 39: Each matches a distinguishing condition and runs a single Add header action.
  • Rule 40: A catch-all with no conditions and a single Send to internet action.

Each message matches at most one header-setting rule, picks up its tenant header, and passes through the catch-all. The effective ceiling is now 39 tenants per rule set with one Send to internet action and one IAM role.

Scale limits of Pattern 3

Pattern 3’s ceiling depends on how you structure the rule set. Two cases:

Case A: Chained structure (39 Add header rules + 1 Send to internet rule): Each rule set uses one Send to internet action, so the 10-action cap isn’t binding. Capacity is 39 tenants per rule set × 40 rule sets per Region = 1,560 tenants per Region.

Case B: Per-tenant Send to internet (each tenant rule has its own Send action): The 10-action cap binds at 10 tenants per rule set. Capacity is 10 tenants per rule set × 40 rule sets per Region = 400 tenants per Region.

The two cases trade off scale against IAM scoping. Case A shares one IAM role across all tenants in the rule set. Case B gives each tenant its own IAM role at the cost of 4× fewer tenants.

Amazon SES supports up to 10,000 tenants per account (adjustable). Workloads that exceed a few hundred tenants, or need runtime tenant changes, can use Pattern 4 or Pattern 5.

Pattern 4: Mail Manager calls Lambda for runtime tenant resolution

Pattern 4, where Mail Manager writes the message to Amazon S3 and invokes a Lambda function that resolves the tenant and delivers through Amazon SES

Figure 5: Pattern 4, where Mail Manager invokes a Lambda function for runtime tenant resolution

Some tenant values require runtime logic, such as a database lookup on the sender IP, an external policy service, or content inspection. For these cases, the Mail Manager Invoke Lambda action runs a Lambda function inside the rule chain.

The Lambda event carries only metadata (headers, envelope sender, recipients, verdicts), not the MIME body. The function also can’t modify the message for downstream actions. Lambda must therefore handle delivery.

The rule writes the raw MIME to Amazon S3 with Write to S3, then invokes the Lambda function with the message ID. The function fetches the object and determines the tenant through the runtime logic your workload requires. That logic might be a database lookup (for example, an Amazon DynamoDB query), a call to an external policy service, or inspection of the message body. It then calls the Amazon SES API v2 SendEmail operation, passing the resolved tenant in the TenantName parameter. Delivery permissions live on the function’s execution role, which carries the ses:TenantName condition key.

The Lambda function is yours to build and maintain. This gives you full control over the tenant resolution logic and everything downstream (retries, dead-letter queues, observability), but it also means you own the operational overhead: code updates, monitoring, and cost management.

Mail Manager can invoke the function synchronously or asynchronously. Synchronous invocation (REQUEST_RESPONSE) keeps Lambda in Mail Manager’s critical path: Mail Manager waits up to 30 seconds for the function to return, and retries on failure. Asynchronous invocation (EVENT) hands control to Lambda instead, so Mail Manager invokes the function and moves on. There are no additional Mail Manager charges for the Lambda invocation beyond standard Lambda pricing.

Pattern 5: Mail Manager stages to Amazon S3, Lambda delivers asynchronously

Pattern 5, where an Amazon S3 event triggers a Lambda function that delivers the message through Amazon SES

Figure 6: Pattern 5, where an Amazon S3 event triggers a Lambda function that delivers the message

In Pattern 4, Mail Manager invokes the function directly through the Invoke Lambda rule action. Pattern 5 removes that direct invocation: the Mail Manager rule ends at Write to S3, and an Amazon S3 event notification triggers the Lambda function instead. Mail Manager’s work finishes at the write, and delivery becomes fully event-driven.

The rule set has two actions: write the raw MIME to Amazon S3, followed by an explicit Drop action. The Drop action prevents accidental duplicate delivery if a Send to internet action is inadvertently added to the rule later. The Lambda function handles delivery through the Amazon SES API, so Mail Manager’s job ends at writing the MIME to Amazon S3. The Amazon S3 event routes to the function directly or through Amazon Simple Queue Service (Amazon SQS) or Amazon EventBridge for fan-out and back-pressure.

The function reads the object, performs the tenant lookup, and calls SendEmail with the TenantName parameter. The Mail Manager critical path is minimal, and Lambda retries use the Lambda retry model with dead-letter queue support. The same Amazon S3 object fans out to multiple consumers (delivery, analytics) without changing the Mail Manager rule.

The Lambda function is yours to build and maintain. The upside is full control over the function and everything after it: tenant resolution logic, retries, dead-letter queues, and observability. The tradeoff is cost and upkeep, since you own code updates, monitoring, and operational overhead.

The other tradeoff is less visibility. After Write to S3, the Mail Manager log no longer records the delivery outcome.

Secure tenant attribution with IAM

Regardless of which pattern sets the X-SES-TENANT header, the Send to internet action’s IAM role should enforce tenant boundaries. Scope the IAM role with a Condition element that includes the ses:TenantName condition key.

Example IAM policy:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "ses:SendEmail",
      "Resource": "*",
      "Condition": {
        "StringEquals": {
          "ses:TenantName": "facilities-printers"
        }
      }
    }
  ]
}

This policy allows the role to send email only when the message is attributed to the facilities-printers tenant. Messages tagged with any other tenant value, or messages with no tenant header, are denied.

In Pattern 1, the sender sets the header, the IAM role on the Send to internet action validates that the claimed tenant matches the role’s permissions. In Pattern 2, the Add header action sets a fixed value, and the IAM role confirms the header matches the expected tenant for that endpoint. In Pattern 3 with a chained structure, a single Send to internet action services all tenants. Scope its role to the set of valid tenant names so untagged messages (those matching no Add header rule) fail authorization. For Patterns 4 and 5, the Lambda function’s execution role carries the ses:TenantName condition key, providing the same enforcement at the API call level.

Paused tenants

Each of the five patterns handles paused tenants the same way. When Amazon SES pauses a tenant (through a reputation policy or manually), sends for that tenant fail with a rejection error. Other tenants keep delivering. The failure surfaces depending on the pattern:

  • Patterns 1 to 3: Mail Manager records the rejection in the rule set log.
  • Patterns 4 and 5: The rejection surfaces in the Lambda function’s Amazon CloudWatch Logs.
  • Patterns 1 to 5: Amazon SES publishes tenant status changes to Amazon EventBridge (such as Sending Status Disabled).

Mail Manager won’t re-route or retry a paused tenant send. Graceful handling (queueing, failover, notification) belongs in the Lambda function in Patterns 4 and 5.

Observability

Observability for these patterns draws on three sources, each answering a different question:

Mail Manager vended log: which rule actions ran, and whether Amazon SES accepted the message from a Send to internet action. Mail Manager delivers this log to a destination you configure: Amazon CloudWatch Logs, Amazon S3, or Amazon Data Firehose. Query CloudWatch Logs with CloudWatch Logs Insights, or query Amazon S3 with Amazon Athena to surface IAM denials, configuration errors, and throttling.

Amazon SES event publishing: the final delivery outcome (delivered, bounced, or complaint), routed through a configuration set. This applies to every pattern.

Lambda Amazon CloudWatch Logs: for Patterns 4 and 5, where delivery runs inside the Lambda function, the acceptance result and any application errors.

To trace a message end to end, correlate these sources. For Patterns 1 to 3, the Mail Manager log and Amazon SES event publishing cover the flow. For Patterns 4 and 5, add the Lambda function’s CloudWatch Logs, since the Mail Manager log ends at Invoke Lambda (Pattern 4) or Write to S3 (Pattern 5).

Limits that shape the architecture

Review the Amazon SES Mail Manager service quotas before committing to a pattern. These quotas most often drive your pattern choice:

Resource Default Where it matters
Maximum message size (SMTP ingress) 40 MB Patterns 1 to 5
Authenticated ingress endpoints per Region 50 Pattern 2 per-tenant endpoints
Rule sets per Region 40 Pattern 2, Pattern 3 partitioning
Rules per rule set 40 Pattern 3
Send to internet action per rule set 10 Pattern 3 tightest constraint
Actions per rule 10
Conditions per rule 10
Addresses per address list 100,000 Pattern 3 consolidation
Tenants per account (Amazon SES) 10,000 (adjustable) Patterns 4 and 5 ceiling
Lambda concurrent executions per Region 1,000 (adjustable) Patterns 4 and 5 throughput ceiling
Lambda timeout (Mail Manager InvokeLambda) 30 seconds Pattern 4 synchronous path
S3 event notification destinations per prefix 1 (use Amazon EventBridge for fan-out) Pattern 5 fan-out design
Lambda invocation payload (synchronous) 6 MB Pattern 4 metadata-only (body in S3)
Sending quota per 24 hours (Amazon SES) 200 in sandbox (adjustable in production) Patterns 1 to 5
Maximum send rate (Amazon SES) 1 message/second in sandbox (adjustable in production) Patterns 1 to 5

Conclusion

The five patterns in this post show how to architect tenant tagging, whether through static endpoints, rule-set headers, or runtime resolution, so you can choose the approach that fits your workload.

Next steps

About the authors

Getting started with Apache Iceberg write support in Amazon Redshift – Part 3

Post Syndicated from Raghu Kuppala original https://aws.amazon.com/blogs/big-data/getting-started-with-apache-iceberg-write-support-in-amazon-redshift-part-3/

Production data is always evolving. Tables gain and lose columns, outgrow their data types, and get re-partitioned as query patterns shift. Multiple engines often need to read the same data. These changes used to mean expensive data rewrites or rebuilt pipelines. Apache Iceberg makes them metadata-only operations, and Amazon Redshift now supports evolving schemas and partitioning layouts through ALTER statements, with no data rewrites and no pipeline rebuilds. You can also create AWS Lake Formation resource links in the catalog of Amazon S3 Tables, a capability of Amazon Simple Storage Service (Amazon S3), for centralized cross-engine governance.

In Part 1, you created Apache Iceberg tables and wrote data directly from Amazon Redshift to your data lake, setting up external schemas, creating tables in both Amazon Simple Storage Service (Amazon S3) and Amazon S3 Tables, and performing INSERT operations with full ACID (Atomicity, Consistency, Isolation, Durability) compliance. In Part 2, you performed DELETE, UPDATE, and MERGE operations to modify data at the row level and synchronize staging and production tables.

In this post, you use the customer and orders datasets from the previous posts to evolve Iceberg table schemas and partitioning with ALTER operations. You also create an AWS Lake Formation resource link in the S3 Tables catalog to share tables with other analytics engines under a single, centralized permission model.

Solution overview

This solution demonstrates ALTER operations for Apache Iceberg tables in Amazon Redshift and Lake Formation resource link creation for the S3 Tables catalog. The walkthrough includes the following key operations:

  • ALTER TABLE RENAME COLUMN – Rename existing columns without changing data types or partition specs.
  • ALTER TABLE ADD/DROP COLUMN – Add new columns or remove existing columns as metadata-only operations.
  • ALTER TABLE ALTER COLUMN – Widen column data types (for example, INT to BIGINT) without rewriting data.
  • ALTER TABLE SET TABLE PROPERTIES – Change compression type for future writes.
  • ALTER TABLE ADD/DROP/REPLACE PARTITION FIELD – Evolve partition specs without re-partitioning existing data.
  • Lake Formation resource link – Create a resource link in the S3 Tables catalog for centralized access governance.

The following diagram shows the end-to-end architecture:

Architecture diagram of Amazon Redshift running ALTER operations on Iceberg tables in S3 Tables, with Lake Formation resource links providing access from Amazon Athena and other engines

Figure 1: Architecture showing Amazon Redshift performing ALTER operations on Iceberg tables in S3 Tables, with Lake Formation resource links providing access from Amazon Athena and other engines

Prerequisites

Complete the setup from Part 1 and Part 2, including:

  • An Amazon Redshift data warehouse (provisioned or Serverless) on patch 201 or higher.
  • The AWS Identity and Access Management (IAM) role (RedshifticebergRole) with permissions for Amazon S3, AWS Glue Data Catalog, and Lake Formation.
  • The customer table in a standard Amazon S3 bucket (AWS Glue catalog: customer_db).
  • The orders table in an Amazon S3 table bucket (iceberg-write-blog@s3tablescatalog).
  • Access to an IAM role that is a Lake Formation data lake administrator.
  • AWS Glue Data Catalog integrated with S3 Tables (s3tablescatalog exists).

Schema evolution with ALTER TABLE

With ALTER TABLE, you can change Iceberg table definitions, including schema, partition specs, and properties, without rewriting stored data. Each operation updates only metadata. The table structure changes instantly while existing data files remain untouched. This helps make schema evolution, partition adjustments, and property updates safe to run on production tables.

Add a column

You can add a new column to an Iceberg table using ALTER TABLE. Each new column is added with a unique field ID that Iceberg uses for column tracking across schema evolution. Existing rows return NULL for the newly added column.

Verify the current schema:

SHOW TABLE dev.demo_iceberg.customer;
SHOW TABLE output listing the current columns of the customer table

Figure 2: SHOW TABLE output showing the current customer table schema

Add the column:

-- Add a loyalty_tier column to the customer table
ALTER TABLE dev.demo_iceberg.customer
ADD COLUMN loyalty_tier VARCHAR;

Verify the schema change:

SHOW TABLE dev.demo_iceberg.customer;
SHOW TABLE output showing the new loyalty_tier column added to the customer schema

Figure 3: SHOW TABLE output showing the loyalty_tier column added to the schema

The following output shows the new loyalty_tier column as NULL for existing rows:

SELECT customer_id, customer_name, city, loyalty_tier
FROM dev.demo_iceberg.customer
ORDER BY customer_id;
Query results showing loyalty_tier as NULL for existing customer rows

Figure 4: Query results showing loyalty_tier as NULL for existing rows

Populate the new column by aggregating order totals from the orders table in S3 Tables:

-- Set loyalty_tier based on total spend from orders
UPDATE dev.demo_iceberg.customer
SET loyalty_tier = CASE
WHEN a.total_spend > 300 THEN 'Gold'
ELSE 'Silver'
END
FROM (
SELECT customer_id, SUM(total_order_amt) AS total_spend
FROM "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders
GROUP BY customer_id
) AS a
WHERE dev.demo_iceberg.customer.customer_id = a.customer_id;

The following output shows customer loyalty tiers after the update:

SELECT customer_id, customer_name, loyalty_tier
FROM dev.demo_iceberg.customer
ORDER BY customer_id;
Customer table query results showing Gold and Silver loyalty tiers

Figure 5: Customer table showing Gold and Silver loyalty tiers

Note: Customer IDs 11, 13, and 15 show NULL for loyalty_tier because they have no matching orders in the orders table.

Drop a column

Remove columns that are no longer needed. The column is removed from the current schema, but data in existing files remains untouched and simply becomes invisible to queries.

Verify the current schema:

SHOW TABLE dev.demo_iceberg.customer;
SHOW TABLE output showing the customer schema before dropping loyalty_tier

Figure 6: SHOW TABLE output showing the current customer table schema before dropping loyalty_tier

Drop the column:

-- Drop the loyalty_tier column
ALTER TABLE dev.demo_iceberg.customer
DROP COLUMN loyalty_tier;

Verify the schema change:

SHOW TABLE dev.demo_iceberg.customer;
SHOW TABLE output showing the customer schema after loyalty_tier is dropped

Figure 7: SHOW TABLE output showing the customer table schema after loyalty_tier is dropped

Verify the column is dropped:

SELECT * FROM dev.demo_iceberg.customer
ORDER BY customer_id;
Query results confirming the loyalty_tier column no longer appears

Figure 8: Query results confirming the loyalty_tier column has been dropped

Note: To drop a column used in the current partition spec, first drop or replace the partition field, then drop the column.

Rename a column

Rename a column without affecting data types or partition specs:

-- Rename city to location
ALTER TABLE dev.demo_iceberg.customer
RENAME COLUMN city TO location;

The following output confirms the column has been renamed to location:

SELECT customer_id, customer_name, location
FROM dev.demo_iceberg.customer
ORDER BY customer_id;
Query results showing the city column renamed to location

Figure 9: Query results showing the renamed column location

Widen a column type

Widen a column’s data type without rewriting data. This is useful when your data outgrows the original precision, for example when order amounts exceed the original decimal range.

Verify the current column type:

SHOW TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders;
SHOW TABLE output showing total_order_amt as DECIMAL(10,2)

Figure 10: SHOW TABLE output showing total_order_amt as DECIMAL(10,2)

Now run the ALTER to widen the column:

-- Widen total_order_amt to support larger order values
ALTER TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders
ALTER COLUMN total_order_amt TYPE DECIMAL(18,2);

Verify the updated column type:

SHOW TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders;
SHOW TABLE output showing total_order_amt widened to DECIMAL(18,2)

Figure 11: SHOW TABLE output confirming total_order_amt widened to DECIMAL(18,2)

Note: Amazon Redshift supports safe type promotions (for example, INT to BIGINT, FLOAT to DOUBLE, DECIMAL(10,2) to DECIMAL(18,2)). Plan column types accordingly for future growth.

Set table properties

Change the compression type for future writes:

Verify the current compression type:

SHOW TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders;
SHOW TABLE output showing the orders table compression type before the change

Figure 12: SHOW TABLE output showing the current compression type before the update

Now run the ALTER to change the compression type:

-- Switch to zstd compression for better ratios
ALTER TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders
SET TABLE PROPERTIES ('compression_type'='zstd');

The following SHOW TABLE output confirms the updated compression setting:

SHOW TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders;
SHOW TABLE output showing compression_type set to zstd

Figure 13: SHOW TABLE output showing compression_type set to zstd

Note: This affects only future writes. Existing data files retain their original compression.

Partition evolution

A powerful feature of Iceberg is partition evolution, the ability to change how a table is partitioned without rewriting existing data. Amazon Redshift writes new data with the updated partition scheme, while existing data remains in the old layout. Query engines handle both layouts transparently.

Adding a partition field

The orders table from Part 1 is partitioned by DAY(order_date). Add an additional bucket partition to distribute data across hash buckets:

Verify the current partition spec:

SHOW TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders;
SHOW TABLE output showing the orders partition spec before adding a field

Figure 14: SHOW TABLE output showing the current partition spec before adding a partition field

Add the partition field:

-- Add bucket partitioning on customer_id
ALTER TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders
ADD PARTITION FIELD bucket(16, customer_id);

After this change, new data is partitioned by both DAY(order_date) and bucket(16, customer_id), while existing data remains in the original day-only layout.

Verify the updated spec:

SHOW TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders;
SHOW TABLE output showing partition spec with DAY(order_date) and bucket(16, customer_id)

Figure 15: SHOW TABLE output showing the updated partition spec with DAY(order_date) and bucket(16, customer_id)

Replacing a partition field

Instead of separately dropping and adding, use REPLACE PARTITION FIELD as a single atomic operation. This is the recommended approach when swapping one transform for another on the same source column, because it makes the intent explicit and avoids a transient state where the table is unpartitioned between operations.

Verify the current partition spec:

SHOW TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders;
SHOW TABLE output showing the current DAY(order_date) partition spec

Figure 16: SHOW TABLE output showing the current partition spec with DAY(order_date) and bucket(16, customer_id)

Replace the partition field:

-- Replace daily partitioning with monthly
ALTER TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders
REPLACE PARTITION FIELD DAY(order_date) WITH MONTH(order_date);

After this change:

  • Existing data remains in day-based partition folders.
  • Amazon Redshift writes new data into month-based partition folders.
  • The query engine reads both layouts transparently.

Confirm the new partition spec:

SHOW TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders;
SHOW TABLE output confirming the partition field replaced with MONTH(order_date)

Figure 17: SHOW TABLE output confirming the partition field replaced with MONTH(order_date)

Insert new data and verify that both partition layouts are queryable:

-- New data follows monthly partitioning
INSERT INTO "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders
(order_date, order_id, customer_id, total_order_amt, total_order_tax_amt,
tax_pct, order_created_at_tz, is_active_ind)
VALUES
('2025-01-15', 1018, 3, 210.00, 16.80, 0.08, '2025-01-15 09:00:00-06:00', true);
-- Query spans both old (daily) and new (monthly) layouts transparently
SELECT order_id, order_date, total_order_amt
FROM "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders
WHERE order_date >= '2024-11-01'
ORDER BY order_date;
Query results spanning both the daily and monthly partition layouts

Figure 18: Query results spanning both partition layouts

Converting to a multi-level partition

Iceberg supports multi-level (composite) partition specs, where data is organized by more than one partition field. You can evolve an existing single-level spec into a multi-level spec by adding partition fields one at a time. Each ADD PARTITION FIELD is a lightweight metadata operation, and no data is rewritten.

The orders table is currently partitioned by MONTH(order_date) and bucket(16, customer_id). Add one more partition field to create a three-level spec:

Verify the current partition spec:

SHOW TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders;
SHOW TABLE output showing the two-level partition spec of MONTH(order_date) and bucket(16, customer_id)

Figure 19: SHOW TABLE output showing the current two-level partition spec of MONTH(order_date) and bucket(16, customer_id)

Add partition field to build the three-level spec:

-- Add a day-level partition field on order_created_at_tz
ALTER TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders
ADD PARTITION FIELD day(order_created_at_tz);

Verify the new multi-level partition spec:

SHOW TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders;
SHOW TABLE output showing the three-level MONTH, bucket, and day partition spec

Figure 20: SHOW TABLE output showing the three-level partition spec of MONTH(order_date), bucket(16, customer_id), and day(order_created_at_tz)

After these changes:

  • Existing data remains in the original single-level layout (month-based folders).
  • Amazon Redshift writes new data into the multi-level layout (month, then bucket, then day folders).
  • The query engine reads both layouts transparently.

Dropping partition fields from a multi-level partition

You can also evolve in the other direction by removing partition fields from a multi-level spec to simplify the partition layout. Like adding fields, dropping a partition field is a metadata-only operation and removes one field per statement.

Verify the current multi-level partition spec:

SHOW TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders;
SHOW TABLE output showing the three-level partition spec before dropping fields

Figure 21: SHOW TABLE output showing the three-level partition spec before dropping fields

Drop the partition fields one at a time:

-- Drop the bucket partition field
ALTER TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders
DROP PARTITION FIELD bucket(16, customer_id);
-- Drop the day-level partition field
ALTER TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders
DROP PARTITION FIELD day(order_created_at_tz);

Verify the table is back to its original single-level spec:

SHOW TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders;
SHOW TABLE output confirming the table back to a single-level MONTH(order_date) spec

Figure 22: SHOW TABLE output confirming the table is back to a single-level MONTH(order_date) partition spec

After dropping a partition field:

  • Data written under the dropped field’s layout stays in place and remains queryable.
  • Amazon Redshift writes new data using only the remaining partition fields.
  • Queries that filtered on the dropped field still work, but they no longer benefit from partition pruning on that field for newly written data.

Supported partition transforms

The following table lists the partition transforms available for Iceberg tables in Amazon Redshift:

Partition transform Syntax example What it does
Year year(order_date) Groups data into yearly partitions based on a date or timestamp column.
Month month(order_date) Groups data into monthly partitions based on a date or timestamp column.
Day day(order_date) Groups data into daily partitions based on a date or timestamp column.
Hour hour(event_ts) Groups data into hourly partitions based on a timestamp column.
Bucket bucket(16, customer_id) Distributes data across N hash buckets for even distribution on high-cardinality columns.
Truncate truncate(3, zip_code) Truncates column values to a fixed width W for grouping similar values together.
Identity identity(region) Partitions by the exact column value with no transformation applied.

Note: A column that is already part of an existing partition field can’t be used in a new partition field. Drop or replace the existing field first.

Accessing S3 Tables with external schemas

Lake Formation resource links provide cross-engine access to your S3 Tables through centralized governance. You create a resource link in the default AWS Glue Data Catalog that points to your S3 Tables database. Amazon Redshift, Amazon Athena, Amazon EMR, and other engines can then discover and query the tables using a single permission model.

Diagram of S3 Tables integration with AWS Glue Data Catalog and Lake Formation

Figure 23: S3 Tables integration with AWS Glue Data Catalog and Lake Formation

For the complete setup walkthrough, including Lake Formation prerequisites, resource link creation, and permission grants, see Optimize Amazon S3 Tables queries with Amazon Redshift. For conceptual details on resource links and S3 Tables catalog integration, see About resource links and Creating an S3 Tables catalog.

The following steps show how to query S3 Tables through a resource link after completing the setup from the referenced blog.

In the Lake Formation console, the resource link appears as a database named iceberg_write_blog_rl (type: Resource link). To grant access to the resource link:

  1. In the Lake Formation console, choose Databases.
  2. Locate iceberg_write_blog_rl (type: Resource link).
  3. Choose Actions, then Grant.
  4. Grant DESCRIBE permission to RedshiftIcebergRole.

Create an external schema

With the resource link in place, create an external schema in Amazon Redshift for two-part notation access.

For IAM federated users:

CREATE EXTERNAL SCHEMA s3tables_iceberg
FROM DATA CATALOG
DATABASE 'iceberg_write_blog_rl'
CATALOG_ID '<ACCOUNT_ID>'
IAM_ROLE 'SESSION';

For database users and business intelligence (BI) tools:

CREATE EXTERNAL SCHEMA s3tables_iceberg
FROM DATA CATALOG
DATABASE 'iceberg_write_blog_rl'
IAM_ROLE 'arn:aws:iam::<ACCOUNT>:role/RedshifticebergRole';

Grant access to specific users or roles:

-- Grant to the IAM role used in this walkthrough
GRANT USAGE ON SCHEMA s3tables_iceberg TO "IAMR:RedshifticebergRole";
Amazon Redshift query showing S3 Tables available through the external schema

Figure 24: S3 Tables available through an external schema

Query with two-part notation

With the external schema created, query S3 Tables using two-part notation:

-- Instead of: "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders
SELECT * FROM s3tables_iceberg.orders;
Query results from S3 Tables through the external schema using two-part notation

Figure 25: Query results from S3 Tables through an external schema using two-part notation

Access methods comparison

The following table compares the available methods for accessing Iceberg tables in Amazon Redshift:

Access method Query syntax Authentication Best for
S3 Tables three-part notation "bucket@s3tablescatalog".namespace.table IAM federated identity only Interactive queries in Query Editor v2 with direct catalog access.
External schema through resource link schema_name.table Any (IAM role defined in schema) BI tools, Data API, JDBC/ODBC applications, and shared team access.
awsdatacatalog awsdatacatalog.database.table IAM federated identity only Multi-database access in a single session without creating external schemas.

Bringing it together

Combine schema evolution with cross-engine access in a single workflow. The following example adds a column to the orders table and immediately queries it through the external schema:

-- 1. Add a column to the S3 Tables orders table
ALTER TABLE s3tables_iceberg.orders
ADD COLUMN fulfillment_status VARCHAR;
-- 2. Update the new column
UPDATE s3tables_iceberg.orders
SET fulfillment_status = 'shipped'
WHERE order_date < '2024-11-01';
UPDATE s3tables_iceberg.orders
SET fulfillment_status = 'pending'
WHERE order_date >= '2024-11-01';
-- 3. Query immediately via the external schema (no schema recreation needed)
SELECT o.order_id, o.order_date, o.fulfillment_status, c.customer_name
FROM s3tables_iceberg.orders o JOIN demo_iceberg.customer c
ON o.customer_id = c.customer_id
ORDER BY o.order_date DESC;
Query results of a cross-catalog join showing the evolved schema through the external schema

Figure 26: Cross-catalog join showing the evolved schema immediately visible through the external schema

The new column is visible through both the three-part notation and the external schema without any additional configuration, because the schema evolution in Iceberg propagates automatically.

Best practices

  • Test ALTER operations in non-production first. While metadata-only, schema changes affect all readers immediately.
  • Use REPLACE PARTITION FIELD instead of DROP + ADD. The atomic operation avoids a transient unpartitioned state.
  • Monitor partition spec changes with SHOW TABLE. Verify the current spec after any partition evolution.
  • Choose partition transforms based on query patterns. Use month() or day() for time-range filters. Use bucket() for high-cardinality join keys.
  • Set table properties before bulk loads. Change compression type (zstd for better ratios, snappy for speed) before large INSERT operations.
  • Run table maintenance after mutations. After performing multiple UPDATE, DELETE, or MERGE operations, run AWS Glue table optimizers to compact deletion files and improve read performance.
  • Use Lake Formation for fine-grained access. Column-level and row-level security can be applied through Lake Formation on tables accessed through resource links.
  • Grant schema access to specific users or roles. Avoid granting to PUBLIC. Use named IAM roles or database users for least-privilege access.
  • Monitor query performance. Use Amazon Redshift query monitoring features to track performance of write operations and optimize partitioning strategies as needed.

Considerations

Keep the following in mind when working with ALTER TABLE and partition evolution on Iceberg tables:

  • Plan for metadata-only behavior. ALTER TABLE operations update metadata instantly, and existing data files remain unchanged. All readers see the new schema immediately after the operation completes.
  • Drop partition fields before dropping partitioned columns. To remove a column used in the current partition spec, first drop or replace the partition field, then drop the column.
  • Use safe type promotions for ALTER COLUMN TYPE. Amazon Redshift supports widening within compatible families (INT to BIGINT, FLOAT to DOUBLE, DECIMAL(10,2) to DECIMAL(18,2)). Plan column types with future growth in mind.
  • Account for mixed partition layouts after evolution. Partition evolution doesn’t re-partition existing data. Old files remain in their original layout, and the query engine reads both layouts transparently.
  • Use external schemas for database user access. The auto-mounted three-part notation ("bucket@s3tablescatalog") requires IAM federated authentication. For database users and BI tools, create an external schema with an explicit IAM role.
  • Use full three-part notation with awsdatacatalog. The USE statement isn’t supported with awsdatacatalog, so always specify the full path.
  • Clean up S3 data separately after dropping tables. Dropping an Iceberg table removes only the catalog entry from AWS Glue Data Catalog. Delete the underlying S3 data files separately, or use AWS Glue table optimizers to remove orphaned files.

Clean up

To avoid ongoing charges, run the following:

-- Drop the fulfillment_status column added during testing
ALTER TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders
DROP COLUMN fulfillment_status;
-- Restore original partition spec (if changed)
ALTER TABLE "iceberg-write-blog@s3tablescatalog".iceberg_write_namespace.orders
REPLACE PARTITION FIELD MONTH(order_date) WITH DAY(order_date);
-- Drop external schema
DROP SCHEMA IF EXISTS s3tables_iceberg;

Conclusion

In this post, you evolved Apache Iceberg table schemas using ALTER TABLE operations. You added, dropped, and renamed columns, widened data types, changed compression, and evolved partition specs, all as metadata-only operations without rewriting data. You also created Lake Formation resource links to provide governed cross-engine access to S3 Tables, and simplified query syntax with external schemas.

This concludes the three-part series on getting started with Apache Iceberg write support in Amazon Redshift:

  1. Part 1: Create Iceberg tables and perform INSERT operations.
  2. Part 2: Run DELETE, UPDATE, and MERGE for row-level modifications.
  3. Part 3: Evolve schemas with ALTER TABLE and add cross-engine access with Lake Formation resource links.

If you have questions or feedback about this series, leave a comment on this post.

Additional resources


About the authors

Raghu Kuppala

Raghu Kuppala

Raghu is an Analytics Specialist Solutions Architect experienced working in the databases, data warehousing, and analytics space. Outside of work, he enjoys trying different cuisines and spending time with his family and friends.

Tanishq Goyal

Tanishq Goyal

Tanishq is a Software Development Engineer at AWS.

Sanket Hase

Sanket Hase

Sanket is an Engineering Manager with the Amazon Redshift team, leading query execution teams in the areas of data lake analytics, hardware-software co-design, and vectorized query execution.

Vlad Ponomarenko

Vlad Ponomarenko

Vlad is a Senior Software Development Engineer with the Amazon Redshift team, working on query processing, serverless, and integrations. Outside of work, he enjoys watching and playing sports and live music.

Sam Wang

Sam Wang

Sam works query processing and data ingestion as a Software Development Engineer on the Amazon Redshift team. When he’s not writing code, you’ll find him on the slopes.

Fahim Chodhury

Fahim Chowdhury

Fahim works on data lake query execution engine and query processing as a Software Development Engineer on the Amazon Redshift team.

Enforce IAM permissions boundaries for Amazon SageMaker Unified Studio Tooling blueprints

Post Syndicated from Sanjana Sekar original https://aws.amazon.com/blogs/big-data/enforce-iam-permissions-boundaries-for-amazon-sagemaker-unified-studio-tooling-blueprints/

Amazon SageMaker Unified Studio now supports custom permissions boundaries for IAM roles created by the Tooling blueprint. Organizations that enforce Service Control Policies (SCPs) requiring permissions boundaries on all AWS Identity and Access Management (IAM) roles can now adopt Amazon SageMaker Unified Studio without modifying their security posture.

Amazon SageMaker Unified Studio is a unified development environment that brings together data engineering, machine learning, and analytics tools into a single workspace. In Amazon SageMaker Unified Studio, a project is a collaborative workspace that bundles people, tools, and access permissions together. It builds every project from a project profile, which defines a list of blueprints. Blueprints are pre-configured infrastructure templates that provision AWS resources at project creation time or on demand, along with their default parameters. The Tooling blueprint is the only mandatory one. Amazon SageMaker Unified Studio deploys it with every project, creating foundational resources such as the project IAM role and security groups.

In this post, you learn how to create a permissions boundary that restricts AI agent capabilities. You then configure it on the Tooling blueprint using the AWS Command Line Interface (AWS CLI). Finally, you validate that the boundary is enforced on all provisioned roles.

The problem

Enterprises in regulated industries use SCPs to require that every IAM role in an account carries a permissions boundary. A well-scoped boundary prevents privilege escalation and verifies no role exceeds the maximum permissions defined by the organization’s security team. Before this feature, Amazon SageMaker Unified Studio Tooling blueprints created IAM roles without permissions boundaries. When an SCP enforced permissions boundaries, project creation failed with an explicit deny:

User: arn:aws:sts::<account-id>:assumed-role/AmazonSageMakerProvisioning-<account-id>/AmazonDataZoneEnvironmentDeployer-<account-id> is not authorized to perform: iam:CreateRole on resource: arn:aws:iam::<account-id>:role/AmazonBedrockServiceRole-<project-id>-<env-id> with an explicit deny in a service control policy

Amazon SageMaker Unified Studio surfaces the blocked role creation as a Tooling environment provisioning failure, as shown in Figure 1.

SMUS project overview showing the Tooling environment in a failed state from a permissions boundary SCP denial

Figure 1: Project creation fails when the SCP requires a permissions boundary that is not attached

The project is marked as failed because its Tooling environment couldn’t deploy in the US East (N. Virginia) AWS Region (us-east-1). The details show a 403 permissions error, while the preceding IAM message identifies the underlying iam:CreateRole SCP denial. This blocked adoption for any organization with SCP-enforced permissions boundaries. The AWS CloudFormation event for the Tooling stack exposes the IAM failure behind the project-level error, as shown in Figure 2.

CloudFormation stack events showing the BedrockServiceRole in CREATE_FAILED from an iam:CreateRole SCP explicit deny

Figure 2: Detailed error showing the SCP denial in the Tooling blueprint AWS CloudFormation stack

The AmazonBedrockServiceRole resource entered CREATE_FAILED because iam:CreateRole was explicitly denied by the SCP, even though AWS CloudFormation surfaced the wrapper error as UnauthorizedTaggingOperation.

Granular control using a permissions boundary: Example use case

Beyond satisfying SCP requirements, permissions boundaries give administrators granular control over what the Tooling blueprint roles can do. For instance, some organizations have SecOps policies that require disabling Data Agent and Data Notebook capabilities across their accounts. These organizations want project members to access data connections and run SQL queries directly, but must block conversational AI, code generation, and notebook cell execution through the agent.

When PermissionsBoundaryArn is configured on the Tooling blueprint, SageMaker Unified Studio attaches the specified customer-managed permissions boundary to all IAM roles provisioned by that blueprint. If your governance requires a boundary, configure it explicitly and verify the resulting roles.

The following permissions boundary policy scopes the roles to the AWS services that Amazon SageMaker Unified Studio uses and then explicitly denies the Amazon DataZone actions that power the AI agent. This permissions boundary is provided for illustrative purposes only and isn’t a recommendation or reference for environment configuration. You should tailor your permissions boundaries to your specific workloads in accordance with the principle of least privilege.

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "AllowSmusServiceScope",
      "Effect": "Allow",
      "Action": [
        "datazone:*",
        "sagemaker:*",
        "glue:*",
        "s3:*",
        "lakeformation:*",
        "redshift:*",
        "redshift-data:*",
        "redshift-serverless:*",
        "athena:*",
        "q:*",
        "elasticmapreduce:*",
        "bedrock:*",
        "lambda:*",
        "kms:*",
        "secretsmanager:*",
        "codecommit:*",
        "logs:*",
        "cloudwatch:*",
        "sts:AssumeRole",
        "iam:PassRole",
        "ec2:Describe*",
        "ec2:CreateNetworkInterface",
        "ec2:DeleteNetworkInterface",
        "ec2:CreateNetworkInterfacePermission",
        "ec2:DeleteNetworkInterfacePermission"
      ],
      "Resource": "*"
    },
    {
      "Sid": "DenyDataNotebookAndDataAgent",
      "Effect": "Deny",
      "Action": [
        "datazone:*Notebook*",
        "datazone:*Cell*",
        "datazone:*Conversation*",
        "datazone:SendMessage",
        "datazone:GenerateCode",
        "datazone:CancelMessage"
      ],
      "Resource": "*"
    }
  ]
}

Warning: validate before using in production. This example scopes the roles to the service namespaces Amazon SageMaker Unified Studio uses, but it is still coarse (it allows each listed service in full) and is provided only for illustration. Because a permissions boundary is a ceiling, it must remain a superset of everything the three Tooling roles (datazone_usr_role, AmazonBedrockServiceRole, and AmazonBedrockLambdaExecutionRole) actually need. If Amazon SageMaker Unified Studio adds a dependency that isn’t listed, provisioning or in-console actions will fail with an access denied error. Validate in a non-production domain first.

With this boundary attached, the Tooling blueprint provisions normally, project members can access data connections and run SQL queries. However, any attempt to invoke the AI assistant or execute notebook cells through the agent returns an access denied error. The boundary acts as a ceiling that no policy attached to the role can override.

How it works

The custom permissions boundary feature operates at the blueprint configuration level. An administrator sets a PermissionsBoundaryArn in the Tooling blueprint’s regional parameters. When a user creates a new project that includes the Tooling blueprint, Amazon SageMaker Unified Studio provisions an AWS CloudFormation stack that creates three IAM roles and attaches the specified boundary to each:

  • datazone_usr_role – the role that all project members assume to access data and resources in that project.
  • AmazonBedrockServiceRole – for Amazon Bedrock operations.
  • AmazonBedrockLambdaExecutionRole – for Amazon Bedrock-related AWS Lambda functions.

Because the boundary is set at the blueprint level, it applies to every project created under that blueprint. No per-project configuration is needed.

Prerequisites

Before you begin, make sure that you have:

If your organization uses AWS Organizations with SCPs that require permissions boundaries, you will also need an organization with the target account as a member and permissions to create and attach SCPs in the management account.

Setting up the SCP (optional)

This section provides instructions to create an SCP and attach it to your AWS Organizations organizational unit or accounts. If your organization already enforces permissions boundaries through SCPs, skip this section. Otherwise, create an SCP in your AWS Organizations management account that denies IAM role creation unless an approved permissions boundary is attached. This also prevents the boundary from being removed, swapped, or weakened afterward:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DenyRoleWithoutApprovedBoundary",
      "Effect": "Deny",
      "Action": [
        "iam:CreateRole",
        "iam:PutRolePermissionsBoundary"
      ],
      "Resource": "*",
      "Condition": {
        "StringNotEquals": {
          "iam:PermissionsBoundary": "arn:aws:iam::${aws:PrincipalAccount}:policy/SMUSToolingBoundary"
        }
      }
    },
    {
      "Sid": "DenyRemovingBoundary",
      "Effect": "Deny",
      "Action": "iam:DeleteRolePermissionsBoundary",
      "Resource": "*"
    },
    {
      "Sid": "ProtectBoundaryPolicy",
      "Effect": "Deny",
      "Action": [
        "iam:DeletePolicy",
        "iam:CreatePolicyVersion",
        "iam:SetDefaultPolicyVersion"
      ],
      "Resource": "arn:aws:iam::${aws:PrincipalAccount}:policy/SMUSToolingBoundary"
    }
  ]
}

This policy does three things:

  • DenyRoleWithoutApprovedBoundary blocks creating a role, or attaching a boundary to an existing role, with anything other than the approved boundary ARN. Denying iam:PutRolePermissionsBoundary in addition to iam:CreateRole stops a privileged principal from swapping in a weaker boundary after the role exists.
  • DenyRemovingBoundary blocks iam:DeleteRolePermissionsBoundary outright, so the boundary cannot be stripped off. (This action doesn’t support the iam:PermissionsBoundary condition key, so it must be denied unconditionally.)
  • ProtectBoundaryPolicy prevents tampering with the boundary policy itself. Deleting it, or publishing and defaulting a new version that quietly widens what it allows.

Note: Scope these denies so you don’t lock yourself out. A broad deny on iam:PutRolePermissionsBoundary and iam:DeleteRolePermissionsBoundary also applies to your own administrators. Add an exception for a break-glass or IAM-admin role (for example, an aws:PrincipalArn StringNotLike condition) so a trusted principal can still manage boundaries.

To create the SCP, sign in to the AWS Organizations console with your management account and go to AWS Organizations → Policies → Service control policies. If SCPs aren’t enabled for your organization yet, choose Enable service control policies first. Choose Create policy, give it a name (for example, test_scp), and replace the default content in the policy editor with the JSON above substituting <account-id> with your account ID. Choose Create policy to save it.

After creating the SCP in the management account, verify its content before attaching it. Figure 3 shows the core create-role control. The full example above adds controls that prevent replacing or removing the boundary and modifying the protected policy.

Figure 3: Service Control Policy defined in the AWS Organizations management account

The AWS Organizations Content tab displays the customer-managed test_scp policy. Its visible statement denies iam:CreateRole unless the request uses the SMUSToolingBoundary policy.

Attach this SCP to the organizational unit or account where your Amazon SageMaker Unified Studio domain and domain-associated accounts reside. To do so, open the test_scp service control policy, choose the Targets tab, and choose Attach. The AWS organization structure appears; select the OU or account where the SCP should apply, then choose Attach policy.

Figure 4 identifies the member account that must inherit the SCP in this example organization. The target member account, datazone-account2, resides under OU2, while datazone-account1 is the organization’s management account. Attaching the SCP to the target account or a parent organizational unit enforces it there.

Figure 4: AWS Organizations account structure showing the management account and the target member account

After attaching the policy, verify the association on the SCP’s Targets tab, as shown in Figure 5.

SCP Targets tab listing datazone-account2 as an account target where test_scp is enforced

Figure 5: Service Control Policy attached to the target account where it should be enforced

The Targets tab lists datazone-account2 as an ACCOUNT target, confirming that test_scp is enforced directly on the intended member account.

Configuring the permissions boundary

In this section you will execute the required steps to create the permissions boundary and enable it in the Tooling blueprint. The example in this walkthrough uses us-east-1. Change it to the Region where your Amazon SageMaker Unified Studio domain is deployed. You must execute the configuration in the account where you plan to create your project. This can be your Amazon SageMaker Unified Studio domain account or accounts associated to your Amazon SageMaker Unified Studio domain.

Step 1: Create the permissions boundary policy

If you haven’t already created the boundary policy, save the following JSON document as a boundary-policy.json file on your workstation:

{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "AllowSmusServiceScope",
            "Effect": "Allow",
            "Action": [
                "datazone:*",
                "sagemaker:*",
                "glue:*",
                "s3:*",
                "lakeformation:*",
                "redshift:*",
                "redshift-data:*",
                "redshift-serverless:*",
                "athena:*",
                "q:*",
                "elasticmapreduce:*",
                "bedrock:*",
                "lambda:*",
                "kms:*",
                "secretsmanager:*",
                "codecommit:*",
                "logs:*",
                "cloudwatch:*",
                "sts:AssumeRole",
                "iam:PassRole",
                "ec2:Describe*",
                "ec2:CreateNetworkInterface",
                "ec2:DeleteNetworkInterface",
                "ec2:CreateNetworkInterfacePermission",
                "ec2:DeleteNetworkInterfacePermission",
                "iam:GetRole",
                "sqlworkbench:*"
            ],
            "Resource": "*"
        },
        {
            "Sid": "DenyDataNotebookAndDataAgent",
            "Effect": "Deny",
            "Action": [
                "datazone:*Notebook*",
                "datazone:*Cell*",
                "datazone:*Conversation*",
                "datazone:SendMessage",
                "datazone:GenerateCode",
                "datazone:CancelMessage"
            ],
            "Resource": "*"
        }
    ]
}

As noted previously, this illustrative policy is scoped to the services Amazon SageMaker Unified Studio uses but is still coarse, and its allow list must stay a superset of what all three Tooling roles need.

Then create the policy using the following command:

aws iam create-policy \
--policy-name SMUSToolingBoundary \
--policy-document file://boundary-policy.json \
--description "Permissions boundary for SMUS Tooling roles - denies Data Agent and Data Notebook capabilities"

Note the policy ARN from the output, because it will be used later in the procedure.

Step 2: Retrieve the ID of your domain

Retrieve the ID of your domain by running the following command. Replace <YOUR_DOMAIN_NAME> with the name of your SageMaker Unified Studio domain.

aws datazone list-domains \
--region us-east-1 \
--query "items[?name=='<YOUR_DOMAIN_NAME>'].id | [0]" \
--output text

Note the returned ID, because it will be used later in the procedure.

Step 3: Identify the Tooling blueprint

Retrieve the Tooling blueprint ID by running the following command. Replace <domain-id> with the ID you noted in Step 2.

aws datazone list-environment-blueprints \
  --domain-identifier <domain-id> \
  --managed \
  --region eu-west-1 \
  --query "items[?name=='Tooling'].id" \
  --output json | jq -r '.[0]'

Note the returned ID, because it will be used later in the procedure.

Step 4: Read the current configuration

Retrieve the current Tooling blueprint configuration by executing the following command. Replace <domain-id> with the ID from Step 2 and <tooling-bp-id> with the ID from Step 3.

aws datazone get-environment-blueprint-configuration \
--domain-identifier <domain-id> \
--environment-blueprint-identifier <tooling-bp-id> \
--region us-east-1 | tee tooling-bp-config-backup.json

Important: Back up the output of get-environment-blueprint-configuration before making any changes. The command above pipes the response to tooling-bp-config-backup.json so you have a restore point if you need to revert.

Note the values of provisioningRoleArn, manageAccessRoleArn, enabledRegions, and all fields inside regionalParameters (AZs, S3Location, Subnets, VpcId). You will need all of these in the next step.

Step 5: Set the permissions boundary

Update the blueprint configuration to include PermissionsBoundaryArn in the regional parameters using the following command.

Important: The put-environment-blueprint-configuration API operates in overwrite mode, it replaces the entire configuration with what you provide. You must include all existing values from the previous step’s output. The only new addition is PermissionsBoundaryArn inside the regional parameters. Omitting any existing parameter removes it.

Make sure to replace <domain-id> with the ID you noted in Step 2, <tooling-bp-id> with the ID you noted in Step 3, and all other <placeholder> values with the corresponding values from Step 4’s output.

aws datazone put-environment-blueprint-configuration \
--domain-identifier <domain-id> \
--environment-blueprint-identifier <tooling-bp-id> \
--enabled-regions '<enabledRegions>' \
--provisioning-role-arn "<provisioningRoleArn>" \
--manage-access-role-arn "<manageAccessRoleArn>" \
--regional-parameters '{
  "<region>": {
    "AZs": "<AZs>",
    "S3Location": "<S3Location>",
    "Subnets": "<Subnets>",
    "VpcId": "<VpcId>",
    "PermissionsBoundaryArn": "arn:aws:iam::<account-id>:policy/SMUSToolingBoundary"
  }
}' \
--region <region>

The following anonymized example is based on an existing Tooling blueprint configuration. Its S3Location reflects the bucket naming pattern used in that environment. Copy the exact S3Location returned in Step 4. Don’t use the following illustrative value. Here’s an example:

aws datazone put-environment-blueprint-configuration \
--domain-identifier <domain-id> \
--environment-blueprint-identifier <tooling-bp-id> \
--enabled-regions '["us-east-1"]' \
--provisioning-role-arn "arn:aws:iam::<account-id>:role/service-role/AmazonSageMakerProvisioning-<account-id>" \
--manage-access-role-arn "arn:aws:iam::<account-id>:role/service-role/AmazonSageMakerManageAccess-us-east-1-<domain-id>" \
--regional-parameters '{
  "us-east-1": {
    "AZs": "us-east-1a,us-east-1b,us-east-1c,us-east-1d",
    "S3Location": "s3://amazon-sagemaker-<account-id>-us-east-1-<suffix>",
    "Subnets": "<subnet-1>,<subnet-2>,<subnet-3>,<subnet-4>",
    "VpcId": "<vpc-id>",
    "PermissionsBoundaryArn": "arn:aws:iam::<account-id>:policy/SMUSToolingBoundary"
  }
}' \
--region us-east-1

Step 6: Verify the configuration was applied

Confirm the permissions boundary ARN is now set in the blueprint configuration using the following command. Make sure to replace <domain-id> with the ID you noted in Step 2 and <tooling-bp-id> with the ID you noted in Step 3.

aws datazone get-environment-blueprint-configuration \
--domain-identifier <domain-id> \
--environment-blueprint-identifier <tooling-bp-id> \
--region us-east-1 \
--query "regionalParameters.\"us-east-1\".PermissionsBoundaryArn"

The output should return your boundary policy ARN:

"arn:aws:iam::<account-id>:policy/SMUSToolingBoundary"

Validating the configuration

After configuring the permissions boundary, in this section you will get instructions to create a new project to verify it works end to end and that the IAM roles created with the project actually include the permissions boundary.

Step 1: Select a project profile in enabled state

Use the following command to list project profiles configured in your domain. Make sure to replace <domain-id> with the ID you noted in Step 2 of the “Configuring the permissions boundary” section.

aws datazone list-project-profiles \
--domain-identifier <domain-id> \
--region us-east-1

Choose a project profile that has "status": "ENABLED". Note the id of any project profile returned in the previous command.

Step 2: Create a test project

Create a new project using the following command. Make sure to replace <domain-id> with the ID you noted in Step 2 of the “Configuring the permissions boundary” section and to replace <profile-id> with the project profile ID noted in Step 1 of this section.

aws datazone create-project \
--domain-identifier <domain-id> \
--name "PB-Validation-$(date +%Y%m%d-%H%M%S)" \
--project-profile-id <profile-id> \
--region us-east-1

Note the id (project ID) returned in the response. Wait for the Tooling blueprint to provision. This typically takes a minute or two. After provisioning completes, confirm that the validation project reaches the Active state, as shown in Figure 6.

SMUS Projects list showing the timestamped PB-Validation project in Active status after successful creation

Figure 6: Project created successfully with the permissions boundary configured

The Projects list shows the timestamped PB-Validation-* project with an Active status, confirming that project creation succeeded with the custom boundary configured.

Step 3: Verify the roles have the boundary attached

In this section you check that the IAM roles created with the project have the permissions boundary attached. Use the following commands to get the configuration for the IAM roles created with the project you just created. Replace <domain-id> and <project-id> with the values from the previous steps.

# Get the environment ID
ENV_ID=$(aws datazone list-environments \
--domain-identifier <domain-id> \
--project-identifier <project-id> \
--region us-east-1 \
--query "items[?name=='Tooling'].id" --output text)

# List IAM roles in the AWS CloudFormation stack
aws cloudformation describe-stack-resources \
--stack-name "DataZone-Env-${ENV_ID}" \
--region us-east-1 \
--query "StackResources[?ResourceType=='AWS::IAM::Role'].PhysicalResourceId" \
--output table

# Verify each role has the boundary
aws iam get-role \
--role-name "<role-name>" \
--query 'Role.PermissionsBoundary'

All three roles should return a response showing the permissions boundary ARN:

{
  "PermissionsBoundaryType": "Policy",
  "PermissionsBoundaryArn": "arn:aws:iam::<account-id>:policy/SMUSToolingBoundary"
}

You can also verify each role in the IAM console. Figure 7 shows the permissions boundary for the project user role.

IAM console Permissions tab showing SMUSToolingBoundary as the permissions boundary on datazone_usr_role

Figure 7: IAM console showing the permissions boundary attached to the datazone_usr_role

The datazone_usr_role Permissions tab displays SMUSToolingBoundary as its customer-managed permissions boundary.

Figure 8 confirms that the same boundary is attached to the Amazon Bedrock service role.

IAM console showing SMUSToolingBoundary as the permissions boundary on AmazonBedrockServiceRole

Figure 8: IAM console showing the permissions boundary attached to the AmazonBedrockServiceRole

The AmazonBedrockServiceRole also displays SMUSToolingBoundary as its customer-managed permissions boundary.

Figure 9 verifies the boundary on the third Tooling role, the Bedrock Lambda execution role.

IAM console showing SMUSToolingBoundary as the permissions boundary on AmazonBedrockLambdaExecutionRole

Figure 9: IAM console showing the permissions boundary attached to the AmazonBedrockLambdaExecutionRole

The AmazonBedrockLambdaExecutionRole likewise displays SMUSToolingBoundary, confirming that all three provisioned roles carry the boundary.

Step 4: Verify the boundary denies AI agent actions

In this section you verify the boundary actually denies AI agent actions. If you configured the boundary from the use case section earlier, the boundary blocks Data Notebooks and messages to the Data Agent, such as the Query Editor assistant. Any such attempt returns an access denied error. The project user role has the boundary attached, so even if the role’s identity policies grant the relevant APIs, the boundary’s explicit deny takes precedence.

To confirm, navigate to your project in SageMaker Unified Studio and test the following actions:

  1. Attempt to create a notebook – In the left sidebar, select Notebooks. Select Create notebook. The operation will fail because the permissions boundary prevents the datazone:CreateNotebook action (Figure 10).

Figure 10: Permissions boundary preventing creation of Data Notebooks

After the create action, Amazon SageMaker Unified Studio reports Failed to create notebook and identifies datazone:CreateNotebook as explicitly denied by SMUSToolingBoundary.

  1. Attempt to use Data Agent in the Query Editor – In the left sidebar, select Query Editor, then select the Chat with AI icon. The agent chat will fail to load because the permissions boundary blocks the APIs required by Data Agent (Figure 11).

Figure 11: Permissions boundary preventing using Data Agent on Query Editor

The Query Editor remains available, but the Agent panel reports “You don’t have access to Data Agent“. In this configured test, that message is the user-visible result of denying the Data Agent APIs. The screenshot itself doesn’t display the denied API or boundary ARN.

Important considerations

  • Immutable after project creation – The permissions boundary is set at provisioning time. Changing the boundary ARN on the blueprint configuration only affects new projects. Existing projects retain their original boundary.
  • Applies to all Tooling-provisioned roles – When PermissionsBoundaryArn is configured on the Tooling blueprint, SageMaker Unified Studio attaches the specified customer-managed permissions boundary to all three IAM roles created by that blueprint. It’s applied uniformly — you can’t selectively apply it to individual roles. No boundary is attached unless you configure one, so if your governance requires a boundary, set it explicitly and verify the resulting roles rather than assuming one is present by default.
  • Policy must exist – The IAM policy referenced by PermissionsBoundaryArn must exist in the account before project creation. If the policy is deleted or the ARN is invalid, provisioning will fail.
  • Tooling blueprint only – Among Amazon SageMaker Unified Studio provided blueprints, only the Tooling blueprint supports custom permissions boundaries. Other provided blueprints that create IAM roles (for example, the EmrOnEc2 blueprint) don’t currently support this feature. If your organization requires permissions boundaries on roles created by additional blueprints, you can build custom blueprints that include a permissions boundary configuration so you can extend this security control across your entire project infrastructure.

Clean up

To remove test resources, delete the test project from the SageMaker Unified Studio UI. On the project’s Overview page, choose the ⋮ (more actions) menu in the top-right and choose Delete project.

Figure 12: Deleting the test project from the project Overview page.

In the Delete project dialog, type confirm in the text box to acknowledge that the action is final, then choose Delete project. This permanently deletes the project and its underlying resources, and triggers an asynchronous AWS CloudFormation stack deletion.

Figure 13: Confirming project deletion.

To remove the boundary from future projects, re-run the put-environment-blueprint-configuration command from Step 5: Set the permissions boundary, but omit the PermissionsBoundaryArn field from the regional parameters. Because you backed up the original configuration in Step 4: Read the current configuration (tooling-bp-config-backup.json), you can reuse the exact same provisioningRoleArn, manageAccessRoleArn, enabledRegions, and regionalParameters values (AZs, S3Location, Subnets, VpcId) — just without PermissionsBoundaryArn — so the blueprint returns to provisioning roles with no permissions boundary.

Conclusion

With the custom permissions boundary feature for Amazon SageMaker Unified Studio, organizations can adopt Amazon SageMaker Unified Studio Tooling blueprints without compromising their IAM governance posture. By configuring a single parameter on the Tooling blueprint, all IAM roles provisioned by future projects automatically carry the specified permissions boundary. This satisfies SCPs that mandate a boundary on every role and gives administrators granular control over what the Tooling roles can do, for example disabling AI agent and notebook capabilities. Remember that the example boundary in this post is illustrative, because it scopes to the services SageMaker Unified Studio uses but is still coarse.

“I just updated the EnvironmentBlueprintConfiguration for the Tooling blueprint to include the new PermissionsBoundaryArn param. After that the blueprint provisioned successfully with the required permissions boundary attached to all the IAM roles, in line with our security policies. In the end it was a one-line change.”

— Nat Noordanus, Data Tech Lead at Nexthink

To get started, create your permissions boundary policy, configure it on the Tooling blueprint using the CLI, and create a project to verify the boundary is attached.

For more information, see the documentation for Amazon SageMaker Unified Studio, IAM permissions boundaries, and Service Control Policies.


About the authors

Sanjana Sekar

Sanjana Sekar

Sanjana is a Software Development Engineer on the Amazon SageMaker Unified Studio team. She is focused on improving Data Agent capabilities and the compute blueprints experience within SageMaker Unified Studio. Outside of work, she enjoys hiking and biking.

Luca Perrozzi

Luca Perrozzi

Luca is a Solutions Architect at AWS, based in Switzerland. He focuses on innovation topics at AWS, especially in Artificial Intelligence. Luca holds a PhD in particle physics and has 15 years of hands-on experience as a research scientist and software engineer.

Ganesh Sambandan

Ganesh Sambandan

Ganesh is a Senior Technical Account Manager at AWS, helping organizations adopt best practices for running secure, reliable and well-architected workloads on AWS. He works closely with strategic customers to accelerate the adoption of AI-driven cloud operations, enabling more effective DevOps practices, automation and operational excellence.

Stefano Sandona

Stefano Sandona

Stefano is a Senior Worldwide Specialist Solutions Architect for Big Data at AWS, helping customers build efficient, secure, and scalable data solutions.

Paolo Romagnoli

Paolo Romagnoli

Paolo is a Senior Solutions Architect at AWS who helps global energy organizations design and build data and AI enterprise solutions at scale.

Configure domain-level VPC networking in Amazon SageMaker Unified Studio

Post Syndicated from Prasad Nadig original https://aws.amazon.com/blogs/big-data/configure-domain-level-vpc-networking-in-amazon-sagemaker-unified-studio/

Enterprise operations teams that run domain-level VPC networking in Amazon SageMaker Unified Studio often support dozens of projects spanning data engineering, analytics, and machine learning (ML) teams. Each project requires private connectivity to internal databases, Amazon Simple Storage Service (Amazon S3) buckets, and AWS services. Without a domain-level Amazon Virtual Private Cloud (Amazon VPC) configuration, project owners coordinate with the networking team individually. This piecemeal approach leads to inconsistent subnet choices, missing VPC endpoints, connectivity failures that are hard to troubleshoot, and a network posture that is difficult to audit.

With domain-level VPC networking, you configure the network once, and all new projects get the right network immediately upon creation. In this post, you learn how to:

  • Configure SageMaker Unified Studio domain-level VPC networking.
  • Select subnets and security groups that provide multi-Availability Zone (multi-AZ) resilience.
  • Update projects that have no VPC to inherit the domain VPC, and understand when a project must be recreated instead.
  • Validate network connectivity from within a project.

In this post, you learn how to configure VPC networking for a SageMaker Unified Studio domain that uses AWS Identity and Access Management (IAM)-based authentication. You see how network components map to domain and project resources, and how to plan a configuration that balances security, connectivity, and operational simplicity.

Solution overview

Domain-level VPC networking provides a single network configuration that applies to all new projects in the domain. Projects automatically inherit the VPC settings, including subnets, security groups, and connectivity to AWS services through VPC endpoints. Existing projects are an exception and are handled separately (see Step 3).

The following diagram shows a single VPC with private subnets across two Availability Zones configured at the domain level, with data engineering, analytics, and ML projects all inheriting that configuration.

Architecture diagram showing domain-level VPC configuration in Amazon SageMaker Unified Studio with private subnets across two Availability Zones.

Figure 1: Domain-level VPC configuration in Amazon SageMaker Unified Studio. A single VPC with private subnets across two Availability Zones is configured at the domain level. All projects (data engineering, analytics, ML) inherit this configuration automatically

Key benefits of this approach:

  • Configure once, apply across projects: New projects inherit the domain VPC without manual intervention.
  • Consistent security posture: A single network boundary covers all data, analytics, and ML workloads.
  • Simplified auditing: One VPC to audit rather than one per project. Turn on VPC Flow Logs and review AWS CloudTrail events for network-level auditing.
  • Reduced operational overhead: Project teams start working immediately without submitting networking requests.

The following AWS services are used in this solution:

Prerequisites

Before configuring domain-level VPC networking, verify you have the following:

  • Domain administrator permissions for Amazon SageMaker Unified Studio.
  • An existing VPC with the following requirements:
    • At least two private subnets in different Availability Zones.
    • DNS hostnames and DNS support enabled.
    • At least five available IP addresses per expected Amazon SageMaker Unified Studio project. This is a baseline minimum. Workloads using AWS Glue, Amazon EMR, or Amazon Redshift Serverless consume additional elastic network interfaces (ENIs) per worker or node. We recommend /24 or larger subnets for production domains and forward-looking capacity planning based on your expected users and compute types. For detailed guidance, see How to set up a network-isolated VPC for Amazon SageMaker Unified Studio.
  • VPC endpoints configured for the AWS services your projects access (for example, Amazon S3, AWS Glue, Amazon SageMaker AI).
  • Private DNS enabled on all interface VPC endpoints (you must enable this so that service DNS names resolve to private IPs). If you use centralized VPC endpoints through AWS Resource Access Manager (AWS RAM) or AWS Transit Gateway, configure Amazon Route 53 Resolver inbound endpoints instead.
  • S3 gateway endpoint route table associations configured for all selected private subnets (without this, S3 access fails in subnets whose route table lacks the prefix-list route).
  • A security group (optional), if not provided, SageMaker Unified Studio creates one automatically.
  • The SageMakerStudioAdminIAMConsolePolicy managed policy (or equivalent permissions including ec2:Describe*, ec2:CreateSecurityGroup, and datazone:* actions) attached to the domain administrator IAM role. See SageMakerStudioAdminIAMConsolePolicy in the AWS Managed Policy Reference for the full permission set.

Note: The VPC must be in the same AWS Region as the domain.

For detailed guidance on VPC networking configuration, see Configure VPC networking for IAM-based domains in the SageMaker Unified Studio Administrator Guide.

VPC endpoint requirements

Because your subnets are private (no internet gateway route), compute resources access AWS services through VPC endpoints. At a minimum, configure the following interface and gateway endpoints (add Amazon Athena, AWS Lake Formation, or Amazon Redshift endpoints if you use them in your projects):

Endpoint Type Purpose
com.amazonaws.region.s3 Gateway S3 access for data storage
com.amazonaws.region.glue Interface AWS Glue job connectivity
com.amazonaws.region.sagemaker.api Interface SageMaker API calls
com.amazonaws.region.sagemaker.runtime Interface Model inference
com.amazonaws.region.logs Interface Amazon CloudWatch Logs
com.amazonaws.region.monitoring Interface Amazon CloudWatch metrics
com.amazonaws.region.sts Interface IAM role assumption
com.amazonaws.region.datazone Interface Amazon SageMaker Unified Studio service connectivity
com.amazonaws.region.ecr.api Interface ECR API calls (container image metadata)
com.amazonaws.region.ecr.dkr Interface ECR image layer pulls (Docker registry)
com.amazonaws.region.kms Interface AWS Key Management Service (AWS KMS) encryption/decryption operations

Note: Interface endpoints incur an hourly charge per Availability Zone plus data processing fees. Gateway endpoints (such as S3) have no hourly charge. Factor endpoint count and AZ spread into your cost estimate.

For a comprehensive list of all mandatory and optional VPC endpoints for a fully network-isolated setup, see How to set up a network-isolated VPC for Amazon SageMaker Unified Studio. For current pricing details, see AWS PrivateLink pricing.

Note: Review your account’s service quotas for interface VPC endpoints per VPC (default 50) and ENIs per Region before scaling. Request increases through Service Quotas if needed.

Solution walkthrough

The following steps walk you through configuring the domain VPC and validating it, from signing in to the console through confirming private connectivity from a project.

Step 1: Sign in and navigate to networking settings

  1. Sign in to the AWS Management Console as your Amazon SageMaker Unified Studio domain administrator (the IAM role designated as the domain login role).
  2. Open the Amazon SageMaker console.
  3. Use the Region selector in the top navigation bar to select the Region where your domain exists.
  4. On the Amazon SageMaker Unified Studio landing page, choose Open to launch your IAM-based domain.

The following screenshot shows the Amazon SageMaker Unified Studio landing page, where you choose Open to launch the domain.

Amazon SageMaker Unified Studio landing page with the Open button to launch the IAM-based domain.

Figure 2: Amazon SageMaker Unified Studio landing page with the Open button to launch the IAM-based domain

  1. From the navigation pane, choose Domain management.

The following screenshot shows Domain management in the navigation pane.

Navigation pane showing Domain management link in Amazon SageMaker Unified Studio.

Figure 3: Domain management on navigation pane

Note: Access to the domain administration page is restricted to the IAM role specified as the domain login role during domain creation.

Step 2: Add VPC configuration

  1. In the navigation pane, choose Settings. In the Networking in this account section, choose Add VPC.

The following screenshot shows the Networking in this account section with the Add VPC button.

Domain management Settings page showing the Networking in this account section with Add VPC button.

Figure 4: Domain management Settings page showing the Networking in this account section to add a VPC

  1. For VPC, select the VPC with connectivity to your compute, database, and storage resources. If no VPC exists, choose Create VPC to provision one using AWS CloudFormation.
  2. For Subnets, select a minimum of two private subnets in different Availability Zones.
  3. (Optional) For Security group, select a security group to control inbound and outbound traffic. If you don’t choose one, SageMaker Unified Studio creates one automatically.
  4. Choose Save.
  5. Verify the VPC configuration status shows Ready in the Networking in this account section.

The following screenshots show the Add VPC dialog and the resulting Ready status in the Networking in this account section.

Add VPC dialog with fields for VPC, subnets, and security group selection.

Figure 5: Add VPC dialog with fields for VPC, subnets, and security group selection

VPC configuration status showing Ready in the Networking in this account section.

Figure 6: VPC configuration status showing Ready in the Networking in this account section

Note: IAM-based domains support only one VPC configuration at a time. AWS IAM Identity Center-based domains can have a VPC per Region. For details, see Configure VPC networking for IAM-based domains in the SageMaker Unified Studio Administrator Guide.

New projects created in the domain now automatically use the saved VPC configuration. Existing projects are an exception. See Step 3 to update them.

Step 3: Update existing projects

Existing projects don’t automatically inherit the domain VPC configuration. How you apply the new settings depends on the project’s current state:

Projects with no VPC configured – Update in place to adopt the domain VPC. See the following steps.

Projects that already have a VPC – These can’t be switched to a different VPC configuration. To adopt the domain VPC:

  1. Create a new project (which inherits the domain VPC automatically).
  2. Recreate connections in the new project.
  3. Migrate assets from the old project.
  4. Back up any data you need, then delete the original project.

Because recreation can disrupt in-progress work and doesn’t migrate project data automatically, schedule this as a planned maintenance window.

To update a project that currently has no VPC configured:

  1. From the domain administration page, choose Projects in the navigation pane.
  2. Choose the project you want to update.
  3. On the project detail page, a banner appears: “Configurations have changed. Please update this project to access the latest configuration.”
  4. In the banner, choose Update.
  5. Confirm the update when prompted.

Repeat this process for each existing project that should use the domain VPC. The following screenshot shows the project detail page with the configuration update banner.

Project detail page showing the update banner for VPC configuration changes.

Figure 7: Project detail page showing the configuration update banner

Step 4: Validate connectivity

After configuring the domain VPC and updating your projects, verify connectivity. Compute resources should have private connectivity to AWS services through the VPC, without any additional project-level network configuration.

Create a notebook in one of your projects as shown in the following figure and run the following code:

Creating a notebook in a SageMaker Unified Studio project to validate VPC connectivity.

Figure 8: Creating a notebook in a SageMaker Unified Studio project to validate VPC connectivity

Requirements: Python 3.8+, Boto3 1.26 or later. Run in a notebook within your SageMaker Unified Studio project.

import boto3
import socket
import ipaddress

def validate_vpc_connectivity():
    """Validate that the project has private connectivity to AWS services
    through the domain-level VPC configuration."""

    results = {}
    region = boto3.session.Session().region_name
    if not region:
        raise RuntimeError('Could not determine AWS Region. Run this notebook inside a SageMaker Unified Studio project.')

    # Test Amazon S3 access via VPC endpoint
    try:
        s3 = boto3.client('s3')
        response = s3.list_buckets()
        results['S3'] = f"[PASS] Accessible ({len(response['Buckets'])} buckets)"
    except Exception as e:
        results['S3'] = f"[FAIL] Failed: {e}"

    # Test AWS Glue access via VPC endpoint
    try:
        glue = boto3.client('glue')
        dbs = glue.get_databases()
        results['Glue'] = f"[PASS] Accessible ({len(dbs['DatabaseList'])} databases)"
    except Exception as e:
        results['Glue'] = f"[FAIL] Failed: {e}"

    # Test STS (role assumption through VPC endpoint)
    try:
        sts = boto3.client('sts')
        identity = sts.get_caller_identity()
        results['STS'] = f"[PASS] Accessible (Account: {identity['Account']})"
    except Exception as e:
        results['STS'] = f"[FAIL] Failed: {e}"

    # Verify interface endpoint resolves to private IP
    try:
        sts_endpoint = f"sts.{region}.amazonaws.com"
        addr_info = socket.getaddrinfo(sts_endpoint, 443, family=socket.AF_INET)
        ip = addr_info[0][4][0]
        is_private = ipaddress.ip_address(ip).is_private
        if is_private:
            results['DNS Resolution'] = f"[PASS] Private IP ({ip}) (traffic stays on AWS network)"
        else:
            results['DNS Resolution'] = f"[WARN] Public IP ({ip}) - check VPC endpoint config"
    except Exception as e:
        results['DNS Resolution'] = f"[FAIL] Failed: {e}"

    # Print results
    print("-" * 40)
    print("Domain VPC Connectivity Validation")
    print("-" * 40)
    for service, status in results.items():
        print(f" {service}: {status}")
    print("-" * 40)
    print(f"\n Region: {region}")

    # Check if all tests passed
    all_passed = all("[PASS]" in status for status in results.values())
    has_warn = any("[WARN]" in status for status in results.values())
    if all_passed:
        print(f"\n [PASS] All services accessible via private VPC endpoints.")
        print(f" This project inherited its network configuration")
        print(f" from the domain without per-project setup.")
    elif has_warn and all("[PASS]" in s or "[WARN]" in s for s in results.values()):
        print(f"\n [WARN] Services are reachable, but DNS resolves to public IPs.")
        print(f" Verify that Private DNS is enabled on your interface VPC endpoints.")
    else:
        print(f"\n [FAIL] Some services are not reachable.")
        print(f" Check that VPC endpoints are configured and security")
        print(f" groups allow outbound traffic on port 443.")

validate_vpc_connectivity()

Expected output when VPC is correctly configured:

Successful validation output showing all services accessible through private VPC endpoints.

Figure 9: Successful validation output showing all services accessible through private VPC endpoints

If any service shows a failure, one common cause is security groups preventing traffic on port 443 to the VPC endpoint. Other causes include missing VPC endpoints, incorrect route table entries, or DNS resolution issues. For more information, see Configure VPC networking for IAM-based domains in the SageMaker Unified Studio Administrator Guide.

Note: An AccessDenied error indicates the request reached the service. Connectivity is working, but IAM permissions need adjustment (for example, the S3 test requires s3:ListAllMyBuckets, which some project roles lack). A timeout or connection error points to a networking problem (missing endpoint, route, or security group rule). The following screenshot shows the validation output when VPC endpoints are missing, where the affected services report timeout errors.

Validation output when VPC endpoints are not configured showing timeout errors.

Figure 10: Validation output when VPC endpoints are not configured. Timeout errors indicate missing endpoints

The security group applied at the domain level controls network access for all projects. To review or tighten the rules:

  1. Navigate to the Amazon VPC console.
  2. Choose Security groups and choose the security group shown in your domain’s Networking settings.
  3. Review the Inbound rules and Outbound rules tabs.

By default, the auto-created security group allows all outbound traffic on port 443 (HTTPS) to reach AWS services through VPC endpoints. Consider restricting outbound rules to only the specific VPC endpoint security groups for least-privilege access. Additionally, make sure your VPC endpoint security groups allow inbound TCP 443 from the domain security group or subnet CIDRs. For distributed compute services (AWS Glue, Amazon EMR), add a self-referencing inbound rule to allow worker-to-worker communication.

Updating VPC configuration

After the initial setup, you can modify the VPC configuration to change the VPC, subnets, or security group:

  1. From the domain administration page, choose Settings in the navigation pane.
  2. In the Networking in this account section, under the Actions column, choose Update.
  3. Update the VPC, subnets, or security group as needed.
  4. Choose Update.

The following screenshot shows the Update VPC dialog, where you modify the VPC, subnets, or security group.

Settings page with the Actions menu showing Update and Remove options for VPC configuration.

Figure 11: Update VPC dialog showing the option to modify VPC, subnets, or security group for the domain

Important: Updating the VPC does not affect already provisioned resources. Newly created resources in projects use the updated VPC. Existing projects that already have a VPC keep their original settings and must be recreated to adopt the change. Projects with no VPC can be updated in place (see Step 3).

Clean up

To remove the VPC configuration from your domain:

  1. From the domain administration page, choose Settings in the navigation pane.
  2. In the Networking in this account section, choose the Actions menu (⋮) and choose Remove.

The following screenshot shows the Actions menu with the Remove option.

Actions menu in the Networking in this account section showing the Remove option.

Figure 12: Actions menu in the Networking in this account section showing the Remove option

If you created a dedicated VPC for this walkthrough and no longer need it:

  • Delete the VPC and associated resources (subnets, VPC endpoints, security groups) from the Amazon VPC console. Before deleting, remove the domain VPC configuration and make sure all project resources are terminated. Active projects create ENIs that block VPC and subnet deletion.
  • If you used an AWS CloudFormation template to create the VPC, delete the stack to remove all resources cleanly. Open the AWS CloudFormation console and delete the stack.

Note: Removing the domain VPC configuration does not retroactively change projects that already have VPC applied. Those projects retain their existing network configuration. New projects created after removal do not have a VPC configured.

Conclusion

In this post, we showed how to configure domain-level VPC networking in Amazon SageMaker Unified Studio. A single domain-level VPC eliminates per-project networking overhead, enforces a consistent security posture, and simplifies compliance auditing.

Key takeaways:

  • Domain-level VPC is a one-time configuration that automatically applies to all new projects.
  • Projects with no VPC can be updated in place. Projects that already have a VPC must be recreated to adopt a changed configuration.
  • Private subnets with VPC endpoints provide secure, private connectivity to AWS services without traversing the public internet.

As next steps, consider:

  • Reviewing your auto-created security group rules and tightening them for least-privilege access.
  • Adding VPC endpoints for additional AWS services as your projects’ needs evolve.
  • Monitoring subnet IP address utilization to plan capacity as you add more projects. Use the AvailableIpAddressCount Amazon CloudWatch metric for your subnets to track utilization and set alarms.

For more information, see Configure VPC networking for IAM-based domains in the Amazon SageMaker Unified Studio Administrator Guide.

 


About the authors

Prasad Nadig

Prasad Nadig

Prasad is a Senior Analytics Specialist Solutions Architect at Amazon Web Services (AWS), specializing in large-scale data analytics and AI. Prasad partners with customers to design, migrate, and modernize their analytics platforms on AWS into scalable, cost-effective solutions, with deep expertise in data lakes, data warehousing, distributed processing, and performance tuning at petabyte scale.

Amit Shyam Jaisinghani

Amit Shyam Jaisinghani

Amit is a Software Engineer on the SageMaker Studio team at Amazon Web Services, and he earned his Master’s degree in Computer Science from Rochester Institute of Technology. Since joining Amazon in 2019, he has built and enhanced several AWS services, including Amazon WorkSpaces and Amazon SageMaker Studio. Outside of work, he explores hiking trails, plays with his two cats, Missy and Minnie, and enjoys playing Age of Empire.

Arun Shanmugam

Arun Shanmugam

Arun is a Senior Analytics Solutions Architect at AWS, with a focus on building modern data architecture. He has been successfully delivering scalable data analytics solutions for customers across diverse industries. Outside of work, Arun is an avid outdoor enthusiast who actively engages in CrossFit, road biking, and cricket.

Announcing Spark Connect on Amazon EMR on EC2: Interactive PySpark anywhere

Post Syndicated from Al MS original https://aws.amazon.com/blogs/big-data/announcing-spark-connect-on-amazon-emr-on-ec2-interactive-pyspark-anywhere/

Today, we’re announcing support for Spark Connect on Amazon EMR on EC2 with the AWS runtime for Apache Spark (emr-spark-8.0, Apache Spark 4.0.2 and later). You can now develop and debug PySpark interactively from Amazon SageMaker Unified Studio Data Notebooks or your own IDE, such as Visual Studio Code, PyCharm, Kiro, or Jupyter. Spark runs on a dedicated Amazon EMR on EC2 cluster while your Python runs locally, so you can set breakpoints and inspect a DataFrame against full-size data from your IDE. In SageMaker Unified Studio Data Notebooks, you connect to your cluster, catalog, and AI tools. Production-scale PySpark and SQL run without leaving the studio. Because each session is isolated with its own permissions, your whole team can share one cluster at the same time. This post shows you how to get started with both SageMaker Unified Studio Data Notebooks and your own IDE.

Previously, developing Spark for an Amazon EMR on EC2 cluster meant working in a notebook tied to that cluster, or packaging your code as a job and submitting it before you could see a result. Local code often behaved differently on the cluster because of version and dependency mismatches, and the slow deploy-and-check loop made those differences hard to find. There was no way to attach your own IDE and debugger and inspect a DataFrame mid-transformation. Spark Connect closes that gap: your code runs against the cluster’s own Spark engine while you develop locally, so the environment you debug in is the one that runs your data.

How Spark Connect works on Amazon EMR on EC2

Spark Connect uses a client-server architecture that separates your application code from the Spark engine. The client is a lightweight PySpark library that runs in your notebook or IDE, and it sends DataFrame and SQL operations over a gRPC/TLS connection to a Spark Connect Server on your cluster. The server runs those operations and returns the results to your local session. Your machine does not need Spark installed and does not need to be sized for the workload.

Spark Connect client-server architecture connecting a local PySpark client to the Spark Connect Server on an Amazon EMR cluster

Figure 1: Spark Connect client-server architecture on Amazon EMR on EC2

When you start a session, Amazon EMR launches the Spark Connect Server as a YARN application on your cluster and hands back an endpoint and a short-lived token. There’s no server for you to stand up or manage. Because that server runs on a cluster you already operate, your session inherits the instance types, libraries, bootstrap actions, and Spark configuration you use in production. What you see while debugging is what runs when the same code is scheduled as a batch job, since both use the same cluster and its configuration.

Share one cluster across your team

Now that you can start sessions, a single dedicated cluster can serve your whole team, because each session is a separate resource with its own execution role, tags, and lifecycle. A single cluster supports up to 1,000 concurrent sessions and 1,000 concurrent execution roles. These values are service maximums, not sizing targets. Actual concurrency depends on cluster size and per-session workload. Because interactive sessions are bursty and rarely all active at once, one cluster typically serves a team larger than its peak concurrent-session count. Enable Amazon EMR managed scaling so that capacity tracks demand. If peak concurrency approaches these maximums, or to isolate cost and data access by group, use multiple clusters—for example, one per team, business unit, or environment. Sharing one cluster gives you:

  • On-demand Spark without extra clusters — Developers get interactive sessions without provisioning a cluster apiece, which keeps utilization high and removes the cost of idle per-person clusters.
  • Consistent environments — Everyone runs the same Spark version, libraries, and security configuration, so results stay consistent, and your platform team patches and monitors one cluster.
  • Isolation and attribution — Per-session execution roles and tags keep each person’s work separate, so you can scope data access by session, track cost by user, and stop one session without disturbing anyone else.
  • Full visibility and control — View active sessions in the Spark UI, review finished ones in the Spark History Server, and manage them from the Amazon EMR console, API, CLI, or SDK.

Getting started

Getting started with Spark Connect on Amazon EMR on EC2 takes three steps: Create an Amazon EMR cluster with Spark Connect session enabled, start a session, and connect from your IDE or SageMaker Unified Studio Data Notebooks.

Note: In SageMaker Unified Studio, on-demand cluster creation is available for domains that use AWS IAM Identity Center. For domains that use AWS Identity and Access Management (IAM), attach an existing cluster. If your cluster runs in a private subnet, make sure that your network configuration allows connectivity between SageMaker Unified Studio and the cluster endpoint.

Prerequisites

You must have the following prerequisites in place.

  • An Amazon EMR cluster running release emr-spark-8.0.0 or later with SessionEnabled set to true.
  • The Spark application is installed on the cluster.
  • Python 3.9 or later with pyspark[connect] installed locally. The PySpark version must match the Spark version on your cluster.
  • For clusters in private subnets, the Amazon EMR service role must include the AmazonEMRServicePolicyForSessions managed policy, which grants permissions to create Network Load Balancers and virtual private cloud (VPC) endpoint services in your account.
  • To use Spark Connect sessions, you need permissions to start and list sessions on the cluster (elasticmapreduce:StartSession, ListSessions), get session details and endpoints and terminate sessions (elasticmapreduce:GetSession, GetSessionEndpoint, TerminateSession), and pass the execution role to the Amazon EMR service (iam:PassRole).

Working with interactive sessions

To create a session-enabled cluster and connect to it, follow these steps.

To start a Spark Connect session

  1. Create a cluster with sessions enabled, running emr-spark-8.0.0 or later. The following is a sample command that you can modify for your needs, such as the instance types and counts:
    aws emr create-cluster \
      --name "spark-connect-cluster" \
      --release-label emr-spark-8.0.0 \
      --applications Name=Spark \
      --service-role EMR_DefaultRole \
      --ec2-attributes InstanceProfile=EMR_EC2_DefaultRole,SubnetId=subnet-id \
      --instance-groups '[
        {"InstanceCount":1,"InstanceGroupType":"MASTER","InstanceType":"m8g.xlarge"},
        {"InstanceCount":2,"InstanceGroupType":"CORE","InstanceType":"m8g.xlarge"}
      ]' \
      --session-enabled \
      --tags Key=for-use-with-amazon-emr-managed-policies,Value=true

    Note: The following steps use the AWS Command Line Interface (AWS CLI) directly. If you develop in SageMaker Unified Studio (Option 1), cluster attachment and session creation are handled for you, so you can skip steps 2 through 6.

  2. After the cluster reaches the WAITING state, start a session and wait for it to reach IDLE:
    aws emr start-session --cluster-id j-XXXXXXXXXXXXX --name "my-session"
    aws emr get-session --cluster-id j-XXXXXXXXXXXXX --session-id is-XXXXXXXXXXXXX

    Note: For runtime role sessions, add the --execution-role-arn parameter to the start-session command.

  3. Retrieve the endpoint and token, and build your connection string from the returned Endpoint value rather than hardcoding a host:
    aws emr get-session-endpoint --cluster-id j-XXXXXXXXXXXXX --session-id is-XXXXXXXXXXXXX

    The response includes the endpoint URL and an authentication token:

    {
      "Endpoint": "https://session-id.emr-spark-connect.region.amazonaws.com",
      "AuthToken": "v2.local.xxx...",
      "AuthTokenExpirationTime": "2026-01-01T01:00:00Z"
    }

  4. Install the matching PySpark client and connect. GetSessionEndpoint returns an https:// URL with no port. Build the connection string by converting it to the sc:// scheme and appending :443. Without the port, the PySpark client defaults to 15002, which isn’t reachable. Your Python code runs locally. The SQL and DataFrame operations run on the cluster:
    pip install 'pyspark[connect]==4.0.2' boto3

    from pyspark.sql import SparkSession
    
    session_id = "is-XXXXXXXXXXXXX"
    auth_token = "<AuthToken from get-session-endpoint>"
    host = "<Endpoint from get-session-endpoint, without https://>"
    
    url = f"sc://{host}:443/;use_ssl=true;x-aws-proxy-auth={auth_token};authorization={session_id}"
    spark = SparkSession.builder.remote(url).getOrCreate()
    spark.sql("SELECT 'Hello from EMR on EC2' AS message").show()

  5. Run a transformation against full-size data. This groups a DataFrame, writes the result to Amazon Simple Storage Service (Amazon S3), and reads it back:
    import pyspark.sql.functions as F
    
    df = spark.range(0, 1000).withColumn(
        "category", F.when(F.col("id") % 2 == 0, "even").otherwise("odd")
    )
    df.groupBy("category").count().show()
    df.write.mode("overwrite").parquet("s3://amzn-s3-demo-bucket/demo/")
    spark.read.parquet("s3://amzn-s3-demo-bucket/demo/").filter("id < 50").orderBy("id").show()

  6. When you finish, terminate the session to release cluster resources. Calling spark.stop() only closes the local connection. The session keeps running until you terminate it or it reaches the idle timeout:
    aws emr terminate-session --cluster-id j-XXXXXXXXXXXXX --session-id is-XXXXXXXXXXXXX

  7. When you’re done with the walkthrough, terminate the cluster you created in step 1 so it stops incurring charges. Terminating the cluster also ends any sessions still running on it:
    aws emr terminate-clusters --cluster-ids j-XXXXXXXXXXXXX

You can start a Spark Connect session in two ways: from SageMaker Unified Studio or from your own IDE client.

Option 1: Develop in SageMaker Unified Studio Data Notebooks

Amazon SageMaker Unified Studio brings your data, catalogs, and analytics and AI tools into one place, and Amazon EMR on EC2 is now one of the Spark runtimes a Data Notebook can use. When you choose that cluster as the notebook runtime, SageMaker Unified Studio connects to it over Spark Connect. The same runtime then drives both your PySpark and SQL cells, so a single notebook can query the AWS Glue Data Catalog and transform the data without switching tools. The built-in AI assistant generates code and execution plans from natural-language prompts, and the Spark UI shows running work alongside your other runtimes.

To start a session from SageMaker Unified Studio:

  1. Open a Data Notebook in SageMaker Unified Studio.
  2. In the Compute panel, do one of the following:
    1. To create a new cluster, choose Create cluster and configure an Amazon EMR on EC2 cluster.
    2. To use an existing cluster, choose Attach cluster and select a running Amazon EMR on EC2 cluster.
  3. Select the cluster as the notebook’s runtime.
  4. Begin writing PySpark or SQL code in the notebook cells.

For a complete example, open the SageMaker Unified Studio Spark Connect example notebook , which connects a Data Notebook to an Amazon EMR on EC2 cluster and runs PySpark and SQL cells against the AWS Glue Data Catalog.

Watch a walkthrough: Develop in a SageMaker Unified Studio Data Notebook. The preceding steps cover the same workflow, so you can complete it from the notebook without the video.

Option 2: Develop in your own IDE

Use the IDE of your choice, such as Visual Studio Code, PyCharm, Kiro, or a local Jupyter notebook. You debug Spark the way you debug any Python program: set a breakpoint, inspect a variable, and step through your code, all while the Spark work runs on the cluster. Your libraries, source control, and continuous integration and continuous delivery (CI/CD) stay on your local machine, and only your Spark operations are sent to the cluster.

To see this end to end, the following example attaches an IDE to a Spark Connect session and steps through a breakpoint against cluster data.

Open the local IDE Spark Connect example notebook then use the connection steps in the preceding Getting started section to attach your client.

Watch a walkthrough: Develop your own IDE with Spark Connect. The written connection steps in Getting started cover the same workflow, so you can complete it without the video.

Use cases

Spark Connect on Amazon EMR on EC2 supports the following interactive workflows:

  • Interactive extract, transform, and load (ETL) development: Build and test pipelines against full-size data on the cluster, then schedule the same transformations as a Spark step on that cluster, where the Spark version, libraries, and configuration already match what you validated.
  • Exploratory data analysis and feature engineering: Analyze production-scale data from your notebook or IDE instead of sampled subsets, so you catch data quality issues earlier.
  • Notebook-driven analytics in SageMaker Unified Studio: Run PySpark and SQL next to your catalogs and AI tools, switching runtimes per notebook.
  • Apache Iceberg lakehouse analytics: Query and manage Iceberg tables through the AWS Glue Data Catalog, with time travel, schema evolution, and partition management.
  • Compute standardization: Point interactive development at the same clusters that run your production batch jobs, so development and production share one engine and configuration.

Release information

Spark Connect on Amazon EMR on EC2 is available with the AWS runtime for Apache Spark (emr-spark-8.0, Apache Spark 4.0.2) and later. It’s available in all AWS Regions where Amazon EMR is available, except the AWS GovCloud (US) Regions and the China Regions. The SageMaker Unified Studio experience is available in its supported Regions. There’s no additional charge for Spark Connect. You pay for the Amazon Elastic Compute Cloud (Amazon EC2) instances in your cluster. Because these sessions run on your own clusters, they use the Amazon EMR on EC2 capabilities you already rely on, including AWS Graviton processors for price-performance and your choice of On-Demand, Reserved, AWS Savings Plans, or Spot capacity.

Considerations for the release are as follows:

  • The PySpark version that you install locally must match the Apache Spark version on your cluster.
  • Spark Connect supports the DataFrame and SQL APIs. RDD-based APIs aren’t supported.
  • Authentication tokens expire after 1 hour, and sessions end after a configurable idle timeout (60 minutes by default, up to 24 hours).
  • High-availability clusters with multiple primary nodes, Trusted Identity Propagation, and fine-grained access control through AWS Lake Formation aren’t supported for Spark Connect sessions in this release.

Conclusion

Spark Connect on Amazon EMR on EC2 brings interactive, debuggable PySpark development to the clusters you already run. Develop on a SageMaker Unified Studio Data Notebook or in your own IDE, debug against full-size data while the cluster runs the work and share a single cluster across your whole team. To get started, see the Interactive sessions with Spark Connect guide or open a Data Notebook in Amazon SageMaker Unified Studio. To learn more about the service, see the Amazon EMR detail page.


About the authors

Al MS

Al MS

Al is a product manager for Amazon EMR at AWS.

Karthik Prabhakar

Karthik Prabhakar

Karthik is a Data Processing Engines Architect for Amazon EMR at AWS, where he specializes in distributed systems architecture and query optimization. He partners with customers to solve complex performance challenges in large-scale data processing workloads. His work centers on engine internals, cost optimization, and architectural patterns for efficient petabyte-scale analytics.

Arun Prabakaran

Arun Prabakaran

Arun is a Senior Software Engineer working at AWS. His expertise spans distributed data processing and large-scale systems. He is passionate about building reliable data platforms and enabling organizations to run analytics and AI workloads at scale.

Rekha Veeraraghavan

Rekha Veeraraghavan

Rekha is a Technical Account Manager at AWS and a Subject Matter Expert in AWS Analytics. She helps enterprise and strategic customers optimize their data analytics solutions with expert guidance and technical support. Drawing deep data engineering expertise, she enables organizations to build scalable, efficient, and cost-effective data processing pipelines on AWS.

Query unstructured data in Amazon SageMaker Catalog using generative AI

Post Syndicated from Nishchai JM original https://aws.amazon.com/blogs/big-data/query-unstructured-data-in-amazon-sagemaker-catalog-using-generative-ai/

Each day, businesses generate massive amounts of unstructured data, such as PDFs, images, email, customer feedback, and medical reports. But knowing data exists isn’t enough. You need to find it, access it, and extract answers from it fast. In Part 1 of this series, you saw how to set up the producer side of the pipeline: using Amazon Textract and Anthropic Claude on Amazon Bedrock to extract and enrich metadata, and then publish those enriched assets to Amazon SageMaker Catalog so your organization can discover them.

In this post, you take the next step: the consumer side. You sign in as a data consumer, search for and subscribe to the enriched unstructured data assets, and then query them using two approaches. The first is a no-code chat agent for natural language queries. The second is Amazon Bedrock model inference for programmatic access. By the end of this post, you will know how to unlock the business knowledge inside your unstructured data and make it available to analysts and application engineers alike.

Solution overview

This post continues the two-part series architecture, where Amazon SageMaker Catalog acts as the central hub connecting data producers and consumers through a publish-subscribe model.

The consumer workflow picks up after the producer has enriched and published the unstructured data assets. As a consumer, you will:

  • Sign in to your SageMaker Unified Studio consumer project and search the catalog using keywords from the enriched metadata README.
  • Subscribe to the published Amazon Simple Storage Service (Amazon S3) asset and get the subscription approved by the producer.
  • Interact with the subscribed data through two options:
    • Option 1 – A no-code chat agent for natural language queries (NLQs), ideal for data analysts and business users.
    • Option 2 – Amazon Bedrock model inference for programmatic NLQ integration, suited for application engineers building data-driven applications.

The following diagram illustrates the consumer workflow in this solution. The consumer (1) signs in to SageMaker Unified Studio, (2) searches the Amazon SageMaker Catalog for enriched unstructured data assets using keywords from the AI-generated metadata, (3) subscribes to the S3 data asset and receives approval from the producer, and then (4) queries the data using either the Amazon Bedrock chat agent app (Option 1) or Amazon Bedrock model inference through a Jupyter notebook (Option 2).

With both a no-code and a programmatic path, consumers across different roles, from analysts to engineers, can query data in the way that fits their workflow, while the SageMaker Catalog approval workflow maintains governed access throughout.

Consumer workflow architecture: sign in to SageMaker Unified Studio, search the SageMaker Catalog, subscribe to the S3 asset with producer approval, then query with the Amazon Bedrock chat agent or model inference

Figure 1: Consumer workflow for the publish-subscribe solution

Prerequisites

Before you begin, make sure you have completed all steps in Part 1 of this series, including:

Consume published data from the consumer project

In this section, you sign in as a consumer user in the SageMaker Unified Studio consumer project. You then subscribe to the S3 bucket by searching for a keyword that is part of the README published in Part 1.

  1. Sign in to the consumer project and search for the keyword emergency, which was added to the README file during publishing. The search returns the enriched asset that the producer published in Part 1.

    SageMaker Unified Studio catalog search for the emergency keyword, returning the enriched asset published in Part 1

    Figure 2: Catalog search results for the emergency keyword

  2. Choose the asset from the results to view its details, including the AI-generated business metadata, glossary terms, and README content. Then choose Subscribe.

    Asset details page showing AI-generated business metadata, glossary terms, and README content, with the Subscribe button

    Figure 3: Asset details with AI-generated metadata and the Subscribe option

  3. Enter analysis as the Reason for request in the Comment section, then choose Request.

    Subscription request dialog with analysis entered as the reason for request in the Comment box

    Figure 4: Subscription request with the reason for request entered

  4. Sign back in to the producer project (unstructured-producer-project) to approve the subscription request.
  5. After approval, return to the consumer project and confirm that the subscribed asset now appears under Manage, Assets, Subscribed assets.

    Consumer project Subscribed assets list confirming the approved subscription

    Figure 5: Approved subscription under the Subscribed assets tab

With the subscription approved, you can now access the enriched unstructured data through two approaches.

Option 1: As a data or business analyst, you can use the Amazon Bedrock chat agent app for natural language queries.

Option 2: As an application engineer, you can use Amazon Bedrock model inference for programmatic natural language queries.

Let’s explore both options.

Option 1: Amazon Bedrock chat agent app

The Amazon Bedrock chat agent app gives you a no-code, conversational interface to query your enriched unstructured data using natural language. As a data analyst or business user, you can ask questions in plain English. You get answers grounded in the documents your organization has ingested, without writing any code. For production workloads, especially in sensitive domains such as healthcare, you can apply Amazon Bedrock Guardrails to add content filtering and grounding validation to your model responses.

Data scientists and application engineers can also extend these capabilities by integrating the chat agent app APIs into custom applications, so users can interact with unstructured Amazon S3 data programmatically.

To set up the Amazon Bedrock chat agent app on your subscribed dataset, complete the following steps.

Prerequisite: Add the S3 data location.

Before creating the chat agent app, you need to add the S3 location of your subscribed data as a registered location in your project.

  1. Choose the Data tab in Overview.
  2. Choose the S3 bucket, and then choose Add to add the S3 location.

    Data tab in the project Overview with the S3 bucket selected and the Add button to register the S3 location

    Figure 6: Adding the S3 location from the Data tab

  3. On the S3 location page, provide the following details:
    • Add a name: producerprojectdata.
    • Add the producer’s S3 path as a new S3 location: s3://amzn-sagemaker-bucket-<domain-id>-<project-id>/medical/.

    Note: You can get the S3 location details from the technical name of your subscribed asset.

    • Choose the AWS Region, and then choose Add data to add this as a new location.

    Note: Make sure the AWS Region you select supports the Amazon Bedrock foundation models used later in this post. For a list of available models by Region, see Supported Regions and models for Amazon Bedrock.

    S3 location page with the location name, producer S3 path, and AWS Region entered before choosing Add data

    Figure 7: S3 location details and AWS Region selection

    Note: Make sure to select only the PDF files within the S3 path for the data source.

    Data source selection showing only the PDF files within the S3 path selected

    Figure 8: Selecting the PDF files as the data source

After the location is added, it appears as a selectable S3 location when creating a knowledge base in AI Apps.

Complete the following steps to configure the chat agent app:

  1. In the left navigation pane, under Generative AI, choose AI Apps.
  2. In the Build section of the page, choose Chat agent.

    AI Apps Build section with Chat agent selected in the left navigation under Generative AI

    Figure 9: Choosing Chat agent in the AI Apps Build section

  3. Expand the Data tab to create a knowledge base with your S3 bucket. On the Create a new knowledge base page, enter the following:
    • Add a name: MedicalKB.
    • Add a description: Knowledge base built from subscribed medical S3 data assets. Contains medical documents used to provide grounded, context-aware responses to medical domain queries.
    • Choose the data source. You will see the S3 bucket that you added in the previous step.
    Create a new knowledge base page with the MedicalKB name, description, and the added S3 bucket as the data source

    Figure 10: Creating the MedicalKB knowledge base from the S3 data source

  4. Choose your embedding model. You can leave the default settings and choose Create. It might take 10–15 minutes to create the knowledge base, depending on file sizes.
  5. After the knowledge base is created, on the Chat agent page:
    • Choose your preferred model from the Model menu (you can switch between different large language models as needed).
    • Under Data, choose your published S3 bucket as the knowledge base.
    • Begin interacting with the agent by entering questions in the Enter prompt field.
    Chat agent page with a model selected and the MedicalKB knowledge base chosen, ready to enter a prompt

    Figure 11: Chat agent page with the model and knowledge base selected

For example, entering “Which age groups had the highest rates of emergency department visits for tooth disorders?” returns an answer grounded in the enriched dental dataset published in Part 1.

The chat agent uses the enriched README metadata along with the underlying documents to surface contextually relevant answers. Analysts can explore unstructured content without needing to know where the data lives or how it’s structured.

Option 2: Natural language queries using Amazon Bedrock model inference

This option demonstrates how to use Amazon Bedrock model inference to query subscribed data using natural language. You can integrate this capability with external chat applications so users can run natural language queries through Amazon Bedrock.

  1. In your consumer project, choose Manage, Assets from the bottom of the left navigation pane. On the Subscribed tab, choose your subscribed S3 asset. Under Actions, choose Open JupyterLab notebook.

    Subscribed S3 asset Actions menu with Open JupyterLab notebook selected in the consumer project

    Figure 12: Opening the JupyterLab notebook from the subscribed asset

  2. This opens the JupyterLab notebook environment. Upload the s3_document_consumer_v2.ipynb notebook and run all the cells. You can download the notebook from s3_document_consumer_v2.ipynb.Note: The project role requires permissions for Amazon S3, Amazon Textract, and Amazon Bedrock. If you followed Part 1, you might already have these policies attached. For details on the required policies and guidance, see the prerequisites in Part 1.
  3. Review the notebook cells.
    JupyterLab notebook cells with the final cell showing a sample question answered by Amazon Bedrock

    Figure 13: Sample question answered by Amazon Bedrock in the notebook

    In the final cell, you find a sample question that Amazon Bedrock answers: “Which primary payer types (Medicare, Medicaid, private insurance, and so on) account for the highest proportion of dental-related emergency department visits?”

    Amazon Bedrock processes the question against the enriched content in the S3 bucket and returns a grounded answer. You can replace this sample question with any query relevant to your documents.

The Amazon Bedrock model inference approach gives you programmatic control, making it possible to embed natural language query capabilities directly into your existing data applications and business intelligence tools.

Clean up

To avoid ongoing charges, make sure to delete the resources used in this solution immediately after completing the walkthrough. The primary cost drivers are SageMaker Unified Studio notebook instances, Amazon Bedrock model inference calls, and Amazon S3 storage.

  1. Stop SageMaker Unified Studio resources:
    • Close running notebooks.
    • Stop running notebook instances.
    • Shut down unused kernels.

    Note: Running notebook instances continue to incur charges even when not in use.

  2. Clean Amazon S3 storage:
    • Delete temporary files created during processing.
    • Remove uploaded test documents that are no longer needed.

    Note: Although Amazon S3 costs are minimal, large volumes of data can accumulate significant charges, so it’s best to remove unneeded data.

Conclusion

In this post, you saw how to consume and query the enriched unstructured data assets published in Part 1 of this series. By subscribing to assets through the Amazon SageMaker Catalog publish-subscribe model, you can discover, access, and interact with your organization’s unstructured data, whether through the no-code chat agent or Amazon Bedrock model inference.

Together, both parts of this series show you how to build a comprehensive pipeline that transforms raw unstructured documents into governed, queryable knowledge assets. The combination of Amazon Textract for extraction, Amazon Bedrock for intelligent summarization and NLQ, and Amazon SageMaker Catalog for governance and discoverability means your teams can focus on extracting business insights rather than managing infrastructure.

To continue your Amazon SageMaker journey, see the following resources:


About the authors

Nishchai JM

Nishchai JM

Nishchai is an Analytics and generative AI Specialist Solutions Architect at Amazon Web Services. He specializes in building larger scale distributed applications and helps customers modernize their workloads on AWS. He thinks Data is new oil and spends most of his time deriving insights from data.

KiKi Nwangwu

KiKi Nwangwu

KiKi is an Analytics and generative AI Specialist Solutions Architect at AWS. She specializes in helping customers architect, build, and modernize scalable data analytics and generative AI solutions. She enjoys traveling and exploring new cultures.

Narendra Gupta

Narendra Gupta

Narendra is a Sr. Specialist Solutions Architect for Data & AI (Analytics) at AWS. He works with customers to design data-driven solutions and has deep expertise in data governance and cataloging.

Aditya Edara

Aditya Edara

Aditya is a Support Engineer at AWS. He serves as a Subject Matter Expert in AWS Analytics services, specializing in Amazon EMR and AWS Glue. Aditya provides expert guidance and technical support to enterprise and strategic customers, helping them optimize data analytics solutions.

Adding custom domains to AWS Lambda MicroVMs with Application Load Balancer

Post Syndicated from Frank Scarfo original https://aws.amazon.com/blogs/compute/adding-custom-domains-to-aws-lambda-microvms-with-application-load-balancer/

AWS Lambda MicroVMs is a serverless compute building block that provides VM-level isolation, near-instant startup performance, and state retention. You can now give each user or job their own execution environment to securely run just-in-time code, whether user or AI-generated. You do this without managing virtualization infrastructure or choosing between isolation, speed, and state retention. Lambda MicroVMs are powered by Firecracker virtualization, the technology underpinning AWS Lambda.

When you run a workload on AWS Lambda MicroVMs, each MicroVM is reachable at a service-generated endpoint that looks like 92cfc7f9-….lambda-microvm-….on.aws. That works, but many teams want to expose their MicroVMs under a domain they own, such as 92cfc7f9-….microvms.example.com. When a browser is the client, they also want to satisfy cross-origin resource sharing (CORS) without changing the application inside the MicroVM.

Both are achievable today, entirely from load-balancing and networking primitives. There is no Amazon CloudFront distribution and no compute in the request path. All you need is an Application Load Balancer (ALB) that terminates TLS with your AWS Certificate Manager (ACM) certificate, rewrites the Host header, and forwards the request over AWS PrivateLink. In this post you’ll deploy that pattern with the AWS Cloud Development Kit (AWS CDK), map a wildcard of custom domains onto your MicroVMs, and let the ALB handle CORS for you.

The complete, deployable example is available as a pattern on Serverless Land. This walkthrough centers on the reusable networking pattern. The sample also includes a small demo application that provisions a MicroVM and mints an access token, which we reference but do not detail here.

What you’ll build

By the end you’ll have:

  • A wildcard custom domain like *.microvms.example.com, where each <uuid>.microvms.example.com maps transparently to the corresponding MicroVM.
  • An internet-facing ALB that rewrites the incoming request’s Host header to the real MicroVM endpoint and forwards requests to it privately over PrivateLink.
  • CORS preflight and response headers handled at the ALB, with no change to the code running in the MicroVM.

Calling https://<uuid>.microvms.example.com/<path> (with the MicroVM access headers described later) reaches the right MicroVM, with your domain intact end to end.

Solution overview

The request flow looks like this:

Request flow from a browser through the Application Load Balancer, which terminates TLS and rewrites the Host header, then forwards over AWS PrivateLink to the Lambda MicroVM service.

The key component is the ALB host header rewrite, introduced in URL and host header rewrite for Application Load Balancers. A listener rule matches the incoming custom host with a regex condition, captures the MicroVM ID from the left-most label, and a host-header-rewrite transform rewrites the Host header to <uuid>.lambda-microvm.<region>.on.aws before forwarding. Because the MicroVM service front-end routes on the Host header, the request lands on the correct MicroVM, while the customer’s domain stays in the browser’s address bar the whole time.

Why not CloudFront? Why not an ALB redirect?

  • CloudFront can also rewrite Host/SNI toward the origin, but a single distribution has static origins. Mapping a wildcard of MicroVM IDs through one distribution would require a CloudFront Function to compute the origin per request. The ALB transform performs the same rewrite for the entire wildcard with zero code.
  • An ALB redirect action only issues an HTTP 301 Moved Permanently/302 Found response. The browser would follow it, and the address bar would then show the .on.aws URL, which breaks our design as it is not a real custom domain. The transform (not a redirect) is what makes the custom domain transparent.

Walkthrough

The example is an AWS CDK application. Configuration lives under the microvm-custom-domains key in cdk.json (hosted zone, wildcard base, the endpoint base to rewrite to, the PrivateLink service name, and the CORS origin). Set those values, then deploy. The sections below explain what the stack creates and why.

Prerequisites

A small VPC (two Availability Zones, which is the minimum for an internet-facing ALB) hosts the ALB and an interface VPC endpoint to the AWS managed MicroVM service. There are no NAT gateways, because nothing here needs egress, which keeps the footprint lean.

// Interface (PrivateLink) endpoint to the AWS managed MicroVM service.
const endpoint = new ec2.InterfaceVpcEndpoint(this, 'MicroVmEndpoint', {
  vpc,
  service: new ec2.InterfaceVpcEndpointService(cfg.microvmVpceServiceName, 443),
  subnets: { subnetType: ec2.SubnetType.PRIVATE_ISOLATED },
});

2. Discover the endpoint’s private IP addresses at deploy time

An ALB IP target group needs the private ENI IP addresses of the interface endpoint (one per Availability Zone). CloudFormation does not expose those IPs as a usable attribute, so the stack resolves them during deployment with an AwsCustomResource that reads the endpoint’s own ENIs by ID (DescribeNetworkInterfaces on vpcEndpointNetworkInterfaceIds).

This is the only compute the package deploys, it runs only during cdk deploy, and it is never in the request path.

3. Request a wildcard TLS certificate

ACM issues a DNS-validated wildcard certificate for *.microvms.example.com, validated through the hosted zone you imported. The ALB presents this certificate for every custom domain under the wildcard.

4. Create the ALB and the MicroVM target group

The internet-facing ALB has an HTTPS:443 listener using the wildcard certificate. The target group holds the endpoint ENI IPs as IP targets, reached over HTTPS:443.

  • Encrypted in transit. A customer-provided AWS Certificate Manager (ACM) certificate is used to securely terminate encryption between the client and the ALB. The ALB re-originates TLS to the MicroVM service so traffic stays encrypted through the network.
  • IP-based targets. The target group uses IP-based targets with the local IP addresses of the VPC endpoints.
  • Health check matcher 200,403,404. The load balancer’s health probes are unauthenticated, so the MicroVM endpoint answers them with 403. A 403 here means “endpoint is reachable,” not “auth is broken,” so the matcher treats it as healthy.
const targetGroup = new elbv2.ApplicationTargetGroup(this, 'MicroVmTargets', {
  vpc,
  protocol: elbv2.ApplicationProtocol.HTTPS,
  port: 443,
  targetType: elbv2.TargetType.IP,
  targets: targetIps.map((ip) => new elbv2t.IpTarget(ip, 443)),
  healthCheck: {
    protocol: elbv2.Protocol.HTTPS,
    path: '/',
    healthyHttpCodes: '200,403,404',
  },
});

const listener = alb.addListener('Https', {
  port: 443,
  protocol: elbv2.ApplicationProtocol.HTTPS,
  certificates: [certificate],
  // Default action for anything that doesn't match our host regex.
  defaultAction: elbv2.ListenerAction.fixedResponse(404, {
    contentType: 'text/plain',
    messageBody: 'Unknown custom domain',
  }),
});

5. Add the host-header rewrite rule

A listener rule matches <uuid>.microvms.example.com with a regex condition and rewrites the Host header to <uuid>.lambda-microvm.<region>.on.aws with a host-header-rewrite transform. The regex captures the left-most label (the MicroVM ID) and reuses it in the replacement.

At the time of writing, the CDK L2 constructs don’t yet model regex host conditions or transforms, so the example reaches the underlying CfnListenerRule to set them:

const escapedBase = customDomainBase.replace(/[.]/g, '\\.');
const matchRegex = `^(.+)\\.${escapedBase}$`;     // capture <uuid>
const replaceWith = `$1.${microvmEndpointBase}`;   // <uuid>.lambda-microvm.<region>.on.aws

const cfnRule = forwardingRule.node.defaultChild as elbv2.CfnListenerRule;

cfnRule.conditions = [{ field: 'host-header', regexValues: [matchRegex] }];

cfnRule.addPropertyOverride('Transforms', [
  {
    Type: 'host-header-rewrite',
    HostHeaderRewriteConfig: { Rewrites: [{ Regex: matchRegex, Replace: replaceWith }] },
  },
]);

6. Point Route 53 at the ALB

Wildcard A and AAAA alias records (*.microvms.example.com) target the ALB, so every MicroVM custom subdomain resolves to it.

7. Deploy

Run the following commands to install the dependencies and then deploy the application.

npm install
npx cdk deploy

Handling CORS at the ALB

If your clients are browsers calling the MicroVM from another origin, CORS is handled entirely at the ALB, with no change to the application inside the MicroVM.

The listener uses ALB header-modification attributes to insert the Access-Control-Allow-* headers on every response. A higher-priority rule answers OPTIONS preflight requests at the edge with a fast 204 response. Otherwise, preflight requests would reach the origin and be rejected without an access token.

// Insert CORS headers on every response on this listener.
const cfnListener = listener.node.defaultChild as elbv2.CfnListener;
cfnListener.addPropertyOverride('ListenerAttributes', [
  { Key: 'routing.http.response.access_control_allow_origin.header_value',  Value: cfg.corsAllowOrigin },
  { Key: 'routing.http.response.access_control_allow_methods.header_value', Value: 'GET,POST,PUT,DELETE,OPTIONS,PATCH,HEAD' },
  { Key: 'routing.http.response.access_control_allow_headers.header_value', Value: 'x-aws-proxy-auth,x-aws-proxy-port,content-type,authorization' },
  { Key: 'routing.http.response.access_control_expose_headers.header_value', Value: 'content-type,content-length' },
  { Key: 'routing.http.response.access_control_max_age.header_value',        Value: '86400' },
]);

// Answer OPTIONS preflights at the ALB.
new elbv2.ApplicationListenerRule(this, 'CorsPreflightRule', {
  listener,
  priority: 10,
  conditions: [elbv2.ListenerCondition.httpRequestMethods(['OPTIONS'])],
  action: elbv2.ListenerAction.fixedResponse(204, { contentType: 'text/plain', messageBody: '' }),
});

Because the ALB adds those headers to both the preflight 204 and the forwarded MicroVM response, a browser’s cross-origin call succeeds without any application change. Set corsAllowOrigin to * for quick testing, and pin it to your own site for anything beyond a demo.

Test it end to end

First, launch a Lambda MicroVM and mint an access token (follow Create your first Lambda MicroVM). When it’s running, the service gives you a generated endpoint that looks like:

012345678-9abc-defg.lambda-microvm.us-east-2.on.aws

To get the custom-domain equivalent, replace the endpoint suffix (.lambda-microvm.<region>.on.aws) with your wildcard base: .microvms.example.com. Everything ahead of that suffix is preserved exactly:

012345678-9abc-defg.microvms.example.com

The ALB’s rewrite rule captures whatever precedes the suffix and re-attaches it to the real endpoint base, so the mapping holds for the entire wildcard. You never register anything per-MicroVM.

With your token in hand, call the custom domain you derived:

curl "https://012345678-9abc-defg.microvms.example.com/<path>" \
  -H "X-aws-proxy-auth: <token>" \
  -H "X-aws-proxy-port: 8080"

The request travels to the ALB, which terminates TLS, rewrites the host header, and forwards over PrivateLink to the MicroVM. The response comes back under your domain.

The reference architecture also includes a single-page demo and a POST /api/provision endpoint that runs or reuses a MicroVM and mints a short-lived token. With it, you can try the flow without wiring up token creation yourself. It even performs this suffix swap for you and hands back a ready-to-click custom-domain URL. See the repository for that piece.

Important considerations

  • Authentication is still the client’s job. This pattern only rewrites Host. The client must still supply a valid, unexpired access token in X-aws-proxy-auth. This is deliberate. MicroVM tokens are per-MicroVM and short-lived, so baking them into infrastructure would be fragile and insecure.
  • Region pinning. PrivateLink is regional, so the ALB, the endpoint, and the MicroVM service must all be in the same Region.
  • Production hardening. If you adapt the sample’s provisioning endpoint, put authentication and rate limiting in front of it, pin CORS to your origin, and scope IAM to the minimum. The sample’s provisioning path is intentionally open for demonstration and is not production-safe as written.
  • Cost. You pay for the ALB and the interface endpoint (hourly plus data processing) in addition to the Lambda MicroVM usage. There is no CloudFront distribution and no per-request compute in the data path.

Clean up

Run the following command in the same directory where you deployed the application from.

npx cdk destroy

This removes the ALB, target groups, endpoint, certificate, VPC, and Route 53 records created by the stack.

Conclusion

You can front AWS Lambda MicroVMs with customer-owned wildcard custom domains using an Application Load Balancer and AWS PrivateLink. The key is the ALB’s host-header rewrite. Because the MicroVM service routes requests based on the Host header, a single rewrite rule can transparently map an entire wildcard of custom domains onto your MicroVMs. CORS is handled at the edge as well. The whole setup relies only on networking primitives, with no CloudFront distribution and no compute in the request path.

To try it yourself, deploy the reference architecture and review the ALB URL and host header rewrite launch post for more on the transform feature.

Running self-hosted AI agent sandboxes with AWS Lambda MicroVMs

Post Syndicated from Brian Krygsman original https://aws.amazon.com/blogs/compute/running-self-hosted-ai-agent-sandboxes-with-aws-lambda-microvms/

Organizations are building AI agents that autonomously write code, query databases, and interact with internal systems on behalf of their teams. These agents handle use cases such as automated code review, data pipeline optimization, and infrastructure troubleshooting. When your AI agent generates a shell command, queries a database, or writes to a file system, that code needs a secure environment to run in. Without isolation, one session’s tool calls can contaminate another session’s state, inadvertently expose sensitive data across tenants, or unintentionally allow untrusted code to reach production resources. Self-hosted sandboxes solve this by keeping agent execution within your own AWS account, giving you full control over networking, secrets, and governance.

Say you’re building an internal AI agent that optimizes database queries for your engineering team. A developer asks it to find the ten slowest queries in your analytics database, rewrite them with better indexing, and test the results. That’s three tool calls in a single session. One hits a live database with real credentials. One generates code. One executes it. Now multiply that by fifty developers using the assistant at the same time. Each session needs its own credentials, its own filesystem, its own network boundary. If credentials or state cross session boundaries, you have inadvertent data exposure.

AWS Lambda MicroVMs is a serverless compute environment that provides general-purpose runtimes with the strong isolation of virtual machines and the rapid scaling of AWS Lambda. Powered by Firecracker virtualization, each MicroVM runs Amazon Linux with full OS access for up to 8 hours. You launch, suspend, resume, and terminate MicroVMs programmatically. You get the serverless benefits of managed infrastructure, responsive scaling, and pay-per-use pricing. Three capabilities make Lambda MicroVMs a strong fit for agent sandboxes:

  • VM-level isolation per environment: Each MicroVM runs in its own Firecracker virtual machine, providing hardware-virtualization-based isolation between sessions without the resource overhead and startup time required of full VMs. One developer cannot see a teammate’s session, even when both run at the same time.
  • Launch from snapshot: Like Lambda SnapStart, MicroVMs boot from a pre-captured memory and disk snapshot, skipping application initialization entirely. Your agent gets a near-instant ready-to-use environment.
  • 4x vertical scaling without re-provisioning: A running MicroVM can scale CPU and memory up to 4x its initial allocation, which can range from 0.25 vCPU/0.5 GB to 4 vCPU/8 GB, without terminating or re-creating the environment. If the agent needs to run a heavy data transformation mid-session, it can get more resources without starting over.

In this post, we show you how to architect and build a self-hosted AI agent that uses Lambda MicroVMs as secure, isolated sandboxes for tool-call execution. Lambda MicroVMs can handle the compute isolation for running tool calls, while the host for production AI agents, such as Amazon Bedrock AgentCore, manages the agent logic, model routing, and session state. A complete reference solution is available in aws-samples.

How self-hosted sandboxes work

A developer asks the agent to “find the ten slowest queries in our analytics database and suggest index improvements.” The agent orchestration system starts a session then breaks the objective into tool calls and distributes them. A worker needs to pick up that session, run the queries, and return results.

Most AI agent orchestration services and frameworks use a work queue model to distribute tool-call execution. The orchestration service enqueues sessions representing tool-call work. A worker, the process that claims a session and executes its tool calls, runs inside a compute environment, posts results, and exits. In this architecture, each Lambda MicroVM is the compute environment, and the worker is the process running inside it. Claude Managed Agents self-hosted sandboxes run those workers inside your own infrastructure rather than on a shared, multi-tenant compute pool. Your database credentials stay in your virtual private cloud (VPC). Your network, introspection, and governance rules apply.

You can trigger workers in two ways:

  • Webhook-triggered: The orchestration application sends a notification when a session is ready. Your control plane launches a worker on demand.
  • Always-on: A long-running process continuously polls the work queue for new sessions.

The Lambda MicroVMs lifecycle aligns with the webhook-triggered pattern, where each session produces one inbound event that launches a fresh MicroVM. Lambda MicroVMs support configurable idle policies. After a configurable idle period, a MicroVM suspends automatically, preserving disk and memory state. It resumes when inbound traffic arrives or when you call the resume API. The MicroVM runs for the duration of the session, the worker exits, and the idle policy suspends then finally terminates the VM. Lifecycle hooks allow you to run custom logic at key steps in the MicroVM lifecycle.

In contrast, the always-on pattern risks breaking the polling loop by suspending the MicroVM when idle, since there’s no inbound traffic between sessions. You could disable the configurable idle period, but then you pay for empty polling. Use the webhook-triggered approach for self-hosted sandboxes on Lambda MicroVMs.

Architecture

The following figure shows the reference solution’s architecture, with the Anthropic agent orchestration service control plane on the left interacting with a self-hosted sandbox environment in AWS on the right.

Reference architecture showing the Anthropic orchestration control plane sending a webhook through API Gateway to a launcher Lambda function that starts a MicroVM worker in your AWS account

Figure 1: Reference architecture for self-hosted AI agent sandboxes on Lambda MicroVMs

The sample architecture is event-driven. The only inbound traffic is the webhook call. When the event arrives, the handler launches a MicroVM. Once launched, the MicroVM pulls its assigned session from the orchestration system’s work queue and runs the task. In our example, the developer’s “find slow queries” request has been queued as a session. The agent now needs to reach your infrastructure, spin up an isolated environment, and hand off the work. The following sequence shows how each component interacts to fulfill a single session.

The orchestration service queues work as sessions. A MicroVM launches to service each session, and the worker is the process running inside that MicroVM that claims the session, executes tool calls, and returns results.

  1. Once the orchestration service marks a session as ready to run, it sends a session.status_run_started webhook to an Amazon API Gateway endpoint, triggering a MicroVM launch.
  2. The launcher verifies the webhook signature using a signing secret from AWS Systems Manager Parameter Store, rejecting invalid or stale deliveries before spending compute.
  3. The launcher calls RunMicrovm, passing the session ID and a secret reference through runHookPayload. It deduplicates on the webhook event ID (backed by Amazon DynamoDB) so retries do not launch duplicate VMs.
  4. The MicroVM boots from a pre-captured Firecracker snapshot and receives the dispatch on its /run lifecycle hook. The worker fetches the environment key from Parameter Store using its execution role. It pulls the matching session from the work queue, claims it, and executes tool calls in an isolated /workspace directory. When finished, it posts results and exits. The idle policy suspends then terminates the VM.

Deduplication. The webhook event ID serves as the idempotency key. The launcher uses Powertools for AWS Lambda (Python) with a DynamoDB persistence layer to verify exactly-once processing. If the orchestration application retries a delivery with the same event ID, Powertools protects the system from launching extra MicroVMs and doing extra work.

Credential boundaries. Each component accesses only the single secret it needs. The launcher reads only the webhook signing secret to verify inbound events. It passes only an ARN reference to the environment key into the MicroVM payload. The MicroVM’s execution role retrieves only that environment key at runtime. No single component holds both secrets.

Component Has access to
Launcher Lambda Webhook signing secret (verify inbound events)
MicroVM worker Environment key (through the execution role, to poll and claim sessions)

Cost model. You pay for MicroVM run time per session, plus standard charges for API Gateway requests, Parameter Store API calls, and Lambda invocations for the launcher. When no sessions are active, no MicroVMs run. Cost scales with concurrent sessions and their duration, avoiding idle compute charges.

Implementation

The following sections explore the reference architecture in more depth.

Project structure

The reference solution uses AWS Serverless Application Model (AWS SAM) for infrastructure-as-code. Alternatively, if you use an AI coding agent such as Claude Code, Kiro, or Cursor, the Agent Toolkit for AWS includes a Lambda MicroVMs skill that gives your agent the procedures to provision, configure, and deploy MicroVM-based sandbox environments on your behalf.

├── template.yaml                    # SAM: launcher, API, WAF, secrets, roles
├── src/
│   ├── functions/launcher.py        # Verify signature, RunMicrovm
│   ├── microvm-image/
│   │   ├── Dockerfile               # AL2023 + Node.js worker
│   │   └── worker/worker.mjs        # Lifecycle hook server
│   └── scripts/build-image.sh       # Package + create MicroVM image

Launcher: verify the webhook before spinning up compute

When the webhook arrives saying a developer’s session is ready, the launcher’s first action is signature verification. If it fails, the function returns 401 immediately. No MicroVM launches. No DynamoDB writes. You don’t pay for fraudulent or replayed requests.

signing_secret = [REDACTED_PASSWORD]  # Verify webhook before spending compute
if not verify_signature(raw_body, headers, signing_secret):
    return {"statusCode": 401, "body": "invalid signature"}

After verification, the launcher builds a dispatch payload containing the session ID, environment ID, region, and an ARN reference to the environment key secret. It passes this to RunMicrovm through runHookPayload:

launched = microvm_client.run_microvm(
    image_identifier="arn:aws:lambda:us-east-1:123456789012:microvm-image:worker",
    run_hook_payload=json.dumps({"session": dispatch}),
    execution_role_arn=config.execution_role_arn,
    maximum_duration_in_seconds=28800,
    ingress_network_connectors=["arn:aws:lambda:::network-connector:aws-network-connector:ALL_INGRESS"],
    egress_network_connectors=["arn:aws:lambda:::network-connector:aws-network-connector:INTERNET_EGRESS"],
)

MicroVM worker: claim one session, execute, exit

The MicroVM image is built from a Firecracker snapshot. The worker process starts during image creation and is captured in the snapshot, so there is no application startup at run time. The /run lifecycle hook delivers the dispatch payload:

// POST /aws/lambda-microvms/runtime/v1/run
case "run": {
    const envelope = JSON.parse(rawBody);
    const dispatch = JSON.parse(envelope.runHookPayload);
    res.writeHead(200); // Acknowledge hook immediately
    res.end();
    const key = await fetchParameter(dispatch.session.ENVIRONMENT_KEY_PARAM_NAME);
    await pollAndHandleSession(dispatch.session.ANTHROPIC_SESSION_ID, key);
    // Session complete; terminate this MicroVM to release all resources
    await terminateMicroVm(envelope.microvmId);
}

The worker acknowledges the hook within its timeout, fetches the environment key, and claims the session. This is where the requested work begins. The worker connects to the analytics database, runs EXPLAIN ANALYZE on the flagged queries, writes optimized alternatives to /workspace/suggestions.sql, and posts the results back to the developer. All of that happens inside this single VM. When the session completes, the worker calls terminate-microvm to release all compute resources.

Deployment

For full deployment instructions, see the reference solution README. Before deploying, make sure you have these prerequisites.

Prerequisites

Four steps

  1. Deploy the control plane. Build and deploy the SAM stack, which creates the launcher Lambda, API Gateway endpoint, WAF WebACL, DynamoDB idempotency table, Parameter Store entries, and MicroVM execution role.
    sam build
    sam deploy --guided --capabilities CAPABILITY_NAMED_IAM

  2. Register the webhook and populate secrets. In the Claude Console, register the stack’s WebhookUrl output as a webhook endpoint subscribed to session.status_run_started. Store the signing secret and environment key in the Parameter Store resources created by the stack.
  3. Build the MicroVM image. Package the Dockerfile and worker code, upload to Amazon S3, and create the image. The service runs your Dockerfile, launches the worker, and captures a Firecracker snapshot. Monitor build progress in Amazon CloudWatch under /aws/lambda/microvms/<image-name>.
    ./src/scripts/build-image.sh

  4. Verify. Create a test session and confirm a MicroVM launches and completes end-to-end. The reference solution includes a verification script that creates a session, triggers the webhook, and validates the full flow.

Using Claude Platform on AWS (CPOA)

The preceding architecture works similarly when you access Claude through Claude Platform on AWS rather than the first-party API. Three things change in the worker:

  1. Client initialization. Replace the first-party client with the AWS client and supply your workspace ID:
    from anthropic import AnthropicAWS
    
    client = AnthropicAWS(aws_region="us-east-1")
    
    # Workspace ID is required on every request
    # Set via ANTHROPIC_AWS_WORKSPACE_ID env var or pass per-call

  2. Authentication options. CPOA supports two modes:
    1. CPOA API key (aws-external-anthropic-api-key-...): Store it in Parameter Store the same way as the first-party environment key. These keys are short-lived (12-hour STS tokens) and must be regenerated when they expire.
    2. SigV4 (IAM): The MicroVM execution role can sign requests directly, so there is no secret to store or rotate. Set the environment key secret to a placeholder value (for example, use-sigv4) and the SDK falls through to IAM credentials automatically. This is the recommended path for production.

In both authentication modes, attach the AWS managed policy AnthropicSelfHostedEnvironmentAccess to the MicroVM execution role. This policy grants the aws-external-anthropic actions needed to poll the work queue, claim sessions, and post results. See IAM actions for Claude Platform on AWS for the full reference.

Prerequisite: Enable outbound web identity federation once per AWS account:

aws iam enable-outbound-web-identity-federation

Everything else, including webhook verification, deduplication, credential separation, and idle policy remains the same.

Security

Earlier we talked about what goes wrong without isolation. Credentials exposed between sessions. Scripts unintentionally reaching production. Agents escaping their sandbox. This architecture implements defense in depth to help prevent these.

Each component accesses a single, scoped secret. The launcher passes only an ARN reference to the worker credential into the MicroVM. The MicroVM’s execution role retrieves only that credential at runtime. The analytics database connection string does not touch the launcher and does not leave your environment.

AWS WAF applies managed rule sets (OWASP, known bad inputs, IP reputation) and per-IP rate limiting. Amazon API Gateway request validation rejects malformed bodies. The launcher performs HMAC signature verification as the true authentication boundary.

Each session runs in its own MicroVM. Sessions do not share memory, disk, or network namespaces. Firecracker provides hardware-virtualization-based isolation. The launcher IAM role reads only the signing secret. The MicroVM execution role reads only the worker credential. Both are scoped to specific Parameter Store ARNs. The Amazon S3 artifact bucket blocks public access, enables versioning, and uses server-side encryption.

Conclusion

This post walked through how to give your internal AI agent a safe place to run database queries, generate code, and execute scripts on behalf of fifty developers without leaking data between sessions or reaching resources it shouldn’t.

AWS Lambda MicroVMs provide ephemeral, VM-isolated compute environments that align with the per-session execution model of AI agent sandboxes. Snapshot-based launch avoids application startup latency. Idle policies terminate VMs once sessions complete. Firecracker isolation verifies that sessions do not share state. You pay only for active execution time and maintain full control over credentials, networking, and governance within your AWS boundary.

You build and operate a serverless control plane. You get per-session VM isolation with no idle compute cost and no shared tenancy.

To get started, explore these resources: