All posts by Antoine Boucherie

Running multi-day AZ evacuation drills with ARC Zonal Shift

Post Syndicated from Antoine Boucherie original https://aws.amazon.com/blogs/architecture/running-multi-day-az-evacuation-drills-with-arc-zonal-shift/

Running a multi-day Availability Zone evacuation drill with ARC Zonal Shift is an effective way to prove your application can withstand a sustained impairment. Multi-Availability Zone deployment is an architectural best practice for building resilient applications on AWS, but there is a gap between deploying multi-AZ and proving it works under sustained stress. Traditional disaster recovery (DR) tests validate the failover mechanism. A typical test shifts traffic, confirms targets respond, and rolls back within minutes. These short exercises don’t surface the issues that only appear over hours or days.

A multi-day AZ evacuation forces time-dependent behaviors to play out completely, exposing failure modes that brief tests miss:

  • Auto Scaling policies that aren’t tuned for sustained N-1 operation over a full day.
  • Deployment pipelines that don’t validate AZ health before placing new workloads.
  • Stale DNS or cached database endpoints.
  • Time-based routine operational processes tested against an N-1 architecture (certificate rotations, credential and secret rotations, maintenance windows, log rotation, backup automation, and cron jobs).
  • Long-lived database connections pinned to a specific AZ that are only used infrequently.
  • Recovery after a sustained multi-day shift, which is a different operational procedure than rolling back to a warm, nearly identical AZ within minutes.

By shifting all traffic away from a single AZ for 48–72 hours using Amazon Application Recovery Controller (ARC) Zonal Shift, you force your architecture to sustain full production load on N-1 zones. This proves capacity sufficiency, database stability, and client reconnection behavior, but most importantly, that your teams can operate normally for days on N-1 capacity.

This post shows you how to plan and run a multi-day AZ evacuation drill across Amazon Elastic Container Service (Amazon ECS), Amazon Elastic Kubernetes Service (Amazon EKS), Amazon Relational Database Service (Amazon RDS) for PostgreSQL, and Amazon Aurora PostgreSQL, with step-by-step CLI commands, a prerequisites section, observability metrics, and restore procedures.

Why financial services institutions are doing this already

Financial services organizations face unique regulatory pressure to demonstrate, not just document, their disaster recovery capabilities. Across the world, there is an increasing focus on building and demonstrating operational resilience within regulated entities. This is shifting the mindset from “show us your runbook” to “show us the evidence”. This uplift in control environments is driving financial services companies to conduct live DR testing under realistic conditions and produce auditable proof of recovery within defined timeframes. Some insurers and banks are now running periodic AZ evacuation drills as part of their operational resilience programs, shifting from “we have multi-AZ” to “we have proven multi-AZ”. Using ARC Zonal Shift, you can shift traffic at the infrastructure layer without changing application code. It works natively across Application Load Balancer, Network Load Balancer, Amazon Elastic Compute Cloud (Amazon EC2) Auto Scaling groups, and Amazon EKS clusters.

Solution overview

In this walkthrough, we demonstrate how to evacuate an Availability Zone for a multi-tier digital platform running on AWS.

We deliberately include both ECS and EKS, and two database engines, to show the evacuation procedure for each major service type readers are likely to run. The architecture is illustrative only.

The following table outlines the architecture:

Layer Components Multi-AZ Configuration
Traffic ingress Application Load Balancer (ALB) fronting ECS Deployed across 3 AZs, cross-zone load balancing activated
Compute (containers) Amazon ECS (Fargate) Stateless tasks distributed across 3 AZ subnets
Traffic ingress Network Load Balancer (NLB) fronting EKS Deployed across 3 AZs, cross-zone load balancing activated
Compute (Kubernetes) Amazon EKS or EKS Auto Mode Stateless services with topology spread constraints across 3 AZs
Database Amazon RDS for PostgreSQL Multi-AZ: primary in AZ A, standby in AZ B
Database Amazon Aurora PostgreSQL Writer in AZ A, reader in AZ B. Storage replicated across all 3 AZs

Note that the Aurora storage layer differs from standard RDS Multi-AZ. Aurora synchronously replicates data to six storage nodes across Availability Zones independently of compute instances, so only the writer or reader instance needs to be failed over as the storage remains fully available throughout the drill.

In this walkthrough, we evacuate AZ A, the zone hosting the RDS primary and Aurora writer.

Multi-tier architecture spanning three Availability Zones: an Application Load Balancer fronting Amazon ECS and a Network Load Balancer fronting Amazon EKS, with Amazon RDS for PostgreSQL and Amazon Aurora PostgreSQL databases, before evacuating AZ A.

How ARC Zonal Shift works

When you initiate a zonal shift, ARC takes two coordinated actions for Amazon Route 53 and Elastic Load Balancers:

  1. DNS removal: The load balancer’s IP address in the affected AZ is removed from DNS. New client queries don’t resolve to that endpoint.
  2. Cross-zone traffic blocking: Load balancer nodes in the remaining AZs stop routing requests to targets in the shifted AZ, even when cross-zone load balancing is enabled.

For Amazon EKS clusters with zonal shift enabled, ARC goes further. It performs the following actions:

  • Cordons all nodes in the impacted AZ, preventing new pod scheduling.
  • Removes pod endpoints in the impacted AZ from EndpointSlice resources, redirecting east-west service-to-service traffic to healthy AZs.
  • Suspends AZ rebalancing for managed node groups and updates ASGs to launch instances only in healthy AZs.
  • Preserves nodes and pods in the shifted AZ (they are not terminated), keeping full capacity available for when the shift ends.

Combined with service-specific procedures for ECS task redistribution and database failover, this creates a complete AZ evacuation across all three traffic dimensions: north-south ingress, east-west service communication, and outbound database connections.

ARC Zonal Shift is a data plane operation by design. Because it works independently of the AWS control plane, it remains available even during an AZ impairment. The other steps in this walkthrough (ECS service updates, manual RDS failovers, manual Aurora failovers, subnet group modifications) are control plane operations. For a planned drill, this distinction has no practical impact because the control plane is healthy. During a real AZ impairment, prioritize the data plane action (start the zonal shift first to stop traffic immediately) and perform control plane operations only after traffic has already been shifted.

Prerequisites

Configure these prerequisites well in advance of your first shift. These are foundational settings that verify that your architecture is shift-ready at all times. For this walkthrough, you should have the following:

  • An AWS account
  • A multi-tier application deployed across 3 Availability Zones with ALB/NLB, ECS or EKS workloads, and RDS or Aurora databases.
  • IAM permissions to manage ARC Zonal Shift, ECS, EKS, and RDS resources.
  • AWS Command Line Interface (AWS CLI) v2 installed and configured.
  • Familiarity with ARC Zonal Shift concepts.
  • Amazon CloudWatch dashboards with per-AZ metric breakdowns (fault rate, latency, and target health).
  • Auto Scaling policies validated for sustained N-1 AZ operation.

Specifically for Elastic Load Balancing (ELB):

  • ALB/NLB deregistration delay set to 60 seconds, which allows existing connections to drain quickly after a shift instead of the default 300 seconds.
  • target_group_health.dns_failover.minimum_healthy_targets.count configured on each target group.

Specifically, for EKS:

  • kubectl installed and configured for your EKS cluster.
  • Turn on Topology Aware Routing on EKS services (or configure Istio locality-aware load balancing).
  • Zonal shift activated on your EKS cluster (one-time setup).

Specifically, for ECS:

  • ECS stopTimeout set to 55 seconds in task definitions, slightly below the ALB deregistration delay so tasks finish in-flight requests before being force-stopped, avoiding 502 errors during the drain window.

Specifically, for RDS:

  • Verify that your RDS primary and standby are provisioned in different Availability Zones.

What changes for a multi-day shift

The mechanics of starting a zonal shift are the same whether you run it for 1 hour or 72 hours. What changes is the operational surface area:

  • Expiry management: Zonal shifts have a maximum duration. You must monitor and extend them before they expire, or traffic returns to the shifted AZ unexpectedly.
  • Scaling drift: Over days, Auto Scaling events in healthy AZs may create capacity imbalances. Monitor and cap scaling so recovery doesn’t overload the returning AZ.
  • Connection pool cycling: After 24+ hours, most client connections will have recycled. This validates DNS TTL compliance across your entire client fleet, something a 1-hour test won’t fully exercise.
  • Operational confidence: Teams will learn to deploy, patch, and troubleshoot in a reduced AZ environment. A multi-day drill forces this to happen naturally rather than in a controlled window.
  • Safe recovery: After days at N-1 capacity, restoring the shifted AZ requires careful ordering. Verify health, scale back gradually, and reintroduce traffic incrementally rather than all at once.

Solution details

Each section below walks through the zonal shift procedure for one layer of the architecture, starting with the compute tier and finishing at the database layer.

Amazon ECS — Zonal Shift with task redistribution

For ECS services behind an ALB, initiating a zonal shift at the load balancer layer stops new traffic from reaching targets in the evacuated AZ. Existing ECS tasks in that AZ remain running but stop receiving requests. To perform a complete evacuation, follow these steps:

Step 0. Before starting, verify ARC Zonal Shift is enabled on the load balancer (disabled by default)

aws elbv2 modify-load-balancer-attributes \
    --load-balancer-arn $ALB_ARN \
    --attributes Key=zonal_shift.config.enabled,Value=true

Step 1. Initiate the zonal shift on the load balancer.

aws arc-zonal-shift start-zonal-shift \
    --resource-identifier $ALB_ARN \
    --away-from $AZ_ID_TO_EVACUATE \
    --expires-in "72h" \
    --comment "Multi-day AZ evacuation drill"

Zonal shifts expire after the duration set in --expires-in. If the shift expires before you cancel it, traffic automatically returns to the shifted AZ. For a multi-day drill, monitor the remaining time and extend before expiry using:

aws arc-zonal-shift update-zonal-shift \
    --zonal-shift-id $SHIFT_ID \
    --resource-identifier $RESOURCE_ARN \
    --expires-in "24h" \
    --comment "Extending drill"

When cross-zone load balancing is enabled (the default for ALB), the shift instructs load balancer nodes in healthy AZs not to route requests to targets in the impaired AZ. Targets are fully isolated regardless of your cross-zone configuration.

Step 2. Restrict new task placement to healthy AZs.

Update the ECS service’s network configuration to exclude subnets in the evacuated AZ. This prevents new tasks from launching in the shifted zone:

aws ecs update-service \
    --cluster $CLUSTER_NAME \
    --service $SERVICE_NAME \
    --network-configuration "awsvpcConfiguration={subnets=[$AZB_SUBNET,$AZC_SUBNET],securityGroups=[$SG_ID],assignPublicIp=DISABLED}"

Step 3. If needed, scale to verify N-1 AZ capacity.

aws ecs update-service \
    --cluster $CLUSTER_NAME \
    --service $SERVICE_NAME \
    --desired-count $N_MINUS_1_COUNT

Note: updating the ECS service network configuration and count that you want are control plane operations. Perform these changes before the drill starts, not during a real AZ impairment when control plane availability may be degraded.

We recommend that you pre-scale your services to handle the loss of an AZ’s worth of capacity before the drill. Your architecture should tolerate AZ loss without needing to scale reactively. See static stability in the Amazon Builders’ library.

In a 3 AZ environment, pre-scaling for N-1 capacity means running approximately 50% more compute than your baseline peak requires. If the cost isn’t justifiable for all workloads, consider scheduled scaling policies that increase capacity during planned drill windows, Auto Scaling with aggressive scale-out thresholds, or load shedding mechanisms. Keep in mind that during an unplanned impairment, you won’t have time to scale reactively. Workloads that aren’t pre-scaled will operate in a degraded state until scaling catches up, which can take minutes under load.

Step 4. Monitor task distribution.

aws ecs describe-tasks \
    --cluster $CLUSTER_NAME \
    --tasks $(aws ecs list-tasks --cluster $CLUSTER_NAME --service-name $SERVICE_NAME --query 'taskArns' --output text) \
    --query 'tasks[].[taskArn,availabilityZone]' --output table

Restore: Revert the network configuration to include all three AZ subnets, then cancel the zonal shift. Tasks will gradually rebalance across all AZs during subsequent deployments.

Important: zonal shift won’t work for single-AZ target groups as the ALB will refuse the shift if healthy targets only exist in one Availability Zone. Verify each target group has targets registered in at least two AZs before proceeding. For more details, refer to Application Load Balancers in the ARC documentation.

Amazon EKS — Zonal Shift with EndpointSlice isolation

Amazon EKS natively supports ARC zonal shift. When you turn on this capability and trigger a shift, ARC handles both the infrastructure and Kubernetes networking layers automatically.

What ARC does when you shift an EKS cluster:

  • Nodes in the impacted AZ are cordoned (no new pod scheduling).
  • The built-in Kubernetes EndpointSlice controller removes pod endpoints in the impacted AZ, so east-west service traffic is automatically redirected to pods in healthy AZs.
  • For managed node groups, AZ rebalancing is suspended and ASGs are updated to only launch in healthy AZs.
  • Nodes and pods in the shifted AZ are not terminated, ensuring full capacity is immediately available when the shift ends.

Step 1. Activate zonal shift for your EKS cluster (one-time setup):

aws eks update-cluster-config \
    --name $CLUSTER_NAME \
    --zonal-shift-config enabled=true

Step 2. Start the zonal shift on both the load balancer and EKS cluster:

# Shift north-south traffic at the load balancer
aws arc-zonal-shift start-zonal-shift \
    --resource-identifier $NLB_ARN \
    --away-from $AZ_ID_TO_EVACUATE \
    --expires-in "72h" \
    --comment "Multi-day AZ evacuation drill - north-south traffic"

# Shift east-west traffic within the EKS cluster
aws arc-zonal-shift start-zonal-shift \
    --resource-identifier $EKS_CLUSTER_ARN \
    --away-from $AZ_ID_TO_EVACUATE \
    --expires-in "72h" \
    --comment "Multi-day AZ evacuation drill - east-west traffic"

Step 3. Verify EndpointSlice update: confirm pods in the shifted AZ are no longer receiving traffic:

# Endpoints should only show AZ B/AZ C
kubectl get endpointslices -l kubernetes.io/service-name=$SERVICE_NAME -o yaml | \
    grep -A2 "zone:"

Step 4. Verify node and pod status:

# Nodes in evacuated AZ should show SchedulingDisabled
kubectl get nodes -l topology.kubernetes.io/zone=$AZ_NAME_TO_EVACUATE

# Confirm traffic distribution across healthy AZs
kubectl get pods -o wide -l app=$APP_LABEL

Verify your pods use topologySpreadConstraints with maxSkew: 1 on topology.kubernetes.io/zone and are pre-scaled to handle N-1 AZ load. The zonal shift doesn’t evict pods or trigger autoscaling by itself.

Note that ARC zonal shift doesn’t control outbound connections from pods to external dependencies like Amazon RDS. If your pods connect to AZ-specific database endpoints, consider using Istio with locality-aware routing. For implementation details, refer to End-to-end recovery from AZ impairments in EKS using Zonal Shift and Istio.

For Aurora, the cluster endpoint automatically routes to the current writer regardless of AZ, so no Istio configuration is needed for writer traffic. However, if you use AZ-specific reader instance endpoints, configure Istio ServiceEntry resources for each endpoint and apply a DestinationRule with localityLbSetting to prefer healthy AZs. This directs outbound database traffic to follow the same shift pattern as your north-south and east-west traffic.

Restore: Cancel both zonal shifts. ARC automatically uncordons nodes, adds pod endpoints back to EndpointSlices, and restores AZ rebalancing. Traffic returns to all three AZs with full capacity already in place.

aws arc-zonal-shift cancel-zonal-shift \
    --zonal-shift-id $EKS_SHIFT_ID \
    --resource-identifier $EKS_CLUSTER_ARN

aws arc-zonal-shift cancel-zonal-shift \
    --zonal-shift-id $NLB_SHIFT_ID \
    --resource-identifier $NLB_ARN

Amazon RDS for PostgreSQL — multi-AZ failover

Regarding Amazon RDS for PostgreSQL in a Multi-AZ deployment, if the primary instance resides in the AZ being evacuated, you must trigger a failover to the synchronous standby. RDS handles this through a reboot with failover.

Prerequisite: verify your RDS primary and standby are provisioned in different Availability Zones. If both are in the same zone, the following steps wouldn’t evacuate the zone as intended.

Step 1. Check current primary location:

aws rds describe-db-instances \
    --db-instance-identifier $RDS_INSTANCE \
    --query 'DBInstances[0].[DBInstanceIdentifier,AvailabilityZone,MultiAZ,SecondaryAvailabilityZone]' \
    --output table

Step 2. If the primary is in the evacuated AZ, manually force failover:

aws rds reboot-db-instance \
    --db-instance-identifier $RDS_INSTANCE \
    --force-failover

Step 3. Wait for availability and verify the new primary AZ:

aws rds wait db-instance-available \
    --db-instance-identifier $RDS_INSTANCE

aws rds describe-db-instances \
    --db-instance-identifier $RDS_INSTANCE \
    --query 'DBInstances[0].AvailabilityZone'

After failover, RDS recreates the standby in the evacuated AZ automatically. This is acceptable for a sustained drill as the standby receives no client traffic. Monitor ReplicaLag to confirm replication health.

Step 4. (Optional) Remove the standby from the evacuated AZ.

If you want zero RDS presence in the evacuated Availability Zone, you can relocate the standby by modifying the DB subnet group:

  1. Create a manual snapshot as a safety net.
  2. Disable Multi-AZ on the instance.
  3. Modify the DB subnet group to include only healthy AZ subnets, removing the evacuated AZ subnet.
  4. Re-enable Multi-AZ so the new standby is created in one of the remaining healthy AZs.

This approach works the same way for Amazon RDS for PostgreSQL as it does for any RDS engine using Multi-AZ deployments. Note that RDS Multi-AZ modifications (disabling/re-enabling Multi-AZ, subnet group changes) can take several minutes to complete. Plan for this during the drill window.

Amazon Aurora PostgreSQL — writer failover & reader management

Aurora provides more control over AZ placement than standard RDS Multi-AZ. You can explicitly choose which reader to promote and manage reader placement across AZs using failover priority tiers.

Step 1. Identify the cluster topology:

aws rds describe-db-clusters \
    --db-cluster-identifier $CLUSTER_ID \
    --query 'DBClusters[0].DBClusterMembers[].{Instance:DBInstanceIdentifier,IsWriter:IsClusterWriter}'

aws rds describe-db-instances \
    --filters Name=db-cluster-id,Values=$CLUSTER_ID \
    --query 'DBInstances[].[DBInstanceIdentifier,AvailabilityZone,DBInstanceStatus]' \
    --output table

Step 2. If the writer is in the evacuated AZ, failover to a reader in a healthy AZ:

aws rds failover-db-cluster \
    --db-cluster-identifier $CLUSTER_ID \
    --target-db-instance-identifier $READER_IN_HEALTHY_AZ

Step 3. Wait for the cluster to stabilize:

aws rds wait db-cluster-available \
    --db-cluster-identifier $CLUSTER_ID

If your Aurora cluster has no pre-existing reader in a healthy AZ, writer promotion requires creating a new instance, which typically takes less than 10 minutes. Pre-provisioning a reader in a separate AZ reduces failover time, often to less than 30 seconds.

Step 4. (Optional) Remove the reader in the evacuated AZ and create one in a healthy AZ.

For a full AZ evacuation where you want zero database presence in the shifted zone:

# Delete the reader instance in the evacuated AZ
aws rds delete-db-instance \
    --db-instance-identifier $INSTANCE_IN_EVACUATED_AZ \
    --skip-final-snapshot

# Create a new reader in a healthy AZ
aws rds create-db-instance \
    --db-instance-identifier ${CLUSTER_ID}-reader-${TARGET_AZ} \
    --db-cluster-identifier $CLUSTER_ID \
    --db-instance-class $INSTANCE_CLASS \
    --engine aurora-postgresql \
    --availability-zone $TARGET_AZ

Step 5. Monitor replication and performance throughout the drill:

aws cloudwatch get-metric-statistics \
    --namespace AWS/RDS \
    --metric-name AuroraReplicaLag \
    --dimensions Name=DBInstanceIdentifier,Value=$READER_INSTANCE \
    --start-time $TIMESTAMP_5MIN_AGO \
    --end-time $TIMESTAMP \
    --period 60 --statistics Average

Monitoring the drill with CloudWatch

A multi-day drill is only as valuable as the evidence it produces. Unlike a brief failover test where you visually confirm targets respond, a 48–72-hour evacuation requires continuous, automated observation, capturing capacity trends, replication health, and latency shifts that only surface under sustained N-1 AZ load.

Before starting the drill, verify you have CloudWatch dashboards with per-AZ metric breakdowns for each layer of your architecture. During the drill, these metrics serve two purposes: real-time operational awareness and post-drill evidence for stakeholders.

Key metrics by layer

The following metrics give you real-time visibility into each layer of the architecture during the drill.

Application Load Balancer / Network Load Balancer

Metric Dimension What to watch
HealthyHostCount Per target group, per AZ Should drop to 0 in evacuated AZ. Stable in healthy AZs
UnHealthyHostCount Per target group, per AZ Targets in evacuated AZ may show unhealthy (expected)
RequestCount Per AZ Zero traffic in shifted AZ. Even distribution in remaining AZs
TargetResponseTime Per AZ Watch for latency increases in healthy AZs under concentrated load
HTTPCode_Target_5XX_Count Per target group Sustained increase signals capacity pressure

Amazon ECS

Metric Dimension What to watch
CPUUtilization Per service Should not exceed 70–80% sustained (indicates capacity headroom)
MemoryUtilization Per service Memory pressure under concentrated load
RunningTaskCount Per service Confirms tasks running only in healthy AZs
DesiredTaskCount vs RunningTaskCount Per service Gap indicates placement failures (check subnet/capacity)

Amazon EKS (using Container Insights)

Metric Dimension What to watch
node_cpu_utilization Per node, filtered by AZ Nodes in healthy AZs absorbing shifted load
pod_cpu_utilization Per pod/namespace Hotspot detection under N-1 operation
node_status_condition Per node Nodes in evacuated AZ should show SchedulingDisabled
pod_number_of_container_restarts Per pod Restart loops may indicate resource pressure

Amazon RDS for PostgreSQL

Metric Dimension What to watch
CPUUtilization Per instance Primary under higher load post-failover
DatabaseConnections Per instance Connection spike after failover (watch for pool exhaustion)
ReadIOPS / WriteIOPS Per instance I/O patterns shift when primary moves AZs
ReplicaLag Per standby Should stabilize within seconds after failover
FreeableMemory Per instance Memory pressure under full client reconnection

Amazon Aurora PostgreSQL

Metric Dimension What to watch
AuroraReplicaLag Per reader instance Establish your cluster baseline during normal operation. Sustained increases from baseline indicate storage pressure. Aurora Replicas typically lag 100 ms or less
CommitLatency Per writer Increased commit latency indicates write contention
BufferCacheHitRatio Per instance Drop below 99% may indicate working set doesn’t fit in memory
DatabaseConnections Per instance Client reconnection behavior after writer promotion
VolumeBytesUsed Per cluster Aurora storage is AZ-independent (should be unaffected)

Export your per-AZ CloudWatch dashboards as snapshots before, during, and after the drill. Combine these with the ARC zonal shift event history (available through list-zonal-shifts) to create an auditable evidence package.

Cleaning up

After completing the drill, restore services carefully. The order matters, especially if Auto Scaling has increased capacity in healthy AZs:

  1. Verify the evacuated AZ is healthy: confirm targets are registered, pods are running, and database instances are available.
  2. Cancel the EKS cluster zonal shift first (east-west traffic resumes). Monitor for errors as internal traffic rebalances.
  3. Cancel the load balancer zonal shift (north-south traffic resumes). Traffic returns gradually as DNS propagates.
  4. If Auto Scaling added capacity in the remaining AZs, scale back gradually over 15 to 30 minutes. Don’t remove capacity before traffic has redistributed evenly.
  5. If cross-zone load balancing is disabled, verify target_group_health.dns_failover.minimum_healthy_targets.count is configured. This allows Route 53 to only route traffic to an AZ once it has enough healthy targets registered, preventing the restored zone from receiving traffic before it’s ready to handle it.
  6. Monitor per-AZ metrics for 30 minutes after restoring to confirm even distribution and no error spikes.

No additional AWS resources are created by ARC Zonal Shift that incur ongoing charges. The zonal shift itself is available at no additional cost.

Conclusion

In this post, you learned how to run a sustained AZ evacuation drill using ARC Zonal Shift across Amazon Elastic Container Service (Amazon ECS), Amazon Elastic Kubernetes Service (Amazon EKS), Amazon RDS for PostgreSQL, and Amazon Aurora. By operating on N-1 Availability Zones for 48–72 hours, rather than a brief failover test, you produce evidence that your multi-AZ architecture delivers genuine, sustained resilience. This is particularly valuable for financial services organizations facing regulatory mandates that require demonstrated recovery capabilities.

To get started, use the prerequisites checklist and step-by-step procedures in this post. Begin in non-production, progress to production during low-traffic windows, and build toward sustained operation under shift. As confidence grows, activate zonal autoshift so AWS can shift traffic automatically when internal telemetry detects an impairment.

You can also use AWS Resilience Hub to assess your application’s resilience posture before and after the drill. It validates that your architecture meets your defined Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets.

Amazon Application Recovery Controller – Zonal Shift

Best practices for zonal shifts in ARC

Using cross-zone load balancing with zonal shift

New AWS Fault Injection Service recovery action for zonal autoshift

End-to-end recovery from AZ impairments in Amazon EKS using EKS Zonal Shift and Istio

Amazon EKS now supports Amazon Application Recovery Controller


About the authors

How Generali Malaysia optimizes operations with Amazon EKS

Post Syndicated from Antoine Boucherie original https://aws.amazon.com/blogs/architecture/how-generali-malaysia-optimizes-operations-with-amazon-eks/

This post is co-authored with Ivan Amemoutou, DevOps and Cloud Lead at Generali Malaysia (“Generali”).

The insurance industry’s shift to cloud computing has accelerated the development and expansion of digital services. To support this transformation, insurers are modernizing their technology stack with solutions that enhance scalability, portability, and operational efficiency. This digital evolution is driven by growing customer expectations for seamless insurance services across all touchpoints. Generali faced this industry-wide challenge head-on, needing both to migrate their legacy applications to the cloud and meet increasing demands for new digital services. To address these needs, they embraced a modern approach by implementing containerized microservices architecture, significantly improving their operational capabilities and service delivery.

Generali started its migration to AWS in 2019. They selected Amazon Elastic Kubernetes Service (Amazon EKS) as the target container service for their modernized applications for its capabilities as an enterprise-grade container management solution and its seamless integration with other AWS services. Previous experience of the Generali DevOps and Cloud team was also a strong factor in selecting Amazon EKS. Although the selection of the target platform was straightforward, the main challenge Generali was facing was to enable the scale of adoption while maintaining a lean operational base.

Today, digital applications and several core insurance solutions are hosted on their EKS clusters, making it an important piece of infrastructure for the company. In this post, we look at how Generali is using Amazon EKS Auto Mode and its integration with other AWS services to enhance performance while reducing operational overhead, optimizing costs, and enhancing security.

Solution overview

Generali strives to implement Amazon EKS best practices and actively align their implementation with the AWS Well-Architected Framework. To that end, they follow the six pillars of Operational Excellence, Security, Reliability, Performance Efficiency, Cost Optimization, and Sustainability to build a robust and scalable platform. By applying Well-Architected principles to their EKS environment, Generali benefits from improved system resilience through automated operations and monitoring, enhanced security through AWS Identity and Access Management (IAM) integration and network policies, optimized costs through right-sizing and automatic scaling, and sustainable practices that minimize their environmental impact while maintaining high performance and reliability.

The following diagram illustrates the architecture of their EKS cluster and some of its integration points with different AWS services.AWS security and monitoring architecture diagram showing integration between Inspection VPC and EKS VPC with multiple AWS services for container workload protection and observability.

This solution offers the following benefits:

  • Simplified management of multiple containerized applications
  • Automated node provisioning and scaling
  • Enhanced security integration
  • Optimized resource utilization and simplified cost management
  • Granular multi-tenant observability

In the following sections, we discuss the integration with AWS services in more detail and how these components align with the AWS Well-Architected Framework.

Operational Excellence, Reliability, and Performance Efficiency with Amazon EKS Auto Mode

Generali faced challenges managing their expanding portfolio of containerized applications. The growth of their containerized services introduced operational inefficiencies and complexities: multiple applications from multiple tenants created operational overhead from manual orchestration and scaling to infrastructure maintenance, making it difficult to optimize costs while enforcing security and compliance across diverse application stacks. These challenges led to over-provisioning of resources and inconsistent security postures across different containerized environments.

To address these pain points, Generali has been adopting Amazon EKS Auto Mode, which automates their cluster infrastructure management, provides production-ready environments with minimal operational overhead, dynamically scales resources based on application demands, and implements consistent security practices with automated upgrades, so their teams can focus on application development rather than infrastructure complexity.

EKS Auto Mode manages the underlying nodes, load balancers, and storage configuration automatically. EKS Auto Mode takes care of scaling the cluster depending on the need of the workloads, while optimizing cost across a set of Amazon Elastic Compute Cloud (Amazon EC2) instances types selected by Generali in the node pools configuration.

With EKS Auto Mode’s expanded Shared Responsibility Model, compared to non-Auto Mode clusters, it also takes care of the patching of the underlying operating system (Bottlerocket), the different Amazon EKS add-ons installed by default, and the upgrade of the cluster, so Generali DevOps and Cloud team can focus on supporting their application teams.

While starting up EKS Auto Mode, the Generali DevOps and Cloud team had to adjust their operations to allow for those new features. For example, EKS Auto Mode releases a new version of its AMI, which automatically upgrades nodes on a regular basis, usually every week. To do so, nodes are terminated to be replaced with upgraded ones. The team had to create disruption control configurations to prevent those disruptions from impacting workloads. For example, they specified a maintenance window during off-peak hours for those upgrades. They also specified Pod Disruption Budgets and Node Disruptions Budgets to make sure critical applications would not see all the pods of a micro-service being terminated at the same time. The team can then focus on monitoring the current services and making sure they stay compliant with upcoming Amazon EKS upgrades, an activity that usually takes a fair amount of time every quarter, which is now automated with EKS Auto Mode.

Finally, the Generali DevOps and Cloud team also follow several principles to maintain reliability of their applications: they only allow stateless micro-services, they treat the underlying pods as immutable, they use Helm chart as a standardize deployment mechanism, and they use Horizontal Pod Autoscaler (HPA) to scale services based on traffic.

Security using Amazon GuardDuty, Amazon Inspector, Amazon Network Firewall, and AWS Secrets Manager

Generali implemented Amazon GuardDuty Extended Threat Detection for their EKS clusters to automatically correlate security signals across Amazon EKS audit logs, runtime behaviors, malware execution, and AWS API activity to identify sophisticated multistage attacks that traditional monitoring approaches often miss. By enabling both Amazon GuardDuty Amazon EKS protection and runtime monitoring, Generali gained comprehensive visibility into complex attack patterns such as container exploitation, privilege escalation, and unauthorized movement within their Kubernetes environment, with detailed timelines mapped to MITRE ATT&CK tactics and techniques. The benefits Generali realizes include reduced investigation time through consolidated security insights, rapid assessment of which containerized infrastructure components require immediate attention, and the ability to prioritize remediation efforts on the most critical affected resources while minimizing the potential blast radius of Amazon EKS targeted attacks.

Generali also uses the new Amazon Inspector capability to map Amazon ECR images to running containers, helping their security teams prioritize vulnerabilities based on containers currently running in their environment rather than just identifying vulnerabilities in repository images. The enhanced service provides Generali with visibility into which container images are actively running across their EKS environments, including cluster Amazon Resource Names (ARNs), the number of EKS pods where images are deployed, and last in-use dates for each vulnerability finding. The key benefits Generali realizes include the ability to prioritize remediation efforts based on actual container usage patterns rather than repository events alone, and comprehensive vulnerability management across container images.

Generali set up AWS Network Firewall to filter outbound HTTPS traffic from applications hosted on their EKS cluster by restricting outbound connections to only a set of hostnames provided by Server Name Indication (SNI) in the allow list, deploying their EKS cluster in private subnets with Network Firewall endpoints in public subnets and NAT gateways in protected subnets. The benefits Generali realizes include enhanced security through egress filtering that monitors and restricts outbound network traffic based on certificate hostnames rather than changing IP addresses, the ability to collect and analyze hostnames accessed by applications through Amazon CloudWatch alert logs for traffic pattern analysis, and improved compliance with security requirements by making sure applications can only access approved external services.

Getting secrets into pods can be done either through environment variables or as mounted volumes. Hard-coding them directly into the deployment template is not recommended, and it is better to store them in AWS Secret Manager and retrieve them dynamically. As a best practice and to reduce operational complexity, Generali choses to only host stateless containers in their cluster, alleviating the need for storage volume. To that end, the best option is to retrieve secrets dynamically and add them as environment variables to the pod. To do so, they implemented the External Secrets Operator on their EKS cluster to use Secrets Manager for centralized secret management, which reads the necessary secrets and automatically stores them as Kubernetes secrets without requiring application code changes or daemonsets. The benefits Generali realizes include improved security, management, and auditability of secret usage through centralized secret management outside their Kubernetes clusters and automatic secret synchronization on a recurring basis to capture credential rotations.

Cost Optimization using tags and Savings Plans

Although EKS Auto Mode already offers some cost optimization features, it’s important for Generali to keep track of resource consumption per business project. To that end, Generali uses AWS Billing split cost allocation data for Amazon EKS to analyze and allocate costs using the AWS Billing Console, gaining insights into Kubernetes costs alongside other AWS spend. The feature allows for split along cost allocation tags for some Kubernetes attributes. These tags include aws:eks:cluster-name, aws:eks:deployment, aws:eks:namespace, and aws:eks:node, so the company can map Amazon EKS consumption against lines of business and applications.

Generali also takes advantage of the following:

Operational Excellence and observability using custom dashboards in Amazon Managed Grafana

Hosting multiple projects from multiple business unit means that different application owners need their own custom analytics dashboards. To provide per-project granularity, Generali uses the integration between CloudWatch and Amazon Managed Grafana to create observability dashboards per EKS namespace. By connecting CloudWatch as a data source in Amazon Managed Grafana, they can visualize Amazon EKS metrics, logs, and traces through Grafana’s powerful visualization capabilities without managing the underlying Grafana infrastructure. Through this integration, Generali can create unified views of cluster health, node performance, pod resource utilization, and application performance indicators, while using Grafana’s advanced alerting and templating features for dynamic dashboard creation.

Lessons learned

Generali’s adoption of EKS Auto Mode, combined with integrated AWS security services and comprehensive observability tools, has transformed their container operations from a complex, manually managed environment to an automated, secure, and efficient platform. The integration with services like GuardDuty, Amazon CloudWatch Container Insights, and Amazon Managed Grafana has created a cohesive ecosystem that maximizes operational efficiency while minimizing management overhead. This transformation has helped the Generali DevOps and Cloud team shift its focus from infrastructure maintenance to strategic application support, resulting in improved security posture, cost optimization, and overall platform reliability.Generali realized the following key benefits:

  • Significant reduction in operational overhead with EKS Auto Mode
  • Enhanced security with automated threat detection and response
  • Reduction in infrastructure costs through optimization
  • Improved mean-time-to-resolution
  • Accelerated application deployment cycles

Conclusion

Amazon EKS Auto Mode has proven to be a transformative service for Generali, helping them build a modern, secure, and efficient container environment that aligns with AWS Well-Architected best practices. With EKS Auto Mode and its integration with AWS services like GuardDuty, Amazon Inspector, and CloudWatch, Generali created a robust foundation that not only enhances their security posture and operational efficiency but also optimizes costs. The Generali DevOps and Cloud team is now able to focus on applications teams’ support with expansion plans to host AI models and upcoming agentic applications.As organizations continue their cloud-based journey, Generali’s experience demonstrates how AWS’s comprehensive container services can help enterprises focus on innovation and business value while maintaining operational excellence, security, and cost-efficiency at scale.

If you’re interested in learning more about Amazon EKS, refer to Amazon EKS Best Practices Guide.

About Generali Malaysia

Generali Malaysia is one of the largest general insurers and an emerging life insurer in the country, dedicated to delivering best in class general and life insurance protection solutions for individuals, families, and businesses. As part of the Generali Group, a global insurance leader with over 190 years of heritage, Generali Malaysia carries forward a deep legacy of protection, service excellence, and innovation.

Today, the company is supported by more than 1,600 employees, over 9,000 agents and partners, and an extensive network of branches nationwide. Guided by its ambition to be a trusted Lifetime Partner, Generali Malaysia is committed to its purpose of empowering lives and dreams. The company continues to drive excellence by leveraging AI, data, and customer centric solutions, while embedding sustainability at the heart of its business.


About the authors

Pilot light with reserved capacity: How to optimize DR cost using On-Demand Capacity Reservations

Post Syndicated from Antoine Boucherie original https://aws.amazon.com/blogs/architecture/pilot-light-with-reserved-capacity-how-to-optimize-dr-cost-using-on-demand-capacity-reservations/

For digital enterprises to remain competitive, resilience is essential for maintaining reliability and building customer trust. End users expect applications to be available 24 hours a day, leading companies to develop increasingly sophisticated methods to provide continuous operation of critical services. Some companies, such as financial services companies, have to meet regulatory requirements such as Digital Operational Resilience Act (DORA) and are expected to manage the risk of outsourcing critical applications. They must design for high availability and plan for potential impairments. By proactively planning for potential disruptions, they’re not just mitigating risks – they’re building trust and delivering unparalleled value to their customers.

When assessing your own applications, you should define a set of objectives, perform a business impact analysis and a risk assessment. This way, you can estimate the impact to the business if the application isn’t available. This results in categorization of the applications and influence their design according to the AWS resilience lifecycle framework. Each application is given a specific Recovery Point Objective (RPO) and Recovery Time Objective (RTO), depending on its criticality for the business.

Not all applications fall in the most critical category. You allocate resources according to the results of the assessment and make trade-offs when designing applications. For example, you’ll have a more stringent RTO and RPO for—and be willing to spend more time and money on—a critical application than on a less critical application. The challenge becomes how to minimize the risk of breaching a specific RTO while optimizing for resources, such as cost and operational complexity.

At AWS, we provide guidance through the Well-Architected Framework and specifically within the Reliability pillar. Disruption can happen at several levels, and we recommend that you explore and prepare for four types of disruptions in the AWS Resilience Hub: application, infrastructure, Availability Zone, and AWS Region.

We recommend that you use managed services and make sure that all production workloads are designed to take advantage of multiple Availability Zones in AWS Regions. If your application also needs to be protected against the unlikely risk of Regional impairment, you should consider a multi-Region disaster recovery (DR) strategy.

You can select from several DR strategies: backup and restore, pilot light, warm standby, and multi-site active-active:

  • Backup and restore – This strategy might not provide the necessary RPO or RTO required for a highly critical application.
  • Multi-site active-active – This strategy increases significantly the cost and operational complexity of your application.
  • Pilot light – This strategy allows for a RPO or RTO in the tens of minutes by having the data asynchronously copied to the secondary Region and ready to be accessed. However, unlike a warm standby, the application servers aren’t deployed and aren’t ready to serve traffic. The pilot light strategy allows for a lower cost but brings a risk that you might not be able to provision the compute capacity you need when you want to fail over to the secondary Region, especially if you require a specific instance type.

In this post, we explore an intermediate strategy between the pilot light and the warm standby strategies: pilot light with reserved capacity. You can use this strategy to reserve compute capacity in a secondary Region while also limiting cost.

The following diagram illustrates where the pilot light with reserved capacity solution lies in the spectrum of disaster recovery strategies.

spectrum of disaster recovery strategies

Reserving capacity, on demand

On-Demand Capacity Reservations were launched in 2018. They make it possible to reserve capacity in the Availability Zone of your choice without a long-term contract. You have the flexibility to create, modify, or cancel reservations at your discretion. It’s especially well-suited if your application is dependent on a specific instance type or size.

Optimizing the cost of On-Demand Capacity Reservations with a Savings Plan

On-Demand Capacity Reservations is a reservation mechanism and doesn’t require a commitment. However, you can optimize your spending by combining the capacity reservation with an AWS Savings Plan. By using Savings Plans, you can achieve up to a 72% discount, a very significant cost reduction for DR instances that have to stay available all year long.

Optimizing the cost of On-Demand Capacity Reservations by sharing Capacity Reservations

To further optimize the cost, you can use your reserved capacity in another account when you don’t need it for DR.

Here’s an example in which we share On-Demand Capacity Reservations with our development and test account:

We have a three-tier application running in production in a primary AWS Region. This application is composed of a load balancer forwarding traffic to a fleet of application servers running on Amazon Elastic Compute Cloud (Amazon EC2) instances, backed by an Amazon Relational Database Service (Amazon RDS) database. All services used by this application are configured to use multiple Availability Zones in this primary Region.

We use the pilot light strategy, so the application data is being replicated to the disaster recovery environment in a secondary Region using Amazon RDS cross-Region read replicas. However, the load balancer and EC2 services aren’t running in DR to limit cost and operational complexity. Following best practices, each environment is running in a different AWS account.

The following diagram illustrates the pilot light strategy setup for our example.

pilot light strategy setup for our example

To reserve capacity in case of failover to the secondary Region, we create an On-Demand Capacity Reservation in the DR account, according to our baseline compute capacity. Because we don’t need this capacity until we fail over the application from the primary to the secondary Region, we share those On-Demand Capacity Reservations with a development and test account hosting our nonproduction environment in the secondary Region. On-Demand Capacity Reservations are Availability Zone specific (and hence Region specific) and can be shared with either AWS accounts or AWS Organizations using AWS Resource Access Manager (AWS RAM).

Best practices are to share those On-Demand Capacity Reservations with a nonproduction organizational unit (OU) within an organization or to directly share with the account(s) hosting the testing environments (for example, user acceptance testing or preproduction). Those environments are usually very similar to the production account in baseline sizing, in order to perform load and performance testing. This is an important point: you want to be able to retrieve those On-Demand Capacity Reservations when needed without impacting other critical applications.

The following diagram illustrates the Capacity Reservations sharing with the development and test account.

Capacity Reservation sharing with development and test account

If an impairment affects our production environment in the primary Region, we can trigger failover to the secondary Region. To reclaim capacity, we need to terminate the EC2 instances running in our development and test account. Capacity becomes available nearly immediately after these instances are successfully terminated. Separately, we can also stop the sharing of On-Demand Capacity Reservations to make sure that the development and test account can’t consume that capacity again. Know that merely unsharing your reservation without terminating development and test instances might not result in complete or immediate capacity retrieval. This is because when you unshare an On-Demand Capacity Reservation, the instances in the consumer account continue to run, and capacity is only returned to the owner account if additional capacity is available in the Amazon EC2 service on-demand pool.

The following diagram illustrates the failover to the DR environment in a secondary Region.

cancellation of the capacity reservations sharing

Steps

Here is a possible approach to take advantage of On-Demand Capacity Reservations to reduce the application’s total infrastructure cost:

  1. Calculate the baseline compute capacity necessary for the DR environment in the secondary Region in the event of failover, including the compute that might already be running in this secondary Region for data stores (for example, a Kafka broker running on Amazon EC2). How much vCPU and RAM is required or what are the exact EC2 instances necessary to host the whole application in case of failover of the production from the primary to the secondary Region.
  2. Create an On-Demand Capacity Reservation for the exact EC2 instances that the application need as a baseline in the DR account. Capacity Reservation Fleet is also a possible choice to reserve capacity across multiple instance types, which is often the case for Amazon Elastic Kubernetes Service (Amazon EKS) or Amazon Elastic Container Service (Amazon ECS) clusters, for example. Creating a Capacity Reservation Fleet will create multiple Capacity Reservations that can be shared independently. It’s also recommended to apply for Savings Plans on those On-Demand Capacity Reservations to save up to 72%.
  3. Share those On-Demand Capacity Reservations from the DR account to one or several accounts, depending on your need. In our example, we share the On-Demand Capacity Reservations with the development and test account, effectively allowing the development and test environment to use compute capacity that has already been reserved.
  4. In case of impairment in the primary Region, terminate the development and test instances first and then stop the On-Demand Capacity Reservation sharing The DR account will recover those reservations. If you want to keep the development and test instances, you will be charged at the on-demand rate.
  5. Redeploy in an automated manner the application in the DR account on new EC2 instances behind a load balancer.

Benefits

By purchasing On-Demand Capacity Reservations in the DR account, you make sure that you always have Amazon EC2 capacity access when required and for as long as you need it. By sharing those On-Demand Capacity Reservations with another AWS account or organization, you can share the cost of the application’s compute capacity with other environments, reducing your application’s total cost of ownership. The additional cost of the DR instances can even reach zero, if your instances are completely consumed by nonproduction environments such as development and testing.

DR savings over On-Demand
Compute Savings Plan – 1 year, no upfront Around 27% (for example, for m7i instance)
Compute Savings Plan – 3 years, all upfront Up to 66%
Instance Savings Plan – 3 years, all upfront Up to 72%
Reservation shared and consumed 100% by development and test environment Up to 100%

Limits

Although you can reserve DR capacity at a minimal cost using the pilot light with reserved capacity solution, there are some limits to keep in mind.

Firstly, we advise looking at this solution only if the Recovery Time Objective of the application, in case of Regional disruption, is in hours because you need to take into account the time needed to:

  • Detect the impairment in the primary Region.
  • Trigger the failover procedure.
  • Terminate the used instances to retrieve capacity (estimated time in minutes)
  • Stop the On-Demand Capacity Reservations sharing and automatically retrieve them in the DR account (estimated time in minutes).
  • Launch the compute infrastructure with the necessary application software in the DR account. You need to make sure that it matches the On-Demand Capacity Reservations according to the criteria used (open or targeted)

If your application requires a lower RTO, we recommend exploring the warm standby strategy.

Secondly, this strategy can only be used for application servers running EC2 instances and ECS or EKS clusters on EC2 because On-Demand Capacity Reservations aren’t available for managed services such as AWS Fargate or AWS Lambda. For those managed services, we recommend having them up and running like in a warm standby strategy, with a minimum baseline capacity that you’re comfortable with.

Thirdly, it requires some nonproduction development and test usage in the selected secondary Region to use the shared On-Demand Capacity Reservation.

Finally, it’s important to consider that this solution brings some complexity and extra operational work. You should plan well ahead, automate the operational tasks where possible, but most importantly, regularly test that the failover of the application works according to plan. We encourage you to perform your own game days to support your operational resilience.

Deciding whether this strategy is a good fit for your application will ultimately be a decision based on your business and regulatory requirements.

Conclusion

In this post, we explained how to reserve capacity in a secondary Region using On-Demand Capacity Reservations. We highlighted how cost can be optimized using Savings Plans and by sharing reserved capacity with noncritical workloads. We saw how we can recover that capacity for the DR environment, in the event of a disaster, to allow the application to continue to serve end users. We looked at the benefits and limits of the pilot light with reserved capacity solution and the necessary steps to put it in place.


About the Authors