Tag Archives: Technical How-to

Configure domain-level VPC networking in Amazon SageMaker Unified Studio

Post Syndicated from Prasad Nadig original https://aws.amazon.com/blogs/big-data/configure-domain-level-vpc-networking-in-amazon-sagemaker-unified-studio/

Enterprise operations teams that run domain-level VPC networking in Amazon SageMaker Unified Studio often support dozens of projects spanning data engineering, analytics, and machine learning (ML) teams. Each project requires private connectivity to internal databases, Amazon Simple Storage Service (Amazon S3) buckets, and AWS services. Without a domain-level Amazon Virtual Private Cloud (Amazon VPC) configuration, project owners coordinate with the networking team individually. This piecemeal approach leads to inconsistent subnet choices, missing VPC endpoints, connectivity failures that are hard to troubleshoot, and a network posture that is difficult to audit.

With domain-level VPC networking, you configure the network once, and all new projects get the right network immediately upon creation. In this post, you learn how to:

  • Configure SageMaker Unified Studio domain-level VPC networking.
  • Select subnets and security groups that provide multi-Availability Zone (multi-AZ) resilience.
  • Update projects that have no VPC to inherit the domain VPC, and understand when a project must be recreated instead.
  • Validate network connectivity from within a project.

In this post, you learn how to configure VPC networking for a SageMaker Unified Studio domain that uses AWS Identity and Access Management (IAM)-based authentication. You see how network components map to domain and project resources, and how to plan a configuration that balances security, connectivity, and operational simplicity.

Solution overview

Domain-level VPC networking provides a single network configuration that applies to all new projects in the domain. Projects automatically inherit the VPC settings, including subnets, security groups, and connectivity to AWS services through VPC endpoints. Existing projects are an exception and are handled separately (see Step 3).

The following diagram shows a single VPC with private subnets across two Availability Zones configured at the domain level, with data engineering, analytics, and ML projects all inheriting that configuration.

Architecture diagram showing domain-level VPC configuration in Amazon SageMaker Unified Studio with private subnets across two Availability Zones.

Figure 1: Domain-level VPC configuration in Amazon SageMaker Unified Studio. A single VPC with private subnets across two Availability Zones is configured at the domain level. All projects (data engineering, analytics, ML) inherit this configuration automatically

Key benefits of this approach:

  • Configure once, apply across projects: New projects inherit the domain VPC without manual intervention.
  • Consistent security posture: A single network boundary covers all data, analytics, and ML workloads.
  • Simplified auditing: One VPC to audit rather than one per project. Turn on VPC Flow Logs and review AWS CloudTrail events for network-level auditing.
  • Reduced operational overhead: Project teams start working immediately without submitting networking requests.

The following AWS services are used in this solution:

Prerequisites

Before configuring domain-level VPC networking, verify you have the following:

  • Domain administrator permissions for Amazon SageMaker Unified Studio.
  • An existing VPC with the following requirements:
    • At least two private subnets in different Availability Zones.
    • DNS hostnames and DNS support enabled.
    • At least five available IP addresses per expected Amazon SageMaker Unified Studio project. This is a baseline minimum. Workloads using AWS Glue, Amazon EMR, or Amazon Redshift Serverless consume additional elastic network interfaces (ENIs) per worker or node. We recommend /24 or larger subnets for production domains and forward-looking capacity planning based on your expected users and compute types. For detailed guidance, see How to set up a network-isolated VPC for Amazon SageMaker Unified Studio.
  • VPC endpoints configured for the AWS services your projects access (for example, Amazon S3, AWS Glue, Amazon SageMaker AI).
  • Private DNS enabled on all interface VPC endpoints (you must enable this so that service DNS names resolve to private IPs). If you use centralized VPC endpoints through AWS Resource Access Manager (AWS RAM) or AWS Transit Gateway, configure Amazon Route 53 Resolver inbound endpoints instead.
  • S3 gateway endpoint route table associations configured for all selected private subnets (without this, S3 access fails in subnets whose route table lacks the prefix-list route).
  • A security group (optional), if not provided, SageMaker Unified Studio creates one automatically.
  • The SageMakerStudioAdminIAMConsolePolicy managed policy (or equivalent permissions including ec2:Describe*, ec2:CreateSecurityGroup, and datazone:* actions) attached to the domain administrator IAM role. See SageMakerStudioAdminIAMConsolePolicy in the AWS Managed Policy Reference for the full permission set.

Note: The VPC must be in the same AWS Region as the domain.

For detailed guidance on VPC networking configuration, see Configure VPC networking for IAM-based domains in the SageMaker Unified Studio Administrator Guide.

VPC endpoint requirements

Because your subnets are private (no internet gateway route), compute resources access AWS services through VPC endpoints. At a minimum, configure the following interface and gateway endpoints (add Amazon Athena, AWS Lake Formation, or Amazon Redshift endpoints if you use them in your projects):

Endpoint Type Purpose
com.amazonaws.region.s3 Gateway S3 access for data storage
com.amazonaws.region.glue Interface AWS Glue job connectivity
com.amazonaws.region.sagemaker.api Interface SageMaker API calls
com.amazonaws.region.sagemaker.runtime Interface Model inference
com.amazonaws.region.logs Interface Amazon CloudWatch Logs
com.amazonaws.region.monitoring Interface Amazon CloudWatch metrics
com.amazonaws.region.sts Interface IAM role assumption
com.amazonaws.region.datazone Interface Amazon SageMaker Unified Studio service connectivity
com.amazonaws.region.ecr.api Interface ECR API calls (container image metadata)
com.amazonaws.region.ecr.dkr Interface ECR image layer pulls (Docker registry)
com.amazonaws.region.kms Interface AWS Key Management Service (AWS KMS) encryption/decryption operations

Note: Interface endpoints incur an hourly charge per Availability Zone plus data processing fees. Gateway endpoints (such as S3) have no hourly charge. Factor endpoint count and AZ spread into your cost estimate.

For a comprehensive list of all mandatory and optional VPC endpoints for a fully network-isolated setup, see How to set up a network-isolated VPC for Amazon SageMaker Unified Studio. For current pricing details, see AWS PrivateLink pricing.

Note: Review your account’s service quotas for interface VPC endpoints per VPC (default 50) and ENIs per Region before scaling. Request increases through Service Quotas if needed.

Solution walkthrough

The following steps walk you through configuring the domain VPC and validating it, from signing in to the console through confirming private connectivity from a project.

Step 1: Sign in and navigate to networking settings

  1. Sign in to the AWS Management Console as your Amazon SageMaker Unified Studio domain administrator (the IAM role designated as the domain login role).
  2. Open the Amazon SageMaker console.
  3. Use the Region selector in the top navigation bar to select the Region where your domain exists.
  4. On the Amazon SageMaker Unified Studio landing page, choose Open to launch your IAM-based domain.

The following screenshot shows the Amazon SageMaker Unified Studio landing page, where you choose Open to launch the domain.

Amazon SageMaker Unified Studio landing page with the Open button to launch the IAM-based domain.

Figure 2: Amazon SageMaker Unified Studio landing page with the Open button to launch the IAM-based domain

  1. From the navigation pane, choose Domain management.

The following screenshot shows Domain management in the navigation pane.

Navigation pane showing Domain management link in Amazon SageMaker Unified Studio.

Figure 3: Domain management on navigation pane

Note: Access to the domain administration page is restricted to the IAM role specified as the domain login role during domain creation.

Step 2: Add VPC configuration

  1. In the navigation pane, choose Settings. In the Networking in this account section, choose Add VPC.

The following screenshot shows the Networking in this account section with the Add VPC button.

Domain management Settings page showing the Networking in this account section with Add VPC button.

Figure 4: Domain management Settings page showing the Networking in this account section to add a VPC

  1. For VPC, select the VPC with connectivity to your compute, database, and storage resources. If no VPC exists, choose Create VPC to provision one using AWS CloudFormation.
  2. For Subnets, select a minimum of two private subnets in different Availability Zones.
  3. (Optional) For Security group, select a security group to control inbound and outbound traffic. If you don’t choose one, SageMaker Unified Studio creates one automatically.
  4. Choose Save.
  5. Verify the VPC configuration status shows Ready in the Networking in this account section.

The following screenshots show the Add VPC dialog and the resulting Ready status in the Networking in this account section.

Add VPC dialog with fields for VPC, subnets, and security group selection.

Figure 5: Add VPC dialog with fields for VPC, subnets, and security group selection

VPC configuration status showing Ready in the Networking in this account section.

Figure 6: VPC configuration status showing Ready in the Networking in this account section

Note: IAM-based domains support only one VPC configuration at a time. AWS IAM Identity Center-based domains can have a VPC per Region. For details, see Configure VPC networking for IAM-based domains in the SageMaker Unified Studio Administrator Guide.

New projects created in the domain now automatically use the saved VPC configuration. Existing projects are an exception. See Step 3 to update them.

Step 3: Update existing projects

Existing projects don’t automatically inherit the domain VPC configuration. How you apply the new settings depends on the project’s current state:

Projects with no VPC configured – Update in place to adopt the domain VPC. See the following steps.

Projects that already have a VPC – These can’t be switched to a different VPC configuration. To adopt the domain VPC:

  1. Create a new project (which inherits the domain VPC automatically).
  2. Recreate connections in the new project.
  3. Migrate assets from the old project.
  4. Back up any data you need, then delete the original project.

Because recreation can disrupt in-progress work and doesn’t migrate project data automatically, schedule this as a planned maintenance window.

To update a project that currently has no VPC configured:

  1. From the domain administration page, choose Projects in the navigation pane.
  2. Choose the project you want to update.
  3. On the project detail page, a banner appears: “Configurations have changed. Please update this project to access the latest configuration.”
  4. In the banner, choose Update.
  5. Confirm the update when prompted.

Repeat this process for each existing project that should use the domain VPC. The following screenshot shows the project detail page with the configuration update banner.

Project detail page showing the update banner for VPC configuration changes.

Figure 7: Project detail page showing the configuration update banner

Step 4: Validate connectivity

After configuring the domain VPC and updating your projects, verify connectivity. Compute resources should have private connectivity to AWS services through the VPC, without any additional project-level network configuration.

Create a notebook in one of your projects as shown in the following figure and run the following code:

Creating a notebook in a SageMaker Unified Studio project to validate VPC connectivity.

Figure 8: Creating a notebook in a SageMaker Unified Studio project to validate VPC connectivity

Requirements: Python 3.8+, Boto3 1.26 or later. Run in a notebook within your SageMaker Unified Studio project.

import boto3
import socket
import ipaddress

def validate_vpc_connectivity():
    """Validate that the project has private connectivity to AWS services
    through the domain-level VPC configuration."""

    results = {}
    region = boto3.session.Session().region_name
    if not region:
        raise RuntimeError('Could not determine AWS Region. Run this notebook inside a SageMaker Unified Studio project.')

    # Test Amazon S3 access via VPC endpoint
    try:
        s3 = boto3.client('s3')
        response = s3.list_buckets()
        results['S3'] = f"[PASS] Accessible ({len(response['Buckets'])} buckets)"
    except Exception as e:
        results['S3'] = f"[FAIL] Failed: {e}"

    # Test AWS Glue access via VPC endpoint
    try:
        glue = boto3.client('glue')
        dbs = glue.get_databases()
        results['Glue'] = f"[PASS] Accessible ({len(dbs['DatabaseList'])} databases)"
    except Exception as e:
        results['Glue'] = f"[FAIL] Failed: {e}"

    # Test STS (role assumption through VPC endpoint)
    try:
        sts = boto3.client('sts')
        identity = sts.get_caller_identity()
        results['STS'] = f"[PASS] Accessible (Account: {identity['Account']})"
    except Exception as e:
        results['STS'] = f"[FAIL] Failed: {e}"

    # Verify interface endpoint resolves to private IP
    try:
        sts_endpoint = f"sts.{region}.amazonaws.com"
        addr_info = socket.getaddrinfo(sts_endpoint, 443, family=socket.AF_INET)
        ip = addr_info[0][4][0]
        is_private = ipaddress.ip_address(ip).is_private
        if is_private:
            results['DNS Resolution'] = f"[PASS] Private IP ({ip}) (traffic stays on AWS network)"
        else:
            results['DNS Resolution'] = f"[WARN] Public IP ({ip}) - check VPC endpoint config"
    except Exception as e:
        results['DNS Resolution'] = f"[FAIL] Failed: {e}"

    # Print results
    print("-" * 40)
    print("Domain VPC Connectivity Validation")
    print("-" * 40)
    for service, status in results.items():
        print(f" {service}: {status}")
    print("-" * 40)
    print(f"\n Region: {region}")

    # Check if all tests passed
    all_passed = all("[PASS]" in status for status in results.values())
    has_warn = any("[WARN]" in status for status in results.values())
    if all_passed:
        print(f"\n [PASS] All services accessible via private VPC endpoints.")
        print(f" This project inherited its network configuration")
        print(f" from the domain without per-project setup.")
    elif has_warn and all("[PASS]" in s or "[WARN]" in s for s in results.values()):
        print(f"\n [WARN] Services are reachable, but DNS resolves to public IPs.")
        print(f" Verify that Private DNS is enabled on your interface VPC endpoints.")
    else:
        print(f"\n [FAIL] Some services are not reachable.")
        print(f" Check that VPC endpoints are configured and security")
        print(f" groups allow outbound traffic on port 443.")

validate_vpc_connectivity()

Expected output when VPC is correctly configured:

Successful validation output showing all services accessible through private VPC endpoints.

Figure 9: Successful validation output showing all services accessible through private VPC endpoints

If any service shows a failure, one common cause is security groups preventing traffic on port 443 to the VPC endpoint. Other causes include missing VPC endpoints, incorrect route table entries, or DNS resolution issues. For more information, see Configure VPC networking for IAM-based domains in the SageMaker Unified Studio Administrator Guide.

Note: An AccessDenied error indicates the request reached the service. Connectivity is working, but IAM permissions need adjustment (for example, the S3 test requires s3:ListAllMyBuckets, which some project roles lack). A timeout or connection error points to a networking problem (missing endpoint, route, or security group rule). The following screenshot shows the validation output when VPC endpoints are missing, where the affected services report timeout errors.

Validation output when VPC endpoints are not configured showing timeout errors.

Figure 10: Validation output when VPC endpoints are not configured. Timeout errors indicate missing endpoints

The security group applied at the domain level controls network access for all projects. To review or tighten the rules:

  1. Navigate to the Amazon VPC console.
  2. Choose Security groups and choose the security group shown in your domain’s Networking settings.
  3. Review the Inbound rules and Outbound rules tabs.

By default, the auto-created security group allows all outbound traffic on port 443 (HTTPS) to reach AWS services through VPC endpoints. Consider restricting outbound rules to only the specific VPC endpoint security groups for least-privilege access. Additionally, make sure your VPC endpoint security groups allow inbound TCP 443 from the domain security group or subnet CIDRs. For distributed compute services (AWS Glue, Amazon EMR), add a self-referencing inbound rule to allow worker-to-worker communication.

Updating VPC configuration

After the initial setup, you can modify the VPC configuration to change the VPC, subnets, or security group:

  1. From the domain administration page, choose Settings in the navigation pane.
  2. In the Networking in this account section, under the Actions column, choose Update.
  3. Update the VPC, subnets, or security group as needed.
  4. Choose Update.

The following screenshot shows the Update VPC dialog, where you modify the VPC, subnets, or security group.

Settings page with the Actions menu showing Update and Remove options for VPC configuration.

Figure 11: Update VPC dialog showing the option to modify VPC, subnets, or security group for the domain

Important: Updating the VPC does not affect already provisioned resources. Newly created resources in projects use the updated VPC. Existing projects that already have a VPC keep their original settings and must be recreated to adopt the change. Projects with no VPC can be updated in place (see Step 3).

Clean up

To remove the VPC configuration from your domain:

  1. From the domain administration page, choose Settings in the navigation pane.
  2. In the Networking in this account section, choose the Actions menu (⋮) and choose Remove.

The following screenshot shows the Actions menu with the Remove option.

Actions menu in the Networking in this account section showing the Remove option.

Figure 12: Actions menu in the Networking in this account section showing the Remove option

If you created a dedicated VPC for this walkthrough and no longer need it:

  • Delete the VPC and associated resources (subnets, VPC endpoints, security groups) from the Amazon VPC console. Before deleting, remove the domain VPC configuration and make sure all project resources are terminated. Active projects create ENIs that block VPC and subnet deletion.
  • If you used an AWS CloudFormation template to create the VPC, delete the stack to remove all resources cleanly. Open the AWS CloudFormation console and delete the stack.

Note: Removing the domain VPC configuration does not retroactively change projects that already have VPC applied. Those projects retain their existing network configuration. New projects created after removal do not have a VPC configured.

Conclusion

In this post, we showed how to configure domain-level VPC networking in Amazon SageMaker Unified Studio. A single domain-level VPC eliminates per-project networking overhead, enforces a consistent security posture, and simplifies compliance auditing.

Key takeaways:

  • Domain-level VPC is a one-time configuration that automatically applies to all new projects.
  • Projects with no VPC can be updated in place. Projects that already have a VPC must be recreated to adopt a changed configuration.
  • Private subnets with VPC endpoints provide secure, private connectivity to AWS services without traversing the public internet.

As next steps, consider:

  • Reviewing your auto-created security group rules and tightening them for least-privilege access.
  • Adding VPC endpoints for additional AWS services as your projects’ needs evolve.
  • Monitoring subnet IP address utilization to plan capacity as you add more projects. Use the AvailableIpAddressCount Amazon CloudWatch metric for your subnets to track utilization and set alarms.

For more information, see Configure VPC networking for IAM-based domains in the Amazon SageMaker Unified Studio Administrator Guide.

 


About the authors

Prasad Nadig

Prasad Nadig

Prasad is a Senior Analytics Specialist Solutions Architect at Amazon Web Services (AWS), specializing in large-scale data analytics and AI. Prasad partners with customers to design, migrate, and modernize their analytics platforms on AWS into scalable, cost-effective solutions, with deep expertise in data lakes, data warehousing, distributed processing, and performance tuning at petabyte scale.

Amit Shyam Jaisinghani

Amit Shyam Jaisinghani

Amit is a Software Engineer on the SageMaker Studio team at Amazon Web Services, and he earned his Master’s degree in Computer Science from Rochester Institute of Technology. Since joining Amazon in 2019, he has built and enhanced several AWS services, including Amazon WorkSpaces and Amazon SageMaker Studio. Outside of work, he explores hiking trails, plays with his two cats, Missy and Minnie, and enjoys playing Age of Empire.

Arun Shanmugam

Arun Shanmugam

Arun is a Senior Analytics Solutions Architect at AWS, with a focus on building modern data architecture. He has been successfully delivering scalable data analytics solutions for customers across diverse industries. Outside of work, Arun is an avid outdoor enthusiast who actively engages in CrossFit, road biking, and cricket.

Query unstructured data in Amazon SageMaker Catalog using generative AI

Post Syndicated from Nishchai JM original https://aws.amazon.com/blogs/big-data/query-unstructured-data-in-amazon-sagemaker-catalog-using-generative-ai/

Each day, businesses generate massive amounts of unstructured data, such as PDFs, images, email, customer feedback, and medical reports. But knowing data exists isn’t enough. You need to find it, access it, and extract answers from it fast. In Part 1 of this series, you saw how to set up the producer side of the pipeline: using Amazon Textract and Anthropic Claude on Amazon Bedrock to extract and enrich metadata, and then publish those enriched assets to Amazon SageMaker Catalog so your organization can discover them.

In this post, you take the next step: the consumer side. You sign in as a data consumer, search for and subscribe to the enriched unstructured data assets, and then query them using two approaches. The first is a no-code chat agent for natural language queries. The second is Amazon Bedrock model inference for programmatic access. By the end of this post, you will know how to unlock the business knowledge inside your unstructured data and make it available to analysts and application engineers alike.

Solution overview

This post continues the two-part series architecture, where Amazon SageMaker Catalog acts as the central hub connecting data producers and consumers through a publish-subscribe model.

The consumer workflow picks up after the producer has enriched and published the unstructured data assets. As a consumer, you will:

  • Sign in to your SageMaker Unified Studio consumer project and search the catalog using keywords from the enriched metadata README.
  • Subscribe to the published Amazon Simple Storage Service (Amazon S3) asset and get the subscription approved by the producer.
  • Interact with the subscribed data through two options:
    • Option 1 – A no-code chat agent for natural language queries (NLQs), ideal for data analysts and business users.
    • Option 2 – Amazon Bedrock model inference for programmatic NLQ integration, suited for application engineers building data-driven applications.

The following diagram illustrates the consumer workflow in this solution. The consumer (1) signs in to SageMaker Unified Studio, (2) searches the Amazon SageMaker Catalog for enriched unstructured data assets using keywords from the AI-generated metadata, (3) subscribes to the S3 data asset and receives approval from the producer, and then (4) queries the data using either the Amazon Bedrock chat agent app (Option 1) or Amazon Bedrock model inference through a Jupyter notebook (Option 2).

With both a no-code and a programmatic path, consumers across different roles, from analysts to engineers, can query data in the way that fits their workflow, while the SageMaker Catalog approval workflow maintains governed access throughout.

Consumer workflow architecture: sign in to SageMaker Unified Studio, search the SageMaker Catalog, subscribe to the S3 asset with producer approval, then query with the Amazon Bedrock chat agent or model inference

Figure 1: Consumer workflow for the publish-subscribe solution

Prerequisites

Before you begin, make sure you have completed all steps in Part 1 of this series, including:

Consume published data from the consumer project

In this section, you sign in as a consumer user in the SageMaker Unified Studio consumer project. You then subscribe to the S3 bucket by searching for a keyword that is part of the README published in Part 1.

  1. Sign in to the consumer project and search for the keyword emergency, which was added to the README file during publishing. The search returns the enriched asset that the producer published in Part 1.

    SageMaker Unified Studio catalog search for the emergency keyword, returning the enriched asset published in Part 1

    Figure 2: Catalog search results for the emergency keyword

  2. Choose the asset from the results to view its details, including the AI-generated business metadata, glossary terms, and README content. Then choose Subscribe.

    Asset details page showing AI-generated business metadata, glossary terms, and README content, with the Subscribe button

    Figure 3: Asset details with AI-generated metadata and the Subscribe option

  3. Enter analysis as the Reason for request in the Comment section, then choose Request.

    Subscription request dialog with analysis entered as the reason for request in the Comment box

    Figure 4: Subscription request with the reason for request entered

  4. Sign back in to the producer project (unstructured-producer-project) to approve the subscription request.
  5. After approval, return to the consumer project and confirm that the subscribed asset now appears under Manage, Assets, Subscribed assets.

    Consumer project Subscribed assets list confirming the approved subscription

    Figure 5: Approved subscription under the Subscribed assets tab

With the subscription approved, you can now access the enriched unstructured data through two approaches.

Option 1: As a data or business analyst, you can use the Amazon Bedrock chat agent app for natural language queries.

Option 2: As an application engineer, you can use Amazon Bedrock model inference for programmatic natural language queries.

Let’s explore both options.

Option 1: Amazon Bedrock chat agent app

The Amazon Bedrock chat agent app gives you a no-code, conversational interface to query your enriched unstructured data using natural language. As a data analyst or business user, you can ask questions in plain English. You get answers grounded in the documents your organization has ingested, without writing any code. For production workloads, especially in sensitive domains such as healthcare, you can apply Amazon Bedrock Guardrails to add content filtering and grounding validation to your model responses.

Data scientists and application engineers can also extend these capabilities by integrating the chat agent app APIs into custom applications, so users can interact with unstructured Amazon S3 data programmatically.

To set up the Amazon Bedrock chat agent app on your subscribed dataset, complete the following steps.

Prerequisite: Add the S3 data location.

Before creating the chat agent app, you need to add the S3 location of your subscribed data as a registered location in your project.

  1. Choose the Data tab in Overview.
  2. Choose the S3 bucket, and then choose Add to add the S3 location.

    Data tab in the project Overview with the S3 bucket selected and the Add button to register the S3 location

    Figure 6: Adding the S3 location from the Data tab

  3. On the S3 location page, provide the following details:
    • Add a name: producerprojectdata.
    • Add the producer’s S3 path as a new S3 location: s3://amzn-sagemaker-bucket-<domain-id>-<project-id>/medical/.

    Note: You can get the S3 location details from the technical name of your subscribed asset.

    • Choose the AWS Region, and then choose Add data to add this as a new location.

    Note: Make sure the AWS Region you select supports the Amazon Bedrock foundation models used later in this post. For a list of available models by Region, see Supported Regions and models for Amazon Bedrock.

    S3 location page with the location name, producer S3 path, and AWS Region entered before choosing Add data

    Figure 7: S3 location details and AWS Region selection

    Note: Make sure to select only the PDF files within the S3 path for the data source.

    Data source selection showing only the PDF files within the S3 path selected

    Figure 8: Selecting the PDF files as the data source

After the location is added, it appears as a selectable S3 location when creating a knowledge base in AI Apps.

Complete the following steps to configure the chat agent app:

  1. In the left navigation pane, under Generative AI, choose AI Apps.
  2. In the Build section of the page, choose Chat agent.

    AI Apps Build section with Chat agent selected in the left navigation under Generative AI

    Figure 9: Choosing Chat agent in the AI Apps Build section

  3. Expand the Data tab to create a knowledge base with your S3 bucket. On the Create a new knowledge base page, enter the following:
    • Add a name: MedicalKB.
    • Add a description: Knowledge base built from subscribed medical S3 data assets. Contains medical documents used to provide grounded, context-aware responses to medical domain queries.
    • Choose the data source. You will see the S3 bucket that you added in the previous step.
    Create a new knowledge base page with the MedicalKB name, description, and the added S3 bucket as the data source

    Figure 10: Creating the MedicalKB knowledge base from the S3 data source

  4. Choose your embedding model. You can leave the default settings and choose Create. It might take 10–15 minutes to create the knowledge base, depending on file sizes.
  5. After the knowledge base is created, on the Chat agent page:
    • Choose your preferred model from the Model menu (you can switch between different large language models as needed).
    • Under Data, choose your published S3 bucket as the knowledge base.
    • Begin interacting with the agent by entering questions in the Enter prompt field.
    Chat agent page with a model selected and the MedicalKB knowledge base chosen, ready to enter a prompt

    Figure 11: Chat agent page with the model and knowledge base selected

For example, entering “Which age groups had the highest rates of emergency department visits for tooth disorders?” returns an answer grounded in the enriched dental dataset published in Part 1.

The chat agent uses the enriched README metadata along with the underlying documents to surface contextually relevant answers. Analysts can explore unstructured content without needing to know where the data lives or how it’s structured.

Option 2: Natural language queries using Amazon Bedrock model inference

This option demonstrates how to use Amazon Bedrock model inference to query subscribed data using natural language. You can integrate this capability with external chat applications so users can run natural language queries through Amazon Bedrock.

  1. In your consumer project, choose Manage, Assets from the bottom of the left navigation pane. On the Subscribed tab, choose your subscribed S3 asset. Under Actions, choose Open JupyterLab notebook.

    Subscribed S3 asset Actions menu with Open JupyterLab notebook selected in the consumer project

    Figure 12: Opening the JupyterLab notebook from the subscribed asset

  2. This opens the JupyterLab notebook environment. Upload the s3_document_consumer_v2.ipynb notebook and run all the cells. You can download the notebook from s3_document_consumer_v2.ipynb.Note: The project role requires permissions for Amazon S3, Amazon Textract, and Amazon Bedrock. If you followed Part 1, you might already have these policies attached. For details on the required policies and guidance, see the prerequisites in Part 1.
  3. Review the notebook cells.
    JupyterLab notebook cells with the final cell showing a sample question answered by Amazon Bedrock

    Figure 13: Sample question answered by Amazon Bedrock in the notebook

    In the final cell, you find a sample question that Amazon Bedrock answers: “Which primary payer types (Medicare, Medicaid, private insurance, and so on) account for the highest proportion of dental-related emergency department visits?”

    Amazon Bedrock processes the question against the enriched content in the S3 bucket and returns a grounded answer. You can replace this sample question with any query relevant to your documents.

The Amazon Bedrock model inference approach gives you programmatic control, making it possible to embed natural language query capabilities directly into your existing data applications and business intelligence tools.

Clean up

To avoid ongoing charges, make sure to delete the resources used in this solution immediately after completing the walkthrough. The primary cost drivers are SageMaker Unified Studio notebook instances, Amazon Bedrock model inference calls, and Amazon S3 storage.

  1. Stop SageMaker Unified Studio resources:
    • Close running notebooks.
    • Stop running notebook instances.
    • Shut down unused kernels.

    Note: Running notebook instances continue to incur charges even when not in use.

  2. Clean Amazon S3 storage:
    • Delete temporary files created during processing.
    • Remove uploaded test documents that are no longer needed.

    Note: Although Amazon S3 costs are minimal, large volumes of data can accumulate significant charges, so it’s best to remove unneeded data.

Conclusion

In this post, you saw how to consume and query the enriched unstructured data assets published in Part 1 of this series. By subscribing to assets through the Amazon SageMaker Catalog publish-subscribe model, you can discover, access, and interact with your organization’s unstructured data, whether through the no-code chat agent or Amazon Bedrock model inference.

Together, both parts of this series show you how to build a comprehensive pipeline that transforms raw unstructured documents into governed, queryable knowledge assets. The combination of Amazon Textract for extraction, Amazon Bedrock for intelligent summarization and NLQ, and Amazon SageMaker Catalog for governance and discoverability means your teams can focus on extracting business insights rather than managing infrastructure.

To continue your Amazon SageMaker journey, see the following resources:


About the authors

Nishchai JM

Nishchai JM

Nishchai is an Analytics and generative AI Specialist Solutions Architect at Amazon Web Services. He specializes in building larger scale distributed applications and helps customers modernize their workloads on AWS. He thinks Data is new oil and spends most of his time deriving insights from data.

KiKi Nwangwu

KiKi Nwangwu

KiKi is an Analytics and generative AI Specialist Solutions Architect at AWS. She specializes in helping customers architect, build, and modernize scalable data analytics and generative AI solutions. She enjoys traveling and exploring new cultures.

Narendra Gupta

Narendra Gupta

Narendra is a Sr. Specialist Solutions Architect for Data & AI (Analytics) at AWS. He works with customers to design data-driven solutions and has deep expertise in data governance and cataloging.

Aditya Edara

Aditya Edara

Aditya is a Support Engineer at AWS. He serves as a Subject Matter Expert in AWS Analytics services, specializing in Amazon EMR and AWS Glue. Aditya provides expert guidance and technical support to enterprise and strategic customers, helping them optimize data analytics solutions.

Adding custom domains to AWS Lambda MicroVMs with Application Load Balancer

Post Syndicated from Frank Scarfo original https://aws.amazon.com/blogs/compute/adding-custom-domains-to-aws-lambda-microvms-with-application-load-balancer/

AWS Lambda MicroVMs is a serverless compute building block that provides VM-level isolation, near-instant startup performance, and state retention. You can now give each user or job their own execution environment to securely run just-in-time code, whether user or AI-generated. You do this without managing virtualization infrastructure or choosing between isolation, speed, and state retention. Lambda MicroVMs are powered by Firecracker virtualization, the technology underpinning AWS Lambda.

When you run a workload on AWS Lambda MicroVMs, each MicroVM is reachable at a service-generated endpoint that looks like 92cfc7f9-….lambda-microvm-….on.aws. That works, but many teams want to expose their MicroVMs under a domain they own, such as 92cfc7f9-….microvms.example.com. When a browser is the client, they also want to satisfy cross-origin resource sharing (CORS) without changing the application inside the MicroVM.

Both are achievable today, entirely from load-balancing and networking primitives. There is no Amazon CloudFront distribution and no compute in the request path. All you need is an Application Load Balancer (ALB) that terminates TLS with your AWS Certificate Manager (ACM) certificate, rewrites the Host header, and forwards the request over AWS PrivateLink. In this post you’ll deploy that pattern with the AWS Cloud Development Kit (AWS CDK), map a wildcard of custom domains onto your MicroVMs, and let the ALB handle CORS for you.

The complete, deployable example is available as a pattern on Serverless Land. This walkthrough centers on the reusable networking pattern. The sample also includes a small demo application that provisions a MicroVM and mints an access token, which we reference but do not detail here.

What you’ll build

By the end you’ll have:

  • A wildcard custom domain like *.microvms.example.com, where each <uuid>.microvms.example.com maps transparently to the corresponding MicroVM.
  • An internet-facing ALB that rewrites the incoming request’s Host header to the real MicroVM endpoint and forwards requests to it privately over PrivateLink.
  • CORS preflight and response headers handled at the ALB, with no change to the code running in the MicroVM.

Calling https://<uuid>.microvms.example.com/<path> (with the MicroVM access headers described later) reaches the right MicroVM, with your domain intact end to end.

Solution overview

The request flow looks like this:

Request flow from a browser through the Application Load Balancer, which terminates TLS and rewrites the Host header, then forwards over AWS PrivateLink to the Lambda MicroVM service.

The key component is the ALB host header rewrite, introduced in URL and host header rewrite for Application Load Balancers. A listener rule matches the incoming custom host with a regex condition, captures the MicroVM ID from the left-most label, and a host-header-rewrite transform rewrites the Host header to <uuid>.lambda-microvm.<region>.on.aws before forwarding. Because the MicroVM service front-end routes on the Host header, the request lands on the correct MicroVM, while the customer’s domain stays in the browser’s address bar the whole time.

Why not CloudFront? Why not an ALB redirect?

  • CloudFront can also rewrite Host/SNI toward the origin, but a single distribution has static origins. Mapping a wildcard of MicroVM IDs through one distribution would require a CloudFront Function to compute the origin per request. The ALB transform performs the same rewrite for the entire wildcard with zero code.
  • An ALB redirect action only issues an HTTP 301 Moved Permanently/302 Found response. The browser would follow it, and the address bar would then show the .on.aws URL, which breaks our design as it is not a real custom domain. The transform (not a redirect) is what makes the custom domain transparent.

Walkthrough

The example is an AWS CDK application. Configuration lives under the microvm-custom-domains key in cdk.json (hosted zone, wildcard base, the endpoint base to rewrite to, the PrivateLink service name, and the CORS origin). Set those values, then deploy. The sections below explain what the stack creates and why.

Prerequisites

A small VPC (two Availability Zones, which is the minimum for an internet-facing ALB) hosts the ALB and an interface VPC endpoint to the AWS managed MicroVM service. There are no NAT gateways, because nothing here needs egress, which keeps the footprint lean.

// Interface (PrivateLink) endpoint to the AWS managed MicroVM service.
const endpoint = new ec2.InterfaceVpcEndpoint(this, 'MicroVmEndpoint', {
  vpc,
  service: new ec2.InterfaceVpcEndpointService(cfg.microvmVpceServiceName, 443),
  subnets: { subnetType: ec2.SubnetType.PRIVATE_ISOLATED },
});

2. Discover the endpoint’s private IP addresses at deploy time

An ALB IP target group needs the private ENI IP addresses of the interface endpoint (one per Availability Zone). CloudFormation does not expose those IPs as a usable attribute, so the stack resolves them during deployment with an AwsCustomResource that reads the endpoint’s own ENIs by ID (DescribeNetworkInterfaces on vpcEndpointNetworkInterfaceIds).

This is the only compute the package deploys, it runs only during cdk deploy, and it is never in the request path.

3. Request a wildcard TLS certificate

ACM issues a DNS-validated wildcard certificate for *.microvms.example.com, validated through the hosted zone you imported. The ALB presents this certificate for every custom domain under the wildcard.

4. Create the ALB and the MicroVM target group

The internet-facing ALB has an HTTPS:443 listener using the wildcard certificate. The target group holds the endpoint ENI IPs as IP targets, reached over HTTPS:443.

  • Encrypted in transit. A customer-provided AWS Certificate Manager (ACM) certificate is used to securely terminate encryption between the client and the ALB. The ALB re-originates TLS to the MicroVM service so traffic stays encrypted through the network.
  • IP-based targets. The target group uses IP-based targets with the local IP addresses of the VPC endpoints.
  • Health check matcher 200,403,404. The load balancer’s health probes are unauthenticated, so the MicroVM endpoint answers them with 403. A 403 here means “endpoint is reachable,” not “auth is broken,” so the matcher treats it as healthy.
const targetGroup = new elbv2.ApplicationTargetGroup(this, 'MicroVmTargets', {
  vpc,
  protocol: elbv2.ApplicationProtocol.HTTPS,
  port: 443,
  targetType: elbv2.TargetType.IP,
  targets: targetIps.map((ip) => new elbv2t.IpTarget(ip, 443)),
  healthCheck: {
    protocol: elbv2.Protocol.HTTPS,
    path: '/',
    healthyHttpCodes: '200,403,404',
  },
});

const listener = alb.addListener('Https', {
  port: 443,
  protocol: elbv2.ApplicationProtocol.HTTPS,
  certificates: [certificate],
  // Default action for anything that doesn't match our host regex.
  defaultAction: elbv2.ListenerAction.fixedResponse(404, {
    contentType: 'text/plain',
    messageBody: 'Unknown custom domain',
  }),
});

5. Add the host-header rewrite rule

A listener rule matches <uuid>.microvms.example.com with a regex condition and rewrites the Host header to <uuid>.lambda-microvm.<region>.on.aws with a host-header-rewrite transform. The regex captures the left-most label (the MicroVM ID) and reuses it in the replacement.

At the time of writing, the CDK L2 constructs don’t yet model regex host conditions or transforms, so the example reaches the underlying CfnListenerRule to set them:

const escapedBase = customDomainBase.replace(/[.]/g, '\\.');
const matchRegex = `^(.+)\\.${escapedBase}$`;     // capture <uuid>
const replaceWith = `$1.${microvmEndpointBase}`;   // <uuid>.lambda-microvm.<region>.on.aws

const cfnRule = forwardingRule.node.defaultChild as elbv2.CfnListenerRule;

cfnRule.conditions = [{ field: 'host-header', regexValues: [matchRegex] }];

cfnRule.addPropertyOverride('Transforms', [
  {
    Type: 'host-header-rewrite',
    HostHeaderRewriteConfig: { Rewrites: [{ Regex: matchRegex, Replace: replaceWith }] },
  },
]);

6. Point Route 53 at the ALB

Wildcard A and AAAA alias records (*.microvms.example.com) target the ALB, so every MicroVM custom subdomain resolves to it.

7. Deploy

Run the following commands to install the dependencies and then deploy the application.

npm install
npx cdk deploy

Handling CORS at the ALB

If your clients are browsers calling the MicroVM from another origin, CORS is handled entirely at the ALB, with no change to the application inside the MicroVM.

The listener uses ALB header-modification attributes to insert the Access-Control-Allow-* headers on every response. A higher-priority rule answers OPTIONS preflight requests at the edge with a fast 204 response. Otherwise, preflight requests would reach the origin and be rejected without an access token.

// Insert CORS headers on every response on this listener.
const cfnListener = listener.node.defaultChild as elbv2.CfnListener;
cfnListener.addPropertyOverride('ListenerAttributes', [
  { Key: 'routing.http.response.access_control_allow_origin.header_value',  Value: cfg.corsAllowOrigin },
  { Key: 'routing.http.response.access_control_allow_methods.header_value', Value: 'GET,POST,PUT,DELETE,OPTIONS,PATCH,HEAD' },
  { Key: 'routing.http.response.access_control_allow_headers.header_value', Value: 'x-aws-proxy-auth,x-aws-proxy-port,content-type,authorization' },
  { Key: 'routing.http.response.access_control_expose_headers.header_value', Value: 'content-type,content-length' },
  { Key: 'routing.http.response.access_control_max_age.header_value',        Value: '86400' },
]);

// Answer OPTIONS preflights at the ALB.
new elbv2.ApplicationListenerRule(this, 'CorsPreflightRule', {
  listener,
  priority: 10,
  conditions: [elbv2.ListenerCondition.httpRequestMethods(['OPTIONS'])],
  action: elbv2.ListenerAction.fixedResponse(204, { contentType: 'text/plain', messageBody: '' }),
});

Because the ALB adds those headers to both the preflight 204 and the forwarded MicroVM response, a browser’s cross-origin call succeeds without any application change. Set corsAllowOrigin to * for quick testing, and pin it to your own site for anything beyond a demo.

Test it end to end

First, launch a Lambda MicroVM and mint an access token (follow Create your first Lambda MicroVM). When it’s running, the service gives you a generated endpoint that looks like:

012345678-9abc-defg.lambda-microvm.us-east-2.on.aws

To get the custom-domain equivalent, replace the endpoint suffix (.lambda-microvm.<region>.on.aws) with your wildcard base: .microvms.example.com. Everything ahead of that suffix is preserved exactly:

012345678-9abc-defg.microvms.example.com

The ALB’s rewrite rule captures whatever precedes the suffix and re-attaches it to the real endpoint base, so the mapping holds for the entire wildcard. You never register anything per-MicroVM.

With your token in hand, call the custom domain you derived:

curl "https://012345678-9abc-defg.microvms.example.com/<path>" \
  -H "X-aws-proxy-auth: <token>" \
  -H "X-aws-proxy-port: 8080"

The request travels to the ALB, which terminates TLS, rewrites the host header, and forwards over PrivateLink to the MicroVM. The response comes back under your domain.

The reference architecture also includes a single-page demo and a POST /api/provision endpoint that runs or reuses a MicroVM and mints a short-lived token. With it, you can try the flow without wiring up token creation yourself. It even performs this suffix swap for you and hands back a ready-to-click custom-domain URL. See the repository for that piece.

Important considerations

  • Authentication is still the client’s job. This pattern only rewrites Host. The client must still supply a valid, unexpired access token in X-aws-proxy-auth. This is deliberate. MicroVM tokens are per-MicroVM and short-lived, so baking them into infrastructure would be fragile and insecure.
  • Region pinning. PrivateLink is regional, so the ALB, the endpoint, and the MicroVM service must all be in the same Region.
  • Production hardening. If you adapt the sample’s provisioning endpoint, put authentication and rate limiting in front of it, pin CORS to your origin, and scope IAM to the minimum. The sample’s provisioning path is intentionally open for demonstration and is not production-safe as written.
  • Cost. You pay for the ALB and the interface endpoint (hourly plus data processing) in addition to the Lambda MicroVM usage. There is no CloudFront distribution and no per-request compute in the data path.

Clean up

Run the following command in the same directory where you deployed the application from.

npx cdk destroy

This removes the ALB, target groups, endpoint, certificate, VPC, and Route 53 records created by the stack.

Conclusion

You can front AWS Lambda MicroVMs with customer-owned wildcard custom domains using an Application Load Balancer and AWS PrivateLink. The key is the ALB’s host-header rewrite. Because the MicroVM service routes requests based on the Host header, a single rewrite rule can transparently map an entire wildcard of custom domains onto your MicroVMs. CORS is handled at the edge as well. The whole setup relies only on networking primitives, with no CloudFront distribution and no compute in the request path.

To try it yourself, deploy the reference architecture and review the ALB URL and host header rewrite launch post for more on the transform feature.

Multi-modal autoscaling with Amazon EC2 Auto Scaling: adding signals for faster, more reliable scaling

Post Syndicated from Shubhendu Dubey original https://aws.amazon.com/blogs/compute/multi-modal-autoscaling-with-amazon-ec2-auto-scaling-adding-signals-for-faster-more-reliable-scaling/

How do you handle unpredictable workload patterns that spike during promotional events or seasonal peaks? Multi-modal autoscaling with Amazon EC2 Auto Scaling combines infrastructure metrics like CPU with application-level signals, so a group scales on the demand its users create and not only on how busy the servers look. Those signals track the load that drives your business outcomes, such as sales or sign-ups.

CPU-based autoscaling works well for many workloads, but some demand does not register as CPU right away. Adding signals such as request counts and application metrics lets a group respond to the load its users create. By publishing Amazon CloudWatch custom metrics and application-driven triggers, you give Auto Scaling more information to act on.

In our testing, a group that scaled only on CPU rejected about 7,000 checkout sessions during a demand spike, and a stronger baseline that added Application Load Balancer request count still rejected about 6,900. A group that added an application metric rejected none, and it held p99 latency to about 0.43 seconds against 2.25 seconds for the CPU-only group. Predictive scaling can add a forecasting layer for cyclical demand, but it needs days of history to be useful, so we treat it as a complement. In this post, we show you how to implement multi-modal autoscaling on EC2 Auto Scaling, with code samples and results from a controlled test.

Prerequisites

To follow along, you need access to the following AWS services with appropriate permissions:

  • EC2 Auto Scaling, for scaling policies and group management.

  • CloudWatch, for metrics, alarms, and dashboards.

  • AWS CloudFormation, for infrastructure deployment.

Expanding beyond single-metric scaling

The default target tracking policy in EC2 Auto Scaling uses average CPU utilization, a practical starting point because CPU usage is a universal characteristic of compute workloads. Adding complementary signals, such as application-level metrics or predictive forecasting, gives Auto Scaling more information to make timely capacity decisions.

For workloads that need a faster response from target tracking alone, see Faster scaling with Amazon EC2 Auto Scaling target tracking.

In distributed architectures, different components can have distinct scaling characteristics. An API gateway might correlate well with request rate, while a background processor scales better on queue depth. With multi-modal scaling, you can match each component’s policy to its actual workload pattern. For containerized workloads, consider event-driven autoscaling with KEDA on Amazon Elastic Kubernetes Service (Amazon EKS).

Multi-modal autoscaling architecture

Multi-modal autoscaling combines three approaches to capacity management. Reactive scaling responds to current CloudWatch metrics, such as CPU utilization, memory, network throughput, response times, and custom application indicators. Application-metric scaling brings workload-specific signals into the decision, using custom CloudWatch metrics like active user sessions, queue depth, or transaction volume. These application metrics are often the closest measurable proxy for business activity such as orders or sign-ups. Predictive scaling uses machine learning in EC2 Auto Scaling to forecast capacity needs from historical patterns, so infrastructure scales before demand increases.

With application-metric scaling, applications can scale on signals that infrastructure metrics miss. An ecommerce platform might scale on active checkout sessions, while a streaming service scales on concurrent stream counts. In the test later in this post, we use active checkout sessions as the custom metric.

Implementing multi-modal autoscaling

This section builds the configuration in layers. Start with CPU target tracking as a baseline that every group keeps, then add a custom application metric that reflects real user load. The test later in this post compares these signals against a request-count baseline. Predictive scaling is an optional forecasting layer described at the end.

Step 1: CPU target tracking

Start with the foundation that most workloads already use: a target tracking policy on average CPU utilization. Target tracking is a managed policy that adjusts capacity to keep a metric at or near a target value. It supports predefined metrics, including CPU utilization and request count per target, and custom CloudWatch metrics. When multiple target tracking policies are active, Auto Scaling coordinates them: it scales out if any policy requires it, but scales in only when all policies agree, which helps prevent oscillation.

# CPU target tracking scaling policy (ASG A)
CPUTargetTrackingPolicy:
  Type: AWS::AutoScaling::ScalingPolicy
  Properties:
    AutoScalingGroupName: !Ref AutoScalingGroupName
    PolicyType: TargetTrackingScaling
    TargetTrackingConfiguration:
      PredefinedMetricSpecification:
        PredefinedMetricType: ASGAverageCPUUtilization
      TargetValue: 70

Our test also included a second infrastructure baseline, a target tracking policy on the load balancer’s request count per target. It uses the same structure with a predefined metric:

# Request count target tracking (ASG B)
RequestCountPTTargetTrackingPolicy:
  Type: AWS::AutoScaling::ScalingPolicy
  Properties:
    AutoScalingGroupName: !Ref AutoScalingGroupName
    PolicyType: TargetTrackingScaling
    TargetTrackingConfiguration:
      PredefinedMetricSpecification:
        PredefinedMetricType: ALBRequestCountPerTarget
        ResourceLabel: !Sub "${Alb.LoadBalancerFullName}/${TargetGroupB.TargetGroupFullName}"
      TargetValue: 300
      DisableScaleIn: false

Step scaling is another option for spike handling. With step scaling, you can define different capacity increments for different alarm thresholds. It keeps evaluating the alarm during scaling activities, which can make it react faster than target tracking’s default evaluation window. Step scaling policies do not coordinate with each other.

Step 2: Add a custom application metric

Next, add a second target tracking policy on a custom CloudWatch metric that reflects application load. In our test, instances publish an active checkout sessions metric at a 10-second resolution. To act on that resolution, set a Period of 10 seconds on the policy. Without it, the policy waits for three 1-minute datapoints like any other and the high-resolution metric only adds publishing cost. With it, a scale-out can begin in about 30 seconds. The policy includes the Auto Scaling group dimension so it tracks the metric for the right group. We set the target to 100 active sessions per instance, about 75 percent of the measured per-instance capacity of 135. This leaves headroom to absorb a spike while new instances boot.

# Custom application metric target tracking (ASG C)
CustomMetricTargetTrackingPolicy:
  Type: AWS::AutoScaling::ScalingPolicy
  Properties:
    AutoScalingGroupName: !Ref AutoScalingGroupName
    PolicyType: TargetTrackingScaling
    TargetTrackingConfiguration:
      CustomizedMetricSpecification:
        MetricName: ActiveCheckoutSessions
        Namespace: ECommerce/CheckoutMetrics
        Dimensions:
          - Name: AutoScalingGroupName
            Value: !Ref AutoScalingGroupName
        Statistic: Average
        Period: 10
      TargetValue: 100
      DisableScaleIn: false

Step 3: Add predictive scaling

Predictive scaling is an optional forecasting layer. It uses machine learning in EC2 Auto Scaling to analyze historical load and scale ahead of recurring, cyclical demand, using customized metric specifications in ForecastAndScale mode. You need to provide several days of history for it to forecast well, so it complements reactive signals rather than replacing them. Start in ForecastOnly mode to watch the forecast before it drives any scaling.

Monitoring

Use CloudWatch dashboards to track how each policy contributes to scaling decisions, and set alarms on the metrics that matter for your workload, such as per-instance load or latency. Enable detailed monitoring on the launch template, with Monitoring set to true, so that the system publishes CPU metrics every minute. Without it, you cannot complete the CPU policy’s scale-in evaluation and your group will stop scaling in. Watching the policies side by side is what surfaced this scale-in behavior.

Performance results

We compared three Auto Scaling groups under an identical load profile in a single 75-minute test in the us-east-1 Region. Each group used c8g.large instances with a minimum of 6 and a maximum of 40 instances, and every group carried the same CPU target tracking policy at 70 percent as a fallback:

  • ASG A: CPU target tracking only. This is the single-signal infrastructure baseline.

  • ASG B: CPU target tracking plus an Application Load Balancer request-count policy. Request rate is a stronger infrastructure baseline than CPU alone.

  • ASG C: CPU target tracking plus the custom checkout-sessions metric at 10-second resolution, published with a Period of 10 seconds.

All the groups received the same load at the same time. During the shared ramp, arrival rate rose and every group scaled correctly, which makes the comparison fair. CPU crossed 70 percent on the CPU group, request count crossed its target of 300 on the request-count group, and all three converged to a similar size.

Ramp phase (arrivals 30% → 85%) A: CPU only B: A + ALB requests C: A + app sessions
Instances (start → peak) 6 → 9 6 → 10 6 → 10
CPU 73.6% 74.0% 69.0%
Requests per target (target 300) 319 323 297

Then arrival rate was held flat while the number of concurrent checkout sessions kept rising, a shape that infrastructure signals cannot see. The next table reports that divergence phase, measured directly from CloudWatch and the load balancer.

Measured metric (divergence) A: CPU only B: A + ALB requests C: A + app sessions
Instances (start → end) 11 → 11 11 → 11 11 → 22
Rejected checkouts 7,064 6,886 0
CPU (start → end) 67.2% → 45.1% 66.5% → 45.5% 66.2% → 36.8%
Peak sessions per instance 135 135 116
Requests per target 285 → 279 282 → 279 277 → 154

The difference is what each group could see. Arrival rate was held flat while the number of concurrent sessions rose, so CPU and request count stayed in range while the application saturated. The CPU-only and request-count groups held at 11 instances and rejected 7,064 and 6,886 checkouts. Their CPU even fell, from about 67 percent to about 45 percent, because a rejected request never reaches the work it would have done, so a policy targeting 70 percent saw spare capacity at the moment the application was failing users. The application-metric group read the rising sessions directly and scaled from 11 to 22 instances, rejecting none.

Effect on latency and errors

We measured latency and rejected checkouts on the load balancer during the test. At rest, all groups were identical. The gap opened only in the divergence phase, when concurrency rose without a matching change in arrival rate. Session slots are the scarce resource here, so sessions per instance is the causal driver of latency. The application group scales on sessions and we report latency as the outcome, rather than scaling on latency directly, which is not recommended for target tracking. The latency figures come from the load balancer’s TargetResponseTime at the end of the divergence phase. A client-side number measured over the internet would reflect network round-trip rather than the service.

Measured metric A: CPU only B: A + ALB requests C: A + app sessions
TargetResponseTime (average), end of divergence 1.122 s 1.120 s 0.284 s
TargetResponseTime (p99), end of divergence 2.254 s 2.235 s 0.431 s
TargetResponseTime (average) at warm-up 0.283 s 0.283 s 0.284 s
Rejected checkouts, drain phase 4,115 3,631 0

The application-metric group, ASG C, kept per-instance load near its target and rejected no checkouts. Its average latency at the end of the divergence phase was 0.284 seconds against 1.122 for the CPU-only group, and its p99 was 0.431 seconds against 2.254. The request-count group, ASG B, tracked its own signal within range the whole time, which is exactly why it could not react: request rate was flat while concurrency climbed.

Once every group has enough capacity, they perform the same. The value of the application signal is in the transition, the gap between when demand arrives and when the fleet is ready, which the infrastructure signals here never detected.

Handling known high-traffic events

For planned events like flash sales, scheduled scaling can pre-scale capacity ahead of time. Multi-modal scaling complements scheduled scaling by handling unplanned spikes and organic traffic that does not follow a fixed schedule.

Understanding cost implications

Running the application signal requires more instances. During the spike it held about 22 instances, against 11 on the infrastructure-only groups. That extra capacity is what kept sessions per instance near the target and stopped the group from turning checkouts away. For your own workload, the question is whether a spike’s worth of extra instances costs less than the checkouts you would otherwise reject.

Conclusion

Multi-modal autoscaling combines infrastructure metrics with application-level signals so a group scales on the demand its users create, not only on how busy its servers look. In our test, a group that scaled only on CPU rejected about 7,000 checkout sessions during a demand spike, and a stronger baseline that added load balancer request count still rejected about 6,900. A group that added a custom application metric rejected none, and held p99 latency near 0.43 seconds against 2.25 seconds for the CPU-only group. Its CPU even fell while the infrastructure groups were failing requests, which shows why an infrastructure signal alone can miss the demand that matters.

Start with CPU target tracking as a fallback. Add a signal that reflects the load your users create, and pick the one that tracks closest to a business outcome like orders or active users. Set a Period on a high-resolution custom metric so the policy can act on it, and enable detailed monitoring so scale-in works. Predictive scaling is worth adding for demand you can forecast, once the group has days of history to learn from.

To implement multi-modal autoscaling, you can open Amazon EC2 Auto Scaling in the AWS Management Console and add a second scaling signal to one of your existing groups, following the configuration steps in this post. For a deeper look at target tracking behavior, see Faster scaling with Amazon EC2 Auto Scaling target tracking. The Amazon EC2 Auto Scaling User Guide covers predictive scaling policies, custom metrics, and scaling cooldowns in detail.

Planning for disaster recovery using AWS Local Zones and AWS Outposts racks

Post Syndicated from Brianna Rosentrater original https://aws.amazon.com/blogs/compute/planning-for-disaster-recovery-using-aws-local-zones-and-aws-outposts-racks/

AWS customers with data residency, low latency, or local data processing requirements can use AWS Hybrid Cloud services to run their workloads either on-premises or within their regulatory boundary. Many of these workloads might be critical to their business, with minimal thresholds for downtime.

This post provides practical design guidance for building highly available architectures that span either two AWS Outposts racks or an Outpost rack and an AWS Local Zone, which are physically designed without single points of failure. By distributing workloads across two geographically and logically independent edge locations, you can achieve high availability while still benefiting from the low-latency, data-residency, and on-premises integration advantages that edge infrastructure provides. To maintain high availability, we recommend that you put a disaster recovery (DR) plan in place and conduct regular DR drills with your applications.

The architectures presented here cover a range of approaches to failure detection and site switching. Each approach offers a different balance between Recovery Time Objective (RTO) and Recovery Point Objective (RPO), operational complexity, and cost. By understanding these trade-offs, you can select the architecture that best aligns to your RPO/RTO targets, data protection and residency requirements, and budget. This helps you achieve the resilience your business requires without over-engineering or over-spending.

Overview

Outposts and Local Zones function as extensions of a single Availability Zone (AZ) within the AWS Region they’re anchored to. For high availability when planning for failover between the two platforms, anchor each to a different parent Region or, at minimum, a different AZ within the same Region. This geographic separation supports the low RPO and RTO targets required for mission-critical workloads. The architectures in this post follow these principles:

  • Shared responsibility: AWS manages the Outposts and Local Zone infrastructure. You provide resilient power, cooling, and network connectivity for Outpost sites, and implement application-level failover logic.
  • Independent failure domains: Treat each site as an independent failure domain. Anchoring each to a different parent AZ (or Region) ensures a failure in one AZ doesn’t affect both sites.
  • Resilient network connectivity: Local Zones connect to their parent Region through the AWS Global Network, designed for maximum resilience. Outpost racks include redundant Outpost Networking Devices (ONDs) with eBGP peering for multipath load balancing and failover.
  • Capacity planning for N+1: Provision additional capacity beyond your expected workload so surviving instances can absorb the load during host failures without degradation.

Building blocks of a disaster recovery strategy

A key design consideration is how quickly the architecture can detect a site failure and redirect traffic, and what layers of your workload need protection. Your RPO and RTO needs govern this requirement. This post covers three approaches to disaster recovery at different layers of your application, each offering a different balance between time-to-recovery and operational complexity:

  1. Active/passive DNS-based failover with Amazon Route 53 health checks.
  2. Active/active architecture using physical or virtual load balancers deployed at each site.
  3. Hybrid database recovery using native database engine replication with Amazon Relational Database Service (Amazon RDS).

Depending on your workload, you can implement a combination of these strategies to support the various layers of compute and storage of your application.

Active/passive DNS-based failover with Amazon Route 53 health checks

If your workload consists of on-premises web servers accessible from the internet or internal network, you can use a DNS-based failover approach to reroute traffic to a healthy web server in the event of a hardware failure or site outage. Although this method supports any DNS service, the following architecture example uses Amazon Route 53.

DNS-based failover supports two primary approaches. The first is health check routing, where DNS resolves requests to the IP address of a known good service endpoint. The second is multi-value routing, where the DNS service returns multiple IP addresses. Clients attempt connection to the first address and automatically fail over to subsequent addresses if the connection times out. Route 53 health checks continuously monitor endpoint availability. When a site becomes unreachable, Route 53 automatically updates DNS responses to route traffic to the surviving site. This approach is globally available and works across both Outposts and Local Zones.

DNS-based failover architecture showing Route 53 health check monitoring, automatically routes to alternate health endpoint if primary endpoint fails health checks. This is an active/passive architecture.

Figure 1: Active/passive DNS-based failover architecture

When a specific application server fails and Route 53 determines it is unreachable, it is dynamically removed from future DNS responses. DNS systems typically have a Time to Live (TTL) of 300 seconds or longer, during which the DNS resolution is cached locally in the client. During this window, the client uses the cached IP address. New requests are automatically directed to active servers. The total recovery time is governed by the combination of the DNS TTL and health check timeout settings, typically resulting in a recovery time of 5 minutes or the TTL setting.

This design pattern works between Outposts, between an Outpost and a third-party provider, between an Outpost and a Local Zone, or between Local Zones. Route 53 can also distribute traffic across these sites, supporting blue/green deployments where you gradually shift traffic from one environment to another.

For AWS Outposts, you can configure Route 53 to monitor an endpoint in the Region. If the Outpost service link disconnects for more than 5 minutes, DNS failover routes traffic to the secondary site. The Outpost and Local Zone can be anchored to the same or different Regions for added resiliency.

As with all architectures using the public internet for replication traffic, configure Transport Layer Security (TLS) encryption in transit, security groups, and network access control lists (NACLs) to secure your data and control access to your subnet resources.

Active/active architecture using physical or virtual load balancers

For Outposts-to-Outposts high availability when your workload must remain on-premises, an alternative to DNS-based failover is an active/active architecture using physical or virtual load balancers deployed at each site. Outposts racks support Application Load Balancer (ALB) as well as third-party L4 and L7 virtual or physical load balancers. Like the DNS-based architecture pattern, you can use this strategy to support workloads that consist of on-premises web servers with low latency, data residency, or continued operations requirements.

In this model, both Outposts can simultaneously serve application traffic, with load balancers continuously monitoring the health of instances. When a failure is detected, the load balancer automatically shifts all traffic to the available Outpost without manual intervention or DNS propagation delays. Typically, the load balancers present a single IP address to service consumers and switch traffic when an endpoint is unavailable. Some load balancers can monitor load and switch traffic based on utilization to maintain response time. This design pattern is specific to Outposts, which support third-party devices connected on premises. It does not work with Local Zones, which are hosted in AWS datacenters.

Active/active architecture using load balancers at each site with data being replicated between sites.

Figure 2: Active/active architecture using physical or virtual load balancers

When you deploy this architecture, make sure the load balancer tier itself does not become a single point of failure. Deploy redundant load balancer instances at each Outpost, with failover between them, so the traffic management layer stays available even if one load balancer instance fails. We also recommend that you configure session persistence and connection draining on your load balancers to minimize disruption to in-flight requests during failover. With this approach, load balancer instances route traffic to your Outpost instances over the local gateway of each Outpost. Traffic continues to be balanced between instances on each Outpost even if one of the Outposts loses its service link connection. You can anchor the Outposts to the same or different Availability Zones or Regions for added resiliency. This approach does require 2N infrastructure and an external load balancer, making it the most resource-intensive to implement.

Some load balancers also support multiple endpoint monitoring. The load balancer monitors both the regional instance and the local application. If the service link fails, based on the administrator’s policy, it can drain connections and route traffic to the other Outpost. This keeps service status and logging fully available on the connected Local Zone or Outpost.

Hybrid database recovery using native database engine replication

If you have two or more logical Outpost racks, you can deploy Amazon RDS on AWS Outposts with Multi-AZ high availability. However, depending on your workload criticality, number of sites, and site locations, a more cost-effective disaster recovery option using one Outpost, one Local Zone, or both might be appropriate. For applications that require a database, you can use your chosen database engine’s native replication features or third-party tooling to create hybrid database architectures across an Outpost and a Local Zone, an Outpost and the Region, or a Local Zone and the Region. If using the Region for failover, this can be the same Region your Outpost or Local Zone is anchored to, or a different Region of your choosing. Limitations based on your chosen database engine and licensing terms apply. In this post, all architecture patterns use a PostgreSQL database. The following three hybrid database strategies expand on the hybrid database with Amazon RDS and AWS Outposts architecture to show how this design pattern supports disaster recovery across Outposts, Local Zones, and AWS Regions.

These architectures use a bring-your-own-license (BYOL) model. The replica instance used for high availability and disaster recovery (HA/DR) is customer-managed, running on Amazon Elastic Compute Cloud (Amazon EC2) and Amazon Elastic Block Store (Amazon EBS). The primary database instance can also be customer-managed, or it can be an RDS-managed database instance so you can use a managed service as your primary operating model. Promoting a replica to primary after a failure is a manual process, but you can automate it with infrastructure as code. Promotion requires updating your DNS entry for the database instance.

Architecture diagram showing database failover from an Outpost rack to a Local Zone. RDS can only run on 1 platform (Outposts, Local Zones, or in Region) and does not support RDS-native read replicas across platforms. Some database engines support native replication features, and customers can implement a self-managed replica using EC2 and EBS.

Figure 3: Database failover from an Outpost rack to a Local Zone

In the preceding diagram (Figure 3), the Outpost and the Local Zone can be in the same or different Regions for added resiliency.

In the following diagram (Figure 4), replication traffic can use either the service link or the local gateway of the Outpost as its network path. Replication continues through the local gateway even if the service link fails. If using the service link, the EC2 replica database instance must be in the Outpost anchor Region. If using the local gateway, the EC2 replica database instance can be in the same Region as or a different Region from the Outpost anchor Region for added resiliency. You need to configure a Virtual Private Gateway, Transit Gateway, or Internet Gateway in the Region to receive the replication traffic from the Outpost.

Architecture showing database failover from an Outpost rack to an AWS Region. RDS can only run on 1 platform (Outposts, Local Zones, or in Region) and does not support RDS-native read replicas across platforms. Some database engines support native replication features, and customers can implement a self-managed replica using EC2 and EBS.

Figure 4: Database failover from an Outpost rack to an AWS Region

In the following diagram (Figure 5), both the primary database instance and the replica are self-hosted on EC2 and EBS. Check your specific Local Zone location for currently supported services to see if the primary database instance can use RDS. The Region used for the EC2 replica DB instance can be the same Region the Local Zone is a part of, or a different Region for added resiliency. If using a different Region, additional networking such as an Internet Gateway is required.

Architecture showing database failover from a Local Zone to an AWS Region. Both the primary and Region database and replica instances are customer-managed using EC2 with EBS.

Figure 5: Database failover from a Local Zone to an AWS Region

In all three architectures, you need to update your DNS records and routing to complete failover to the secondary location. If your workload requires data residency, consider whether you can use an AWS Region as a failover destination.

Disaster recovery overview

The strategies discussed in this post support different RTO/RPO objectives. Recovery time depends on the amount of effort to redeploy or reroute to an alternate environment, and whether this process is manual or automated. Recovery point depends on whether the workload has persistent data that needs to be replicated, whether that replication happens synchronously or asynchronously, and whether you use a backup and restore approach. The following table is a high-level overview of the RTO/RPO you can expect for each approach based on these factors:

Architecture RTO RPO
Active/passive DNS-based failover Total failover time = DNS TTL + (health check interval x failure threshold) Equal to replication schedule, or backup interval
Active/active with load balancers Seconds, traffic is already being routed to both environments Seconds, data is already being synchronously replicated between sites
Hybrid database (same anchor Region) Minutes, time needed to reroute to replica instance Equal to replication schedule, faster replication window expected for data traveling less distance
Hybrid database (different anchor Region) <1 hour, time needed to reroute to replica instance Equal to replication schedule, longer replication window expected for data traveling a greater distance

Table 1: RTO/RPO disaster recovery overview for each architecture

For the active/passive DNS-based failover architecture, DNS TTL, Route 53 health check interval, and failure threshold are all settings you configure to your preferences. The default Route 53 health check interval is 30 seconds, but can be set as low as 10 seconds. The default Route 53 failure threshold is 3 failed checks, but can be set to any number between 1 to 10. Generally, active/active architectures provide the lowest RTO/RPO for your workloads, whereas active/passive architectures incur some downtime during a disaster when rerouting user traffic to your passive standby environment. Review your workload RTO/RPO objectives to determine which approach is right for you. You might require different strategies for different tiers of workload based on your threshold for downtime at each tier.

Considerations

When choosing a disaster recovery strategy, consider:

  • Latency impact based on the location of your failover site and where your application users are.
  • Resilient network connectivity between your primary and secondary failover locations, or between your on-premises site and the AWS Region. Architecture-specific guidance is included in each section.
  • If your workload requires data residency, evaluate if a particular disaster recovery approach can be used.
  • Promoting a replica (either RDS-managed or customer-managed) is a manual process that you can automate with infrastructure as code, and it requires updating your DNS entry for the database instance.
  • Database replicas might support synchronous or asynchronous replication depending on the database engine. Consider your RPO objectives when evaluating the hybrid database architectures.
  • Limitations based on your chosen database engine and licensing terms apply. Consult your licensing terms and conduct failover drills to test these architecture patterns with your workloads before implementing into production.

Conclusion

This post showed different architecture patterns for disaster recovery using both Outposts and Local Zones. See Building highly resilient applications with on-premises interdependencies using AWS Local Zones for additional guidance. Reach out to your AWS account team to learn more about the hybrid edge architectures discussed in this post. To discuss Outposts with an expert on any of these topics, submit the AWS Outposts contact form. To begin using Local Zones, enable a Local Zone from your account and start experimenting.

Deploying regulated workloads on AWS Local Zones and AWS Outposts

Post Syndicated from Brianna Rosentrater original https://aws.amazon.com/blogs/compute/deploying-regulated-workloads-on-aws-local-zones-and-aws-outposts/

Customers in many industries and geographic locations have specific data sovereignty and residency objectives. AWS Local Zones and AWS Outposts are fully managed infrastructure solutions for customers that need to keep data within specific geographic boundaries and also want the scalability and innovation of cloud services. The challenge lies not only in where data resides, but in how to architect, secure, and audit these deployments effectively. Whether you’re architecting a new solution or migrating existing regulated workloads to AWS hybrid edge infrastructure, this post provides an overview of key technologies to help you build auditable and secure architectures that support your organization’s data residency objectives.

Solution framework

Building a solution for data residency deployments on AWS hybrid infrastructure requires a thoughtful, layered approach. Rather than a prescriptive solution, this post presents a flexible framework that you can adapt to your specific operational requirements.

The AWS Shared Responsibility Model clearly delineates where the responsibilities of AWS end and yours begin. This model provides a critical separation: AWS controls the management infrastructure, while your data remains inaccessible to AWS operators, as enforced by the hardware-based isolation of the Nitro System. There is no operator access to the instances, applications, or data. This architectural separation provides the foundation for implementing stringent data residency controls.

To build upon this foundation, you can implement security best practices by following the guidance in the AWS Well-Architected security pillar, which helps you strengthen application-level protections and data security controls. For deeper guidance, see the Data Residency with Hybrid Cloud Services Lens, which covers considerations for operations, security, cost, performance, and reliability for regulated workloads.

When implementing data residency controls, you might need auditable evidence of traffic patterns for your internal governance processes. By using third-party monitoring tools combined with port mirroring capabilities, you can generate reports that show all traffic between your applications and databases remains within your Outpost environment. This visibility provides auditable evidence that traffic remains within your designated boundaries. You can also use AWS Artifact to access audit reports for your hybrid infrastructure.

Governance tools form the final layer of this regulatory framework, establishing guardrails around your deployment. These tools continuously monitor and enforce configuration policies, verifying that your environment stays aligned with your security and governance policies, operates within required parameters, and alerts you proactively when issues arise. This shift from reactive to proactive management helps you maintain consistent governance of your environment at scale.

Together, these layered technologies create a framework for deploying regulated workloads designed to support your data residency objectives while benefiting from the innovation and scalability of AWS services.

Shared responsibility model

When extending workloads to Local Zones and Outposts, the shared responsibility model adapts to these hybrid cloud environments while maintaining the same core principles. AWS continues to manage the underlying infrastructure and services, while you retain control over your data, applications, and configurations. This supports consistent security postures whether workloads run in AWS Regions, Local Zones, or on Outposts infrastructure. You deploy Outposts in a data center or colocation facility of your choice. Under the shared responsibility model, you are responsible for meeting site requirements for power, cooling, on-premises networking, and the Outpost service link connection to the Region. All traffic between the Outpost and the parent Region traverses an encrypted set of VPN connections over the service link, protecting communications in transit without requiring additional configuration. AWS continues to be responsible for maintaining the Outposts hardware as a managed service.

This partnership approach to security means you can build auditable solutions with data residency controls without compromising on the innovation and scalability that AWS provides.

AWS Shared Responsibility Model showing AWS responsibility for infrastructure and customer responsibility for data and configurations

Figure 1: The AWS Shared Responsibility Model in a hybrid edge deployment

AWS Nitro System

The AWS Nitro System is the virtualization platform that powers Amazon Elastic Compute Cloud (Amazon EC2) instances. It uses dedicated hardware and software to offload virtualization functions from the server CPU and delivers near-bare-metal performance. Both Outposts and Local Zones also use the Nitro System. By design, the Nitro System has no operator access. There is no way for AWS or any entity to log into the EC2 Nitro hosts, access compute resources, or reach encrypted customer data remotely. The following diagram shows the purpose-built hardware components of the Nitro System.

AWS Nitro System stack showing the Nitro Card, Nitro Security Chip, and Nitro Hypervisor components

Figure 2: The AWS Nitro System hardware and software stack

The Nitro System combines purpose-built hardware consisting of the following key security components:

  • The Nitro Card – provides I/O interfaces used for Amazon Virtual Private Cloud (Amazon VPC) network virtualization, Amazon Elastic Block Store (Amazon EBS), and instance storage, freeing up host CPU resources. Nitro Cards are logically isolated from the system main board that runs customer workloads and can be live-updated, reducing the need for maintenance windows and workload disruption.
  • The Nitro Security Chip – provides the link between the Nitro Controller (used for orchestration) and the system main board. It intercepts and controls all firmware updates, preventing the main CPUs from being used to modify system firmware. This is particularly important when running bare metal EC2 instances. This chip is also used for boot control to validate system firmware integrity.
  • The Nitro Hypervisor – designed to receive EC2 instance management commands sent by the Nitro Controller, provide compute virtualization and logical instance isolation, and assign SR-IOV virtual functions as needed. It includes no general-purpose operating system features, only the features absolutely necessary for its function, and works with other purpose-built Nitro components to maintain its small size and bare-metal-like performance. This simple design reduces the risk for remote networking attacks and driver-based privilege escalations.
  • The Nitro Security Key (Outposts only) – a removable device that stores the external key required to decrypt all data at rest on your Outpost. At the end of your Outposts commitment, after migrating your data off the Outpost, you can destroy this key to cryptographically shred any remaining data on the Outpost.

These components work together to provide a layered security approach that doesn’t compromise performance. By designing each component to have a specific function decoupled from the main system board, the Nitro System provides non-disruptive firmware updates and reduces classes of security issues often found in other hypervisor systems.

AWS Organizations Service Control Policies

AWS Organizations Service Control Policies (SCPs) are a governance tool that helps you enforce data residency requirements by controlling where resources can be created and where data can be stored or processed. SCPs function as permission guardrails that define the maximum available permissions for IAM users and roles across your organization’s accounts. By implementing deny guardrails through SCPs, you can prevent resource provisioning in unwanted locations by restricting access to AWS APIs at the infrastructure level.

When deploying regulated workloads on Local Zones and Outposts, SCPs work in conjunction with AWS Control Tower landing zones to create custom guardrails that control data movement, processing, and storage. These policies can be designed with either preventative rules (blocking actions before they occur) or detective rules (identifying compliance violations after the fact). SCPs can restrict data transfer, saving, or snapshot creation outside a specified AWS location, and they can isolate workloads to a specific location. You can apply SCPs across accounts and organizational units (OUs) within your organization. For more information, see Best practices for managing data residency in AWS Local Zones using landing zone controls and Architecting for data residency with AWS Outposts rack and landing zone guardrails.

Here’s an example SCP that restricts EC2 instance launches and network interface creation to only specified AWS Local Zone subnets:

{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "DenyNotLocalZonesSubnet",
            "Effect": "Deny",
            "Action": [
                "ec2:RunInstances",
                "ec2:CreateNetworkInterface"
            ],
            "Resource": [
                "arn:aws:ec2:*:*:network-interface/*"
            ],
            "Condition": {
                "ForAllValues:ArnNotEquals": {
                    "ec2:Subnet": [
                        "arn:aws:ec2:us-west-2:123456789012:subnet/subnet-localzone1",
                        "arn:aws:ec2:us-west-2:123456789012:subnet/subnet-localzone2"
                    ]
                }
            }
        }
    ]
}

Compliance monitoring

After you implement the security and governance best practices described in the preceding sections, you can demonstrate that traffic remains within your designated boundaries by using Amazon VPC Traffic Mirroring (also called port mirroring outside of AWS). This mirrors traffic between your application servers and databases. You can use a mirror target report to show that the traffic does not transit the AWS Region. For step-by-step instructions, see Get started using Traffic Mirroring to monitor network traffic. The key configuration steps include the following:

  1. Configure security groups – Allow inbound UDP port 4789 only from the security group of the source instances being mirrored, or from specific private CIDR ranges within the VPC. Do not open this port to 0.0.0.0/0.
  2. Create a traffic mirror target – Use the elastic network interface (ENI) of your monitoring instance.
  3. Create a traffic mirror filter – Define which traffic to capture, either all traffic or specific traffic.
  4. Create mirror sessions – Create one for each source instance you want to monitor. Lower session numbers are evaluated first when multiple sessions exist.
  5. Capture traffic – Use tcpdump on the target instance to analyze mirrored packets.
Amazon VPC Traffic Mirroring architecture on Outposts, mirroring traffic between application servers and databases to a monitoring instance

Figure 3: Amazon VPC Traffic Mirroring architecture on an Outpost

All instances must be in the same VPC, or connected through VPC peering or an AWS Transit Gateway. Traffic Mirroring encapsulates the mirrored traffic using VXLAN on UDP port 4789. Traffic Mirroring might impact network performance on source instances, so test in a development environment before deploying to production. The following image shows a sample traffic mirroring report that uses NetFlow Analyzer. For this post, all network traffic shown is simulated.

NetFlow Analyzer sample report showing traffic captured from an environment with VPC Traffic Mirroring configured

Figure 4: Sample traffic mirroring report in NetFlow Analyzer

Clean up

If you tested the VPC Traffic Mirroring architecture described in the preceding section, terminate any unnecessary resources to avoid ongoing costs. Remove the resources in the following order to avoid dependency errors:

  1. Delete the traffic mirror sessions – In the Amazon VPC console, navigate to Traffic Mirroring, Mirror Sessions. Select each mirror session you created and choose Actions, Delete. Repeat for all sessions associated with your source instances.
  2. Delete the traffic mirror filter – Navigate to Traffic Mirroring, Mirror Filters. Select the filter you created and choose Actions, Delete. You must delete all associated mirror sessions before you can delete the filter.
  3. Delete the traffic mirror target – Navigate to Traffic Mirroring, Mirror Targets. Select the target pointing to the ENI of your monitoring instance and choose Actions, Delete.
  4. Revoke security group rules – Navigate to Security Groups and select the security group attached to your monitoring instance. Remove the inbound rule that allows UDP port 4789 from the security group or CIDR range of the source instances.
  5. Terminate the monitoring instance (optional) – If you launched a dedicated EC2 instance solely for traffic capture and analysis, navigate to the EC2 console and terminate the instance. This also releases the associated ENI used as the mirror target.
  6. Delete any stored packet captures (optional) – If you saved tcpdump output to Amazon Simple Storage Service (Amazon S3) or local storage on the instance, delete those files if they are no longer needed for audit reporting.

You can verify that all Traffic Mirroring resources have been removed by running the following AWS Command Line Interface (AWS CLI) commands:

aws ec2 describe-traffic-mirror-sessions
aws ec2 describe-traffic-mirror-targets
aws ec2 describe-traffic-mirror-filters

Each command should return an empty list, confirming that no mirroring resources remain active in your account.

Conclusion

In this post, we covered how the AWS Nitro System, AWS Organizations SCPs with an AWS Control Tower landing zone, and VPC Traffic Mirroring provide capabilities for governing workloads with data residency requirements. Apply the SCP example in this post to test restricting instance launches and network interface creation to specific subnets. To learn more about Outposts for hybrid deployments, review the Getting started with AWS Outposts guide and submit the AWS Outposts contact form. To get started with Local Zones, review the Getting started with AWS Local Zones guide, opt in to a Local Zone, and begin trying some of the architecture patterns described in this post.

Architecting a secure landing zone in the AWS European Sovereign Cloud

Post Syndicated from Pablo Pagani original https://aws.amazon.com/blogs/security/architecting-a-secure-landing-zone-in-the-aws-european-sovereign-cloud/

The AWS European Sovereign Cloud is a new, independent cloud for Europe, physically and logically separate from existing AWS Regions and operated within the European Union (EU). It provides the same services, features, and APIs as AWS commercial Regions, but runs as a distinct AWS partition (aws-eusc), with its own control plane, AWS Identity and Access Management (IAM), billing, console, and service endpoints. Understanding the partition boundary is the key that unlocks correct answers to questions about billing roll-ups, single sign-on (SSO), cross-account roles, AWS Direct Connect, and image distribution. In this post, we show you how to architect a secure, scalable landing zone in the AWS European Sovereign Cloud. We cover account structure and governance, identity managed as infrastructure as code (IaC), centralized logging to a security and event management (SIEM) tool, data protection, network and perimeter design, secure continuous integration and delivery (CI/CD) and artifact distribution, and incident response. Throughout, we map the design to the AWS Security Reference Architecture (AWS SRA) and the AWS Well-Architected Framework, and we call out which behaviors are platform boundaries of a sovereign partition and which are configuration choices you can adapt.

If you are evaluating compliance readiness alongside your landing zone build-out, see the companion post Landing Zone Accelerator Independent Assessment Report for C5:2020 now available on AWS Artifact. This post covers how to align with C5:2020 criteria and provides an independent assessment report and compliance workbook, resources that complement the architectural patterns described here.

The foundational concept: EUSC is a partition

AWS groups Regions into partitions. Every Region is in exactly one partition, and each partition has one or more Regions. Partitions have independent instances of AWS Identity and Access Management (IAM) and provide a hard boundary between Regions in different partitions. AWS commercial Regions are in the aws partition, Regions in China are in the aws-cn partition, and AWS GovCloud Regions are in the aws-us-gov partition. The AWS European Sovereign Cloud is the aws-eusc partition, with its first Region in Brandenburg, Germany (eusc-de-east-1).

Some AWS services provide cross-Region functionality, such as Amazon S3 Cross-Region Replication or AWS Transit Gateway Inter-Region peering. These capabilities work only between Regions in the same partition. You can’t use IAM credentials from one partition to interact with resources in a different partition. There are practical differences that impact your architecture, shown in the following table:

Dimension Commercial AWS (aws) AWS European Sovereign Cloud (aws-eusc)
ARN prefix arn:aws: arn:aws-eusc:
Console or endpoint domain amazonaws.com amazonaws.eu
AWS Organizations One organization in the partition A separate, independent organization
AWS IAM Identity Center Instance in the partition A separate instance in the partition
Billing Consolidated in the partition’s payer A separate payer and billing system (EUR currency)
Cross-partition features and services such as: sts:AssumeRole, VPC peering, Transit Gateway, AWS RAM, Amazon S3 replication Cross-Region features and functionality Not available across the aws and aws-eusc boundary

Every AWS Region is sovereign by design: if you find yourself architecting across Regions, note that this partition boundary means that the centralization—one organization, one logging account, one identity source, one billing roll-up—is achievable within each partition. In the EUSC you operate an independent landing zone in aws-eusc that mirrors your commercial operating model. Where you need to bridge the two clouds (for example, a standard application CI/CD system in an AWS commercial Region deploying into EUSC), you integrate at the network or API layer with separate credentials for each partition, not with cross-partition trust. With these partition fundamentals in mind, the remainder of this post walks you through the considerations to build a production-ready landing zone in the EUSC. Each section addresses a critical layer of the architecture, starting with how to write partition-aware infrastructure code that works across both aws and aws-eusc, then moving into the organizational and governance controls that underpin everything else.

Cross-partition IaC

These Terraform and AWS CloudFormation IaC snippets demonstrate partition-aware Amazon Resource Name (ARN) construction—a pattern that helps ensure your infrastructure code works unchanged across AWS partitions (such as standard aws, GovCloud aws-us-gov, or European Sovereign Cloud aws-eusc).

In the case of using the same Terraform script from a commercial Region, ensure that arn:aws isn’t hard coded. This Terraform script derives the partition at deploy time so the same modules work in both AWS commercial Regions and the EUSC.

# Terraform: partition-aware ARNs (works unchanged in aws and aws-eusc)data "aws_partition""current" {}
data "aws_partition" "current" {}
data "aws_region" "current" {}
data "aws_caller_identity" "current" {}
data "aws_organizations_organization" "current" {}

locals {
partition = data.aws_partition.current.partition# "aws" or "aws-eusc"
account_id = data.aws_caller_identity.current.account_id
org_id= data.aws_organizations_organization.current.id
ecs_task_execution_policy_arn = "arn:${local.partition}:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy"
central_logs_bucket_arn= "arn:${local.partition}:s3:::${local.org_id}-central-logs"
}

resource "aws_iam_role" "amazon_ecs_role" {
  name = "AmazonECSrole"
  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [
      {
        Action = "sts:AssumeRole"
        Effect = "Allow"
        Sid= ""
        Principal = {
          # IAM service principals are "amazonaws.com" across all partitions,
          # this stays literal (do NOT use ${AWS::URLSuffix} here).
          Service = "ecs-tasks.amazonaws.com"
        }
      },
    ]
  })
}

resource "aws_iam_role_policy_attachment" "amazon_ecs_role_attach" {
  role= "AmazonECSrole"
  # Use ${local.partition } rather than hardcoding "aws" in the ARN.
  policy_arn = "arn:${local.partition}:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy"
}

# CloudFormation: use the AWS::Partition pseudo parameter, never a literal "aws"

AWSTemplateFormatVersion: "2010-09-09"

Description: >-
Creates an ECS task execution role. Demonstrates using the
${AWS::Partition} pseudo parameter in ARNs instead of hardcoding "aws",
so the template works across partitions (aws, aws-cn, aws-us-gov, aws-eusc).

Resources:
  ExecRole:
    Type: AWS::IAM::Role
    Properties:
      RoleName: AmazonECSroleCF
      AssumeRolePolicyDocument:
        Version: "2012-10-17"
        Statement:
          - Effect: Allow
            Principal:
              Service:
                # IAM service principals are "amazonaws.com" across all partitions,
                # so this stays literal (do NOT use ${AWS::URLSuffix} here).
                - "ecs-tasks.amazonaws.com"
            Action: "sts:AssumeRole"
      ManagedPolicyArns:
        # Use ${AWS::Partition} rather than hardcoding "aws" in the ARN.
        - !Sub "arn:${AWS::Partition}:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy"

Outputs:
  ExecRoleArn:
    Description: ARN of the created ECS task execution role
    Value: !GetAtt ExecRole.Arn

Account structure and governance

AWS Control Tower offers a straightforward way to set up and govern an AWS multi-account environment, following prescriptive best practices. AWS Control Tower orchestrates the capabilities of several other AWS services, including AWS Organizations, AWS Service Catalog, and AWS IAM Identity Center, to build a landing zone in less than an hour. Resources are set up and managed on your behalf.

We recommend following the AWS Security Reference Architecture (AWS SRA) multi-account model structure for a EUSC deployment. Use the management account only for governance, deploy universal security guardrails through service control policies (SCPs), resource control policies (RCPs), and service deployments (such as AWS CloudTrail) that will affect all member accounts in the organization.

Region-deny SCPs are commonly applied in commercial Regions, but aren’t required (at this time) in the EUSC because of the physically and logically separated nature of its design.

Other possible SCPs for the management OU:

  • Service-level guardrails – Restrict which AWS services can be used, based on your compliance posture.
  • Network perimeter controls – Enforce virtual private cloud (VPC) endpoints, deny public access patterns, and restrict egress.
  • Encryption and key management – Require AWS Key Management Service (AWS KMS) managed keys for all data-at-rest services and enforce key policies aligned with your sovereignty requirements.

Note: As additional EUSC Regions or Local Zones become available, the partition boundary continues to enforce isolation from non-EUSC Regions. If you need to restrict usage to a subset of EUSC Regions (for example, only eusc-de-east-1 but not a future eusc-de-west-1), a Region-deny SCP would become relevant at that point.

Identity: IAM Identity Center as IaC, no direct payer access

IAM Identity Center is available in the AWS European Sovereign Cloud as an independent instance within the partition. You can connect it to your external identity provider (IdP)—Microsoft Entra ID, Okta, and so on—using SAML/SCIM, exactly as in AWS commercial Regions. If you already use Identity Center to federate in the commercial partition, you can point a second Identity Center integration at the same corporate IdP, so users keep one set of credentials. You manage permission sets, groups, and account assignments separately for each partition.

Manage permission sets and assignments as code

The following Terraform defines a permission set with both an AWS managed policy and an inline least-privilege policy, then assigns a group to a target account. Reproduce the aws_ssoadmin_account_assignment for each account or organizational unit (OU) mapping.

Because group membership comes from your IdP over SCIM, the IdP handles joiner, mover, and leaver, and access in EUSC updates automatically. No one receives direct access to the management account; all human access flows through IAM Identity Center permission sets assigned to non-management accounts.

data "aws_ssoadmin_instances" "this" {}
data "aws_partition" "current" {}

# Workload account the group is assigned to.
variable "analytics_workload_account_id" {
  type        = string
  description = "Account ID of the workload account to assign the permission set to"
}

locals {
  sso_instance_arn  = tolist(data.aws_ssoadmin_instances.this.arns)[0]
  identity_store_id = tolist(data.aws_ssoadmin_instances.this.identity_store_ids)[0]
}

resource "aws_ssoadmin_permission_set" "analytics_operator" {
  name             = "AnalyticsOperator"
  description      = "Operate analytics workloads; no IAM or billing"
  instance_arn     = local.sso_instance_arn
  session_duration = "PT4H"
}

# Attach an AWS managed policy
resource "aws_ssoadmin_managed_policy_attachment" "analytics_ro" {
  instance_arn       = local.sso_instance_arn
  permission_set_arn = aws_ssoadmin_permission_set.analytics_operator.arn
  managed_policy_arn = "arn:${data.aws_partition.current.partition}:iam::aws:policy/ReadOnlyAccess"
}

# Add a least-privilege inline policy (note the partition-aware ARNs)
resource "aws_ssoadmin_permission_set_inline_policy" "analytics_inline" {
  instance_arn       = local.sso_instance_arn
  permission_set_arn = aws_ssoadmin_permission_set.analytics_operator.arn
  inline_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Sid      = "OperateAnalyticsData"
      Effect   = "Allow"
      Action   = ["s3:GetObject", "s3:PutObject", "s3:ListBucket"]
      Resource = [
        "arn:${data.aws_partition.current.partition}:s3:::analytics-*",
        "arn:${data.aws_partition.current.partition}:s3:::analytics-*/*"
      ]
      Condition = { StringEquals = { "aws:RequestedRegion" = "eusc-de-east-1" } }
    }]
  })
}

# Group for analytics operators.
# In production this is typically synced from your IdP via SCIM; here we
# manage it directly so the config is self-contained.
resource "aws_identitystore_group" "analytics" {
  identity_store_id = local.identity_store_id
  display_name      = "analytics-operators"
  description       = "Analytics operators"
}

# Assign the group to a workload account with the permission set
resource "aws_ssoadmin_account_assignment" "analytics_to_workload" {
  instance_arn       = local.sso_instance_arn
  permission_set_arn = aws_ssoadmin_permission_set.analytics_operator.arn
  principal_id       = aws_identitystore_group.analytics.group_id
  principal_type     = "GROUP"
  target_id          = var.analytics_workload_account_id
  target_type        = "AWS_ACCOUNT"
}

Cross-account roles for governance, logging, and tooling

Cross-account roles within the EUSC partition work normally; this is how the logging and security-tooling accounts collect from workload accounts. Scope each trust policy to a specific principal and harden it with an external ID (for third-party tooling) and partition-aware ARNs.

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": {
      "AWS": "arn:${AWS::Partition}:iam::<SECURITY_TOOLING_ACCOUNT_ID>:role/SecurityAuditCollector"
    },
    "Action": "sts:AssumeRole",
    "Condition": {
      "StringEquals": { "sts:ExternalId": "eusc-sec-audit" },
      "ArnLike": { "aws:PrincipalArn": "arn:${AWS::Partition}:iam::*:role/SecurityAuditCollector" }
    }
  }]
}

A role in the aws partition can’t assume a role in aws-eusc (or the reverse).

Logging and monitoring: centralized in EUSC, exported to your SIEM

A sovereign logging architecture requires three things:

  • A single, immutable store for all audit and operational logs
  • A central security account that runs detective controls and correlates findings
  • A reliable, in-partition path that feeds everything into your SIEM without data ever leaving the boundary.

In the subsections that follow, we walk through each layer: Centralized log collection in the Log Archive account, Amazon GuardDuty and AWS Security Hub administration through the Security Tooling account, and the pull-based SIEM integration pattern that keeps telemetry inside the EUSC partition.

Centralize logs in the Log Archive account

The Log Archive account holds the organization trail and a central log bucket as part of the landing zone. In the commercial AWS partition, global services like IAM route their CloudTrail events to us-east-1. In the EUSC, global services events are logged within the EUSC partition because the control plane is independent and located entirely within the EU.

Organization level detective services

GuardDuty and Security Hub are available in EUSC, but organization-wide auto-enable and some newer features might differ from commercial AWS features at any given time. Design the Security Tooling account as the delegated administrator where supported. If org-level auto-enable isn’t yet available, enable per-account through your IaC (AWS CloudFormation StackSets) so coverage is complete and code-managed. Treat the EUSC service and feature list as the source of truth and gate optional features behind a partition flag.

Network security and perimeter, including AWS Direct Connect

The AWS European Sovereign Cloud has its own sovereign AWS Direct Connect points of presence (PoPs), with dedicated networking infrastructure and connectivity from European providers, providing customers an autonomous network path into the partition. You terminate Direct Connect in a dedicated Network account and share connectivity to workload VPCs using Transit Gateway (with AWS RAM). A Direct Connect connection or Direct Connect gateway in the commercial partition can’t be extended into aws-eusc. To reach EUSC VPCs, you provision a separate Direct Connect connection that lands in the EUSC partition’s Network account. If your on-premises network already backhauls to AWS commercial Regions, you connect that network to EUSC with its own virtual interface or connection, or a site-to-site VPN. You don’t bridge the two AWS partitions through a shared Direct Connect gateway.

The following figure shows the recommended perimeter design in EUSC.

Figure 1: Recommended perimeter design in EUSC

Figure 1: Recommended perimeter design in EUSC

The perimeter design includes:

  • Centralized egress and inspection – Route workload egress through an inspection VPC in the Network account (gateway load balancer with your firewall of choice, or AWS Network Firewall. Keep workload VPCs private with no internet gateway.
  • Private service access – Use VPC interface endpoints (VPCe) for AWS service calls so traffic stays on the AWS network within the partition. VPCe doesn’t cross partitions; expose any commercial-partition service to EUSC consumers over DX/VPN and an in-EUSC load balancer.
  • DNS – EUSC has its own Amazon Route 53. For names that must resolve across clouds, use subdomain delegation or Resolver forwarding rules over your DX or VPN link rather than expecting hosted zones to be visible across partitions.
  • Segmentation as code – Express segmentation with security groups referencing prefix lists and keep the EUSC IP ranges current from the partition’s published ip-ranges file in your firewall automation.

Data protection

AWS Key Management Service (AWS KMS) is available in EUSC; use customer managed keys for all sensitive data stores and enforce their use with SCPs and key policies. Where your residency or operational-autonomy requirements call for it, evaluate AWS KMS external and imported key material options available in the partition.

For workloads where regulation mandates that key material never resides within the cloud provider’s infrastructure, configure an AWS KMS External Key Store (XKS) in the EUSC Region. The XKS proxy connects AWS KMS to your EU-based hardware security module (HSM) (on-premises or hosted with an EU trust service provider); all encrypt and decrypt operations are performed by your external key manager. Note the trade-offs: increased latency, reduced availability SLA, and added operational burden. Reserve XKS for the subset of data where regulatory or contractual obligations explicitly require it.

The EUSC Region has achieved SOC 2, BSI C5 Type 1 attestation, and seven ISO certifications, including ISO 27001, 27017, 27018, and 27701. Reference these in your data protection evidence packages when demonstrating encryption-at-rest and key management controls to EU regulators.

Secure CI/CD and distributing images across the partition boundary

If you need to deploy existing images or binaries into EUSC (aws-eusc) from existing AWS commercial (aws) accounts, you can’t use cross-partition Amazon Elastic Container Registry (Amazon ECR) replication, Amazon Machine Image (AMI) copy, or Amazon Simple Storage Service (Amazon S3) replication. Instead, treat EUSC as an independent supply-chain destination:

  • Container images – Build (or re-tag and re-sign) images and push to an Amazon ECR registry inside EUSC. ECR cross-Region replication works within the partition (useful as EUSC adds Regions or Local Zones), but the initial crossing from commercial is an explicit pipeline push using EUSC credentials. Sign images with a sovereign signing key and verify at deploy time.
  • AMIs and images – Rebuild golden images in EUSC with EC2 Image Builder (run the pipeline natively in EUSC), or import virtual machine (VM) images using Amazon S3 in EUSC and aws ec2 import-image. There is no direct cross-partition AMI copy.
  • Binaries and artifacts – Stage in an artifact bucket in the EUSC Shared Services account. Move packages across the boundary with aws s3 sync or AWS DataSync over your DX or VPN, or using controlled export, then distribute within the partition using in-partition S3 replication to other EUSC Regions or Local Zones as they come online.
# Push a container image to ECR inside EUSC (note the .eu endpoint).
# Credentials/profile must be for the aws-eusc partition.
aws ecr get-login-password --region eusc-de-east-1 --profile eusc-shared-services \
  | docker login --username AWS --password-stdin \
    111122223333.dkr.ecr.eusc-de-east-1.amazonaws.eu

docker tag company/runtime:7.x \
  111122223333.dkr.ecr.eusc-de-east-1.amazonaws.eu/company/runtime:7.x
docker push \
  111122223333.dkr.ecr.eusc-de-east-1.amazonaws.eu/company/runtime:7.x

Remember that endpoints and ARNs use amazonaws.eu in the EUSC partition. IAM service principals always use amazonaws.com regardless of partition.

Replicating deployment code and pipelines

Run a native deployment plane in EUSC (AWS CodePipeline, AWS CodeBuild, AWS CodeDeploy, or your existing tool deployed in-partition) in the Shared Services account, with cross-account deploy roles into workload accounts. If a commercial-partition continuous-integration system must deploy into EUSC, give it separate credentials for each partition; the clean pattern is OIDC federation with two trust configurations, one for each partition, because no cross-partition role assumption exists.

// Deploy role in an EUSC workload account, trusted by the EUSC Shared Services pipeline role
{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": { "AWS": "arn:aws-eusc:iam::<SHARED_SERVICES_ACCT>:role/PipelineDeployRole" },
    "Action": "sts:AssumeRole",
    "Condition": { "StringEquals": { "sts:ExternalId": "EUSC-deploy" } }
  }]
}

For account vending and landing-zone-as-code, use Account Factory for Terraform (AFT) deployed in EUSC. AFT pipelines create accounts through AWS Control Tower, apply baseline guardrails, bootstrap the preceding partition-aware modules, and register OUs, giving you the accounts and account groups as code. Keep Terraform state for each partition in an in-EUSC Amazon S3 backend with an Amazon DynamoDB lock table; don’t share state across partitions.

The Landing Zone Accelerator on AWS (LZA) solution is an alternative deployment method that provisions a baseline security architecture and includes customizations for each partition, with consideration for service availability. A customized configuration baseline for European Sovereign Cloud was recently released and is accompanied by the LZA Compliance Workbook, which maps regional European security standards and international frameworks to over 200 security settings deployed by LZA.

Supported compared to by-design boundaries: A quick reference

Capability Status in EUSC What to do
AWS Control Tower account vending, controls Supported in-partition Govern the EUSC Region; drive vending with AFT; re-register OUs after Region changes
AWS Control Tower–managed or self-managed IAM Identity Center Configuration choice Choose self-managed to own permission sets as code
Permission set creation and assignment as IaC Supported Manage with SCIM groups from your IdP
Identity Center single home and delegated admin for each partition By-design behavior Administer from one Region; non-issue in single-Region EUSC
Cross-account roles (governance, logging, tooling) Supported within partition Scope trust to specific principals and ExternalId
Cross-partition AssumeRole, VPC peering, TGW, RAM, Amazon S3 replication Not available (security boundary) Integrate at network or API layer with separate per-partition credentials
Billing roll-up across accounts and Regions Supported within the EUSC org Aggregate in a finance or governance account in-partition
Billing roll-up across the aws and aws-eusc boundary Separate billing systems (EUR payer) Keep cost analysis in-partition or in an EU-resident tool
Multi-Region image distribution (Amazon ECR, AMI, and Amazon S3) Supported within partition Push into EUSC first, then replicate in-partition
GuardDuty and Security Hub Available—some org-auto-enable and features vary Delegate admin where supported; per-account enable using IaC otherwise
CloudFront, Shield Advanced, Firewall Manager, Inspector In planning at time of publication Follow on AWS Builder Center (capabilities) for release updates

Billing and cost governance

Roll up within the EUSC organization, not across partitions. Enable consolidated billing in the EUSC management account and deliver AWS Data Exports (Cost and Usage Report 2.0) to an S3 bucket in a dedicated finance or governance account in the Security or Infrastructure OU.

Set permissions so workload teams can query the curated data in that account; no one should be able to access the management or payer account directly. You can’t replicate billing data into the commercial partition; the EUSC has a separate payer (billed in EUR through the EU contracting entity).

Conclusion

Architecting in the AWS European Sovereign Cloud is, in most respects, architecting a second well-run AWS landing zone with one organizing principle that resolves nearly every design question: it’s an independent partition. Centralization of governance, identity, logging, and billing is fully achievable, but within the EUSC partition. The boundaries you encounter between commercial AWS and EUSC—no cross-partition roles, peering, replication, or billing roll-up—are the sovereignty guarantees doing their job.

Build the foundation as code. Use an AWS Control Tower landing zone driven by AFT, IAM Identity Center permission sets and assignments in Terraform federated to your corporate IdP, an immutable central log store, customer managed encryption keys constrained to the sovereign Region, and a Network account terminating a dedicated Direct Connect. Add a CI/CD plane that pushes images and artifacts into the partition with per-partition credentials. Keep every ARN partition-aware and every optional service behind a feature flag, and the same modules will serve both clouds.

To accelerate your build with additional enablement from AWS, explore the LZA Universal Configuration for European Sovereign Cloud on GitHub, which packages many of the patterns described in this post into a ready-to-deploy baseline. To complement your deployment with compliance readiness, the LZA Independent Assessment Report for C5:2020 evaluates how LZA’s security baseline maps to C5:2020 technical requirements, and you can download the report and the LZA Compliance Workbook from AWS Artifact.

Further reading

If you have feedback about this post, submit comments in the Comments section below or start a thread on AWS re:Post.


Pablo Pagani

Pablo Pagani

Pablo is a Systems Development Manager for AWS European Sovereign Cloud, based in Madrid, Spain. He has previously held roles within Enterprise Support and Professional Services. An active member of the Security Technical Field Community, he helps customers build a secure journey on AWS. Pablo developed his passion for computers while writing his first lines of code in BASIC on an MSX computer with 64 KB of RAM.

Margo Cronin

Margo is an EMEA Principal Solutions Architect specializing in Security & Compliance and is based out of Zurich Switzerland. Her interests include security, privacy, cryptography, and compliance. She is passionate about her work unblocking security challenges for AWS customers, enabling their successful cloud journeys. She is an author of the “AWS User Guide to Financial Services Regulations and Guidelines in Switzerland”.

Discover and govern Snowflake data using SageMaker Unified Studio

Post Syndicated from Marco Duarte original https://aws.amazon.com/blogs/big-data/discover-and-govern-snowflake-data-using-sagemaker-unified-studio/

Many organizations operate in hybrid data environments where critical assets live in Snowflake while analytics workloads run on AWS, which can create governance gaps, discovery friction, and duplicated efforts when the two aren’t connected.

With Amazon SageMaker Unified Studio, you can govern data across Snowflake and AWS through its integrated catalog and AWS Glue Data Quality, a capability of AWS Glue. You connect directly to Snowflake tables without moving data, apply quality rules using AWS Glue Visual ETL, and publish validated assets to Amazon SageMaker Catalog, maintaining consistent governance across your entire distributed data estate.

Without this integration, cataloging Snowflake data requires building extraction pipelines, often taking days. With SageMaker Unified Studio connected to Snowflake, you can query, catalog, and validate the quality of federated data in 5–15 minutes. No data replication or custom ETL code required.

In this post, we show you how to connect Snowflake to Amazon SageMaker Unified Studio, register data assets in Amazon SageMaker Catalog, configure data quality validation using AWS Glue Visual ETL, and publish assets for unified collaboration. By following these steps, you enrich federated assets with data quality scores so that consumers across your organization can discover and trust the data, all while keeping it in Snowflake.

Solution overview

This solution integrates Snowflake with Amazon SageMaker Unified Studio for centralized data cataloging and quality validation.

The architecture uses an AWS Glue connection to federate the Snowflake catalog into Amazon SageMaker Unified Studio. Tables become available in the project catalog without complex storage configurations. You can query data directly using SQL analytics, publish datasets to Amazon SageMaker Catalog for organization-wide discovery, and apply data quality rules through AWS Glue Visual ETL pipelines.

The workflow consists of the following steps:

Architecture diagram: Snowflake federated into SageMaker Unified Studio through AWS Glue, with data quality validation and publishing to SageMaker Catalog

Figure 1: Architecture for federating Snowflake into SageMaker Unified Studio and validating data quality

  1. Snowflake connection creation on Amazon SageMaker Unified Studio — Amazon SageMaker Unified Studio uses an AWS Glue connection to federate Snowflake tables and views into its open data lakehouse architecture. The federated catalog entry is registered in AWS Glue Data Catalog and governed by AWS Lake Formation for centralized access control, without moving data out of Snowflake.
  2. Federate Snowflake tables into the Amazon SageMaker publisher project — The Amazon SageMaker publisher project discovers the federated Snowflake tables through the AWS Glue Data Catalog integration.
  3. Publish the dataset to Amazon SageMaker Catalog — The publisher project publishes the dataset as a governed asset to the Amazon SageMaker Catalog, making it discoverable for data consumers across the organization.
  4. Validate data quality — AWS Glue Data Quality runs validation rules against the federated Snowflake data and publishes the data quality results directly to the corresponding asset in Amazon SageMaker Catalog.
  5. Consume data — Users access Snowflake data through two paths:
    1. Publisher project users — Query data with SQL Analytics — Users in the publisher project can query the Snowflake data directly using Amazon SageMaker Unified Studio SQL Analytics for interactive exploration and analysis, without copying or moving data.
    2. Consumer project users — Discovery and subscription through SageMaker Catalog — Other Amazon SageMaker consumer projects discover the published asset in the Amazon SageMaker Catalog, subscribe to it, and consume the data for their analytics and machine learning workloads.

Prerequisites

To follow along, you need:

  • An active Snowflake account with administrator access.
  • Tables or views created within a schema inside a Snowflake database.
  • An Amazon SageMaker Unified Studio and project created.
  • An Amazon Simple Storage Service (Amazon S3) bucket for AWS Glue assets.
  • Appropriate AWS Identity and Access Management (IAM) permissions configured (Amazon SageMaker Catalog is built on Amazon DataZone, so the IAM actions use the datazone: prefix.)

Your AWS Glue job execution role requires specific permissions to interact with Amazon SageMaker Catalog.

Required IAM policies for the AWS Glue job role

1. Amazon SageMaker Catalog search and listing permissions: Attach a policy that allows the AWS Glue job to search and list assets in Amazon SageMaker Catalog.

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "datazone:SearchListings",
        "datazone:GetListing",
        "datazone:ListDomains",
        "datazone:GetDomain"
      ],
      "Resource": "arn:aws:datazone:<REGION>:<ACCOUNT_ID>:domain/<DOMAIN_ID>"
    }
  ]
}

2. Amazon SageMaker Catalog time series data posting permissions: Add permissions to post data quality metrics:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "datazone:PostTimeSeriesDataPoints",
        "datazone:GetAsset",
        "datazone:ListAssetRevisions"
      ],
      "Resource": "arn:aws:datazone:<REGION>:<ACCOUNT_ID>:domain/<DOMAIN_ID>"
    }
  ]
}

Configure the AWS Glue job role as an Amazon SageMaker domain user

Configure the IAM role used by your AWS Glue job as a domain user. In the Amazon SageMaker console, navigate to your domain, choose Access management, and add the AWS Glue job execution IAM role as a domain user.

Project-level permissions

Add the AWS Glue job execution role as a project member with Owner permissions. Navigate to your project, go to Project settings > Members, and add the role.

For more information about IAM roles for AWS Glue, see the AWS Glue security documentation. For Amazon SageMaker Unified Studio permissions, refer to the Amazon SageMaker Unified Studio administrator guide.

Querying Snowflake datasets from Amazon SageMaker Unified Studio

The following sections walk you through connecting Snowflake to Amazon SageMaker Unified Studio and running data quality validation with results displayed in Amazon SageMaker Catalog.

Identifying information in Snowflake

First, gather your Snowflake connection details. You need a Snowflake account with tables or views created at the schema level within a database.

To obtain Snowflake connection information:

  1. Navigate to your Snowflake environment and sign in with administrator credentials.

    Snowflake sign-in screen for administrator credentials
  2. Choose your user account and choose Connect a tool to Snowflake.

  3. Note the Account/Server URL displayed on the screen.
  4. Choose the Config File tab, select values for Warehouse, Database, and Schema, and copy these values for use in the next section.

Creating the connection in Amazon SageMaker Unified Studio

The Add Connection feature stores Snowflake connectivity details including credentials, server, and database information. Amazon SageMaker Unified Studio uses this connection to federate the Snowflake catalog through AWS Glue, so you can query data within minutes of setup.

You need an Amazon SageMaker Unified Studio domain and a project, which acts as a data producer project.

To create the Snowflake connection:

  1. In your Amazon SageMaker Unified Studio project, go to Overview.

    SageMaker Unified Studio project Overview page
  2. Choose Data.

    Data option in the SageMaker Unified Studio project navigation
  3. Choose + Add, then choose Add Connection.

    Add menu in SageMaker Unified Studio with the Add Connection option
    Add Connection panel in SageMaker Unified Studio
  4. Choose Next.
  5. Select Snowflake and choose Next.

    Connection type selection showing Snowflake in SageMaker Unified Studio
  6. Complete the connection details:
    • Name: snowflake-connection.
    • Description (Optional): Enter a description for your connection.
    • Host: Your Snowflake account URL (for example, XXXXXXXXX-XXX000000.snowflakecomputing.com).
    • Port: 443.
    • Database: Your database name (for example, sm_demo).
    • Warehouse: Your warehouse name (for example, COMPUTE_WH).
    • Schema: Your schema name (for example, demo).
    • Additional Properties:
      • Register in AWS Glue Data Catalog: Turn on checkbox.
      • Case conflict handling: Select the option based on Snowflake naming syntax.
    • Authentication:
      • Username: Your Snowflake username.
      • Password: Your Snowflake password.
    Snowflake connection details form with name, host, port, database, warehouse, and schema fields
    Connection form showing authentication and AWS Glue Data Catalog registration options
  7. Choose Add Data.

After creating the connection, wait a few minutes for the federated connection to be established. Search within Amazon SageMaker Unified Studio for the database and created objects.

Federated Snowflake database and objects appearing in SageMaker Unified Studio search

Federated Snowflake tables registered in the AWS Glue Data Catalog

With the Snowflake connection established and the federated tables registered in AWS Glue Catalog, you’re now ready to query Snowflake data directly from Amazon SageMaker Unified Studio, without moving or replicating any data.

Query results from a federated Snowflake table in the SageMaker Unified Studio query editor

How federated queries work

When you run a query in the Amazon SageMaker Unified Studio query editor against a federated Snowflake table, Amazon Athena runs the request. Athena is the underlying query engine integrated into Amazon SageMaker Unified Studio. Athena reads the table definition from AWS Glue Catalog, connects to Snowflake through the established connection, and pushes the query down for execution. Athena returns results directly to the query editor while Snowflake processes the data in place, and only the query results travel across the connection. Amazon SageMaker Unified Studio doesn’t copy data to S3 or any intermediate storage.

After you’ve validated that queries return the expected results, the next step is to publish this dataset to Amazon SageMaker Catalog, making it discoverable and shareable across your organization.

Publishing Snowflake datasets to the SageMaker Catalog

Now that your Snowflake connection is configured, you can publish your datasets to the Amazon SageMaker Catalog, making them discoverable and shareable across your organization.

Creating data assets in SageMaker Catalog

Data assets in Amazon SageMaker Catalog are the cataloged representation of your data resources. They help teams discover, govern, and share data across your organization.

In this section, you create a data asset associated with a Snowflake table. This process transforms a technical Snowflake table into a cataloged resource enriched with business metadata.

To create a data source:

  1. In your Amazon SageMaker Unified Studio project, go to Manage.

    Manage tab in the SageMaker Unified Studio project
  2. Choose Data Sources.
  3. Choose Create Data Source.

  4. Select the AWS Glue option.

    Data source type selection showing the AWS Glue option
  5. Turn on the Import data lineage checkbox and select the connection: project.default_lakehouse.

    Data source configuration with Import data lineage and the project.default_lakehouse connection selected
  6. Complete the form and choose Next:
    • Catalog: Select Enter the catalog name and enter snowflake-connection.
    • Database name: Enter your database name (for example, movies).
    • Table selection criteria: Enter * for all tables in the database, or enter a specific table name.
    Data source form showing catalog name, database name, and table selection criteria
  7. Keep the default options and choose Next until you reach the summary screen.

    SageMaker Unified Studio data source configuration summary screen
    Data source review screen before creation
  8. Review your settings and choose Create.

To extract metadata and publish assets:

  1. Choose Run to start extracting metadata from AWS Glue Data Catalog.

    Data source detail page with the Run option to extract metadata from the AWS Glue Data Catalog
  2. Wait for the run to complete.
  3. Go to Assets to view the Asset Inventory.

    Asset inventory in SageMaker Catalog after the data source run completes

The following screenshot shows the asset inventory after the data source run completes.

  1. Choose an asset to view its details.

    Asset detail page in SageMaker Catalog showing the Snowflake table metadata

At this point, you can enrich the business context by choosing Generate Descriptions. Amazon SageMaker Catalog analyzes the asset’s technical structure and generate:

  • Business descriptions in natural language for the asset.
  • Contextual definitions for each field/column.
  • Suggested glossary terms that could be applied.
  1. After your asset has been enriched with the necessary business metadata, you can publish it to the Amazon SageMaker Catalog by choosing Publish Asset.

Publish Asset option on the enriched Snowflake asset in SageMaker Catalog

The Snowflake enriched asset is now available to data consumers across your organization. Other users can discover it, subscribe to it, and consume it without data replication.

Implementing data quality rules with AWS Glue Data Quality

This section explains how to apply data quality validations to Snowflake data using AWS Glue Data Quality and visualize results in Amazon SageMaker Catalog.

Setting up the custom transform

Upload two files to an Amazon S3 bucket in the same AWS account where you run AWS Glue:

Copy both files to your AWS Glue assets S3 bucket in the transforms folder (s3://aws-glue-assets-<account-id>-<region>/transforms). AWS Glue Studio reads all JSON files from this folder to register custom visual transforms.

Custom transform files uploaded to the transforms folder in the AWS Glue assets S3 bucket

In the following sections, we walk you through the steps of building an ETL pipeline for data quality validation using AWS Glue Studio.

Creating the AWS Glue Visual ETL job

AWS Glue for Spark provides built-in support for reading from Snowflake data sources.

To create a new visual ETL job:

  1. Open the AWS Glue console at https://console.aws.amazon.com/glue/. Choose ETL jobs, then Visual ETL.

    AWS Glue console showing ETL jobs and the Visual ETL option

Establishing the Snowflake connection

To add a Snowflake source:

  1. In the job pane, choose Snowflake as your source. For Snowflake connection, select the connection that you created earlier. Specify the relevant schema and table for data quality checks.

    Snowflake source node configured in the AWS Glue visual ETL job

The visual editor displays the Data source properties panel where you select your connection, database, and enter a custom query targeting your Snowflake table.

Applying data quality rules

After establishing the Snowflake connection, configure the data quality evaluation step using the Data Quality Definition Language (DQDL).

To add data quality validation:

  1. Choose Transform and choose Evaluate Data Quality.
  2. Define domain-specific data quality rules using DQDL. For more information, see the AWS DQDL documentation.

    Evaluate Data Quality transform with DQDL rules in AWS Glue Studio
  3. Choose to output the data quality results. Optionally, store outcomes in Amazon S3 or publish to Amazon CloudWatch with alert notifications.

The preview of the data quality results from the ruleOutcomes node shows the outcomes of each rule.

Preview of the data quality rule outcomes from the ruleOutcomes node

Post the data quality results to Amazon SageMaker Catalog

To configure the custom transform:

  1. Add the Datazone DQ Result Sink transform to your job.
  2. Connect the ruleOutcomes node output to this transform.
  3. Complete the parameters:
    • Role to assume (Optional): Only needed for associated accounts.
    • Domain ID: Your Amazon SageMaker Unified Studio domain ID (found in the Amazon SageMaker Unified Studio portal).
    • Table name and Schema name: Same values used when creating the Snowflake source transform.
    • Data quality ruleset name: The name you want to give to the ruleset in Amazon SageMaker Catalog.
    • Max results: Maximum number of assets to return in case of multiple matches.

The following image shows the complete job graph with the Datazone DQ Result Sink transform configured.

AWS Glue visual ETL job graph with Snowflake source, Evaluate Data Quality, ruleOutcomes, and Datazone DQ Result Sink nodes

The visual editor displays four nodes connected sequentially: the Snowflake data source, the Evaluate Data Quality transform, the ruleOutcomes SelectFromCollection transform, and the Datazone DQ Result Sink transform.

To configure job parameters:

  1. Choose Job details.
  2. In Job parameters, add the following key-value pair:
    • --additional-python-modules
    • boto3>=1.34.105
  3. Save and run the job.

AWS Glue job parameters with the additional-python-modules key set to boto3

Visualizing data quality results in the SageMaker Catalog

After the AWS Glue ETL job completes, you can view the data quality information directly in Amazon SageMaker Catalog. This is the key outcome of running data quality on a federated source: the asset gains quality scores and metadata without ever leaving Snowflake. This makes it trustworthy and ready for other teams across your organization to use. Data consumers can now discover this asset in Amazon SageMaker Catalog and evaluate its quality before subscribing, without needing direct access to Snowflake or running their own validation.

To view data quality results:

  1. Open the Amazon SageMaker Unified Studio console.
  2. Navigate to your project.
  3. Go to Assets.
  4. Choose the Snowflake data asset.
  5. View the data quality information displayed on the asset page.

The following image shows the asset page in Amazon SageMaker Catalog with the data quality score populated.

SageMaker Catalog asset page showing a populated data quality score for the Snowflake asset

Data Quality tab in SageMaker Catalog showing an overall score of 100 with the movies rule set passed

The Data Quality tab shows an overall score of 100 and lists the rule set movies with a Passed result (1/1). This confirms that the data quality checks from AWS Glue posted successfully to Amazon SageMaker Catalog.

Clean up

To avoid ongoing charges, remove the resources you created during this walkthrough:

  1. Delete the AWS Glue ETL job — Open the AWS Glue console, choose ETL jobs, select your job, and then choose Delete.
  2. Remove the AWS Glue connection — In the AWS Glue console, go to Connections, select the Snowflake connection, and then choose Delete.
  3. Delete the data source in SageMaker Catalog — In your Amazon SageMaker Unified Studio project, go to Data Sources, select the data source you created, and then choose Delete.
  4. Remove S3 assets — Delete the custom transform files from your s3://aws-glue-assets-<account-id>-<region>/transforms/ bucket.
  5. Remove IAM policies — Detach and delete the IAM policies you attached to the AWS Glue job execution role. Remove the role as a domain user and project member.

Conclusion

In this post, we showed you how to connect Snowflake to Amazon SageMaker Unified Studio for centralized data cataloging and quality validation. This approach maintains consistent governance without replicating data. Key benefits include:

  • Query without data movement: Access Snowflake data directly from Amazon SageMaker Unified Studio through federated queries, using the interoperable data architecture of AWS and eliminating time-consuming data replication.
  • Centralized governance: Maintain a single source of truth for data discovery, quality metrics, and governance policies across your distributed data estate.
  • Automated quality validation: Apply consistent data quality rules using AWS Glue Data Quality and visualize results directly in Amazon SageMaker Catalog.
  • Unified collaboration: Support data discovery and sharing across your organization through the publishing capabilities of Amazon SageMaker Catalog.

To get started, open the Amazon SageMaker Unified Studio console. To learn more about related topics, see Cross-account lakehouse governance with Amazon S3 Tables and SageMaker Catalog and Get started with AWS Glue Data Quality dynamic rules for ETL pipelines.


About the authors

Marco Duarte López

Marco Duarte López

Marco is a Data Specialist Solutions Architect at AWS, based in Santiago, Chile. He works with organizations across the region to design modern data architectures and governance frameworks that enable trusted, scalable data consumption. He is a member of the AWS Technical Field Community (TFC) for Analytics, where he specializes in Data & AI Governance, and has led data transformation programs for some of the largest enterprises in the region.

Diego Ortiz

Diego Ortiz

Diego is a Senior Data Strategy Solutions Architect for Latin America based in San Juan, Puerto Rico, with 14+ years of experience in technology roles. He supports organizations across countries and industries to develop data and AI strategies aligned with their business objectives, combining strategic vision with deep technical expertise in data and AI technologies. He is a core member of the Data Governance global community at AWS and leads the analytics technical community in the Spanish-speaking countries of Latin America.

AWS STS simplifies session token size limits and adds session token size monitoring

Post Syndicated from Rishi Tripathy original https://aws.amazon.com/blogs/security/aws-sts-simplifies-session-token-size-limits-and-adds-session-token-size-monitoring/

AWS Security Token Service (AWS STS) has simplified session token size limits, giving you more room for your session policies and session tags. STS has replaced the packed policy size and the overall session token size limits with a single token size limit of 4,096 bytes. STS now reports session token size in API responses, Amazon CloudWatch metrics, and AWS CloudTrail events. By using STS, you can also generate session tokens of different sizes, so you can find the maximum token size that your infrastructure can support.

The 4,096-byte limit is the current maximum, not a permanent ceiling. AWS might increase the limit as new capabilities are added that require session tokens to carry more information.

In this post, you learn what has changed, what this change means for you, and what to do next.

What has changed

AWS STS session-vending APIs, such as AssumeRole, AssumeRoleWithSAML, AssumeRoleWithWebIdentity, GetSessionToken, and GetFederationToken, return temporary security credentials: an access key ID, a secret access key, and a session token. This change governs the session token, the opaque string that STS creates from the session policies and tags you pass plus the context that AWS adds.

Three things have changed.

  • A single limit: Previously, STS enforced two size limits on the session token. It serialized and compressed your session policies and tags into a form called the packed policy, which had its own limit. The assembled token, which included the packed policy, had a separate overall limit. A request could fail against either limit, and both failures returned the same PackedPolicyTooLargeException, so you couldn’t tell which one you exceeded. STS now enforces a single limit: the assembled session token must fit within 4,096 bytes. The separate packed policy limit, which made failures hard to predict, has been removed. When a token exceeds the assembled session token limit, STS returns PackedPolicyTooLargeException. STS continues to use the same exception, so existing error-handling code works without an SDK update.
  • Session token size is now reported. Every successful response from an STS session-vending API includes SessionTokenSize (which reports the session token size in bytes) and SessionTokenUtilization (which reports the percentage of the 4,096-byte limit consumed). STS also returns PackedPolicySize in every successful response for backward compatibility. PackedPolicySize now reports the same value as SessionTokenUtilization, enabling applications that use older AWS SDK versions to monitor utilization through this field. These response fields are also recorded in CloudTrail events. In CloudWatch, SessionTokenSize and SessionTokenMaxSize (the enforced limit) are published in the AWS/STS namespace.
  • Testing is more straightforward: MinimumSessionTokenSize is a new optional parameter on the STS session-vending APIs. You can use it to increase a session token to at least the size you specify, up to 4,096 bytes. Use the parameter to find the maximum token size your infrastructure can handle.
Behavior Previously Now
Limits enforced Two: Packed policy size and assembled token size One: Assembled session token size (4096 bytes)
Error on failure PackedPolicyTooLargeException: The error message didn’t identify which of the two limits was exceeded. PackedPolicyTooLargeException: The updated message reports your session token size and the maximum allowed size, both in bytes.
Session token size visibility Not reported

API Response and AWS CloudTrail:

SessionTokenSize, SessionTokenUtilization, and PackedPolicySize. PackedPolicySize reports the same percentage as SessionTokenUtilization for backward compatibility.

Amazon CloudWatch:

SessionTokenSize and SessionTokenMaxSize

Infrastructure testing No mechanism MinimumSessionTokenSize: Parameter on sesssion-vending APIs

What this change means for you?

How this affects you depends on your situation. The following scenarios cover the most common cases.

  • If you have never hit a token size error: You’re unlikely to notice a change. Your tokens stay their current size and gain headroom. Over time they could become larger than your systems have handled before. We recommend you use MinimumSessionTokenSize to find the maximum token size your systems can handle. See the What to do next section for more details.
  • If you’ve hit PackedPolicyTooLargeException before: Some requests that previously failed now succeed under the single limit. Review any workarounds you put in place specifically to avoid token size errors and decide whether you still need them. General best practices still apply: consistent tag casing and reused tag values compress more efficiently, and concise session policies keep the assembled token smaller. No code change is required for error handling. AWS STS still returns PackedPolicyTooLargeException when the assembled session token exceeds the limit, the same exception STS returned before this change.
  • If your systems enforce their own size limits on credentials: If your application uses an AWS SDK to obtain temporary credentials and make AWS API calls, the SDK handles the session token internally, so token size doesn’t affect your code. Focus instead on systems that store or forward session tokens, such as load balancers, proxies, caches, and databases. These systems might have size limits that smaller tokens didn’t reach. For example, a database column defined as varchar(2048) can’t hold a 4,096-byte token. Review where you persist or pass session tokens, and identify the maximum token size each system supports. The next section shows how to test this.

What to do next

We recommend the following three steps to prepare your systems for this change.

  1. Validate the maximum token size your systems can handle. Use MinimumSessionTokenSize to find the maximum session token size each system in your infrastructure can handle. Knowing these limits helps you identify systems that might reject or truncate larger tokens. The 4,096-byte limit reflects today’s needs, not a permanent ceiling. It might grow as AWS introduces new capabilities such as additional context keys for new services, richer audit metadata, and larger cryptographic signatures as the industry transitions to post-quantum algorithms. Avoid hard-coding the current maximum into your systems and revisit any fixed size assumptions if the limit changes.

    Tip: AWS STS serializes and compresses your session policies and tags when assembling the token. Compression results vary based on the actual content, not just its length. Two sets of tags with identical character counts can produce different token sizes. This is why MinimumSessionTokenSize is a more reliable way to test your infrastructure than estimating from input length.

    aws sts assume-role \
      --role-arn arn:aws:iam::123456789012:role/MyRole \
      --role-session-name validation-test \
      --minimum-session-token-size 4096

    Start at 4,096 bytes to test against the largest possible token. If a system truncates or rejects it, lower the value to find the size your infrastructure supports, then raise that limit where you can. MinimumSessionTokenSize is available in the latest AWS SDK, AWS Command Line Interface (AWS CLI), and Tools for PowerShell versions. See the STS API Reference for details. If your AWS SDK or AWS CLI predates the parameter, update it to use this feature.

  2. Monitor your session token size (recommended). If your infrastructure has size constraints, you can use monitoring to see tokens that are approaching your limit and act before a request fails. AWS STS reports size through three channels, each suited to a different need.
    • In the API response: Reading SessionTokenUtilization and SessionTokenSize from the response requires the latest AWS SDK version. You can also monitor token size through CloudWatch and CloudTrail without updating your SDK.
    {
      "Credentials": {
        "AccessKeyId": "REDACTED",
        "SecretAccessKey": "REDACTED",
        "SessionToken": "REDACTED",
        "Expiration": "2026-06-30T12:00:00Z"
      },
      "AssumedRoleUser": { "...": "..." },
      "PackedPolicySize": 61,
      "SessionTokenSize": 2532,
      "SessionTokenUtilization": 61
    }

    • In CloudWatch: STS publishes SessionTokenSize and SessionTokenMaxSize in the AWS/STS namespace. Use them to build dashboards and set alarms. Set your alarm against the size limit you found during testing, not the 4,096-byte maximum. The maximum is the same for every account, so your own infrastructure limit is the one that matters.

    The following figure shows the SessionTokenMaxSize and SessionTokenSize metrics graphed in the CloudWatch console.

    Figure 1: SessionTokenMaxSize and SessionTokenSizemetrics in the CloudWatch console

    Figure 1: SessionTokenMaxSize and SessionTokenSizemetrics in the CloudWatch console

    • In CloudTrail: Each STS session-vending event records SessionTokenUtilization and SessionTokenSize for successful calls.
    {
      "eventName": "AssumeRole",
      "responseElements": {
        "credentials": { "...": "..." },
        "assumedRoleUser": { "...": "..." },
        "packedPolicySize": 61,
        "sessionTokenUtilization": 61,
        "sessionTokenSize": 2532
      }
    }

  3. Use appropriate fields for monitoring session token utilization. AWS STS still returns PackedPolicySize in session-vending API responses and CloudTrail records for backward compatibility. The field now reports the same value as SessionTokenUtilization: the percentage of the 4,096-byte session token size limit consumed by the token. As a result, PackedPolicySize values might appear lower even when your token content has not changed.

    If your SDK exposes SessionTokenUtilization, use that field because its name reflects the value’s current meaning. If an earlier SDK does not expose SessionTokenUtilization, use PackedPolicySize to monitor the same utilization percentage without updating the SDK. We recommend you monitor SessionTokenSize for the token size in bytes.

Conclusion

You now have more room for session tags, tag values, and session policies in your AWS sessions. AWS STS enforces a single 4,096-byte session token limit, returns a clearer error message when a token exceeds it, and reports token size so you can track growth proactively. Validate your token-handling systems with MinimumSessionTokenSize, and watch SessionTokenUtilization and SessionTokenSize for ongoing visibility.

References

If you have feedback about this post, submit comments in the Comments section below.


Rishi Tripathy

Rishi Tripathy

Rishi is a Principal Product Manager on the AWS Identity and Access Management (IAM) team. He focuses on access control mechanisms that help enterprises secure their AWS environments at scale. He is passionate about building security primitives that are straightforward to adopt and hard to misconfigure.

Tanmay Baid

Tanmay Baid

Tanmay is a Senior Software Development Engineer on the AWS Identity and Access Management (IAM) team. He works on the core identity systems behind the credentials and tokens customers rely on to access AWS at massive scale. He enjoys working on the hard problems at the intersection of distributed systems, identity, and security.

Architecting resilient authentication with Amazon Cognito multi-Region replication

Post Syndicated from Abrom Douglas original https://aws.amazon.com/blogs/security/architecting-resilient-authentication-with-amazon-cognito-multi-region-replication/

Your consumer identity and access management (CIAM) system is the foundation of your customer experience. It’s how users sign in, access services, and engage with your applications. As your business scales across geographies, ensuring authentication is always available becomes a core architectural requirement. However, building multi-Region authentication has traditionally required complex custom replication solutions that synchronize user data, manage consistency, and handle failover, all adding significant operational overhead. Amazon Cognito simplifies this with multi-Region replication (MRR), which automatically replicates user pools across AWS Regions with near-real-time synchronization, built-in failover, and seamless sign-in, while keeping operational complexity and costs optimized.

In this post, we show you how to prepare your user pool for MRR, provide architectural decisions and reference architectures for business to consumer (B2C), business to business (B2B), and machine to machine (M2M) use cases, and practical guidance on implementing failover strategies.

Amazon Cognito MRR at a glance

Amazon Cognito MRR creates a replica user pool in another AWS Region (a replica Region) that shares the same user pool ID as your primary user pool. The primary user pool (the user pool in your primary Region) remains authoritative, and its configurations (app client IDs, client secrets), user data (attributes, hashed credentials, group memberships), and external identity provider (IdP) settings are replicated to the replica with eventual consistency.

The user pool in the replica Region (replica user pool) supports user authentication operations (such as sign-in, token generation and revocation) and read-only operations towards user pool configurations and user attributes (such as list users and groups and describe user pool configurations). Write operations against user pool configurations and updating user attributes aren’t enabled in the replica user pool and can only be made in the primary user pool. Amazon Cognito returns an Action temporarily unavailable error when using managed login, or an OperationNotEnabledException when using an AWS SDK for those operations. See Supported API operations in secondary Regions for a list of API operations supported in replica Regions.

JSON web tokens (JWTs) and active sessions are interoperable between Regions; for example, a refresh token issued by the primary Region is accepted in the replica Region to retrieve new ID and access tokens.

While this post primarily focuses on MRR architecture patterns and considerations, you can visit the following posts to learn more about MRR basics and the next-generation infrastructure behind it:

Prepare for multi-Region replication

In this section, we show you architectural decisions and preparation work for a successful MRR deployment.

Apply a multi-Region customer managed key

Without MRR enabled, data is encrypted at rest with an AWS owned AWS Key Management Service (AWS KMS) key and encrypted in transit with TLS 1.2 and TLS 1.3 with hybrid post-quantum key exchange. Before enabling MRR, you must configure your user pool to use a customer managed key. This must be a symmetric multi-Region AWS KMS customer managed key.

Architectural considerations for your KMS key:

  • You only need to set up one replica multi-Region key for your customer managed key because Amazon Cognito MRR supports only one additional replica Region.
  • You own the administration of the customer managed key, including key policies, rotation, and deletion. You can also consider a key rotation strategy before enabling MRR or enable automatic key rotation.
  • Follow least-privilege principles in KMS key policy and scope the KMS key to your user pool only. You can do so by applying a condition statement: kms:EncryptionContext:aws:cognito-idp:<userpool-arn>. See the data encryption section in the Amazon Cognito developer guide for a full example key policy.

Choose a multi-Region OIDC issuer

In each ID and access tokens, Amazon Cognito includes a default Issuer claim in the JWT payload, referred as iss, to represent the identity provider that issued the token. The OpenID Connect (OIDC) specification dictates that the iss format must be a URL that uses https scheme and publishes a JSON metadata document about the identity provider available at the <iss>/.well-known/openid-configuration path. The metadata document must also include the JSON Web Key (JWK) document in the <iss>/.well-known/jwks.json path, which contains the signing keys to validate the token signatures for its integrity.

The original issuer type follows the format as https://cognito-idp.<region>.amazonaws.com/<userpool_id>. However, this issuer URL format and the OIDC well-known metadata are regional resources. As part of the MRR capability, Cognito introduces a new multi-Region OIDC issuer type, the updated issuer, and follows the format as https://issuer-cognito-idp.<region>.amazonaws.com/<userpool_id>. This new updated issuer type replaces the original single Region type and maintains availability of the issuer endpoint regardless of the state of primary or replica Region.

Based on the issuer URL format you select, your OpenID Connect discovery endpoint is hosted at  <iss>/.well-known/openid-configuration and your JSON Web Key Set (JWKS) endpoint at <iss>/.well-known/jwks.json. Both original type and updated type are supported with the Amazon Cognito MRR capability. You can change the issuer type at any stage in your MRR journey, and the newly issued tokens, including those generated by refresh tokens, will reflect the most current issuer type configurations.

We recommend adopting the updated issuer type. With the updated issuer type, the OpenID Connect discovery document and JWKS endpoint remain consistent and available regardless of which Region is servicing requests. This means your applications can always fetch signing keys for token verification, even during a regional impairment.

To adopt the updated issuer type, update your applications and downstream dependencies to validate against the new updated iss value. If you use the aws-jwt-verify library, update to v5.2.1 or later that supports updated issuer type. Plan this as a coordinated deployment; existing ID and access tokens with the original issuer type remain valid and accepted by Amazon Cognito endpoints until they expire. When using an existing refresh token to exchange for a new set of ID and access tokens, new tokens always carry the current issuer format configuration at the time of token refresh operation, providing interoperability across two issuer formats.

If you can’t immediately adopt the multi-Region issuer—for example, because downstream services or third-party integrations validate the iss claim against a hard-coded original format pattern—you can enable MRR while continuing to use the original issuer type. However, in this configuration the OIDC discovery endpoint and JWKS endpoint are tied to a single Region and might be unavailable during a regional impairment. Your multi-Region application might not be able to fetch public keys dynamically and validate token signatures. To mitigate this, it’s a good practice to implement a JWKS caching strategy in your token verification layer. Cache the signing keys locally (respecting the Cache-Control headers) so your applications can continue to validate tokens using cached keys when the JWKS endpoint is unreachable. This approach lets you benefit from MRR for user authentication while maintaining token verification continuity until you’re ready to complete the issuer migration. To learn more about the original and updated issuer types, see the Amazon Cognito user pools as an OIDC issuer section of the developer guide.

Configure regional service dependencies

Amazon Cognito user pools support several integrations with AWS services for extended customization functionalities. Those AWS services are regional services and must be configured independently in the replica Region, including:

  • AWS Lambda – Lambda triggers (for example, pre-authentication, pre-token generation, and others) are invoked in different authentication stages and should be deployed in the replica Region and attached to the replica user pool to match customized behaviors in the primary user pool. When deploying Lambda triggers, you can adopt the same logic for both primary and replica user pools and access to downstream resources or set up a different logic to characterize different behaviors when requests are served in the replica Region.
  • AWS WAF – WAF web access control lists (web ACLs) are associated to protect the user pool from unwanted requests. When accepting traffic to the replica user pool, create matching WAF web ACLs in the replica Region.
  • Amazon Simple Notification Service (Amazon SNS) – If you send text messages (for example, SMS-based multi-factor authentication (MFA), passwordless authentication, or SMS notifications), configure Amazon SNS in the replica Region. SNS requires additional set up (origination identities, spending limits) in each Region, and sender ID registration time depends on several factors.
  • Amazon Simple Email Service (Amazon SES) – If you use Amazon SES for email delivery, verify sending domains and email addresses in the replica Region and configure your replica user pool accordingly.
  • Amazon CloudWatch – If you export user activity logs from Amazon Cognito to a CloudWatch log group, or monitor service quotas in CloudWatch, configure alarms and analytics accordingly.

Use infrastructure-as-code tools like AWS CloudFormation or AWS Cloud Development Kit (AWS CDK) to maintain consistent configurations and deployments across Regions and environments. You should also monitor for any configuration drifts between assets.

Consider automatic domain failover

For authentication use cases that rely on managed login and OAuth 2.0 endpoints—including federated authentication and M2M authorization—Amazon Cognito supports automatic failover to the replica Region with an Amazon Route 53 health check. Cognito uses the health status of Route 53 health check to control whether traffic routes to the primary or replica user pool. The health check can be set up to monitor the health of an endpoint, a CloudWatch alarm, or a calculated number of other health checks, so you determine what triggers a healthy or unhealthy state and can adjust traffic routing as needed.

Both the Amazon Cognito prefix domain (for example, auth.us-east-1.amazoncognito.com) and custom domain (for example, auth.example.com) support automatic domain failover. Your domain serves as the single entry point for the user pool OAuth 2.0 endpoints and directs traffic to the managed login pages. Cognito automatically fails over domain traffic to the replica Region when a Route 53 health check becomes unhealthy and fails back to the primary Region when the check is healthy. You don’t need to create another prefix domain in the replica user pool for failover use cases.

With the automatic failover capability, you can use a single domain to serve external IdP configurations, including redirect URIs and SAML assertion consumer URLs. For example, use https://auth.example.com/saml2/logout to send SAML 2.0 sign-out responses. Because the domain can serve traffic to both the primary and replica Regions and remains unchanged across Regions, your external IdP configurations stay consistent across Regions, and existing federated users continue to authenticate without disruption. This means that you can enable MRR without having to contact external IdP admins to update configurations; all existing configurations will continue to work.

For SDK-based authentication use cases without managed login, a custom domain isn’t strictly required. We recommend configuring a custom endpoint for SDK requests to simplify failover orchestration, so you don’t have to modify the Region parameter in the SDK configuration. Behind your custom endpoint, you can use the same Route 53 health check or a custom load balancing strategy to proxy API requests to primary or replica Region endpoints. You might also consider load balancing user authentication traffic, by referring to an X-Amz-Target HTTP header (for example, X-Amz-Target: AWSCognitoIdentityProviderService.InitiateAuth), to both the primary and replica Regions, while keeping user sign-up operations in the primary Region. If you use both managed login and SDK authentication in the same user pool, you can consider using the custom domain as the custom endpoint of the SDK for a streamlined operation, where Route 53 health check initiates failover and failback between the primary and replica Regions.

Plan for TOTP MFA alternatives

Time-Based One-Time Password (TOTP) MFA isn’t supported in replica user pools. Users configured to use TOTP MFA must authenticate through the primary Region. If your application relies on TOTP as a second factor, this limitation requires careful planning because you want to enable an alternative MFA for your users, such as SMS OTP, email OTP, or passkey.

Review quotas

When you activate a replica user pool, you gain a separate set of default quotas in the replica Region. Previously reserved higher quotas for your user pool in the primary Region aren’t carried over to the replica Region.

Data sovereignty

When selecting a replica Region for your user pool, consider your organization’s data sovereignty and residency requirements, as user identity data will be replicated to and stored in that Region. For guidance on navigating compliance, continuity, and control obligations that may influence your Region selection, see Practical digital sovereignty: Navigating the pillars of compliance, continuity, and control.

Reference architectures

In this section, we show you reference architectures for common authentication patterns using the Amazon Cognito MRR capability. Each architecture demonstrates how Cognito MRR works with different authentication use cases.

Managed login and federation

Amazon Cognito managed login provides a fully managed authentication UI that handles sign-in, sign-up, and federation flows. With MRR, managed login endpoints are served from the healthy user pool based on your Route 53 health check configuration. Managed login also includes OAuth 2.0 endpoints and can be used with local Cognito accounts and federated users. Figure 1 depicts a reference architecture for using managed login to authenticate Cognito users.

Figure 1: Cognito MRR reference architecture for managed login and federation use cases

Figure 1: Cognito MRR reference architecture for managed login and federation use cases

When using Amazon Cognito with managed login, the process flow is:

  1. The user visits the application and is redirected to the managed login to begin the authentication flow.
  2. Managed login uses the Route 53 health check to control traffic routing.
  3. If the health check returns a healthy status, all traffic to the managed login flows to the primary Region user pool for user authentication.
  4. For a federated user, the primary Region user pool redirects the user to a federated IdP or social IdP for authentication. After successful authentication, Amazon Cognito creates or updates user attributes depending on whether it’s a new user signing in for first time or an existing user.
  5. If the health check returns an unhealthy status, all traffic to the managed login flows to the replica Region user pool. Cognito users will authenticate against the replica user pool.
  6. The replica Region user pool endpoint redirects federated users to external IdPs. However, any user creation or attribute update against replica user pool will fail until the health check returns healthy and traffic routes back to the primary Region.

M2M architecture

In an M2M architecture, services authenticate using the OAuth 2.0 client credentials grant. This flow doesn’t involve users; instead, backend services exchange client credentials for access tokens.

Figure 2: Cognito MRR reference architecture for machine-to-machine use case

Figure 2: Cognito MRR reference architecture for machine-to-machine use case

The authentication flow is:

  1. Application clients send a POST request to the Amazon Cognito /token endpoint with client credentials.
  2. Managed login uses the Route 53 health check to determine whether traffic should flow to the primary or replica user pool.
  3. If the health check returns a healthy status, traffic to the /token endpoint will flow to the primary Region user pool.
  4. If the health check returns an unhealthy status, traffic to the /token endpoint will flow to the replica Region user pool. After the health check returns to a healthy status, traffic will return to routing to the primary user pool.

SDK-based architecture

For applications that use AWS SDK or Amazon Cognito APIs directly (rather than through managed login), the authentication flow is embedded in your application code. This gives you more control over the user experience but requires additional considerations for failover.

Figure 3: Cognito MRR reference architecture for SDK use cases

Figure 3: Cognito MRR reference architecture for SDK use cases

The process shown in Figure 3 is:

  1. The user visits the application and signs in through a custom UI (using APIs or SDKs).
  2. (Optional) An Amazon Route 53 health check is configured to perform a health check against regional proxy endpoints and determine traffic routing. You can also use a custom health check or your DNS resolver to make traffic routing determinations.
  3. If the health check returns a healthy status, all traffic to the proxy endpoints will flow to the primary Region proxy for user authentication. You can also choose to load balance user authentication traffic across both the primary and backup Regions.
  4. The primary Region Amazon API Gateway proxy forwards user requests to the Amazon Cognito regional endpoint.
  5. If the health check returns an unhealthy status, all traffic will flow to the replica Region proxy.
  6. The replica Region API Gateway proxy begins forwarding user requests to the Amazon Cognito regional endpoint until the health check returns to healthy status.

In an SDK-based architecture, Amazon Cognito regional endpoints can also be called directly. You can also set up custom routing to use replica Region endpoints to load balance user authentication requests by routing read-only requests to both the primary and replica Region endpoints while keeping write requests in the primary Region.

Failover strategies

Now that you’ve set up multi-Region replication with Amazon Cognito, the next step is to test and monitor your multi-Region configuration. In this section, we walk through strategies for monitoring your endpoints, determining when to trigger failover, and testing your failover readiness.

Monitor with Route 53 health checks

Failover for Managed Login and all OAuth 2.0 flows is driven by Amazon Route 53 health checks associated with your Amazon Cognito prefix or custom domain. You’re responsible for what determines the state of this health check. The health check isn’t tied to your DNS CNAME record but is the signal that tells Amazon Cognito whether to route traffic to the primary or replica Region for all managed login endpoints. When the health check fails, Amazon Cognito routes traffic to the replica user pool. When the health check recovers, traffic is restored to the primary user pool.

A practical approach to get started to build a health check:

  1. Create a synthetic canary – Use Amazon CloudWatch Synthetics to run a canary that periodically exercises an actual authentication flow against your primary Region. For example, the canary can perform a client credentials token request against your Amazon Cognito domain’s /oauth2/token endpoint or execute a full AdminInitiateAuth API call with test credentials. This validates that the end-to-end authentication path is functional, not just that an endpoint is responding.
  2. Tie the canary to a CloudWatch alarm – Configure a CloudWatch alarm on the canary’s SuccessPercent CloudWatch metric. Set a threshold that accounts for transient errors (for example, alarm when success drops below 90% for three consecutive evaluation periods).
  3. Connect the alarm to your Route 53 health check (optional) – Create a Route 53 health check that monitors the CloudWatch alarm. When the alarm enters the ALARM state, the health check fails, and Amazon Cognito routes traffic to the replica user pool. If you prefer to rely on human intervention, skip this step and instead configure the CloudWatch alarm alert your operations team to manually invert the health check.

After you have your health check, associate it with your Amazon Cognito domain using the UpdateUserPoolDomain API or the Amazon Cognito console.

Authentication-only compared to full-stack failover

Before implementing failover, consider how your authentication layer relates to the rest of your application stack. There are two common patterns:

  • Authentication-only failover – Your application remains in a single Region, but authentication traffic fails over to the Amazon Cognito replica if only the primary Region’s authentication service is impaired. This works when your application can continue operating with tokens already issued (for example, cached JWTs, active sessions) and when downstream APIs don’t depend on the same Region as your user pool. Consider this option when the rest of your stack has its own availability model.
  • Full-stack failover – Your entire application—compute, data stores, APIs, and authentication—fails over to a replica Region. In this model, Amazon Cognito MRR is one component of a broader multi-Region architecture where authentication flows have tight dependencies on regional resources (such as Lambda triggers calling regional Amazon DynamoDB tables, or post-authentication logic writing to a regional event bus) that must be co-located with the user pool.

Use Amazon Application Recovery Controller (ARC) to coordinate failover across all components with a single action. ARC provides three capabilities that are particularly relevant for multi-Region authentication architectures:

  • Routing controls – Extremely reliable data plane controls that let you shift DNS traffic across Regions, with safety rules that prevent partial or unintended failovers (for example, preventing you from failing over authentication without also failing over the dependent API layer).
  • Readiness checks – Continuous monitoring of resource quotas, capacity, and network routing policies in your secondary Region, so you have confidence that the replica environment—including your Amazon Cognito replica user pool and its regional dependencies—can handle production traffic before you failover.
  • Region switch – Centralized, automated, and observable multi-Region recovery orchestration across multiple AWS accounts and resources, so you can execute a coordinated failover of your Cognito user pool alongside databases, compute, and APIs in a single recovery plan.

ARC is particularly valuable when your Amazon Cognito Lambda triggers, WAF rules, SNS and SES configurations, and downstream services all need to switch Regions in lockstep. Rather than managing failover for each component independently, you can use ARC to define a single recovery group that treats your authentication stack and application stack as one unit. To learn more about the capabilities and use cases of ARC, see Introducing Amazon Route 53 Application Recovery Controller.

The right choice depends on your recovery scope. Map the dependencies in your authentication flow: if your Lambda triggers call regional DynamoDB tables or your post-authentication logic writes to a regional event bus, those tight couplings point to full-stack failover. If your application validates tokens independently and doesn’t make real-time calls back to Amazon Cognito after token issuance, authentication-only failover keeps both your blast radius and operational overhead smaller.

Determine when to failover

Triggering failover too aggressively risks unnecessary disruptions; too conservatively risks a drop in desired availability. Here are the factors to balance:

  • Monitor authentication flow health – Validate that critical flows are functioning, including managed login endpoint availability and token endpoint responses.
  • Use composite health checks – Combine multiple signals. For example, require both the managed login and token endpoints to be healthy.
  • Set appropriate thresholds – Configure failure thresholds (for example, three consecutive failures) to distinguish transient errors from genuine impairments.
  • Consider downstream dependencies – Factor in Lambda triggers, external IdPs, and other regional services.
  • Client side retry logic – For SDK-based single-page application (SPA) architectures, consider implementing client-side retry logic with Region failover. When the primary Region is unavailable, your application should detect the failure and redirect authentication of API calls to the replica Region’s Amazon Cognito endpoint.

Understanding and determining the recovery time objective (RTO) and recovery point objective (RPO) should also be the key factor in determining when and why to failover. See the Establishing RPO and RTO Targets for Cloud Applications blog post to learn more.

Test failover readiness

If using Route 53 health check, start by manually inverting your Route 53 health check during a maintenance window. In the Route 53 console, enable Invert health check status to force the health check into a failed state; this triggers failover to the replica Region without requiring any infrastructure changes. While traffic is routing to the replica Region, validate that your critical authentication flows (sign-in, token refresh, federation) work correctly, then disable the inversion to restore traffic to the primary. This test confirms your end-to-end failover path is functional.

When you’re confident in the basic failover path, graduate to more realistic failure simulations with AWS Fault Injection Service (FIS). Create FIS experiment templates that disrupt your primary Region’s Amazon Cognito dependencies; for example, block network access to a dependent resource or inject latency into downstream API calls. Use FIS stop conditions (guardrails) to automatically halt experiments if unexpected impacts are detected. These experiments validate not just that failover triggers correctly, but that your replica Region handles real authentication load under degraded conditions.

We recommend conducting failover tests on a predefined and regular cadence and after any significant changes to your authentication architecture. Document your runbooks and make sure your operations team is familiar with both the failover and recovery procedures.

Conclusion

In this post, we built on the foundational knowledge of the Amazon Cognito MRR capability and showed you how to architect resilient authentication for real-world use cases:

  • Preparation considerations – Multi-Region KMS keys, OIDC issuer transitions, regional dependencies, and TOTP MFA considerations
  • Reference architectures – B2C, B2B, and M2M patterns using managed login, plus SDK-based approaches
  • Failover strategies – Route 53 health checks, ARC integration, and testing with health check inversion and AWS FIS

To get started, make sure your user pool is on the Essentials or Plus feature plan, configure your multi-Region KMS key and OIDC issuer, and create your first replica. For step-by-step setup instructions, see Multi-Region replication for user pools

If you have feedback or thoughts about this post, submit comments below. If you have questions, start a new thread on Amazon Cognito re:Post or contact AWS Support.


Abrom-Douglas-author

Abrom Douglas III

Abrom is a Senior Solutions Architect within AWS Identity with over 20 years of software engineering and security experience, specializing in identity and access management. He loves speaking with customers about how identity and access management can provide secure outcomes that enable both business and technology initiatives. In his free time, he enjoys cheering for Arsenal FC, photography, travel, volunteering, and competing in duathlons.

Edward Sun

Edward Sun

Edward is a Senior Security Specialist Solutions Architect focused on identity and access management. He loves helping customers throughout their cloud transformation journey with architecture design, security best practices, migration, and cost optimizations. Outside of work, Edward enjoys hiking, golfing, and cheering for his alma mater, the Georgia Bulldogs.

Operationalizing least privilege: Automate IAM remediation through your CI/CD pipeline

Post Syndicated from Luis Pastor original https://aws.amazon.com/blogs/security/operationalizing-least-privilege-automate-iam-remediation-through-your-ci-cd-pipeline/

The principle of least privilege is straightforward to articulate but challenging to maintain at scale. When teams first deploy applications to AWS, they often grant broader permissions than strictly necessary; it’s faster to get things working, and the plan is always to tighten permissions later. But later rarely comes. Permissions accumulate, AWS Identity and Access Management (IAM) principals that once needed broad access for initial deployment retain those permissions long after they’re necessary, and some principals stop being used entirely. Even small teams face this challenge—permission reviews aren’t a one-time task but an ongoing operational burden that demands automation.

AWS IAM Access Analyzer addresses detection and recommendation. It identifies unused permissions across IAM roles and users: actions that haven’t been exercised, services that haven’t been accessed, and principals that aren’t being assumed at all. For each finding, it generates a recommended policy with the excess permissions removed. Security teams can see exactly what to fix, but manual remediation doesn’t persist. A security engineer can right-size a role today, but if that role is defined in an AWS CloudFormation template or AWS Cloud Development Kit (AWS CDK) stack, the next deployment restores the original permissions. The fix must live where the role is defined, and not every role starts in the same place. Some are managed through infrastructure-as-code (IaC), where remediation means updating source code and deploying through a pipeline. Others were created manually through the AWS Management Console and have no code representation. And some principals aren’t being used at all and need a controlled decommission path. Each scenario requires a different remediation strategy.

This post walks through an automated remediation workflow that bridges the gap between detection and action. Instead of findings accumulating in a dashboard waiting for someone to investigate, the automation classifies each role by how it was created and produces a ready-to-review remediation artifact: a pull request with production-ready CDK code and a plain-English explanation for IaC-managed roles, an issue with the recommended policy and step-by-step IaC migration guidance for manually created roles, or a soft-disable issue with a monitored decommission plan for unused principals. Each output flows through your existing code review and issue tracking processes—the same workflows your teams already follow. By the end of this post, you’ll have a pattern that converts IAM Access Analyzer findings into tested, deployable code changes rather than a growing backlog of security tickets.

Understanding the problem

Unused IAM permissions increase the attack surface. Removing unused permissions limits the actions available to any compromised credentials, reducing potential impact. Roles that aren’t being assumed represent unused resources; removing them simplifies your IAM inventory and reduces potential access paths that aren’t actively monitored.

The challenge isn’t knowing what to fix. As we said earlier, Access Analyzer provides both the findings and the recommended policies. The challenge is acting on that knowledge consistently across your environment. Each finding requires context:

  • What the role does
  • Who created the role
  • Determining if the permission is unused or used infrequently
  • If the role is managed in a CloudFormation stack, or was created through the console

Multiply this by hundreds of roles and security teams face a backlog that grows faster than they can address it.

Manual remediation compounds the problem. A security engineer can right-size a role directly in the console, but that fix is fragile. If the role is defined in an IaC template, the next deployment restores the original permissions. If it was created manually, there’s no record of what changed or why, and no easy way to revert if the change causes issues.

This is where IaC changes the equation. When roles are defined in code, remediation means updating that code. Changes flow through pull requests, are reviewed by the team that owns the role, and deploy consistently across environments. The fix becomes permanent, not a point-in-time correction that drifts back on the next deployment. And because every change is tracked in version control, teams can confidently remove permissions knowing they can revert if something breaks. That safety net matters; it’s often the difference between a team acting on a finding and leaving it in the backlog.

Solution overview

The solution automates remediation by connecting four capabilities: IAM Access Analyzer for detection and policy recommendations, CloudTrail for role attribution, Amazon Bedrock for CDK code generation and plain-English explanations, and your existing continuous integration and delivery (CI/CD) pipeline for remediation execution. The workflow operates on a core principle: every IAM role has an origin, and that origin determines the remediation path.

Figure 1 shows the solution architecture: Amazon EventBridge triggers an AWS Lambda orchestrator on a daily schedule. The Lambda orchestrator integrates with IAM Access Analyzer, CloudTrail, Amazon Bedrock, and Amazon CloudWatch. Each finding is routed to one of three remediation paths: a pull request for IaC-managed roles, an issue for manually created roles, and a soft-disable issue for unused roles.

Figure 1: The daily remediation workflow; from scheduled trigger to the three role-based remediation paths

Figure 1: The daily remediation workflow; from scheduled trigger to the three role-based remediation paths

On each scheduled run, the automation retrieves active findings from IAM Access Analyzer and queries CloudTrail to determine how each role was created. Roles created through CloudFormation or AWS CDK have a traceable origin: the service principal, stack name, and originating repository. Roles created manually through the console have a different origin: the IAM user who created them and the timestamp. This distinction drives the remediation strategy.

For IaC-managed roles, the automation retrieves the IAM Access Analyzer-recommended policy and uses Amazon Bedrock to wrap it in production-ready CDK code that includes the role definition and policy statements and imports what your CI/CD pipeline needs to deploy the update. It then creates a pull request in the originating repository. The pull request (PR) includes the updated CDK code, a policy diff showing exactly which permissions are being removed, and a plain-English explanation of the changes, for example, “This change removes write access to S3, keeping only read and list permissions.” Your existing code review process evaluates the change, and after being merged, the fix deploys consistently across environments.

For manually created roles, the automation creates an issue that includes the IAM Access Analyzer-recommended policy with unused permissions removed, a diff highlighting the changes, and an Amazon Bedrock-generated explanation of what the permission changes accomplish. The issue also provides guidance on importing the role into your IaC codebase. This gives teams an immediate remediation path while encouraging long-term governance through IaC adoption.

For roles that aren’t being assumed at all, the automation takes a more cautious approach. Instead of taking direct action, it creates an issue recommending a soft-disable workflow: attach a deny-all policy to the role, monitor for 30 days to confirm no workload depends on it, then delete. The issue provides the steps and context, the team executes the decommission through their preferred process, whether that’s a console change, an AWS Command Line Interface (AWS CLI) script, or a PR removing the role from the IaC. This controlled decommission path reduces the risk of removing a role that’s used infrequently or seasonally.

The solution supports both single-account and organization-wide deployment. In single-account mode, it uses an ACCOUNT_UNUSED_ACCESS analyzer to process findings for one account. In organization mode, it uses an ORGANIZATION_UNUSED_ACCESS analyzer deployed in a delegated administrator account, which generates findings across all member accounts from a single vantage point. The Lambda function automatically detects which analyzer type is available and extracts the account ID from each finding’s resource Amazon Resource Name (ARN), so role attribution and remediation routing work the same way regardless of scope.

This three-path strategy acknowledges operational reality. Not all roles start in IaC, not all unused roles are safe to delete immediately, and forcing immediate migration isn’t always practical. The solution provides a clear path forward for each scenario: remediate IaC roles through code, give teams actionable recommendations for manually created roles, and safely decommission what’s no longer needed. Over time, your infrastructure becomes increasingly code-driven, and remediation becomes a routine part of your CI/CD process rather than a manual security task.

Technical details

Consider a company—call them AnyCompany—running 200 IAM roles across three AWS accounts. Some roles were created through AWS CDK stacks during initial deployment. Others were created manually through the console by engineers who needed quick access during incident response or prototyping. A handful haven’t been assumed in over 6 months. AnyCompany’s security team wants to act on their IAM Access Analyzer findings, but each role requires different handling. The solution’s architecture addresses this by routing each finding through a classification and remediation pipeline.

Figure 2 shows how each IAM Access Analyzer finding is processed:

  1. The finding is first checked against exclusions and excluded findings are skipped.
  2. Remaining findings are split by type: UnusedPermission findings retrieve a recommended policy from IAM Access Analyzer and then query CloudTrail for role origin, while UnusedIAMRole findings follow the unused role path.
  3. By origin, IaC-managed roles generate AWS CDK code using Amazon Bedrock and create a pull request.
  4. Manually created or unknown-origin roles create an issue with the recommended policy and IaC migration guidance.
  5. Unused roles create a soft-disable issue to deny-all, monitor for 30 days, then delete.
  6. All paths publish CloudWatch metrics.
Figure 2: Detailed component interactions—the orchestrator’s five steps, its four service integrations, and the three remediation paths

Figure 2: Detailed component interactions—the orchestrator’s five steps, its four service integrations, and the three remediation paths

The rest of this section walks through each component using AnyCompany’s roles as examples.

Exclusion filtering

Before processing any finding, the Lambda function loads an exclusion configuration and checks whether the role should be skipped. This prevents the automation from creating remediation items for roles that legitimately need broad permissions.

{
  "excluded_roles": [
    "arn:aws:iam::123456789012:role/BreakGlassRole",
    "arn:aws:iam::123456789012:role/ServiceLinkedRole"
  ],
  "excluded_permissions": [
    "iam:*",
    "sts:AssumeRole"
  ],
  "excluded_by_tag": {
    "NoRemediation": ["true"],
    "CriticalService": ["true"]
  },
  "min_unused_days": 30
}

AnyCompany excludes their break-glass role (used only during incidents), any service-linked roles, and roles tagged CriticalService. The min_unused_days threshold prevents false positives from seasonal workloads; a role that ran a quarterly batch job 25 days ago won’t generate a finding.

Detection and analysis

IAM Access Analyzer generates two types of findings relevant to this solution. UnusedPermission findings identify roles with permissions that haven’t been exercised within the analysis period. UnusedIAMRole findings identify roles that haven’t been assumed at all. The Lambda function queries both finding types separately because they follow different remediation paths.

The Lambda function auto-detects the analyzer type at startup. When ANALYZER_SCOPE is set to organization, it checks for an ORGANIZATION_UNUSED_ACCESS analyzer first and falls back to ACCOUNT_UNUSED_ACCESS if none exists. If multiple analyzers of the same type exist in the account, the Lambda function selects the first active analyzer returned by the API. To target a specific analyzer, set the ANALYZER_ARN environment variable explicitly. With an organization-level analyzer, findings include roles from all member accounts. The Lambda function extracts the account ID from each finding’s resource ARN (for example, account 111122223333 from arn:aws:iam::111122223333:role/MyRole) and carries that context through the entire pipeline: attribution, remediation, and issue or PR creation all include the originating account.

For UnusedPermission findings, the Lambda function calls GenerateFindingRecommendation to initiate policy generation, then retrieves the IAM Access Analyzer-recommended policy through the GetFindingRecommendation API. This is a key integration point: IAM Access Analyzer provides the right-sized policy with unused permissions removed, so the automation doesn’t need to generate policies itself.

Here’s what a typical finding looks like for one of AnyCompany’s application roles:

{
  "id": "a1b2c3d4-5678-90ab-cdef-example11111",
  "resource": "arn:aws:iam::123456789012:role/AnyCompanyOrderProcessorRole",
  "findingType": "UnusedPermission",
  "analyzedAt": "2026-03-01T00:00:00Z",
  "unusedPermissions": [
    { "action": "s3:PutObject", "lastAccessed": null },
    { "action": "s3:DeleteObject", "lastAccessed": null },
    { "action": "s3:PutBucketPolicy", "lastAccessed": null },
    { "action": "dynamodb:DeleteItem", "lastAccessed": null }
  ],
  "activePermissions": [
    { "action": "s3:GetObject", "lastAccessed": "2026-02-28T14:30:00Z" },
    { "action": "s3:ListBucket", "lastAccessed": "2026-02-28T14:30:00Z" },
    { "action": "dynamodb:Query", "lastAccessed": "2026-02-28T12:00:00Z" }
  ]
}

The OrderProcessorRole has write and delete permissions for Amazon Simple Storage Service (Amazon S3) and Amazon DynamoDB, but only uses read operations. The IAM Access Analyzer recommendation removes the four unused actions while preserving the three active ones.

For UnusedIAMRole findings, no recommendation is needed: the role isn’t being assumed at all, so the remediation is to disable or delete it. The Lambda function caps the number of unused role issues per run (configurable using MAX_UNUSED_ROLE_ISSUES, default 10) to avoid overwhelming teams with a flood of issues on the first execution.

Role attribution using CloudTrail

For each finding, the Lambda function queries CloudTrail to determine how the role was created. The CreateRole event contains the information needed to classify the role’s origin.

An IaC-created role looks like this in CloudTrail:

{
  "eventName": "CreateRole",
  "userIdentity": {
    "type": "AWSService",
    "invokedBy": "cloudformation.amazonaws.com"
  },
  "requestParameters": {
    "roleName": "AnyCompanyOrderProcessorRole"
  },
  "userAgent": "cloudformation.amazonaws.com"
}

The cloudformation.amazonaws.com service principal and user agent tell the automation this role was created through a CloudFormation or AWS CDK deployment. The Lambda function then looks up the role’s tags to find the originating repository (stored in a Repository tag set during deployment).

A manually-created role looks different:

{
  "eventName": "CreateRole",
  "userIdentity": {
    "type": "IAMUser",
    "userName": "jstiles"
  },
  "requestParameters": {
    "roleName": "AnyCompanyIncidentResponseRole"
  },
  "userAgent": "console.amazonaws.com"
}

Here, the IAMUser type and console.amazonaws.com user agent indicate someone created this role through the console. Roles created through the AWS CLI show a similar pattern: the IAMUser type with a user agent like aws-cli/2.x.x. The automation classifies both console and AWS CLI-created roles as manually created, because neither has an IaC origin that can be updated programmatically. The automation captures the username and timestamp for the remediation issue.

Cross-account role attribution

When the Lambda function processes findings from an organization-level analyzer, the role might live in a different account than the one running the function. The automation handles this by assuming a cross-account role (configurable using CROSS_ACCOUNT_ROLE_NAME, defaulting to OrganizationAccountAccessRole) in the member account, then querying that account’s CloudTrail and IAM APIs for the CreateRole event. If the cross-account assume fails—because the role doesn’t exist in that account or permissions aren’t configured—the automation falls back gracefully, classifying the role as unknown origin and creating an issue with the account ID and available context. This approach helps the automation produce an actionable output for findings even when attribution is incomplete.

Policy recommendations and AWS CDK code generation

For IaC-managed roles with UnusedPermission findings, the Lambda function retrieves the IAM Access Analyzer-recommended policy and sends it to Amazon Bedrock to generate production-ready AWS CDK code. This is an important distinction: IAM Access Analyzer decides what the policy should be, and Amazon Bedrock wraps that policy in the AWS CDK constructs, imports, and resource definitions that the CI/CD pipeline needs to deploy the update.

The prompt instructs Amazon Bedrock to convert the recommended policy to AWS CDK code exactly as provided, with no modifications:

Generate Python CDK code that creates/updates the role with the
RECOMMENDED policy exactly as provided. Include proper imports
(aws_cdk, aws_iam), use CDK best practices (PolicyStatement,
proper resource ARNs), and add tags: ManagedBy=CDK,
RemediatedBy=AccessAnalyzer.

IAM Access Analyzer generates recommendations for both inline policies and customer managed policies. When a managed policy has partially unused permissions, the recommendation contains the full right-sized policy. The automation wraps this in AWS CDK code as an iam.ManagedPolicy construct. Note that if a managed policy is shared across multiple roles, the recommendation applies to the specific role’s usage pattern. In this case, the automation generates an issue for manual review rather than a PR, because modifying a shared policy could affect other roles.

The generated code goes through a validation step before inclusion in any PR. The Lambda function compiles the Python code to check for syntax errors and verifies that required AWS CDK patterns (iam, PolicyStatement) are present. If validation fails, the finding is logged as an error rather than creating a broken PR.

The solution doesn’t currently invoke the IAM Access Analyzer ValidatePolicy API to check the generated policy for errors or overly permissive statements. However, this is a natural extension point. Teams can add a validation step that calls ValidatePolicy on the Amazon Bedrock-generated policy before including it in a PR, detecting issues like missing resource constraints or invalid action names.

Amazon Bedrock also generates a plain-English explanation of the policy changes. For AnyCompany’s OrderProcessorRole, the explanation might read:

“The role currently has full S3 write access and DynamoDB delete permissions, but only uses read operations. Removing s3:PutObject, s3:DeleteObject, s3:PutBucketPolicy, and dynamodb:DeleteItem reduces the scope of impact if credentials are compromised, while preserving the s3:GetObject, s3:ListBucket, and dynamodb:Query permissions the application needs.”

The solution uses the Anthropic Claude Sonnet model on Amazon Bedrock for CDK code generation (where accuracy matters) and Claude Haiku on Amazon Bedrock for explanations (where speed and cost efficiency matter more).

Three-path remediation

The Lambda function evaluates each finding’s origin and routes it to one of three remediation paths.

Path 1: IaC-managed roles (pull request) – For AnyCompany’s OrderProcessorRole, the automation creates a PR in the originating repository. The PR includes:

  • The Amazon Bedrock-generated AWS CDK code implementing the IAM Access Analyzer-recommended policy
  • A policy diff showing exactly which permissions are being removed
  • The plain-English explanation of what the changes accomplish
  • Labels (security, iam-remediation, automated) for filtering and tracking

The team that owns the role reviews the PR through their normal code review process. Once merged, the fix deploys consistently across environments through the existing CI/CD pipeline.

Path 2: Manually-created roles (issue) – For AnyCompany’s IncidentResponseRole, the automation creates an issue that includes the Access Analyzer-recommended policy with unused permissions removed, a diff highlighting the changes, an Amazon Bedrock-generated explanation, and step-by-step guidance on importing the role into IaC. This gives the team an immediate remediation path (apply the recommended policy) while encouraging long-term governance through IaC adoption.

Path 3: Unused roles (soft-disable issue) – For roles that haven’t been assumed at all, the automation creates an issue recommending a three-stage decommission workflow: attach a deny-all policy to the role, monitor for 30 days to confirm no workload depends on it, then delete. This controlled approach reduces the risk of removing a role that’s used infrequently or seasonally – if something breaks during the monitoring period, removing the deny-all policy restores access immediately.

Dry-run mode

Before creating real PRs and issues, you can run the automation in dry-run mode by setting “dry_run": true in the CI/CD configuration or setting the CI_CD_PLATFORM environment variable to dryrun. In this mode, the Lambda function processes findings, classifies roles, and generates remediation data, but logs what it would create instead of making actual API calls to your repository platform. You can use the log to validate the automation’s behavior, review the classification accuracy, and tune exclusions before going live.

Operational metrics

The Lambda function publishes CloudWatch metrics after each run:

findings_processed Total UnusedPermission findings evaluated
iac_roles_found Roles classified as IaC-managed
manual_roles_found Roles classified as manually created
unused_roles_found Roles with no assume activity (UnusedIAMRole findings)
prs_created Pull requests created for IaC roles
issues_created Issues created (manual roles and unused roles)
errors Processing errors (failed classifications, API failures)

These metrics feed into dashboards and alarms. AnyCompany sets an alarm on errors > 5 to catch API throttling or configuration issues, and tracks prs_created + issues_created over time to measure remediation velocity.

Implementation

The solution ships as two AWS CDK stacks and deploys in minutes. The accompanying GitHub repository contains the complete source code, AWS CDK stacks, configuration templates, and step-by-step deployment instructions.

At a high level, deployment involves:

  1. Prerequisites: An AWS account with an ACCOUNT_UNUSED_ACCESS or ORGANIZATION_UNUSED_ACCESS analyzer enabled, Python 3.11 or later, AWS CDK v2, a CI/CD platform API token stored in AWS Secrets Manager, and Amazon Bedrock model access for the Anthropic Claude models you plan to use. The model IDs are configurable environment variables (BEDROCK_CODEGEN_MODEL and BEDROCK_EXPLANATION_MODEL); Amazon Bedrock retires older foundation models over time, so if the shipped defaults stop working, set these variables to current models you have enabled and redeploy. The repository README documents this.
  2. Configuration: Two files in the config/ directory control behavior. exclusions.json defines which roles and permissions to skip (break-glass roles, service-linked roles, tagged exceptions), and ci_cd_config.json configures your repository platform integration (GitLab or GitHub), labels, and throttling limits.
  3. Deploy: Run cdk deploy --all to create the Lambda function, EventBridge schedule, IAM roles, and CloudWatch alarms.
  4. Validate in dry-run mode: Start with “dry_run": true to see how the automation classifies your roles without creating real PRs or issues. Review the CloudWatch logs to confirm attribution accuracy and tune exclusions.
  5. Go live: Set “dry_run": false and redeploy. The Lambda function runs on schedule (daily by default) and begins creating PRs and issues.

The repository README covers each step in detail, including organization-wide deployment, cross-account configuration, and platform-specific setup for GitLab and GitHub.

Operational considerations

Deploying the automation is only the starting point. Running it in production means making decisions about how roles are retired, how the volume of findings is managed at scale, which roles warrant human review before any change is proposed, and how you measure the automation’s impact over time. The following practices keep remediation sustainable as your IAM footprint grows, so the automation reduces operational burden rather than adding to it.

Unused role lifecycle

Unused roles follow a three-stage decommission workflow. When the automation identifies a role that hasn’t been assumed within the analysis period, it creates an issue with the recommended decommission steps; the automation doesn’t modify the role directly. The team then follows the soft-disable approach:

  1. Attach a deny-all inline policy to the role. This blocks all actions without deleting the role or its existing policies.
  2. Monitor for 30 days. If a workload depends on the role (seasonal jobs, infrequent batch processes), the deny-all policy surfaces the dependency quickly. Removing the deny-all policy restores full access immediately; no need to recreate the role or reattach policies.
  3. Delete the role after the monitoring period confirms no impact.

This approach is deliberately conservative. Deleting a role is irreversible; you lose the trust policy, attached policies, and any resource-based policies that reference it. The soft-disable step gives teams a safety net while still making progress on reducing their unused role inventory.

Scaling and throttling

On AnyCompany’s first run, the automation found 47 unused permission findings and 4 unused roles. That’s manageable. But organizations with hundreds of accounts and thousands of roles might see significantly more findings on initial deployment.

This is especially true with an organization-level analyzer. A single-account deployment might surface dozens of findings; an organization-level analyzer across multiple accounts could surface hundreds or thousands on the first run. The throttling controls become critical at this scale.

Two throttling controls prevent the automation from overwhelming teams:

  • max_findings_per_run (default 50): Caps the total UnusedPermission findings processed per Lambda function execution. Remaining findings are picked up on the next scheduled run.
  • MAX_UNUSED_ROLE_ISSUES (default 10): Caps unused role issues per run. This is especially important during initial deployment when you might have a large backlog of roles that haven’t been assumed in months.

Start with conservative limits and increase them as your team builds confidence in the review process. A team that can review 10 PRs per week shouldn’t receive 50 on Monday morning.

Approval workflows for sensitive roles

Not every role should receive automated PRs. Roles with administrative permissions or access to sensitive data might warrant manual review before any remediation is created. The exclusion configuration supports this through the approval_required_for_tags field:

{
  "approval_required_for_tags": {
    "Sensitive": ["true"],
    "Admin": ["true"]
  }
}

Roles matching these tags generate issues for manual review instead of automated PRs, regardless of whether they’re IaC-managed. This gives security teams a checkpoint for high-risk roles while still automating remediation for standard application roles.

Monitoring and alerting

The metrics published after each Lambda function run (covered in the Technical details section) feed into CloudWatch dashboards and alarms. A few patterns worth setting up:

  • Alert on errors > 5 per run to catch API throttling, expired CI/CD tokens, or Amazon Bedrock availability issues.
  • Track prs_created + issues_created over time. A healthy trend shows this number decreasing as your environment converges toward least privilege.
  • Monitor unused_roles_found as a leading indicator. A sudden increase might signal a team spinning up roles for a project and not cleaning up afterward.
  • Compare iac_roles_found to manual_roles_found over time. As teams adopt IaC, the ratio should shift toward IaC-managed roles, which means more automated remediation and less manual work.

Cost

The solution uses Lambda (minimal cost at daily execution), CloudTrail (typically already enabled), IAM Access Analyzer (charges per IAM role or user analyzed per month for the unused access analyzer), and Amazon Bedrock (pay-per-token for AWS CDK code generation and explanations). For most organizations the ongoing cost is low, and Amazon Bedrock token usage is the largest variable, scaling with the number of findings processed per day and the complexity of each policy. Review the pricing pages for each service for current rates.

For organization-level deployments, the IAM Access Analyzer cost scales with the number of IAM roles analyzed across all member accounts. The ORGANIZATION_UNUSED_ACCESS analyzer charges per role per month across the organization, so an organization with 500 roles across 20 accounts will see higher analyzer costs than a single account with 50 roles. Review the IAM Access Analyzer pricing page for current rates.

Cleanup

To remove the solution, run cdk destroy --all from the infrastructure/ directory. This removes the Lambda function, EventBridge rule, CloudWatch alarms, and IAM roles created by the stacks.

If you stored a CI/CD platform API token in Secrets Manager as part of deployment, delete it with aws secretsmanager delete-secret --secret-id <your-secret-name> --recovery-window-in-days 7. The 7-day recovery window lets you restore the secret if the deletion was accidental. After 7 days, the secret is permanently deleted and can’t be recovered. To delete immediately without a recovery window, add --force-delete-without-recovery.

Lambda automatically creates a CloudWatch Logs log group at /aws/lambda/<function-name> that persists after cdk destroy --all and continues to incur log storage charges. To remove it, run aws logs delete-log-group --log-group-name /aws/lambda/<function-name>. WARNING: This permanently deletes all execution logs.

The IAM Access Analyzer isn’t created by the AWS CDK stacks. WARNING: Deleting the analyzer permanently removes all findings, analysis history, and unused permission data. Export any findings you need to retain before deletion. After exporting, run aws accessanalyzer delete-analyzer --analyzer-name <your-analyzer-name> to delete it. The ACCOUNT_UNUSED_ACCESS and ORGANIZATION_UNUSED_ACCESS analyzer types incur charges based on the number of IAM roles and users analyzed per month.

If you deployed in organization mode and created cross-account roles (default name: OrganizationAccountAccessRole) in member accounts solely for this solution, remove them from those accounts.

Any PRs or issues already created in your CI/CD platform remain after stack deletion; they’re artifacts in your repository, not AWS resources. See the repository README for detailed cleanup instructions.,

Conclusion

Automating IAM permission remediation turns least privilege from a periodic compliance exercise into an operational practice. By connecting IAM Access Analyzer findings and recommendations to your CI/CD pipeline, remediation shifts from manual security tasks to code review processes that your teams already follow.

The three-path strategy acknowledges how infrastructure evolves. IaC-managed roles receive pull requests with production-ready AWS CDK code and plain-English explanations. Manually created roles receive actionable issues with recommended policies and IaC migration guidance. Unused roles are put on a controlled decommission path that protects against accidental disruption. Over time, the manual role count decreases as teams adopt IaC, and remediation becomes a routine part of your deployment pipeline.

Start with a pilot. Choose 10–20 non-production roles, deploy in dry-run mode, and review the classification results. Tune your exclusions, confirm the CloudTrail attribution is accurate for your environment, and then enable live remediation. Expand to production roles after your team is comfortable with the review cadence.

When you’re ready to scale beyond a single account, switch to an organization-level analyzer and the same Lambda function will process findings across all member accounts with no architectural changes required, only a configuration toggle.

The complete source code, AWS CDK stacks, and configuration templates are available in the accompanying GitHub repository.

If you have feedback about this post, submit comments in the Comments section below.


Luis Pastor

Luis E Pastor

Luis is a Senior Security Solutions Architect at AWS specializing in infrastructure security, compliance, and generative AI security. He leads technical field communities focused on security and compliance while contributing to AWS Well-Architected Framework guidance. Before AWS, he helped clients across financial services, healthcare, and retail industries improve their security posture in hybrid environments. Outside of work, Luis enjoys staying active and culinary adventures.

Rodolfo Brenes

Rodolfo Brenes

Rodolfo is a Principal Solutions Architect focused on Cloud Governance and Compliance. With over 18 years of experience, he currently leads a technical field community in AWS helping customers scale and improve their security and governance frameworks. Besides work, Rodolfo enjoys video games, playing with his four cats, and won’t say no to a good outdoor adventure.

Sowjanya Rajavaram

Sowjanya Rajavaram

Sowjanya is a Sr Solution Architect who specializes in Identity and Security in AWS. Her entire career has been focused on helping customers of all sizes solve their identity and access management problems. She enjoys traveling and experiencing new cultures and food.

Satish Uppalapati

Satish is an Associate Assurance Consultant with AWS Security Assurance Services (SAS) and has more than 8 years of experience in IT risk, governance, and regulatory assurance. He works with AWS customers to align cloud environments with multiple frameworks. Satish helps organizations build security and governance programs that meet regulatory objectives while supporting business operations. He also focuses on advancing governance for AI systems, including emerging standards.

Validating multi-agent decisions with Step Functions and Bedrock AgentCore

Post Syndicated from Ben Freiberg original https://aws.amazon.com/blogs/compute/validating-multi-agent-decisions-with-step-functions-and-bedrock-agentcore/

For an airline operations team, a single flight cancellation sets off a chain reaction. Hundreds of passengers need new itineraries within minutes, and no two cases are alike. They have different loyalty tiers, sit on different fare rules, and have downstream connections that may not wait. Passengers have varying cabin and seat preferences and might fall under different regulatory entitlements depending on where they booked and where they are flying.

Most airlines handle this with a layered system: rule-based automation covers the simple, one-hop rebooks, and everything else flows to a manual queue staffed by service agents. That works when disruptions are isolated. When they are not, the queue overwhelms, waiting times spike, and passengers booked alternatives themselves that create downstream knock-on disruptions.

This is exactly where AI agents become compelling. An agent can reason across seat availability, fare rules, loyalty entitlements, and connection timing the way an experienced desk agent would, but at machine speed and across hundreds of cases in parallel. Multi-agent collaboration typically lets a supervisor agent route work to collaborator sub-agents, with the model itself deciding which sub-agent runs and in what order. But an unconstrained agent might optimize for the passenger’s preference while ignoring a codeshare restriction, rebook onto a flight that meets minimum connection time on paper but not at that specific airport, or calculate compensation under the wrong regulatory regime because it misread the ticket’s point of sale.

Orchestrating specialized Amazon Bedrock AgentCore agents with AWS Step Functions gives you the reasoning power of generative AI with the guardrails of deterministic validation. Step Functions adds native fan-out across thousands of passengers, a callback pattern that pauses a case for human review at zero compute cost, and a durable execution history that serves as your audit trail. The principle is that agents propose, and deterministic code validates. The pattern is demonstrated here for airline rebooking, but it applies anywhere automated decisions can have real financial or regulatory consequences.

Solution overview

The design is a Step Functions state machine where deterministic steps that map to the business processes wrap each agent’s non-deterministic behavior. The following diagram shows the end-to-end flow. At a high level, the workflow proceeds through these stages:

  1. The workflow starts when a flight-cancellation event arrives, for example through an Amazon EventBridge integration.
  2. An enrichment step pulls additional data such as the passenger manifest, current bookings, loyalty status, and stored preferences.
  3. The workflow fans out to run agents in parallel for each affected passenger.
  4. Two agents then run for each passenger: a find-alternatives agent proposes the top three rebooking options, and a compensation agent determines entitlement based on route, delay duration, and cause.
  5. A deterministic validation step runs after each agent, confirming flights are actually bookable and entitlement rules are followed before either result is used.
  6. The workflow checks whether the case can be auto-confirmed, or needs human review.
  7. Bookings are confirmed, compensation issues, and confirmations are sent. Unresolved cases go to human agents.

The key principle: no agent Task state writes to the reservation system or issues a payment. Only deterministic Task states do that, and only after a deterministic validation step has passed.

Integrating AgentCore harness with Step Functions

AgentCore harness is a managed agent loop. You specify a model, system prompt, and tools, and the harness runs the reasoning cycle (model calls, tool execution, memory management, and response generation) end-to-end in a single API call. It handles the intra-agent orchestration so that Step Functions can focus on inter-agent orchestration: fan-out, sequencing, validation gates, and exception routing. Step Functions provides a native optimized integration for AgentCore harness, which calls InvokeHarness against a target HarnessArn. The optimized integration gives you an extended per-Task timeout of 15 minutes (900 seconds), so agents have enough time to reason through complex proposals. The trade-off is that the agent call is request-response only. There is no .sync and no .waitForTaskToken on the agent step, and only the final assistant message is returned to the state machine.

The following Amazon States Language snippet shows the optimized harness invocation inside a Distributed Map. For the full definition, see the sample on Serverless Land.

{
  "Comment": "Illustrative - per-passenger rebooking fan-out",
  "StartAt": "RebookPassengers",
  "States": {
    "RebookPassengers": {
      "Type": "Map",
      "ItemProcessor": {
        "ProcessorConfig": { "Mode": "DISTRIBUTED", "ExecutionType": "STANDARD" },
        "StartAt": "FindAlternatives",
        "States": {
          "FindAlternatives": {
            "Type": "Task",
            "Resource": "arn:aws:states:::bedrockagentcore:invokeHarness",
            "Parameters": {
              "HarnessArn": "<HARNESS_ARN>",
              "RuntimeSessionId.$": "$.passenger.sessionId",
              "Messages": [{ "Role": "user", "Content": [{ "Text.$": "States.JsonToString($.passenger)" }] }]
            },
            "TimeoutSeconds": 900,
            "ResultPath": "$.proposal",
            "Next": "ValidateRebooking"
          },
          "ValidateRebooking": { "Type": "Task", "Resource": "arn:aws:states:::lambda:invoke", "End": true }
        }
      },
      "MaxConcurrency": 1000,
      "End": true
    }
  }
}

Note: the service name is spelled bedrockagentcore (no hyphen) in the Step Functions resource string, but bedrock-agentcore (with a hyphen) in the AgentCore ARN.

MaxConcurrency is set to 1000 to bound fan-out and protect downstream booking and inventory systems. If you omit it or set it to 0, you get the default behavior, which runs up to 10,000 parallel child executions. The agent Task flows directly into a deterministic validation Task.

How it differs from managed multi-agent collaboration

Multi-agent collaboration typically means that a supervisor agent decides which sub-agent runs and which tools it calls. Step Functions moves those decisions out of the agent layer entirely.

This design puts orchestration, fan-out, validation, routing, retries, and the audit trail into Step Functions instead. Routing is a deterministic state you define and can test in isolation, not a model classification you hope will be consistent. You get a per-state execution history (every transition recorded with input and output), whereas agent-layer traces require opt-in and provide reasoning rationale rather than a durable, always-on event log.

Design walkthrough of the reference app

The following image shows the Step Functions state machine implemented by the sample application.

Step Functions state machine showing the rebooking workflow: trigger, enrich, a Distributed Map fan-out with agent and deterministic validation stages, choice routing to human review, and execute stages

Figure 1: The Step Functions state machine for the airline rebooking workflow

Stage 1, Trigger. An Amazon EventBridge rule starts the workflow on a flight-cancellation event.

Stage 2, Enrich. A deterministic Task pulls the passenger manifest, bookings, loyalty status, and preferences into the execution state.

Stage 3, Map fan-out. A Distributed Map iterates affected passengers in parallel. The choice of Map type matters at scale. An inline Map runs up to 40 concurrent iterations, which is the documented threshold for choosing Distributed mode. A Distributed Map runs up to 10,000 parallel child executions by default, the right tool when a hub event affects thousands of passengers.

Stage 4, Agent 1 find alternatives. An AgentCore Task proposes the top three options, reasoning over the passenger’s preferences and constraints.

Stage 5, Deterministic validation of the rebooking proposal. An AWS Lambda Task confirms each proposed flight is bookable by checking live availability, fare rules, and route validity, and it rejects hallucinated options. An agent might confidently propose a flight that does not exist. This stage is where that proposal is caught before it can become a ticket.

Stage 6a, Agent 2 draft compensation. A second AgentCore Task drafts personalized, customer-facing notification text only. It does not compute entitlement and it does not move money.

Stage 6b, Deterministic entitlement check. A Lambda Task computes and validates the entitlement against rule tables before any compensation issues. Consumer-protection frameworks such as EU Regulation 261/2004 (EU261) and US Department of Transportation refund rules are referenced here illustratively, to show why deterministic, auditable computation matters. The specific bands, triggers, and amounts are configuration you own and validate against current legal guidance, not something an agent should infer.

Stage 7, Choice routing and human-in-the-loop. A Choice state auto-confirms rebookings for some passengers and routes the rest to a human. For the cases that need review, the workflow waits on a separate .waitForTaskToken Task, backed by Lambda, Amazon Simple Notification Service (Amazon SNS), or Amazon Simple Queue Service (Amazon SQS), with a 4-hour timeout. The wait happens on this separate callback Task, never on the agent step.

{
  "Comment": "Illustrative - route and wait on a human, not on the agent",
  "RouteDecision": {
    "Type": "Choice",
    "Choices": [
      {
        "Variable": "$.passenger.autoConfirmEligible",
        "BooleanEquals": true,
        "Next": "ExecuteBooking"
      }
    ],
    "Default": "AwaitHumanApproval"
  },
  "AwaitHumanApproval": {
    "Type": "Task",
    "Resource": "arn:aws:states:::sqs:sendMessage.waitForTaskToken",
    "Parameters": {
      "QueueUrl": "https://sqs.us-east-1.amazonaws.com/123456789012/approvals",
      "MessageBody": {
        "taskToken.$": "$$.Task.Token",
        "passengerId.$": "$.passenger.id",
        "options.$": "$.proposal.validatedOptions"
      }
    },
    "TimeoutSeconds": 14400,
    "Next": "ExecuteBooking"
  }
}

Stage 8, Execute. Deterministic Task states confirm the booking, issue compensation, and send confirmation. Each execution Task derives an idempotency token from the passenger ID combined with the decision ID (the child execution name, or a hash of the validated option set) and passes it to the booking and payment APIs, so a retry or redrive is a no-op instead of a duplicate booking or a second payment.

Stage 9, Aggregate and exception routing. The workflow summarizes outcomes and routes any unresolved cases to human agents.

The validation step itself is ordinary deterministic code. A simplified rebooking validator in Python looks like the following.

# Illustrative - reject any option the agent proposed that is not bookable
def handler(event, context):
    passenger = event["passenger"]
    proposed = event["proposal"]["options"]

    validated = []
    for option in proposed:
        flight = lookup_flight(option["flightId"])
        if flight is None:
            continue  # hallucinated or stale flight, reject
        if flight["seatsAvailable"] < 1:
            continue  # no inventory, reject
        if not fare_rules_allow(passenger["fareClass"], flight):
            continue  # fare rule violation, reject
        if not route_is_valid(passenger["origin"], passenger["destination"], flight):
            continue  # invalid route, reject
        validated.append(option)

    return {
        "passengerId": passenger["id"],
        "validatedOptions": validated,
        "autoConfirmEligible": passenger["loyaltyTier"] == "top" and len(validated) > 0,
    }

Best practices and guardrails

Reject hallucinations through validations. No agent proposal is applied without a deterministic validation step passing first. This minimizes the impact of hallucinations, prompt injections, or bugs on your workflow.

Keep a complete audit trail. Step Functions execution history records every state transition, input, and output, and pairing that with durable persistence gives you a per-decision record. You can show exactly which proposal was made, which validation passed or failed, and who approved the exception.

Surface only true exceptions to humans. Humans handle only what validation or the agent cannot resolve. Auto-confirmation handles the clear cases, and people spend their attention on the genuinely ambiguous ones.

Hold executions open cheaply. The .waitForTaskToken callback holds the execution open with no compute charges while the execution is paused. For example, you can cost-efficiently park thousands of pending approvals overnight. Refer to the AWS Step Functions pricing page for current details.

Make execution idempotent. Guard reservation execution and compensation issuance against retries and double-sends, as shown in Stage 8. Derive the idempotency token from the passenger ID and decision ID, and pass it to your booking and payment APIs so that a replay is a no-op.

Respect cost and timeouts. Keep each per-agent Task timeout within the 15-minute quota, bound your Map concurrency to protect downstream systems, and track the token usage returned in the agent response so you can attribute and forecast cost.

Handle errors deliberately. Apply Retry and Catch on the agent Tasks for conditions such as BedrockAgentCore.ThrottlingException and BedrockAgentCore.ResourceNotFoundException, and on the Lambda validation Tasks for their own failure modes. A Catch on an agent Task can route a stuck passenger straight to the human queue rather than failing the whole child execution.

Confirm availability and Region support. Check the current availability status and supported AWS Regions for AgentCore and the Step Functions integration at the AWS Capabilities by Region on Builder Center.

Conclusion

A flight-cancellation event is a challenging test of automated decision-making, because the output can have immediate financial impact. The way to use AI agents safely in that setting is to let them do what they are good at, proposing options and drafting language, while never letting a proposal become an action until deterministic code has approved it. In this design, orchestration, fan-out, validation, routing, and retries are implemented in Step Functions rather than inside an agent’s reasoning. Agents do not make changes directly, and their output is only applied after deterministic validation. You get a per-decision record for review, and you hold exceptions open on a callback that adds no compute or storage cost while it waits.

To get started, deploy the reference pattern from Serverless Land and adapt the validation layer to your own workflow.

Building resilient real-time streaming workers with Amazon DynamoDB leases

Post Syndicated from Siddhesh Tiwari original https://aws.amazon.com/blogs/architecture/building-resilient-real-time-streaming-workers-with-amazon-dynamodb-leases/

Consider a real-time transcription service processing 500 concurrent meetings. Each worker processing these meetings requires a dedicated outbound WebSocket connection to an upstream streaming source. When a single worker fails, it drops 100+ connections, causing 2 to 3 minutes of data loss per connection until operators manually restart services.

Building real-time streaming workers that maintain hundreds of persistent WebSocket connections presents a coordination challenge: when a worker stops unexpectedly, its connections become unmanaged and data stops flowing. Exactly one worker must own each connection, yet workers fail, redeploy, and scale independently. Without a mechanism to track ownership and automatically transfer connections that healthy workers can claim, operators must intervene manually for every failure.

This pattern reduces manual intervention during failures, reduces connection recovery time from minutes to seconds, and helps minimize downtime during deployments without requiring external coordination services.

In this post, you learn how to build a WebSocket fleet management system on Amazon Elastic Container Service (Amazon ECS) and AWS Fargate. Amazon DynamoDB is the primary service that manages distributed lease ownership, coordination, and failover in this solution. For the compute layer, this post uses Amazon ECS on AWS Fargate to run the worker fleet. However, you can adapt this pattern to any compute layer of your choice, such as Amazon Elastic Kubernetes Service (Amazon EKS) or Amazon Elastic Compute Cloud (Amazon EC2) with Auto Scaling groups, without changing the core lease logic. You learn how to implement lease-based ownership with conditional writes, automatic failover through orphan reconciliation, and low downtime deployments through graceful shutdown.

The challenge: managing long-lived WebSocket connections

WebSocket connections are fundamentally different from HTTP requests. An HTTP request arrives, gets processed, and returns a response. The server holds no state between requests. A WebSocket connection, by contrast, is a persistent bidirectional channel. The worker must maintain an open TCP connection, process messages the upstream source sends, and respond to keep-alive pings from the upstream source.

This statefulness introduces several operational challenges:

Worker failures. When a worker process stops unexpectedly or its container terminates, the worker drops its WebSocket connections. The upstream source might buffer data briefly, but without a mechanism to detect the failure and reassign the connection to a healthy worker, the system loses data.

Rolling deployments. ECS rolling deployments terminate old tasks and start new ones. Each terminated task drops its connections. Without coordination, there’s a window where connections have no owner.

Horizontal scaling. Adding workers is straightforward. New tasks start and pick up work. Removing workers is harder. You need to drain connections from departing workers and verify other workers take over before the task exits.

Double-claiming. If two workers both believe they own the same connection, they both attempt to connect to the same upstream source. This can cause duplicate data processing, protocol errors, or connection rejection by the upstream service.

Because the workers are WebSocket clients that initiate outbound connections to upstream sources, you need a coordination mechanism that operates at the application layer rather than the network layer.

Solution overview

The architecture uses six AWS services to coordinate a fleet of WebSocket workers:

Architecture of the WebSocket fleet: API Gateway and Lambda write events to DynamoDB and SQS, and ECS Fargate workers claim leases and publish metrics to CloudWatch

Figure 1: WebSocket fleet management architecture

  1. Amazon API Gateway: You use this to receive START and STOP events from external systems through a REST API. A START event signals that a new streaming session (for example, a meeting or live feed) has begun and requires a dedicated WebSocket connection. A STOP event signals that the streaming session has ended and the connection should be released.
  2. AWS Lambda (event router): You use this to write connection state to Amazon DynamoDB and enqueue a notification to Amazon Simple Queue Service (Amazon SQS).
  3. Amazon DynamoDB: You use this to store connection state and lease ownership. Conditional writes (atomic operations that succeed only if specified conditions are met) can provide distributed locking capabilities without external coordination services.
  4. Amazon SQS: You use this to distribute work notifications to workers for fast pickup of new connections.
  5. Amazon ECS on AWS Fargate: You use this to run the worker fleet. Each worker polls Amazon SQS, manages WebSocket connections, and renews leases through heartbeats.
  6. Amazon CloudWatch: You use this to collect custom metrics (active connection count) that drive ECS automatic scaling.

The key insight is that DynamoDB conditional writes act as a distributed lock without requiring a separate coordination service. Each connection has a lease: a time-bounded ownership claim. Workers must continuously renew their lease. If a worker stops unexpectedly, the lease expires and another worker takes over.

Why not SQS alone or an existing lock client?

SQS plays an important role in this architecture as a fast notification channel, but it cannot serve as the sole coordination mechanism. SQS is designed for task execution, delivering a unit of work to one consumer. WebSocket connection ownership is not a one-time task. It is a continuous state that must be maintained and renewed for the lifetime of the connection. SQS has no mechanism to track who currently owns a connection, query for connections with no active owner, or represent the domain state (desired_state, ws_url, last_seq) needed to manage a connection. DynamoDB provides all these capabilities through persistent items, conditional writes, and secondary indexes.

The amazon-dynamodb-lock-client library published by AWS implements similar distributed locking primitives on DynamoDB. However, it is designed for Java environments and does not integrate domain-specific connection state into the lock record. This solution is implemented in async Python to match the worker architecture, combines lock ownership and connection metadata in a single DynamoDB item to reduce read operations, and uses a GSI to enable fleet-wide reconciliation queries that a general-purpose lock client does not provide.

The lease pattern

A lease is a row in DynamoDB that tracks who owns a connection and when that ownership expires. The table uses the following schema:

Attribute Type Description
Pk String (Partition Key) Connection ID, for example, CONN#meeting-123
desired_state String STARTED or STOPPED
ws_url String Upstream WebSocket URL to connect to
lease_owner String Worker ID that currently owns this connection
lease_expires_at_ms Number Epoch milliseconds when the lease expires
last_seq Number Last processed sequence number (for resumption)

A global secondary index (GSI), a secondary lookup structure that you can use to query on non-primary-key attributes, on desired_state (partition key) and lease_expires_at_ms (sort key) allows efficient queries for unmanaged connections: those with desired_state = STARTED and an expired lease.

A note on clock accuracy

The lease expiration mechanism relies on epoch millisecond timestamps generated by worker processes using their local system clocks. DynamoDB evaluates lease expiration conditions against the now value supplied by the calling worker, not against a DynamoDB server-side clock. This means all workers must have reasonably synchronized clocks for the lease pattern to behave correctly.

AWS Fargate tasks running in the same AWS region receive clock synchronization through the Amazon Time Sync Service, which keeps clock skew between tasks to within a few milliseconds. This is well within the safety margin provided by the default 20-second lease duration and 5-second heartbeat interval. If you deploy this pattern on compute infrastructure outside of AWS Fargate, verify that NTP synchronization is configured and monitor for clock drift. For environments where clock accuracy cannot be guaranteed, increase the lease duration by the maximum expected clock skew to prevent false lease expirations.

The lease lifecycle has four states. Figure 2 shows the lease state machine.

State machine showing the lease lifecycle transitions between the Acquire, Renew, Release, and Expired states

Figure 2: Lease lifecycle

Acquire

A worker claims a connection by writing its worker ID (lease_owner) and a future expiration timestamp (lease_expires_at_ms) to the DynamoDB lease record. The conditional expression ensures that only one worker can succeed: it checks that either no lease exists yet (attribute_not_exists) or the existing lease has already expired (lease_expires_at_ms < :now). If two workers attempt to acquire the same connection simultaneously, DynamoDB evaluates this condition atomically and only one worker succeeds. The other receives a ConditionalCheckFailedException and gracefully backs off.

The following code example is from the worker application (worker.py), which initializes the Amazon DynamoDB table client, worker ID, and configuration at startup. The complete implementation is available in the GitHub repository.

async def try_acquire_lease(pk: str) -> Optional[dict]:
    """Attempt to acquire lease on a connection."""
    try:
        resp = table.update_item(
            Key={"pk": pk},
            UpdateExpression=(
                "SET lease_owner = :w, "
                "lease_expires_at_ms = :exp, "
                "updated_at_ms = :now"
            ),
            ConditionExpression=(
                "attribute_not_exists(lease_expires_at_ms) "
                "OR lease_expires_at_ms < :now"
            ),
            ExpressionAttributeValues={
                ":w": WORKER_ID,
                ":exp": now_ms() + LEASE_SECONDS * 1000,
                ":now": now_ms(),
            },
            ReturnValues="ALL_NEW",
        )
        return resp["Attributes"]
    except ClientError as e:
        if e.response["Error"]["Code"] == "ConditionalCheckFailedException":
            return None  # Another worker already owns this connection
        raise

The ConditionExpression is the critical piece: it succeeds when the lease does not exist yet (attribute_not_exists) or has already expired (lease_expires_at_ms < :now).

Renew

The owning worker renews its lease every few seconds (the heartbeat). The conditional expression verifies the worker still owns the lease:

async def renew_lease(pk: str) -> bool:
    """Renew lease for owned connection."""
    try:
        table.update_item(
            Key={"pk": pk},
            UpdateExpression=(
                "SET lease_expires_at_ms = :exp, "
                "updated_at_ms = :now"
            ),
            ConditionExpression="lease_owner = :w",
            ExpressionAttributeValues={
                ":w": WORKER_ID,
                ":exp": now_ms() + LEASE_SECONDS * 1000,
                ":now": now_ms(),
            },
        )
        return True
    except ClientError:
        return False  # Lost ownership

If renewal returns False, the worker knows it has lost ownership (perhaps another worker acquired the expired lease) and exits cleanly.

Release

During graceful shutdown, the worker explicitly releases its leases so other workers can acquire them immediately rather than waiting for expiration:

async def release_lease(pk: str):
    """Release lease on connection."""
    try:
        table.update_item(
            Key={"pk": pk},
            UpdateExpression=(
                "SET lease_owner = :empty, "
                "lease_expires_at_ms = :zero"
            ),
            ConditionExpression="lease_owner = :w",
            ExpressionAttributeValues={
                ":w": WORKER_ID,
                ":empty": "",
                ":zero": 0,
            },
        )
    except ClientError:
        pass  # Already released or taken by another worker

Expired

When a worker stops unexpectedly, because of a container crash, network partition, or process failure, it can no longer renew its lease. Unlike graceful shutdown, the worker has no opportunity to explicitly release ownership. The lease remains in DynamoDB with the crashed worker’s lease_owner value, but the lease_expires_at_ms timestamp passes without renewal.

This expired lease represents a connection with no active owner: desired_state remains STARTED (the connection should be active) but no healthy worker is managing it. The connection is now an orphan.

The reconciliation loop detects this condition by querying the GSI for records where desired_state = STARTED and lease_expires_at_ms < now. Any healthy worker that finds such a record can attempt to acquire it using the same conditional write used during initial acquisition. Because lease_expires_at_ms < :now is one of the valid conditions for acquisition, the expired lease is treated identically to an unclaimed one.

The Expired state is transient: it exists between the moment a lease stops being renewed and the moment the reconciliation loop runs and a new worker successfully acquires it. The maximum time a connection spends in the Expired state is bounded by the reconciliation interval (default: 60 seconds).

Technical implementation

The following sections walk through each component of the system, starting with how events enter the pipeline and ending with how the fleet scales.

Event ingestion

When an external system needs to start or stop a streaming connection, it sends an event to the Lambda event router through API Gateway. The Lambda function writes the connection state to DynamoDB and enqueues a notification to SQS:

def handler(event, context):
    payload = json.loads(event.get("body", "{}"))
    event_type = payload["event_type"].upper()
    connection_id = payload["connection_id"]
    pk = f"CONN#{connection_id}"

    if event_type == "START":
        table.put_item(Item={
            "pk": pk,
            "desired_state": "STARTED",
            "ws_url": payload["ws_url"],
            "last_seq": 0,
            "lease_owner": "",
            "lease_expires_at_ms": 0,
            "updated_at_ms": now_ms(),
        })
        sqs.send_message(
            QueueUrl=QUEUE_URL,
            MessageBody=json.dumps({"pk": pk})
        )

    elif event_type == "STOP":
        table.update_item(
            Key={"pk": pk},
            UpdateExpression="SET desired_state = :s, updated_at_ms = :t",
            ExpressionAttributeValues={
                ":s": "STOPPED", ":t": now_ms()
            },
        )

    return {"statusCode": 200, "body": "OK"}

DynamoDB is the source of truth for connection state. Amazon SQS serves as a fast notification channel. When a START event arrives, the SQS message immediately notifies available workers that they can claim a new connection, so workers do not need to wait for the next reconciliation cycle (default: 60 seconds) to discover and acquire the new connection. Without SQS, new connections would only be picked up when the reconciliation loop queries the GSI for unmanaged connections on its next scheduled run.

Worker polling

Each ECS Fargate worker runs a continuous SQS polling loop to pick up new connection notifications. The loop follows four steps before starting a new WebSocket connection:

1. Capacity check

Before accepting any new work, the worker checks whether it has reached its maximum connection limit (MAX_CONNECTIONS). If the worker is at capacity, it pauses for 5 seconds and skips the current polling cycle. This prevents a single worker from being overwhelmed while other workers in the fleet remain underutilized.

2. Deduplication

If the worker already manages the connection referenced in the SQS message (tracked in its local connections dictionary), it deletes the message and moves on. This handles cases where the same connection generates multiple SQS notifications, for example during retries or redeliveries.

3. Lease acquisition before WebSocket start

The SQS message is a hint, not a guarantee of ownership. Before starting a WebSocket connection, the worker must successfully acquire the DynamoDB lease using try_acquire_lease. If another worker has already claimed the connection, try_acquire_lease returns None and this worker skips it. This ensures exactly one worker owns each connection at any time.

4. Task creation

If the lease is acquired and desired_state is STARTED, the worker creates an async task to manage the WebSocket connection. The SQS message is then deleted regardless of whether the lease was acquired, preventing repeated reprocessing of the same notification.

The following code shows the full polling loop implementation:

async def poll_sqs():
    while not shutdown_event.is_set():
        if len(connections) >= MAX_CONNECTIONS:
            await asyncio.sleep(5)
            continue

        resp = await asyncio.to_thread(
            sqs.receive_message,
            QueueUrl=QUEUE_URL,
            MaxNumberOfMessages=1,
            WaitTimeSeconds=10,
            VisibilityTimeout=30,
        )

        for msg in resp.get("Messages", []):
            body = json.loads(msg["Body"])
            pk = body["pk"]

            if pk in connections:
                sqs.delete_message(
                    QueueUrl=QUEUE_URL,
                    ReceiptHandle=msg["ReceiptHandle"]
                )
                continue

            conn_data = await try_acquire_lease(pk)
            if conn_data and conn_data.get("desired_state") == "STARTED":
                asyncio.create_task(
                    manage_websocket(
                        pk, conn_data["ws_url"],
                        conn_data.get("last_seq", 0)
                    )
                )
            sqs.delete_message(
                QueueUrl=QUEUE_URL,
                ReceiptHandle=msg["ReceiptHandle"]
            )

Connection management

Once a worker acquires a lease, it opens a WebSocket connection to the upstream source and runs three concurrent async tasks for the lifetime of that connection. These three tasks work together to keep the connection alive, process incoming data, and detect when the connection should stop.

1. Heartbeat loop

The heartbeat loop calls renew_lease every HEARTBEAT_EVERY seconds. If renewal fails, meaning another worker has taken ownership or the lease record has changed, the loop exits immediately. This is the mechanism by which a worker detects that it has lost ownership of a connection mid-flight.

2. Receive loop

The receive loop processes every incoming message from the upstream WebSocket source. Each message is written to a separate DynamoDB messages table with the connection ID, a timestamp, the message data, and the worker ID. The loop runs continuously until the WebSocket connection closes or an error occurs.

3. Desired state checker

Every 10 seconds, the desired state checker reads the connection record from DynamoDB. If desired_state has been set to STOPPED, meaning an external system sent a STOP event through the API, the loop exits, signaling that this connection should be closed even though the WebSocket itself is still open.

How the three tasks interact

All three tasks run concurrently using asyncio.gather. When any one of the three tasks returns or raises an exception, asyncio.gather completes and execution moves to the finally block. This means a single trigger, lease loss, WebSocket closure, or a STOP event, is sufficient to cleanly end the connection regardless of the state of the other two tasks.

Cleanup

The finally block always runs, regardless of how the connection ended. It releases the DynamoDB lease so other workers can acquire the connection immediately and removes the connection from the worker’s local tracking dictionary.

The following code shows the full connection management implementation:

async def manage_websocket(pk: str, ws_url: str, last_seq: int):
    connections[pk] = {"pk": pk, "ws_url": ws_url, "ws": None}

    try:
        async with websockets.connect(ws_url) as ws:
            connections[pk]["ws"] = ws

            async def heartbeat_loop():
                while not shutdown_event.is_set():
                    await asyncio.sleep(HEARTBEAT_EVERY)
                    if not await renew_lease(pk):
                        print(f"[{pk}] Lost lease, closing")
                        return

            async def receive_loop():
                async for msg in ws:
                    data = json.loads(msg)
                    messages_table.put_item(Item={
                        "pk": pk,
                        "sk": str(now_ms()),
                        "message_data": data.get("data", str(data)),
                        "timestamp_ms": now_ms(),
                        "worker_id": WORKER_ID,
                    })

            async def check_desired_state():
                while not shutdown_event.is_set():
                    await asyncio.sleep(10)
                    resp = table.get_item(Key={"pk": pk})
                    if resp.get("Item", {}).get("desired_state") == "STOPPED":
                        return

            await asyncio.gather(
                heartbeat_loop(),
                receive_loop(),
                check_desired_state()
            )

    except Exception as e:
        print(f"[{pk}] WebSocket error: {e}")
    finally:
        await release_lease(pk)
        connections.pop(pk, None)

Production note: The code samples use print() for clarity. In production, replace these with structured logging (the Python logging module or Amazon CloudWatch Logs) and emit CloudWatch metrics for lease acquisition failures and reconnection events to support operational alerting.

Scaling note: The per-connection check_desired_state() loop shown here works for small fleets. At scale, replace individual GetItem calls with a single centralized loop that uses BatchGetItem to check the state of all active connections in one call, reducing DynamoDB reads from N calls every 10 seconds to 1 batched call.

Orphan reconciliation

The reconciliation loop is the safety net of the system. It runs on every worker periodically, independent of the SQS polling loop. Its sole purpose is to find connections that should be active but have no current owner, and reacquire them.

The loop queries the GSI for all records where desired_state = STARTED and lease_expires_at_ms is less than the current time. These are connections that an external system has requested as active, but whose lease has either never been claimed or has expired without renewal, indicating the previous owner is no longer running.

For each orphaned connection found, the worker calls try_acquire_lease. Because try_acquire_lease uses a DynamoDB conditional write, multiple workers can safely run reconciliation concurrently without risk of double-claiming. Exactly one worker succeeds for each connection. The others receive a ConditionalCheckFailedException and move on.

The reconciliation interval (default: 60 seconds) determines the maximum recovery time for unexpected worker terminations. A worker that crashes without running its graceful shutdown handler leaves its leases to expire naturally after LEASE_SECONDS (default: 20 seconds). The reconciliation loop then picks up those connections within the next 60-second cycle, giving a worst-case recovery time of approximately 80 seconds (20 seconds lease expiry plus up to 60 seconds reconciliation interval).

The following code shows the full implementation:

async def reconcile_orphaned_connections():
    while not shutdown_event.is_set():
        await asyncio.sleep(RECONCILE_EVERY)

        if len(connections) >= MAX_CONNECTIONS:
            continue

        resp = table.query(
            IndexName=GSI_NAME,
            KeyConditionExpression=(
                "desired_state = :state "
                "AND lease_expires_at_ms < :now"
            ),
            ExpressionAttributeValues={
                ":state": "STARTED",
                ":now": now_ms()
            },
            Limit=RECONCILE_PAGE_SIZE,
        )

        for item in resp.get("Items", []):
            pk = item["pk"]
            if pk not in connections and len(connections) < MAX_CONNECTIONS:
                conn_data = await try_acquire_lease(pk)
                if conn_data:
                    asyncio.create_task(
                        manage_websocket(
                            pk, conn_data["ws_url"],
                            conn_data.get("last_seq", 0)
                        )
                    )
Failover sequence in which a crashed worker’s lease expires and another worker reacquires the connection through orphan reconciliation

Figure 3: Automatic failover through orphan reconciliation

Graceful shutdown

When ECS sends a SIGTERM signal during a rolling deployment or scale-in event, the worker has a limited window to clean up before the container is forcibly terminated. Rather than dropping connections abruptly and waiting for leases to expire naturally, the worker performs a coordinated shutdown in three steps.

Step 1: Signal propagation

The signal_handler function sets a shared shutdown_event when SIGTERM is received. This event is checked by every running loop across all active connections. The heartbeat loop, the desired state checker, and the reconciliation loop all exit their while not shutdown_event.is_set() loops as soon as the event is set. No additional per-connection shutdown logic is needed. The shared event propagates the shutdown signal automatically to all concurrent tasks.

Step 2: Parallel cleanup

Rather than closing connections and releasing leases sequentially, which would take longer as the number of active connections grows, the worker closes all WebSocket connections and releases all leases concurrently using asyncio.gather. For a worker managing hundreds of connections, this keeps the total shutdown time roughly constant regardless of connection count.

Step 3: Immediate lease release

During graceful shutdown, the worker sets lease_expires_at_ms = 0 for each released connection. A value of 0 means the lease appears already expired to any worker running a reconciliation query. Other workers in the fleet pick up the released connections on their next reconciliation cycle rather than waiting for the original lease duration (default: 20 seconds) to elapse naturally.

Contrast with unexpected termination

Graceful shutdown is the fast path. When a worker exits cleanly through SIGTERM, connections are available for reacquisition within one reconciliation cycle. When a worker crashes unexpectedly without running the shutdown handler, leases expire naturally after LEASE_SECONDS (default: 20 seconds) and are then picked up by the reconciliation loop. Both paths converge on the same outcome, another worker acquires the connection, but graceful shutdown is significantly faster.

The following code shows the full graceful shutdown implementation:

shutdown_event = asyncio.Event()

def signal_handler(signum, frame):
    shutdown_event.set()

async def graceful_shutdown():
    await shutdown_event.wait()
    tasks = []
    for pk, conn in list(connections.items()):
        if conn.get("ws"):
            tasks.append(conn["ws"].close())
        tasks.append(release_lease(pk))
    await asyncio.gather(*tasks, return_exceptions=True)

Setting shutdown_event causes the heartbeat loops and state checkers to exit their while not shutdown_event.is_set() loops. The graceful_shutdown function then closes the active WebSocket connections and releases its leases in parallel. Released leases have lease_expires_at_ms = 0, which means the reconciliation loop on other workers picks them up on its next cycle rather than waiting for the original lease to expire.

Scaling the fleet

Each worker publishes a custom CloudWatch metric with its active connection count:

async def publish_metrics():
    while not shutdown_event.is_set():
        await asyncio.sleep(30)
        cw.put_metric_data(
            Namespace="WsFleet",
            MetricData=[{
                "MetricName": "ActiveConnections",
                "Value": len(connections),
                "Unit": "Count",
                "Dimensions": [
                    {"Name": "ServiceName", "Value": SERVICE_NAME}
                ],
            }],
        )

An AWS Application Auto Scaling target tracking policy scales the fleet based on the average ActiveConnections metric across all workers. When the average exceeds the target (for example, 700 connections per task), ECS launches additional tasks. New tasks start their SQS polling and reconciliation loops, picking up new connections and rebalancing the fleet.

Application Auto Scaling adds and removes ECS tasks based on the average ActiveConnections CloudWatch metric across the worker fleet

Figure 4: Automatic scaling based on active connection count

Scale-in is safe because of the lease pattern. When ECS terminates a task, the worker receives SIGTERM, releases its leases, and other workers acquire the freed connections through reconciliation.

Configuration Value Rationale
Lease duration 20 seconds Long enough to survive brief network hiccups, short enough for fast failover
Heartbeat interval 5 seconds Renew well before expiration (4x safety margin)
Reconciliation interval 60 seconds Balance between recovery speed and DynamoDB read cost
Max connections per task 700 Based on memory and CPU profiling per connection
Scale-out cool down 2 minutes Prevent thrashing during traffic spikes
Scale-in cool down 15 minutes Allow connections to stabilize before removing capacity

Tuning guidance. These values represent a starting point. Adjust based on your requirements:

  • Lease duration: Start with 20s. Reduce for faster failover, increase if network hiccups cause false expirations.
  • Heartbeat interval: Keep below lease duration. A 4:1 ratio (lease:heartbeat) gives 4 renewal attempts before expiry.
  • Reconciliation interval: Start with 60s. Reduce for faster recovery from unexpected terminations, increase to lower DynamoDB read cost.
  • Max connections per task: Start with 100 and increase while monitoring memory and CPU utilization in CloudWatch Container Insights. Each WebSocket connection typically consumes 2-5 MB of memory depending on message throughput.

DynamoDB cost considerations

The dominant cost driver in this architecture is heartbeat writes. Each active connection generates one update_item call per heartbeat interval, consuming 1 WCU. At the default 5-second heartbeat interval:

Active connections WCUs/second Approx. monthly cost (on demand) Approx. monthly cost (provisioned)
100 20 ~$65 ~$10
500 100 ~$325 ~$47
2,000 400 ~$1,300 ~$190

For production deployments at sustained high connection counts, use provisioned capacity with Auto Scaling rather than on-demand pricing. Heartbeat writes are predictable and consistent, which makes them well-suited to provisioned throughput. Configure Auto Scaling on your provisioned capacity to track connection count changes as the fleet scales.

To reduce cost, consider the following adjustments:

  1. Increase the heartbeat interval. Doubling the heartbeat interval from 5 seconds to 10 seconds halves WCU consumption. Maintain the 4:1 lease-to-heartbeat ratio by also doubling the lease duration. This increases the failover window proportionally.
  2. Increase the reconciliation interval. Increasing from 60 seconds to 120 seconds halves RCU consumption from reconciliation queries. This slows recovery from unexpected terminations.
  3. Use BatchGetItem for desired state checks. Replace the per-connection get_item calls in the check_desired_state loop with a single BatchGetItem call covering all active connections. This reduces RCU consumption from N reads per cycle to 1 batched read per cycle.

    GSI queries during reconciliation use eventually consistent reads by default, which halves the RCU cost compared to strongly consistent reads. Monitor your GSI read consumption in the DynamoDB console and adjust the reconciliation page size and interval to stay within your cost targets.

Conclusion

Managing long-lived WebSocket connections at scale requires explicit ownership tracking, automatic failover, and coordination across a fleet of workers. This post showed you a pattern that addresses these challenges using DynamoDB conditional writes as a distributed lease mechanism.

Key takeaways:

  • You can use DynamoDB conditional writes for atomic distributed coordination without external lock services. The ConditionExpression on update_item helps confirm one worker owns each connection at a time.
  • The heartbeat and reconciliation pattern handles the full failure spectrum. Lease expiration detects unexpected worker terminations. Graceful shutdown handles rolling deployments. New workers acquire leases and departing workers release them, making scaling safe.
  • This pattern applies to systems that manage long-lived WebSocket connections at scale: real-time transcription, IoT data ingestion, financial feed processing, or live event streaming.

Getting started

The complete implementation, including the worker application, Lambda event router, and Terraform templates for the DynamoDB table, SQS queue, and ECS cluster, is available in the GitHub repository. Follow the instructions in the repository README to deploy the infrastructure and validate the lease lifecycle with a small set of test connections.

For further enhancements, add distributed tracing with AWS X-Ray for end-to-end visibility across workers, and implement reconnection logic with upstream replay or offset-based resumption to handle data gaps between worker failure and recovery.

Further reading

How to migrate from Amazon CloudSearch to Amazon OpenSearch Serverless

Post Syndicated from Prasad Nadig original https://aws.amazon.com/blogs/big-data/how-to-migrate-from-amazon-cloudsearch-to-amazon-opensearch-serverless/

If you run search on Amazon CloudSearch, now is the time to plan your migration to Amazon OpenSearch Serverless. Modern search has moved on to capabilities beyond what CloudSearch provides: semantic and hybrid search, Retrieval Augmented Generation (RAG), and agentic search. OpenSearch Serverless gives you all of these with automatic scaling on a pay-for-what-you-use basis. You don’t need to choose or maintain infrastructure. OpenSearch Serverless maintains the hands-off, operational simplicity of CloudSearch.

This post shows you how to migrate your CloudSearch domain to an Amazon OpenSearch Serverless collection. We walk you through assessing your CloudSearch configuration, creating an OpenSearch Serverless collection with explicit index mappings, converting your documents and queries, configuring security policies, loading your data with Amazon OpenSearch Ingestion, and validating the migration before cutting over.

Key differences to note

Prerequisites

To follow along with this post, you need the following:

  • An AWS account.
  • An existing Amazon CloudSearch domain with indexed data.
  • Source data available in a durable store such as Amazon Simple Storage Service (Amazon S3) or Amazon DynamoDB (CloudSearch doesn’t provide a built-in export or backup feature, so your original source data is required to re-ingest into OpenSearch).
  • AWS Identity and Access Management (IAM) permissions to create and manage Amazon OpenSearch Serverless collections, encryption policies, network policies, and data access policies.
  • An Amazon OpenSearch Ingestion pipeline (or alternative ingestion method) for loading data.

Plan the migration

Planning is where you decide what success means: minimal downtime, no data loss, current functionality preserved, and custom configurations carried over. You don’t need to plan for infrastructure because OpenSearch Serverless provisions and scales compute for you. Your main planning task is to assess your current CloudSearch configuration so you can reproduce its behavior on the target.

Document your existing setup from the Amazon CloudSearch console. Record the current instance type, the partition count, and the replication count. Capture the total document count and overall data size, and record every field definition, including field types and the search, facet, and sort settings for each field. Note any analyzers, synonyms, stopwords, or custom rank expressions. Note whether you use the 2011 or the 2013 CloudSearch API version, because the 2013 API added faceting and filtering features that change how you model the target.

OpenSearch Serverless is the right target for most CloudSearch workloads, but not all of them. If your workload needs very low read-after-write latency (a short refresh interval), tight and predictable query response times, or direct control over instance configuration, choose an Amazon OpenSearch Service managed clusters deployment instead and size it from your workload profile.

The migration involves four main concerns: your source data format, your queries, your field definitions, and your access policies. Before you plan the details, it helps to see the whole migration at once. The following diagram maps the migration across four phases: your source CloudSearch environment, the migration pipeline that converts and moves your data, the OpenSearch Serverless target, and cutover and operations.

Migration workflow across four phases: source CloudSearch, migration pipeline, OpenSearch Serverless target, and cutover and operations

Figure 1: The migration workflow across four phases

In the source environment, you assess your CloudSearch configuration and back up your source data (Amazon S3, Amazon DynamoDB, or another store). Note the Source Data Format (SDF), the URL-based query syntax, and the IAM access policies you need to carry over. In the migration pipeline, you map field types, convert the data format from CloudSearch JSON to OpenSearch-compatible JSON, convert your queries to the OpenSearch query domain-specific language (DSL), configure security, bulk-ingest the data, and validate the result. The OpenSearch Serverless target holds the collection, index mappings, ingested documents, and the encryption, network, and data access policies, and it scales with your workload on a pay-per-use basis. In cutover and operations, you update your application to the new endpoint and clients, monitor with Amazon CloudWatch, and decommission CloudSearch once no traffic remains.

Model your data in OpenSearch Service

OpenSearch Service uses index mappings to define the fields and data types in an index. Because you know your CloudSearch schema, define the target mapping explicitly when you create the index. Create the index and set its mapping in a single request, and set dynamic to strict so OpenSearch rejects any document that contains a field you did not define. Strict mapping catches schema drift at ingest time, avoiding the default OpenSearch behavior of creating new mappings for undefined fields.

PUT /imdb_movies
{
  "mappings": {
    "dynamic": "strict",
    "properties": {
      "title": {
        "type": "text",
        "fields": {
          "keyword": { "type": "keyword" }
        }
      },
      "genres": { "type": "keyword" },
      "rating": { "type": "float" },
      "release_date": { "type": "date" }
      ...
    }
  }
}

Field type mapping

The following table maps CloudSearch field types to their OpenSearch Service equivalents.

CloudSearch OpenSearch Service equivalent Notes
text text Text is tokenized. Stemming, synonyms, and stopwords apply. Good for matching user terms.
literal keyword Not tokenized. Good for exact-match search.
int integer Use for ranking, faceting, and narrowing.
double float or double .
date date .
boolean boolean .
latlon geo_point .
text-array text OpenSearch handles arrays natively, so map to the base text type.
literal-array keyword OpenSearch handles arrays natively, so map to the base keyword type.
multi-value nested or object .
long long .
binary binary .

Two mapping details deserve attention. First, pick the smallest numeric type that fits your data rather than copying the widths CloudSearch uses. CloudSearch stores integers as 64-bit values, but few datasets hold numbers that large. A long or a double consumes more disk than an integer, a short, or a float with no benefit when the values are small. Evaluate the actual range of each field and choose the narrowest type that holds it. Reserve long for values that genuinely exceed the roughly 2.1 billion ceiling of integer, and use float instead of double unless you need double precision. Smaller types shrink your index and speed up queries.

Second, if you sort or aggregate on a text field, add a keyword sub-field. The preceding example mapping has a keyword subfield for the title field. You access the field using dot notation: title.keyword. OpenSearch doesn’t sort or aggregate analyzed text fields by default.

As noted earlier, if you run several CloudSearch domains, model each one as a separate index within a single OpenSearch Serverless collection to consolidate them.

Move your data

Migrating to OpenSearch Service is a re-ingestion: you convert your source documents and index them into the collection you created. CloudSearch doesn’t provide a built-in backup or snapshot feature. It relies on the documents you send through the indexing process, so before you migrate, make sure your source data is available in a durable store such as Amazon S3, Amazon DynamoDB, or another database.

The conversion is a format translation. CloudSearch accepts data in SDF as JSON or XML, where a document batch is a collection of add and delete operations. The JSON that CloudSearch uses differs from the JSON that OpenSearch Service expects, so you must transform each source document into an OpenSearch document whose fields match the index mapping you defined earlier. Handle the same details the mapping calls out: emit each numeric value so it fits the narrow type you chose for its field rather than a wide long or double, format dates to match your date mapping, and drop or rename any field that your strict mapping doesn’t define.

CloudSearch batch format showing add and delete operations in JSON OpenSearch bulk batch format showing index operations in JSON

Figure 2: CloudSearch batch format (left) compared to OpenSearch batch format (right)

You can write a small conversion script. Have the script write its output to an Amazon S3 bucket so the converted documents live in a durable store you can re-ingest from as many times as you need.

With your converted documents in Amazon S3, use Amazon OpenSearch Ingestion to load them. Amazon OpenSearch Ingestion is a feature of Amazon OpenSearch Service that you can use to ingest, filter, transform, enrich, and route data to an Amazon OpenSearch Service domain or an OpenSearch Serverless collection. Configure an OpenSearch Ingestion pipeline with an Amazon S3 source (you can use an OpenSearch Ingestion blueprint to get started) that reads your converted documents. Let its built-in processors apply any final transformation before the pipeline writes to your collection. A managed pipeline reading from Amazon S3 gives you a repeatable, restartable load without operating ingestion infrastructure, which makes it the recommended path for most migrations.

If you prefer to load data directly, OpenSearch Service exposes a REST API, so you can index documents with a standard client such as curl or with the OpenSearch client libraries for many languages. Direct indexing is convenient for a small dataset or a quick test, but an Amazon S3 source with OpenSearch Ingestion is the better choice for a production migration.

Convert your queries

CloudSearch uses a URL-based query format. You pass a query parameter in the URL and submit either a simple string search or a JSON-formatted query. OpenSearch Service uses a REST API and the OpenSearch query DSL in the request body, which gives you compound queries, function scoring, and richer relevance control. You can use generative AI coding assistants to help with this translation. Provide your CloudSearch query patterns, and the model generates the equivalent OpenSearch query DSL, which you then validate against your test cases.

Query syntax changes

CloudSearch appends parameters such as sort to the query URL, while OpenSearch expresses sorting, filtering, and boosting as explicit elements of the request body. For example, a title search for “shakespeare” in CloudSearch looks like the following.

https://my-cloudsearch-domain.us-east-1.cloudsearch.amazonaws.com/2013-01-01/search?q=shakespeare&size=10

The equivalent query in OpenSearch Service uses the query DSL.

GET /imdb_movies/_search
{
  "query": {
    "match": { "title": "shakespeare" }
  }
}

To keep result sets consistent after migration, set the default operator to AND in OpenSearch to match the default query behavior of CloudSearch. The following table shows common CloudSearch query patterns and their OpenSearch Service equivalents, using a sample IMDB movies dataset.

Query type CloudSearch (Lucene syntax) OpenSearch Service query DSL
Compound AND title:"Inception" AND genres:"Sci-Fi" {"query":{"bool":{"must":[{"match":{"title":"Inception"}},{"match":{"genres":"Sci-Fi"}}]}}}
Compound NOT title:"Star Wars" AND NOT genres:"Comedy" {"query":{"bool":{"must":[{"match":{"title":"Star Wars"}}],"must_not":[{"match":{"genres":"Comedy"}}]}}}
Wildcard title:Batman* {"query":{"wildcard":{"title":{"value":"batman*"}}}}
Numeric range rating:[7 TO 9] {"query":{"range":{"rating":{"gte":7,"lte":9}}}}
Date range (after) release_date:[2015-01-01T00:00:00Z TO *] {"query":{"range":{"release_date":{"gte":"2015-01-01T00:00:00Z"}}}}
Boosting title:"The Matrix"^6 OR genres:"Sci-Fi"^4 {"query":{"bool":{"should":[{"query_string":{"query":"title": \"The Matrix\"^6","fields":["title"]}},{"query_string":{"query":"genres:\"Sci-Fi\"^4","fields":["genres"]}}]}}}
Sorting title:"Batman" sort=release_date desc {"query":{"match":{"title":"Batman"}},"sort":[{"release_date":{"order":"desc"}}]}

Sorting and boosting

Boosting is useful when you want certain fields or terms to carry more weight in relevance scoring. A higher boost value means the term contributes more to the score. OpenSearch also supports sorting by _score (relevance), which is the default when you specify no sort. For the full query language, see the OpenSearch query DSL documentation.

Configure security

CloudSearch uses AWS Identity and Access Management policies to control access to its configuration and domain service APIs. You attach user-based policies to an IAM role, user, or group, and the document, search, and suggest actions in those policies control access to the CloudSearch APIs.

OpenSearch Serverless applies security through policies at several layers.

  • Collections: Encrypted at rest by default, using either an AWS owned key or a customer managed key defined in an encryption policy.
  • Network policies: Define whether a collection is reachable privately through a virtual private cloud (VPC) endpoint or over the internet.
  • Data access policies: Control which IAM principals and Security Assertion Markup Language (SAML) identities can create indexes and read or write data in the collection.

Amazon OpenSearch Service provisioned domains also offer fine-grained access control, with role-based access control and security at the index, document, and field level. For OpenSearch Serverless, data access policies provide collection-level and index-level permissions, controlling which IAM principals and SAML identities can create, read, or write data within a collection.

Validate the migration

Validation confirms that the migration is complete and correct before you send production traffic to OpenSearch Serverless. Work through five kinds of validation.

  • Documents: Check your document count. Your OpenSearch Serverless indexes should have the same count as your CloudSearch indexes.
  • Queries: Translate your most important queries and run them manually against your collection. Spot check the output for the presence of important results.
  • Ranking: Check the order of results, especially for queries with custom rank functions or field weighting. Results might not match exactly, so look for anything that’s incorrect.
  • Latency: Ideally you should tee your production traffic to your Serverless collection to get real latency metrics. Worst case, generate at least 100,000 synthetic queries across all your query types and run them. Monitor OpenSearch Compute Unit (OCU) consumption with Amazon CloudWatch to understand your cost profile.

To validate search functionality, run the same query against both systems and compare the results. Reuse the query pairs from the conversion step so you exercise the syntax differences directly. For example, to check a numeric range against the sample IMDB movies dataset, run the following query in CloudSearch.

https://my-cloudsearch-domain.us-east-1.cloudsearch.amazonaws.com/2013-01-01/search?q=rating: [7 TO 9]&size=10

Run the equivalent query DSL against your OpenSearch Serverless collection.

GET /imdb_movies/_search
{
  "query": {
    "range": { "rating": { "gte": 7, "lte": 9 } }
  }
}

Confirm that both queries return the same set of movies. Then repeat the comparison for a query that exercises relevance, such as the boosted query from the conversion step, and confirm the top results appear in the same order.

GET /imdb_movies/_search
{
  "query": {
    "bool": {
      "should": [
        { "match": { "title": { "query": "The Matrix", "boost": 6 } } },
        { "match": { "genres": { "query": "Sci-Fi", "boost": 4 } } }
      ]
    }
  }
}

Cut over and operate

When validation passes, update your application to use the OpenSearch Serverless endpoint and the query DSL, and switch from the CloudSearch SDK to the OpenSearch client libraries. After cutover, confirm that no application still points to a CloudSearch endpoint, retain your source data backups in Amazon S3 for rollback, and then delete the CloudSearch domain.

Operating OpenSearch Serverless in production is lighter than operating a domain, because OpenSearch Serverless scales compute for you and you do not tune shards, instance types, or capacity. Your focus shifts to cost and search quality. Monitor OCU consumption and search latency with Amazon CloudWatch, and set alarms on the thresholds that matter to you. Review OCU usage patterns to understand cost and find optimization opportunities, and set capacity limits on the collection to cap the maximum OCUs it can consume. For guidance, see Managing capacity limits for Amazon OpenSearch Serverless and Monitoring Amazon OpenSearch Serverless.

Cost considerations

With OpenSearch Serverless, you pay only for the compute and storage your workload consumes, and OpenSearch Serverless charges for compute and storage separately. OpenSearch Serverless scales indexing compute and search compute independently, so a write-heavy or a read-heavy workload scales only the dimension it needs, and compute can scale to zero when a collection is idle, in which case you pay only for storage. To share hardware across workloads, place collections in a collection group so they draw from the same compute rather than each provisioning its own. For pricing and unit details, see Amazon OpenSearch Service pricing.

Clean up

Because you’re migrating to OpenSearch Serverless, the resources that you’ve created will likely become your production resources. If not, delete any OpenSearch Serverless collections and S3 buckets you created to avoid incurring ongoing cost.

Conclusion

In this post, you saw how Amazon CloudSearch and Amazon OpenSearch Serverless compare, and how the concepts you rely on in CloudSearch (field types, query syntax, autoscaling, and access control) translate into OpenSearch Service. You assess your CloudSearch configuration, model your data with explicit OpenSearch mappings, move your converted documents into the collection with OpenSearch Ingestion, convert your URL-based queries into the OpenSearch query DSL, configure security, and validate before cutover. OpenSearch Serverless gives you the hands-off operational model you have with CloudSearch, and adds richer query capabilities, granular data access policies, and automatic scaling. To get started, create an OpenSearch Serverless collection on the AWS Management Console and follow the steps in this post.

To learn more, see the following resources:


About the authors

Prasad Nadig

Prasad Nadig

Prasad is a Senior Analytics Specialist Solutions Architect at Amazon Web Services (AWS), specializing in large-scale data analytics and AI. Prasad partners with customers to design, migrate, and modernize their analytics platforms on AWS into scalable, cost-effective solutions, with deep expertise in data lakes, data warehousing, distributed processing, and performance tuning at petabyte scale.

Jon Handler

Jon Handler

Jon is a Senior Principal Solutions Architect for Search Services at Amazon Web Services. Jon works closely with OpenSearch and Amazon OpenSearch Service, providing help and guidance to a broad range of customers who have search and log analytics workloads. Prior to joining AWS, Jon’s career as a software developer included four years of coding a large-scale, eCommerce search engine.

Accelerating Spark queries with Iceberg materialized views

Post Syndicated from Yuzhou Sun original https://aws.amazon.com/blogs/big-data/accelerating-spark-queries-with-iceberg-materialized-views/

In this post, you learn how to reduce Apache Spark query execution time with Apache Iceberg materialized views without changing a single SQL query.

Organizations running analytical workloads on their data lakes often hit a common wall: queries that are slow and costly, yet difficult to rewrite by hand. Multi-table joins, heavy aggregations, and window functions over large fact tables all drive up execution times, but the SQL behind them often can’t be changed. It might come from business intelligence (BI) dashboards, packaged independent software vendor (ISV) applications, or legacy reports, where editing the source introduces regression risk that outweighs the performance gain.

Starting with Amazon EMR 7.12.0 and AWS Glue 5.1, you can accelerate these queries without rewriting them. Automatic query rewrite analyzes the logical plan of each incoming query and compares it against a metadata cache of available MVs. When the optimizer finds a materialized view (MV) that satisfies all or part of a query, it rewrites the plan to read from that MV instead of the base tables. Matches can be structural (aggregations and joins) or exact (more complex patterns like window functions). If no MV matches, the original query runs unchanged with no impact on correctness.

If you have previously tried to speed up slow analytical queries, you might have considered one of the following alternatives. Here is how automatic query rewrite compares:

Query modification approach Stored results Refreshes Modification to existing queries
Standard views in AWS Glue No (re-runs each time) n/a Required
Custom ETL pipeline Yes Manual Required
Hand-rolled rewrite Yes Manual Required
Materialized views with automatic rewrite enabled Yes Automatically through AWS Glue Data Catalog on a schedule when configured Not required when supported

In this post, we:

  • Give a high-level overview of how automatic query rewrite works in Apache Spark.
  • Walk through a concrete example, showing how the same query can benefit from MVs at different levels of coverage.
  • Discuss the trade-offs so you can choose the right MV shape for your workload.

Prerequisites

To use automatic query rewrite with Iceberg materialized views, you need:

  • Amazon EMR release 7.12.0 or later, or AWS Glue 5.1 or later.
  • Source tables in Apache Iceberg or Parquet format, registered in the AWS Glue Data Catalog, in the same AWS Region and account as the materialized view. Parquet source tables are supported for automatic query rewrite starting with Amazon EMR 7.14.0 and AWS Glue 8.1.
  • An Amazon Simple Storage Service (Amazon S3) Tables (a capability of Amazon S3) bucket, or an S3 general purpose bucket, for the materialized view data.
  • Permissions for the definer role. You can use AWS Identity and Access Management (IAM) policies or AWS Lake Formation.
  • Automatic query rewrite turned on in your Spark session: --conf spark.sql.optimizer.answerQueriesWithMVs.enabled=true.
  • For Parquet source tables, set spark.sql.materializedView.v1SourceTables.enabled=true and spark.sql.materializedView.v1ETagVersioning.enabled=true.

For more Spark configurations, see Introducing Apache Iceberg materialized views in AWS Glue Data Catalog.

How it works

Here is how MVs and automatic query rewrite work together:

  • You define a SQL query with aggregations, joins, or filters across your supported source tables.
  • AWS Glue Data Catalog stores the precomputed results as an Apache Iceberg table in your Amazon S3 bucket. You can store it in a general purpose S3 bucket or in Amazon S3 Tables. Any Apache Iceberg-compatible query engine can read the materialized view, including Amazon Athena, Amazon EMR, AWS Glue, Amazon Redshift, and Iceberg-compatible third-party query engines. Automatic query rewrite is available on the AWS optimized Spark runtime in Amazon Athena, Amazon EMR, and AWS Glue. Other engines can query the materialized view directly, but they don’t rewrite queries to use it automatically.
  • Automatic refresh keeps the MV current on a schedule that you define, for example SCHEDULE REFRESH EVERY 1 DAY. You set it at creation time or later with ALTER MATERIALIZED VIEW ... ADD SCHEDULE REFRESH. At that scheduled time, the refresh process checks the current Apache Iceberg snapshot ID or Parquet file ETags and refreshes the MV when it detects source-table changes.
  • Automatic query rewrite redirects matching queries to the MV at query optimization time. Automatic query rewrite in Apache Spark uses two matching strategies:
    • Structural rewrite (adapted from Amazon Redshift) handles an MV defined as a single SELECT-FROM-WHERE-GROUP-BY block over INNER joins. The optimizer can roll up an MV’s aggregates to a coarser grain and pull extra query predicates up onto the MV scan.
    • Exact-match rewrite handles MVs defined as other shapes, such as window functions and outer joins, by matching a canonicalized form of the MV body against subtrees of the query plan.

When the optimizer evaluates a query, it consults a metadata cache of MVs from the configured catalogs and chooses the best match. It also checks MV staleness during optimization. It skips stale MVs, so rewrite won’t return stale results. If no MV matches, the original query runs unchanged.

Note that automatic query rewrite is opt-in: set spark.sql.optimizer.answerQueriesWithMVs.enabled=true when creating the Apache Spark session.

Example: One query with three potential MVs

An MV doesn’t need to cover an entire query to help it. Automatic query rewrite in Apache Spark operates on subtrees: when an MV matches a portion of your query plan, the rewriter substitutes that subtree and lets the rest of the query run on the rewrite output unchanged. The same query can therefore be served by many possible MV designs, each making a different trade-off between per-query speedup, storage cost, and reuse across other queries.

To make this concrete, consider a typical analytics query: “Top 100 preferred US customers by total store spending.” It joins fact and dimension tables, applies two selective filters on the customer dimension, aggregates per customer, ranks the result with a window function, and keeps only the top 100:

SELECT c_customer_id, total_revenue, num_transactions, avg_purchase, revenue_rank
FROM (
    SELECT cust.c_customer_id,
        SUM(sales.ss_quantity * sales.ss_sales_price) AS total_revenue,
        COUNT(*) AS num_transactions,
        AVG(sales.ss_quantity * sales.ss_sales_price) AS avg_purchase,
        RANK() OVER (ORDER BY SUM(sales.ss_quantity * sales.ss_sales_price) DESC) AS revenue_rank
    FROM base_catalog.base_db.store_sales sales
    INNER JOIN base_catalog.base_db.customer cust
        ON sales.ss_customer_sk = cust.c_customer_sk
    WHERE cust.c_birth_country = 'UNITED STATES'
        AND cust.c_preferred_cust_flag = 'Y'
    GROUP BY cust.c_customer_id
) ranked
WHERE revenue_rank <= 100
ORDER BY revenue_rank;

Query 1: The original query. Top 100 preferred US customers by total store spending, before any materialized view.

Three MV designs cover progressively more of this query, from a single-table pre-aggregate to the full query body itself:

Tier 1: Pre-aggregate store_sales only, no join, no filter. This tier is a single-table aggregate of store_sales at customer-surrogate-key grain. The query still must join the customer table, apply both filters, re-aggregate at c_customer_id grain, and run the window function.

CREATE MATERIALIZED VIEW mv_catalog.mv_db.customer_tier_1 AS
SELECT ss_customer_sk,
    SUM(ss_quantity * ss_sales_price) AS sum_revenue,
    COUNT(ss_quantity * ss_sales_price) AS count_revenue,
    COUNT(*) AS num
FROM base_catalog.base_db.store_sales
GROUP BY ss_customer_sk;

Tier 1 MV: Single-table pre-aggregate of store_sales by customer surrogate key (no join, no filter).

The following plans compare the original query plan to the rewritten plan:

Window, filter, Sort
+- Aggregate by c_customer_id
:  total_revenue = SUM(ss_quantity * ss_sales_price)
:  num_transactions = COUNT(*)
:  avg_purchase = AVG(ss_quantity * ss_sales_price)
+- Project
   +- Join Inner ON ss_customer_sk = c_customer_sk
      :- BatchScan store_sales <- reads the large store_sales table
      +- Filter c_birth_country='UNITED STATES' AND c_preferred_cust_flag='Y'
         +- BatchScan customer

Plan 1: Original plan. Scans the large store_sales table.

Window, filter, Sort
+- Aggregate by c_customer_id <- rolls up pre-aggregated sums
:  total_revenue = SUM(sum_revenue) <- sum of sum_revenue
:  num_transactions = SUM(num) <- sum of num
:  avg_purchase = SUM(sum_revenue) / SUM(count_revenue) <- sum of sum_revenue / sum of count_revenue
+- Project
   +- Join Inner ON ss_customer_sk = c_customer_sk
      :- BatchScan customer_tier_1 <- reads pre-aggregated MV
      +- Filter c_birth_country='UNITED STATES' AND c_preferred_cust_flag='Y'
         +- BatchScan customer

Plan 2: Rewritten plan (Tier 1). Reads the pre-aggregated customer_tier_1 MV.

Tier 2: Pre-join store_sales x customer, pre-apply one filter (c_preferred_cust_flag = ‘Y’). The middle tier pre-joins both tables and bakes in the preferred-customer filter. The query still must apply the country filter as a residual on the MV scan and run the RANK() window.

CREATE MATERIALIZED VIEW mv_catalog.mv_db.customer_tier_2 AS
SELECT cust.c_customer_id, cust.c_birth_country,
    SUM(sales.ss_quantity * sales.ss_sales_price) AS sum_revenue,
    COUNT(sales.ss_quantity * sales.ss_sales_price) AS count_revenue,
    COUNT(*) AS num
FROM base_catalog.base_db.store_sales sales
INNER JOIN base_catalog.base_db.customer cust
    ON sales.ss_customer_sk = cust.c_customer_sk
WHERE cust.c_preferred_cust_flag = 'Y'
GROUP BY cust.c_customer_id, cust.c_birth_country;

Tier 2 MV: Pre-joins store_sales and customer, with the preferred-customer filter applied.

Rewritten query plan:

Window, filter, Sort
+- Aggregate by c_customer_id <- rolls up pre-aggregated sums
:  total_revenue = SUM(sum_revenue) <- sum of sum_revenue
:  num_transactions = SUM(num) <- sum of num
:  avg_purchase = SUM(sum_revenue) / SUM(count_revenue) <- reads pre-aggregated MV
+- Filter c_birth_country='UNITED STATES' [residual filter on MV scan]
   +- BatchScan customer_tier_2 <- reads pre-aggregated MV

Plan 3: Rewritten plan (Tier 2). Country filter applied as a residual on the MV scan.

Tier 3: Match the entire query, including the window function and top N filter. This is the most specific tier. The MV body is the target query verbatim (minus the top-level ORDER BY, which is meaningless for a stored set). The MV stores the top-ranked rows the query asks for (rank ≤ 100).

CREATE MATERIALIZED VIEW mv_catalog.mv_db.customer_tier_3 AS
SELECT c_customer_id, total_revenue, num_transactions, avg_purchase, revenue_rank
FROM (
    SELECT cust.c_customer_id,
        SUM(sales.ss_quantity * sales.ss_sales_price) AS total_revenue,
        COUNT(*) AS num_transactions,
        AVG(sales.ss_quantity * sales.ss_sales_price) AS avg_purchase,
        RANK() OVER (ORDER BY SUM(sales.ss_quantity * sales.ss_sales_price) DESC) AS revenue_rank
    FROM base_catalog.base_db.store_sales sales
    INNER JOIN base_catalog.base_db.customer cust
        ON sales.ss_customer_sk = cust.c_customer_sk
    WHERE cust.c_birth_country = 'UNITED STATES'
        AND cust.c_preferred_cust_flag = 'Y'
    GROUP BY cust.c_customer_id
) ranked
WHERE revenue_rank <= 100;

Tier 3 MV: Stores the exact ranked output of the query (exact-match path).

This tier exercises the exact-match rewrite path: the rewriter canonicalizes the MV body and matches it against the query’s logical plan.

Rewritten plan:

Sort revenue_rank ASC
+- BatchScan customer_tier_3 <- reads around 100 stored rows

Plan 4: Rewritten plan (Tier 3). Reads around 100 stored rows.

The trade-off

The three tiers trade per-query speedup against reuse and storage. In our testing on TPC-DS 3 TB, we observed the following:

MV design Pre-computed Reuse Per-query speedup MV size
Baseline (no MV) nothing n/a 1x n/a
Tier 1: store_sales agg by customer surrogate key aggregate of all sales per customer broadest: any per-customer aggregation ~5x faster 0.07% of store_sales for TPC-DS 3 TB
Tier 2: store_sales x customer agg, one filter pre-applied join + aggregate, preferred customers only medium: any country filter, preferred customers ~10x faster 0.04% of store_sales for TPC-DS 3 TB
Tier 3: entire query body verbatim (exact-match) exact ranked output of this query narrowest: only this exact query shape 20x+ faster negligible (only 100 rows)

Performance measured on TPC-DS 3 TB. Speedup is the ratio of baseline execution time to MV-accelerated execution time. Results might vary based on data characteristics, cluster size, and query complexity.

In addition, MVs incur additional cost. Each one runs a query against your source tables once and stores the result. The more pre-computation it does (joining more tables, applying more filters), the more time it takes.

The following chart plots per-query speedup and creation time for the three tiers in our testing on TPC-DS 3 TB. Per-query speedup rises steadily, from about 5x at Tier 1 to over 20x at Tier 3. Creation time doesn’t follow the same pattern: it peaks at Tier 2. Tier 2 pre-joins and aggregates all preferred customers across every country, so it materializes the most data work. Tier 3 applies both filters, so it processes far fewer rows and costs less to create.

Chart comparing three materialized view designs. In our testing with TPC-DS 3 TB, we observed per-query speedup rises from about 5x (Tier 1) to over 20x (Tier 3), while creation time peaks at Tier 2, which materializes the most data work. Stacked bars show creation time split into catalog setup, data work, and commit.

Figure 1: Per-query speedup and creation time across the three materialized view tiers, measured on TPC-DS 3 TB

Start by identifying one expensive query that runs repeatedly with stable filters. It is likely a good candidate for an exact-match MV.

Validating automatic query rewrite

To confirm that your query benefited from automatic rewrite:

  1. Query plan inspection: Check the query’s optimized logical plan or physical plan for a leaf scan node referencing the MV (for example, BatchScan mv_catalog.mv_db.your_mv_name). If the MV appears as a scan source, rewrite succeeded.
  2. Log confirmation (Amazon EMR 7.14.0+): Look for INFO-level log entries such as AQMV outcome: rewritten=true, mvs=[mv_name], duration=12ms.
  3. No-rewrite diagnostics (Amazon EMR 7.14.0+): If rewrite didn’t occur, check the MVRewriteMetricsEvent in the Apache Spark Event Log for the specific reason the optimizer skipped the MV.

If you have set spark.sql.optimizer.answerQueriesWithMVs.enabled=true but your query still runs against the base tables, check the following common causes:

  1. Write commands block rewrite by default. INSERT and MERGE statements don’t trigger rewrite. Set spark.sql.optimizer.answerQueriesWithMVs.commandBlockingEnabled=false to turn on rewrite within write command subqueries.
  2. The MV is stale. Rewrite skips the MV when one or more source tables have changed since its last refresh. Wait for the next scheduled refresh, or force an immediate refresh with REFRESH MATERIALIZED VIEW <mv_name>.
  3. Heuristic candidate filtering. The optimizer uses heuristic checks to narrow the set of MV candidates before attempting a full match. In some cases, an MV that could benefit the query might be filtered out early by these heuristics.
  4. Spark version mismatch (Amazon EMR 7.13.0+). Automatic query rewrite skips MVs whose stored IMV_sparkVersion does not match the cluster’s current Apache Spark version. To bypass this check, set spark.sql.materializedView.sparkVersionCompatibilityCheck.enabled=false.
  5. MV metadata cache not loaded. The metadata cache loads lazily during optimization of the first rewritable query in a Spark session. If your critical query fires before the cache is warm, the MV will not be available. Run a small warm-up query (for example, SELECT 1 FROM <some_iceberg_table>) at session start to pay this cost off the critical path.
  6. MV metadata cache memory limit reached. If the cache was disabled or stopped loading MVs because of reaching its memory limit, increase spark.driver.memory.
  7. Too many tables in configured catalogs. If there are many tables or MVs in the configured catalogs, the cache might not finish loading before your query starts. Place MVs in a dedicated catalog, add it to spark.sql.materializedViews.additionalCatalogs, and set spark.sql.materializedViews.scanCurrentCatalog=false to skip scanning the current catalog.
  8. Parquet base tables have additional limitations and configuration requirements. For automatic query rewrite with Parquet base tables, set spark.sql.materializedView.v1SourceTables.enabled=true and spark.sql.materializedView.v1ETagVersioning.enabled=true. Without ETag versioning, Spark can’t determine a usable source-table version and skips the MV. Partitioned Parquet base tables are also subject to additional validation limits.

Performance considerations

Turning on automatic query rewrite has overhead: it introduces trade-offs that might affect some queries negatively:

  1. Optimization overhead. Enabling rewrite adds processing time during query optimization as the optimizer evaluates MV candidates against the query plan. This overhead applies to every query in the session, including those that ultimately don’t match any MV.
  2. Reduced task parallelism. Reading from an MV instead of the original base table might produce fewer tasks or introduce data skew, depending on the MV’s data layout. This reduces parallelism compared to a direct scan of the larger, more evenly distributed source table.

Conclusion

In this post, we showed how automatic query rewrite can accelerate your existing Apache Spark workloads. It uses Apache Iceberg materialized views in the AWS Glue Data Catalog, without changing a single line of SQL. By storing precomputed results as managed Apache Iceberg tables, the AWS Glue Data Catalog lets the Apache Spark optimizer transparently substitute matching query plans. You get the performance benefit of pre-aggregation without the application-level rewiring. BI dashboards, ISV-generated reports, and legacy pipelines all benefit the moment a matching MV exists.

We walked through three MV designs for the same analytical query, each striking a different balance between per-query speedup, storage footprint, and reuse across your workload. As the trade-off table shows, our testing found that a narrow, exact-match MV delivered 20x+ acceleration for a single query shape. A broader pre-aggregate served an entire family of queries at a more modest ~5x gain. The right choice depends on how many queries share the same join-and-aggregate pattern and how frequently your source data changes.

To get started:

  1. Launch an Amazon EMR 7.12.0+ cluster or an AWS Glue 5.1+ job.
  2. Create an MV over your most expensive repeating query using CREATE MATERIALIZED VIEW in the AWS Glue Data Catalog.
  3. Turn on automatic query rewrite by setting spark.sql.optimizer.answerQueriesWithMVs.enabled=true in your Spark session configuration.
  4. Verify the rewrite by inspecting the optimized query plan for an MV scan node, or by checking INFO-level logs on Amazon EMR 7.14.0+.

Queries with multi-table joins, heavy aggregations, or window functions over large fact tables are strong initial candidates. Start with one high-cost, frequently executed query. Validate the speedup, then expand to broader MVs as you identify shared patterns across your workload.

Special thanks to everyone who contributed to the automatic query rewrite feature and this blog: Andre Hernich, Leon Lin, Yiyang Chen, Geeta Krishna Panda, Ashok Chintalapati, Muhammad Malik, Rishabh Bhatia, and Giovanni Fumarola.

References

For more detail, see the following resources:


About the authors

Yuzhou Sun

Yuzhou Sun

Yuzhou is a software development engineer for Open Data Analytics Engines at Amazon Web Services.

Srishti Mittal

Srishti Mittal

Srishti is a product manager for Open Data Analytics Engines at Amazon Web Services.

Kinshuk Pahare

Kinshuk Pahare

Kinshuk serves as Head of Product for Analytics Engines at AWS, where he leads the product teams responsible for Amazon Redshift, AWS Glue, Amazon EMR, and Amazon Athena. With over six years at AWS, he brings deep expertise in building and scaling cloud-native analytics platforms that help organizations unlock the value of their data at any scale.

Henry Laih

Henry Laih

Henry is a software development engineer for Open Data Analytics Engines at Amazon Web Services.

Srikanth Kandula

Srikanth Kandula

Srikanth is an engineer who works in analytics and distributed systems at Amazon Web Services.

Shahryar Baki

Shahryar Baki

Shahryar is a software development engineer for Open Data Analytics Engines at Amazon Web Services.

Testing application resilience with Amazon SQS and AWS Fault Injection Service

Post Syndicated from Richard Whitworth original https://aws.amazon.com/blogs/architecture/testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service/

When your application can no longer send or receive messages through an Amazon Simple Queue Service (Amazon SQS) queue, downstream processing can stall. The cause might be a misconfigured Identity and Access Management (IAM) policy, a network partition, a bad deployment, or a transient service event. Your application sees much the same thing regardless: SQS operations start failing. How your services handle those failures (failing fast on what won’t succeed, opening circuit breakers, buffering on the producer side) can be the difference between a brief disruption and a cascading outage.

If you’ve never tested those mechanisms under failure, you’re relying on assumptions. With AWS Fault Injection Service (AWS FIS), you can find out first. The goal isn’t to verify that SQS works, it’s to learn what your application does when SQS operations fail, and whether you’d notice. A resilience experiment tests your recovery mechanisms and your observability at once.

In this post, you’ll learn how to:

  • Structure a resilience experiment with a clear, measurable hypothesis and success criteria.
  • Use AWS FIS and AWS Systems Manager (SSM) Automation to simulate progressive access disruption to SQS queues.
  • Interpret Amazon CloudWatch metrics to determine whether your resilience mechanisms are working, distinguishing producer-side from consumer-side behavior.
  • Identify and fix gaps in your application’s failure handling.

Solution overview

In this experiment, you block your application’s access to SQS with a scoped deny resource policy, then restore access and observe recovery. The policy you apply rejects the data-plane operations your application depends on (sending, receiving, deleting, and changing message visibility, plus purging) while leaving queue management untouched. Disruption duration increases across four phases to surface different classes of failure. To learn more, see Control planes and data planes.

Important: Don’t deny sqs:*. In IAM policy evaluation, an explicit Deny overrides every Allow, including the automation’s own permission to remove the policy later. A deny covering sqs:SetQueueAttributes, sqs:AddPermission, and sqs:RemovePermission can lock the queue so that even the role that applied it can’t clean it up. Scope the deny to data-plane actions only. See Configuring the experiment for the safe policy shape, SQS troubleshooting: access denied, and this re:Post article on deny-policy lockout.

You can find the fault-injection code (the SSM Automation document, the FIS experiment template, and example IAM policies for both roles) in the FIS template library on GitHub.

Progressive experiment phases

Short disruptions can reveal whether your failure-handling mechanisms activate. Longer ones can expose systemic issues that might only appear under sustained failure.

Note: This experiment tests your application’s resilience patterns, not SQS itself. The scoped deny simulates what your application would experience during a network partition, a permission change, or another access disruption.

You can watch two distinct failure surfaces at once:

Producer side: The component that calls SendMessage. When sends are denied, you’re testing how the producer handles failed enqueues: does it fail fast, open a circuit breaker, buffer locally, or drop messages?

Consumer side: The component that calls ReceiveMessage and DeleteMessage. When receives are denied, you’re testing backlog growth during the outage and, on recovery, redelivery and whether a consumer working through the accumulated backlog keeps up or starts pushing messages toward the DLQ.

To isolate a consumer outage instead, where a bad deployment or scaling issue stops your consumers while producers keep sending, deny only sqs:ReceiveMessage and sqs:DeleteMessage. This can be done by editing the SSM Automation document.

Producer service sends to an SQS queue, a consumer service reads from it, a dead-letter queue attaches to the source queue, and CloudWatch collects metrics

Architecture diagram: producer service → SQS queue → consumer service, with a dead letter queue attached to the source queue and CloudWatch collecting metrics from the producer, the consumer, and both queues.

A note on the example workload. The behaviors here (circuit breakers, local buffering, thread pools) assume long-running producer and consumer services rather than short-lived Lambda invocations.

Define your hypothesis

Start with one question: can you reason about what your system should do when SQS access disappears? That question, not whether this is your first experiment, determines what kind of hypothesis you write.

If you can, state the expectation and the metrics you’ll judge it by: When our application loses access to SQS for [duration], we expect [specific, observable behavior]. Our system will [recovery expectation] within [time] of access being restored, as measured by [metric(s)].

A team with resilience patterns already in place might write: When our order processing application loses access to SQS for 5 minutes, we expect the producer to open its circuit breaker within 30 seconds, fail fast, and buffer messages in local durable storage rather than dropping them. On recovery it will replay the buffer and return to normal processing rates within 2 minutes, as measured by NumberOfMessagesSent returning to baseline and ApproximateNumberOfMessagesVisible draining to near zero within 15 minutes.

If this is your first test or this failure mode has never been exercised, that isn’t a prerequisite. You don’t necessarily need to read the code and settings first, though a basic understanding of the implementation and its normal load helps you set guardrails that bound the test’s impact. Frame the hypothesis as discovery, stating what you’ll observe instead of what you predict:

Our order processing application has never been tested under SQS access loss. We’ll block access for 2 minutes and observe how the producer handles failed sends and whether the consumer recovers unaided, as measured by NumberOfMessagesSent, ApproximateNumberOfMessagesVisible, ApproximateAgeOfOldestMessage, and application error rates.

Either way, write it down before you proceed. The gaps between what you wrote and what happens are where your system needs work.

Prerequisites

The GitHub repo ships working examples. The following bullets note which file to start from. You’ll need:

  • An instrumented producer and consumer: the application under test. This is the one prerequisite with no example in the repo. The library ships the fault injection, not the workload. The observation tables in this post assume your application emits circuit-breaker state, failed-send and dropped-message counters, fallback-store writes, and duplicate-processing metrics. Without that instrumentation you’ll watch the queue metrics move and learn little about your application.
  • An IAM role for AWS FIS, trusted by fis.amazonaws.com and able to run the SSM Automation document (ssm:StartAutomationExecution and related, plus iam:PassRole). The repo provides both pieces: sqs-queue-impairment-tag-based-fis-role-iam-policy.json for the permissions and fis-iam-trust-relationship.json for the trust policy. Add the Amazon CloudWatch Logs permissions only if you enable experiment logging. See Logging for AWS FIS.
  • An IAM role for the SSM Automation document, able to read and modify the target queues’ policies (sqs:GetQueueAttributes, sqs:SetQueueAttributes, sqs:ListQueues, sqs:ListQueueTags). Start from sqs-queue-impairment-tag-based-ssm-automation-role-iam-policy.json and ssm-iam-trust-relationship.json in the repo. The example policy conditions the write on aws:ResourceTag/FIS-Ready, which helps prevent the automation from touching untagged queues. Keep that condition.
  • SQS queues tagged FIS-Ready: True. This scopes which queues the automation targets. Tag only non-production queues or use planned test windows.
  • A CloudWatch dashboard and alarms combining those application metrics with the queue metrics across the producer, consumer, and queue (see Monitoring strategy). The repo’s README includes an example put-metric-alarm command for a customer-impact alarm you can adapt as your stop condition.
  • A documented rollback plan in case the automation can’t remove the deny policy (if a deny ever locks out queue management, see the re:Post article on deny-policy lockout).

Important: Run these experiments in a non-production environment first. In production, confirm you have change management approvals.

Configuring the experiment

AWS Systems Manager Automation applies and removes the deny policy; AWS FIS orchestrates the sequence.

Systems Manager Automation document

The SSM Automation document follows four steps:

  1. getTargetQueues: finds SQS queues tagged with FIS-Ready: True. It calls ListQueues once, which returns at most 1,000 queue URLs, so in an account with more queues than that, add pagination or a QueueNamePrefix filter before you rely on it to find every tagged queue.
  2. applyDenyAllPolicyToQueues: adds a scoped deny statement to each queue’s resource policy. Deny only the data-plane actions your application uses, never the management actions, so the automation can remove its own statement during cleanup. If you adapt the automation, consider adding a validation step that refuses to apply any deny covering management actions. A lockout would then require changing both the policy and the validation.

Tip: You can make the deny self-expiring by adding a DateLessThan condition on aws:CurrentTime to the statement, so the deny stops applying at a set time even if the cleanup step never runs. See IAM condition operators for date and time.

  1. waitForDuration: sleeps for the specified impairment duration (ISO 8601 format, for example PT2M).
  2. removeDenyAllPolicyFromQueues: removes the FISTemporaryDeny statement, restoring normal access. The document routes onFailure and onCancel to this step so that an aborted run attempts to clean up, and the step raises if it can’t restore a policy rather than reporting success.

Choose your blast radius with Principal. "Principal": "*" denies the data plane to every caller: the application under test, but also any admin, canary, or other consumer of that queue. That faithfully simulates a service partition, but on a shared queue it impairs more than your app. To impair only the application, the more common “my app lost access” case, scope the deny to its IAM role:

"Principal": { "AWS": "arn:aws:iam::<account-id>:role/<application-role>" }

A role arn matches all sessions of that role, catching the app’s calls without applying the deny to other callers. On a shared queue, a principal-scoped deny also changes what you measure: queue-level metrics blend impaired and healthy traffic, so lean on your application’s client-side metrics and read recovery as a return to pre-event levels rather than to zero and back. Set the automation’s optional targetPrincipalArn parameter to scope the deny to one principal, or leave it empty to deny all. The rest of this post assumes the full-queue deny ("Principal": "*").

FIS experiment template

The FIS template chains the four impairment phases with recovery periods between each, calling the SSM Automation document with an increasing duration; startAfter fields enforce sequential execution. The escalation is the point. You watch cause and effect at increasing severity:

Phase Duration What this duration tends to surface
Impair 1 2 minutes Fail-fast behavior and circuit-breaker activation
Recover 3 minutes Buffered messages replay. Metrics return to baseline
Impair 2 5 minutes Backlog accumulation as the queue fills undrained
Recover 3 minutes Backlog burndown
Impair 3 7 minutes Thread-pool and memory pressure from sustained failure
Recover 2 minutes Recovery under a larger backlog. Whether the consumer keeps up
Impair 4 15 minutes Systemic limits under prolonged loss of access
The FIS experiment template with four impairment actions of 2, 5, 7, and 15 minutes chained by recovery waits

Figure: The FIS experiment template: four impairment actions (2, 5, 7, and 15 minutes) chained with recovery waits between them.

Stop conditions. A stop condition halts the experiment automatically if a specified CloudWatch alarm fires, an essential control for an escalating experiment. A triggered stop condition also unwinds what it can: FIS cancels the run, and the automation’s onCancel step removes the deny, restoring access. That rollback is a property of this experiment’s design, not of stop conditions in general: an action like EC2 instance termination does not support rollback, so check each action’s rollback behavior before relying on a stop condition to help limit damage. The library template ships with "stopConditions": [{"source":"none"}], because the right alarm depends on health signals the template can’t assume.

Metric choice matters: alarming on a queue metric like ApproximateAgeOfOldestMessage or NumberOfMessagesSent would be incorrect, as those are supposed to move during impairment. So, the alarm would trip in the first 2-minute phase and abort the run before the longer phases surface anything interesting. You’d be alarming on the effect you’re injecting.

Instead, tie the stop condition to a signal that should stay healthy if your resilience mechanisms are working. This would reflect real customer impact. If that signal degrades more than you’ll tolerate (error rates that don’t recover within the 2 minutes your hypothesis allows), your resilience has already failed and continuing risks further customer impact. Some metrics to consider alarming on:

  • An application error rate or transaction-success metric (a custom CloudWatch metric your app emits, for example failed orders per minute), the most direct measure of customer impact and independent of the SQS metrics you’re perturbing.
  • Load balancer 5xx count or target response time (for example HTTPCode_Target_5XX_Count on an Application Load Balancer), a good proxy when you don’t yet emit a business metric.
  • DLQ depth: ApproximateNumberOfMessagesVisible on the dead-letter queue crossing a threshold, which signals messages are failing permanently rather than only backing up recoverably.

See Stop conditions for AWS FIS for more information.

Deriving the threshold from your hypothesis. The preceding hypothesis expects recovery within 2 minutes of access being restored. That number is also your alarm. If failed orders per minute is your customer-impact metric and its baseline is near zero, set the alarm to failed orders per minute > 10 for 2 consecutive 1-minute periods: long enough that a spike while the circuit breaker opens shouldn’t abort the run, short enough that failing to recover inside your hypothesis window stops it. Design the alarm for how the metric behaves during failure rather than for the test: when the circuit breaker opens, a low-volume custom metric might stop emitting data points entirely. Tighten it as you approach production.

AWS FIS experiment in Stopped state after the customer-impact alarm breached and halted the run

Figure: When the customer-impact alarm breached, AWS FIS halted the experiment automatically (State: Stopped)

Running the experiment

To start the experiment with AWS FIS you can use the console or the AWS CLI:

aws fis start-experiment --experiment-template-id <YOUR_TEMPLATE_ID> --region <YOUR_REGION>

What to observe during impairment

During each phase, SQS operations return AccessDenied errors. Note what that does and doesn’t exercise: a 403 is non-retryable, so this experiment validates that your code recognizes it and stops, not your backoff path. To exercise retries and backoff, inject a retryable fault such as throttling or timeouts. The producer and the consumer fail differently, so watch them separately.

SQS queue access policy showing the scoped FISTemporaryDeny statement blocking SendMessage and ReceiveMessage

Figure: During impairment, the queue’s access policy carries the scoped FISTemporaryDeny statement; SendMessage and ReceiveMessage return AccessDenied while management actions still work.

Producer side (the component calling SendMessage):

Stage Healthy response Unhealthy response Signal to watch
First failed SendMessage Recognizes AccessDenied and fails fast Crashes, hangs, or blocks the calling thread NumberOfMessagesSent drops to ~0. Producer error rate rises
After 3 to 5 consecutive failures Circuit breaker opens. Sheds or buffers load Continues retrying indefinitely Circuit-breaker state metric. Producer CPU / threads / connections
Send gives up (non-retryable error, or retry budget exhausted) Fails fast and persists the payload to durable fallback storage, or alerts, does not silently drop Drops the message silently (permanent loss) Producer “failed send / dropped” counter. Fallback-store writes
Application state Stays responsive. Degrades gracefully Returns 500s to callers. Unbounded in-memory queueing Producer health checks, request latency
Resource usage Bounded by backoff and circuit breaker CPU/memory/connections climb (tight retry loops) Producer CPU, memory, connection-pool usage

Note: a producer that gives up on a send does not route anything to the DLQ.

Consumer side (the component calling ReceiveMessage / DeleteMessage):

Stage Healthy response Unhealthy response Signal to watch
First failed ReceiveMessage / DeleteMessage Backs off its poll loop rather than hammering. Any in-flight message returns to the queue after the visibility timeout Crashes or hangs the consumer loop NumberOfMessagesReceived / NumberOfMessagesDeleted drop
Backlog accumulates (consumers can’t drain) Backlog alarm fires. Scaling responds if keyed to queue depth (for example, backlog per worker) Backlog grows unbounded. Consumers idle-loop ApproximateNumberOfMessagesVisible stops draining (goes flat or climbs); ApproximateAgeOfOldestMessage climbs
Application state Idempotent processing. Safe to retry Duplicate side effects on redelivery Downstream idempotency / duplicate-write metrics
Resource usage Bounded by visibility timeout and backoff In-flight messages pile up. Consumer saturation ApproximateNumberOfMessagesNotVisible. Consumer CPU/memory

Don’t expect the DLQ to fill during impairment. Redrive is driven by maxReceiveCount: a message moves to the DLQ only after a consumer has received it that many times without deleting it. With ReceiveMessage denied, nothing is delivered, the receive count doesn’t increment, and nothing redrives. The DLQ depends on the very call that’s blocked, so it’s something to watch for during recovery, not during the outage.

Key CloudWatch metrics, and how to read them:

  • NumberOfMessagesSent: drops to zero when the deny policy takes effect and producers can no longer enqueue.
NumberOfMessagesSent dropping to zero during each impairment window and spiking on recovery, ending at the stop-condition halt

Figure: NumberOfMessagesSent drops to zero during every impairment window (red) and spikes on recovery (green) as buffered messages replay. The final phase ends at the stop-condition halt (orange).

  • ApproximateNumberOfMessagesVisible: the current backlog of messages available for retrieval. During impairment this often stops changing, a signal that tells you something is wrong precisely because it goes flat (nothing is being sent or drained).
  • ApproximateAgeOfOldestMessage: increases as unprocessed messages age, but only if the queue already held a message when the deny took effect. On an empty queue it won’t climb, which is why you read it alongside the visible-message count.
ApproximateAgeOfOldestMessage climbing while the visible backlog stays undrained during the 15-minute phase, then both collapsing on recovery

Figure: Consumer-side impact: ApproximateAgeOfOldestMessage climbs while the visible backlog sits undrained during the 15-minute phase, then both collapse the moment access is restored.

Application error rate: spikes initially, then stabilizes if circuit breakers engage.

Producer circuit breaker opening within seconds of each impairment and closing on recovery as a square wave

Figure: The producer’s circuit breaker opens (1) within seconds of each impairment and closes (0) on recovery, a clean square wave that lags each fault window slightly because it opens only after a few sustained failures.

Count-based metrics (NumberOfMessagesSent/Received/Deleted) reflect system-level activity and can include retries and duplicates, so treat them as trend indicators rather than exact unique-message counts.

What to observe during recovery

When the deny policy is removed, the producer and consumer recover on different timelines.

Producer side:

What to observe Healthy response Unhealthy response
Send resumes NumberOfMessagesSent climbs back to baseline. Circuit breaker half-opens, then closes within ~30 seconds Circuit breaker stays open (stale failure state). Manual restart needed
Buffered / fallback payloads Replayed from durable fallback storage and re-sent idempotently Lost permanently (if silently dropped during impairment)
Producer buffering to durable fallback storage during impairment and replaying the buffer on recovery

Figure: The producer buffers to durable fallback storage during impairment (no dropped messages) and replays the buffer on recovery: send success, buffered writes, and replays over the run.

Consumer side:

What to observe Healthy response Unhealthy response
Receive / delete resumes NumberOfMessagesReceived / NumberOfMessagesDeleted recover Consumers stay wedged. No auto-recovery
Backlog burndown ApproximateNumberOfMessagesVisible drains steadily; ApproximateAgeOfOldestMessage falls Drain stalls: the rate spikes, then drops to zero and stays there (consumer overwhelmed or stuck)
DLQ contents Genuinely-poison messages redriven and reprocessed in controlled batches within the DLQ retention period Reprocessed all at once (overwhelming downstream), or left to age out of the DLQ and be deleted

Recovery is when the DLQ can move. A consumer overwhelmed by the accumulated backlog can re-fail messages and push some to the DLQ. If healthy messages land there, your maxReceiveCount is too low or your consumer isn’t keeping up.

Analyzing results

After the experiment completes, compare what happened against your hypothesis. Focus on these questions:

  • Did your circuit breakers activate? Measure from the first AccessDenied to when your application stopped attempting SQS operations. Over your target (typically 30 seconds) means your detection threshold is too high.
  • Did your system preserve messages? Reconcile attempted sends against messages processed after recovery, plus the DLQ and producer-side fallback storage. If the numbers don’t add up you have message loss, and the gap tells you which side lost them.
  • Are the recovered messages still worth processing? Preservation and relevance are different questions. After a long outage, some buffered sends and backlogged messages represent requests the client has already given up on, and processing them spends recovery capacity acting on stale intent. Compare each message’s timestamp to the current time as you consume it, and drop or sideline anything no longer actionable, deliberately rather than by letting it age out. See REL05-BP04: Fail fast and limit queues.
  • How did recovery behave? Look at ApproximateNumberOfMessagesVisible after each recovery period. A healthy system drains steadily. If the drain stalls (the rate spikes, then drops to zero and stays there), your consumer is overwhelmed or stuck.
  • Did longer disruptions reveal new failure modes? Compare the 2-minute phase against the 15-minute one. What tends to surface only under sustained failure:
    • Thread pool exhaustion from accumulated retry threads.
    • Memory pressure from buffered messages.
    • Connection pool starvation.
    • DLQ messages aging out: messages that sit in the DLQ longer than its retention period are deleted (see Best practices).

Results that match your hypothesis are evidence your resilience mechanisms work. Results that don’t are your work list.

Best practices

The sections below cover the resilience patterns that turn the gaps this experiment surfaces into fixes: retry logic, circuit breakers, dead-letter queues, and monitoring.

Retry logic with exponential backoff

Don’t retry everything. Retry only errors that might succeed on a repeat, such as throttling, timeouts, and transient 5xxs, and fail fast on non-retryable ones like the AccessDenied (403) this experiment injects. For the errors worth retrying, use exponential backoff with jitter: each failure increases the wait exponentially (1s, 2s, 4s, 8s, and so on) with a random offset that prevents producers who failed together from retrying together and spiking a recovering dependency.

The AWS SDKs have configurable retry behavior built in, so configure it rather than rolling your own. See Timeouts, retries, and backoff with jitter.

Circuit breakers

A circuit breaker stops attempting operations after a threshold of consecutive failures, then lets a single test call through after a recovery timeout. That saves resources on calls that are likely to fail and gives the dependency room to recover. Choose the open-state behavior deliberately: shedding or buffering load is safer than silently switching to an alternate path, because fallback paths are exercised only during failures and tend to fail with them. See Using load shedding to avoid overload and Avoiding fallback in distributed systems.

Dead letter queues

Configure a DLQ for every queue. It’s a consumer-side safety net for poison messages, not a producer overflow buffer. Set maxReceiveCount to the number of processing attempts that make sense for your workload (typically 3 to 5). Because redelivery is what feeds a DLQ, every consumer must tolerate seeing a message twice. See Making retries safe with idempotent APIs. For what a large post-recovery backlog can do, see Avoiding insurmountable queue backlogs.

A DLQ has no depth limit. The constraint is the retention period, after which SQS deletes the message (default 4 days, maximum 14). Set the DLQ’s retention longer than the source queue’s to provide more investigation time before messages are deleted. For standard queues, note that the retention clock runs from the original enqueue time and does not reset on the move to the DLQ, so time in the source queue counts against it. (FIFO queues do reset it.) After the experiment, confirm messages were preserved rather than aged out.

Monitoring strategy

For what to emit and at what granularity, see Instrumenting distributed systems for operational visibility. Build a CloudWatch dashboard combining these across the producer, the consumer, and the queue:

Queue-level metrics: NumberOfMessagesSent, NumberOfMessagesReceived, NumberOfMessagesDeleted, ApproximateNumberOfMessagesVisible, ApproximateNumberOfMessagesNotVisible, and ApproximateAgeOfOldestMessage, read as described in what to observe during impairment.

Application-level metrics:

  • Error rates by type (distinguish AccessDenied from other failures), tagged by producer vs. consumer.
  • Circuit breaker state changes (open / closed / half-open transitions).
  • DLQ message count.
  • End-to-end message processing latency.

Alarm on ApproximateAgeOfOldestMessage exceeding your SLA threshold as a production alert, but not as an experiment stop condition, since the metric is supposed to rise during impairment. Use a customer-impact signal there instead (see Stop conditions).

Clean up your environment

  • Verify the deny policy is gone. Check each queue’s access policy on the console or run aws sqs get-queue-attributes --queue-url <URL> --attribute-names Policy. If FISTemporaryDeny is still there, retrieve the policy, delete the statement, and reapply with aws sqs set-queue-attributes.
  • Process messages that landed in your DLQs during the experiment.
  • Review CloudWatch metrics to confirm your queues have returned to normal operation.
  • Document your findings: What matched your hypothesis, what didn’t, and what you’re fixing.

Expand your resilience testing

Once the basics hold, extend the experiment:

Partial failure: Impair only a subset of your queues to test whether your application handles mixed healthy/unhealthy dependencies.

Note: Don’t run two impairment experiments against the same queue concurrently. The automation reads the policy, modifies it, and writes it back. Concurrent runs can overwrite each other and leave a stale deny behind. Target distinct queues, or run them in sequence.

Consumer-side only: Block only ReceiveMessage and DeleteMessage while allowing SendMessage, to simulate a consumer outage while producers keep filling the queue (the most common real-world scenario).

Combine with other failures: Run the SQS experiment alongside EC2 instance termination or network latency injection to test compound failure scenarios.

Explore AWS Resilience Hub: Use AWS Resilience Hub to assess your application’s resilience posture and get recommendations for improvement.

Using FIS scenarios

A scenario is an AWS-authored template bundling the actions, targets, and duration for a recognizable event, so you start from a reviewed definition instead of assembling actions yourself. While AWS provides multiple scenarios in the library, here are two that are a good place to start.

AZ Availability: Power Interruption induces the symptoms of losing power in one Availability Zone: zonal EC2, ECS, and EKS compute stops, new launches in that AZ fail, and subnet connectivity is lost. It’s the sharper test of the queue-based decoupling this post exercises, because producers and consumers lose capacity while the queue itself is not targeted. You learn whether surviving consumers absorb the backlog, whether Auto Scaling replaces capacity in the remaining AZs rather than retrying in the impaired one, and whether the backlog drains inside your hypothesis window. It defaults to 30 minutes of impairment plus 30 of recovery, twice this post’s longest phase.

AZ: Application Slowdown introduces additional latency between resources within a single Availability Zone (AZ). This latency creates many of the symptoms of an application slowdown, a partial disruption, sometimes known as a gray failure. It adds latency to network flows between target resources. Network flows represent the traffic between computing resources: the data packets carrying requests, responses, and other communications between your servers, containers, and services. The scenario can help to validate observability setups, tune alarm thresholds, discover application sensitivity to slowdowns, and practice critical operational decisions like AZ evacuation.

Scenarios carry the same obligations: write the hypothesis first and set the stop condition on a customer-impact metric rather than one the scenario is designed to move. Your derived threshold works unchanged. Copy a scenario into your own template to narrow the targets or change the duration. See Working with the AWS FIS scenario library.

Conclusion

In this post, you learned how to discover what your application does when SQS operations fail, and whether you’d notice. Every gap between your hypothesis and the results is an opportunity to improve your system’s resilience and its observability.

Start with the 2-minute phase in a non-production environment. Fix what breaks. Then run the full sequence and keep running it as the application evolves. Each phase of growth brings failure modes you might only find under load.

For more information, see:


About the authors

Validating multi-Region DR for Terraform Enterprise with AWS FIS

Post Syndicated from Frenil Randeria original https://aws.amazon.com/blogs/architecture/validating-multi-region-dr-for-terraform-enterprise-with-aws-fis/

In October 2025, Athenahealth, a major North American Electronic Health Record (EHR) provider, discovered a gap. An AWS regional service event in us-east-1 made their single-Region HashiCorp Terraform Enterprise (TFE) deployment inaccessible to their developers. This post shares the architecture, the AWS Fault Injection Service (AWS FIS) validation approach, HashiCorp best practices, and lessons learned from the collaboration between AWS and HashiCorp. Together, these help increase resiliency and verify that the customer’s critical workloads remain active during regional service events.

Currently, Terraform Enterprise (TFE) deployments are only supported within a single AWS Region. This means that, for TFE customers without a well-tested DR plan, a regional service event can block your engineering teams from deploying, modifying, or recovering infrastructure. A multi-Region disaster recovery (DR) strategy addresses this risk. The architecture in this post is a customer-operated DR pattern: HashiCorp supports TFE within a single Region and the HVD module targets single-Region deployments, so the multi-Region failover described here is designed, operated, and tested by the customer rather than provided as a supported product configuration. That strategy only works if you validate your regional failover workflow before you need it. Reacting during an event that impairs your primary Region costs developer productivity and business continuity. AWS FIS exposes hidden dependencies and configuration issues by injecting real failures into your AWS environment.

The following sections walk you through how to design three-phase AWS FIS experiments for TFE, expose hidden dependencies in failover automation, and validate both failover and failback for a multi-Region TFE deployment. You can help prevent extended downtime that impacts your infrastructure deployment capabilities and achieve 12-14 minute recovery times.

Prerequisites

To follow the validation approach in this post, you should have the following:

Starting architecture

If you’re running TFE in a single Region within AWS, this section describes the starting point. Athenahealth’s deployment used HashiCorp’s Terraform Enterprise Validated Design (HVD) module with the following components:

  • Amazon Elastic Compute Cloud (Amazon EC2) instances running TFE application servers.

  • Amazon Aurora PostgreSQL-Compatible Edition for application state.

  • Amazon Simple Storage Service (Amazon S3) for Terraform workspace state files.

Athenahealth hosted these components in the us-east-1 Region. The deployment provided Availability Zone-level resilience but lacked regional failover capabilities. The October 2025 event highlighted what was missing: no cross-Region database replication, no secondary Region compute capacity, no Terraform state file backup outside us-east-1, and DNS pointing exclusively to the primary Region.

Following the October 2025 event, Athenahealth engaged both their AWS and HashiCorp account teams for guidance on protecting not only TFE, but other critical workloads as well. The three organizations worked as a single team to design and implement a multi-Region DR strategy for the TFE environment. By combining the AWS Well-Architected Framework guidance regarding operational excellence and reliability, along with HashiCorp’s best practices regarding DR strategies using Terraform, the team came up with the multi-Region architecture (Figure 1) that would replace the customer’s current single-Region deployment.

Multi-Region DR solution

With an active-passive multi-Region design across us-east-1 (primary) and us-west-2 (DR), Athenahealth achieved a 12-14 minute Recovery Time Objective (RTO) and less than 1 minute Recovery Point Objective (RPO). In the AWS disaster recovery taxonomy, this is a pilot light strategy: data replicates continuously to the DR Region while compute stays at zero. A warm standby variant (DR minimum capacity of 1) trades higher cost for faster RTO. This section covers the architecture components that make this possible and the four-step failover process you can follow during a regional service event.

Multi-Region active-passive DR architecture for Terraform Enterprise with bidirectional S3 replication and a Route 53 health check.

Figure 1: Multi-Region active-passive DR architecture for Terraform Enterprise on AWS. Note the bidirectional S3 replication arrows between Regions and the Amazon Route 53 health check that determines the active Region.

The example Terraform code used to configure and manage the core architecture components, along with the failover process can be found in this sample GitHub repository. You can use this code to test a similar pattern for your TFE workload.

Core architecture components

You can route traffic to the active Region with Amazon Route 53 DNS alias records pointing to an Elastic Load Balancing (ELB) Network Load Balancer. Alias records for ELB targets use a 60-second time-to-live (TTL), which limits how long DNS resolvers cache the record. Clients begin resolving to the DR Region within about a minute of failover rather than waiting for longer cached entries to expire.

In each Region, an Amazon Virtual Private Cloud (Amazon VPC) spans three Availability Zones, and Amazon EC2 Auto Scaling groups manage the TFE instances. To help minimize cost, Athenahealth scaled DR Region compute to zero during normal operations by setting the Auto Scaling group minimum capacity to 0.

For cross-Region database replication, Athenahealth uses Aurora PostgreSQL-Compatible global databases, which provide sub-second replication lag and managed failover. The primary cluster runs one writer and two readers across three Availability Zones. The secondary cluster maintains an inactive writer that’s ready for promotion.

You can replicate Terraform workspace state files bidirectionally between primary and DR Amazon S3 buckets with S3 cross-Region replication. This design supports failback without data resynchronization.

AWS Secrets Manager and AWS Key Management Service (AWS KMS) provide cross-Region credential and encryption key management. Two TFE-specific dependencies deserve attention when you design for multi-Region. First, the TFE encryption password protects the internal Vault unseal key and root token. DR instances configured with a different value cannot start or decrypt existing data, so verify that this secret is replicated to your DR Region and referenced by your DR launch configuration. Second, if you run TFE in Active/Active mode, external Redis holds the job queue and cache. Account for a Redis equivalent in the DR Region and decide what in-flight job loss is acceptable at failover. Amazon CloudWatch alarms in each Region monitor the TFE instances, Auto Scaling groups, and Aurora clusters in that Region. Detection of a primary Region impairment does not depend on the primary Region itself: the Amazon Route 53 health check shown in Figure 1 probes the TFE endpoint from a globally distributed checker fleet, and alerts publish through Amazon Simple Notification Service (Amazon SNS) topics in both Regions.

Failover process

The architecture uses a four-step failover sequence:

1. Activate DR Auto Scaling group (2-5 minutes). Scale from 0 to 1 instance and validate health checks. TFE exposes a health check endpoint (/_health_check) that returns a 200 OK response when the application is running. The Network Load Balancer target group and the Route 53 health check probe this endpoint to determine instance health.

2. Promote Aurora PostgreSQL-Compatible global database (~1 minute). Promote the DR writer. Complete this step before DNS failover shifts traffic to the DR Region, to help prevent both Regions from accepting writes simultaneously (known as a split-brain scenario in distributed databases).

3. Confirm Amazon Route 53 DNS failover (~60 seconds). The pre-configured Route 53 failover routing policy detects the unhealthy primary endpoint and routes traffic to the DR Region’s Network Load Balancer. This happens in the Route 53 data plane, with no record modifications at failover time.

[Important: Your failover process should not depend on control plane API calls during an event. Modifying Route 53 records to perform failover is a documented anti-pattern because the Route 53 control plane operates from a single Region. Athenahealth avoided this dependency with pre-configured health check-based failover routing. For manually initiated Region switches through a highly available data plane, consider Amazon Application Recovery Controller (ARC) Region switch, which Athenahealth plans to evaluate in a future phase.]

4. Scale out for production load (5-10 minutes). Increase Auto Scaling group capacity while monitoring Amazon CloudWatch.

Note: Run failover scripts from outside the primary Region—for example, from the DR Region, a separate management Region, or a CI/CD system that is not dependent on the primary Region. If your failover automation runs in the primary Region, it might be unreachable during the event you are trying to recover from.

The ordering between Aurora promotion and traffic shift is enforced procedurally rather than by an automated control. The DR Auto Scaling group runs at zero capacity during normal operations, so the DR endpoint cannot pass health checks until an operator executes the runbook, and the runbook sequences promotion ahead of scaling for traffic. For an orchestrated Region switch with explicit sequencing controls, consider Amazon Application Recovery Controller Region switch, which Athenahealth plans to evaluate in a future phase.

Total failover execution time: 12-14 minutes, meeting the established RTO.

Validating with AWS Fault Injection Service

With the multi-Region architecture now in place, we still needed to confirm it would work under real failure conditions. This is where AWS FIS was introduced into the DR workflow. AWS FIS injects controlled failures into your AWS environment that can be used to measure actual recovery times and catch configuration issues before an actual disruption. The following three phases show how Athenahealth validated their architecture, and you can apply the same approach to your TFE deployment as well.

Progressive experiment approach

Rather than testing full regional failover immediately, the team validated resilience in three progressive phases starting with individual compute failures, then database failover, and finally simulated S3 connectivity loss. Each phase built confidence in a specific layer of the architecture before combining them, and each surfaced issues that manual review had missed.

Phase A: Amazon EC2 and Auto Scaling group failure injection

Athenahealth hypothesized that if TFE instances were stopped or Amazon EC2 capacity became unavailable, the Auto Scaling group would launch replacement instances within five minutes without manual intervention. To test this, they ran the following AWS FIS actions: aws:ec2:stop-instances, aws:ec2:asg-insufficient-instance-capacity-error, and Auto Scaling group suspend and resume operations. You can use these same actions to validate your own Auto Scaling group recovery behavior.

The results confirmed the hypothesis. The Auto Scaling group detected failed instances and launched replacements within 2-3 minutes. Network Load Balancer health checks removed failed instances from rotation within 30 seconds.

These experiments also revealed an outdated Amazon Machine Image (AMI) reference in the DR Region’s Auto Scaling group launch template. Athenahealth builds custom AMIs and copies them to the DR Region, but the DR launch template still referenced an older version. This configuration drift only surfaced when AWS FIS forced the Auto Scaling group to launch new instances. If you’re running similar experiments, check the launch template AMI references in both Regions as part of your validation.

AWS FIS console experiments list showing one experiment in the Running state.

Figure 2. FIS experiments list showing experiment EXP5vVgVbYvM7G7CFk in Running state, created July 27, 2026 at 12:29:20 IST.

AWS FIS experiment details page showing the Suspend-ASG action completed and Stop-TFE-Instances running.

Figure 3. FIS experiment details showing template TFE-Primary-Region-Full-Outage, CloudWatch log destination /tfe/lab/fis/logs, Suspend-ASG completed, and Stop-TFE-Instances running.

Amazon EC2 console showing the primary instance in us-east-1 in the stopped state.

Figure 4. Primary EC2 instance in us-east-1 stopped after the FIS stop-instances action.

AWS FIS action summary showing Stop-TFE-Instances completed and Wait-For-Failover-Test running.

Figure 5. FIS action summary showing Stop-TFE-Instances completed at 12:35:17 IST and Wait-For-Failover-Test running.

AWS FIS experiment running during the wait window, with the stop action completed and the resume action pending.

Figure 6. FIS experiment still running during the wait window, with stop action completed and resume action pending.

AWS FIS experiment completed with all four actions showing completed, including Resume-ASG-Launch-via-Automation.

Figure 7. FIS experiment completed at 12:51:00 IST. All four actions show completed, including Resume-ASG-Launch-via-Automation.

Phase B: Aurora PostgreSQL-Compatible database cluster failover

Athenahealth hypothesized that if the Aurora PostgreSQL-Compatible global database cluster failed over, TFE would resume writes within one minute without manual intervention. To test this, they ran the aws:rds:failover-db-cluster AWS FIS action. The results confirmed the hypothesis for the database layer. Aurora promoted the secondary cluster’s writer in 58 seconds. During the promotion, TFE experienced approximately 15 seconds of write unavailability.

The Amazon Relational Database Service (Amazon RDS) Global Endpoint automatically redirected connections to the new writer. Athenahealth also discovered that the application layer did not meet the hypothesis. TFE connection pooling settings caused extended reconnection delays.

They reduced the connection pool timeout from 60 seconds to 10 seconds, which improved recovery time significantly. If you’re running TFE with Aurora PostgreSQL-Compatible, review your connection pool settings as part of your AWS FIS validation.

Note that this experiment exercised the coordinated failover path of Aurora, which requires the primary Region to be reachable to synchronize before promotion. During an actual event impairing the primary Region, you would instead use Aurora Global Database managed failover (the failover-global-cluster command with the --allow-data-loss option) or a manual detach-and-promote. These paths do not wait for replication to synchronize, so promotion timing differs and the RPO is bounded by the replication lag at the time of the event rather than zero. Treat the 58-second promotion and sub-second lag measured here as coordinated-path results, and plan unplanned-path expectations using the Aurora Global Database disaster recovery documentation.

Phase C: Amazon S3 connectivity disruption

Athenahealth hypothesized that if TFE lost connectivity to Amazon S3 in the primary Region, the DR Region bucket would hold current replicated state files without data loss. Testing this required a workaround, because AWS FIS doesn’t provide a direct action to disrupt Amazon S3 access. You can use the aws:network:disrupt-connectivity action instead to inject network ACL rules that block S3 traffic at the subnet level.

The aws:network:disrupt-connectivity action targets subnets, not S3 buckets directly. AWS FIS injects network ACL rules on compute private subnets, which blocks egress traffic to S3 service endpoints and simulates Regional S3 disruption for TFE instances in those subnets.

This experiment validates the S3 consumer, meaning TFE losing access to S3. It does not disrupt S3 cross-Region replication, because service-side replication between buckets does not traverse your subnet network ACLs. To test delayed or paused replication between Regions, use the Cross-Region: Connectivity scenario described in Next steps. Note that this approach assumes your TFE instances reach S3 over an in-VPC path, such as a gateway VPC endpoint, so the injected network ACL rules sit on the egress path to S3.

To configure this AWS FIS experiment, use the following template. Replace <your-tfe-compute-private-subnet-prefix> with your actual subnet name prefix, which you can find in the Amazon VPC console under Subnets.

{
    "actions": {
        "DisruptS3Connectivity": {
            "actionId": "aws:network:disrupt-connectivity",
            "parameters": {
                "duration": "PT10M",
                "scope": "all"
            },
            "targets": {
                "Subnets": "TFE-Compute-Private-Subnets"
            }
        }
    },
    "targets": {
        "TFE-Compute-Private-Subnets": {
            "resourceType": "aws:ec2:subnet",
            "resourceTags": {
                "Name": "<your-tfe-compute-private-subnet-prefix>-*"
            },
            "selectionMode": "ALL"
        }
    }
}

The results confirmed the hypothesis for data durability. TFE detected S3 connectivity loss within 5 seconds. The DR Region S3 bucket contained replicated state files with less than 30 seconds of replication lag, and bidirectional replication prevented state file loss. The experiment also exposed a failure mode Athenahealth had not anticipated: the state file dependency issue detailed in the following section.

Combined experiment: primary Region impairment

After validating each layer individually, the team combined the faults into a single AWS FIS experiment template (shown in Figures 2-7). The experiment suspends the primary Region Auto Scaling group, stops the TFE instances, holds the faults in place during a wait window while the team executed the four-step failover runbook, and then resumes the Auto Scaling group. This end-to-end run validated the complete failover process under simultaneous compute impairment. The combined run surfaced no new failure modes beyond those found in the individual phases, which was itself the confirmation the team wanted.

Measuring recovery times

Across the three AWS FIS experiment phases, Amazon CloudWatch measured the following recovery times:

  • Amazon EC2 failure recovery: 2-3 minutes (automated Auto Scaling group replacement)

  • Aurora PostgreSQL-Compatible failover: 1-2 minutes (managed promotion)

  • Failover execution time: 12-14 minutes (operator-triggered four-step process)

  • Aurora replication lag: less than 1 second (99th percentile)

  • S3 replication lag: less than 30 seconds (99th percentile)

  • Data loss during failover: 0 bytes (across each experiment)

[Note: These measurements reflect controlled testing conditions. Aurora Global Database and S3 cross-Region replication are both asynchronous. During an actual event, writes committed within the replication lag window (sub-second for Aurora, up to 30 seconds for S3) may not yet be available in the DR Region. Plan for near-zero rather than zero data loss when setting RPO expectations.]

  • Failback RTO: approximately 20 minutes (including approximately 5 minutes for Aurora PostgreSQL-Compatible global database re-establishment)

The 12-14 minutes measure failover execution time, from the operator triggering the runbook to full recovery. End-to-end recovery from event onset also includes detection time and the decision to fail over, so plan for a larger overall RTO.

Lesson learned: the state file dependency pitfall

Athenahealth first identified this risk during production failover and failback testing: their automation scripts depended on Terraform S3 state file outputs from both Regions. Subsequent AWS FIS experiments (Phase C) confirmed the severity. When S3 access is lost, those scripts fail entirely.

How the dependency breaks failover

Athenahealth’s failover and failback scripts automate the four-step process described earlier: scaling the Auto Scaling group in the target Region, promoting the Aurora global database writer, and verifying application health. An operator triggers them as part of the manual failover runbook, and they run from outside the primary Region. In their original form, the scripts retrieved infrastructure identifiers such as the Amazon RDS global cluster ID and Auto Scaling group name from state files stored in Amazon S3:

# Failover script excerpt (problematic approach)
# Retrieve RDS Global Cluster ID from primary Region state file
RDS_GLOBAL_CLUSTER_ID=$(terraform output \
    -state=s3://<primary-region-bucket>/terraform.tfstate \
    rds_global_cluster_id)

# Retrieve DR Auto Scaling group name from primary Region state file
DR_ASG_NAME=$(terraform output \
    -state=s3://<primary-region-bucket>/terraform.tfstate \
    dr_asg_name)

# Run Aurora failover
aws rds failover-global-cluster \
    --global-cluster-identifier $RDS_GLOBAL_CLUSTER_ID \
    --region us-west-2

For an unplanned Regional impairment, add the --allow-data-loss option to this command to perform a managed failover instead of a switchover, because a switchover requires the primary Region to be healthy.

During a Regional service impairment, the primary Region’s S3 state file may be unreachable. The same applies in reverse during failback. The failover script tries to read the state file, S3 times out, and the script stops. This creates a circular dependency: you can’t run the failover without access to the infrastructure you’re trying to recover from.

How to remove the dependency

The underlying principle is to remove every recovery dependency on the Region you are recovering from. Failover automation must not read configuration from the control plane or data plane of the impaired Region. Athenahealth implemented this principle by hardcoding infrastructure identifiers directly in their failover scripts. You can obtain these values from your Terraform outputs during normal operations:

# Failover script excerpt (resilient approach)
# Infrastructure identifiers hardcoded, not from dynamic lookups
RDS_GLOBAL_CLUSTER_ID="<your-tfe-global-cluster>"
DR_ASG_NAME="<your-tfe-dr-asg-us-west-2>"
ROUTE53_HOSTED_ZONE_ID="<your-hosted-zone-id>"

# Run Aurora failover with no state file dependency
aws rds failover-global-cluster \
    --global-cluster-identifier $RDS_GLOBAL_CLUSTER_ID \
    --region us-west-2

The trade-off is maintenance: hardcoded values require manual updates when infrastructure changes. Athenahealth addressed this with a CI/CD pipeline that compares hardcoded values against Terraform outputs and alerts on drift.

Hardcoding is one implementation of the principle. A Region-independent configuration source outside the primary Region achieves the same resilience with less drift risk, such as an AWS Systems Manager Parameter Store parameter replicated across Regions, an Amazon DynamoDB global table, or values committed to the repository that holds your failover scripts.

Testing failback

The state file dependency affected both failover and failback. Athenahealth validated failback by running full failover to DR, operating there for 30 minutes, then returning to primary. Bidirectional S3 replication prevented state file loss.

Collaboration model

If you’re planning a multi-Region DR project for TFE, consider a cross-functional approach. Athenahealth’s five-month engagement combined AWS resilience and AWS FIS expertise, HashiCorp TFE architecture knowledge, Terraform DR best practices, and HVD modules. This combination helped the customer reach production-validated DR faster than working independently. You can engage AWS Support or AWS Professional Services for similar guidance.

Conclusion

You can maintain infrastructure deployment capabilities during events that impact a single Region with a validated multi-Region DR architecture for TFE. This architecture achieved a validated RTO of 12-14 minutes and an RPO of less than 1 minute.

Multi-Region DR is not the right choice for every TFE deployment. Athenahealth chose this approach because TFE manages infrastructure for critical healthcare workloads. The October 2025 event showed that losing the ability to deploy during a regional event was a risk the business could not accept. Costs vary based on your configuration, but Athenahealth observed costs approximately 30-40% higher than their single-Region deployment, primarily from Aurora PostgreSQL-Compatible global database replication and S3 cross-Region replication. Weigh this cost against your own RTO and RPO requirements. For less critical workloads, a single-Region deployment with regular backups and a tested restore process may meet your needs.

Key takeaways

  1. Give your infrastructure as code (IaC) tools the same resilience as production workloads. When your TFE deployment becomes unavailable during a Regional service event, you can’t deploy fixes or recover infrastructure.

  2. Validate DR with controlled failure injection. AWS FIS experiments simulating real S3 connectivity loss exposed the state file circular dependency, a failure mode that only surfaces under actual disruption conditions.

  3. Remove recovery dependencies on the Region you are recovering from. Dynamic lookups from state files tie failover to the infrastructure being recovered. Hardcoded identifiers or a Region-independent configuration source both work. Use drift detection to keep values current.

  4. Test failback, not only failover. Without failback validation, you risk getting stuck in the DR Region or causing data loss when returning to primary.

  5. Use subnet-level network disruption to simulate S3 connectivity disruption. The aws:network:disrupt-connectivity action targeting compute subnets simulates Regional S3 connectivity loss, which is the recommended approach because AWS FIS doesn’t offer a direct S3 disruption action.

Next steps

You can implement this solution in your environment with the following steps:

1. Run Phase A AWS FIS experiments on non-production TFE instances to validate Auto Scaling group recovery.

2. Review the HashiCorp’s Terraform Enterprise Validated Design module and DR guidance.

3. Establish your RTO and RPO targets before designing your Aurora replication strategy.

4. Create your first AWS FIS experiment to validate your DR architecture.

After you validate these three phases, extend your testing with additional AWS FIS scenarios. The AZ Availability: Power Interruption scenario validates recovery from the loss of an Availability Zone. The Cross-Region: Connectivity scenario simulates disrupted network connectivity between Regions, including paused S3 replication, which would delay state file replication to your DR Region.

If you need help designing or validating a multi-Region DR strategy, contact AWS Support or AWS Professional Services.

Cleanup

If you deploy this architecture for testing, delete the following resources in both Regions to avoid ongoing charges:

  • Aurora PostgreSQL-Compatible global database clusters.

  • Amazon S3 buckets with cross-Region replication.

  • Amazon EC2 instances in the DR Region Auto Scaling group.

If you used the sample GitHub repo to set up a test multi-Region TFE environment, verify that you also run terraform destroy to avoid any additional charges.

Resources:


About the authors

Build declarative ETL pipelines with AWS Glue 6.0

Post Syndicated from Syed Humair original https://aws.amazon.com/blogs/big-data/build-declarative-etl-pipelines-with-aws-glue-6-0/

Data teams commonly build the extract, transform, and load (ETL) pipelines that turn raw order events into analyst-ready aggregates as a bronze, silver, and gold sequence, the medallion architecture. Bronze holds raw ingested records, silver holds cleaned and validated data, and gold holds the business-level aggregates that analysts query. Today you build this on AWS Glue with an orchestrator such as Amazon Managed Workflows for Apache Airflow (Amazon MWAA) or AWS Step Functions coordinating the stages. Many teams run production pipelines exactly this way. As a pipeline grows, the coordination work grows with it: you wire job dependencies, manage intermediate checkpoints, and add retry logic stage by stage.

AWS Glue 6.0, powered by Apache Spark 4.1, introduces Spark Declarative Pipelines (SDP), which simplifies this further. Instead of orchestrating jobs by hand, you declare what each dataset should contain and let the declarative framework resolve dependencies, manage checkpoints, and orchestrate execution order automatically. The result runs as a single declarative job, with no manual directed acyclic graph (DAG) wiring or imperative orchestration code.

In this post, you build a single AWS Glue 6.0 job that turns raw order records into validated, aggregated, analytics-ready tables through the bronze, silver, and gold sequence. You do this without writing any orchestration logic. This walkthrough uses the AWS Command Line Interface (AWS CLI), and the same operations are available through the AWS SDKs.

Solution overview

You build a single AWS Glue 6.0 job that reads raw order records from a CSV file in Amazon Simple Storage Service (Amazon S3). The job flows them through three declared datasets. These are a bronze materialized view (ingest as-is), a silver materialized view (type, validate, and classify), and a gold SQL materialized view (aggregate by region). With AWS Glue Data Catalog integration turned on, all three land as Data Catalog tables, queryable with standard SQL tooling such as Amazon Athena. SDP resolves the dependency order from the dataset references in your code, so you never orchestrate the steps yourself.

Two ways to build the pipeline

Before you build the pipeline, let’s understand this new way of writing ETL pipelines with a quick comparison of the imperative and declarative approaches.

With the imperative approach, you need three AWS Glue jobs, plus an orchestrator to handle sequencing and error handling. A typical pipeline therefore has two layers: an orchestration layer and the ETL processing layer. The following diagram shows this two-layer imperative pipeline.

Two-layer imperative pipeline: three AWS Glue jobs coordinated by an orchestrator.

Figure 1: The two-layer imperative pipeline, with three AWS Glue jobs coordinated by an orchestrator.

Compared to that, the declarative approach runs as a single ETL job with SDP. The following diagram mirrors the previous one, but here it is a single AWS Glue ETL job instead of three jobs plus an orchestrator.

Declarative pipeline: a single AWS Glue job running the bronze, silver, and gold layers with Spark Declarative Pipelines.

Figure 2: The declarative pipeline, a single AWS Glue job running the bronze, silver, and gold layers with SDP.

The declarative approach reduces more than the number of jobs. It removes the boilerplate that surrounds them. An orchestrator such as Amazon MWAA or AWS Step Functions already handles retries and parallelism, but only at the granularity of a whole job. To get finer control, teams often split a pipeline into several jobs and then hand-wire the dependencies between them. With SDP, you no longer hand-wire a DAG, manage per-stage checkpoints, or split the pipeline into separate jobs for retries and parallelism. SDP derives the dependency graph from your table references and coordinates execution at the level of individual tables. You can still invoke an SDP job from an orchestrator when a broader workflow calls for it, but the pipeline’s internal coordination is no longer code you write and maintain.

SDP separates the what from the how: you declare datasets (the outputs you want), and SDP builds the flows that produce them and runs them as one pipeline, resolving dependencies and execution order automatically.

You declare these abstractions through Python decorators. This post covers three of them, @dp.table, @dp.materialized_view, and @dp.temporary_view, each with its own purpose:

  • @dp.table defines a streaming table, which processes new data incrementally on each run. Typical use cases are raw event ingestion and change data capture (CDC) feeds.
  • @dp.materialized_view defines a materialized view for batch use cases. Today, this dataset type fully recomputes on each run. Common uses include parsing, aggregations, and machine learning (ML) feature engineering.
  • @dp.temporary_view is for temporary computations and aggregations. It’s pipeline-scoped and isn’t persisted outside the pipeline. Use it for enrichment lookups and subqueries.

Streaming tables append only new arrivals. Materialized views fully recompute. This post uses @dp.materialized_view for all three layers to keep the walkthrough focused. In production, you would typically use @dp.table for the bronze layer to process only new files as they arrive rather than re-reading the full source each run.

Running and refreshing the pipeline

When you rerun a pipeline, you don’t always want the same work to happen. Sometimes you only want to confirm the pipeline is well-formed before spending compute. Other times you want to run it but recompute only the datasets that changed rather than the entire graph. SDP handles both cases through two independent controls, and it helps to keep them separate:

  • Execution mode (the spark.glue.sdp.jobMode key) answers run or only validate?
  • Refresh scope (the spark.glue.sdp.runMode key) answers given that I’m running, what do I recompute?

Execution mode. VALIDATE runs the pipeline in dry-run mode: SDP checks the YAML syntax, dependency resolution, and SQL and Python compilation without writing any data. Use it to verify your pipeline is well-formed before committing compute. RUN (the default) executes the pipeline normally, resolving the dependency graph and materializing datasets.

# Dry run: validate the graph, write nothing
aws glue start-job-run \
  --job-name "${JOB_NAME}" \
  --arguments '{"--conf":"spark.glue.sdp.jobMode=VALIDATE"}' \
  --region "${AWS_REGION}"

# Normal execution
aws glue start-job-run \
  --job-name "${JOB_NAME}" \
  --arguments '{"--conf":"spark.glue.sdp.jobMode=RUN"}' \
  --region "${AWS_REGION}"

Refresh scope. By default, a RUN recomputes every materialized view. You can narrow or widen that with spark.glue.sdp.runMode:

  • --refresh <datasets> updates only the named datasets (comma-separated, no spaces).
  • --full-refresh <datasets> resets and recomputes only the named datasets (for streaming tables, this also clears their checkpoints).
  • --full-refresh-all resets and recomputes every dataset in the pipeline.
# Selective refresh of named datasets
aws glue start-job-run \
  --job-name "${JOB_NAME}" \
  --arguments '{"--conf":"spark.glue.sdp.jobMode=RUN --conf spark.glue.sdp.runMode=--refresh silver_orders,gold_sales_summary"}' \
  --region "${AWS_REGION}"

# Full reset and recompute of the entire pipeline
aws glue start-job-run \
  --job-name "${JOB_NAME}" \
  --arguments '{"--conf":"spark.glue.sdp.jobMode=RUN --conf spark.glue.sdp.runMode=--full-refresh-all"}' \
  --region "${AWS_REGION}"

Selective refresh is useful during development, so you can iterate on a single layer without reprocessing the entire graph. Note that --refresh and --full-refresh each take an explicit list of datasets. To reset the whole pipeline, use --full-refresh-all. Because materialized views hold no incremental state, resetting a materialized view and refreshing it both fully recompute it. The reset-versus-refresh distinction matters for streaming tables, where a refresh processes only new data and a reset clears the checkpoint and reprocesses from scratch.

The multiple values are passed as a single --conf argument string ("spark.glue.sdp.jobMode=RUN --conf spark.glue.sdp.runMode=..."). This is the serialization the AWS Glue SDP mode expects for the run.

Materialized views: Batch transforms with automatic dependency resolution

Materialized views recompute their full result set on each run. SDP infers dependencies from table references: in this pipeline, silver_orders references bronze_orders, so SDP runs bronze first, as shown in the following diagram.

Dependency graph showing Spark Declarative Pipelines running the bronze layer before the silver layer.

Figure 3: SDP infers the dependency order from table references and runs bronze before silver.

The core pattern is a decorated function that returns a DataFrame:

@dp.materialized_view(comment="Raw orders loaded from CSV")
def bronze_orders() -> DataFrame:
    return spark.read.schema(ORDERS_SCHEMA).option("header", "true").csv(ORDERS_PATH)

The silver layer references bronze_orders through spark.table("bronze_orders"), with no explicit dependency declaration. SDP builds the DAG by analyzing table references in your code and runs bronze first automatically.

Bronze reads every column as a string by design: the bronze layer preserves raw source data without coercion. Type casting, validation, and filtering happen in the silver layer.

SQL and Python coexistence

SDP supports both Python and SQL definitions in the same pipeline project. A SQL materialized view can reference a Python-defined table directly, for example the gold layer aggregating the silver table:

CREATE MATERIALIZED VIEW gold_sales_summary
COMMENT 'Completed-order metrics by region'
AS
SELECT
  region,
  COUNT(*) AS order_count,
  CAST(ROUND(SUM(amount), 2) AS DECIMAL(10, 2)) AS total_sales,
  CAST(ROUND(AVG(amount), 2) AS DECIMAL(10, 2)) AS average_order_value
FROM silver_orders
GROUP BY region;

In this post, Python files define ingestion and validation logic, and SQL files define reporting views and aggregations. SDP discovers both through the libraries glob pattern in the pipeline specification and resolves the cross-language dependencies automatically. The complete source for all three layers follows in the step-by-step walkthrough.

Build the pipeline: Step by step

The rest of this post is a hands-on walkthrough. You build a single AWS Glue 6.0 job that reads orders.csv and processes it through the bronze, silver, and gold layers. The steps are:

  1. Prerequisites: AWS account, AWS Identity and Access Management (IAM) role, and S3 bucket.
  2. Set up sample data: create orders.csv and upload it to Amazon S3.
  3. Build the pipeline files (the spark-pipeline.yml specification plus the three transformation files).
  4. Package the pipeline into a zip and upload it to Amazon S3.
  5. Create the database: a Data Catalog database with an S3 location.
  6. Configure the job: create the AWS Glue 6.0 job with the SDP flag.
  7. Validate: run in dry-run mode to verify the graph.
  8. Run the pipeline to materialize all datasets.
  9. Query results: inspect the tables with Amazon Athena.
  10. Clean up: delete the resources you created.

Step 1 – Prerequisites

To follow along, you need:

  • An AWS account with access to AWS Glue 6.0.
  • A dedicated IAM role trusted by glue.amazonaws.com (set up in the following section).
  • A private, encrypted Amazon S3 bucket with Block Public Access enabled.
  • The AWS CLI configured with credentials for a non-production account.

IAM role for the pipeline

Create a role that AWS Glue can assume, with the following trust policy:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": { "Service": "glue.amazonaws.com" },
    "Action": "sts:AssumeRole"
  }]
}

Attach the AWS managed policy AWSGlueServiceRole, which grants the AWS Glue Data Catalog and Amazon CloudWatch Logs access the job needs. Then add an inline policy that scopes Amazon S3 access to your bucket, covering the input data, the pipeline zip, the pipeline storage (state) path, and the warehouse location:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject", "s3:ListBucket"],
    "Resource": [
      "arn:aws:s3:::amzn-s3-demo-bucket",
      "arn:aws:s3:::amzn-s3-demo-bucket/*"
    ]
  }]
}

For a full breakdown of the baseline permissions, see Setting up IAM permissions for AWS Glue.

Set the walkthrough variables

Set the following variables, replacing the example values (us-east-1, amzn-s3-demo-bucket, the account ID 111122223333, and the role name) with your own:

export AWS_REGION="us-east-1"
export BUCKET="amzn-s3-demo-bucket"
export PREFIX="simple-sdp-demo"
export DATABASE="simple_sdp_demo_db"
export ROLE_ARN="arn:aws:iam::111122223333:role/AWSGlueServiceRole-sdp-demo"
export JOB_NAME="simple-sdp-demo"

Step 2 – Set up sample data

The pipeline reads a CSV of order records. Save the following as orders.csv:

order_id,customer_id,region,amount,status,order_ts
O-1001,C-101,EMEA,120.50,COMPLETE,2026-07-23T08:00:00Z
O-1002,C-102,AMER,750.00,COMPLETE,2026-07-23T08:15:00Z
O-1003,C-103,EMEA,-10.00,INVALID,2026-07-23T08:30:00Z
O-1004,C-104,APAC,320.25,COMPLETE,2026-07-23T09:00:00Z
O-1005,C-105,AMER,250.00,COMPLETE,2026-07-23T09:15:00Z
O-1006,C-106,EMEA,90.00,COMPLETE,2026-07-23T09:30:00Z

Upload the file to the input/ location under your project prefix, which is where the bronze layer reads it (the ORDERS_PATH in 01_bronze.py, shown in Step 3). Use the variables you exported in Step 1:

aws s3 cp orders.csv \
  "s3://${BUCKET}/${PREFIX}/input/orders.csv" \
  --region "${AWS_REGION}"

The file includes one invalid order (O-1003, a negative amount), which the silver layer filters out to demonstrate the validation step. The AMER and EMEA regions each have two completed orders, so the gold layer’s order_count and average_order_value are meaningful aggregations rather than single-row passthroughs.

Step 3 – Build the pipeline files

The pipeline project uses the structure introduced earlier: a transformations/ folder holding the three layer definitions (01_bronze.py, 02_silver.py, 03_gold.sql), plus the spark-pipeline.yml specification. The following screenshot shows this layout in a code editor.

Pipeline project layout in a code editor, showing the transformations folder and the spark-pipeline.yml file.

Figure 4: The pipeline project layout in a code editor.

The complete contents of each file follow.

3a. spark-pipeline.yml

The specification names the pipeline, points to the Data Catalog database, configures state storage, and discovers transformation files. As with the transformation files, it uses the __DATABASE__, __BUCKET__, and __PREFIX__ tokens, which you substitute at packaging time in Step 4:

name: simple_sdp_demo
catalog: spark_catalog
database: __DATABASE__
storage: s3://__BUCKET__/__PREFIX__/state/
libraries:
  - glob:
      include: transformations/**
configuration:
  spark.sql.shuffle.partitions: "4"

3b. transformations/01_bronze.py

Bronze preserves the raw source as strings. No coercion, no filtering:

"""Bronze layer: preserve source order records as strings."""
from pyspark import pipelines as dp
from pyspark.sql import DataFrame, SparkSession
from pyspark.sql.types import StringType, StructField, StructType

spark = SparkSession.active()

ORDERS_PATH = "s3://__BUCKET__/__PREFIX__/input/orders.csv"

ORDERS_SCHEMA = StructType([
    StructField("order_id", StringType(), True),
    StructField("customer_id", StringType(), True),
    StructField("region", StringType(), True),
    StructField("amount", StringType(), True),
    StructField("status", StringType(), True),
    StructField("order_ts", StringType(), True),
])


@dp.materialized_view(comment="Raw orders loaded from CSV")
def bronze_orders() -> DataFrame:
    return (
        spark.read
        .schema(ORDERS_SCHEMA)
        .option("header", "true")
        .csv(ORDERS_PATH)
    )

The path uses the tokens __BUCKET__ and __PREFIX__ rather than hardcoded values. AWS Glue reads these files from the packaged zip at runtime, so shell variables like ${BUCKET} are not expanded inside them. You substitute the tokens with your real values when you package the project in Step 4, which keeps every file consistent with the variables you exported in Step 1.

3c. transformations/02_silver.py

Silver casts types, filters to complete orders with positive amounts, and derives an amount_band classification:

"""Silver layer: type, validate, and classify complete orders."""
from pyspark import pipelines as dp
from pyspark.sql import DataFrame, SparkSession
from pyspark.sql.functions import col, to_timestamp, trim, when

spark = SparkSession.active()


@dp.materialized_view(comment="Validated complete orders with typed values")
def silver_orders() -> DataFrame:
    typed = (
        spark.table("bronze_orders")
        .select(
            trim(col("order_id")).alias("order_id"),
            trim(col("customer_id")).alias("customer_id"),
            trim(col("region")).alias("region"),
            col("amount").cast("double").alias("amount"),
            trim(col("status")).alias("status"),
            to_timestamp("order_ts", "yyyy-MM-dd'T'HH:mm:ss'Z'").alias("order_ts"),
        )
        .filter(
            col("order_id").isNotNull()
            & col("region").isNotNull()
            & col("order_ts").isNotNull()
            & (col("status") == "COMPLETE")
            & (col("amount") > 0)
        )
    )
    return typed.select(
        "*",
        when(col("amount") >= 500, "large")
        .when(col("amount") >= 100, "medium")
        .otherwise("small")
        .alias("amount_band"),
    )

Silver reads bronze with spark.table("bronze_orders"), so SDP infers the dependency and runs bronze first. Two details matter here:

  • The to_timestamp call passes an explicit format, "yyyy-MM-dd'T'HH:mm:ss'Z'". The source timestamps are ISO 8601 with a Z suffix. Giving the format treats Z as a literal and produces the same wall-clock value regardless of the job’s session time zone, which keeps the result deterministic.
  • The transformation runs in two projections: the first casts and filters, and the second derives amount_band from the already-typed amount column. Deriving columns with .select(...) rather than a separate .withColumn(...) step keeps SDP’s reference to bronze_orders resolvable as a pipeline dependency. This way, SDP consistently orders the bronze layer before the silver layer. The order matters here too. Spark 4.1 enables ANSI mode by default, so comparing the raw string amount against a number would fail. amount_band therefore reads the already-cast amount.

3d. transformations/03_gold.sql

The gold layer aggregates order metrics by region using SQL:

CREATE MATERIALIZED VIEW gold_sales_summary
COMMENT 'Completed-order metrics by region'
AS
SELECT
  region,
  COUNT(*) AS order_count,
  CAST(ROUND(SUM(amount), 2) AS DECIMAL(10, 2)) AS total_sales,
  CAST(ROUND(AVG(amount), 2) AS DECIMAL(10, 2)) AS average_order_value
FROM silver_orders
GROUP BY region;

Step 4 – Package the project

Substitute the __BUCKET__, __PREFIX__, and __DATABASE__ tokens with the values you exported in Step 1. Then package spark-pipeline.yml and the transformations/ folder into a zip with both at the zip root. Because AWS Glue reads these files from the zip at runtime, the substitution has to happen now, at packaging time, not through shell variables at run time:

# Render the tokens into a build/ copy, leaving your source files untouched
rm -rf build/package && mkdir -p build/package/transformations

sed -e "s|__BUCKET__|${BUCKET}|g" \
    -e "s|__PREFIX__|${PREFIX}|g" \
    -e "s|__DATABASE__|${DATABASE}|g" \
    spark-pipeline.yml > build/package/spark-pipeline.yml

sed -e "s|__BUCKET__|${BUCKET}|g" \
    -e "s|__PREFIX__|${PREFIX}|g" \
    transformations/01_bronze.py > build/package/transformations/01_bronze.py
cp transformations/02_silver.py transformations/03_gold.sql build/package/transformations/

# Zip with the spec and transformations at the zip root
(cd build/package && zip -r -q ../simple-sdp-demo.zip spark-pipeline.yml transformations)

# Upload
aws s3 cp build/simple-sdp-demo.zip "s3://${BUCKET}/${PREFIX}/pipeline/simple-sdp-demo.zip" --region "${AWS_REGION}"

Only spark-pipeline.yml and 01_bronze.py carry tokens, so the other files are copied as-is. The uploaded object is named simple-sdp-demo.zip, which is the same name the job references in Step 6.

Step 5 – Create the database

The database named in spark-pipeline.yml must already exist in the AWS Glue Data Catalog, with an S3 location URI, before the pipeline runs. SDP does not create it automatically:

aws glue get-database --name "${DATABASE}" --region "${AWS_REGION}" >/dev/null 2>&1 \
|| aws glue create-database \
--database-input "{\"Name\":\"${DATABASE}\",\"LocationUri\":\"s3://${BUCKET}/${PREFIX}/warehouse/\"}" \
--region "${AWS_REGION}"

Step 6 – Configure the job

Create an AWS Glue 6.0 job with the zip as ScriptLocation and the SDP flag enabled:

aws glue create-job \
--name "${JOB_NAME}" \
--role "${ROLE_ARN}" \
--command "{\"Name\":\"glueetl\",\"ScriptLocation\":\"s3://${BUCKET}/${PREFIX}/pipeline/simple-sdp-demo.zip\",\"PythonVersion\":\"3\"}" \
--glue-version "6.0" \
--worker-type "G.1X" \
--number-of-workers 2 \
--default-arguments "{\"--enable-spark-declarative-pipeline\":\"true\",\"--enable-glue-datacatalog\":\"true\"}" \
--region "${AWS_REGION}"

Key arguments:

Argument Purpose
--enable-spark-declarative-pipeline Activates the SDP executor (required)
--enable-glue-datacatalog Uses the AWS Glue Data Catalog as the Spark Hive metastore, so the pipeline’s output tables register in the catalog
ScriptLocation Points to the pipeline zip, not a .py file

Table 2: Key arguments for the create-job command.

The create-job command sets ScriptLocation to the pipeline zip. You can also point it to an Amazon S3 prefix: upload the unzipped spark-pipeline.yml and transformations/ to a prefix and set ScriptLocation to that prefix (with a trailing /). No other change is needed, and the --enable-spark-declarative-pipeline flag stays the same. The zip keeps the upload to a single object.

Step 7 – Validate (dry run)

Run the job in validation mode first to verify the dependency graph without materializing data:

aws glue start-job-run \
  --job-name "${JOB_NAME}" \
  --arguments '{"--conf":"spark.glue.sdp.jobMode=VALIDATE"}' \
  --region "${AWS_REGION}"

Validation analyzes the project structure, dependency graph, and SQL and Python compilation without creating tables, executing transforms, or writing data. Confirm that the database has no tables after validation completes.

On AWS Glue, validation runs as a job (jobMode=VALIDATE), so you create the job in Step 6 and then validate it here. If you develop locally with the open source spark-pipelines CLI, you can run its dry-run against the project before packaging and uploading.

Step 8 – Run the pipeline

Start the pipeline in normal execution mode:

aws glue start-job-run \
  --job-name "${JOB_NAME}" \
  --arguments '{"--conf":"spark.glue.sdp.jobMode=RUN"}' \
  --region "${AWS_REGION}"

After the run completes, list the materialized tables:

aws glue get-tables \
  --database-name "${DATABASE}" \
  --region "${AWS_REGION}" \
  --query 'TableList[].Name' \
  --output table

Expected tables: bronze_orders, silver_orders, gold_sales_summary.

After the run, the AWS Glue console shows the three output tables in the simple_sdp_demo_db database. The database’s Location is the warehouse path you configured, s3://amzn-s3-demo-bucket/simple-sdp-demo/warehouse/, and each table stores its data under that prefix. The following screenshot shows the database properties and the three tables (bronze_orders, silver_orders, and gold_sales_summary), each registered in the AWS Glue Data Catalog.

The bronze_orders, silver_orders, and gold_sales_summary tables in the AWS Glue Data Catalog.

Figure 5: The three output tables in the AWS Glue Data Catalog.

Step 9 – Query results

Query the tables with Amazon Athena. If this is your first time using Athena in this Region, set an Amazon S3 query-results location for your workgroup first (Athena console, Settings). Also make sure your identity can read the simple_sdp_demo_db tables in the Data Catalog and the underlying S3 data.

-- Bronze preserves all 6 source rows
SELECT * FROM simple_sdp_demo_db.bronze_orders ORDER BY order_id;

-- Silver retains the 5 complete orders with positive amounts
SELECT * FROM simple_sdp_demo_db.silver_orders ORDER BY order_id;

-- Gold aggregates by region
SELECT * FROM simple_sdp_demo_db.gold_sales_summary ORDER BY region;

Expected gold result:

region order_count total_sales average_order_value
AMER 2 1000.00 500.00
APAC 1 320.25 320.25
EMEA 2 210.50 105.25

Table 3: Gold layer aggregation results by region.

Running the query in the Amazon Athena console returns the aggregated result. The following screenshot shows the gold query and its three result rows (AMER, APAC, and EMEA), matching the values in the preceding table.

Amazon Athena console showing the gold query and its AMER, APAC, and EMEA result rows.

Figure 6: The gold table results in the Amazon Athena console.

Cost considerations

AWS Glue 6.0 bills ETL jobs by the data processing unit (DPU)-hour, per second, with a 1-minute minimum per run. AWS Glue 6.0 is also priced 30 percent lower per DPU-hour than AWS Glue 5.1, with no change to your workload, so the same job costs less to run on 6.0. This walkthrough runs on 2 G.1X workers (2 DPUs), reads a 6-row CSV, and completes each run in about 2 minutes. It produces three tables in one AWS Glue Data Catalog database.

To estimate the cost of a run, multiply the 2 DPUs by the run time in hours by your Region’s AWS Glue 6.0 DPU-hour rate. You can find that rate on the AWS Glue pricing page, and rates differ by AWS Region. The Amazon S3 objects created are the 6-row CSV, the pipeline zip, and the three tables’ data. To stop further charges, delete the resources when you finish, as shown in the next step.

Step 10 – Clean up

To avoid ongoing charges, delete the resources you created:

# Delete the AWS Glue job
aws glue delete-job --job-name "${JOB_NAME}" --region "${AWS_REGION}"

# Delete the Data Catalog database and its table metadata
aws glue delete-database --name "${DATABASE}" --region "${AWS_REGION}"

# Remove the S3 objects
aws s3 rm "s3://${BUCKET}/${PREFIX}/" --recursive --region "${AWS_REGION}"

What’s next

You now have a single pipeline that turns raw order records into validated, aggregated analytics tables, without writing orchestration logic. From here you can:

  • Extend: Add transformation stages (additional @dp.materialized_view functions) and connect them by referencing upstream tables. The pipeline picks up the new dependency automatically.
  • Scale: This walkthrough uses materialized views throughout, so every layer fully recomputes on each run (materialized views don’t support incremental refresh). To process only new data as it arrives, convert the bronze layer to a streaming table, which maintains state across runs with checkpoints. For that cross-run state to persist, a streaming table’s data and checkpoint state must not be stored locally. Hive or AWS Glue managed tables require the database’s LocationUri to point to an Amazon S3 path, while Apache Iceberg tables manage their table metadata themselves.
  • Govern: Protect the Data Catalog tables SDP produces with AWS Lake Formation fine-grained access control. It enforces table-, row-, column-, and cell-level permissions on read queries in AWS Glue Spark jobs (Glue 5.0 and later, for Hive and Iceberg tables). Because this enforcement covers batch reads, it applies to SDP’s materialized views but not to streaming tables, which read through Spark Structured Streaming.
  • Automate: Store the pipeline project in source control. Have your continuous integration and continuous delivery (CI/CD) pipeline package and upload it to Amazon S3 so each job run maps to a known build. Version the zip by object key, or upload the unzipped project to an S3 prefix and turn on Amazon S3 bucket versioning.
  • Monitor: Use Amazon CloudWatch metrics and AWS Glue job run insights for pipeline observability, latency tracking, and failure alerting.

Conclusion

In this post, you used Spark Declarative Pipelines, the declarative alternative to explicitly orchestrated ETL, now available in AWS Glue 6.0. Two decorated Python functions and one SQL file define the bronze, silver, and gold datasets, and SDP resolves the dependencies and manages execution order for you.

With SDP, you declare what each dataset should contain and the declarative framework handles ordering and execution. A three-layer pipeline that would otherwise need separate transform and orchestration logic runs as one job that you can ship and maintain.

To get started, open the AWS Glue console and build the walkthrough pipeline, or adapt the pattern to your own bronze, silver, and gold datasets. For the full set of features, see the AWS Glue 6.0 launch announcement. To move existing jobs to the Spark 4.1 runtime, see Upgrade AWS Glue jobs to AWS Glue 6.0 with AI-powered Spark upgrades. For job configuration details, see the AWS Glue Developer Guide.


About the authors

Syed Humair

Syed Humair

Syed is a Senior Analytics Specialist Solutions Architect at Amazon Web Services, based in Dubai. He has nearly 20 years of experience in data strategy, data engineering, AI, and enterprise architecture across industries including financial services, retail, telecom, and healthcare. At AWS, he works with enterprise customers to build AI-ready data foundations, from lakehouse architectures and open data formats to real-time analytics and data governance. He is the co-author of the AWS Certified Data Engineer Study Guide (Wiley, 2025).

Shrey Malpani

Shrey Malpani

Shrey is a Senior Product Manager Technical at Amazon Web Services (AWS), where he works at the intersection of distributed data processing and data integration. He helps customers build AI-ready data platforms for analytics and machine learning. His focus is scaling data integration and data management across services like AWS Glue, Amazon EMR, and Amazon Redshift.

Bo Li

Bo Li

Bo is a Senior Software Development Engineer on the AWS Glue team. He is devoted to designing and building end-to-end solutions to address customers’ data analytic and processing needs with cloud-based, data-intensive and generative AI technologies.

Kartik Panjabi

Kartik Panjabi

Kartik is a Software Development Manager on the AWS Glue team. His team builds generative AI features for data integration and distributed systems for data integration.

From silos to insights: Federated data access patterns for AI agents

Post Syndicated from James Wu original https://aws.amazon.com/blogs/big-data/from-silos-to-insights-federated-data-access-patterns-for-ai-agents/

Enterprise data today is scattered across specialized systems, each with its own tools and expertise. Querying a database requires SQL. Accessing batch data on Amazon Simple Storage Service (Amazon S3) requires compute engines such as Amazon Athena and Trino. Consuming real-time streams from Amazon Kinesis requires streaming expertise. Each software as a service (SaaS) application has its own API, authentication model, and query language. Today, only data engineers can navigate this landscape, and business users file tickets, wait for reports, or rely on dashboards that answer yesterday’s questions. When a leader needs a one-time answer spanning multiple systems, they’re back in the ticket queue.

Consider a streaming media company: customer profiles, content catalogs, and ad campaign performance are stored as batch data on Amazon S3. Viewership telemetry such as device type, stream quality, watch duration, and buffering events flows in real time through Amazon Kinesis. Subscriber management and support tickets live in a relational customer relationship management (CRM) database. Leaders routinely ask questions like:

  • Which titles drove the most subscriber growth last quarter?
  • How does marketing spend correlate with viewing completion rates?
  • Is churn spiking among users who haven’t engaged with new content?

Answering these questions faces two challenges:

The data silo problem. The data lives in multiple places with batch stores on S3, real-time streams in Kinesis, and an online transaction processing (OLTP) database, each with its own access patterns, query language, and authentication model. Organizations traditionally solve this by building data lakes or adopting a data mesh, but both require significant data engineering investment and ongoing maintenance.

The access gap. The expertise to navigate the enterprise systems is concentrated in the hands of few data engineers, creating a bottleneck that no dashboard or business intelligence (BI) tool fully resolves. Every new one-time requirement means more engineering work, and it’s not self-service.

A fundamentally different approach is emerging: instead of moving all data to one place or building bespoke integrations for each source, let AI agents talk directly to the systems where data lives. Model Context Protocol (MCP) makes this possible, an open protocol that standardizes how AI applications connect to external data sources and tools. MCP servers wrap diverse systems behind a uniform interface for tool discovery, invocation, and response handling. Any user can ask a question in natural language and the agent reaches the right data without knowing which system holds it, what API to use, or what query language is required.

In this post, we propose reference architectures for accessing data stored in different systems and datastores using MCP and Amazon Bedrock AgentCore. The patterns apply to enterprises with mixed data sources, but we ground the narrative in our streaming media company example described earlier to make the problem concrete.

Solution overview

Our solution is a federated data foundation for a streaming media company. It supports real-time and batch analytics using MCP servers and Amazon Bedrock AgentCore, and it makes analytics accessible across the organization. The following reference architecture shows the complete picture from data ingestion through governance and compute layers to the generative AI layer where agents orchestrate across MCP servers. The demo uses synthetic data: batch datasets are generated with Python scripts, and streaming telemetry is produced by AWS Lambda. The complete source code is available in the accompanying GitHub repository, so you can deploy and try it yourself.

Reference architecture showing data ingestion, governance and compute layers, and the generative AI layer where agents orchestrate across MCP servers

Figure 1: Reference architecture for federated data access across batch, streaming, and relational sources

Walkthrough

This section covers the prerequisites and then walks through how a user request flows end to end through the reference architecture.

Prerequisites

Request flow

  1. User request: A user submits a natural-language question through a React application served by Amazon CloudFront with static assets on Amazon S3.
  2. Authentication: Amazon Cognito authenticates the user and issues an identity token that travels with the request to the agent layer.
  3. Agent orchestration: The request reaches a Strands agent running on AgentCore runtime, a capability of Amazon Bedrock AgentCore. The agent reasons over the question and determines which data sources to query.
  4. Gateway routing: Amazon Bedrock AgentCore Gateway, a capability of Amazon Bedrock AgentCore, aggregates all three MCP servers behind a single endpoint, handling tool discovery, authentication, and routing.
  5. MCP server execution: The agent routes the query to the appropriate MCP server(s), each running on Amazon Bedrock AgentCore runtime behind Amazon Bedrock AgentCore Gateway. The Data Processing MCP server queries AWS Glue Data Catalog and Amazon Athena for batch and streaming data on S3, the Amazon Aurora MCP server translates tool calls into SQL against the Amazon Aurora MySQL CRM database, and the AWS Documentation MCP server provides AWS service context.
  6. Data sources: The architecture deliberately spans multiple storage systems to reflect how enterprise data is typically fragmented across teams and technologies. Batch data (customer profiles, content titles, and ad campaigns) is generated by AWS Lambda on an Amazon EventBridge schedule and lands as Parquet files on Amazon S3. Streaming viewership telemetry (what users watch, when they pause, where they drop off) flows through Amazon Kinesis Data Streams and Amazon Data Firehose to S3. CRM records (subscriber plans, support tickets, account status) live in an Amazon Aurora MySQL database. AWS Glue Data Catalog registers the S3-based sources under a unified metadata layer, and AWS Lake Formation enforces fine-grained access policies across the catalog. This mix of batch, streaming, and relational sources is what makes federated access essential. No single query engine can reach all datasets natively.
  7. Response: Results flow back through Amazon Bedrock AgentCore Gateway to the agent, which composes a natural-language answer and delivers it to the user through the front end.

For deploying our reference architecture, follow the instructions in the code repository.

Design patterns for federated data access

Within our architecture, we propose three design patterns for federated data access, each on a spectrum between centralized governance and direct access flexibility.

Pattern 1: Catalog-first access

AWS Glue Data Catalog registers all S3 sources under a unified metadata layer: schemas, business context, data quality metrics, and lineage. The AWS Data Processing MCP server, hosted on Amazon Bedrock AgentCore runtime, wraps AWS Glue Catalog metadata and Amazon Athena query capabilities behind standard MCP tool calls. So when a user asks “Which ad campaigns drove the most subscriber activations last quarter?”, the agent discovers tables through catalog tools and resolves business terms from column metadata. It then executes the join through Athena without ever calling a Glue API directly.

The following diagram traces how a single user request flows through the federated data access architecture: from the agent, through the MCP server, and down to the data in Amazon S3.

Request flow for the catalog-first access pattern, from the agent through the MCP server to data in Amazon S3

Figure 2: Request flow for the catalog-first access pattern

Internally, our agent built using Strands Agent framework has three components: a system prompt, a large language model (LLM), and a set of MCP tools. We use Claude Haiku 4.5 powered by Amazon Bedrock as the foundation LLM with tools discovered through the Amazon Bedrock AgentCore Gateway. The system prompt teaches the agent how to use those tools not by listing every column in every table, but by providing intent-based routing rules and a mandatory schema discovery workflow. Here’s an extract from the system prompt:

TOOL DISCOVERY & ROUTING:

You access tools via the MCP Gateway. Use x_amz_bedrock_agentcore_search
to find the right tool by keyword when unsure.

Routing by intent:
- Telemetry/streaming/viewing data → Glue catalog tools, then Athena query tools
- CRM/support tickets/ratings → MySQL tools (run_query, get_table_schema)
- AWS service questions → documentation search tools

SCHEMA DISCOVERY (MANDATORY before writing SQL):

Before writing any Athena query, retrieve the table schema:
→ Use manage_aws_glue_tables with operation='get-table',
database_name='acme_telemetry', table_name='<table>'

This returns all columns, data types, partition keys, and storage details.

To see this in action, consider what happens when a user asks “How many streaming events in February 2026 by event type?”:

  1. The agent’s routing rules match “streaming events” to the AWS Glue Catalog and Athena query path. If unsure which tool to use, the Gateway’s semantic search discovers tools by keyword rather than requiring exact names.
  2. The agent calls manage_aws_glue_tables exposed by the Data Processing MCP server to retrieve the full schema: column names and types, partition keys (year, month, day, hour), and storage format.
  3. With the schema in hand, the agent writes Presto/Trino SQL with partition filters (WHERE year='2026' AND month='02').
  4. The agent executes the query, retrieves results, and composes a natural-language answer. The user never sees SQL, Glue APIs, or partition strategies.

This discover-then-query workflow is what makes the pattern self-service. The Amazon Bedrock AgentCore Gateway provides unified tool discovery as new MCP servers appear without updating routing logic. The AWS Glue Data Catalog provides a live metadata layer for new tables and columns to appear immediately.

This pattern isn’t unique to AWS. Other platforms adopt the same model. For example, Databricks offers managed MCP servers for Unity Catalog, letting agents discover and query governed datasets, AI models, and functions registered in Unity Catalog. The common trade-off across all of them: all data must be cataloged before agents can access it, which can bottleneck rapidly changing environments.

Catalog-first access where the agent uses AWS Glue Data Catalog and Amazon Athena to query governed data on Amazon S3

Figure 3: Catalog-first access with AWS Glue Data Catalog and Amazon Athena

Pattern 2: Direct source access

Agents access source systems directly through dedicated MCP servers (no intermediate catalog). The Aurora MCP server, hosted on Amazon Bedrock AgentCore runtime, queries the Amazon Aurora CRM database directly. Therefore, a question like “How many open support tickets from premium subscribers?” routes to the MCP server, which translates the tool call into SQL against Aurora. The agent never constructs a database connection or manages credentials. The MCP server handles authentication through AWS Secrets Manager and exposes only two tools: run_query for SQL execution and get_table_schema for schema inspection.

Direct source access where the Aurora MCP server queries the Amazon Aurora CRM database without an intermediate catalog

Figure 4: Direct source access to the Amazon Aurora CRM database

Internally, the same agent architecture as Pattern 1 applies: a system prompt, an LLM, and a set of MCP tools. We use Claude Haiku 4.5 powered by Amazon Bedrock as the foundation LLM with tools discovered through the Amazon Bedrock AgentCore Gateway. There’s no catalog layer to query first. The system prompt provides lightweight schema hints: table names and key enum values needed for WHERE clauses so the agent can route correctly and write valid filters without a round trip:

MYSQL CRM DATA (Aurora MySQL via RDS Data API):

Database: acme_crm

Tables:
- support_tickets: status (open|in_progress|resolved|closed),
  priority (low|medium|high|critical),
  category (billing|technical|content|account)
- content_ratings: rating (1-5), review_text

Use get_table_schema to verify full column details before complex queries.
Use run_query(sql='SELECT...') to execute. Default to read-only SELECT.
Use standard MySQL syntax (not Presto/Trino).

For straightforward queries, the agent writes SQL directly from these hints. For complex queries such as multi-table joins or unfamiliar columns, the agent calls get_table_schema first to verify the full schema, mirroring the discover-then-query discipline from Pattern 1 but against the source database rather than a catalog. To see this in action, consider “Show me open critical support tickets by category”:

  1. The agent’s routing rules match “support tickets” to the MySQL CRM path and call run_query with a SELECT against support_tickets filtered by status='open' and priority='critical'.
  2. The Aurora MCP server translates this into a query against Amazon Aurora through the RDS Data API.
  3. Results return through the AgentCore Gateway and the agent composes a formatted answer with ticket counts, categories, and so on.

The direct access pattern trades catalog governance for simplicity. There’s no metadata registration step. The MCP server queries the database as-is, which means schema changes in Aurora are immediately visible. This makes it ideal for operational databases where the schema is stable and well-understood, and where the overhead of cataloging every table would slow down access without adding value.

Earlier this year, the AWS MCP Server became generally available. It’s part of the Agent Toolkit for AWS, a suite of tooling that includes the MCP Server, skills, and plugins that help coding agents build more effectively and efficiently on AWS. Rather than exposing a fixed set of per-service tools, the server provides generic AWS API access: aws___run_script executes Python in a sandboxed environment with credentialed access to the AWS APIs, authenticated with SigV4 and authorized by your existing AWS Identity and Access Management (IAM) policies. Because that reaches most of AWS APIs, you can connect your agents to relational data in Aurora through the RDS Data API or to real-time streaming data in Kinesis Data Streams, using boto3 calls such as GetShardIterator and GetRecords.

Pattern 3: Hybrid access

In practice, most organizations won’t pick only one pattern because the data landscape is too diverse. That’s exactly the case for our streaming media company: batch and streaming data on S3 benefits from catalog-first governance (Pattern 1), while the Aurora CRM database is better served by direct access (Pattern 2). Our reference architecture combines both patterns under a single orchestrator agent. Governed sources route through the catalog. Operational sources are accessed directly and both paths coexist behind the same agent. The key insight: both paths use the same protocol. Amazon Bedrock AgentCore runtime hosts the MCP servers, and AgentCore Gateway handles tool discovery, authentication, and routing. Organizations can start with whichever pattern fits their current data maturity and grow into unified access as they onboard more sources.

Validate the deployment

Access the CloudFront URL from the stack outputs, log in with your test user credentials, and try these queries:

Query 1 – Customer analytics with visualization:

“Build a chart on customer breakup by subscription type?”

The agent queries the customers table in Athena and generates bar and pie charts showing the distribution across subscription tiers.

Bar and pie charts showing customer distribution across subscription tiers

Figure 5: Customer distribution across subscription tiers

Query 2 – CRM operational breakdown:

“Show me the breakdown of support tickets by category and priority.”

This routes entirely to the MySQL MCP server, querying the Aurora CRM database for ticket distribution without touching S3 or Athena.

Support ticket breakdown by category and priority returned from the Aurora CRM database

Figure 6: Support ticket breakdown by category and priority

Query 3 – Federated cross-source query:

“What are the top five highest-rated titles and how many streaming hours do they have?”

This requires the agent to query content_ratings from Aurora for ratings, then correlate with streaming_events and titles in Athena.

Query results listing the top five highest-rated titles alongside their streaming hours

Figure 7: Top five highest-rated titles and their streaming hours

Things to consider

Consider these additional factors when you deploy the preceding architecture patterns to production:

  • Application security: Our architecture patterns use Amazon Cognito for identity access and control. However, you should carefully review the identity used by the agent to interact with backend systems.
  • Data lineage and access control: Consider using AWS Lake Formation for data governance, authentication, and authorization of data assets in the agentic AI application.
  • Semantic layer for agents: Agentic response quality can be improved by providing agents with the right business context and building an independent semantic layer. AWS has recently announced support for business context and semantic search. This can help the agent discover and understand data by semantic meaning, improve response quality and avoid hallucination, and many other issues.

Clean up

To avoid ongoing charges, destroy both AWS Cloud Development Kit (AWS CDK) stacks (agent stack first, then data stack) and remove any orphaned resources such as Kinesis streams and Amazon CloudWatch log groups. For detailed clean-up instructions, visit the repository’s README.

Conclusion

Enterprise data stays locked behind silos and an access gap. Every one-time question routes through a handful of data engineers while the insight goes stale. MCP flips the model. Instead of centralizing data or wiring bespoke integrations, you deploy MCP servers that wrap each source behind a standardized protocol and let AI agents query them on behalf of the user. Whether you choose catalog-first access, direct access, or both unified behind a single agent, the agent navigates the complexity so the user doesn’t have to. Adding a new data source means deploying a new MCP server, not redesigning the pipeline.

Open questions remain, for example, data lineage across agent-composed outputs, identity and authorization when agents are the primary data consumers, and audit trails that capture not only what an agent accessed but why. This landscape is growing fast: AWS Labs MCP Servers, AWS MCP documentation, and the MCP Gateway Registry.

Deploy the reference architecture, experiment with the patterns, and contribute back what you learn.

Acknowledgements

We would like to thank Yadgiri Pottabathini for his effort in testing the repository.


About the authors

James Wu

James Wu

James is a Principal GenAI/ML Specialist Solutions Architect at AWS, helping enterprises design and execute AI transformation strategies. Specializing in generative AI, agentic systems, and media supply chain automation, he is a featured conference speaker and technical author. Prior to AWS, he was an architect, developer, and technology leader for over 10 years, with experience spanning engineering and marketing industries.

Rahul Sharma

Rahul Sharma

Rahul is a Sr. Specialist Solutions Architect at Amazon Web Services. He is passionate about the data technologies that help leverage data as a strategic asset and is based out of New York.

Amit Kalawat

Amit Kalawat

Amit is a Principal Solutions Architect at Amazon Web Services based out of New York. He works with enterprise customers as they transform their business and journey to the cloud.

Anirudha Joshi

Anirudha Joshi

Anirudha is a Principal Customer Solutions Manager at AWS. A firm believer in working backwards from customer problems, AJ partners with AWS Media & Entertainment (M&E) customers to guide them through their unique technology transformation journeys. He is a member of the AWS Serverless and Machine Learning/Artificial Intelligence TFCs, with a focus on Agentic AI. Outside of work, AJ coaches and runs marathons, hits the trails hiking, and plays golf.

Network connectivity patterns for the next generation of Amazon OpenSearch Serverless

Post Syndicated from Salman Ahmed original https://aws.amazon.com/blogs/big-data/network-connectivity-patterns-for-the-next-generation-of-amazon-opensearch-serverless/

Network connectivity patterns for private access to Amazon OpenSearch Serverless used to require considerable setup. You had to create virtual private cloud (VPC) endpoints in every consumer VPC and configure Amazon Route 53 Profiles for cross-account DNS. You also had to maintain custom private hosted zones with CNAME records and deploy resolver inbound endpoints for on-premises connectivity. The next generation of OpenSearch Serverless changes this. It uses standard AWS PrivateLink interface endpoints with native private DNS support. Connectivity patterns that previously required multi-step DNS orchestration now work with the same endpoint mechanics you already use for other AWS services.

Collections use resource-based endpoints on the on.aws domain in two formats. The per-collection endpoint (<collectionId>.aoss.<region>.on.aws) reaches a single collection, and the hostname itself identifies which collection you want, so no additional routing information is needed. The per-account Regional endpoint (<accountId>.aoss.<region>.on.aws) reaches any collection in your account through one hostname. Because the hostname alone does not identify a specific collection, you add the x-amz-aoss-collection-name header (or x-amz-aoss-collection-id) to each request to name the target collection. The AWS SDKs include this header automatically when they sign the request with Signature Version 4 (SigV4).

Both formats use standard AWS PrivateLink. You create the VPC endpoint from the Amazon Virtual Private Cloud (Amazon VPC) console or the Amazon Elastic Compute Cloud (Amazon EC2) CreateVpcEndpoint API, using the service name com.amazonaws.<region>.aoss-data. It is the same interface endpoint you create for any other AWS service.

In this post, each pattern shows the architecture, the DNS resolution flow, and the data traffic path. Patterns 1 through 8 operate within a single Region across one or more accounts, labeled Region A in the diagrams, so the repeated Region A boxes in a cross-account pattern are the same Region. Only Pattern 9 spans Regions, shown as Region A and Region B.

These patterns apply to the collection (data) endpoint only. When you create a collection, you also receive an OpenSearch UI endpoint. That endpoint uses a separate PrivateLink mechanism today, with its own VPC endpoint and access policy, and is on a path to move to the standard PrivateLink model. OpenSearch UI connectivity is out of scope for this post.

Prerequisites

DNS resolution

When you create a standard VPC endpoint for com.amazonaws.<region>.aoss-data with private DNS enabled, AWS creates a private hosted zone for *.aoss.<region>.on.aws and associates it with your VPC. This zone maps collection hostnames to the endpoint’s private elastic network interface (ENI) IP addresses. Your compute’s DNS query reaches the VPC’s Amazon Route 53 Resolver at VPC+2, which resolves the hostname to ENI IPs.

One endpoint serves every collection hostname in the Region. The following AWS CLI command creates that interface endpoint, and the --private-dns-enabled flag turns on the private DNS resolution described here.

aws ec2 create-vpc-endpoint \
  --vpc-id vpc-abc123 \
  --service-name com.amazonaws.us-east-1.aoss-data \
  --vpc-endpoint-type Interface \
  --subnet-ids subnet-111 subnet-222 \
  --security-group-ids sg-xxx \
  --private-dns-enabled

In Regions that support Federal Information Processing Standards (FIPS), the same endpoint also resolves *.aoss-fips.<region>.on.aws for FIPS-compliant access.

OpenSearch Serverless has no per-collection Dashboards endpoint. Use OpenSearch UI applications to explore and visualize collection data.

The diagrams in the following patterns use an Amazon EC2 instance to represent the compute client. Any compute in the VPC reaches a collection the same way, including EC2 instances, AWS Lambda functions attached to the VPC, and containers on Amazon Elastic Container Service (Amazon ECS) or Amazon Elastic Kubernetes Service (Amazon EKS). The connectivity, DNS resolution, and access policies are the same regardless of the compute type.

Pattern 1: Private access from a single VPC

Compute in a VPC needs private access to collections in the same account. The following diagram shows the architecture for private access from a single VPC.

Compute in a single VPC reaches a collection through a VPC interface endpoint with private DNS enabled

Figure 1: Private access from a single VPC

Create a standard VPC endpoint in the VPC where your compute runs, then reference its ID in the collection’s network policy.

For the DNS resolution flow, (1) compute queries <collectionId>.aoss.<region>.on.aws, and the VPC Route 53 Resolver at VPC+2 returns the endpoint ENI IP addresses because private DNS is enabled on the endpoint.

For the data traffic path, (2) compute connects to the ENI, and (3) PrivateLink forwards the request to the service, which routes to the collection by hostname.

Pattern 2: Multiple VPCs in the same account

Several VPCs, split by environment, tier, or team, need private access to the same collections. The following diagram shows how each VPC uses its own endpoint to reach the same collections.

Three VPCs in one account, each with its own aoss-data interface endpoint reaching the same collections

Figure 2: Multiple VPCs in the same account

Each VPC needs exactly one aoss-data endpoint with private DNS enabled, and that single endpoint already reaches every collection in the Region. DNS resolves independently within each VPC, so there is no cross-VPC DNS dependency. Adding a new VPC takes two steps. Create the endpoint, then add its endpoint ID to the collection’s network policy. Do not create a second aoss-data endpoint with private DNS enabled in the same VPC. Both endpoints share the same private hosted zone, which causes a conflict and the creation fails.

For the DNS resolution flow, (1) compute in each VPC queries the collection hostname, and that VPC’s Route 53 Resolver at VPC+2 returns the endpoint ENI IP addresses because private DNS is enabled on the endpoint.

For the data traffic path, (2) compute connects to its local ENI, and (3) PrivateLink forwards the request to the service, which routes to the collection by hostname.

Pattern 3: On-premises access from a single account

On-premises clients reach collections over AWS Direct Connect or AWS Site-to-Site VPN, which connect to the VPC through AWS Transit Gateway or AWS Cloud WAN. The following diagram shows the DNS and data path for on-premises access.

Figure 3: On-premises access from a single account

On-premises DNS servers sit outside the VPC and cannot resolve PrivateLink private DNS names directly. Place an Amazon Route 53 Resolver inbound endpoint in the VPC that holds the aoss-data VPC endpoint. On-premises DNS forwards queries for aoss.<region>.on.aws to that inbound endpoint. The inbound endpoint resolves them against the private hosted zone. The inbound endpoint’s security group must allow TCP/UDP port 53 from your on-premises resolver ranges.

For the DNS resolution flow, (1) the client queries the on-premises resolver. (2) The on-premises conditional forwarder for *.aoss.<region>.on.aws sends the query over Direct Connect or VPN, through Transit Gateway or Cloud WAN, to the inbound endpoint. The inbound endpoint uses the VPC Route 53 Resolver to return the endpoint’s private ENI IPs.

For the data traffic path, (3) the client sends an HTTPS request with the Transport Layer Security (TLS) Server Name Indication (SNI) header set to the collection hostname, over Direct Connect or VPN through Transit Gateway or Cloud WAN. (4) Traffic crosses the VPC’s attachment ENI, (5) reaches the VPC endpoint ENIs, and (6) PrivateLink forwards the request to the service.

Pattern 4: Cross-account access with an endpoint in each consumer VPC

A central account hosts collections, and compute in spoke accounts needs private access. Many enterprises start here. The following diagram shows the cross-account endpoint architecture.

Spoke accounts each with their own interface endpoint reaching collections in a central account over PrivateLink

Figure 4: Cross-account access with an endpoint in each consumer VPC

Each spoke creates its own endpoint. The collection owner’s network policy references the spoke’s endpoint ID. The data access policy grants the spoke’s IAM role. PrivateLink carries the traffic end to end, with no Transit Gateway and no peering.

The endpoint lives in the spoke account, not the collection account. The spoke team creates a standard interface VPC endpoint in the spoke VPC for the service name com.amazonaws.<region>.aoss-data with private DNS enabled. The collection owner does not create this endpoint. After the endpoint is ready the spoke shares its endpoint ID with the collection owner, who adds that ID to the collection network policy under SourceVPCEs. A network policy accepts endpoint IDs from accounts across your organization. Each spoke creates its own endpoint and shares the ID rather than peering VPCs or routing through another account’s endpoint.

Network access and data access stay separate. The network policy authorizes the endpoint, and the data access policy authorizes the identity. A serverless data access policy grants principals from the collection’s own account. For a spoke in another account, you create an IAM role in the collection account and grant that role in the data access policy. The spoke role then assumes it to sign requests.

The following network access policy lists the two spoke endpoint IDs under SourceVPCEs and sets AllowFromPublic to false, so only those endpoints reach the collection and the policy denies public access.

[
  {
    "Description": "Cross-account access from spoke",
    "Rules": [
      {
        "ResourceType": "collection",
        "Resource": [
          "collection/my-collection"
        ]
      }
    ],
    "AllowFromPublic": false,
    "SourceVPCEs": [
      "vpce-spoke-b-id",
      "vpce-spoke-c-id"
    ]
  }
]

For the DNS resolution flow, (1) spoke compute queries the VPC Route 53 Resolver at VPC+2, which returns the local endpoint ENI IPs because private DNS is enabled on the endpoint.

For the data traffic path, (2) compute connects to the local ENI. (3) PrivateLink forwards the request to the service, which checks the network policy for the endpoint ID and the data access policy for the IAM role before routing. Adding a spoke takes one API call and two policy edits.

Pattern 5: Centralized shared endpoint with Amazon Route 53 Profiles over Transit Gateway

You want fewer PrivateLink endpoints, so you run one shared endpoint in a networking VPC and reach it from spoke accounts over Transit Gateway or AWS Cloud WAN, with no endpoint in each spoke. The following diagram shows this centralized architecture.

Figure 5: Centralized shared endpoint with Amazon Route 53 Profiles over Transit Gateway

Pattern 5 consolidates access through a single shared endpoint in a central networking VPC rather than creating one per spoke. Because spoke VPCs have no local endpoint, they cannot resolve *.aoss.<region>.on.aws on their own. You share the endpoint’s private DNS with spoke VPCs using Amazon Route 53 Profiles, shared through AWS Resource Access Manager (AWS RAM). This is the one pattern where you still manage DNS propagation.

For the DNS resolution flow, (1) the spoke resolves the hostname through the shared Route 53 Profile, which returns the networking-VPC endpoint ENI IPs.

For the data traffic path, (2) traffic leaves the compute through the spoke VPC’s attachment ENI, (3) crosses Transit Gateway or Cloud WAN into the networking VPC’s attachment ENI, (4) reaches the shared endpoint ENIs, and (5) PrivateLink forwards the request to the service.

Pattern 6: Cross-account centralized networking with on-premises

A central account hosts collections. A separate networking account owns Direct Connect or VPN and Route 53. On-premises clients reach the collections through the networking account. The following diagram shows this architecture.

Figure 6: Cross-account centralized networking with on-premises

The networking account runs the standard VPC endpoint and a Route 53 Resolver inbound endpoint. The collection owner’s network policy references the networking account’s endpoint ID.

For the DNS resolution flow, (1) the on-premises client queries its resolver. (2) The on-premises conditional forwarder sends the query over Direct Connect or VPN, through Transit Gateway or Cloud WAN, to the inbound endpoint. The inbound endpoint uses the VPC Route 53 Resolver to return the endpoint’s private ENI IPs.

For the data traffic path, (3) the client sends HTTPS over Direct Connect or VPN, through Transit Gateway or Cloud WAN. (4) Traffic crosses the networking VPC’s attachment ENI, (5) reaches the networking-VPC endpoint ENIs, and (6) PrivateLink forwards the request to the service in the central account. The two teams coordinate through one artifact, the endpoint ID.

Pattern 7: Distributed multi-business-unit with spoke-account access

Spoke accounts such as analytics or application teams need collections spread across several business unit accounts, and each unit manages its own collections. The following diagram shows the distributed multi-business-unit architecture.

Spoke accounts reaching collections spread across several business unit accounts, each spoke with its own endpoint

Figure 7: Distributed multi-business-unit with spoke-account access

Each spoke creates one standard endpoint, which resolves every collection hostname in the Region. Each business unit’s network policy lists the spoke endpoint IDs. Access control decides which collections a spoke reaches. DNS does not.

For the DNS resolution flow, (1) spoke compute queries the VPC Route 53 Resolver at VPC+2, which returns the endpoint ENI IPs because private DNS is enabled on the endpoint.

For the data traffic path, (2) compute connects to the local ENI, and (3) PrivateLink forwards the request to the service, which routes to the correct business unit collection by hostname.

Action Required change
New collection in any BU No networking change is needed because in the network policy collection/*  wildcard, already covers any new collection
New spoke account Spoke creates an endpoint, and BUs add its ID to their policies
Remove spoke access BUs remove the endpoint ID and the IAM principal

Pattern 8: Distributed multi-business-unit with on-premises access

Several business units own collections in separate accounts. On-premises clients reach collections across all of those accounts through a central networking account. The following diagram shows this architecture.

Figure 8: Distributed multi-business-unit with on-premises access

The networking account runs one standard endpoint that resolves *.aoss.<region>.on.aws hostnames, regardless of which account owns the collection. Each business unit’s network policy includes the networking endpoint ID.

For the DNS resolution flow, (1) the on-premises client queries its resolver. (2) The on-premises conditional forwarder for *.aoss.<region>.on.aws sends the query over Direct Connect or VPN, through Transit Gateway or Cloud WAN, to the networking VPC’s inbound endpoint. The inbound endpoint uses the VPC Route 53 Resolver to return the shared endpoint’s private ENI IPs.

For the data traffic path, (3) the client sends HTTPS over Direct Connect or VPN, through Transit Gateway or Cloud WAN. (4) Traffic crosses the networking VPC’s attachment ENI, (5) the request arrives at the shared endpoint ENIs, and (6) the service routes to business unit 1 or business unit 2 by hostname, as long as that business unit’s policy lists the networking endpoint ID. Adding a collection in any business unit needs no networking change if the network policy uses a collection/* wildcard, since the wildcard already covers it.

Pattern 9: Cross-Region access strategies

Consumers in Region B need data that lives in collections in Region A. The following diagram shows cross-Region access strategies.

Independent collections in Region A and Region B, each with its own endpoint, with a replication arrow showing cross-Region data sync

Figure 9: Cross-Region access strategies

Collections are Regional. No built-in cross-Region endpoint or replication exists. Deploy independent collections in each Region, each with its own endpoint and policies, then synchronize data with one of these approaches.

  • Dual-write. The application writes to both Regions at ingestion time.
  • Amazon OpenSearch Ingestion pipeline. A pipeline replicates index operations to the secondary Region with near real-time lag. The pipeline creates its own PrivateLink endpoint to the destination collection. It adds the endpoint to that collection’s network policy automatically. You only need to name the network policy and grant the pipeline role.
  • Amazon Simple Storage Service (Amazon S3) Cross-Region Replication with re-ingestion. Cross-Region Replication copies objects, and an OpenSearch Ingestion pipeline loads them into the local collection. Lag runs in minutes, at the lowest cost of these approaches.

For the DNS resolution flow, DNS resolves locally in each Region, the same as Pattern 1. Each collection hostname carries its Region, so a hostname in Region A resolves through Region A’s own endpoint and a hostname in Region B resolves through Region B’s own endpoint, with no cross-Region DNS.

For the data traffic path, (1) compute in each Region uses that Region’s own endpoint to reach its local collection. Writes land in the primary Region and the sync approach you choose replicates them to the secondary Region, where local readers query the replica. The replicate arrow shows that cross-Region movement, such as an OpenSearch Ingestion pipeline that writes into the secondary-Region collection.

Scale-to-zero changes the economics. An idle secondary-Region collection costs only storage until requests arrive.

Summary

Pattern Components
1. Same VPC Standard endpoint and network policy
2. Multiple VPCs Endpoint per VPC and a policy listing all IDs
3. On-premises Endpoint, Route 53 inbound endpoint, on-premises forwarder, and Transit Gateway or Cloud WAN
4. Cross-account Endpoint per consumer, network policy, and data policy
5. Centralized shared endpoint Shared endpoint, Route 53 Profiles through RAM, and Transit Gateway or Cloud WAN
6. Central networking with on-premises Networking endpoint, Route 53 inbound, forwarder, Transit Gateway or Cloud WAN, and policies
7. Multi-BU with spoke access Endpoint per spoke, and each BU policy lists spoke IDs
8. Multi-BU with on-premises One networking endpoint reached through Transit Gateway or Cloud WAN, and each BU policy lists its ID
9. Cross-Region Independent collections per Region and a data-sync approach

Across each private pattern, the VPC endpoint resolves all *.aoss.<region>.on.aws hostnames through standard PrivateLink private DNS. Network policies control which endpoints reach a collection, and data access policies control which principals operate on the data. Only Pattern 5 asks you to manage DNS.

Cost considerations

The connectivity pattern you choose drives recurring cost, so match it to your scale instead of adding infrastructure you do not need. The two charges that come up most often, a Route 53 Resolver inbound endpoint and Route 53 Profiles, are both optional for access that stays inside AWS.

A Route 53 Resolver inbound endpoint is needed only for the on-premises patterns (3, 6, and 8), where an on-premises resolver forwards queries into the VPC. Traffic that stays inside AWS never uses it. Route 53 Profiles apply only when a VPC has no endpoint of its own, as in Pattern 5, where the profile carries the shared endpoint’s private DNS to the spoke. When each VPC runs its own interface endpoint, DNS resolves locally through the VPC Route 53 Resolver at no extra charge, so neither the inbound endpoint nor a profile is required.

For most multi-account and multi-Region deployments, an interface endpoint in each consumer VPC (Pattern 4) is the least complex and often the least expensive option. You pay for the interface endpoints you already need for private access, and local DNS resolution adds nothing. Because collections are Regional and each Region resolves on its own, this scales across Regions with no cross-Region DNS.

Centralizing on one shared endpoint (Pattern 5) lowers the number of interface endpoints. However, it adds Transit Gateway or Cloud WAN data processing charges and the cost of sharing DNS. You share that DNS either through Route 53 Profiles or through a private hosted zone that you associate across accounts and maintain yourself. A smaller endpoint count is not automatically cheaper because transit data processing can exceed the savings. Compare both designs against your own traffic before you decide.

Scale to zero also shapes cost. An idle collection, such as a secondary-Region replica in Pattern 9, releases its compute and bills only for storage until requests arrive. For current rates, see AWS PrivateLink pricing, Amazon Route 53 pricing, and Amazon OpenSearch Service pricing.

Conclusion

OpenSearch Serverless uses standard AWS PrivateLink for private connectivity. You create a VPC endpoint, enable private DNS, and reference the endpoint ID in your network policy. The model scales from single-VPC access to multi-account and multi-business-unit designs, and only Pattern 5 adds DNS infrastructure, where you share the endpoint’s private DNS with Route 53 Profiles. The per-account regional endpoint goes further and serves any collection in an account through one hostname and connection pool. To get started, create your first collection in the OpenSearch Serverless console, or explore the OpenSearch Serverless documentation for detailed API references and tutorials.


About the authors

Salman Ahmed

Salman Ahmed

Salman is a Senior Technical Account Manager at AWS, specializing in helping customers design, implement, and optimize their AWS environments. He combines deep networking expertise with a passion for exploring emerging technologies to help organizations get the most out of their cloud investments. Outside of work, he enjoys photography, traveling, and watching his favorite sports teams.

Ankush Goyal

Ankush Goyal

Ankush is a Senior Technical Account Manager at AWS Enterprise Support, specializing in helping customers in the travel and hospitality industries optimize their cloud infrastructure. With over 20 years of IT experience, he focuses on using AWS networking services to drive operational efficiency and cloud adoption. Ankush is passionate about delivering impactful solutions and helping clients to streamline their cloud operations.

Ravi Bhatane

Ravi Bhatane

Ravi is a Software Engineer at AWS working on Amazon OpenSearch Serverless. He builds the gateway layer that fronts the service, handling private connectivity, authentication, and request routing for customer traffic into collections. He’s drawn to distributed systems and the challenge of keeping them secure, highly available, and low latency as they grow. Outside of work, he enjoys photography and hiking.