Tag Archives: announcements

ICYMI: August 2026 @AWS Security

Post Syndicated from Rodolfo Brenes original https://aws.amazon.com/blogs/security/icymi-august-2026-aws-security/

Read all about the latest AWS security features, compliance updates, and hands-on resources in our monthly digest posts. You’ll find expert blog posts, new service capabilities, code samples, and workshops.

AWS Security Blog posts

August brought 20 AWS Security Blog posts organized across seven categories. Identity and access management led the month with five posts covering self-service rate limits for Amazon Cognito, a decade of AWS Managed Microsoft AD, a redesigned sign-in experience, console Private Access for isolated VPCs, and automated IAM Identity Center governance. Data protection followed with four posts on AWS KMS data key caching, ACME protocol support in AWS Certificate Manager, Amazon S3 over-permissioned access remediation, and the upcoming deprecation of email-based domain validation. AI security continued to grow with four posts on custom authentication in Amazon Bedrock AgentCore Gateway, user authorization propagation in AI agents, and extending Bedrock Guardrails to tool interactions. Threat detection, governance and networking.

Identity

From 2 weeks to 2 minutes: Amazon Cognito launches provisioned limits for self-service rate limit management

Authors: Kiran Dongara, Howie Li | Published: August 5, 2026

Learn to use Amazon Cognito provisioned limits for on-demand authentication rate limit adjustments, replacing the previous 10–14 day support ticket process with self-service capacity scaling in minutes.

A decade of enterprise identity in the cloud with AWS Managed Microsoft AD

Authors: Vladimir Provorov, Tekena Orugbani, Rodney Underkoffler | Published: August 7, 2026

AWS Managed Microsoft AD celebrates 10 years of fully managed Active Directory in the cloud, now offering Standard, Enterprise, and Hybrid editions with multi-Region replication and 20+ AWS service integrations.

Updates to your AWS sign-in experience

Authors: Vaibhav Chowla, Ella Segura | Published: August 17, 2026

AWS is gradually rolling out a redesigned sign-in page with a unified email entry point, social identity provider options, and an updated session selection experience for managing multiple active sessions.

Extend your data perimeter to the AWS Management Console with Private Access

Authors: Madhur Kulkarni, Abhijit Barde, Sujay Ghosh, Mateusz Jaworski | Published: August 28, 2026

AWS Management Console Private Access now supports VPCs without internet connectivity, routing all console traffic – authentication, static assets, and service API calls – through AWS PrivateLink endpoints to strengthen your data perimeter.

Automate IAM Identity Center governance with continuous discovery and reporting

Author: Jonathan Nguyen | Published: August 31, 2026

Learn to deploy automated discovery and reporting for AWS IAM Identity Center applications and assignments across your organization, with event-driven monitoring that validates naming conventions and enables near real-time enforcement of governance policies.

Data Protection

Caching KMS data keys in multi-thread environments: per-tenant encryption for event-driven systems at scale

Authors: Maria Gutovsky, Hemmy Yona | Published: August 6, 2026

Learn to solve the cache stampede problem in multi-tenant envelope encryption using the AWS-recommended hierarchical keyring pattern or a custom Caffeine-based caching approach to reduce AWS KMS costs.

Automate certificates with ACME support in AWS Certificate Manager

Authors: Anthony Harvey, Chandan Kundapur | Published: August 6, 2026

Learn to use ACME protocol support in AWS Certificate Manager to automate public certificate issuance and renewal using standard clients like Certbot and cert-manager, with enterprise controls for domain scoping and centralized visibility.

Securing your Amazon S3 buckets: identifying and remediating over-permissioned access

Authors: Hetal Kolekar, Fernando Chiera di Vasco Freitas, Manonmayi Vedam | Published: August 7, 2026

Learn to detect and fix over-permissioned Amazon S3 buckets across multi-account environments using AWS Lambda, AWS Config, and AWS Security Hub, with automation for continuous monitoring.

AWS Certificate Manager will discontinue email validation to prove domain validation for certificates

Authors: Adam Aboudi, Poojil Tripathi | Published: August 13, 2026

ACM will discontinue email-validated public certificates by September 30, 2027, aligning with CA/B Forum standards – learn the timeline and how to migrate to DNS validation in place.

AI Security

Implement custom authentication for tools integration using request Lambda interceptor in AgentCore Gateway

Authors: Nishant Mainro, Ram Ramani | Published: August 18, 2026

Learn to use a request Lambda interceptor in Amazon Bedrock AgentCore Gateway to bridge legacy authentication mechanisms like Basic Auth, isolating credentials from AI agents using AWS Secrets Manager.

Propagate user authorization context in AI agents with Amazon Bedrock AgentCore

Authors: Anshu Bathla, Prafful Gupta, Rohit Verma | Published: August 19, 2026

Learn to enforce least-privilege access in AI agents by propagating user identity through Amazon Bedrock AgentCore to Amazon DynamoDB, Knowledge Bases, and Salesforce, without embedding authorization logic in agent code.

Extend Amazon Bedrock Guardrails to tool interactions using the Strands Agents SDK

Authors: Stephan Traub | Published: August 27, 2026

Learn to extend Amazon Bedrock Guardrails beyond the model boundary to tool calls, external data, and MCP server interactions using three validation checkpoints built with Strands Agents SDK lifecycle hooks.

Threat detection and incident response

Security Hub Extended adds supply chain security as its tenth category

Author: Michael Fuller | Published: August 18, 2026

AWS Security Hub Extended now includes supply chain security with Chainguard and Socket as curated partners, helping you verify open source dependencies and block malicious packages through a single AWS billing relationship.

Detecting multi-stage attacks on AWS: a guide to cross-service signal correlation

Authors: Nisha Kashyap | Published: August 26, 2026

Learn to correlate signals across AWS CloudTrail, VPC Flow Logs, and Route 53 Resolver logs to detect multi-stage attacks by layering your business context, data classification, access norms, and change windows – on top of Amazon GuardDuty Extended Threat Detection.

AWS partners with Anthropic and OpenAI to bring AWS Continuum into developer workflows

Author: Chet Kapoor | Published: August 5, 2026

AWS Continuum for code vulnerabilities now integrates with Anthropic Claude Code, OpenAI Codex, and Kiro, enabling developers to discover, prioritize, validate, and remediate vulnerabilities within their coding environments.

We invited a direct competitor into Security Hub Extended. Here’s why.

Authors: Michael Fuller | Published: August 31, 2026

AWS Security Hub Extended now includes Upwind, a runtime-first cloud security company, offering eBPF-based workload protection with pay-as-you-go pricing through a single AWS bill, reinforcing customer choice even where capabilities overlap with AWS offerings.

Infrastructure security

AWS Network Firewall now supports rule hit count

Authors: Preetkumar Shah, Amit Gaur, Cheriyan Mundapuzha, Santosh Shanbhag, Srivalsan Mannoor Sudhagar | Published: August 20, 2026

AWS Network Firewall now tracks how often stateful rules match traffic, helping you identify unused rules, validate security controls for compliance, and accelerate incident response at no additional cost.

Governance and compliance

Landing Zone Accelerator independent assessment report for C5:2020 now available on AWS Artifact

Authors: Kevin Donohue, Michael Wahlers | Published: August 11, 2026

An independent assessment by Schellman evaluates how Landing Zone Accelerator on AWS aligns to C5:2020 requirements, implementing 325 security controls to accelerate your compliance journey.

Fast track ISM-ready cloud environments and IRAP assessments with Landing Zone Accelerator on AWS

Authors: Kevin Donohue, Dave Connell, Dan Friebe | Published: August 25, 2026

A new independent assessment by gwi.digital evaluates Landing Zone Accelerator against 1,081 ISM controls, achieving 91% coverage of addressable scope to help Australian customers accelerate IRAP assessment readiness.

How Moeve scales AWS governance with automated AWS Organization Service Control Policies

Authors: Gonzalo Guerrero, Jonatan De Martín, Rayco Martinez | Published: August 26, 2026

Learn how Moeve manages 150 SCPs across 300+ accounts in three AWS Organizations using a governance-as-code model; policies live in Git, deploy through GitHub Actions pipelines, and attach dynamically based on account metadata during onboarding, with Amazon EventBridge and AWS Lambda providing real-time observability of every organizational change.

August Security Bulletins

In August 2026, AWS published 23 security bulletins addressing vulnerabilities across open-source SDKs, MCP servers, developer tools, and the OpenSearch ecosystem. A dominant theme was the AI agent tool surface: prompt-injection consent bypasses in Strands Agents Tools enabled command and code execution, an insecure direct object reference exposed cross-tenant agent memory, and credential disclosure and authorization flaws affected the Amazon MQ, DocumentDB, and AWS Transform MCP servers and the Bedrock AgentCore harness. Remote code execution recurred throughout, from prototype pollution and Java deserialization in OpenSearch to path-traversal-to-root in amazon-ssm-agent, Zip Slip in awsdac, and an uncontrolled search path in the Kiro IDE and CLI on Windows.

Other notable issues include disabled SSH host key verification in the AWS CLI, memory-safety flaws in the AWS SDK for C++, privilege escalation in the FreeRTOS-Kernel, and memory-amplification denial of service in Amazon ion-java. OpenSearch accounted for a large share of the month, spanning authorization, input validation, SSRF, stored XSS, and denial of service, while Athena Federated Query connectors exposed Secrets Manager secrets. A common thread: insufficient authorization and input validation at the boundary between AI agents and the systems they reach. All patches are available, upgrade promptly. For more information, see AWS Security Bulletins.

AWS Samples

In August 2026, we published 17 new code samples organized into five categories: governance and compliance (7), AI security (3), data protection and privacy (3), identity and access management (2), and infrastructure security (2). This month’s collection reflects the rapid growth of agentic AI workloads: most samples focus on governing, auditing, and securing AI agents built on Amazon Bedrock AgentCore, from platform-level governance and telemetry to fraud investigation and biosecurity screening.

Governance and compliance

Agentic Governance Platform

Learn to deploy an AWS-native control plane for governing AI agents across your enterprise with centralized registry, Microsoft Entra ID single sign-on, Cedar tool policies, multi-vendor agent inventory, and Langfuse observability, all self-hosted on Amazon Bedrock AgentCore.

Enterprise Agentic AI Platform Accelerator

Learn to deploy a secure, governed foundation for production AI agents on Amazon Bedrock AgentCore with modular CDK stacks covering identity, gateway, memory, runtime, and observability; supporting Strands Agents, LangGraph, and Claude Agent SDK with opt-in security controls.

Video Compliance Agent

Learn to deploy an end-to-end pipeline that automatically verifies video content against broadcast compliance guidelines such as Ofcom; extracting frames, audio transcripts, and OCR text shot by shot, then using Amazon Bedrock to flag potential violations and produce a structured per-shot compliance report.

Intelligent Security for Healthcare APIs

Learn to add behavioral anomaly detection, automated data sensitivity classification, and HIPAA compliance reporting to your FHIR API using Amazon Bedrock; running asynchronously so clinical workflows are never blocked, with Amazon Comprehend Medical and Bedrock Guardrails anonymizing PHI throughout the monitoring path.

Governed Agentic Companion

Learn to deploy a governed, orchestrator-driven multi-agent builder companion on Amazon Bedrock AgentCore, reachable from Kiro, Claude Code, or any MCP client; the kit enforces 13 codified tenets through an always-on governance gate with no off switch, routing each request to a specialist while blocking deploys, secret leaks, and ungrounded answers by construction.

Audit the Agent

Learn to deploy a serverless daily executive audit pipeline for AWS AI agents (AWS DevOps Agent, AWS Security Agent) using AWS Step Functions and AWS Lambda; the report answers five questions: what the agent accessed, who authorized it, what it cost, its risk posture across five trust dimensions, and whether you should be concerned, all sourced deterministically from AWS CloudTrail, CUR, and IAM with AI-generated summaries bounded by layered guardrails.

AI Security

Telemetry Enablement for AgentCore CloudFormation

Learn to deploy a single AWS CloudFormation stack that enables Amazon CloudWatch logs and X-Ray traces for every Amazon Bedrock AgentCore resource type – Runtime, Gateway, Memory, Browser, CodeInterpreter, and WorkloadIdentity – using native and custom telemetry rules.

Bedrock Guardrails to OCSF on CloudWatch

Learn to transform Amazon Bedrock Guardrails intervention events into OCSF Detection Finding records and land them in the Amazon CloudWatch unified data store; enabling you to query guardrail violations alongside AWS CloudTrail, Amazon VPC Flow Logs, and other sources with Amazon Athena or CloudWatch Logs Insights.

Bedrock Readiness Agent

Learn to deploy a read-only assessment agent built with the Strands Agents SDK that evaluates your Amazon Bedrock environment across six dimensions: IAM governance, data retention, quota headroom, model selection fitness, cost projection, and operational observability; generating severity-rated findings with AWS CloudFormation and Terraform remediation templates you can apply directly.

Infrastructure security

DDoS Guardian

Learn to install an agent skill that reviews an AWS WAF web ACL as a system: evaluation order, rule interactions, and L7 DDoS posture; then delivers a severity-ranked HTML report with ready-to-apply remediation. Offline, read-only, no AWS resources modified.

Biosecurity Screening Policy on Amazon Bedrock AgentCore Gateway

Learn to use Policy in Amazon Bedrock AgentCore to deterministically screen AI agent tool requests for biosecurity risks, combining Cedar policies with three independent screening layers: MMseqs2 sequence alignment, ESMC-600M embedding similarity, and Foldseek structural homology; enabling defense-in-depth controls that block high-risk protein sequences before they reach downstream tools.

MCP Fraud Investigation Agent

Learn to deploy an end-to-end AI-powered e-commerce fraud investigation agent built with the Strands SDK on Amazon Bedrock AgentCore, connecting through an AgentCore Gateway over MCP to query transaction history, customer profiles, login activity, support cases, and fraud playbooks; a React dashboard on AWS Amplifystreams the agent’s reasoning token by token as it works each case.

Identity

IAM Account Access Manager with ABAC

Learn to implement workforce access using IAM account access manager and attribute-based access control, where one IAM role per project shares a single policy document and access decisions are made by comparing role tags against resource tags at request time; onboarding a new project requires only tagging and entitlement configuration with no policy authoring.

Operationalizing Least Privilege: Automate IAM Remediation through Your CI/CD Pipeline

Learn to automate remediation of unused IAM permissions using AWS IAM Access Analyzer, AWS CloudTrail, and Amazon Bedrock for AI-generated AWS CDK code; the solution attributes each role to its origin (IaC or manual), then creates pull requests for IaC-managed roles or issues for manually created roles in GitLab or GitHub, with configurable exclusion rules and policy diffs for human review.

Data Protection

Data Residency Chatbot with Amazon Bedrock AgentCore

Learn to deploy a data-residency-compliant natural-language chatbot on Amazon Bedrock AgentCore where all data and AI inference stay within a single AWS Region; the solution uses a Strands agent that answers plain-English questions from Aurora PostgreSQL through governed, whitelist-validated read-only tools exposed via an AgentCore Gateway, with a residency guard that rejects any cross-region inference profile, demonstrated with a rooftop-solar subsidy program and adaptable to any sector or geography.

LLM-based PII Detection

Learn to detect personally identifiable information in conversational text using Amazon Bedrock; the solution prompts any Converse-compatible model to return PII spans with category, value, and character offsets, handles long inputs via word-boundary chunking, recovers near-miss labels, and supports custom categories, few-shot examples, and pluggable backends beyond Bedrock.

Agentic Data Classification and Redaction

Learn to build a conversational AI research assistant that automatically classifies documents for MNPI, PII, and security sensitivity, then enforces per-user redaction at query time using Amazon Bedrock AgentCore, Guardrails, and Amazon OpenSearch Serverless vector search.

Conclusion

August 2026 provides comprehensive guidance and runnable code for governing agentic AI workloads at enterprise scale, from orchestrator-driven governance gates and biosecurity screening policies to daily executive audit pipelines and OCSF-normalized guardrail telemetry. The posts and samples provide patterns for console Private Access in isolated VPCs, attribute-based access control with IAM account access manager, automated least-privilege remediation through CI/CD, and cross-service signal correlation for multi-stage threat detection. Each resource includes deployment steps or runnable code so you can validate in your own environment before adopting. Subscribe to the AWS Security Blog RSS feed to receive updates as they publish, and revisit this digest monthly for a consolidated view of what changed and what to act on.

If you have feedback about this post, submit comments in the Comments section below.


Rodolfo Brenes

Rodolfo Brenes

Rodolfo is a Principal Solutions Architect focused on Cloud Governance and Compliance. With over 18 years of experience, he currently leads a technical field community in AWS helping customers scale and improve their security and governance frameworks. Besides work, Rodolfo enjoys video games, playing with his four cats, and won’t say no to a good outdoor adventure.

Anna Brinkmann

Anna has 18 years of experience in the technical content space and has spent the last 6 years managing the AWS Security Blog. Outside of work, she enjoys spending time with her family.

Announcing Spark Connect on Amazon EMR on EC2: Interactive PySpark anywhere

Post Syndicated from Al MS original https://aws.amazon.com/blogs/big-data/announcing-spark-connect-on-amazon-emr-on-ec2-interactive-pyspark-anywhere/

Today, we’re announcing support for Spark Connect on Amazon EMR on EC2 with the AWS runtime for Apache Spark (emr-spark-8.0, Apache Spark 4.0.2 and later). You can now develop and debug PySpark interactively from Amazon SageMaker Unified Studio Data Notebooks or your own IDE, such as Visual Studio Code, PyCharm, Kiro, or Jupyter. Spark runs on a dedicated Amazon EMR on EC2 cluster while your Python runs locally, so you can set breakpoints and inspect a DataFrame against full-size data from your IDE. In SageMaker Unified Studio Data Notebooks, you connect to your cluster, catalog, and AI tools. Production-scale PySpark and SQL run without leaving the studio. Because each session is isolated with its own permissions, your whole team can share one cluster at the same time. This post shows you how to get started with both SageMaker Unified Studio Data Notebooks and your own IDE.

Previously, developing Spark for an Amazon EMR on EC2 cluster meant working in a notebook tied to that cluster, or packaging your code as a job and submitting it before you could see a result. Local code often behaved differently on the cluster because of version and dependency mismatches, and the slow deploy-and-check loop made those differences hard to find. There was no way to attach your own IDE and debugger and inspect a DataFrame mid-transformation. Spark Connect closes that gap: your code runs against the cluster’s own Spark engine while you develop locally, so the environment you debug in is the one that runs your data.

How Spark Connect works on Amazon EMR on EC2

Spark Connect uses a client-server architecture that separates your application code from the Spark engine. The client is a lightweight PySpark library that runs in your notebook or IDE, and it sends DataFrame and SQL operations over a gRPC/TLS connection to a Spark Connect Server on your cluster. The server runs those operations and returns the results to your local session. Your machine does not need Spark installed and does not need to be sized for the workload.

Spark Connect client-server architecture connecting a local PySpark client to the Spark Connect Server on an Amazon EMR cluster

Figure 1: Spark Connect client-server architecture on Amazon EMR on EC2

When you start a session, Amazon EMR launches the Spark Connect Server as a YARN application on your cluster and hands back an endpoint and a short-lived token. There’s no server for you to stand up or manage. Because that server runs on a cluster you already operate, your session inherits the instance types, libraries, bootstrap actions, and Spark configuration you use in production. What you see while debugging is what runs when the same code is scheduled as a batch job, since both use the same cluster and its configuration.

Share one cluster across your team

Now that you can start sessions, a single dedicated cluster can serve your whole team, because each session is a separate resource with its own execution role, tags, and lifecycle. A single cluster supports up to 1,000 concurrent sessions and 1,000 concurrent execution roles. These values are service maximums, not sizing targets. Actual concurrency depends on cluster size and per-session workload. Because interactive sessions are bursty and rarely all active at once, one cluster typically serves a team larger than its peak concurrent-session count. Enable Amazon EMR managed scaling so that capacity tracks demand. If peak concurrency approaches these maximums, or to isolate cost and data access by group, use multiple clusters—for example, one per team, business unit, or environment. Sharing one cluster gives you:

  • On-demand Spark without extra clusters — Developers get interactive sessions without provisioning a cluster apiece, which keeps utilization high and removes the cost of idle per-person clusters.
  • Consistent environments — Everyone runs the same Spark version, libraries, and security configuration, so results stay consistent, and your platform team patches and monitors one cluster.
  • Isolation and attribution — Per-session execution roles and tags keep each person’s work separate, so you can scope data access by session, track cost by user, and stop one session without disturbing anyone else.
  • Full visibility and control — View active sessions in the Spark UI, review finished ones in the Spark History Server, and manage them from the Amazon EMR console, API, CLI, or SDK.

Getting started

Getting started with Spark Connect on Amazon EMR on EC2 takes three steps: Create an Amazon EMR cluster with Spark Connect session enabled, start a session, and connect from your IDE or SageMaker Unified Studio Data Notebooks.

Note: In SageMaker Unified Studio, on-demand cluster creation is available for domains that use AWS IAM Identity Center. For domains that use AWS Identity and Access Management (IAM), attach an existing cluster. If your cluster runs in a private subnet, make sure that your network configuration allows connectivity between SageMaker Unified Studio and the cluster endpoint.

Prerequisites

You must have the following prerequisites in place.

  • An Amazon EMR cluster running release emr-spark-8.0.0 or later with SessionEnabled set to true.
  • The Spark application is installed on the cluster.
  • Python 3.9 or later with pyspark[connect] installed locally. The PySpark version must match the Spark version on your cluster.
  • For clusters in private subnets, the Amazon EMR service role must include the AmazonEMRServicePolicyForSessions managed policy, which grants permissions to create Network Load Balancers and virtual private cloud (VPC) endpoint services in your account.
  • To use Spark Connect sessions, you need permissions to start and list sessions on the cluster (elasticmapreduce:StartSession, ListSessions), get session details and endpoints and terminate sessions (elasticmapreduce:GetSession, GetSessionEndpoint, TerminateSession), and pass the execution role to the Amazon EMR service (iam:PassRole).

Working with interactive sessions

To create a session-enabled cluster and connect to it, follow these steps.

To start a Spark Connect session

  1. Create a cluster with sessions enabled, running emr-spark-8.0.0 or later. The following is a sample command that you can modify for your needs, such as the instance types and counts:
    aws emr create-cluster \
      --name "spark-connect-cluster" \
      --release-label emr-spark-8.0.0 \
      --applications Name=Spark \
      --service-role EMR_DefaultRole \
      --ec2-attributes InstanceProfile=EMR_EC2_DefaultRole,SubnetId=subnet-id \
      --instance-groups '[
        {"InstanceCount":1,"InstanceGroupType":"MASTER","InstanceType":"m8g.xlarge"},
        {"InstanceCount":2,"InstanceGroupType":"CORE","InstanceType":"m8g.xlarge"}
      ]' \
      --session-enabled \
      --tags Key=for-use-with-amazon-emr-managed-policies,Value=true

    Note: The following steps use the AWS Command Line Interface (AWS CLI) directly. If you develop in SageMaker Unified Studio (Option 1), cluster attachment and session creation are handled for you, so you can skip steps 2 through 6.

  2. After the cluster reaches the WAITING state, start a session and wait for it to reach IDLE:
    aws emr start-session --cluster-id j-XXXXXXXXXXXXX --name "my-session"
    aws emr get-session --cluster-id j-XXXXXXXXXXXXX --session-id is-XXXXXXXXXXXXX

    Note: For runtime role sessions, add the --execution-role-arn parameter to the start-session command.

  3. Retrieve the endpoint and token, and build your connection string from the returned Endpoint value rather than hardcoding a host:
    aws emr get-session-endpoint --cluster-id j-XXXXXXXXXXXXX --session-id is-XXXXXXXXXXXXX

    The response includes the endpoint URL and an authentication token:

    {
      "Endpoint": "https://session-id.emr-spark-connect.region.amazonaws.com",
      "AuthToken": "v2.local.xxx...",
      "AuthTokenExpirationTime": "2026-01-01T01:00:00Z"
    }

  4. Install the matching PySpark client and connect. GetSessionEndpoint returns an https:// URL with no port. Build the connection string by converting it to the sc:// scheme and appending :443. Without the port, the PySpark client defaults to 15002, which isn’t reachable. Your Python code runs locally. The SQL and DataFrame operations run on the cluster:
    pip install 'pyspark[connect]==4.0.2' boto3

    from pyspark.sql import SparkSession
    
    session_id = "is-XXXXXXXXXXXXX"
    auth_token = "<AuthToken from get-session-endpoint>"
    host = "<Endpoint from get-session-endpoint, without https://>"
    
    url = f"sc://{host}:443/;use_ssl=true;x-aws-proxy-auth={auth_token};authorization={session_id}"
    spark = SparkSession.builder.remote(url).getOrCreate()
    spark.sql("SELECT 'Hello from EMR on EC2' AS message").show()

  5. Run a transformation against full-size data. This groups a DataFrame, writes the result to Amazon Simple Storage Service (Amazon S3), and reads it back:
    import pyspark.sql.functions as F
    
    df = spark.range(0, 1000).withColumn(
        "category", F.when(F.col("id") % 2 == 0, "even").otherwise("odd")
    )
    df.groupBy("category").count().show()
    df.write.mode("overwrite").parquet("s3://amzn-s3-demo-bucket/demo/")
    spark.read.parquet("s3://amzn-s3-demo-bucket/demo/").filter("id < 50").orderBy("id").show()

  6. When you finish, terminate the session to release cluster resources. Calling spark.stop() only closes the local connection. The session keeps running until you terminate it or it reaches the idle timeout:
    aws emr terminate-session --cluster-id j-XXXXXXXXXXXXX --session-id is-XXXXXXXXXXXXX

  7. When you’re done with the walkthrough, terminate the cluster you created in step 1 so it stops incurring charges. Terminating the cluster also ends any sessions still running on it:
    aws emr terminate-clusters --cluster-ids j-XXXXXXXXXXXXX

You can start a Spark Connect session in two ways: from SageMaker Unified Studio or from your own IDE client.

Option 1: Develop in SageMaker Unified Studio Data Notebooks

Amazon SageMaker Unified Studio brings your data, catalogs, and analytics and AI tools into one place, and Amazon EMR on EC2 is now one of the Spark runtimes a Data Notebook can use. When you choose that cluster as the notebook runtime, SageMaker Unified Studio connects to it over Spark Connect. The same runtime then drives both your PySpark and SQL cells, so a single notebook can query the AWS Glue Data Catalog and transform the data without switching tools. The built-in AI assistant generates code and execution plans from natural-language prompts, and the Spark UI shows running work alongside your other runtimes.

To start a session from SageMaker Unified Studio:

  1. Open a Data Notebook in SageMaker Unified Studio.
  2. In the Compute panel, do one of the following:
    1. To create a new cluster, choose Create cluster and configure an Amazon EMR on EC2 cluster.
    2. To use an existing cluster, choose Attach cluster and select a running Amazon EMR on EC2 cluster.
  3. Select the cluster as the notebook’s runtime.
  4. Begin writing PySpark or SQL code in the notebook cells.

For a complete example, open the SageMaker Unified Studio Spark Connect example notebook , which connects a Data Notebook to an Amazon EMR on EC2 cluster and runs PySpark and SQL cells against the AWS Glue Data Catalog.

Watch a walkthrough: Develop in a SageMaker Unified Studio Data Notebook. The preceding steps cover the same workflow, so you can complete it from the notebook without the video.

Option 2: Develop in your own IDE

Use the IDE of your choice, such as Visual Studio Code, PyCharm, Kiro, or a local Jupyter notebook. You debug Spark the way you debug any Python program: set a breakpoint, inspect a variable, and step through your code, all while the Spark work runs on the cluster. Your libraries, source control, and continuous integration and continuous delivery (CI/CD) stay on your local machine, and only your Spark operations are sent to the cluster.

To see this end to end, the following example attaches an IDE to a Spark Connect session and steps through a breakpoint against cluster data.

Open the local IDE Spark Connect example notebook then use the connection steps in the preceding Getting started section to attach your client.

Watch a walkthrough: Develop your own IDE with Spark Connect. The written connection steps in Getting started cover the same workflow, so you can complete it without the video.

Use cases

Spark Connect on Amazon EMR on EC2 supports the following interactive workflows:

  • Interactive extract, transform, and load (ETL) development: Build and test pipelines against full-size data on the cluster, then schedule the same transformations as a Spark step on that cluster, where the Spark version, libraries, and configuration already match what you validated.
  • Exploratory data analysis and feature engineering: Analyze production-scale data from your notebook or IDE instead of sampled subsets, so you catch data quality issues earlier.
  • Notebook-driven analytics in SageMaker Unified Studio: Run PySpark and SQL next to your catalogs and AI tools, switching runtimes per notebook.
  • Apache Iceberg lakehouse analytics: Query and manage Iceberg tables through the AWS Glue Data Catalog, with time travel, schema evolution, and partition management.
  • Compute standardization: Point interactive development at the same clusters that run your production batch jobs, so development and production share one engine and configuration.

Release information

Spark Connect on Amazon EMR on EC2 is available with the AWS runtime for Apache Spark (emr-spark-8.0, Apache Spark 4.0.2) and later. It’s available in all AWS Regions where Amazon EMR is available, except the AWS GovCloud (US) Regions and the China Regions. The SageMaker Unified Studio experience is available in its supported Regions. There’s no additional charge for Spark Connect. You pay for the Amazon Elastic Compute Cloud (Amazon EC2) instances in your cluster. Because these sessions run on your own clusters, they use the Amazon EMR on EC2 capabilities you already rely on, including AWS Graviton processors for price-performance and your choice of On-Demand, Reserved, AWS Savings Plans, or Spot capacity.

Considerations for the release are as follows:

  • The PySpark version that you install locally must match the Apache Spark version on your cluster.
  • Spark Connect supports the DataFrame and SQL APIs. RDD-based APIs aren’t supported.
  • Authentication tokens expire after 1 hour, and sessions end after a configurable idle timeout (60 minutes by default, up to 24 hours).
  • High-availability clusters with multiple primary nodes, Trusted Identity Propagation, and fine-grained access control through AWS Lake Formation aren’t supported for Spark Connect sessions in this release.

Conclusion

Spark Connect on Amazon EMR on EC2 brings interactive, debuggable PySpark development to the clusters you already run. Develop on a SageMaker Unified Studio Data Notebook or in your own IDE, debug against full-size data while the cluster runs the work and share a single cluster across your whole team. To get started, see the Interactive sessions with Spark Connect guide or open a Data Notebook in Amazon SageMaker Unified Studio. To learn more about the service, see the Amazon EMR detail page.


About the authors

Al MS

Al MS

Al is a product manager for Amazon EMR at AWS.

Karthik Prabhakar

Karthik Prabhakar

Karthik is a Data Processing Engines Architect for Amazon EMR at AWS, where he specializes in distributed systems architecture and query optimization. He partners with customers to solve complex performance challenges in large-scale data processing workloads. His work centers on engine internals, cost optimization, and architectural patterns for efficient petabyte-scale analytics.

Arun Prabakaran

Arun Prabakaran

Arun is a Senior Software Engineer working at AWS. His expertise spans distributed data processing and large-scale systems. He is passionate about building reliable data platforms and enabling organizations to run analytics and AI workloads at scale.

Rekha Veeraraghavan

Rekha Veeraraghavan

Rekha is a Technical Account Manager at AWS and a Subject Matter Expert in AWS Analytics. She helps enterprise and strategic customers optimize their data analytics solutions with expert guidance and technical support. Drawing deep data engineering expertise, she enables organizations to build scalable, efficient, and cost-effective data processing pipelines on AWS.

Supporting ASD’s multi-factor authentication campaign: Why MFA matters more than ever

Post Syndicated from Grace Zhang original https://aws.amazon.com/blogs/security/supporting-asds-multi-factor-authentication-campaign-why-mfa-matters-more-than-ever/

The Australian Signals Directorate (ASD) has this month issued a clear call to action through its Multi-factor authentication: Switch it on campaign, urging businesses, organisations, and individuals to enable multi-factor authentication (MFA) across their online accounts. At AWS, we strongly support this message.

As threat actors continue to target credentials through phishing, credential stuffing, and social engineering, passwords alone are no longer enough. MFA is one of the most effective security controls available. It’s a cornerstone of ASD’s Essential Eight maturity model and a recognized component of major cybersecurity frameworks worldwide. ASD’s campaign reinforces what the security community has long advocated: switching on MFA is one of the simplest and most impactful steps any organization or individual can take to protect themselves online, and we encourage all to heed ASD’s call.

How AWS enforces MFA across every account type

At AWS, we’ve put this principle into practice at scale. In June 2025, AWS Identity and Access Management (IAM) achieved comprehensive MFA enforcement for root users across all account types, a significant milestone and the first of its kind among major cloud providers. This was the culmination of a deliberate, phased security journey: beginning with requiring MFA for AWS Organizations management account root users in May 2024, expanding to standalone account root users in June 2024, introducing centralized root access management in November 2024, and completing enforcement across all account types including member accounts. MFA prevents over 99 percent of password-related attacks and is available to all AWS customers at no additional cost, with support for FIDO2 passkeys and FIDO-certified security keys for phishing-resistant authentication. This milestone reflects our ongoing commitment to secure-by-design principles, setting a high bar for our customers’ default security posture and demonstrating that organizations of any scale can, and should, make MFA the standard rather than the exception.

ASD’s Multi-factor authentication: Switch it on campaign banner

Extending MFA beyond your AWS environment

A compromised email account can be used to reset AWS passwords. A breached source control system can expose infrastructure-as-code secrets. Enable MFA on your email, collaboration tools, source control, and other services that support it. Visit the ASD Multi-factor authentication campaign page for broader guidance.

Getting started with MFA on AWS

AWS enforces MFA automatically for root users. To extend that same protection to your IAM users—the identities your team members and applications use daily—you can configure MFA individually through the AWS Management Console for IAM. To learn more, see Security best practices in IAM. For phishing-resistant authentication with FIDO2 passkeys, see Passkeys and security keys in IAM.

If you have questions or feedback about MFA on AWS, leave a comment below or reach out on AWS re:Post. If you haven’t already, heed ASD’s call and switch on MFA across every account you own.

This post was written in support of ASD’s Multi-factor authentication: Switch it on campaign. For more AWS security content, visit the AWS Security Blog.

If you have feedback about this post, submit comments in the Comments section below.


Grace Zhang

Grace Zhang

Grace is the Regulatory and Security Compliance Lead for Australia and New Zealand (ANZ), based in Sydney. She supports security assurance and compliance initiatives across the ANZ region, helping customers navigate regulatory requirements and build confidence in the security of the AWS Cloud.

Running self-hosted AI agent sandboxes with AWS Lambda MicroVMs

Post Syndicated from Brian Krygsman original https://aws.amazon.com/blogs/compute/running-self-hosted-ai-agent-sandboxes-with-aws-lambda-microvms/

Organizations are building AI agents that autonomously write code, query databases, and interact with internal systems on behalf of their teams. These agents handle use cases such as automated code review, data pipeline optimization, and infrastructure troubleshooting. When your AI agent generates a shell command, queries a database, or writes to a file system, that code needs a secure environment to run in. Without isolation, one session’s tool calls can contaminate another session’s state, inadvertently expose sensitive data across tenants, or unintentionally allow untrusted code to reach production resources. Self-hosted sandboxes solve this by keeping agent execution within your own AWS account, giving you full control over networking, secrets, and governance.

Say you’re building an internal AI agent that optimizes database queries for your engineering team. A developer asks it to find the ten slowest queries in your analytics database, rewrite them with better indexing, and test the results. That’s three tool calls in a single session. One hits a live database with real credentials. One generates code. One executes it. Now multiply that by fifty developers using the assistant at the same time. Each session needs its own credentials, its own filesystem, its own network boundary. If credentials or state cross session boundaries, you have inadvertent data exposure.

AWS Lambda MicroVMs is a serverless compute environment that provides general-purpose runtimes with the strong isolation of virtual machines and the rapid scaling of AWS Lambda. Powered by Firecracker virtualization, each MicroVM runs Amazon Linux with full OS access for up to 8 hours. You launch, suspend, resume, and terminate MicroVMs programmatically. You get the serverless benefits of managed infrastructure, responsive scaling, and pay-per-use pricing. Three capabilities make Lambda MicroVMs a strong fit for agent sandboxes:

  • VM-level isolation per environment: Each MicroVM runs in its own Firecracker virtual machine, providing hardware-virtualization-based isolation between sessions without the resource overhead and startup time required of full VMs. One developer cannot see a teammate’s session, even when both run at the same time.
  • Launch from snapshot: Like Lambda SnapStart, MicroVMs boot from a pre-captured memory and disk snapshot, skipping application initialization entirely. Your agent gets a near-instant ready-to-use environment.
  • 4x vertical scaling without re-provisioning: A running MicroVM can scale CPU and memory up to 4x its initial allocation, which can range from 0.25 vCPU/0.5 GB to 4 vCPU/8 GB, without terminating or re-creating the environment. If the agent needs to run a heavy data transformation mid-session, it can get more resources without starting over.

In this post, we show you how to architect and build a self-hosted AI agent that uses Lambda MicroVMs as secure, isolated sandboxes for tool-call execution. Lambda MicroVMs can handle the compute isolation for running tool calls, while the host for production AI agents, such as Amazon Bedrock AgentCore, manages the agent logic, model routing, and session state. A complete reference solution is available in aws-samples.

How self-hosted sandboxes work

A developer asks the agent to “find the ten slowest queries in our analytics database and suggest index improvements.” The agent orchestration system starts a session then breaks the objective into tool calls and distributes them. A worker needs to pick up that session, run the queries, and return results.

Most AI agent orchestration services and frameworks use a work queue model to distribute tool-call execution. The orchestration service enqueues sessions representing tool-call work. A worker, the process that claims a session and executes its tool calls, runs inside a compute environment, posts results, and exits. In this architecture, each Lambda MicroVM is the compute environment, and the worker is the process running inside it. Claude Managed Agents self-hosted sandboxes run those workers inside your own infrastructure rather than on a shared, multi-tenant compute pool. Your database credentials stay in your virtual private cloud (VPC). Your network, introspection, and governance rules apply.

You can trigger workers in two ways:

  • Webhook-triggered: The orchestration application sends a notification when a session is ready. Your control plane launches a worker on demand.
  • Always-on: A long-running process continuously polls the work queue for new sessions.

The Lambda MicroVMs lifecycle aligns with the webhook-triggered pattern, where each session produces one inbound event that launches a fresh MicroVM. Lambda MicroVMs support configurable idle policies. After a configurable idle period, a MicroVM suspends automatically, preserving disk and memory state. It resumes when inbound traffic arrives or when you call the resume API. The MicroVM runs for the duration of the session, the worker exits, and the idle policy suspends then finally terminates the VM. Lifecycle hooks allow you to run custom logic at key steps in the MicroVM lifecycle.

In contrast, the always-on pattern risks breaking the polling loop by suspending the MicroVM when idle, since there’s no inbound traffic between sessions. You could disable the configurable idle period, but then you pay for empty polling. Use the webhook-triggered approach for self-hosted sandboxes on Lambda MicroVMs.

Architecture

The following figure shows the reference solution’s architecture, with the Anthropic agent orchestration service control plane on the left interacting with a self-hosted sandbox environment in AWS on the right.

Reference architecture showing the Anthropic orchestration control plane sending a webhook through API Gateway to a launcher Lambda function that starts a MicroVM worker in your AWS account

Figure 1: Reference architecture for self-hosted AI agent sandboxes on Lambda MicroVMs

The sample architecture is event-driven. The only inbound traffic is the webhook call. When the event arrives, the handler launches a MicroVM. Once launched, the MicroVM pulls its assigned session from the orchestration system’s work queue and runs the task. In our example, the developer’s “find slow queries” request has been queued as a session. The agent now needs to reach your infrastructure, spin up an isolated environment, and hand off the work. The following sequence shows how each component interacts to fulfill a single session.

The orchestration service queues work as sessions. A MicroVM launches to service each session, and the worker is the process running inside that MicroVM that claims the session, executes tool calls, and returns results.

  1. Once the orchestration service marks a session as ready to run, it sends a session.status_run_started webhook to an Amazon API Gateway endpoint, triggering a MicroVM launch.
  2. The launcher verifies the webhook signature using a signing secret from AWS Systems Manager Parameter Store, rejecting invalid or stale deliveries before spending compute.
  3. The launcher calls RunMicrovm, passing the session ID and a secret reference through runHookPayload. It deduplicates on the webhook event ID (backed by Amazon DynamoDB) so retries do not launch duplicate VMs.
  4. The MicroVM boots from a pre-captured Firecracker snapshot and receives the dispatch on its /run lifecycle hook. The worker fetches the environment key from Parameter Store using its execution role. It pulls the matching session from the work queue, claims it, and executes tool calls in an isolated /workspace directory. When finished, it posts results and exits. The idle policy suspends then terminates the VM.

Deduplication. The webhook event ID serves as the idempotency key. The launcher uses Powertools for AWS Lambda (Python) with a DynamoDB persistence layer to verify exactly-once processing. If the orchestration application retries a delivery with the same event ID, Powertools protects the system from launching extra MicroVMs and doing extra work.

Credential boundaries. Each component accesses only the single secret it needs. The launcher reads only the webhook signing secret to verify inbound events. It passes only an ARN reference to the environment key into the MicroVM payload. The MicroVM’s execution role retrieves only that environment key at runtime. No single component holds both secrets.

Component Has access to
Launcher Lambda Webhook signing secret (verify inbound events)
MicroVM worker Environment key (through the execution role, to poll and claim sessions)

Cost model. You pay for MicroVM run time per session, plus standard charges for API Gateway requests, Parameter Store API calls, and Lambda invocations for the launcher. When no sessions are active, no MicroVMs run. Cost scales with concurrent sessions and their duration, avoiding idle compute charges.

Implementation

The following sections explore the reference architecture in more depth.

Project structure

The reference solution uses AWS Serverless Application Model (AWS SAM) for infrastructure-as-code. Alternatively, if you use an AI coding agent such as Claude Code, Kiro, or Cursor, the Agent Toolkit for AWS includes a Lambda MicroVMs skill that gives your agent the procedures to provision, configure, and deploy MicroVM-based sandbox environments on your behalf.

├── template.yaml                    # SAM: launcher, API, WAF, secrets, roles
├── src/
│   ├── functions/launcher.py        # Verify signature, RunMicrovm
│   ├── microvm-image/
│   │   ├── Dockerfile               # AL2023 + Node.js worker
│   │   └── worker/worker.mjs        # Lifecycle hook server
│   └── scripts/build-image.sh       # Package + create MicroVM image

Launcher: verify the webhook before spinning up compute

When the webhook arrives saying a developer’s session is ready, the launcher’s first action is signature verification. If it fails, the function returns 401 immediately. No MicroVM launches. No DynamoDB writes. You don’t pay for fraudulent or replayed requests.

signing_secret = [REDACTED_PASSWORD]  # Verify webhook before spending compute
if not verify_signature(raw_body, headers, signing_secret):
    return {"statusCode": 401, "body": "invalid signature"}

After verification, the launcher builds a dispatch payload containing the session ID, environment ID, region, and an ARN reference to the environment key secret. It passes this to RunMicrovm through runHookPayload:

launched = microvm_client.run_microvm(
    image_identifier="arn:aws:lambda:us-east-1:123456789012:microvm-image:worker",
    run_hook_payload=json.dumps({"session": dispatch}),
    execution_role_arn=config.execution_role_arn,
    maximum_duration_in_seconds=28800,
    ingress_network_connectors=["arn:aws:lambda:::network-connector:aws-network-connector:ALL_INGRESS"],
    egress_network_connectors=["arn:aws:lambda:::network-connector:aws-network-connector:INTERNET_EGRESS"],
)

MicroVM worker: claim one session, execute, exit

The MicroVM image is built from a Firecracker snapshot. The worker process starts during image creation and is captured in the snapshot, so there is no application startup at run time. The /run lifecycle hook delivers the dispatch payload:

// POST /aws/lambda-microvms/runtime/v1/run
case "run": {
    const envelope = JSON.parse(rawBody);
    const dispatch = JSON.parse(envelope.runHookPayload);
    res.writeHead(200); // Acknowledge hook immediately
    res.end();
    const key = await fetchParameter(dispatch.session.ENVIRONMENT_KEY_PARAM_NAME);
    await pollAndHandleSession(dispatch.session.ANTHROPIC_SESSION_ID, key);
    // Session complete; terminate this MicroVM to release all resources
    await terminateMicroVm(envelope.microvmId);
}

The worker acknowledges the hook within its timeout, fetches the environment key, and claims the session. This is where the requested work begins. The worker connects to the analytics database, runs EXPLAIN ANALYZE on the flagged queries, writes optimized alternatives to /workspace/suggestions.sql, and posts the results back to the developer. All of that happens inside this single VM. When the session completes, the worker calls terminate-microvm to release all compute resources.

Deployment

For full deployment instructions, see the reference solution README. Before deploying, make sure you have these prerequisites.

Prerequisites

Four steps

  1. Deploy the control plane. Build and deploy the SAM stack, which creates the launcher Lambda, API Gateway endpoint, WAF WebACL, DynamoDB idempotency table, Parameter Store entries, and MicroVM execution role.
    sam build
    sam deploy --guided --capabilities CAPABILITY_NAMED_IAM

  2. Register the webhook and populate secrets. In the Claude Console, register the stack’s WebhookUrl output as a webhook endpoint subscribed to session.status_run_started. Store the signing secret and environment key in the Parameter Store resources created by the stack.
  3. Build the MicroVM image. Package the Dockerfile and worker code, upload to Amazon S3, and create the image. The service runs your Dockerfile, launches the worker, and captures a Firecracker snapshot. Monitor build progress in Amazon CloudWatch under /aws/lambda/microvms/<image-name>.
    ./src/scripts/build-image.sh

  4. Verify. Create a test session and confirm a MicroVM launches and completes end-to-end. The reference solution includes a verification script that creates a session, triggers the webhook, and validates the full flow.

Using Claude Platform on AWS (CPOA)

The preceding architecture works similarly when you access Claude through Claude Platform on AWS rather than the first-party API. Three things change in the worker:

  1. Client initialization. Replace the first-party client with the AWS client and supply your workspace ID:
    from anthropic import AnthropicAWS
    
    client = AnthropicAWS(aws_region="us-east-1")
    
    # Workspace ID is required on every request
    # Set via ANTHROPIC_AWS_WORKSPACE_ID env var or pass per-call

  2. Authentication options. CPOA supports two modes:
    1. CPOA API key (aws-external-anthropic-api-key-...): Store it in Parameter Store the same way as the first-party environment key. These keys are short-lived (12-hour STS tokens) and must be regenerated when they expire.
    2. SigV4 (IAM): The MicroVM execution role can sign requests directly, so there is no secret to store or rotate. Set the environment key secret to a placeholder value (for example, use-sigv4) and the SDK falls through to IAM credentials automatically. This is the recommended path for production.

In both authentication modes, attach the AWS managed policy AnthropicSelfHostedEnvironmentAccess to the MicroVM execution role. This policy grants the aws-external-anthropic actions needed to poll the work queue, claim sessions, and post results. See IAM actions for Claude Platform on AWS for the full reference.

Prerequisite: Enable outbound web identity federation once per AWS account:

aws iam enable-outbound-web-identity-federation

Everything else, including webhook verification, deduplication, credential separation, and idle policy remains the same.

Security

Earlier we talked about what goes wrong without isolation. Credentials exposed between sessions. Scripts unintentionally reaching production. Agents escaping their sandbox. This architecture implements defense in depth to help prevent these.

Each component accesses a single, scoped secret. The launcher passes only an ARN reference to the worker credential into the MicroVM. The MicroVM’s execution role retrieves only that credential at runtime. The analytics database connection string does not touch the launcher and does not leave your environment.

AWS WAF applies managed rule sets (OWASP, known bad inputs, IP reputation) and per-IP rate limiting. Amazon API Gateway request validation rejects malformed bodies. The launcher performs HMAC signature verification as the true authentication boundary.

Each session runs in its own MicroVM. Sessions do not share memory, disk, or network namespaces. Firecracker provides hardware-virtualization-based isolation. The launcher IAM role reads only the signing secret. The MicroVM execution role reads only the worker credential. Both are scoped to specific Parameter Store ARNs. The Amazon S3 artifact bucket blocks public access, enables versioning, and uses server-side encryption.

Conclusion

This post walked through how to give your internal AI agent a safe place to run database queries, generate code, and execute scripts on behalf of fifty developers without leaking data between sessions or reaching resources it shouldn’t.

AWS Lambda MicroVMs provide ephemeral, VM-isolated compute environments that align with the per-session execution model of AI agent sandboxes. Snapshot-based launch avoids application startup latency. Idle policies terminate VMs once sessions complete. Firecracker isolation verifies that sessions do not share state. You pay only for active execution time and maintain full control over credentials, networking, and governance within your AWS boundary.

You build and operate a serverless control plane. You get per-session VM isolation with no idle compute cost and no shared tenancy.

To get started, explore these resources:

AWS reimagines the getting started experience

Post Syndicated from Micah Walter original https://aws.amazon.com/blogs/aws/aws-reimagines-the-getting-started-experience/

Amazon Web Services (AWS) started with a handful of foundational infrastructure services such as Amazon Simple Storage Service (Amazon S3), Amazon Elastic Compute Cloud (Amazon EC2), and Amazon Simple Queue Service (Amazon SQS), so that anyone with an idea could start building. As the world’s largest companies and governments adopted AWS, they asked for features to optimize their configuration for a range of global business contexts, security requirements, and operational needs. To meet these needs, AWS expanded globally through new Regions and added breadth and depth of services in security, networking, governance, and cost controls, so those customers could operate wherever they needed and at the scale they require. That combination of global reach, breadth, and depth remains essential for those customers, but if you are at the start of a new idea, every configuration option is effort standing in the way of shipping your dream product fast.

Today, we’re announcing a new simplified experience on AWS for builders who are working at the pace of AI. Instead of having to complete configuration tasks before you can work on your project, you start with sensible defaults and simple administration. You sign up using an existing identity from providers including Google, GitHub, and Apple. For most new customers, no credit card is required to start and you receive $100 in free credits as part of the AWS Free Tier. You can build immediately in your first project. As you continue to work, you can invite collaborators with just an email address, without learning about AWS Identity and Access Management (IAM) or AWS IAM Identity Center. When your project grows beyond the free credits, you can set a spend limit so you stay within your budget on the paid plan. If you grow to need additional customization, you can activate advanced AWS features to access the full breadth and depth of AWS without migrating.

How it works
When you sign up, AWS organizes your work in a project. A project contains an AWS account, where you create resources, and settings for sharing with team members. AWS creates that structure for you and applies additional security controls so you can start building your idea. After signing in, you get a prompt to paste into your coding agent that configures it to work with your new AWS environment. From there, your agent can deploy resources, run workloads, and iterate on your application following best practices for working with AWS.

You can create another project with a click. When you want to work with an additional team member, you send an invitation to their email address. Identity permissions are handled for you, so there are no IAM users to create; each person you invite only gets access to the projects you specify. Console workflows and coding agents also configure permissions between supported services and resources automatically, so you do not have to set up or troubleshoot resource permissions by hand.

When you’re ready to move beyond free credits, you can upgrade to a paid plan by entering your payment method. You can set a monthly spend limit on a project based on your usage trends, starting at $20 per month. The spend limit is the ceiling for that project’s costs, and you pay for what you actually use up to that amount. For example, if you set a $50 spend limit and your project incurs $32 in charges that month, you pay $32 (plus taxes). AWS will suggest a spend limit based on your usage, and you can accept that recommendation or set a custom amount if you are planning to further scale your usage. If your project approaches the limit, you first receive notifications. If spend reaches the limit, AWS pauses your project rather than accumulating charges, and you can resume working on it when you raise the limit. Each project has its own spend limit so you can give a larger budget to a workload that is gaining traction while keeping a smaller budget on an experimental idea.

Let’s try it out
To get started, I went to aws.amazon.com and chose Create account. I signed in with my Google account and within seconds had a new project ready to go, as shown in the following screenshot.

The Sign up for AWS page, with options to continue with email or sign in using Google, GitHub, Apple, or Amazon.

The first thing I saw was a prompt to configure my coding agent. I copied the prompt and pasted it into my agent. The agent set up the AWS Command Line Interface (AWS CLI) and the Agent Toolkit for AWS, logged me into AWS, and created a CLAUDE.md file in my project with guidance for the new experience.

The Setup Agent Toolkit for AWS dialog, with a prompt to copy and paste into your coding agent.

With the agent connected, I gave it a short prompt: build an API that returns a new unique sequential ID on every request. The agent created an AWS Lambda function, an Amazon DynamoDB table, and an Amazon API Gateway API, then deployed them for me. I did not have to configure resource permissions by hand. Within a few minutes I had a public endpoint that returned a newly minted ID on each request. My project started with $100 in free credits, and I received an additional $20 when the Lambda function was deployed.

A coding agent prompt to build an API that returns a unique sequential ID on every request.

The coding agent presents architecture options for the sequential ID API, with AWS Lambda and Amazon DynamoDB selected.

The coding agent confirms the API is live and lists the Amazon DynamoDB table, AWS Lambda function, and Amazon API Gateway API it deployed.

From the project, I could manage settings, invite team members by email, and monitor billing, as shown in the following screenshots.

The Projects page, showing remaining free-plan days, credits, and a project.

The project Members page, with the option to invite a new team member by email.

The Billing page, showing a $0.00 balance on the free plan, remaining credits, and cost by project.

Activating advanced features
If you reach the point where you need multiple Regions, or governance features like custom policies in AWS Organizations, you can activate advanced features at no additional cost. You’ll find yourself in a fully configured AWS Organization built according to best practices, with no migration and no downtime. Everything you configured previously is preserved and reflected in the underlying AWS services.

Now rolling out
We’ve heard from builders that they do not want to spend their first hours configuring an AWS environment. They want to build what they came to build, and we listened. AWS began as a place where anyone with an idea could start building, and this new simplified experience brings that starting point back, with sensible defaults so you can begin immediately, and with the global reach, breadth, and depth of AWS still there when your idea needs it. We are gradually rolling this experience out to new customers. We cannot wait to see what you build, and we want your feedback on the experience.

To try the new experience, create a new AWS account. To learn more, see Sign up for AWS (new).

Connect Amazon SageMaker Unified Studio to Microsoft Power BI – Part 1: IAM Identity Center (IDC)-based domains

Post Syndicated from Ramesh H Singh original https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-1-iam-identity-center-idc-based-domains/

Connecting Power BI to your Amazon SageMaker Unified Studio data catalogs typically required third-party bridges. These bridges added complexity and licensing costs. In this post, you create a direct connection using new authentication modes in the Amazon Athena ODBC driver, removing those dependencies entirely. If your organization uses Power BI as its business intelligence (BI) tool, your analysts can configure access to governed data in Amazon SageMaker Unified Studio without changing their tools or workflows. As an AWS alternative, Amazon Quick Sight provides serverless BI integration with Amazon SageMaker Unified Studio at pay-per-session pricing.

A previous post showed the connection method using a third-party ODBC-JDBC bridge. The Amazon Athena ODBC driver (version 2.2.0 and later) now supports Amazon SageMaker Unified Studio authentication directly, eliminating the need for customers to configure third-party bridge components previously required for this connection. This bridge also created additional components and required ongoing maintenance. The native connection simplifies the architecture by reducing these requirements.

UC Irvine, a top-ten U.S. public research university, consolidates student data from systems across multiple departments into a single governed repository that supports reporting, research, and analytics for decision-making at the strategic, tactical, and operational levels. Many of their analysts rely on Power BI to explore and visualize this governed data.

“Our users rely on Power BI for data visualization and reporting, but connecting to governed data in AWS previously required workarounds. The ODBC connection feature gives a direct path from Power BI into our SageMaker Unified Studio projects—no bridge software, no extra licensing, just a connection string and we’re ready to go.”

— Bernadette Theologidy, Manager, Student Analytics, UC Irvine

The Athena ODBC driver introduces two new authentication modes for SageMaker Unified Studio:

  1. SageMakerBrowserIdc (for IDC-based domains): The driver opens a browser window and authenticates through AWS IAM Identity Center (and your external identity provider, if configured). No local AWS credentials are needed.
  2. SageMakerIam (for AWS Identity and Access Management (IAM)-based and IDC-based domains): The driver uses AWS credentials from the default credential provider chain. For this walkthrough, we use AWS IAM Identity Center to provide those credentials.

You connect Microsoft Power BI to Amazon SageMaker Unified Studio through Athena. The Athena ODBC driver supports using two connection methods that use these authentication modes:

Method 1: DSN-based (Athena Power BI connector): You configure an ODBC Data Source Name (DSN) and use the Athena connector in Power BI. This method supports DirectQuery and Import mode with both SageMakerBrowserIdc and SageMakerIam authentication.

Method 2: DSN-less (Power BI ODBC connector): You use the Power BI ODBC connector with a connection string, requiring no DSN configuration. This method supports Import mode only with SageMakerIam authentication. DirectQuery isn’t available because the Power BI ODBC connector doesn’t support it. The connection string in Power BI Desktop must match exactly the one on Power BI Service. Because the gateway runs as a Windows service without interactive browser access, both ends must use SageMakerIam.

Feature Method 1: DSN-based Method 2: DSN-less
Power BI Connector Amazon Athena connector ODBC connector
Data connectivity mode DirectQuery and Import Import only
Requires DSN configuration Yes No
Data freshness Real-time (DirectQuery) or scheduled (Import) Scheduled refresh only
Authentication types SageMakerIam and SageMakerBrowserIdc SageMakerIam only
Domain types supported IAM-based and IDC-based IAM-based and IDC-based
Best for Dashboards requiring live data Scenarios where DSN management is not possible or scheduled refresh is acceptable

This is Part 1 of a two-part series. This post covers IDC-based domains using both connection methods. Part 2 covers IAM-based domains.

Solution overview

In this walkthrough, you take the role of a data analyst at an energy company. You need to understand the current state and future direction of the U.S. power generation fleet using the Public Utility Data Liberation Project, available on the Registry of Open Data on AWS. Our goal is to analyze generation capacity and identify where new investment is flowing. We connect Power BI to Athena through Amazon SageMaker Unified Studio and query the EIA-860 generators dataset directly from our data catalog. The result is a single visualization that reveals the energy transition.

The following diagram illustrates the solution architecture for connecting Power BI to Amazon SageMaker Unified Studio through Amazon Athena.

Architecture diagram showing Power BI connecting to Amazon Athena through Amazon SageMaker Unified Studio, with a Microsoft on-premises data gateway on Amazon EC2

Figure 1: Architecture diagram

The following architecture demonstrates a six-step workflow.

  1. Data engineers and analysts connect Power BI Desktop to Athena as a data source.
  2. They build their reports locally.
  3. They then publish them to the Power BI Service.
  4. Microsoft On-Premises Data Gateway on an Amazon Elastic Compute Cloud (Amazon EC2) instance connects to Athena using the instance’s attached IAM role.
  5. The Power BI Service then uses this gateway connection.
  6. Report viewers access the published reports through Power BI Service to make data-driven decisions.

On the AWS side, Athena queries the data catalog managed by AWS Glue Data Catalog. The catalog references data stored in Amazon Simple Storage Service (Amazon S3). An Amazon SageMaker Unified Studio project governs all access.

In an IDC-based domain (covered in this post), Power BI Desktop uses SageMakerBrowserIdc for Method 1 and SageMakerIam for Method 2. Power BI Desktop can run on-premises or on an EC2 instance. The gateway always uses SageMakerIam (it runs as a Windows service without browser access) and authenticates using instance profile credentials, which rotate automatically. The gateway can only query data within projects where its IAM role has been added as a member. For IAM-based domains, see Part 2.

Prerequisites

Before connecting Power BI to Amazon SageMaker Unified Studio, verify that your environment meets these requirements:

  • Athena ODBC driver – The latest Amazon Athena ODBC driver (version 2.2.0 or more recent) for Windows 64-bit.
  • Microsoft Power BI Desktop – The latest version installed on your Windows machine.
  • Microsoft Power BI Pro License – Required for publishing reports and configuring the on-premises data gateway.
  • Microsoft Power BI on-premises data gateway – The latest version installed on the EC2 instance.
  • Amazon SageMaker Unified Studio – An Amazon SageMaker Unified Studio IDC-based domain.

You need an Amazon SageMaker Unified Studio project with data assets. For detailed instructions, refer to the Amazon SageMaker Unified Studio User Guide.

The following screenshot shows the Amazon SageMaker Unified Studio project Query Editor interface, which runs a preview query against the EIA-860 generators dataset.

SageMaker Unified Studio Query Editor previewing the EIA-860 generators dataset

Figure 2: SageMaker Unified Studio project with the EIA-860 generators dataset available in the data catalog

Method 1: DSN-based connection (Athena Power BI connector)

This method uses the Amazon Athena Power BI connector with an ODBC Data Source Name (DSN), supporting DirectQuery and Import mode.

You configure Power BI Desktop to connect to your data assets in Amazon SageMaker Unified Studio using the SageMakerBrowserIdc authentication mode. The driver opens a browser window and authenticates through IAM Identity Center (and your external identity provider, if configured).

Add your SSO user as a member of your SageMaker Unified Studio project

Your single sign-on (SSO) user needs project-level access to query data with Athena. Verify your user is listed as a project member or add it by following Add project members in the Amazon SageMaker Unified Studio User Guide.

The following screenshot shows the SageMaker Unified Studio project user management page, where project owners can add or remove project users and roles.

SageMaker Unified Studio project members page listing users and roles

Figure 3: Members of a SageMaker Unified Studio project

Gather configuration values to configure your Amazon Athena ODBC DSN

Gather the following values from your Amazon SageMaker Unified Studio project:

  1. Open your Amazon SageMaker Unified Studio project.
  2. In the top right, select the three dots.
  3. Choose Project details.
  4. Select JDBC and ODBC details.
  5. Under ODBC connection details copy the following information: IDC issuer URL, domain ID, project ID, Athena workgroup name and AWS Region.

The following screenshot shows the Amazon SageMaker Unified Studio project overview page, where you can copy these details.

SageMaker Unified Studio project overview showing ODBC connection details

Figure 4: ODBC connection details

Configure the ODBC DSN

Create a System DSN using the Amazon Athena ODBC driver. For the general DSN creation steps, see Configuring a data source name on Windows in the Amazon Athena User Guide.

Enter the following values:

Field Value
Data Source Name Name your datasource (for example, pbi-idcdomain)
Region The AWS Region where your Amazon SageMaker domain is provisioned (for example, us-east-1)
Catalog AwsDataCatalog
Database default
Workgroup Your Athena workgroup name (for example, workgroup-abcdefghij-klmexample)

In the Authentication Options, configure the following values:

Field Value
Authentication Type SageMakerBrowserIdc
SSO Start URL IAM Identity Center entry point (for example, https://identitycenter.amazonaws.com/ssoins-0example)
SSO Region Region of IAM Identity Center (for example, us-east-1)
SageMaker Domain ID dzd-123456example
SageMaker Project ID abcd12example
SageMaker Domain Region Region of your Amazon SageMaker Unified Studio project (for example, us-east-1)

Choose OK, then Test to verify the connection. Choose Allow Access when prompted by the browser.

The following screenshot shows the consent prompt.

Browser consent prompt requesting access approval during authentication

Figure 5: Browser consent prompt

The following screenshot shows the successful connection test.

ODBC DSN configuration showing a successful connection test with SageMakerBrowserIdc

Figure 6: Successful connection test in the ODBC DSN configuration with SageMakerBrowserIdc authentication

Connect Power BI Desktop to your data

With the DSN configured, you can connect Power BI Desktop to your data catalog and load the generators dataset.

  1. Open Power BI Desktop.
  2. Open the Get Data menu and select More.
  3. Search for and select Amazon Athena and choose Connect.
  4. For Data Source Name (DSN), enter pbi-idcdomain.
  5. Select DirectQuery.
  6. Choose OK.
  7. Choose Use Data Source Configuration and then Connect.
  8. In the AwsDataCatalog folder, navigate to your database.
  9. Select the core_eia860__scd_generators table.
  10. Choose Load.

The following screenshot shows Power BI Desktop successfully connected to the AWS data catalog.

Power BI Desktop connected to the data catalog with the generators table loaded

Figure 7: Power BI Desktop connected to the data catalog with the generators table loaded using SageMakerBrowserIdc authentication

Create your dashboard and publish it

You can create a dashboard to visualize U.S. power generation data. To create a visualization, complete the following steps:

  1. In the Visualizations pane, choose the Stacked bar chart.
  2. Assign the Y-Axis: Drag technology_description to the Y-Axis.
  3. Assign the X-Axis (Values): Drag capacity_mw to the X-Axis (automatically summed).
  4. Assign the Legend (Stack): Drag operational_status to the Legend field.
  5. Choose Publish.
  6. Give your report a name (for example, generation-idcdomain) and choose Save.
  7. Sign in and choose a destination workspace.
Power BI Desktop stacked bar chart of generation capacity by technology and operational status

Figure 8: Power BI Desktop report using the EIA-860 generators dataset

After publishing, the report structure is available on Power BI Service.

Method 2: DSN-less connection (Power BI ODBC connector)

In this method, you use the Power BI ODBC connector with a connection string (no DSN required). This method supports Import mode only and SageMakerIam authentication. Because the gateway cannot perform browser authentication, both Desktop and gateway must use SageMakerIam. If your workflow requires SageMakerBrowserIdc, use Method 1.

If your machine already has AWS credentials through another method in the default credential provider chain, skip the following setup.

Administrator setup

Create a custom permission set named SageMakerDataAnalyst in IAM Identity Center with the following inline policy. For detailed steps, see Create a permission set in the AWS IAM Identity Center User Guide.

{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "SageMakerAccess",
            "Effect": "Allow",
            "Action": [
                "datazone:GetConnection",
                "datazone:ListConnections",
                "datazone:GetDomain",
                "datazone:GetProject"
            ],
            "Resource": "*"
        },
        {
            "Sid": "STSForDriver",
            "Effect": "Allow",
            "Action": [
                "sts:GetCallerIdentity"
            ],
            "Resource": "*"
        }
    ]
}

Assign your user to this permission set for the AWS account containing your SageMaker Unified Studio domain. Then configure your AWS Command Line Interface (AWS CLI) SSO profile by running aws configure sso. For the full CLI configuration walkthrough with detailed steps, see Part 2. After your profile is configured, run aws sso login to authenticate.

Add the IAM identity as a member of SageMaker Unified Studio project

The IAM identity providing credentials needs both domain-level and project-level access to query data through Athena.

  1. Add AWSReservedSSO_SageMakerDataAnalyst_1234example as a domain IAM user: see Managing users in the Amazon SageMaker Unified Studio Admin Guide. Choose Current account.
SageMaker Unified Studio domain users list including the IAM identity

Figure 9: List of users of your SageMaker Unified Studio domain including the IAM identity

  1. Add AWSReservedSSO_SageMakerDataAnalyst_1234example as a project member: see Add project members in the Amazon SageMaker Unified Studio User Guide.
SageMaker Unified Studio project members list including the IAM identity

Figure 10: Members of a SageMaker Unified Studio project including the IAM identity

Gather configuration values

Gather the following connection values from your Amazon SageMaker Unified Studio project:

  1. Open your Amazon SageMaker Unified Studio Project.
  2. On the navigation pane, choose Overview.
  3. Select JDBC and ODBC details.
  4. Select the Using IAM auth toggle.
  5. Copy the ODBC connection string.
SageMaker Unified Studio project overview showing the ODBC connection string for IAM auth

Figure 11: ODBC connection string on the SageMaker Unified Studio project overview

Connect Power BI Desktop to your data and publish

With the configuration parameters of your project, you can connect Power BI Desktop to your data catalog and load the generators dataset.

  1. Open Power BI Desktop.
  2. Open the Get Data menu and select More.
  3. Search for and select ODBC and choose Connect.
  4. For Data Source Name (DSN), select (None).
  5. Expand Advanced Options.
  6. In the Connection string field, enter your connection string. For example, Driver={Amazon Athena ODBC (x64)};AwsRegion=us-east-1;Catalog=AwsDataCatalog;Schema=default;Workgroup=workgroup-abcdefghij-klmexample;SageMakerDomainId= dzd-123456example;SageMakerProjectId= abcd12example;SageMakerDomainRegion=us-east-1;AuthenticationType=SageMakerIam;
  7. Choose OK.
  8. Choose Default or Custom and then Connect.
  9. In the AwsDataCatalog folder, navigate to your database.
  10. Select the core_eia860__scd_generators table.
  11. Choose Load.

When publishing, name your report generation-idcdomain-dsnless.

Configure the on-premises data gateway and view your report on Power BI Service

After creating your reports in Power BI Desktop, configure the on-premises data gateway to view your report on Power BI Service.

You can configure the gateway using either a DSN or a DSN-less connection string, matching the method you used in Power BI Desktop.

Create and attach an IAM role to the Power BI Gateway EC2 instance

Create an IAM role for the EC2 instance that will host your Power BI gateway. Name the role pbi-gateway-role (or a name of your choice). The role must use EC2 as the trusted entity and include the following inline policy:

{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "SageMakerAccess",
            "Effect": "Allow",
            "Action": [
                "datazone:GetConnection",
                "datazone:ListConnections",
                "datazone:GetDomain",
                "datazone:GetProject"
            ],
            "Resource": "*"
        },
        {
            "Sid": "STSForDriver",
            "Effect": "Allow",
            "Action": [
                "sts:GetCallerIdentity"
            ],
            "Resource": "*"
        }
    ]
}

Attach this role to your Power BI Gateway EC2 instance. For detailed steps on creating and attaching an IAM role to an EC2 instance, refer to IAM roles for Amazon EC2 in the Amazon EC2 User Guide.

Add the Power BI Gateway IAM role as a member of SageMaker Unified Studio project

The gateway IAM role needs project-level access to query data through Athena.

  1. Add the IAM pbi-gateway-role role as a domain IAM user: see Managing users in the Amazon SageMaker Unified Studio Admin Guide. Choose Current account (or Associated account if your gateway is deployed in a different account).

The following screenshot, from the Amazon SageMaker page of the AWS Management Console, shows the list of users of your Amazon SageMaker Unified Studio domain, including the IAM gateway role.

SageMaker Unified Studio domain users list including the Power BI gateway IAM role

Figure 12: List of users of your SageMaker Unified Studio domain including the IAM gateway role

Add the IAM pbi-gateway-role role as a project member: see Add project members in the Amazon SageMaker Unified Studio User Guide.

The following screenshot shows the Amazon SageMaker Unified Studio project user management page listing the project members.

SageMaker Unified Studio project members list including the Power BI gateway IAM role

Figure 13: Members of a SageMaker Unified Studio project including the IAM gateway role

Configure the data source on Power BI Gateway

How you configure the data source depends on the method you used in Power BI Desktop.

Method 1 (DSN-based)

Configure a System DSN on the gateway EC2 instance following the same ODBC DSN steps described in Method 1. When configuring, make sure that:

  • You use the System DSN tab (not User DSN) because the gateway runs as a Windows service under a separate account.
  • The authentication type is set to SageMakerIam regardless of what you used on Desktop.
  • The DSN name matches exactly the one configured on Power BI Desktop (for example, pbi-idcdomain)

Method 2 (DSN-less)

No configuration is needed on the gateway machine itself. You configure the data source directly in Power BI Service.

Configure the data source and view your report on Power BI Service

To view your report, complete the following steps:

  1. Open the workspace where you saved your report.
  2. Search the Semantic Model which has the same name as your report (for example, generation-idcdomain) and choose the More options icon (three dots).
  3. Choose Settings.
  4. Expand Gateway and Cloud Connection.
  5. Choose View Datasources (play icon) on your gateway.
  6. Choose Manually add to gateway.
  7. Add a connection name (for example, pbi-idcdomain).

The next step depends on the method that you chose:

Method 1 (DSN-based)

  1. Add the DSN (for example, pbi-idcdomain) that matches exactly the one configured on Power BI Desktop.

Method 2 (DSN-less)

  1. In the Connection string field, enter the connection string that matches exactly the one used in Power BI Desktop.

Next, continue with the configuration:

  1. Select Anonymous as Authentication Method.
  2. Choose Create.
  3. Expand again Gateway and Cloud Connection.
  4. For Maps to, choose the connection that you created (for example, pbi-idcdomain).
  5. Choose Apply.
  6. Return to the workspace where you saved your report.
  7. On the Content section, choose your report (for example, generation-idcdomain).

The following screenshot shows a Power BI report on Power BI Service.

Published Power BI report rendering on Power BI Service

Figure 14: Power BI report on Power BI Service

You can now see your report online with the data from your Amazon SageMaker Unified Studio project.

Clean up

To avoid additional charges after testing, delete the Amazon SageMaker Unified Studio domain and EC2 instances. Refer to Delete domains and Terminate Instances for instructions.

Conclusion

In this post, you connected Microsoft Power BI to Amazon SageMaker Unified Studio using an IDC-based domain with both DSN-based and DSN-less methods. This provides a direct connection, with no third-party licensing, that maintains data governance. In Part 2, we cover IAM-based domains.

You can automate many steps of this process. For information about automating DSN creation on the Power BI Gateway or Service, refer to How ENGIE automates the deployment of Amazon Athena data sources on Microsoft Power BI. If you don’t want users adding the gateway IAM role directly, you can create a custom blueprint as a self-service tool for gateway role addition. The blueprint uses a ProjectMembership resource with a configurable parameter that project owners can activate at project creation, automatically adding the gateway role as a project contributor.

For additional best practices, refer to the Using Microsoft Power BI with the AWS Cloud Whitepaper. To learn more, visit Amazon SageMaker Unified Studio and Amazon Athena.


About the authors

Ramesh Singh

Ramesh Singh

Ramesh is a Senior Product Manager Technical (External Services) at AWS in Seattle, Washington, currently with the Amazon SageMaker team. He is passionate about building high-performance ML/AI and analytics products that help enterprise customers achieve their critical goals.

Armando Segnini

Armando Segnini

Armando is a Senior Analytics Specialist Solutions Architect at AWS, partnering with enterprise customers to architect scalable data, analytics, and AI platforms. He helps organizations turn complex data challenges into business value through expertise in streaming, BI integration, and generative AI. Outside of work, Armando enjoys traveling with his family, exploring new cultures, photography, and functional fitness competitions.

Gaurav Sharma

Gaurav is a Specialist Solutions Architect (Analytics) at AWS, supporting US public sector customers on their cloud journey. Outside of work, Gaurav enjoys spending time with his family and reading books.

Krishna Atluru

Krishna Atluru

Krishna is an Enterprise Support Lead TAM at AWS. He provides customers with in-depth guidance on improving security posture and operational excellence for their workloads, helping them build secure, resilient, and cost-effective solutions. His areas of expertise include building serverless architectures, and data and analytics solutions. Outside of work, Krishna enjoys cooking, swimming, and traveling.

Saushthav Saxena

Saushthav Saxena

Saushthav is a Software Development Engineer at AWS on the Amazon Athena team, where he has spent the past few years working on distributed systems and data analytics at scale. Based in the San Francisco Bay Area, his background spans full-stack development, high performance computing, and large-scale infrastructure. Outside of work, he enjoys reading sci-fi novels, swimming, and traveling with family and friends.

Connect Amazon SageMaker Unified Studio to Microsoft Power BI – Part 2: IAM-based domains

Post Syndicated from Ramesh H Singh original https://aws.amazon.com/blogs/big-data/connect-amazon-sagemaker-unified-studio-to-microsoft-power-bi-part-2-iam-based-domains/

In Part 1 of this series, we connected Microsoft Power BI to Amazon SageMaker Unified Studio using an IAM Identity Center (IDC)-based domain. The Amazon Athena ODBC driver (version 2.2.0 and later) supports Amazon SageMaker Unified Studio authentication natively, removing the third-party ODBC-JDBC bridge previously required. We walked through both the DSN-based connection and the DSN-less connection, from Power BI Desktop through the on-premises data gateway to Power BI Service, where report viewers access published dashboards.

In this post, you create the same direct connection using an AWS Identity and Access Management (IAM)-based domain. The walkthrough covers the same two connection methods. The differences are the Amazon SageMaker Unified Studio console navigation paths, the configuration values, and an additional administrator setup that provides AWS credentials through AWS IAM Identity Center. This is Part 2 of a two-part series. For a detailed comparison of the two connection methods, see Part 1.

Solution overview

The architecture is the same as the previous post (see the architecture diagram and walkthrough scenario in Part 1). Power BI Desktop connects to Amazon Athena through the ODBC driver and the Amazon SageMaker Unified Studio project governs all data access. At the same time, the on-premises data gateway on an Amazon Elastic Compute Cloud (Amazon EC2) instance bridges the connection to Power BI Service so report viewers can access published dashboards.

The difference is in authentication: An IAM-based domain uses SageMakerIam authentication for both connection methods. The driver retrieves credentials from the AWS default credential provider chain. For this walkthrough, AWS IAM Identity Center provides those credentials through a custom permission set. Power BI Desktop can run on-premises or on an EC2 instance in the AWS Cloud. The gateway EC2 instance authenticates using its attached IAM role.

Prerequisites

Complete the prerequisites from Part 1. Additionally, you need:

  • AWS Command Line Interface (AWS CLI) – The latest version of the AWS CLI installed on your Windows machine. In this post series, the ODBC driver uses the AWS IAM Identity Center profile configured through the CLI for authentication.
  • Amazon SageMaker Unified Studio – An Amazon SageMaker Unified Studio IAM-based domain with AWS IAM Identity Center single sign-on (SSO) enabled.

The following screenshot shows the Amazon SageMaker Unified Studio (IAM-based domain) project Query Editor interface. It runs a preview query on the EIA-860 generators dataset.

SageMaker Unified Studio Query Editor previewing the EIA-860 generators dataset in an IAM-based domain

Figure 1: SageMaker Unified Studio (IAM-based domain) project with the EIA-860 generators dataset available in the data catalog

Administrator setup

This section configures AWS IAM Identity Center to provide credentials for the SageMakerIam authentication mode. It applies to Method 1 (IAM-based domain) and Method 2 (both domain types). If your machine already has AWS credentials available through another method in the default credential provider chain, you can skip this section and proceed directly to the method of your choice. For the full list of credential sources, refer to Credential providers in the AWS SDKs and Tools Reference Guide.

Create a permission set in IAM Identity Center

Create a custom permission set named SageMakerDataAnalyst in IAM Identity Center with the following inline policy. For detailed steps, see Create a permission set in the AWS IAM Identity Center User Guide.

{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "SageMakerAccess",
            "Effect": "Allow",
            "Action": [
                "datazone:GetConnection",
                "datazone:ListConnections",
                "datazone:GetDomain",
                "datazone:GetProject"
            ],
            "Resource": "*"
        },
        {
            "Sid": "STSForDriver",
            "Effect": "Allow",
            "Action": [
                "sts:GetCallerIdentity"
            ],
            "Resource": "*"
        }
    ]
}

The "Resource": "*" is required because these API actions do not support resource-level permissions. For more information, see Actions, resources, and condition keys for Amazon DataZone.

This doesn’t grant broad access to your data. These are read-only metadata actions that allow the ODBC driver to discover connection details and retrieve temporary Athena credentials. The actual data access is governed by Amazon SageMaker Unified Studio project membership: Users can only query data within projects where they have been explicitly added as members. The Amazon SageMaker Unified Studio project IAM role provides Athena and Amazon S3 permissions separately.

Assign users to the permission set

To assign users or groups to the target AWS account, complete the following steps:

  1. In the IAM Identity Center console, choose AWS accounts.
  2. Select the target account where your Amazon SageMaker Unified Studio IAM-based domain is deployed.
  3. Choose Assign users or groups.
  4. Select the SSO users or groups that need access.
  5. Select the SageMakerDataAnalyst permission set.
  6. Choose Submit.

Configure AWS IAM Identity Center profile

To configure the AWS IAM Identity Center profile, run the following command in your terminal on Windows:

aws configure sso

When prompted, enter the following values:

Prompt Value
SSO session name For example, smus
SSO start URL The IDC issuer URL. For example, https://identitycenter.amazonaws.com/ssoins-0example
SSO region The SSO Region. For example, us-east-1
SSO registration scopes sso:account:access

A browser window opens for authentication. After authentication, select your account and the SageMakerDataAnalyst role.

The following screenshots show the consent window and the successful authentication message.

Browser consent prompt requesting access approval during AWS CLI SSO authentication

Figure 2: Browser consent prompt

Browser page confirming successful AWS CLI SSO authentication

Figure 3: Browser authentication successful message

When prompted, enter the following values:

Prompt Value
Default client Region None
CLI default output format None
Profile Name Change value by default

The resulting ~/.aws/config file should look like the following:

[default]
sso_session = smus
sso_account_id = 1234example
sso_role_name = SageMakerDataAnalyst

[sso-session smus]
sso_start_url = https://identitycenter.amazonaws.com/ssoins-0example
sso_region = us-east-1
sso_registration_scopes = sso:account:access

Verify authentication and daily use

To verify that your SSO profile is working correctly, run the following command:

aws sts get-caller-identity

You should receive a response like the following:

{
    "UserId": "AROARHJJNFBQD6EXAMPLE:[email protected]",
    "Account": "111122223333",
    "Arn": "arn:aws:sts::111122223333:assumed-role/AWSReservedSSO_SageMakerDataAnalyst_1234example/[email protected]"
}

For daily use, no passwords or EC2 instance roles are required. When your SSO session expires, run the following command to quickly refresh it:

aws sso login

Add your IAM identity as a member of your Amazon SageMaker Unified Studio project

The IAM identity providing credentials to the ODBC driver needs project-level access to query data through Athena. If you completed the administrator setup, this is the SSO role associated with your permission set (for example, AWSReservedSSO_SageMakerDataAnalyst_1234example). If you’re using another credential source, add the IAM role or user that provides those credentials. For detailed steps, see Managing users for IAM-based domains in the Amazon SageMaker Unified Studio Administrator Guide.

The following screenshot shows the Amazon SageMaker Unified Studio domain management page, which lists the members in a project.

SageMaker Unified Studio project members list

Figure 4: List of members of your SageMaker Unified Studio project

Gather the information to authenticate

To get the parameters that you need to authenticate, complete these steps:

  1. Open your Amazon SageMaker Unified Studio Project.
  2. Open Domain Management.
  3. Choose Users.
  4. Choose View SSO connection.
  5. Copy the end of the Instance ARN, so we can build the Instance URL like https://identitycenter.amazonaws.com/ssoins-0example

The following screenshot shows the Amazon SageMaker Unified Studio domain management page with SSO connection details.

SageMaker Unified Studio domain SSO connection details showing the IAM Identity Center instance ARN

Figure 5: AWS IAM Identity Center information

  1. Choose the user icon and copy the Region as shown in the following screenshot.
SageMaker Unified Studio user menu showing the Region

Figure 6: User icon with the Region information

Method 1: DSN-based connection (Athena Power BI connector)

In this method, you configure an ODBC Data Source Name (DSN) and use the Amazon Athena connector in Power BI. This method uses SageMakerIam authentication mode and supports both DirectQuery and Import mode.

This section covers IAM-based domains. For IDC-based domains, see Part 1.

Gather configuration values to configure your Amazon Athena ODBC DSN

Before configuring the ODBC DSN, gather the following connection values from your Amazon SageMaker Unified Studio project:

  1. Open your Amazon SageMaker Unified Studio Project.
  2. Top right, select the three dots.
  3. Choose Project details.
  4. Select JDBC and ODBC details.
  5. Copy the following values: domain ID, Amazon SageMaker project ID, AWS Region, and Athena workgroup.

The following screenshot shows the Amazon SageMaker Unified Studio project overview page, which provides the project details to copy.

SageMaker Unified Studio project details showing domain ID, project ID, Region, and Athena workgroup

Figure 7: Project details with SageMaker domain ID, SageMaker project ID, Region, and Athena workgroup

Configure the ODBC DSN

Create a System DSN using the Amazon Athena ODBC driver. For the general DSN creation steps, see Configuring a data source name on Windows in the Amazon Athena User Guide. Enter the following values:

Field Value
Data Source Name Name your datasource (for example, pbi-iamdomain)
Region The AWS Region where your Amazon SageMaker domain is provisioned (for example, us-east-1)
Catalog AwsDataCatalog
Database default
Workgroup Your Athena workgroup name (for example, workgroup-abcdefghij-klmexample)

In the Authentication Options, configure the following values:

Field Value
Authentication Type SageMakerIam
SageMaker Domain ID dzd-123456example
SageMaker Project ID abcd12example
SageMaker Region Region of your SageMaker Unified Studio project (for example, us-east-1)

Choose OK, then Test to verify the connection. Choose Allow Access when prompted by the browser.

The following screenshot shows the successful connection test.

ODBC DSN configuration showing a successful connection test with SageMakerIam

Figure 8: Successful connection test in the ODBC DSN configuration with SageMakerIam authentication

Connect Power BI Desktop to your data

With the DSN configured, you can connect Power BI Desktop to your data catalog and load the generators dataset.

  1. Open Microsoft Power BI Desktop.
  2. Open the Get Data menu and select More.
  3. Search for and select Amazon Athena and choose Connect.
  4. For Data Source Name (DSN), enter pbi-iamdomain.
  5. Select DirectQuery.
  6. Choose OK.
  7. Choose Use Data Source Configuration and then Connect.
  8. In the AwsDataCatalog folder, navigate to your database.
  9. Select the core_eia860__scd_generators table.
  10. Choose Load.

The following screenshot shows Power BI Desktop successfully connected to the data catalog.

Power BI Desktop connected to the data catalog with the generators table loaded

Figure 9: Power BI Desktop connected to the data catalog with the generators table loaded using SageMakerIam authentication

Create your dashboard and publish it

You can create a dashboard to visualize U.S. power generation data. To create a visualization, complete the following steps:

  1. In the Visualizations pane, choose the Stacked bar chart.
  2. Assign the Y-Axis: Drag technology_description to the Y-Axis.
  3. Assign the X-Axis (Values): Drag capacity_mw to the X-Axis (automatically summed).
  4. Assign the Legend (Stack): Drag operational_status to the Legend field.
  5. Choose Publish.
  6. Give your report a name (for example, generation-iamdomain) and choose Save.
  7. Sign in and choose a destination workspace.

The following screenshot shows the Power BI dashboard with U.S. power generation data.

Power BI stacked bar chart of U.S. generation capacity by technology and operational status

Figure 10: Power BI dashboard with U.S. power generation data

After you publish, the report structure becomes available on Microsoft Power BI Service.

Method 2: DSN-less connection (Power BI ODBC connector)

In this method, you use the Power BI ODBC connector with a connection string (no DSN required). This method supports Import mode only and SageMakerIam authentication. Because the gateway can’t perform browser authentication and connection strings need to match, both Desktop and gateway must use SageMakerIam.

This section covers IAM-based domains. For IDC-based domains, see Part 1.

Gather configuration values to configure your DSN-less connection

Gather the following connection values from your Amazon SageMaker Unified Studio project:

  1. Open your Amazon SageMaker Unified Studio Project.
  2. Top right, select the three dots.
  3. Choose Project details.
  4. Select JDBC and ODBC details.
  5. Copy the ODBC connection string.

The following screenshot shows the Amazon SageMaker Unified Studio project overview page with the ODBC connection string to copy.

SageMaker Unified Studio project overview showing the ODBC connection string

Figure 11: Project details with ODBC connection string

Connect Power BI Desktop to your data and publish

With the configuration parameters of your project, you can connect Power BI Desktop to your data catalog and load the generators dataset.

  1. Open Power BI Desktop.
  2. Open the Get Data menu and select More.
  3. Search for and select ODBC and choose Connect.
  4. For Data Source Name (DSN), select (None).
  5. Expand Advanced Options.
  6. In the Connection string field, enter your connection string. For example, Driver={Amazon Athena ODBC (x64)};AwsRegion=us-east-1;Catalog=AwsDataCatalog;Schema=default;Workgroup=workgroup-abcdefghij-klmexample;SageMakerDomainId= dzd-123456example;SageMakerProjectId= abcd12example;SageMakerDomainRegion=us-east-1;AuthenticationType=SageMakerIam;
  7. Choose OK.
  8. Choose Default or Custom and then Connect.
  9. In the AwsDataCatalog folder, navigate to your database.
  10. Select the core_eia860__scd_generators table.
  11. Choose Load.

When publishing, name your report generation-iamdomain-dsnless.

Configure the gateway and view your report on Power BI Service

After creating your reports in Power BI Desktop, configure the on-premises data gateway to view your report on Power BI Service.

You can configure the gateway using either a DSN or a DSN-less connection string, matching the method you used in Power BI Desktop.

Create and attach an IAM role to the Power BI Gateway EC2 instance

Create an IAM role for the EC2 instance that will host your Power BI gateway. Name the role pbi-gateway-role (or a name of your choice). The role must use EC2 as the trusted entity and include the following inline policy:

{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "SageMakerAccess",
            "Effect": "Allow",
            "Action": [
                "datazone:GetConnection",
                "datazone:ListConnections",
                "datazone:GetDomain",
                "datazone:GetProject"
            ],
            "Resource": "*"
        },
        {
            "Sid": "STSForDriver",
            "Effect": "Allow",
            "Action": [
                "sts:GetCallerIdentity"
            ],
            "Resource": "*"
        }
    ]
}

Attach this role to your Power BI Gateway EC2 instance. For detailed steps on creating and attaching an IAM role to an EC2 instance, refer to IAM roles for Amazon EC2 in the Amazon EC2 User Guide.

Add the Power BI Gateway IAM role as a member of SageMaker Unified Studio project

The gateway IAM role needs project-level access to query data through Athena. The steps to add the role differ depending on your domain type.

IAM-based domain

  1. Open your Amazon SageMaker Unified Studio Project.
  2. Open Domain Management.
  3. Choose your Project Name.
  4. Choose Members.
  5. Choose Add members.
  6. Select the IAM role of your Power BI gateway (for example, pbi-gateway-role).
  7. Choose Add.

The following screenshot shows the Amazon SageMaker Unified Studio project domain management page with options to add members to a project.

SageMaker Unified Studio project members list including the Power BI gateway IAM role

Figure 12: List of members of a SageMaker Unified Studio project with the IAM gateway role

Configure the data source on Power BI Gateway

How you configure the data source depends on the method you used in Power BI Desktop.

Method 1 (DSN-based)

Configure a System DSN on the gateway EC2 instance following the same ODBC DSN steps described in Method 1. When configuring, make sure that:

  • You use the System DSN tab (not User DSN) because the gateway runs as a Windows service under a separate account.
  • The authentication type is set to SageMakerIam.
  • The DSN name matches exactly the one configured on Power BI Desktop (for example, pbi-iamdomain).

Method 2 (DSN-less)

No configuration is needed on the gateway machine itself. You configure the data source directly in Power BI Service.

Configure the data source and view your report on Power BI Service

To view your report, complete the following steps:

  1. Open the workspace where you saved your report.
  2. Search the Semantic Model which has the same name as your report (for example, generation-iamdomain) and choose the More options icon (three dots).
  3. Choose Settings.
  4. Expand Gateway and Cloud Connection.
  5. Choose View Datasources (play icon) on your gateway.
  6. Choose Manually add to gateway.
  7. Add a connection name (for example, pbi-iamdomain).

The next step depends on the method that you chose:

Method 1 (DSN-based)

  1. Add the DSN (for example, pbi-iamdomain) that matches exactly the one configured on Power BI Desktop.

Method 2 (DSN-less)

  1. In the Connection string field, enter the connection string that matches exactly the one used in Power BI Desktop.

Next, continue with the configuration:

  1. Select Anonymous as Authentication Method.
  2. Choose Create.
  3. Expand again Gateway and Cloud Connection.
  4. For Maps to, choose the connection that you created (for example, pbi-iamdomain).
  5. Choose Apply.
  6. Return to the workspace where you saved your report.
  7. On the Content section, choose your report (for example, generation-iamdomain).

The following screenshot shows a report on Power BI Service.

Published Power BI report rendering on Power BI Service

Figure 13: Power BI report on Power BI Service

You can now see your report online with the data from your Amazon SageMaker Unified Studio project.

Clean up

To avoid additional charges after testing, delete the Amazon SageMaker Unified Studio domain and EC2 instances. Refer to Delete domains and Terminate Instances for instructions.

Conclusion

In this two-part series, you connected Power BI to Amazon SageMaker Unified Studio through Amazon Athena. Part 1 covered IDC-based domains. This post covered IAM-based domains using SageMakerIam authentication. This provides a direct connection path, with no third-party licensing, while maintaining data governance and security.

You can automate many steps of this process. For information about automating DSN creation on the Power BI Gateway or Service, refer to How ENGIE automates the deployment of Amazon Athena data sources on Microsoft Power BI. If you don’t want users adding the gateway IAM role directly, you can create a custom blueprint as a self-service tool for gateway role addition. The blueprint uses a ProjectMembership resource with a configurable parameter that project owners can activate at project creation, automatically adding the gateway role as a project contributor.

For additional best practices, refer to the Using Microsoft Power BI with the AWS Cloud Whitepaper. To learn more, visit Amazon SageMaker Unified Studio and Amazon Athena.


About the authors

Ramesh H Singh

Ramesh H Singh

Ramesh is a Senior Product Manager Technical at AWS in Seattle, focused on Amazon SageMaker. He’s passionate about building analytics and AI products that help enterprise customers unlock real value from their data. Away from work, he spends his time hiking with family and exploring spirituality. Connect with him on LinkedIn.

Armando Segnini

Armando Segnini

Armando is a Senior Analytics Specialist Solutions Architect at AWS, partnering with enterprise customers to architect scalable data, analytics, and AI platforms. He helps organizations turn complex data challenges into business value through expertise in streaming, BI integration, and generative AI. Outside of work, Armando enjoys traveling with his family, exploring new cultures, photography, and functional fitness competitions.

Gaurav Sharma

Gaurav is a Specialist Solutions Architect (Analytics) at AWS, supporting US public sector customers on their cloud journey. Outside of work, Gaurav enjoys spending time with his family and reading books.

Krishna Atluru

Krishna Atluru

Krishna is an Enterprise Support Lead TAM at AWS. He provides customers with in-depth guidance on improving security posture and operational excellence for their workloads, helping them build secure, resilient, and cost-effective solutions. His areas of expertise include building serverless architectures, and data and analytics solutions. Outside of work, Krishna enjoys cooking, swimming, and traveling.

Saushthav Saxena

Saushthav Saxena

Saushthav is a Software Development Engineer at AWS on the Amazon Athena team, where he has spent the past few years working on distributed systems and data analytics at scale. Based in the San Francisco Bay Area, his background spans full-stack development, high-performance computing, and large-scale infrastructure. Outside of work, he enjoys reading sci-fi novels, swimming, and traveling with family and friends.

AWS Security Reference Architecture: A deep dive into PCI DSS compliance

Post Syndicated from Avik Mukherjee original https://aws.amazon.com/blogs/security/aws-security-reference-architecture-a-deep-dive-into-pci-dss-compliance/

Amazon Web Services (AWS) is excited to announce the publication of the AWS Security Reference Architecture (AWS SRA) Payment Card Industry (PCI) Data Security Standard (DSS) Deep Dive. This new guide extends the core AWS SRA to provide prescriptive, architecture-level guidance for organizations that store, process, or transmit cardholder data on AWS.

Organizations subject to PCI DSS have long asked for a comprehensive reference that bridges the gap between general AWS security best practices and the specific technical and organizational controls required to achieve and maintain PCI DSS compliance. This guide answers that need by showing how AWS SRA patterns address PCI DSS intent from account scoping and network segmentation to encryption, logging, and access control.

What is the AWS SRA PCI DSS Deep Dive?

The AWS SRA is a holistic, prescriptive security architecture guide that describes how AWS security services fit together across a multi-account AWS environment. It’s built around a modular, three-tier web architecture and is intentionally designed to be adapted as needed. Not every workload needs every service, but the AWS SRA provides the full range of options and their architectural relationships.

This PCI DSS deep dive doesn’t replace the AWS SRA, it extends it:

  • Mapping AWS SRA account types to PCI DSS scoping boundaries showing which accounts are in scope, connected-to or security-impacting, or out-of-scope in a typical payment architecture.
  • Layering PCI-specific controls onto existing AWS SRA service configurations. For example, additional logging granularity, encryption requirements, or network restrictions that go beyond the AWS SRA baseline.

The architectural patterns and controls described in this guide apply equally to merchants and service providers.

Who should use the guide

The guide is intended for:

  • Security architects designing or extending an AWS multi-account landing zone for workloads under PCI DSS scope.
  • Compliance engineers mapping AWS controls to PCI DSS requirements during assessment.
  • Cloud platform teams building shared security services that must accommodate a cardholder data environment (CDE).
  • Qualified Security Assessors (QSAs) and Internal Security Assessors (ISAs) who want to understand how AWS SRA patterns address PCI DSS intent.

Key AWS SRA design principles for PCI DSS

The guide applies six foundational AWS SRA design principles that are particularly relevant to PCI DSS compliance:

  • Implement a strong identity foundation: Enforce least privilege, separation of duties, and centralized identity management. Eliminate reliance on long-term static credentials.
  • Enable traceability: Monitor, alert, and audit actions in real time. Integrate log and metric collection with automated investigation and response systems.
  • Apply security at all layers: Defense-in-depth with preventive and detective controls at edge, virtual private cloud (VPC), load balancing, compute, OS, application, and code layers.
  • Protect data in transit and at rest: Classify data by sensitivity and apply encryption, tokenization, and access control mechanisms.
  • Keep people away from data: Reduce or eliminate direct access to cardholder data through automation and tooling.
  • Prepare for security events: Establish incident management processes, run simulations, and implement automated detection and recovery.

How to use the guide

The AWS SRA PCI DSS deep dive can be consumed in two ways:

  • As a narrative: Read the guide from beginning to end, starting with the PCI DSS primer, through the architecture and account scoping model, to the detailed requirement mappings. This approach gives you a complete understanding of how AWS SRA and PCI DSS intersect.
  • As a reference: Navigate directly to specific PCI DSS requirements or AWS SRA account types relevant to your current project. The guide includes architecture diagrams, requirement mapping tables, and service-specific configurations that you can use independently.

The guide includes downloadable architecture diagrams and detailed control mapping tables that complement the narrative content, making it straightforward to reference during security reviews and PCI DSS assessments.

Next steps

Security is a journey, not a destination. Review the AWS SRA PCI DSS Deep Dive guide and begin mapping its patterns to your own cardholder data environment and then validate existing environments against SRA best practices using SRA verify.

If you need assistance, contact AWS Professional Services, your AWS account team, or the AWS Partner Network, who can work with you to translate the reference architecture into a customized AWS environment that you can then operate.

If you have feedback about this post, submit comments in the Comments section below. If you need assistance architecting or implementing a PCI DSS-compliant AWS environment, contact the AWS Security Assurance Services team.


Author

Avik Mukherjee

Avik is a Senior Security Solutions Architect with more than a decade of experience in IT governance, security, risk, and compliance across retail, financial, and technology industries. He’s one of the original authors of the AWS Security Reference Architecture and leads the effort to extend the AWS SRA to compliance frameworks.

Akanksha Chaturvedi

Akanksha is a Senior Security Assurance Consultant with over 10 years of specialized experience in risk-based security assessments and regulatory compliance across highly regulated industries. She is an expert practitioner in HIPAA, PCI-DSS, GDPR, FedRAMP, and IRAP frameworks, with demonstrated success in architecting and deploying enterprise security programs from conception through full implementation.

Nimesh Ravas

Nimesh Ravasa

Nimesh is a Senior Assurance Consultant at AWS who focuses on security assurance and compliance for cloud-centered architectures. He brings extensive experience in PCI DSS assessments and security architecture reviews, helping organizations build and maintain compliant environments on AWS. He is passionate about translating complex compliance requirements into actionable technical guidance.

Omner Barajas

Omner Barajas

Omner is a Senior Security Solutions Architect at AWS with deep expertise in designing secure, scalable architectures for regulated industries. He specializes in network security, identity management, and security automation, and works with customers across financial services and payments to implement defense-in-depth strategies aligned with industry standards.

Viktor Mu

Viktor Mu

Viktor is a Senior Assurance Consultant at AWS with a strong background in information security governance, risk, and compliance. He specializes in helping organizations achieve and maintain PCI DSS compliance in complex cloud environments, with particular expertise in scoping, segmentation, and continuous compliance monitoring.

Scale down Kinesis Data Streams on-demand capacity with ODA warm throughput

Post Syndicated from Pratik Patel original https://aws.amazon.com/blogs/big-data/scale-down-kinesis-data-streams-on-demand-capacity-with-oda-warm-throughput/

Customers have been using Amazon Kinesis Data Streams to stream data at any scale. Some use On-demand Standard to let the service manage capacity, while others with predictable traffic patterns use On-demand Advantage and warm throughput to ensure streams can handle instant throughput increases. Streaming workloads rarely run at peak volume all the time: flash sales end, batch migrations complete, and telemetry bursts subside. However, manual intervention is often required to scale back down after the burst subsides. Amazon Kinesis Data Streams now supports scaling down ingest capacity for on-demand Advantage streams with warm throughput, which optimizes downstream compute costs and performance by removing excess capacity. You configure this by turning on On-demand Advantage mode (ODA) and setting a new warm throughput value that is equal to or smaller than the existing amount.

With this launch, you can now proactively reduce write throughput capacity, optimizing costs while maintaining performance and giving you more control over your stream’s provisioning.

In this post, we explore the warm throughput scale-down capability. We cover the challenge it addresses, how it works, how to monitor stream behavior with Amazon CloudWatch metrics, and best practices for using it effectively.

The challenge: Excess capacity after traffic spikes

Amazon Kinesis Data Streams on-demand mode automatically scales to handle increases in data throughput. When your stream experiences a traffic spike, Kinesis Data Streams splits shards to accommodate the higher volume. This automatic scaling helps your applications keep pace with data during surges.

However, many real-world workloads experience transient bursts that don’t represent sustained throughput needs. Consider a retail platform that processes a flash sale event, a healthcare system that ingests a large batch of patient records during a migration window, or an Internet of Things (IoT) fleet that transmits a high-volume firmware update telemetry burst. In each scenario, the stream scales up to accommodate the spike, but the elevated capacity remains long after the burst has subsided. Although Kinesis on-demand Advantage doesn’t charge for the elevated capacity, your consuming applications may see a higher cost and lower performance.

Consider a Kinesis data stream running with 100 MB/s ingest throughput that requires 100 shards. A traffic spike of an additional 50 MB/s forces on-demand mode to scale streams to 150 shards. The spike subsides within minutes, but those 150 shards remain.

If your AWS Lambda consumer uses a parallelization factor of 2, you go from 200 concurrent invocations (2 × 100 shards) to 300 (2 × 150 shards). This is a 50 percent jump in concurrent Lambda execution, even though ingest throughput has returned to 100 MB/s. Those extra 100 AWS Lambda invocations consume compute, count against your concurrent execution quota, and add cost while processing data with small batch sizes.

Kinesis Client Library (KCL) consumers incur operational overhead. KCL tracks one lease per shard in Amazon DynamoDB, so 50 additional shards mean 50 more leases to scan, renew, and checkpoint every heartbeat cycle. The result is more Amazon DynamoDB overhead for lease management and reduced consumption performance overall.

Before this launch, you had limited options to address this excess capacity:

  • Switch to provisioned mode to manually set shard count, losing the benefits of automatic scaling.
  • Accept the higher capacity and associated costs until the stream self-adjusted.

These approaches either introduced operational overhead or resulted in paying for capacity that exceeded your workload’s actual requirements.

The solution: Warm throughput scale-down

With on-demand capacity reduction, you can now set a lower or equal warm throughput value on your on-demand stream to trigger a capacity reduction. The stream adjusts to the requested capacity or the amount needed to support peak data ingest usage within the last hour, whichever is higher. This safeguard helps your stream retain sufficient capacity for current traffic while releasing the excess you no longer need.

This capability is available at no additional cost for all on-demand streams that have On-demand Advantage mode turned on.

How it works

Warm throughput provides bidirectional capacity management for on-demand streams:

  • Scale up (existing capability): If you forecast an upcoming traffic event, you can configure warm throughput to a higher value to prepare the stream in advance so that capacity is available when data arrives without throttling.
  • Scale down (new capability): If a transient burst has caused the stream to scale significantly beyond its steady-state needs, you can trigger a scale-down by setting warm throughput to a lower value.

When you set a warm throughput value that is equal to or lower than the current value on an on-demand stream, Kinesis Data Streams evaluates the request against your stream’s recent traffic. The resulting capacity is the greater of:

  1. The warm throughput value you requested.
  2. The capacity needed to support peak data ingest usage within the last hour.

This mechanism prevents you from accidentally reducing capacity below what your current workload demands. If data traffic increases after a scale-down has completed, on-demand mode can still expand stream ingest capacity through reactive scaling to avoid rate limiting.

Getting started

Prerequisites

To follow along, you need the following:

  1. An existing Kinesis data stream in on-demand mode.
  2. On-demand Advantage mode turned on.
  3. AWS Command Line Interface (AWS CLI) installed and configured.
  4. AWS Identity and Access Management (IAM) permissions for kinesis:UpdateStreamMode.

To trigger a scale-down, set a lower warm throughput value on your on-demand stream using the AWS CLI:

aws kinesis update-stream-mode \
--stream-arn arn:aws:kinesis:us-east-1:111122223333:stream/my-stream/my-stream \
--warm-throughput-in-mb 50

Monitoring stream behavior with Amazon CloudWatch

To observe the effects of a scale-down operation and understand your stream’s capacity and shard count, Amazon CloudWatch provides several key metrics. Monitoring these metrics helps you make informed decisions about when and how much to scale down.

Key metrics to monitor

The following table summarizes the CloudWatch metrics most relevant to warm throughput scale-down:

Metric Namespace Description
IncomingBytes AWS/Kinesis Total bytes ingested per period. Use the Sum statistic to see aggregate throughput across all shards.
IncomingRecords AWS/Kinesis Total records ingested per period. Helps identify traffic patterns and burst frequency.
WriteProvisionedThroughputExceeded AWS/Kinesis Number of records rejected because of throttling. A non-zero value after scale-down indicates capacity is set too low.

Observing shard count behavior during scale-down

To track shard count changes resulting from a scale-down, use the DescribeStreamSummary API, which returns the OpenShardCount field in its response. Note that OpenShardCount is not a CloudWatch metric. It’s available through the API and is also displayed on the Kinesis Data Streams console. You can poll this value periodically or build a custom CloudWatch metric using an AWS Lambda function to track shard count over time.

Here is how you can expect the stream to behave:

  1. Before the burst: Your stream operates at steady-state with a baseline shard count appropriate for your normal traffic. For example, a stream handling 20 MiB/s of write throughput might have approximately 67 open shards.
  2. During the burst: As traffic spikes, Kinesis Data Streams automatically splits shards to accommodate the increased load. The OpenShardCount rises, and IncomingBytes increases correspondingly.
  3. After the burst (before scale-down): Traffic returns to baseline, but the OpenShardCount remains elevated because the stream retains capacity for up to double the recently observed peak.
  4. After triggering scale-down: After you set a lower warm throughput, the OpenShardCount decreases as Kinesis Data Streams merges shards to match the requested capacity (subject to the one-hour peak safeguard). You can observe this transition by polling DescribeStreamSummary or on the Kinesis console.
Chart showing Kinesis Data Streams shard count rising during a traffic spike and decreasing after a warm throughput scale-down

Figure 1: Amazon Kinesis Data Streams shard count over time during a scale-down event, showing the incoming-data spike and the resulting change in shard count

Best practices

When using warm throughput scale-down, consider the following recommendations:

  1. Analyze traffic patterns before scaling down. Review at least 24 hours of IncomingBytes and IncomingRecords CloudWatch metrics to understand your baseline throughput before setting a lower warm throughput value. This helps you avoid setting capacity below your actual steady-state needs.
  2. Set warm throughput above your observed steady-state peak. Because on-demand streams accommodate up to double the observed peak, set your target warm throughput at or above your typical peak rather than your average. This maintains headroom for normal traffic variability without throttling.
  3. Monitor throttling after scale-down. Watch WriteProvisionedThroughputExceeded closely in the hours following a scale-down. If throttling occurs, increase the warm throughput value. The stream will automatically scale back up, but proactive monitoring reduces the duration of any impact.
  4. Use scale-down after known transient events. The feature is most effective when you can identify that a traffic spike was temporary, for example, after a planned batch migration, marketing event, or scheduled data backfill. Avoid scaling down during periods of uncertain or growing traffic.
  5. Use the one-hour safeguard. The system won’t reduce capacity below what’s needed to serve peak ingest from the last hour. If you’re unsure about the right target, you can set a low warm throughput value and rely on this safeguard to prevent under-provisioning for active traffic.

Conclusion

Amazon Kinesis Data Streams now supports scaling down ingest capacity with warm throughput, giving you elastic control over On-demand Advantage stream capacity. With this capability, you can release excess capacity after transient traffic bursts, improving cost efficiency while maintaining the automatic scaling benefits of on-demand mode.

To get started, turn on On-demand Advantage mode for your stream and use the warm throughput setting to manage capacity. Track shard count with DescribeStreamSummary to observe capacity changes and confirm your stream keeps appropriate headroom for your workload. Try warm throughput scale-down today in the Amazon Kinesis console, and to learn more, see Amazon Kinesis Data Streams on-demand capacity mode in the Developer Guide.


About the authors

Pratik Patel

Pratik Patel

Pratik is Sr Technical Account Manager and streaming analytics specialist. He works with AWS customers and provides ongoing support and technical guidance to help plan and build solutions using best practices and proactively helps in keeping customers’ AWS environments operationally healthy.

Priyanka Chaudhary

Priyanka Chaudhary

Priyanka is Senior Solutions Architect at AWS. She is specialized in data lake and analytics services and helps many customers in this area. As a Solutions Architect, she plays a crucial role in guiding strategic customers through their cloud journey by designing scalable and secure cloud solutions. Outside of work, she loves spending time with friends and family, watching movies, and traveling.

Varsha Palepu

Varsha Palepu

Varsha is a Solutions Architect and an analytics specialist on the AWS streaming team. She helps small and medium businesses innovate on AWS and creates technical streaming content to empower customers in their cloud journey.

Every team is a data team — bring Amazon Redshift analytics to ChatGPT Work

Post Syndicated from Naresh Chainani original https://aws.amazon.com/blogs/big-data/every-team-is-a-data-team-bring-amazon-redshift-analytics-to-chatgpt-work/

Today, AWS is announcing the AWS Data Analytics plugin for the new Data agent in ChatGPT Work. The plugin helps teams across an organization ask questions in natural language, analyze governed data across their Amazon Redshift data warehouse and data lakes, and create shareable dashboards. All this happens from a conversation in ChatGPT Work.

Tens of thousands of customers choose Amazon Redshift every day to run their most demanding workloads, because it delivers analytics at scale with industry-leading price performance. They love how Amazon Redshift provides access to their data warehouses and data lakes together in one place. Teams can combine curated business data with the broader operational, historical, and third-party data stored in open formats like Apache Iceberg in their data lakes. This gives them a complete picture to make business-critical decisions across their data.

Customers have asked AWS for a way to put that trusted data in the hands of more of their people. That means not only the analysts and engineers who write SQL, but also the sales leaders, operations managers, and finance teams who depend on the results. A sales leader wants to know how the customer pipeline has changed this quarter. An operations manager wants to understand why fulfillment times changed over the past month. That’s why we built the AWS Data Analytics plugin, bringing the power of Amazon Redshift and AWS analytics to ChatGPT Work.

“Business teams can make decisions faster when they can source their own analytics and build the dashboards they need. Our work with AWS gives more people that ability, helping them understand changes in performance and decide where to focus. The AWS Data Analytics plugin connects Amazon Redshift to the Data agent in ChatGPT Work, so employees can analyze trusted company data simply by asking, with their organization’s existing access controls in place.”

— Arpan Shah, General Manager, Technology at OpenAI

The new plugin helps shorten the path from question to decision for everyone. Using the Data agent in ChatGPT Work, employees can explore the data they are authorized to access in Amazon Redshift by asking questions in everyday language. They can then refine the analysis, investigate changes, and turn the results into a dashboard without leaving ChatGPT Work. The plugin works with both Amazon Redshift provisioned clusters and Serverless workgroups. Customers can integrate it into their existing multi-cluster or multi-workgroup environments and benefit from the cost and security controls they’ve already set up.

Consider Maya, a business analyst supporting a revenue operations team. She wants to understand the revenue performance across various segments and regions.

Maya starts by loading the AWS Data Analytics plugin in ChatGPT Work, and then asking:

What are the revenue metrics for the past 30 days compared to the previous 30-day period?

ChatGPT Work conversation asking for revenue metrics over the past 30 days compared to the previous 30-day period

Figure 1: Asking for revenue metrics in ChatGPT Work using the AWS Data Analytics plugin

The plugin translates her question into SQL, or a sequence of queries if needed, and runs them against the relevant data in Amazon Redshift. It returns key revenue performance metrics based on the same curated revenue data that her analytics team maintains.

Table of revenue performance metrics the plugin returned from Amazon Redshift

Figure 2: Revenue performance metrics returned from Amazon Redshift

Maya notices that gross margin is declining and asks a follow-up question:

What is my revenue breakdown by product category and region for the past 90 days?

Revenue results segmented by product category and region for the past 90 days in ChatGPT Work

Figure 3: Revenue breakdown by product category and region for the past 90 days

The plugin carries the context forward, segments the results, and helps Maya understand each segment’s performance for the past 90 days. She can inspect the analysis and ask additional questions to drill down even further to understand why certain regions are lagging or why certain segments are outperforming others.

This conversational workflow doesn’t replace the data models, metric definitions, or governance practices that the analytics team has established. It helps more employees use that data directly, giving analysts more time for high-value work.

The AWS Data Analytics plugin connects ChatGPT Work to Amazon Redshift and uses the context of the connected analytics environment to help answer questions with the Data agent. During a conversation, it can:

  • Discover the schemas, tables, columns, and data types available to the user.
  • Translate a natural-language question into Amazon Redshift SQL.
  • Run the query against the customer’s Amazon Redshift environment.
  • Present the results in a table or concise explanation.
  • Use follow-up questions to filter, compare, or drill into the results.
  • Turn an analysis into an interactive dashboard that teams can share and explore.

Because the analysis runs against the customer’s existing data, teams can continue to use the curated datasets and business definitions they already maintain in Amazon Redshift. Customers whose Amazon Redshift environments query data in both a warehouse and a data lake can also make that data available through the governed datasets exposed to the plugin. The AWS Data Analytics plugin also supports our broader AWS data and analytics services. This includes the ability to work with AWS Glue Data Catalog, Amazon S3 Tables (a capability of Amazon Simple Storage Service (Amazon S3)), Amazon Athena, and vector search on AWS.

Natural-language analytics requires more than passing a prompt to a database. The agent needs to understand SQL specific to Amazon Redshift, discover metadata, choose the right tables and columns, and construct queries that follow service best practices. The plugin was built using Amazon Redshift skills from the Agent Toolkit for AWS. These skills provide tested procedures and service-specific guidance that agents can use when working with Amazon Redshift.

To get started, install the AWS Data Analytics plugin in ChatGPT Work to connect it to Amazon Redshift. Give your teams a conversational path to governed insights across your data warehouse and data lake today.

To learn more, see the following resources:


About the author

Naresh Chainani

Naresh Chainani

Naresh is a Director of Engineering at AWS, where he leads Amazon Redshift, one of the world’s most widely used cloud data warehouses. With over 20 years of experience across IBM and AWS, he is a recognized leader in high-performance database systems, holding more than a dozen patents and numerous publications at top venues including SIGMOD and VLDB. Naresh is passionate about advancing the state of the art in analytics and developing the next generation of engineering talent.

Customize Amazon API Gateway destinations for execution logs

Post Syndicated from Giedrius Praspaliauskas original https://aws.amazon.com/blogs/compute/customize-amazon-api-gateway-destinations-for-execution-logs/

Amazon API Gateway execution logs help you trace request processing step by step through your REST API stages. They capture authorization results, integration latency, mapping template output, and error details that are otherwise invisible at the API surface. When a production request fails in a way the access log cannot explain, the execution log is usually where you find the explanation.

Until now, execution logs had two constraints. Every log event was truncated at 1 KB, so a request carrying a moderately sized JSON body would exceed that limit and the remainder was dropped. Logs could only go to the auto-managed log group that API Gateway creates for you (API-Gateway-Execution-Logs_{rest-api-id}/{stage_name}).

With Amazon CloudWatch Logs delivery for REST API execution logs, you can now route execution logs to Amazon CloudWatch Logs, Amazon Simple Storage Service (Amazon S3), or Amazon Data Firehose. Log events can be up to 1 MB per entry, and you benefit from vended logs pricing.

In this post, you learn how CloudWatch Logs delivery works with API Gateway execution logs, how to configure it, and what patterns work best for common observability scenarios.

Understanding API Gateway execution logs

API Gateway produces two categories of logs: access logs and execution logs. Access logs record a summary line per request, similar to an HTTP server access log. You configure the format and destination yourself.

Execution logs are different. They capture the internal processing of each request as it moves through the API Gateway pipeline: authorizer evaluation, request validation, integration dispatch, response mapping, and error handling. These logs exist so you can answer questions such as “why did my authorizer reject this token?” or “what did the mapping template produce before it reached my backend integration?”

API Gateway manages execution log creation automatically. When you set loggingLevel to INFO or ERROR in your stage’s method settings, the service writes execution log events to a CloudWatch Logs log group it manages on your behalf. You do not choose the log group name or configure retention directly on it.

The auto-managed model works for many customers but may create friction for teams with specific observability requirements. Compliance frameworks that require logs in S3 with a particular prefix structure need an extra subscription filter and delivery mechanism. Sending execution logs into a security information and event management (SIEM) tool through a Firehose stream requires a forwarding layer.

Configurable log delivery with CloudWatch Logs

CloudWatch Logs delivery separates log routing from log content. Two concepts control the behavior:

DeliverySource is scoped to your API Gateway stage ARN. It defines where logs go. You create a delivery source, then attach one or more delivery destinations (CloudWatch Logs log group, S3 bucket, or Firehose stream).

MethodSettings controls what gets logged. The loggingLevel setting (INFO, ERROR, or OFF) and dataTraceEnabled flag still determine which log events API Gateway produces. These settings work the same way regardless of whether you use the auto-managed log group or CloudWatch Logs delivery.

When you create a delivery using the CloudWatch Logs APIs, CloudWatch Logs activates your log delivery on your API Gateway stage. When you delete the delivery, CloudWatch Logs disables it accordingly. You do not need to flip any flags on the API Gateway side, and the execution logs automatically resume flowing to the auto-managed log group.

Your existing method settings keep their meaning. The loggingLevel and dataTraceEnabled values continue to control log content. If loggingLevel is already INFO or ERROR, creating a delivery redirects those logs to your chosen destination with no further configuration.

The following diagram shows how the pieces fit together.

Diagram showing one API Gateway stage delivery source fanning out to CloudWatch Logs, Amazon S3, and Firehose destinations.

Figure 1 — A single delivery source scoped to an API Gateway stage feeds one or more deliveries, each of which writes to a delivery destination backed by CloudWatch Logs, Amazon S3, or Amazon Data Firehose

The following table summarizes what changes when log delivery is active.

Aspect Standard execution logging Log delivery
Destination Auto-managed CloudWatch Logs log group CloudWatch Logs, Amazon S3, or Firehose
Multi-destination No Yes
Pricing Standard CloudWatch Logs ingestion Vended logs pricing
Log event size Truncated at 1 KB Up to 1 MB
Setup Set loggingLevel in MethodSettings Create delivery through CloudWatch Logs APIs
Teardown Set loggingLevel to OFF Delete delivery

What stays the same

Only execution log routing changes. Access logs continue to flow through accessLogSettings to whatever log group you configure, and unrelated stage features such as AWS X-Ray tracing, detailed CloudWatch metrics, throttling, and caching behave exactly as they did before.

Configuration and integration options

Before you create a delivery, confirm the following requirements:

  • The API Gateway REST API is deployed to a stage.
  • loggingLevel is set to INFO or ERROR in MethodSettings.
  • The account-level CloudWatch Logs IAM role is configured. For setup steps, see Set up CloudWatch logging for REST APIs in API Gateway.
  • For cross-account delivery, the destination has an appropriate resource policy attached through PutDeliveryDestinationPolicy.

Sending logs to a custom CloudWatch Logs log group

The most common starting point is redirecting execution logs to a log group you own. You get direct control over retention policies, metric filters, and subscription filters. The following steps use the AWS Command Line Interface (AWS CLI) with the fictitious REST API ID abc123, stage prod, Region us-east-1, and account 111122223333.

  1. Create a delivery source referencing your stage ARN. The log type for REST API execution logs is EXECUTION_LOGS:
    aws logs put-delivery-source \
        --name my-apigw-execution-logs \
        --resource-arn arn:aws:apigateway:us-east-1:111122223333:/restapis/abc123/stages/prod \
        --log-type EXECUTION_LOGS

  2. Create a delivery destination pointing to your custom (existing) log group, then create the delivery that connects them:
    aws logs put-delivery-destination \
        --name my-execution-log-destination \
        --delivery-destination-configuration \
            destinationResourceArn=arn:aws:logs:us-east-1:111122223333:log-group:/my-api/execution-logs

    aws logs create-delivery \
        --delivery-source-name my-apigw-execution-logs \
        --delivery-destination-arn arn:aws:logs:us-east-1:111122223333:delivery-destination:my-execution-log-destination

  3. Verify that the delivery is active by listing deliveries for the source:
    aws logs describe-deliveries

The response includes the delivery ID, source, and destination ARN after delivery is established. Execution logs flow to /my-api/execution-logs instead of the auto-managed group.

Note: Log delivery adds structured fields (resource_arn, event_timestamp, api_id, stage, resource_path, http_method, and payload) to each event, so a new delivery emits more than your previous logs. To keep the traditional execution log format with nothing extra, set output format and record fields while creating delivery destination and creating delivery:

aws logs put-delivery-destination \
    --output-format "plain" ...

aws logs create-delivery \
    --record-fields "payload" \
    --field-delimiter "" ...

Routing logs to Amazon S3

S3 works well for long-term retention at lower cost, or for feeding logs into analytics tools such as Amazon Athena. The bucket must be in the same region as your API. Create a delivery destination pointing to your bucket:

aws logs put-delivery-destination \
    --name s3-archive-destination \
    --delivery-destination-configuration \
        destinationResourceArn=arn:aws:s3:::amzn-s3-demo-apigw-logs

Then create a delivery using the same source name. CloudWatch Logs delivers the events to your bucket, where you can query them with Athena or catalog them with AWS Glue.

Streaming to Amazon Data Firehose

For real-time analytics pipelines or third-party SIEM integration, Firehose delivery sends execution log events directly to your stream. The setup is identical: create a delivery destination with your Firehose stream ARN, then create a delivery. With direct Firehose delivery, you no longer need to maintain CloudWatch Logs subscription filters and AWS Lambda forwarders to route execution logs to external analytics systems.

Multi-destination delivery and per-destination shaping

A single delivery source supports multiple destinations. You can route the same execution logs to CloudWatch Logs for real-time alerting, S3 for long-term compliance retention, and Firehose for your SIEM, all from one stage. Create additional deliveries using the same delivery source with different destination ARNs.

Each destination receives identical log events. To shape what reaches each destination, apply a CloudWatch Logs subscription filter on the CloudWatch Logs destination. For example, you can forward only ERROR-level events to a Lambda function that pushes alerts to a SIEM, while the same delivery source writes the full event stream to S3 for compliance.

Management console experience

You can also add a log delivery destination in the management console after you enable logging for the stage.

API Gateway console showing the option to add a log delivery destination after logging is enabled for the stage.

You can specify multiple destinations, both in the current or in a different account:

API Gateway console showing multiple delivery destinations configured, including cross-account options.

Keeping existing monitoring intact

If you have dashboards or alarms on the auto-managed log group, use that same log group as one of your delivery destinations. Your existing monitoring keeps working, and you gain the ability to send logs to additional destinations such as S3 or Firehose in parallel.

Best practices

Update dashboards and alarms before enabling log delivery. When you activate log delivery, the auto-managed log group stops receiving logs. Any CloudWatch alarms, dashboards, or Contributor Insights rules pointing to API-Gateway-Execution-Logs_{rest-api-id}/{stage_name} stop working. Migrate these references to your new log group before creating the delivery.

Keep loggingLevel at INFO or ERROR. Log delivery controls routing, not content. If loggingLevel is OFF, no execution log events are produced regardless of whether a delivery exists. Verify your method settings before troubleshooting missing logs.

Treat the 1 MB log event capacity as a security decision, not only a debugging convenience. With dataTraceEnabled set to true, execution logs include complete request and response payloads up to 1 MB. Those payloads might contain personally identifiable information (PII) or other sensitive data. Confirm your log destinations have appropriate access controls, encryption, and retention policies. Mask or filter sensitive fields in mapping templates upstream of logging and enable data tracing selectively per method or only in non-production stages.

Start with a single destination, then expand. Validate that your log group or bucket receives events correctly before adding Firehose or additional destinations.

Log delivery is best-effort. In rare cases, some log events might not be delivered. For audit-critical workloads, build retention and reconciliation that account for occasional missing events rather than treating execution logs as the system of record.

Cleaning up

To avoid ongoing charges from the resources you created while following this post, delete the delivery and then remove the destinations and any example S3 bucket or Data Firehose delivery stream you no longer need. Deleting the delivery returns the stage to standard auto-managed logging.

aws logs delete-delivery --id <delivery-id>

When the delivery is deleted, CloudWatch Logs disables log delivery on the API Gateway stage automatically. The delivery source and delivery destination remain as independent objects. Delete them with delete-delivery-source and delete-delivery-destination if you do not plan to reuse them.

Conclusion

CloudWatch Logs delivery for API Gateway REST API execution logs helps address the 1 KB event truncation and single managed destination constraints. You can now route full execution logs to CloudWatch Logs, Amazon S3, or Amazon Data Firehose, use multiple destinations from a single stage, and pay vended logs pricing.

The feature works alongside existing method settings. No changes to your current logging configuration are required beyond creating the delivery itself.

To get started, refer to Route execution logs with Amazon CloudWatch Logs delivery in the API Gateway documentation. For more about CloudWatch Logs delivery configuration, see Enable logging from AWS services. For pricing details, review the Amazon CloudWatch pricing page. Try it on a test stage and share your experience in the comments.

Announcing 90-minute function timeout on AWS Lambda Managed Instances

Post Syndicated from Tarun Rai Madan original https://aws.amazon.com/blogs/compute/announcing-90-minute-function-timeout-on-aws-lambda-managed-instances/

AWS Lambda now supports a 90-minute function timeout for asynchronous and event source mapping (ESM) invocations on AWS Lambda Managed Instances (LMI), a capability of AWS Lambda. This is a 6x increase from the previous 15-minute limit. Customers running data processing, media transcoding, financial calculations, AI inference, and batch workloads can now use Lambda functions for jobs that require longer continuous execution, without re-architecting their applications. This also applies to invocations within Lambda durable functions, which use checkpoints to track progress and automatically recover from failures through replay, skipping completed work. When invoked asynchronously, a multi-step durable execution can run for up to 1 year.

Evolution of function timeout on Lambda

Lambda’s function timeout has increased over time, from 5 minutes at launch in 2014 to 15 minutes in 2018. As customers sought to use the simplicity of Lambda for data-intensive workloads, the 15-minute timeout limit forced architectural tradeoffs for applications where customers needed longer continuous execution time. Several patterns emerged:

  • Media processing: Speech-to-text transcription and video transcoding that routinely need longer than 15 minutes of continuous execution.

  • Financial calculations: Monte Carlo simulations, bond pricing, and portfolio risk analysis that are memory-intensive and often require longer than 15 minutes.

  • Data processing and ETL pipelines: Batch jobs processing multi-gigabyte datasets or aggregating data from external sources that exceed 15 minutes during peak volumes.

  • AI inference: Model testing and inference jobs (for example, reasoning tasks) that fit Lambda’s memory and CPU profile but exceed its timeout.

  • Web scraping and file transfer: Crawling external sites or pulling large file sets from vendors that exceed 15 minutes when sources respond slowly.

In each case, customers preferred Lambda’s simplicity but had to re-architect when jobs hit the 15-minute limit.

Fast forward to 2026, Lambda supports two form factors: functions (event-driven, 15-minute timeout), and MicroVMs for user or AI-generated just-in-time code (HTTP-driven, 8-hour duration). To allow customers to benefit from the simplicity of serverless compute with the flexibility and pricing model of EC2 for steady-state workloads, we extended the on-demand capacity mode of Lambda to add Lambda Managed Instances (LMI). With Lambda Managed Instances, you can process multiple concurrent requests per instance, access specialized compute configurations, and drive cost efficiency through EC2 pricing advantages, without managing infrastructure.

We also support Lambda durable functions (powered by the durable execution SDK) on both on-demand and LMI capacity modes. Durable functions provide application-level checkpointing and workflow-as-code: your code saves execution state using step() and wait() operations, and gracefully recovers from infrastructure failures by resuming from the last checkpoint rather than restarting from scratch. While a durable execution (the complete lifecycle of a durable function) can run for up to a year, each invocation was still limited to 15 minutes.

As customers onboard more workloads to serverless compute to benefit from its simplicity, they need longer continuous execution for data-intensive use cases like AI inference, media transcoding, scientific modeling, and financial calculations that do not fit Lambda’s 15-minute duration constraints. Today, we are extending the function timeout on Lambda Managed Instances to 90 minutes for asynchronous and ESM invocations. This includes invocations within a durable function, where a multi-step application can continue to run for up to 1 year when invoked asynchronously.

Activating 90-minute function timeout

You can now configure any Lambda function running on a Managed Instance with a timeout of up to 90 minutes (5,400 seconds) for async and ESM invocations. Synchronous invocations retain the existing 15-minute maximum. The function executes exactly as before: same runtime, same handler, same IAM execution role, same virtual private cloud (VPC) configuration. The only difference is that your function now supports longer continuous execution. Your initialization code (Init phase) is still limited to 15 minutes on Lambda Managed Instances.

To set the timeout, update your function configuration using the AWS CLI:

aws lambda update-function-configuration \
    --function-name my-data-processor \
    --timeout 5400 \
    --region us-east-1

Or in AWS CloudFormation / AWS Serverless Application Model (AWS SAM):

MyFunction:
  Type: AWS::Serverless::Function
  Properties:
    FunctionName: my-data-processor
    Runtime: python3.12
    Handler: app.handler
    Timeout: 5400
    MemorySize: 10240

You do not need to change any code. You can also update the timeout from the Lambda console under Configuration > General Configuration (Figure 1), or configure it through natural language prompts in your AI coding assistants (like Claude Code or Kiro) by installing the Agent Toolkit for AWS.

Lambda console General configuration page showing the function timeout field set to 90 minutes

Figure 1: Configuring Lambda function timeout

The change takes effect on subsequent invocations after the function timeout is updated. For event source mappings, allow a few minutes for the new configuration to propagate. Your existing observability setup continues to work as expected: Amazon CloudWatch metrics, AWS CloudTrail, and AWS X-Ray capture the full invocation lifecycle without any changes. For details, see monitoring Lambda functions and monitoring durable functions.

There is no additional charge for using the 90-minute timeout. Standard Lambda Managed Instances pricing applies.

90-minute timeout and durable functions

The 90-minute function timeout and durable functions are complementary. The function timeout (--timeout) controls how long each individual invocation can run, while the durable execution timeout (ExecutionTimeout in --durable-config) controls the total elapsed time from execution start to completion. Durable functions use checkpoints to track progress and automatically recover from failures through replay, re-executing from the beginning while skipping completed work. With today’s launch, each asynchronous invocation in a durable function running on a Managed Instance can now execute for up to 90 minutes continuously, while the corresponding durable execution can run for up to 1 year. For synchronous and event source mapping invocations, both the invocation and the corresponding durable execution are limited to 90 minutes.

For idempotent jobs (for example, an ETL pipeline step triggered by SQS), the extended timeout alone might be sufficient. If the host fails, the message returns to the queue and a fresh invocation starts. For jobs where re-execution is expensive (for example, a 40-minute inference run already 30 minutes in), combine both. Enable durable functions to checkpoint periodically, so a failure at minute 35 resumes from the last checkpoint rather than restarting from zero.

Invocation behavior: asynchronous, event source mappings, and synchronous

Asynchronous invocations (up to 90 minutes): If the function fails or times out, Lambda applies your configured retry policy (up to two retries by default) and routes failed events to your dead-letter queue or on-failure destination.

Event source mappings (up to 90 minutes): For SQS, configure your queue’s visibility timeout to be at least six times the function timeout. This gives Lambda enough time to retry if a function is throttled while processing a previous batch. Lambda validates this at event source mapping creation time, but does not prevent subsequent changes to queue or function settings that might create a mismatch.

For Amazon Kinesis and Amazon DynamoDB Streams, configure the maximum batching window and parallelization factor to account for longer processing times per batch.

If your batch contains multiple records and you want to avoid re-processing the entire batch when one record fails, enable partial batch failure reporting. This is available for SQS, Kinesis, DynamoDB Streams, Amazon Managed Streaming for Apache Kafka (Amazon MSK), and self-managed Apache Kafka event source mappings. With partial batch failures enabled, only the failed records are retried, not the entire batch.

Note that invocations for Amazon MQ ESM and Amazon DocumentDB (with MongoDB compatibility) ESM remain limited to 15 minutes.

Synchronous invocations (15 minutes maximum): Synchronous invocations retain the existing 15-minute maximum timeout. If you set your function timeout to greater than 15 minutes and invoke it synchronously, Lambda continues to apply the 15-minute timeout. The GetFunctionConfiguration API reports the configured timeout value.

To see which event sources invoke Lambda functions synchronously or asynchronously, refer to Lambda documentation.

Considerations and best practices

Because your functions now support longer continuous execution, consider these best practices for components that might be ephemeral in nature, such as network connections and credentials.

Networking: Make sure idle connection timeouts on downstream services (RDS, Amazon ElastiCache, external APIs) accommodate the full function duration. If your function routes traffic through a NAT Gateway, send keep-alive packets to prevent idle connections from being dropped (350-second idle timeout). Respect DNS TTL values for external hostname resolution. The AWS SDK handles this automatically, but custom HTTP clients might cache DNS records beyond their TTL.

Credentials: If your function acquires temporary credentials or tokens, verify they remain valid for the full execution duration or refresh them in the background.

Idempotency: Lambda does not guarantee exactly-once processing. With longer-running functions, the window for retries and duplicate deliveries increases. You can use Powertools for AWS Lambda to implement idempotency in your function code so that operations like payments or database writes produce the same result even if executed more than once. If you use Lambda durable functions, steps have at-least-once execution semantics by default. The SDK skips completed steps during replay, but steps that fail before checkpointing may re-execute. You can use execution names as idempotency keys for durable functions.

Conclusion

The 90-minute function timeout on Lambda Managed Instances addresses one of the most common customer needs for building data-intensive applications on AWS Lambda. Data processing, media transcoding, AI inference, and financial computation workloads that exceed 15 minutes can now run on Lambda without code changes or architectural workarounds. We look forward to hearing from you if you need a longer timeout for synchronous invocations, or for the on-demand capacity mode, on our AWS Lambda Roadmap GitHub page.

To get started, update your function’s timeout configuration and deploy. For a step-by-step walkthrough, see Getting started with Lambda Managed Instances. For sample code demonstrating long-running functions with durable checkpointing, see durable functions examples. To learn more about AWS Lambda, visit aws.amazon.com/lambda.

Bring your own client certificate for backend mTLS in Amazon API Gateway

Post Syndicated from Biswanath Mukherjee original https://aws.amazon.com/blogs/compute/bring-your-own-client-certificate-for-backend-mtls-in-amazon-api-gateway/

Enterprises that use Amazon API Gateway in front of internal or partner backends often want to bring their own client certificate for backend mutual TLS (mTLS) authentication. During mTLS, the backend presents its own server certificate and also requests the caller to present a client certificate to validate it against a trusted certificate authority (CA). Until now, you could use only an API Gateway-generated, self-signed SSL certificate for the outbound connection, because there was no CA behind it for the backend to trust. Backends that enforce a specific corporate or partner CA reject that self-signed certificate, and the mutual TLS handshake fails. Bringing your own CA-signed certificate is necessary for scenarios such as migrating APIs off legacy gateways or meeting your internal PKI mandates that require certificates from an approved CA.

With API Gateway, you can now bring your own client certificate for backend mutual TLS (mTLS) authentication. You can either use a third-party certificate or a certificate issued by AWS Private Certificate Authority. If you’re using a third-party certificate, you must import the certificate in AWS Certificate Manager (ACM). Then you configure the ACM certificate ARN in your REST API stage. API Gateway presents that certificate during the backend mTLS handshake.

Solution overview

In this post, you build a REST API with an outbound mTLS connection using this newly launched API Gateway feature. This solution demonstrates an outbound mTLS connection between Amazon API Gateway and a backend application running on Amazon Elastic Container Service (Amazon ECS).

The following diagram shows the solution architecture.

Architecture diagram showing API Gateway presenting an ACM client certificate to a Network Load Balancer that forwards traffic to an NGINX sidecar and validator app on Amazon ECS Fargate, with certificates issued by AWS Private CA through ACM

The solution uses:

  1. AWS Private Certificate Authority with a root-subordinate CA hierarchy to issue both the client and server certificates through AWS Certificate Manager (ACM).

  2. Amazon API Gateway REST API stage configured with the ACM client certificate ARN (ClientCertificateId), so that the API Gateway presents the certificate during the outbound TLS handshake.

  3. Amazon ECS on AWS Fargate running an NGINX sidecar that holds the server certificate and validates the incoming client certificate against a CA bundle (root and subordinate chain).

A request goes through the following steps:

  1. Client application invokes the REST API exposed by API Gateway. The API Gateway stage is configured with an ACM client certificate ARN.

  2. API Gateway opens an outbound connection to the Network Load Balancer (NLB) to begin the TLS handshake. The API Gateway presents an ACM client certificate configured at the stage level when the backend requests one.

  3. The NLB listens for the incoming TCP request on port 443 and forwards the call to Amazon ECS Fargate. The NLB acts as a passthrough and does not terminate the TLS connection.

  4. The NGINX sidecar container running on Amazon ECS performs the inbound mTLS handshake:

  • NGINX presents the backend server certificate and verifies the client certificate against a mounted CA bundle (root and subordinate chain).

  • After verification, NGINX forwards the request and parsed certificate details to the validator app container over local HTTP.

  • The validator app re-checks the certificate validity window, matches the common name against an allowlist, and returns a structured JSON response.

Note: The NGINX sidecar is not mandatory for this flow. It demonstrates separation of concerns: NGINX handles the mTLS handshake, and the validator app contains the business logic.

Prerequisites for demo

To follow along, you need the following:

Environment setup

Run the following commands to set up the demo environment:

  1. Create a new folder and clone the GitHub repository:

    git clone https://github.com/aws-samples/sample-api-backend-mtls
    cd sample-api-backend-mtls

  2. Set the environment variables after replacing the placeholders:

    ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
    REGION=<Your AWS Region, for example, us-east-1>
    STACK_NAME=<Your stack name e.g. outbound-mtls-backend>

Build the container images

Run the following commands to create container images of NGINX sidecar container and the validator app containers:

  1. Create two Amazon Elastic Container Registry (Amazon ECR) repositories, one for NGINX and another for validator app containers respectively:

    NGINX_REPO_URI=$(aws ecr create-repository \
      --repository-name $STACK_NAME-nginx-sidecar \
      --image-tag-mutability IMMUTABLE \
      --image-scanning-configuration scanOnPush=true \
      --region "$REGION" \
      --query "repository.repositoryUri" --output text)
    
    VALIDATOR_REPO_URI=$(aws ecr create-repository \
      --repository-name $STACK_NAME-validator-app \
      --image-tag-mutability IMMUTABLE \
      --image-scanning-configuration scanOnPush=true \
      --region "$REGION" \
      --query "repository.repositoryUri" --output text)
    
    aws ecr get-login-password --region "$REGION" \
      | docker login --username AWS --password-stdin "${ACCOUNT_ID}.dkr.ecr.${REGION}.amazonaws.com"

  2. Build and push the NGINX and validator app containers:

    docker build --platform linux/amd64 -t $STACK_NAME-nginx-sidecar nginx/
    docker tag $STACK_NAME-nginx-sidecar:latest "${NGINX_REPO_URI}:latest"
    docker push "${NGINX_REPO_URI}:latest"
    docker build --platform linux/amd64 -t $STACK_NAME-validator-app validator_app/
    docker tag $STACK_NAME-validator-app:latest "${VALIDATOR_REPO_URI}:latest"
    docker push "${VALIDATOR_REPO_URI}:latest"

Deploy and test the solution

You first deploy the stack without the client certificate configured in the API Gateway and perform negative testing. The mTLS handshake will fail because of a missing client certificate in the request. Then you update the stack to configure client certificate in API Gateway stage and retest mTLS.

  1. Run the following command to build and deploy the overall stack without client certificate configured at API Gateway stage:

    sam build
    sam deploy \
      --stack-name $STACK_NAME \
      --resolve-s3 \
      --capabilities CAPABILITY_IAM \
      --region "$REGION" \
      --parameter-overrides \
      NginxRepositoryUri="$NGINX_REPO_URI" \
      ValidatorRepositoryUri="$VALIDATOR_REPO_URI" \
      EnableOutboundMtls=false

  2. Wait for the task to reach RUNNING and pass its target group health check:

    EcsClusterName=$(aws cloudformation describe-stacks \
      --stack-name $STACK_NAME --region "$REGION" \
      --query "Stacks[0].Outputs[?OutputKey=='EcsClusterName'].OutputValue" \
      --output text)
    
    TargetGroupArn=$(aws cloudformation describe-stacks \
      --stack-name $STACK_NAME --region "$REGION" \
      --query "Stacks[0].Outputs[?OutputKey=='TargetGroupArn'].OutputValue" \
      --output text)
    
    aws ecs list-tasks --cluster "$EcsClusterName" --region "$REGION"
    
    aws elbv2 describe-target-health \
      --target-group-arn "$TargetGroupArn" --region "$REGION"

  3. Capture the front API invoke URL from the stack outputs:

    FRONT_API_URL=$(aws cloudformation describe-stacks \
      --stack-name $STACK_NAME --region "$REGION" \
      --query "Stacks[0].Outputs[?OutputKey=='FrontApiUrl'].OutputValue" \
      --output text)
    
    NLB_DNS_NAME=$(aws cloudformation describe-stacks \
      --stack-name $STACK_NAME --region "$REGION" \
      --query "Stacks[0].Outputs[?OutputKey=='NlbDnsName'].OutputValue" \
      --output text)
    
    FRONT_CLIENT_CERT_ARN=$(aws cloudformation describe-stacks \
      --stack-name $STACK_NAME --region "$REGION" \
      --query "Stacks[0].Outputs[?OutputKey=='FrontClientCertArn'].OutputValue" \
      --output text)
    
    FRONT_API_ID=$(aws cloudformation describe-stacks \
      --stack-name $STACK_NAME --region "$REGION" \
      --query "Stacks[0].Outputs[?OutputKey=='FrontApiId'].OutputValue" \
      --output text)

  4. Wait a minute or two after the stack finishes, then invoke the front API:

    curl -v "$FRONT_API_URL"

    The following is the NGINX configuration for mTLS:

    ...
    server {
        listen 443 ssl;
        # Server identity (issued by the private CA).
        ssl_certificate /etc/nginx/certs/server.crt;
        ssl_certificate_key /etc/nginx/certs/server.key;
        # Inbound mutual TLS: require and validate the client certificate
        # against the CA bundle (root + subordinate CA chain).
        ssl_client_certificate /etc/nginx/certs/ca_bundle.pem;
        ssl_verify_client on;
        ssl_verify_depth 2;
        ssl_protocols TLSv1.2;
    ...}

    The curl command returns HTTP/2 400, with a response body containing 400 No required SSL certificate was sent. Because the API Gateway is not presenting a client certificate on the outbound handshake, the NGINX sidecar container in Amazon ECS rejects the mTLS connection. The following screenshot shows the response:

    Terminal response showing an HTTP/2 400 error with the message No required SSL certificate was sent
  5. Now redeploy the solution with outbound mTLS enabled:

    sam deploy \
      --stack-name $STACK_NAME \
      --resolve-s3 \
      --capabilities CAPABILITY_IAM \
      --region "$REGION" \
      --parameter-overrides \
      NginxRepositoryUri="$NGINX_REPO_URI" \
      ValidatorRepositoryUri="$VALIDATOR_REPO_URI" \
      EnableOutboundMtls=true

  6. Wait a minute or two after the stack finishes, then invoke the front API again:

    curl -v "$FRONT_API_URL"

    Because the client certificate is now presented during the mTLS handshake, the handshake completes successfully, as shown in the following response snippet:

    Terminal response showing a successful mTLS handshake and an HTTP 200 response from the backend

Automatic certificate renewal

When a certificate changes in ACM, API Gateway detects the update and propagates the new certificate automatically. You do not redeploy the stage, and the API experiences no downtime during rotation. Certificate propagation is eventually consistent. During an update, the backend might briefly receive either the old or the new certificate. ACM also emits certificate expiration notifications through Amazon EventBridge, which you can use to set alarms before a certificate expires.

Clean up

If you followed along only for demonstration purposes, to avoid incurring future charges, run the following commands to delete the resources created in this demo:

  1. Clean up the S3 buckets:

    NLB_LOGS_BUCKET=$(aws cloudformation describe-stacks \
      --stack-name $STACK_NAME --region $REGION \
      --query "Stacks[0].Outputs[?OutputKey=='NlbAccessLogsBucketName'].OutputValue" \
      --output text)
    
    aws s3api list-object-versions --bucket "$NLB_LOGS_BUCKET" \
      --query '{Objects: Versions[].{Key:Key,VersionId:VersionId}}' \
      --output json | \
      jq -c '.Objects[]? // empty' | \
      while read -r obj; do
        aws s3api delete-object --bucket "$NLB_LOGS_BUCKET" \
          --key "$(echo "$obj" | jq -r .Key)" \
          --version-id "$(echo "$obj" | jq -r .VersionId)" \
          --region $REGION
      done
    
    aws s3api list-object-versions --bucket "$NLB_LOGS_BUCKET" \
      --query '{Objects: DeleteMarkers[].{Key:Key,VersionId:VersionId}}' \
      --output json | \
      jq -c '.Objects[]? // empty' | \
      while read -r obj; do
        aws s3api delete-object --bucket "$NLB_LOGS_BUCKET" \
          --key "$(echo "$obj" | jq -r .Key)" \
          --version-id "$(echo "$obj" | jq -r .VersionId)" \
          --region $REGION
      done

  2. Delete the stack:

    sam delete --stack-name $STACK_NAME --region $REGION --no-prompts

  3. Delete the ECR repository:

    aws ecr delete-repository --repository-name $STACK_NAME-nginx-sidecar --force --region $REGION
    aws ecr delete-repository --repository-name $STACK_NAME-validator-app --force --region $REGION

Conclusion

In this post, you configured a REST API with an outbound mTLS connection using Amazon API Gateway and an ECS Fargate backend. With this new feature launch in API Gateway, you can now bring your own client certificate for outbound mTLS handshake for your REST APIs. You can now meet your internal PKI mandates to authenticate backends that pin a specific certificate issuer.

To get started, import a certificate from your own PKI into ACM and configure your API Gateway REST API stage for outbound mTLS authentication. For more information, see Present client certificates to backend services with mutual TLS in API Gateway. If you have feedback about this post, leave it in the comments section. For technical questions, you can start a thread on AWS re:Post.

Further reading

OSPAR 2026 report now available with 167 services in scope

Post Syndicated from James Chang original https://aws.amazon.com/blogs/security/ospar-2026-report-now-available-with-167-services-in-scope/

We’re pleased to confirm the successful completion of our annual Amazon Web Services (AWS) Outsourced Service Provider’s Audit Report (OSPAR) assessment on July 29, 2026, in line with the OSPAR version 2.0 framework.

The Association of Banks in Singapore (ABS) established the Guidelines on Control Objectives and Procedures for Outsourced Service Providers (ABS Guidelines) to set out baseline control criteria for outsourced service providers (OSPs) operating in Singapore. These guidelines cover key areas such as cyber hygiene, technology risk management, business continuity, data security, cryptography, and software application development and management, drawing on regulatory direction from the Monetary Authority of Singapore (MAS).

This year’s certification cycle broadens the scope with five additional services, covering the 167 AWS services within the AWS Asia Pacific (Singapore) Region. The newly added services are:

This latest certification reinforces our commitment to the security standards expected of cloud providers within Singapore’s financial services industry. For customers, OSPAR offers a way to ease due diligence efforts typically associated with compliance reviews.

You can download the latest OSPAR report from AWS Artifact, a self-service portal for on-demand access to AWS compliance reports. Sign in to AWS Artifact in the AWS Management Console, or learn more at Getting Started with AWS Artifact. The list of services in scope for OSPAR is available in the report and is also available at AWS Services in Scope by Compliance Program.

We remain committed to expanding the OSPAR program’s scope over time, guided by customer architectural and regulatory needs. For any questions regarding the OSPAR report, reach out to your AWS account team.

If you have feedback about this post, submit comments in the Comments section below.

James Chang

James Chang

James is part of the Global Security Assurance team and has taken on major audit programs across the Asia Pacific Japan (APJ) region, including Japan’s ISMAP certification. He has also delivered APJ regulatory assessments and customer assurance engagements.

Ignatius Lee

Ignatius Lee

Ignatius is a Security Assurance professional based in Singapore, covering audits across the Asia Pacific Japan (APJ) region. Since joining Security Assurance in early 2025, he has contributed to key audit programs across the region, including Hong Kong, Singapore, Australia, Japan, Indonesia, and Korea.

Joseph Goh

Joseph Goh

Joseph is the APJ ASEAN Lead at AWS, based in Singapore. He leads security audits, certifications, and compliance programs across the Asia Pacific region. Joseph is passionate about delivering programs that build trust with customers and providing them assurance on cloud security.

Agentic security: Detection and response at machine speed

Post Syndicated from Gee Rittenhouse original https://aws.amazon.com/blogs/security/agentic-security-detection-and-response-at-machine-speed/

After talking with enterprise security leaders over the past year, one thing has become clear: the rise of autonomous AI agents is the most significant shift in security posture since the move to cloud. Organizations across every industry are adopting AI agents that authenticate on behalf of users, execute multistep workflows, and make decisions across infrastructure, often without waiting for human approval. Security operations need to keep pace.

At Amazon Web Services (AWS), we believe security should evolve ahead of AI adoption, not behind it. That belief drove our team to collaborate with the SANS Institute on a new chapter in the 2026 Cloud Security Exchange eBook, where we lay out a practical framework for securing agentic workloads at enterprise scale.

The challenge: Threats now move at machine speed

Traditional security was built for deterministic systems with predictable inputs and outputs. Agentic workloads break those assumptions. The same prompt can produce a compliant response on one request and a policy-violating response on the next. Agents adapt their behavior over time as they interact with users, data, and tools and operate with genuine autonomy: connecting to APIs, chaining actions together, and making independent decisions.

These properties mean that security controls designed for one-time assessments no longer suffice. Detection and response need to operate continuously and at machine speed.

What makes this urgent is the gap between adoption velocity and security maturity. Although 80% of organizations have adopted AI, only 10% govern it. Agents are being built by an expanding population of developers—including those using low-code tools—creating governance challenges that existing security programs must be extended to address.

Extending what already works

The good news, agentic security isn’t a blank slate. It builds on the same principles security teams already apply: identity governance, least privilege, defense in depth, and backup and recovery. What changes is how those principles are implemented when workloads are autonomous and probabilistic. In our eBook chapter, we cover four foundational areas:

  • Agent identity and governance: Every agent needs its own identity with temporary, scoped credentials rather than persistent, broad access. This extends zero trust principles to AI agents, where every request is authenticated and authorized independently, and every action has a traceable authorization chain. When a single agent combines access to sensitive data, the ability to communicate externally, and exposure to untrusted content, the risk profile changes significantly. Design patterns that prevent any single component from combining all three reduce that risk substantially.
  • Evolving detection for agentic workloads: Static, rule-based detection designed for human activity patterns can’t keep up with agent behavior. Organizations need continuous behavioral monitoring, living baselines that adapt as agents evolve, and instrumented observation that surfaces anomalies in real time. Amazon GuardDuty delivers this today, analyzing security signals continuously to detect threats as they emerge.
  • Response that balances speed with precision: When threats move at machine speed, response must be automated and tiered: some agent behaviors should be contained immediately, others require human judgment. The response framework we outline distinguishes between actions that can be automated safely and those that need escalation.
  • From single agents to multiagent ecosystems: Agents are already composing into teams, delegating subtasks, negotiating access, and coordinating across organizational boundaries. Each stage of this evolution inherits every security requirement that came before it, meaning organizations securing today’s basic chat agents are already laying the foundation for tomorrow’s multiagent ecosystems.

Security as an enabler of agentic AI adoption

The security leaders I speak with aren’t asking whether to adopt AI agents. They’re asking how to adopt them responsibly, at speed, and without slowing down the business.

AWS approaches this challenge by building security into the platform at every layer. Agentic AI built on AWS inherits nearly two decades of experience securing mission-critical workloads. Amazon GuardDuty, Amazon Inspector, and AWS Security Hub work together to provide continuous threat detection, vulnerability management, and unified security operations, all adapting to the unique characteristics of agentic workloads.

This isn’t about building new security from scratch. It’s about extending the security foundations your teams already trust into an environment where AI operates with increasing autonomy.

Read the full framework

Our chapter in the 2026 Cloud Security Exchange eBook goes deeper on each of these areas, with specific architectural patterns, implementation guidance, and frameworks for security teams at every stage of agentic AI maturity, whether you’re evaluating, piloting, or operating at scale.

Read the 2026 Cloud Security Exchange eBook: Agentic Security: Detection and Response at Machine Speed

You can learn more about AWS security services at AWS Cloud Security, or explore our AI Security Framework for a comprehensive view of how AWS secures AI workloads with the right controls, at the right layers, at the right phases.

If you have feedback about this post, submit comments in the Comments section below.


Gee Rittenhouse

Gee Rittenhouse

Gee is the Vice President of Agentic Security at AWS. He holds a PhD from MIT and brings extensive leadership experience across enterprise security and cloud. He previously served as CEO of Skyhigh Security and Senior Vice President and General Manager of Cisco’s Security Business Group, where he was responsible for Cisco’s worldwide cybersecurity business.

MCP went stateless: Is your AWS MCP server deployment well-architected?

Post Syndicated from Anand Komandooru original https://aws.amazon.com/blogs/architecture/mcp-went-stateless-is-your-aws-mcp-server-deployment-well-architected/

On July 28, 2026, MCP published its largest revision since launch, making the protocol core stateless and bringing remote MCP servers into alignment with AWS Well-Architected Framework best practices. The initialize handshake is gone, and so is the Mcp-Session-Id header that clients had to echo on every later request. Every request now carries its own protocol version and client context. A client’s first message can be the actual tool call, and any server instance can respond to it. If your MCP server was built for the session-based protocol, the sticky sessions, shared session stores, and custom observability plumbing it required are no longer necessary. If you run behind Amazon Bedrock AgentCore Gateway, protocol management and backward compatibility are handled for you. This post is for teams managing the full deployment stack themselves.

If a client wants to know what a server supports before calling it, a new server/discover method returns the supported protocol versions, capabilities, and identity in a single response. Servers must implement it per the MCP 2026-07-28 specification, but calling it is optional for the client.

This matters on AWS because the old design fought horizontal scaling. A session lived on whichever instance issued it. Running more than one instance meant either pinning clients with sticky routing or externalizing session state to a shared store. Both were correct for that protocol. With the new protocol, neither is required. This post maps the MCP 2026-07-28 specification against the Well-Architected Agentic AI Lens and recommends migrating, because the new protocol achieves natively what the old one could only achieve through compensating infrastructure.

One thing to settle up front, because it drives everything else: stateless describes the protocol, not your application. Stateful use cases still work.

Think of it as a coat check. Under the old protocol the server was a valet who remembered your face, which meant you had to keep dealing with that same valet and nobody else could help you. Now you get a numbered ticket, and any attendant can serve you because the ticket carries the reference. When a server needs continuity across calls, a tool returns an identifier for the stored state. The model includes that identifier on the calls that follow. The state stays in your datastore. The model carries only the key. This is ordinary REST discipline. It has an advantage over the old model. The identifier sits in the model’s context rather than hidden in a header. The model can reason about it and thread it across tools.

What changes in your architecture

The following table compares the deployment patterns the session-based protocol required against the patterns the stateless core now supports.

Before (session-based) After (2026-07-28 stateless)
Elastic Load Balancing Application Load Balancer (ALB) stickiness so each session reaches the same instance. Plain round-robin. Delete the stickiness configuration.
Session state in Amazon DynamoDB or Amazon ElastiCache. No session store. Server-minted identifiers passed as tool arguments.
Parse request bodies at the gateway to route by method. Route and throttle on the Mcp-Method and Mcp-Name headers.
AWS Lambda required workarounds for the stateful handshake. AWS Lambda is a natural fit. Request in, response out.
Refetch tool lists per session. No caching story. Cache with ttlMs and cacheScope, the protocol’s built-in freshness fields.
Bolt-on tracing per implementation. Proprietary protocol logging channel. W3C Trace Context in _meta for distributed tracing. stderr or OpenTelemetry for logging. Protocol logging is deprecated.
Rely on stream resumption (Last-Event-ID) for broken responses. Make tools idempotent. Clients re-issue broken calls.

⚠ Don’t delete yet if you serve 2025-era clients. The 2026-07-28 spec includes a backward-compatible lane that preserves session semantics for older clients. Your ALB stickiness rules and session store (DynamoDB/ElastiCache) must remain in place until you stop serving pre-2026-07-28 clients.
Action: Instrument your gateway to log protocol version per request. Set a sunset date for the legacy lane and communicate it to client teams. Only decommission session infrastructure after traffic on the old version reaches zero. This guidance applies to session infrastructure built to compensate for the old protocol’s requirements. Managed hosts that offer session features by design for specific use cases are not in scope.

One behavioral change to plan for. Servers can no longer push a request to a client mid-call, which is how confirmations, sampling, and root queries used to work over a held-open stream. The spec replaces that pattern with Multi Round-Trip Requests (MRTR). A server that needs input returns an input_required result containing an inputRequests map. This map holds elicitations, sampling calls, or root queries, and an opaque requestState token. The client fulfills the requests, then re-sends the original call with inputResponses and the echoed requestState. Any instance can pick that up because requestState carries all the context the server needs to resume. No shared session store is required. The server does not hold the connection open. This is what makes the pattern work on AWS Lambda.

The Well-Architected view

The AWS Well-Architected Agentic AI Lens already prescribes standardized protocol-based integration as a best practice. For more detail, refer to Establish standardized tool integration protocols (MCP, A2A). What follows is not new guidance but a reading of how the MCP 2026-07-28 specification makes those best practices genuinely achievable for a remote MCP server, pillar by pillar.

Diagram mapping MCP 2026-07-28 protocol changes to the six Well-Architected Agentic AI Lens pillars

Figure 1: How the MCP 2026-07-28 specification maps to the Well-Architected Agentic AI Lens pillars

Operational excellence. The Lens identifies observability as the foundation for operating agents. If you cannot trace a decision end to end, you cannot debug, optimize, or audit it. The 2026-07-28 spec builds observability into the protocol itself. Three changes make this concrete:

  1. Tracing. Every request carries W3C Trace Context keys in _meta (traceparent, tracestate, baggage), so it traces end to end through any OpenTelemetry-compatible backend, including Amazon CloudWatch. The Lens prescribes end-to-end tracing and telemetry for agent operations.
  2. Operational signals without body parsing. The Mcp-Method and Mcp-Name headers expose the operation type on every POST, and every response carries a required resultType field (complete or input_required). Gateways and observability tools get unambiguous per-operation signals for metrics, alarms, and AWS WAF rules without inspecting payloads. The result directly addresses the Lens recommendation for implementing metrics and monitoring for agent-specific patterns.
  3. Standardized logging. MCP’s proprietary protocol logging is deprecated in favor of stderr and OpenTelemetry. The Lens makes the same recommendation: implement structured logging through standardized, queryable formats.

Security. The Lens treats agent security as harder than traditional service security: agents act autonomously with delegated credentials, and their inputs (including state identifiers) are visible to, and potentially manipulable by, the model. MCP’s 2026-07-28 spec hardens the protocol surface against these risks. Five changes strengthen the security posture:

  1. Issuer validation. Clients must validate the iss parameter per RFC 9207, confirming which authorization server produced a response. The Lens calls for the same discipline under strong authentication for agent identities.
  2. Client type declaration. Clients must declare application_type at registration so a desktop or CLI client is not mistaken for a web app, verifying authentication mechanisms match the client’s security profile. The same Lens best practice applies: strong authentication for agent identities. (Note: Dynamic Client Registration itself is now deprecated in favor of Client ID Metadata Documents.)
  3. Bounded human interaction. A server can prompt a user only while it is handling that user’s request, through the Multi Round-Trip Requests pattern. This is a protocol-enforced constraint that bounds when human interaction can occur, aligning with the Lens’s human-in-the-loop controls for critical decisions.
  4. Ownership enforcement. Because state identifiers are visible to the model, servers must enforce ownership on every call. The protocol will not stop a caller from presenting an identifier that is not theirs, so the Lens best practice for tool authorization at the gateway applies: validate that the requesting identity owns the resource it references. The same discipline applies to requestState tokens: the spec requires servers to treat them as untrusted input and protect their integrity with HMAC or AEAD, rejecting any token that fails verification.
  5. Schema validation. Tool input and output schemas are now validated against JSON Schema 2020-12, giving servers a formal contract for rejecting malformed or injected arguments before execution. This maps to the Lens requirement to validating tool inputs at the boundary.

Reliability. Agents hold multi-step context that is expensive to reconstruct after failure, making reliability harder than in traditional services. MCP’s 2026-07-28 spec addresses this at the protocol layer. Four changes reduce that fragility:

  1. Stateless transport. The spec removes protocol-level sessions, so any instance can serve any request. Instance loss is a non-event. Retries need no session affinity, and scale-in never drains sessions. The protocol embodies the failure-isolation philosophy at the protocol layer without additional infrastructure.
  2. Continuation tokens. Interrupted multi-step interactions resume through requestState, an opaque continuation token the server returns and the client echoes on retry. This embodies the Lens principle of designing workflows in stages with incremental recovery.
  3. Idempotent retry. Stream resumability was removed, so a broken response stream loses the in-flight payload and the client must re-issue the call. The mitigation is the same idempotent task execution pattern the Lens prescribes for retryable agent actions: make tools idempotent so re-issued requests produce no duplicate side effects.
  4. Standardized error codes. The spec allocates error code ranges (-32000 to -32019 implementation-defined, -32020 to -32099 reserved for MCP), giving clients and gateways a canonical signal set for retry, backoff, and circuit-breaking decisions. Gateways can now implement standardized communication protocols.

Performance efficiency. Redundant data fetches and per-interaction protocol overhead are the two main performance drags the Lens identifies in agentic workloads. MCP’s 2026-07-28 spec addresses both at the protocol layer. Three changes reduce that overhead:

  1. Protocol-declared caching. Two fields are now required on list and resource-read results: ttlMs (how many milliseconds a response stays fresh) and cacheScope (whether shared intermediaries can cache it or only the requesting client). Tool lists now return in deterministic order, allowing LLM prompt-cache hits across calls. The protocol now delivers what the Lens recommends under optimizing inference-time performance for agent workloads.
  2. Freshness semantics. Clients and MCP-aware gateways can cache responses using protocol-declared freshness (ttlMs + cacheScope), the same data-type-specific TTL discipline the Lens recommends under protocol-declared freshness semantics, without guessing at staleness.
  3. Header-based routing. Routing and throttling decisions now live in HTTP headers (Mcp-Method, Mcp-Name) rather than parsed message bodies, reducing per-interaction overhead in line with what the Lens prescribes for efficient protocol-based agent communications.

Cost optimization. The Lens identifies always-on infrastructure serving bursty agent traffic as the highest source of idle cost in an agent stack. MCP’s stateless architecture eliminates an entire category of that cost: session infrastructure.

  1. Delete session infrastructure. Audit for anything that exists only to preserve sessions (ElastiCache clusters, sticky-routing rules, session-replication logic) and delete it. This follows the same principle the Lens applies to cost-optimizing tool serving through serverless and resource sharing. Infrastructure that runs constantly to serve unpredictable traffic should be replaced with consumption-based patterns that scale to zero. A two-node Amazon ElastiCache (cache.t4g.micro) session store is about $23/month (AWS Pricing Calculator, July 2026). The larger saving is eliminating an entire class of infrastructure and the operational burden around it. Sticky routing costs capacity too by distributing load unevenly, and the savings scale with the size of your fleet.
  2. Serverless as first-class pattern. AWS Lambda has no sticky routing and no persistent connections. A session-based MCP server meant externalizing state to a shared store. Even a “session-free” mode still paid for the mandatory handshake. With the 2026-07-28 stateless core, request in, response out is exactly what AWS Lambda does natively. Serverless MCP moves from workaround to first-class pattern, delivering what the Lens recommends for cost-optimizing tool serving through serverless and resource sharing.

Sustainability. The Lens identifies static provisioning for bursty agent traffic as the primary source of wasted infrastructure capacity. The 2026-07-28 spec’s stateless architecture eliminates the structural reasons for that over-provisioning.

  1. No more pinned-session capacity. The spec’s stateless design means no instance holds a session, so no instance needs to stay warm for one. Right-size against your actual traffic pattern rather than a theoretical peak, the same principle the Lens applies to appropriately scaling compute, networking, and data dependencies for agent workloads. Instance-agnostic routing means the fleet you do keep can run closer to its real utilization, instead of padding for the instances that happened to hold long-lived sessions.

The AWS Well-Architected Agentic AI Lens articulated these best practices as general principles for agentic workloads. The fact that a major protocol revision, designed independently, converges on the same architectural shape is evidence that the framework captures something real about how reliable distributed systems need to work.

What to watch

The architectural shift creates its own operational surface. These are the areas where the new defaults need deliberate attention rather than passive adoption.

Long-lived streams did not disappear. The subscriptions/listen method consolidates change notifications into a single opt-in POST-response stream, so check idle timeouts across your load balancer, proxy, and compute tier if your servers use it.

Deprecations with a clock. The spec deprecated Roots, Sampling, Logging, and the HTTP+SSE transport with a twelve-month floor before removal. The earliest any of these can be removed is July 2027. It also removed ping, logging/setLevel, and notifications/roots/list_changed outright, and moved log level into per-request _meta. The suggested migration paths:

  • Pass directories through tool parameters or resource URIs instead of Roots.
  • Integrate directly with LLM provider APIs instead of Sampling.
  • Log to stderr or OpenTelemetry instead of protocol-level Logging.
  • Migrate HTTP+SSE to Streamable HTTP.

Plan the exits now rather than at the deadline.

MCP Apps puts server-supplied HTML inside your host. Pre-declared UI resource templates, mandatory iframe sandboxing, and auditable JSON-RPC communication between the iframe and host all help. But treat template review as mandatory before deployment, and decide deliberately which servers in your fleet can ship UI at all.

cacheScope is a multi-tenant disclosure risk. Setting cacheScope: "public" on a response that contains tenant-specific data lets shared intermediaries serve one tenant’s list to another. Default to "private" and widen deliberately only for responses that are genuinely identical across callers.

Built-in protection against future breaks

Three mechanisms shipped alongside the stateless core to prevent a repeat of this kind of breaking change.

A feature lifecycle policy gives every feature an Active, Deprecated, or Removed state. Nothing can be removed until at least twelve months after it is deprecated. An extensions framework lets new capabilities ship as opt-in extensions that prove themselves outside the core. That is where Tasks landed after its experimental version needed a redesign. And no Standards Track proposal can reach Final status without a matching scenario in the conformance suite. This is the same suite the official SDKs are validated against.

The handshake and session removal were a deliberate, one-time break to fix the foundation. From here, what you build against 2026-07-28 comes with documented notice periods.

Self-check

Run these ten questions against your own deployment before you decide whether, and how, to migrate.

  1. Can any instance of your server handle any request, with no session affinity at the load balancer?
  2. Have you deleted everything that existed only to preserve a protocol session?
  3. Do your list responses set ttlMs and cacheScope deliberately, and does your gateway route on headers rather than parsed bodies?
  4. Does every client validate iss, and does every server enforce ownership per identifier rather than trusting the identifier itself?
  5. Do you have a firm date to stop supporting 2025-11-25 clients?
  6. Have you replaced server-initiated pushes with Multi Round-Trip Requests so no instance holds a connection open for client input?
  7. Are your tools idempotent so clients can safely re-issue any broken call?
  8. Do you propagate W3C Trace Context end-to-end and emit logs through stderr or OpenTelemetry instead of MCP protocol logging?
  9. Are you still paying for session infrastructure (DynamoDB, ElastiCache, sticky routing) that nothing uses?
  10. Do you have a governance policy for MCP Apps before any server in your fleet exposes one?

A “no” to any of these is where the new spec pays off. Each maps to the pillar sections earlier in this post. Start with the migration path that follows, run your server against the official conformance suite, and use the related AWS resources at the end to plan the change.

Migration path

You do not need to move immediately. Protocol versions are frozen snapshots, and a client and server only need to share one, so 2025-11-25 servers keep working with clients that still speak it. But hosts retire old versions on their own timeline, the community is already moving (GitHub’s MCP Server shipped support ahead of the release), and 2025-11-25 is now frozen. Future capabilities and fixes land on 2026-07-28 or later.

For a new server, target 2026-07-28 directly: stateless from the start, explicit identifiers, and no dependence on Roots, Sampling, or MCP Logging.

For an existing server, work through these steps in order:

  1. Upgrade the SDK and opt in. Speaking the new revision is never automatic.
  2. Audit for session assumptions and migrate off the experimental Tasks API if you used it (Tasks is now an official extension with a redesigned interface).
  3. Plan the deprecation exits (Roots, Sampling, Logging, HTTP+SSE) and change the resource-not-found error code from -32002 to -32602.
  4. Collect the infrastructure savings by deleting session stores, sticky-routing rules, and handshake infrastructure.

For a platform or gateway team: add header-based routing and per-operation throttling on Mcp-Method, honor ttlMs and cacheScope in your caching layer. Also propagate W3C Trace Context, and set a policy for MCP Apps before the first server in your fleet ships one.

Validate before you ship. The official conformance suite covers the new behaviors, and protocol inspectors can pin 2026-07-28 to test your server against exactly what clients will send. Start in a test environment, then promote to production once the suite passes.

Conclusion

The session-based protocol was correct for the constraints it operated under, but those constraints are gone. If you are deploying MCP servers on AWS, the 2026-07-28 specification is the Well-Architected path forward. Migrate your servers, sunset your legacy lane, and delete the infrastructure that existed only to compensate for a protocol limitation that no longer applies.


About the authors

Deliver real-time data to streaming tables for Apache Iceberg with Amazon Kinesis Data Streams

Post Syndicated from Nikit Pednekar original https://aws.amazon.com/blogs/big-data/deliver-real-time-data-to-streaming-tables-for-apache-iceberg-with-amazon-kinesis-data-streams/

Amazon Kinesis Data Streams now supports streaming tables, a fully managed capability that continuously delivers your streaming data as queryable Apache Iceberg tables on Amazon S3 Tables. Amazon S3 Tables is a capability of Amazon Simple Storage Service (Amazon S3). Streaming tables reduce data delivery costs to S3 Tables by up to 50% compared to self-managed alternatives and reduce downstream query costs by up to 30% through intelligent inline compaction that eliminates the small file problem. You need no custom applications, no self-managed compute, and no operational overhead.

Customers increasingly want to unify streaming data with Apache Iceberg for near-real-time analytics, fraud detection, personalization, and artificial intelligence and machine learning (AI/ML) feature pipelines. But integrating the two has meant operating complex custom connectors, managing format conversions, and contending with the performance impact of many small Parquet files that slow queries and increase costs. Streaming tables solve this: configure delivery in a few steps from the console or through APIs, and your data becomes queryable from Amazon Athena, Amazon Redshift, and Apache Spark within minutes. Tables are automatically registered in AWS Glue Data Catalog, making them immediately discoverable for analytics engines and AI agents.

For workloads that don’t require Iceberg table format, you can also deliver streaming data to Amazon S3 general purpose buckets. Delivery is in the source data format, ideal for archival, backup, and ML training data pipelines, with the same serverless, fully managed delivery and no infrastructure to operate.

Challenges with delivering streaming data to Apache Iceberg

Customers today face three challenges when integrating streaming data with Apache Iceberg.

Operational complexity: Connecting Kinesis Data Streams to Iceberg tables today requires deploying and maintaining custom connectors, Apache Flink jobs, or consumer applications. Teams must manage pipeline failures, handle format conversions, scale infrastructure, and monitor delivery reliability. These operational tasks consume significant engineering time and introduce ongoing risk of downtime.

Resiliency and the small file problem: Without proper coordination, simultaneous writes from multiple high-throughput shards can conflict, leading to failed commits, data freshness delays, and degraded performance. Streaming ingestion of high-volume data creates large numbers of small Parquet files in Iceberg tables, forcing a difficult trade-off between data freshness and query efficiency.

Cost: Customers typically spend up to $28/TB operating streaming extract, transform, and load (ETL) pipelines from Kinesis Data Streams using self-managed alternatives based on internal analysis. This creates a high price barrier to getting streaming data into queryable formats and makes cost unpredictable as volume grows.

How delivery to streaming tables solves these challenges

Streaming tables are a native capability built directly into Amazon Kinesis Data Streams. There is no separate service to deploy, no connector to version, and no consumer application to maintain. You enable delivery in a few steps from the console or through APIs.

Zero operational overhead: Streaming tables remove the need to build and operate custom consumer applications for data delivery. No pipeline infrastructure to provision, no scaling logic to write, no failure handling to implement. The capability automatically scales to process gigabytes per second of throughput.

Built-in resiliency: Streaming tables provide write coordination and exactly once delivery semantics across all shards in your stream, resolving concurrent writer conflicts and ensuring data integrity without manual intervention.

Intelligent compaction, no trade-offs: During ingestion, streaming tables perform inline compaction that produces query-optimized Parquet files, eliminating the small file problem while maintaining minute-level data freshness. This reduces downstream query costs by up to 30 percent compared to uncompacted delivery.

Consumption-based pricing: You pay only for data delivered: $14/TB for Iceberg delivery to S3 Tables in US East (N. Virginia) Region (us-east-1) (50% savings compared to self-managed alternatives) and $11/TB for general purpose S3 delivery (60% savings compared to self-managed alternatives). When your stream is idle, you pay nothing for delivery. Combined with Kinesis Data Streams On-Demand Advantage pricing, which eliminates per-shard charges and scales automatically, the entire path from ingestion to queryable Iceberg tables operates on a pure consumption model.

End-to-end managed streaming analytics architecture

With delivery to streaming tables, you now have a fully managed end-to-end real-time data architecture from data ingestion through storage to analytics. Your producers publish events to a Kinesis Data Stream, which continuously delivers data as optimized Iceberg read-only tables in S3 Tables. From there, you can query your streaming data using analytics engines like Amazon Athena, Amazon Redshift, Amazon EMR (Apache Spark), or Apache Flink. You can also let AI agents discover and reason over your data through Glue Data Catalog semantic search. This managed experience removes the intermediate infrastructure that customers previously assembled: separate connector clusters, compaction jobs, and custom consumers. It replaces them with a single, serverless pipeline from stream to insight.

The following diagram illustrates this end-to-end architecture.

End-to-end architecture from Kinesis Data Streams producers to Iceberg tables in S3 Tables, queryable by Athena, Redshift, EMR, and Flink

Figure 1: End-to-end managed streaming analytics architecture from ingestion to query

Getting started

To get started, sign in to the Amazon Kinesis Data Streams console, navigate to your streams, and enable delivery to streaming tables in a few steps. Specify the stream you want to deliver, configure your schema settings using AWS Glue Schema Registry, and choose your destination S3 Tables location. After you enable it, delivery to streaming tables immediately begins materializing your streaming data as queryable Iceberg tables in S3 with no further intervention required. There’s no infrastructure to provision and no minimum commitment. You pay only for data delivered.

Additionally, you can use Amazon Kinesis Data Streams APIs to programmatically set up, update, or delete delivery to streaming tables configurations for your data streams. With these APIs, teams can build agentic workflows and infrastructure-as-code patterns to manage configurations across multiple data streams at scale.

Getting started with the Kinesis Data Streams Agent Skill

The Kinesis Data Streams Agent Skill provides AI-assisted guidance for setting up streaming tables integrations for your existing or new data streams. The skill helps you configure delivery to S3 Tables (Iceberg) or S3, including schema registry setup, AWS Identity and Access Management (IAM) role configuration, and validation.

Installing as an Agent Skill

Agent Skills are discovered automatically by compatible tools through the SKILL.md file. Refer to the Agent Toolkit for AWS Skill Installation Guide to install the managing-amazon-kinesis-data-streams Agent Skill. We also recommend you install the AWS MCP Server in your developer tool of choice, which exposes tools for searching AWS documentation, blogs, and Skills dynamically at runtime. These capabilities make agents more accurate and powerful for AWS related development and operational tasks, and make skill discovery and installation more flexible. Refer to Setting up the AWS MCP Server for guidance on installing the AWS MCP Server in your environment.

For example:

aws configure agent-toolkit
aws agent-toolkit add-skill --skill-name managing-amazon-kinesis-data-streams

To verify the installation, interact with the skill in your preferred tool.

To start delivering data from your data streams to Apache Iceberg tables in real time, prompt “Create me a streaming table on my events data stream” to your agent of choice:

Agent chat showing a prompt to create a streaming table on the events data stream

Figure 2: Prompting the agent to create a streaming table

The agent dynamically loads the managing-amazon-kinesis-data-streams skill and starts by gathering the available resources in your AWS account for the streaming tables integration. After it gathers that data, it confirms the resources to use or create, and creates the integration:

Agent confirming the AWS resources to use or create for the streaming tables integration

Figure 3: The agent confirming resources before creating the integration

After creating the integration, the agent summarizes the status and can then help with any other operational tasks with your data. For example, the agent can help you set up AWS Lake Formation permissions to query the data in S3 Tables with Athena, or configure your table maintenance behavior in S3 Tables:

Agent summarizing integration status and offering Lake Formation permissions or table maintenance setup

Figure 4: The agent offering follow-up operational tasks

Conclusion

Streaming tables are available in all AWS Regions where Amazon Kinesis Data Streams is offered. Pricing is $14/TB for delivery to S3 Tables (Apache Iceberg) and $11/TB for delivery to general purpose S3 buckets. To learn more, visit the documentation and pricing pages.


About the authors

Nikit Pednekar

Nikit Pednekar

Nikit is Principal Product Manager for Amazon Kinesis Data Streams. He leads product vision, strategy, and the P&L for AWS’s real-time data streaming portfolio- Amazon Kinesis Data Streams and related services. Working backwards from customer needs, he drives the streaming roadmap to help AWS customers build scalable, low-latency, real time data architectures.

Mazrim Mehrtens

Mazrim Mehrtens

Mazrim is a Sr. Specialist Solutions Architect for messaging and streaming workloads. Mazrim works with customers to build and support systems that process and analyze terabytes of streaming data in real time, run enterprise machine learning (ML) pipelines, and create systems to share data across teams seamlessly with varying data toolsets and software stacks.

Ren Liu

Ren Liu

Ren is a Solutions Architect at AWS in Seattle, working across the full stack from landing zone design and cloud governance to real-time streaming and ML inference. He works with ISV customers in cybersecurity, FinOps, and healthcare to architect secure, scalable solutions powered by generative AI.

Amazon EC2 R9g and R9gd instances powered by AWS Graviton5 processors are now generally available

Post Syndicated from Daniel Abib original https://aws.amazon.com/blogs/aws/amazon-ec2-r9g-and-r9gd-instances-powered-by-aws-graviton5-processors-are-now-generally-available/

Today, Amazon EC2 R9g and R9gd instances are generally available, powered by AWS Graviton5 processors. R9g instances are memory-optimized and deliver up to 25% better compute performance compared to Graviton4-based R8g instances, powered by the most energy efficient processor AWS has ever built.

R9g instances are ideal for memory-intensive workloads including databases, in-memory caches (Valkey, Redis, MemCached), real-time big data analytics, Linux-based workloads including containerized and micro-service-based applications (e.g. Kubernetes, Docker, EKS, ECS), as well as applications written in popular programming languages such as C/C++, Rust, Go, Java, Python, .NET Core, Node.js, Ruby, and PHP.

R9gd instances include local NVMe-based SSD block-level storage, ideal for memory-intensive workloads requiring fast, low-latency local storage such as open-source databases, distributed real-time big data analytics, large in-memory databases, and large caching workloads.

If you’re running workloads on R8g instances today, R9g gives you more performance per vCPU with faster memory, higher network and Amazon EBS bandwidth, and a larger L3 cache, all while using less energy.

What makes R9g different
Graviton5 processors bring several hardware improvements over Graviton4:

  • Up to 25% higher compute performance per vCPU
  • DDR5 8800 MT/s memory (up from 5600 MT/s in Graviton4), the fastest memory available in the cloud
  • 5x larger L3 cache for better data locality
  • Up to 2x higher network and EBS bandwidth for the largest instance sizes (up to 100 Gbps network, up to 72 Gbps EBS on the 48xlarge)
  • Up to 3x higher packet-processing performance

R9g and R9gd instances support Instance Bandwidth Configuration (IBC), which lets you adjust the allocation of bandwidth between Amazon EBS and Amazon VPC networking by 25%. This helps optimize performance for workloads with specific bandwidth requirements such as databases and caching.

All R9g and R9gd instances run on the AWS Nitro System, which offloads virtualization, storage, and networking to dedicated hardware. This gives your applications near-bare-metal performance while maintaining strong security isolation between instances.

R9g and R9gd instances feature the Nitro Isolation Engine (NIE), the same enhancement to the Nitro System introduced with C9g and M9g instances earlier this year, which enforces isolation of instances and harnesses formal verification to provide assurances of isolation with mathematical precision. Nitro Isolation Engine is a purpose-built component that is responsible for enforcing isolation between virtual machines, including mediation of all access to virtual machine memory, CPU register state, and I/O devices through a minimal set of APIs. Nitro Isolation Engine leverages formal verification, a technique to mathematically demonstrate that the hardware or software behaves as intended, and not just in specific test cases. This intensive verification technique establishes Nitro as the first formally verified cloud hypervisor, pioneering a new standard for mathematically proven cloud security. To learn more about the Nitro Isolation Engine, visit the blog post. For details on the formal verification results, including scope and assumptions, see the technical white paper.

EC2 R9g and R9gd instance specifications
R9g and R9gd instances are each available in 11 sizes, from medium to metal-48xl. The following tables show the full specifications for each size.

Instance size vCPUs Memory (GiB) Instance Storage Network Bandwidth (Gbps) EBS Bandwidth (Gbps)
r9g.medium 1 8 EBS-Only Up to 15 Up to 12
r9g.large 2 16 EBS-Only Up to 15 Up to 12
r9g.xlarge 4 32 EBS-Only Up to 15 Up to 12
r9g.2xlarge 8 64 EBS-Only Up to 17 Up to 12
r9g.4xlarge 16 128 EBS-Only Up to 17 Up to 12
r9g.8xlarge 32 256 EBS-Only 17 12
r9g.12xlarge 48 384 EBS-Only 25 18
r9g.16xlarge 64 512 EBS-Only 34 24
r9g.24xlarge 96 768 EBS-Only 50 36
r9g.48xlarge 192 1536 EBS-Only 100 72
r9g.metal‑48xl 192 1536 EBS-Only 100 72

R9gd instances offer the same compute and networking performance as R9g, with the addition of local NVMe-based SSD storage for workloads that need fast, low-latency scratch space or temporary caches.

Instance size vCPUs Memory (GiB) Instance Storage (NVMe SSD) Network Bandwidth (Gbps) EBS Bandwidth (Gbps)
r9gd.medium 1 8 1 x 59 GB Up to 15 Up to 12
r9gd.large 2 16 1 x 118 GB Up to 15 Up to 12
r9gd.xlarge 4 32 1 x 237 GB Up to 15 Up to 12
r9gd.2xlarge 8 64 1 x 474 GB Up to 17 Up to 12
r9gd.4xlarge 16 128 1 x 950 GB Up to 17 Up to 12
r9gd.8xlarge 32 256 1 x 1900 GB 17 12
r9gd.12xlarge 48 384 3 x 950 GB 25 18
r9gd.16xlarge 64 512 1 x 3800 GB 34 24
r9gd.24xlarge 96 768 3 x 1900 GB 50 36
r9gd.48xlarge 192 1536 3 x 3800 GB 100 72
r9gd.metal‑48xl 192 1536 3 x 3800 GB 100 72

Getting started
You can launch R9g and R9gd instances from the Amazon EC2 console using any supported Arm-based AMI. R9g instances support Amazon Linux 2023, Amazon Linux 2, Ubuntu 22.04+, RHEL 8.4+, SUSE Linux Enterprise Server 15 SP3+, Debian 12+, and other major Linux distributions.

If you’re migrating from R8g, no code changes are required for most applications. Select the equivalent R9g instance size and your application runs with better performance. For containerized workloads, R9g works with Amazon EKS, Amazon ECS, and standard Kubernetes deployments. Multi-arch container images built for Arm64 run without changes.

Several resources help you get started: the AWS Graviton Getting Started Guide covers how to build, run, and optimize workloads on Graviton-based instances. The Graviton Savings Dashboard helps you track cost savings. AWS Transform automates code transformations for migrating Java applications from x86 to Graviton. To learn more, visit AWS Graviton Processors or Level up your compute with AWS Graviton.

Pricing and availability
Amazon EC2 R9g and R9gd instances are available in US East (N. Virginia, Ohio), US West (Oregon), and Europe (Frankfurt) Regions.

R9g and R9gd instances are available for purchase through Savings Plans, On-Demand, Spot Instances, Dedicated Instances, or Dedicated Hosts. For detailed pricing, visit the Amazon EC2 pricing page.

Ready to get started? Launch R9g instances from the Amazon EC2 console. For more details, visit the Amazon EC2 R9g instances page.

If you want to call APIs, search documentation, find regional availability, and check troubleshooting about this feature, try using the AWS MCP Server and plugins with your preferred AI tool. Share your feedback on AWS re:Post for Amazon EC2 or reach out through your usual AWS Support contacts.

— Daniel Abib

We invited a direct competitor into Security Hub Extended. Here’s why.

Post Syndicated from Michael Fuller original https://aws.amazon.com/blogs/security/we-invited-a-direct-competitor-into-security-hub-extended-heres-why/

When customers keep pointing you to a solution that overlaps with parts of your own offering, you have a choice to make. This post is about the choice we made with Upwind, and why we’d make it again.

AWS Security Hub Extended exists because customers told us what was working for them in enterprise security and asked us to simplify adoption and integration. Upwind was one of the solutions customers kept naming, so we brought them in. Upwind didn’t only agree to participate, they committed fully to integration. They brought their full solution portfolio into Extended with aggressive pay-as-you-go pricing from day one. They got their field organization fully aligned on joint deal flow and have driven more customer activity and closed deals through Security Hub Extended than any other partner in the program.

Giving customers choice, even when it overlaps

Multiple best-of-breed options in cloud security—including one that overlaps with our own capabilities—are straightforward when you start with what customers need. Some will choose Security Hub Essentials for cloud security posture management and vulnerability scanning. Some will choose Upwind for runtime-first protection. Some will run both and get stronger outcomes from the combination. The customer decides, not us. That principle applies to every partner in Security Hub Extended. We listen to what’s working, and we simplify adoption through the same AWS relationship customers already have.

“Our customers run on AWS, and Security Hub is where their security operations live,” said Amiram Shachar, Co-Founder and CEO of Upwind. “Being inside Security Hub means customers get Upwind’s cloud workload protection with the same billing, the same support path, and the same operational model they already know. We’re here because it’s a better outcome for the customers we share.”

Who is Upwind?

Upwind is a cloud security company trusted by Siemens, Peloton, Roku, Wix, Nextdoor, and Nubank. Fast Company named them one of the Most Innovative Companies of 2026.

What makes them different is runtime. Most cloud security solutions scan configurations periodically and report what could be a risk based on static posture. Upwind deploys an eBPF-based sensor directly in the Linux kernel that sees what workloads are doing in real time, including process behavior, network connections, API calls, and container interactions. All observed continuously. That means Upwind can tell you not only what could theoretically be exploited, but what is actively at risk right now. That distinction cuts alert noise dramatically and lets security teams focus on what genuinely matters.

Better together. Not only with AWS, but with each other

Now extend that to the rest of your security stack. If you’re already running other Security Hub Extended solutions, they work together without you building the integrations.

A customer running Chainguard for supply chain security, Upwind for runtime protection, and Splunk for security operations gets a connected experience. Chainguard helps ensure clean, malware-resistant dependencies at build time. Upwind validates workload behavior at runtime and enriches those findings with real-time context. Everything flows into Splunk through Security Hub for unified triage. One experience, one bill, no custom integration work. The security team sees the full lifecycle from build to production without stitching tools together.

That same pattern applies with 7AI, where AI-driven automation can triage and investigate Upwind’s runtime events alongside endpoint, identity, and network signals, all without manual pipeline work.

This is the multi-way partnership that Security Hub Extended was designed to enable. These solutions aren’t only easier to buy together, they’re building toward each other. The findings flow into Security Hub in OCSF (Open Cybersecurity Schema Framework), get correlated and prioritized together, and route to the downstream tools your team already uses. Your security stack gets stronger as a whole, not only solution by solution.

How it works commercially

This isn’t a paper partnership. We’re closing multi-million dollar deals together through Security Hub Extended. Upwind has engaged faster than any other partner in the program, bringing their own customer opportunities and joining AWS-originated deals to close them jointly. One enterprise customer recently replaced their incumbent CNAPP with Upwind through a Security Hub Extended Private Offer. The deciding factors were runtime visibility that their previous solution couldn’t deliver and a single predictable commercial model that replaced complex per-module pricing across multiple vendors. The commercial model has momentum, and it’s because Upwind invested not only in signing an agreement but in the engineering and go-to-market work that makes joint success real.

Upwind is available through Security Hub Extended with pay-as-you-go pricing, one AWS bill, and no required long-term commitment. For enterprises that prefer committed-pricing agreements, Security Hub Extended Private Offers are also available with deeper discounts and the ability to aggregate spend across partners. You choose the path that fits how you buy. If you’re already running Security Hub for posture management and vulnerability scanning, adding Upwind gives you runtime visibility alongside what you already see. No new tooling to stand up, no new workflow to learn. It shows up in your existing prioritized view of risk.

What Upwind is building next

Upwind continues to expand. AI workload protection that monitors model behavior and agent tool calls at runtime. Windows Server VM coverage across AWS, Azure, and GCP. Deeper integration with the Security Hub correlation engine so runtime context enriches attack-path intelligence automatically. The partnership deepens as both sides invest.

“We believe runtime context and AWS-native signals together produce stronger outcomes than either alone,” said Amiram Shachar, Co-Founder and CEO of Upwind. “As Security Hub deepens its correlation and Upwind extends its runtime fabric, customers who use both will have a view of risk that no single solution can replicate. That’s the future we’re building toward together.”

What this means for you

Security Hub Extended exists to give you access to the solutions your peers are already succeeding with through the AWS relationship you already have. Upwind is what that philosophy looks like when applied to a category where AWS has an existing offering. We listened to customers, saw what was working for them, and made it available with the same commercial model as everything else.

Enable Upwind through the AWS Security Hub console. Pay-as-you-go. No commitment required. If you want to understand what consolidation looks like with Security Hub Extended, talk to your AWS account team.

We’re just getting started, but the momentum is real.

If you have feedback about this post, submit comments in the Comments section below.


Michael Fuller

Michael has been with AWS for 16 years and led product for AWS Security Services for 11 years. Michael has 29 years in the industry and held several roles in product management, business development, and software development for IBM, Cisco, and Amazon. Michael has a Bachelor’s of Science in Computer Engineering from the University of Arizona and an MBA from the University of Washington.

Integrate Amazon Redshift and IAM Identity Center with enhanced VPC routing

Post Syndicated from Maneesh Sharma original https://aws.amazon.com/blogs/big-data/integrate-amazon-redshift-and-iam-identity-center-with-enhanced-vpc-routing/

You can now use AWS IAM Identity Center authentication with enhanced VPC routing on Amazon Redshift clusters and Amazon Redshift Serverless workgroups. Your users get single sign-on with their existing corporate credentials, and the authentication traffic originates from within your virtual private cloud (VPC) through a VPC endpoint, staying on the AWS private network.

We covered the IAM Identity Center integration end to end in a previous post, Integrate Identity Provider (IdP) with Amazon Redshift Query Editor V2 and SQL Client using AWS IAM Identity Center for seamless Single Sign-On. That post shows how users sign in through Query Editor V2 and third-party SQL clients, and how their identity is propagated to the AWS analytics services.

Many organizations also require that this traffic doesn’t traverse the public internet. Enhanced VPC routing sends everything between your cluster and other AWS services through your VPC, where you can govern it with security groups, network ACLs, and endpoint policies, and observe it in VPC Flow Logs. For teams with data residency, regulatory, or network isolation requirements, it’s often mandatory.

In this post, we show how the new IAM Identity Center VPC endpoints provide a private network path for authentication traffic when enhanced VPC routing is enabled. We walk through the endpoint setup and validate the flow from Query Editor V2 and a SQL client. To use this feature, your Amazon Redshift cluster must be running patch 204 or later, and you must create the VPC endpoints described in the following steps.

Solution overview

When a user signs in with IAM Identity Center from Query Editor V2 or a SQL client, Amazon Redshift doesn’t simply accept the token the client presents. It validates the token with IAM Identity Center and resolves the caller’s identity before the session is established. These calls originate from Amazon Redshift, not from your client, and enhanced VPC routing changes the network path they take.

Authentication flow with enhanced VPC routing

With enhanced VPC routing enabled, the calls Amazon Redshift makes to IAM Identity Center traverse your VPC and follow your networking configuration. The flow is as follows:

  1. The user signs in through Query Editor V2 or a SQL client and authenticates against your identity provider through IAM Identity Center.
  2. IAM Identity Center issues an access token, which the client presents to Amazon Redshift on the database connection.
  3. Amazon Redshift validates the access token against the IAM Identity Center OpenID Connect (OIDC) endpoint, confirming the token’s scopes and the user’s entitlement to the Amazon Redshift application. It doesn’t trust the token the client presented without verification.
  4. Amazon Redshift calls the same OIDC endpoint again to exchange that token for one scoped to Amazon Redshift.
  5. Amazon Redshift calls the IAM Identity Center identity store to resolve the user and their group membership.
  6. Amazon Redshift maps the resolved identity to a database identity, applies role-based access control, and establishes the session.

The following diagram illustrates this authentication flow, showing how each call from Amazon Redshift to IAM Identity Center traverses the VPC through interface endpoints.

Authentication flow showing Amazon Redshift reaching the IAM Identity Center OIDC and identity store endpoints through VPC endpoints

Figure 1: IAM Identity Center authentication flow with enhanced VPC routing enabled

As shown in the diagram, steps 3–5 represent calls that Amazon Redshift makes through your VPC to the IAM Identity Center OIDC and identity store endpoints.

Because a cluster with no public IP address doesn’t use an internet gateway route, Amazon Redshift has no path to IAM Identity Center by default. You provide one with interface VPC endpoints for the two services it needs, the IAM Identity Center OIDC endpoint and the identity store endpoint, which keep the traffic on the AWS network over AWS PrivateLink.

This solution covers the following steps:

  1. Enable enhanced VPC routing.
  2. Verify the DNS attributes on your VPC.
  3. Create the interface VPC endpoints required for IAM Identity Center authentication.
  4. Create interface VPC endpoints for AWS Glue and AWS Lake Formation (optional, if you query a data lake or lakehouse).
  5. Create an Amazon Simple Storage Service (Amazon S3) gateway endpoint.
  6. Validate that the endpoints are available and using private DNS.
  7. Test single sign-on with Amazon Redshift Query Editor V2.
  8. Test single sign-on with a SQL client using the Amazon Redshift JDBC driver.
  9. Verify the calls on AWS CloudTrail.

Prerequisites

You should have the following prerequisites:

  • An AWS account with an Amazon Redshift provisioned cluster. Amazon Redshift Serverless also supports enhanced VPC routing, and the same endpoints apply, but you substitute the equivalent workgroup commands and settings.
  • A working IAM Identity Center integration with Amazon Redshift, as described in Integrate Identity Provider (IdP) with Amazon Redshift Query Editor V2 and SQL Client using AWS IAM Identity Center for seamless Single Sign-On.
  • A cluster running patch 204 or later, which is the minimum maintenance version that supports IAM Identity Center authentication with enhanced VPC routing.
  • Permissions to create VPC endpoints in the VPC where the cluster runs, specifically ec2:CreateVpcEndpoint and ec2:DescribeVpcEndpoints.
  • Optionally, an Amazon Elastic Compute Cloud (Amazon EC2) instance inside the same VPC with SQL Workbench/J and the Amazon Redshift JDBC driver, version 2.1.0.30 or later with its dependent libraries, to test the SQL client flow.

Walkthrough

The examples in this post use the Canada (Central) AWS Region (ca-central-1). Replace all placeholder values with your own.

Step 1: Enable enhanced VPC routing and turn off public access

To control network traffic with Amazon Redshift enhanced VPC routing, you enable enhanced VPC routing in Amazon Redshift. The cluster or workgroup must also not be publicly accessible, so that traffic to IAM Identity Center and other services goes through your VPC endpoints rather than an internet gateway. Follow the instructions in Enable enhanced VPC routing to enable it for a new provisioned cluster or serverless workgroup. For an existing cluster or workgroup, follow these steps:

  1. Sign in to the AWS Management Console and open the Amazon Redshift console at https://console.aws.amazon.com/redshiftv2/.
  2. Open the provisioned cluster or serverless workgroup you want to modify:
    1. For an existing provisioned cluster – choose the Properties tab.
    2. For an existing serverless workgroup – choose the Data access tab.
  3. In the Network and security section, choose Edit.
    1. Select Turn on enhanced VPC routing to route network traffic through the VPC.
    2. If Turn on Publicly accessible is enabled, clear it so the cluster or workgroup is not publicly accessible.
  4. Choose Save changes.

The following screenshot shows the Network and security section with enhanced VPC routing enabled and public accessibility turned off.

Amazon Redshift Network and security section with enhanced VPC routing turned on and public access turned off

Figure 2: Enable enhanced VPC routing in Amazon Redshift

Note: Amazon Redshift restarts the cluster automatically when you change enhanced VPC routing. Make this change during a maintenance window.

Step 2: Verify the DNS attributes on your VPC

Private DNS is what redirects the public AWS service hostnames to your interface endpoints, and it depends on two VPC attributes. Follow these steps:

  1. Navigate to Amazon Redshift and choose the Properties tab for Amazon Redshift provisioned, or the Data access tab for Amazon Redshift Serverless.
  2. Under Network and security setting, choose the associated VPC.
  3. Your VPC details open in a new browser tab.
  4. Review the Details section, where the attributes appear as DNS hostnames and DNS resolution. Make sure that both properties are set to Enabled. The following screenshot shows the VPC Details page with both DNS attributes set to Enabled.
VPC Details page showing DNS hostnames and DNS resolution both set to Enabled

Figure 3: DNS hostnames and DNS resolution enabled on the VPC Details page

  1. If either of the properties is Disabled, choose Actions, choose Edit VPC settings, select Enable on the attribute you need, and choose Save. The following screenshot shows the Edit VPC settings page where you enable these DNS attributes.
Edit VPC settings page with the DNS hostnames and DNS resolution attributes being enabled

Figure 4: Enable DNS hostnames and DNS resolution

Step 3: Create the interface VPC endpoints for IAM Identity Center authentication

Create the two interface endpoints that the authentication flow needs. Each corresponds to one of the two IAM Identity Center calls in the authentication flow described earlier:

Service endpoint Used for
com.amazonaws.<region>.sso-oauth Validating and exchanging the IAM Identity Center access token
com.amazonaws.<region>.identitystore Resolving the user and their group membership

To create an interface endpoint for an AWS service

  1. Open the Amazon Virtual Private Cloud (Amazon VPC) console at https://console.aws.amazon.com/vpc/.
  2. In the navigation pane, choose Endpoints.
  3. Choose Create endpoint.
  4. For Type, choose AWS services.
  5. IAM Identity Center is a Regional service, so these endpoints must reach the AWS Region where your IAM Identity Center instance is available. If you’re using IAM Identity Center multi-Region replication (your instance is replicated to the Region where your Amazon Redshift cluster runs), leave Enable Cross Region endpoint unchecked.
  6. For Service name, search for sso-oauth and select the service for your Region (com.amazonaws.<region>.sso-oauth). The following screenshot shows the top section of the Create endpoint page with the sso-oauth service selected.
Create endpoint page with the sso-oauth service selected for the Region

Figure 5: Create an interface VPC endpoint, part 1

  1. For VPC, select the VPC from which you will access the AWS service. In our use case, we choose the Amazon Redshift VPC.
  2. To enable private DNS support, select Additional settings and choose Enable private DNS name.
  3. For Subnets, select the subnets in which to create endpoint network interfaces. You can select one subnet per Availability Zone. You can’t select multiple subnets from the same Availability Zone. For more information, see Subnets and Availability Zones.
  4. For IP address type, choose IPv4. This assigns IPv4 addresses to the endpoint network interfaces. This option is supported only if all selected subnets have IPv4 address ranges and the service accepts IPv4 requests.
  5. For Security groups, select the security groups to associate with the endpoint network interfaces. For this post, we have selected default security group associated with Redshift. The following screenshot shows the VPC, subnet, and security group selections for the endpoint.
Create endpoint page showing the VPC, subnet, and security group selections

Figure 6: Create an interface VPC endpoint, part 2

  1. For Policy, to allow all operations by all principals on all resources over the interface endpoint, select Full access. To restrict access, select Custom and enter a policy. This option is available only if the service supports VPC endpoint policies. For more information, see Endpoint policies.
  2. (Optional) To add a tag, choose Add new tag and enter the tag key and the tag value.
  3. Choose Create endpoint. The following screenshot shows the policy and tag settings before you create the endpoint.
Create endpoint page showing the policy set to Full access and the tag settings

Figure 7: Create an interface VPC endpoint, part 3

Repeat steps 1–14 for the identity store endpoint, search for identitystore and select the service for your Region (com.amazonaws.<region>.identitystore).

Two settings in the preceding steps are important:

  • Enable private DNS name is required. Amazon Redshift resolves the public service hostname, for example, oidc.<region>.amazonaws.com. Private DNS is what points that hostname at your interface endpoint, so the traffic stays inside your VPC.
  • Use the cluster’s security group, because the cluster is the caller. The endpoint’s security group must allow inbound HTTPS on port 443 from the cluster. Reusing the cluster’s own security group is the simplest approach when it already allows traffic from itself. A dedicated security group needs an explicit port 443 inbound rule from the cluster’s security group.

(Optional) To create an interface endpoint using the command line

Step 4 (optional): Create endpoints for AWS Glue and AWS Lake Formation

Complete this step only if your cluster queries external data through the AWS Glue Data Catalog and AWS Lake Formation. Common examples include Amazon S3 Tables, a capability of Amazon S3, and data lakes registered with Lake Formation. If you only need single sign-on, you can skip to Step 5. Amazon Redshift calls the AWS Glue Data Catalog to enumerate databases and tables, and calls AWS Lake Formation to check permissions and vend temporary credentials for the underlying data. Like the authentication calls, these are made by the cluster, so with enhanced VPC routing enabled they travel through your VPC and need a path of their own.

Repeat steps 1–14 from Step 3 for:

  • com.amazonaws.<region>.glue.
  • com.amazonaws.<region>.lakeformation.

With these endpoints in place, Amazon Redshift routes external catalog operations such as listing external tables through the VPC endpoints rather than the public internet, keeping metadata traffic on the AWS network.

Step 5: Create an Amazon S3 gateway endpoint

With enhanced VPC routing enabled, anything the cluster does against Amazon S3 (COPY, UNLOAD, and Amazon S3 Tables) also travels through your VPC. Create a gateway endpoint and associate it with the route table(s) used by your cluster’s subnets:

  1. Open the Amazon VPC console at https://console.aws.amazon.com/vpc/.
  2. In the navigation pane, choose Endpoints, then choose Create endpoint.
  3. For Type, choose AWS services.
  4. For Service name, search for s3 and select the service for your Region with Type: Gateway (com.amazonaws.<region>.s3). The following screenshot shows the Create endpoint page with the Amazon S3 gateway service selected.
Create endpoint page with the Amazon S3 gateway service selected

Figure 8: Create an Amazon S3 gateway endpoint, part 1

  1. For VPC, choose your Amazon Redshift VPC.
  2. For Route tables, select the route table(s) associated with the subnets your cluster runs in.
  3. Choose Create endpoint. The following screenshot shows the VPC and route table selections for the S3 gateway endpoint.
Create endpoint page showing the VPC and route table selections for the S3 gateway endpoint

Figure 9: Create an Amazon S3 gateway endpoint, part 2

Step 6: Validate the endpoints

Confirm in the Amazon VPC console that every endpoint you created is available, and that private DNS is enabled on the interface endpoints.

  1. Open the Amazon VPC console at https://console.aws.amazon.com/vpc/.
  2. In the navigation pane, choose Endpoints.
  3. In the endpoints list, use the filter bar to filter by VPC ID (choose VPC ID and select your Amazon Redshift VPC). Then locate the endpoints you created for this walkthrough, sso-oauth, identitystore, the Amazon S3 gateway endpoint, and (if you created them) glue and lakeformation.
  4. Confirm each endpoint shows a Status of Available.
  5. Select each interface endpoint (sso-oauth, identitystore, glue, lakeformation) and, on the Details tab, confirm Private DNS names enabled is Yes. The following screenshot shows the completed endpoints list with each endpoint in the Available state.
VPC endpoints list showing the interface and gateway endpoints in the Available state

Figure 10: VPC endpoints created in this walkthrough, in the Available state

Step 7: Test single sign-on with Amazon Redshift Query Editor V2

  1. On the Amazon Redshift console, choose Query editor v2.
  2. Choose your cluster and then choose IAM Identity Center as the connection method.
  3. Sign in with your corporate credentials when prompted.
  4. Expand the cluster in the tree view to list databases, schemas, and tables.

The database list populates within a few seconds. Confirm the login on the server side by querying the connection log. Run this as a user who does not use IAM Identity Center, for example a database user with a password, or through the Amazon Redshift Data API:

SELECT record_time, user_name, auth_method, driver_version, remote_host, event
FROM sys_connection_log
WHERE auth_method LIKE '%Idc%'
AND record_time > dateadd(minute, -15, getdate())
ORDER BY record_time DESC;

A successful sign-in shows user_name as <idc_namespace>:<[email protected]> with event of authenticated, which confirms that the identity was resolved through the endpoints you created. The following screenshot shows the sys_connection_log query results, where each IAM Identity Center sign-in appears with a user_name in the <idc_namespace>:<[email protected]> format.

Query Editor V2 connected through IAM Identity Center, showing the expanded database list

Figure 11: Query Editor V2 connected with IAM Identity Center, showing the database list

Step 8: Test single sign-on with a SQL client

Testing from a SQL client on an EC2 instance inside your VPC is the stronger validation, and we recommend doing both. Query Editor V2 connects through an Amazon Redshift managed proxy, so its connections are recorded with a loopback address. A client running inside your VPC connects to the cluster endpoint directly, which is exactly the path the endpoints you created are there to serve.

Set up SQL Workbench/J

SQL Workbench/J connects through the Amazon Redshift JDBC driver. On an EC2 instance in the same VPC as your cluster, download and install SQL Workbench/J.

  1. Download the latest Amazon Redshift JDBC driver together with its dependent libraries, and extract the archive to a folder on the instance.
  2. Start SQL Workbench/J, and choose File, then Manage Drivers.
  3. Choose the Create a new entry icon, and for Name, enter Amazon Redshift.
  4. For Library, choose the folder icon, and select the driver JAR file along with every JAR file in the dependent libraries folder. Keep only one version of the driver in the list, and remove any previous entries.
  5. Choose File, then Connect window, and choose the Create a new connection profile icon. Enter a name for the profile, such as redshift-idc.
  6. For Driver, choose the Amazon Redshift driver that you created.
  7. For URL, enter your cluster endpoint in the form jdbc:redshift://<cluster endpoint>:5439/<database>, for example jdbc:redshift://my-redshift-cluster.abc123xyz789.ca-central-1.redshift.amazonaws.com:5439/dev.
  8. Leave Username and Password empty. The browser plugin obtains the identity interactively.
  9. Choose Extended Properties, and add the following three properties:
Property Value
plugin_name com.amazon.redshift.plugin.BrowserIdcAuthPlugin
issuer_url https://identitycenter.amazonaws.com/ssoins-<instance-id>
idc_region The Region of your IAM Identity Center instance, such as ca-central-1
  1. Clear Separate connection per tab, so that each editor tab reuses the same physical connection rather than prompting you to sign in again.
  2. Choose Test. Your default browser opens. Sign in with your corporate credentials, and then choose Allow access so that the Amazon Redshift JDBC driver can access your data.
  3. If the connection succeeds, you see a prompt confirming the connection to your Amazon Redshift endpoint, as shown in the following screenshot.
SQL Workbench/J connection profile for Amazon Redshift using the IAM Identity Center browser plugin

Figure 12: SQL Workbench/J connection for Amazon Redshift using the IAM Identity Center browser plugin

In the browser, you will see the following message once the authentication is successful.

Congratulations! You have IAM Identity Center single sign-on working on an Amazon Redshift cluster with enhanced VPC routing enabled.

Step 9: Verify the calls on AWS CloudTrail

You can confirm from AWS CloudTrail that these calls travel through your interface endpoints rather than the internet. Each event includes a vpcEndpointId field naming the endpoint the call traversed, along with a vpcEndpointAccountId field identifying the account that owns it.

The following table maps each interface endpoint to the CloudTrail event you’ll see:

Service endpoint Event source CloudTrail event
com.amazonaws.<region>.sso-oauth sso-oauth.amazonaws.com CreateTokenWithIAM
com.amazonaws.<region>.identitystore identitystore.amazonaws.com DescribeUser, ListGroupMembershipsForMember, BatchDescribeGroup

The following screenshot shows the snippet from the CloudTrail logs showing the CreateTokenWithIAM event that Amazon Redshift generates when it exchanges the IAM Identity Center access token. The eventSource is sso-oauth.amazonaws.com, and the vpcEndpointId field confirms the call traversed your interface VPC endpoint rather than the public internet. The invokedBy field shows the call originated from Amazon Redshift (redshift.amazonaws.com), not from the client.

CloudTrail CreateTokenWithIAM event with event source sso-oauth.amazonaws.com and a vpcEndpointId field

Figure 13: CloudTrail CreateTokenWithIAM event traversing the sso-oauth interface endpoint

Similarly, the following screenshot shows a DescribeUser event (event source identitystore.amazonaws.com) generated when Amazon Redshift resolves the authenticated user against the identity store. As with the previous event, the invokedBy field shows the call originated from Amazon Redshift, and the vpcEndpointId field confirms it traversed the identitystore interface endpoint.

CloudTrail DescribeUser event with event source identitystore.amazonaws.com and a vpcEndpointId field

Figure 14: CloudTrail DescribeUser event traversing the identitystore interface endpoint

Note: where the events appear depends on how your IAM Identity Center instance is deployed:

  • CreateTokenWithIAM (event source sso-oauth.amazonaws.com) is recorded in the same account as your Amazon Redshift cluster.
  • Identity Store API calls (DescribeUser, ListGroupMembershipsForMember, BatchDescribeGroup) are recorded in the account that owns your IAM Identity Center instance. The event source is identitystore.amazonaws.com. If you use a centralized instance in a delegated administrator or management account, these events appear in that account and not in the account running your cluster. Searching the cluster’s own account returns nothing, even when authentication is working normally. To confirm which account to look in, run aws sso-admin list-instances and check OwnerAccountId.

Clean up

To avoid incurring future charges, delete the resources you created for this walkthrough. These endpoints provide the network path for single sign-on while enhanced VPC routing is enabled, so remove them only if you no longer need the integration.

  • On the Amazon VPC console, choose Endpoints.
  • Select the sso-oauth and identitystore interface endpoints you created, and choose Actions, then Delete VPC endpoints.
  • Select the glue and lakeformation interface endpoints, if you created them, and delete them.
  • Select the Amazon S3 gateway endpoint and delete it. This also removes its route table entries.
  • Terminate the EC2 instance you used to test the SQL client connection, if you created one for this walkthrough.
  • If you no longer need the integration, remove the IAM Identity Center application assignment for Amazon Redshift and delete the associated IAM role and policy.

Conclusion

In this post, we showed you how to enable AWS IAM Identity Center authentication for Amazon Redshift on clusters with enhanced VPC routing enabled. Your users get single sign-on with their corporate credentials, and the authentication traffic stays private to your VPC. The key concept is that Amazon Redshift, not your client, validates the access token. Because enhanced VPC routing is enabled, Amazon Redshift routes that validation call through your VPC. On a cluster running patch 204 or later, interface endpoints for sso-oauth and identitystore give Amazon Redshift a private path over AWS PrivateLink. Adding endpoints for AWS Glue, AWS Lake Formation, and Amazon S3 extends the same benefit to data lake and lakehouse queries.

Try this setup in your own environment and let us know what you think in the comments. For more information, see the following resources:


About the authors

Maneesh Sharma

Maneesh Sharma

Maneesh is a Senior Analytics Specialist Solutions Architect at AWS with more than 15 years of experience designing and implementing large-scale data warehouse and analytics solutions. He works with FSI and Enterprise customers to implement modern analytics architectures using Amazon Redshift, Amazon SageMaker Unified Studio, Amazon S3 Tables, AWS Glue, AWS Lake Formation, and AWS IAM Identity Center.

Laura Reith

Laura Reith

Laura is an Identity Solutions Architect at AWS, where she thrives on helping customers overcome security and identity challenges. In her free time, she enjoys wreck diving and traveling around the world.

Suchintya Dandapat

Suchintya Dandapat

Suchintya is a Principal Product Manager for AWS where he partners with enterprise customers to solve their toughest identity challenges, enabling secure operations at global scale.

Jonathan Glaser

Jonathan is a Software Development Engineer on the Amazon Redshift Connectivity team, where he works on authentication and the Redshift client drivers. He focuses on secure authentication for Redshift, including its integration with IAM Identity Center in Enhanced VPC Routing environments. Jonathan holds master’s degrees in biotechnology and computer science, and in his spare time enjoys reading fiction and exploring NYC’s food scene.

Nishtha Mehrotra

Nishtha is a Senior Software Development Engineer on the Redshift Connectivity team at AWS. With over 10 years of software engineering experience, she specializes in database driver development, performance optimization, and identity integration. Nishtha works on Amazon Redshift’s ODBC, JDBC, and Python drivers, ensuring reliable and performant connectivity for customers at scale. She is passionate about improving developer experience and building secure, high-performance data infrastructure that powers analytics workloads across AWS.