Tag Archives: Technical How-to

Automating the Experimentation Lifecycle with Kiro, AWS DevOps Agent, and LaunchDarkly

Post Syndicated from Greg Eppel original https://aws.amazon.com/blogs/devops/automating-the-experimentation-lifecycle-with-kiro-aws-devops-agent-and-launchdarkly/

Introduction

Continuous improvement depends on experimentation. Teams know that the fastest path to better outcomes is to test changes against real user behavior, measure results, and iterate. In practice, sustaining that cycle is slow and costly because the overhead compounds with each attempt.

Three barriers slow teams down:

1. Planning cost — Turning a proposed change into a testable experiment requires defining a feature flag strategy, coordinating implementation, and wiring everything together before any user sees new behavior.

2. Measurement disconnected from action — Once live, teams must configure metrics, define success criteria, monitor, and interpret results. When metrics regress, remediation traditionally depends on a human merging a fix or rolling back a deployment.

3. Stalled iteration — Without a record of which change caused which outcome, the next hypothesis is a guess, so iteration often does not happen and the goal stalls.

This post introduces a reference solution that closes the gap between defining a goal and reaching it. A team states an improvement goal (for example, increase add-to-cart rate by 10%), and agents plan the experiment, implement the change, deploy it behind a feature flag, measure its impact, and iterate on the result, all within defined safety boundaries. The solution connects Kiro for code generation, AWS DevOps Agent for orchestration and release readiness review, and LaunchDarkly for feature flag governance, experiments, and Guarded Releases for safe, metric-driven rollouts with automatic rollback. The architecture described here is a reference implementation you can build today. A more turnkey experience is planned for the future.

Pre-requisites

Step 1. Enable AWS DevOps Agent and Create an Agent Space. AWS DevOps Agent is available in the AWS regions listed here. Follow these steps to create your AWS DevOps Agent and create an Agent Space.

Step 2. Create your LaunchDarkly account. Create your LaunchDarkly account using the AWS Marketplace or through LaunchDarkly website.

Step 3. Enable the LaunchDarkly MCP Server in the Agent Space. AWS DevOps Agent connects to LaunchDarkly’s hosted MCP server as a client, giving it the ability to query flag state, read targeting rules, and list flags by project or environment.

Step 4 — Register the LaunchDarkly MCP server (account-level). MCP servers are registered at the AWS account level and shared among all Agent Spaces in that account.

  • Sign in to the AWS DevOps Agent console.
  • Navigate to the Capability Providers page (side navigation).
  • Find MCP Server under the Available providers section and choose Register.
  • Enter the MCP server details (see table below).
  • Choose Next.
    • Name: LaunchDarkly
    • Endpoint URL: https://mcp.launchdarkly.com/mcp/launchdarkly
    • Description: LaunchDarkly feature flag management MCP server
    • Enable Dynamic Client Registration: Select this checkbox to allow DevOps Agent to automatically register with LaunchDarkly’s authorization server

Step 4a — Configure the authorization flow

LaunchDarkly’s hosted MCP server uses OAuth for authentication:

  • Select OAuth 3LO (Three-Legged OAuth).
  • Choose Next.
  • Complete the OAuth authorization — you will be redirected to LaunchDarkly’s consent page to authorize the connection.
  • Choose Next.
  • Tip: Refer to the LaunchDarkly MCP server documentation for specific OAuth scope and credential details.

Step 4b — Review and submit

  • Review the MCP server configuration details.
  • Choose Submit.
  • AWS DevOps Agent validates the connection to LaunchDarkly’s MCP server.
  • On successful validation, the MCP server is registered at the account level.

Step 5 — Add the MCP server to your Agent Space

After the account-level registration, connect it to your specific Agent Space:

  • In the AWS DevOps Agent console, select your Agent Space (created in Section 1).
  • Go to the Capabilities tab.
  • In the MCP Servers section, choose Add.
  • Select the LaunchDarkly MCP server you just registered.
  • Configure tool access:
    • Allow all tools — makes all LaunchDarkly MCP tools available to the agent
    • Select specific tools — allowlist only the tools you need (recommended for production)
  • Choose Add.

Step 5 — Validate the connection. Run a test query to confirm the integration is working. In the DevOps Agent console, start a new investigation or chat session and ask: “List the feature flags in the <your-project-key> project in the production environment.” If the agent returns flag data from LaunchDarkly, the connection is active.

Solution overview

The automated experimentation lifecycle operates as a closed loop. A team states an improvement goal, and the system moves through a continuous cycle: decide what to try next, implement the change behind a feature flag, validate and deploy it, run an experiment to measure impact, roll it out safely, and feed the outcome back into the next iteration. The loop continues until the goal is met or the team decides to stop.

Flowchart showing the Plan-Prove-Iterate continuous improvement loop for AWS DevOps Agent. The Plan phase covers steps 1 through 5: generate hypothesis, create feature flag, implement behind flag, release readiness review, and merge PR to deploy. The Prove phase has two sub-phases: Experiment (50/50 split on 10% traffic measuring business KPIs) and Guarded Release (ramp from 20% to 40% with auto-rollback on regression). The Iterate phase covers steps 6 through 8: record outcome, generate report, and feed into next hypothesis, ending with a goal-met decision gate. An improvement goal banner reads "Increase add-to-cart rate by 15%."

End-to-end Plan-Prove-Iterate workflow showing how AWS DevOps Agent orchestrates hypothesis generation, feature-flagged implementation, experimentation, guarded rollout, and outcome recording in a continuous improvement loop.

Each component has a distinct responsibility. AWS DevOps Agent orchestrates the cycle: it runs on a schedule as a Custom Agent which is a user-defined agent with its own instructions, skills, and connected tools that executes autonomously without pausing for input unless something fails. AWS DevOps Agent supports Custom Agents as a way to encode a specific workflow, including its decision logic, safety constraints, and cadence, into an agent that runs end-to-end on its own. In this solution, the Custom Agent reviews goals, generates hypotheses informed by prior outcomes, coordinates implementation and validation, and drives iteration across multiple experiment cycles.”. Kiro CLI runs in headless mode inside the Experiment MCP Server container on Amazon Bedrock AgentCore, implementing code changes behind LaunchDarkly feature flags and opening pull requests without a human operating an IDE.

LaunchDarkly hosts feature flags, experiments, and Guarded Releases, monitors metrics in real time, and reverts flag state when a threshold is breached. It also exposes a hosted MCP server with tools the agent calls directly. The Experiment MCP Server (custom, built for this solution) exposes the remaining operations over MCP: code implementation through Kiro, PR merge, and deployment triggering.

The agent acts as an MCP client connected to these two servers. LaunchDarkly’s hosted MCP server provides flag management, experiment lifecycle, Guarded Release, and observability tools. The Experiment MCP Server provides code implementation, PR merging, and deployment tools. This design separates decision-making from execution: the agent decides what to do, the MCP servers handle how.

Plan / Prove / Iterate

The lifecycle operates in three phases.

Plan — The agent decides the next action for a goal, generates a hypothesis informed by prior outcomes when iterating, and creates a feature flag in LaunchDarkly. It then invokes Kiro CLI to implement the change behind the flag and open a pull request. AWS DevOps Agent validates the change through release readiness review. After a green review, the PR is merged and a GitHub Actions workflow deploys the application through AWS Amplify.

Prove — Two sequential phases run after deployment. First, a 50/50 experiment splits 10% of traffic on a business KPI (for example, add-to-cart rate) until statistical significance selects a winning variation. Then a Guarded Release ramps the winning variation from 20% to 30% to 40% and eventually to 100% while LaunchDarkly monitors operational guardrails (error rate, page-load-time-p95). If a guardrail threshold is breached, LaunchDarkly reverts the flag state automatically, requiring no redeployment. The experiment measures value (does the change improve the goal metric?); the Guarded Release measures safety (does the change hold up at scale?).

Iterate — After a rollout concludes, the agent queries LaunchDarkly’s Change History API to associate specific flag modifications with outcomes. The recorded outcome informs the next hypothesis, and the cycle repeats until the goal is met or the agent recommends waiting.

Extending the agent with a custom MCP server

AWS DevOps Agent reads code, reviews changes, and decides what to do next. It does not take action on its own. To move from decision to execution, you connect it to MCP servers that expose operations as tools.

LaunchDarkly’s hosted MCP server covers flags, experiments, and Guarded Releases. We needed operations it doesn’t cover — writing code, merging PRs, and deploying — so we built the Experiment MCP Server. It runs on Amazon Bedrock AgentCore and exposes five tools: create_task and get_task_status (invoke Kiro CLI to implement changes and open a PR), merge_pr, trigger_deployment, and get_deployment_status.

These are mutation operations. When the agent calls create_task, Kiro writes real code. When it calls merge_pr, that code lands in main. You are responsible for this server — what it exposes, which repos it can touch, which branches it can merge to. We scoped ours to one repository, one branch, and one Amplify application. Those constraints live in the MCP server’s code, not the agent’s prompt, because API-level scoping cannot be misinterpreted.

The Experiment MCP Server [CG1] is a Python application built on FastMCP, packaged as a container and deployed to Amazon Bedrock AgentCore over stateless HTTP so the platform can restart or replace the container without breaking in-flight requests. At startup, the container pulls credentials from AWS Secrets Manager, clones the target repository, and makes Kiro CLI available as a local binary. This single-container design keeps everything colocated: when the agent calls create_task, the server spawns Kiro CLI as a headless subprocess with direct filesystem access to the cloned repo rather than making a network call to a separate code-generation service. Kiro CLI receives a structured prompt containing the task description, the LaunchDarkly flag key, and the variation details, then writes the change, commits to a new branch, and pushes. The server opens a pull request through the GitHub API and returns the task ID immediately without waiting for Kiro to finish. The caller polls get_task_status, which long-polls against an S3-backed state store so task progress survives container restarts. Deployment tracking follows a similar pattern: trigger_deployment dispatches a GitHub Actions workflow and returns the real GitHub run ID, and get_deployment_status reads live status directly from GitHub, so there is nothing to lose if the container cycles between calls. The overall design principle is that the MCP server coordinates work and delegates persistence to external systems (S3 for task state, GitHub for deployment state, Secrets Manager for credentials) rather than holding anything in memory that a restart would erase.

How the agent works

The agent runs on a schedule. Each run, it evaluates the current state of each goal and picks one of three actions: create a new experiment (no active rollout exists), iterate on a prior result (a rollout completed and the goal is not yet met), or wait (an experiment or rollout is still in progress).

The entry point for the system is an outcome, not a task list. The team picks a business metric from the available set — add-to-cart rate, checkout conversion, bounce rate, or page-load-time-p95 — and sets a target improvement, for example “increase add-to-cart rate by 10%.” Error rate is reserved as a safety guardrail during the Guarded Release phase and cannot be chosen as the primary success metric, because the system needs an independent operational signal to decide whether a winning variation is safe to scale. Beyond the metric and the target, all other inputs are optional. The agent infers the current baseline, the areas of the application in scope for changes, and any constraints from the codebase and production data. If those assumptions are off, the team corrects them before any code is written. The team states where they want to end up, and the agent works backward from there.

Ecommerce demo store product listing page showing a grid of six products: Wireless Headphones at $149.99, Bluetooth Speaker at $79.99, USB-C Hub at $49.99, Mechanical Keyboard at $129.99, Leather Wallet at $59.99, and Canvas Backpack at $89.99. Each product card displays a product photo with name and price below. The page header shows "Demo Store" with Products and cart navigation links.

Demo Store product listing page used as the test surface for the add-to-cart experimentation cycles. Product cards currently show the control layout (no inline Add to Cart button).

For new goals, the agent explores the target repository and proposes a code change likely to move the metric. For iterations, it reads prior outcomes and adjusts its approach based on what worked and what did not. Before any code change, the agent creates a feature flag in LaunchDarkly (boolean, OFF by default, named with a convention like exp-add-to-cart-*) so every change ships behind a flag from the start.

Implementation runs through Kiro CLI in headless mode. The agent calls create_task, Kiro clones the repository, writes the change behind the feature flag, and opens a pull request.

GitHub merged pull request titled "feat: add inline Add to Cart button on listing page (atc-on-listing)" by gteppel. The PR merged 1 commit into main from experiment/atc-on-listing with 2 files changed. The description lists changes to page.tsx and a new ListingAddToCartButton.tsx component, explains flag-true and flag-false behavior, documents the atc-on-listing feature flag key with two variations, and notes TypeScript verification.

Merged GitHub PR implementing the feature-flagged inline Add to Cart button on the product listing page, controlled by the atc-on-listing LaunchDarkly flag.

AWS DevOps Agent then runs a release readiness review on the PR. If the review fails, the agent retries up to three times before stopping to ask for help. After a green review, the PR is merged and a GitHub Actions workflow deploys through AWS Amplify.

AWS DevOps Agent Release Readiness Review report for "Add to Cart Urgency Boost," completed on August 26, 2026. The report shows a recommended action of Standard Deployment, zero critical issues, commit d60dc4c, and 3 detected changes (all additions). The analysis section confirms all new behavior is gated behind the LaunchDarkly flag exp-add-to-cart-urgency-boost with a safe default of false. Recommendations include guarded rollout starting at a small treatment percentage, confirming the flag exists in LaunchDarkly, monitoring add-to-cart and checkout conversion metrics, and verifying treatment audience overlap.

AWS DevOps Agent Release Readiness Review for the Add to Cart Urgency Boost experiment. The automated review found zero critical issues and recommended standard deployment with a guarded rollout.

Proving the change

Once deployed, the flag is toggled on and the experiment begins. The agent creates a 50/50 experiment across 10% of traffic, splitting on the goal’s business KPI. In production, experiment data comes from real users interacting with your application, with metrics emitted through OpenTelemetry to LaunchDarkly. For this reference implementation, we built a synthetic traffic generator that simulates user sessions across both treatment and control variations, producing the conversion events and operational metrics that drive experiment decisions. It runs alongside the demo application and generates enough volume to reach statistical significance within minutes rather than days. The synthetic traffic generator is a demo convenience, not a production requirement. Any application that emits the right events to LaunchDarkly will work with this architecture.

The agent checks for results on each Custom Agent execution until statistical significance is reached. In an interactive chat session, you prompt the agent to check when you are ready. If the treatment wins, the agent proceeds to the Guarded Release. If it loses, the agent archives the flag and records the outcome for the next iteration.

LaunchDarkly experiment results dashboard showing Exposures and Summary panels. Exposures panel shows 17,977 user contexts over 1 hour with a 50/50 split between Control (no listing CTA) and Treatment (listing CTA). Summary panel shows a Healthy status, 1-day duration on August 26 2026, Treatment shipped as the winning variation with a relative difference of plus 1.0 and 100% probability to beat control. The experiment was stopped because Treatment beat control with plus 98.7% relative lift, statistically significant.

LaunchDarkly experiment summary for the inline Add to Cart listing CTA test. Treatment won decisively with 98.7% relative lift in add-to-cart conversion and 100% probability to beat control.

The Guarded Release ramps the winning variation from 20% to 30% to 40% while LaunchDarkly [1] applies sequential testing to the operational guardrail metric, halting the rollout as soon as the data shows a statistically significant regression against the original variation.. If a guardrail threshold is breached at any stage, LaunchDarkly reverts flag state at runtime without a redeployment. Guarded Releases and automatic rollback serve as the runtime safety net: if something goes wrong after deployment, the system reverts flag state without waiting for a human to intervene.

To validate the safety net in the reference implementation, we triggered a simulated error-rate spike during the ramp. LaunchDarkly detected the regression within the monitoring window, halted the rollout, and reverted the flag to its pre-rollout state automatically. No human intervened, no redeployment ran, and the application returned to the control behavior within seconds. The screenshot below shows the Guarded Release dashboard after the rollback.

LaunchDarkly Guarded Release dashboard showing an automatic rollback triggered by an error rate regression. A red banner states the default rule rolled back automatically after detecting a regression for Error Rate, ended August 27 at 10:23 AM. The error rate chart shows the treatment (true) variation at 0.507% versus control (false) at 0.498% with a sample size of approximately 500 per variation. The system rolled back to serving the false variation.

LaunchDarkly Guarded Release auto-rollback event. The system detected an error rate regression during the ramp phase and automatically rolled traffic back to the control variation.

After recording the rollback and feeding the outcome into the next iteration, the agent adjusted its approach and proposed a revised implementation that avoided the latency regression. The second attempt followed the same pipeline: hypothesis, feature flag, implementation, review, deployment, experiment, and Guarded Release. This time, monitoring completed with no regressions detected. LaunchDarkly rolled the winning variation forward to full traffic, with add-to-cart conversion lifting from 20.1% to 37.9% across the treatment population, confirming the experiment result held at scale.

LaunchDarkly Guarded Release dashboard showing successful monitoring completion. A green banner states monitoring completed on the default rule, ended August 27 at 10:48 AM. The Add to Cart metric chart shows the treatment (true) variation at 37.9% conversion versus control (false) at 20.1%, a lift of plus 17.7 percentage points. No regressions were detected, and the default rule rolled forward to serve the true variation. Sample sizes are 821 (true) and 864 (false).

LaunchDarkly Guarded Release monitoring completion. The Add to Cart metric showed a 17.7 percentage point lift with no regressions, so the system graduated the treatment to 100% of traffic.

After each cycle, the agent generates a report documenting the hypothesis, experiment results, rollout outcome, and a recommendation for the next iteration. This report feeds into the next decision, so no context is lost between cycles.

Add-to-Cart Experimentation Log showing a cycle summary table with three experiment cycles. Goal is to increase the add-to-cart metric by 10% in the default project and production environment. Cycle C tested adding an Add to Cart button directly to the listing page, resulted in a Winner outcome with plus 22.6% lift (significant, probability to beat baseline 98%), and was rolled out to 100%. Cycle B tested changing the button color from blue to green/orange, resulted in Inconclusive with plus 2.1% lift (not significant, approximately 120 units). Cycle A tested changing button placement on the detail page, resulted in Inconclusive with minus 1.4% lift (not significant, approximately 98 units).

Experimentation cycle summary showing three hypothesis-test iterations. Only Cycle C (inline Add to Cart on listing page) reached statistical significance and was promoted to production. The two cosmetic experiments (button color and placement) were inconclusive.

Safety boundaries

The system operates within defined constraints. The agent validates every change through release readiness review before merge. It creates a feature flag before writing any code, so every change can be toggled off without a redeployment. Guarded Releases enforce operational guardrails at runtime with automatic rollback. The agent retries failed validations up to three times, then stops and asks for help rather than proceeding. All credentials are stored in AWS Secrets Manager and referenced by name only, never exposed in agent logs or tool calls.

Getting started

To implement this workflow, you need AWS DevOps Agent enabled in your AWS account, a LaunchDarkly account (start with a free 30-day AWS trial), and a target application and repository. The reference uses a Next.js app deployed through AWS Amplify. Experiments are available on every LaunchDarkly plan, including the free Developer plan. Guarded Releases, which automate progressive rollouts with automatic rollback, require a LaunchDarkly Enterprise plan with the Guardian add-on. Without Guarded Releases, the workflow still runs experiments and reports results. You manage the rollout manually instead. If your plan does not include Guarded Releases, update the agent skill definition below to remove the Guarded Release actions.

Setup requires three steps. First, add the LaunchDarkly remote MCP server to your AWS DevOps Agent space. Second, deploy the Experiment MCP Server container to an AgentCore runtime, storing API keys and tokens in AWS Secrets Manager. Third, create your custom agent with the orchestration skill. Use the experimentation skill in AWS DevOps Agent to guide you through defining goals, connecting the MCP servers, and writing the orchestration instructions. The full orchestration skill is included below.

---
name: "experiment-orchestration"
description: "Orchestrates automated experimentation lifecycle using LaunchDarkly Guarded Rollouts, an AI coding agent for implementation, and GitHub Actions for deployment."
---
 
# Automated Experimentation
 
Use this skill when you have a goal you want to move through experimentation (e.g., "increase checkout conversion by 15%", "decrease page load time by 20%").
 
**Core principle: experiment first, then guarded rollout.** Always prove a change on a small, fixed slice of traffic via an A/B experiment before ramping it up through a guarded rollout. Never start a guarded rollout blind — it exists only to scale a change the experiment has already shown to work.
 
**Execution mode:** once the goal is confirmed (Step 1), run Steps 2–8 end-to-end. Async operations (code implementation, release review, deployment, experiment monitoring, rollout monitoring) should be checked periodically, not tight-polled — see the waiting note in each step. Only stop and ask the user something if a step fails unrecoverably (repeated failed release reviews, deployment failure, or an inconclusive/losing experiment result).
 
**The final report (Step 8) is mandatory, not optional.** The moment an experiment or rollout reaches a terminal outcome — winner, loser, inconclusive, or rollback — produce the full report in the same turn you announce the outcome. Don't let a casual "it worked! ????" substitute for the structured report.
 
## Step 1: Goal Clarification
 
Before doing anything, get answers to:
 
1. **What metric measures success?** *(Required)* e.g. conversion rate, page load time, bounce rate. Reserve your error-rate metric as a safety guardrail — never use it as the primary success metric.
2. **What's the target improvement?** *(Required)* e.g. 15% increase, 200ms decrease.
3. **What's the current baseline?** *(Optional — infer from production metrics if not given)*
4. **What parts of the app are in scope?** *(Optional — infer from the codebase if not given)*
5. **Any constraints?** *(Optional)* e.g. no changes to the payment flow.
 
Questions 1–2 are required before proceeding; infer 3–5 where possible and confirm your assumptions with the user before implementing.
 
## Step 2: Hypothesis Generation
 
Explore the target repository/codebase to find a plausible change:
 
1. Search and read the relevant code paths.
2. Think through what UI/UX or logic change could plausibly move the chosen metric.
3. Check whether this hypothesis (or something close to it) has already been tried and failed — look at flag history or archived flags with similar naming. Avoid repeating a known failure.
4. Present the hypothesis to the user before proceeding, along with your reasoning and any inferred assumptions from Step 1.
 
**Before finalizing a flag key, check for collisions:** look up any candidate flag key first.
- Already fully shipped (100% one variation, no split) → already decided, pick a different hypothesis.
- Actively running an experiment → mid-flight, don't compete with it, pick a different hypothesis.
- Doesn't exist → safe to create.
 
## Step 3: Implementation
 
1. Create a boolean feature flag, OFF by default in all environments. Name it with a clear pattern like `exp-<metric>-<short-description>` (e.g. `exp-checkout-conversion-cta-color`), lowercase with hyphens, ~50 chars max.
2. Hand off implementation to your coding agent/tool of choice, with clear instructions to gate the change behind the exact flag key from step 1.
3. This step is asynchronous — check status periodically rather than looping tightly on it.
4. Once implementation completes, move to Step 4 with the resulting branch/PR. If it fails, report the error and stop.
 
## Step 4: Release Readiness
 
Run your standard release/risk review on the PR before merging.
 
- If it passes: merge the PR.
- If it fails: feed the review's specific feedback back into implementation and retry. Cap retries at a small fixed number (e.g. 3 attempts total) — if it still hasn't passed, stop and report the last failure to the user rather than retrying indefinitely.
 
*(If your environment genuinely has no review capability available — e.g., a fully unattended automation context — you can skip straight to merge, but treat that as a deliberate, narrow exception you call out explicitly, not a default. Skipping review removes your only gate against shipping broken code.)*
 
## Step 5: Deployment
 
Deployment typically won't fire automatically on merge if your workflow is manually-triggered (`workflow_dispatch`-only) — you'll need to trigger it explicitly.
 
1. Trigger the deploy workflow on the merge target branch. Treat "already an in-progress deployment for this ref" as expected de-duplication, not an error — don't re-trigger.
2. Poll for status, but let your polling tool's own internal long-poll do the waiting rather than looping tightly yourself.
3. Watch for a "stale" status specifically: if a deployment reports "running" for far longer than normal, cross-check the actual CI run history by commit SHA/timing before assuming it's still in progress — a background poll process may have died without updating the record.
4. **Trigger a deployment at most once per attempt.** If you're unsure whether a previous trigger succeeded, check status first — never re-trigger just because you're unsure.
5. On timeout: stop, check the CI run directly, report the situation, ask how to proceed.
6. On explicit failure: stop and report — do not proceed to the experiment.
7. On success: proceed immediately to Step 6.
 
## Step 6: Experiment Phase (fixed 10%)
 
Prove the change on a small, fixed slice of traffic. Do **not** start a guarded rollout here — that's Step 7, and only after this proves out.
 
1. Turn the flag ON.
2. Configure a fixed 50/50 split across 10% of traffic (a flat allocation, not a staged ramp) on your chosen randomization unit (typically "user"). The remaining 90% of traffic is excluded from the experiment entirely. 
3. Create an experiment with:
   - Exactly one primary metric: the success metric from Step 1.
   - Guardrail metric(s): always include your error-rate metric; add a performance metric (e.g. p95 page load time) too if this is a performance-focused change.
   - Treatments: control (off) at 50%, treatment (on) at 50%, allocated to 10% of total traffic.
4. Start the experiment/data collection.
5. Move to Step 7 to monitor toward a decision.
 
## Step 7: Monitoring & Outcome
 
Check status periodically — don't tight-loop. In an interactive session, check once and report progress, then pick back up later. In an unattended/scheduled context, check once per invocation and persist your progress somewhere durable between runs.
 
**Phase 1 — Prove the experiment at 10% (gate before any rollout):**
 
Watch for statistical significance on the primary metric:
 
- **Significant + positive lift** → experiment proven. Stop the experiment iteration and move to Phase 2.
- **Significant + negative lift** → declare a loser, archive the flag, skip Phase 2, go straight to the Step 8 report.
- **No significance after a reasonable ceiling (e.g. 30 minutes)** → report "inconclusive, need more traffic" and stop; don't proceed to Phase 2.
 
Never declare a winner off a single data point or before your stats engine confirms significance.
 
**Phase 2 — Guarded rollout ramp (only after Phase 1 proves the change):**
 
Start a guarded rollout with:
- The winning ("on") variation as the test, the original as control.
- Same randomization unit as the experiment.
- **Exactly 3 monitored stages, capped well below 100%** — e.g. 20% → 30% → 40%, ~60 minutes monitoring each. Don't add a stage at or above 100%; Guarded-rollout implementations reject stages above 50% audience allocation, and the rollout auto-promotes to 100% itself once the final monitored stage completes cleanly — no explicit 100% stage needed.
- The same primary + guardrail metrics as the experiment, each configured to notify and auto-rollback on regression.
 
Track stage progression. If the rollout rolls back or stops at any point, treat it as a regression: declare failed, clean up the flag (deprecate/archive it), and go to the Step 8 report.
 
Once the final stage completes cleanly and auto-promotes to 100%, declare a winner and go to the Step 8 report.
 
**Retrying after a rollback:** a rollback isn't always caused by your monitored metrics genuinely regressing — it can also be triggered by an unrelated application error surfacing mid-ramp. Before blindly restarting after the user says they've fixed something:
1. Confirm the flag's current state (should be back to 100% control, nothing stuck mid-rollout).
2. Check the change history timing between "advanced to next stage" and "reverted." A rollback within seconds of advancing is inconsistent with a full metric-window regression and points to an external cause instead.
3. If the flag is cleanly reverted and the external cause is confirmed fixed, it's safe to restart the guarded rollout from scratch with the same parameters.
4. Don't silently retry without this check, and don't refuse to retry just because a prior attempt rolled back — a genuinely fixed external cause is a legitimate reason to retry. A metric-driven loser is not — don't retry that.
 
**On any terminal outcome, immediately produce the Step 8 report in the same turn** — a one-line "it worked!" note is fine as a lead-in, but the structured report must follow, not wait for a follow-up request.
 
## Step 8: Report
 
Runs automatically the instant Step 7 reaches a terminal outcome (winner + auto-promoted to 100%; loser; inconclusive; or rollback/failure). Use this exact structure:
 
```
## Experiment Report: [Goal Description]
 
**Date:** [YYYY-MM-DD]
**Goal:** [metric] [direction] by [target]%
**Status:** [achieved / in progress / stalled]
 
### Hypothesis
[What we tried and why]
 
### Implementation
- Flag: [flag_key]
- Files modified: [list]
- Branch: [branch name]
 
### Release Readiness
- [reviewed, passed after N attempt(s) / skipped, per your environment's process]
 
### Experiment Phase (10% fixed split)
- Status: [proven / loser / inconclusive]
- Duration: [time]
- Metric change: [before] → [after] ([+/-]%)
- Statistical significance: [value, confidence interval]
 
### Guarded Rollout Phase (if reached)
- Status: [completed / rolled_back / not started]
- Duration: [time]
- Stages reached: [N of 3 monitored stages]
- If rolled back and retried: [root cause, outcome of retry]
 
### Safety Metrics
- error-rate: [baseline] → [final] ([no regression / regression detected])
- [other guardrails]: [baseline] → [final] ([status])
 
### Next Steps
[What to do next based on the outcome]
```
 
## Safety Rules (the non-negotiables)
 
- Always present the hypothesis before implementing.
- Always run a release/risk review before merging, unless your environment has a deliberate, explicitly-called-out exception.
- **Always prove a change via a fixed small-percentage experiment before starting any guarded rollout** — never ramp blind.
- Always include an error-rate (or equivalent "don't break prod") metric as a guardrail, separate from your success metric.
- Add a performance guardrail (e.g. p95 latency) for performance-focused changes.
- Every rollout metric should be configured to both notify AND auto-rollback on regression — don't rely on notification alone.
- **Cap guarded rollout stages well below 100%** (most platforms reject stages ≥50% audience allocation) and let the platform auto-promote to 100% after the final stage — don't try to add an explicit 100% stage.
- Distinguish a metric-driven rollback (don't retry) from an external-cause rollback (safe to retry once fixed) before restarting a rolled-back rollout.
- The final report is automatic and mandatory on every terminal outcome — never defer it to a follow-up ask.

Conclusion

This post described how AWS DevOps Agent, Kiro CLI, and LaunchDarkly connect into a closed-loop system that turns an improvement goal into a series of measured, safe experiments. The agent runs autonomously on a schedule: it generates hypotheses informed by prior outcomes, creates feature flags before any code change, invokes Kiro CLI in headless mode to implement changes behind those flags, validates through release readiness review, deploys through GitHub Actions and AWS Amplify, and hands off to LaunchDarkly for experiment measurement and guarded rollout. If a guardrail is breached at any point during the rollout, LaunchDarkly reverts flag state at runtime without a redeployment. After each cycle, the agent records what happened and feeds it into the next decision.

This directly addresses the three barriers that slow experimentation:

● Planning cost is reduced because the agent handles hypothesis generation, flag creation, implementation coordination, and validation. The team defines the goal; the system handles the wiring.

● Measurement disconnected from action is addressed because LaunchDarkly monitors metrics in real time and reverts flag state automatically when a guardrail is breached, requiring no redeployment and no waiting for a human to notice.

● Stalled iteration is solved because every outcome is recorded and fed into the next hypothesis automatically. The system does not forget what it learned, and it does not stall between iterations.

The architecture is available to implement today as a reference. The orchestration skill included in this post encodes the full 8-step workflow: goal clarification, hypothesis generation, implementation, release readiness, deployment, experiment, monitoring, guarded rollout, and reporting. Teams define their improvement goal, connect the LaunchDarkly MCP server and the Experiment MCP Server to a DevOps Agent custom agent, and let the system iterate toward the target within the safety boundaries they configure. A more turnkey experience is planned for the future.

Authors

Greg Eppel

Greg Eppel is a Principal Specialist for DevOps Agent and has spent the last several years focused on Cloud Operations and helping AWS customers on their cloud journey.

Jonathan Nolen

Jonathan Nolen is the CPO at LaunchDarkly, the leading platform for Runtime Control for AI software development.

He first joined LaunchDarkly in 2018 and has led the Product, Engineering and Design teams. Currently he is leading the team to build critical infrastructure that helps thousands of customers deliver at agentic speed and still ship software safely. Jonathan was at Atlassian from 2005 until 2018. He helped grow the company from 25 employees to over 2,500 and contributed to multiple Atlassian products. Jonathan and his team also built the Atlassian Marketplace, which in 2024 had done over $3 billion of business for the Atlassian community.

Query Amazon S3 Tables from Amazon EMR Trino using the Iceberg REST endpoint

Post Syndicated from Shubham Purwar original https://aws.amazon.com/blogs/big-data/query-amazon-s3-tables-from-amazon-emr-trino-using-the-iceberg-rest-endpoint/

Organizations running analytics on Amazon Simple Storage Service (Amazon S3) data lakes often struggle with the operational overhead of managing Apache Iceberg tables, including compaction, snapshot expiration, and metadata tracking, while still needing fast, interactive SQL access across large volumes of data. Amazon S3 Tables, a capability of Amazon S3, addresses this by providing a purpose-built storage layer with native Apache Iceberg support and automated table maintenance. When you query S3 Tables from Amazon EMR using Trino and the Iceberg REST endpoint, you get a fully managed, open-standards-based analytics stack without the undifferentiated heavy lifting of table upkeep.

When paired with Amazon EMR running Trino, organizations gain access to a high-performance distributed SQL query engine capable of processing large-scale datasets. Trino’s ability to query data across multiple sources, combined with the automated optimization features of S3 Tables, creates a flexible analytics platform. The integration uses Apache Iceberg’s REST catalog specification, providing a standardized interface that supports compatibility across different compute engines while maintaining full control over query execution and data processing logic.

This architectural pattern is particularly valuable for organizations seeking to modernize their data platforms without vendor lock-in, as it relies on open standards and formats. The solution delivers high-throughput query performance with distributed SQL execution while significantly reducing the operational burden of managing table metadata, compaction, and snapshot lifecycle management. In this post, we show you how to create and query Amazon S3 Tables using Trino on Amazon EMR through the Apache Iceberg REST catalog endpoint.

Solution overview

This implementation demonstrates a complete integration between the Trino distribution on Amazon EMR and Amazon S3 Tables through the Apache Iceberg REST catalog endpoint. The architecture uses several key AWS services working in concert:

Amazon EMR serves as the managed compute layer, providing a scalable Hadoop framework that hosts the Trino query engine. Amazon EMR handles cluster provisioning, configuration management, and automatic scaling, allowing teams to focus on analytics rather than infrastructure management.

Apache Trino acts as the distributed SQL query engine, offering ANSI SQL compatibility and the ability to process queries across massive datasets with low latency for interactive workloads. Its connector architecture supports integration with various data sources, including the Iceberg REST catalog.

Amazon S3 Tables provides the storage and catalog layer, managing Apache Iceberg tables with built-in optimization. The service automatically handles compaction, snapshot expiration, and metadata management, reducing operational overhead while maintaining query performance. S3 Tables exposes a REST API endpoint that conforms to the Apache Iceberg REST catalog specification, which provides standardized integration with any Iceberg-compatible engine.

Apache Iceberg REST endpoint serves as the communication protocol between Trino and S3 Tables. This RESTful interface handles catalog operations including namespace management, table creation, metadata retrieval, and transaction coordination. The endpoint supports AWS Signature Version 4 authentication for secure access to table resources.

The data flow follows this pattern: Users submit SQL queries through the Trino CLI or JDBC interface. Trino’s Iceberg connector communicates with the S3 Tables REST endpoint to retrieve table metadata and plan query execution. The query engine then reads data directly from S3 using optimized file formats (Parquet, ORC) while using Iceberg’s metadata layer for partition pruning and predicate pushdown. Write operations follow a similar path, with Trino coordinating with S3 Tables to commit new data files and update table metadata atomically.

This architecture delivers several key benefits: separation of compute and storage for independent scaling, automated table maintenance reducing operational costs, open-source format compatibility preventing vendor lock-in, and fine-grained access control through AWS Identity and Access Management (IAM) and AWS Lake Formation integration.

Architecture diagram showing Trino on Amazon EMR querying Amazon S3 Tables through the Apache Iceberg REST catalog endpoint

Figure 1: Solution architecture for querying Amazon S3 Tables from Trino on Amazon EMR

Prerequisites

Before getting started, make sure that you have the following:

  • An active AWS account with billing enabled.
  • An AWS Identity and Access Management (IAM) user with specific permissions to create and manage resources, such as a virtual private cloud (VPC), subnet, security group, IAM roles, Amazon EMR, Interface VPC endpoints, S3 Tables bucket and S3 buckets.
  • Sufficient VPC capacity in your chosen AWS Region.

For this post, we create the solution resources in the US East (N. Virginia) Region (us-east-1) using AWS CloudFormation templates. In the following sections, we show you how to configure your resources and implement the solution.

Note: Querying Amazon S3 Tables through Trino on Amazon EMR requires Trino version 475 or later, available in Amazon EMR 7.11 and later.

Part A: Configure Amazon S3 Tables integration with Trino on Amazon EMR using AWS CloudFormation

In this post, you use the CloudFormation template emr-trino-s3tables.yaml.

  • This template deploys the following resources: a VPC with one private subnet, an S3 Tables interface VPC endpoint for private access, and an Amazon EMR cluster running Trino integrated with Amazon S3 Tables through the Apache Iceberg REST catalog endpoint.
  • It also creates an S3 Tables bucket, a general-purpose S3 bucket, IAM roles, and security groups.
  • At deploy time, it dynamically generates the Trino catalog configuration and bootstrap script.

To create the solution resources, complete the following steps:

  1. Launch the stack emr-trino-s3tables.yaml using the CloudFormation template.

Launch Cloudformation Stack

  1. Provide the parameter values as listed in the following table.
Parameters Description Sample value
Stack Name Name of CloudFormation stack emr-s3tables-trino
VPC CIDR block IP range (CIDR notation) for this VPC. 10.0.0.0/16
Private Subnet CIDR block IP range (CIDR notation) for the private subnet in the second Availability Zone. 10.0.1.0/24
Resource name Prefix Short prefix applied to every resource name emr-s3tables
S3 Tables bucket name Name of S3 table Bucket trinoemrs3tablebuck
EMR release Release version of Amazon EMR EMR 7.12

The stack creation process can take approximately 15 minutes to complete. You can check the Outputs tab for the stack after the stack is created, as shown in the following screenshot.

Figure 3: CloudFormation stack outputs

Figure 3: CloudFormation stack outputs

Understanding the deployment

The CloudFormation template performs several key tasks:

  1. Infrastructure provisioning: Sets up the Amazon EMR cluster with Trino, VPC, subnet, security group, and S3 table bucket.
  2. Configuration: Creates necessary Trino configuration files.
  3. Integration configuration: Sets up the Iceberg REST connector for S3 Tables.

Part B: Connecting Trino to Amazon S3 Tables with Iceberg REST endpoint

The CloudFormation template automatically configures the S3 Tables catalog in Trino on Amazon EMR. In the next section, we examine the configuration that drives this integration.

1. Catalog configuration details

A catalog in Trino on Amazon EMR is the configuration that grants access to a specific data source. Each Trino on Amazon EMR cluster can have multiple catalogs configured, allowing access to different data sources simultaneously.

As part of this setup, the CloudFormation template creates a catalog properties file at /etc/trino/conf/catalog/s3tables_irc.properties with the following configuration:

connector.name=iceberg
iceberg.catalog.type=rest
iceberg.rest-catalog.uri=https://s3tables.<REGION>.amazonaws.com/iceberg
iceberg.rest-catalog.warehouse=arn:aws:s3tables:AwsRegion:<ACCOUNT-ID>:bucket/<BUCKET-NAME>
iceberg.rest-catalog.sigv4-enabled=true
iceberg.rest-catalog.signing-name=s3tables
iceberg.rest-catalog.view-endpoints-enabled=false
fs.hadoop.enabled=false
fs.native-s3.enabled=true
s3.region=us-east-1
s3.iam-role=arn:aws:iam::<ACCOUNT-ID>:role/service-role/<ROLE-NAME>

2. S3 Tables Iceberg REST endpoint configuration properties

The following table lists the key properties in the catalog configuration on Trino:

Property name Description
iceberg.rest-catalog.uri REST server API endpoint URI (necessary).
iceberg.rest-catalog.warehouse Warehouse ID or location for the catalog (necessary). For S3 Tables, this is the ARN for the S3 table bucket as shown in the preceding properties example.
iceberg.rest-catalog.sigv4-enabled Must be set to ‘true’ (necessary)
iceberg.rest-catalog.signing-name Must be set to ‘s3tables’ (necessary)
iceberg.rest-catalog.view-endpoints-enabled Must be set to ‘false’ (necessary)
fs.hadoop.enabled Must be set to ‘false’
fs.native-s3.enabled Must be set to ‘true’
s3.iam-role Amazon Resource Name (ARN) of the IAM role with permissions to S3 Tables. In this post, we use the same role, which is the service role for Amazon EMR.
s3.region AWS Region, for example us-east-1

This configuration establishes a connection between Trino and the S3 Tables REST endpoint. You can have multiple catalogs registered, one per S3 table bucket, which is determined by the iceberg.rest-catalog.warehouse property.

3. Configure Amazon EMR service IAM role trust relationships

The Amazon EMR service role requires proper trust relationships to function correctly. Navigate to the IAM console and configure the trust policy for your Amazon EMR service role:

{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Principal": {
                "Service": "elasticmapreduce.amazonaws.com"
            },
            "Action": "sts:AssumeRole"
        },
        {
            "Effect": "Allow",
            "Principal": {
                "AWS": "arn:aws:iam::<ACCOUNT-ID>:role/service-role/AmazonEMR-InstanceProfile"
            },
            "Action": "sts:AssumeRole"
        }
    ]
}

This trust policy establishes two critical relationships:

  1. The Amazon EMR service can assume the role to manage cluster operations.
  2. The EC2 instance profile can assume the role to access S3 Tables with elevated permissions.

4. Working with S3 Tables in Trino on Amazon EMR

Now that you have Trino on Amazon EMR set up and configured to work with S3 Tables, you can explore how to work with this integration.

4.1. Connecting to Trino on Amazon EMR

Navigate to Amazon EMR and select Connect to the primary node using AWS Systems Manager Session Manager for passwordless SSH.

Figure 4: Connecting to the primary node with Session Manager

When you’re connected, you can use the Trino CLI with your S3 Tables catalog:

sudo su - hadoop
trino-cli --catalog s3tables_irc

This connects you to the Trino on Amazon EMR using the S3 Tables integration you configured.

Trino CLI connected to the s3tables_irc catalog on Amazon EMR

Figure 5: Trino CLI connected to the S3 Tables catalog

4.2. Examples: Creating and querying tables

In this section you run through some example queries to demonstrate the functionality.

4.2.1 Creating a namespace

First, you create a namespace (schema) in S3 Tables. A namespace in S3 Tables is a logical container or organizational unit that helps group related tables and objects together.

CREATE SCHEMA blog_namespace;
USE blog_namespace;

4.2.2 Creating a table

Create a table with various data types. You don’t need to specify the table type as Iceberg explicitly because you’re connecting to the Iceberg catalog. You can use all standard Iceberg capabilities, such as partitioning and sorting. Furthermore, some of the important Iceberg table properties that support table maintenance operations are configured with default values. You also have the option to edit the configurations using S3 Tables maintenance APIs.

CREATE TABLE IF NOT EXISTS customers (
customer_sk INT,
customer_id VARCHAR,
salutation VARCHAR,
first_name VARCHAR,
last_name VARCHAR,
preferred_cust_flag VARCHAR,
birth_day INT,
birth_month INT,
birth_year INT,
birth_country VARCHAR,
login VARCHAR
) WITH (
format = 'PARQUET',
sorted_by = ARRAY['customer_id']
);

Table property explanation:

  • format = 'PARQUET': Specifies Parquet as the file format for optimal compression and query performance.
  • sorted_by = ARRAY['customer_id']: Defines sort order within data files, improving query performance for customer_id filters.

Verify the table creation:

SHOW TABLES;

You should see customers in the output, confirming the table exists in the S3 Tables catalog.

4.2.3 Inserting data

You can insert some sample data into your table. You can also use an existing table in any of the catalogs configured in Trino on Amazon EMR to read data and write into the S3 table with an INSERT INTO ... SELECT statement.

INSERT INTO customers VALUES
(1, 'AAAAA', 'Mrs', 'Martha', 'Rivera', 'Y', 8, 4, 1984, 'US', 'mrivera'),
(2, 'AAAAB', 'Mr', 'Mateo', 'Jackson', 'N', 22, 6, 2001, 'US', 'mjackson'),
(3, 'BAAAA', 'Ms', 'Mary', 'Major', 'Y', 16, 2, 1999, 'US', 'mmajor'),
(4, 'BBAAA', 'Mr', 'Paulo', 'Santos', 'N', 30, 3, 1973, 'US', 'psantos'),
(5, 'AACAA', 'Ms', 'Ana', 'Silva', 'N', 2, 6, 1982, 'CA', 'asilva'),
(6, 'ABAAA', 'Mr', 'Alejandro', 'Rosalez', 'N', 5, 12, 1988, 'US', 'arosalez'),
(7, 'BBAAA', 'Ms', 'Nikki', 'Wolf', 'N', 6, 1, 2006, 'MX', 'nwolf'),
(8, 'ACAAA', 'Mr', 'Arnav', 'Desai', 'N', 15, 7, 1976, 'US', 'adesai');

This INSERT operation demonstrates Trino’s ability to write data to S3 Tables. Behind the scenes, Trino:

  1. Writes data files in Parquet format to S3.
  2. Communicates with the S3 Tables REST endpoint to register the new files.
  3. Atomically commits the transaction, updating table metadata.

4.2.4 Querying data

Execute a SELECT query to retrieve and verify the inserted data:

SELECT * FROM customers LIMIT 10;

The query should return all eight customer records with proper formatting. You can also execute more complex analytical queries:

-- Count customers by country
SELECT birth_country, COUNT(*) as customer_count
FROM customers
GROUP BY birth_country
ORDER BY customer_count DESC;

-- Find customers born after 1990
SELECT first_name, last_name, birth_year
FROM customers
WHERE birth_year > 1990
ORDER BY birth_year;

These queries demonstrate Trino’s SQL capabilities and the integration with S3 Tables for both read and write operations.

4.3 Explore advanced features

S3 Tables with Iceberg provides several features for data management:

4.3.1 Time travel queries

Step 1: Check available snapshots.

-- Query table as of a specific timestamp. Check available snapshots
SELECT * FROM "customers$snapshots";

Step 2: Query the table as of a specific snapshot.

SELECT * FROM customers FOR VERSION AS OF <snapshot_id_from_step1>;

4.3.2 Schema evolution

-- Add a new column
ALTER TABLE customers ADD COLUMN email VARCHAR;

-- Rename a column
ALTER TABLE customers RENAME COLUMN login TO username;

Cleaning up

To clean up the resources, navigate to CloudFormation and delete the stack that you created.

Conclusion

This solution demonstrates an integration between Amazon EMR Trino and Amazon S3 Tables using the Apache Iceberg REST catalog specification. In this post, we showed you how to create and query S3 Tables from Trino on Amazon EMR. The architecture delivers several advantages for modern data platforms:

Operational simplicity: S3 Tables eliminates the complexity of managing Iceberg table metadata, compaction schedules, and snapshot lifecycle policies. The service handles these operations automatically, allowing data teams to focus on analytics rather than infrastructure maintenance.

Performance at scale: The architecture is designed for large-scale workloads. Trino distributes query execution across the cluster while Iceberg’s metadata layer helps the engine locate only the relevant data files. Features like partition pruning, predicate pushdown, and columnar file formats can help improve performance for both interactive and batch workloads.

Cost efficiency: This architecture separates compute and storage, so you can scale each independently based on workload requirements. S3 Tables automatically compacts small files to help reduce storage overhead, and Amazon EMR clusters can scale dynamically so you pay for compute only when needed.

Open standards and portability: By using Apache Iceberg’s open table format and REST catalog specification, this solution avoids vendor lock-in. Other Iceberg-compatible engines can access tables created in S3 Tables including Apache Spark, Apache Flink, and Dremio, providing flexibility in tool selection.

Fine-grained access control: Integration with IAM and resource-based policies provides access control at the table bucket, namespace, and table level. For fine-grained access at the column and row level, you can integrate with AWS Lake Formation. AWS Signature Version 4 authentication supports secure communication between Trino and S3 Tables.

ACID transactions: Iceberg’s transaction model guarantees atomicity, consistency, isolation, and durability for all table operations. This supports reliable concurrent reads and writes, making the platform suitable for production workloads requiring data consistency.

This architectural pattern is particularly well-suited for organizations building modern data lakehouses, migrating from traditional data warehouses, or consolidating multiple analytics platforms. The combination of the managed compute of Amazon EMR, Trino’s versatile query engine, and the automated table management of S3 Tables creates a strong foundation for data-driven decision making.

To learn more about the services and features discussed in this post, see the following resources:


About the authors

Shubham Purwar

Shubham Purwar

Shubham is an AWS Analytics Specialist Solution Architect. He helps organizations unlock the full potential of their data by designing and implementing scalable, secure, and high-performance analytics solutions on AWS. In his free time, Shubham loves to spend time with his family and travel around the world.

Anirudh Chawla

Anirudh Chawla

Anirudh is an AWS Analytics Specialist Solution Architect. He helps organizations empower businesses to harness their data effectively through the analytics services of AWS. His interest lies in building highly available distributed systems.

Nitin Kumar

Nitin Kumar

Nitin is a Solutions Architect at AWS. He partners with customers to transform their cloud journey through innovative, scalable solutions. In his free time, he likes to watch movies and spend time with his family.

Prashanthi Chinthala

Prashanthi Chinthala

Prashanthi is a Cloud Engineer (DIST) at AWS. She helps customers overcome Amazon EMR challenges and develop scalable data processing and analytics pipelines on AWS.

Automate planned lifecycle upgrades with AWS DevOps Agent and Kiro

Post Syndicated from Nehal Sangoi original https://aws.amazon.com/blogs/devops/automate-planned-lifecycle-upgrades-with-aws-devops-agent-and-kiro/

AWS uses Planned Lifecycle Events (PLEs) for AWS Health to signal that a managed service version is approaching end of standard support. Several AWS services such as Amazon Elastic Kubernetes Service (Amazon EKS), Amazon Relational Database Service (Amazon RDS), Amazon OpenSearch Service, and Amazon ElastiCache publish these events through AWS Health when a running resource needs to move to a newer version before a published deadline. For the team receiving that alert, the work that follows is remarkably similar regardless of which service triggered it. Engineers must identify every affected resource across accounts and AWS Regions, determine the correct target version, and assess compatibility constraints for dependencies and consumers. They then update infrastructure-as-code (IaC) definitions to reflect the new versions, validate that no breaking changes are introduced, and deploy within the deadline. When multiple services reach end-of-support on overlapping timelines, each with dozens of affected resources, this per-service effort compounds into a sustained operational burden for engineering and operations teams.

AWS DevOps Agent is a frontier agent that resolves and proactively helps prevent incidents, continuously improving reliability and performance of applications on AWS and hybrid environments. AWS DevOps Agent helps review software changes for production risks while investigating incidents and identifying operational improvements as an experienced DevOps engineer.

AWS DevOps Agent and Kiro are transforming how organizations manage version upgrades across AWS managed services and turn these into a governed, event-driven workflow. The AWS DevOps Agent automates the investigation: it discovers impacted resources, analyzes upgrade paths, and produces a structured change specification. Kiro provides the agentic development environment to apply those changes, validate safety constraints, and open a pull request (PR) for human review. The engineer’s role shifts from executing the upgrade to reviewing a PR that has already been investigated, coded, and validated. The engineers can even write the upgrade logic as a custom AWS DevOps Agent skill, and the framework handles orchestration, validation, and delivery.

This post and the sample code demonstrates the approach with an end-to-end Amazon EKS upgrade example. The underlying pattern of event detection, agent-driven investigation, automated code changes, and a failure retry loop applies to other AWS managed services that publish AWS Health PLEs.

In this post, you will learn how to:

  • Automate planned lifecycle upgrade events detection using AWS Health and Amazon EventBridge
  • Use AWS DevOps Agent to investigate the upgrade path and produce a structured change spec.
  • Run Kiro CLI (headless mode) in a continuous integration and continuous delivery (CI/CD) pipeline to apply code changes, validate safety constraints, and open a pull request.
  • Close the loop with automatic upgrade deployment failure detection where a failed deployment triggers root-cause analysis, mitigation planning, operator notification, and a code fix pull request without human initiation.

Solution overview

The following diagram shows the end-to-end flow, from the initial AWS Health event through to the pull request and the pipeline upgrade loop.

Architecture diagram showing the end-to-end EKS upgrade pipeline. AWS Health publishes a Planned Lifecycle Event to Amazon EventBridge. An Amazon EventBridge rule triggers a Health Lambda function that signs and posts a webhook payload to AWS DevOps Agent. The agent runs the eks-upgrade-planning skill and emits an Investigation Completed event to Amazon EventBridge. A second Amazon EventBridge rule triggers the Trigger Lambda, which fetches journal records through ListJournalRecords, detects a CDK Change Spec, retrieves the GitHub PAT from AWS Secrets Manager, and dispatches the eks-upgrade.yml GitHub Actions workflow. GitHub Actions runs Kiro CLI in headless mode to apply CDK changes, validate with cdk synth, and open a pull request for human review. After merge, the eks-deploy.yml workflow runs cdk deploy and tags the CloudFormation stack with the originating investigation ID.

Figure 1: Architecture diagram of the automated upgrade pipeline

There are five main phases in this flow. Let’s walk through each phase.

Phase 1: Detection

a. The pipeline starts when AWS Health publishes an AWS_EKS_PLANNED_LIFECYCLE_EVENT to the default Amazon EventBridge bus with the following event details:

service: EKS
eventTypeCategory: scheduledChange
eventTypeCode: AWS_EKS_PLANNED_LIFECYCLE_EVENT
affectedEntities: <array of cluster ARNs with status: PENDING>
eventRegion: <region of the affected cluster>

b. An Amazon EventBridge rule named eks-health-planned-lifecycle matches this event and invokes the AWS Lambda function devops-agent-health-event.

c. The Lambda function extracts the relevant information (cluster name and region), builds a webhook payload with eventType: incident and priority: HIGH, and POSTs to AWS DevOps Agent webhook endpoint, instructing the agent to follow the eks-upgrade-planning skill for the specific cluster and region. The Lambda function does not validate those values, so a failed extraction can leave the investigation running against placeholder data.

Phase 2: Investigation

a. AWS DevOps Agent uses the eks-upgrade-planning skill to discover cluster topology, validate the version increment, check addon compatibility, scan for deprecated APIs, and determine upgrade sequence.

b. The agent outputs a structured AWS Cloud Development Kit (AWS CDK) Change Spec containing target version strings for every component, a rollback readiness assessment (confirming the 7-day rollback window will be available post-upgrade), a feasibility assessment (READY, BLOCKED, or NEEDS_REMEDIATION), and a risk rating.

c. When AWS DevOps Agent completes its investigation, it emits an Investigation Completed event to Amazon EventBridge with the following event details:

source: aws.aidevops
detail-type: Investigation Completed
detail.metadata.agent_space_id: <the agent space ID>
detail.metadata.task_id: <the backlog task ID>
detail.metadata.execution_id: <the execution ID>
detail.data.status: <investigation result status>

Phase 3: Code and validation

a. A second Amazon EventBridge rule devops-agent-investigation-events matches this event, filtered by agent_space_id so that only events from the specific agent space trigger the pipeline.

b. The rule invokes the Trigger Upgrade Lambda function (devops-agent-trigger-upgrade). This Lambda function fetches the investigation’s journal records through ListJournalRecords and scans the output for content markers to determine the next action. Markers are checked in a fixed priority order so that a failure investigation quoting upstream CLUSTER_VERSION context cannot accidentally re-trigger an upgrade workflow. When either a CDK Change Spec heading or a resolved CLUSTER_VERSION line is present, the Lambda function treats the investigation as having produced an actionable upgrade plan. It retrieves the GitHub Personal Access Token (PAT) from AWS Secrets Manager, builds the investigation metadata into a summary JSON, and dispatches the eks-upgrade.yml GitHub Actions workflow through the GitHub API. The dispatched payload is a compact summary record (~3.8 KB) containing the CDK Change Spec, not the full investigation transcript, which exceeds GitHub’s workflow dispatch size limit.

c. Before the workflow lets a coding agent near the code, it validates what the investigation produced. An extraction step scans the received payload for fenced code blocks containing CLUSTER_VERSION. Each candidate block is held to a strict format contract:

  • No leftover placeholder markers.
  • A Kubernetes version matching X.Y.
  • A kubectl layer package matching @aws-cdk/lambda-layer-kubectl-vNN.
  • Every addon version matching vX.Y.Z-eksbuild.N unless explicitly marked NOT_INSTALLED.

The workflow also enforces the agent’s own feasibility verdict. If the investigation concluded BLOCKED or NEEDS_REMEDIATION, the run stops and the coding agent is not invoked. When validation passes, the single deduplicated spec block is written to a temporary file for the coding step. The workflow stops with an error if no spec block is found, no block passes validation, or multiple conflicting specs are present. The pipeline fails closed rather than handing an ambiguous instruction to a coding agent.

d. GitHub Actions then installs Kiro CLI, gated on a minimum tested version, with anything newer allowed through but flagged as untested. The installer is downloaded and executed as two discrete steps rather than piped from curl, and Kiro is then invoked in headless mode:

kiro-cli chat --no-interactive --trust-tools=read,write,glob,grep \
"Read kiro-cdk-instructions.md for context on the CDK patterns. Then read /tmp/cdk-change-spec.txt — it contains the validated CDK Change Spec extracted from the DevOps Agent investigation. Apply those values exactly. Modify lib/iteration3-stack.ts ONLY. Do NOT derive or guess version numbers — use only the values from the spec file. Make only the file edits — do not run any build or shell commands, and do not commit."

e. Two things are worth noting about this invocation. Kiro is trusted with file tools only (read, write, glob, grep) with no shell or command execution, so the scope of the agent step is limited to file edits in the checked-out working tree. And it is told explicitly not to derive version numbers: every value comes from the validated spec file, so a model that misreads the investigation cannot substitute a version of its own. Kiro reads kiro-cdk-instructions.md, a standalone reference that prescribes the CDK modification procedure for EKS upgrades, then modifies lib/iteration3-stack.ts and nothing else. The kubectl layer dependency is handled separately, by npm, in a later step. Neither the AWS DevOps Agent nor Kiro can query a package registry, so neither can know which versions of that layer actually exist. The spec carries only the package name and npm resolves the version. It is the pipeline’s own principle applied to itself: identify what the model cannot know, and move it out of the model’s reach rather than letting it guess.

f. Two independent gates run after Kiro exits. The first diffs the working tree against a single-file allowlist and fails the run if anything other than lib/iteration3-stack.ts was touched. That diff is a containment check on the agent’s write access and only after that audit passes, a separate step updates the kubectl layer dependency in package.json. The second gate runs the full build and CDK synthesis pipeline, so a change that does not compile or synthesize does not create a pull request.

Phase 4: Review and deploy

a. After Kiro exits, the workflow opens a GitHub Pull Request (PR) on a branch named upgrade/eks-automated-<run_id>. Kiro’s role ends at file edits. It does not interact with Git or GitHub. The PR body includes a rollback window advisory documenting the 7-day reversal deadline, a reviewer checklist, and a machine-readable investigation-context block containing the agent space ID and task ID. The post-merge deploy workflow parses that block to tag the AWS CloudFormation stack, so a future upgrade failure carries a record of which investigation produced the deployed plan. The tag is informational only and the failure investigation is not linked to the upgrade investigation, keeping the two workstreams independent.

b. The automated pipeline pauses at the pull request. The Site Reliability Engineering (SRE) team reviews the changes using their existing approval process.

c. After merge, the team deploys using their standard CI/CD pipeline. The investigation-context tags on the stack enable traceability back to the originating event if issues arise.

Phase 5: Failure detection and automated mitigation

The pipeline includes a closed-loop failure path. If a deployed upgrade fails, the system automatically investigates the root cause, generates a mitigation plan, notifies the SRE team, and opens a code fix pull request, all without human initiation. The pipeline attempts this automated recovery once. If the failure investigation itself does not produce actionable results, the pipeline stops and we recommend manually reviewing the cluster upgrade failure through the AWS DevOps Agent console or standard operational runbooks.

With EKS version rollbacks now available, the eks-failure-root-cause skill evaluates whether a rollback is the faster recovery before recommending a code fix. In case a deployment failure occurs within the 7-day rollback window, the root-cause investigation first evaluates whether a version rollback would resolve the issue faster than a code fix. When rollback readiness checks pass and the root cause is version-related (not a code or configuration error), the skill directs the agent to recommend version rollback (aws eks update-cluster-version --kubernetes-version <previous-version>) as the primary recovery action, with the code fix PR as a follow-up hardening measure. If rollback is not viable (outside the window, node skew, forward-only addon changes), the pipeline continues to the existing code fix workflow.

The following diagram shows the failure path from CloudFormation rollback through to the code fix pull request and operator notification.

Architecture diagram showing the closed-loop failure path. An AWS CloudFormation rollback emits a stack status change event to Amazon EventBridge. An Amazon EventBridge rule triggers the Failure Lambda, which posts a signed webhook to the same AWS DevOps Agent space requesting root-cause analysis. A triage skill prevents linking to upgrade investigations. The agent produces a Root Cause section and emits an Investigation Completed event. The Trigger Lambda fetches journal records, detects the Root Cause marker without a Mitigation Plan, and calls UpdateBacklogTask to activate the Mitigation Agent. It then schedules a one-time Amazon EventBridge Scheduler check to poll for completion. When the Mitigation Agent finishes, the Trigger Lambda detects the Mitigation Plan marker and produces two parallel outputs: it dispatches the next-steps.yml GitHub Actions workflow where Kiro CLI implements the agent-ready specification as a code fix pull request, and it publishes the execution plan with immediate recovery steps to an Amazon SNS topic for operator notification.

Figure 2: Architecture diagram of the failure mitigation loop

a. When cdk deploy fails after merge, CloudFormation emits a stack status change event (such as ROLLBACK_FAILED, ROLLBACK_COMPLETE, UPDATE_ROLLBACK_FAILED, or UPDATE_ROLLBACK_COMPLETE) to Amazon EventBridge. An Amazon EventBridge rule (eks-cfn-stack-failure) matches one of these terminal rollback statuses and invokes the Failure Lambda function.

One point deserves emphasis before a responder acts on this event: a CloudFormation stack rollback does not revert an EKS control plane version. Reverting the template to one that specifies a lower Kubernetes version is not a cluster version rollback. That has to be initiated explicitly through the UpdateClusterVersion API, the AWS CLI, or the console. If CloudFormation had already updated the control plane before failing on a later resource, the stack can report a completed rollback while the cluster remains on the new version. Confirm the cluster’s actual Kubernetes version rather than inferring it from the stack status.

b. The Failure Lambda function opens a new investigation on the same agent space (eks-upgrade-poc) used for upgrade planning. The prompt instructs the agent to analyze the failure and produce a root-cause assessment. Using the scoping controls for agent sessions, a single agent space can handle both investigation types safely:

  • Global Instructions (applied to all agent types) enforce hard rules: “never reference findings from an upgrade-planning investigation when performing failure root-cause analysis” and vice versa. These always-on rules are the primary isolation boundary.
  • A triage skill (eks-investigation-triage-rules, scoped to Incident Triage) adds explicit “never link” rules that prevent the agent from correlating failure investigations with upgrade investigations, even when they involve the same cluster.
  • Scoped RCA skills activate based on incident context: eks-upgrade-planning triggers for Health events, eks-failure-root-cause triggers for CloudFormation rollbacks. The agent selects the correct skill automatically.

c. When the root-cause investigation completes, it emits the Investigation Completed event to Amazon EventBridge. The same Trigger Lambda function that handles upgrade completions picks up this event (filtered by agent_space_id).

d. The Trigger Lambda function (devops-agent-trigger-upgrade) fetches the investigation’s journal records through ListJournalRecords and scans for content markers. If a Root Cause heading is present in the content markers but no Mitigation Plan heading exists, the Lambda function knows the root-cause phase is complete but mitigation hasn’t run yet. It programmatically activates the Mitigation Agent by calling UpdateBacklogTask with status PENDING_START, instructing AWS DevOps Agent to generate a recovery plan based on the root-cause findings. It then schedules a one-time check by using Amazon EventBridge Scheduler, set for five minutes later, to poll for mitigation completion. The Mitigation Agent does not reliably emit a second completion event. If mitigation is still running when the check fires, the Lambda function reschedules at three-minute intervals. If the execution has finished but its journal records are not yet fully written, it retries at one-minute intervals until they appear. Polling is capped at thirty attempts so a stuck mitigation cannot loop indefinitely. If the mitigation execution ends in a terminal failure status (FAILED, CANCELED, or TIMED_OUT), the Lambda function publishes an Amazon Simple Notification Service (Amazon SNS) alert and stops polling rather than retrying indefinitely. Because a native Investigation Completed event and a scheduled poll can both reach the Trigger Lambda function for the same task, dispatches are guarded by a lock built on deterministic Amazon EventBridge Scheduler schedule names, so the same recovery is not dispatched twice.

e. The Mitigation Agent produces up to two outputs depending on what the failure requires: an execution plan with immediate recovery steps if manual intervention is needed, and an agent-ready specification with CDK code changes if an infrastructure fix can prevent recurrence. Either output may be omitted if the mitigation does not call for it.

f. When the scheduled poll detects the mitigation output, the Trigger Lambda function delivers both results:

  1. Operator notification: The SRE team receives an SNS notification with the immediate recovery steps so they can recover the cluster without waiting for a code review.
  2. Code fix pull request: If the mitigation includes a CDK change spec, a GitHub Actions workflow runs Kiro CLI to implement the agent-ready specification and opens a pull request for human review. When the root cause lies outside the CDK stack, such as an application-level API deprecation or a custom admission webhook, the pipeline delivers the execution plan with manual remediation steps only and does not generate a PR.

The responder acts on the urgent manual steps immediately while the automated code fix goes through the normal review process.

Why a closed loop matters

Even with thorough investigation and validation, real-world upgrades can fail because of conditions the agent couldn’t observe pre-deployment: workload-specific API deprecations, custom admission webhooks that reject updated resources, or transient control plane issues during the upgrade window. A pipeline that only handles the happy path leaves the team scrambling manually when things go wrong. The closed loop is designed to apply the same agent-driven rigor to failure recovery.

Keeping skills current: Daily skill review

AWS services evolve continuously, new EKS versions ship, addon defaults change, and API deprecation timelines shift. A skill written today may contain outdated version constraints or miss a new upgrade path within weeks. The pipeline includes an automated daily review that keeps the agent’s skills current without manual monitoring.

An Amazon EventBridge rule triggers a Skill Review Lambda function daily. The Lambda function fetches all four skill files (eks-upgrade-planning, eks-failure-root-cause, eks-investigation-triage-rules, and eks-skill-review itself) from the GitHub repository’s main branch and posts them, embedded in the incident description, to the agent space as a new signed-webhook investigation. The agent runs a dedicated review skill (eks-skill-review) that verifies each claim in the embedded content against authoritative AWS sources. It queries AWS APIs for current EKS version availability, addon defaults, and deprecation schedules, then compares what it finds against the embedded skill content.

When the review identifies gaps, outdated constraints, or missing upgrade paths, the Trigger Lambda function dispatches a skill-update.yml GitHub Actions workflow. Kiro CLI applies the recommended edits to the skill files and opens a pull request. The team receives an SNS notification on the eks-skill-update-notifications topic, reviews the PR, and after merging, re-uploads the updated skill zips to the agent space. If no changes are needed, the pipeline logs the result and exits silently. A third path guards against silent failure: if the agent’s output carries the spec heading but no parse-able spec can be isolated from it, the Lambda function dispatches the workflow with the full findings so the run fails visibly rather than reporting a false no-change result.

This self-maintenance loop means the pipeline’s knowledge stays aligned with EKS capabilities, including changes like the recently announced version rollback feature, without requiring the team to manually track service announcements and update skills.

Two caveats apply. First, skill-based triage routing relies on model judgment and can vary between runs on identical input. Treat the daily review as a best-effort maintenance loop, not a guaranteed daily gate. Second, while the review inspects its own skill file, edits to the review procedure still require the same human merge-and-re-upload cycle as any other skill change.

Safety constraints: What the pipeline enforces and why

Amazon EKS upgrades carry risks that make automated safety checks essential. The pipeline enforces constraints at every stage, from the agent’s investigation through to the final CDK diff validation.

Only one minor version at a time. EKS does not support skipping Kubernetes versions. For example, you can move from 1.30 to 1.31, but not from 1.30 to 1.32. The agent validates this in Step 2 of its investigation and stops with an error if a version skip is detected. This constraint means that clusters that are multiple versions behind require sequential upgrades, each with its own investigation and validation cycle.

Control plane upgrades are reversible for 7 days. EKS supports Kubernetes version rollbacks, so you can revert a control plane upgrade to the previous minor version within seven days. EKS evaluates rollback readiness through cluster insights under the ROLLBACK_READINESS category, checking API usage compatibility, cluster health, kubelet and kube-proxy version skew, and EKS-managed add-on compatibility. Insights with ERROR or UNKNOWN status block the rollback until resolved, so rollback can be unavailable even within the 7-day window if readiness checks fail. After the window closes, rollback is no longer offered regardless of cluster state. Rolling back from a version under standard support into one under extended support resumes extended support charges. The upgrade-planning skill checks rollback readiness during its investigation and documents the window in the PR body, so reviewers know their safety net and its constraints.

Rollback is not always viable. Even within the 7-day window, rollback may be unavailable or inappropriate when:

  • Resources were created during the 7-day window using APIs or fields that exist only in the newer version, which must be removed before rolling back.
  • Add-on versions are not rolled back automatically, and a downgrade can fail if the current configuration settings are incompatible with the target add-on version. Rollback readiness insights evaluate only EKS managed add-ons.
  • Nodes were already upgraded and now have version skew. Managed node groups must be rolled back before the control plane, the inverse of the upgrade sequence.
  • Workloads have adopted features available only in the newer Kubernetes version.
  • The cluster uses AWS Fargate worker nodes. Fargate pods running the current version must be deleted before rollback, or the kubelet version skew check bypassed with --force.
  • The cluster was automatically upgraded at the end of extended support (rollback unavailable), or at the end of standard support (rollback requires changing the cluster’s upgrade policy to EXTENDED first)
  • The cluster was created at its current Kubernetes version rather than upgraded into it, so there is no prior version to return to.
  • Rollback supports only N to N-1. You cannot roll back across multiple minor versions.

The agent’s risk assessment flags the conditions the pipeline actually encodes (deprecated API usage, add-on version incompatibility, and node version skew) and records them in the PR body alongside its ROLLBACK_AVAILABLE verdict. The remaining conditions above are documented AWS behavior that reviewers should confirm manually. The pipeline does not check them. Note too that the --force flag bypasses insight checks only. It does not bypass the prerequisite validations (the 7-day window, the created-at-version check, or the single-minor-version rule) and it cannot override an incompatible Amazon EKS feature enabled at the current version.

vpc-cni must be updated before node groups. New Amazon Machine Images expect the updated CNI plugin, so the Amazon Virtual Private Cloud (Amazon VPC) CNI add-on upgrade must precede any node group update. If the add-on has not been updated first, pods on the new nodes lose networking. The CDK stack declares this ordering explicitly: the managed node group carries a CloudFormation DependsOn the Amazon VPC CNI add-on, so an update cannot reach the node group before the add-on has been updated. The sequence is also declared non-negotiable in the upgrade-planning skill and the Global Instructions, and the agent reproduces the required order in its investigation output and the PR body. The remaining add-on order (kube-proxy, then Coredns) is documented operational sequence rather than a synthesized dependency.

A Replace means cluster destruction. A Replace action deletes the resource and recreates it. For an Amazon EKS cluster, that means the control plane, all workloads, and all state are destroyed and rebuilt from scratch, which makes the cdk diff the single most important thing a reviewer looks at. The pipeline reduces the chance of a destructive change reaching that review through layered gates rather than a single check:

  • Version values are taken verbatim from the validated spec file rather than derived by the model.
  • Kiro CLI is restricted to file tools only (read, write, glob, grep) and cannot run shell commands.
  • A file-change allowlist fails the run if anything other than lib/iteration3-stack.ts was modified.
  • A separate step updates the kubectl layer dependency, and a final validation step runs the build and CDK synthesis so that only changes that compile and synthesize successfully can reach a pull request.

The PR body’s reviewer checklist then requires a cdk diff showing Modify and not Replace, alongside version-correctness and add-on compatibility checks. That is a human gate, not an automated one, and it is the final defense before the separately triggered deploy workflow runs after merge.

These constraints are enforced at multiple points: during the agent’s investigation, during Kiro’s code modification and validation, and again at the human review gate on the pull request. Redundant checks at the earlier stages reduce the risk of a single point of failure allowing a destructive change through.

With the safety model clear, here’s what you need before deploying.

Getting started

Follow these steps to deploy the whole solution into your own account, from the Amazon EKS cluster through to the agent space, skills, and event routing.

Important: This solution deploys billable AWS resources including an Amazon EKS cluster, AWS Lambda functions, Amazon EventBridge rules, AWS Identity and Access Management (IAM) roles, and AWS Secrets Manager secrets. You will incur charges while these resources are running. We recommend deploying in a development account and following the Clean up section after completing the walkthrough to avoid ongoing charges.

Prerequisites

To deploy this pipeline in your own environment, you need the following:

AWS account and tooling

  • An AWS account in a region where AWS DevOps Agent is available, with AWS CDK bootstrapped and AWS Command Line Interface (AWS CLI) v2 configured.
  • Permissions to create Amazon EKS clusters, AWS Identity and Access Management (IAM) roles, Lambda functions, Amazon EventBridge rules, and Secrets Manager secrets. The walkthrough uses administrative credentials for brevity. Scope them down for anything beyond a sandbox account.
  • Node.js 20.x or later and npm.

GitHub

  • A GitHub repository (fork or clone https://github.com/aws-samples/sample-automate-planned-lifecycle-upgrades-with-aws-devops-agent-and-kiro).
  • A GitHub fine-grained Personal Access Token (PAT) granting Read and write on Actions, Contents, and Pull requests for your fork, which you will store on AWS Secrets Manager.
  • A KIRO_API_KEY repository secret holding your Kiro CLI API key.
  • For the optional post-merge deploy workflow only: an IAM role that trusts GitHub’s OpenID Connect (OIDC) provider, with its ARN stored as the AWS_DEPLOY_ROLE_ARN repository secret. The sample does not create this role, and the upgrade pipeline through pull request creation works without it.

Kiro

  • A Kiro CLI API key, which requires a Kiro Pro, Pro+, or Power subscription.

Step 1: Clone the repository

git clone https://github.com/aws-samples/sample-automate-planned-lifecycle-upgrades-with-aws-devops-agent-and-kiro.git
cd sample-automate-planned-lifecycle-upgrades-with-aws-devops-agent-and-kiro

Step 2: Run the bootstrap script to provision the Amazon EKS cluster, AWS DevOps Agent space, Lambda functions, and Amazon EventBridge rules:

./bootstrap.sh

Step 3: Follow the README to configure the webhook credentials, GitHub PAT, and Kiro API key.

Step 4: Upload the AWS DevOps Agent skills and configure agent instructions

Operations teams use AWS DevOps Agent Space web apps for daily incident response activities. This standalone application provides an interface where SREs can launch investigations, interact with the agent through natural language chat, view application topologies, and review incident prevention recommendations.

  1. Access the AWS DevOps Agent space web app
    1. In the AWS DevOps Agent console, select your agent space (eks-upgrade-poc).
    2. Select Launch web app from the top right, choosing IAM or AWS IAM Identity Center option based on your setup. This opens the dedicated web app that the operations teams use to conduct investigations and review recommendations within that space.

The single agent space uses Global Instructions, agent-type-scoped instructions, and four skills to route investigations correctly and enforce isolation between upgrade and failure paths.

  1. Configure Global Instructions
    1. In the AWS DevOps Agent web app navigate to Knowledge > Instructions > All agents
    2. Paste the contents of instructions/global-instructions.md from the repository and select Save.

The Instructions page groups global instructions with the agent-type-scoped instructions, as the following screenshot shows.

Fig 3: AWS DevOps Agent web app showing the Knowledge Base section with Instructions tab open, displaying Global Instructions and agent-type-scoped instructions configuration

Figure 3: The Instructions page showing Global Instructions and agent-type-scoped instructions

  1. Configure Incident Mitigation instructions
    1. In the same agent space, navigate to Knowledge > Instructions > Incident Mitigation
    2. Paste the contents of instructions/mitigation-agent-instructions.md from the repository and select Save.
  1. Upload the agent skills
    1. Zip the skill folder from the repository:
cd skills
zip -r eks-upgrade-planning.zip eks-upgrade-planning
zip -r eks-failure-root-cause.zip eks-failure-root-cause
zip -r eks-investigation-triage-rules.zip eks-investigation-triage-rules
zip -r eks-skill-review.zip eks-skill-review
    1. In the AWS DevOps Agent web app, navigate to Settings > Skills > Custom Skills and select Add Skill.

The Skills page separates the custom skills you upload from AWS managed skills, as the following screenshot shows.

Fig 4: AWS DevOps Agent web app showing the Skills Management page with Custom Skills and Managed Skills tabs

Figure 4: The Skills Management page with the Custom Skills and Managed Skills tabs

    1. Select Upload Skill from the pop-up.
    2. For each skill, upload the zip file.
    3. Under agent type scope, select the agent type listed in the following table and choose Upload.

Note: Each skill must be scoped to the correct agent type so the agent activates it in the right context.

Skill Scope Purpose
eks-upgrade-planning Incident RCA 7-step EKS upgrade investigation producing a CDK Change Spec
eks-failure-root-cause Incident RCA Root-cause analysis for CloudFormation rollback failures
eks-investigation-triage-rules Incident Triage Prevents linking between upgrade and failure investigations
eks-skill-review Incident RCA Daily review of skills for gaps and outdated information

The Upload Skill dialog takes the zip file and the agent type scope together, as the following screenshot shows.

Fig 5: Upload Skill dialog on AWS DevOps Agent, showing fields for uploading a skill zip file and selecting the agent type scope

Figure 5: The Upload Skill dialog for choosing a skill zip file and agent type scope

Step 5: Subscribe to SNS topics

Subscribe your on-call email to both SNS topics the stack creates: eks-upgrade-failure-mitigation (mitigation plans and pipeline failure alerts) and eks-skill-update-notifications (daily skill review findings).

Step 6: Test the pipeline end-to-end

The README includes a step-by-step walkthrough, end-to-end test instructions, and optional configuration for the failure mitigation SNS notifications.

Clean up

To avoid ongoing charges, delete the resources deployed during this walkthrough. The repository includes a cleanup script that removes everything in reverse order.

Run the cleanup script:

./cleanup.sh

The script deletes the CloudFormation stack (agent space, Lambda functions, Amazon EventBridge rules, Secrets Manager secrets) and the CDK stack (EKS cluster, node group, VPC). See the repository README for pre-cleanup steps and details on resources that require manual removal.

Security best practices

Security and compliance is a shared responsibility between AWS and the customer, as outlined in the Shared Responsibility Model. We encourage you to review this model for a comprehensive understanding of the respective responsibilities.

In this solution, we implemented the following security measures:

  • Secrets management. Webhook HMAC credentials and the GitHub PAT are stored on AWS Secrets Manager and are not hard-coded or passed as environment variables. Lambda functions retrieve secrets at invocation time using least-privilege IAM policies scoped to only the specific secret ARNs they require.
  • Least-privilege IAM. Each Lambda function operates with a dedicated IAM role granting only the minimal permissions required for its specific function. The Health Lambda function can only read webhook credentials and invoke the AWS DevOps Agent webhook. The Trigger Lambda function can only read journal records, update backlog tasks, create and delete the Amazon EventBridge Scheduler schedules it uses for mitigation polling, dispatch GitHub workflows, and publish to the two designated SNS topics (eks-upgrade-failure-mitigation for operator notifications and eks-skill-update-notifications for daily skill review alerts).
  • Webhook authentication. Communications between Lambda functions and the AWS DevOps Agent webhook use HMAC-SHA256 signed payloads. The agent validates the signature on every request, rejecting payloads with an invalid or missing signature.
  • GitHub token scoping. The GitHub Personal Access Token uses fine-grained permissions scoped to a single repository with only the Actions, Contents, and Pull Requests permissions required for workflow dispatch and PR creation.
  • No long-lived credentials in CI/CD. The post-merge deploy workflow (eks-deploy.yml) uses GitHub Actions OIDC federation to assume a short-lived IAM role, removing long-lived access keys from the GitHub environment.
  • Encryption. All data at rest in Amazon Simple Storage Service (Amazon S3) (CloudFormation template uploads, CDK assets) is encrypted using server-side encryption. Secrets Manager secrets are encrypted with a customer-managed AWS Key Management Service (AWS KMS) key created by the template. All API communications use TLS encryption in transit.
  • Constrained agent tooling. Kiro CLI runs with file tools only (read, write, glob, grep), with no shell or command execution, so the scope of the agent step is limited to file edits in the checked-out working tree. After Kiro exits, a separate workflow step diffs the working tree against a single-file allowlist (lib/iteration3-stack.ts) and fails the run if any other file was modified. The mitigation path’s workflow uses a wider three-file allowlist (adding package.json and package-lock.json), since a code fix can legitimately require other dependency changes. The agent cannot execute commands, alter workflow definitions, or touch IAM policies or the CloudFormation template.
  • Pinned, verified CI tooling. Kiro CLI is pinned to a minimum tested version. The workflow fails on anything older and warns on anything newer, so an untested release cannot be silently adopted. The installer is downloaded and executed as two discrete steps rather than piped directly from curl to a shell.

We recommend applying these additional security practices:

  • Enable AWS CloudTrail logging for the devops-agent API calls to maintain an audit trail of agent interactions.
  • Restrict the Amazon EventBridge rules to accept events only from expected sources and account IDs.
  • Rotate the GitHub PAT and webhook HMAC secret on a regular cadence.
  • Review the OWASP Top 10 for LLMs for guidance on securing AI-driven pipelines.

Looking ahead: Additional AWS DevOps Agent capabilities

Two recently released AWS DevOps Agent capabilities could further strengthen this pipeline, though they are not included in our solution:

Release management: AWS DevOps Agent can automatically review code changes for standards adherence, cross-repository dependency risks, and access-control correctness before deployment. In the context of this pipeline, Release management could evaluate the Kiro-generated CDK pull request against your organization’s policies and flag cross-service breaking changes that CDK diff alone would miss. It can also generate and execute change-specific tests against a running environment, catching integration failures before merge. For more information, see Release management.

Improvements (proactive incident prevention): AWS DevOps Agent analyzes patterns across your incident investigations and delivers prioritized recommendations to help prevent recurring failures. For the EKS upgrade pipeline, this means the agent can identify systemic patterns across multiple failed upgrades, such as a recurring addon incompatibility or a misconfigured node group setting, and generate agent-ready specifications to address the root cause proactively. Recommendations are categorized across observability, infrastructure, governance, and code optimization, and can be handed directly to a coding agent for implementation. Access this capability through the Improvements page in the AWS DevOps Agent web app. For more information, see Proactive incident prevention.

Conclusion

This pipeline shifts end-of-support upgrades from a reactive, manual process to a proactive, event-driven workflow. The investigation, code changes, and validation that an engineer previously performed per cluster now arrive as a reviewed pull request, with no human intervention until the approval step. When AWS Health detects an approaching end-of-support milestone, the system investigates, codes, validates, and delivers a pull request. This reduces mean time to remediation from days to minutes and frees engineers to focus on architecture decisions rather than repetitive upgrade mechanics.

The pipeline’s separation of investigation from delivery means that onboarding a new AWS managed service, such as Amazon RDS engine versions, Amazon ElastiCache engine upgrades, or Lambda runtime deprecations, requires only a new investigation skill. The event routing, code modification, validation, and PR infrastructure remains unchanged.

To get started, clone the repository and run bootstrap.sh, which deploys the CDK stack first (VPC, EKS cluster, managed addons, and the AWS Load Balancer Controller) and then the devops-agent-space.yaml CloudFormation template that creates the agent space, IAM roles, Amazon EventBridge rules, Lambda functions, and Secrets Manager secrets. Configure your webhook credentials and GitHub PAT on AWS Secrets Manager, point the GitHub Actions workflow at your CDK repository, and the pipeline is live. The next Planned Lifecycle Event that fires for your Amazon EKS clusters will produce a validated, reviewable pull request with no human intervention required until the review step.

Next steps

Whether you are exploring, prototyping, or ready to deploy, here is where to go next:

Just evaluating? Read the event workflow walkthrough, which traces every event, Lambda function invocation, and decision point traced end to end, with nothing to deploy. Pair it with the upgrade-planning skill to see the investigation logic that produces the CDK Change Spec.

Ready to run it? Clone the repository and follow the deployment guide in a development account. Roughly 25 minutes for bootstrap.sh, plus 10–15 minutes of configuration, and the synthetic health event in the README produces your first agent-generated pull request. Run cleanup.sh when you are finished to stop the charges.

Ready to adapt it? The investigation logic lives entirely in skills/eks-upgrade-planning/SKILL.md. The routing, validation, and PR machinery is service-agnostic. Onboarding another service that publishes lifecycle events means a new skill and a matching Amazon EventBridge pattern, not a new pipeline. Start with that skill’s output contract, since it is what the validation gate enforces.

To go deeper on the solution, see the AWS DevOps Agent documentation for how investigations, skills, and agent types work, the AWS DevOps Agent Skills reference for the SKILL.md format, and the Kiro CLI documentation for headless-mode options.


About the authors

Nehal Sangoi

Nehal Sangoi

Nehal is a Senior Technical Account Manager at Amazon Web Services (AWS). She provides strategic technical guidance to Independent Software Vendors in the security space, helping them architect resilient, scalable solutions using AWS best practices. Nehal specializes in Generative AI workloads, partnering with ISV customers to accelerate innovation and deliver secure, cloud-native outcomes. Connect with Nehal on LinkedIn.

Tipu Qureshi

Tipu Qureshi

Tipu is a Senior Principal Technologist in AWS Agentic AI, focusing on operational excellence and incident response automation. He works with AWS customers to design resilient, observable cloud applications and autonomous operational systems.

Ben Peterson

Ben Peterson

Ben is a Senior Solutions Architect with AWS. He is passionate about enhancing the developer experience and driving customer success. In his role, he provides strategic guidance on using the comprehensive AWS suite of services to modernize legacy systems, optimize performance, and unlock new capabilities. Connect with Ben on LinkedIn.

Akshay Singhal

Akshay Singhal

Akshay is a Principal Technical Account Manager at Amazon Web Services supporting Enterprise Support customers focusing on the Security ISV segment. He provides technical guidance for customers to implement AWS solutions, with expertise spanning serverless architectures and GenAI workloads. Connect with Akshay on LinkedIn.

Building medallion architecture with Iceberg materialized views in Amazon SageMaker

Post Syndicated from Gaurav Sharma original https://aws.amazon.com/blogs/big-data/building-medallion-architecture-with-iceberg-materialized-views-in-amazon-sagemaker/

Building a Medallion Architecture today typically means that you must build three separate systems working in concert: extract, transform, and load (ETL) jobs to transform data between layers, an orchestrator (such as Apache Airflow or AWS Step Functions) to sequence those jobs in the correct order, and custom change-data-capture (CDC) logic to make sure that each job processes only new or modified records. Each component must be authored, tested, deployed, and maintained independently and when one breaks, the entire pipeline stalls.

In this post, we show how Apache Iceberg materialized views in Amazon SageMaker collapse transformation, orchestration, and incremental processing into a single SQL definition per layer. You declare what each layer should contain, and the system handles when and how it refreshes based on your refresh configuration. With this approach, you can build a Bronze → Silver → Gold pipeline with three SQL statements. This reduces the complexity of maintaining separate orchestration code, CDC logic, and job artifacts.

What is medallion architecture

The medallion architecture organizes data into three progressive layers:

  • Bronze layer – Captures raw data as-is from source systems, preserving the original format for auditability and replay.
  • Silver layer – Applies cleaning, deduplication, type casting, and business logic to produce validated, query-ready datasets.
  • Gold layer – Aggregates Silver data into business-level metrics, key performance indicators (KPIs), and dimensional models optimized for analytics and reporting.

Each layer builds on the previous one, creating clear lineage from raw ingestion to business insight.

Traditional versus declarative approach

The two approaches differ in how much infrastructure you build and maintain.

Traditional approach

You write an ETL job such as Apache Spark script for Bronze to Silver layer and another for Silver to Gold layer. You build a directed acyclic graph (DAG) in Apache Airflow or a Step Functions state machine to run them in order. You implement CDC logic like tracking high watermarks, comparing snapshots, or consuming change streams such that each job processes only new data.

Declarative approach with Iceberg materialized views

You write one CREATE MATERIALIZED VIEW statement per layer with a SCHEDULE REFRESH EVERY N HOURS clause. The AWS Glue managed Spark compute executes the refresh, but you don’t author, version, or deploy a job artifact. Iceberg’s row-level change tracking (position-delete and equality-delete files) identifies which rows changed since the last refresh and AWS Glue processes only those rows. The dependency chain is implicit in the SQL definitions. The only code you maintain is the SQL transformation logic itself.

Apache Iceberg and materialized views

Apache Iceberg is an open-source, high-performance table format designed for petabyte-scale analytic datasets in data lakes. It provides ACID transactions, time travel, schema evolution, and hidden partitioning.

With an Iceberg materialized view, you can define each layer of a medallion architecture as a SQL statement. Under the hood, AWS Glue uses Iceberg’s change-tracking metadata to identify which rows changed since the last refresh, then processes only those rows using managed Spark compute. You configure scheduling and incremental processing through SQL definitions, and the system executes atomic refreshes without requiring you to write pipeline code.

When refreshed, the Gold materialized view reads incrementally from the Silver materialized view, which in turn reads from the Bronze table. This creates a declarative dependency chain: each layer’s definition points to the layer below it, and the system resolves which data to reprocess at each refresh.

Service support for Iceberg materialized views

At time of publication, the following services support creating and refreshing Iceberg materialized views:

For the latest version requirements, see the AWS Glue materialized views documentation.

Technical architecture

The architecture uses Amazon S3 Tables, a capability of Amazon Simple Storage Service (Amazon S3), as the storage layer. Amazon S3 Tables is a managed Apache Iceberg offering that alleviates the administrative overhead of maintaining Iceberg tables. AWS Glue Data Catalog manages table metadata, and Amazon SageMaker Unified Studio provides the AI-powered notebook environment with AWS Glue 5.1 for authoring and executing materialized view definitions.

The diagram illustrates a three-tier data lakehouse pipeline built on Apache Iceberg. The Bronze layer contains raw trip data (trips_bronze table on S3 Tables with fields: trip_id, city, vehicle_type, fare, status) that you ingest through INSERT/Append operations.

An incremental REFRESH feeds the Silver layer, where a materialized view (mv_trips_silver) performs timestamp conversion, null filtering, and computes derived columns like revenue_per_mile and rating_category. It processes only new or changed rows.

The Silver layer then refreshes two Gold layer materialized views on a daily schedule: mv_city_daily_metrics (city, date, trips, drivers, revenue, tips) and mv_vehicle_performance (vehicle_type, city, trips, revenue, distance). The Gold layer serves downstream consumers including Amazon Athena, Amazon Quick Sight, Amazon Redshift, and first-party (1P) or third-party (3P) compute engines supporting the Iceberg REST API.

The pipeline flows as follows:

Diagram of the medallion pipeline: a Bronze table feeds a Silver materialized view that feeds two Gold materialized views consumed by analytics engines

Figure 1: The three-tier medallion pipeline from the Bronze table through Silver and Gold materialized views to analytics consumers

Prerequisites

Before starting, verify that you have the following:

  • An AWS account with permissions for Amazon SageMaker Unified Studio, AWS Glue, S3 Tables, and AWS Lake Formation.
  • An Amazon SageMaker Unified Studio domain.

Step 1: Initialize the environment

Open the AWS Management Console and navigate to Amazon SageMaker.

Amazon SageMaker console landing page

Figure 2: The Amazon SageMaker console landing page

Choose Get Started to set up Amazon SageMaker Unified Studio.

SageMaker Unified Studio Get Started setup page

Figure 3: The Get Started page for setting up SageMaker Unified Studio

Choose Open to launch Amazon SageMaker Unified Studio.

Button to open and launch SageMaker Unified Studio

Figure 4: The option to open and launch SageMaker Unified Studio

After you’re in SageMaker Unified Studio, choose Data in the left pane to create the S3 Tables bucket (a managed Apache Iceberg feature of Amazon S3) and a database. Choose Add, then choose Create S3 Tables Catalog, and provide a catalog and a database name. Finally, choose Create Catalog.

Create S3 Tables Catalog dialog with catalog and database name fields

Figure 5: The Create S3 Tables Catalog dialog with catalog and database name fields

After the catalog creation is complete, in the left navigation pane, choose Notebooks.

Notebooks option in the SageMaker Unified Studio left navigation pane

Figure 6: The Notebooks option in the SageMaker Unified Studio navigation pane

Choose Create Notebook.

Create Notebook button in SageMaker Unified Studio

Figure 7: The Create Notebook button in SageMaker Unified Studio

Before using the notebook, select either Athena Spark or Glue Spark compute connection as the runtime engine for your notebook.

Runtime engine selection showing Athena Spark and Glue Spark compute connections

Figure 8: Selecting Athena Spark or Glue Spark as the notebook runtime engine

Use the following code samples in individual notebook cells. You can also provide transformation requirements in natural language, and the SageMaker Data Agent will generate SQL code for you.

SageMaker Data Agent generating SQL from a natural language prompt

Figure 9: The SageMaker Data Agent generating SQL from a natural language request

Add each code block in a new cell by choosing the SQL button:

SQL cell-type button in the notebook toolbar

Figure 10: The SQL button for adding a code block to a notebook cell

Choose Athena Spark or Glue Spark as your compute from the cell menu.

Compute connection selection in the notebook cell menu

Figure 11: The compute selection in the notebook cell menu

If you encounter errors after cell execution, use the data agent chatbot or the Fix with AI button to resolve them.

Fix with AI button and data agent chatbot for resolving cell errors

Figure 12: The Fix with AI button for resolving cell execution errors

Step 2: Ingest data into Bronze

Generate 300 realistic ride-sharing trips and insert them directly into the Bronze Iceberg table. This simulates a raw data ingestion layer. In production, you generally configure a streaming source or batch load based on your requirements.

Copy the following code into the first notebook cell (use a Python cell type).

import random
from datetime import datetime, timedelta

CITIES = {
    "San Francisco": {"lat_range": (37.70, 37.82), "lon_range": (-122.52, -122.38), "surge_prob": 0.3},
    "Austin": {"lat_range": (30.22, 30.40), "lon_range": (-97.80, -97.68), "surge_prob": 0.15},
    "Chicago": {"lat_range": (41.85, 41.95), "lon_range": (-87.70, -87.60), "surge_prob": 0.2},
    "Seattle": {"lat_range": (47.55, 47.68), "lon_range": (-122.40, -122.28), "surge_prob": 0.25},
}
VEHICLE_TYPES = ["UberX", "Comfort", "XL", "Black"]
PAYMENT_METHODS = ["credit_card", "debit_card", "apple_pay", "google_pay", "cash"]
STATUSES = ["completed"] * 4 + ["cancelled_rider", "cancelled_driver"]
BASE_FARES = {"UberX": 2.50, "Comfort": 3.50, "XL": 4.00, "Black": 7.00}
PER_MILE = {"UberX": 1.75, "Comfort": 2.25, "XL": 2.50, "Black": 3.75}
PER_MIN = {"UberX": 0.35, "Comfort": 0.45, "XL": 0.50, "Black": 0.65}

rows = []
for i in range(300):
    city_name = random.choice(list(CITIES.keys()))
    city = CITIES[city_name]
    vehicle = random.choice(VEHICLE_TYPES)
    duration = random.randint(5, 45)
    distance = round(random.uniform(1.0, 20.0), 1)
    surge = round(random.uniform(1.0, 2.5), 1) if random.random() < city["surge_prob"] else 1.0
    base = BASE_FARES[vehicle]
    fare = round((base + distance * PER_MILE[vehicle] + duration * PER_MIN[vehicle]) * surge, 2)
    tip = round(fare * random.choice([0, 0, 0.1, 0.15, 0.2, 0.25]), 2)
    status = random.choice(STATUSES)
    day = random.randint(0, 2)
    hour = random.choices(range(24),
        weights=[1,1,1,1,1,2,4,8,10,8,6,5,6,5,5,5,6,8,10,8,6,4,2,1])[0]
    trip_time = datetime(2025, 12, 1) + timedelta(days=day, hours=hour, minutes=random.randint(0, 59))

    rows.append((
        f"TRIP-{i+1:06d}",
        f"DRV-{random.randint(1000, 5000)}",
        f"RDR-{random.randint(10000, 99999)}",
        city_name, vehicle,
        round(random.uniform(*city["lat_range"]), 6),
        round(random.uniform(*city["lon_range"]), 6),
        round(random.uniform(*city["lat_range"]), 6),
        round(random.uniform(*city["lon_range"]), 6),
        trip_time.isoformat(),
        (trip_time + timedelta(minutes=duration)).isoformat(),
        duration, distance, surge, base, fare, tip, round(fare + tip, 2),
        random.choice(PAYMENT_METHODS),
        random.choice([None, 3, 4, 4, 5, 5, 5]) if status == "completed" else None,
        status,
    ))

schema = ("trip_id STRING, driver_id STRING, rider_id STRING, city STRING, "
    "vehicle_type STRING, pickup_lat DOUBLE, pickup_lon DOUBLE, "
    "dropoff_lat DOUBLE, dropoff_lon DOUBLE, trip_start_time STRING, "
    "trip_end_time STRING, duration_minutes INT, distance_miles DOUBLE, "
    "surge_multiplier DOUBLE, base_fare DOUBLE, trip_fare DOUBLE, "
    "tip_amount DOUBLE, total_amount DOUBLE, payment_method STRING, "
    "rating INT, status STRING")

df = spark.createDataFrame(rows, schema)
df.writeTo("{CATALOG_NAME}.{NAMESPACE_NAME}.trips_bronze").createOrReplace()

print(f"Created Table and Inserted {len(rows)} trips into Bronze layer")

Step 3: Explore Bronze

Run a preview on the bronze table. The output should look like the following screenshot:

Preview of raw Bronze table trip records with string timestamps and nullable fields

Figure 13: A preview of raw trip records in the Bronze table

You should see raw, unprocessed trip records with string timestamps and nullable fields. This is exactly what the Silver layer will clean up.

Now, verify the ingested data by querying the Bronze table for basic statistics.

SELECT COUNT(*) as total_trips, COUNT(DISTINCT city) as cities,
COUNT(DISTINCT vehicle_type) as vehicle_types,
MIN(trip_start_time) as earliest, MAX(trip_start_time) as latest
FROM ({CATALOG_NAME}.{NAMESPACE_NAME}.trips_bronze

The output should look like the following screenshot:

Query results showing total trips, distinct cities, and vehicle types in the Bronze table

Figure 14: Bronze table statistics showing total trips, distinct cities, and vehicle types

Step 4: Create the Silver materialized view

This SQL statement defines the Silver layer as a materialized view that cleans, transforms, and derives new columns from the Bronze table. Note that this is only a definition. The system processes the data at refresh time.

CREATE MATERIALIZED VIEW IF NOT EXISTS {CATALOG_NAME}.{DATABASE}.mv_trips_silver
COMMENT 'Silver layer: Cleaned trip data with proper types and derived columns'
SCHEDULE REFRESH EVERY 1 DAY
AS
SELECT
trip_id, driver_id, rider_id, city, vehicle_type,
pickup_lat, pickup_lon, dropoff_lat, dropoff_lon,
CAST(trip_start_time AS TIMESTAMP) as trip_start_timestamp,
CAST(trip_end_time AS TIMESTAMP) as trip_end_timestamp,
duration_minutes, distance_miles, surge_multiplier,
base_fare, trip_fare, tip_amount, total_amount,
payment_method, rating, status,
CASE WHEN distance_miles > 0 THEN total_amount / distance_miles ELSE 0 END as revenue_per_mile,
CASE WHEN rating >= 4 THEN 'High' WHEN rating >= 3 THEN 'Medium' ELSE 'Low' END as rating_category
FROM {CATALOG_NAME}.{DATABASE}.trips_bronze
WHERE trip_id IS NOT NULL AND driver_id IS NOT NULL AND rider_id IS NOT NULL
AND total_amount >= 0 AND distance_miles >= 0

print("Silver MV created: urbanride.mv_trips_silver")

Verify the Silver layer output:

SELECT trip_id, city, vehicle_type, total_amount,
ROUND(revenue_per_mile, 2) as rev_per_mile, rating_category
FROM {CATALOG_NAME}.{DATABASE}.mv_trips_silver LIMIT 5

Notice how the Silver layer now has proper timestamps, derived revenue_per_mile, and rating categories: clean, typed, and ready for you to aggregate.

The output should look like the following screenshot:

Silver materialized view results with typed timestamps, revenue_per_mile, and rating_category columns

Figure 15: Silver materialized view results with typed timestamps and derived columns

Step 5: Create Gold materialized views

Gold materialized views read incrementally from the Silver materialized view. This is a nested materialized view pattern: a materialized view built on top of another materialized view.

Gold 1: City daily metrics

With this materialized view, you can aggregate trip data by city and date with a scheduled daily refresh.

CREATE MATERIALIZED VIEW IF NOT EXISTS {CATALOG_NAME}.urbanride.mv_city_daily_metrics
COMMENT 'Gold layer: Daily aggregated metrics by city'
SCHEDULE REFRESH EVERY 1 DAY
AS
SELECT
city, DATE(trip_start_timestamp) as trip_date,
COUNT(*) as total_trips,
COUNT(DISTINCT driver_id) as active_drivers,
COUNT(DISTINCT rider_id) as active_riders,
SUM(total_amount) as total_revenue,
SUM(distance_miles) as total_distance,
SUM(tip_amount) as total_tips
FROM {CATALOG_NAME}.{DATABASE}.mv_trips_silver
WHERE status = 'completed'
GROUP BY city, DATE(trip_start_timestamp)

print("Gold MV created: mv_city_daily_metrics (reads from Silver MV, refreshes daily)")

Gold 2: Vehicle performance

With this materialized view, you can aggregate performance metrics by vehicle type and city.

CREATE MATERIALIZED VIEW IF NOT EXISTS {CATALOG_NAME}.{DATABASE}.mv_vehicle_performance
COMMENT 'Gold layer: Vehicle type performance metrics'
SCHEDULE REFRESH EVERY 1 DAY
AS
SELECT
vehicle_type, city,
COUNT(*) as trip_count,
SUM(total_amount) as total_revenue,
SUM(distance_miles) as total_distance,
SUM(tip_amount) as total_tips
FROM {CATALOG_NAME}.{DATABASE}.mv_trips_silver
WHERE status = 'completed'
GROUP BY vehicle_type, city

print("Gold MV created: mv_vehicle_performance (reads from Silver MV, refreshes daily)")

Dependency chain

The complete pipeline dependency is:

trips_bronze (table)
└── mv_trips_silver (materialized view)
    ├── mv_city_daily_metrics (MV on MV, daily schedule)
    └── mv_vehicle_performance (MV on MV, daily schedule)

Each layer is defined by a single SQL statement. There are no DAGs to maintain, no job definitions to deploy, and no watermark tracking to implement.

Step 6: Query the Gold layer

Query the Gold materialized views to see aggregated business metrics.

City daily metrics Gold table

SELECT city, trip_date, total_trips, active_drivers,
ROUND(total_revenue, 2) as revenue,
ROUND(total_revenue / total_trips, 2) as avg_per_trip
FROM {CATALOG_NAME}.{DATABASE}.mv_city_daily_metrics
ORDER BY trip_date DESC, revenue DESC LIMIT 15

The output should look like the following screenshot:

City daily metrics results with trips, active drivers, and revenue per city

Figure 16: City daily metrics from the Gold materialized view

Vehicle performance Gold table

SELECT vehicle_type, city, trip_count,
ROUND(total_revenue, 2) as revenue,
ROUND(total_revenue / trip_count, 2) as avg_per_trip
FROM {CATALOG_NAME}.{DATABASE}.mv_vehicle_performance
ORDER BY revenue DESC

The output should look like the following screenshot:

Vehicle performance results with trip counts and revenue by vehicle type and city

Figure 17: Vehicle performance metrics from the Gold materialized view

The Gold layer gives you pre-aggregated, business-ready metrics without writing aggregation jobs.

Step 7: Data propagation demo

This section demonstrates how changes propagate through the layers using INSERT, UPDATE (MERGE), and DELETE operations followed by incremental refresh. In production, the scheduled refresh handles this automatically. We trigger it manually here for demonstration purposes.

INSERT new records

Insert new trip records into the Bronze table.

INSERT INTO {CATALOG_NAME}.{DATABASE}.trips_bronze VALUES
('DEMO_TRIP_001', 'DRIVER_999', 'RIDER_888', 'Seattle', 'UberX',
47.6062, -122.3321, 47.6205, -122.3493,
'2024-12-15 14:30:00', '2024-12-15 14:50:00',
20, 5.2, 1.0, 10.0, 15.0, 3.0, 18.0, 'credit_card', 5, 'completed'),
('DEMO_TRIP_002', 'DRIVER_888', 'RIDER_777', 'Seattle', 'XL',
47.6101, -122.3300, 47.6550, -122.3080,
'2024-12-15 15:00:00', '2024-12-15 15:35:00',
35, 8.5, 1.5, 15.0, 30.0, 5.0, 35.0, 'cash', 4, 'completed'),
('DEMO_TRIP_003', 'DRIVER_777', 'RIDER_666', Portland, 'Comfort',
30.2672, -97.7431, 30.2800, -97.7400,
'2024-12-15 16:00:00', '2024-12-15 16:15:00',
15, 3.0, 1.0, 8.0, 12.0, 2.0, 14.0, 'credit_card', 5, 'completed')

print("Inserted 3 new trips into Bronze")

Refresh Silver (incremental)

Refresh the Silver materialized view. Iceberg materialized view processes only three new records.

REFRESH MATERIALIZED VIEW {CATALOG_NAME}.{DATABASE}.mv_trips_silver"

Verify the new records propagated

SELECT trip_id, city, total_amount, ROUND(revenue_per_mile, 2) as rev_per_mile, rating_category
FROM {CATALOG_NAME}.{DATABASE}.mv_trips_silver
WHERE trip_id LIKE 'DEMO_TRIP_%' ORDER BY trip_id

The output should look like the following screenshot:

Silver materialized view showing three newly inserted demo trips

Figure 18: The Silver materialized view showing the three newly inserted demo trips

Refresh Gold (cascading from the Silver materialized view)

Refresh the Gold materialized view. It reads from the refreshed Silver materialized view and processes only the incremental changes.

REFRESH MATERIALIZED VIEW {CATALOG_NAME}.{DATABASE}.mv_city_daily_metrics

Verify the Gold layer reflects the new trips

SELECT city, trip_date, total_trips, ROUND(total_revenue, 2) as revenue
FROM {CATALOG_NAME}.{DATABASE}.mv_city_daily_metrics
WHERE trip_date = '2024-12-15' ORDER BY city

The output should look like the following screenshot:

City daily metrics reflecting the newly added trips for December 15, 2024

Figure 19: City daily metrics reflecting the new trips for 2024-12-15

UPDATE through MERGE

Use MERGE to update existing records in Bronze, then refresh incrementally.

MERGE INTO {CATALOG_NAME}.{DATABASE}.trips_bronze AS target
USING (SELECT 'DEMO_TRIP_002' as trip_id, 5 as new_rating, 20.0 as new_tip) AS source
ON target.trip_id = source.trip_id
WHEN MATCHED THEN UPDATE SET
target.rating = source.new_rating,
target.tip_amount = source.new_tip,
target.total_amount = target.trip_fare + source.new_tip

Refresh Silver and verify

REFRESH MATERIALIZED VIEW {CATALOG_NAME}.{DATABASE}.mv_trips_silver")

SELECT trip_id, rating, rating_category, tip_amount, total_amount,
ROUND(revenue_per_mile, 2) as rev_per_mile
FROM {CATALOG_NAME}.{DATABASE}.mv_trips_silver WHERE trip_id = 'DEMO_TRIP_002'

print("UPDATE propagated: rating 4->5, tip $5->$20, total $35->$50")

The output should look like the following screenshot:

Silver materialized view showing the updated rating and tip for DEMO_TRIP_002

Figure 20: The Silver materialized view showing the updated rating and tip for the demo trip

Step 8: Cleanup

Drop materialized views, tables, the namespace, and delete the S3 Tables bucket to fully clean up resources.

# Drop MVs (Gold first, then Silver, due to dependency order)
spark.sql(f"DROP MATERIALIZED VIEW IF EXISTS {CATALOG_NAME}.{DATABASE}.mv_city_daily_metrics")
spark.sql(f"DROP MATERIALIZED VIEW IF EXISTS {CATALOG_NAME}.{DATABASE}.mv_vehicle_performance")
spark.sql(f"DROP MATERIALIZED VIEW IF EXISTS {CATALOG_NAME}.{DATABASE}.mv_trips_silver")
print("All materialized views dropped")

# Drop base table
spark.sql(f"DROP TABLE IF EXISTS {CATALOG_NAME}.{DATABASE}.trips_bronze")
print("Base table dropped")

# Drop the namespace
spark.sql(f"DROP NAMESPACE IF EXISTS {CATALOG_NAME}.{DATABASE} ")
print("Namespace dropped")

# Delete the S3 table bucket
import boto3
s3tables_client = boto3.client("s3tables")

# List and delete all remaining tables in the bucket
tables_response = s3tables_client.list_tables(
    tableBucketARN=TABLE_BUCKET_ARN, namespace="{DATABASE}"
)
for table in tables_response.get("tables", []):
    s3tables_client.delete_table(
        tableBucketARN=TABLE_BUCKET_ARN, namespace="{DATABASE}", name=table['name']
    )
    print(f" Deleted table: {table['name']}")

# Delete the namespace and bucket
s3tables_client.delete_namespace(tableBucketARN=TABLE_BUCKET_ARN, namespace="urbanride")
s3tables_client.delete_table_bucket(tableBucketARN=TABLE_BUCKET_ARN)
print(f"S3 table bucket deleted: {TABLE_BUCKET_NAME}")

Limitations and considerations

While materialized views remove most orchestration code, note the following:

  1. No sub-hour freshness. The minimum schedule granularity is one hour (SCHEDULE REFRESH EVERY 1 HOUR).
  2. Cascading refresh isn’t automatic. Refreshing Silver doesn’t trigger Gold in the same operation. Each layer refreshes on its own schedule or must be triggered sequentially.
  3. Deletes require a FULL refresh. An incremental REFRESH that feeds the Silver layer detects inserts and updates through Iceberg metadata but cannot detect row removals. Use REFRESH ... FULL when delete propagation is needed.
  4. SQL subset only. Some window functions, user-defined functions (UDFs), and complex expressions might not be supported in materialized view definitions.
  5. Schema evolution requires recreation. If the source schema changes in a way that affects the materialized view definition, you must drop and recreate it.
  6. AWS-specific extension. Iceberg materialized views are not part of the open-source Apache Iceberg specification. They aren’t portable to non-AWS environments.

Pricing

AWS bills materialized view auto-refresh at USD $0.44 per DPU-hour (4 vCPU, 16 GB memory), billed per second with a 1-minute minimum. When you configure scheduled refresh, the AWS Glue Data Catalog uses managed Spark compute to incrementally update the materialized view. You pay only for the compute time of each refresh run.

There are no separate charges for storing materialized view metadata in the Data Catalog (covered under standard catalog pricing: first million objects at no additional cost, then $1.00 per 100K objects/month). The materialized view data itself is stored as Iceberg files in S3 Tables or Amazon S3, charged at standard Amazon S3 storage rates.

Manual refreshes triggered from Spark (through Amazon Athena, Amazon EMR, or AWS Glue notebooks) are billed under those services’ respective compute pricing rather than the materialized view auto-refresh rate. For the latest pricing details, see the AWS Glue pricing page.

Estimated cost for this tutorial: Running through all steps once with 300 records typically consumes less than 0.5 DPU-hours total (~$0.22 in AWS Glue compute plus negligible Amazon S3 storage).

Summary

In this post, you built a Bronze → Silver → Gold medallion architecture using three SQL statements with nested materialized views and no orchestration code. The full pipeline creation took under 2 minutes, and incremental refreshes processed only changed data with no watermarks, no DAGs, no CDC plumbing.

To get started with your own data, create an Amazon SageMaker Unified Studio project, define your Bronze table, and express your transformation logic as Iceberg materialized views. For more information, see the Apache Iceberg materialized views documentation in the AWS Glue Developer Guide.

References

Using materialized views with AWS Glue

Query AWS Glue Data Catalog materialized views

Using materialized views with Amazon EMR

Working with Amazon S3 Tables and table buckets


About the authors

Gaurav Sharma

Gaurav Sharma

Gaurav is a Specialist Solutions Architect (Analytics) at AWS, supporting US public sector customers on their cloud journey. Outside of work, Gaurav enjoys spending time with his family and staying informed on technology, politics, and history through books, videos, and podcasts.

Matt David

Matt David

Matt is a Product Marketing Manager at AWS, specializing in helping data teams with AI-powered analytics. His areas of interest include self-service analytics, data democratization, and preparing organizations for the age of AI agents. He brings extensive experience from his roles at Atlassian, Hex, and DataCamp.

Build a dynamic streaming data lake with Apache Iceberg and Apache Flink

Post Syndicated from Francisco Morillo original https://aws.amazon.com/blogs/big-data/build-a-dynamic-streaming-data-lake-with-apache-iceberg-and-apache-flink/

Handling upstream schema changes is a common operational challenge in streaming data pipelines that write to a data lake. When a source schema changes, teams often face a difficult choice: restart the pipeline or perform a manual migration. A restart can pause ingestion and delay or lose in-flight data. A manual migration consumes engineering time and introduces the risk of schema inconsistencies while the data lake falls behind the source.

For example, consider an Apache Flink job that ingests order_events and writes to an Iceberg table. On Monday, the pipeline runs normally. By Wednesday, the upstream team adds a new loyalty_tier field and introduces a new interaction_events event type. Traditionally, you would need to stop the Flink job, update your schema definitions, and redeploy. With Apache Iceberg’s Dynamic Iceberg Sink on Amazon Managed Service for Apache Flink, the pipeline can handle both changes at the record level without disruption. The DynamicSink routes each event to the right Iceberg table and evolves table schemas as new columns appear, with no operator intervention.

Managed Service for Apache Flink is a fully managed AWS service that you can use to build and deploy streaming applications without setting up infrastructure and managing resources. Apache Flink’s distributed processing engine with exactly once processing guarantees through checkpointing paired with Apache Iceberg’s two-phase commit provides end-to-end consistency without duplications or data loss.

In this post, we show you how to build a dynamic streaming data lake that adapts to new event types and schema changes without stopping the pipeline. Using Apache Flink 2.3 and Apache Iceberg 1.11.0 on Managed Service for Apache Flink, we walk through the DataStream API patterns for per-record table routing and automatic schema evolution. The complete implementation is available in this GitHub repository.

Apache Iceberg dynamic sink

The Dynamic Iceberg Sink allows Flink to dynamically route records to multiple Iceberg tables based on user-defined logic. It also creates and updates tables on the fly and evolves both table schemas and partition specs during streaming execution, controlled through the DynamicRecord class, which eliminates the need for Flink job restarts when requirements change.

Per-record table routing with DynamicIcebergSink

The DynamicIcebergSink resolves the target table at the record level rather than at pipeline configuration time. Records flow through a DynamicRecordGenerator that, for each input, emits one or more DynamicRecord values. Each DynamicRecord carries its own target table ID, schema, partition spec, and row payload, so the sink knows where to write and how the table should look:

DynamicIcebergSink.forInput(events)
    .generator(generator)
    .catalogLoader(catalogLoader)
    .immediateTableUpdate(true)
    .cacheMaxSize(cacheMaxSize)
    .cacheRefreshMs(cacheRefreshMs)
    .append();

The generator receives each record and emits a DynamicRecord targeting a resolved table that looks as follows:

return new DynamicRecord(
    tableId,
    tableBranch,
    icebergSchema,
    rowData,
    partitionSpec,
    distributionMode,
    1);

The sink creates the table if it does not exist and evolves its schema when a record carries new columns. cacheMaxSize and cacheRefreshMs bound the sink’s per-table metadata cache, so a job that writes to many tables does not reload metadata on every record. immediateTableUpdate(true) controls how those catalog changes are applied, which the following section on automatic schema evolution explains. A single Flink job can ingest and route order_events, interaction_events, user_events, and future event types without additional sink definitions.

However, the sink also needs to know what the table looks like. That is why every DynamicRecord also carries the Iceberg schema so that DynamicIcebergSink can create the table on first sight and evolve it as new fields appear. The schema information can be inferred from the data or read from a schema registry.

Automatic schema evolution

Streaming sources add new fields over time, and DynamicIcebergSink handles them without a restart. Before writing each record, it compares the record’s schema against the target table. If the record has a new field, Iceberg adds it as an optional column and commits the change with the next data file. Existing files stay valid and no table rewrite is needed. When you query older files, the new column returns null.

The immediateTableUpdate setting controls where the catalog change happens. The GitHub sample repository sets immediateTableUpdate=true, so the writer subtask that sees the new schema applies the create or alter inline, before it emits the record. This gives the lowest latency but makes more concurrent calls to the catalog. When set to false, records that require a table change take a detour. Records whose table, schema, and partition spec already match the sink’s cached metadata go straight to the writers. Records that do need a change are routed, keyed by table name, to an update operator, so updates for the same table apply one at a time. Once the update commits and the cache refreshes, subsequent records match again and skip the detour. In steady state, with no schema changes arriving, this path adds no extra shuffle. Either way, the schema comparison and the resulting table change are the same.

Schema changes are non-destructive by default. The sink can add new columns, widen existing types (for example, int to long or float to double), relax a required column to optional, and drop columns. Importantly, DynamicIcebergSink does not support renaming columns at the time of writing.

Source schemas are identified in two ways: inferring the schema from source records (for example, JSON inference) and reading serialized records from a schema registry (for example, AWS Glue Schema Registry (GSR)). Schema evolution behavior for the Iceberg sink table depends on the schema source. JSON inference adds any new field it sees, with no contract. For example, this allows the job to initially infer a schema as an integer, and later expand to a long when larger values are detected. Schema registry serialized records define the policy using the registry’s compatibility rules (for example, BACKWARD). This means that incompatible producer changes are rejected when the schema is registered rather than at write time.

The partition spec travels on each DynamicRecord, so the sink applies it when it creates or updates the table. How our sample derives that spec is covered in the partitioning section.

Solution overview

The following diagram illustrates the solution architecture. A data generator (a local Java application) writes events to an Amazon Kinesis Data Stream. In Avro mode it also registers each event schema in the AWS Glue Schema Registry. A Managed Service for Apache Flink application consumes the stream, resolves a target Iceberg table for each record, and writes to Iceberg tables in Amazon S3, cataloged either in the AWS Glue Data Catalog or, for fully managed tables, in Amazon S3 Tables, a capability of Amazon S3.

Data generator sends events to Amazon Kinesis Data Streams, and Managed Service for Apache Flink routes each record to an Iceberg table in Amazon S3

Figure 1: Solution architecture for routing streaming records to per-event Iceberg tables on Managed Service for Apache Flink

At a high level, a single Managed Service for Apache Flink application reads raw records from Kinesis and resolves a target Iceberg table for each record. It uses the DynamicIcebergSink to create and evolve tables on demand. The same job handles many event types because the destination is decided per record, not per sink.

A note on stream topology: the examples assume one Kinesis stream carrying multiple event types, which keeps the walkthrough focused. This is not a requirement for the pattern. If your events arrive on separate streams (for example, one stream per producer or per domain), create one KinesisStreamsSource per stream and union them into a single DataStream before the sink. The routing generator chooses the destination table from the record itself, so many sources can fan into one DynamicIcebergSink and still land in the correct tables.

Unioning does not add shuffle cost. The sink always re-distributes records by an internal per-table writer key, so a unioned stream and N separate pipelines incur the same per-record exchange. The distribution mode each DynamicRecord carries only changes which writer subtask a row lands on, not whether a shuffle occurs. The real tradeoff is isolation. All tables share one writer pool, one commit aggregator, and one committer. A hot stream’s backpressure and checkpoint alignment therefore couple to every other stream, and writer parallelism is a single job-wide setting. Prefer one unioned pipeline when you have many small-to-medium event types that should pool capacity. Split into separate applications when one stream is high-volume enough to need its own writer parallelism and failure isolation.

DynamicIcebergSink needs a schema for every record. The sample provides two interchangeable ways to obtain it, implemented as two generator variants: Option 1 infers the schema from each JSON record at runtime. Option 2 reads the registered schema from AWS Glue Schema Registry. Everything downstream (routing, table creation, and schema evolution) is identical, and only the generator changes.

Option 1: Infer the schema from the JSON record

SchemaAgnosticRoutingGenerator implements Iceberg’s DynamicRecordGenerator. Its generate method maps the routing field to a table name, infers the schema, derives a partition spec, and emits a DynamicRecord through the collector:

@Override
public void generate(JsonNode json, Collector<DynamicRecord> out) {
    String tableName = determineTableName(json); // routing field -> table name
    TableIdentifier tableId = TableIdentifier.of(database, tableName);
    Schema schema = inferSchemaFromJson(json); // cached by schema signature
    RowData rowData = convertJsonToRowData(json, schema);
    PartitionSpec spec = buildPartitionSpec(schema); // cached per schema
    out.collect(new DynamicRecord(
        tableId, "main", schema, rowData, spec, DistributionMode.NONE, 4));
}

The table name comes from an explicit table-name field when present, otherwise from the routing field (event_type by default).

For schemaless or semi-structured JSON, the generator infers an Iceberg schema directly from each record. This is convenient, but inference is fundamentally lossy because JSON does not carry type information. The generator therefore applies deliberately conservative rules and selects a stable type rather than the narrowest one:

JSON value Iceberg type
Integer LongType (all integral values are widened to long)
String StringType
Floating-point values DoubleType
Boolean BooleanType
ISO-8601 timestamps TimestampType (microseconds)
Nested JSON object StructType (with fields inferred recursively)
JSON array ListType (with element type inferred from array contents)

Partitioning the routed tables

Partitioning is decided by our generator, not by the sink, and the same mechanism applies to both schema options: the JSON-inference and schema-registry generators share the partition-candidate logic. The open source DynamicIcebergSink applies whatever PartitionSpec each DynamicRecord carries. Our sample’s SchemaAgnosticRoutingGenerator builds that spec at runtime: it reads a list of candidate partition fields from the partition.candidates application property and derives a per-table spec from the fields it observes. For each table, buildPartitionSpec walks that list and keeps only the candidates present in the table’s schema.

The same list adapts to each table. A table with event_date and region is partitioned by identity(event_date) and identity(region). A table with none of the candidates is created unpartitioned. The resulting spec travels on each DynamicRecord, so the sink applies it when it first creates the table.

For example, with partition.candidates = event_time,region,product: a table whose schema has event_time and product is created partitioned by those two. A table with only event_time gets identity(event_time). A table with none of the candidates is created unpartitioned. Partition specs are not frozen at creation time either: the sink evolves them through Iceberg partition-spec evolution, adding a candidate field when it later appears in the table’s schema and removing one that disappears. This is a metadata-only change, so existing data files keep the spec they were written with.

Two operational practices follow. First, always include your event-time field among the candidates so every table is at least time-partitioned, and monitor for unpartitioned tables through the table’s $partitions metadata or its spec in the catalog: a producer that emits create_timestamp instead of event_time will silently create unpartitioned tables until the candidate list is updated. Second, be deliberate with generic fields like region. If a source produces high-cardinality values for a candidate field, you can correct the spec later. Evolution applies to newly written files only, so the small files already written remain until compaction rewrites them.

Note that the candidate list is global, not per table. It tracks every field you might partition on, and each table takes only the ones it has.

Option 2: Read the schema from a schema registry

Inference is convenient but lossy, and it offers no contract: nothing stops a producer from silently changing a field’s type or meaning. The second option removes the guesswork by reading the schema from a registry instead of the data. Many production streaming platforms standardize on strongly typed Avro schemas managed through AWS Glue Schema Registry. With GSR, producers register schemas explicitly, each record on Kinesis is Avro-encoded and prefixed with a schema-version ID, and the consumer decodes against the exact registered schema. That gives you three things JSON inference cannot: precise types (a long stays a long, a timestamp-micros stays a timestamp-micros), a governed evolution policy enforced at registration, and a single source of truth shared across producers and consumers.

The pattern works with any schema registry that gives consumers the writer’s schema per record. The sample implements it with AWS Glue Schema Registry, but the same generator shape applies to other registries.

The dynamic-sink-avro-sample module applies GSR-managed Avro schemas to the same dynamic routing and schema evolution pattern. For each record, AvroToDynamicRecordGenerator reads the schema-version ID and fetches the writer schema from GSR, caching it after the first lookup. It then converts that schema to an Iceberg schema, decodes the payload into RowData, and emits a DynamicRecord, exactly as the JSON generator does:

The sink wiring is identical to option 1. Only the generator changes, and because the source carries raw Avro bytes the input stream is byte[] rather than parsed JSON:

AvroToDynamicRecordGenerator generator = new AvroToDynamicRecordGenerator(
    awsRegion, registryName, database, partitionCandidates, branch);
DynamicIcebergSink.forInput(eventBytes)
    .generator(generator)
    // identical catalogLoader, immediateTableUpdate(true), cache, and write settings as option 1
    .append();

Because the schema comes from GSR rather than from inspecting bytes, the Avro-to-Iceberg type mapping is exact:

Category Avro type Iceberg type
Primitive int IntegerType
Primitive long LongType
Primitive float FloatType
Primitive double DoubleType
Primitive string StringType
Primitive boolean BooleanType
Logical timestamp-millis TimestampType (preserves millisecond precision)
Logical timestamp-micros TimestampType (preserves microsecond precision)
Logical decimal DecimalType
Complex record StructType (nested fields mapped recursively)
Complex array ListType (element type inferred from items schema)
Complex map MapType (keys are always StringType)

The GSR integration handles schema versioning transparently. As soon as a producer registers a new schema version containing additional fields, the Flink consumer deserializes the updated payload and evolves the Iceberg table to match, with no job restart.

Prerequisites

To follow along, you need the following:

  • An AWS account with permissions to create Amazon Kinesis Data Streams, Managed Service for Apache Flink applications, AWS Glue resources, and Amazon S3 buckets (plus Amazon S3 Tables if you choose that catalog).
  • The AWS Command Line Interface (AWS CLI) configured with credentials.
  • Node.js 18 or later and the AWS Cloud Development Kit (AWS CDK) CLI.
  • Java 17 or later and Apache Maven 3.9 or later, to build the data generator.
  • Docker running locally. The CDK build bundles the Flink application jars inside a Maven image.

Deploy and test the solution

The accompanying repository provisions everything through a single parameterized AWS CDK stack.

  1. Install the CDK dependencies and bootstrap your environment (first time only):
    cd cdk-infrastructure && npm install
    npx cdk bootstrap aws://<account>/<region>

  2. Deploy the variant you want to try:
    npx cdk deploy -c appType=dynamic -c tableFormatVersion=2 # JSON inference variant
    npx cdk deploy -c appType=dynamic-avro -c tableFormatVersion=2 # GSR Avro variant

    Add -c catalogType=s3tables to either command to use Amazon S3 Tables instead of the AWS Glue Data Catalog. The walkthrough sets tableFormatVersion=2 so you can query the results with a broad range of engines. Omit it to use the default, Iceberg format version 3, when you query with a v3-aware engine such as Spark on Amazon EMR 7.12+ or AWS Glue ETL.

  3. Start the application using the ApplicationName value from the stack outputs:
    aws kinesisanalyticsv2 start-application --application-name <ApplicationName> --run-configuration 'ApplicationRestoreConfiguration={ApplicationRestoreType=SKIP_RESTORE_FROM_SNAPSHOT}'

  4. Send test events with the included data generator. Start with the v1 payloads, which create the tables without the optional fields:
    java -jar data-generator/target/data-generator-1.0-SNAPSHOT.jar <stream-name> <region> 100 60 v1

    Then send v2 payloads, which add the userAgent and scrollDepth fields. This second run is the schema evolution you observe in the next step:

    java -jar data-generator/target/data-generator-1.0-SNAPSHOT.jar <stream-name> <region> 100 60 v2

    For the Avro variant, the generator registers each schema version in the AWS Glue Schema Registry as it sends:

    java -jar data-generator/target/data-generator-1.0-SNAPSHOT.jar avro <stream-name> <region> <registry-name> 100 60

  5. Query the routed tables in Amazon Athena. You should see one Iceberg table per event type appear in the database within a checkpoint interval, and after sending v2 events, the new fields (userAgent, scrollDepth) show up as optional columns on the same tables. The Iceberg metadata tables (for example, SELECT * FROM "db"."table$snapshots") show each commit the sink makes.

Clean up

When you finish testing, delete the resources to stop incurring charges:

cd cdk-infrastructure && npx cdk destroy

CDK removes the Kinesis Data Stream, the Managed Service for Apache Flink application, and the stack-created AWS Identity and Access Management (IAM) roles. Additionally, empty and delete the S3 warehouse bucket to remove the Iceberg data and metadata files, delete any schemas the Avro variant registered in the AWS Glue Schema Registry, and delete the table bucket contents if you used the S3 Tables catalog.

Conclusion

With Apache Iceberg 1.11.0 and Flink 2.3, you can build streaming data lake architectures that adapt to change without stopping the pipeline. With per-record routing, a single Flink application can write multiple event types to separate Iceberg tables, while automatic schema evolution keeps table definitions aligned with changing source data. Choosing AWS Glue Schema Registry over runtime JSON inference adds precise types and a governed evolution contract, and a configurable partition-candidate list keeps each routed table partitioned correctly without pre-declaring its schema.

The result is fewer pipeline redeployments, reduced operational overhead, and a data lake that remains synchronized with evolving application schemas.

To get started, follow the deploy and test section, then adapt the routing field and partition candidates to your own event types.

The full sample code is available in the accompanying GitHub repository.


About the authors

Francisco Morillo

Francisco Morillo

Francisco is a Sr. Streaming Solutions Architect at AWS, specializing in real-time analytics architectures. With over five years in the streaming data space, Francisco has worked as a data analyst for startups and as a big data engineer for consultancies, building streaming data pipelines. He has deep expertise in Amazon Managed Streaming for Apache Kafka (Amazon MSK) and Amazon Managed Service for Apache Flink.

Felix John

Felix John

Felix is a Global Solutions Architect and data & AI expert at AWS, based out of Germany. He focuses on supporting AWS’ strategic global automotive & manufacturing customers on their data & AI transformation journey.

Observing and evaluating production agents using OpenSearch Agent Health

Post Syndicated from Ulrich Hinze original https://aws.amazon.com/blogs/big-data/observing-and-evaluating-production-agents-using-opensearch-agent-health/

As AI agents are moving from experimental prototypes to production workloads, teams need visibility into what agents are doing and a systematic way to measure whether they’re doing it well. Traditional testing methodologies like unit and integration tests fall short for this task, as measuring an agent’s quality isn’t a straightforward true/false decision. Instead, agent observability and evaluations (evals for short) provide a two-legged solution to this problem. Agent observability captures the details of an agent’s behavior, and evals compare this behavior to the behavior that you want. With this approach, teams can monitor their agent’s quality over time and introduce agent-specific quality gates in their software development lifecycle.

In this post, we show how to combine an AI agent running on AWS with OpenSearch Agent Health for observability and evals. You will deploy an agent and its observability data pipeline to AWS, then use Agent Health as a local development tool connecting to your cloud resources.

Overview of solution

Agent observability and evaluations rely on OpenTelemetry traces to understand agent behavior. Traces describe the flow of a request through components of a system. OpenSearch Agent Health is a purpose-built tool for analyzing agent traces and running evaluations against an agent for quality control. Although Agent Health works with any open source OpenSearch installation, many AWS customers choose Amazon OpenSearch Ingestion and Amazon OpenSearch Service for ingesting and storing their OpenTelemetry data. You can connect OpenSearch Agent Health to these AWS resources to fetch live data and store its own configuration and evaluation history.

The following diagram shows the overall architecture of the solution presented in this post: Architecture diagram showing the agent, Amazon OpenSearch Ingestion, Amazon OpenSearch Service, and OpenSearch Agent Health observability and evaluation flow

Figure 1: Solution overview

The individual parts are:

  1. AWS Amplify for hosting an assistant-ui chat interface. Connects to the agent backend using the Agent-User Interaction (AG-UI) protocol.
  2. Sample ecommerce AI agent using Strands Agents SDK, deployed to Amazon Bedrock AgentCore runtime, exposing an AG-UI Server-Sent Events (SSE) endpoint. This agent has access to multiple tools, such as product search and shopping basket operations. For this sample project, the tool calls are all simulated within the agent runtime rather than including API calls to other systems. The agent emits messages, reasoning steps, and tool calls as OpenTelemetry traces.
  3. Large language models (LLMs) on Amazon Bedrock. One model (Amazon Nova 2 Lite) is used to power the agent, the other model (Anthropic Claude Opus 4.6) is used to evaluate the agent behavior.
  4. Amazon OpenSearch Ingestion for collecting and transforming the raw agent traces and loading them into an Amazon OpenSearch Service domain. Agent traces have the same structure as regular OpenTelemetry traces, with the addition of generative AI semantics (for example, tool calls and token usage). This means a regular OpenTelemetry pipeline configuration can be used to process agent traces.
  5. OpenSearch Agent Health for analyzing traces and running evaluation test cases and benchmarks against the agent. Agent Health uses the same AG-UI endpoint as the front-end application. It authenticates to the application, to Amazon Bedrock for model functionality, and to Amazon OpenSearch Service using AWS SigV4 authentication.

Walkthrough

In this walkthrough, we showcase how you can use Agent Health and Strands to measure and improve your agent’s quality over time.

We follow these steps:

  • Deploy solution to AWS and test the application.
  • Start Agent Health locally and connect it to cloud resources.
  • Explore agent traces and run evaluations.

We have created a GitHub repository for you to follow along.

Prerequisites

For this walkthrough, you should have the following prerequisites:

  • An AWS account
  • Git
  • Node.js
  • AWS Cloud Development Kit (AWS CDK)

Deploy solution to AWS and test the application

In this section, you check out the repository and deploy the infrastructure to AWS. Be aware that these steps create AWS resources that incur cost. We cover cleanup steps at the end of this post.

First, clone the repository to a local directory:

git clone https://github.com/aws-samples/sample-agent-health-with-amazon-opensearch-service && cd sample-agent-health-with-amazon-opensearch-service

Switch to the infra folder and install dependencies:

cd infra && npm install

Before you can start the deployment, determine the AWS Identity and Access Management (IAM) user or role that you will use to start Agent Health later on. In many cases, this will be the same role that you use to deploy the infrastructure. Set this ARN in your environment by issuing the following command:

export AGENT_HEALTH_READER_ARN=arn:aws:iam::<YOUR_ACCOUNT_ID>:role/<YOUR_ROLE_NAME>

Bootstrap your AWS account for use with AWS CDK:

cdk bootstrap -c agentHealthReaderArn=$AGENT_HEALTH_READER_ARN

Run the infrastructure deployment. Review and acknowledge IAM statement changes when prompted. This takes around 25 minutes to complete:

cdk deploy -R -c agentHealthReaderArn=$AGENT_HEALTH_READER_ARN

The -R parameter defines that if something fails during this deployment, the successfully provisioned resources are retained. Be aware that this command creates AWS resources, incurring cost. Review the cleanup section at the end of this post for removing all created resources.

When the deploy command finishes successfully, you should see an output like the following:

...
AgentObservabilityStack

✨ Deployment time: 1402.42 s

Outputs:
AgentObservabilityStack.AgentEndpoint = https://bedrock-agentcore.us-east-1.amazonaws.com/runtimes/arn%3Aaws%3Abedrock-agentcore%3Aus-east-1%3A123456789012%3Aruntime%2Fretail_agent-abcdefghij/invocations?qualifier=AgentObservabilityStAgentRuntimeEndpointABCDEFGH
AgentObservabilityStack.ChatUrl = https://main.abcdefghijklmn.amplifyapp.com
...

Next, create a user for your application. Retrieve the CDK output value for AgentObservabilityStack.UserPoolId. Create a user for the application using the user pool ID, an email address, and a strong password (minimum eight characters including uppercase, lowercase, letter, and digit):

export COGNITO_EMAIL=<YOUR_EMAIL>
export COGNITO_PASSWORD=<YOUR_PASSWORD>
export USER_POOL=<YOUR_USER_POOL_ID>
aws cognito-idp admin-create-user --user-pool-id $USER_POOL --username $COGNITO_EMAIL --message-action SUPPRESS --user-attributes Name=email_verified,Value=true
aws cognito-idp admin-set-user-password --user-pool-id $USER_POOL --username $COGNITO_EMAIL --password "$COGNITO_PASSWORD" --permanent

You can now access the retail agent application. From the CDK output values, retrieve the value for AgentObservabilityStack.ChatUrl. Copy and paste this URL into your browser. Log in with your email and password. You should now see the agent interface:

Sample retail agent chat interface showing the ecommerce assistant ready for queries

Figure 2: Sample retail agent user interface

Experiment with the application. Here is an example sequence of queries you can put in:

  • Do you have books on Python?
  • Is this in stock?
  • Put it into my basket.
  • What else can you do for me?

Start Agent Health

Now that you have the infrastructure running, you can start OpenSearch Agent Health locally and connect it to your cloud resources.

The CDK infrastructure deployment created a file cdk-output.json, which contains all relevant configuration values for Agent Health. We’ve already created a file agent-health/agent-health.config.ts that pulls these values dynamically in your environment, so you can start Agent Health without any further configuration.

Open a terminal and start Agent Health by running the following command:

cd ../agent-health && npm install
npx @opensearch-project/agent-health

Open http://localhost:4001 in your browser to access Agent Health UI. Choose Agent Traces in the sidebar menu to access your agent’s traces. You should see traces from your previous interactions:

Agent Health Traces view listing agent traces captured from previous interactions

Figure 3: Agent traces. As Agent Health is in active development, this interface might have changed since the time of writing

Expand the trace and explore the information it contains, such as token count and agent trajectory (sequence of messages, reasoning steps, and tool calls).

If you’re unable to access the application or see any traces, verify the following:

  • Check Agent Health logs in your terminal for any errors. Also check whether Agent Health is running on an alternative port, like 4002 instead of 4001.
  • If there are permission errors when accessing traces from OpenSearch, verify that the AWS credentials in your terminal match the principal (user or role) that you specified under the agentHealthReaderArn CDK parameter during cdk deploy. This principal must have ESHttpGet:* IAM permissions. Agent Health uses your current AWS credentials to access the OpenSearch API for querying traces. The OpenSearch API is guarded by both IAM and OpenSearch fine-grained access control.

Create and run a test

Choose Test Cases and New Test Case. Fill out the required fields with the following information:

  • Name: Should add to cart.
  • Initial Prompt: Add some wireless headphones to my cart. Take any that you have in stock.
  • Expected Outcomes: PROD-001 added to cart.

Back in the test cases overview, select the created test case and choose Run Test. In the Configure Run dialog, choose Retail Assistant (production) for Agent, Tool Usage Efficiency for Evaluator, Claude Opus 4.8 for Judge Model, and choose Start Run.

Agent Health now runs the configured initial prompt against the agent. The agent completes the task and sends execution traces to OpenSearch. Agent Health uses an evaluation model to check both agent responses and traces on successful execution, according to the defined expected outcomes. After the test is completed, go through the different tabs to check the test results.

Agent Health evaluation report showing test results across multiple tabs

Figure 4: Agent Health evaluation report

If you’re unable to run the test, check the following:

  • Agent Health automatically creates an Amazon Cognito token for your user upon start, but this token can expire. Restarting Agent Health creates a new token. Verify that both the COGNITO_EMAIL and COGNITO_PASSWORD variables are still set in your terminal environment.

Beyond test cases

After running a single test case, choose Benchmarks in the sidebar menu. With Benchmarks, you can run multiple test cases in parallel and summarize their results. You can compare benchmark runs by choosing Evaluation Runs in the sidebar, where you can analyze trends in pass rate, cost, and duration over time. Lastly, choose Evaluators to define your own evaluation logic beyond the predefined ones.

You can also run Agent Health tests with its command-line interface, which is handy for automation and continuous integration (CI). The equivalent command of running the preceding test is:

npx @opensearch-project/agent-health run -t <TEST_CASE_ID> -a "Retail Assistant" -e system-tool-usage --judge-model claude-opus-4.8 -e system-tool-usage

where TEST_CASE_ID can be retrieved from the browser URL when you visit the Agent Health UI (test case IDs start with tc-).

Agent Health stores all test cases, other configuration, and reports locally on disk in the agent-health/agent-health-data directory.

Cleaning up

To avoid incurring future charges, delete the resources:

cd ../infra && cdk destroy -c agentHealthReaderArn=$AGENT_HEALTH_READER_ARN

Conclusion

In this post, you learned to set up and use OpenSearch Agent Health for production agent observability and evaluations. To discover more features, see the Agent Health documentation pages. You can discuss and request additional features, and get help with setup, through the issues in the GitHub project. For more information, see the observability documentation for Amazon OpenSearch Service, where you can learn about the features available to build observability for both agents and traditional systems using OpenSearch. To investigate issues in production AI agents, see the recent post Unified observability in Amazon OpenSearch Service.


About the authors

Ulli Hinze

Ulli is a Solutions Architect based in Berlin, Germany. He focuses on SaaS, agentic AI, and OpenSearch, and helps customers build and modernize their solutions on AWS. His previous roles included software development, platform engineering, and architecture.

Megha Goyal

Megha is a Senior Software Engineer at AWS OpenSearch. For the past year she has focused on AI agent observability and evaluations, building Agent Health — an open-source developer tool for agents. Previously, she worked on data integrations with Amazon CloudWatch and Amazon Security Lake for the observability and security space. When she’s not building software, she enjoys designing and 3D-printing models at home.

Rekha Thottan

Rekha Thottan

Rekha is a Senior Product Manager Technical on the Amazon OpenSearch Service team.

Accelerate Apache Spark debugging on Amazon EMR with AWS DevOps Agent

Post Syndicated from Kalyan Janaki original https://aws.amazon.com/blogs/big-data/accelerate-apache-spark-debugging-on-amazon-emr-with-aws-devops-agent/

When an Apache Spark job fails on Amazon EMR, the root cause can hide in executor logs, memory profiles, or application code. As data pipelines grow in complexity, correlating logs, metrics, and traces across multiple services requires significant operational effort. AWS DevOps Agent handles this investigation autonomously while keeping operators in the loop to review findings and approve fixes. From a single chat prompt, it produces a root cause and mitigation plan, often without any human involvement beyond the initial question.

The native AWS API tools in AWS DevOps Agent don’t extend into Spark-internal artifacts. Sometimes those tools can’t reach the evidence that pins down the root cause: a Spark History Server event log, executor Python worker memory, or a line of code that allocated too much. In these cases, AWS DevOps Agent can describe symptoms (“the executor exited with code 1”) but can’t identify the actual antipattern that caused them.

This post shows how to extend AWS DevOps Agent to investigate failures in Apache Spark workloads on Amazon EMR. You register the Apache Spark Troubleshooting Agent for Amazon EMR, a managed Model Context Protocol (MCP) server hosted by AWS, as a custom capability provider in your AWS DevOps Agent space. You route the traffic over AWS PrivateLink so MCP calls never traverse the public internet. Then you watch a single agent chat session investigate a deliberately failing Spark job, from Amazon CloudWatch alarm to line-numbered root cause, in about two minutes.

Prerequisites

Before you begin, make sure you have the following:

How AWS DevOps Agent discovers custom tools through MCP

Model Context Protocol (MCP) is an open standard that defines how AI agents discover and invoke external tools. AWS DevOps Agent supports connecting to custom MCP servers, which means you can expose new capabilities to it without modifying the agent itself. When you connect an MCP server to AWS DevOps Agent, the agent automatically discovers the available tools, understands their schemas, and calls them as part of its investigation workflow. You build and connect the MCP server, and the agent handles the rest.

MCP tools sit alongside the agent’s built-in AWS API tools. During a single investigation, the agent can interleave calls to cloudwatch.describe-alarms, emr-serverless.get-job-run, and a custom MCP tool such as analyze_spark_workload. The agent picks the right one for each subtask. You augment the agent’s reach without replacing what it already does.

For this integration, you don’t build an MCP server. The Apache Spark Troubleshooting Agent for Amazon EMR is itself a managed MCP server, hosted by AWS at a regional endpoint. Your job is to register that endpoint with AWS DevOps Agent and authorize the agent to call it. This requires a network path from the agent to the endpoint, plus an IAM role for AWS Signature Version 4 request signing.

Why Spark internals visibility matters

The actual root cause for a Spark failure usually lives somewhere none of those APIs (such as Amazon CloudWatch Logs Insights, AWS CloudTrail, or Amazon EMR step-status calls) can reach:

The Apache Spark Troubleshooting Agent for Amazon EMR reads the following sources.

The Spark History Server event log is a per-job archive in Amazon Simple Storage Service (Amazon S3) with stage timings, task-level metrics, executor utilization, shuffle read/write volumes, and garbage-collection pauses. Amazon EMR exposes this data through the Spark UI on Amazon EMR Serverless, Amazon EMR on Amazon Elastic Compute Cloud (Amazon EC2), and Amazon EMR on Amazon Elastic Kubernetes Service (Amazon EKS), but interpreting signals like data skew, executor memory pressure, or stages that take significantly longer than expected requires familiarity with Spark internals.

  • The Spark query plan — the logical and physical plan the driver compiled. Without it, you can’t identify antipatterns such as unnecessary data repartitioning or missing broadcast hints that trigger expensive shuffles.
  • The application source code in Amazon S3 — the .py or .jar code artifact the job ran. Without it, you can’t quote the offending line of a mapPartitions user-defined function or an inefficient collect().
  • The Python worker process telemetry — the PySpark worker is a separate Python subprocess outside the Java Virtual Machine’s (JVM) managed memory. When it crashes from spark.executor.pyspark.memory exhaustion, the JVM driver sees a generic “executor exited unexpectedly” message. The actual cause is invisible to standard JVM-level logs.

When the agent invokes analyze_spark_workload during an investigation, it returns a structured analysis with the antipattern identified at the line level, the offending stage isolated, and a concrete fix: both code changes and configuration changes.

Integrating AWS DevOps Agent with Apache Spark Troubleshooting MCP

This section explains how AWS DevOps Agent connects to the Apache Spark Troubleshooting Agent through a private MCP endpoint and orchestrates the investigation workflow.

How it works

Architecture diagram showing AWS DevOps Agent connecting to the Apache Spark Troubleshooting Agent over AWS PrivateLink

Figure 1: Integration architecture between AWS DevOps Agent and the Apache Spark Troubleshooting Agent for Amazon EMR over AWS PrivateLink

  1. You submit an investigation prompt in AWS DevOps Agent.
  2. AWS DevOps Agent sends a SigV4-signed MCP call into your Amazon VPC through the AWS DevOps Agent private connection.
  3. The private connection forwards the request to the Interface VPC Endpoint.
  4. The endpoint routes the request over AWS PrivateLink to the Apache Spark Troubleshooting Agent for Amazon EMR, which AWS manages.
  5. The MCP service reads from your data sources (Amazon EMR, Amazon S3, Amazon CloudWatch Logs) using the same IAM role AWS DevOps Agent assumed for the call.
  6. When a CloudWatch alarm transitions to ALARM state (for example, a failed-jobs alarm for your Amazon EMR Serverless application), AWS DevOps Agent automatically triggers an investigation without manual intervention.
  7. AWS DevOps Agent decides which tools to call based on the prompt. For a Spark failure, that includes the Apache Spark Troubleshooting MCP server you registered as a capability provider.
  8. Each MCP request is signed with AWS Signature Version 4 using the IAM role assigned to the capability provider. The request travels from AWS DevOps Agent into your Amazon VPC through the private connection. This private connection is a managed VPC Lattice resource gateway you created during setup.
  9. From the resource gateway, the request flows to the Interface VPC Endpoint for the Amazon SageMaker Unified Studio MCP service, then on to the Apache Spark Troubleshooting Agent. The traffic stays entirely on the AWS network.
  10. The MCP server reads the inputs it needs from your AWS account using the IAM role that you assigned to the capability provider during MCP server registration. This role grants access to the Spark History Server event log and application source code in Amazon S3, the driver and executor stdout streams in Amazon CloudWatch Logs, and the job-run metadata from Amazon EMR Serverless.
  11. The MCP server returns its diagnostic findings to AWS DevOps Agent. The agent then analyzes the results, identifies the root cause, and presents recommended fixes both code-level and configuration-level in your chat.

Setting up the demo

As part of this demo, this post includes a sample AWS CloudFormation template, tested in the us-east-1 Region, that provisions the following resources for the walkthrough:

  • A dedicated Amazon Virtual Private Cloud (Amazon VPC) with two private subnets in Availability Zones supported by the Apache Spark Troubleshooting Agent for Amazon EMR.
  • An Interface VPC Endpoint for the Apache Spark Troubleshooting Agent for Amazon EMR.
  • An IAM role that AWS DevOps Agent assumes to invoke the Apache Spark Troubleshooting MCP server with AWS Signature Version 4.
  • A deliberately failing PySpark workload running on Amazon EMR Serverless, including the Amazon EMR Serverless application, the Spark execution role, and the demo logs stored in Amazon S3 bucket.
  • An Amazon CloudWatch alarm that fires when the demo job fails. This alarm is used as the trigger for the agent investigation later in this section.

Step 1: Clone the repository

Clone the git repository for the CloudFormation template, PySpark script, and Parquet data.

git clone https://github.com/aws-samples/sample-aws-data-processing-and-analytics.git

Step 2: Deploy the AWS CloudFormation stack

Deploy the template using the following AWS CLI command.

cd sample-aws-data-processing-and-analytics/blogs/devops-agent-spark-mcp-integration

aws cloudformation create-stack \
  --stack-name spark-troubleshooting-demo \
  --template-body file://cloudformation/spark-troubleshooting-devops-agent-blog.yaml \
  --capabilities CAPABILITY_NAMED_IAM \
  --region us-east-1

The stack reaches CREATE_COMPLETE in approximately 4–6 minutes. When it does, capture the following stack outputs, which you paste into the AWS DevOps Agent console in the next two steps:

  • DemoVpcId — the VPC ID for the AWS DevOps Agent private connection.
  • DemoSubnetIds — the two subnet IDs for the AWS DevOps Agent private connection.
  • SMUSVpcEndpointSecurityGroupId — the security group ID.
  • TroubleshootingRoleArn — the IAM role Amazon Resource Name (ARN).
  • MCPEndpointURL — the MCP endpoint URL to register.
  • FailedJobsAlarmName — the CloudWatch alarm name to reference in your investigation prompt.
  • DemoBucket — the S3 bucket name where you copy the demo script and Parquet data.

To retrieve all outputs at once, use the following AWS CLI command.

aws cloudformation describe-stacks \
  --region us-east-1 \
  --stack-name spark-troubleshooting-demo \
  --query "Stacks[0].Outputs" --output table
# Get your bucket name from the stack outputs
DEMO_BUCKET=$(aws cloudformation describe-stacks --stack-name spark-troubleshooting-demo --region us-east-1 --query 'Stacks[0].Outputs[?OutputKey==`DemoBucket`].OutputValue' --output text)

# Copy the script
cd sample-aws-data-processing-and-analytics/blogs/devops-agent-spark-mcp-integration
aws s3 cp scripts/customer_events_aggregator.py s3://$DEMO_BUCKET/customer_events_aggregator.py

# Copy the Parquet data
aws s3 cp data/ s3://$DEMO_BUCKET/data/ --recursive

Step 3: Create an agent space

The agent space defines which AWS account and Region the agent monitors, which IAM role it assumes, and which capability providers, including MCP servers, it can call.

Follow the steps in Creating an Agent Space in the AWS DevOps Agent User Guide. When completing those steps, use the following values:

Parameter Value
Name data-pipeline-troubleshooting
Region us-east-1
Agent Space role Choose Auto-create a new DevOps Agent role — the console generates a DevOpsAgentRole-AgentSpace* role with AIOpsAssistantPolicy attached
Optional integrations Not required

After the agent space reaches Active status, proceed to create the private connection.

Step 4: Create the AWS DevOps Agent private connection

AWS DevOps Agent uses the private connection to reach into your Amazon VPC. Follow the steps in Connecting to privately hosted tools in the AWS DevOps Agent User Guide. You can use either the console or the AWS CLI command documented under Create a private connection.

When completing those steps, use the following values from your CloudFormation stack outputs:

Parameter Value
Name A descriptive name (for example, spark-private)
VPC DemoVpcId from your stack outputs
Subnets Both subnet IDs from DemoSubnetIds
Security group SMUSVpcEndpointSecurityGroupId
TCP port ranges (Advanced configuration) 443
Host address (Service target details) sagemaker-unified-studio-mcp.us-east-1.api.aws
DNS resolution In VPC (private DNS)
Certificate public key None

After the connection reaches Active status, proceed to Step 5.

Step 5: Register the Apache Spark Troubleshooting MCP server as a capability provider

With the private connection in place, register the MCP server as a capability provider. Follow the steps in Registering an MCP server at the account level in the AWS DevOps Agent User Guide.

When completing those steps, use the following values:

Parameter Value
Name spark-troubleshooting
Endpoint URL MCPEndpointURL from your stack outputs
Connect to endpoint using a private connection Selected

Step 6: Add the MCP server to the agent space

With the MCP server registered, you need a workspace where investigations run. The agent space defines which AWS account and Region the agent monitors, which IAM role it assumes, and which capability providers, including MCP servers, it can call.

  1. In the MCP Server section, choose Add.
MCP Server section of the agent space detail page with an Add button

Figure 2: MCP Server section of the agent space detail page

  1. In the Add a capability dialog, locate spark-troubleshooting in the list of registered MCP servers and choose Add.
Add a capability dialog listing the spark-troubleshooting MCP server

Figure 3: Add a capability with the spark-troubleshooting MCP server listed

  1. On the Select MCP server tools page, both tools that the Apache Spark Troubleshooting Agent for Amazon EMR publishes are listed: analyze_spark_workload and analyze_spark_history_server_endpoint. Select both checkboxes, then choose Save.
Select MCP server tools page with both Spark troubleshooting tools checked

Figure 4: The Select MCP server tools page with both Spark troubleshooting tools selected

The agent space connects to the MCP server, lists its tools, and displays 2 Available / 2 Connected. Both tools are now part of your agent’s catalog.

MCP Server section showing spark-troubleshooting connected with two available and two connected tools

Figure 5: MCP Server section showing spark-troubleshooting connected with both tools available

Seeing it in action

To see the integration end to end, you submit a PySpark job, watch the CloudWatch alarm move to ALARM, and then ask AWS DevOps Agent to investigate using the alarm name.

The failing workload

The CloudFormation template provisioned an Amazon EMR Serverless application called analytics-events-platform and configured a sample PySpark job, customer_events_aggregator.py. The script simulates a common Python-side memory bug: a mapPartitions user-defined function accumulates 11 copies of every input row in an in-memory Python list before yielding results, while the job runs with spark.executor.pyspark.memory=256m. The Python worker process exceeds the 256 MB cap, the kernel kills it, Spark retries four times, and the stage is marked failed.

Submit the failing job

Run the DemoSubmitJobCommand from your stack outputs in your terminal. It looks like this:

aws emr-serverless start-job-run \
  --region us-east-1 \
  --application-id <DemoApplicationId> \
  --execution-role-arn <DemoExecutionRoleArn> \
  --name daily-customer-events-rollup \
  --job-driver '{"sparkSubmit":{"entryPoint":"s3://<DemoBucket>/customer_events_aggregator.py","entryPointArguments":["<DemoBucket>"],"sparkSubmitParameters":"--conf spark.executor.cores=2 --conf spark.executor.memory=1g --conf spark.executor.pyspark.memory=256m --conf spark.executor.instances=2"}}' \
  --configuration-overrides '{"monitoringConfiguration":{"s3MonitoringConfiguration":{"logUri":"s3://<DemoBucket>/logs/"}}}'

The command returns a jobRunId. Note it down. You will see it later in the agent’s investigation.

The job goes through PENDING to SCHEDULED to RUNNING to FAILED and reaches FAILED state in roughly four minutes.

Watch the CloudWatch alarm fire

The CloudFormation template also created a CloudWatch alarm named <DemoApplicationId>-FailedJobs (the exact name is in the FailedJobsAlarmName stack output). The alarm watches the FailedJobs metric in the AWS/EMRServerless namespace, scoped to your demo application, and flips to ALARM within a minute or two of the job failing.

Open the Amazon CloudWatch console, choose Alarms in the left navigation pane, and confirm the alarm is in In alarm state.

Amazon CloudWatch console alarm detail page showing the FailedJobs alarm in alarm state

Figure 6: The Amazon CloudWatch alarm detail page showing the FailedJobs alarm in the In alarm state

Ask AWS DevOps Agent to investigate

  1. Open your AWS DevOps Agent space.
  2. In the left navigation pane, choose Operator Access, then choose Incidents.
  3. Choose Start an investigation.
  4. Paste the following prompt, replacing <FailedJobsAlarmName> with the value from your stack outputs:

CloudWatch alarm in us-east-1 just went into ALARM state. Investigate why and recommend a fix

AWS DevOps Agent Start an investigation panel with the alarm prompt entered

Figure 7: AWS DevOps Agent Start an investigation panel with the Amazon CloudWatch alarm investigation prompt

The agent’s investigation chains together native AWS API tools and the Apache Spark Troubleshooting MCP tool you registered:

  1. use_aws cloudwatch describe-alarms — fetches the alarm definition and reads its metric dimensions, identifying that the alarm is scoped to Amazon EMR Serverless application <DemoApplicationId>.
  2. use_aws emr-serverless list-job-runs — finds the most recent FAILED job run on that application.
  3. use_aws emr-serverless get-job-run — pulls the FAILED run’s metadata and last-known error.
  4. spark-troubleshooting analyze_spark_workload — invokes the Apache Spark Troubleshooting Agent for Amazon EMR through the MCP capability provider, passing the application ID and job run ID. This is where the deep analysis happens.

Review the root cause and fix

When the investigation completes, AWS DevOps Agent presents the results across two tabs: Investigation timeline and Root cause.

The Investigation timeline shows every step the agent took: skills loaded, native AWS API calls made, and the moment it called the analyze_spark_workload MCP tool to analyze the failed Spark job. Each entry is expandable so you can audit the inputs and outputs.

Investigation timeline listing the agent tool calls and the MCP invocation

Figure 8: Investigation timeline tab showing the sequence of agent tool calls and the spark-troubleshooting MCP invocation

The Root cause tab is where the answer lands. It is organized into three sections that mirror what an experienced engineer would write in an incident report:

Root cause tab showing impact, root causes, and key findings for the memory failure

Figure 9: The Root cause tab showing the impact summary, identified root causes, and key findings for the Spark memory exhaustion failure

  • Impact — what failed, when, and for how long. For our demo, this calls out that the daily-customer-events-rollup job on the analytics-events-platform application failed with a MemoryError and that the alarm transitioned to ALARM state at the time of the failure.
  • Root causes — the actual antipattern. The agent identifies that customer_events_aggregator.py combines three compounding issues: an expand_event function (line 23) that amplifies each input row 11×, a repartition(1) that funnels all data into a single partition on a single executor, and a collect() (line 31) that pulls the amplified dataset back to the driver. All three run with only 1 GB of executor memory.
  • Key findings — supporting facts behind the diagnosis, including the executor memory configuration, the application’s maximum capacity, and how the agent confirmed each fact from the analyzed artifacts.

Both the antipattern identification and the supporting evidence come from artifacts the agent could only reach through the MCP tool: the application source code in Amazon S3, the Spark History Server event log, and the query plan. Without the Apache Spark Troubleshooting Agent for Amazon EMR plugged in, AWS DevOps Agent would have stopped at “the executor exited with a memory error.”

Clean up

To avoid ongoing charges, delete the resources you created. Some resources are managed by the AWS DevOps Agent console and must be removed there first. Otherwise, the CloudFormation stack deletion fails.

  1. In the AWS DevOps Agent console, open your data-pipeline-troubleshooting agent space, choose the MCP Server section, select spark-troubleshooting, and choose Remove.
  2. From the Agent spaces list, select data-pipeline-troubleshooting and choose Delete.
  3. In Capability Providers, select spark-troubleshooting and choose Deregister.
  4. In Capability Providers → Private connections, select smus-spark-private and choose Delete.
  5. Delete the AWS CloudFormation stack. This removes the Amazon VPC, the Interface VPC Endpoint, the security group, the IAM role, the Amazon EMR Serverless application, the Spark execution role, the Amazon CloudWatch alarm, and the demo logs bucket.
aws cloudformation delete-stack \
  --region us-east-1 \
  --stack-name spark-troubleshooting-demo

Conclusion

In this post, you connected the Apache Spark Troubleshooting Agent for Amazon EMR to AWS DevOps Agent as a custom MCP capability provider. You kept the traffic on the AWS network with AWS PrivateLink, and ran a failing PySpark job to see the integration end to end. A CloudWatch alarm fired, you asked the agent to investigate, and a single chat session returned the root cause along with code and configuration fixes.

You can extend this pattern beyond the demo scenario. Consider connecting the MCP server to agent spaces that monitor your production Amazon EMR environment. Any Spark job that writes a History Server event log becomes diagnosable through the same workflow.

To continue learning, explore the following resources:

If you’ve already integrated the Apache Spark Troubleshooting Agent into your operational workflow, or if you’re exploring other MCP-based extensions for AWS DevOps Agent, we want to hear about your experience. Share your thoughts and questions in the comments.


About the authors

Kalyan Janaki

Kalyan Janaki

Kalyan is Senior Big Data & Analytics Specialist with Amazon Web Services. He helps customers architect and build highly scalable, performant, and secure cloud-based solutions on AWS.

Aneesh Varghese

Aneesh Varghese

Aneesh is a Senior Technical Account Manager at AWS with more than 20 years of Information Technology industry experience. Aneesh supports enterprise customers in cost optimization strategies, Cloud operations, MLOps, providing advocacy and strategic technical guidance to help plan and build solutions using AWS best practices. Outside of work, Aneesh likes to spend time with family, play Basketball and Badminton

Scheduling email campaigns at scale with Amazon EventBridge Scheduler

Post Syndicated from Oluwaseun Ademuwagun original https://aws.amazon.com/blogs/compute/scheduling-email-campaigns-at-scale-with-amazon-eventbridge-scheduler/

Scheduling email campaigns becomes more complex when you need to send email to millions of recipients at the unique time best suited for each customer. Consider these examples:

  • A flash sale might need to hit inboxes at 9 AM local time across every time zone.
  • A follow-up email (often known as a drip sequence) might need to send a second message exactly 3 days after the first message per subscriber.
  • A re-engagement campaign might target users who haven’t logged in for 30 days.

The scheduling requirements involve multiple considerations. You’re sending hundreds of millions of messages, each at its own optimal moment personalized to the recipient’s time zone and behavior.

In this post, we walk through how to use Amazon EventBridge Scheduler to personalize email notifications to each recipient. We create one schedule per recipient to deliver each email at its individually optimal moment, with zero idle compute cost. We also show how Amazon EventBridge Scheduler handles higher volumes. Amazon EventBridge Scheduler supports billions of schedules. By default, you have a quota of 10 million schedules.

Solution overview

When every recipient has their own ideal delivery time, you need a scheduling layer that can hold billions of individual send intents and fire each one at the right moment. Most teams reach for one of three familiar patterns, each with tradeoffs that become painful at scale.

  1. Batch cron jobs: A job runs every hour, queries for all messages due in the next window, and sends them out. Recipients get email in imprecise hourly batches. At scale, the batch job itself becomes a bottleneck, processing millions of rows per run, competing for database connections, and creating a sudden spike in load on the email provider.
  2. Delay queues: You can use Amazon Simple Queue Service (Amazon SQS) as a delay queue. A delay queue postpones the delivery of new messages to a customer for a set time. A limitation of this approach is that Amazon SQS caps delays at 15 minutes.
  3. Third-party campaign tools: Offload to a SaaS email platform. This works until you need tight integration with your application data, custom send-time optimization, or control over delivery infrastructure. You’re also paying per-recipient fees that compound at scale.

All three approaches either sacrifice precision (batching), hit architectural limits (delay queues), or surrender control (third-party tools).

The building block approach

Amazon EventBridge Scheduler treats each email send as a discrete scheduled action. Instead of “process all messages due this hour,” you express the intent directly: “send this email to this person at this time.” Amazon EventBridge Scheduler holds that intent with zero compute cost until the moment arrives, then triggers the scheduled action. See the Amazon EventBridge Scheduler User Guide for the full API reference and current service quotas.

For email campaigns, Amazon EventBridge Scheduler becomes the send-time dispatcher, the component that schedules every email in a campaign for its individually optimal moment, whether that’s timezone-adjusted, behavior-triggered, or sequence-driven.

Architecture diagram

The architecture follows an event-driven, per-recipient scheduling pattern for an email campaign. To start the campaign, you first define the target audience and the content they receive. Next, you need a way to create the per-recipient schedule. To do that for a campaign that can contain millions of recipients, you need a scalable mechanism to create the schedules. You can achieve this with an AWS Step Functions state machine, a serverless workflow service that coordinates multiple AWS services into structured, visual workflows called state machines. In this solution, we orchestrate the creation of the schedules by using a Distributed Map state within the state machine, which lets us fan out and accelerate schedule creation. It does this by splitting a large dataset into chunks and processing them across thousands of parallel child executions. It reads the recipient list from Amazon Simple Storage Service (Amazon S3), applies time zone logic per recipient, and creates an individual Amazon EventBridge Scheduler resource for each recipient in parallel. After the workflow creates all schedules, the execution completes.

The actual email delivery happens later, entirely decoupled from the campaign creation step. At the scheduled time, Amazon EventBridge Scheduler invokes Amazon Simple Email Service (Amazon SES) directly, passing the template name and personalization data as template variables. For campaigns requiring complex personalization logic (conditional content, real-time suppression checks, or data enrichment), you can optionally route through an AWS Lambda function before SES. If you need to adjust timing or content for specific recipients, you can update their individual schedules directly without reprocessing the entire campaign.

Figure 1: Per-recipient email scheduling architecture with Amazon EventBridge Scheduler

Walkthrough

The solution uses four core components that work together: a campaign manager to define send-time rules, Step Functions Distributed Map to fan out and accelerate schedule creation, Amazon EventBridge Scheduler to hold each per-recipient intent and deliver through Amazon SES directly, and automatic cleanup through schedule self-deletion.

How it works

  1. Create the campaign: A marketer defines the campaign: audience segment, email template, and send-time rules (for example, “9 AM in each recipient’s local time zone” or “24 hours before a Black Friday sale”).
  2. Campaign manager fans out: An AWS Step Functions workflow uses Distributed Map to iterate over the recipient list and create one Amazon EventBridge Scheduler schedule per recipient per campaign step directly through SDK integration. Each schedule encodes the exact send time for that individual.
  3. Amazon EventBridge Scheduler fires at the right moment: At each recipient’s scheduled time, Amazon EventBridge Scheduler invokes Amazon SES directly through a universal target, passing the template name and personalization data (recipient name and attributes) as template variables.
  4. SES personalizes and sends: Amazon SES renders the email template with the provided data and delivers the message.
  5. Schedule self-deletes: ActionAfterCompletion='DELETE' prevents the accumulation of spent schedules.

Prerequisites

To follow along with this walkthrough, you need the following:

  • AWS account and permissions: An active AWS account with permissions to create Amazon EventBridge Scheduler schedules, AWS Step Functions state machines, and Amazon SES identities, along with an AWS Identity and Access Management (IAM) role for Amazon EventBridge Scheduler to invoke Amazon SES.
  • Development environment: Python 3.13 or later, AWS SDK for Python (Boto3) version 1.26 or later, and AWS Command Line Interface v2 (AWS CLI v2).
  • Amazon SES configuration: Move your Amazon SES account out of sandbox mode to allow sending to arbitrary recipients.

Scaling the fan-out with Step Functions

For campaigns with millions of recipients, use AWS Step Functions Distributed Map to parallelize schedule creation. When you want to activate a campaign, you trigger a Step Functions workflow. This workflow fans out and creates schedules across the recipient list by using a Distributed Map with direct SDK integration. The direct SDK integration between Step Functions and Amazon EventBridge Scheduler lets each child execution call CreateSchedule directly. The following state machine definition reads recipients from an Amazon S3 CSV file and creates schedules in parallel:

{
  "Comment": "Fan out campaign schedule creation via direct SDK integration",
  "StartAt": "EnsureScheduleGroup",
  "States": {
    "EnsureScheduleGroup": {
      "Type": "Task",
      "Resource": "arn:aws:states:::aws-sdk:scheduler:createScheduleGroup",
      "Parameters": {
        "Name.$": "States.Format('campaign-{}', $.campaign_id)"
      },
      "ResultPath": null,
      "Catch": [
        {
          "ErrorEquals": [
            "Scheduler.ConflictException"
          ],
          "ResultPath": null,
          "Next": "FanOutRecipients"
        }
      ],
      "Next": "FanOutRecipients"
    },
    "FanOutRecipients": {
      "Type": "Map",
      "ItemProcessor": {
        "ProcessorConfig": {
          "Mode": "DISTRIBUTED",
          "ExecutionType": "STANDARD"
        },
        "StartAt": "BuildScheduleInput",
        "States": {
          "BuildScheduleInput": {
            "Type": "Pass",
            "Parameters": {
              "schedule_name.$": "States.Format('campaign-{}-{}', $.campaign_id, $.recipient.id)",
              "group_name.$": "States.Format('campaign-{}', $.campaign_id)",
              "schedule_expression.$": "States.Format('at({}T{}:00:00)', $.send_date_date, $.send_hour)",
              "timezone.$": "$.recipient.timezone",
              "target_input": {
                "FromEmailAddress": "[email protected]",
                "Destination": {
                  "ToAddresses.$": "States.Array($.recipient.email)"
                },
                "Content": {
                  "Template": {
                    "TemplateName.$": "$.template_id",
                    "TemplateData.$": "States.JsonToString($.recipient.attributes)"
                  }
                }
              }
            },
            "Next": "CreateSchedule"
          },
          "CreateSchedule": {
            "Type": "Task",
            "Resource": "arn:aws:states:::aws-sdk:scheduler:createSchedule",
            "Retry": [
              {
                "ErrorEquals": [
                  "Scheduler.SdkClientException"
                ],
                "IntervalSeconds": 2,
                "MaxAttempts": 3,
                "BackoffRate": 2
              }
            ],
            "Parameters": {
              "Name.$": "$.schedule_name",
              "GroupName.$": "$.group_name",
              "ScheduleExpression.$": "$.schedule_expression",
              "ScheduleExpressionTimezone.$": "$.timezone",
              "FlexibleTimeWindow": {
                "Mode": "FLEXIBLE",
                "MaximumWindowInMinutes": 5
              },
              "Target": {
                "Arn": "arn:aws:scheduler:::aws-sdk:sesv2:sendEmail",
                "RoleArn": "arn:aws:iam::976764934189:role/CampaignFanOutRole-dev",
                "Input.$": "States.JsonToString($.target_input)",
                "RetryPolicy": {
                  "MaximumEventAgeInSeconds": 7200,
                  "MaximumRetryAttempts": 5
                }
              },
              "ActionAfterCompletion": "DELETE"
            },
            "ResultPath": null,
            "End": true
          }
        }
      },
      "ItemReader": {
        "Resource": "arn:aws:states:::s3:getObject",
        "ReaderConfig": {
          "InputType": "CSV",
          "CSVHeaderLocation": "FIRST_ROW"
        },
        "Parameters": {
          "Bucket.$": "$$.Execution.Input.recipient_bucket",
          "Key.$": "$$.Execution.Input.recipient_key"
        }
      },
      "ItemSelector": {
        "campaign_id.$": "$$.Execution.Input.campaign_id",
        "template_id.$": "$$.Execution.Input.template_id",
        "send_date_date.$": "$$.Execution.Input.send_date_date",
        "send_hour.$": "$$.Execution.Input.send_hour",
        "recipient": {
          "id.$": "$$.Map.Item.Value.id",
          "email.$": "$$.Map.Item.Value.email",
          "timezone.$": "$$.Map.Item.Value.timezone",
          "attributes": {
            "first_name.$": "$$.Map.Item.Value.first_name",
            "signup_date.$": "$$.Map.Item.Value.signup_date"
          }
        }
      },
      "MaxConcurrency": 1000,
      "ResultPath": null,
      "End": true
    }
  }
}

Concurrency alignment with Amazon EventBridge Scheduler API limits

Step Functions Distributed Map supports up to 10,000 concurrent child workflows. Each child calls the CreateSchedule API directly, which has a default rate limit of 5,000 TPS. This limit is sufficient for most campaigns. If your campaign volumes require higher throughput, check your current quotas in the Service Quotas console and request an increase.

To avoid throttling, set MaxConcurrency below the CreateSchedule TPS quota. A value of 2,500 provides a comfortable buffer to account for bursts and retries without requiring a quota change. For larger campaigns, request an increase through AWS Service Quotas (adjustable to tens of thousands) and raise MaxConcurrency to match.

Canceling a campaign

A schedule group is an Amazon EventBridge Scheduler resource used to organize schedules. For this use case, we have a schedule group per campaign. If you need to pull a campaign (error in content, legal issue, or strategy change), you can cancel all scheduled sends for that campaign by deleting the entire schedule group. The following code shows how to cancel all pending sends for a campaign:

def cancel_campaign(campaign_id):
    """Cancel all pending sends for a campaign by deleting its schedule group."""
    scheduler.delete_schedule_group(
        Name=f'campaign-{campaign_id}'
    )

Operational considerations

Moving to production introduces a few scaling and reliability concerns to plan for.

Handling invocation spikes at delivery time

When a mass campaign schedules millions of messages for the same time, this creates cascading pressure across two limits:

  • Amazon EventBridge Scheduler invocations throttle limit: The default is 1,000 TPS per AWS Region, and it is adjustable to tens of thousands of TPS through AWS Service Quotas. Amazon EventBridge Scheduler queues invocations internally and retries with exponential backoff when the downstream target throttles.
  • Amazon SES sending quotas: Your SES account has a per-second sending rate. If the effective invocation rate exceeds this, messages fail with throttling errors. Align Amazon SES sending quotas with your campaign volume. Check your current SES quota in the Service Quotas console and request an increase before launching large campaigns. See Amazon SES best practices for deliverability at scale.

To handle an invocation spike, we recommend using the FlexibleTimeWindow feature of Amazon EventBridge Scheduler. Setting MaximumWindowInMinutes lets Amazon EventBridge Scheduler spread invocations across a time window rather than firing them all at the exact second. Size the window based on your campaign: divide the total schedules by your effective TPS to determine the minimum spread needed. For example, 500,000 schedules at 5,000 TPS need at least a 2-minute window.

Cost model

You pay for Amazon EventBridge Scheduler on a per-invocation basis.

Cleanup

To avoid ongoing charges, delete the resources created during this walkthrough:

  1. Delete any runtime-created schedule groups.
    aws scheduler delete-schedule-group --name campaign-<campaign-id>

  2. Delete the Step Functions state machine.
    aws stepfunctions delete-state-machine \
        --state-machine-arn arn:aws:states:us-east-1:<account-id>:stateMachine:CampaignFanOut

Note: If you have active schedules still waiting to fire, deleting the schedule group will cancel all pending sends.

IAM role for Amazon EventBridge Scheduler and Step Functions

The Step Functions state machine needs an execution role with permissions to create schedules, send email, and pass the role to the Amazon EventBridge Scheduler service. Amazon EventBridge Scheduler needs permissions to call SES. The following policy shows the combined permissions for both scenarios:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "AllowPassRoleToScheduler",
      "Effect": "Allow",
      "Action": "iam:PassRole",
      "Resource": "arn:aws:iam::<ACCOUNT_ID>:role/CampaignFanOutRole",
      "Condition": {
        "StringEquals": {
          "iam:PassedToService": "scheduler.amazonaws.com"
        }
      }
    },
    {
      "Sid": "AllowSESSend",
      "Effect": "Allow",
      "Action": [
        "ses:SendEmail",
        "ses:SendTemplatedEmail"
      ],
      "Resource": "arn:aws:ses:<REGION>:<ACCOUNT_ID>:identity/[email protected]"
    },
    {
      "Sid": "DistributedMapExecution",
      "Effect": "Allow",
      "Action": [
        "states:StartExecution",
        "states:DescribeExecution",
        "states:StopExecution"
      ],
      "Resource": [
        "arn:aws:states:<REGION>:<ACCOUNT_ID>:stateMachine:CampaignFanOut",
        "arn:aws:states:<REGION>:<ACCOUNT_ID>:execution:CampaignFanOut:*"
      ]
    },
    {
      "Sid": "ReadRecipientsBucket",
      "Effect": "Allow",
      "Action": [
        "s3:GetObject",
        "s3:ListBucket"
      ],
      "Resource": [
        "arn:aws:s3:::campaign-recipients-<ACCOUNT_ID>",
        "arn:aws:s3:::campaign-recipients-<ACCOUNT_ID>/*"
      ]
    },
    {
      "Sid": "CreateSchedules",
      "Effect": "Allow",
      "Action": "scheduler:CreateSchedule",
      "Resource": "arn:aws:scheduler:<REGION>:<ACCOUNT_ID>:schedule/campaign-*"
    },
    {
      "Sid": "CreateScheduleGroups",
      "Effect": "Allow",
      "Action": "scheduler:CreateScheduleGroup",
      "Resource": "arn:aws:scheduler:<REGION>:<ACCOUNT_ID>:schedule-group/campaign-*"
    },
    {
      "Sid": "PassRoleToScheduler",
      "Effect": "Allow",
      "Action": "iam:PassRole",
      "Resource": "arn:aws:iam::<ACCOUNT_ID>:role/SchedulerCampaignRole",
      "Condition": {
        "StringEquals": {
          "iam:PassedToService": "scheduler.amazonaws.com"
        }
      }
    }
  ]
}

This policy scopes the scheduler:CreateSchedule and scheduler:CreateScheduleGroup actions to resources prefixed with campaign-*, following least-privilege principles.

A condition restricts the iam:PassRole permission so that it can only pass the role to the Amazon EventBridge Scheduler service.

Conclusion

In this post, we walked through how to use Amazon EventBridge Scheduler to personalize email campaign delivery for each recipient. An email campaign system has two core problems: deciding what to send and deciding when to send it. Most teams over-engineer the “when” with polling infrastructure, batch jobs, and queue chains. Amazon EventBridge Scheduler collapses that into a single CreateSchedule API call per recipient.

To get started, explore Amazon EventBridge Scheduler on the AWS Management Console. Browse Serverless Land patterns for more than 20 Amazon EventBridge Scheduler patterns and other use cases beyond email campaigns.

Suggested tags: Amazon EventBridge, architecture, events, modernization, serverless.

Measuring and improving search quality with Amazon OpenSearch Service

Post Syndicated from Aruna Govindaraju original https://aws.amazon.com/blogs/big-data/measuring-and-improving-search-quality-with-amazon-opensearch-service/

Search is the front door of many applications, yet most teams struggle to answer a deceptively simple question: “Is my search actually returning relevant results?” Query logs tell you what users typed, not what they saw, what they selected, or why they left. When search feels broken, the culprit is rarely the engine. It’s the lack of deliberate signal collection, measurement, and a feedback loop to act on it.

You can close this gap on Amazon OpenSearch Service using User Behavior Insights (UBI), an open schema standard for capturing search behavior, and Search Relevance Workbench (SRW), a toolkit for measuring and evaluating search quality. Your application generates the UBI-formatted records. Together, UBI and SRW give you a repeatable framework: collect signals, turn them into relevance judgments, and validate every change before it ships.

In this post, we show you how to capture UBI data on an Amazon OpenSearch Service domain and use those signals to evaluate search quality. This is the first post in a two-part series. We build the foundation here, and Part 2 covers automating the workflow end to end.

The challenge: You can’t improve what you can’t measure

Consider a shopper searching for “handbag” on an ecommerce site. The catalog has 16 products (tote bags, duffel bags, laptop bags), but every title only says “bag.” The search returns zero results. Most shoppers leave. A patient one retries with “bag” and finds what they were looking for.

Your server log recorded that first query as a clean sub-second response: no error, no alert, no signal. What it missed entirely was a customer with purchase intent. That customer hit a vocabulary gap between how they search and how you write your catalog. Zoom out and apply this lens to misspelled queries, poor handling of long-tail searches, and abandoned sessions. The blind spot is larger than you think.

There’s a second problem: click signals are position biased. Users select the first result far more than the fifth, regardless of relevance, so raw click counts reflect where results appeared, not whether they deserved to be there. Any judgment derived from clicks must correct for this bias. We return to it when generating judgments.

Capturing behavioral data with UBI

UBI defines two indices. The ubi_queries index holds one record per executed query: the text the user typed, the full query that ran (filters and facets included), and the IDs of the documents returned. The ubi_events index holds every subsequent user action: impressions, hovers, clicks, add-to-carts, each stamped with the result position and the product’s business identifier (object_id). A shared query_id links every event back to the query that triggered it. Two additional identifiers complete the picture: client_id tracks the browser across visits, and session_id scopes events to a single visit.

A query record captures what the user asked and which document IDs the engine returned, including zero-result cases like the handbag search, which appears as a record with an empty result list. Here’s the shopper’s follow-up search for “bag”:

{
  "query_id": "1bf736d4-d673-4763-9193-4bc8a2282115",
  "client_id": "9a9968ac-664b-42d7-9a9e-96f412b5ab49",
  "user_query": "bag",
  "query": "{"multi_match": {"query": "bag", "fields": ["title", "description", "category", "brand"]}}",
  "query_response_hit_ids": [
    "3760170840499",
    "8400000000042"
  ],
  "timestamp": "2026-07-23T07:53:35.264Z",
  "application": "retail-shop"
}

The UBI queries schema reference documents the complete query schema, including the mandatory attributes.

The event record captures what the user did next. For each result rendered, emit an impression event. When the user selects a result, emit a click event. Here is the impression event for the first result of the bag search:

{
  "action_name": "impression",
  "query_id": "1bf736d4-d673-4763-9193-4bc8a2282115",
  "client_id": "9a9968ac-664b-42d7-9a9e-96f412b5ab49",
  "session_id": "0f2e6f2a-8f4e-4f60-9f6e-2a1b3c4d5e6f",
  "user_query": "bag",
  "timestamp": "2026-07-23T07:53:41.112Z",
  "event_attributes": {
    "position": {
      "ordinal": 1
    },
    "object": {
      "object_id": "3760170840499",
      "object_id_field": "object_id"
    }
  }
}

event_attributes also accepts custom fields of your own alongside the standard position and object structures. The action_name attribute is critical: The judgment model you use later consumes only impression and click events. Treat a paginated results page as the same logical query: reuse the query_id and record absolute positions. The UBI events schema reference documents the complete event schema.

Collecting UBI data on Amazon OpenSearch Service

Behavioral data (what results ranked, what users saw, what they selected) exists only in the application layer. Your application owns the records, and Amazon OpenSearch Ingestion (OSI), a fully managed, serverless data collector powered by Data Prepper, provides the managed delivery path. Your application sends the records as SigV4-signed HTTP POST requests to the OSI pipeline endpoints. Route browser events through your backend for signing. One thing to understand before you write any code: Your application generates and owns the query_id attribute. The application creates the ID when it runs a search and stamps it on every subsequent event the user produces, until the user issues a new search or the session ends.

Prerequisites

To follow along, you need an Amazon OpenSearch Service domain running OpenSearch 3.5 or later with the OpenSearch UI application, permissions to create OpenSearch Ingestion pipelines with an AWS Identity and Access Management (IAM) pipeline role, and a search application you can instrument to emit behavioral records.

Create the UBI indices

Before you start collecting user metrics, you need the two indices in place with the right mappings. Field types matter here: query_id as keyword supports exact joins between queries and events, timestamp as date supports time-range queries, and event_attributes as dynamic means you can extend events with custom fields without schema changes.

Create ubi_queries first in Dev Tools. It holds the query-side records. We abbreviated the mappings here. Refer to the published queries-mapping.json file for the complete version:

PUT ubi_queries
{
  "mappings": {
    "properties": {
      "query_id": { "type": "keyword" },
      "client_id": { "type": "keyword" },
      "user_query": { "type": "keyword" },
      "query_response_hit_ids": { "type": "keyword" },
      "timestamp": {
        "type": "date",
        "format": "strict_date_time"
      },
      "application": { "type": "keyword" }
    }
  }
}

Then create ubi_events. It holds every user action that follows (refer to the full events-mapping.json file):

PUT ubi_events
{
  "mappings": {
    "properties": {
      "query_id": { "type": "keyword", "ignore_above": 100 },
      "action_name": { "type": "keyword", "ignore_above": 100 },
      "client_id": { "type": "keyword", "ignore_above": 100 },
      "session_id": { "type": "keyword", "ignore_above": 100 },
      "user_query": { "type": "keyword" },
      "timestamp": {
        "type": "date",
        "format": "strict_date_time"
      },
      "event_attributes": {
        "dynamic": true,
        "properties": {
          "position": {
            "properties": {
              "ordinal": { "type": "integer" }
            }
          },
          "object": {
            "properties": {
              "object_id": { "type": "keyword" },
              "object_id_field": { "type": "keyword" }
            }
          }
        }
      }
    }
  }
}

With both indices created, the next step is routing data into them. You can deliver UBI data to your domain in several ways. This post uses OSI pipelines, shown end to end in the diagram that follows the setup.

Set up the OSI pipelines

Create two OSI pipelines: one for queries and another for events. Each pipeline exposes an HTTP source endpoint that your application writes to (shown on each pipeline’s console page) and sinks data to the corresponding index. The following configuration defines the events pipeline:

version: '2'
ubi-events:
  source:
    http:
      path: /ubi/events
      max_request_length: 10mb
  processor:
    - date:
        from_time_received: true
  sink:
    - opensearch:
        hosts: ["https://<domain-endpoint>"]
        aws:
          serverless: false
          region: <region>
          sts_role_arn: <pipeline-role-arn>
        index_type: custom
        index: ubi_events
    - s3:
        aws:
          region: <region>
          sts_role_arn: <pipeline-role-arn>
        object_key:
          path_prefix: 'ubi_events/%{yyyy}/%{MM}/%{dd}'
        bucket: <bucket-name>
        threshold:
          maximum_size: 50mb
          event_collect_timeout: 60s
        codec:
          ndjson:

Note: the queries pipeline follows the same pattern, with /ubi/queries as the path and ubi_queries as the sink index and S3 prefix. Create the pipeline role yourself or let OpenSearch Ingestion create it. If your domain uses fine-grained access control, also map the pipeline role to a backend role so the domain accepts the pipeline’s writes. Refer to the tutorial Collecting UBI-formatted data in Amazon OpenSearch Service for detailed steps.

With the pipelines running, your application can start sending data. The following diagram illustrates the end-to-end flow:

UBI collection flow from the search application through OpenSearch Ingestion into the ubi_queries and ubi_events indices

Figure 1: The UBI collection pattern on Amazon OpenSearch Service

The workflow consists of the following steps:

  1. Users interact with your search application.
  2. The application sends signed query records to the OSI HTTP endpoint.
  3. OSI writes queries to the ubi_queries index.
  4. Users interact with the results, viewing and selecting documents.
  5. The application sends signed event records, carrying the same query_id, to the OSI HTTP endpoint.
  6. OSI writes events to the ubi_events index.
  7. Optionally, both pipelines archive records to Amazon Simple Storage Service (Amazon S3).
  8. Search Relevance Workbench (OpenSearch UI) works with the collected data in the ubi_queries and ubi_events indices.

Note: if you’re already collecting site analytics through an existing third-party tool, you don’t need to replace it. Map your search-related events (queries, clicks, and conversions) into the UBI schema and store them in OpenSearch. That’s enough to unlock the out-of-the-box evaluation framework, implicit judgment generation, and the full SRW metrics pipeline, without defining a single custom metric from scratch.

Visualize the data collected

After the UBI behavior metrics start to trickle in, you can review the data in the Discover tab on the OpenSearch UI dashboard. Filtering ubi_queries for empty result lists ranks your vocabulary gaps. You can also visualize the data collected through the sample User Behavior Insights (UBI) dashboards in OpenSearch.

OpenSearch Discover view of UBI records for the zero-result handbag query and the follow-up bag query

Figure 2: UBI records in Discover, showing the zero-result handbag query and the follow-up bag query with its impressions and pagination events

With data flowing into your indices, keep these things in mind as you scale to production:

  • Keep telemetry off the search critical path – Queue records and forward them asynchronously. Losing a fraction of behavioral data is statistically harmless. Blocking users isn’t.
  • Manage volume deliberately – Batch impression events, and if you sample, sample whole queries rather than individual events to preserve the click-through ratios that drive judgments.
  • Isolate analytical load for larger deployments – Route pipelines to a separate analysis domain with the same engine version, mappings, and analyzers as production. This keeps behavioral writes from touching live search latency.
  • Plan for retention and integrity – Register the UBI mappings as an index template and apply an Index State Management (ISM) retention policy as your indices grow. You should validate and rate-limit the event write path, and cover query text and client identifiers with your data retention policy.

Evaluating search quality with Search Relevance Workbench

With ubi_queries and ubi_events collecting data, you now have the signals needed to evaluate search quality. Search Relevance Workbench, generally available in the OpenSearch UI from Amazon OpenSearch Service 3.5, turns those signals into structured experiments: comparing query configurations, scoring results against relevance judgments, and surfacing metrics that guide iterative tuning.

The Search Relevance Workbench home screen in the OpenSearch UI

Figure 3: Search Relevance Workbench in the OpenSearch UI

SRW experiments rely on three components. You set them up once, then reuse them across every experiment you run: a query set (the fixed queries you evaluate against), search configurations (the query structures you want to compare), and a judgment list (the relevance ground truth). The following sections walk through each one.

Step 1: Create a query set

A query set is the fixed collection of queries you evaluate against. Keeping it fixed makes results comparable across experiments. Effective query sets reflect real traffic, not intuition. You can seed one from your top queries, a random sample, or a hand-picked mix that includes long-tail and low-performing queries. Alternatively, SRW can sample directly from ubi_queries using Probability-Proportional-to-Size (PPS) sampling, which selects queries in proportion to how often users issue them. This approach represents frequent queries like “bag”, so your metrics reflect search quality as users experience it.

Query set creation screen sampling queries from real traffic in the ubi_queries index

Figure 4: Creating a query set sampled from real traffic in ubi_queries

Step 2: Define search configurations

A search configuration defines how a search executes: the index, the query structure, and a %SearchText% placeholder that SRW replaces with each query in your set. Creating two configurations and running them against the same query set and judgment list is how you validate a change before any user sees it.

As an example, here we define two configurations: a baseline multi_match query (retail_query) and a variant that boosts title matches (retail_boosted_query), so we can measure whether the boost actually helps ranking.

retail_query retail_boosted_query
{
  "query": {
    "multi_match": {
      "query": "%SearchText%",
      "fields": [
        "title",
        "description",
        "category",
        "brand"
      ]
    }
  }
}
{
  "query": {
    "multi_match": {
      "query": "%SearchText%",
      "fields": [
        "title^2",
        "description",
        "category",
        "brand"
      ]
    }
  }
}

Configurations go beyond query variants: a candidate can be an entirely different retrieval strategy, like hybrid search combining keyword and neural retrieval. You can use judgments to rate query-document pairs independently of your retrieval approach. You can test a semantic or hybrid approach offline against your existing traffic before shipping it.

Step 3: Create the judgment list

A judgment is a relevance rating for a query-document pair: the ground truth that quality metrics measure against. You can create judgments that are explicit (from stakeholders or a large language model acting as judge), imported, or implicit (derived from behavior). Here we use implicit judgments derived from UBI selection behavior, scored using the Clicks Over Expected Clicks (COEC) model. The COEC model helps correct position bias by comparing each document’s actual click rate against the expected rate for its rank position. Documents that outperform their position score as relevant. Those that users select because they ranked first score near average.

Judgment list creation screen with the Implicit click-based type and COEC click model selected

Figure 5: Creating an implicit judgment list with the Implicit (Click based) type and the COEC click model

Three things to get right before you run experiments:

  1. object_id in your events must match the document _id from your product catalog. The search configurations you define return this _id, which lets SRW join judgments to results.
  2. Implicit judgments are statistical. They need volume and query coverage. As a working rule of thumb, aim for hundreds to thousands of real sessions per query to separate signal from noise.
  3. Max Rank controls how deep in the result list events count. If users paginate, set it beyond a single page. We use 20 here.

Step 4: Run experiments

This post uses three SRW capabilities: Query Analysis, Query Set Comparison, and Search Evaluation. Query Analysis is a quick eyeball check: compare two configurations side by side for a specific query to see exactly what changed and why the metrics moved. The other two answer harder questions with numbers: how good a configuration is, and how two configurations compare against real relevance signals.

Query Set Comparison (also called pairwise comparison) takes two configurations and computes ranking similarity. Jaccard overlap measures how much the two result lists share, while Rank-Biased Overlap (RBO) weights agreement at the top of the list more heavily. Near-identical scores mean the change will barely register with users. Low overlap means a real ranking shift worth reviewing carefully before shipping. In this run, the two configurations score 0.93 Jaccard and 0.92 RBO, a modest but real shift. SRW cannot score zero-result queries like “handbag”: They show zero similarity in a comparison and Failed in an evaluation, a signal they need a different fix than ranking adjustments.

Query Set Comparison results showing Jaccard and Rank-Biased Overlap scores for the two configurations

Figure 6: Query Set Comparison showing Jaccard and Rank-Biased Overlap between the two configurations

Search Evaluation (also called pointwise evaluation) scores one configuration against your query set and judgment list across four metrics, each computed over the top k results (k=10 by default):

Metric What it measures What it tells you
Coverage@k Proportion of returned documents that have judgments How much to trust the other three metrics. Low Coverage means many results were never judged
Precision@k Fraction of the top k results that are relevant How many irrelevant results appear on the first page
MAP@k (Mean Average Precision) Precision averaged across ranks, rewarding relevant documents placed early Whether relevant results appear early, even when Precision ties
NDCG@k (Normalized Discounted Cumulative Gain) Graded judgment values, discounted by position (rank 1 counts more than rank 9) Whether the best results appear first. The primary comparison metric

Each pointwise experiment evaluates one configuration. To compare candidates, run one experiment per configuration and compare the results. In this run, the baseline (retail_query) scores Coverage@10 of 1.0, Precision@10 of 1.0, MAP@10 of 0.95, and NDCG@10 of 0.93, with the zero-result “handbag” query showing as Failed in the per-query detail.

Search evaluation results showing Coverage, Precision, MAP, and NDCG at 10 with per-query detail

Figure 7: Search evaluation results for one configuration: Coverage, Precision, MAP, and NDCG at 10, with per-query detail

From measurement to improvement

The preceding experiments are the harness. The following are common levers to test with it. Express each as a new search configuration, evaluate it against the same query set and judgment list, and adopt it only if the metrics move:

  • Synonyms – One option for addressing known vocabulary gaps is to build synonyms. A search-time synonym token filter treats “handbag” and “bag” as equivalent, and with Amazon OpenSearch Service, you can hot deploy custom synonym packages without reindexing.
  • Field weights – Adjust the fields and boosts in a multi_match query, like the title^2 variant tested earlier.
  • Semantic retrieval – A hybrid query combines keyword and neural scores, addressing vocabulary mismatch as a class rather than term by term. Judgments evaluate it offline exactly like a lexical candidate.
  • Reranking – A rerank processor in a search pipeline reorders the top results using a cross-encoder model.

Clean up

To avoid future charges, delete the resources you created for this walkthrough:

  • Delete the two OpenSearch Ingestion pipelines. To reuse them later, stop them instead. A stopped pipeline keeps its configuration and incurs no OpenSearch Compute Unit (OCU) hour charges.
  • If you configured the optional Amazon S3 archive, delete the archived objects (or the bucket).
  • If you keep the domain, optionally delete the ubi_queries and ubi_events indices and the query sets, judgment lists, and experiments you created. These live on the domain and incur no separate charges.
  • If you created the domain specifically for this post, delete it to remove everything, including the resources in the previous step. Deleting a domain is irreversible. Don’t delete a domain that serves other workloads.

Conclusion

UBI collects the evidence, COEC turns it into judgments, and SRW experiments deliver the verdict: Coverage, Precision, MAP, and NDCG in place of guesswork. Ship the winning configuration, keep collecting, and the next round of judgments shows whether the improvement holds with real behavior. Where there used to be an opinion, there is now a number.

Everything here follows a repeatable pattern, and repeatable patterns lend themselves to automation. Part 2 walks through the Search Relevance Agent, available through the AI Assistant chat (the Ask AI button) in the OpenSearch UI. The agent analyzes your UBI signals, generates tuning hypotheses, and validates them offline before recommending changes. The pipeline you built in this post is the foundation. Stay tuned for Part 2.

To go deeper on the evaluation features, refer to the Search Relevance Workbench documentation.


About the authors

Aruna Govindaraju

Aruna Govindaraju

Aruna is an Amazon OpenSearch Specialist Solutions Architect and has worked with many commercial and open source search engines. She is passionate about search, relevancy, and user experience. Her expertise with correlating end-user signals with search engine behavior has helped many customers improve their search experience.

Sean Bjurstrom

Sean Bjurstrom

Sean is an Enterprise Support Lead in ISV accounts at Amazon Web Services, where he specializes in Analytics technologies and draws on his background in consulting to support customers on their analytics and cloud journeys. Sean is passionate about helping businesses harness the power of data to drive innovation and growth. Outside of work, he enjoys running and has participated in several marathons.

Utkarsh Agarwal

Utkarsh Agarwal

Utkarsh is a Cloud Support Engineer in the Support Engineering team at AWS. He provides guidance and technical assistance to customers, helping them build scalable, highly available, and secure solutions in the AWS Cloud. In his free time, he enjoys watching movies, TV series, and, of course, cricket! Lately, he has also been attempting to master foosball.

Automate IAM Identity Center governance with continuous discovery and reporting

Post Syndicated from Jonathan Nguyen original https://aws.amazon.com/blogs/security/automate-iam-identity-center-governance-with-continuous-discovery-and-reporting/

AWS IAM Identity Center integrates with external identity provider (IdP) to provide customers with a centralized authentication and authorization solution for AWS resources across AWS Organizations. AWS continues to invest into IAM Identity Center with a growing number of AWS services that natively integrate with IAM Identity Center. As your AWS organization scales, maintaining visibility into who has access to which applications and enforcing governance policies across accounts and Regions becomes increasingly complex. Identity Center helps address this by centralizing authentication and authorization for AWS resources across your organization, integrating with your external identity provider and a growing number of AWS services. However, as adoption scales, tracking access assignments and enforcing governance policies consistently becomes its own challenge.

This blog post focuses on planning your integration between an identity provider and IAM Identity Center for managed applications in your organization. We also walk through deploying and using an automated Identity Center discovery and reporting sample solution to help answer the governance and security questions:

  1. Which users or groups have access to which AWS applications?
  2. Who last accessed a specific AWS application and when?
  3. Which users and groups are assigned to which IAM Identity Center applications across organization and AWS Regions?
  4. How can you quickly generate reports to assist with compliance audits or security reviews?

The sample solution will identify associated AWS applications and the corresponding user and group assignments for the IAM Identity Center instances within your organization. The output is stored in a queryable format and generates CSV files for downstream analysis or reporting.

Plan identity governance for Identity Center application assignments

There are four key areas to start on when planning how to manage delegation and provisioning access across IAM Identity Center managed AWS applications. Bring together key stakeholders across security, governance, application, and business teams to make sure the implementation and integration will fit into the overall identity governance strategy.

  1. Who can provision managed AWS applications: You can implement the IAM restrictions for creation of new AWS resources within AWS accounts in your organization. For example, if you restrict provisioning into a production AWS account to only infrastructure as code (IaC) IAM roles, you would continue implementing restrictions using AWS identity policies, service control policies (SCP), resource control policies (RCP), or IaC policy evaluation tools like Open Policy Agent (OPA) or Checkov.
  2. Who manages user and group assignments: The managed application administrator handles authorization to managed applications within an AWS account. It’s recommended to clearly define roles and responsibilities across the workflow. You would have an IaC pipeline manage the integrated AWS resource provisioning with IAM Identity Center, then another workflow to allow requests to manage user and group membership for the managed application.
  3. How authentication flows from the IdP to AWS resources: Users will authenticate into Identity Center, then be authorized to access AWS managed applications. From there, they will be authorized to access the associated AWS service and resources tied to the managed application. Depending on the AWS service, the associated downstream resources might have their own IAM principals that the users can access.
  4. Mapping IdP identities to AWS resource access: There needs to be a link for workforce users and groups in your IdP, to Identity Center managed applications, and to downstream resources and permissions. Identifying the relationship will help you understand access within your AWS environment. Trusted identity propagation (TIP) is an additional feature of Identity Center that provides an end to end trail of the identity to the downstream service.

Create and manage an Identity Center application assignment lifeycle

As a security best practice, you should enable delegated administration when managing Identity Center within an AWS organization instances.

After you have IAM Identity Center set up within an organization instance, your member AWS accounts can start creating associated AWS resources. Within each member AWS account, the IAM principals that provision AWS resources will need two types of service-specific IAM permissions:

  • The first type of IAM permissions will be specific to the AWS service you want to provision. For example, to create an Amazon SageMaker AI domain, you would need the same IAM permissions to create the SageMaker AI domain and the downstream AWS resources SageMaker AI might use.
  • The second type of IAM permissions is specific to IAM Identity Center. The IAM principal used to create the resource, in this example SageMaker AI, will also need permissions to manage applications within the Identity Center instance.
{
	"Version": "2012-10-17",
	"Statement":
	[
		{
            "Effect": "Allow",
            "Action": [
                "sso:CreateManagedApplicationInstance",
                "sso:GetManagedApplicationInstance",
                "sso:DeleteManagedApplicationInstance",
                "sso:DescribeRegisteredRegions"
            ],
            "Resource": "*"
        },
        {
            "Effect": "Allow",
            "Action": [
                "sso:CreateApplication",
                "sso:DescribeApplication",
                "sso:DeleteApplication",
                "sso:PutApplicationGrant",
                "sso:PutApplicationAuthenticationMethod",
                "sso:PutApplicationAccessScope"
            ],
            "Resource": 
            [
                "arn:aws:sso::<INSERT-ACCOUNT-ID>:application/ssoins-<INSERT-INSTANCE-ID>/apl-*"
            ]
        }
    ]
}

IAM Identity Center application Amazon Resource Names (ARNs) follow a different standard naming convention that isn’t based on the original resource name that was provided during resource creation. For example, when a user creates an Amazon Simple Storage Service (Amazon S3) bucket and sets a specific bucket name, that bucket name is included in the ARN: arn:[partition]:s3:::[bucket-name]. Identity Center application ARNs use unique identifiers (GUIDs) generated at creation time.

Manage access for an Identity Center application

After the IAM Identity Center application is created, you will need to manage access to the Identity Center application and associated AWS resources. To continue with the SageMaker AI domain example, after the domain is created, an authorized IAM principal will need to assign Identity Center users or groups from the Identity Center instance to the domain. For Identity Center, you will need two types of Identity Center IAM permissions.

The first type of IAM permissions is used to list IAM Identity Center users and groups within the Identity Center instance. This is needed to read and select specific IAM users or groups to assign to an Identity Center application.

{
    "Version": "2012-10-17",
    "Statement":
    [
        {
            "Sid": "ListIdentityCenterUsers",
            "Effect": "Allow",
            "Action":
            [
                "identitystore:ListUsers",
                "identitystore:DescribeUser",
                "identitystore:ListGroups",
                "identitystore:DescribeGroup",
                "identitystore:ListGroupMemberships"
            ],
            "Resource": "*"
        }
    ]
}

Although IAM Identity Center users and groups have a GUID, the GUIDs aren’t clearly linked to the resource friendly names. For example, a group name could be Read-Only and the resource GUID could be 1234567890-abcdef12-3456-7890-abcd-ef1234567890 in the identity store. Additionally, the IAM actions to list users or groups require the AllUsers or AllGroups parameter. Because List actions require access to users and groups, a restrictive IAM policy can’t be used to prevent IAM principals from seeing a subset of users or groups within the identity store. The second type of IAM permission is used to create and manage application assignments for the Identity Center application within the Identity Center instance.

{
    "Version": "2012-10-17",
    "Statement":
    [
        {
            "Sid": "ManageApplicationAssignments",
            "Effect": "Allow",
            "Action": 
            [
                "sso:CreateApplicationAssignment",
                "sso:DeleteApplicationAssignment",
                "sso:ListApplicationAssignments",
                "sso:PutApplicationAssignmentConfiguration"
            ],
            "Resource":
            [
                "arn:aws:sso::<INSERT-ACCOUNT-ID>:application/ssoins-<INSERT-INSTANCE-ID>/apl-*"
            ]
        }
    ]
}

Because the IAM Identity Center application ARN is created using a unique application ID during creation, it’s not recommended to implement an IAM policy restricting authorized IAM principals to manage specific Identity Center applications. For example, to limit the application assignments to only a specific set of applications, you would need to:

  1. Create the AWS resource with IAM Identity Center as the authentication mechanism
  2. Query the Identity Center application ARN for the associated AWS resource
  3. Identify the IAM principal that will be used for application assignments
  4. Create or update an IAM policy associated to that IAM principal to allow application assignments for that specific application
  5. Create or update an SCP to restrict application assignment to that specific IAM principal

In lieu of implementing resource restrictions within identity policies, you should limit management of IAM Identity Center application and application assignments to a limited number of authorized IAM principals. In addition, it is recommended to implement detective and reactive capabilities to manage Identity Center application assignments.

Plan your naming conventions and automation strategy

IAM Identity Center provides several APIs to capture information about your AWS organization instances, applications, and assignments. Before implementing automation or guardrails, you should develop a methodical approach and understand what outcome you’re working backwards from. Start by defining naming conventions and deciding what parts of the workflow you want to centralize.

  1. Determine a naming convention for groups within your IdP: For example: AWS_<ACCT#>_<AWS_Service>_<LOB>_<ENV>_<AppName>. The IdP group name would look like: AWS_123412341234_SageMaker_Data_PROD_GTLabel.
  2. Define the naming convention for AWS resources for your Identity Center integrated applications: For example: <AWS_Service>_<LOB>_<AppName>. The AWS resource name would look like: SageMaker_Data_GTLabel.
  3. Define the naming convention for Identity Center application names: For example: <AWS_Service>_<LOB>_<ENV>_<AppName>. The Identity Center application name would look like: SageMaker_Data_PROD_GTLabel.
  4. Decide on the restrictions that you want to implement within your AWS environment. Depending on your enterprise’s security standard, you can implement specific restrictions based on mapping of a similar combination of ENV (environment), AWS service, LOB (line of business), or application name.
  5. Choose the portions of the application workflow that you want to centralize. This could include creating the application, making application assignments, or remediating issues.

As more configurations and permissions are centralized, additional overhead and bottlenecks can be introduced. It’s important to find the right balance for your enterprise. For example, if you centralize application assignments, each application team will need to submit a request to modify assignments that will be reviewed by a centralized team and could result in a delayed response. Conversely, if each application team handles their own assignments, there’s a risk that application assignments won’t align to enterprise security standards.

By understanding your goals and how you want to reach them, you can tailor the sample solution accordingly. Getting alignment on this requires planning and coordination across multiple teams within your organization. When thinking about more customized authorization logic—such as using provisioned AWS resource metadata—you should review how the specific AWS service integrates with IAM Identity Center managed applications. For example, if you want to find the Identity Center application ARN for a specific AWS resource, such as a SageMaker AI domain, use the following approach. A reverse lookup is necessary because AWS services create Identity Center applications with GUID-based ARNs that aren’t easily discoverable.

#!/bin/bash

DOMAIN_ID="d-xxxxxxxxxxxx"

REGION="xx-xxxx-x"

# Step 1: Get SageMaker domain details
echo "=== SageMaker Domain Details ==="
DOMAIN_INFO=$(aws sagemaker describe-domain \
--region $REGION \
--domain-id $DOMAIN_ID)

# Step 2: Extract Identity Center application ARN
SSO_APP_ARN=$(echo $DOMAIN_INFO | jq -r '.SingleSignOnApplicationArn')
echo "Identity Center App ARN: $SSO_APP_ARN"

IAM Identity Center automation sample solutions

The sample-iam-idc-application-discovery-reporting solution hosted on GitHub consists of two separate AWS CDK stacks:

  1. IAM Identity Center governance reporting stack (/identity-center-reporting directory) – Provides automated discovery and report generation (using CSV files)
  2. IAM Identity Center remediation stack (/identity-center-remediation directory) – Provides real-time enforcement and notifications

The recommendation is to deploy the reporting stack first to establish baseline visibility, then deploy the remediation stack for enforcement.

The following diagram depicts that IAM Identity Center governance architecture.

The reporting sample deploys the following resources:

  1. Amazon EventBridge – Rule invokes the discovery workflow daily at 2:00 AM UTC (configurable)
  2. AWS Step Functions – Orchestrates the multi-stage discovery workflow across instances, applications,and assignments
  3. AWS Lambda – Takes the following actions:
    1. Discovers IAM Identity Center instances across the organization and member accounts
    2. Application discovery that enumerates the applications configured in each Identity Center instance
    3. Assignment discovery maps users and groups to applications, resolving friendly names from the Identity Store
  4. Amazon DynamoDB – Stores the discovered instances, applications, and assignments, encrypted with an AWS Key Management Service (AWS KMS) customer-managed key
  5. Amazon API Gateway – Provides an IAM-authenticated REST API for a Lambda function to generate and export reports as CSV files
  6. Amazon S3 – Stores the encrypted CSV file exports, with lifecycle policies and time-limited Amazon S3 presigned download URLs

Deploy the IAM Identity Center reporting sample

The following procedure deploys the automated discovery and reporting infrastructure using AWS Cloud Development Kit (AWS CDK). Make sure you have the following prerequisites in place, then continue with the steps to set up the solution.

Prerequisites

You need to have the following to test the solution in this post.

  1. An AWS organization with an IAM Identity Center organization instance with delegated administrator access configured
  2. IAM Identity Center configured with at least one instance
  3. AWS Command Line Interface (AWS CLI) configured with appropriate credentials
  4. Python 3.12 & Node.js 18 or later installed for CDK deployment

To deploy the IAM Identity Center reporting solution, run the following commands:

  1. Clone the solution repository:
    git clone https://github.com/aws-samples/sample-iam-idc-application-discovery-reporting
    cd identity-center-reporting

  2. Install dependencies:
    python3.12 -m venv .venv && source .venv/bin/activate
    pip install -r requirements.txt

  3. Bootstrap the CDK (if not already done):
    cdk bootstrap aws://<INSERT-ACCOUNT-ID>/<INSERT-REGION>

  4. Deploy the sample solution:
    export IDC_EXTERNAL_ID="$(uuidgen)" # alternatively you can set this value — member-account roles need the same value
    
    cdk deploy --parameters AllowedIpRange=10.0.0.0/8 --parameters CrossAccountExternalId="$IDC_EXTERNAL_ID"

    Note: AllowedIPRange is optional but recommended as a security best practice. The parameter will add a network restriction to download the Amazon S3 presigned URL export.

  5. Optional: For AWS account-level Identity Center instance discovery, a cross-account IAM role is required.
    python scripts/deploy-cross-account-roles.py --external-id "$IDC_EXTERNAL_ID"

Figure 2: Successful AWS CDK deployment of the reporting stack

Figure 2: Successful AWS CDK deployment of the reporting stack

After the stack is successfully deployed, obtain the CDK output values for the API Gateway URL and S3 bucket name. If using a command line to deploy, these values will be displayed after the stack successfully deploys. It can also be found in the AWS Management Console as AWS CloudFormation stack output. The output will be used for generating reports in the following sections.

Note that this stack is for the reporting stack only. Reactive monitoring and deployment are described in the next section.

After the reporting stack is successfully deployed, the automation will run on a daily schedule. The first discovery run executes immediately after deployment. You can monitor discovery execution history and detailed logs through the the AWS Step Functions console. Review the detailed Lambda function logs in Amazon CloudWatch Logs. Query discovered instances, applications, and assignments through the DynamoDB console for one-time analysis.

Generate reports for Identity Center application assignments

To generate on-demand reports as CSV files from the REST API:

  1. Set env variables for Sigv4 authentication
      export AWS_REGION="<REPLACE-REGION>"
      export API_ID="<REPLACE-API-ID>"
      eval "$(aws configure export-credentials --profile "<YOUR-PROFILE>" --format env)"

  2. Export applications
      curl -sS --fail-with-body \
        --aws-sigv4 "aws:amz:${AWS_REGION}:execute-api" \
        --user "${AWS_ACCESS_KEY_ID}:${AWS_SECRET_ACCESS_KEY}" \
        --header "x-amz-security-token: ${AWS_SESSION_TOKEN}" \
        "https://${API_ID}.execute-api.${AWS_REGION}.amazonaws.com/prod/export/applications" \
        -o applications.json

  3. Export assignments with user and group names
      curl -sS --fail-with-body \
        --aws-sigv4 "aws:amz:${AWS_REGION}:execute-api" \
        --user "${AWS_ACCESS_KEY_ID}:${AWS_SECRET_ACCESS_KEY}" \
        --header "x-amz-security-token: ${AWS_SESSION_TOKEN}" \
        "https://${API_ID}.execute-api.${AWS_REGION}.amazonaws.com/prod/export/assignments" \
        -o assignments.json

  4. The API returns a JSON response with a presigned Amazon S3 URL that’s valid for 15 minutes:
    {
        "message": "CSV export generated successfully",
        "download_url": "https://<bucket>.s3.amazonaws.com/exports/applications/2026/06/22/applications_export_20260722_184538.csv?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Expires=900&...",
        "filename": "applications_export_20260722_184538.csv",
        "s3_key": "exports/applications/2026/07/22/applications_export_20260722_184538.csv",
        "file_size_bytes": 6514,
        "export_type": "applications",
        "generated_at": "2026-07-22T18:45:38Z",
        "expires_at": "2026-07-22T19:00:38Z",
        "request_id": "a1b2c3d4-...."
    }

    {
        "message": "CSV export generated successfully",
        "download_url": "https://<bucket>.s3.amazonaws.com/exports/applications/2026/07/22/applications_export_20260722_184538.csv?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Expires=900&...",
        "filename": "applications_export_20260722_184538.csv",
        "s3_key": "exports/applications/2026/07/22/applications_export_20260722_184538.csv",
        "file_size_bytes": 6514,
        "export_type": "applications",
        "generated_at": "2026-07-22T18:45:38Z",
        "expires_at": "2026-07-22T19:00:38Z",
        "request_id": "a1b2c3d4-...."
    }

The generated CSV files include enriched data with friendly names:

Instance ARN Account ID Application name Principal type Principal name Status
arn:aws:sso:::instance/… 123456789012 SageMaker_PROD GROUP Engineering-Team-Dev ACTIVE
arn:aws:sso:::instance/… 123456789012 OpenSearch_PROD USER [email protected] ACTIVE

You can use the generated CSV files to help identify anomalies or non-compliant assignments, such as:

  1. Each PROD application should only have GROUP assignments. OpenSearch_PROD has a USER principal type and so is non-compliant.
  2. Each PROD application should only allow PROD groups assigned. SageMaker_PROD has a DEV group name (Engineering-Team-Dev) assigned and so is non-compliant.

Based on the testing and analysis of the output from the IAM Identity Center governance reporting sample solution, it’s important to start thinking about what restrictions to put in place for application assignments. It’s also important to conduct this exercise before taking action within the Identity Center remediation sample solution in the next section.

IAM Identity Center remediation

The following diagram shows the IAM Identity Center remediation architecture.

The IAM Identity Center remediation sample solution deploys the following resources:

  1. Amazon EventBridge – Matches IAM Identity Center assignment and profile events from CloudTrail (sso.amazonaws.com) and invokes the monitor function across the following IAM actions:
    1. CreateApplicationAssignment
    2. DeleteApplicationAssignment
    3. PutApplicationAssignmentConfiguration
    4. AssociateProfile
    5. DisassociateProfile
    6. CreateProfile
    7. UpdateProfile
    8. DeleteProfile
  2. Lambda – Resolves the application and group names, validates the assignment against your naming convention, and notifies or remediates based on the configured mode
  3. Amazon Simple Notification Service (Amazon SNS) – Publishes alerts for non-compliant assignments to subscribers (for example, email)
  4. Amazon Simple Queue Service (Amazon SQS) – Captures events the Lambda function fails to process for later inspection
  5. AWS KMS – Customer-managed key to encrypt the Lambda environment variables, CloudWatch logs, SNS topic, and dead-letter queue
  6. Amazon CloudWatch – Log group stores the function’s structured, encrypted logs as an audit trail

Flexible naming policies support regex-based pattern matching for specific organizational requirements. The automation actions are logged to CloudWatch with structured JSON for additional analysis and reporting.

The following procedure deploys the remediation infrastructure using AWS Cloud Development Kit (AWS CDK). Make sure you have the following prerequisites in place, then continue with the steps to set up the solution.

Prerequisites

You need the following to run the remediation solution:

  1. An AWS organization with an IAM Identity Center organization instance with delegated administrator access configured
  2. IAM Identity Center configured with at least one instance and an IdP
  3. Access to create groups within the integrated IdP
  4. AWS Command Line Interface (AWS CLI) configured with appropriate credentials
  5. Python 3.12 & Node.js 18 or later installed for CDK deployment

Provide your IAM instance ARN and the account ID where IAM Identity Center is administered:

git clone https://github.com/aws-samples/sample-iam-idc-application-discovery-reporting # only needed if you did not clone in the previous reporting section
cd identity-center-remediation
cdk deploy --context enableAutoDeletion=false --parameters IdentityCenterInstanceArn=arn:aws:sso:::instance/ssoins-<INSERT-ORG-INSTANCE-ID> --parameters ManagementAccountId=<INSERT-MANAGEMENT-ACCOUNT>

Note: If you don’t pass a parameter for GroupNameRegex, the default action of the sample solution is to verify the group name appears as a whole word in the application name: Case-insensitive, splitting on -, _, and spaces, so ReadOnly matches sagemaker_readonly but read does not. If different validation is needed, the sample can be deployed with the regex value for GroupNameRegex.

After the solution is deployed, we will walk through testing both a compliant and non-compliant application assignment.

Gather Identity Center and application information

For this blog, we have already created two groups within the IdP that is integrated into an IAM Identity Center instance. We also already created two applications within Identity Center instance to use. Next, we’ll need to gather information specific to the environment to run through each example.

  1. Obtain the IAM Identity Center instance ARN and set the value.
    INSTANCE_ARN=$(aws sso-admin list-instances --region <REPLACE-REGION> --query "Instances[0].InstanceArn" --output text)
    
    echo "$INSTANCE_ARN"

  2. Obtain the IAM Identity Center identity store ID

    IDENTITY_STORE_ID=$(aws sso-admin list-instances --region <REPLACE-REGION> --query "Instances[0].IdentityStoreId" --output text)
    
    echo "$IDENTITY_STORE_ID"

  3. Get existing groups in IAM Identity Center
    aws identitystore list-groups --identity-store-id $IDENTITY_STORE_ID --query "Groups[].{Name:DisplayName,Id:GroupId}" --output table

  4. Get existing enabled applications in IAM Identity Center
    aws sso-admin list-applications --instance-arn $INSTANCE_ARN --query "Applications[?Status=='ENABLED'].{Name:Name,ARN:ApplicationArn}" --output table

After you have the output for IAM Identity Center groups and applications, select two groups and one application that you want to test with. You will need to set additional variables for each group GUID and application ARN. In this example, I select the following two groups (ReadOnly and Developer) for testing and set the environment variables using export:

  1. Group #1 Name: ReadOnly
    export GRP_READONLY=abc12345-1234-1234-1234-abcdef123456
    export GRP_DEVELOPER=abc12345-1234-1234-1234-abcdef123457
    export APP_READONLY="arn:aws:sso::<INSERT-ACCOUNT-ID>:application/<INSERT-INSTANCE-ARN>/<INSERT-APPLICATION-ARN>"

    • Group #2 Name: Developer
    • Application Name: sagemaker_readonly

    As part of this validation, the sample solution verifies the group name appears as a whole word in the application name: Case-insensitive, splitting on the -, _, characters and spaces, so ReadOnly matches sagemaker_readonly but read does not. For different use-cases, The GroupNameRegex parameter can be used during deployment.

    Test compliant and non-compliant assignments

    Run the following command to add the ReadOnly group assignment to the sagemaker_readonly application:

    aws sso-admin create-application-assignment --application-arn $APP_READONLY --principal-id $GRP_READONLY --principal-type GROUP

    The group assignment request meets the validation criteria because the application name sagemaker_readonly contains the group name ReadOnly. The output logs for this validation exist within the associated lambda function CloudWatch log group /aws/lambda/identity-center-app-monitor”.

    In this example, the logs will show:

    ✓ COMPLIANT - Group name found in application name

    applicationName="sagemaker_readonly” groupName="ReadOnly”

    Remediation action determined: NONE

    Run the following command to try to add the Developer group assignment to the sagemaker_readonly application:

    aws sso-admin create-application-assignment --application-arn $APP_READONLY --principal-id $GRP_DEVELOPER --principal-type GROUP

    The group assignment request doesn’t meet the validation criteria because the application name sagemaker_readonly doesn’t contain the group name Developer. The output logs for this validation exists within the associated lambda function CloudWatch log group /aws/lambda/identity-center-app-monitor. Note that the remediation action listed shows NOTIFICATION_ONLY, meaning it only sent a notification to the configured SNS topic and did not take action. If you want the group assignment to be deleted, the value should be set to enableAutoDeletion=true.

    In this example, the logs will show:

    ✗ NON-COMPLIANT - Group name not found in application name

    applicationName="sagemaker_readonly” groupName="Developer”

    Remediation action determined: NOTIFICATION_ONLY

    SNS notification sent successfully

    The SNS message will look like:

    {
    	"eventType": "NON_COMPLIANT_ASSIGNMENT",
    	"applicationName": "sagemaker_readonly",
    	"groupName": "Developer",
        "action": "NOTIFICATION_ONLY",
        "status": "SUCCESS",
        "applicationArn": "arn:aws:sso::1234:application/ssoins-1234/apl-1234",
        "groupId": "abc12345-1234-1234-1234-abcdef123457",
        "initiatedBy": { 
        	"type": "AssumedRole", 
        	"arn": "arn:aws:sts::1234:assumed-role/.../you" 
    	}
    }

    Scheduled reporting gives baseline visibility into IAM Identity Center managed applications. Event-driven monitoring can provide near real-time notification or enforcement. Together, these sample solutions can help align and scale Identity Center with your governance and security standards through both historical analysis and immediate response.

    Clean up

    For each deployed CDK stack, run the following commands in the AWS account where it was deployed.

    To delete the remediation stack, run the following commands:

    cd sample-iam-idc-application-discovery-reporting/identity-center-remediation
    cdk destroy

    To delete the reporting stack, run the following commands:

    cd sample-iam-idc-application-discovery-reporting/identity-center-reporting
    cdk destroy

    IAM governance automation at scale

    Achieving effective IAM Identity Center governance at scale requires moving beyond manual processes to automated, continuous monitoring and reporting. The following high-level steps can provide an Identity center governance framework:

    1. Deploy the sample automation with an IAM principal that has access in your delegated administrator account.
    2. Establish a baseline by running your first discovery and reviewing the generated reports.
    3. Configure naming policies to match your organization’s security conventions.
    4. Deploy the event-driven monitoring capabilities to enable real-time policy enforcement and automated response based on the security policies.
    5. Start in notification mode to establish a baseline before enabling auto-remediation.
    6. Integrate with your governance tools by connecting the API endpoints to your compliance dashboards or ITSM tools.
    7. Move to auto-remediation once you have validated policies are working as expected.

    Conclusion

    Managing AWS IAM Identity Center at scale doesn’t have to be a manual, time-consuming process. By implementing automated discovery and reporting combined with real-time event-driven monitoring, you can maintain continuous visibility into your organization’s identity and access landscape, respond immediately to policy violations, and enforce governance policies consistently across your organization. Automation reduces operational work, strengthens security, speeds up incident response, and maintains compliance. Start by deploying these solutions to gain visibility and enable real-time enforcement.

    If you have feedback about this post, submit comments in the Comments section below. If you have questions about this post, start a new thread on AWS re:Post or contact AWS Support.


    Author

    Jonathan Nguyen

    Jonathan is a Principal WWSO AI Security Solution Architect at AWS. He helps customers develop a comprehensive AI security strategy so they can deploy secure AI workloads at scale, integrate AI-powered security services, and defend against AI-powered threats.

    Build a real-time event pipeline with Spark Real-Time Mode on AWS Glue 6.0

    Post Syndicated from Shoukat Ghouse original https://aws.amazon.com/blogs/big-data/build-a-real-time-event-pipeline-with-spark-real-time-mode-on-aws-glue-6-0/

    Real-time event pipelines rarely get to work with a uniform schema. Whether it’s IoT metrics, ecommerce clickstreams, or financial pricing vectors, each event type brings its own schema. An equity trade and a rates trade, for instance, carry almost entirely different fields. Ingesting these multi-schema streams has traditionally forced suboptimal architectural choices. You build separate tables for each event type or maintain a wide STRUCT where every possible field across all event types must be declared upfront (fast reads, but sparse and rigid). The other option is to flatten everything into an unwieldy schema with hundreds of columns. To sidestep that maintenance burden, many teams dump events into a plain JSON string column that introduces significant performance penalty. Querying a single nested field requires your engine to deserialize the entire JSON blob for every row. At scale, you burn compute and budget scanning terabytes of raw text to extract a few bytes of data.

    Adding to the challenge, these pipelines typically demand mixed processing speeds. You need a real-time path (not near-real-time) to flag anomalies or high-risk events with sub-second latency, while simultaneously pushing those same events into analytical storage for deep historical analysis in batch.

    With AWS Glue 6.0, you can tackle all of these challenges (schema heterogeneity, JSON scanning overhead, and mixed-latency requirements) from a single pipeline. Built on Apache Spark 4.1 with Apache Iceberg v3 support, AWS Glue 6.0 brings Variant columns, Variant shredding, Spark Real-Time Mode (RTM), and Arrow-native user-defined functions (UDFs) to a fully managed, serverless environment.

    In this post, we walk you through how to build this multi-layer architecture using a financial services use case: a market risk pipeline processing trade pricing vectors. While the example is finance, the patterns apply wherever you deal with heterogeneous schemas, expensive JSON parsing, and mixed real-time/batch requirements such as IoT device fleets, multi-tenant SaaS platforms, logistics tracking, and beyond. We will show you how to flag high-risk trades with sub-second latency, stream everything into an Iceberg v3 data lake as Variants, and run batch Value at Risk (VaR) computations efficiently using Arrow-native UDFs.

    Solution overview

    A bank’s Market Risk team receives a continuous stream of trade pricing vectors from front-office systems. Each trade event carries:

    1. A trade ID and book/desk IDs.
    2. A pricing vector as semi-structured data. The schema varies by asset class (for example, equities carry risk sensitivities known as Greeks such as delta/gamma, foreign exchange (FX) carries volatility surfaces, rates carry curve sensitivities).
    3. A region ID for jurisdictional reporting (uses a column DEFAULT value, so rows that omit it get auto-populated).

    The team needs three things from this stream, each at a different speed. We build the pipeline in three layers, each addressing a distinct requirement with a purpose-built AWS Glue 6.0 capability.

    Layer 1: Real-time trade position breach detection (sub-second latency)

    Positions must be updated in sub-second time, not seconds of traditional micro-batch streaming. For a team monitoring position limits, those seconds mean trades can breach limits before the system reacts. Spark Real-Time Mode (RTM) eliminates the micro-batch boundary entirely, letting records flow continuously through the pipeline so that high-risk trades trigger alerts within sub-second latency of arrival.

    Layer 2: Near-real-time analytical lakehouse (seconds latency)

    Every trade must land in a queryable data lake within seconds, with heterogeneous pricing vectors stored without declaring a fixed schema upfront. The Iceberg v3 Variant type handles this natively. The raw semi-structured payload goes into a single column regardless of asset class schema. At write time, Variant shredding automatically extracts fields observed in the data into typed Parquet columns, so downstream analytical queries read only the columns they need without deserializing the full blob. Trade amendments and cancellations are handled at a lower cost with deletion vectors (merge-on-read), and column DEFAULT values reduce boilerplate in ingestion code.

    Layer 3: Batch risk computation (minutes to hours latency)

    Risk metrics like Value at Risk (VaR) must be computed in Python across millions of trades. Traditional row-by-row pickle serialization between the Java Virtual Machine (JVM) and Python is the bottleneck. Arrow-native UDFs process data as vectorized columnar batches, eliminating serialization overhead and accelerating Python-based risk calculations.

    The solution uses three separate AWS Glue 6.0 jobs, each independently scalable:

    • Real-time path (Scala, gluestreaming): Reads trades from Amazon Managed Streaming for Apache Kafka (Amazon MSK), enriches them with risk scores and breach flags, and writes alerts to a downstream Kafka topic. It runs with a fixed set of workers that are always on. Downstream fraud detection and position limit systems consume the alerts topic for real-time blocking decisions.
    • Near-real-time path (PySpark, gluestreaming): Reads from the same MSK topic and lands the full trade history into an Iceberg v3 table. It uses Glue auto scaling and can scale down between batches, keeping costs lower.
    • Batch analytics (PySpark, glueetl): Reads from the Iceberg v3 table, extracts fields using variant_get, computes VaR across the portfolio, and writes aggregated risk reports to a downstream summary table.

    The following diagram illustrates the solution architecture.

    Architecture diagram of a real-time market risk pipeline on AWS Glue 6.0. All compute runs inside a VPC within an AWS Account. A Sample Trades Producer (AWS Glue job, simulating front office trading systems) publishes to an Amazon MSK topic named trade-risk-vectors. From MSK, three processing paths branch out. The Real-Time Path uses an AWS Glue 6.0 Spark Real-Time Mode job in Scala for continuous processing (JSON extraction, risk scoring, breach detection), writing alerts to an Amazon MSK trade-alerts topic that feeds CloudWatch Alarms, SNS notifications, and position limit systems. The Near-Real-Time Path uses an AWS Glue 6.0 micro-batch PySpark job that applies PARSE_JSON to Variant, TIMESTAMP_NTZ with nanosecond precision, and shredding, writing to an Amazon S3 Apache Iceberg v3 table named trade_risk_vectors. The Batch Consumption path reads that table with an AWS Glue 6.0 PySpark job using variant_get extraction and an Arrow UDF for Value at Risk computation and jurisdiction classification, writing to an Amazon S3 Iceberg v3 table named daily_risk_summary that feeds downstream analytics. Amazon S3, AWS Glue Data Catalog, and CloudWatch are regional services shown outside the VPC but inside the AWS Account, accessed privately through VPC endpoints.

    Figure 1: Real-time market risk pipeline on AWS Glue 6.0

    Prerequisites

    To follow along with this post, you need the following:

    1. An AWS account in a Region where AWS Glue 6.0 is available.
    2. An AWS Identity and Access Management (IAM) role with permissions to deploy AWS CloudFormation stacks and create resources including AWS Glue, Amazon MSK, AWS Lambda, Amazon Simple Storage Service (Amazon S3), and the AWS Glue Data Catalog.

    Deploy the CloudFormation stack

    We provide an AWS CloudFormation template that provisions all the resources needed for this walkthrough.

    The stack provisions the following resources:

    • An Amazon MSK cluster with two topics: trade-risk-vectors (input) and trade-alerts (real-time alerts output).
    • An Amazon S3 bucket for Iceberg table storage and streaming checkpoints.
    • An AWS Glue database (risk_analytics_<account-id>_glue6b1).
    • An IAM role (GlueRole-<account-id>-glue6b1) with permissions for Glue, MSK, S3, and CloudWatch.
    • Virtual private cloud (VPC) networking: A Glue network connection (connection-<account-id>-glue6b1), S3 gateway endpoint, and Glue interface endpoint.
    • AWS Glue job rtm-alerts-<account-id>-glue6b1 (Scala): This job reads trades from MSK, scores risk in real time using Spark RTM, writes alerts to the trade-alerts topic.
    • AWS Glue job nrt-ingestion-<account-id>-glue6b1 (PySpark): This job reads trades from MSK, writes to Iceberg v3 table with Variant + shredding enabled.
    • AWS Glue job batch-var-<account-id>-glue6b1 (PySpark): This job reads from Iceberg v3 table, computes VaR with Arrow UDF, demonstrates deletion vectors.
    • AWS Glue job producer-<account-id>-glue6b1-helper (PySpark): This job generates sample trade events (equities, FX, rates) to the trade-risk-vectors topic.

    Deploy the CloudFormation stack:

    1. Download the CloudFormation template from the GitHub repository.
    2. Sign in to the AWS CloudFormation console
    3. Choose Create stack > With new resources > Upload a template file, and upload the downloaded template.
    4. Enter the following parameters:
      • VpcId: Your VPC ID.
      • SubnetIds: At least two subnets in different Availability Zones.
      • SecurityGroupId: A dedicated security group that allows all inbound TCP traffic from itself (self-referencing rule).
      • RouteTableId: The main route table for your VPC.
    5. Acknowledge the IAM capabilities and choose Create stack.

    Stack creation takes approximately 20 minutes.

    After the stack completes, open the AWS Glue console and start the jobs in this order:

    1. Start rtm-alerts-<account-id>-glue6b1 and nrt-ingestion-<account-id>-glue6b1.
    2. Once both show RUNNING, start producer-<account-id>-glue6b1-helper.
    3. After the producer finishes (~3.5 minutes), run batch-var-<account-id>-glue6b1 for risk aggregation.

    The consumers must be running before the producer starts so that trades are scored in real time and landed in the Iceberg table as they arrive. The batch job runs last because it reads from the Iceberg table that the near-real-time path populates.

    Understand the Iceberg v3 table design

    The CloudFormation stack provisions Glue jobs that create two Iceberg v3 tables, trade_risk_vectors (primary trade store) and daily_risk_summary (batch VaR output), using new data types and features:

    1. VARIANT: Stores semi-structured pricing vectors without requiring a fixed schema.
    2. DEFAULT values: Automatically applies provided defaults when fields aren’t provided.
    3. Deletion vectors (merge-on-read): Enables fast row-level updates and deletes.

    Open the AWS Glue console under Data Catalog > Tables > trade_risk_vectors.

    AWS Glue Data Catalog console showing the trade_risk_vectors table with its Variant and default-valued columns

    Figure 2: The trade_risk_vectors table in the AWS Glue Data Catalog

    The following is the Create Table command:

    CREATE TABLE {TABLE} (
        trade_id STRING, book_id STRING, desk STRING,
        asset_class STRING DEFAULT 'UNKNOWN',
        execution_time STRING,
        pricing_vector VARIANT,
        var_contribution DOUBLE DEFAULT 0.0,
        risk_weight DOUBLE DEFAULT 1.0,
        trade_date DATE, region STRING DEFAULT 'EMEA'
    ) USING iceberg
    TBLPROPERTIES ('format-version'='3', 'write.delete.mode'='merge-on-read',
        'write.update.mode'='merge-on-read',
        'write.parquet.shred-variants'='true')
    PARTITIONED BY (trade_date, asset_class)

    Note the use of DEFAULT values for asset_class, var_contribution, risk_weight, and region. This is an Iceberg v3 feature that applies defaults automatically when values aren’t provided during writes, reducing boilerplate in ingestion code. The pricing_vector column is defined as a Variant type, and write.parquet.shred-variants='true' automatically extracts Variant fields into separate typed Parquet columns at write time for faster downstream queries.

    Sample trade event generator

    The CloudFormation stack includes a Glue job (producer-<accountid>-glue6b1-helper) that produces realistic trade events to the trade-risk-vectors MSK topic. Each event carries a pricing_vector with a completely different schema per asset class. This is exactly the problem Variant solves.

    Equity trade (greeks, scenarios with sector/region breakdowns):

    JSON pricing vector for an equity trade showing greeks and per-sector and per-region scenario breakdowns

    Figure 3: Sample equity trade pricing vector

    Rates trade (curve sensitivities per tenor, calibration params):

    JSON pricing vector for a rates trade showing curve sensitivities per tenor and calibration parameters

    Figure 4: Sample rates trade pricing vector

    Completely different structures: greeks vs curve sensitivities, BlackScholes vs HullWhite. Both land in the same pricing_vector VARIANT column with no schema changes required.

    Ingest trades with Spark Real-Time Mode

    Traditional Spark Structured Streaming uses micro-batches: collect records, schedule a job, process, commit, wait. Even with small batches, the fixed overhead of planning and scheduling adds noticeable latency per batch. For a risk team monitoring position limits, the delay can let a trade breach a limit before the system reacts.

    The following Scala job reads trade events from Amazon MSK, applies lightweight risk rules based on data directly available in the event, and writes alerts to a Kafka topic, all with sub-second latency. The real-time path intentionally avoids external lookups (market data, volatility surfaces) to stay fast. The full VaR computation happens later in the batch layer where latency is less critical.

    You can view the complete job code in the AWS Glue console under the rtm-alerts-<accountid>-glue6b1 job. Additionally, all the scripts are available in the GitHub repository.

    Scala real-time job code that reads from Kafka, scores risk, and writes alerts to a Kafka topic

    Figure 5: Scala real-time job that scores trades and writes alerts

    The Trigger.RealTime("1 minute") is what distinguishes this from a traditional micro-batch. Records flow through the pipeline continuously. Records are processed the instant they arrive. The 1-minute parameter controls how often Spark checkpoints its progress for recovery. It does not control how often records are processed. RTM on AWS Glue 6.0 currently supports Kafka-source, stateless, Scala workloads with fixed workers (no auto scaling) and update output mode only. This makes it ideal for stateless transformations that require sub-second latency, such as the filter, enrich, score, and route pattern shown here. The heavier computation (VaR, aggregations) runs in the micro-batch/batch layer where sub-second latency is less critical.

    The real-time path acts as a circuit breaker: trades over $50M notional are flagged CRITICAL, over $25M flagged HIGH. Downstream systems consume the trade-alerts topic and can block or escalate before the next trade executes. The detailed VaR computation (which requires market data, volatility surfaces, and the full pricing vector) runs in the batch consumption layer where latency is less sensitive.

    After the streaming phase completes, the job reads back from the trade-alerts topic and measures end-to-end latency. It compares two MSK timestamps: when the trade was received by MSK from the producer, and when the alert was received by MSK from RTM.

    To verify the alerts and latency, open the Amazon CloudWatch console > Log groups > /aws-glue/jobs/output and select the RTM job’s log stream. You will see the alert summary showing each flagged trade with its end-to-end latency. The following is a sample.

    CloudWatch log output listing flagged trades with CRITICAL and HIGH labels and their end-to-end latency

    Figure 6: CloudWatch output showing flagged trades and end-to-end latency

    Store trades in Iceberg v3 with Variant shredding enabled

    The near-real-time path reads from the same MSK topic but writes to an Iceberg v3 table using standard micro-batch streaming. This job runs separately with auto scaling enabled, scaling between batches, keeping costs lower than the always-on real-time path.

    You can view the complete job code in the AWS Glue console under the nrt-ingestion-<accountId>-glue6b1 job. The critical aspects are the Variant conversion and the Iceberg write:

    PySpark code applying PARSE_JSON to build a Variant column and writing to the Iceberg v3 table

    Figure 7: Near-real-time PySpark job writing trades to Iceberg v3 as a Variant

    The PARSE_JSON() function converts the raw pricing vector into a native Variant, regardless of the asset class schema. Whether the incoming trade is an equity with greeks, an FX option with a volatility surface, or a rates swap with curve sensitivities, it all goes into the same column. Since the table has write.parquet.shred-variants enabled, fields observed in the initial sample are automatically extracted into typed Parquet columns for fast downstream queries.

    How shredding works

    During the write process, Spark automatically extracts the Variant fields it observes into separate typed Parquet columns at write time, a feature called shredding. At the start of each write, the engine buffers a sample of rows (controlled by write.parquet.variant-inference-buffer-size), infers which fields exist and their types, then uses that schema to shred all subsequent rows in the file. Every field observed in that sample gets its own typed column, including nested objects. Rows that lack a particular field simply store NULL in that shredded column. For our risk table, fields like $.greeks.delta, $.dv01, and $.model all live in their own typed Parquet columns, even if only one asset class carries a specific field. The result: faster read performance because queries access only the typed columns they need, skipping the rest of the document entirely. Shredding is transparent to queries. variant_get() calls work the same way whether the field is shredded or not. The query engine automatically routes to the shredded column when available, falling back to the binary Variant blob for fields that aren’t part of the inferred schema.

    Note: Shredding adds write latency because the engine must infer the schema and write additional typed columns. In this pipeline, we enable shredding on the near-real-time path and absorb that cost, since the downstream read benefits (batch VaR, ad-hoc queries, audit) far outweigh the write penalty. For latency-sensitive pipelines where every millisecond on the write path matters, you can disable shredding on the streaming table and instead write shredded data in a separate batch job that reads from the unshredded table and inserts into a shredded copy. This approach trades architectural simplicity for lower ingestion latency.

    Build the batch consumption layer

    The third AWS Glue 6.0 job reads from the Iceberg v3 table, extracts risk metrics from the Variant column, computes VaR using an Arrow-native UDF, and writes aggregated results to a summary table.

    You can view the complete job code in the AWS Glue console under the batch-var-<accountid>-glue6b1 job. The key aspects are the variant_get extraction from deeply nested structures and the Arrow-native UDF:

    PySpark code using variant_get to extract deeply nested fields from the Variant column

    Figure 8: Extracting nested Variant fields with variant_get

    Notice how variant_get reaches into arbitrarily nested structures: $.greeks.delta (2 levels), $.scenarios[0].breakdown.by_sector.financials (5 levels), $.model_params.calibration.fit_error (4 levels). All with the same function call. No pre-flattening, no schema-per-asset-class tables, no ETL to restructure the data before querying.

    Once the risk metrics are extracted, we need to run a Monte Carlo-style Historical VaR that simulates 1,000 daily profit and loss (P&L) scenarios per trade and returns the 99th percentile loss. This is where the @arrow_udf decorator comes in.

    Python Arrow UDF code running a Monte Carlo Historical VaR simulation for each trade

    Figure 9: Arrow-native UDF computing Historical VaR

    The @arrow_udf decorator is new in Spark 4.1. Your function receives and returns pyarrow.Array directly, operating on the entire batch of rows at once. There is no pickle serialization, no row-by-row invocation, and no Pandas conversion. Data flows as native Arrow columnar arrays between the JVM and Python. For compute-heavy operations like VaR across hundreds of thousands of rows, this can be significantly faster than traditional scalar UDFs. Additionally, you can use the built-in UDF profilers to identify performance and memory bottlenecks in compute-heavy UDFs such as VaR calculations.

    Handle late trade corrections with deletion vectors

    In financial markets, trade amendments and cancellations are common. The batch VaR job demonstrates this after completing the risk computation. It amends one trade and cancels another.

    PySpark code amending one trade and cancelling another in the Iceberg table

    Figure 10: Amending and cancelling trades with merge-on-read

    Prior to Iceberg v3, row-level deletes required either rewriting entire data files (copy-on-write) or maintaining separate positional delete files that store (file_path, row_position) pairs as Parquet rows (merge-on-read). Both approaches are expensive at scale. Copy-on-write rewrites gigabytes for a single amendment, and positional deletes degrade read performance as delete files accumulate (each read must parse and hash-join all delete records against the data file).

    Because we configured the table with write.delete.mode='merge-on-read', UPDATEs and DELETEs write deletion vectors instead of positional delete files used in Iceberg v2. A deletion vector is a Roaring Bitmap stored in a Puffin file (.puffin), one per affected data file, marking which row positions are deleted. At read time, the engine loads a single bitmap and skips flagged positions with a bit check. No file joins, no linear scan through multiple delete files. The bitmap is compact regardless of how many rows are deleted, and read performance remains predictable as amendments accumulate.

    The batch job also verifies the deletion vectors were created. You can see the results in the job’s output logs.

    Job output confirming deletion vector Puffin files were created for the affected data files

    Figure 11: Output verifying deletion vectors were created

    Clean up

    To avoid incurring further charges, delete the CloudFormation stack. This removes all resources provisioned as part of this post, including the S3 bucket, Glue jobs, Iceberg tables, MSK cluster, and IAM roles.

    Conclusion

    In this post, we built a multi-layer market risk pipeline using AWS Glue 6.0 (real-time alerting, near-real-time ingestion, and batch analytics):

    • Spark Real-Time Mode (RTM) on the real-time path delivers sub-second trade scoring and breach alerting, eliminating the micro-batch boundary so position limits are enforced before the next trade executes.
    • Iceberg v3 Variant on the near-real-time path stores heterogeneous pricing vectors without schema flattening. One table handles equities, FX, and rates with different schemas per row.
    • Variant shredding delivers faster reads by automatically extracting fields into separate typed Parquet columns at write time with no manual tuning required.
    • Arrow-native UDFs eliminate pickle serialization overhead for Python-based risk calculations, processing data as vectorized columnar batches on the batch layer.
    • Deletion vectors handle trade amendments and cancellations without costly data file rewrites, using compact Roaring Bitmaps instead of accumulating positional delete files.
    • Default values reduce boilerplate in ingestion code.

    To get started with AWS Glue 6.0, see the AWS Glue documentation. For more information about Apache Iceberg v3, see the Iceberg specification.


    About the authors

    Shoukat Ghouse

    Shoukat Ghouse

    Shoukat is a Senior Specialist Solutions Architect for Big Data, Analytics, and Data Governance at Amazon Web Services (AWS). He partners with enterprise and financial services customers worldwide to design and scale production-grade data lakehouse platforms on Apache Spark, Apache Iceberg, AWS Glue, Amazon Athena, Amazon EMR and Amazon SageMaker Unified Studio. His focus spans distributed data processing, fine-grained data governance, and helping organizations build AI-ready data foundations that power analytics and machine learning at scale.

    Shrey Malpani

    Shrey Malpani

    Shrey is a Senior Product Manager Technical at Amazon Web Services (AWS), where he works at the intersection of distributed data processing and data integration. He is focused on building and scaling data integration and data management capabilities across services like AWS Glue, Amazon EMR, and Amazon Redshift that help customers build AI-ready data platforms for their analytics and machine learning workflows.

    Danylo Prozorov

    Danylo Prozorov

    Danylo is a Software Development Engineer at Amazon Web Services (AWS), where he works at the intersection of distributed data processing, AI-powered Spark troubleshooting, and AI-driven engineering automation. He focuses on the AWS Glue data integration libraries and AI-powered Spark troubleshooting capabilities across AWS Glue and Amazon EMR, delivering scalable and reliable data integration for customers’ ETL and analytics workloads.

    Kartik

    Kartik

    Kartik is a Software Development Manager on the AWS Glue team. His team builds generative AI features for the Data Integration and distributed system for data integration.

    Optimize EKS operations with agents: Reduce MTTR with AWS DevOps Agent and a Kubernetes Operator

    Post Syndicated from HoSeong Lee original https://aws.amazon.com/blogs/devops/optimize-eks-operations-with-agents-reduce-mttr-with-aws-devops-agent-and-a-kubernetes-operator/

    Introduction

    Running workloads on Amazon Elastic Kubernetes Service (Amazon EKS) can involve managing failures like OOMKilled or IP exhaustion. Engineers must repeatedly collect pod logs, trace events, and check node logs—a process that slows at night/weekends, with critical data lost when pods are deleted or nodes become unhealthy. This collection phase is pure overhead on mean time to resolution (MTTR): the incident stays open while an engineer gathers data that a machine could have captured the moment the failure occurred. Automating it shortens MTTR and lets the on-call engineer start at the analysis step instead of the data-gathering step.

    Existing AI tools have limitations: K8sGPT only analyzes current resource state, and Amazon Bedrock Agents requires manual tool integration and pipeline setup. Neither provides end-to-end automated incident investigation.

    AWS DevOps Agent addresses these gaps—a frontier agent that connects code repositories, observability tools, CI/CD pipelines, and skills to autonomously analyze root causes. This post shows how to build an automated incident response pipeline using the DevOps Agent Operator, a Kubernetes Operator that detects EKS failures and triggers DevOps Agent investigations automatically.

    Solution overview

    AWS DevOps Agent provides powerful incident analysis. However, it does not detect pod failures inside an EKS cluster on its own. To start an investigation, an external source must trigger DevOps Agent through a webhook. When this trigger occurs, two conditions must be met:

    1. Immediate failure detection: You must detect the failure before the pod is rescheduled or deleted.
    2. Sufficient context: You must send the data that the analysis needs, such as the manifest, logs, events, and node information.

    The DevOps Agent Operator is a Kubernetes Operator that meets both conditions automatically.

    Why use an Operator?

    DevOps Agent runs only when something calls it through a webhook or a manual trigger. In 24/7 operations, doing this manually is not practical. Kubernetes keeps events for only about an hour, restarted containers overwrite their logs, and deleted pods lose them entirely. If you do not collect data right after a failure, the key evidence is gone for good.

    DevOps Agent can already run describe and logs with kubectl, and tools like Datadog can detect failures and trigger it.

    A separate Operator still adds value for three reasons:

    1. Proactive preservation of volatile data: The Operator detects state changes in milliseconds via watch and preserves data to S3/CloudWatch instantly—before external tool delays (metric collection, alert evaluation, webhook delivery) let evidence disappear.
    2. Selective collection of node-level data: kubectl exposes only container-level and event data, but root causes often live deeper in the node—for example, OOMKilled traces to node dmesg, and IP exhaustion details are in IPAMD introspection. Because the Operator knows the real-time pod-to-node mapping, it collects only what each failure type needs from the exact node.
    3. Encoding operational knowledge in code: The Operator pattern captures human expertise in code, applying different strategies per failure type—dmesg/memory for OOMKilled, previous logs/restart history for CrashLoopBackOff, IPAMD/ENI mappings for IP exhaustion—directly improving analysis accuracy.

    In short, the Operator captures evidence at the failure site before it disappears and collects data beyond the reach of kubectl, giving DevOps Agent the best possible material to analyze.

    Note: If data collection or an upload to Amazon S3 or CloudWatch Logs fails, the reconcile returns an error and the pod is requeued with exponential backoff rather than dropped, and throttled AWS API requests are retried automatically. The Operator also runs a single reconcile worker and marks each pod with a processed annotation, so a mass failure—for example, 100 replicas crashing at once—is handled one pod at a time and each pod is reported only once. For noisy clusters, WEBHOOK_MIN_SEVERITY and WEBHOOK_SKIP_CATEGORIES let you narrow which failures trigger an investigation.

    Architecture

    Architecture diagram of the DevOps Agent Operator solution. Inside an Amazon EKS cluster in a VPC, the Operator watches pods and detects failures, collects pod data and node logs, stores the data in CloudWatch Logs and Amazon S3, and triggers AWS DevOps Agent through a webhook. DevOps Agent investigates using skills, the stored logs, and GitHub code changes, then notifies the DevOps engineer in Slack.

    Figure 1. End-to-end flow from failure detection to investigation.

    The preceding diagram shows the full flow. The DevOps Agent Operator detects a failure inside the EKS cluster and sends the context to AWS DevOps Agent.

    Getting started

    Prerequisites

    • Region availability: AWS DevOps Agent is available in six AWS Regions—US East (N. Virginia), US West (Oregon), Europe (Frankfurt), Europe (Ireland), Asia Pacific (Sydney), and Asia Pacific (Tokyo). Create your Agent Space in one of these Regions.
    • Node type: Node-level log collection uses AWS Systems Manager Run Command against the EC2 instance that ran the failed pod, so it requires Amazon EKS managed node groups or self-managed EC2 nodes. On AWS Fargate, the Operator still collects Kubernetes-level data—the pod manifest, events, and container logs—but node-level data such as dmesg output and IPAMD introspection is not available.
    • Systems Manager registration: Attach the AmazonSSMManagedInstanceCore policy to your node group’s IAM role so the nodes appear as managed nodes. Without it, node-level collection is skipped and only Kubernetes-level data is collected.

    Setting up this solution involves two steps.

    The first step is to configure the Agent Space for DevOps Agent. You connect the sources that DevOps Agent needs to analyze an incident, such as code repositories and observability tools. You also set up a generic webhook to receive failure information from the Operator.

    The second step is to deploy the DevOps Agent Operator to the EKS cluster. When the Operator detects a pod failure, it collects the context and sends it automatically to the webhook that you set up in the first step.

    After you complete these steps, you have an end-to-end pipeline. When a pod failure occurs, DevOps Agent starts an investigation automatically.

    Step 1: Configure the Agent Space for DevOps Agent

    Configure the webhook

    DevOps Agent supports two types of webhooks:

    • Integration-specific webhooks: Created automatically when you set up an integration with an external solution, such as Slack or Datadog.
    • Generic webhooks: Created manually to trigger an investigation from sources that an external integration does not cover.

    The DevOps Agent Operator uses a generic webhook. It maintains security through HMAC-SHA256 authentication.

    For detailed setup instructions, see the following documentation. This post creates a generic webhook as an example.

    Configure the pipeline

    You connect GitHub or GitLab so that DevOps Agent can track deployment events and correlate code changes with failures.

    1. Register GitHub or GitLab at the AWS account level.
    2. Connect the repositories that you want to monitor to the Agent Space.

    With this connection, DevOps Agent can analyze the recent deployment history and code changes when a failure occurs. DevOps Agent currently supports GitHub and GitLab. For GitLab, you can use both the managed instance and a self-managed instance that is reachable from outside.

    For detailed setup instructions, see the following documentation. This post uses GitHub as an example.

    Configure communication

    DevOps Agent joins your team’s existing communication channels to share its investigation activity. When you connect Slack, you can follow the full process in real time, from failure detection to completed analysis.

    For detailed setup instructions, see the following documentation. This post uses Slack as an example.

    Step 2: Deploy the DevOps Agent Operator

    To install the DevOps Agent Operator, complete the prerequisite steps and the Operator deployment steps in order.

    Before you continue, download the source code. You can find the source code at the following link: DevOps Agent Operator source code

    1. Prerequisite steps

    Before you deploy the Operator to an existing EKS cluster, complete the following prerequisite steps.

    1.1. Create an IAM policy for SSM, Amazon S3, and CloudWatch
    cat >devops-agent-operator-permission.json <<EOF
    {
      "Version": "2012-10-17",
      "Statement": [
        {
          "Sid": "SSMCommandExecution",
          "Effect": "Allow",
          "Action": [
            "ssm:SendCommand",
            "ssm:GetCommandInvocation"
          ],
          "Resource": [
            "arn:aws:ec2:<aws-region>:*:instance/*",
            "arn:aws:ssm:<aws-region>:*:*"
          ]
        },
        {
          "Sid": "S3LogStorage",
          "Effect": "Allow",
          "Action": [
            "s3:PutObject"
          ],
          "Resource": "arn:aws:s3:::<s3-bucket-name>/*"
        },
        {
          "Sid": "S3BucketAccess",
          "Effect": "Allow",
          "Action": [
            "s3:ListBucket"
          ],
          "Resource": "arn:aws:s3:::<s3-bucket-name>"
        },
        {
          "Sid": "CloudWatchLogsIncidentStorage",
          "Effect": "Allow",
          "Action": [
            "logs:CreateLogStream",
            "logs:PutLogEvents"
          ],
          "Resource": "arn:aws:logs:<aws-region>:*:log-group:/<cloudwatch-log-group-name>:*"
        }
      ]
    }
    EOF
    

    Note: To keep this example readable, the policy allows ssm:SendCommand on EC2 instances in the account. In production, restrict it to your cluster’s nodes with an IAM condition key—for example, a StringEquals condition on ssm:resourceTag/eks:cluster-name in a statement that targets only the instance ARN—so that the Operator cannot run commands on unrelated instances. Keep the AWS-RunShellScript document ARN in a separate statement without the condition, because a document carries no instance tags and a single combined statement would deny the call.

    Next, create the policy from this file.

    aws iam create-policy \
        --policy-name devops-agent-operator-policy \
        --policy-document file://devops-agent-operator-permission.json
    
    1.2. Create a trust policy
    cat >devops-agent-operator-trust-policy.json <<EOF
    {
        "Version": "2012-10-17",
        "Statement": [
            {
                "Sid": "AllowEksAuthToAssumeRoleForPodIdentity",
                "Effect": "Allow",
                "Principal": {
                    "Service": "pods.eks.amazonaws.com"
                },
                "Action": [
                    "sts:AssumeRole",
                    "sts:TagSession"
                ]
            }
        ]
    }
    EOF
    
    1.3. Create an IAM role
    aws iam create-role \
        --role-name devops-agent-operator-role \
        --assume-role-policy-document file://devops-agent-operator-trust-policy.json
    
    aws iam attach-role-policy --role-name devops-agent-operator-role --policy-arn=arn:aws:iam::<aws-account-id>:policy/devops-agent-operator-policy
    
    1.4. Associate Pod Identity

    EKS Pod Identity associates Kubernetes service accounts directly with IAM roles, enabling pods to access AWS services like Amazon CloudWatch under the principle of least privilege. For more information, see Learn how EKS Pod Identity grants pods access to AWS services.

    Pod Identity requires the eks-pod-identity-agent add-on, which is not installed on existing clusters by default. If your cluster does not have it yet, add it first:

    aws eks create-addon \
        --cluster-name <eks-cluster-name> \
        --addon-name eks-pod-identity-agent
    

    Then create the association:

    aws eks create-pod-identity-association
      --cluster-name <eks-cluster-name>
      --namespace devops-agent-operator-system
      --service-account devops-agent-operator
      --role-arn arn:aws:iam::<aws-account-id>:role/devops-agent-operator-role
    

    2. Build the image

    Because the Operator is a reference implementation, the sample provides source code only—no prebuilt container image.
    Build the image with the Dockerfile at the following location and push it to a registry that you control, which also keeps the image that runs in your cluster inside your own supply chain. Then use that image to deploy the Operator.
    Building the image locally requires Go 1.25 or later. The Operator is built against the Kubernetes 1.35 client libraries and uses only the core Pod, Node, and Event APIs.

    For example, suppose that you create a separate repository from all the files under Devops Agent Operator – Sample

    You can then build the image through CI/CD with the following GitHub Action as a reference.

    name: Build and Push container images to GitHub Container Registry
    jobs:
      ...
      build-and-push:
        name: Build and Push Image
        runs-on: ubuntu-latest
        needs: create-tag
        steps:
          - name: Checkout
            uses: actions/checkout@v4
            with:
              fetch-depth: 0
          - name: Setup Go
            uses: actions/setup-go@v5
            with:
              go-version-file: go.mod
          - name: Login to GitHub Container Registry
            uses: docker/login-action@v3
            with:
              registry: ghcr.io
              username: ${{ github.repository_owner }}
              password: ${{ secrets.WRITE_REGISTRY_TOKEN }}
          - name: Set up Docker Buildx
            uses: docker/setup-buildx-action@v3
          - name: Build and Push
            uses: docker/build-push-action@v6
            with:
              context: .
              file: Dockerfile
              push: true
              provenance: false
              no-cache: true
              tags: |
                "ghcr.io/${{ github.repository_owner }}/devops-agent-operator:${{ needs.create-tag.outputs.sha_short }}"
                "ghcr.io/${{ github.repository_owner }}/devops-agent-operator:latest"
    

    3. Deploy the Operator

    The following steps are based on the example YAML files for the DevOps Agent Operator. Download the repository, or run the following command to download the files, and then continue.

    curl -s https://api.github.com/repos/aws-samples/kr-tech-blog-sample-code/contents/containers/devops-agent-operator/examples?ref=main | jq -r '.[].download_url' | xargs -n1 curl -O
    
    3.1. Set the environment variables

    Open the 05-deployment.yaml file, and then change the following variables to values that match your environment.

    containers:
    - name: manager
        # Use the image that you built in step 2
        image: <operator-image>:latest
        ...
        env:
        # Required settings
        - name: DEVOPS_AGENT_WEBHOOK_URL
            value: "<devops-agent-webhook-url>"
        ...
        - name: EKS_CLUSTER_NAME
            value: "<eks-cluster-name>"
        - name: AWS_REGION
            value: "<aws-region>"
        - name: AWS_ACCOUNT_ID
            value: "<aws-account-id>"
        # Optional settings
        - name: ENABLE_SSM_COLLECTION
            value: "true"
        - name: CLOUDWATCH_LOG_GROUP
            value: "<cloudwatch-log-group-name>"
    

    Also change the 04-configmap.yaml file to values that match your environment.

    data:
      # Comma-separated list of namespaces to watch (empty = all namespaces)
      WATCH_NAMESPACES: ""
      # Comma-separated list of namespaces to exclude
      EXCLUDE_NAMESPACES: "kube-system,kube-public,kube-node-lease"
      # Enable AWS SSM node log collection (requires IAM permissions)
      ENABLE_SSM_COLLECTION: "true"
      # AWS region for SSM and S3
      AWS_REGION: "<aws-region>"
      ...
    

    In a shared or multi-tenant cluster, set WATCH_NAMESPACES to the namespaces that your team owns so that the Operator does not collect data from other teams’ workloads. If you leave it empty, the Operator watches every namespace except those listed in EXCLUDE_NAMESPACES.

    Note: DevOps Agent references the collected data only while it investigates the incident, so you do not need to retain it long-term. Keeping a short retention period on the CloudWatch log group—and a matching S3 Lifecycle expiration rule on the bucket—keeps the storage cost of this solution minimal.

    # Expire the incident logs in CloudWatch Logs after 14 days
    aws logs put-retention-policy \
        --log-group-name <cloudwatch-log-group-name> \
        --retention-in-days 14
    
    # Expire the incident objects in Amazon S3 after 14 days
    aws s3api put-bucket-lifecycle-configuration \
        --bucket <s3-bucket-name> \
        --lifecycle-configuration '{"Rules":[{"ID":"expire-incident-data","Status":"Enabled","Filter":{"Prefix":"incidents/"},"Expiration":{"Days":14}}]}'
    
    3.2. Create the webhook secret

    Edit the 06-webhook-secret.yaml file:

    stringData:
      webhook-secret: "<webhook-secret>"
    
    3.3. Deploy the Kubernetes resources
    kubectl apply -f .
    

    The example deployment runs a single replica with leader election enabled, so you can raise the replica count for availability without two Operators processing the same failure.

    3.4. Verify the deployment
    # Check the pod status
    kubectl get pods -n devops-agent-operator-system
    
    # Check the logs
    kubectl logs -f deployment/devops-agent-operator \
      -n devops-agent-operator-system
    

    When the Operator works correctly, it produces the following logs:

    Configuration loaded
    Log collector initialized (sinceMinutes: 15)
    Webhook client initialized
    CloudWatch Logs client initialized
    S3 client initialized
    Starting workers (worker count: 1)
    

    Use case: Automated analysis of an OOMKilled failure

    The following scenario shows how the DevOps Agent Operator and DevOps Agent work together. In this environment, Slack is connected as the notification channel for DevOps Agent, and GitHub is connected as the pipeline.

    Scenario

    In this scenario, a developer pushed a code change to add a new feature to the web-python service and built a new container image. The developer then updated the running web-python deployment in the EKS cluster with the newly built image.

    After the new version rolled out successfully, the developer verified that other services were unaffected. Shortly after, a Slack notification arrived. DevOps Agent reported that the pod that was just deployed had terminated with an OOMKilled status, and that it was investigating the related incident.

    kubectl get pods -o custom-columns='NAME:.metadata.name,READY:.status.containerStatuses[*].ready,STATUS:.status.phase,RESTARTS:.status.containerStatuses[*].restartCount,IMAGE:.spec.containers[*].image'
    NAME READY STATUS RESTARTS IMAGE
    web-python-56b9874b88-tdljd true Running 0 <your-registry>/web-python:sha-96cd2b0
    
    # Deploy the new version
    kubectl get pods -o custom-columns='NAME:.metadata.name,READY:.status.containerStatuses[*].ready,STATUS:.status.phase,RESTARTS:.status.containerStatuses[*].restartCount,IMAGE:.spec.containers[*].image'
    NAME READY STATUS RESTARTS IMAGE
    web-python-645b4f7867-lvqgr true Running 0 <your-registry>/web-python:sha-15d1398
    

    The following steps describe what happens after the pod with the new image is deployed.

    Step-by-step flow

    1. Failure detection

    The kubelet detects the OOM termination of the web-python container and updates the pod status. The informer in the DevOps Agent Operator receives this change in real time. It detects the change from the previous state (Running) to the current failure state (OOMKilled).

    kubectl describe po web-python-645b4f7867-lvqgr
    Name: web-python-645b4f7867-lvqgr
    Namespace: default
    ...
    Annotations: devops-agent.io/failure-type: OOMKilled
                      devops-agent.io/processed: true
                      devops-agent.io/processed-at: 2026-05-30T07:24:43Z
    

    2. Kubernetes-level data collection

    As soon as the Operator detects the failure, it collects Kubernetes-level data including pod manifests, pod logs, previous crash logs, and OOM-related event timelines.

    3. Node-level data collection

    It then gathers node-level data such as kubelet, containerd, and ipamd logs, disk/memory/network usage, and the kernel OOM killer log from dmesg output.

    4. Data storage

    Based on your configuration, the Operator stores the collected data in CloudWatch Logs and Amazon S3. DevOps Agent can reference the data in CloudWatch Logs during the investigation when it needs to.

    5. DevOps Agent trigger

    The Operator sends a webhook request that includes an HMAC-SHA256 signature to DevOps Agent. The payload includes investigation instructions for the AI agent.

    The DevOps Agent Operator handles steps 1 through 5. You can also see these steps in the logs of the Operator pod.

    # 1. Failure detection
    2026-05-30T07:24:05Z INFO Failure detected {"controller": "pod", ... "pod": {"name":"web-python-645b4f7867-lvqgr","namespace":"default"}, "failureType": "OOMKilled", "container": "web-python", "exitCode": 137}
    
    # 2-3. Data collection
    2026-05-30T07:24:06Z INFO ssm-collector Collecting node logs via SSM {"node": "ip-192-168-1-10.ec2.internal", "instanceID": "i-0123456789abcdef0"}
    ...
    
    # 4. Data storage
    2026-05-30T07:24:08Z INFO cloudwatch CloudWatch Logs upload completed {"logGroup": "cw-log-group-devops-agent-operator", "logStream": "incidents/2026-05-30T07-24-05Z/default/web-python-645b4f7867-lvqgr", "eventsCount": 13}
    ...
    
    # 5. DevOps Agent trigger
    ...
    2026-05-30T07:24:08Z INFO webhook Webhook request with S3 reference successful {"incidentId": "2026-05-30T07-24-05Z/default/web-python-645b4f7867-lvqgr", "status": 200}
    

    6-7. DevOps Agent Investigation

    As DevOps Agent starts the investigation, it shares the incident and its investigation status in the Slack channel that you configured for communication. Through this notification, the engineer can open the Agent Space and follow the investigation in real time.

    Slack message from the AWS DevOps Agent app reading "Investigation started: Pod OOMKilled: default/web-python-645b4f7867-lvqgr", with a link to view the investigation.

    Figure 2. DevOps Agent announces the OOMKilled incident in Slack.

    Skill-based investigation: DevOps Agent automatically selects the skill that matches the incident type. Following the OOMKilled skill, it systematically performs the steps to check the memory configuration, analyze usage patterns, and review the code change history.

    The Investigation timeline tab showing the payload sent by the Operator—cluster, pod name, node, failure type OOMKilled, exit code 137, and the Amazon S3 data location—followed by the skill file that DevOps Agent read.

    Figure 3. The investigation timeline opens with the payload the Operator sent.

    Correlation analysis: In addition to the troubleshooting data that it receives, DevOps Agent connects the following sources for its analysis:

    • GitHub: Checks recent code changes for memory-related modifications.
    • CloudWatch: Checks memory usage trends in Container Insights.

    In this scenario, you can see that DevOps Agent starts its analysis from the data that the Operator uploaded to CloudWatch Logs, as the skill specifies.

    The timeline showing the OOMKilled symptom and four investigation tasks running in parallel: search-app-logs, search-performance-metrics, search-code-repos, and check-cloudtrail-changes.

    Figure 4. DevOps Agent runs four investigation tasks in parallel.

    The code repository task listing the files DevOps Agent read from the connected repository, including the Kubernetes deployment manifest, the application source, the Dockerfile, and the requirements file.

    Figure 5. DevOps Agent reads the manifest and application code from the connected repository.

    The skill also specifies the relationship between the GitHub repository that you connected as a pipeline and the container image. DevOps Agent uses this information to review the code changes that occurred recently.

    This information helps DevOps Agent identify the root cause of the incident.

    Two findings on the timeline: an unbounded processed_records list leaking about 20 Mi per minute, and a deployment updated to the image that contains the leaking code.

    Figure 6. Two findings: the unbounded list and the deployment that introduced it.

    8. Analysis results

    DevOps Agent organizes the analysis results:

    • Investigation Timeline: This tab shows the agent’s investigation steps—which skills it referenced and what data it analyzed.
      This view helps you optimize the skill to guide investigations more efficiently.
    • Root causes: This section summarizes the root cause from the overall investigation.
    Unbounded `processed_records` list in web-python application causes memory leak at ~20Mi/min
    The Python Flask application in image `<your-registry>/web-python:sha-15d1398` contains a background worker thread (`_cache_worker`) that generates 500 records every 2 seconds and appends processed results to an in-memory list called `processed_records`. Unlike the `cache` list which has eviction logic capped at 80MB (`CACHE_SIZE_MB`), the `processed_records` list has NO eviction or size limit — it grows unboundedly. With Python/Flask overhead (~30MB) + the cache growing toward its 80MB cap, the remaining headroom within the 200Mi container memory limit is exhausted in approximately 10 minutes. This was confirmed by two consecutive pod instances (lvqgr and 7rwdj) both being OOMKilled after exactly ~10 minutes of runtime.
    

    With the investigation from DevOps Agent, the engineer can identify the cause of the problem.

    In the preceding example, you can see how the agent identifies a critical memory leak in the recently changed service code. It then reasons about the cause of the OOM event together with the commit ID.

    The Root cause tab showing the memory leak summary with two supporting observations: the pod being OOMKilled twice within 30 minutes under a 200Mi limit, and the audit log entry for the image change.

    Figure 7. The Root cause tab with its supporting observations.

    9. Analysis and mitigation plan through chat

    The engineer reviews the results and, when needed, can ask DevOps Agent follow-up questions:

    • “Check whether other services show a similar memory growth pattern.”
    • “Will fixing it with approach A help solve the problem?”

    In the following example, the engineer asks whether increasing the pod memory limit will help solve the problem. The agent responds based on its investigation.

    A chat panel where the engineer asks whether raising the pod memory limit to 250Mi would mitigate the issue, and DevOps Agent answers that it would only add about 2.5 minutes before the same OOMKill.

    Figure 8. Follow-up chat on whether a higher memory limit would help.

    As this shows, DevOps Agent goes beyond simple problem analysis. It uses the context that it accumulated during the investigation to respond to the engineer’s follow-up questions with detailed explanations.

    In this scenario, the problem is a logic issue in the source code. For that reason, DevOps Agent could not provide a clear plan at the Kubernetes or AWS infrastructure level. However, based on the root cause, you can receive a mitigation plan related to a rollback.

    The Mitigation plan tab proposing a rollback of the web-python deployment to the previous image, with numbered preparation steps and the AWS CLI commands to verify the cluster first.

    Figure 9. The Mitigation plan tab proposes a rollback.

    Conclusion

    In this post, we introduced the DevOps Agent Operator – a Kubernetes Operator that automatically detects EKS workload failures, collects diagnostic data, and triggers AWS DevOps Agent for root cause analysis.

    By combining these two tools, engineers gain the following benefits:

    • Faster response: Automatic data collection and analysis as soon as a failure occurs, even during nights and weekends.
    • No loss of information: Immediate preservation of all troubleshooting data before a pod is rescheduled or deleted.
    • Comprehensive analysis: DevOps Agent analyzes code repositories, observability tools, and CI/CD pipelines together to trace root causes that are hard to find with a single tool.
    • Organizational knowledge: Through skills, the solution reflects your team’s operational knowledge, enabling incident response with consistent quality.
    • Continuous improvement: Proactive recommendations based on accumulated incident data help prevent future incidents.

    Looking ahead, there are several ways to extend this solution:

    • Support for more resource types: Extend monitoring beyond pods to Job, CronJob, Deployment, and StatefulSet.
    • MCP server integration: DevOps Agent supports Model Context Protocol (MCP) servers, enabling advanced workflows such as querying additional resources during analysis or performing pattern analysis on past incidents.
    • Proactive pattern analysis: As incident data accumulates in Amazon S3 and CloudWatch Logs, DevOps Agent can identify recurring patterns – such as “OOMKilled repeats every Monday morning” – and recommend preventive measures.

    The DevOps Agent Operator project is open source on GitHub. It is a reference implementation rather than a supported product: use it as a working example of how to encode your own detection conditions and collection strategy for the failures your team actually sees.

    To try it yourself, clone the repository, follow the deployment steps in this post, and point the Operator at your own Agent Space webhook. Start with a non-production cluster and a narrow WATCH_NAMESPACES list, then widen the scope once you see the investigations that DevOps Agent produces.

    References

    HoSeong Lee

    HoSeong Lee

    HoSeong is a Cloud Support Engineer at AWS, specializing in containers, infrastructure as code, and CI/CD. With a background in web development and DevOps, he helps customers troubleshoot issues and keep their AWS workloads running reliably. He has deep expertise in Amazon EKS and is interested in applying agentic AI to automate day-to-day operations.

    Boyoung Kim

    Boyoung Kim

    Boyoung is a Cloud Support Engineer at AWS, focusing on containers, infrastructure as code, and CI/CD. She analyzes recurring customer issues and turns proven support patterns into reusable guidance, helping customers build more stable and efficient production workloads.

    YoungJoon Jeong

    YoungJoon Jeong

    YoungJoon is a Specialist Solutions Architect at AWS, specializing in Kubernetes platform engineering and AI/ML infrastructure. He works with enterprises across APJC to design and build production Amazon EKS environments spanning agentic AI platforms, GPU scheduling, hybrid infrastructure, and security governance. He also maintains an open source engineering playbook covering EKS best practices, AI platform architecture, performance benchmarks, and cloud-native operations.

    Extend your data perimeter to the AWS Management Console with Private Access

    Post Syndicated from Madhur Kulkarni original https://aws.amazon.com/blogs/security/extend-your-data-perimeter-to-the-aws-management-console-with-private-access/

    Organizations in regulated industries such as financial services, government, defense, and healthcare restrict their sensitive workloads to isolated network environments with no access to the public internet. Until now, customers could restrict AWS Management Console access to authorized AWS accounts and corporate networks, but the console itself required internet connectivity. This was creating tension between operational convenience and network security controls.

    We’re happy to announce that AWS Management Console Private Access is now generally available with support for virtual private clouds (VPCs) without internet connectivity. Organizations in regulated industries that restrict workloads to isolated network environments can now route all traffic for supported service consoles—including authentication flows, static assets (JavaScript, CSS, images), console-only APIs, and AWS service API calls—through AWS PrivateLink VPC endpoints, eliminating the need for an internet gateway, NAT gateway, or any route to the public internet. This capability is available in all AWS commercial Regions for a select set of supported service consoles.

    In 2023, we launched AWS Management Console Private Access, which you can use to connect to the console by routing console, sign-in, and service API calls through VPC endpoints. However, accessing the console required internet connectivity for static assets and console-only APIs. This meant security teams faced a choice: allow internet connectivity to use the console or deny console access to operators working in network-isolated environments.

    With this launch, AWS Management Console Private Access addresses two common scenarios:

    • Console traffic over internet restricted networks: Traffic for supported service consoles now flows entirely through your VPC endpoints—no proxy allowlists to maintain, no TLS-intercepting proxies to operate, and no CLI-only workflows to accept as a compromise. The same path works seamlessly from Amazon WorkSpaces, Amazon Elastic Compute Cloud (Amazon EC2) instances, and on-premises networks connected through AWS Direct Connect or AWS Site-to-Site VPN. Combined with sign-in resource control policies (RCPs) and sign-in resource policies, you can ensure that console authentication only succeeds from expected networks—even if valid credentials are presented elsewhere, the session is denied. Teams that previously relied on restricted egress rules or manual domain allowlists now get full console access with the same network controls they already trust.
    • Data-exfiltration prevention: Private Access enables you to restrict which AWS accounts and organizational identities can use the AWS Management Console from within your VPC. This prevents access from personal accounts and from accounts outside your organization. Attach a VPC endpoint policy with an aws:ResourceOrgID condition, and console actions are automatically scoped to resources inside your organization. Sign-in RCPs add a second layer by ensuring authentication only succeeds from networks within your perimeter. Together, these controls prevent supported service consoles from being used to access resources in accounts outside your organization—such as personal accounts—without requiring complex network-layer workarounds.

    In this post, you will learn how AWS Management Console Private Access works in environments without internet connectivity, and how to layer access controls using VPC endpoint policies and sign-in resource control policies (RCPs) to strengthen your data perimeter.

    Solution overview

    AWS Management Console Private Access and sign-in resource control policies are a natural extension of the service control policies (SCPs), resource control policies, and VPC endpoint policies you already use for API traffic; now applied to the console session itself. The same data perimeter controls for identity, resource, and network that protect your programmatic access now protect interactive browser sessions too.

    Perimeter Control objective Policy construct Implementation Steps
    Identity Only trusted identities can access my resources Sign-In RCPs and RBPs Restrict which principals can sign in to the console. Before authentication, signin:PrincipalArn is available for exemptions only. After authentication, RCPs restrict at the organization, account, or principal level (aws:PrincipalOrgID, aws:PrincipalAccount, aws:PrincipalArn); RBPs restrict at the account or principal level.
    Identity Only trusted identities are allowed from my network Console VPC endpoint policy and Sign-In VPC endpoint policy Console endpoint: aws:PrincipalOrgID or aws:PrincipalAccount on signed-in identities. Sign-In endpoint: aws:ResourceOrgID or aws:ResourceAccount before authentication, principal and resource keys after authentication. Blocks sign-in to accounts outside your organization, such as personal accounts, from your network.
    Resource My identities can access only trusted resources SCP Resource perimeter SCP with aws:ResourceOrgID follows your principals into every console session; each service API call the console makes on their behalf is denied if the target resource is outside your organization.
    Resource Only trusted resources can be accessed from my network Console VPC endpoint policy and service VPC endpoint policies Console endpoint policy with aws:ResourceOrgID and aws:ResourceAccount scopes what the console can reach through your network.
    Network My identities can access resources only from expected networks SCP Network perimeter SCPs that use aws:SourceVpc deny your principals’ service calls from outside expected networks. With Private Access, requests proxied by the console to supported services carry aws:SourceVpc set to the VPC hosting your Private Access endpoints. Direct browser requests carry VPC context only when the service has its own VPC endpoint, so configure endpoints for every service you use. AWS recommends conditioning on aws:SourceVpc rather than specific aws:SourceVpce values.
    Network My resources can only be accessed from expected networks Sign-In RBPs and RCPs and network perimeter RCPs Sign-In policies deny console authentication from unexpected networks using aws:SourceIp, aws:SourceVpc, aws:SourceVpce, and aws:VpcSourceIp in both pre-authentication and post-authentication statements. Network perimeter RCPs apply the same network conditions to your data resources for any access path.

    With this launch, Console Private Access routes browser traffic for supported service consoles through VPC endpoints, including:

    • Authentication flows – Sign-in, credential exchange, and session token requests
    • Static assets – JavaScript, CSS, and images that render the console UI
    • Service console API calls – The backend requests made when users interact with service consoles

    How traffic flows from a workload in a private VPC through the three Private Access endpoints, with no path to the public internet (shown in Figure 1):

    1. The operator’s browser requests <region>.console.aws.amazon.com.
    2. The corporate DNS forwarder forwards the query to an Amazon Route 53 Resolver inbound endpoint configured within the VPC, which forwards the traffic to the console VPC endpoint.
    3. Browser traffic flows from on-premises through Direct Connect (or AWS Site-to-Site VPN) to the VPC, and the VPC endpoint routes traffic to the console service over the AWS private network.
    4. The console service redirects to the SignIn endpoint <region>.signin.aws.amazon.com to establish a browser session.
    5. The DNS now resolves to the SignIn VPC endpoint’s private IP addresses, and browser traffic flows to the SignIn service over the AWS private network.
    6. After entering credentials, the SignIn service evaluates VPC endpoint policies, in addition to resource-based policies (RBPs) and RCPs, then redirects back to the console.
    7. The console evaluates VPC endpoint policies, loads static content from the console API VPC endpoint, and enforces identity and resource restrictions when making calls to AWS service APIs.
    8. Users can now access the AWS Management Console over Private Access.
    Figure 1: Network isolation architecture

    Figure 1: Network isolation architecture

    Deploy a pilot of AWS Management Console Private Access

    This high-level walkthrough sets up AWS Management Console Private Access for a single AWS Region within one organizational unit (OU). We recommend rolling out incrementally; validate each step before you expand to additional Regions and OUs.

    If you want to validate the mechanics of a Private Access deployment before you build out the full solution, the Getting started with a test environment guide walks you through a minimal configuration: a single VPC with the three Private Access endpoints and a permissive policy. This gives you a working setup to experiment with, independent of the deployment described in the rest of this post. To understand how sign-in policies can verify a user’s network location when they access the console, see Controlling console access with resource-based policies and resource control policies.

    Prerequisites

    You must have the following prerequisites:

    Step 1: Baseline current console access

    Before changing anything, use CloudTrail to map how your users access the console today. Search for eventName = ConsoleLogin over a representative window (we recommend 30 days) and review the sourceIPAddress, vpcEndpointId, and awsRegion fields. Identify which identity types are in use: root user, IAM user, SAML federation, and AWS IAM Identity Center. Decide which OU or account you will pilot with.

    Note: A misconfigured sign-in policy can lock users out of the console. Avoid piloting in a production or shared account. Instead, use a dedicated test account and configure a break-glass principal (covered in Step 5) before enabling access enforcement.

    Step 2: Create the Private Access VPC endpoints

    In your chosen Region, create or identify a VPC to host the endpoints, then create three interface VPC endpoints in that VPC:

    • com.amazonaws.<region>.console for the console.
    • com.amazonaws.<region>.signin for AWS Sign-In.
    • com.amazonaws.<region>.console-static for console-only APIs. This endpoint is required only if your VPC has no internet path.

    Step 3: Configure private DNS for AWS Management Console Private Access

    To use AWS Management Console Private Access, you must configure private DNS so that the console domains—.console.aws.amazon.com, .signin.aws.amazon.com, and the associated static-content domains—resolve to your interface endpoints’ network interfaces within your VPC.

    • For workloads inside your VPC: If the workloads in your VPC use the default Amazon Route 53 Resolver, no additional DNS configuration is required. When you create each interface endpoint, enable the private DNS name option (set PrivateDnsEnabled = true). The public console domains will then resolve automatically to the endpoint network interfaces inside your VPC. If you use a custom DNS resolver or a private hosted zone, you must configure it explicitly to map the console domains to the endpoint addresses. See Working with private hosted zones for more information. For the complete list of domains and detailed DNS configuration steps, see the AWS Management Console Private Access required endpoints documentation.
    • For workloads outside your VPC: For workloads that reach the endpoints from outside the VPC—such as corporate offices connecting over AWS Direct Connect or AWS Site-to-Site VPN—ensure that your corporate DNS resolver returns the endpoint addresses for these domains. See Simplify DNS management in a multi-account environment with Route 53 Resolver for more information.

    Step 4: Verify private connectivity

    Sign in to the console from a workload inside your VPC. The console should load normally. To confirm that traffic is routing through your VPC endpoints, look for the lock icon in the console navigation bar, shown in Figure 2.

    Figure 2: Console Private Access

    Figure 2: Console Private Access

    You can also verify in CloudTrail that recent ConsoleLogin events show the vpcEndpointId field populated with one of your endpoint IDs. Here’s an example CloudTrail ConsoleLogin event snippet showing the vpcEndpointId field:

    {
      "eventVersion": "1.08",
      "userIdentity": {
        "type": "AssumedRole",
        "principalId": "AROA3XFRBF23EXAMPLE:john.doe",
        "arn": "arn:aws:sts::123456789012:assumed-role/Admin/john.doe",
        "accountId": "123456789012"
      },
      "eventTime": "2026-07-08T19:15:32Z",
      "eventSource": "signin.amazonaws.com",
      "eventName": "ConsoleLogin",
      "awsRegion": "us-east-1",
      "sourceIPAddress": "10.0.1.47",
      "userAgent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7)...",
      "requestParameters": null,
      "responseElements": {
        "ConsoleLogin": "Success"
      },
      "additionalEventData": {
        "LoginTo": "https://console.aws.amazon.com/console/home",
        "MobileVersion": "No",
        "MFAUsed": "Yes",
        "vpcEndpointId": "vpce-0abc123def456789a"
      },
      "eventID": "a1b2c3d4-5678-90ab-cdef-EXAMPLE11111",
      "eventType": "AwsConsoleSignIn",
      "recipientAccountId": "123456789012"
    }
    

    If the AWS Management Console doesn’t load, work through the following checks.

    Private DNS is enabled on each interface endpoint (or your custom resolver returns the endpoint addresses)

    When Private DNS is enabled, AWS automatically creates the DNS entries that resolve the service domains (such as console.aws.amazon.com) to the private IP addresses of your VPC endpoints. Without it, your browser still routes to the public AWS endpoints, bypassing your private access setup entirely.

    1. Confirm that each endpoint shows PrivateDnsEnabled:
      aws ec2 describe-vpc-endpoints \
      --filters "Name=vpc-endpoint-type,Values=Interface" \
      --query "VpcEndpoints[].{Id:VpcEndpointId,Service:ServiceName,PrivateDns:PrivateDnsEnabled}" \
      --output table
      

    2. Then, from within your VPC, verify that the domains resolve to private addresses:
      nslookup console.aws.amazon.com
      nslookup signin.aws.amazon.com
      
      # Should return a private IP (e.g., 10.x.x.x), not a public one
      

    Each query should return a private IP address from your VPC CIDR range. If you use a custom DNS resolver instead of the Amazon-provided DNS, ensure your forwarding rules direct the AWS domain queries to the Route 53 Resolver inbound endpoints in your VPC.

    The endpoint security groups allow HTTPS (TCP 443) from your workload subnets

    Each VPC endpoint creates elastic network interfaces (ENIs) in your subnets, and these ENIs are governed by security groups. If those security groups don’t permit inbound HTTPS traffic from your workloads, the connection fails silently.

    1. Identify the security groups attached to your endpoints:
      aws ec2 describe-vpc-endpoints --vpc-endpoint-ids vpce-0abc123def456789a \
        --query "VpcEndpoints[].Groups[].GroupId" --output text
      

    2. Then verify that each security group allows inbound TCP 443 from your workload subnets:
      aws ec2 describe-security-groups --group-ids sg-xxxxxxxx \
        --query "SecurityGroups[].IpPermissions[?ToPort==\`443\`]" \
        --output json
      

    For traffic from outside the VPC (Direct Connect or Site-to-Site VPN), corporate DNS returns the endpoint IPs and the route propagates correctly

    If you access the console from an on-premises workstation connected over AWS Direct Connect or AWS Site-to-Site VPN, two additional conditions must be met.

    1. Your corporate DNS must resolve the AWS domains to the VPC endpoint private IPs. From your on-premises machine, run:

      nslookup console.aws.amazon.com

      If this returns public AWS IPs, your corporate DNS isn’t forwarding queries through Route 53 Resolver. Configure conditional forwarding for the aws.amazon.com and amazonaws.com domains to your Resolver inbound endpoint IPs.

    2. Second, network routes must propagate correctly. Ensure the route table associated with your endpoint subnets has propagated routes from your virtual private gateway (VGW) or transit gateway, so return traffic can reach your on-premises network. Verify this with:
      aws ec2 describe-route-tables \
        --filters "Name=association.subnet-id,Values=subnet-xxxxx" \
        --query "RouteTables[].PropagatingVgws"
      

    A quick end-to-end validation: Run traceroute console.aws.amazon.com from your workstation and confirm the path uses private hops only—no traffic should traverse the public internet.

    If the console loads but the lock icon is missing

    If the console loads but the connection isn’t private (for example, the lock icon is missing), the browser is reaching the console over the public internet instead of through your VPC endpoints.

    • Run nslookup console.aws.amazon.com from a workload inside the VPC. The result should be a private IP from your VPC CIDR range. A public IP means DNS is bypassing the endpoint, which usually happens because Private DNS has not been enabled on the interface endpoint (set PrivateDnsEnabled = true).
    • For workloads outside the VPC, make sure your corporate DNS forwards the console domains into the VPC, for example, through an Amazon Route 53 Resolver inbound endpoint.

    Step 5: Apply VPC endpoint policies

    Attach an endpoint policy to the console and AWS Sign-In endpoints that limits access to identities in your organization. The static-content endpoint doesn’t support endpoint policies.

    Begin with a permissive Allow * policy and confirm that traffic routes through the endpoints (you should see the vpcEndpointId field populated in CloudTrail console events). After confirming the routing, add restrictions to your VPC endpoint policy and observe the traffic.

    A starter policy uses two condition keys: aws:PrincipalOrgID to restrict identities to your organization and aws:ResourceOrgID to restrict the resources the console can reach to your organization’s resources. The full reference, including additional condition keys and resource-restriction patterns, is in the AWS Management Console Private Access user guide.

    For policies beyond the pilot, see Data perimeters on AWS. The data perimeter policy examples GitHub repository covers service-specific considerations for implementing data perimeters in your environment.

    Step 6: Apply a Sign-In policy

    Sign-In policies deny console authentication requests that don’t match your network or principal conditions. The policy is composed of a pre-authentication statement covering signin:Authenticate and a post-authentication statement covering signin:AuthorizeOAuth2Access and signin:CreateOAuth2Token. Include both statements.

    For your pilot, deploy the policy as an RCP from your AWS Organizations management account. When enabled, the RCP applies to all accounts in your organization, so we recommend piloting in a dedicated test organization before rolling it out broadly. Activate enforcement by calling the signin:PutConsoleAuthorizationConfiguration API for the organization in the us-east-1 Region (AWS Sign-In replicates policies globally from there). Resource permission statements have no effect until console authorization is enabled.

    Important: Configure at least one excluded principal as a break-glass path before you enable the RCP. The recommended principal is a dedicated IAM role.

    1. Write the permission statements that define the network conditions:
      Example – Restrict access to corporate VPC:

      aws signin put-resource-permission-statement \
        --source-vpc vpc-0abc123def456789 \
        --requested-region us-west-2 \
        --excluded-principal "arn:aws:iam::123456789012:user/EmergencyAdmin" \
        --region us-east-1
      

      Example – Restrict access to specific IP range:

      aws signin put-resource-permission-statement \
        --source-ip "IP_ADDRESS" \
        --excluded-principal "arn:aws:iam::123456789012:role/BreakGlassRole" \
        --region us-east-1
      

    2. Enable console authorization for the organization to start enforcing the policy.
      aws signin put-console-authorization-configuration \
        --target-id <your-target-id> \
        --region us-east-1
      

    3. Review the consolidated policy that’s now in effect: 
      aws signin get-resource-policy --region us-east-1
      

    For policy examples, the AWS Command Line Interface (AWS CLI) reference, and the lockout-recovery procedure, see the sign-in RBP blog post and the Controlling console access with resource-based policies documentation. The same documentation also covers the per-account alternative, which uses an RBP attached to a single account instead of an organization-wide RCP.

    Step 7: Add a service VPC endpoint

    So far, the console shell loads, the lock icon appears, and your Sign-In policy lets approved identities through. If you sign in to a service console such as the AWS Key Management Service (AWS KMS) console, the page might fail to load resources or hang. The Private Access endpoints carry the console shell, not the service API calls that the console makes on your behalf. In a VPC without an internet gateway, those calls have nowhere to go.

    Add a VPC endpoint for the service itself. For the pilot, create an AWS KMS interface endpoint (com.amazonaws.<region>.kms) in the same VPC, with Private DNS enabled. Open the AWS KMS console from inside the VPC and confirm the list of keys loads. Repeat for each service your users need on day one. Please note that a single service console often calls more than one AWS service API. If a console loads but parts of the page show errors or stay empty, the most common cause is a missing endpoint for one of the services it depends on.

    The current list of services that support PrivateLink is in the AWS PrivateLink documentation. Service consoles whose services don’t support PrivateLink will not work in a no-internet VPC and need to be handled separately.

    Step 8: Hide Regions and services you haven’t configured (optional)

    Console links to a service or Region that you don’t have endpoints for will fail inside your VPC. To prevent users from navigating to broken pages, use User Experience Customization (UXC) to hide Regions and services that aren’t part of your Private Access deployment. UXC is configured at the account level and applies to navigation, search results, and service-selection drop-downs.

    Step 9: Validate, then expand

    After applying the endpoint policies and the Sign-In RCP to one pilot account:

    1. Sign in from inside the corporate network. The session should succeed.
    2. Sign in from outside the corporate network. The session should be denied at the Sign-In step, before reaching the console.
    3. In CloudTrail, confirm ConsoleLogin events show vpcEndpointId populated for traffic from inside the network.
    4. For unexpected denials, look in CloudTrail for ConsoleLogin events with the error message Authorization denied because of a resource-based policy or Authorization denied because of a resource control policy to identify which statement was responsible.

    Considerations

    A few items worth mentioning before you commit to this design:

    • AWS IAM Identity Center: IAM Identity Center sign-in support isn’t yet available through a VPC endpoint. Initial single sign-on (SSO) authentication must still transit over the internet.
    • Programmatic access: Sign-In RBPs and RCPs gate interactive console sign-in. AWS SDK and AWS CLI requests signed with SigV4 aren’t affected. This is also your recovery path: a principal with signin:DeleteConsoleAuthorizationConfiguration permission can disable enforcement programmatically if console authorization is misconfigured.
    • Apps integrated with AWS Sign-In: Sign-In policies also apply to Amazon Connect, Amazon WorkSpaces, Amazon QuickSight, AWS Health Dashboard, Amazon AppStream 2.0, and Amazon Lightsail when those applications use AWS Sign-In to authenticate.
    • AWS Management Console Private Access is available in all commercial AWS Regions but supports only a subset of AWS service consoles. See Supported AWS Regions, service consoles, and features in Private Access documentation for more information.
    • For services that aren’t supported, you can still navigate to other consoles, but will require internet connectivity for the unsupported service consoles and console-only APIs.
    • Costs: You pay regular AWS PrivateLink endpoint pricing and data processing for each endpoint and each Region you deploy in. The three Private Access endpoints (console, signin, and console-static) plus the service endpoints you already use are the relevant line items.

    Conclusion

    In this post, we showed you how to extend the AWS data perimeter framework to the AWS Management Console. You routed console traffic through VPC endpoints with AWS Management Console Private Access, restricted console sign-in by network and organization with Sign-In RBPs and RCPs, and configured the console to operate in a VPC without an internet gateway. The four control objectives that you already enforce for API traffic now also apply to the console.

    To get started, see the AWS Management Console Private Access documentation. For deployment patterns and Region-by-Region considerations, see the AWS Management Console Private Access reference architectures. For background on the broader pattern, see Establishing a data perimeter on AWS.

    If you have feedback about this post, submit comments in the Comments section below. If you have questions about this post, start a new thread on the IAM forum on AWS re:Post or contact AWS Support.


    Madhur Kulkarni

    Madhur Kulkarni

    Madhur is a Sr. Customer Solutions Manager at AWS, working with Strategic Accounts customers to accelerate cloud adoption and drive business outcomes. He partners with cross-functional teams across customer engineering, AWS service teams, and specialists organizations to deliver enterprise-scale cloud solutions.

    Mateusz Jaworski

    Mateusz Jaworski

    Mateusz is a Principal Engineer at AWS, where he works on AWS Management Console.

    Sujay Ghosh

    Sujay Ghosh

    Sujay is a Software Development Manager at AWS, where he leads a team responsible for enabling secure, reliable access to the AWS Management Console. He is passionate about building scalable infrastructure that helps millions of customers manage their cloud resources safely and efficiently.

    Abhijit Barde

    Abhijit Barde

    Abhijit is a Principal Product Manager at AWS, where he focuses on making it straightforward for all AWS users to discover, monitor, and operate their AWS infrastructure using conversational assistants and generative AI.

    How to build a serverless mass email solution with Amazon SES

    Post Syndicated from Brad Watson original https://aws.amazon.com/blogs/messaging-and-targeting/how-to-build-a-serverless-mass-email-solution-with-amazon-ses/

    Sending mass email campaigns presents significant challenges for many organizations. Enterprises often spend millions annually on proprietary email systems that are inflexible and expensive to maintain. These legacy platforms can restrict sending capacity, offer limited control, and require costly licensing agreements. The challenges intensify when handling large-scale communications like automated notifications, bulk marketing campaigns, and system-generated alerts. These scenarios create reliability issues, scaling limitations, and rising costs that impact teams’ ability to communicate effectively with customers.

    Recently, a large federal organization faced similar challenges, spending over a million dollars annually on their email campaigns. By building a custom email solution on AWS, they sent a 2 million email campaign for approximately $300. This cost includes Amazon Simple Email Service (Amazon SES) and other AWS services. This transformation cut costs while providing the scalability and flexibility they needed for their growing campaign needs.

    This transformation succeeded because building a cloud-native serverless mass email solution offers several advantages:

    • Cost optimization.
      • Pay only for email sent and actual compute resources used.
      • Remove costs associated with managing email servers.
      • Remove expensive licensing fees and maintenance overhead.
    • Scalability and reliability.
      • Automatically handle varying email volumes without infrastructure changes.
      • Support reliable delivery through built-in retry mechanisms and error handling.
      • Perform consistently during peak sending periods.
    • Security and compliance.
      • Secure access control through AWS Identity and Access Management (IAM) roles with least-privilege principles.
      • Comprehensive audit trails for all email campaigns with detailed logging to support your reporting requirements.
      • Detailed logging that customers can use for their compliance and reporting requirements.
      • Data encryption in transit and at rest that you can configure.

    In this post, we explore the architecture of a cloud-native serverless mass email solution that integrates Amazon SES with AWS Step Functions, Amazon API Gateway, and Amazon DynamoDB. You will learn how these services work together to process email campaigns at scale while minimizing cost. Let’s get started!

    Solution overview

    The serverless mass email solution consists of two main components: a user-friendly frontend interface and a scalable serverless backend. The frontend operates completely independently from the backend processing system, communicating through RESTful APIs from Amazon API Gateway. With this architecture, you can use the provided frontend interface as-is. Alternatively, you can integrate your own custom UI or existing applications while using the same backend email processing infrastructure.

    The following diagram shows the complete architecture of the serverless mass email solution, including how the frontend and backend components connect through API Gateway to process email campaigns.

    Complete serverless mass email architecture, with the frontend and backend connected through Amazon API Gateway

    Figure 1: Complete architecture

    Frontend architecture and user flow

    The frontend of the solution prioritizes usability while providing email campaign capabilities. Here’s how the components work together:

    Frontend architecture: web interface, Amazon Cognito authentication, and requests through API Gateway to AWS Lambda

    Figure 2: Frontend architecture of the SES email application

    1. Login – Users navigate to the web interface URL (hosted on Amazon Simple Storage Service (Amazon S3)) which prompts them to authenticate.
    2. User authentication – Amazon Cognito handles authentication, providing secure user management and restricting access to authorized users.
    3. User interface – After successful authentication, users are redirected to a graphical user interface (GUI) where they can design and save email templates and launch large-scale campaigns (refer to figures 3 and 4).
      1. Templates.
        1. Amazon SES supports two types of templates: stored and inline. Stored templates live in SES, and you can reuse them across campaigns. With inline templates, you define the content and variables directly in the email sending request. Both approaches support dynamic personalization by replacing variables with recipient-specific data when the email is sent. For example, you can create a template that personalizes each email with the recipient’s name, custom offers, or any other dynamic content. For detailed information about template capabilities and personalization options, refer to the Amazon SES template documentation.

    The following screenshots show the campaign interface, the template creation interface, and the campaign monitoring interface.

    Email template creation interface of the mass email application

    Figure 3: Email template creation interface

    Mass email campaign interface of the application

    Figure 4: Mass email campaign interface

    Campaign monitoring interface showing the delivery status of a mass email campaign

    Figure 5: Campaign monitoring interface

    1. Request processing – Each user action triggers a secure request through Amazon API Gateway to AWS Lambda functions, which then coordinate with our backend processing system.

    From the user’s perspective, the experience is similar to using any standard email platform, with the added capability of handling campaigns at scale. This interface helps marketing teams, customer success managers, and business operations staff create and launch email campaigns directly through their browser, without needing to understand complex email protocols.

    Backend architecture

    After a user initiates an email campaign, our backend orchestrates a series of steps to facilitate reliable, large-scale email delivery. Let’s follow how an email campaign flows through the system:

    Backend architecture: Step Functions orchestrates batching, Lambda sends email through Amazon SES, and DynamoDB logs delivery attempts

    Figure 6: Backend architecture of the SES email application

    As shown in the preceding figure, the backend processes email campaigns through the following steps:

    1. Email campaign processor – When a user creates a new campaign through the GUI, a Lambda function processes the initial request, taking the user’s selected email template and campaign parameters. The function then triggers an AWS Step Functions workflow.
    2. Workflow orchestration – The Step Functions workflow acts as the conductor and coordinates the entire email sending process. It initializes the campaign, sets up necessary configurations, and organizes the campaign into manageable batches.
    3. Recipient processing – Before sending email, the Step Functions workflow retrieves recipient information, including the recipient’s name and email address, from DynamoDB and checks it for accurate delivery details.
    4. Batch email processing – The Step Functions workflow begins organizing the email into manageable batches. The workflow queues these batches in Amazon Simple Queue Service (Amazon SQS), preparing them for processing.
    5. Batch monitoring – As batches move through the system, Step Functions actively monitors their progress, tracking the status of each batch throughout the sending process.
    6. Email sending – When SQS receives a message, it invokes a Lambda function that sends the email to Amazon SES for delivery. The function logs each delivery attempt in DynamoDB, with failed deliveries automatically returning to the SQS queue for retry attempts. It also records successful deliveries to support idempotency and prevent duplicate sends.
    7. Record management – DynamoDB stores an audit trail that tracks both successful and failed delivery attempts, providing detailed logs to support reporting, campaign performance assessments, and compliance efforts.

    Using these AWS services, the solution automatically scales from sending a few email to millions without manual intervention or infrastructure provisioning. You pay only for what you use, with no idle server costs. To demonstrate the cost-effectiveness of this architecture: sending 10,000 email costs approximately USD $4, including all AWS service charges. For current pricing details, refer to Amazon SES pricing.

    To deploy this solution in your AWS account, refer to the source code on GitHub.

    Conclusion

    In this post, we explored the architecture of a scalable email sending solution using Amazon SES and other AWS serverless services. This architecture removes the complexity of traditional email infrastructure while providing capabilities for handling large-scale email campaigns. Whether you’re looking to modernize your existing email infrastructure or stand up a new solution, this serverless approach offers the ideal combination of streamlined design, scalability, and cost-effectiveness.

    Additional resources


    About the authors

    Build your own continuous modernization pipeline with AWS Transform custom

    Post Syndicated from Janardhan Molumuri original https://aws.amazon.com/blogs/devops/build-your-own-continuous-modernization-pipeline-with-aws-transform-custom/

    Introduction

    Development velocity has reached new heights with AI-driven development tools and practices. Organizations are generating code faster than ever before. But that speed carries risk. Researchers Anderson, Parker, and Tan warned in MIT Sloan Management Review, “Legacy systems tend to carry hidden debt; layering AI-generated code on top of them creates additional tangled dependencies.” The faster you generate code, the faster technical debt compounds — especially in brownfield environments where outdated frameworks, deprecated libraries, and undocumented services already carry years of accumulated risk.

    As organizations accelerate their software development, manual or periodic processes to synchronize dependencies and update documentation no longer keep pace, and technical debt piles up faster than ever. Continuous modernization built into your pipeline enables you to maintain up-to-date dependencies and documentation across repositories on every commit, preventing future tech debt and improving AI agent accuracy and accountability.“

    You can embed AI-powered code transformations directly into your CI/CD pipelines, turning modernization from a periodic project into an automated, ongoing practice. AWS gives you two ways to get there. AWS Transform – continuous modernization is the fully managed option, delivering continuous modernization automatically with no pipeline for you to build or maintain. The Do-It-Yourself (DIY) approach assembles the same practices yourself using AWS Transform custom and your existing CI/CD platform. Choose DIY when you need to fit modernization into a specific pipeline (GitHub Actions, AWS CodePipeline, Jenkins, GitLab CI, and so on), or want to customize the workflow with existing tools like Dependabot.

    In this post, we cover the DIY approach on how to set up a continuous modernization pipeline using AWS Transform custom and demonstrate it in action.

    The Do It Yourself (DIY) path – continuous modernization pipeline with AWS Transform custom

    Sample application: instrumentShop

    For this walkthrough, we use a dated Java application called instrumentShop (Figure 1) — a Java microservices application built with Spring Boot that simulates an online instrument shop to demonstrate four practices: automated dependency remediation, auto-documentation on every commit, scaling transformations across repositories, and continual learning.

    Architecture overview
    instrumentShop Java application architecture: a Spring Gateway routing traffic to four REST services (Agents, Instruments, Consumers, Products), with a Thymeleaf client, PostgreSQL persistence, and Hystrix circuit breaking.

    Figure 1: instrumentShop Java application architecture

    The instrumentShop application is a Spring Boot microservices application with a Spring Gateway (v1.5.19) routing traffic from a single HTTP/8010 entry point to four REST services: Agents, Instruments, Consumers, and Products. A Thymeleaf client provides server-side rendering, PostgreSQL 13.1 handles persistence via JDBC, and Hystrix provides circuit-breaking for inter-service calls. A ShopTester utility generates HTTP traffic for testing.

    This application is a strong candidate for continuous modernization:

    • Spring Boot 1.5.19 is years past end of life and carries known CVEs
    • Hystrix has been in maintenance mode since Netflix deprecated it in 2018
    • Cross-service coordination — dependency updates must propagate across multiple microservices
    • Transitive dependency risk — PostgreSQL JDBC drivers and other transitive dependencies accumulate security advisories over time

    A typical workflow for the continuous modernization pipeline is shown below (Figure 2):

    • A developer pushes code to main — GitHub Actions triggers the auto-documentation workflow, generating updated architecture docs and technical debt reports.
    • Dependabot detects a vulnerable dependency — A PR opens automatically. GitHub Actions triggers the dependency remediation workflow, runs AWS Transform custom to remediate the code, validates with tests, and pushes the result back to the PR.
    • A platform team defines a new transformation (e.g., “Upgrade Spring Boot to the latest stable release “) — The scheduled GitHub Actions workflow runs the transformation weekly in non-interactive mode across all instrumentShop microservices and other repositories in the portfolio.
    • The agent learns — Knowledge items from each execution improve future runs, reducing manual intervention over time.

    AWS Transform continuous code modernization workflow
    Figure 2: AWS Transform continuous code modernization workflow

    Prerequisites

    • Before setting up the continuous modernization pipeline, ensure you have the following:
    • An active AWS account with permissions for AWS Transform custom
    • AWS Transform CLI installed and configured in your development environment
    • Authentication with AWS credentials configured locally and proper IAM permissions to call AWS Transform
    • Git installed for cloning sample repositories
    • GitHub Dependabot enabled on your repository for automated vulnerability detection

    Continuous modernization through CI/CD in action

    Continuous modernization shifts code transformation from a periodic project into an automated, pipeline-driven practice. Instead of scheduling a “modernization sprint” once a year, your CI/CD pipeline identifies and remediates technical debt on every commit, every dependency alert, and across every repository.

    We implement this through four practices, each powered by AWS Transform custom running as a step in GitHub Actions workflows.

    Note: This post uses GitHub Actions because the instrumentShop demo repository is built with it. The same AWS Transform CLI (atx) commands work with AWS CodePipeline, Jenkins, GitLab CI, CircleCI, or any CI/CD system that runs shell commands. Continuous modernization is a practice, not a tool choice.

    Important: Every atx custom def exec invocation in this post uses the –trust-all-tools flag, which allows the agent to execute tools without interactive confirmation. This is required for non-interactive CI/CD execution. Review your organization’s security policies before enabling this flag in production pipelines.

    1. Dependency analysis and remediation

    GitHub Dependabot scans your repository for known vulnerabilities and generates alerts when a new vulnerability is added or your dependency graph changes—for example, when you push commits that update packages or versions. However, resolving these alerts requires more than bumping a version number. Upgrading a dependency can introduce breaking API changes, require code modifications, or demand configuration updates.

    AWS Transform custom helps handle the code changes needed to resolve the alerts. It runs via a GitHub Actions workflow that triggers automatically to:

    • Fetch the list of latest Dependabot alerts
    • Run AWS Transform custom to analyze the alerts and apply code transformations
    • Run your build and test suite to validate the changes
    • Create a new pull request for each resolved alert

    The workflow calls a shell script that invokes the AWS Transform CLI in headless mode with retry logic. Place this script at the root of your repository:

    run_dependabot_alert_fixes.sh:

    #!/usr/bin/env bash
    set -euo pipefail
    
    # -------------------------------------------------------------------
    # run_dependabot_alert_fixes.sh
    # Runs the Dependabot alert remediation transformation in headless mode.
    # Retries up to MAX_RETRIES times on failure.
    #
    # Usage:
    #   ./run_dependabot_alert_fixes.sh [-n <transformation-name>] [-p <path>] [-c <build-command>]
    #
    # Defaults:
    #   -n  Remediate-Critical-GitHub-Dependabot-Alerts-Java-Maven
    #   -p  .                   (current directory)
    #   -c  mvn clean install   (Maven build)
    # -------------------------------------------------------------------
    
    TRANSFORMATION_NAME="Remediate-Critical-GitHub-Dependabot-Alerts-Java-Maven"
    CODE_PATH="."
    BUILD_CMD="mvn clean install"
    MAX_RETRIES=3
    
    while getopts "n:p:c:" opt; do
      case $opt in
        n) TRANSFORMATION_NAME="$OPTARG" ;;
        p) CODE_PATH="$OPTARG" ;;
        c) BUILD_CMD="$OPTARG" ;;
        *) echo "Usage: $0 [-n <transformation-name>] [-p <path>] [-c <build-command>]" && exit 1 ;;
      esac
    done
    
    echo "=== AWS Transform Custom ==="
    echo "Transformation: $TRANSFORMATION_NAME"
    echo "Code path:      $CODE_PATH"
    echo "Build command:  $BUILD_CMD"
    echo "============================"
    
    attempt=1
    while [ $attempt -le $MAX_RETRIES ]; do
      echo "--- Attempt $attempt of $MAX_RETRIES ---"
    
      if atx custom def exec \
        -n "$TRANSFORMATION_NAME" \
        -p "$CODE_PATH" \
        -c "$BUILD_CMD" \
        -x -t; then
        echo "=== Transformation completed successfully ==="
        exit 0
      fi
    
      echo "Attempt $attempt failed."
      attempt=$((attempt + 1))
    
      if [ $attempt -le $MAX_RETRIES ]; then
        echo "Retrying in 10 seconds..."
        sleep 10
      fi
    done
    
    echo "=== All $MAX_RETRIES attempts failed ==="
    exit 1

    This script accepts optional flags to override the transformation name (-n), code path (-p), and build command (-c). The -x flag enables non-interactive mode and -t enables --trust-all-tools, both required for CI/CD execution. On failure, it retries up to three times with a 10-second backoff.

    Your CI/CD workflow must configure AWS credentials and install the AWS Transform CLI before invoking this script. With this setup, Dependabot alerts are reviewed continuously for any changes — not just a version bump, but the complete code adaptation required to make the upgrade work.

    2. Auto documentation

    Documentation is one of the most neglected aspects of modern software development. Documentation increases accuracy and acts as a contract between requirements and implementation. AWS Transform custom codebase analysis capability generates structured documentation covering architecture, technical debt, code metrics, and migration planning on every incremental update ensuring every Agent or human that modifies the codebase is working from a true “current state”.

    By embedding this as a post-push step in your CI/CD pipeline, your documentation stays current automatically. The workflow triggers on every pull request to main, runs your build and test suite, then calls a shell script that invokes AWS Transform custom to generate documentation and commits it back to the PR branch.

    Place this script at the root of your repository:

    run_code_analysis.sh:

    #!/usr/bin/env bash
    set -euo pipefail
    
    # -------------------------------------------------------------------
    # run_code_analysis.sh
    # Runs an AWS Transform custom transformation in headless mode.
    # Retries up to MAX_RETRIES times on failure.
    #
    # Usage:
    #   ./run_code_analysis.sh [-n <name>] [-p <path>] [-c <build-cmd>] [-U <pr-url>]
    #
    # Defaults:
    #   -n  GitHub-PR-Context-Codebase-Analysis
    #   -p  .                   (current directory)
    #   -c  mvn clean install   (Maven build)
    #   -U  (empty)             PR URL
    # -------------------------------------------------------------------
    
    TRANSFORMATION_NAME="GitHub-PR-Context-Codebase-Analysis"
    CODE_PATH="."
    BUILD_CMD="mvn clean install"
    PR_URL=""
    MAX_RETRIES=3
    
    while getopts "n:p:c:U:" opt; do
      case $opt in
        n) TRANSFORMATION_NAME="$OPTARG" ;;
        p) CODE_PATH="$OPTARG" ;;
        c) BUILD_CMD="$OPTARG" ;;
        U) PR_URL="$OPTARG" ;;
        *) echo "Usage: $0 [-n <name>] [-p <path>] [-c <build-cmd>] [-U <pr-url>]" && exit 1 ;;
      esac
    done
    
    echo "=== AWS Transform Custom ==="
    echo "Transformation: $TRANSFORMATION_NAME"
    echo "Code path:      $CODE_PATH"
    echo "Build command:  $BUILD_CMD"
    echo "PR URL:         $PR_URL"
    echo "============================"
    
    attempt=1
    while [ $attempt -le $MAX_RETRIES ]; do
      echo "--- Attempt $attempt of $MAX_RETRIES ---"
    
      if atx custom def exec \
        -n "$TRANSFORMATION_NAME" \
        -p "$CODE_PATH" \
        -c "$BUILD_CMD" \
        -g "additionalPlanContext=$PR_URL" \
        -x -t; then
        echo "=== Transformation completed successfully ==="
        exit 0
      fi
    
      echo "Attempt $attempt failed."
      attempt=$((attempt + 1))
    
      if [ $attempt -le $MAX_RETRIES ]; then
        echo "Retrying in 10 seconds..."
        sleep 10
      fi
    done
    
    echo "=== All $MAX_RETRIES attempts failed ==="
    exit 1

    This script accepts optional flags for the transformation name (-n), code path (-p), build command (-c), and PR URL (-U). Pass the PR URL to the agent via the -g flag as additionalPlanContext, giving it awareness of the pull request context when generating documentation. On failure, it retries up to three times with a 10-second backoff.

    Your CI/CD workflow must configure AWS credentials and install the AWS Transform CLI before invoking this script. The workflow commits the generated documentation back to the PR branch automatically, keeping your architecture docs and technical debt reports current with every code change.

    Every push now updates the documentation (Figures 3 and 4) — reducing knowledge silos and preserving institutional knowledge.

    A GitHub pull request triggering the auto-documentation workflow.

    Figure 3: PR triggering auto-documentation

    Generated documentation output showing architecture and technical debt reports.

    Figure 4 – Generated documentation output

    3. Scale across repositories

    For organizations with hundreds of microservices, transforming one repository at a time doesn’t scale. AWS Transform custom non-interactive mode combined with GitHub Actions matrix strategy allows you to orchestrate transformations across your entire portfolio in parallel. You can run them on demand or on a recurring schedule, so modernization runs as a continuous practice rather than a one-time project.

    # .github/workflows/scale-modernization.yml
    name: Scale Modernization
    on:
      schedule:
        - cron: '0 6 * * 1'
      workflow_dispatch:
    
    jobs:
      transform-repos:
        runs-on: ubuntu-latest
        strategy:
          matrix:
            repo:
              - magnefique-studios/instrumentShop
              - magnefique-studios/orderService
              - magnefique-studios/paymentGateway
        steps:
          - name: Checkout ${{ matrix.repo }}
            uses: actions/checkout@v4
            with:
              repository: ${{ matrix.repo }}
              token: ${{ secrets.GH_PAT }}
    
          - name: Configure AWS credentials
            uses: aws-actions/configure-aws-credentials@v4
            with:
              role-to-assume: ${{ secrets.AWS_ROLE_ARN }}
              aws-region: us-east-1
    
          - name: Install ATX CLI
            run: curl -fsSL https://transform-cli.awsstatic.com/install.sh | bash
    
          - name: Run transformation
            run: |
              atx custom def exec \
                --transformation-name "spring-boot-3-upgrade" \
                --code-repository-path "." \
                --build-command "mvn clean install" \
                --non-interactive \
                --trust-all-tools

    Tip: GitHub Actions matrix strategy runs each repository in parallel automatically — no separate orchestration layer needed. For larger portfolios, you can also wrap this in AWS Batch or AWS Fargate for large-scale parallel execution. The AWS Transform web console tracks progress across all repositories in a single view.

    4. Continual learning

    Each time AWS Transform custom completes a transformation, a memory agent scans the full execution trajectory and extracts lessons. Lessons include patterns that the agent learned, decisions that the agent made during planning, and feedback you provide during execution. AWS Transform custom automatically attaches these lessons to your transformation definition, which improves accuracy in subsequent runs.

    AWS Transform custom applies lessons automatically, and each lesson belongs to a category that groups related lessons for review. You can browse and archive any lesson you do not want AWS Transform custom to apply to future runs.This keeps a human in the loop on what the agent “remembers” which matters when the same transformation runs across many repositories with different conventions.

    In practice, this means your “Spring Boot 3 Upgrade” transformation gets sharper with each execution. The first repository surfaces the edge cases; once you review the resulting lessons and archive the ones that do not fit, subsequent runs handle those edge cases without intervention.

    For production use, you can combine these practices into a single workflow file:

    Note: The individual workflows shown in Practices 1–3 are presented separately for clarity. Combine them into a single workflow file as shown here, or keep them as separate workflow files depending on your team’s preference.

    # .github/workflows/continuous-modernization.yml
    name: Continuous Modernization
    on:
      push:
        branches: [main]
      pull_request:
        types: [opened]
      schedule:
        - cron: '0 6 * * 1'
    
    jobs:
      dependency-remediation:
        if: github.actor == 'dependabot[bot]'
        runs-on: ubuntu-latest
        steps:
          - uses: actions/checkout@v4
            with:
              ref: ${{ github.head_ref }}
          - name: Configure AWS credentials
            uses: aws-actions/configure-aws-credentials@v4
            with:
              role-to-assume: ${{ secrets.AWS_ROLE_ARN }}
              aws-region: us-east-1
          - name: Install ATX CLI
            run: curl -fsSL https://transform-cli.awsstatic.com/install.sh | bash
          - name: Remediate dependency changes
            run: |
              atx custom def exec \
                --transformation-name "dependency-remediation" \
                --code-repository-path "." \
                --build-command "mvn clean install" \
                --non-interactive \
                --trust-all-tools
    
      auto-documentation:
        if: github.event_name == 'push'
        runs-on: ubuntu-latest
        steps:
          - uses: actions/checkout@v4
          - name: Configure AWS credentials
            uses: aws-actions/configure-aws-credentials@v4
            with:
              role-to-assume: ${{ secrets.AWS_ROLE_ARN }}
              aws-region: us-east-1
          - name: Install ATX CLI
            run: curl -fsSL https://transform-cli.awsstatic.com/install.sh | bash
          - name: Generate documentation
            run: |
              atx custom def exec \
                --transformation-name "codebase-documentation" \
                --code-repository-path "." \
                --build-command "echo 'docs-only'" \
                --non-interactive \
                --trust-all-tools
    
      weekly-modernization:
        if: github.event_name == 'schedule'
        runs-on: ubuntu-latest
        steps:
          - uses: actions/checkout@v4
          - name: Configure AWS credentials
            uses: aws-actions/configure-aws-credentials@v4
            with:
              role-to-assume: ${{ secrets.AWS_ROLE_ARN }}
              aws-region: us-east-1
          - name: Install ATX CLI
            run: curl -fsSL https://transform-cli.awsstatic.com/install.sh | bash
          - name: Run modernization scan
            run: |
              atx custom def exec \
                --transformation-name "tech-debt-analysis" \
                --code-repository-path "." \
                --build-command "mvn clean install" \
                --non-interactive \
                --trust-all-tools

    Conclusion

    Continuous modernization moves code transformation out of periodic sprints and into your CI/CD pipeline. By combining GitHub Dependabot’s vulnerability detection with AWS Transform custom agent, orchestrated through GitHub Actions, you can:

    • Remediate dependency vulnerabilities automatically — beyond version bumps to full code adaptation
    • Keep documentation current with every commit, preserving institutional knowledge
    • Scale transformations across hundreds of repositories with consistent quality
    • Improve continuously as the agent accumulates knowledge items from each execution

    The instrumentShop sample application demonstrates that even a moderately complex microservices architecture — with end-of-life Spring Boot versions, deprecated libraries like Hystrix, and multiple interconnected services — can be continuously modernized without dedicated modernization sprints.

    Ready to get started? This post walked through the do-it-yourself path with AWS Transform custom. If you would rather have continuous modernization delivered as a fully managed service, explore AWS Transform continuous modernization. Either way, visit the AWS Transform documentation to start your continuous modernization journey.

    Janardhan Molumuri

    Janardhan Molumuri is a Principal Technical Leader at AWS with over two decades of engineering leadership experience, advising customers on cloud and AI Adoption strategies and emerging technologies including generative AI. He has passion for thought leadership, speaking, writing, and enjoys exploring technology trends to solve problems at scale.

    Maxine Rosa

    Maxine Rosa is a Sr World Wide Generative AI Specialist at AWS focused on developer tooling including AWS Transform and Kiro. With a background in Software Engineering, Solution Engineering and Go-to-Market strategy, she helps AWS customers adopt Generative AI tooling into their current Software Development Lifecycle.

    Kola Akinnibi

    Kola Akinnibi is an Associate Solutions Architect at AWS focused on observability, partnering with ISVs and large enterprises to bring end-to-end monitoring to AI agents and modern applications. He helps customers design observability solutions that scale, and has a passion for sharing technical content.

    Renuka Krishnan

    Renuka Krishnan is a Senior Specialist Solutions Architect at AWS, specializing in code modernization using agentic AI and AWS services. She has over 15 years of experience architecting and implementing solutions, and works with customers to accelerate application development and modernization through AI-powered solutions.

    Venugopalan Vasudevan

    Venugopalan Vasudevan (Venu) is a Principal Specialist Solutions Architect at AWS, where he leads Generative AI initiatives focused on Amazon Q Developer, Kiro, and AWS Transform. He helps customers adopt and scale AI-powered developer and modernization solutions to accelerate innovation and business outcomes.

    Creating and testing an End User Messaging RCS agent with AWS CLI

    Post Syndicated from Bruno Giorgini original https://aws.amazon.com/blogs/messaging-and-targeting/creating-and-testing-an-end-user-messaging-rcs-agent-with-aws-cli/

    A step-by-step walkthrough for setting up a Rich Communication Services (RCS) test agent, from brand assets to verified inbound messaging.

    If you’re still sending plain SMS, you’re leaving a significant experience gap on the table. SMS gives you 160 characters of unformatted text, no branding, and zero confirmation that your message was even read. Rich Communication Services (RCS) changes that entirely. It delivers branded carousels, read receipts, typing indicators, high-resolution images, and verified sender identity, all through the native messaging app your customers already use. No app download required, no new account to create.

    Compared to over-the-top (OTT) platforms like WhatsApp or iMessage for Business, RCS doesn’t fragment your audience. It works on an Android’s default messaging app with RCS enabled or an iPhone on iOS 18 or later, which means you reach users where they already are. You are not limited to the ones who happen to have a specific app installed. And compared to building a custom in-app messaging experience, RCS requires no SDK, no UI work, and no convincing users to enable notifications.

    With AWS End User Messaging, standing up an RCS agent is surprisingly fast. You configure your brand assets, submit a registration, and within minutes you have a test agent sending branded messages through production APIs. This is real infrastructure, not a sandbox. That means you can prototype, validate your integration, and show stakeholders a working demo before committing to a full build.

    This post walks through the entire process of creating an RCS test agent using only the AWS Command Line Interface (AWS CLI). Using the CLI means every step is a repeatable, scriptable command. Need to spin up another agent in a different account or Region? Run the same script and you’re done in minutes. By the end, you will have a working agent that can send branded messages to verified testers and receive inbound messages with automatic responses.

    What you will build

    In this walkthrough, you will:

    1. Create an RCS agent and configure its brand identity (logo, banner, accent color).
    2. Submit a test registration for automated approval.
    3. Add a verified tester device.
    4. Send your first branded RCS message.
    5. Configure and verify inbound messaging with an automatic keyword response.

    Prerequisites

    Before you begin, confirm you have:

    • An AWS account with access to AWS End User Messaging (Amazon Pinpoint SMS and Voice v2 API)
    • AWS CLI v2.35.12 or later installed and configured with credentials that have pinpoint-sms-voice-v2:* permissions. Version 2.35.12 adds the send-rcs-message command, which you will need for rich media messages (rich cards, carousels, and suggestion chips) beyond this walkthrough. For production deployments, scope the IAM policy down to only the specific actions your application requires. The pinpoint-sms-voice-v2:* scope is convenient for testing but broader than necessary.
    • rsvg-convert for generating brand asset images from SVG (install with brew install librsvg on macOS)
    • A test phone that supports RCS messaging.

    Verify your setup:

    # Confirm AWS credentials are working
    aws sts get-caller-identity
    # Verify EUM access
    aws pinpoint-sms-voice-v2 describe-spend-limits --region us-east-1
    # Confirm rsvg-convert is installed
    which rsvg-convert

    If you use a named AWS CLI profile, append --profile <your-profile> to every AWS command in this walkthrough.

    Step 1: Create the RCS agent

    The first step is to create an empty RCS agent container. The agent’s display name and branding come from the registration you will configure in Step 2.

    aws pinpoint-sms-voice-v2 create-rcs-agent \
      --region us-east-1

    Expected output:

    {
      "RcsAgentArn": "arn:aws:sms-voice:us-east-1:123456789012:rcs-agent/rcs-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4",
      "RcsAgentId": "rcs-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4",
      "Status": "CREATED",
      "DeletionProtectionEnabled": false,
      "CreatedTimestamp": "2026-07-15T10:00:01.000000-07:00"
    }

    Save the RcsAgentId and RcsAgentArn values. You will use them throughout this walkthrough.

    Next, enable deletion protection to prevent accidental removal. This is especially important once carrier approvals are in place, since re-creating an agent requires a new registration and approval cycle:

    aws pinpoint-sms-voice-v2 update-rcs-agent \
      --rcs-agent-id rcs-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4 \
      --deletion-protection-enabled \
      --region us-east-1

    Step 2: Generate brand assets

    Your RCS agent needs a logo (224×224 px, must be under 50 KB as PNG) and a banner (1440×448 px, must be under 200 KB as PNG). Both must be JPEG or PNG format. You can use your own designs as long as they meet these dimension and size requirements. In this example, we generate them as SVGs and convert to PNG.

    Create the logo SVG

    Create a file named brand-assets/logo.svg:

    <svg xmlns="http://www.w3.org/2000/svg" width="224" height="224" viewBox="0 0 224 224"><defs><linearGradient id="bg" x1="0%" y1="0%" x2="100%" y2="100%"><stop offset="0%" style="stop-color:#0D47A1"/><stop offset="100%" style="stop-color:#1565C0"/></linearGradient></defs><rect width="224" height="224" rx="40" fill="url(#bg)"/><g transform="translate(112,90)"><path d="M-48,-36 L48,-36 C54,-36 58,-32 58,-26 L58,16 C58,22 54,26 48,26             L10,26 L0,42 L-10,26 L-48,26 C-54,26 -58,22 -58,16 L-58,-26             C-58,-32 -54,-36 -48,-36 Z" fill="white" opacity="0.95"/><path d="M-20,-12 C-14,-20 14,-20 20,-12" stroke="#0D47A1" stroke-width="4" fill="none" stroke-linecap="round"/><path d="M-14,-2 C-9,-8 9,-8 14,-2" stroke="#0D47A1" stroke-width="4" fill="none" stroke-linecap="round"/><circle cx="0" cy="6" r="4" fill="#0D47A1"/></g><text x="112" y="168" text-anchor="middle" font-family="Arial, Helvetica, sans-serif" font-size="16" font-weight="bold" fill="white">AWS EUM</text><text x="112" y="188" text-anchor="middle" font-family="Arial, Helvetica, sans-serif" font-size="12" fill="white" opacity="0.85">DEMO</text></svg>

    Create the banner SVG

    Create a file named brand-assets/banner.svg:

    <svg xmlns="http://www.w3.org/2000/svg" width="1440" height="448" viewBox="0 0 1440 448"><defs><linearGradient id="bannerBg" x1="0%" y1="0%" x2="100%" y2="100%"><stop offset="0%" style="stop-color:#0D47A1"/><stop offset="50%" style="stop-color:#1565C0"/><stop offset="100%" style="stop-color:#0D47A1"/></linearGradient></defs><rect width="1440" height="448" fill="url(#bannerBg)"/><circle cx="200" cy="224" r="300" fill="white" opacity="0.03"/><circle cx="1300" cy="100" r="250" fill="white" opacity="0.04"/><text x="720" y="190" text-anchor="middle" font-family="Arial, Helvetica, sans-serif" font-size="56" font-weight="bold" fill="white">    AWS End User Messaging  </text><text x="720" y="250" text-anchor="middle" font-family="Arial, Helvetica, sans-serif" font-size="48" font-weight="bold" fill="white" opacity="0.9">    Demo  </text><text x="720" y="320" text-anchor="middle" font-family="Arial, Helvetica, sans-serif" font-size="24" fill="white" opacity="0.7">    Rich messaging experiences, powered by AWS  </text></svg>

    Convert to PNG

    rsvg-convert -w 224 -h 224 brand-assets/logo.svg -o brand-assets/logo.png
    rsvg-convert -w 1440 -h 448 brand-assets/banner.svg -o brand-assets/banner.png

    Verify the file sizes. The logo must be under 50 KB and the banner under 200 KB:

    ls -la brand-assets/*.png
    # logo.png   ~9 KB
    # banner.png ~79 KB

    Step 3: Create and configure the registration

    RCS agents require a registration that contains all brand details. For testing, use the TEST_RCS_LAUNCH_REGISTRATION type.

    Create the registration

    aws pinpoint-sms-voice-v2 create-registration \
      --registration-type TEST_RCS_LAUNCH_REGISTRATION \
      --region us-east-1

    Expected output:

    {
      "RegistrationId": "registration-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4",
      "RegistrationType": "TEST_RCS_LAUNCH_REGISTRATION",
      "RegistrationStatus": "CREATED",
      "CurrentVersionNumber": 1
    }

    Save the RegistrationId.

    aws pinpoint-sms-voice-v2 create-registration-association \
      --registration-id registration-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4 \
      --resource-id rcs-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4 \
      --region us-east-1

    Upload brand assets

    Upload the logo and banner as registration attachments. Note that --attachment-body and --attachment-url cannot be used together. Use --attachment-body with the fileb:// prefix:

    # Upload logo
    aws pinpoint-sms-voice-v2 create-registration-attachment \
      --attachment-body fileb://brand-assets/logo.png \
      --region us-east-1
    # Save: RegistrationAttachmentId (e.g., attachment-1111aaaa2222bbbb3333cccc4444dddd)
    # Upload banner
    aws pinpoint-sms-voice-v2 create-registration-attachment \
      --attachment-body fileb://brand-assets/banner.png \
      --region us-east-1
    # Save: RegistrationAttachmentId (e.g., attachment-5555eeee6666ffff7777aaaa8888bbbb)

    Set registration fields

    The registration has 23 fields. Each field has a specific type that determines which CLI parameter to use:

    Field type CLI parameter Example
    TEXT --text-value --text-value "My Brand"
    SELECT --select-choices --select-choices "MULTI_USE"
    ATTACHMENT --registration-attachment-id --registration-attachment-id "attachment-abc123"

    Do not use --field-values. That parameter does not exist in this CLI.

    Set all the TEXT fields:

    REG_ID="registration-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4"
    REGION="us-east-1"
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.brandName" \
      --text-value "AWS End User Messaging Demo" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.senderDisplayName" \
      --text-value "AWS End User Messaging Demo" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.agentDescription" \
      --text-value "Experience the power of rich messaging with AWS End User Messaging" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.accentColor" \
      --text-value "#0D47A1" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.contactPhoneNumber" \
      --text-value "+12065550100" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.contactPhoneLabel" \
      --text-value "Call Us" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.contactEmailAddress" \
      --text-value "[email protected]" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.contactEmailLabel" \
      --text-value "Email Us" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.contactWebsite" \
      --text-value "https://www.example.com" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.contactWebsiteLabel" \
      --text-value "Visit Website" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.privacyPolicyUrl" \
      --text-value "https://www.example.com/privacy" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.privacyPolicyLabel" \
      --text-value "Privacy Policy" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.termsAndConditionsUrl" \
      --text-value "https://www.example.com/terms" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.termsAndConditionsLabel" \
      --text-value "Terms and Conditions" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.serviceName" \
      --text-value "AWS End User Messaging Demo RCS Agent" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.monthlyRcsVolume" \
      --text-value "1000" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "complianceKeywords.helpResponse" \
      --text-value "Reply STOP to opt out. For help, contact [email protected]" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "complianceKeywords.stopResponse" \
      --text-value "You have been unsubscribed. No more messages will be sent." \
      --region $REGION

    Set the SELECT fields. These use --select-choices instead of --text-value:

    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.useCase" \
      --select-choices "MULTI_USE" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.billingCategory" \
      --select-choices "CONVERSATIONAL" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.averageMonthlyRcsFrequency" \
      --select-choices "10" \
      --region $REGION

    Set the ATTACHMENT fields. These use --registration-attachment-id:

    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.logoImage" \
      --registration-attachment-id "attachment-1111aaaa2222bbbb3333cccc4444dddd" \
      --region $REGION
    
    aws pinpoint-sms-voice-v2 put-registration-field-value \
      --registration-id $REG_ID \
      --field-path "agentDetails.bannerImage" \
      --registration-attachment-id "attachment-5555eeee6666ffff7777aaaa8888bbbb" \
      --region $REGION

    A note on accent color

    The accent color must meet a 4.5:1 contrast ratio against white. This is the WCAG AA accessibility standard, enforced to make sure the text is readable for users with visual impairments. Colors with an HSL lightness value above ~45% will typically fail this threshold and be rejected with ACCENT_COLOR_CONTRAST_INSUFFICIENT. Safe choices include #0D47A1 (blue), #1B5E20 (green), #BF360C (orange), #B71C1C (red), and #4A148C (purple). If you are using a custom brand color, verify it passes before submitting using the WebAIM Contrast Checker.

    Submit the registration

    aws pinpoint-sms-voice-v2 submit-registration-version \
      --registration-id $REG_ID \
      --region $REGION

    Expected output:

    {
      "RegistrationId": "registration-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4",
      "VersionNumber": 1,
      "RegistrationVersionStatus": "SUBMITTED"
    }

    Step 4: Wait for approval

    Poll the registration and agent status. Test registrations typically complete within a few minutes.

    # Check registration status
    aws pinpoint-sms-voice-v2 describe-registrations \
      --registration-ids $REG_ID \
      --query 'Registrations[0].{Status:RegistrationStatus,Version:CurrentVersionNumber}' \
      --region $REGION
    
    # Check agent status
    aws pinpoint-sms-voice-v2 describe-rcs-agents \
      --query "RcsAgents[?RcsAgentId=='rcs-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4'].{Status:Status,TestingStatus:TestingAgent.Status}" \
      --region $REGION

    You will see the status progress through these stages:

    Registration status Agent status Testing status Meaning
    SUBMITTED PENDING PENDING Under review
    REVIEWING PENDING PENDING Automated checks in progress
    COMPLETE TESTING ACTIVE Ready to use

    Wait until TestingAgent.Status shows ACTIVE before proceeding.

    NOTE: If the registration returns REQUIRES_UPDATES, run describe-registration-field-values to find fields with a DeniedReason. Create a new registration version with create-registration-version, re-populate all 23 fields (new versions do not inherit values), fix the issue, and re-submit.

    Step 5: Add a verified tester

    Wait at least 120 seconds after agent creation before adding testers. Then register your test device:

    aws pinpoint-sms-voice-v2 create-verified-destination-number \
      --destination-phone-number +12065550199 \
      --rcs-agent-id rcs-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4 \
      --region $REGION

    You will receive a tester invitation on your phone within 2 to 20 minutes from “RBM Tester Management.” On iPhone, check the Unknown Senders folder. Tap “Make me a tester” to accept.

    After accepting, verify the status:

    aws pinpoint-sms-voice-v2 describe-verified-destination-numbers \
      --filters Name=rcs-agent-id,Values=rcs-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4 \
      --region $REGION \
      --query 'VerifiedDestinationNumbers[].{Phone:DestinationPhoneNumber,Status:Status}'

    Expected output once accepted:

    [
      {
        "Phone": "+12065550199",
        "Status": "VERIFIED"
      }
    ]

    Step 6: Send your first RCS message

    Before sending, check for potential blockers.

    Check the protect configuration

    Verify that the US is not blocked in your account’s default protect configuration:

    # List protect configurations
    aws pinpoint-sms-voice-v2 describe-protect-configurations --region $REGION
    
    # Check US status on the default (account-default) protect configuration
    aws pinpoint-sms-voice-v2 get-protect-configuration-country-rule-set \
      --protect-configuration-id <your-protect-config-id> \
      --number-capability SMS \
      --query 'CountryRuleSet.US' \
      --region $REGION

    If the US status is BLOCK, update it to ALLOW:

    aws pinpoint-sms-voice-v2 update-protect-configuration-country-rule-set \
      --protect-configuration-id <your-protect-config-id> \
      --country-rule-set-updates '{"US":{"ProtectStatus":"ALLOW"}}' \
      --number-capability SMS \
      --region $REGION

    Check the opt-out list

    aws pinpoint-sms-voice-v2 describe-opted-out-numbers \
      --opt-out-list-name Default \
      --region $REGION

    If your test number appears in the list, remove it:

    aws pinpoint-sms-voice-v2 delete-opted-out-number \
      --opt-out-list-name Default \
      --opted-out-number +12065550199 \
      --region $REGION

    Now, send the test message:

    aws pinpoint-sms-voice-v2 send-text-message \
      --destination-phone-number +12065550199 \
      --origination-identity rcs-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4 \
      --message-body "Hello from AWS End User Messaging Demo! This is your first RCS test message." \
      --message-type TRANSACTIONAL \
      --region $REGION

    Expected output:

    {"MessageId": "a1b2c3d4-e5f6-7890-abcd-ef1234567890"}

    Check your phone. You should see a branded message from your agent with the logo and accent color you configured. On iPhone, check the Unknown Senders folder.

    Step 7: Configure and test inbound messaging

    With inbound messaging, your agent can respond to messages that testers send back. Configure an automatic keyword response, then verify it end to end.

    Set up an automatic keyword response

    The put-keyword API configures an automatic reply when someone sends a specific keyword to your agent. With it, you can verify inbound messaging without writing any backend code:

    aws pinpoint-sms-voice-v2 put-keyword \
      --keyword RCSINBOUNDTESTING \
      --keyword-action AUTOMATIC_RESPONSE \
      --keyword-message "Inbound test successful! Your message was received." \
      --origination-identity rcs-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4 \
      --region $REGION

    Test inbound messaging

    While the previous steps used the CLI exclusively, the inbound testing deep link is most easily accessed through the console. Navigate to your agent and use the Testing tab to generate the deep link:

    1. Open the AWS End User Messaging console: https://console.aws.amazon.com/sms-voice/home?region=[REGION]#/rcs-agents.
    2. Select your agent and choose the Testing tab.
    3. Choose Inbound deep link.
    4. Enter RCSINBOUNDTESTING in the message body field.
    5. Choose Generate link.
    6. Scan the QR code with your test phone. The message is pre-filled.
    7. Send the message.

    You should receive the automatic response: “Inbound test successful! Your message was received.”

    Clean up

    To avoid unexpected charges, remove the resources created during this walkthrough when you are finished testing. You must delete resources in the following order. Attempting to delete the agent before its registration results in a ConflictException: RESOURCE_NOT_EMPTY error.

    # 1. Remove the keyword
    aws pinpoint-sms-voice-v2 delete-keyword \
      --keyword RCSINBOUNDTESTING \
      --origination-identity rcs-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4 \
      --region $REGION
    
    # 2. Remove verified tester
    aws pinpoint-sms-voice-v2 delete-verified-destination-number \
      --verified-destination-number-id <your-verified-number-id> \
      --region $REGION
    
    # 3. Disable deletion protection
    aws pinpoint-sms-voice-v2 update-rcs-agent \
      --rcs-agent-id rcs-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4 \
      --no-deletion-protection-enabled \
      --region $REGION
    
    # 4. Delete the registration
    aws pinpoint-sms-voice-v2 delete-registration \
      --registration-id registration-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4 \
      --region $REGION
    
    # 5. Delete the agent
    aws pinpoint-sms-voice-v2 delete-rcs-agent \
      --rcs-agent-id rcs-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4 \
      --region $REGION

    If you modified the protect configuration (changed US from BLOCK to ALLOW), revert it to its original state if your account does not need US messaging enabled.

    Summary

    You now have a working RCS test agent that can send and receive branded messages. Here is a recap of the resources created:

    Resource Value
    Agent ID rcs-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4
    Registration ID registration-a1b2c3d4e5f6a1b2c3d4e5f6a1b2c3d4
    Region us-east-1
    Console https://us-east-1.console.aws.amazon.com/sms-voice/home?region=us-east-1#/rcs-agents

    Registration field reference

    For reference, here is the complete list of registration fields and their types:

    Field Type Requirement
    agentDetails.brandName TEXT Required
    agentDetails.serviceName TEXT Required
    agentDetails.senderDisplayName TEXT Required
    agentDetails.useCase SELECT Required
    agentDetails.agentDescription TEXT Required
    agentDetails.bannerImage ATTACHMENT Required
    agentDetails.logoImage ATTACHMENT Required
    agentDetails.accentColor TEXT Required
    agentDetails.contactPhoneNumber TEXT Conditional
    agentDetails.contactPhoneLabel TEXT Conditional
    agentDetails.contactEmailAddress TEXT Conditional
    agentDetails.contactEmailLabel TEXT Conditional
    agentDetails.contactWebsite TEXT Conditional
    agentDetails.contactWebsiteLabel TEXT Conditional
    agentDetails.privacyPolicyUrl TEXT Required
    agentDetails.privacyPolicyLabel TEXT Optional
    agentDetails.termsAndConditionsUrl TEXT Required
    agentDetails.termsAndConditionsLabel TEXT Optional
    agentDetails.averageMonthlyRcsFrequency SELECT Required
    agentDetails.billingCategory SELECT Required
    agentDetails.monthlyRcsVolume TEXT Required
    complianceKeywords.helpResponse TEXT Conditional
    complianceKeywords.stopResponse TEXT Conditional

    Troubleshooting

    Error Resolution
    ACCENT_COLOR_CONTRAST_INSUFFICIENT Use a darker accent color with 4.5:1 contrast ratio against white. Create a new registration version and re-populate all fields.
    DESTINATION_COUNTRY_BLOCKED_BY_PROTECT_CONFIGURATION Update the protect configuration to set the US to ALLOW for SMS capability.
    DESTINATION_PHONE_NUMBER_OPTED_OUT Remove the number from the Default opt-out list with delete-opted-out-number.
    Registration REQUIRES_UPDATES Run describe-registration-field-values to find fields with DeniedReason. Create a new version, re-populate all 23 fields, fix the issue, and re-submit.
    No tester invitation received Wait up to 20 minutes. Check the Unknown Senders folder on iPhone. Verify the agent status is ACTIVE.
    Message delivered as SMS instead of RCS Confirm the agent is ACTIVE, the device supports RCS, and you used the correct origination identity.

    Next steps

    With your test agent running, you can explore richer message types such as cards and carousels, set up event destinations for programmatic inbound message handling, or add more verified testers. For production use, submit a full launch registration instead of a test registration.

    For an overview of the business case for RCS and implementation strategy, see Upgrade business messaging with RCS on AWS. For sample code and scripts that automate this walkthrough, see the sample-rcs-agent-setup-and-send-messages repository on GitHub. For more information, see the AWS End User Messaging service page and the RCS documentation.


    About the authors

    Detecting multi-stage attacks on AWS: A guide to cross-service signal correlation

    Post Syndicated from Nisha Kashyap original https://aws.amazon.com/blogs/security/detecting-multi-stage-attacks-on-aws-a-guide-to-cross-service-signal-correlation/

    A single alert from one security service tells you something happened. Read that signal alongside activity from other services and your own business context, and you will know whether what happened is part of a multi-stage attack.

    Consider a short sequence. An identity calls GetCallerIdentity from a source address it hasn’t previously used. Within minutes, that same identity runs a burst of List and Describe calls across several services, and some of them fail with AccessDenied. Soon after, a large volume of data leaves your environment toward a domain that was registered last week. Amazon GuardDuty might already flag pieces of this, such as the reconnaissance from an unfamiliar source, through finding types like Recon:IAMUser/* or Discovery:S3/*. What you gain from correlating the pieces yourself is a single view of the sequence, tied to your own business context, so you can act on the whole rather than triaging findings one at a time.

    This post is for security engineers and security operations teams who run Amazon Web Services (AWS) detection services and want to catch patterns specific to their environment. You will see how AWS detection and your business context fit together, and how to build correlations that use that context. The examples run in Amazon CloudWatch Logs Insights so you can try them today, and the closing section describes how to grow them into an automated pipeline. The walkthrough later in this post lists the prerequisites for these queries.

    Start with AWS detection services

    Begin with the AWS detection services. They cover the threats common across customers, and everything in this post is built on them.

    Turn these on and tune them before you build anything custom. Tuning means adjusting sensitivity to reduce false positives for your environment, choosing which data sources each service monitors, and suppressing findings for known-good patterns.

    GuardDuty correlates multi-stage attacks for you

    Before you build anything by hand, see what GuardDuty already does for you. Amazon GuardDuty Extended Threat Detection correlates signals across multiple data sources including AWS CloudTrail, Amazon S3 data events, runtime monitoring, Amazon Elastic Kubernetes Service (Amazon EKS) audit logs, and more, then raises a single critical severity attack sequence finding when it spots a multi-stage pattern. It recognizes sequences such as credential compromise followed by data exfiltration, maps them to MITRE ATT&CK tactics, and attaches a timeline and remediation guidance. If you have GuardDuty enabled today, then GuardDuty Extended Threat Detection is already enabled by default and needs no queries from you. For details on how GuardDuty charges apply, see Amazon GuardDuty pricing.

    The credential compromise sequence in the opening example is the kind of universal pattern GuardDuty Extended Threat Detection is built to catch, so rely on it for those. Attack sequence findings show up in the GuardDuty console next to your other findings, and they route to Security Hub and your response workflows the same way.

    GuardDuty handles the threats that look the same in every account. What it doesn’t have is the context that makes a given action suspicious in your account. That’s what you provide.

    Add your business context

    Business context is what only you know about your environment: which buckets hold sensitive data, which principals have a reason to touch which resources, which role chains your policy permits, and when your production change windows open. GuardDuty Extended Threat Detection learns from patterns common across customers, but it can’t answer these environment-specific questions. Express them as correlations and you add a detection layer tuned to your environment. Each of the following four patterns turns one of these facts into a query.

    Run these queries in the AWS Management Console for CloudWatch by choosing Logs, then Logs Insights, using the CloudWatch Logs Insights query language. Most read CloudTrail events from a CloudWatch Logs log group that your trail delivers to. If your trail writes only to Amazon S3, add CloudWatch Logs delivery on the trail, or run equivalent queries in Amazon Athena (a serverless query service for analyzing data in Amazon S3 using SQL).

    Note: The queries and code in this post use placeholder values. Replace them with your own before running: your-sensitive-bucket (your S3 bucket name), your-key-id (your AWS KMS key ID), region (your AWS Region, such as us-east-1), account-id (your 12-digit AWS account ID), and aws-cloudtrail-logs-my-trail (your CloudTrail log group name).

    A note on multi-account environments. In AWS Organizations, an organization trail delivers every account’s events to one log group, so these queries work as-is but return cross-account results. Filter by recipientAccountId for account-scoped views. Without an organization trail, run queries per account or use Amazon Security Lake as a central query surface.

    The attack chain mapped to AWS services

    Multi-stage attacks move through five phases, and each phase leaves a signal in a different service. These signals surface across three log sources: CloudTrail, which records API activity in your account; Amazon VPC Flow Logs, which capture network connection metadata; and Amazon Route 53 Resolver query logs, which record DNS queries from your VPCs.

    • Initial access – Stolen credentials reach your environment. CloudTrail records GetCallerIdentity, GetSessionToken, or AssumeRole from an unfamiliar source.
    • Discovery – The threat actor enumerates with List, Describe, and Get calls, often triggering AccessDenied responses.
    • Privilege escalation – The threat actor chains roles or edits policies. CloudTrail records AssumeRole sequences, PutRolePolicy, or CreateAccessKey.
    • Lateral movement – The threat actor moves across accounts or AWS Regions, assuming roles and creating resources in unfamiliar places.
    • Exfiltration – Data leaves through GetObject calls at scale, large outbound transfers in VPC Flow Logs, and DNS queries in Route 53 Resolver query logs to recently registered domains.

    Figure 1 shows the five attack phases mapped to the AWS log source that records each one.

    Figure 1: Attack chain mapped to AWS services

    Figure 1: Attack chain mapped to AWS services

    GuardDuty Extended Threat Detection watches this chain for universal patterns. The four patterns that follow add the dimension you supply: your business context.

    Pattern one: Sensitive data access by an unexpected principal

    Your data classification and access norms drive this detection. One bucket holds customer records, another holds public web assets, and you know which principals have a reason to read the customer records, which are sensitive. Encode that knowledge and an ordinary looking read turns into something worth chasing.

    Three signals converge here. CloudTrail shows GetObject at volume on a bucket you’ve classified as sensitive. The principal isn’t on your list of expected readers for that bucket. And VPC Flow Logs show a large outbound transfer from the same source in the same window, while DNS query logs show a recently registered destination domain, which together increase your confidence that there’s a potential threat.

    CloudTrail management events don’t record GetObject. You must turn on CloudTrail data events for the buckets you care about to capture GetObject. Many teams miss GetObject because data events weren’t enabled on the relevant buckets.

    This query shows bulk reads on a sensitive bucket, grouped by principal. Run it in CloudWatch Logs Insights with your CloudTrail log group selected.

    fields @timestamp, userIdentity.arn, requestParameters.bucketName
    | filter eventSource = "s3.amazonaws.com" and eventName = "GetObject"
    | filter requestParameters.bucketName = "your-sensitive-bucket"
    | stats count(*) as objectReads,
            count_distinct(requestParameters.key) as distinctObjects
            by userIdentity.arn, bin(10m)
    | filter objectReads > 100
    | sort objectReads desc
    

    The threshold of 100 is a placeholder. Run the query over a week of normal activity, find the ninety-fifth percentile read count for that bucket, and set the threshold above it. Then check each principal the query returns against your expected reader list. A principal that isn’t on the list, reading at volume, is the result to investigate.

    To corroborate, look for a matching outbound transfer. Switch the log group selector to your VPC Flow Logs log group and run this.

    fields @timestamp, srcAddr, dstAddr, bytes
    | filter action = "ACCEPT"
    # exclude RFC 1918 private ranges so only external destinations remain
    | filter dstAddr not like /^10\./
            and dstAddr not like /^192\.168\./
            and dstAddr not like /^172\.(1[6-9]|2[0-9]|3[0-1])\./
    | stats sum(bytes) as totalBytes by srcAddr, dstAddr, bin(10m)
    | filter totalBytes > 1000000000
    | sort totalBytes desc

    The Amazon S3 query returns a principal, and the Flow Logs query works on IP addresses, so you translate one into the other. The worked example later in this post covers that translation in full.

    Picture an analytics role that reads a reporting bucket all day. One afternoon, it reads a thousand objects from your customer records bucket instead. GuardDuty stays quiet, because an authenticated role making valid GetObject calls isn’t suspicious anywhere else. Your query flags it, because that role isn’t on the expected reader list for that bucket. The classification you applied is what turns silence into a signal.

    Figure 2 shows a bulk read from a sensitive bucket in CloudTrail, a large outbound transfer in VPC Flow Logs, and a young domain resolution in Route 53 Resolver logs.

    Figure 2: Three signals converging within a single time window to indicate exfiltration

    Figure 2: Three signals converging within a single time window to indicate exfiltration

    Pattern two: A role chain that crosses your access policy

    Picture a deployment that assumes one role to build, then a second to release. For one principal, that two-hop AssumeRole chain is routine; for a different principal it’s a policy violation. This pattern relies on your trust topology—the chains your organization permits—so put that knowledge in the query.

    This pattern needs three conditions:

    • CloudTrail shows several AssumeRole calls from the same source inside a short window
    • The chain ends in a sensitive action such as CreateAccessKey, PutRolePolicy, or AttachUserPolicy
    • The starting identity isn’t one your policy expects to run that chain

    In CloudWatch Logs Insights, select your CloudTrail log group and run this query, which surfaces chains of two or more hops.

    fields @timestamp, userIdentity.arn, requestParameters.roleArn, sourceIPAddress
    | filter eventName = "AssumeRole"
    | stats count(*) as assumeCount,
            count_distinct(requestParameters.roleArn) as rolesAssumed
            by sourceIPAddress, bin(5m)
    | filter assumeCount >= 2 and rolesAssumed >= 2
    | sort assumeCount desc

    Two hops is the minimum for a chain; raise the count if your environment chains roles often. Your deployment pipeline probably assumes several roles an hour, as do AWS service principals such as AWS Security Hub. Exclude the identities you expect to see assuming multiple roles, including your pipeline role and known AWS service principals. What’s left is the set to investigate, such as a person assuming several roles at an odd hour and ending in a new access key. Treat that distinction as data: list the identities and actions you consider normal, and review the chains that fall outside the list.

    Pattern three: An encryption key used outside its owning workload

    Resource ownership is the signal here. A given AWS Key Management Service (AWS KMS) key creates and controls the encryption keys for a workload, and a single key should serve a single workload, such as a payments service. A Decrypt call against it is a valid, authorized API action, so nothing about the call itself looks wrong. The ownership rule you set is what makes another principal’s use of the key worth a second look.

    This pattern applies only to customer-managed keys scoped to one workload. It doesn’t apply to AWS-managed keys (alias/aws/*) or to customer-managed keys intentionally shared across services. Confirm single-workload intent from the key policy’s Principal block before deploying this rule.

    Two conditions indicate misuse:

    • CloudTrail shows Decrypt or GenerateDataKey calls on a key that’s tied to one workload
    • The calling principal isn’t the role that owns that workload

    Against your CloudTrail log group, run this query to list the principals that called a specific key.

    fields @timestamp, userIdentity.arn, eventName
    | filter eventSource = "kms.amazonaws.com"
    | filter eventName in ["Decrypt", "GenerateDataKey", "Encrypt"]
    | filter resources.0.ARN = "arn:aws:kms:region:account-id:key/your-key-id"
    | stats count(*) as keyUses by userIdentity.arn, eventName
    | sort keyUses desc

    Compare what comes back against the one workload role you expect. A principal you don’t recognize on that key is the signal. Because key misuse is an early move in data theft, this correlation catches activity that only your ownership knowledge can flag.

    Consider a key that wraps your payments database. The payments service role calls it in normal operation, and nothing else should. If a developer role or a freshly created role runs Decrypt against it, the call succeeds and reads as ordinary in isolation. The reason it matters is the ownership rule you hold in your head and now state in this query.

    Pattern four: A privileged action outside your change window

    Start with the query, then read what it means.

    fields @timestamp, userIdentity.arn, eventName, sourceIPAddress
    | filter eventName in ["PutRolePolicy", "AttachRolePolicy",
            "CreateAccessKey", "AuthorizeSecurityGroupIngress", "PutBucketPolicy"]
    | stats count(*) as sensitiveChanges by userIdentity.arn, eventName, sourceIPAddress
    | sort sensitiveChanges desc

    Run it against your CloudTrail log group, scoped to your off-hours window when you schedule it, so it returns only activity outside the change window. Your change process defines what normal looks like here: production security and identity changes flow through a pipeline during defined hours, run by a known actor. A console-driven policy change at 2:00 AM, made by a person rather than the pipeline, doesn’t fit those expectations. The signal is a sensitive change such as PutRolePolicy or AuthorizeSecurityGroupIngress, made outside the window, by a person rather than your pipeline role.

    Exclude the actors you expect, such as your deployment pipeline role, your patch automation role, and AWS service principals like AWS CloudFormation and AWS Systems Manager. What remains is privileged change made outside your process, which is both what an attacker does to establish persistence and what your own change discipline says shouldn’t happen.

    Your pipeline might open security group rules during a deployment every weekday afternoon. A person opening a security group rule at midnight on a weekend is the same API call carrying a very different meaning. The schedule and the actor, both facts you define, are what separate the two.

    Build your first correlation rule

    The following walkthrough uses pattern one as a complete example. The other three patterns follow the same design with their own queries.

    Prerequisites

    These prerequisites feed the queries in this walkthrough. Confirm each one before you start:

    • A CloudTrail trail logging management events to a CloudWatch Logs log group
    • CloudTrail data events enabled for your sensitive S3 buckets
    • GuardDuty enabled, with its protection plans and Extended Threat Detection
    • VPC Flow Logs on for your production VPCs
    • Amazon Route 53 Resolver query logging on

    CloudTrail, GuardDuty, VPC Flow Logs, and Route 53 Resolver query logging provide the raw signals that your correlations connect. Without them, the queries in this post return empty results.

    Step 1: Record the bucket and its expected readers

    Choose one sensitive bucket to monitor, and write down the principals allowed to read it. Store the list where your automation can reach it, such as a configuration file in version control or an Amazon DynamoDB table (a managed NoSQL database).

    {
      "customer-records-prod": [
        "arn:aws:iam::123456789012:role/AnalyticsPipeline",
        "arn:aws:iam::123456789012:role/ComplianceAudit"
      ],
      "financial-data-archive": [
        "arn:aws:iam::123456789012:role/FinanceReporting"
      ]
    }

    This example hardcodes the list for simplicity. In production, load it from a DynamoDB table or Parameter Store so you can update it without redeploying.

    Step 2: Baseline before you set a threshold

    Run the pattern one query over one week of normal activity. Find the 95th percentile read count for the bucket and use a value greater than that as your alert threshold. This step keeps legitimate high-volume access from generating false positives later.

    Set the THRESHOLD_READS environment variable to this value when you configure the function in Step 5.

    Step 3: Run the access query

    In the CloudWatch console:

    1. Choose Logs, then choose Logs Insights.
    2. In the Select log group(s) dropdown, select your CloudTrail log group.
    3. Set the time range to 3h (the last three hours).
    4. In the query editor, paste the pattern one query.
    5. Replace your-sensitive-bucket with your bucket name.
    6. Choose Run query.
    7. Review the principals in the results table.
    8. Compare each principal against your expected reader list from step 1, and flag any that are not on it.

    Each result includes a principal that step 4 translates into an IP address.

    Step 4: Correlate with network activity

    CloudTrail logs actions by AWS Identity and Access Management (IAM) principal, while VPC Flow Logs record traffic by IP address. To connect the two signals, translate the principal into its address.

    For a role attached to an Amazon Elastic Compute Cloud (Amazon EC2) instance, the userIdentity.principalId field includes the instance ID after the colon, in the form AROAEXAMPLE:i-1234567890abcdef0. Copy the instance ID and look up its private IP address.

    aws ec2 describe-instances \
      --instance-ids i-1234567890abcdef0 \
      --query "Reservations[0].Instances[0].PrivateIpAddress" \
      --output text

    Other compute types differ. A VPC-connected AWS Lambda function sends traffic through elastic network interfaces in your subnets, so correlate on those interface addresses. An Amazon Elastic Container Service (Amazon ECS) task records its network interface in task metadata. For a plain assumed-role session with no instance behind it, the sourceIPAddress field in CloudTrail already holds the caller’s address, so you correlate on it directly.

    Run the Flow Logs query from pattern one, filtering srcAddr to that address within 10 minutes of the Amazon S3 read timestamp. A match places the same source behind both the sensitive read and a large external transfer in one window. CloudTrail events reach CloudWatch Logs 5–15 minutes after the API call, so correlate on eventTime rather than query time. Query a wider lookback than your correlation window: for example, look back 30 to 60 minutes but correlate on a 10-minute eventTime window. Steps 3 and 4 are manual validation; step 5 automates them.

    Figure 2 shows DNS resolution as a third corroborating signal. This walkthrough implements the CloudTrail and VPC Flow Logs correlation. To add DNS, apply the same run_query() pattern against your Route 53 Resolver query log group.

    Step 5: Automate the check

    Move the query into a Lambda function (serverless compute that runs your code without a server to manage), send results to a notification channel, and schedule regular runs. Work through the following sub-procedures.

    To create the notification channel

    1. Open the Amazon Simple Notification Service (Amazon SNS) console. Amazon SNS is a managed messaging service that delivers notifications to subscribers.
    2. In the navigation pane, choose Topics.
    3. Choose Create topic.
    4. For Type, select Standard.
    5. For Name, enter security-correlation-alerts.
    6. Choose Create topic.
    7. Note the topic Amazon Resource Name (ARN) at the top of the topic details page. You will use it in the function.
    8. Choose Create subscription.
    9. For Protocol, select Email.
    10. For Endpoint, enter your email address or incident management endpoint.
    11. Choose Create subscription, then confirm the subscription from the email AWS sends.

    To create the EventBridge Scheduler execution role

    The schedule needs a role that lets it invoke your function, and its trust policy needs conditions that pin the role to the schedule you own. Without those conditions, another account with access to the scheduler service could theoretically call this role; a class of misuse known as the confused deputy problem.

    1. Create a trust policy file named scheduler-trust-policy.json.

    {
      "Version": "2012-10-17",
      "Statement": [
        {
          "Effect": "Allow",
          "Principal": { "Service": "scheduler.amazonaws.com" },
          "Action": "sts:AssumeRole",
          "Condition": {
            "StringEquals": {
              "aws:SourceAccount": "ACCOUNT-ID"
            },
            "ArnLike": {
              "aws:SourceArn": "arn:aws:scheduler:REGION:ACCOUNT-ID:schedule/*/s3-access-correlation-hourly"
            }
          }
        }
      ]
    }
    

    2. Create the role, then attach permission to invoke the function. Scope Resource to the specific function ARN so this role can’t invoke anything else.

    aws iam create-role \
      --role-name EventBridgeSchedulerRole \
      --assume-role-policy-document file://scheduler-trust-policy.json
    
    aws iam put-role-policy \
      --role-name EventBridgeSchedulerRole \
      --policy-name LambdaInvokePolicy \
      --policy-document '{
        "Version": "2012-10-17",
        "Statement": [
          {
            "Effect": "Allow",
            "Action": "lambda:InvokeFunction",
            "Resource": "arn:aws:lambda:REGION:ACCOUNT-ID:function:CorrelationFunction"
          }
        ]
      }'

    When you create the function, Lambda automatically creates an execution role. You will attach the permissions this function needs to that role in a later step.

    To deploy the correlation function

    1. Open the Lambda console.
    2. Choose Create function.
    3. For Function name, enter CorrelationFunction.
    4. For Runtime, select the latest Python runtime.
    5. Choose Create function.
    6. On the Code tab, replace the default code with the following function, then choose Deploy.
    import os
    import time
    import logging
    import boto3
    from botocore.exceptions import ClientError
    
    logger = logging.getLogger()
    logger.setLevel(logging.INFO)
    
    logs = boto3.client("logs")
    sns = boto3.client("sns")
    ec2 = boto3.client("ec2")
    
    CLOUDTRAIL_LOG_GROUP = os.environ["CLOUDTRAIL_LOG_GROUP"]
    FLOWLOGS_LOG_GROUP = os.environ["FLOWLOGS_LOG_GROUP"]
    SNS_TOPIC = os.environ["SNS_TOPIC_ARN"]
    BUCKET = os.environ["SENSITIVE_BUCKET"]
    THRESHOLD = int(os.environ.get("THRESHOLD_READS", "100"))
    
    # Expected readers per bucket
    EXPECTED_READERS = {
        "customer-records-prod": [
            "arn:aws:iam::123456789012:role/AnalyticsPipeline",
            "arn:aws:iam::123456789012:role/ComplianceAudit",
        ],
    }
    
    
    def run_query(log_group, query, start, end):
        """Start a Logs Insights query and wait for it to finish."""
        started = logs.start_query(
            logGroupName=log_group,
            startTime=start,
            endTime=end,
            queryString=query,
        )
        query_id = started["queryId"]
        while True:
            outcome = logs.get_query_results(queryId=query_id)
            if outcome["status"] in ("Complete", "Failed", "Cancelled"):
                break
            time.sleep(1)
        if outcome["status"] != "Complete":
            raise RuntimeError(f"Query did not complete: {outcome['status']}")
        return [{f["field"]: f["value"] for f in row} for row in outcome["results"]]
    
    
    def private_ip_for_principal(principal_id):
        """Resolve an EC2 instance role principalId to its private IP."""
        if ":" not in principal_id:
            return None
        instance_id = principal_id.split(":", 1)[1]
        if not instance_id.startswith("i-"):
            return None
        reservations = ec2.describe_instances(InstanceIds=[instance_id])
        for reservation in reservations["Reservations"]:
            for instance in reservation["Instances"]:
                return instance.get("PrivateIpAddress")
        return None
    
    
    def egress_bytes(src_addr, start, end):
        """Sum external egress bytes for one source address."""
        query = f"""
        fields srcAddr, dstAddr, bytes
        | filter action = "ACCEPT" and srcAddr = "{src_addr}"
        | filter dstAddr not like /^10\\./
                and dstAddr not like /^192\\.168\\./
                and dstAddr not like /^172\\.(1[6-9]|2[0-9]|3[0-1])\\./
        | stats sum(bytes) as totalBytes
        """
        rows = run_query(FLOWLOGS_LOG_GROUP, query, start, end)
        if rows and rows[0].get("totalBytes"):
            return int(rows[0]["totalBytes"])
        return 0
    
    
    def lambda_handler(event, context):
        try:
            # 1-hour lookback absorbs CloudTrail's 5-15 min delivery latency;
            # correlation happens on eventTime via 10-min bins in the query below.
            end = int(time.time())
            start = end - 3600  # 1 hour lookback
            allowed = EXPECTED_READERS.get(BUCKET, [])
    
            access_query = f"""
            fields userIdentity.arn, userIdentity.principalId
            | filter eventSource = "s3.amazonaws.com" and eventName = "GetObject"
            | filter requestParameters.bucketName = "{BUCKET}"
            | stats count(*) as objectReads
                    by userIdentity.arn, userIdentity.principalId, bin(10m)
            | filter objectReads > {THRESHOLD}
            """
    
            for row in run_query(CLOUDTRAIL_LOG_GROUP, access_query, start, end):
                principal = row.get("userIdentity.arn")
                if not principal or principal in allowed:
                    continue
    
                message = (
                    f"Principal {principal} read {row.get('objectReads')} "
                    f"objects from {BUCKET}."
                )
    
                ip = private_ip_for_principal(row.get("userIdentity.principalId", ""))
                if ip and egress_bytes(ip, start, end) > 1_000_000_000:
                    message += (
                        f" The same source ({ip}) also sent a large volume of "
                        f"data to external destinations in the same window."
                    )
    
                sns.publish(
                    TopicArn=SNS_TOPIC,
                    Subject="Unexpected S3 access detected",
                    Message=message,
                )
        except ClientError as error:
            logger.error(f"AWS API error: {error}")
            raise
        except Exception as error:
            logger.error(f"Unexpected error: {error}")
            raise
        finally:
            logger.info("Correlation check completed")

    1. On the Configuration tab, choose General configuration, then choose Edit. Set Timeout to 5 minutes (300 seconds). CloudWatch Logs Insights queries run asynchronously and can take 30 to 60 seconds against large log groups. Choose Save.
    2. On the Configuration tab, choose Environment variables, then choose Edit, and add CLOUDTRAIL_LOG_GROUP, FLOWLOGS_LOG_GROUP, SNS_TOPIC_ARN, SENSITIVE_BUCKET, and THRESHOLD_READS.
    3. On the Configuration tab, choose Permissions, open the execution role, and attach the following least-privilege policy.
    {
      "Version": "2012-10-17",
      "Statement": [
        {
          "Effect": "Allow",
          "Action": ["logs:StartQuery", "logs:GetQueryResults"],
          "Resource": [
            "arn:aws:logs:REGION:ACCOUNT-ID:log-group:aws-cloudtrail-logs-my-trail:*",
            "arn:aws:logs:REGION:ACCOUNT-ID:log-group:vpc-flow-logs:*"
          ]
        },
        {
          "Effect": "Allow",
          "Action": "ec2:DescribeInstances",
          "Resource": "*"
        },
        {
          "Effect": "Allow",
          "Action": "sns:Publish",
          "Resource": "arn:aws:sns:REGION:ACCOUNT-ID:security-correlation-alerts"
        }
      ]
    }

    Replace REGION, ACCOUNT-ID, and the log-group names with your values. The ec2:DescribeInstances action doesn’t support resource-level permissions, so Resource: "*" is required for that statement; the other statements are scoped to specific ARNs.

    To schedule automated runs

    Amazon EventBridge (a serverless event bus that connects applications using events) runs targets on a schedule. Create one from the command line, using the role you made earlier.

    aws scheduler create-schedule \
      --name s3-access-correlation-hourly \
      --schedule-expression "rate(1 hour)" \
      --target "Arn=arn:aws:lambda:REGION:ACCOUNT-ID:function:CorrelationFunction,RoleArn=arn:aws:iam::ACCOUNT-ID:role/EventBridgeSchedulerRole" \
      --flexible-time-window "Mode=OFF"

    Step 6: Add enrichment context (optional)

    Enrichment cuts triage time by adding an independent signal, but it isn’t required for the correlation to work. This step adds costs. You pay your geolocation provider for API calls, and the additional Lambda execution time increases your Lambda charges. To add IP geolocation, sign up for a geolocation API, add this function to the code, and call it where the handler resolves an IP.

    import urllib.request
    import json
    
    def geo_context(ip_address):
        """Enrich an IP address with geolocation data from your provider."""
        try:
            url = f"https://your-geolocation-api.example/json/{ip_address}"
            with urllib.request.urlopen(url, timeout=5) as response:
                data = json.load(response)
            return {
                "country": data.get("country_name"),
                "city": data.get("city"),
                "org": data.get("org"),
            }
        except Exception as error:
            logger.warning(f"Geolocation lookup failed for {ip_address}: {error}")
            return None

    Inside the handler’s loop, after you resolve ip, append the location to the alert.

                if ip:
                    geo = geo_context(ip)
                    if geo:
                        message += (
                            f" Source location: {geo['city']}, "
                            f"{geo['country']} ({geo['org']})."
                        )

    Step 7: Scale to additional patterns and accounts

    As your library grows, move the logic into automated pipelines with EventBridge, Lambda, and AWS Step Functions (a serverless orchestration service that coordinates multiple services into workflows), and surface correlations next to findings in Security Hub. For cross-service correlation at scale, CloudWatch unified data and telemetry capabilities can convert security and compliance data into the OCSF format and let you query sources such as CloudTrail, VPC Flow Logs, and DNS logs from one interface. Security Lake with Athena is a strong option for long-term analysis. Choose the endpoint that fits your retention and query needs.

    Figure 3 shows a correlation pipeline built on AWS services including EventBridge, Lambda, Step Functions, and AWS Security Hub. The pipeline runs from data sources through scheduled queries and enrichment to automated response and centralized visibility.

    Figure 3: A correlation pipeline built on AWS services

    Figure 3: A correlation pipeline built on AWS services

    Conclusion

    You now have four correlation patterns that layer your business context on top of GuardDuty Extended Threat Detection to catch attacks specific to your environment. A few principles carry across every correlation you build.

    • Identity is your primary correlation key: Track the same principal across services.
    • Time windows matter, but they depend on the attack: Events minutes apart are usually related for fast, automated sequences; the ten-minute bins here work for that pattern. Slow or manual reconnaissance can stretch across hours or days, so widen the window when the pattern is deliberate rather than automated.
    • Context is what you add: Your data classification, access norms, resource ownership, and change windows are signals you bring to detection.
    • Start with one rule: A single well-tuned correlation catches more significant activity than a wall of uncorrelated alerts.

    GuardDuty Extended Threat Detection handles the multi-stage patterns common across customers. The correlations in this post add the layer that only your business context can supply. Start with one pattern this week, validate it against your own traffic, and add the next pattern after the first proves reliable.

    Have you built correlation rules for patterns not covered here? Share your experience in the Comments section below.

    Further reading

     

    Nisha Kashyap

    Nisha Kashyap

    Nisha Kashyap is a Senior Support Security Engineer at AWS. She works on threat detection and security operations, helping customers investigate security events and build detection that connects signals across AWS services and reflects their own environment.

    AI-driven software delivery with Kiro, AWS DevOps Agent and Bluebox by Dynatrace

    Post Syndicated from Philipp Ushiromiya original https://aws.amazon.com/blogs/devops/ai-driven-software-delivery-with-kiro-aws-devops-agent-and-bluebox-by-dynatrace/

    This post was co-written with Michael Stephan, Senior Principal Product Manager, and Christian Kreuzberger, Principal Software Engineer, at Dynatrace.

    AI-driven software delivery changes how code gets written, but not what production demands of it. A generated change still has to fit the traffic your service receives, the dependencies it calls, and the capacity limits it runs within. Without that context, you validate the change after it ships, which adds rework and deployment risk.

    Kiro turns intent into specifications, code, and pull requests. AWS DevOps Agent investigates incidents and proposes mitigations. Bluebox by Dynatrace supplies the runtime topology, dependency, and traffic data that both draw on, so each change and each investigation is grounded in how the system behaves rather than how it’s expected to behave. In this post, we will follow a travel-booking example from feature design through post-deployment remediation. You’ll see how telemetry from Bluebox shapes a change in Kiro, how AWS DevOps Agent investigates an incident, and where human review and existing CI/CD controls remain in the process.

    What are Kiro and AWS DevOps Agent?

    Kiro is an agentic development environment that applies AI across the software development lifecycle. Its spec-driven workflow organizes a feature request into requirements, design, and implementation tasks before generating any code.

    AWS DevOps Agent is a frontier agent for software delivery and operations across AWS, multicloud, and on-premises environments. It investigates incidents, identifies likely root causes, and recommends mitigations. Its release management capability (Preview) reviews code for release readiness and runs release tests before deployment.

    Bluebox by Dynatrace: Helps agents ship the code you trust to production

    To close the loop between code generation and production context, Kiro and AWS DevOps Agent rely on real-time production intelligence. This is where Bluebox by Dynatrace fits in. Bluebox provides the observability foundation that detects problems, measures their impact, and surfaces the runtime application topology, service dependencies, and actual traffic patterns that make AI-generated code and autonomous investigations truly production-aware.

    Without production telemetry, AI-generated code operates in a vacuum – it cannot know that an endpoint handles 40:1 read-to-write ratios, that a service dependency has specific latency characteristics, or how API traffic fluctuates throughout the day. Bluebox grounds actions taken by Kiro and AWS DevOps Agent in how the system actually behaves, not in assumptions about how it should behave.

    How the closed loop works

    The combination of Kiro, AWS DevOps Agent, and Bluebox creates a continuous cycle from development through production and back:

    • Production-aware code generation: Before code is written, Kiro retrieves runtime context from Bluebox – service topology, traffic patterns, and resource utilization. Kiro’s spec-driven workflow translates this context into requirements and generates code that aligns with real production conditions from the first commit.
    • Confident code review: Kiro generates pull requests with production evidence attached. The release management capability in AWS DevOps Agent reviews the change for dependency impacts, drifts from internal standards, and production readiness – running autonomous tests in isolated environments.
    • Continuous monitoring: After deployment, Dynatrace continuously monitors application behavior. When an anomaly occurs, Bluebox detects it and surfaces full production context.
    • Autonomous investigation: Bluebox triggers AWS DevOps Agent with the relevant observability and topology data. AWS DevOps Agent performs a deep investigation, correlating telemetry, logs, infrastructure changes, and deployment history to pinpoint the root cause.
    • Automated remediation: AWS DevOps Agent generates the mitigation plan from the observability and runtime data that Bluebox provides. Bluebox adds that plan to the investigation report and files it as a GitHub issue. Kiro then proposes a production-aware fix as a pull request for your review, completing the loop.

    Figure 1: Bluebox supports the closed loop from feature build to operations.

    Next, we walk through a concrete example of this workflow in action.

    Walkthrough

    We follow a travel-booking application through two connected scenarios: shipping a new feature with production context, then responding to a production incident after it deploys.

    Building a production-aware feature

    Consider a team enhancing a travel booking application to improve customer experience. You begin by describing a new feature in Kiro, such as updating how products are displayed or adjusting backend logic to support new capabilities. In this case, we are using Kiro IDE.

    Figure 2. A feature request in Kiro, with the project’s steering documents loaded for context.

    Kiro’s spec-driven workflow expands this request into structured requirements before writing code. You connect Kiro to the Bluebox CLI to retrieve the full production context from Dynatrace: service dependencies, runtime topology, and observed traffic. The following figure shows how Kiro queries current load data for the flight-search path, including the ratio of Amazon DynamoDB reads to writes. Kiro composes and runs the CLI command on your behalf, so you don’t have to type it or set environment variables by hand. The command and its output stay visible in the session, so you can approve it before it runs and check what was retrieved before acting on it. In this case, the command queries the Bluebox API for the requested metrics. The output returns read and write counts per second for the DynamoDB table behind flight search, along with the services calling it.

    Figure 3. Kiro runs the Bluebox CLI, then reads the codebase with production context before proposing changes.

    The telemetry shows the flight-search endpoint is read-heavy. Users repeatedly query the same routes, at roughly 40 reads for every write against the DynamoDB table. Repeated identical reads are what a cache absorbs, so Kiro proposes an Amazon ElastiCache layer in front of the table, sized to the active working set derived from the observed request distribution. Without the read-to-write ratio, the same request could have produced a larger provisioned table or an added read replica, neither of which addresses repeated identical queries.

    Kiro generates the code that implements the change and opens a pull request in GitHub for review. Nothing reaches production until a reviewer approves and merges it. The pull request carries the code changes and the Bluebox telemetry that justified them, so reviewers assess the decision against the same telemetry Kiro retrieved.

    Figure 4. Kiro pushes a feature branch and opens a pull request in GitHub.

    After review and approval through standard processes, a reviewer merges the pull request, and the existing CI/CD pipeline deploys the change.

    Figure 5. The pull request is reviewed and merged through the standard GitHub workflow.

    Responding to a production incident

    With the feature live, Dynatrace continues monitoring the application. A marketing promotion then drives traffic above the observed baseline, and failed requests start to appear. The loop now runs from operations back to development.

    Figure 6. Dynatrace detects a spike in failed requests, surfacing the production incident.

    Bluebox collects the relevant observability and topology data, runs an initial root-cause analysis, then opens an autonomous investigation in AWS DevOps Agent. The AWS DevOps Agent multi-agent reasoning architecture decomposes the investigation across specialized capabilities that each examine one class of evidence: telemetry, logs, infrastructure configuration, and recent deployment activity.

    Figure 7. Bluebox delegates an autonomous investigation to AWS DevOps Agent.

    AWS DevOps Agent locates the cause in the DynamoDB table rather than the new cache. The table’s billing mode had been changed to PROVISIONED, with 5 read capacity units (RCU) and 5 write capacity units (WCU) and no auto scaling. The ElastiCache layer absorbs repeated reads, but cache misses and all writes still reach DynamoDB, and at promotion traffic that residual load exceeds 5 RCU and 5 WCU. AWS DevOps Agent produces a mitigation plan with specific remediation steps. This plan and the full investigation context from Bluebox, is documented as a GitHub issue.

    Figure 8. GitHub issue is created with results from Bluebox and AWS DevOps Agent.

    Kiro proposes a production-aware fix as a new pull request – including the root-cause analysis, supporting telemetry, and recommended configuration changes.

    Figure 9. The Kiro coding session works on the GitHub issue and creates a remediation Pull Request.

    The fix is reviewed, merged, and deployed like any other change. Dynatrace then confirms that error rates and response times return to baseline, which closes the loop.

    Conclusion

    In this post, we showed how Kiro, AWS DevOps Agent, and Bluebox by Dynatrace connect production telemetry with feature development and incident remediation. The travel-booking example keeps human review and existing CI/CD controls in the process while passing operational context from production back to development.

    To get started pick one application and define a measurable outcome, such as investigation time, change-failure rate, or pull-request review time. Then:

    1. Download Kiro and start building with spec-driven development
    2. Enable AWS DevOps Agent for autonomous incident investigation and remediation
    3. Get started with Bluebox by Dynatrace to complete the loop with production intelligence

    Simone Pomata

    Simone is a Principal Solutions Architect at AWS. He has worked enthusiastically in the tech industry for more than 10 years. At AWS, he helps customers succeed in building new technologies every day.

    Philipp Ushiromiya

    Philipp Ushiromiya is a Solutions Architect at AWS. He helps customers drive organizational modernization through cloud-native solutions and DevOps practices. His passion for GenAI enables teams to accelerate development with cutting-edge technology.

    Michael Stephan

    Michael Stephan is a Senior Principal Product Manager at Dynatrace with over 15 years of experience in the IT industry. He specializes in helping Dynatrace customers effectively monitor and optimize their cloud environments.

    Christian Kreuzberger

    Christian Kreuzberger is a Principal Software Engineer at Dynatrace, with over 20 years of experience in the IT industry. At Dynatrace, he builds software that helps cloud-native and AI-native organizations automate their operations.

    Setting up an RCS agent with an AI coding assistant and AWS End User Messaging

    Post Syndicated from Bruno Giorgini original https://aws.amazon.com/blogs/messaging-and-targeting/setting-up-an-rcs-agent-with-an-ai-coding-assistant-and-aws-end-user-messaging/

    Clone a repo, open it in your AI coding assistant, type “go,” and walk away with a working RCS agent.

    Creating an RCS agent on AWS End User Messaging normally means juggling 23 registration fields, three different CLI parameter types, brand asset requirements, and a multi-step approval process. An AI coding assistant can handle all of that for you. With AWS End User Messaging, you can create RCS agents that send and receive rich messages complete with your brand’s logo, colors, and verified identity.

    Setting up an RCS agent involves creating an agent container, uploading brand assets, configuring a 23-field registration, submitting for approval, adding verified testers, and testing both outbound and inbound messaging. Each field has a specific type (TEXT, SELECT, or ATTACHMENT) that requires a different CLI parameter, and getting any of them wrong means starting over.

    We built an open-source sample repository that encodes all of this knowledge into an AGENTS.md file. When you open the repo in an AI coding assistant like Kiro, Cursor, or Windsurf, the assistant reads the instructions and walks you through the entire setup interactively. You provide a brand name and your phone number. The AI handles everything else.

    How it works

    The repository aws-samples/sample-rcs-agent-setup-and-send-messages contains:

    • AGENTS.md — A structured instruction file that AI coding assistants read automatically. It contains the complete RCS agent setup workflow: credential checks, brand asset generation, registration field configuration, tester management, and message testing.
    • brand-assets/ — Template SVG files for the agent logo (224×224 px) and banner (1440×448 px), ready to be customized and converted to PNG.
    • .kiro/steering/rcs-agent-setup.md — A Kiro-specific steering file with the same instructions, using the inclusion: always frontmatter so Kiro loads it automatically.

    The AGENTS.md file is the key. It defines six skills that the AI assistant executes in sequence:

    1. Create RCS agent — Creates the agent container, generates brand assets (logo and banner SVGs), converts them to PNG, creates a test registration, sets all 23 fields with the correct parameter types, and submits for approval.
    2. Add verified testers — Registers test phone numbers and guides you through accepting the tester invitation.
    3. Send a test message — Checks for blockers (protect configuration, opt-out lists) and sends your first branded RCS message.
    4. Set up inbound keyword — Configures an automatic response keyword so you can test inbound messaging without writing backend code.
    5. Verify inbound messaging — Walks you through the console deep link flow to confirm two-way messaging works.
    6. Delete an RCS agent — Removes an agent cleanly by disabling deletion protection, deleting the associated registration, then deleting the agent itself.

    Prerequisites

    Before you start, you need:

    • An AWS account with access to AWS End User Messaging.
    • AWS Command Line Interface (AWS CLI) v2.35.12 or later installed and configured with credentials that have pinpoint-sms-voice-v2:* permissions.
    • An AI coding assistant that reads AGENTS.md files (Kiro, Cursor, Windsurf, or similar).
    • librsvg for SVG to PNG conversion (brew install librsvg on macOS).
    • A test phone that supports RCS messaging.

    Getting started

    Follow these steps to go from zero to a working RCS agent. The entire process takes about five minutes.

    Step 1: Clone the repository

    git clone https://github.com/aws-samples/sample-rcs-agent-setup-and-send-messages.git
    cd sample-rcs-agent-setup-and-send-messages

    Step 2: Open in your AI coding assistant

    Open the cloned directory in your preferred AI coding assistant. The assistant will automatically detect the AGENTS.md file (or .kiro/steering/rcs-agent-setup.md if you are using Kiro).

    Step 3: Type “go”

    In the chat panel, type go. The AI assistant will:

    1. Check your AWS credentials — It runs aws sts get-caller-identity and asks how you authenticate if credentials are not configured. It supports named profiles, SSO, IAM user credentials, and environment variables.
    2. Verify EUM access — It confirms your account can use AWS End User Messaging.
    3. Check tooling — It verifies rsvg-convert is installed for brand asset generation.
    4. Ask for your preference — Quick mode (provide a brand name) or interactive mode (you specify every detail).

    Step 4: Provide a brand name

    In quick mode, you provide a brand name and the AI generates everything else: a description, an accessible accent color, contact information with placeholder values, privacy and terms URLs, and custom SVG brand assets with your brand name and colors.

    In interactive mode, the AI asks for each detail one section at a time: brand name, accent color, logo description, banner description, contact information, and policy URLs.

    Step 5: Watch it work

    The AI assistant executes every AWS CLI command in sequence:

    1. Creates the RCS agent container.
    2. Enables deletion protection.
    3. Creates a test registration and links it to the agent.
    4. Generates and converts brand asset SVGs to PNG.
    5. Uploads the logo and banner as registration attachments.
    6. Sets all 23 registration fields using the correct parameter type for each (TEXT, SELECT, or ATTACHMENT).
    7. Submits the registration and polls for approval.
    8. Reports when the agent is active.

    Step 6: Add a tester and send a message

    Once the agent is approved, the AI asks for your test phone number, registers it as a verified tester, and waits for you to accept the invitation. After verification, it checks for blockers (protect configuration and opt-out lists), then sends your first branded RCS message.

    Step 7: Test inbound messaging

    The AI configures an automatic keyword response and walks you through the console deep link flow to verify two-way messaging. When you send RCSINBOUNDTESTING to your agent, you receive an automatic reply confirming inbound messaging works.

    What the AI handles for you

    The AGENTS.md file encodes several non-obvious behaviors that would otherwise require trial and error:

    Challenge How the repo handles it
    create-rcs-agent takes no --display-name parameter The brand name comes from the registration, not the agent creation call. The instructions reflect this.
    Three different field parameter types The instructions include a field reference table mapping each of the 23 fields to its correct CLI parameter: --text-value, --select-choices, or --registration-attachment-id.
    --field-values does not exist The instructions explicitly warn against this non-existent parameter and use the correct alternatives.
    --attachment-body and --attachment-url conflict The instructions use --attachment-body only.
    Accent color contrast requirements The instructions include pre-validated color choices with 4.5:1 contrast ratio against white.
    Field paths differ from what you might expect The correct paths are agentDetails.logoImage and agentDetails.bannerImage, not logoAttachmentId or bannerAttachmentId.
    New registration versions do not inherit field values The troubleshooting section warns that all 23 fields must be re-populated when creating a new version.

    Customizing the repo

    You can modify the AGENTS.md file to fit your workflow:

    • Change default values — Update placeholder contact information, privacy URLs, or terms URLs to match your organization.
    • Add custom brand assets — Replace the template SVGs in brand-assets/ with your own designs. Keep the logo at 224×224 px and the banner at 1440×448 px.
    • Extend the skills — Add new skills for richer message types (cards, carousels), event destinations for programmatic inbound handling, or integration with other AWS services.

    Cleanup

    To remove the resources created during testing:

    # 1. Disable deletion protection
    aws pinpoint-sms-voice-v2 update-rcs-agent \
      --rcs-agent-id <your-agent-id> \
      --no-deletion-protection-enabled \
      --region us-east-1
    
    # 2. Delete the associated registration (required before deleting the agent)
    aws pinpoint-sms-voice-v2 delete-registration \
      --registration-id <your-registration-id> \
      --region us-east-1
    
    # 3. Delete the agent
    aws pinpoint-sms-voice-v2 delete-rcs-agent \
      --rcs-agent-id <your-agent-id> \
      --region us-east-1

    Note: You must delete the registration before the agent. Skipping this step results in a ConflictException: RESOURCE_NOT_EMPTY error.

    Conclusion

    The aws-samples/sample-rcs-agent-setup-and-send-messages repository turns a multi-step, error-prone CLI workflow into a guided conversation. Clone the repo, open it in your AI coding assistant, type “go,” and you have a working RCS agent that can send and receive branded messages to verified testers.

    The AGENTS.md pattern is reusable. Any complex AWS workflow with non-obvious API behavior can be encoded the same way: document the correct commands, parameter types, and pitfalls in a structured file, and let the AI assistant execute it interactively.

    For a detailed manual walkthrough of the same process, see Creating and testing an RCS agent with AWS End User Messaging. For an overview of the business case for RCS, see Upgrade business messaging with RCS on AWS. For more information, see the AWS End User Messaging service page and the RCS documentation.


    About the author

    Enable cross-cloud analytics with Amazon S3 Tables and Google BigQuery, Part 1: IAM-based access control

    Post Syndicated from Lakshmi Nair original https://aws.amazon.com/blogs/big-data/enable-cross-cloud-analytics-with-amazon-s3-tables-and-google-bigquery-part-1-iam-based-access-control/

    Organizations running analytics workloads across multiple clouds often hit the same friction: the data lives on one cloud, but the engine querying it lives on another. Copying data across the boundary creates a second dataset that must be kept in sync, adding cost, latency, and reconciliation overhead. In this post, we address a specific instance of that pattern: your Google BigQuery users need to work with data that lives in Amazon S3 Tables, a capability of Amazon Simple Storage Service (Amazon S3), on AWS. The ideal outcome is a single, governed dataset that serves teams in both clouds without a standing replication pipeline between them.

    With Amazon S3 Tables, you get managed Apache Iceberg tables with built-in compaction, snapshot management, and an integration with the AWS Glue Data Catalog. Because S3 Tables stores data in the open Iceberg format, supported external engines can read it directly if the right access path exists.

    This two-part blog series demonstrates how you can connect Google BigQuery to Amazon S3 Tables using the cross-cloud lakehouse with AWS Glue. We cover two access control approaches:

    1. AWS Identity and Access Management (IAM): You can define a single policy that uses IAM permissions to set up access to both table metadata and data.
    2. AWS Lake Formation: You can use temporary vended credentials for data access, with metadata access managed by Lake Formation permissions.

    This post focuses on the IAM-based approach. Part 2 covers the Lake Formation approach for organizations that need credential-vended access across multiple engines.

    By the end, you will have BigQuery querying Iceberg tables stored on S3 Tables without data copy or duplication, providing live access to Iceberg data.

    Cross-cloud analytics scenarios

    There are several scenarios where organizations benefit from cross-cloud querying capabilities. Here are some of the common patterns this architecture addresses:

    Schema evolution across cloud boundaries

    When source schemas change frequently, streaming pipelines writing to BigQuery-managed store require coordinated DDL changes on the BigQuery table and downstream views. Teams often work around this challenge by storing payloads as untyped columns and parsing them later.

    With Iceberg on S3 Tables, schema evolution is tracked in table metadata. When the writing engine adds a new column, BigQuery’s Lakehouse refresh picks up the updated schema automatically on the next sync cycle.

    Multi-cloud analytics without data duplication

    A company has its production data environment on AWS (data lakes, warehouses, streaming) but acquired a business unit that runs analytics exclusively on BigQuery. In-place querying from BigQuery keeps your data in Amazon S3 Tables, so you pay for one copy, work from live data, and avoid the operational overhead of a synchronized second store.

    Cost optimization for infrequently queried datasets

    An organization has hundreds of datasets on AWS, but only a fraction is queried daily from BigQuery. Replicating all of them to Google Cloud Storage drives unnecessary storage and transfer costs. With Lakehouse catalog federation, you keep your data on S3 Tables. BigQuery reads data only when queried, so you pay per query rather than per-copy storage.

    Decoupled compute across engines

    Data team wants storage on AWS with the flexibility for multiple engines to read the same data: BigQuery and Amazon Redshift for data warehousing use cases, Amazon Athena for interactive ad-hoc querying, Amazon SageMaker AI for machine learning (ML). With Apache Iceberg’s open format, you can use one storage layer, many compute engines, no data copies between them.

    Solution overview

    You use the AWS Glue Iceberg REST Catalog (IRC) as the bridge between BigQuery and S3 Tables. BigQuery’s cross-cloud Lakehouse creates a federated catalog that syncs metadata from the Glue IRC, then uses the synced metadata to read Iceberg data files directly.

    Architecture diagram showing BigQuery connecting to Amazon S3 Tables through the AWS Glue Iceberg REST Catalog

    Figure 1: Architecture diagram showing BigQuery connecting to Amazon S3 Tables through the AWS Glue Iceberg REST Catalog

    The key components in this architecture:

    1. Amazon S3 Tables: With Amazon S3 Tables, you get a fully managed Apache Iceberg table experience in Amazon S3, optimized for analytics workloads. You can register table metadata in the AWS Glue Data Catalog for discovery and governance.
    2. AWS Glue Data Catalog: With AWS Glue Data Catalog, you can access the federated s3tablescatalog catalog that maps S3 Tables resources (table buckets, namespaces, tables) into a catalog hierarchy from supported analytics engines. The standard Iceberg REST endpoint of Glue Data Catalog serves table metadata to external engines. BigQuery connects through this endpoint.
    3. Google Cross-Cloud Lakehouse: With Google Cross-Cloud Lakehouse, you can connect BigQuery to external Iceberg catalogs. It assumes an AWS IAM role using OpenID Connect (OIDC), calls the Glue Iceberg REST endpoint, and syncs metadata on a configurable refresh interval.

    Prerequisites

    Before you begin, you need:

    • An AWS account with Amazon S3 Tables available in your AWS Region.
    • A Google Cloud project with billing enabled and the BigLake API activated.
    • AWS Command Line Interface (AWS CLI) and gcloud CLI installed and configured.
    • An S3 table bucket with at least one namespace and table containing data.

    Setting up Amazon S3 Tables

    If you already have S3 Tables with data, skip to the next section. Otherwise, create a table bucket, namespace, and populate a table.

    Create a table bucket and namespace

    Use the AWS CLI to create resources as follows:

    # Create a Table bucket
    aws s3tables create-table-bucket \
        --name <TABLE_BUCKET_NAME> \
        --region <REGION>
    
    # Create a Namespace (Database)
    aws s3tables create-namespace \
        --table-bucket-arn "arn:aws:s3tables:<REGION>:<AWS_ACCOUNT_ID>:bucket/<TABLE_BUCKET_NAME>" \
        --namespace <NAMESPACE> \
        --region <REGION>

    Integrating S3 Tables with the Glue Data Catalog

    For BigQuery to access S3 Tables, the tables must be discoverable through the Glue Data Catalog. S3 Tables integrates with Glue through a federated catalog called s3tablescatalog.

    Set up S3 Tables integration with the Glue Data Catalog using IAM mode

    Open the Amazon S3 console:

    1. In the navigation pane, choose Table buckets.
    2. Choose Enable integration, and then choose Enable integration again to confirm.

    This creates the s3tablescatalog federated catalog in Glue, where access is controlled entirely by IAM policies on the calling role. This is a one-time setup per account and Region. After you enable it, the analytics integration applies to all table buckets in your account.

    The Enable integration option on the table buckets page of the Amazon S3 console

    Figure 2: Enabling the S3 Tables integration in the Amazon S3 console

    Alternatively, create the catalog using the AWS CLI:

    aws glue create-catalog --region <REGION> --cli-input-json '{
      "Name": "s3tablescatalog",
      "CatalogInput": {
        "FederatedCatalog": {
          "Identifier": "arn:aws:s3tables:<REGION>:<AWS_ACCOUNT_ID>:bucket/*",
          "ConnectionName": "aws:s3tables"
        },
        "CreateDatabaseDefaultPermissions": [
          { "Principal": {"DataLakePrincipalIdentifier": "IAM_ALLOWED_PRINCIPALS"}, "Permissions": ["ALL"] }
        ],
        "CreateTableDefaultPermissions": [
          { "Principal": {"DataLakePrincipalIdentifier": "IAM_ALLOWED_PRINCIPALS"}, "Permissions": ["ALL"] }
        ]
      }
    }'

    Create a table and insert data

    Now, to create the table and insert data, open the Amazon Athena console. In the query editor, select s3tablescatalog/<TABLE_BUCKET_NAME> as your data source and <NAMESPACE> as the database. Then run the following SQL statements one by one:

    CREATE TABLE `<NAMESPACE>`.orders (
        order_id STRING,
        customer_id STRING,
        amount BIGINT,
        order_date DATE,
        region STRING
    )
    TBLPROPERTIES ('table_type' = 'iceberg');
    
    INSERT INTO orders
    VALUES
        ('ORD-001', 'C100', 4500, DATE '2024-06-01', 'EMEA'),
        ('ORD-002', 'C200', 8900, DATE '2024-06-01', 'EMEA'),
        ('ORD-003', 'C100', 3200, DATE '2024-06-02', 'NAMER'),
        ('ORD-004', 'C300', 12000, DATE '2024-06-02', 'NAMER'),
        ('ORD-005', 'C400', 6700, DATE '2024-06-03', 'APJ'),
        ('ORD-006', 'C200', 4100, DATE '2024-06-03', 'APJ'),
        ('ORD-007', 'C500', 9500, DATE '2024-06-04', 'EMEA'),
        ('ORD-008', 'C100', 2800, DATE '2024-06-04', 'LATAM'),
        ('ORD-009', 'C600', 15000, DATE '2024-06-05', 'NAMER'),
        ('ORD-010', 'C300', 7200, DATE '2024-06-05', 'LATAM');

    Configuring cross-cloud access

    BigQuery assumes an AWS IAM role via OIDC federation to access the Glue IRC. This section walks through creating the role, OIDC provider, and permissions.

    Create the OIDC identity provider

    Register Google as an OIDC identity provider in your AWS account. This allows AWS to validate tokens issued by Google’s identity service:

    aws iam create-open-id-connect-provider \
        --url https://accounts.google.com \
        --client-id-list accounts.google.com \
        --thumbprint-list 08745487e891c19e3078c1f2a07e452950ef36f6

    The –thumbprint-list parameter is optional. When omitted, IAM automatically retrieves the thumbprint from the OIDC provider’s certificate. See AWS documentation for details.

    Create the cross-cloud IAM role

    Login into AWS Console, and  create the role with a placeholder trust policy. You will update it with the actual BigLake service account ID after you create the federated catalog in Google Cloud.

    aws iam create-role \
        --role-name bigquery-cross-cloud-role \
        --max-session-duration 43200 \
        --assume-role-policy-document '{
          "Version": "2012-10-17",
          "Statement": [{
            "Effect": "Allow",
            "Principal": {
              "Federated": "arn:aws:iam::<AWS_ACCOUNT_ID>:oidc-provider/accounts.google.com"
            },
            "Action": "sts:AssumeRoleWithWebIdentity",
            "Condition": {
              "StringEquals": {
                "accounts.google.com:sub": ["PLACEHOLDER"],
                "accounts.google.com:aud": ["PLACEHOLDER"]
              }
            }
          }]
        }'

    The --max-session-duration 43200 allows sessions up to 12 hours, which is needed for long-running BigQuery queries.

    Attach permissions

    The permissions policy differs based on your access control approach. For the IAM-based approach, attach the following policy:

    {
      "Version": "2012-10-17",
      "Statement": [
        {
          "Sid": "GlueRead",
          "Effect": "Allow",
          "Action": [
            "glue:GetCatalog", "glue:GetDatabase", "glue:GetDatabases",
            "glue:GetTable", "glue:GetTables", "glue:GetPartition", "glue:GetPartitions"
          ],
          "Resource": [
            "arn:aws:glue:<AWS_REGION>:<AWS_ACCOUNT_ID>:catalog",
            "arn:aws:glue:<AWS_REGION>:<AWS_ACCOUNT_ID>:catalog/s3tablescatalog",
            "arn:aws:glue:<AWS_REGION>:<AWS_ACCOUNT_ID>:catalog/s3tablescatalog/<TABLE_BUCKET>",
            "arn:aws:glue:<AWS_REGION>:<AWS_ACCOUNT_ID>:database/s3tablescatalog/<TABLE_BUCKET>/<NAMESPACE>",
            "arn:aws:glue:<AWS_REGION>:<AWS_ACCOUNT_ID>:table/s3tablescatalog/<TABLE_BUCKET>/<NAMESPACE>/*"
          ]
        },
        {
          "Sid": "S3TablesRead",
          "Effect": "Allow",
          "Action": [
            "s3tables:GetTableBucket", "s3tables:ListTableBuckets",
            "s3tables:ListNamespaces", "s3tables:GetNamespace",
            "s3tables:ListTables", "s3tables:GetTable",
            "s3tables:GetTableMetadataLocation", "s3tables:GetTableData"
          ],
          "Resource": [
            "arn:aws:s3tables:<AWS_REGION>:<AWS_ACCOUNT_ID>:bucket/<TABLE_BUCKET>",
            "arn:aws:s3tables:<AWS_REGION>:<AWS_ACCOUNT_ID>:bucket/<TABLE_BUCKET>/*"
          ]
        },
        {
          "Sid": "S3TablesListBuckets",
          "Effect": "Allow",
          "Action": ["s3tables:ListTableBuckets"],
          "Resource": "*"
        }
      ]
    }

    Connecting BigQuery to S3 Tables

    With the AWS side configured, create the federated catalog in Google Cloud that connects BigQuery to the Glue IRC.

    Create the federated catalog

    Authenticate to Google Cloud using gcloud auth login, or use Cloud Shell, which is pre-authenticated. Verify that the BigLake API is enabled:

    gcloud services enable biglake.googleapis.com --project="<GCP_PROJECT_ID>"

    For IAM mode:

    gcloud alpha biglake iceberg catalogs create <FEDERATED_CATALOG_NAME> \
        --project="<GCP_PROJECT_ID>" \
        --catalog-type=federated \
        --federated-catalog-type=glue \
        --glue-aws-region=<AWS_REGION> \
        --glue-aws-role-arn=arn:aws:iam::<AWS_ACCOUNT_ID>:role/bigquery-cross-cloud-role \
        --glue-warehouse=<AWS_ACCOUNT_ID>:s3tablescatalog/<TABLE_BUCKET> \
        --primary-location=<GCP_REGION>

    The --glue-warehouse parameter uses the format <AWS_ACCOUNT_ID>:s3tablescatalog/<TABLE_BUCKET>. This tells the Glue IRC to scope requests to your specific S3 Tables bucket within the federated catalog hierarchy.

    The --primary-location refers to the Google Cloud region where the federated catalog metadata is stored. Use the AWS to Google Cloud region mapping to find the corresponding GCP region for your AWS Region. For example, AWS us-east-1 maps to GCP us-east4.

    Retrieve the BigLake service account ID

    After catalog creation, Google provisions a dedicated service account for your federated catalog. Retrieve its numeric ID:

    BIGLAKE_SA_ID=$(gcloud alpha biglake iceberg catalogs describe <FEDERATED_CATALOG_NAME> \
        --project="<GCP_PROJECT_ID>" \
        --format="value(biglake-service-account-id)")
    echo $BIGLAKE_SA_ID

    Update the AWS trust policy

    Back on AWS, replace the placeholder in the IAM role’s trust policy with the actual service account ID:

    aws iam update-assume-role-policy \
        --role-name bigquery-cross-cloud-role \
        --policy-document '{
          "Version": "2012-10-17",
          "Statement": [{
            "Effect": "Allow",
            "Principal": {
              "Federated": "arn:aws:iam::<AWS_ACCOUNT_ID>:oidc-provider/accounts.google.com"
            },
            "Action": "sts:AssumeRoleWithWebIdentity",
            "Condition": {
              "StringEquals": {
                "accounts.google.com:sub": ["<BIGLAKE_SA_ID>"],
                "accounts.google.com:aud": ["<BIGLAKE_SA_ID>"]
              }
            }
          }]
        }'

    Register the service account ID in the OIDC provider’s audience list. Without this step, AWS rejects the token because the aud claim doesn’t match any registered client:

    aws iam add-client-id-to-open-id-connect-provider \
        --open-id-connect-provider-arn "arn:aws:iam::<AWS_ACCOUNT_ID>:oidc-provider/accounts.google.com" \
        --client-id "<BIGLAKE_SA_ID>"

    Set up metadata sync

    Wait 3–5 minutes for IAM changes to propagate globally, then set up background refresh:

    gcloud alpha biglake iceberg catalogs update <FEDERATED_CATALOG_NAME> \
        --project="<GCP_PROJECT_ID>" \
        --refresh-interval=300s

    The --refresh-interval (300 seconds in this example) determines how often BigQuery syncs metadata from the Glue IRC. New tables and schema changes appear in BigQuery within this interval.

    Querying from BigQuery

    After the catalog refresh completes, BigQuery automatically creates external datasets corresponding to the synced namespaces. No manual CREATE SCHEMA is required.

    Verify the sync:

    gcloud alpha biglake iceberg namespaces list \
        --catalog="<FEDERATED_CATALOG_NAME>" \
        --project="<GCP_PROJECT_ID>"

    Run a query in BigQuery:

    SELECT * FROM `<GCP_PROJECT_ID>.<FEDERATED_CATALOG_NAME>.<NAMESPACE>.orders` LIMIT 1000

    Sample Query Output:

    SELECT
        customer_id,
        COUNT(*) as order_count,
        SUM(amount) as total_spend
    FROM `<PROJECT_ID>.<FEDERATED_CATALOG_NAME>.<NAMESPACE>.orders`
    GROUP BY customer_id
    ORDER BY total_spend DESC
    BigQuery query results showing order count and total spend per customer from the Amazon S3 Tables data

    Figure 3: BigQuery query results returned directly from the Amazon S3 Tables data

    BigQuery reads the Iceberg metadata to identify which Parquet data files contain relevant data. It also applies partition pruning where applicable, and fetches only the necessary files from S3 Tables managed storage.

    Schema evolution

    When new columns are added to an Iceberg table on the AWS side (through Spark, Athena, or the Glue IRC), the schema change is captured in Iceberg’s metadata. On the next Lakehouse refresh cycle, BigQuery picks up the new columns automatically. No DDL changes are needed in BigQuery.

    Metadata freshness

    The s3tablescatalog in Glue is a federated catalog that resolves table metadata live from the S3 Tables service on each request. When a streaming job commits new data to an S3 Table, the latest metadata is immediately available through the AWS Glue IRC. BigQuery sees the update on its next refresh cycle (as configured by --refresh-interval).

    OIDC identity federation

    The trust relationship between Google Cloud and AWS uses OpenID Connect. When BigQuery Lakehouse needs to access your data, it presents a signed JWT token containing:

    • iss: accounts.google.com (the issuer)
    • sub: The BigLake service account ID (identifies which catalog is making the request)
    • aud: The same service account ID (the intended audience)

    AWS validates this token against the registered OIDC provider and trust policy conditions before issuing temporary credentials. Each federated catalog receives a unique service account ID, providing per-catalog isolation and auditability through AWS CloudTrail.

    Network path

    By default, traffic between BigQuery and AWS travels over the public internet. For workloads requiring private connectivity, Google Cloud supports Cross-Cloud Interconnect or Partner Interconnect. This helps routing queries over a dedicated network path. Refer to the Google Cloud documentation for private interconnect configuration.

    Clean up

    To avoid ongoing charges, remove the resources created in this walkthrough.

    On AWS:

    # Delete the table (if created for this walkthrough)
    aws s3tables delete-table \
        --table-bucket-arn "arn:aws:s3tables:<AWS_REGION>:<AWS_ACCOUNT_ID>:bucket/<TABLE_BUCKET>" \
        --namespace analytics --name orders --region <AWS_REGION>
    
    # Delete namespace and table bucket
    aws s3tables delete-namespace \
        --table-bucket-arn "arn:aws:s3tables:<AWS_REGION>:<AWS_ACCOUNT_ID>:bucket/<TABLE_BUCKET>" \
        --namespace <NAMESPACE> --region <AWS_REGION>
    
    aws s3tables delete-table-bucket --name <TABLE_BUCKET> --region <AWS_REGION>
    
    # Delete IAM role and OIDC provider (if no longer needed)
    aws iam delete-role --role-name bigquery-cross-cloud-role

    On Google Cloud:

    gcloud alpha biglake iceberg catalogs delete <FEDERATED_CATALOG_NAME> \
        --project="<GCP_PROJECT_ID>" --location=<GCP_REGION>

    Conclusion

    This post demonstrated how to query Amazon S3 Tables from Google BigQuery using the open Apache Iceberg format and the AWS Glue Iceberg REST Catalog as the metadata bridge. Using Apache Iceberg’s open format, you can write data once on AWS and read it from supported engines that speak Iceberg, including BigQuery. We used IAM-based access control to govern access to both Glue Data Catalog metadata and the underlying Amazon S3 Tables data. This is the simpler configuration path with fewer components. In Part 2, we walk through configuring AWS Lake Formation to vend temporary, scoped credentials to BigQuery for data access.

    To get started with this pattern in your environment:


    About the authors

    Lakshmi Nair

    Lakshmi Nair

    Lakshmi is a Principal Analytics Specialist Solutions Architect at AWS. She specializes in designing advanced analytics systems across industries. She focuses on crafting cloud-based data platforms, enabling real-time streaming, big data processing, and robust data governance.

    Srividya Parthasarathy

    Srividya Parthasarathy

    Srividya was a Senior Big Data Architect on the AWS Lake Formation team. She works with product team and customer to build robust features and solutions for their analytical data platform. She enjoys building data mesh solutions and sharing them with the community.