Amazon S3 Vectors now supports metadata pre-filtering for higher recall on filtered searches

Post Syndicated from Daniel Abib original https://aws.amazon.com/blogs/aws/amazon-s3-vectors-now-supports-metadata-pre-filtering-for-higher-recall-on-filtered-searches/

Today, we’re announcing metadata pre-filtering for Amazon S3 Vectors, which delivers higher recall on filtered queries by evaluating your metadata filter before the similarity search. You can filter on attributes such as tenant, category, status, or time, and pre-filtering adds prefix matching with $startsWith for paths, URLs, and hierarchical keys. Each vector carries up to 2 KB of filterable metadata, and a single query supports up to 100 filter constraints. There is no additional cost, no re-ingestion, and no change to your queries.

Most applications never search a whole index. They search the part of it that belongs to a particular user, account, or category, and they express that scope as a metadata filter. Semantic search, retrieval-augmented generation (RAG), and agentic applications all need the same thing from a filtered query: a similarity search that covers the vectors matching the filter, and returns the closest of them. With pre-filtering, a filtered query returns more of the relevant matches your index contains, giving you higher recall on filtered searches.

Common use cases

Pre-filtering applies wherever results have to be both relevant and correctly scoped:

  • Legal and professional services: A law firm or e-discovery platform searches documents scoped to a single client, and with $startsWith narrows further by matter number, folder path, or document ID prefix. A single client is a small share of a firm-wide archive, and filters this narrow are where pre-filtering improves recall most.
  • Financial services: An investment research platform searches analyst notes, filings, and call transcripts scoped by issuer, document type, and publication date.
  • Media and entertainment: A streaming service filters by content rating and regional licensing before the semantic search, finding similar titles restricted to G and PG content licensed in one territory.
  • Agentic applications: An agent working within a user’s session filters on fields such as owner, document set, and timestamp so its searches cover the material relevant to the task at hand. Higher recall means more of that material reaches the agent, which improves task reliability

How pre-filtering works

Each vector in an S3 Vectors index can carry application-defined metadata, and a query can filter on those fields.

Every vector index has an index mode. On an index whose index mode is ENHANCED, S3 Vectors resolves your filter first, then searches only the vectors that match. On an index whose index mode is CLASSIC, S3 Vectors performs the vector search and filter evaluation in tandem, validating each candidate vector against your filter as it searches. Existing indexes use CLASSIC until you update them.

Consider a support knowledge base of 8 million tickets, where an agent searches one customer’s history for a recurring error. If that customer accounts for 400 of those tickets, resolving customer_id first means the similarity search runs across all 400 of them, so the agent sees that customer’s prior occurrences. Before the index was updated, the same query drew its candidates from the full 8 million, and the result set contained fewer of that customer’s matching tickets.

On highly selective filters, pre-filtering returns up to 5x more of the matching vectors than the same query returned before on CLASSIC indexes.

Getting started

Before you start, make sure your IAM policy grants permissions for the new actions.

You can get started in three steps. The walkthrough below builds a small product-catalog index and runs a selective filter against it, the same pattern you would use for a multi-tenant RAG store or a document search scoped to one client.

First, create a vector index:

aws s3vectors create-index \
  --index-name product-catalog \
  --vector-bucket-name my-vector-bucket \
  --dimension 1536 \
  --distance-metric cosine

The dimension must match the output size of your embedding model, and distance-metric should match how that model was trained (cosine is common for text embeddings). Second, write vectors with the PutVectors API, attaching up to 2 KB of filterable metadata to each vector:

aws s3vectors put-vectors \
  --index-name product-catalog \
  --vector-bucket-name my-vector-bucket \
  --vectors '[{
    "key": "doc-001",
    "data": {"float32": [0.1, 0.2, 0.3, ...]},
    "metadata": {
      "tenant_id": "t-10428",
      "category": "legal",
      "created_date": "2026-03-15",
      "active": true
    }
  }]'

Each vector carries the attributes your application filters on. In this example, tenant_id scopes results to a single customer, category narrows by document type, created_date records when the document was created, and active is a boolean flag. By default every metadata field is filterable, so you can query on any of them without declaring a schema up front.

Third, run a filtered similarity query with the QueryVectors API. The filter uses a compact JSON syntax where a bare key-value pair is an equality match, and operators such as $and, $or, and $gt combine or refine conditions. Pass --return-metadata so the query returns each vector’s metadata:

aws s3vectors query-vectors \
  --index-name product-catalog \
  --vector-bucket-name my-vector-bucket \
  --query-vector '{"float32": [0.1, 0.2, 0.3, ...]}' \
  --top-k 50 \
  --return-metadata \
  --filter '{"$and": [
    {"tenant_id": "t-10428"},
    {"category": "legal"},
    {"active": true}
  ]}'

The expected result is a single vector, doc-001, the only one matching all three filter conditions (tenant_id, category, and active):

{
  "vectors": [
    {
      "distance": 0.9717477560043335,
      "key": "doc-001",
      "metadata": {
        "tenant_id": "t-10428",
        "category": "legal",
        "created_date": "2026-03-15",
        "active": true
      }
    }
  ],
  "distanceMetric": "cosine"
}

S3 Vectors first narrows the search space to vectors matching all three filter conditions, then returns the 50 most similar vectors from that subset. Because the filter is applied before the search, those results are drawn from across all the vectors that match it.

Prefix matching with $startsWith

Pre-filtering adds a prefix match operator for filtering on paths, URLs, and hierarchical keys. A document store that encodes case and folder structure into a document ID can scope a search to a subtree in one condition:

--filter '{"$startsWith": {"document_id": "matter-4417/exhibits/"}}'

$startsWith joins the existing operators: equality, numeric range, set membership, existence checks, and boolean logic with $and and $or.

Turning on pre-filtering for existing indexes

Call UpdateIndexMode on an existing index to turn on pre-filtering:

aws s3vectors update-index-mode \
  --vector-bucket-name my-vector-bucket \
  --index-name product-catalog \
  --index-mode ENHANCED

Pre-filtering takes effect in place. Your existing vectors are not re-ingested, your queries do not change, and the new filter operators are available immediately.

Here is the difference on the same index and the same query. Before the update, a query scoped to one tenant returns two of the ten results requested:

aws s3vectors query-vectors \
  --vector-bucket-name my-vector-bucket \
  --index-name product-catalog \
  --query-vector '{"float32": [0.1, 0.2, 0.3, ...]}' \
  --top-k 10 \
  --return-metadata \
  --filter '{"tenant_id": "t-10428"}'
{
  "vectors": [
    { "key": "doc-114", "distance": 0.41 },
    { "key": "doc-322", "distance": 0.55 }
  ],
  "distanceMetric": "cosine"
}

After the update, the same query returns a full result set drawn from across that tenant’s documents:

{
  "vectors": [
    { "key": "doc-018", "distance": 0.09 },
    { "key": "doc-207", "distance": 0.13 },
    { "key": "doc-114", "distance": 0.41 },
    ... 7 more
  ],
  "distanceMetric": "cosine"
}

Rolling out across your indexes

Once you have validated pre-filtering on an index, set the default index mode on the vector bucket so that new indexes use ENHANCED without a follow-up call:

aws s3vectors put-vector-bucket-default-index-mode \
  --vector-bucket-name my-vector-bucket \
  --default-index-mode ENHANCED

To bring the rest of your existing indexes across, list them and check the index mode on each one, then call UpdateIndexMode on the ones still using CLASSIC:

aws s3vectors list-indexes \
  --vector-bucket-name my-vector-bucket

aws s3vectors get-index \
  --vector-bucket-name my-vector-bucket \
  --index-name product-catalog

Things to know 

  • Indexes created in vector buckets created on or after September 30, 2026 use index mode ENHANCED. Indexes in buckets that existed before that date use CLASSIC until you set the bucket default, including indexes created in those buckets afterward.
  • A single query supports up to 100 filter constraints, counted per value the filter evaluates. If a query exceeds that, you can usually consolidate the filter, replacing a 300-value $in over legal cases with a single caseId field, for example, or split it into smaller queries, run them in parallel, and merge the results by distance.

Get started today

Metadata pre-filtering is available at no additional cost in all commercial AWS Regions where Amazon S3 Vectors is available, and in the AWS China Regions. You pay standard S3 Vectors pricing for storage, PUT requests, and queries. For full pricing details, visit the Amazon S3 pricing page. For regional availability, visit Amazon S3 Vectors Regions and quotas.

Whether you’re scoping a RAG application to one tenant, scoping an agent’s searches to one user’s documents, or narrowing a catalog search to a licensing window, pre-filtering lets you apply those filters without trading away recall. To learn more and get started, visit the Amazon S3 Vectors documentation. Send feedback to AWS re:Post for S3 or through your usual AWS Support contacts.

— Daniel Abib

Running production experiments with AWS AppConfig experimentation

Post Syndicated from Aparna Krishnamoorthy original https://aws.amazon.com/blogs/devops/running-production-experiments-with-aws-appconfig-experimentation/

A redesigned checkout button is meant to lift sales. A longer cache time to live (TTL) is meant to cut backend load. But until you test a change against real production traffic, decisions come down to intuition and whoever argues hardest, not evidence. With AWS AppConfig experimentation, you can make data-driven calls instead: using A/B testing, you expose a change to a slice of real users, measure what happens, and let the results decide. It’s part of AWS AppConfig, a capability of AWS Systems Manager, so there’s no separate platform to stand up.

Testing in a development environment confirms a change works, but not how real users or production workloads respond. Releasing to everyone at once answers that but exposes every user to any negative effects, such as broken checkout flows or degraded performance. An experiment is the middle ground: real production traffic, but only a controlled slice of it.

AWS AppConfig experimentation builds on AWS AppConfig feature flags: you define a hypothesis and eligible audience, and a feature flag assigns each participant to the control or a treatment. The AWS AppConfig Agent delivers the right value to each participant, and you keep full control of your data, you join treatment-assignment records with your existing analytics platform.

In this post, you build the following experiments:

  • A frontend experiment that tests a redesigned Add to cart button
  • A backend experiment that compares cache TTL settings in a service running on Amazon Elastic Container Service (Amazon ECS)

You also learn how to record treatment assignments, monitor application health, analyze the results, and promote the winning treatment.

Prerequisites

This post assumes you’ve completed the experimentation prerequisites in the AWS AppConfig User Guide: a feature flag deployed to your environment, the AWS AppConfig Agent installed and configured in your compute environment (including the IAM permissions it needs), and experiment assignment logging enabled on the Agent. To follow the examples here, you also need:

1. AWS Command Line Interface (AWS CLI) 2.35.12 or later, configured with credentials for the account and Region you use. Verify your version with aws –version.

2. A data warehouse or analytics destination for assignment and outcomes data. This post uses Amazon Athena over data in Amazon S3, but AWS AppConfig experimentation works with Amazon CloudWatch or any warehouse you already use, such as Amazon Redshift or Snowflake.

Architecture Overview

AWS AppConfig experimentation adds A/B testing on top of the AWS AppConfig workflow you already use. The control plane lives in AWS AppConfig, treatments are delivered at the edge by the AWS AppConfig Agent, and analysis stays in your existing data warehouse. AWS AppConfig itself provides real-time aggregate traffic metrics; it doesn’t own results analytics, which keeps your metric definitions and data under your control.

The flow looks like this:

AppConfig Experimentation Architecture

Figure 1: AWS AppConfig experimentation — the control plane in AWS AppConfig plus three runtime responsibilities (delivery, measurement, and safety), with the Agent’s assignment log reaching Amazon S3 through Amazon CloudWatch Logs and Amazon Data Firehose 

Beyond the control plane in AWS AppConfig, where you define the experiment and its treatments, three things happen at runtime:

  • Delivery (the data plane). The AWS AppConfig Agent retrieves the feature flag configuration from AWS AppConfig, caches it locally, and asynchronously polls for updates. Your application asks the agent for a flag over the local HTTP endpoint (http://localhost:2772/...), passing an entity Id and any relevant request context. The agent returns the assigned treatment. The Agent returns the same treatment for the same entity for the life of the run, so a user or instance never flips treatments mid-experiment.
  • Measurement. Two data sets meet here. The first is treatment assignments. The Agent writes one JSON record per assignment to standard error. Your log driver ships that record to Amazon CloudWatch Logs, and a subscription filter (matched on the record type) forwards it through Amazon Data Firehose into Amazon S3. The second is your metric events — conversions, latency, cost, and errors — which keep flowing through whatever pipeline you already run.

    Note: We recommend not using sensitive information such as personally identifiable information (PII) for the entity ID, the Agent logs it verbatim. If you must, hash or pseudonymize it identically in both your assignment and metric data – the entity ID is the key that joins the two.
  • Safety. Amazon CloudWatch alarms watch operational and experiment metrics. If an alarm fires during a run, you stop the run, which ends exposure and returns users to the deployed configuration.

One mechanism serves a frontend team, an AI team, and a backend team, each keeping its own metrics and tooling.

Example 1 — Frontend UI Experiment

Scenario. Your team believes a redesigned “Add to cart” button will lift conversion, but you only have a hypothesis. You want to expose it to a slice of production traffic and measure real behavior.

Create the experiment (console) 

The experiment definition ties an application, environment, configuration profile, and feature flag to a control and one or more treatments. Audience rules determine who is eligible, and launch criteria define what counts as success. In the AWS AppConfig console, choose Experiments, then Create experiment, and work through five steps. AI-assisted experiment design in the console can validate your setup against Amazon’s experimentation best practices, helping you catch design gaps before you start a run.

Step 1 — Document your hypothesis. Give the experiment a descriptive Experiment name (add-to-cart-redesign), state the Experiment hypothesis, and use Launch criteria to record the evidence required before you promote a winner — a minimum sample per treatment, the lift you’re looking for, and the guardrails that must not regress. Writing it down now is what makes the result interpretable weeks later, and Validate my hypothesis and launch criteria will review both before you continue. Select StoreFront as the Application name.

Figure 2: Documenting the hypothesis and launch criteria for the add-to-cart experiment



Step 2 — Specify target audience. Describe the audience, then build the rule. The Rule builder tab composes conditions from an attribute, an operator, and a value: $platform equals “web”, joined with And to $geo in [“USA”,”CAN”]. Sample blueprints offers pre-built rules to start from, and the Editor tab shows the same rule as an expression — the form to use if you later automate this:

(and 

  (eq $platform "web") 

  (in $geo ["USA","CAN"]) 

) 



Figure 3: Building the audience rule from two conditions 

That notation is an S-expression — a prefix, function-style form, (operator arg1 arg2 …). The notation is generic; AppConfig defines the operators and the $-prefixed attribute references, which are populated from the caller context your application sends. Here $geo is an attribute your application supplies — the visitor’s country, not an AWS Region.

Note: AWS AppConfig is Regional, so an experiment lives in one AWS Region — don’t use AWS Region as an audience attribute. Segment on caller properties (geography, platform, plan tier, app version); to test across Regions, replicate the experiment definition in each and measure independently.

Step 3 — Select experiment feature flag. Choose the Environment (prod), the Configuration profile holding the flag (Features), and the Feature flag itself (add_to_cart_button). The list shows flags already deployed to that environment.

Figure 4: Selecting the deployed feature flag the experiment will control



Step 4 — Add treatments. Describe the Control as your known-good baseline, confirm its Flag value is toggled ON, and under Attribute values set button_style to classic and button_color to #232F3E. The list includes every attribute the flag defines, including ones this experiment doesn’t vary.

Figure 5: The control treatment, serving the current button 

Then describe Treatment 1 the same way, setting button_style to prominent and button_color to #FF9900. Keep each treatment to a single change so the result stays interpretable. AppConfig allocates traffic evenly across treatments automatically — an even 50/50 here, which is also what maximizes statistical power. Advanced settings offers custom weights, but the even split is the recommended default.

Figure 6: The treatment variant and the even traffic split 

Step 5 — Review and complete. Check the summary and save. Users who aren’t assigned to the experiment continue to receive the default flag value deployed to their environment; only users in the control or a treatment are measured. AppConfig generates a treatment key for each treatment rather than deriving it from your description — it’s the value the Agent returns as _variant — and you use these keys later in treatment overrides and analysis queries.

Figure 7: The saved experiment definition, ready to start a run



Retrieve the treatment (Node.js) 


Your web tier asks the local AWS AppConfig Agent for the flag, passing the visitor identity as Entity-Id so the same visitor always gets the same experience, plus any context the audience rule needs. Request the flag by name with the flag parameter – if you request the whole configuration without it, no experiment assignment is recorded.

// Node.js 18 or later: fetch and Headers are globals, no imports needed. 

 

const AGENT = "http://localhost:2772"; 

const PATH = 

  "/applications/StoreFront/environments/prod/configurations/Features"; 

 

export async function getButtonTreatment(visitorId, geo) { 

  // Multiple context values are sent as repeated "Context" headers, 

  // so use a Headers object with append (an object literal would drop 

  // all but the last "Context" entry). 

  const headers = new Headers(); 

  headers.set("Entity-Id", visitorId); // consistent assignment for the run 

  headers.append("Context", "platform=web"); 

  headers.append("Context", `geo=${geo}`); 

 

  const res = await fetch(`${AGENT}${PATH}?flag=add_to_cart_button`, { 

    headers, 

  }); 

 

  // The request names a single flag, so the agent returns that flag’s 

  // object directly — there is no wrapper keyed by flag name: 

  // { "_variant": "__t1__", "enabled": true, 

  //   "button_style": "prominent", "button_color": "#FF9900" } 

  return await res.json(); 

}

Capture treatment assignments (AWS AppConfig Agent) 

The Agent logs each assignment for you – you don’t write the exposure-logging code. With EXPERIMENT_ASSIGNMENT_LOG_DESTINATION set to stderr, the Agent emits an assignment record the first time it assigns a visitor to a treatment, and that record’s timestamp is what makes clean post-exposure attribution possible later (see the Analyzing experiment results section). Your application’s job is narrower: keep producing the outcome data you already produce — add-to-cart clicks, checkouts, orders. The one thing to get right is the join key:as the Entity-Id, pass an identifier your outcomes dataset already carries — a customer ID, account ID, or session ID — so the records join with no change on your side.

# Amazon ECS, Amazon EKS, or Amazon EC2 - set this on the agent container 

# or host. To collect the files yourself, use a base directory instead of 

# "stderr", in the form file:/tmp/aws-appconfig/assignments/ 

EXPERIMENT_ASSIGNMENT_LOG_DESTINATION=stderr 

 

# AWS Lambda - set this on the function. The extension reads the same 

# setting under a prefixed name. 

AWS_APPCONFIG_EXTENSION_EXPERIMENT_ASSIGNMENT_LOG_DESTINATION=stderr 

 

# One record per assignment. Shown formatted for readability; the agent 

# writes it on a single line so log shippers treat it as one event. 

{ 

  "type": "AWS.AppConfig.TreatmentAssignment", 

  "timestamp": "2026-07-29T16:27:55Z", 

  "region": "us-east-1", 

  "accountId": "111122223333", 

  "applicationId": "dn32rvt", 

  "experimentDefinitionId": "uioedbc", 

  "experimentRunNumber": "5", 

  "treatmentKey": "__t1__", 

  "entityId": "visitor-8f2c" 

} 

Note: The Agent’s ordinary application logs go to the same place as the assignment records, so your forwarder must match on type. Forward everything and you also forward application logs alongside the assignments and break the schema downstream. 

Getting the assignment log from STDERR into Amazon S3 

Three hops take you from the Agent’s standard error to something you can query. The runtime ships STDERR to Amazon CloudWatch Logs, which AWS Lambda and Amazon ECS do for you. A subscription filter forwards only the assignment records. Amazon Data Firehose writes them to Amazon S3. The first hop needs no work, so these two commands are the whole pipeline for the frontend example:

# 1. Firehose stream that lands assignment records in Amazon S3. 

#    The three processors are the step people miss - see the note below. 

aws firehose create-delivery-stream \ 

  --delivery-stream-name experiment-assignments \ 

  --delivery-stream-type DirectPut \ 

  --extended-s3-destination-configuration '{ 

    "RoleARN":   "arn:aws:iam::111122223333:role/FirehoseToS3", 

    "BucketARN": "arn:aws:s3:::my-experiment-data", 

    "Prefix":    "assignments/", 

    "ProcessingConfiguration": { 

      "Enabled": true, 

      "Processors": [ 

        {"Type": "Decompression", 

         "Parameters": [{"ParameterName": "CompressionFormat", 

                         "ParameterValue": "GZIP"}]}, 

        {"Type": "CloudWatchLogProcessing", 

         "Parameters": [{"ParameterName": "DataMessageExtraction", 

                         "ParameterValue": "true"}]}, 

        {"Type": "AppendDelimiterToRecord"} 

      ] 

    } 

  }' 

 

# 2. Forward only assignment records from the agent’s log group to that stream. 

aws logs put-subscription-filter \ 

  --log-group-name "/ecs/storefront" \ 

  --filter-name "appconfig-treatment-assignments" \ 

  --filter-pattern '{ $.type = "AWS.AppConfig.TreatmentAssignment" }' \ 

  --destination-arn \ 

    "arn:aws:firehose:us-east-1:111122223333:deliverystream/experiment-assignments" \ 

  --role-arn "arn:aws:iam::111122223333:role/CWLtoFirehose" 

 

Why the processors matter. CloudWatch Logs doesn’t forward events one at a time — it batches them, gzips each batch, and wraps it in an envelope. With no processing configured, your S3 objects hold compressed JSON, with each assignment record buried as an escaped string inside logEvents[].message.

Three built-in processors undo that. Decompression unzips the batch, CloudWatchLogProcessing with DataMessageExtraction discards the envelope and keeps only the message contents, and AppendDelimiterToRecord puts a newline between records so each lands on its own line. What arrives in Amazon S3 is then exactly the JSON the Agent emitted, one record per row. Leave Firehose compression off, because CloudWatch Logs has already gzipped the payload on the way in. To store Parquet, turn on Firehose data format conversion — it needs decompression enabled too.

Check it before you ramp. Treatment-assignment overrides produce no assignment records (overridden entities would pollute your results), so the 0% window can’t exercise this pipeline. Ramp to a small exposure instead, 1% is sufficient, let real assignments flow, and confirm a record lands under s3://my-experiment-data/assignments/. If the object is gzipped or the record is nested under logEvents, the processors are not configured correctly, and the analysis query later in this post will return nothing. Fix that before you ramp any further.

Amazon ECS without CloudWatch Logs. You can also skip the middle hop entirely. Run FireLens with Fluent Bit as the task’s log router, filter on the same type field, and write straight to Amazon S3. That is fewer moving parts and no envelope to unwrap, in exchange for owning the Fluent Bit configuration yourself.

Either route ends the same way: point an AWS Glue table at the S3 prefix, using the fields from the sample record above. That table is the treatment_assignments source the query in Analyzing experiment results reads, and joining it to your business data is then an ordinary SQL join on the identifier both sides already share. One naming detail to watch when you write that table definition: timestamp is a reserved word in Athena DDL, so the column has to be backtick-quoted there. Queries against the table need no quoting.

The same setup works for AI experiments. Prompt text is configuration rather than code, so a system prompt and its model parameters can live in flag attributes and be varied exactly like the button styling above — no redeploy to reword a prompt. Key the assignment on a session ID so a single conversation doesn’t switch prompts mid-thread, and treat token cost and latency as first-class guardrails, since a “better” prompt that quietly doubles spend isn’t a win.

Example 2 — Backend Cache TTL Experiment

Scenario. You suspect a longer cache TTL will cut database load without noticeably hurting freshness. This is a backend experiment, so the natural unit of assignment isn’t a user — it’s the service instance. Entity-Id set to the instance/task ID gives you stable, instance-level segmentation.

Example 1 used the console, which is the quickest way to get a first experiment running. This example uses the AWS CLI — the same definition expressed as JSON, which is what you’d reach for to script experiment creation or keep it in source control.

Create the experiment (CLI, instance-level segmentation) 

aws appconfig create-experiment-definition \ 

  --application-identifier "Catalog" \ 

  --environment-identifier "prod" \ 

  --configuration-profile-identifier "Features" \ 

  --flag-key "cache_config" \ 

  --name "cache-ttl-tuning" \ 

  --hypothesis "A longer cache TTL reduces DB load without hurting freshness" \ 

  --audience-rule '(eq $service "product-catalog")' \ 

  --control file://control.json \ 

  --treatments file://cache-treatments.json 

Define the variations (JSON) 

The control and each treatment are TreatmentInput objects: a FlagValue (the flag’s Enabled state plus its AttributeValues) and a Weight that sets traffic allocation. Where the console offered a plain Value field, the API takes a typed object — NumberValue for a number, StringValue for a string.

control.json:

{ 

  "Description": "Current 60s cache TTL (baseline)", 

  "Weight": 50.0, 

  "FlagValue": { "Enabled": true, "AttributeValues": { "ttl_seconds": { "NumberValue": 60 } } } 

} 

 

cache-treatments.json:

[ 

  { 

    "Description": "Increase cache TTL to 300s to reduce DB load", 

    "Weight": 50.0, 

    "FlagValue": { "Enabled": true, "AttributeValues": { "ttl_seconds": { "NumberValue": 300 } } } 

  } 

] 

 

Type numeric flag attributes deliberately. Define ttl_seconds as a number attribute on the feature flag with minimum and maximum constraints, and express it in whole seconds. AWS AppConfig validates attribute values when you save the configuration profile, so an out-of-range TTL fails there rather than in production.



Apply in an Amazon ECS service (Java / Spring Boot) 

The AWS AppConfig Agent runs as a sidecar container in the same Amazon ECS task and is reachable at localhost:2772. Use the Amazon ECS task ID as the Entity-Id so each instance holds a consistent treatment for the whole run.

@Component 

public class CacheConfigProvider { 

 

    private static final String AGENT_URL = 

        "http://localhost:2772/applications/Catalog/environments/prod" 

      + "/configurations/Features?flag=cache_config"; 

 

    private final HttpClient http = HttpClient.newHttpClient(); 

    private final String entityId = resolveTaskId(); // instance-level unit 

 

    public CacheConfig getCacheConfig() throws Exception { 

        HttpRequest request = HttpRequest.newBuilder() 

            .uri(URI.create(AGENT_URL)) 

            .header("Entity-Id", entityId) 

            .header("Context", "service=product-catalog") 

            .GET() 

            .build(); 

 

        HttpResponse<String> response = 

            http.send(request, HttpResponse.BodyHandlers.ofString()); 

 

        // single-flag request: no wrapper keyed by flag name 

        JsonNode flag = new ObjectMapper().readTree(response.body()); 

 

        String treatment = flag.get("_variant").asText(); 

        int ttl = flag.get("ttl_seconds").asInt(); 

 

        emitMetrics(entityId, treatment); // the agent logs the assignment 

        return new CacheConfig(ttl, treatment); 

    } 

 

    private String resolveTaskId() { 

        // ECS injects ECS_CONTAINER_METADATA_URI_V4. A GET on 

        // $ECS_CONTAINER_METADATA_URI_V4/task returns the task metadata, 

        // whose TaskARN ends with the task ID. 

        try { 

            String metadataUri = System.getenv("ECS_CONTAINER_METADATA_URI_V4"); 

            HttpRequest metadata = HttpRequest.newBuilder() 

                .uri(URI.create(metadataUri + "/task")) 

                .GET() 

                .build(); 

            String body = 

                http.send(metadata, HttpResponse.BodyHandlers.ofString()).body(); 

            String taskArn = 

                new ObjectMapper().readTree(body).get("TaskARN").asText(); 

            return taskArn.substring(taskArn.lastIndexOf('/') + 1); 

        } catch (Exception e) { 

            // Fail fast. A hard-coded fallback would hand every task the same 

            // Entity-Id, put the whole fleet in one treatment, and quietly 

            // invalidate the experiment. 

            throw new IllegalStateException("Could not resolve the ECS task ID", e); 

        } 

    } 

} 

Your service then emits its guardrail metrics tagged with the same entity_id, so you can compare the 60s and 300s TTL directly.

Configuring Safety Guardrails

An experiment is a production change, so treat it like one: define what “bad” looks like before you ramp. Two controls limit the damage: gradual exposure keeps the exposure small, and Amazon CloudWatch alarms that tell you when to stop the run.

Start safe, ramp gradually. Start every run at 0% audience exposure. At 0%, no traffic is assigned unless you add treatment-assignment overrides — specific entity IDs pinned to a treatment. Use that window to validate the treatment before any real users are exposed: confirm the flag renders the expected experience for each treatment (including the control), confirm your outcome metric logging works, and share the overrides with stakeholders for a preview. For the full validation checklist, see About running and monitoring an experiment.

One thing that window can’t cover: overrides produce no assignment records, so the assignment log pipeline is only exercised once real traffic is being assigned. Clear the overrides, increase exposure in small steps, and treat that first step as the point to confirm assignment records are landing in your warehouse before you ramp further. Watch metrics at each level so a regression hits a small blast radius, not your full audience. Treat overrides as a validation tool, not audience targeting — leave production segmentation to the audience rule.

# Start the run at 0% exposure to validate before exposing any production users. 

# Billing for the run begins with this call and continues until the run is stopped. 

# The overrides pin named entities to a treatment so you can exercise the 

# application path. Overridden entities are deliberately left out of the 

# assignment log, so this window cannot validate that pipeline. 

aws appconfig start-experiment-run \ 

  --application-identifier "StoreFront" \ 

  --experiment-definition-identifier "add-to-cart-redesign" \ 

  --exposure-percentage 0 \ 

  --treatment-overrides '[{"TreatmentKey": "__t1__", 

                            "EntityIds": ["qa-jane", "qa-raj"]}, 

                           {"TreatmentKey": "__control__", 

                            "EntityIds": ["qa-sam"]}]' 

 

Define rollback triggers with Amazon CloudWatch alarms. Before you start a run, decide which metrics indicate unacceptable behavior and create Amazon CloudWatch alarms to watch them. Monitor those alarms throughout the run. If an alarm fires:

  • Evaluate the impact and scope of the regression.
  • Stop the experiment run — this ends audience exposure immediately and returns users to the currently deployed feature flag configuration.

Note that after you increase exposure, it cannot be decreased within the same run. This is intentional to prevent data corruption. To reduce exposure, stop the run and start a new one at a lower percentage.

# Example: alarm on elevated 5xx error rate to watch during the experiment. 

# The dimension is not optional: without it the alarm watches a metric that 

# never receives data, so it sits in INSUFFICIENT_DATA instead of firing. 

aws cloudwatch put-metric-alarm \ 

  --alarm-name "exp-add-to-cart-5xx" \ 

  --namespace "AWS/ApplicationELB" \ 

  --metric-name "HTTPCode_Target_5XX_Count" \ 

  --statistic Sum \ 

  --period 60 \ 

  --evaluation-periods 3 \ 

  --threshold 50 \ 

  --comparison-operator GreaterThanThreshold \ 

  --dimensions Name=LoadBalancer,Value=app/storefront-alb/50dc6c495c0c9188 \ 

  --treat-missing-data notBreaching 

A note on automatic rollback. AWS AppConfig environment monitors (alarms associated with an AppConfig environment) automatically roll back an unhealthy configuration deployment. They are scoped to deployments, not experiment runs:

  • While a run is active, AWS AppConfig manages the flag value for assigned entities. Ending exposure requires an explicit stop-experiment-run call.
  • Keep the monitors in place — you still deploy configurations during and after a run (promoting the winner is a deployment).
  • Treat the alarm-and-stop-the-run pattern above as the guardrail for the experiment itself.

Choose guardrail metrics by experiment type. The right alarm depends on what you’re testing:

  • Frontend/UI: page load time, client-side error rate, rendering failures, 5xx rate. A conversion lift means nothing if the page is throwing errors.
  • Backend: p99 latency, throughput, error rate, and resource-specific signals (for the cache example, cache hit ratio and database load). A treatment can look neutral on business metrics while degrading system health.

Operational hygiene. Follow these rules to keep your results valid:

  • Do not change treatment behavior mid-run — stop, modify, and start a new run instead, or you invalidate the data.
  • Avoid shipping unrelated changes or overlapping experiments on the same audience while a run is active.
  • Monitor operational metrics alongside your experiment metrics — a positive result on the headline metric can still hide a latency or error regression.

Analyzing Experiment Results

AWS AppConfig provides aggregate real-time metrics — exposure levels, treatment allocation, traffic distribution — but it doesn’t compute your results. Your metric definitions and raw data stay in your warehouse (Amazon S3 + Amazon Athena, Amazon Redshift, Snowflake, Databricks, or other) where you control exactly how success is measured.

The core principle: post-exposure attribution. Only count a user’s outcomes after the moment they were assigned to their treatment. Events before assignment don’t attribute to the experiment and bias your results. Concretely, you join your metric events to the Agent’s assignment records on entity ID, and keep only metric events whose timestamp is at or after that entity’s assignment timestamp.

Assuming the Agent’s assignment records and your metric events have landed in Amazon S3 and are queryable through Amazon Athena:

WITH assignments AS ( 

    SELECT 

        entityid                                AS entity_id, 

        treatmentkey                            AS treatment, 

        MIN(from_iso8601_timestamp(timestamp))  AS assigned_at 

    FROM treatment_assignments   -- AWS AppConfig Agent records from STDERR 

    WHERE type = 'AWS.AppConfig.TreatmentAssignment' 

      AND experimentdefinitionid = 'uioedbc' 

      AND experimentrunnumber = '5' 

    GROUP BY entityid, treatmentkey 

), 

attributed_conversions AS ( 

    SELECT 

        a.treatment, 

        a.entity_id, 

        COUNT(m.entity_id) AS conversions 

    FROM assignments a 

    LEFT JOIN experiment_events m          -- your existing outcomes table 

        ON  m.entity_id = a.entity_id      -- or m.customer_id, whichever column it already has 

        AND m.event_type = 'conversion' 

        -- post-exposure attribution: outcomes only count after assignment 

        AND from_iso8601_timestamp(m.timestamp) >= a.assigned_at 

    GROUP BY a.treatment, a.entity_id 

) 

SELECT 

    treatment, 

    COUNT(DISTINCT entity_id)                         AS assigned_users, 

    SUM(CASE WHEN conversions > 0 THEN 1 ELSE 0 END)  AS converters, 

    ROUND( 

        SUM(CASE WHEN conversions > 0 THEN 1 ELSE 0 END) * 100.0 

        / COUNT(DISTINCT entity_id), 2 

    )                                                 AS conversion_rate_pct 

FROM attributed_conversions 

GROUP BY treatment 

ORDER BY conversion_rate_pct DESC; 

This returns assigned users, converters, and conversion rate per treatment so you can compare each treatment against the control.

The only requirement on your outcomes data is that it carries the same identifier you passed as Entity-Id and a timestamp. Whatever table structure, column names, or warehouse you already use works — the join is on that shared identifier with a timestamp filter.

Adapting the query for other metrics. The assignments CTE and the post-exposure join are reusable; only the metric aggregation changes:

  • Backend/cache experiments: aggregate AVG(db_query_count), cache hit ratio, or approx_percentile(latency_ms, 0.99) per treatment to confirm the longer TTL cut load without hurting p99.
  • Continuous metrics generally: replace the converter count with AVG(...), SUM(...), or approx_percentile(...) over the attributed rows.

Before you declare a winner, check three things. First, confirm each treatment reached the sample size you set in your launch criteria. Second, confirm the split matches the configured weights; a 50/50 experiment that lands at 46/54 points to an assignment or logging bug, not a result. Third, run a significance test in your statistics tooling, such as a two-proportion z-test for conversion rate. Act on the result only when all three checks pass.

Stopping an experiment and promoting a winner

Stop an experiment run when you have a clear result, when something goes wrong, or when priorities shift. Stopping ends exposure immediately. AWS AppConfig stops managing the flag and your application serves whatever configuration is currently deployed to the environment.

Promoting the winner without a gap. The order matters. If you stop first, users briefly revert to the pre-experiment default while you redeploy. To avoid that:

  • While the experiment is still running, update the feature flag to match the winning treatment and deploy it. Assigned entities see no change — AWS AppConfig is still serving them their treatment.
  • Stop the experiment. AppConfig releases the flag, and your application picks up the configuration you just deployed: the winner, at 100%, with no gap.

Mind what you change in step 1. Add the winning values as a variant gated by the same audience rule the experiment uses, and nobody sees a change until you stop the run. Change the flag’s default value instead and everyone outside the experiment’s audience picks up the winning value the moment the deployment lands — choose this when you want a full release.

Cost considerations and cleaning up

You pay for experiment-run hours. Billing starts when you call start-experiment-run until you stop the run. Defining experiments and treatments is free.

What drives your bill:

  • Run duration. Stop the run once you have enough data to make a decision.
  • Concurrent runs. Each active run bills independently. Three simultaneous experiments means three times the hourly rate.
  • Your data pipeline. AWS AppConfig doesn’t charge for the assignment records the Agent emits; the pipeline that carries them does. CloudWatch Logs, Amazon Data Firehose, Amazon S3, Athena, and your guardrail alarms each bill at their normal rates — see each service’s pricing page, and AWS Systems Manager Pricing for experiment runs.

To keep costs down, estimate sample size upfront so you know roughly how long a run needs to last, and validate what you can in the 0% window before you ramp. The assignment-log pipeline is the exception: it produces records — and bills — only once real traffic is being assigned.

See AWS AppConfig experimentation pricing details here.

Clean up what you created. The run-hour charge stops only when you stop the run, and the assignment pipeline keeps billing for as long as it stays in place. When you have finished with the examples in this post, remove what you created in this order:

  • Stop any running experiment run. Exposure ends immediately and the run-hour charge stops. If you are promoting a winner, deploy the winning flag value first, as described above.
  • Delete the experiment definitions for both examples. ARCHIVE hides a definition but keeps its run history; DESTROY removes the definition and the history permanently.
  • Take down the assignment pipeline and the alarms: the Amazon CloudWatch Logs subscription filter, the Amazon Data Firehose delivery stream, and the guardrail alarms. Then unset EXPERIMENT_ASSIGNMENT_LOG_DESTINATION on the Agent and redeploy so it stops writing assignment records.
  • Decide what to do with the data. The assignment records in Amazon S3, the AWS Glue table over them, and your Athena query-results location all keep incurring storage charges. Delete them unless you want to keep the audit trail, along with the IAM roles and the log group you created only for this walkthrough.

Leave the feature flag and its configuration profile in place if your application still reads them — deleting the flag removes configuration your code depends on. Only the experiment definition has to go.

1. Stop the run – this ends exposure and the run-hour charge. 

aws appconfig stop-experiment-run \ 

  --application-identifier "StoreFront" \ 

  --experiment-definition-identifier "add-to-cart-redesign" \ 

  --run 5 

2. Delete both definitions. Use ARCHIVE instead of DESTROY to keep the run history for future reference

#    the run history for future reference. 

aws appconfig delete-experiment-definition \ 

  --application-identifier "StoreFront" \ 

  --experiment-definition-identifier "add-to-cart-redesign" \ 

  --delete-type DESTROY 

 

aws appconfig delete-experiment-definition \ 

  --application-identifier "Catalog" \ 

  --experiment-definition-identifier "cache-ttl-tuning" \ 

  --delete-type DESTROY 

3. Remove the assignment pipeline and the guardrail alarm. 

aws logs delete-subscription-filter \ 

  --log-group-name "/ecs/storefront" \ 

  --filter-name "appconfig-treatment-assignments" 

 

aws firehose delete-delivery-stream \ 

  --delivery-stream-name "experiment-assignments" 

 

aws cloudwatch delete-alarms --alarm-names "exp-add-to-cart-5xx" 

4. Optional and irreversible – drop the queryable copy of the assignment data. Substitute your own AWS Glue database name. 

aws glue delete-table \ 

  --database-name "experiments" \ 

  --name "treatment_assignments" 

 

aws s3 rm "s3://my-experiment-data/assignments/" --recursive 

Conclusion

In this post, we took a single idea — “we have a theory, but no production evidence” — and turned it into two concrete experiments using AWS AppConfig experimentation: a frontend button redesign and a backend cache-TTL change. In each case, we created an experiment definition and expressed the control and treatments as feature-flag variants. We delivered them through the Agent with gradual exposure and alarm guardrails, and analyzed results with post-exposure attribution in our own data warehouse.

Experimentation becomes part of the AWS AppConfig workflow you already use, and you keep ownership of your metrics, analysis, and data. You pay per experiment-run hour, so your cost grows with how much you test.

To go deeper, start with the AWS AppConfig experimentation documentation, review running and monitoring an experiment for guardrail best practices and try the hands-on workshop.

If you have questions or feedback, leave a comment on this post. To get started, open the AWS AppConfig console and create your first experiment definition.

Running multi-day AZ evacuation drills with ARC Zonal Shift

Post Syndicated from Antoine Boucherie original https://aws.amazon.com/blogs/architecture/running-multi-day-az-evacuation-drills-with-arc-zonal-shift/

Running a multi-day Availability Zone evacuation drill with ARC Zonal Shift is an effective way to prove your application can withstand a sustained impairment. Multi-Availability Zone deployment is an architectural best practice for building resilient applications on AWS, but there is a gap between deploying multi-AZ and proving it works under sustained stress. Traditional disaster recovery (DR) tests validate the failover mechanism. A typical test shifts traffic, confirms targets respond, and rolls back within minutes. These short exercises don’t surface the issues that only appear over hours or days.

A multi-day AZ evacuation forces time-dependent behaviors to play out completely, exposing failure modes that brief tests miss:

  • Auto Scaling policies that aren’t tuned for sustained N-1 operation over a full day.
  • Deployment pipelines that don’t validate AZ health before placing new workloads.
  • Stale DNS or cached database endpoints.
  • Time-based routine operational processes tested against an N-1 architecture (certificate rotations, credential and secret rotations, maintenance windows, log rotation, backup automation, and cron jobs).
  • Long-lived database connections pinned to a specific AZ that are only used infrequently.
  • Recovery after a sustained multi-day shift, which is a different operational procedure than rolling back to a warm, nearly identical AZ within minutes.

By shifting all traffic away from a single AZ for 48–72 hours using Amazon Application Recovery Controller (ARC) Zonal Shift, you force your architecture to sustain full production load on N-1 zones. This proves capacity sufficiency, database stability, and client reconnection behavior, but most importantly, that your teams can operate normally for days on N-1 capacity.

This post shows you how to plan and run a multi-day AZ evacuation drill across Amazon Elastic Container Service (Amazon ECS), Amazon Elastic Kubernetes Service (Amazon EKS), Amazon Relational Database Service (Amazon RDS) for PostgreSQL, and Amazon Aurora PostgreSQL, with step-by-step CLI commands, a prerequisites section, observability metrics, and restore procedures.

Why financial services institutions are doing this already

Financial services organizations face unique regulatory pressure to demonstrate, not just document, their disaster recovery capabilities. Across the world, there is an increasing focus on building and demonstrating operational resilience within regulated entities. This is shifting the mindset from “show us your runbook” to “show us the evidence”. This uplift in control environments is driving financial services companies to conduct live DR testing under realistic conditions and produce auditable proof of recovery within defined timeframes. Some insurers and banks are now running periodic AZ evacuation drills as part of their operational resilience programs, shifting from “we have multi-AZ” to “we have proven multi-AZ”. Using ARC Zonal Shift, you can shift traffic at the infrastructure layer without changing application code. It works natively across Application Load Balancer, Network Load Balancer, Amazon Elastic Compute Cloud (Amazon EC2) Auto Scaling groups, and Amazon EKS clusters.

Solution overview

In this walkthrough, we demonstrate how to evacuate an Availability Zone for a multi-tier digital platform running on AWS.

We deliberately include both ECS and EKS, and two database engines, to show the evacuation procedure for each major service type readers are likely to run. The architecture is illustrative only.

The following table outlines the architecture:

Layer Components Multi-AZ Configuration
Traffic ingress Application Load Balancer (ALB) fronting ECS Deployed across 3 AZs, cross-zone load balancing activated
Compute (containers) Amazon ECS (Fargate) Stateless tasks distributed across 3 AZ subnets
Traffic ingress Network Load Balancer (NLB) fronting EKS Deployed across 3 AZs, cross-zone load balancing activated
Compute (Kubernetes) Amazon EKS or EKS Auto Mode Stateless services with topology spread constraints across 3 AZs
Database Amazon RDS for PostgreSQL Multi-AZ: primary in AZ A, standby in AZ B
Database Amazon Aurora PostgreSQL Writer in AZ A, reader in AZ B. Storage replicated across all 3 AZs

Note that the Aurora storage layer differs from standard RDS Multi-AZ. Aurora synchronously replicates data to six storage nodes across Availability Zones independently of compute instances, so only the writer or reader instance needs to be failed over as the storage remains fully available throughout the drill.

In this walkthrough, we evacuate AZ A, the zone hosting the RDS primary and Aurora writer.

Multi-tier architecture spanning three Availability Zones: an Application Load Balancer fronting Amazon ECS and a Network Load Balancer fronting Amazon EKS, with Amazon RDS for PostgreSQL and Amazon Aurora PostgreSQL databases, before evacuating AZ A.

How ARC Zonal Shift works

When you initiate a zonal shift, ARC takes two coordinated actions for Amazon Route 53 and Elastic Load Balancers:

  1. DNS removal: The load balancer’s IP address in the affected AZ is removed from DNS. New client queries don’t resolve to that endpoint.
  2. Cross-zone traffic blocking: Load balancer nodes in the remaining AZs stop routing requests to targets in the shifted AZ, even when cross-zone load balancing is enabled.

For Amazon EKS clusters with zonal shift enabled, ARC goes further. It performs the following actions:

  • Cordons all nodes in the impacted AZ, preventing new pod scheduling.
  • Removes pod endpoints in the impacted AZ from EndpointSlice resources, redirecting east-west service-to-service traffic to healthy AZs.
  • Suspends AZ rebalancing for managed node groups and updates ASGs to launch instances only in healthy AZs.
  • Preserves nodes and pods in the shifted AZ (they are not terminated), keeping full capacity available for when the shift ends.

Combined with service-specific procedures for ECS task redistribution and database failover, this creates a complete AZ evacuation across all three traffic dimensions: north-south ingress, east-west service communication, and outbound database connections.

ARC Zonal Shift is a data plane operation by design. Because it works independently of the AWS control plane, it remains available even during an AZ impairment. The other steps in this walkthrough (ECS service updates, manual RDS failovers, manual Aurora failovers, subnet group modifications) are control plane operations. For a planned drill, this distinction has no practical impact because the control plane is healthy. During a real AZ impairment, prioritize the data plane action (start the zonal shift first to stop traffic immediately) and perform control plane operations only after traffic has already been shifted.

Prerequisites

Configure these prerequisites well in advance of your first shift. These are foundational settings that verify that your architecture is shift-ready at all times. For this walkthrough, you should have the following:

  • An AWS account
  • A multi-tier application deployed across 3 Availability Zones with ALB/NLB, ECS or EKS workloads, and RDS or Aurora databases.
  • IAM permissions to manage ARC Zonal Shift, ECS, EKS, and RDS resources.
  • AWS Command Line Interface (AWS CLI) v2 installed and configured.
  • Familiarity with ARC Zonal Shift concepts.
  • Amazon CloudWatch dashboards with per-AZ metric breakdowns (fault rate, latency, and target health).
  • Auto Scaling policies validated for sustained N-1 AZ operation.

Specifically for Elastic Load Balancing (ELB):

  • ALB/NLB deregistration delay set to 60 seconds, which allows existing connections to drain quickly after a shift instead of the default 300 seconds.
  • target_group_health.dns_failover.minimum_healthy_targets.count configured on each target group.

Specifically, for EKS:

  • kubectl installed and configured for your EKS cluster.
  • Turn on Topology Aware Routing on EKS services (or configure Istio locality-aware load balancing).
  • Zonal shift activated on your EKS cluster (one-time setup).

Specifically, for ECS:

  • ECS stopTimeout set to 55 seconds in task definitions, slightly below the ALB deregistration delay so tasks finish in-flight requests before being force-stopped, avoiding 502 errors during the drain window.

Specifically, for RDS:

  • Verify that your RDS primary and standby are provisioned in different Availability Zones.

What changes for a multi-day shift

The mechanics of starting a zonal shift are the same whether you run it for 1 hour or 72 hours. What changes is the operational surface area:

  • Expiry management: Zonal shifts have a maximum duration. You must monitor and extend them before they expire, or traffic returns to the shifted AZ unexpectedly.
  • Scaling drift: Over days, Auto Scaling events in healthy AZs may create capacity imbalances. Monitor and cap scaling so recovery doesn’t overload the returning AZ.
  • Connection pool cycling: After 24+ hours, most client connections will have recycled. This validates DNS TTL compliance across your entire client fleet, something a 1-hour test won’t fully exercise.
  • Operational confidence: Teams will learn to deploy, patch, and troubleshoot in a reduced AZ environment. A multi-day drill forces this to happen naturally rather than in a controlled window.
  • Safe recovery: After days at N-1 capacity, restoring the shifted AZ requires careful ordering. Verify health, scale back gradually, and reintroduce traffic incrementally rather than all at once.

Solution details

Each section below walks through the zonal shift procedure for one layer of the architecture, starting with the compute tier and finishing at the database layer.

Amazon ECS — Zonal Shift with task redistribution

For ECS services behind an ALB, initiating a zonal shift at the load balancer layer stops new traffic from reaching targets in the evacuated AZ. Existing ECS tasks in that AZ remain running but stop receiving requests. To perform a complete evacuation, follow these steps:

Step 0. Before starting, verify ARC Zonal Shift is enabled on the load balancer (disabled by default)

aws elbv2 modify-load-balancer-attributes \
    --load-balancer-arn $ALB_ARN \
    --attributes Key=zonal_shift.config.enabled,Value=true

Step 1. Initiate the zonal shift on the load balancer.

aws arc-zonal-shift start-zonal-shift \
    --resource-identifier $ALB_ARN \
    --away-from $AZ_ID_TO_EVACUATE \
    --expires-in "72h" \
    --comment "Multi-day AZ evacuation drill"

Zonal shifts expire after the duration set in --expires-in. If the shift expires before you cancel it, traffic automatically returns to the shifted AZ. For a multi-day drill, monitor the remaining time and extend before expiry using:

aws arc-zonal-shift update-zonal-shift \
    --zonal-shift-id $SHIFT_ID \
    --resource-identifier $RESOURCE_ARN \
    --expires-in "24h" \
    --comment "Extending drill"

When cross-zone load balancing is enabled (the default for ALB), the shift instructs load balancer nodes in healthy AZs not to route requests to targets in the impaired AZ. Targets are fully isolated regardless of your cross-zone configuration.

Step 2. Restrict new task placement to healthy AZs.

Update the ECS service’s network configuration to exclude subnets in the evacuated AZ. This prevents new tasks from launching in the shifted zone:

aws ecs update-service \
    --cluster $CLUSTER_NAME \
    --service $SERVICE_NAME \
    --network-configuration "awsvpcConfiguration={subnets=[$AZB_SUBNET,$AZC_SUBNET],securityGroups=[$SG_ID],assignPublicIp=DISABLED}"

Step 3. If needed, scale to verify N-1 AZ capacity.

aws ecs update-service \
    --cluster $CLUSTER_NAME \
    --service $SERVICE_NAME \
    --desired-count $N_MINUS_1_COUNT

Note: updating the ECS service network configuration and count that you want are control plane operations. Perform these changes before the drill starts, not during a real AZ impairment when control plane availability may be degraded.

We recommend that you pre-scale your services to handle the loss of an AZ’s worth of capacity before the drill. Your architecture should tolerate AZ loss without needing to scale reactively. See static stability in the Amazon Builders’ library.

In a 3 AZ environment, pre-scaling for N-1 capacity means running approximately 50% more compute than your baseline peak requires. If the cost isn’t justifiable for all workloads, consider scheduled scaling policies that increase capacity during planned drill windows, Auto Scaling with aggressive scale-out thresholds, or load shedding mechanisms. Keep in mind that during an unplanned impairment, you won’t have time to scale reactively. Workloads that aren’t pre-scaled will operate in a degraded state until scaling catches up, which can take minutes under load.

Step 4. Monitor task distribution.

aws ecs describe-tasks \
    --cluster $CLUSTER_NAME \
    --tasks $(aws ecs list-tasks --cluster $CLUSTER_NAME --service-name $SERVICE_NAME --query 'taskArns' --output text) \
    --query 'tasks[].[taskArn,availabilityZone]' --output table

Restore: Revert the network configuration to include all three AZ subnets, then cancel the zonal shift. Tasks will gradually rebalance across all AZs during subsequent deployments.

Important: zonal shift won’t work for single-AZ target groups as the ALB will refuse the shift if healthy targets only exist in one Availability Zone. Verify each target group has targets registered in at least two AZs before proceeding. For more details, refer to Application Load Balancers in the ARC documentation.

Amazon EKS — Zonal Shift with EndpointSlice isolation

Amazon EKS natively supports ARC zonal shift. When you turn on this capability and trigger a shift, ARC handles both the infrastructure and Kubernetes networking layers automatically.

What ARC does when you shift an EKS cluster:

  • Nodes in the impacted AZ are cordoned (no new pod scheduling).
  • The built-in Kubernetes EndpointSlice controller removes pod endpoints in the impacted AZ, so east-west service traffic is automatically redirected to pods in healthy AZs.
  • For managed node groups, AZ rebalancing is suspended and ASGs are updated to only launch in healthy AZs.
  • Nodes and pods in the shifted AZ are not terminated, ensuring full capacity is immediately available when the shift ends.

Step 1. Activate zonal shift for your EKS cluster (one-time setup):

aws eks update-cluster-config \
    --name $CLUSTER_NAME \
    --zonal-shift-config enabled=true

Step 2. Start the zonal shift on both the load balancer and EKS cluster:

# Shift north-south traffic at the load balancer
aws arc-zonal-shift start-zonal-shift \
    --resource-identifier $NLB_ARN \
    --away-from $AZ_ID_TO_EVACUATE \
    --expires-in "72h" \
    --comment "Multi-day AZ evacuation drill - north-south traffic"

# Shift east-west traffic within the EKS cluster
aws arc-zonal-shift start-zonal-shift \
    --resource-identifier $EKS_CLUSTER_ARN \
    --away-from $AZ_ID_TO_EVACUATE \
    --expires-in "72h" \
    --comment "Multi-day AZ evacuation drill - east-west traffic"

Step 3. Verify EndpointSlice update: confirm pods in the shifted AZ are no longer receiving traffic:

# Endpoints should only show AZ B/AZ C
kubectl get endpointslices -l kubernetes.io/service-name=$SERVICE_NAME -o yaml | \
    grep -A2 "zone:"

Step 4. Verify node and pod status:

# Nodes in evacuated AZ should show SchedulingDisabled
kubectl get nodes -l topology.kubernetes.io/zone=$AZ_NAME_TO_EVACUATE

# Confirm traffic distribution across healthy AZs
kubectl get pods -o wide -l app=$APP_LABEL

Verify your pods use topologySpreadConstraints with maxSkew: 1 on topology.kubernetes.io/zone and are pre-scaled to handle N-1 AZ load. The zonal shift doesn’t evict pods or trigger autoscaling by itself.

Note that ARC zonal shift doesn’t control outbound connections from pods to external dependencies like Amazon RDS. If your pods connect to AZ-specific database endpoints, consider using Istio with locality-aware routing. For implementation details, refer to End-to-end recovery from AZ impairments in EKS using Zonal Shift and Istio.

For Aurora, the cluster endpoint automatically routes to the current writer regardless of AZ, so no Istio configuration is needed for writer traffic. However, if you use AZ-specific reader instance endpoints, configure Istio ServiceEntry resources for each endpoint and apply a DestinationRule with localityLbSetting to prefer healthy AZs. This directs outbound database traffic to follow the same shift pattern as your north-south and east-west traffic.

Restore: Cancel both zonal shifts. ARC automatically uncordons nodes, adds pod endpoints back to EndpointSlices, and restores AZ rebalancing. Traffic returns to all three AZs with full capacity already in place.

aws arc-zonal-shift cancel-zonal-shift \
    --zonal-shift-id $EKS_SHIFT_ID \
    --resource-identifier $EKS_CLUSTER_ARN

aws arc-zonal-shift cancel-zonal-shift \
    --zonal-shift-id $NLB_SHIFT_ID \
    --resource-identifier $NLB_ARN

Amazon RDS for PostgreSQL — multi-AZ failover

Regarding Amazon RDS for PostgreSQL in a Multi-AZ deployment, if the primary instance resides in the AZ being evacuated, you must trigger a failover to the synchronous standby. RDS handles this through a reboot with failover.

Prerequisite: verify your RDS primary and standby are provisioned in different Availability Zones. If both are in the same zone, the following steps wouldn’t evacuate the zone as intended.

Step 1. Check current primary location:

aws rds describe-db-instances \
    --db-instance-identifier $RDS_INSTANCE \
    --query 'DBInstances[0].[DBInstanceIdentifier,AvailabilityZone,MultiAZ,SecondaryAvailabilityZone]' \
    --output table

Step 2. If the primary is in the evacuated AZ, manually force failover:

aws rds reboot-db-instance \
    --db-instance-identifier $RDS_INSTANCE \
    --force-failover

Step 3. Wait for availability and verify the new primary AZ:

aws rds wait db-instance-available \
    --db-instance-identifier $RDS_INSTANCE

aws rds describe-db-instances \
    --db-instance-identifier $RDS_INSTANCE \
    --query 'DBInstances[0].AvailabilityZone'

After failover, RDS recreates the standby in the evacuated AZ automatically. This is acceptable for a sustained drill as the standby receives no client traffic. Monitor ReplicaLag to confirm replication health.

Step 4. (Optional) Remove the standby from the evacuated AZ.

If you want zero RDS presence in the evacuated Availability Zone, you can relocate the standby by modifying the DB subnet group:

  1. Create a manual snapshot as a safety net.
  2. Disable Multi-AZ on the instance.
  3. Modify the DB subnet group to include only healthy AZ subnets, removing the evacuated AZ subnet.
  4. Re-enable Multi-AZ so the new standby is created in one of the remaining healthy AZs.

This approach works the same way for Amazon RDS for PostgreSQL as it does for any RDS engine using Multi-AZ deployments. Note that RDS Multi-AZ modifications (disabling/re-enabling Multi-AZ, subnet group changes) can take several minutes to complete. Plan for this during the drill window.

Amazon Aurora PostgreSQL — writer failover & reader management

Aurora provides more control over AZ placement than standard RDS Multi-AZ. You can explicitly choose which reader to promote and manage reader placement across AZs using failover priority tiers.

Step 1. Identify the cluster topology:

aws rds describe-db-clusters \
    --db-cluster-identifier $CLUSTER_ID \
    --query 'DBClusters[0].DBClusterMembers[].{Instance:DBInstanceIdentifier,IsWriter:IsClusterWriter}'

aws rds describe-db-instances \
    --filters Name=db-cluster-id,Values=$CLUSTER_ID \
    --query 'DBInstances[].[DBInstanceIdentifier,AvailabilityZone,DBInstanceStatus]' \
    --output table

Step 2. If the writer is in the evacuated AZ, failover to a reader in a healthy AZ:

aws rds failover-db-cluster \
    --db-cluster-identifier $CLUSTER_ID \
    --target-db-instance-identifier $READER_IN_HEALTHY_AZ

Step 3. Wait for the cluster to stabilize:

aws rds wait db-cluster-available \
    --db-cluster-identifier $CLUSTER_ID

If your Aurora cluster has no pre-existing reader in a healthy AZ, writer promotion requires creating a new instance, which typically takes less than 10 minutes. Pre-provisioning a reader in a separate AZ reduces failover time, often to less than 30 seconds.

Step 4. (Optional) Remove the reader in the evacuated AZ and create one in a healthy AZ.

For a full AZ evacuation where you want zero database presence in the shifted zone:

# Delete the reader instance in the evacuated AZ
aws rds delete-db-instance \
    --db-instance-identifier $INSTANCE_IN_EVACUATED_AZ \
    --skip-final-snapshot

# Create a new reader in a healthy AZ
aws rds create-db-instance \
    --db-instance-identifier ${CLUSTER_ID}-reader-${TARGET_AZ} \
    --db-cluster-identifier $CLUSTER_ID \
    --db-instance-class $INSTANCE_CLASS \
    --engine aurora-postgresql \
    --availability-zone $TARGET_AZ

Step 5. Monitor replication and performance throughout the drill:

aws cloudwatch get-metric-statistics \
    --namespace AWS/RDS \
    --metric-name AuroraReplicaLag \
    --dimensions Name=DBInstanceIdentifier,Value=$READER_INSTANCE \
    --start-time $TIMESTAMP_5MIN_AGO \
    --end-time $TIMESTAMP \
    --period 60 --statistics Average

Monitoring the drill with CloudWatch

A multi-day drill is only as valuable as the evidence it produces. Unlike a brief failover test where you visually confirm targets respond, a 48–72-hour evacuation requires continuous, automated observation, capturing capacity trends, replication health, and latency shifts that only surface under sustained N-1 AZ load.

Before starting the drill, verify you have CloudWatch dashboards with per-AZ metric breakdowns for each layer of your architecture. During the drill, these metrics serve two purposes: real-time operational awareness and post-drill evidence for stakeholders.

Key metrics by layer

The following metrics give you real-time visibility into each layer of the architecture during the drill.

Application Load Balancer / Network Load Balancer

Metric Dimension What to watch
HealthyHostCount Per target group, per AZ Should drop to 0 in evacuated AZ. Stable in healthy AZs
UnHealthyHostCount Per target group, per AZ Targets in evacuated AZ may show unhealthy (expected)
RequestCount Per AZ Zero traffic in shifted AZ. Even distribution in remaining AZs
TargetResponseTime Per AZ Watch for latency increases in healthy AZs under concentrated load
HTTPCode_Target_5XX_Count Per target group Sustained increase signals capacity pressure

Amazon ECS

Metric Dimension What to watch
CPUUtilization Per service Should not exceed 70–80% sustained (indicates capacity headroom)
MemoryUtilization Per service Memory pressure under concentrated load
RunningTaskCount Per service Confirms tasks running only in healthy AZs
DesiredTaskCount vs RunningTaskCount Per service Gap indicates placement failures (check subnet/capacity)

Amazon EKS (using Container Insights)

Metric Dimension What to watch
node_cpu_utilization Per node, filtered by AZ Nodes in healthy AZs absorbing shifted load
pod_cpu_utilization Per pod/namespace Hotspot detection under N-1 operation
node_status_condition Per node Nodes in evacuated AZ should show SchedulingDisabled
pod_number_of_container_restarts Per pod Restart loops may indicate resource pressure

Amazon RDS for PostgreSQL

Metric Dimension What to watch
CPUUtilization Per instance Primary under higher load post-failover
DatabaseConnections Per instance Connection spike after failover (watch for pool exhaustion)
ReadIOPS / WriteIOPS Per instance I/O patterns shift when primary moves AZs
ReplicaLag Per standby Should stabilize within seconds after failover
FreeableMemory Per instance Memory pressure under full client reconnection

Amazon Aurora PostgreSQL

Metric Dimension What to watch
AuroraReplicaLag Per reader instance Establish your cluster baseline during normal operation. Sustained increases from baseline indicate storage pressure. Aurora Replicas typically lag 100 ms or less
CommitLatency Per writer Increased commit latency indicates write contention
BufferCacheHitRatio Per instance Drop below 99% may indicate working set doesn’t fit in memory
DatabaseConnections Per instance Client reconnection behavior after writer promotion
VolumeBytesUsed Per cluster Aurora storage is AZ-independent (should be unaffected)

Export your per-AZ CloudWatch dashboards as snapshots before, during, and after the drill. Combine these with the ARC zonal shift event history (available through list-zonal-shifts) to create an auditable evidence package.

Cleaning up

After completing the drill, restore services carefully. The order matters, especially if Auto Scaling has increased capacity in healthy AZs:

  1. Verify the evacuated AZ is healthy: confirm targets are registered, pods are running, and database instances are available.
  2. Cancel the EKS cluster zonal shift first (east-west traffic resumes). Monitor for errors as internal traffic rebalances.
  3. Cancel the load balancer zonal shift (north-south traffic resumes). Traffic returns gradually as DNS propagates.
  4. If Auto Scaling added capacity in the remaining AZs, scale back gradually over 15 to 30 minutes. Don’t remove capacity before traffic has redistributed evenly.
  5. If cross-zone load balancing is disabled, verify target_group_health.dns_failover.minimum_healthy_targets.count is configured. This allows Route 53 to only route traffic to an AZ once it has enough healthy targets registered, preventing the restored zone from receiving traffic before it’s ready to handle it.
  6. Monitor per-AZ metrics for 30 minutes after restoring to confirm even distribution and no error spikes.

No additional AWS resources are created by ARC Zonal Shift that incur ongoing charges. The zonal shift itself is available at no additional cost.

Conclusion

In this post, you learned how to run a sustained AZ evacuation drill using ARC Zonal Shift across Amazon Elastic Container Service (Amazon ECS), Amazon Elastic Kubernetes Service (Amazon EKS), Amazon RDS for PostgreSQL, and Amazon Aurora. By operating on N-1 Availability Zones for 48–72 hours, rather than a brief failover test, you produce evidence that your multi-AZ architecture delivers genuine, sustained resilience. This is particularly valuable for financial services organizations facing regulatory mandates that require demonstrated recovery capabilities.

To get started, use the prerequisites checklist and step-by-step procedures in this post. Begin in non-production, progress to production during low-traffic windows, and build toward sustained operation under shift. As confidence grows, activate zonal autoshift so AWS can shift traffic automatically when internal telemetry detects an impairment.

You can also use AWS Resilience Hub to assess your application’s resilience posture before and after the drill. It validates that your architecture meets your defined Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets.

Amazon Application Recovery Controller – Zonal Shift

Best practices for zonal shifts in ARC

Using cross-zone load balancing with zonal shift

New AWS Fault Injection Service recovery action for zonal autoshift

End-to-end recovery from AZ impairments in Amazon EKS using EKS Zonal Shift and Istio

Amazon EKS now supports Amazon Application Recovery Controller


About the authors

How MHK built a HIPAA-eligible agentic AI solution on Amazon Bedrock

Post Syndicated from Deepti Tirumala original https://aws.amazon.com/blogs/architecture/how-mhk-built-a-hipaa-eligible-agentic-ai-solution-on-amazon-bedrock/

Healthcare organizations face an increasingly complex challenge: processing vast volumes of medical documents, including clinical records, claims, prior authorizations, appeals, and pharmacy data, while maintaining strict HIPAA compliance and security standards. Traditional approaches require dedicated engineering teams to build individual AI systems for each use case, each needing its own compliance infrastructure, audit trails, and security controls. Agentic frameworks that can scale across use cases can reduce lengthy development cycles and high operational overhead.

MHK, a Hearst Health company ranked #1 in payer care management solutions in the 2024 Best in KLAS: Software & Services Report, faced this exact challenge. Their medical management solution serves health plans across multiple workflows (medical, pharmacy, grievance, and appeals), each requiring intelligent document processing and decision support. Building separate AI systems for each workflow was unsustainable as demand grew.

To solve this, they developed the SmartProminence AI Orchestrator, a HIPAA-eligible agentic workflow framework built on AWS that reduced manual medical review effort by 90%. New AI features that previously took 3+ months to deploy now ship in 2 weeks.

In this post, we walk through how MHK architected this solution using Amazon Bedrock, Amazon Elastic Container Service (Amazon ECS), and event-driven patterns to create a reusable, multi-tenant orchestrator for healthcare AI.

MHK uses Amazon Bedrock exclusively for foundation model inference. They built their own orchestration, retrieval, and validation layers because healthcare workflows require domain-specific controls: DAG-based multi-step execution, clinical document retrieval tied to case context, and HIPAA-specific validation logic that goes beyond general-purpose guardrails. This approach keeps Bedrock focused on scalable model access while MHK retains full control over workflow behavior and compliance enforcement.

Background

MHK provides healthcare cost management and compliance solutions to health plans across the United States. Their medical management system supports the full lifecycle of care decisions, from the moment a provider submits a request for service through final resolution. This includes prior authorization, claims adjudication, appeals processing, pharmacy benefit verification, and medical director reviews.

Each workflow involves analyzing unstructured medical documents such as clinical notes, lab results, imaging reports, and multi-page faxed records, against structured policy criteria. Before MHK’s SmartProminence AI Orchestrator, case managers spent 5 to 10 minutes manually processing each incoming document, while medical directors spent longer reviewing complex cases that required policy adherence determinations.

MHK needed to automate this research while maintaining healthcare’s audit trail and compliance requirements, and to do so across all product modules without building separate AI infrastructure for each one.

Business challenge

As MHK evaluated how to bring AI capabilities across their entire product suite, three core challenges emerged.

  • Fragmented AI infrastructure. Each AI-powered feature would require its own deployment pipeline, HIPAA compliance certification, security controls, and monitoring. Every new AI roadmap item meant a new cluster, a new compliance engagement, and a new operational burden. For a company serving multiple health plans across multiple modules, this approach could not scale.
  • Lengthy development cycles. Deploying a new AI workflow through traditional engineering took 3 to 6 months, not including requirements gathering. The engineering team could not keep pace with the product roadmap.
  • Manual effort at premium cost. Case managers, nurses, pharmacists, and medical directors spent hours per case manually searching through patient records. The cost was especially acute for medical directors (physicians) and pharmacists, whose hourly rates make even small-time savings translate into significant ROI.

Solution overview: SmartProminence AI Orchestrator

MHK built the SmartProminence AI Orchestrator, a multi-tenant, agentic workflow framework running entirely on AWS. Rather than building separate AI systems for each use case, MHK created a single orchestrator where various AI workflows can be deployed through configuration. Define your prompts, specify your input/output schemas, and register the agent. The solution handles everything else including HIPAA compliance, encryption, audit trails, scaling, and orchestration.

The solution is architected around a controller-agent pattern in which a Workflow Engine Controller resolves workflow dependencies and dispatches individual steps to LLM processing agents. The entire system is stateless, event-driven, and independently scalable.

The following diagram illustrates the high-level architecture of the solution.

Architecture of the MHK SmartProminence AI Orchestrator on AWS, showing the orchestration core, workflow controllers, and processing agents communicating through Amazon SQS queues

Figure 1: MHK SmartProminence AI Orchestrator architecture on AWS

At the core of the architecture, the Agent Orchestration Core serves as the central nervous system. Built on Spring Boot and running on AWS Fargate, it exposes a REST API that handles job submission, workflow management, LLM proxying, and token management. Critically, it is the only component that directly accesses the database: controllers and agents interact exclusively through the orchestration core’s API, enforcing strict data access boundaries.

Architecture overview

This section examines the key architectural patterns the orchestrator uses to process diverse healthcare workflows at scale.

Controller-agent pattern with Amazon Bedrock

MHK selected Amazon Bedrock for its multi-model access through a single API, letting them choose the best model per workflow step without separate integrations. As a managed AWS service, Bedrock inherits existing AWS Identity and Access Management (IAM), Amazon Virtual Private Cloud (Amazon VPC), and encryption controls, avoiding a new trust boundary. Built-in content filtering and invocation logging satisfy healthcare auditability requirements, and its model-agnostic architecture lets MHK adopt newer models without rearchitecting the solution.

The orchestrator enforces a strict separation between workflow orchestration and LLM processing. The Workflow Engine Controller determines what needs to happen and in what order, while LLM processing agents execute individual steps. This separation lets agent processing scale independently from workflow logic, and it makes the workflow the single source of truth while agents operate only on specific, actionable steps.

When a job arrives, the controller loads the version-pinned workflow definition, resolves step dependencies into a DAG using Kahn’s algorithm, and pre-creates step executions in a WAITING state. It uses conditional Spring Expression Language (SpEL) expressions to decide which steps to run versus skip, then dispatches agents layer by layer. Steps at the same depth run in parallel, and the controller polls for completion before advancing to the next depth.

Each agent runs a standardized pipeline: input binding (resolving expressions to gather prior step results), optional vision processing for scanned documents, prompt assembly with enriched context, LLM invocation to Amazon Bedrock (Claude), and post-processing for field extraction, type coercion, and structured output.

For parallel workloads within a single step, agents use Java virtual threads for each execution. This lets them process multiple items concurrently, such as extracting data from each page of a multi-page document simultaneously.

Event-driven orchestration with Amazon SQS

Communication between the orchestration core, controllers, and agents flows through Amazon Simple Queue Service (Amazon SQS) queues. To trigger a workflow, the orchestration core places a ControllerTaskMessage on the Controller Invoke Queue. To dispatch an individual step, it places an AgentTaskMessage on the Agent Invoke Queue. Each message is secured with a capability token scoped to only that operation’s data.

This design delivers four properties. Stateless processing means available instances can pick up pending messages. Independent scaling lets agents scale horizontally through ECS Fargate. Fault isolation keeps a failed task from blocking parallel steps, and dead letter queues capture failures. Decoupled deployment ships new agent versions without system-wide restarts.

The only blocking call in the pipeline is the LLM invocation to Amazon Bedrock. Everything else is asynchronous and event-driven, so the system can process hundreds of concurrent jobs without resource contention.

DAG-based parallel execution

The workflow engine uses depth-based parallel execution to maximize throughput. Consider a medical policy review workflow: at Depth 0, agents simultaneously extract patient demographics and pull claims history. At Depth 1, once both are complete, a policy lookup agent identifies the relevant criteria. At Depth 2, an evidence-gathering agent searches through the patient’s clinical history for documentation that satisfies each policy criterion. The controller only advances to the next depth when steps at the current depth have completed.

Conditional expressions can dynamically skip steps based on upstream results. For example, if the initial classification step determines that a case does not involve prescription drugs, the pharmacy verification step at the next depth is automatically skipped, saving both time and token costs. This conditional logic is evaluated by the controller using Spring Expression Language (SpEL) against the structured outputs of completed steps.

Dynamic agent registry

When MHK needs a new agent type, whether for a new medical management module or a new kind of analysis, the process is configuration-driven rather than engineering-driven.

A developer defines the agent configuration (prompt templates, input/output schemas, and model selection), then registers it through the orchestration core’s API. Terraform automatically provisions the supporting infrastructure: SQS queues, IAM roles, and ECS task definitions. The agent immediately becomes available for workflow step assignments, with no new compliance certification needed, since it runs within the already certified orchestrator.

This transformed MHK’s development velocity. The engineering team focuses on prompt design and workflow logic rather than infrastructure scaffolding.

Conversational memory and case association

The orchestrator maintains context across multiple workflow executions for the same patient case. Each execution returns a job ID that the upstream system associates with the case record. Over a case’s lifetime there may be three or more executions (initial intake, policy review, and appeal processing), each producing structured outputs that stay available for later executions.

Once a 30-page clinical record has been analyzed, its structured output is available for future queries on that case without re-running ingestion. When a medical director reviews an appeal weeks later, the patient’s history is already organized and searchable.

Prior context remains in Amazon Simple Storage Service (Amazon S3), encrypted with the client’s dedicated AWS Key Management Service (AWS KMS) key.

Responsible AI controls

MHK enforces safe LLM outputs through application-layer validation built into each agent’s processing pipeline. Every agent post-processes model responses against expected output schemas, cross-references extracted data with source documents to detect hallucinations, and rejects responses that fail confidence thresholds. Domain-specific checks verify that outputs reference only the patient’s own clinical records and match policy-specific medical criteria. LLM inputs and outputs are logged with full audit trails, which supports compliance review and reproducibility for every AI-assisted decision.

AWS services used

The following table summarizes the AWS services that compose the SmartProminence AI Orchestrator and the role each plays in the architecture.

Service Role in architecture
Amazon Bedrock Foundation model inference with IAM role-based authentication
Amazon ECS (Fargate) Containerized orchestration core, workflow controllers, and processing agents
Amazon SQS Event-driven inter-component communication with dead letter queues for fault tolerance
Amazon RDS (MySQL 8.4) Workflow definitions, execution state tracking, multi-AZ for high availability
Amazon S3 Job artifacts, document storage, immutable workflow configurations (KMS encrypted)
AWS KMS

Capability token signing and validation

Per-client encryption keys for multi-tenant data isolation

Amazon Cognito OAuth2/JWT authentication for API access and user identity
Amazon CloudWatch Logging, metrics, token usage tracking, and alerting (no PHI)
Elastic Load Balancing TLS 1.3-terminated application load balancer
Amazon VPC Network isolation with private subnets, VPC endpoints for service access

Security and compliance

Healthcare data demands the highest security standards, and MHK’s architecture implements defense-in-depth across every layer. The orchestrator processes protected health information (PHI) for multiple health plan clients simultaneously, making multi-tenant data isolation a foundational feature.

  • Per-client encryption. Every client has their own AWS KMS key. Documents stored in Amazon S3 are double-encrypted: S3 server-side encryption plus client-specific KMS encryption. Even if a job were somehow misrouted (which the token system helps prevent), the receiving agent could not decrypt another client’s data because it would not have access to that client’s KMS key. The database layer adds row-level encryption on top of Amazon Relational Database Service (Amazon RDS) storage-level encryption, providing defense-in-depth for data at rest.
  • Capability token model. A least-privilege token system limits what each component can access. A controller-scoped token can read workflow definitions and job data, dispatch agent tasks, and create step executions. An agent-scoped token can only read its step’s input, write its own result, call the LLM through the proxy, and upload artifacts. Tokens are generated per-dispatch through KMS, so even a compromised agent cannot reach data from other steps, workflows, or clients.
  • Network isolation. The database subnets have no internet access. AWS service communication (Amazon S3, Amazon SQS, AWS KMS, AWS Secrets Manager, Amazon CloudWatch, Amazon Elastic Container Registry (Amazon ECR)) flows through VPC endpoints, meaning no data ever traverses the public internet. Connections use TLS 1.3 for encryption in transit.
  • Compliance controls. LLM request and response bodies are not logged. Only token counts and content hashes are recorded. Workflow configurations are stored immutably in S3 for complete version history. Agents receive only the minimum context needed for their step, following the principle of data minimization.

Results and impact

The SmartProminence AI Orchestrator delivered measurable business outcomes across both MHK’s internal operations and their health plan clients.

90% reduction in manual review effort. For document intake workflows, processing time dropped from 5–10 minutes per document (manual) to under 1 minute (automated with human-in-the-loop verification). For complex medical director reviews, the system pre-gathers the relevant evidence and presents a structured summary, reducing the physician’s task from hours of document searching to a 30-second approval or denial decision.

85% faster AI feature deployment. New AI capabilities that previously required a full 3–6 month engineering release cycle now deploy in approximately 2 weeks. The engineering team defines workflow configuration and prompt logic without building custom infrastructure, compliance pipelines, or security controls for each feature.

Unified compliance posture. Instead of attesting each AI feature independently, MHK maintains a single orchestrator-level HIPAA and SOC 2 attestation that covers the agents. New agents inherit the orchestrator’s security controls automatically: per-client encryption, audit logging, token-based access, and data minimization.

Multi-tenant extensibility. Health plan clients can run AI workflows through the framework without building their own HIPAA-eligible infrastructure. Because the orchestrator is configuration-driven, MHK can onboard new use cases for existing clients or deploy entirely new health plan customers with minimal engineering effort.

Conclusion

The orchestrator’s controller-agent architecture provides a blueprint for organizations that need to scale AI capabilities across multiple use cases without multiplying their compliance burden. The key insight is that compliance infrastructure should be an orchestrator-level concern, not a per-feature concern, and that agentic orchestration patterns can be both powerful and auditable when designed with healthcare-grade security from the ground up.

Looking ahead, MHK is extending the orchestrator with conversational interfaces so case managers and medical directors can interactively query case data, using the same workflow memory and security infrastructure. The dynamic agent registry continues to grow as new medical management modules adopt AI-powered decision support. MHK is also exploring AWS Marketplace as a distribution channel to bring their HIPAA-eligible agentic framework to organizations beyond healthcare that require similar compliance thresholds.

Share your experience building HIPAA-eligible AI workflows in the comments or reach out if you’re exploring agentic architectures for regulated industries.

To learn more, get started with Amazon Bedrock and explore the Amazon Bedrock code samples to build your own agentic AI solutions on AWS.

 


About the authors

Amazon Aurora PostgreSQL now supports direct querying of Apache Iceberg and Parquet data in your data lake

Post Syndicated from Esra Kayabali original https://aws.amazon.com/blogs/aws/amazon-aurora-postgresql-now-supports-direct-querying-of-apache-iceberg-and-parquet-data-in-your-data-lake/

Today, we’re announcing a new capability for Amazon Aurora PostgreSQL that you can use to directly query operational data together with data stored in your data lake in Apache Iceberg and Apache Parquet formats, using your existing PostgreSQL applications and tools. By eliminating the need to extract, transform, and load (ETL) structured data from data lakes into your operational database, you can reduce operational complexity and simplify application development. You can also use Aurora PostgreSQL to query data from data lakes managed in Iceberg REST Catalog (IRC)-compatible catalogs, giving you access to data across a breadth of analytics systems without moving or duplicating it. Whether you’re powering real-time dashboards, enriching transactions with historical context, or building AI agents that reason over both live and archived data, you can now do it all through a single, familiar interface.

Previously, if your application needed to combine recent transactional data in Aurora with historical records stored in Amazon S3, a common approach was to build reverse ETL pipelines that duplicated data, increased infrastructure costs, and required ongoing engineering effort to keep everything synchronized. This challenge only grows as you increasingly embed AI agents into your applications, where it is impractical to predict and pre-replicate every dataset an agent might need.

DuckLabs, the team that maintains the DuckDB project, recently joined Amazon, and this capability is an example of how the efficiency of DuckDB is being integrated into our services. DuckDB is now embedded directly within Aurora PostgreSQL, so you can query live operational data (including uncommitted writes) alongside your data lake in a single query. Query processing stays within Aurora, with no additional network hops and no ETL pipelines that duplicate data. You can query Apache Iceberg tables managed through the AWS Glue Data Catalog, as well as Parquet and Iceberg data stored in Amazon S3 and S3 Tables. You do all of this using familiar PostgreSQL syntax and your existing applications and tools.

We’re excited to bring the speed and simplicity of DuckDB directly into Aurora PostgreSQL, so you and your agents can query and combine operational and Iceberg data using the familiar PostgreSQL applications, tools, and endpoints already in use. By building this capability around DuckDB, future improvements to the open source engine can continue to bring performance and functionality gains to Aurora and other AWS services.

What is new

This capability is supported on two Aurora PostgreSQL major versions: 17 (starting with 17.11) and 18 (starting with 18.6). To use it, you create an Aurora PostgreSQL cluster, attach an IAM role with the AuroraAnalytics feature, and enable the aurora_analytics extension. The IAM role is what gives Aurora access to your data in Amazon S3 and the AWS Glue Data Catalog. You then create foreign tables that point to your Iceberg or Parquet data in the data lake, and query them using familiar PostgreSQL syntax. You can complete this setup through the Amazon RDS console, or with any PostgreSQL client such as psql. The process is well documented in the Aurora PostgreSQL documentation.

You can query data across external IRC-compatible catalogs through AWS Glue Data Catalog federation. You register the external catalog once with Glue, and then create foreign tables for the tables you want to query, the same way you would for any Glue-native table. A single query can then join data stored in Aurora with Iceberg tables registered across multiple catalogs, so applications get a unified view without moving data or replacing your existing catalog investments.

Aurora also applies optimizations such as predicate pushdown and column pruning so that only the relevant data is read. This keeps queries efficient even as the underlying data grows. Frequently accessed data is also cached in your Aurora instance, so subsequent queries against the same data return faster. You can inspect this behavior per query using aurora_analytics_stat_statements(), which reports metrics such as rows scanned, bytes read from Amazon S3, and cache hits.

To see how direct querying works, I connected to my Aurora PostgreSQL database using psql and created the extension:

CREATE EXTENSION aurora_analytics;

For my walkthrough, I set up a simple financial scenario. I have a recent_transactions table in Aurora with the last 7 days of customer transactions, and a Parquet file in Amazon S3 containing 5 years of historical transaction data. To make Aurora aware of the historical data, I created a foreign table pointing at the Parquet file in S3:

CREATE FOREIGN TABLE transaction_history ()
SERVER aurora_analytics_server
OPTIONS (
    location 's3://<my-bucket>/finance/transaction_history.parquet',
    format 'parquet'
);

Notice the empty parentheses in the CREATE FOREIGN TABLE statement. Aurora automatically reads the schema from the Parquet file metadata, so you do not need to define columns manually. For workloads with many tables, you can skip creating them one at a time: a single IMPORT FOREIGN SCHEMA statement bulk-creates foreign tables for every Iceberg or Parquet table in an AWS Glue Data Catalog database, inferring schemas automatically.

With both tables in place, I ran a single query that combines the recent operational data in Aurora with the historical data in S3:

SELECT merchant, category, amount, transaction_date, 'recent' AS source
FROM recent_transactions
WHERE customer_id = 'C-1001'
UNION ALL
SELECT merchant, category, amount, transaction_date, 'historical' AS source
FROM transaction_history
WHERE customer_id = 'C-1001'
  AND transaction_date >= CURRENT_DATE - INTERVAL '5 years'
ORDER BY transaction_date DESC
LIMIT 15;

The result shows both recent and historical transactions in a single result set. The 7 most recent rows come from Aurora, and the rest come directly from the Parquet file in S3. DuckDB handles the analytical scan of the Parquet data under the hood, while Aurora handles the operational data. That single query would have previously required a pipeline to move the historical data into the database first.

If a query pattern needs single-digit-millisecond latency, you can materialize data from the data lake into a native Aurora PostgreSQL table using familiar commands such as CREATE TABLE AS SELECT, INSERT INTO ... SELECT, or MERGE INTO. The materialized table lives in Aurora and is queried like any other PostgreSQL table, giving you a low-latency path for hot data without operating a separate ingestion pipeline. The read queries can run on any Aurora PostgreSQL instance in your cluster, whether the writer or a read replica, so you can offload analytical scans from your operational workload. The materialization commands write data into Aurora, so they run on the writer instance.

Get started today

Direct querying of Apache Iceberg and Parquet data from Amazon Aurora PostgreSQL is available today in all commercial AWS Regions and AWS GovCloud (US) Regions, at no additional charge. You pay only for the incremental Aurora compute the queries consume and Amazon S3 request costs for reading data lake files.

To learn more, visit the Amazon Aurora features page, read the Aurora PostgreSQL documentation, or try it in the Amazon RDS console. We welcome your feedback through AWS re:Post or through your usual AWS Support contacts.

— Esra

Celebrating Our Newest AWS Heroes – September 2026

Post Syndicated from Taylor Jacobsen original https://aws.amazon.com/blogs/aws/celebrating-our-newest-aws-heroes-september-2026/

Today, we’re excited to introduce the newest members of the AWS Heroes program. AWS Heroes are a vibrant, worldwide group of AWS experts who go above and beyond to share knowledge, mentor others, and build thriving communities. These individuals make a real difference in helping developers and organizations succeed with AWS.

This month, we welcome three exceptional community leaders from across the globe, each bringing unique expertise and a deep commitment to empowering builders everywhere.

Avinash Shashikant Dalvi – Bengaluru, India

Serverless Hero Avinash Shashikant Dalvi is a tech architect and co-organizer of AWS User Group Bengaluru who is focused on serverless, containers, and production-ready applications on AWS. He has delivered over 40 community talks, publishes the AWS for Product Builders newsletter, and creates technical content covering Amazon ECS, AWS Fargate, AWS Lambda, and AWS Amplify.

Joanne Skiles – Orlando, USA

Serverless Hero Joanne Skiles is an engineering leader and educator with over 16 years of experience building full-stack systems, including serverless architecture and AI systems on AWS. She organizes the Orlando AWS User Group and teaches cloud and AI concepts through her YouTube channel, conference talks, and her podcasts Chaotic Commits and Her Career Unplugged. Joanne is also a professor in the Computer Science department at Rollins College, where she runs the Transparent Systems lab.

Xiaofei Li – Shanghai, China

Community Hero Xiaofei Li is an AWS Golden Jacket holder and is an active community leader in the Greater China Region, leading the Kiro, Amazon Quick, and Tokyo Chinese AWS communities. He founded the Kiro Chinese User Community (5,000+ members) and initiated the Chinese localization of AWS Builder Cards across 15 cities and 3,000+ participants. Xiaofei also mentors underserved students and supports Women in Tech initiatives.

Learn More

Visit the AWS Heroes webpage if you’d like to learn more about the AWS Heroes program, or to connect with a Hero near you. To learn more about how to get involved with the AWS community, visit our AWS Builder Center.

— Taylor

Accelerating AS/400 business rule extraction with Kiro: Step-by-step guide

Post Syndicated from Daniel Gray original https://aws.amazon.com/blogs/devops/accelerating-as-400-business-rule-extraction-with-kiro-step-by-step-guide/

AS/400 business rule extraction no longer requires months of manual effort. With Kiro, an agentic AI-powered development environment (spanning IDE, CLI, web, and mobile surfaces, along with the Kiro Crew workspace), you can compress the process into days. This step-by-step guide walks through the approach. Organizations face a common challenge: critical business logic embedded in extensive RPG and COBOL code bases, often maintained by a declining number of developers and subject matter experts (SMEs) with RPG expertise. The fulfillment rules and shipping logic are scattered across interconnected programs that no single person fully understands.

In this post, we walk you through a step-by-step approach for using Kiro to extract business rules from AS/400 RPG and COBOL programs, generate technical specifications, and produce modernization-ready documentation.

Extraction process challenges

Before this engagement, one of our customers faced several challenges with their existing business rules extraction process. They were planning to modernize their AS/400 order fulfillment workflow, which handled inventory validation, shipping document generation, and warehouse operations.

  • Significant consulting costs for specialized AS/400 consultants.
  • Time-intensive manual analysis, typically 4–6 weeks of dedicated effort.
  • Documentation that becomes outdated before the team finishes writing it.
  • Risk of overlooking critical business logic during modernization.

The following is the sample system flow considered to walk through the step-by-step guide.

PROG001 (Interactive Validation)
   │  Validates orders, checks inventory, resolves periods
   ▼
PROG002 (Batch Control)
   │  Manages batch processing of validated orders
   ▼
PROG003 (File Management)
   │  Handles file splitting for large shipment batches
   ▼
PROG004 (Content Generation)
   │  Generates shipping manifests, allocates stock by warehouse priority
   ▼
PROG005 (Encoding and Transmission)
Converts EBCDIC to UTF-8, transmits to external Carrier Gateway

Each program has embedded business rules, including order validation and stock allocation with warehouse priority. These programs also handle shipping weight calculations, character encoding conversion, and integration with external carrier systems. Traditional analysis would have taken 4–6 weeks per system. The effort required across consultants, technical writers, and reviewers would have been 40–80 person-hours per system.

Solution

With Kiro, an agentic AI-powered development environment, you can extract comprehensive business rules, generate technical specifications, and create modernization-ready documentation in hours, not months (as detailed in the Outcomes section).

Working autonomously across your code base, Kiro analyzes dependencies, traces execution paths, and produces detailed documentation.

The approach relies on two core Kiro capabilities:

  • Steering files: Persistent instructions that guide the AI’s behavior, including project context, naming conventions, analysis standards. Configure them once and they apply to all subsequent sessions. Steering files can reference documentation templates that define the exact output format. Each subsequent analysis follows the same repeatable structure.
  • Specs: A structured way to define requirements, design, and implementation tasks. Kiro executes tasks autonomously with progress tracking. Spec tasks tell Kiro which templates to use and where to save the output.

The workflow has three phases:

Phase 1 – Configure steering files to define project context, directory structure, and technical standards. Examples include “extract 10–20 lines of code context around business rules” and “map abbreviated DDS field names to business terms.” Create documentation templates that the steering files reference. These templates specify the exact output format for business rules with code snippets, pseudocode equivalents, DDS field mappings, and integration specifications.

---
inclusion: always
---
# AS/400 Business Rule Extraction Project
## Goal
Analyze a legacy AS/400 order fulfillment system and extract all business
rules to produce modernization-ready documentation. Discover the program
workflow, data architecture, and business logic by reading the source code.
## Source File Locations
- sourcefiles/rpg/ — RPG IV programs (.RPGLE)
- sourcefiles/cl/ — CL programs (.CLLE)
- sourcefiles/dds/ — DDS definitions: physical files (.PF), logical files (.LF), display files (.DSPF)
- sourcefiles/data/ — DB2 table exports (.csv), one per physical file
## What to Discover
- What each program does and how they relate to each other (trace CALL statements and SBMJOB commands)
- Which files each program accesses and how (read the F-specs at the top of each RPG program)
- Business rules embedded in RPG subroutines (validation, processing, calculation logic)
- How configuration tables drive runtime behavior (trace CHAIN lookups and conditional branching)
- External system integration points (identify calls to programs outside this codebase)
- Data flow between programs (trace parameters passed via CALL/PARM and shared files)
- The meaning of cryptic DDS field names (map them to business terms using TEXT keywords and program context)
## What to Produce
- Business rules with original RPG code snippets (10-20+ lines of context)
- Pseudocode equivalents for every business rule
- DDS field-to-business-term mappings for all physical files
- File dependencies matrix (which programs access which files and how)
- Inter-program parameter passing documentation
- Configuration-to-behavior mapping (trace config table values to subroutine invocations)
- Integration specifications for any external system calls
Use the template at templates/Technical_Implementation_Spec.md for output format.
Save all generated documentation to the output/ directory.
## Constraints
- Analysis only — never create executable programs or modify source files
- Read-only operations on all source files
- Every business rule must trace back to specific program, subroutine, and line numbers
- Discover the system's behavior from the source code — do not assume what the programs do

Figure 1: Steering files provide persistent instructions that guide the analysis behavior of Kiro across sessions, configured once and applied to subsequent analyses

Phase 2 – Build a Kiro Spec with discrete, actionable tasks: analyze source files, parse DDS definitions, extract business rules, generate pseudocode, create consolidated documentation using the templates, and verify business rules against source code.

# Implementation Plan: AS/400 Business Rule Extraction
## Overview
This implementation plan extracts business rules and technical specifications from a legacy AS/400 order fulfillment system. The analysis workflow reads RPG programs, CL programs, and DDS file definitions to discover business logic, data architecture, program workflows, and integration points. All findings will be documented using the provided template and saved to the output directory.
## Tasks
- [ ] 1. Analyze DDS physical and logical file definitions
- Read all .PF files in sourcefiles/dds/ and extract field definitions (name, type, length, decimals, TEXT, COLHDG, VALUES)
- Read all .LF files and document key structures and access paths
- Read any .DSPF files and document screen layouts and field mappings
- Map every cryptic field name to a business term using TEXT keywords, column headers, or literal values
- Document key structures and file relationships (which LF belongs to which PF)
- Save intermediate analysis to output/
- [ ] 2. Analyze each program and extract business rules
- Read all RPG programs (.RPGLE) in sourcefiles/rpg/
- Read all CL programs (.CLLE) in sourcefiles/cl/
- For each program, extract F-spec file declarations with access modes
- Identify all subroutines and document their boundaries (line numbers)
- Extract business rules from subroutines, mainline code, and CL logic
- Include 10-20+ lines of original source code context for each rule
- Generate pseudocode equivalents using common programming constructs
- Categorize each rule (validation, processing, calculation, error handling, integration)
- [ ] 3. Map the program workflow and data flow
- Trace all CALL statements and SBMJOB/QCMDEXC invocations across programs
- Document parameters passed at each inter-program call point
- Build the complete program-to-program workflow chain
- Document how data flows between programs via shared files and parameters
- Map logical file usage back to underlying physical files
- [ ] 4. Analyze configuration-driven behavior
- Identify patterns where programs CHAIN to a table and branch based on values read
- Read the CSV data exports in sourcefiles/data/ to see current configuration values
- Trace each configuration value to the code path it triggers
- Flag any inactive or dead configuration entries
- Produce a configuration-to-behavior mapping
- [ ] 5. Document integration specifications
- Identify all calls to programs outside this codebase
- Document parameters, data formats, and protocols for each external interface
- Document any character encoding conversions (CCSID values and transformations)
- Document file paths, naming conventions, and transmission mechanisms
- [ ] 6. Generate consolidated Technical Implementation Specification
- Load the template from templates/Technical_Implementation_Spec.md
- Populate all template sections with the analysis from tasks 1-5
- Include business rules with original code snippets and pseudocode
- Include DDS field mappings, file dependencies, configuration mappings, and integration specs
- Ensure every claim traces to specific program, subroutine, and line numbers
- Save to output/
- [ ] 7. Validate documentation completeness and accuracy
- Verify all programs have been analyzed
- Verify all DDS physical files have field-to-business-term mappings
- Verify all business rules have both source code snippets and pseudocode
- Verify all inter-program calls are documented with parameters
- Verify the consolidated document follows the template structure
- Cross-check source code references for accuracy (correct line numbers)
## Notes
- This is a read-only analysis workflow — no source files will be modified
- Every business rule must trace to specific program, subroutine, and line numbers
- DDS field mappings use TEXT keywords, COLHDG, and VALUES to determine business terms
- Configuration-driven behavior is identified by CHAIN + conditional branching patterns
- All generated documentation will be saved to the output/ directory
- The template at templates/Technical_Implementation_Spec.md defines the output format
## Task Dependency Graph
```json
{
  "waves": [
    { "id": 0, "tasks": ["1"] },
    { "id": 1, "tasks": ["2", "3"] },
    { "id": 2, "tasks": ["4", "5"] },
    { "id": 3, "tasks": ["6"] },
    { "id": 4, "tasks": ["7"] }
  ]
}

```

Figure 2: The Kiro Spec, showing discrete tasks that Kiro executes autonomously with progress tracking

Phase 3 – Execute the Spec and let Kiro work autonomously. Monitor progress as tasks complete, then review the generated documentation.

Important: AI-extracted rules should be reviewed by an AS/400 SME. Automated extraction might occasionally misinterpret complex or ambiguous business logic, so human validation remains essential before acting on extracted rules.

Here’s an example of what Kiro produces. Given this RPG subroutine that validates orders against the master file, Kiro generates a plain-language business rule and its pseudocode equivalent:

Rule 1.3.16: Stock Allocation

Category: Processing Subroutine: ALLCST (lines 3820-3960) Description: Allocates stock from warehouse inventory. Looks up inventory by item key, verifies sufficient available quantity, then decrements available quantity and increments reserved quantity by the order amount. Updates the inventory record.

Source Code (lines 3820-3960):

   3820      C     ALLCST        BEGSR
   3830      C     ITEMKY        CHAIN     INVSTCK1                           42
   3840      C     *IN42         IFEQ      '0'
   3850      C     QTYAV         IFGE      ORDQTY
   3860      C     QTYAV         SUB       ORDQTY        QTYAV
   3870      C     QTYRS         ADD       ORDQTY        QTYRS
   3880      C                   UPDATE    INVFMT
   3890      C                   Z-ADD     0             ALLERR            1 0
   3900      C                   ELSE
   3910      C                   Z-ADD     1             ALLERR
   3920      C                   END
   3930      C                   ELSE
   3940      C                   Z-ADD     2             ALLERR
   3950      C                   END
   3960      C                   ENDSR

Pseudocode:

function allocateStock():
    inventory = findByKey(InventoryStock, itemKey)
    if inventory found:
        if inventory.quantityAvailable >= orderQuantity:
            inventory.quantityAvailable -= orderQuantity
            inventory.quantityReserved += orderQuantity
            update inventoryStock
            allocationError = 0 // OK
        else:
            allocationError = 1 // Insufficient stock
    else:
        allocationError = 2 // Item not found

Figure 3: Kiro extracts business rules with original RPG code, pseudocode equivalents, and plain-English descriptions

The following is the DDS field mapping that translates abbreviated AS/400 field names into business terms:

1.1 ORDERMST — Order Master

Field Type Length Dec TEXT (Business Term) COLHDG VALUES Used By
ZIORCD A 8 — Order Code Order / Code — PROG001, PROG002, PROG003, PROG004
ZIPERD P 6 0 Fulfillment Period Fulfill / Period — PROG001
CURPER P 6 0 Current Period Current / Period — PROG001
STATUS A 1 — Order Status Order / Status ‘A’ ‘H’ ‘C’ ‘X’ ’ ’ PROG001, PROG002
CUSTNAME A 40 — Customer Name Customer / Name — PROG001, PROG002, PROG003, PROG004
WHSCD A 4 — Warehouse Code Warehouse / Code — PROG001, PROG002, PROG003, PROG004
ORDDTE P 8 0 Order Date Order / Date — PROG001
ORDQTY P 7 0 Order Quantity Order / Quantity — PROG001
SHPTYP A 2 — Shipment Type Shipment / Type — PROG001
PRIORT A 1 — Priority Code Priority ‘1’ ‘2’ ‘3’ PROG001

Record Format: ORDERMST — TEXT(‘Order Master Record’)

Key: ZIORCD (unique)

STATUS Values: A = Active, H = Hold, C = Complete, X = Canceled, ’ ’ = New/Blank

PRIORT Values: 1 = High (requires MGR session), 2 = Medium, 3 = Low

1.2 INVSTOCK — Inventory Stock Levels

Field Type Length Dec TEXT (Business Term) COLHDG VALUES Used By
ITEMCD A 10 — Item Code Item / Code — PROG001, PROG004
WHSCD A 4 — Warehouse Code Warehouse / Code — PROG001, PROG004
QTYOH P 9 0 Quantity On Hand Qty / On Hand — PROG001
QTYAV P 9 0 Quantity Available Qty / Available — PROG001
QTYRS P 9 0 Quantity Reserved Qty / Reserved — PROG001
UNITWT P 7 2 Unit Weight KG Unit / Weight — PROG001, PROG004
UNITLN P 5 2 Unit Length CM Unit / Length — PROG001, PROG004

Figure 4: DDS field mapping translates abbreviated AS/400 field names into business terms

This mapping is essential for modernization. Without it, developers building the replacement system are guessing at what Z1ORDCD means.

Deployment

The following steps walk you through setting up and running the extraction workflow.

Prerequisites

Before you begin, make sure that you have the following in place:

  • Kiro installed on your workstation (download from https://kiro.dev/).
  • Access to the AS/400 source code you plan to analyze (RPG/RPGLE, CL/CLLE, and DDS definitions), exported as text files.
  • Optionally, DB2 configuration tables exported to CSV for configuration-driven behavior analysis.
  • Familiarity with your organization’s business domain, plus access to an AS/400 SME to validate the extracted rules.
  • A local project directory where Kiro can read the source files and write generated documentation.

The complete setup is available in the companion GitHub repository listed in the Resources section. This includes steering files, templates, sample AS/400 source code, and Spec definitions.

Kiro project structure showing the sourcefiles, templates, output, and .kiro steering and specs folders

Figure 5: Project structure in Kiro, showing source files, steering configuration, templates, and output directory

The setup has five steps:

Step 1: Project setup

Create the directories that you will be working from for source files, data, output, and other artifacts:

mkdir my-as400-analysis
cd my-as400-analysis
mkdir -p sourcefiles/rpg sourcefiles/cl sourcefiles/dds sourcefiles/data
mkdir -p templates output .kiro/steering .kiro/specs

Step 2: Configure steering files

Create steering files to define your analysis standards. For example, .kiro/steering/product.md:

# Project Context
This project analyzes AS/400 RPG and COBOL programs to extract business rules.
# Analysis Standards
- Extract 10--20+ lines of code context around each business rule
- Map abbreviated DDS field names to business terms
- Document inter-program dependencies and parameter passing
- Identify configuration-driven behavior patterns

Step 3: Add your source files

Copy your AS/400 source code into the sourcefiles/ subdirectories: RPGLE files in rpg/, CLLE files in cl/, and DDS definitions in dds/. Optionally, export DB2 tables to CSV in sourcefiles/data/ for configuration table analysis if you have programs with conditional logic that use those tables to hold runtime configuration options.

Step 4: Create a Kiro Spec

In Kiro, use the command palette: Create New Spec. Define tasks like:

Example Spec definition:

Spec Name: AS/400 Business Rule Extraction
Task 1: Analyze DDS physical and logical file definitions in sourcefiles/dds/
Task 2: For each RPG program in sourcefiles/rpg/, extract business rules with 10--20 lines of surrounding code context
Task 3: Map DDS field names to business terms using templates/field-mapping-template.md
Task 4: Document inter-program data flow and parameter passing
Task 5: Generate consolidated Technical Implementation Specification using templates/tis-template.md
Task 6: Validate that all extracted rules reference valid source line numbers

Step 5: Execute

  • Open the Spec in Kiro, choose Start, and monitor progress as tasks complete autonomously. Review the generated documentation in the output/ directory.
  • For detailed instructions, templates, and example outputs, see the GitHub repository.

What the workflow looks like

Figure 6: Kiro executing the Spec, with real-time progress as each task completes

When you execute the Spec, Kiro processes tasks in sequence with real-time progress tracking. Here is what happens during execution:

  1. Opening the Spec with all tasks listed.
  2. Kiro autonomously reading RPG source files and DDS definitions.
  3. Business rules being extracted with code snippets and pseudocode.
  4. DDS field names being mapped to business terms.
  5. The final consolidated documentation in the output directory.

Outcomes

This section summarizes the measured results from the customer engagement described earlier in this post (a five-program AS/400 order fulfillment system with approximately 40,000 lines of RPG/COBOL). Traditional estimates sourced from the customer’s prior modernization planning documents. Results vary by code base complexity.

Time and effort savings

Using Kiro reduced both elapsed time and total person-hours by an order of magnitude compared to the customer’s traditional manual approach. The following table compares the two approaches:

Metric Traditional Approach Kiro-Assisted Savings
Total effort 40-80 person-hours 12 person-hours 70-85% reduction (measured against the customer’s planning estimates)
Timeline 4-6 weeks 3 days ~90% reduction (measured against the customer’s planning estimates)

Breakdown of Kiro-assisted effort

The 12-hour total breaks down as follows, showing that most of the time is spent on human review rather than setup or execution:

  • Setup (steering + templates + spec): 2 hours.
  • Kiro autonomous execution: 30 minutes.
  • Review and validation: 9.5 hours (reflective of iterative refinement of steering, template, spec, and execution).
  • Total: approximately 12 hours per system of 5 programs with approximately 40,000 lines of code (measured during the customer engagement described earlier in this post).

What Kiro produced

Kiro autonomously generated a complete documentation package for the five-program system, including:

  • Business rules catalog with original RPG code snippets and pseudocode equivalents.
  • DDS field-to-business-term mappings across seven physical files.
  • File dependencies matrix showing which programs access which files.
  • Inter-program parameter passing documentation.
  • Configuration-to-behavior mapping (tracing DB2 config table values to RPG subroutine invocations).
  • Integration specifications for the external carrier gateway (CCSID conversion, transmission parameters).
  • Over 50 pages of structured, template-aligned documentation (measured output from this engagement).

Multiplier effect

The setup cost (templates, steering, Specs) is one-time and is not repeated for additional systems. The following projections extrapolate the per-system effort (approximately 10 hours) from the single-system measured results and add the one-time setup only once:

Scale Traditional Kiro-Assisted Savings
1 system 40-80 hrs / 4-6 weeks 12 hrs / 3 days 28-68 hrs
10 systems 400-800 hrs / 40-60 weeks 102 hrs / 30 days 298-698 hrs

Key quality improvements

  • Consistent, template-driven output across every system analyzed.
  • Exact line number references back to source code for every business rule.
  • Cross-referencing between DDS definitions and RPG program usage alleviates guesswork.
  • Reusable templates and Specs can often be reused for similar systems with minimal reconfiguration.

Conclusion

Legacy AS/400 business rule extraction doesn’t need to take months. With the steering files and Specs in Kiro, you can extract business logic from RPG code bases and produce developer-ready documentation in days.

You still need AS/400 knowledge, business context, and architectural judgment to validate, prioritize, and plan the modernization. But you don’t need to spend months manually reading code and writing specifications. With Kiro handling extraction, you can focus on strategy and decision-making.

To get started, download Kiro, clone the companion repository, and try it on a legacy system this week. For more on AS/400 modernization patterns, refer to the AWS Mainframe Modernization documentation.

If you have questions or want to share your experience, leave a comment on this post. If you’re an AWS customer working on AS/400 or mainframe modernization, reach out through your AWS account team.


About the authors

Daniel Gray

Daniel Gray

Daniel is a Senior Solutions Architect at AWS in the Worldwide Public Sector GovTech organization, where he partners with independent software vendors (ISVs) serving state and local government and public safety markets. He helps these ISVs architect, migrate, and modernize their platforms on AWS — spanning cloud migrations, AI/GenAI adoption, security, and resilience. He is also a member of the Mainframe Modernization Technical Field Community (TFC), contributing expertise on AS400 (IBM i, iSeries) topics

Jasmine Rasheed Syed

Jasmine Rasheed Syed

Jasmine is a Sr. Customer Solutions Manager at AWS, focused on accelerating time to value for customers on their cloud and AI journey by adopting best practices, mechanisms, and AI-powered solutions to transform their business at scale. He partners with customers to identify high-impact AI/ML use cases and helps them move from experimentation to production faster. Jasmine is a seasoned, results-oriented leader with 22+ years of experience in Insurance, Retail & CPG, and Media & Entertainment. He brings a unique ability to bridge the gap between cutting-edge AI capabilities and real-world business outcomes, enabling organizations to harness the full potential of generative AI, machine learning, and data-driven decision-making.

Oscar Hernandez

Oscar Hernandez

Oscar is a Senior Account Executive at AWS, focused on driving AI workload adoption and cloud strategy for global enterprises. He works with executive leaders to identify high-impact AI opportunities and build long-term technology roadmaps that deliver sustained business value. With over 15 years of experience in cloud and enterprise technology, Oscar specializes in helping customers navigate rapid technological change and accelerate production AI deployments at scale.

[$] The year in Plasma and what’s ahead

Post Syndicated from jzb original https://lwn.net/Articles/1096518/

A lot has happened in the KDE
Plasma desktop environment
in the last year. Marco Martin, a KDE contributor
who spends most of his time working on Plasma, took the stage at Akademy 2026 in Graz, Austria to give
an update on Plasma’s major new features, some of the minor-but-interesting
ones, and a preview of what’s coming soon. The biggest upcoming change, dropping
X11 support from Plasma, has been well-advertised; but there are also plans
afoot to further improve remote-desktop support and more.

Critical Cisco Catalyst SD-WAN Manager API authentication bypass exploited in the wild (CVE-2026-76504)

Post Syndicated from Rapid7 original https://www.rapid7.com/blog/post/etr-critical-cisco-catalyst-sd-wan-manager-api-authentication-bypass-exploited-in-the-wild-cve-2026-76504

Overview

On September 30, 2026, Cisco published a security advisory for CVE-2026-76504, a critical API authentication bypass vulnerability affecting Cisco Catalyst SD-WAN Manager. The vulnerability has a CVSSv3.1 score of 9.8 and results from improper handling of URL encoding (CWE-177). An unauthenticated, remote attacker can send a crafted HTTP request that bypasses an authentication rule for a specific API endpoint, gaining access to the API with the privileges of the admin user.

According to Cisco, CVE-2026-76504 is being actively exploited in the wild; Cisco PSIRT became aware of the activity in September 2026. Cisco Catalyst SD-WAN Manager systems with ports exposed to the internet are at risk of compromise. The vulnerability affects the product regardless of system configuration, and Cisco has not provided a workaround, however vendor supplied updates are available. Rapid7 strongly recommends that organizations upgrade affected systems to a fixed release on an emergency basis, outside of normal patch cycles, and investigate internet-facing systems for signs of exploitation.

Cisco Catalyst SD-WAN Manager was also affected by two critical, unauthenticated peering authentication flaws earlier in 2026: CVE-2026-20127 and Rapid7-discovered CVE-2026-20182. Both were distinct issues in the vdaemon service and similar parts of its networking stack. CVE-2026-76504 targets a separate API authentication path, but the recurrence of authentication bypasses in internet-facing Catalyst SD-WAN control components reinforces the need for emergency remediation.

Mitigation guidance

Cisco has released software updates that remediate CVE-2026-76504. Organizations running affected instances of Cisco Catalyst SD-WAN Manager should upgrade to an appropriate fixed release listed below without waiting for a regular patch cycle:

Cisco Catalyst SD-WAN Software release

First fixed release

Earlier than 20.9

Migrate to a fixed release

20.9

20.9.10.1

20.12

20.12.8.2

20.15

20.15.6.1

20.18

20.18.4.1

26.1

26.1.2.1

26.2

26.2.1

Cisco has addressed the vulnerability in the cloud-based Cisco SD-WAN Cloud (Cisco Managed) release 20.15.605, and indicates that no customer action is required for that service.

There are no workarounds. As a temporary mitigation, Cisco recommends that on-premises customers prevent access to the system from unsecured networks. If internet access is required, restrict access to known, trusted hosts and protect Cisco Catalyst SD-WAN control components behind a filtering device. Cisco indicates that this mitigation is already deployed in Cisco Catalyst SD-WAN Cloud Hosted environments. Organizations should apply updates even when the mitigation is in place.

Because active exploitation has occurred, Rapid7 strongly recommends that organizations audit affected systems for compromise. For help assessing a potentially compromised system, Cisco customers may open a Severity 3 TAC case with CVE-2026-76504 in the title and provide an admin-tech file generated with the request admin-tech command.

For the latest mitigation guidance and release compatibility information, please refer to the vendor’s security advisory.

Rapid7 customers

Exposure Command, Vulnerability Management, and Nexpose

Exposure Command, Vulnerability Management, and Nexpose customers can assess exposure to CVE-2026-76504 with vulnerability checks expected to be available in the October 1 content release.

Indicators of compromise

Cisco recommends reviewing the following logs for requests related to j_security_check from unknown or unauthorized IP addresses:

  • /var/log/nms/containers/service-proxy/serviceproxy-access.log: Requests with an encoded character in the j_security_check path, such as POST /%6a_security_check HTTP/1.1.

  • /var/log/nms/vmanage-server.log: Requests to j_security_check associated with usernames beginning with viptela-reserved-.

The %6a value, which URI-encodes the character j, is only an example. According to Cisco, an attacker can exploit the vulnerability by encoding any single character in the request. The vendor cautions that these log entries can also occur during standard operations and should be evaluated against normal network posture to avoid false positives.

Updates

  • September 30, 2026: Initial publication.

Higher education is under siege, and fragmented security is making it harder to respond

Post Syndicated from Rapid7 original https://www.rapid7.com/blog/post/it-higher-education-under-siege-fragmented-security

Higher education faces a difficult security equation. Universities hold large volumes of sensitive student, financial, health, and research data while supporting open networks, distributed users, legacy infrastructure, and increasingly complex cloud environments. Attackers have taken notice, and the pressure on security teams continues to grow.

In Q2 2025, universities faced an average of 4,388 cyberattacks per organization per week, up 24% from the same period in 2024. Nine in ten universities reported experiencing a breach or security incident during the previous 12 months, while the average cost of a data breach in education reached $10.22 million. Confirmed attacks against higher education institutions exposed more than 3.9 million records in 2025, with ransomware continuing to disrupt teaching, research, financial aid, and administrative operations.

Those figures are concerning on their own, but they only explain part of the problem. For university systems with multiple campuses, the way security is organized can create an additional layer of risk.

Why is higher education so difficult to secure?

Universities operate differently from most commercial organizations. Open access, collaboration, and academic freedom are central to their mission, which means security teams must protect environments where students, faculty, researchers, guests, and third parties connect from almost anywhere.

That openness sits alongside an unusually broad mix of sensitive data. A single university may hold student PII, financial aid and tax records, health information, proprietary research, government-funded projects, and intellectual property. Many institutions also rely on legacy systems that have been connected over time to modern cloud applications, APIs, learning platforms, and research networks, creating visibility gaps that can be difficult to manage. 

Resource pressure adds to the challenge. The draft cites 94% of higher education IT leaders as saying they lack enough personnel to defend their environments adequately, leaving relatively small teams responsible for sprawling networks with large numbers of users, devices, applications, and third-party services. 

Why multi-campus fragmentation increases cyber risk

For multi-campus university systems, many of these pressures are compounded by decentralized security operations. Individual campuses often maintain their own infrastructure, security tools, teams, incident response processes, vendor relationships, and renewal cycles. 

The result can be limited visibility across the wider institution. If ransomware is detected at one campus, teams elsewhere may have no immediate view of the same attacker activity. If a zero-day is exploited in one research environment, another campus may remain exposed because the intelligence and response process stay local. 

Fragmentation also affects efficiency. When each campus independently buys, deploys, and manages its own security stack, the wider university system can carry duplicated costs, additional management overhead, and inconsistent coverage. Fragmentation can also slow the spread of threat intelligence across a university system. If one campus detects a new attack pattern, an unusual intrusion technique, or previously unseen malware, that insight may remain local rather than reaching security teams elsewhere in time to act. A suspicious login sequence identified at Campus B, for example, could be the early signal of activity already moving toward Campus A or Campus C, but without shared visibility each team may investigate the same threat independently and at different speeds.

The same problem can appear during vulnerability response. If one campus confirms active exploitation of a newly disclosed vulnerability in a research environment, another campus may still be exposed because patching decisions, asset inventories, and remediation workflows are managed separately. What should become a system-wide priority can remain a local incident until someone connects the dots.

Attackers do not necessarily respect those organizational boundaries. A smaller or less-resourced campus can provide an entry point into relationships, systems, and data connected to the wider institution, while defenders may still be working with a campus-by-campus view.

What should university systems change?

Higher education security needs to preserve the autonomy individual campuses require while improving visibility and coordination across the broader institution.

That means giving security teams a shared view of exposure, threats, and active incidents across campuses, along with the ability to coordinate detection and response when activity in one part of the university may affect another. It also creates an opportunity to reduce duplicated tooling and processes, share threat intelligence more effectively, and make better use of limited security resources.

The objective is a model where a local security team can continue managing the needs of its own campus without losing access to the wider context of what is happening across the university system.

As the threat landscape becomes more connected, higher education security architecture needs to become more connected with it.

In Part 2 of this series, we’ll look at another pressure making that shift more urgent: the growing compliance burden across FERPA, GLBA, HIPAA, and CMMC, and why fragmented security can make regulatory readiness harder to manage across a university system. 

Rapid7 helps more than 11,000 organizations worldwide take command of their security. Learn more at rapid7.com/sled.

[$] Comparing Chromium development at Google and Igalia

Post Syndicated from jake original https://lwn.net/Articles/1094721/

Sharon Yang is a Chromium developer who worked at Google on the browser
and now works on it at Igalia. On the final day of FOSSY 2026, she gave a
presentation on her experiences with both of those companies, comparing and
contrasting the ways the each operates and how that affects work on the
code base. She enjoyed working at Google and feels the same about Igalia,
so the talk was not aimed at complaints—instead it was meant to give a feel
for two companies that are rather different.

The Linux Foundation Technical Advisory Board 2026 election approaches

Post Syndicated from corbet original https://lwn.net/Articles/1097758/

The election for members of the Linux Foundation Technical Advisory Board
will be held electronically after the close of the upcoming Linux Plumbers Conference. The call for
candidates
is is open, with a nomination deadline of October 7.
There are five seats to fill this time, including the one vacated by the
unfortunate passing of Dan Williams.

Serving on the TAB is a good way to help the kernel-development community.
Please see this article from last year for
an overview of what the TAB does and why membership is rewarding, then
consider putting in your nomination.

The collective thoughts of the interwebz