Tag Archives: Customer Solutions

Accelerating airline retailing innovation: how Datalex modernized with AWS Experience-Based Acceleration and agentic AI

Post Syndicated from Kanniah Vagathupatti Jaikumar original https://aws.amazon.com/blogs/architecture/accelerating-airline-retailing-innovation-how-datalex-modernized-with-aws-experience-based-acceleration-and-agentic-ai/

Datalex, a leader in airline ecommerce solutions, set out to answer a question facing every established product-based business: how do you build for where your industry is going, not only where it is today? For more than two decades, Datalex has powered digital retailing for many of the world’s leading airlines, capability built up over years and encoded in a substantial, mission-critical system that runs shopping, pricing, and booking at scale. That depth is a considerable asset, and it is also what makes evolution demanding. The airline industry is moving decisively toward Modern Airline Retailing, an offers-and-orders model with richer integrations and AI-native experiences, and Datalex set out to build the system for that future while carrying forward the proven retail logic its customers rely on every day. The system’s foundations had served that mission reliably for years. The goal now was to modernize the runtime and delivery model so the team could ship the next generation of retailing capability faster, without disrupting the airline operations running on it today. Datalex framed a considered roadmap, Project Phoenix, to get there, and the open question was how much of that journey could be accelerated.

The company’s CTO, Brian Lewis, sponsored the modernization effort and brought together teams across engineering, product, and operations. As Brian Lewis put it, “This is Datalex’s most important project.” The system’s richness was precisely what made the task substantial: years of sophisticated, tightly integrated retail logic that airlines depend on around the clock, built on a mature Java and EJB2 architecture. The team’s central question was never whether the system had value to carry forward, it clearly did, but how to evolve a system of this depth incrementally, at speed and without disruption to airline customers’ operations.

In December 2025, Datalex partnered with AWS for a three-day Experience-Based Acceleration (EBA) workshop. The pace surprised even the system’s own engineers, a measure of how much sophisticated logic they knew sat beneath the surface. As Eric Pitkeathly, Tech Lead, put it: “I did not believe going into the EBA that a migration from EJB/Java 8 to Spring/Java 21 was possible in 3 days! But it was.” This breakthrough did not happen by chance. Following a Modernization Assessment (MODA), the AWS team identified that most of the technological challenges could be accelerated through comprehensive support and the strategic use of agentic AI tools like Kiro and AWS Transform Custom.

This post shares how Datalex used the AWS Experience-Based Acceleration (EBA) methodology to prove modernization feasibility, establish repeatable migration patterns, and integrate generative AI capabilities, all while maintaining their commitment to serving airline customers without disruption.

Solution overview

The AWS Experience-Based Acceleration workshop brought together 16 Datalex engineers with 6 AWS specialists for an intensive three-day engagement at the AWS Dublin offices. Rather than attempting to modernize the entire system at once, the teams focused on proving feasibility through four parallel workstreams, each tackling an important aspect of the modernization journey.

  1. Stream 1: Microservice prototyping extracted a slice of the Reservation component from the existing n-tier architecture and refactored it into a modern Spring Boot microservice running on Java 21. This workstream proved that migration was technically feasible and established reusable patterns for the remaining code base. The Modernization Assessment revealed a decisive insight: by building a compatible runtime environment, the team could run the system’s existing code on modern technologies with minimal modifications. This was clear evidence that Datalex’s foundations were fundamentally sound and could serve as the stepping stone to the modern system rather than something to be rebuilt from scratch. The remaining code changes were then automated through AI-powered coding assistants like Amazon Q Developer and Kiro, accelerating the transformation.
  2. Stream 2: DevSecOps pipeline built an end-to-end continuous integration and continuous delivery (CI/CD) pipeline using AWS services including Amazon Elastic Container Service (Amazon ECS), Amazon Elastic Container Registry (Amazon ECR), and AWS Security Hub. The pipeline embedded security scanning at every stage, from code commit through container deployment. This implemented a shift-left security approach that validates infrastructure as code before deployment.
  3. Stream 3: QA and observability added proactive, real-time monitoring across the system. The team implemented comprehensive observability using Amazon CloudWatch Container Insights, CloudWatch Logs, and Datadog for application performance monitoring. Using the Strangler Fig pattern, the AWS team advised implementing a gateway that could route requests to either the existing REST API or the modernized API through a simple parameter change. This architectural approach supported rapid non-regression testing and real-time validation of the modernization strategy, all within the three-day timeframe. The QA team also created custom dashboards that provide real-time visibility into system health and performance metrics.
  4. Stream 4: Agentic AI proof of concept demonstrated how generative AI could enhance the system. Using Amazon Bedrock AgentCore, the team built an AI-powered natural language interface for booking retrieval integrated with the existing REST API. The multi-agent orchestration system included specialized agents for authentication, data retrieval, and reporting, all secured through Amazon Cognito and integrated with Kong API Gateway.

The target architecture uses Amazon ECS for container orchestration, with Kong API Gateway providing protocol translation between REST and SOAP while supporting dynamic routing between existing and modernized services. With this approach, Datalex can modernize incrementally without disrupting existing airline operations.

Architecture overview

Modernizing an airline retail system requires integrating new capabilities while maintaining existing operations. Datalex’s architecture shows how to layer generative AI agents, modern microservices, and enhanced observability onto an established, proven system without disrupting customer-facing services. The architecture uses a business service proxy to route traffic between existing and modernized components while maintaining backward compatibility.

Datalex structured their modernization around the following components:

  • Demo application: Angular-based frontend demonstrating the modernized user experience.
  • Agent orchestrator: Amazon Bedrock coordinates multiple specialized agents for different workflows.
  • Business service proxy: Routes requests between existing n-tier architecture and new microservices.
  • Modernized services: Spring Boot microservices (SOAP connector and core services) replacing EJB components.
  • Current n-tier architecture: Existing services continue operating while being incrementally replaced.
  • Event-driven messaging: Apache Kafka enables asynchronous communication between components.
  • Observability stack: Amazon CloudWatch, Amazon Managed Grafana, and AWS X-Ray provide monitoring across the layers.
  • Security infrastructure: AWS Secrets Manager and AWS Identity and Access Management (IAM) handle authentication and authorization.

Datalex’s AWS cloud environment serves as the foundation, with the business service proxy acting as the request traffic controller. The proxy routes incoming requests to either the existing n-tier system or the new Spring Boot microservices based on migration status. This approach lets Datalex move services incrementally without requiring a big-bang cutover.

The agent orchestrator integrates with Amazon Bedrock AgentCore to manage three specialized agents: authentication, data retrieval, and reporting. These agents handle specific workflows, calling into both existing and modernized services through the API Gateway and business service proxy. An event-driven architecture using Kafka decouples components and enables real-time data processing.

The architecture confirms that Datalex can modernize individual services independently while maintaining system stability. Existing components remain fully operational until their replacements are tested and ready for production traffic.

Datalex modernization architecture showing the business service proxy routing traffic between the existing n-tier system and new Spring Boot microservices, with Amazon Bedrock agent orchestration, Kafka messaging, and a CloudWatch, Grafana, and X-Ray observability stack

Figure 1: Datalex target architecture with the business service proxy routing between existing and modernized services

The high-level workflow is summarized as follows:

  1. Route traffic intelligently: Business service proxy directs requests to existing or modernized services based on component status.
  2. Deploy modernized microservices: Spring Boot services run alongside the existing n-tier architecture in AWS.
  3. Integrate agent orchestration: Amazon Bedrock manages specialized agents that call both old and new services.
  4. Enable event-driven patterns: Kafka handles asynchronous messaging between decoupled components.
  5. Implement comprehensive monitoring: CloudWatch, Grafana, and X-Ray track performance across components.
  6. Manage secrets centrally: AWS Secrets Manager handles credentials for both existing and modern services.
  7. Support multiple deployment sources: CI/CD pipelines from GitHub, ECR, and Terraform provision infrastructure.
  8. Maintain backward compatibility: API Gateway preserves existing interfaces while routing to new implementations.
  9. Validate incrementally: Each migrated service is tested before the next migration begins.

Technical implementation

Modernizing the system

The prototyping workstream tackled one of the most daunting aspects of the modernization: extracting business logic from a tightly coupled code base. The team selected the Reservation component as their proof of concept because it represented typical complexity found throughout the code base.

The migration involved several technical transformations:

Runtime modernization: Moving from Java 8 to Java 21 brought immediate benefits. Virtual threading capabilities improved concurrent processing, while optimized garbage collection reduced the memory footprint. The team measured a 35% reduction in memory usage compared to the previous JBOSS deployment.

Framework transition: Replacing EJB2 with Spring Boot streamlined the architecture and improved developer productivity. Spring’s extensive testing support improved feature test coverage, while the framework’s modular design allowed for smaller, locally testable service components.

Containerization: Packaging the microservice as a Docker container supported deployment flexibility. The team configured Amazon ECS Fargate to handle container orchestration, avoiding the operational overhead of managing Amazon Elastic Compute Cloud (Amazon EC2) instances running JBOSS.

The migration pattern established during the workshop provides a repeatable approach for the remaining services. Teams can now identify bounded contexts within the system, extract business logic with dependency analysis, refactor to Spring framework patterns, containerize with security hardening, and deploy through an automated pipeline, all while running in parallel with the existing system during transition.

Building production-grade CI/CD

The DevSecOps workstream transformed deployment from a manual, hours-long process into an automated pipeline that completes in under 10 minutes. The pipeline architecture integrates security at every stage:

Build stage: Code commits trigger automated builds using AWS CodeBuild. The build process includes dependency scanning and static code analysis, catching security vulnerabilities before they reach production.

Container security: Images pushed to Amazon ECR undergo automated security scanning. The pipeline validates that the images meet security standards before deployment, with findings aggregated in AWS Security Hub.

Infrastructure validation: Terraform modules defining infrastructure undergo security scanning to verify compliance with organizational policies. This infrastructure-as-code approach provides consistency across environments while maintaining security guardrails.

Deployment automation: The pipeline supports multiple deployment strategies including blue/green deployments for zero-downtime releases, canary deployments for gradual rollout validation, and rolling updates for incremental changes. Native rollback capabilities in Amazon ECS support rapid recovery without operator intervention, improving mean time to recovery.

Implementing comprehensive observability

The observability workstream extended the system with real-time insight into system behavior. The team implemented a multi-layered monitoring approach:

Infrastructure monitoring: Amazon CloudWatch Container Insights provides visibility into container-level metrics including CPU, memory, network, and disk utilization. CloudWatch Logs aggregates logs from the containers for centralized troubleshooting.

Application performance monitoring: Datadog integration provides distributed tracing across microservices, so teams can track requests as they flow through the system. Custom business metrics dashboards surface key performance indicators relevant to airline retail operations.

Proactive monitoring: CloudWatch Synthetics runs automated tests against critical endpoints, alerting teams to issues before customers experience them. This proactive monitoring reduces mean time to detection and resolution.

The observability system also includes an artificial booking generator that streamlines testing for engineering teams, so they can validate performance without requiring production-like data.

Integrating agentic AI

The AI workstream demonstrated how generative AI capabilities could layer onto the modernized architecture without requiring complete system rewrites. The implementation uses Amazon Bedrock AgentCore to orchestrate multiple specialized agents:

Agent architecture: An orchestrator coordinates three specialized agents: an authentication agent for identity verification, a data retrieval agent for accessing business information, and a reporting agent for generating insights. Each agent communicates with backend services through Amazon Bedrock AgentCore Gateway, which translates between the agent’s natural language interface and the system’s REST APIs.

Security implementation: Amazon Cognito provides user authentication and authorization, so airline customers can own agent configuration while maintaining security boundaries. The Model Context Protocol (MCP) Gateway acts as an intermediary between AI agents and REST APIs to support secure communication.

Use case validation: The team built a conversational interface for booking retrieval, so users can query reservation data using natural language. The agent translates conversational queries into API calls, retrieves data from existing Datalex REST APIs, and presents results in a user-friendly format.

This proof of concept validated the technical feasibility of AI integration and established patterns for future AI-enabled features. The architecture provides a foundation for intelligent automation and conversational interfaces that could differentiate Datalex’s product offerings in the airline retail landscape.

Architecture considerations

The target architecture balances modernization goals with operational realities. Amazon Bedrock AgentCore Gateway serves as an important integration layer, supporting gradual migration by routing traffic between existing and modernized services based on configurable rules. With this strangler fig pattern, Datalex can modernize incrementally while maintaining system continuity.

Multi-AZ deployment across Amazon ECS provides high availability, while auto scaling based on CPU and memory metrics makes sure the system can handle traffic variations without manual intervention. The containerized architecture reduces hosting costs through more efficient resource utilization compared to the previous Amazon EC2-based deployment.

Benefits and results

The three-day Experience-Based Acceleration (EBA) delivered outcomes that exceeded expectations. The workshop achieved a 4.9 out of 5.0 customer satisfaction score, with 98% of participants rating their experience as “extremely satisfied.”

Accelerated feasibility proof: What Datalex estimated would take weeks or months to validate independently was accomplished in three days. As the Tech Refresh Dev Manager noted, the AWS team provided a “force multiplier” effect, bringing specialized expertise across modernization, DevOps, and AI/ML domains.

Established migration patterns: The workshop created reusable patterns for migrating the remaining 4 million lines of code. Teams now have documented approaches for extracting services from the n-tier architecture, building secure CI/CD pipelines, implementing observability, and integrating AI capabilities.

Measurable performance improvements: The modernized architecture delivers tangible benefits including a 35% reduction in memory footprint, deployment time reduced from hours to under 10 minutes, 60% faster startup time with Java 21 optimizations, and automated scaling without manual intervention.

Competitive advantage: The modernization positions Datalex to meet growing customer demand for a modern, extensible system. The proven migration path and AI integration capabilities provide competitive differentiation in the airline retail technology landscape.

Cost optimization: Lower hosting costs result from the smaller memory footprint of Spring services compared to existing JBOSS instances. Reduced operational overhead through automation and removal of manual scaling further decreases the total cost of ownership.

Developer productivity: As CTO Brian Lewis observed, “Things that would have taken weeks have been completed in a day.” The modern tooling and frameworks improve the developer experience, while the microservices architecture supports parallel team development and faster iteration cycles.

Looking ahead

Datalex plans to build on the Experience-Based Acceleration (EBA) outcomes through a phased approach. The immediate focus involves maturing the Spring framework to run in parallel with existing JBOSS infrastructure, so teams can gain operational experience before migrating production services.

The company will identify the first production candidate service for migration using the established patterns. The team will implement comprehensive health checks across microservices to support production readiness. They will also quantify cost savings from containerization to provide concrete data for sales teams and executive decision-making.

The AI agent proof of concept opens new product opportunities. Datalex’s product management team will evaluate whether to offer AI-enabled features as product add-ons for airline customers. The conversational interface could improve customer self-service capabilities and reduce operational overhead through intelligent automation.

A follow-up workshop will maintain momentum and address additional modernization challenges. The ongoing partnership with AWS provides access to expertise and best practices as Datalex continues the transformation journey.

Business outcome

The Datalex board approved a significant investment for their technology modernization work, and their prototype tiger team expanded into a fully working Agile team. The Experience-Based Acceleration (EBA) engagement yielded a significant budget for the re-systeming scope of work, with executive-facing KPIs for each quarter mapped against their internal deliverables.

How the Experience-Based Acceleration approach made the difference

The EBA shortened discovery. Datalex estimated 8–12 weeks to validate whether the Reservation component could be extracted without breaking dependencies. With AWS specialists working alongside their engineers, the team had a working response by the end of day one.

It removed cross-cutting blockers. The DevSecOps pipeline required expertise across container scanning, infrastructure validation, and deployment patterns spanning multiple AWS services. The AWS team brought that knowledge into the room, saving weeks of trial and error.

It created evidence for investment decisions. After three days, the team had working code and measurable results they could present to the board. These were proof points that would otherwise have taken three to four months of part-time effort.

Conclusion

Datalex migrated their core services from EJB/Java 8 to Spring Boot/Java 21 in three days during the EBA workshop. The migration established patterns the team now uses across their system, reducing what would have been months of uncertainty into a repeatable process.

The workshop addressed a specific problem: Datalex needed to know if modernization was practical for their code base. By working through one service end-to-end, the team got their response. They also got working code that handles authentication, implements observability, and can run generative AI features, all without rewriting the entire system.

Three lessons from this engagement apply to other modernization projects. First, prove it works on one service before planning the full migration. Second, your team needs to be in the room when the migration happens, because documentation alone will not capture the decisions that matter. Third, add security and monitoring during the migration, not after.

Datalex can now respond to market changes faster and ship features their airline customers are asking for. Their CTO calls this their most important project, and the three-day EBA gave them the technical proof they needed to commit.

If you are planning a similar modernization, AWS Professional Services offers EBA workshops that can help validate your approach. Contact your account team to discuss how this model might work for your system.

Learn more

To learn more about AWS Experience-Based Acceleration programs, visit the AWS Professional Services page. For information about modernizing Java applications, see the AWS Modernization Hub.


About the authors

How Mirelo AI brought sound design to the IDE with MCP and Kiro powers

Post Syndicated from Florian Breton original https://aws.amazon.com/blogs/devops/how-mirelo-ai-brought-sound-design-to-the-ide-with-mcp-and-kiro-powers/

Mirelo AI set out to fix how sound design works in the integrated development environment (IDE). For most developers, sound design has always meant leaving the IDE: opening a browser, digging through stock libraries, and trimming and syncing clips by hand. Sound is the last creative layer most developers reach, and the one they most often get wrong without specialist help. For teams building games, apps, and interactive products, digital audio workstations (DAWs) live outside the development workflow, so most ship with placeholder audio or nothing at all.

Mirelo AI, a Europe-based generative AI lab, builds models that turn a text prompt or a video clip into production-ready sound effects, synced to the picture when video is provided. Feed a video clip to the model and it returns audio matched to the action on screen. Mirelo built a hosted server on the Model Context Protocol (MCP), the open standard that lets AI assistants discover and call external tools. With Mirelo’s hosted MCP server, developers can reach those models from the tools they already use.

In this post, we describe how that MCP server became a power in Kiro, the agentic development environment from AWS. We also cover how Mirelo’s AWS Enterprise Support account team helped bring it to the Kiro powers marketplace. The result: Developers using Kiro can generate sound design from a natural-language prompt without leaving their editor.

Solution overview

With a Kiro power, you get Mirelo’s hosted MCP server bundled with Agent Skills that tell the AI assistant when and how to use it. When a developer’s prompt mentions sound, audio, or effects, Kiro activates the power, connects to Mirelo’s server, and loads its tools into the conversation. The developer describes the sound they want. Kiro calls Mirelo’s models and returns a finished audio file into the project.

The design has three parts:

  • Mirelo’s audio models, exposed through their existing HTTP API.
  • The hosted MCP server (https://mcp.mirelo.ai/mcp), which wraps that API as callable tools.
  • The Kiro power, which bundles the MCP configuration with an Agent Skill and publishes it to the marketplace for one-action install.

The problem: Sound is the last mile of creative development

Visual assets have mature tooling inside IDEs and design systems. Audio does not. A game developer prototyping a level generates textures, writes shaders, and tests physics in the editor. But the moment they need a matching footstep sound, the flow breaks. They open a browser, search a stock library, download candidates, trim them to length, and manually sync timing.

Content creators working with generative video hit the same wall from the other side. AI models now produce visual content in seconds, but each clip ships silent, so adding sound means switching tools, breaking flow, and spending more time on audio than the video itself took to create.

Mirelo’s thesis: Sound generation should live where developers already work, not in a separate application.

From API to agent tool: The Mirelo MCP server

Mirelo already had an API powering their Studio product. The question was how to make those capabilities reachable inside AI-powered development environments without asking developers to write integration code.

MCP answers that. By wrapping their API as an MCP server, Mirelo exposed their full audio generation pipeline to any compatible AI assistant. The Mirelo MCP is hosted and remote: developers add a single URL, authenticate through their browser, and start generating audio from conversation. There’s no local install and no API key to manage.

The server covers the full sound design loop:

  • Generate from text: describe a sound effect in natural language and receive a finished audio file.
  • Generate from video: pass a video clip and get synced sound effects for every action.
  • Extend: lengthen an audio clip that is too short for the scene.
  • Inpaint: replace a selected region of a clip while leaving the rest untouched.

A preflight tool estimates credits and runtime before generation runs, which matters when an agent works through a batch of files rather than a single effect.

Comparing direct API access with a Kiro power

Without the power, a developer who wants a sound effect works against the API directly. For anything longer than a short clip, that means the asynchronous path: submit the job, get a job ID back, poll for status, then download the result once it completes. A minimal version looks like this:

# 1. Submit the job
# The code samples in this post are provided for demonstration and educational purposes only and
# are not intended for production use without additional security review and testing. In
# particular, store API keys in a secrets manager rather than inline, and add error handling and
# retry limits before deploying.
JOB=$(curl -s https://api.mirelo.ai/v2/text-to-sfx/v1.6/jobs \
  --request POST \
  --header 'Authorization: Bearer sk-<your-api-key>' \
  --header 'Content-Type: application/json' \
  --data '{
    "prompt": "Heavy rain on a metal roof with distant thunder",
    "duration_ms": 45000,
    "output_format": "mp3"
  }' | jq -r '.job_id')

# 2. Poll until the job leaves the "processing" state
while true; do
  STATUS=$(curl -s "https://api.mirelo.ai/v2/text-to-sfx/v1.6/jobs/$JOB" \
    --header 'Authorization: Bearer sk-<your-api-key>' | jq -r '.status')
  [ "$STATUS" = "succeeded" ] && break
  [ "$STATUS" = "failed" ] && echo "generation failed" && exit 1
  sleep 3
done

# 3. Fetch the result URL and download the clip
curl -s "https://api.mirelo.ai/v2/text-to-sfx/v1.6/jobs/$JOB" \
  --header 'Authorization: Bearer sk-<your-api-key>' \
  | jq -r '.result_urls[0]' \
  | xargs curl -o rain.mp3

This works, but the developer owns every step. They store the API key, choose the sync endpoint for short clips and the async endpoint for longer ones, and call preflight to estimate credits. They also run the poll loop with sensible backoff, handle the failure state, download the result, and retry on transient errors. Each surface, whether a web app, a game editor, or a batch script, reimplements the same glue.

With the power, the developer describes the sound and the agent assembles that same request. Kiro authenticates through browser sign-in and reads the tool schema Mirelo published. It fills in prompt, duration_ms, and output_format from the conversation, runs the preflight tool when the job is large, and waits on the async job when generation runs long. The developer writes:

Generate 45 seconds of heavy rain on a metal roof with distant thunder, as an mp3.

The file is added to the project. The underlying API call is the same. What changes is who assembles and operates it.

The following table compares the two paths:

Concern Direct API Kiro power/MCP tool call
Authentication Store and send Bearer sk-… on every call Browser sign-in, managed by Kiro
Cost check Call /preflight yourself Agent calls the preflight tool when the job warrants it
Sync compared to async You choose the endpoint and implement polling Agent selects based on job size
Response handling Parse result_urls and download Agent returns the file into the project
Reuse across tools Reimplement the glue per surface One power, available in any Kiro session

Why package it as a Kiro power

Mirelo’s MCP server already worked in several AI assistants. With a Kiro power, you get discoverability, so you find the integration while browsing the marketplace, and automatic activation, so there is no URL to paste or configuration to write.

A power bundles an MCP server with Agent Skills, the structured instructions that guide an AI assistant through a specific workflow. The AI assistant learns what tools exist and when to reach for them. Just as important is how a power loads. A traditional MCP setup registers every tool definition upfront. Connecting a handful of servers can burn tens of thousands of tokens, a large share of the context window, before your first prompt. Kiro powers load dynamically instead. Installed powers sit dormant until your conversation mentions relevant keywords, at which point Kiro activates only that power’s tools and skills and deactivates them when you move on. Skills load the same way, on-demand, so the AI assistant pulls in a specific workflow’s instructions only when it’s working on that task. The result is near-zero baseline context cost and a Mirelo integration that surfaces its sound-design tools exactly when they’re needed, without crowding out the rest of your work.

The structure of a power is small. The following layout shows the three files that define it:

example-power/
|-- plugin.json          # Manifest: name, keywords, and metadata
|-- mcp.json             # Remote MCP server configuration
+-- skills/
    +-- sound-design/
        +-- SKILL.md     # Guides the agent through audio workflows

The plugin.json manifest declares the keywords that trigger activation. The mcp.json file points to the provider’s hosted server. The skill teaches the agent the difference between generating a one-shot effect and sound-designing an entire sequence.

The path from idea to marketplace

The Kiro powers connection came from Mirelo’s AWS Enterprise Support account team. The team spotted the fit between Mirelo’s MCP server and the powers marketplace. They built a proof-of-concept power to show how the integration would work and connected Mirelo with the submission process. Mirelo then packaged their official hosted server as the published power.

From the first conversation to a live power took about two weeks, most of it marketplace review. The engineering itself fit into a single afternoon. The impact is easiest to see in the developer’s workflow. Finding a single sound effect that matches the video is slow, manual work. That includes searching a stock library, auditioning candidates, trimming, and syncing. With the power, a single prompt returns a usable, synced clip, replacing a lengthy manual workflow with one step.

What developers can do with it

After the Mirelo power is active, sound design becomes part of the conversation. A developer polishing a web app might ask:

Generate a soft, satisfying click for this submit button. Short, no metallic ring.

On a game prototype, the request could be:

Here’s my gameplay clip. Generate footstep and impact sounds that match the character’s movement.

For a video project that needs a longer bed:

Extend this forest ambience to forty-five seconds so it covers the full scene transition.

And to fix a single moment:

The glass-break sound at 0:03 is too harsh. Inpaint that region with something more subtle, like thin crystal.

Each request calls Mirelo’s models and returns audio ready to use, without the developer leaving the editor.

A distribution channel for AI model companies

For Mirelo, the Kiro powers marketplace is a new kind of distribution. API businesses have historically reached developers through documentation sites, SDKs, and marketing. A power puts the capability inside the tool developers already use. Developers reach it by intent rather than by integration work.

The model fits AI services that augment creative workflows. Developers don’t plan to use a sound API the way they plan to use a database. They need sound the moment they realize their project is silent. A power meets them at that point of intent.

Conclusion

In this post, we described how Mirelo AI turned a hosted MCP server into a Kiro power, and how AWS Enterprise Support helped move it into the marketplace. For developers, sound design is now a prompt away inside Kiro. For AI model companies, the same path turns an existing MCP server into a distribution channel that reaches developers at the point of intent.

To get started:


About the authors

Florian Breton

Florian Breton

Florian is a Technical Account Manager (TAM) at AWS Enterprise Support based in EMEA, where he helps generative AI startups run and scale their workloads on AWS. Outside of direct customer work, he builds internal AWS tooling and contributes to open source projects that improve the AWS customer experience.

David Kernert

David Kernert

David is a former engineer at AWS, now working at Mirelo AI to scale the training and serving infrastructure behind Mirelo’s generative audio models.

Tofig Hasanov

Tofig Hasanov

Tofig is former Amazon software engineer, now working at Mirelo AI where he is working on building user facing products that expose Mirelo model capabilities, including the MCP.

How MHK built a HIPAA-eligible agentic AI solution on Amazon Bedrock

Post Syndicated from Deepti Tirumala original https://aws.amazon.com/blogs/architecture/how-mhk-built-a-hipaa-eligible-agentic-ai-solution-on-amazon-bedrock/

Healthcare organizations face an increasingly complex challenge: processing vast volumes of medical documents, including clinical records, claims, prior authorizations, appeals, and pharmacy data, while maintaining strict HIPAA compliance and security standards. Traditional approaches require dedicated engineering teams to build individual AI systems for each use case, each needing its own compliance infrastructure, audit trails, and security controls. Agentic frameworks that can scale across use cases can reduce lengthy development cycles and high operational overhead.

MHK, a Hearst Health company ranked #1 in payer care management solutions in the 2024 Best in KLAS: Software & Services Report, faced this exact challenge. Their medical management solution serves health plans across multiple workflows (medical, pharmacy, grievance, and appeals), each requiring intelligent document processing and decision support. Building separate AI systems for each workflow was unsustainable as demand grew.

To solve this, they developed the SmartProminence AI Orchestrator, a HIPAA-eligible agentic workflow framework built on AWS that reduced manual medical review effort by 90%. New AI features that previously took 3+ months to deploy now ship in 2 weeks.

In this post, we walk through how MHK architected this solution using Amazon Bedrock, Amazon Elastic Container Service (Amazon ECS), and event-driven patterns to create a reusable, multi-tenant orchestrator for healthcare AI.

MHK uses Amazon Bedrock exclusively for foundation model inference. They built their own orchestration, retrieval, and validation layers because healthcare workflows require domain-specific controls: DAG-based multi-step execution, clinical document retrieval tied to case context, and HIPAA-specific validation logic that goes beyond general-purpose guardrails. This approach keeps Bedrock focused on scalable model access while MHK retains full control over workflow behavior and compliance enforcement.

Background

MHK provides healthcare cost management and compliance solutions to health plans across the United States. Their medical management system supports the full lifecycle of care decisions, from the moment a provider submits a request for service through final resolution. This includes prior authorization, claims adjudication, appeals processing, pharmacy benefit verification, and medical director reviews.

Each workflow involves analyzing unstructured medical documents such as clinical notes, lab results, imaging reports, and multi-page faxed records, against structured policy criteria. Before MHK’s SmartProminence AI Orchestrator, case managers spent 5 to 10 minutes manually processing each incoming document, while medical directors spent longer reviewing complex cases that required policy adherence determinations.

MHK needed to automate this research while maintaining healthcare’s audit trail and compliance requirements, and to do so across all product modules without building separate AI infrastructure for each one.

Business challenge

As MHK evaluated how to bring AI capabilities across their entire product suite, three core challenges emerged.

  • Fragmented AI infrastructure. Each AI-powered feature would require its own deployment pipeline, HIPAA compliance certification, security controls, and monitoring. Every new AI roadmap item meant a new cluster, a new compliance engagement, and a new operational burden. For a company serving multiple health plans across multiple modules, this approach could not scale.
  • Lengthy development cycles. Deploying a new AI workflow through traditional engineering took 3 to 6 months, not including requirements gathering. The engineering team could not keep pace with the product roadmap.
  • Manual effort at premium cost. Case managers, nurses, pharmacists, and medical directors spent hours per case manually searching through patient records. The cost was especially acute for medical directors (physicians) and pharmacists, whose hourly rates make even small-time savings translate into significant ROI.

Solution overview: SmartProminence AI Orchestrator

MHK built the SmartProminence AI Orchestrator, a multi-tenant, agentic workflow framework running entirely on AWS. Rather than building separate AI systems for each use case, MHK created a single orchestrator where various AI workflows can be deployed through configuration. Define your prompts, specify your input/output schemas, and register the agent. The solution handles everything else including HIPAA compliance, encryption, audit trails, scaling, and orchestration.

The solution is architected around a controller-agent pattern in which a Workflow Engine Controller resolves workflow dependencies and dispatches individual steps to LLM processing agents. The entire system is stateless, event-driven, and independently scalable.

The following diagram illustrates the high-level architecture of the solution.

Architecture of the MHK SmartProminence AI Orchestrator on AWS, showing the orchestration core, workflow controllers, and processing agents communicating through Amazon SQS queues

Figure 1: MHK SmartProminence AI Orchestrator architecture on AWS

At the core of the architecture, the Agent Orchestration Core serves as the central nervous system. Built on Spring Boot and running on AWS Fargate, it exposes a REST API that handles job submission, workflow management, LLM proxying, and token management. Critically, it is the only component that directly accesses the database: controllers and agents interact exclusively through the orchestration core’s API, enforcing strict data access boundaries.

Architecture overview

This section examines the key architectural patterns the orchestrator uses to process diverse healthcare workflows at scale.

Controller-agent pattern with Amazon Bedrock

MHK selected Amazon Bedrock for its multi-model access through a single API, letting them choose the best model per workflow step without separate integrations. As a managed AWS service, Bedrock inherits existing AWS Identity and Access Management (IAM), Amazon Virtual Private Cloud (Amazon VPC), and encryption controls, avoiding a new trust boundary. Built-in content filtering and invocation logging satisfy healthcare auditability requirements, and its model-agnostic architecture lets MHK adopt newer models without rearchitecting the solution.

The orchestrator enforces a strict separation between workflow orchestration and LLM processing. The Workflow Engine Controller determines what needs to happen and in what order, while LLM processing agents execute individual steps. This separation lets agent processing scale independently from workflow logic, and it makes the workflow the single source of truth while agents operate only on specific, actionable steps.

When a job arrives, the controller loads the version-pinned workflow definition, resolves step dependencies into a DAG using Kahn’s algorithm, and pre-creates step executions in a WAITING state. It uses conditional Spring Expression Language (SpEL) expressions to decide which steps to run versus skip, then dispatches agents layer by layer. Steps at the same depth run in parallel, and the controller polls for completion before advancing to the next depth.

Each agent runs a standardized pipeline: input binding (resolving expressions to gather prior step results), optional vision processing for scanned documents, prompt assembly with enriched context, LLM invocation to Amazon Bedrock (Claude), and post-processing for field extraction, type coercion, and structured output.

For parallel workloads within a single step, agents use Java virtual threads for each execution. This lets them process multiple items concurrently, such as extracting data from each page of a multi-page document simultaneously.

Event-driven orchestration with Amazon SQS

Communication between the orchestration core, controllers, and agents flows through Amazon Simple Queue Service (Amazon SQS) queues. To trigger a workflow, the orchestration core places a ControllerTaskMessage on the Controller Invoke Queue. To dispatch an individual step, it places an AgentTaskMessage on the Agent Invoke Queue. Each message is secured with a capability token scoped to only that operation’s data.

This design delivers four properties. Stateless processing means available instances can pick up pending messages. Independent scaling lets agents scale horizontally through ECS Fargate. Fault isolation keeps a failed task from blocking parallel steps, and dead letter queues capture failures. Decoupled deployment ships new agent versions without system-wide restarts.

The only blocking call in the pipeline is the LLM invocation to Amazon Bedrock. Everything else is asynchronous and event-driven, so the system can process hundreds of concurrent jobs without resource contention.

DAG-based parallel execution

The workflow engine uses depth-based parallel execution to maximize throughput. Consider a medical policy review workflow: at Depth 0, agents simultaneously extract patient demographics and pull claims history. At Depth 1, once both are complete, a policy lookup agent identifies the relevant criteria. At Depth 2, an evidence-gathering agent searches through the patient’s clinical history for documentation that satisfies each policy criterion. The controller only advances to the next depth when steps at the current depth have completed.

Conditional expressions can dynamically skip steps based on upstream results. For example, if the initial classification step determines that a case does not involve prescription drugs, the pharmacy verification step at the next depth is automatically skipped, saving both time and token costs. This conditional logic is evaluated by the controller using Spring Expression Language (SpEL) against the structured outputs of completed steps.

Dynamic agent registry

When MHK needs a new agent type, whether for a new medical management module or a new kind of analysis, the process is configuration-driven rather than engineering-driven.

A developer defines the agent configuration (prompt templates, input/output schemas, and model selection), then registers it through the orchestration core’s API. Terraform automatically provisions the supporting infrastructure: SQS queues, IAM roles, and ECS task definitions. The agent immediately becomes available for workflow step assignments, with no new compliance certification needed, since it runs within the already certified orchestrator.

This transformed MHK’s development velocity. The engineering team focuses on prompt design and workflow logic rather than infrastructure scaffolding.

Conversational memory and case association

The orchestrator maintains context across multiple workflow executions for the same patient case. Each execution returns a job ID that the upstream system associates with the case record. Over a case’s lifetime there may be three or more executions (initial intake, policy review, and appeal processing), each producing structured outputs that stay available for later executions.

Once a 30-page clinical record has been analyzed, its structured output is available for future queries on that case without re-running ingestion. When a medical director reviews an appeal weeks later, the patient’s history is already organized and searchable.

Prior context remains in Amazon Simple Storage Service (Amazon S3), encrypted with the client’s dedicated AWS Key Management Service (AWS KMS) key.

Responsible AI controls

MHK enforces safe LLM outputs through application-layer validation built into each agent’s processing pipeline. Every agent post-processes model responses against expected output schemas, cross-references extracted data with source documents to detect hallucinations, and rejects responses that fail confidence thresholds. Domain-specific checks verify that outputs reference only the patient’s own clinical records and match policy-specific medical criteria. LLM inputs and outputs are logged with full audit trails, which supports compliance review and reproducibility for every AI-assisted decision.

AWS services used

The following table summarizes the AWS services that compose the SmartProminence AI Orchestrator and the role each plays in the architecture.

Service Role in architecture
Amazon Bedrock Foundation model inference with IAM role-based authentication
Amazon ECS (Fargate) Containerized orchestration core, workflow controllers, and processing agents
Amazon SQS Event-driven inter-component communication with dead letter queues for fault tolerance
Amazon RDS (MySQL 8.4) Workflow definitions, execution state tracking, multi-AZ for high availability
Amazon S3 Job artifacts, document storage, immutable workflow configurations (KMS encrypted)
AWS KMS

Capability token signing and validation

Per-client encryption keys for multi-tenant data isolation

Amazon Cognito OAuth2/JWT authentication for API access and user identity
Amazon CloudWatch Logging, metrics, token usage tracking, and alerting (no PHI)
Elastic Load Balancing TLS 1.3-terminated application load balancer
Amazon VPC Network isolation with private subnets, VPC endpoints for service access

Security and compliance

Healthcare data demands the highest security standards, and MHK’s architecture implements defense-in-depth across every layer. The orchestrator processes protected health information (PHI) for multiple health plan clients simultaneously, making multi-tenant data isolation a foundational feature.

  • Per-client encryption. Every client has their own AWS KMS key. Documents stored in Amazon S3 are double-encrypted: S3 server-side encryption plus client-specific KMS encryption. Even if a job were somehow misrouted (which the token system helps prevent), the receiving agent could not decrypt another client’s data because it would not have access to that client’s KMS key. The database layer adds row-level encryption on top of Amazon Relational Database Service (Amazon RDS) storage-level encryption, providing defense-in-depth for data at rest.
  • Capability token model. A least-privilege token system limits what each component can access. A controller-scoped token can read workflow definitions and job data, dispatch agent tasks, and create step executions. An agent-scoped token can only read its step’s input, write its own result, call the LLM through the proxy, and upload artifacts. Tokens are generated per-dispatch through KMS, so even a compromised agent cannot reach data from other steps, workflows, or clients.
  • Network isolation. The database subnets have no internet access. AWS service communication (Amazon S3, Amazon SQS, AWS KMS, AWS Secrets Manager, Amazon CloudWatch, Amazon Elastic Container Registry (Amazon ECR)) flows through VPC endpoints, meaning no data ever traverses the public internet. Connections use TLS 1.3 for encryption in transit.
  • Compliance controls. LLM request and response bodies are not logged. Only token counts and content hashes are recorded. Workflow configurations are stored immutably in S3 for complete version history. Agents receive only the minimum context needed for their step, following the principle of data minimization.

Results and impact

The SmartProminence AI Orchestrator delivered measurable business outcomes across both MHK’s internal operations and their health plan clients.

90% reduction in manual review effort. For document intake workflows, processing time dropped from 5–10 minutes per document (manual) to under 1 minute (automated with human-in-the-loop verification). For complex medical director reviews, the system pre-gathers the relevant evidence and presents a structured summary, reducing the physician’s task from hours of document searching to a 30-second approval or denial decision.

85% faster AI feature deployment. New AI capabilities that previously required a full 3–6 month engineering release cycle now deploy in approximately 2 weeks. The engineering team defines workflow configuration and prompt logic without building custom infrastructure, compliance pipelines, or security controls for each feature.

Unified compliance posture. Instead of attesting each AI feature independently, MHK maintains a single orchestrator-level HIPAA and SOC 2 attestation that covers the agents. New agents inherit the orchestrator’s security controls automatically: per-client encryption, audit logging, token-based access, and data minimization.

Multi-tenant extensibility. Health plan clients can run AI workflows through the framework without building their own HIPAA-eligible infrastructure. Because the orchestrator is configuration-driven, MHK can onboard new use cases for existing clients or deploy entirely new health plan customers with minimal engineering effort.

Conclusion

The orchestrator’s controller-agent architecture provides a blueprint for organizations that need to scale AI capabilities across multiple use cases without multiplying their compliance burden. The key insight is that compliance infrastructure should be an orchestrator-level concern, not a per-feature concern, and that agentic orchestration patterns can be both powerful and auditable when designed with healthcare-grade security from the ground up.

Looking ahead, MHK is extending the orchestrator with conversational interfaces so case managers and medical directors can interactively query case data, using the same workflow memory and security infrastructure. The dynamic agent registry continues to grow as new medical management modules adopt AI-powered decision support. MHK is also exploring AWS Marketplace as a distribution channel to bring their HIPAA-eligible agentic framework to organizations beyond healthcare that require similar compliance thresholds.

Share your experience building HIPAA-eligible AI workflows in the comments or reach out if you’re exploring agentic architectures for regulated industries.

To learn more, get started with Amazon Bedrock and explore the Amazon Bedrock code samples to build your own agentic AI solutions on AWS.

 


About the authors

Accelerating AS/400 business rule extraction with Kiro: Step-by-step guide

Post Syndicated from Daniel Gray original https://aws.amazon.com/blogs/devops/accelerating-as-400-business-rule-extraction-with-kiro-step-by-step-guide/

AS/400 business rule extraction no longer requires months of manual effort. With Kiro, an agentic AI-powered development environment (spanning IDE, CLI, web, and mobile surfaces, along with the Kiro Crew workspace), you can compress the process into days. This step-by-step guide walks through the approach. Organizations face a common challenge: critical business logic embedded in extensive RPG and COBOL code bases, often maintained by a declining number of developers and subject matter experts (SMEs) with RPG expertise. The fulfillment rules and shipping logic are scattered across interconnected programs that no single person fully understands.

In this post, we walk you through a step-by-step approach for using Kiro to extract business rules from AS/400 RPG and COBOL programs, generate technical specifications, and produce modernization-ready documentation.

Extraction process challenges

Before this engagement, one of our customers faced several challenges with their existing business rules extraction process. They were planning to modernize their AS/400 order fulfillment workflow, which handled inventory validation, shipping document generation, and warehouse operations.

  • Significant consulting costs for specialized AS/400 consultants.
  • Time-intensive manual analysis, typically 4–6 weeks of dedicated effort.
  • Documentation that becomes outdated before the team finishes writing it.
  • Risk of overlooking critical business logic during modernization.

The following is the sample system flow considered to walk through the step-by-step guide.

PROG001 (Interactive Validation)
   │  Validates orders, checks inventory, resolves periods
   ▼
PROG002 (Batch Control)
   │  Manages batch processing of validated orders
   ▼
PROG003 (File Management)
   │  Handles file splitting for large shipment batches
   ▼
PROG004 (Content Generation)
   │  Generates shipping manifests, allocates stock by warehouse priority
   ▼
PROG005 (Encoding and Transmission)
Converts EBCDIC to UTF-8, transmits to external Carrier Gateway

Each program has embedded business rules, including order validation and stock allocation with warehouse priority. These programs also handle shipping weight calculations, character encoding conversion, and integration with external carrier systems. Traditional analysis would have taken 4–6 weeks per system. The effort required across consultants, technical writers, and reviewers would have been 40–80 person-hours per system.

Solution

With Kiro, an agentic AI-powered development environment, you can extract comprehensive business rules, generate technical specifications, and create modernization-ready documentation in hours, not months (as detailed in the Outcomes section).

Working autonomously across your code base, Kiro analyzes dependencies, traces execution paths, and produces detailed documentation.

The approach relies on two core Kiro capabilities:

  • Steering files: Persistent instructions that guide the AI’s behavior, including project context, naming conventions, analysis standards. Configure them once and they apply to all subsequent sessions. Steering files can reference documentation templates that define the exact output format. Each subsequent analysis follows the same repeatable structure.
  • Specs: A structured way to define requirements, design, and implementation tasks. Kiro executes tasks autonomously with progress tracking. Spec tasks tell Kiro which templates to use and where to save the output.

The workflow has three phases:

Phase 1 – Configure steering files to define project context, directory structure, and technical standards. Examples include “extract 10–20 lines of code context around business rules” and “map abbreviated DDS field names to business terms.” Create documentation templates that the steering files reference. These templates specify the exact output format for business rules with code snippets, pseudocode equivalents, DDS field mappings, and integration specifications.

---
inclusion: always
---
# AS/400 Business Rule Extraction Project
## Goal
Analyze a legacy AS/400 order fulfillment system and extract all business
rules to produce modernization-ready documentation. Discover the program
workflow, data architecture, and business logic by reading the source code.
## Source File Locations
- sourcefiles/rpg/ — RPG IV programs (.RPGLE)
- sourcefiles/cl/ — CL programs (.CLLE)
- sourcefiles/dds/ — DDS definitions: physical files (.PF), logical files (.LF), display files (.DSPF)
- sourcefiles/data/ — DB2 table exports (.csv), one per physical file
## What to Discover
- What each program does and how they relate to each other (trace CALL statements and SBMJOB commands)
- Which files each program accesses and how (read the F-specs at the top of each RPG program)
- Business rules embedded in RPG subroutines (validation, processing, calculation logic)
- How configuration tables drive runtime behavior (trace CHAIN lookups and conditional branching)
- External system integration points (identify calls to programs outside this codebase)
- Data flow between programs (trace parameters passed via CALL/PARM and shared files)
- The meaning of cryptic DDS field names (map them to business terms using TEXT keywords and program context)
## What to Produce
- Business rules with original RPG code snippets (10-20+ lines of context)
- Pseudocode equivalents for every business rule
- DDS field-to-business-term mappings for all physical files
- File dependencies matrix (which programs access which files and how)
- Inter-program parameter passing documentation
- Configuration-to-behavior mapping (trace config table values to subroutine invocations)
- Integration specifications for any external system calls
Use the template at templates/Technical_Implementation_Spec.md for output format.
Save all generated documentation to the output/ directory.
## Constraints
- Analysis only — never create executable programs or modify source files
- Read-only operations on all source files
- Every business rule must trace back to specific program, subroutine, and line numbers
- Discover the system's behavior from the source code — do not assume what the programs do

Figure 1: Steering files provide persistent instructions that guide the analysis behavior of Kiro across sessions, configured once and applied to subsequent analyses

Phase 2 – Build a Kiro Spec with discrete, actionable tasks: analyze source files, parse DDS definitions, extract business rules, generate pseudocode, create consolidated documentation using the templates, and verify business rules against source code.

# Implementation Plan: AS/400 Business Rule Extraction
## Overview
This implementation plan extracts business rules and technical specifications from a legacy AS/400 order fulfillment system. The analysis workflow reads RPG programs, CL programs, and DDS file definitions to discover business logic, data architecture, program workflows, and integration points. All findings will be documented using the provided template and saved to the output directory.
## Tasks
- [ ] 1. Analyze DDS physical and logical file definitions
- Read all .PF files in sourcefiles/dds/ and extract field definitions (name, type, length, decimals, TEXT, COLHDG, VALUES)
- Read all .LF files and document key structures and access paths
- Read any .DSPF files and document screen layouts and field mappings
- Map every cryptic field name to a business term using TEXT keywords, column headers, or literal values
- Document key structures and file relationships (which LF belongs to which PF)
- Save intermediate analysis to output/
- [ ] 2. Analyze each program and extract business rules
- Read all RPG programs (.RPGLE) in sourcefiles/rpg/
- Read all CL programs (.CLLE) in sourcefiles/cl/
- For each program, extract F-spec file declarations with access modes
- Identify all subroutines and document their boundaries (line numbers)
- Extract business rules from subroutines, mainline code, and CL logic
- Include 10-20+ lines of original source code context for each rule
- Generate pseudocode equivalents using common programming constructs
- Categorize each rule (validation, processing, calculation, error handling, integration)
- [ ] 3. Map the program workflow and data flow
- Trace all CALL statements and SBMJOB/QCMDEXC invocations across programs
- Document parameters passed at each inter-program call point
- Build the complete program-to-program workflow chain
- Document how data flows between programs via shared files and parameters
- Map logical file usage back to underlying physical files
- [ ] 4. Analyze configuration-driven behavior
- Identify patterns where programs CHAIN to a table and branch based on values read
- Read the CSV data exports in sourcefiles/data/ to see current configuration values
- Trace each configuration value to the code path it triggers
- Flag any inactive or dead configuration entries
- Produce a configuration-to-behavior mapping
- [ ] 5. Document integration specifications
- Identify all calls to programs outside this codebase
- Document parameters, data formats, and protocols for each external interface
- Document any character encoding conversions (CCSID values and transformations)
- Document file paths, naming conventions, and transmission mechanisms
- [ ] 6. Generate consolidated Technical Implementation Specification
- Load the template from templates/Technical_Implementation_Spec.md
- Populate all template sections with the analysis from tasks 1-5
- Include business rules with original code snippets and pseudocode
- Include DDS field mappings, file dependencies, configuration mappings, and integration specs
- Ensure every claim traces to specific program, subroutine, and line numbers
- Save to output/
- [ ] 7. Validate documentation completeness and accuracy
- Verify all programs have been analyzed
- Verify all DDS physical files have field-to-business-term mappings
- Verify all business rules have both source code snippets and pseudocode
- Verify all inter-program calls are documented with parameters
- Verify the consolidated document follows the template structure
- Cross-check source code references for accuracy (correct line numbers)
## Notes
- This is a read-only analysis workflow — no source files will be modified
- Every business rule must trace to specific program, subroutine, and line numbers
- DDS field mappings use TEXT keywords, COLHDG, and VALUES to determine business terms
- Configuration-driven behavior is identified by CHAIN + conditional branching patterns
- All generated documentation will be saved to the output/ directory
- The template at templates/Technical_Implementation_Spec.md defines the output format
## Task Dependency Graph
```json
{
  "waves": [
    { "id": 0, "tasks": ["1"] },
    { "id": 1, "tasks": ["2", "3"] },
    { "id": 2, "tasks": ["4", "5"] },
    { "id": 3, "tasks": ["6"] },
    { "id": 4, "tasks": ["7"] }
  ]
}

```

Figure 2: The Kiro Spec, showing discrete tasks that Kiro executes autonomously with progress tracking

Phase 3 – Execute the Spec and let Kiro work autonomously. Monitor progress as tasks complete, then review the generated documentation.

Important: AI-extracted rules should be reviewed by an AS/400 SME. Automated extraction might occasionally misinterpret complex or ambiguous business logic, so human validation remains essential before acting on extracted rules.

Here’s an example of what Kiro produces. Given this RPG subroutine that validates orders against the master file, Kiro generates a plain-language business rule and its pseudocode equivalent:

Rule 1.3.16: Stock Allocation

Category: Processing Subroutine: ALLCST (lines 3820-3960) Description: Allocates stock from warehouse inventory. Looks up inventory by item key, verifies sufficient available quantity, then decrements available quantity and increments reserved quantity by the order amount. Updates the inventory record.

Source Code (lines 3820-3960):

   3820      C     ALLCST        BEGSR
   3830      C     ITEMKY        CHAIN     INVSTCK1                           42
   3840      C     *IN42         IFEQ      '0'
   3850      C     QTYAV         IFGE      ORDQTY
   3860      C     QTYAV         SUB       ORDQTY        QTYAV
   3870      C     QTYRS         ADD       ORDQTY        QTYRS
   3880      C                   UPDATE    INVFMT
   3890      C                   Z-ADD     0             ALLERR            1 0
   3900      C                   ELSE
   3910      C                   Z-ADD     1             ALLERR
   3920      C                   END
   3930      C                   ELSE
   3940      C                   Z-ADD     2             ALLERR
   3950      C                   END
   3960      C                   ENDSR

Pseudocode:

function allocateStock():
    inventory = findByKey(InventoryStock, itemKey)
    if inventory found:
        if inventory.quantityAvailable >= orderQuantity:
            inventory.quantityAvailable -= orderQuantity
            inventory.quantityReserved += orderQuantity
            update inventoryStock
            allocationError = 0 // OK
        else:
            allocationError = 1 // Insufficient stock
    else:
        allocationError = 2 // Item not found

Figure 3: Kiro extracts business rules with original RPG code, pseudocode equivalents, and plain-English descriptions

The following is the DDS field mapping that translates abbreviated AS/400 field names into business terms:

1.1 ORDERMST — Order Master

Field Type Length Dec TEXT (Business Term) COLHDG VALUES Used By
ZIORCD A 8 — Order Code Order / Code — PROG001, PROG002, PROG003, PROG004
ZIPERD P 6 0 Fulfillment Period Fulfill / Period — PROG001
CURPER P 6 0 Current Period Current / Period — PROG001
STATUS A 1 — Order Status Order / Status ‘A’ ‘H’ ‘C’ ‘X’ ’ ’ PROG001, PROG002
CUSTNAME A 40 — Customer Name Customer / Name — PROG001, PROG002, PROG003, PROG004
WHSCD A 4 — Warehouse Code Warehouse / Code — PROG001, PROG002, PROG003, PROG004
ORDDTE P 8 0 Order Date Order / Date — PROG001
ORDQTY P 7 0 Order Quantity Order / Quantity — PROG001
SHPTYP A 2 — Shipment Type Shipment / Type — PROG001
PRIORT A 1 — Priority Code Priority ‘1’ ‘2’ ‘3’ PROG001

Record Format: ORDERMST — TEXT(‘Order Master Record’)

Key: ZIORCD (unique)

STATUS Values: A = Active, H = Hold, C = Complete, X = Canceled, ’ ’ = New/Blank

PRIORT Values: 1 = High (requires MGR session), 2 = Medium, 3 = Low

1.2 INVSTOCK — Inventory Stock Levels

Field Type Length Dec TEXT (Business Term) COLHDG VALUES Used By
ITEMCD A 10 — Item Code Item / Code — PROG001, PROG004
WHSCD A 4 — Warehouse Code Warehouse / Code — PROG001, PROG004
QTYOH P 9 0 Quantity On Hand Qty / On Hand — PROG001
QTYAV P 9 0 Quantity Available Qty / Available — PROG001
QTYRS P 9 0 Quantity Reserved Qty / Reserved — PROG001
UNITWT P 7 2 Unit Weight KG Unit / Weight — PROG001, PROG004
UNITLN P 5 2 Unit Length CM Unit / Length — PROG001, PROG004

Figure 4: DDS field mapping translates abbreviated AS/400 field names into business terms

This mapping is essential for modernization. Without it, developers building the replacement system are guessing at what Z1ORDCD means.

Deployment

The following steps walk you through setting up and running the extraction workflow.

Prerequisites

Before you begin, make sure that you have the following in place:

  • Kiro installed on your workstation (download from https://kiro.dev/).
  • Access to the AS/400 source code you plan to analyze (RPG/RPGLE, CL/CLLE, and DDS definitions), exported as text files.
  • Optionally, DB2 configuration tables exported to CSV for configuration-driven behavior analysis.
  • Familiarity with your organization’s business domain, plus access to an AS/400 SME to validate the extracted rules.
  • A local project directory where Kiro can read the source files and write generated documentation.

The complete setup is available in the companion GitHub repository listed in the Resources section. This includes steering files, templates, sample AS/400 source code, and Spec definitions.

Kiro project structure showing the sourcefiles, templates, output, and .kiro steering and specs folders

Figure 5: Project structure in Kiro, showing source files, steering configuration, templates, and output directory

The setup has five steps:

Step 1: Project setup

Create the directories that you will be working from for source files, data, output, and other artifacts:

mkdir my-as400-analysis
cd my-as400-analysis
mkdir -p sourcefiles/rpg sourcefiles/cl sourcefiles/dds sourcefiles/data
mkdir -p templates output .kiro/steering .kiro/specs

Step 2: Configure steering files

Create steering files to define your analysis standards. For example, .kiro/steering/product.md:

# Project Context
This project analyzes AS/400 RPG and COBOL programs to extract business rules.
# Analysis Standards
- Extract 10--20+ lines of code context around each business rule
- Map abbreviated DDS field names to business terms
- Document inter-program dependencies and parameter passing
- Identify configuration-driven behavior patterns

Step 3: Add your source files

Copy your AS/400 source code into the sourcefiles/ subdirectories: RPGLE files in rpg/, CLLE files in cl/, and DDS definitions in dds/. Optionally, export DB2 tables to CSV in sourcefiles/data/ for configuration table analysis if you have programs with conditional logic that use those tables to hold runtime configuration options.

Step 4: Create a Kiro Spec

In Kiro, use the command palette: Create New Spec. Define tasks like:

Example Spec definition:

Spec Name: AS/400 Business Rule Extraction
Task 1: Analyze DDS physical and logical file definitions in sourcefiles/dds/
Task 2: For each RPG program in sourcefiles/rpg/, extract business rules with 10--20 lines of surrounding code context
Task 3: Map DDS field names to business terms using templates/field-mapping-template.md
Task 4: Document inter-program data flow and parameter passing
Task 5: Generate consolidated Technical Implementation Specification using templates/tis-template.md
Task 6: Validate that all extracted rules reference valid source line numbers

Step 5: Execute

  • Open the Spec in Kiro, choose Start, and monitor progress as tasks complete autonomously. Review the generated documentation in the output/ directory.
  • For detailed instructions, templates, and example outputs, see the GitHub repository.

What the workflow looks like

Figure 6: Kiro executing the Spec, with real-time progress as each task completes

When you execute the Spec, Kiro processes tasks in sequence with real-time progress tracking. Here is what happens during execution:

  1. Opening the Spec with all tasks listed.
  2. Kiro autonomously reading RPG source files and DDS definitions.
  3. Business rules being extracted with code snippets and pseudocode.
  4. DDS field names being mapped to business terms.
  5. The final consolidated documentation in the output directory.

Outcomes

This section summarizes the measured results from the customer engagement described earlier in this post (a five-program AS/400 order fulfillment system with approximately 40,000 lines of RPG/COBOL). Traditional estimates sourced from the customer’s prior modernization planning documents. Results vary by code base complexity.

Time and effort savings

Using Kiro reduced both elapsed time and total person-hours by an order of magnitude compared to the customer’s traditional manual approach. The following table compares the two approaches:

Metric Traditional Approach Kiro-Assisted Savings
Total effort 40-80 person-hours 12 person-hours 70-85% reduction (measured against the customer’s planning estimates)
Timeline 4-6 weeks 3 days ~90% reduction (measured against the customer’s planning estimates)

Breakdown of Kiro-assisted effort

The 12-hour total breaks down as follows, showing that most of the time is spent on human review rather than setup or execution:

  • Setup (steering + templates + spec): 2 hours.
  • Kiro autonomous execution: 30 minutes.
  • Review and validation: 9.5 hours (reflective of iterative refinement of steering, template, spec, and execution).
  • Total: approximately 12 hours per system of 5 programs with approximately 40,000 lines of code (measured during the customer engagement described earlier in this post).

What Kiro produced

Kiro autonomously generated a complete documentation package for the five-program system, including:

  • Business rules catalog with original RPG code snippets and pseudocode equivalents.
  • DDS field-to-business-term mappings across seven physical files.
  • File dependencies matrix showing which programs access which files.
  • Inter-program parameter passing documentation.
  • Configuration-to-behavior mapping (tracing DB2 config table values to RPG subroutine invocations).
  • Integration specifications for the external carrier gateway (CCSID conversion, transmission parameters).
  • Over 50 pages of structured, template-aligned documentation (measured output from this engagement).

Multiplier effect

The setup cost (templates, steering, Specs) is one-time and is not repeated for additional systems. The following projections extrapolate the per-system effort (approximately 10 hours) from the single-system measured results and add the one-time setup only once:

Scale Traditional Kiro-Assisted Savings
1 system 40-80 hrs / 4-6 weeks 12 hrs / 3 days 28-68 hrs
10 systems 400-800 hrs / 40-60 weeks 102 hrs / 30 days 298-698 hrs

Key quality improvements

  • Consistent, template-driven output across every system analyzed.
  • Exact line number references back to source code for every business rule.
  • Cross-referencing between DDS definitions and RPG program usage alleviates guesswork.
  • Reusable templates and Specs can often be reused for similar systems with minimal reconfiguration.

Conclusion

Legacy AS/400 business rule extraction doesn’t need to take months. With the steering files and Specs in Kiro, you can extract business logic from RPG code bases and produce developer-ready documentation in days.

You still need AS/400 knowledge, business context, and architectural judgment to validate, prioritize, and plan the modernization. But you don’t need to spend months manually reading code and writing specifications. With Kiro handling extraction, you can focus on strategy and decision-making.

To get started, download Kiro, clone the companion repository, and try it on a legacy system this week. For more on AS/400 modernization patterns, refer to the AWS Mainframe Modernization documentation.

If you have questions or want to share your experience, leave a comment on this post. If you’re an AWS customer working on AS/400 or mainframe modernization, reach out through your AWS account team.


About the authors

Daniel Gray

Daniel Gray

Daniel is a Senior Solutions Architect at AWS in the Worldwide Public Sector GovTech organization, where he partners with independent software vendors (ISVs) serving state and local government and public safety markets. He helps these ISVs architect, migrate, and modernize their platforms on AWS — spanning cloud migrations, AI/GenAI adoption, security, and resilience. He is also a member of the Mainframe Modernization Technical Field Community (TFC), contributing expertise on AS400 (IBM i, iSeries) topics

Jasmine Rasheed Syed

Jasmine Rasheed Syed

Jasmine is a Sr. Customer Solutions Manager at AWS, focused on accelerating time to value for customers on their cloud and AI journey by adopting best practices, mechanisms, and AI-powered solutions to transform their business at scale. He partners with customers to identify high-impact AI/ML use cases and helps them move from experimentation to production faster. Jasmine is a seasoned, results-oriented leader with 22+ years of experience in Insurance, Retail & CPG, and Media & Entertainment. He brings a unique ability to bridge the gap between cutting-edge AI capabilities and real-world business outcomes, enabling organizations to harness the full potential of generative AI, machine learning, and data-driven decision-making.

Oscar Hernandez

Oscar Hernandez

Oscar is a Senior Account Executive at AWS, focused on driving AI workload adoption and cloud strategy for global enterprises. He works with executive leaders to identify high-impact AI opportunities and build long-term technology roadmaps that deliver sustained business value. With over 15 years of experience in cloud and enterprise technology, Oscar specializes in helping customers navigate rapid technological change and accelerate production AI deployments at scale.

How Property Finder automated incident management with AWS DevOps Agent

Post Syndicated from Nada Tlohi original https://aws.amazon.com/blogs/devops/how-property-finder-automated-incident-management-with-aws-devops-agent/

When a production service starts saturating the CPU at 1 AM, every minute counts for incident management. For Property Finder, a production incident could mean failed searches, frustrated users, and direct revenue impact. Property Finder is the leading property portal in the Middle East and North Africa (MENA), serving millions of property seekers across five markets.

Before adopting AWS DevOps Agent, incident response followed a familiar pattern: an alert fires, an on-call engineer wakes up, spends 20–40 minutes correlating metrics across tools, manually documents findings, and opens a fix. Mean Time to Resolution stretched to 2–3 days for non-critical issues.

Today, that entire workflow runs autonomously. From alert to root cause analysis, Slack notification, Jira ticket, on-call phone call with context, and auto-remediation pull request (PR), the full lifecycle completes in 14 minutes. This post walks through the implementation and shows how a separate custom agent that automatically generates code fixes is the key differentiator.

The business problem

Property Finder runs a distributed microservices architecture on Amazon Elastic Container Service (Amazon ECS) fronted by Application Load Balancers (ALBs). When infrastructure issues occur, the impact is immediate: users see failed searches, agents cannot update listings, and revenue is directly impacted during peak hours.

The traditional workflow had three gaps:

  1. Detection lag. Non-critical anomalies could go undetected for days.
  2. Context switching. Engineers bounced between five or more tools per incident.
  3. Knowledge silos. Runbooks lived in people’s heads, not automation.

Solution architecture

Property Finder’s implementation connects AWS DevOps Agent at the center of a three-tier pipeline: Detection and Trigger, Autonomous Investigation, and Event-Driven Output.

Three-tier incident pipeline from a CloudWatch alarm through AWS DevOps Agent investigation to Slack, Jira, and GitHub outputs

Figure 1: End-to-end autonomous incident management architecture

The numbered steps correspond to the data flow in Figure 1:

  1. ECS CPU spike triggers an Amazon CloudWatch Alarm. CloudWatch Metrics Insights monitors service health across all ECS clusters. When sustained CPU exceeds 98%, the alarm transitions to ALARM state.
  2. AWS Lambda formats and HMAC-signs the payload. Triggered directly by the CloudWatch alarm action (which fires only on ALARM state transitions), AWS Lambda enriches the payload with service metadata, signs it with HMAC-SHA256 using credentials from AWS Secrets Manager, and POSTs to the webhook.
  3. The agent begins autonomous investigation. Parallel subagents query ECS metrics, AWS CloudTrail, ALB traffic patterns, and Grafana telemetry (Prometheus, Loki, Pyroscope). The agent reads relevant source code from GitHub for correlation.
  4. Findings post to Slack in real time. The native Slack integration posts investigation progress to #incidents. The full root cause analysis, impact assessment, and mitigation plan appear at the end of the thread.
  5. Investigation Completed event fires to Amazon EventBridge. Amazon EventBridge triggers an orchestrator Lambda that fans out to three independent targets simultaneously.
  6. Lambda creates a Jira ticket with the full root cause analysis. The Lambda retrieves the investigation summary from journal records and creates a prioritized ticket with root cause, severity, and affected service.
  7. Grafana IRM pages the on-call engineer by phone. A Lambda posts a Grafana Alerting-compatible payload to the IRM webhook. The escalation chain calls the engineer with full investigation context: what broke, why, and the recommended fix.
  8. The remediation agent opens a GitHub PR with the auto-fix. It receives the root cause, generates a Terraform or code fix, and opens a Draft PR through a GitHub Model Context Protocol (MCP) server. Engineers review before merging.

A real incident

The example-service, Property Finder’s core property search microservice serving millions of queries per day across five MENA markets, experienced CPU saturation at 99.11%. The pipeline resolved it end-to-end in 14 minutes.

1:21 AM │ Alarm fires (ECS CPU > 98%)

1:22 AM │ Investigation starts + Slack posted

1:22 AM │ 4 parallel subagents launched

1:32 AM │ Root cause identified

1:33 AM │ Jira ticket [redacted] created

1:34 AM │ On-call paged via phone call

1:35 AM │ GitHub PR [redacted] opened with fix

The detection Lambda handles three tasks: (1) retrieves the webhook secret from AWS Secrets Manager, (2) enriches the CloudWatch alarm event with ECS service metadata (cluster name, service name, task count), and (3) HMAC-signs the payload before POSTing to the webhook. The key authentication pattern:

# HMAC-SHA256 signing for webhook authentication
ts = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%S.000Z")
body = json.dumps(incident)
sig = hmac.new(webhook_secret.encode("utf-8"),
               f"{ts}:{body}".encode(), hashlib.sha256).digest()
http.request("POST", webhook_url, body=body,
             headers={"x-amzn-event-timestamp": ts,
                      "x-amzn-event-signature": base64.b64encode(sig).decode()})
Four parallel subagents querying ECS, CloudTrail, ALB, and Grafana data sources during the investigation

Figure 2: Four parallel subagents investigating ECS, CloudTrail, ALB, and Grafana data sources simultaneously

Root cause: Conflicting CPU and memory target-tracking autoscaling policies combined with an insufficient capacity floor. The service had both a CPU policy (target 70%) and a memory policy (target 75%). Actual memory usage sat at 3–8%, creating a persistent conflict between the two policies.

With MinCapacity set too low, the service could not sustain the task count needed to absorb CPU load. The resulting instability (22+ scaling flips observed) prevented stable scale-out, leaving the service effectively pinned at two tasks with no CPU headroom.

This is a common organizational issue: teams configure both scaling dimensions without realizing the interaction, especially when the capacity floor is not sized for baseline traffic. The agent identified the pattern in 10 minutes, a task that typically requires senior engineers with deep scaling expertise and hours of CloudWatch metric correlation.

Investigation output naming conflicting CPU and memory autoscaling policies as the root cause

Figure 3: Root cause analysis identifying the conflicting autoscaling policy

Slack incidents channel message linking to the running investigation at 1:22 AM

Figure 4: Slack notification with investigation link posted at 1:22 AM

Auto-created Jira ticket showing priority, root cause, and affected service

Figure 5: Jira ticket [redacted] auto-created with priority, root cause, and affected service

At 1:34 AM, the on-call engineer received a phone call through Grafana IRM with the complete investigation context. No need to wake up and hunt for root cause across dashboards.

Grafana IRM escalation chain routing the alert to the on-call engineer

Figure 6: Grafana IRM escalation chain routing the alert and calling the on-call engineer

Incoming on-call phone call at 1:34 AM carrying the investigation context

Figure 7: Incoming phone call at 1:34 AM with investigation context

Mitigation plan generated: (1) Remove the memory-based scaling policy, (2) raise MinCapacity to handle baseline traffic, (3) implement CPU-only target tracking at 70%. This plan was passed to a separate custom agent for remediation.

Remediation

Remediation is the key differentiator in this pipeline. It is a dedicated remediation agent (pr-creation-agent) invoked only after investigation completes. AWS DevOps Agent enforces read-only access to infrastructure through a per-session permission guardrail. Effective permissions are the intersection of the execution role’s IAM policy and the guardrail, and write actions are excluded.

The split separates concerns: investigation stays within that read-only envelope, whereas the remediation agent is scoped to a GitHub MCP server as its only external integration. Safety at the remediation layer does not rely on the agent’s built-in directed actions approval mechanism. Instead, two controls enforce the boundary. First, the remediation agent is a separate, narrowly scoped agent with access limited to GitHub MCP. Second, every output is a Draft pull request that requires human review and merge before taking effect. The GitHub MCP connection is authenticated with a fine-grained personal access token scoped to the specific infrastructure repositories, with an expiration and rotation policy. No elevated IAM role or additional agent permissions are required.

How it works

When the “Investigation Completed” Amazon EventBridge event fires, a Lambda orchestrator invokes the remediation agent with the investigation ID. The agent then:

  1. Reads findings from journal records to understand the root cause and recommended fix.
  2. Maps the AWS account to the correct repository. Property Finder has six infrastructure repos for different teams (B2B, B2C, core-platform, growth, data-engineering, shared-infra). The agent extracts the account ID from resource ARNs and routes to the right repo. This mapping is validated through automated tests and updated as new accounts or repositories are onboarded.
  3. Checks for duplicate PRs by searching existing PR titles and bodies for the investigation ID. If a matching PR exists, it reports the URL and exits without creating a duplicate.
  4. Reads the relevant Terraform files through GitHub MCP (GITHUB-MCP_get_file_contents), identifies the exact changes required, and plans the fix.
  5. Creates a feature branch (fix/{investigation_id}), commits the changes, and opens a Draft PR with a structured template including problem summary, root cause, changes made, and a testing checklist.

AWS also supports remediation through Kiro CLI with AWS CodeBuild or Kiro-ready prompts. Property Finder chose an approach that fits their multi-team repository structure: the remediation agent runs entirely within the Agent Space (the managed environment where custom agents execute), uses GitHub MCP for repository access, and maps multiple repositories to different teams automatically.

The orchestrator Lambda is triggered by the “Investigation Completed” Amazon EventBridge event. It first retrieves the investigation findings from journal records, then fans out to three targets simultaneously. Target one creates a Jira ticket with the full root cause analysis, severity, and affected service. Target two posts a Grafana Alerting-compatible payload to the Grafana IRM webhook to trigger phone call escalation. Target three invokes the remediation agent through the CreateChat and SendMessage API, passing the investigation ID and root cause context so it can generate the appropriate code fix.

Draft GitHub pull request with a problem summary, root cause, and changes template

Figure 8: GitHub PR [redacted] generated by the remediation agent with a structured problem, root cause, and changes template

Terraform diff replacing the memory scaling policy with a CPU-only target-tracking policy

Figure 9: Terraform diff showing the new CPU-only scaling policy replacing the conflicting memory configuration

The PR is always opened as Draft. Engineers review, run terraform plan, validate in staging, and merge. The agent never auto-merges.

Results

Metric Before After Improvement
End-to-end time Hours to days 14 minutes >88% reduction
Investigation 20 to 40 min (manual) 10 min (autonomous) 50–75% reduction
Documentation Manual, incomplete Auto-generated root cause analysis + Jira 100% documented
Remediation Manual PR by engineer Auto-fix PR + review Minutes to code fix

Cost considerations: Each incident invokes two agent sessions (investigation + remediation) with up to four parallel subagents. Billing is based on agent minutes. For detailed pricing, see the AWS DevOps Agent pricing page. We recommend reviewing pricing for all services used in this architecture.

“We now rely fully on AWS DevOps Agent to identify infrastructure-related issues. It has helped us identify multiple complex issues without even opening a support ticket. Even if we had raised tickets, it would likely have taken support engineers hours to find the root cause, whereas we resolved these issues in minutes.”

— Yasitha Bogamuwa, Cloud Engineering Manager, Property Finder

Getting started

Prerequisites:

  1. An Agent Space configured in your account.
  2. Amazon CloudWatch and AWS CloudTrail enabled for observability.
  3. Slack, Grafana, and GitHub connected as capabilities.
  4. Infrastructure resources tagged for topology mapping.

Step 1: Configure the webhook trigger. Set up CloudWatch Alarm action to invoke a Lambda function. The Lambda enriches the payload, HMAC-signs it, and POSTs to your Agent Space webhook endpoint.

Step 2: Set up event-driven outputs. Create an Amazon EventBridge rule for “Investigation Completed” events (source: aws.aidevops). Add Lambda targets for Jira, Grafana IRM, and optionally a remediation custom agent.

Step 3: Test end-to-end. Trigger a test alarm and verify the full pipeline: investigation starts, Slack posts, Jira ticket created, on-call paged, and PR opened.

For a similar integration pattern with Salesforce, see Automating Incident Investigation with AWS DevOps Agent and Salesforce MCP Server on the AWS DevOps Blog.

Clean up

This post describes an architecture pattern implemented by Property Finder. If you deployed test resources while following along, remember to delete any CloudWatch Alarms, Lambda functions, Amazon EventBridge rules, and Agent Space configurations to avoid ongoing charges. For a full list of resources and associated costs, review the pricing pages for each AWS service used in this architecture.

Conclusion

Property Finder’s implementation shows that autonomous incident management works in production today, with their pipeline running since early 2026. The agent never auto-merges. Human review remains in the loop by design: the agent accelerates, the engineer decides. The on-call engineer wakes up to a phone call with the root cause already identified, a Jira ticket filed, and a PR ready for review.

Explore the AWS DevOps Agent documentation to get started with your own autonomous pipeline.

  1. Getting Started with AWS DevOps Agent.
  2. Automating Incident Investigation with Salesforce MCP.
  3. Building an End-to-End Agentic SRE.
  4. Amazon EventBridge User Guide.
  5. Grafana IRM Documentation.

About the authors

Nada Tlohi

Nada Tlohi

Nada is a Technical Account Manager at AWS based in Dubai, UAE. She helps strategic enterprise customers across the MENA region transform their cloud operations and improve system reliability by adopting AIOps, incident automation, and DevOps best practices.

Conor Manton

Conor Manton

Conor is a Principal Technical Account Manager at AWS, based in San Francisco. He works with strategic enterprise customers to accelerate their cloud journey, with a focus to operationalize AI-powered workflows to drive business outcomes.

Jaydeep Singh

Jaydeep Singh

Jaydeep is a Senior DevOps Engineer at Property Finder. He specializes in designing and operating scalable cloud infrastructure, containerized platforms, and Kubernetes ecosystems. He leads platform reliability, infrastructure automation, and continuous integration and continuous delivery (CI/CD) initiatives, so engineering teams can build and deploy applications securely, efficiently, and at scale.

How Delivery Hero rebuilt real-time ad measurement with Apache Flink

Post Syndicated from Kirill Tishenkov original https://aws.amazon.com/blogs/big-data/how-delivery-hero-rebuilt-real-time-ad-measurement-with-apache-flink/

This post is co-written with Kirill Tishenkov, Alexandru Pisarenco, Upendra Kambhampati, and Sabariesh Ganesan from Delivery Hero.

Real-time ad measurement is one of the harder streaming problems in advertising. Every impression and click has to be accurate enough to bill a vendor for, and fresh enough for the ad server to act on. In this post, we describe how Delivery Hero moved its ad measurement pipeline from hourly batch processing to real time on Amazon Managed Service for Apache Flink. Delivery Hero, based in Berlin, Germany, is one of the world’s leading local delivery platforms, operating across Asia, Europe, Latin America, the Middle East, and North Africa. Working with more than 1.5 million restaurant partners and local vendors in around 65 countries, Delivery Hero handles millions of orders for food, groceries, and everyday essentials daily.

At the center of Delivery Hero’s business sits an advertising platform that connects vendors and brands with millions of active consumers. The platform handles tens of thousands of messages per second and processes billions of ad events per day, supporting an advertising revenue stream that reached almost EUR 1.5 billion in 2025. Every impression served and every click recorded must satisfy two requirements at once. The data must be accurate enough to bill vendors fairly, and fresh enough for the ad server to act on in real time. Delivery Hero replaced its batch-oriented measurement system with a fully real-time pipeline built on Amazon Managed Service for Apache Flink. The new pipeline cut infrastructure costs by more than half and reached a level of data quality the previous system could not.

Challenges with the legacy system

The legacy ads measurement system consumed impression, click, and order events from message queues. It enriched them through synchronous API calls for campaign metadata and product lookups, then wrote hourly aggregated metrics to a reporting database. This design worked at a modest scale, but five structural problems emerged as traffic grew.

No event-time semantics, and slow processing. The pipeline bucketed events by the time it processed them rather than the time they occurred, because most events arrived without a usable event timestamp. Results were internally consistent, but they skewed whenever ingestion lagged or events arrived out of order. That widened the error bar on every time-sensitive metric, including return on ad spend (ROAS). The bigger cost was speed. Metrics were assembled in hourly batches, so the average gap between when an event occurred and when it was recorded was 61 minutes. The platform was reacting to clicks and impressions up to an hour after the fact, far too late for budget pacing or ad serving.

Synchronous enrichment capped how far the system could scale. Enrichment is the step that attaches business context to a raw ad event: which campaign it belongs to, which vendor owns it, and which product was advertised. In the legacy system, every event triggered a chain of blocking external API calls to fetch that context. During traffic spikes, such as a flash sale or a back-to-school surge, exhausted connection pools cascaded into billing, ad serving, and reporting simultaneously. There was no back-pressure mechanism and no way to scale enrichment independently of event ingestion.

The database behind the pipeline was built for a very different access pattern. The pipeline kept its working data in a NoSQL document database: deduplication keys, attribution history, and running totals. The platform inherited that database from its pre-streaming era, when ad measurement looked like document storage and retrieval. The workload then evolved into continuous deduplication, multi-day attribution lookups, and rolling aggregation. Every event ended up triggering a full document read and write against a database designed for occasional access, not per-event mutation. Read/write amplification stored far more data than the logic needed, every write triggered index updates and collection scans, and storage costs grew in lockstep with query latency. At peak load, this often tipped into production outages.

Reprocessing was a project, not a capability. Recovery from a bug, a traffic spike, or a corrupted upstream batch required different tooling for every consuming system. Billing replay was a hand-rolled combination of Google Cloud BigQuery tables, Pub/Sub topics, and custom CLI scripts. Reporting replay ran as a separate daily Airflow job with a one-hour-per-day cost and a six-month horizon. Campaigns and credits events had no replay path at all. Every recovery was a coordination exercise across teams. Every event type that could not be replayed was a class of problems that could only be patched manually after the fact.

Incomplete event context corrupted downstream data quality. Enrichment was synchronous and best-effort, so the pipeline still wrote through events that failed a lookup or arrived malformed, leaving their fields blank. The pipeline had no mechanism to recover the missing context later. Three gaps mattered most:

  • Missing session rate: the share of events that landed without a usable session ID, leaving the interaction unattached to the user browsing session it belonged to. At 30–40 percent, roughly a third of all events could not be tied back to a session, breaking any session-scoped analysis or feature.
  • Missing customer identifiers (IDs): the share of events with no customer ID, severing the link between an ad interaction and the customer who generated it and weakening attribution and personalization.
  • Missing impression timestamps: the share of impression events lacking a reliable event-time timestamp (the same root cause as the processing-time fallback described earlier). At 91 percent, most impressions had no trustworthy event time, forcing the processing-time approximation and widening the error bar on every time-based metric.

These omissions propagated silently into the reporting metrics and into the session-scoped features consumed by machine learning (ML) models for campaign ranking, conversion-rate estimation, and anomaly detection.

The team set three non-negotiable requirements. First, fault-tolerant data processing, to eliminate data loss. Second, stateful stream processing that could hold multiple days of interaction history in low-cost, low-latency storage. Third, fully managed infrastructure, so engineers could focus on application logic rather than cluster operations.

The team selected Apache Flink because it satisfies all three requirements natively, without bolting on external systems. Its event-time watermark model helps place out-of-order events in the correct time window even when they arrive late. Its RocksDB state backend holds large keyed state on disk without Java Virtual Machine (JVM) heap pressure.

The team chose Amazon Managed Service for Apache Flink over self-hosted Flink on Amazon Elastic Kubernetes Service (Amazon EKS) to eliminate the operational burden of managing JobManagers, TaskManagers, and checkpoint storage. Amazon Kinesis Data Streams serves as the upstream event bus, with two streams: one for user event actions (impressions and clicks) and one for orders. The team chose Kinesis Data Streams over Amazon Managed Streaming for Apache Kafka (Amazon MSK) for cost efficiency at this topology.

Amazon DynamoDB holds campaign and product reference data, queried through Flink’s Async I/O API to enrich events without blocking the processing pipeline. AWS Secrets Manager stores ad event decryption keys, retrieved once at job startup. Amazon Simple Storage Service (Amazon S3) stores granular event logs in Avro format and serves as the incremental checkpoint store for Flink state. Amazon EventBridge Pipes bridged Amazon Simple Queue Service (Amazon SQS) to Kinesis in the minimum viable product (MVP) phase without any custom code, cutting time-to-production by two weeks.

Solution architecture

The following diagram shows the end-to-end pipeline.

Architecture diagram. Two Amazon Simple Notification Service (Amazon SNS) topics receive user event actions and order events. Amazon SQS buffers them, and Amazon EventBridge Pipes or AWS Fargate forwards them into two Amazon Kinesis Data Streams. Amazon Managed Service for Apache Flink then decrypts, deduplicates, enriches from Amazon DynamoDB, attributes, and aggregates the events. It writes granular events and checkpoints to Amazon S3, aggregated metrics to the reporting database, and billing events to Apache Kafka topics consumed by the ad server and budget service.

Figure 1: End-to-end architecture of the real-time ad measurement pipeline

Two Amazon Simple Notification Service (Amazon SNS) topics ingest events: one receives user event actions (compressed, encrypted ad tokens containing campaign, vendor, and placement metadata), the other receives order events. Amazon SQS buffers both before Amazon EventBridge Pipes (MVP) or an AWS Fargate service (production) forwards them into Kinesis.

Amazon Managed Service for Apache Flink runs a five-stage Java pipeline:

  1. Decompress and decrypt. The pipeline decrypts the ad event token using keys from AWS Secrets Manager.
  2. Deduplicate. The pipeline keys events on a composite of entity, ad, event, and customer identifiers. Flink’s RocksDB state tracks seen events over a 30-hour window (approximately 20 GB of state), filtering duplicates while preserving them in Amazon S3 for audit.
  3. Enrich. Flink’s Async I/O API queries Amazon DynamoDB concurrently for campaign metadata and product master codes, populated continuously from upstream Kafka topics by an AWS Fargate consumer.
  4. Attribute. A multi-day keyed interval join matches user event actions to subsequent orders on entity, customer, vendor, and campaign dimensions (approximately 100 GB of state). This stage emits attributed orders to Amazon S3.
  5. Aggregate. The pipeline accumulates impression, click, order, revenue, and ad spend metrics in RocksDB state, then batch-upserts them to the reporting database every 5 minutes.

The pipeline emits billing events (cost per mille (CPM) impressions and valid cost per click (CPC) clicks) to Apache Kafka topics. The ad server and budget service consume those topics in real time. Flink checkpoints all state incrementally to Amazon S3, so the job restores from the last checkpoint after a failure. Kinesis Data Streams and the upstream sources deliver at-least-once, and the deduplication stage in step 2 drops any event replayed during recovery. Billing is therefore effectively exactly-once, even though the transport underneath it is at-least-once.

Results and impact

The redesigned architecture achieved quantifiable performance gains across data fidelity, processing throughput, and operational expenditure, while introducing capabilities that were not feasible under the legacy model.

Processing latency: From hourly windows to real time

The average gap between when an event was published and when it was recorded dropped from 61 minutes to 1.2 seconds. Budget pacing and aggregated metrics now reflect activity within seconds rather than the following hour. Downstream ad serving and budget pacing systems act on real-time signals instead of reconciling after the fact.

Cost efficiency

The migration reduced monthly operational costs by approximately 57 percent, which more than halves the annual run rate for the pipeline. The saving came alongside stronger reliability, not at its expense.

System reliability

Durable attribution window. The multi-day attribution window lives in RocksDB-backed keyed state, roughly 100 GB on local TaskManager disks, checkpointed incrementally to Amazon S3. Per-key lookups stay in the low-millisecond range regardless of state size, and a crash or shard rebalance restores state from the last checkpoint rather than triggering a reconciliation job.

Elasticity replacing fragility. Async I/O against DynamoDB removed the synchronous enrichment chain that previously gated every event. The pipeline sustains 20,000 messages per second at peak without back-pressure leaking into ad serving or billing, and enrichment scales independently of ingestion. Flash sales and seasonal surges no longer threaten upstream systems.

Replayable history. The pipeline persists every raw event to Amazon S3 in Avro format the moment it lands, and Kinesis Data Streams retains the source stream for up to 7 days. When a logic bug surfaces or a downstream contract changes, the team reprocesses the affected time range deterministically against the original inputs. There is no bespoke backfill job and no reconciliation against external systems. Past data is a first-class input, not a frozen artifact.

Data quality at the source

The following table compares the three data quality gaps before and after the migration.

Metric Before After
Missing session rate 30–40% 0%
Missing customer IDs 5% 0.8%
Missing impression timestamps 91% 0.2%

Downstream applications now receive fully enriched transactional and session context. Machine learning models use session-scoped features for campaign ranking, conversion-rate estimation, and anomaly detection. The pipeline now computes those features from a complete event stream, rather than one in which roughly a third of events were missing session context and 91 percent of impressions were missing a reliable timestamp.

What’s next

The pipeline described here is the first of several planned migrations to Amazon Managed Service for Apache Flink. The team is extending the same architecture to additional ad formats, and connecting real-time Flink aggregations directly to the ad serving layer for sub-second budget pacing. The real-time data layer built for measurement also serves as the foundation for AI-driven use cases. The team plans to explore live user interaction streams feeding personalization ranking models and grounded large language model (LLM) recommendations, which were impractical with batch-oriented infrastructure.

Conclusion

Delivery Hero’s migration to Amazon Managed Service for Apache Flink shows that effectively exactly-once billing, multi-day stateful attribution, and manageable operational complexity are not competing goals. The combination that made it work: Kinesis Data Streams for ingestion, DynamoDB for low-latency enrichment, Amazon S3 for event storage and checkpointing, and Amazon EventBridge Pipes for rapid MVP delivery. Together they produced a system that is more accurate, more resilient, and less expensive than the one it replaced. For advertising platforms where billing accuracy and attribution correctness are commercial imperatives, this architecture offers a replicable path from batch approximation to real-time measurement.

To get started with Apache Flink on AWS, see the Amazon Managed Service for Apache Flink Developer Guide.

Additional resources


About the authors

Kirill Tishenkov

Kirill Tishenkov

Kirill is a Senior Software Engineer at Delivery Hero specializing in distributed stream processing and large-scale state management.

Alexandru Pisarenco

Alexandru Pisarenco

Alexandru is a Senior Software Engineer at Delivery Hero focusing on real-time data pipelines, backfill strategies, and multi-market rollouts.

Upendra Kambhampati

Upendra Kambhampati

Upendra is an Engineering Manager at Delivery Hero leading the AdTech Data Engineering team.

Sabariesh Ganesan

Sabariesh Ganesan

Sabariesh is a Senior Engineering Manager at Delivery Hero responsible for the Vendor AdTech Data platform and Ads measurement domain.

Joseph Idicula Watasseril

Joseph Idicula Watasseril

Joseph (he/him) is a Senior Solutions Architect at AWS, based in Berlin. With over 15 years of experience in tech consulting and software development, Joseph works with Delivery Hero to apply cloud solutions to their business challenges.

Francisco Morillo

Francisco Morillo

Francisco is a Senior Streaming Solutions Architect at AWS, specializing in real-time analytics architectures. With over five years in the streaming data space, Francisco has worked as a data analyst for startups and as a big data engineer for consultancies, building streaming data pipelines. He has deep expertise in Amazon Managed Streaming for Apache Kafka (Amazon MSK) and Amazon Managed Service for Apache Flink.

How Moeve standardized dbt runs across data lakes with Amazon Athena

Post Syndicated from Rubén Romero Córdoba original https://aws.amazon.com/blogs/big-data/how-moeve-standardized-dbt-runs-across-data-lakes-with-amazon-athena/

As organizations grow, data processing often becomes fragmented across teams, environments, and orchestration tools. This fragmentation leads to inconsistent patterns, duplicated logic, limited cost visibility, and operational overhead.

At Moeve we were no exception. Our analytics teams build their transformations with dbt, an open source tool that defines transformations as SQL models, resolves the references between them, and works out the order in which they run. dbt describes what to transform, but it does not define where or how a project runs. We left that decision to each team, and as the number of projects grew we found fragmented pipelines, inconsistent compute engines, and limited cost visibility slowing every project down. Standardizing our dbt runs on Amazon Athena was how we worked our way out of that.

This post describes the architecture of the centralized, serverless solution we built on Athena, which reduced onboarding for a new dbt project from days to about 15 minutes. The solution centralizes how dbt runs across our data lakes while staying loosely coupled from orchestration. It uses Amazon Athena as the default processing engine, a centralized dbt launcher, and a shared event bus for downstream orchestration.

In the sections that follow we explain how we decoupled dbt runs from orchestration using AWS Step Functions and AWS Fargate, why Athena fits our workloads from a cost and operational perspective, how storing run parameters in Amazon DynamoDB rather than in pipeline code removed the infrastructure deployment step from onboarding, and how publishing results to Amazon EventBridge lets our run and orchestration layers evolve independently.

The challenge of running dbt at scale

Before the dbt launcher, our dbt runs had grown in different directions. Run logic was embedded in project-specific pipelines, orchestration and processing were tightly coupled, teams selected different compute engines for comparable workloads, and we had no consistent governance over run parameters and retries.

As the number of dbt projects increased, this made it difficult to enforce consistent standards and to evolve the solution without touching every pipeline. We needed a way to standardize dbt runs across our data lakes, decouple running a project from deciding what to run, improve cost control and observability, and deliver faster and safer continuous integration and continuous delivery (CI/CD) iterations.

Why Amazon Athena as the default dbt engine

Choosing the processing engine for dbt is a foundational architectural decision, so we made it first.

Serverless processing

Athena is fully serverless. There are no clusters to provision, scale, or maintain. Our teams run their queries with the default Athena pricing, which charges for the data a query scans and gives us elasticity with no capacity planning.

Because each data lake lives in its own account, each team makes its own decision about Athena payment. A team whose workload grows into continuous, high-concurrency usage can move to Athena capacity reservations with no change to the launcher, to their dbt profiles, or to their project configuration. None of our teams have needed to do so yet, and the architecture keeps that choice independent per team.

Our workload is predictable but not continuous. Each dbt project runs for a few minutes when its schedule fires or when its upstream data lands, then stays idle until the next trigger. That shape is what made serverless the right fit for us.

Why Athena fit our solution

For Moeve, the decision to standardize on dbt and Amazon Athena was driven by our goal of creating a common transformation solution that could be adopted across multiple teams and AWS accounts while keeping operations lightweight.

Our data was already stored in Amazon S3 and registered in the AWS Glue Data Catalog, making Athena a natural processing layer. Athena allowed us to run transformations without managing clusters, capacity, or infrastructure, which was particularly important for a small central team supporting multiple domains.

We evaluated alternative processing engines, but for our workload profile and data volumes, Athena provided the best balance between scalability, operational simplicity, and maintainability. Adapter maturity was another important factor. The dbt Athena adapter offered strong integration with testing, CI/CD workflows, and the broader dbt ecosystem, reducing the operational risk of maintaining custom solutions.

Athena also aligned naturally with our architecture. Transformations run in the AWS account that owns the data through cross-account role assumption, while orchestration remains centralized. As a result, we standardized how projects run across teams while keeping compute close to the data.

Finally, the Apache Iceberg support in Athena underpins the idempotent incremental processing model described in this post, so incremental loads and historical reprocessing follow the same path with minimal operational overhead.

Optimized incremental processing

Our largest cost was not reading source data. It was merging into it.

Our fact tables are Apache Iceberg tables in the data lake, and most of our models are incremental. Each run brings in new or corrected records and merges them into a target that can hold several years of history. A merge has to locate the rows it is about to update, and without a predicate on the target the query reads far more of the table than the incoming data can affect. The common approach is a static filter such as the last 30 days, which is wrong in both directions: too wide for an ordinary daily load, and too narrow as soon as a correction arrives for an older partition.

Instead of a fixed window, the platform derives the predicate from the data. Before the merge runs, it reads the distinct values of the partition column present in the incoming dataset and builds the target predicate from them. For a single partition it applies an equality predicate, for a small set an IN list, and for a larger set a bounded range. The merge then reads only the partitions the incoming data can affect.

The same principle applies on the source side. The launcher builds the source filter from the parameters given for that run: an explicit range, an arbitrary SQL condition, or, when neither is supplied, a default window taken from the project configuration. Input is therefore bounded to the subset each run needs.

Two results mattered to us. Because Athena charges for the data a query scans, narrowing both ends of the merge reduces cost without any team hand-tuning individual models. More importantly, a daily load and a full historical reprocess became the same operation with different inputs. Every run is idempotent, so the same input always produces the same result regardless of how many times it runs. That removed the distinction between processing and reprocessing from our runbooks and simplified incident response.

Table design is what makes this pruning possible. Partitioning, columnar formats, and compression all contribute, and the AWS Big Data Blog post Top 10 performance tuning tips for Amazon Athena covers the general techniques. We have deliberately not published a before and after figure here, because the saving depends so heavily on partition design and data distribution that a single number would mislead without extensive context.

Architecture overview

Moeve built a centralized dbt launcher that runs dbt jobs in a uniform way, regardless of the project or the target data lake.

Architecture diagram

Architecture of the centralized dbt launcher running cross-account dbt jobs on Step Functions, Fargate, and Amazon Athena

Figure 1: Centralized dbt launcher and cross-account run flow

At a high level, the architecture consists of:

  • AWS Step Functions to control the run lifecycle.
  • AWS Fargate to run dbt in an isolated, ephemeral container.
  • Amazon Athena as the default dbt processing engine.
  • Amazon DynamoDB to store dbt project configuration.
  • Amazon EventBridge to publish run results.

Every dbt run follows the same contract, which gives us consistency and reduces the operational surface we must maintain.

Cross-account processing model

The solution operates in a centralized account while running transformations in domain-specific data lake accounts: corporate, marketing, and manufacturing. The Fargate container assumes a dbt-child IAM role in the target account, so the container processes data where it lives while governance stays centralized.

Each data lake account keeps control of its own IAM permissions and manages its own storage and catalog without affecting the solution. This also puts costs in the right account. We could have attributed Athena spend using Athena workgroups, but Athena is only part of what a query costs. The Amazon Simple Storage Service (Amazon S3) requests it makes and the AWS Key Management Service (AWS KMS) operations it triggers are real costs as well, and running in the owning account attributes all of them to the team that owns the data, per project and per run.

Centralizing dbt runs with the dbt launcher

To stop every team inventing its own way of running dbt, we built a single launcher that all of them go through. Instead of dbt logic living inside multiple pipelines, every run is triggered through one well-defined path.

Run lifecycle

The launcher is an AWS Step Functions state machine. It receives a run request with its parameters, starts an AWS Fargate task from our dbt container image, and the container assumes the IAM role of the target data lake account. dbt then runs its SQL transformations in Athena, reading the source tables and materializing the targets. Alongside the run, Elementary, an open source dbt package, records model-level results and data quality test outcomes. When the run finishes, the launcher publishes a completion event to Amazon EventBridge.

The state machine can be started in different ways depending on the scenario. Some projects run on a schedule, others are triggered when upstream data lands, and teams can request a run on demand. Those decisions are made by our orchestration layer, which submits a standardized run request to the launcher. The launcher therefore stays focused on running dbt projects, regardless of how the run was initiated.

This sequence is identical for all projects and environments. The launcher is responsible only for running the project it was asked to run. It doesn’t decide what should run next, and that responsibility is intentionally delegated to downstream consumers through Amazon EventBridge.

Configuration-driven runs with Amazon DynamoDB

We had to decide where a project’s run parameters would live. In the pipeline definition, changing a timeout would be a code change, a review, a build, and a deployment, for a value we sometimes need to change while an incident is open. We put the parameters in DynamoDB instead, keyed by project, and the launcher reads them at the start of every run. That lets us decouple code deployment from run behavior, update parameters without redeploying services, and enforce consistent defaults across all dbt projects.

The parameters themselves are modest. They cover the target environment and AWS Identity and Access Management (IAM) role, the default processing engine, how long to allow a run to take, how many times to retry on failure, and how many days of data to process by default. These values change for operational reasons rather than logical ones, which is why we did not want them coupled to a release cycle.

The effect on onboarding was larger than we expected. Deploying a new dbt project is now a merge of the dbt models and one configuration entry. There is no Terraform change, no infrastructure review, and nothing to provision, because the compute the project needs already exists and is shared. What used to be a multi-step infrastructure pipeline is now a single CI workflow that validates the project’s SQL and lineage locally, then merges and registers it. We run that local validation with DuckDB, which returns feedback in under two minutes without consuming cloud resources.

Governance did not weaken as a result. Who may change a configuration entry is controlled the same way as any other production change. What changed is that the change no longer has to travel through an infrastructure deployment to take effect.

CI/CD pipeline diagram

A pull request triggers local validation of the project’s SQL and lineage. On merge, the dbt models are deployed and the project’s configuration entry is registered in DynamoDB, after which the launcher can run the project.

Figure 2: CI workflow for onboarding and updating a dbt project

Athena remains the engine of record. Local validation catches Jinja errors, unresolved references, and obvious SQL mistakes, but Athena-specific behavior, cross-account permissions, and AWS Glue Data Catalog interactions are only proven in the target environment. It’s important to be explicit about that boundary with the teams, so that a green CI run is not read as a guarantee.

Publishing run results with Amazon EventBridge

After a dbt run finishes, the launcher publishes a structured event to a central Amazon EventBridge bus recording whether the run succeeded, which project and source it covered, when it started and finished, how long it took, and which datasets it updated. The launcher does not know which consumers are subscribed.

This event-driven approach gives us loose coupling between running a project and orchestrating what comes next, multiple downstream consumers for the same signal, and independent evolution of both layers.

At Moeve the main consumer is our orchestration layer, which models the dependencies between datasets as a graph. Each completion event tells it that a node is now up to date, so it can determine which downstream projects have all their inputs ready and start them. That consumer has no special status. An AWS Lambda function, a monitoring dashboard, or a notification integration can subscribe to the same events without any change to the launcher.

Validation and observability

Because every project follows the same lifecycle, we get validation and observability in one place instead of per pipeline. Step Functions shows the state of any run and the step at which it failed, and error handling and retries are defined once. Fargate logs carry the container runtime detail. dbt and Elementary report model-level results and data quality test outcomes. The completion event on the bus is the auditable record that a project finished and what it produced. Together these layers give us operational visibility without coupling the components to each other.

Cost control and resource cleanup

The platform is serverless end to end, which keeps idle cost close to zero and removes a class of operational mistake. Fargate tasks are created for a run and destroyed when it ends, so no long-running container needs maintenance. Athena has no persistent compute. An idle Step Functions state machine costs nothing. Nothing is left running unintentionally, which matters when the number of projects on the platform keeps growing.

Results

The metrics in the following table are the ones our teams notice day to day, and the reason the solution is maintainable by a small central team.

Metric Before After
Onboarding time for a new dbt project Days, including pipeline and infrastructure setup About 15 minutes, configuration only
Run consistency Varied by team Same lifecycle and contract for every project
CI feedback time 8 to 10 minutes on Jenkins Under 2 minutes with local validation
Cost visibility Per-account aggregate Per-project and per-run attribution
Operational overhead One pipeline per project One solution for all projects

Conclusion

Standardizing how we run dbt turned out to depend less on dbt than on defining two boundaries clearly.

The first is the contract of the launcher: parameters in, event out. We defined that interface before building the internals, and it has stayed stable while the implementation changed several times. The second is the separation between running a project and deciding what to run next. Making the launcher publish events without knowing its consumers is why our orchestration layer could be rebuilt while the launcher stayed as it was, and the launcher has never been modified to accommodate a new orchestration requirement.

Two smaller decisions carried more weight than we expected. Keeping run parameters in DynamoDB rather than in pipeline definitions means timeouts, retries, and engine selection can be changed without a deployment, which is valuable during incident response. Making every run idempotent removed the distinction between processing and reprocessing, so a daily load and a full historical reprocess are the same operation with different inputs.

Athena is serverless, so our teams did not need to set up infrastructure of their own. That is what made it practical for everyone to run dbt in a standardized way, and why onboarding a new project went from days to about 15 minutes.

If your organization runs dbt across multiple accounts and teams, pair Amazon Athena with a clear contract for the component that runs your projects: fixed parameters in, a published event out.

Resources


About the authors

Rubén Romero Córdoba

Rubén Romero Córdoba

Rubén is a Cloud and Data Architect at Keepler with 10+ years spanning AI research, software engineering, and AWS data architecture. He designs secure, scalable, and maintainable data platforms focused on lakehouse architectures, governance, and operational efficiency. Curious and pragmatic, he studies how systems work, explores emerging tech, shares knowledge, and favors simple solutions that deliver real value without unnecessary complexity.

Ricardo Bravo Panes

Ricardo Bravo Panes

Ricardo is a Data Architect at Moeve, where he designs cloud-native data solutions on AWS, turning complex challenges into scalable and sustainable solutions. He is passionate about data, distributed systems, and simplifying complexity to help teams move faster.

Álvaro Ponce Cabrera

Álvaro Ponce Cabrera

Álvaro is a Data Engineer and Data Platform Lead at Moeve, focused on building scalable data products and cloud-native solutions that connect industrial and business data. His interests span data architecture, governance, AI, and developer experience, always seeking pragmatic solutions that maximize business value while keeping complexity under control.

Gonzalo Guerrero Leon

Gonzalo Guerrero Leon

Gonzalo is a TAM at AWS who empowers enterprise customers through strategic technical guidance. Throughout his 10-year tenure at Amazon, he’s contributed to multiple cornerstone divisions, including HR, IT, Alexa, Amazon Business, and AWS, gaining insight into the technology landscape. Outside of work, Gonzalo enjoys playing volleyball with his wife and exploring the world alongside their adventurous Boston Terrier, Tigre.

How CSIRO built scalable, cost-optimized genomic variant querying on AWS

Post Syndicated from Prof. Denis Bauer original https://aws.amazon.com/blogs/architecture/how-csiro-built-scalable-cost-optimized-genomic-variant-querying-on-aws/

This is a guest post by Denis Bauer, Yatish Jain, Anuradha Wickramarachchi, Brendan Hosking, and Nick Edwards of CSIRO, in collaboration with the ASP Prototyping and Scaling Team at AWS.

In this post, we describe how researchers at CSIRO, Australia’s national science agency, built Serverless Beacon (sBeacon), a scalable serverless solution for securely querying genomic variant data on AWS, underpinning production-scale clinical and research applications.

The Beacon protocol is the widely adopted standard for exchanging genomic and phenotypic data developed by the Global Alliance for Genomics and Health (GA4GH). It uses an API to define how data is shared, with the goal of enabling efficient and secure data discovery across international research and clinical networks.

sBeacon is a production-ready implementation of this standard, built using AWS services: Amazon Simple Storage Service (Amazon S3), AWS Lambda, Amazon DynamoDB, and Amazon Athena. By using these foundational AWS serverless services, sBeacon is able to provide the following benefits to researchers and clinicians needing to perform genomic variant querying:

  • Highly scalable for large cohorts: sBeacon can scale to support hundreds of millions of individuals (and billions of genomic locations), which makes it suitable even for mega-biobank-scale datasets.
  • Low cost to run: Because it uses a serverless, cloud-native architecture, sBeacon can operate for approximately USD 0.40 per month for a 1000 Genomes-scale dataset. The following case study breaks down ingestion, query, and storage costs in detail.
  • High performance and fast query response: Real-world queries return in seconds (about 5 seconds) because of the serverless compute and efficient architecture, for near real-time data lookups.
  • No heavy data ingestion or transformation needed: sBeacon can directly consume standard VCF files (a common format for genomic variant data), which reduces the need to load data into databases or transform it to different data structures.
  • Rapid onboarding of new data: Genomic data generation is accelerating because it underpins clinical diagnosis and treatment, and because its complexity demands ever-larger cohorts to study complex traits. As a result, both clinical services and research cohorts must continuously onboard new data, and Beacon supports real-time generation-to-use life cycles (about 18 seconds).
  • Improved privacy, data ownership, and decentralization: Because sBeacon doesn’t require central databases and supports federated networks, data stays under the control of original holders, which can help data custodians address privacy and ethical considerations in sensitive genomic and medical data sharing.
  • Lower barrier to entry for broader participation: Its affordability, simplicity, and small operational footprint can help make it more accessible for smaller or resource-limited institutions and countries, which can increase participation from underrepresented populations and improve data diversity.
  • Zero trust model: sBeacon enforces explicit authentication, least-privilege data access, ephemeral compute isolation, and strict cloud-native boundary controls that help confirm no component, user, or request is implicitly trusted.

Prerequisites

sBeacon is deployed as a container that sets up the necessary development environment, with Terraform defining the resources for the deployment. To get started, clone the terraform-aws-serverless-beacon repository on the GitHub website.

git clone https://github.com/aehrc/terraform-aws-serverless-beacon.git

Make sure that your development environment contains Docker and has the necessary permissions for you to use it without super user access. Press Ctrl+Shift+P (Cmd+Shift+P on macOS) to open the command palette in VS Code, and then choose Reopen in Container. This opens the workspace in the container environment that we have defined.

Now, run the following command to initialize the necessary libraries and Lambda layers.

bash init.sh

Next, run the following command to initialize the Terraform environment.

terraform init

Optionally, you can define a backend by following the instructions in the repository. After the preceding command runs successfully, you can run the deployment command.

terraform apply

Enter yes when prompted to proceed with the deployment. After the deployment is complete, you receive information such as the API URL and the command to sign in as the admin or guest user. To shut down the entire service, run terraform destroy. Any created datasets are lost (but not the VCFs on which they are based).

Solution walkthrough

CSIRO developed sBeacon for sharing and querying genomic and medical data. sBeacon uses AWS serverless technology for the elastic scaling of compute resources.

The architecture of sBeacon performs two broad processes:

  1. Data onboarding: the ingestion and indexing of genomic metadata into sBeacon.
  2. Data querying: the querying of the genomic metadata by end users.

Data onboarding

During the onboarding process, you define where the genomic data and the metadata (such as disease status, age, and location) is located. Note that genomic data is not copied out of its original location but rather is referenced when needed. In contrast, metadata is loaded to sBeacon’s storage mechanisms because it is necessary to perform indexing that allows efficient querying. The user will need to ensure no sensitive or privacy-revealing data is disclosed. The example details the approach using CSIRO’s Ontoserver. However, sBeacon supports the API schema of the Ensembl OLS V4 specification.

Data onboarding architecture for sBeacon, showing genomic data location submitted to an API Gateway endpoint, AWS Lambda functions handling indexing, metadata written to Amazon S3 in ORC format, CSIRO Ontoserver building the ontology index, and Amazon Athena building the metadata tables.

Figure 1. Data onboarding.

The data onboarding process is summarized by the following steps:

  1. The onboarding starts with the user submitting the location of the genomic data as request payloads to an API Gateway endpoint.
  2. The request payloads are forwarded to an AWS Lambda function that handles the data indexing.
  3. The metadata is written to an Amazon S3 bucket in the ORC format, to allow future querying and processing by Athena.
  4. An AWS Lambda function is called to orchestrate the indexing process.
  5. The CSIRO Ontoserver is called to build the ontology index for advanced metadata queries.
  6. The resulting index files are written to Amazon S3.
  7. CREATE TABLE AS SELECT (CTAS) queries are run on Amazon Athena to build the metadata tables.
  8. Athena loads the metadata from Amazon S3 into the metadata tables.
  9. The metadata tables are written back to Amazon S3 in ORC format.

Data querying

Querying in sBeacon is flexible, catering to a wide range of applications from human genetic disease to pathogen queries. We achieved this by designing the query architecture modularly. This approach let us separate the querying logic into several Lambda functions based on their querying scope, while maintaining a similar architecture.

The following architecture diagram describes the workflow for metadata querying, which uses the Variant Querying Module described later in this section.

Metadata querying architecture for sBeacon, showing a user query sent to an API Gateway endpoint, the Microservice Lambda function looking up ontology terms in Amazon DynamoDB, querying metadata tables in Amazon Athena, and querying the Variant Querying Module before returning a Beacon-formatted result.

Figure 2. Data querying.

  1. The user submits their query to the API Gateway endpoint.
  2. API Gateway calls the Microservice Lambda function.
  3. The Microservice Lambda function looks up the relevant query ontology terms in an Amazon DynamoDB table.
  4. The matching ontology descendent terms (and their codes) are returned to the Microservice Lambda function. The descendent terms are those that match a hierarchical descendent of each term, or each term itself, from the query.
  5. Using the ontology codes from step 3, the metadata tables on Athena are queried.
  6. The metadata associated with the query is returned from Athena.
  7. If required by the query, the Microservice Lambda function queries the Variant Querying Module.
  8. The variant data associated with the genomic conditions in the query is returned to the Microservice Lambda function.
  9. The result is formatted according to the Beacon protocol and is returned to the user through Amazon API Gateway.
  10. The response is received by the user.
Variant Querying Module architecture for sBeacon, showing an Initiator Lambda function fanning out the splitQuery and performQuery Lambda functions across VCF files in Amazon S3, and optionally querying metadata from Amazon Athena before returning results to the Microservice Lambda function.

Figure 3. Variant Querying Module.

Genomic variant queries are performed using the Variant Querying Module. The workflow of this module is as follows:

  1. The Microservice Lambda function calls an Initiator Lambda function.
  2. The Initiator Lambda function fans out the splitQuery Lambda function across the VCF files.
  3. The performQuery Lambda function is then fanned out across the VCF regions in each of the files involved in the query.
  4. The performQuery Lambda function fetches the VCF files from Amazon S3.
  5. The query results are synchronously returned to the parent Initiator Lambda function.
  6. If requested by the user, metadata can optionally be queried, where the Initiator Lambda function queries the metadata from Athena.
  7. Athena queries the metadata from Amazon S3 (through an external table).
  8. The metadata results are returned to Athena.
  9. The Initiator Lambda function receives the metadata from Athena.
  10. All the query results, including any optional metadata, are returned to the calling Microservice Lambda function.

Case study: 1000 Genomes dataset

We demonstrate sBeacon on chromosome 1 of the 1000 Genomes Project to report how it handles large-scale variant queries. We measure ingestion efficiency, query scalability, and cost for typical population-scale analyses, such as identifying SNP variants across defined genomic regions. The case study uses chromosome 1 (chr1, 8% of the genome) from the 1000 Genomes Project, which contains 2504 samples. This multi-sample VCF is approximately 1.1 GB compressed, with data stored in Amazon S3. Note that sBeacon can also process cohorts of single-sample VCF files. All costs in this section are for the Asia Pacific (Sydney) Region (ap-southeast-2), exclude applicable taxes, and reflect pricing at the time of writing.

sBeacon can ingest chromosome 1 from the 2504 individuals in 18 seconds, for less than 1 cent (USD 0.00052). This is because sBeacon does not copy the large genomic information but instead creates index files that enable random access. Cost is therefore driven predominantly by storing the copied metadata. After ingestion, sBeacon can be maintained for USD 0.000025 per month (1 MB of compressed metadata stored for 2504 samples in ORC format, plus genomic index files). If you store the genomic data as well, this would be USD 0.032 for chr1 (at USD 0.025 per GB in ap-southeast-2) or about USD 0.425 for the whole genome.

Query time is similarly near real time. For example, querying across a region of 10,000 base pairs to determine the genotypes in this region takes 1.52 seconds across the 2504 individuals. This would serve a query such as “Fetch all individuals with a specific BRCA1 mutation who have stage 3 cancer.” The cost for such a query is USD 0.00013. Note how the query time stays constant even with an increasing number of variants returned (for example, from 4 to 400).

Table 1. Query example costing and times (whole chromosome 1).

Query region size (bases) Number of variants found Average Time Compute Cost (per query in USD)
10 4 1.51 s (+- 0.26) 0.00013
100 18 1.52 s (+- 0.25) 0.00013
1,000 84 1.62 s (+- 0.24) 0.00014
5,000 229 1.65 s (+- 0.29) 0.00014
10,000 400 1.52 s (+- 0.11) 0.00013

Table 2. Cost for ingestion, querying, and idling (whole chromosome 1 for 2504 genomes with less than 10 MB of metadata).

Scenario Metric Cost (USD) per month
Ingestion Cost per 1000 ingestions 0.53 (32.82 GB seconds of Lambda)
Query compute cost per 1000 queries 0.28 (9.8 GB seconds of Lambda)
Query Athena Cost per 1000 queries 0.05
Idle Cost (Storage Cost) 1.1 GB 0.03
Query DynamoDB Cost Per 1000 queries 0.0005

Security features

Security and compliance is a shared responsibility between AWS and the customer. AWS is responsible for protecting the infrastructure that runs the AWS services described in this post, and you are responsible for your use of those services, including how you configure them, which identities you grant access to, and which data you choose to onboard. Consider the services you choose carefully, because your responsibilities vary depending on the services used, how you integrate those services into your IT environment, and applicable laws and regulations. For more information, see the AWS Shared Responsibility Model.

Zero trust model

  • Explicit authentication and authorization – Every API request must carry a valid JWT issued by the Amazon Cognito user pool (aws_api_gateway_authorizer.BeaconUserPool-authorizer, type COGNITO_USER_POOLS). The authorizer runs at API Gateway before any Lambda function is invoked, so requests do not reach a handler without Cognito validation. Token validation includes signature, expiry, and audience (Cognito app client ID). You can disable authentication during the first deployment with BEACON_ENABLE_AUTH = false for intentionally public or open beacons. This is an explicit operator decision, not a default.

Authorization (what a valid user can do) is enforced inside the Lambda layer, not in Amazon API Gateway:

  • Group membership (sbeacon-record-access-user-group, and so on) controls the maximum granularity returned.
  • Admin-only operations (dataset submission, deletion) check for sbeacon-admin-group membership before proceeding.
  • Least-privilege data access – sBeacon implements role-based access control (RBAC) through Cognito groups that map directly to disclosure tiers. You assign each user one or more of the following:
Cognito group Maximum disclosure
sbeacon-boolean-access-user-group exists: true/false only
sbeacon-count-access-user-group aggregate counts
sbeacon-record-access-user-group full variant details and sample names
sbeacon-admin-group preceding tiers plus dataset management

The JWT carries the user’s group memberships as claims. The query Lambda function reads these claims to determine requested_granularity and include_details, then passes both flags to performQuery. performQuery computes only what was requested. A boolean-tier user’s request does not cause sample-level data to be computed or returned, even if it exists in the VCF.

  • Ephemeral compute isolation – Lambda execution environments are stateless by design. Each cold start is a fresh container, /tmp (1,024 MB for performQuery) is cleared between cold starts, and concurrent invocations run in separate sandboxes with no shared memory. The bcftools subprocess inside performQuery runs and exits within the Lambda function lifetime (10 second timeout). No state persists after invocation.
  • Cloud-native boundary controls – API Gateway is the public entry point in this architecture. Amazon S3 buckets, DynamoDB tables, Athena, and Amazon SNS topics have no public resource policies. Amazon S3 buckets are created with private ACLs and BucketOwnerPreferred ownership controls. Lambda functions run on AWS-managed VPCs with no inbound network access. Amazon SNS topics are account-private (no external principal grants).

Privacy and data ownership

Each institution deploys the entire Terraform stack into its own AWS account, so there is no shared infrastructure, no central data lake, and no cross-account trust. VCF files live in the deploying institution’s Amazon S3 bucket and do not leave it. performQuery passes the Amazon S3 URL directly to bcftools as a subprocess argument, which uses htslib HTTP byte-range requests to read only the tabix-indexed region of interest (about 1 KB per query). The raw genomic sequence bytes do not pass through Lambda memory as returnable data. What the query returns upstream (exists as a boolean, call_count as an integer, and variant representations) is aggregate result data, not source sequence.

Decentralization in sBeacon is achieved at the storage layer, not the compute layer. The _vcfLocations registered for a dataset are Amazon S3 URIs, and these can point to buckets owned by entirely different organizations. When a query runs, performQuery passes each URI directly to bcftools, and htslib issues HTTP byte-range requests (Range: bytes=X-Y) against the Amazon S3 REST API of whichever organization owns that bucket. The raw VCF bytes do not leave the source organization’s Amazon S3 bucket. Only the query result (exists, count, or variant record) is returned.

Data onboarding privacy

The submitDataset endpoint sits behind the same API Gateway Cognito authorizer as all other endpoints. An unauthenticated request receives a 401 response before reaching any Lambda function. Beyond authentication, the handler also checks that the caller is a member of sbeacon-admin-group. A valid token from a user in only record-access or count-access is rejected. This means the beacon operator explicitly controls the set of people who can introduce data into the system, so onboarding is not a self-service capability.

Further considerations

We chose AWS Lambda over AWS Step Functions in this architecture because it can process much larger payloads. Given the size and complexity of genomic data and the fan-in and fan-out architecture for parallel handling, AWS Lambda emerged as the lower-cost and more flexible approach for this workload.

As demonstrated in the sBeacon publication, the architecture can cater to population-scale datasets. However, if you accidentally attempt to run a range query of the entire genome, the architecture times out at the Amazon API Gateway level. Applying functional operations over the whole genome requires further architectural considerations.

Because a single fan-out query spawns many parallel Lambda invocations, you need to monitor concurrency consumption to confirm that burst queries do not exhaust the account’s concurrency pool and starve other functions. Tracking the ConcurrentExecutions metric at both the account and function level provides early visibility into capacity pressure.

Similarly, because synchronous Lambda invoke does not automatically retry on throttle, a 429 response from a performQuery invocation means the result is silently lost unless the application handles it explicitly. Setting Amazon CloudWatch alarms on the Throttles metric for performQuery allows you to take corrective action, such as requesting a concurrency limit increase, before throttles affect query accuracy. Alternatively, we have produced a separate architecture that sends alert email with diagnostic information when Lambda functions fail, available in the error-catcher repository on the GitHub website. You can implement this in the repository or set it up as a standalone service to catch Lambda errors thrown by sBeacon.

After idle periods, simultaneous performQuery invocations might encounter cold starts that add latency to query responses. Enabling provisioned concurrency on the query-path Lambda functions helps reduce this cold-start latency during burst fan-out scenarios at the price of increasing the idle cost.

Conclusion

In this post, we described how CSIRO built sBeacon, a fast, scalable, and low-cost way to run genomics workloads on AWS. sBeacon implements the GA4GH Beacon standard with a fully serverless and modular architecture. This publicly available solution supports near real-time querying of standard VCF data, scales to mega-biobank cohorts, minimizes ingestion effort, and supports privacy and zero-trust security. If you are considering genomics on AWS, you can deploy sBeacon on existing Amazon S3-hosted VCF data, integrate it with clinical or research workflows through the Beacon API, and progressively federate with other Beacons for secure, cross-institutional genomic data discovery. Set up sBeacon to query your genomic data and explore the possibilities of securely sharing insights with your collaborators. You can read more about sBeacon in our publication: Scalable genomic data exchange and analytics with sBeacon. The source code for sBeacon can be downloaded from our GitHub repository.


About the authors

How Equinix cut operational overhead with a shared services architecture on Amazon EKS

Post Syndicated from Chhavi Kaushik original https://aws.amazon.com/blogs/architecture/how-equinix-cut-operational-overhead-with-a-shared-services-architecture-on-amazon-eks/

This post is cowritten by Manikandan Vasu, Vanji Sivajothy, and Ramchandra Koty from Equinix.

Equinix is the world’s digital infrastructure company, operating over 260 data centers across more than 70 metros globally. To address the operational sprawl that had grown from its earlier self-managed Kubernetes environment, Equinix built a shared services architecture on Amazon EKS. Previously, Equinix operated a self-managed Kubernetes environment on Amazon EC2 instances. In this model, they provisioned EC2 instances serving as etcd, control plane, and worker nodes, relying on open-source tooling to automate the installation and configuration of Kubernetes components. While this approach provided control and flexibility in the early stages of their Kubernetes journey, it also introduced a structural problem: individual application teams independently provisioning and managing their own clusters, each with its own lifecycle, configuration, and operational patterns.

Over time, this decentralized ownership model created significant operational sprawl. With no shared infrastructure layer, the cloud operations team had no consistent mechanism to enforce governance, standardize configurations, or provide common infrastructure services across the organization. Each application team operated in isolation, making decisions about networking, observability, security policies, and deployment pipelines independently. This resulted in a fragmented environment that was increasingly difficult to manage, secure, and scale.

Key operational challenges included:

  • Operational sprawl – Independently managed clusters led to duplicated infrastructure and no unified operational baseline.
  • No centralized governance – The cloud operations team had no mechanism to enforce network isolation, security policies, or deployment standards across team-owned clusters.
  • Lack of shared services model – Common needs (CI/CD, observability, data services, networking) were solved differently by each team, creating redundancy.
  • Cluster lifecycle complexity – Upgrading and patching control planes across multiple self-managed clusters introduced compounding risk with every cycle.

These challenges made it clear that Equinix needed a fundamental architectural shift, from a model where every team managed its own cluster. They moved to a shared services architecture where the cloud operations team centrally owns and governs the infrastructure while application teams focus purely on their business services. This led to their migration to Amazon EKS with a multi-account strategy designed for scalability, security, and centralized control.

In this post, we walk through how Equinix designed and implemented their shared services architecture on Amazon EKS after migrating from their self-managed Kubernetes environment. We also cover the operational results they achieved, including 4x faster deployments and significantly reduced operational overhead.

The solution: A shared services architecture on Amazon EKS (North Star architecture)

Equinix North Star architecture: a multi-account Amazon EKS design with workload isolation, centralized shared services, and hybrid connectivity through AWS Transit Gateway and AWS Direct Connect

Figure 1: Equinix North Star architecture, a multi-account Amazon EKS architecture with workload isolation, centralized shared services, and hybrid connectivity through AWS Transit Gateway and AWS Direct Connect

Key architecture capabilities

The North Star architecture implements a multi-account, shared services model on Amazon EKS that cleanly separates concerns between application teams and the cloud operations team. The diagram illustrates the user acceptance testing (UAT) environment, with identical patterns replicated across non-production (system integration testing and development) and production, built around three core AWS accounts:

  1. Workloads Account (Workloads-VPC): The Workloads account hosts all application team services running on Amazon EKS in a dedicated VPC deployed across two Availability Zones (AZ-1 and AZ-2) in us-west-1 for high availability. Traffic ingress is managed through Application Load Balancer (ALB) with Kubernetes Gateway API (GatewayClass and Gateway resources) providing precise routing to application namespaces. Cilium serves as the Container Network Interface (CNI), enforcing network policies that isolate workloads at the pod level, while Hubble provides real-time network observability across all application traffic flows.
  2. Platform Account (Platform-VPC): The Platform account is owned and operated by the cloud operations team, housing all shared infrastructure services in a separate VPC. This includes:
    • Managed data services: Amazon RDS, Amazon MSK, Amazon OpenSearch Service, Amazon MQ, and Amazon S3 consumed by application workloads across the account boundary.
    • CI/CD infrastructure: GitHub Runners deployed as EKS workloads, providing a centralized, self-service pipeline for all application teams.
    • Shared services: Capabilities managed centrally and accessible to all application teams.

Traffic ingress to platform services is handled through Network Load Balancer (NLB), with the same
Cilium/Hubble networking and observability stack as the Workloads account.

  1. Network Account (Network-VPC): A dedicated Network account acts as the connectivity hub, implementing an AWS Transit Gateway architecture that spans two Regions (us-west-1 and us-east-2) with Transit Gateway peering between them. Key networking capabilities include:
    • Cross-account connectivity: AWS Transit Gateway attachments connect both the Workloads and Platform VPCs to the central hub, enabling controlled communication between accounts.
    • DNS resolution Amazon Route 53 Resolver endpoints (inbound and outbound) in each Availability Zone, with a private hosted zone providing service discovery across the platform.
    • Hybrid connectivity: AWS Direct Connect Gateway with dual circuits connecting back to on-premises Equinix border routers through network firewalls for security enforcement.

Results: Measurable impact across the organization

The migration to Amazon EKS delivered clear, measurable outcomes that validated Equinix’s North Star architecture strategy:

40%

Reduction in operational overhead by eliminating the need to manage Kubernetes infrastructure. AWS now handles upgrades, patching, and high availability automatically.

4x

Increase in deployment frequency, enabling engineering teams to ship features and updates faster, freed from the constraints of infrastructure bottlenecks.

100%

Unified architecture adopted across multiple business organizations, establishing a single, consistent operating model for all containerized workloads.

Additional operational improvements include:

  • Enhanced developer productivity: Standardized CI/CD workflows through centralized GitHub Runners and self-service namespace provisioning reduced friction across the development lifecycle, allowing application teams to deploy independently without cloud operations team intervention.
  • Improved observability: Hubble provides unified network flow visibility across both clusters, replacing fragmented, team-specific monitoring.
  • Strengthened security posture: Multi-account isolation between application and shared services workloads, combined with Cilium network policies for pod-level segmentation, reduced the scope of potential security incidents and simplified compliance enforcement.
  • Accelerated onboarding: New application teams onboard to the North Star architecture in days rather than weeks, deploying into pre-configured namespaces with access to shared data services, CI/CD pipelines, and observability. They do this without provisioning or managing cluster infrastructure.

“Amazon EKS gave us the foundation we needed to establish our North Star architecture – a scalable, standardized infrastructure that lets our engineers focus on what matters most: delivering innovation for our customers.”

Conclusion

Equinix’s migration to Amazon EKS from a self-managed environment demonstrates how leading digital infrastructure companies are using managed services from AWS to eliminate undifferentiated operational work and redeploy engineering talent toward higher-value innovation.

The North Star architecture now serves as Equinix’s blueprint for scaling containerized workloads across additional business organizations and geographies. It is an architecture that grows with the company while maintaining the governance, security, and operational consistency that enterprise-scale infrastructure demands.

If you are managing complex, distributed infrastructure, the broader takeaway is this: when you consolidate operational ownership onto a well-architected infrastructure and remove the burden of cluster management from your application teams, you can achieve a markedly different pace of innovation.

To get started with your own Amazon EKS deployment, visit the Amazon EKS product page or follow the Getting started with Amazon EKS guide. You can also explore the EKS Best Practices Guide for recommendations on multi-tenancy, networking, and security.


About the authors

How DHI Group accelerates generative AI workloads from idea to production using hackathons

Post Syndicated from Umesh Kalaspurkar original https://aws.amazon.com/blogs/architecture/how-dhi-group-accelerates-generative-ai-workloads-from-idea-to-production-using-hackathons/

With the advent of generative AI, organizations across industries face a common challenge: how do you move from the experimentation and ideation phase to production-ready workloads quickly and confidently? Many teams get stuck in a cycle of proofs of concept that never ship. DHI Group, a leader in talent acquisition services, was evaluating options to accelerate its generative AI adoption in an effort to roll out features at an accelerated pace. The traditional software development lifecycle (SDLC) approach involved months of requirements gathering, architecture reviews, and phased development that wouldn’t deliver the speed DHI needed. They needed a mechanism that would simultaneously validate technical feasibility, build organizational AI literacy, and produce shippable code.

In this post, explore how AWS partnered with DHI Group using a structured Hackathon Acceleration Package (HAP) to quickly generate production-grade artifacts, accelerate organizational AI confidence, and create a repeatable framework for innovation.

Hackathon Acceleration Package

In this section, review how DHI and AWS collaborated to plan and host hackathons to achieve the key business outcomes defined by DHI leadership. The entire process can be split into four phases:

Phase 1: Preparation

In the initial phase, the AWS team and DHI leadership collaborated to define the key outcomes the participants would work toward. The hackathon themes included:

  • Interpreting Job Descriptions Better: Enhancing the system’s parsing and presentation of job requirements.
  • Premium Candidate Experience: Defining what “Premium” means from the candidate’s perspective.
  • Onboarding That Sticks: Guiding new users through uncertainty to realize value sooner.
  • Candidate Engagement & Stickiness: Sustaining candidate engagement and return visits.
  • AgileATS Network: Streamlining the ClearanceJobs–AgileATS integration.
  • Streamlining Recruiter Experience: Reducing friction across the recruiter workflow.

Phase 2: Enablement

To support these outcomes, the AWS team curated and delivered training sessions and hands-on workshops covering generative AI concepts across Amazon Bedrock AgentCore and the AI-driven development lifecycle (AI-DLC). DHI has embraced Kiro as its productivity tool of choice, so AWS tailored the workshops around Kiro, giving participants prescriptive guidance on applying it across the full software development lifecycle.

Phase 3: Hackathon

The three-day hackathon was hosted by DHI at their headquarters in Des Moines, Iowa, and was attended by 20 DHI participants split across 3 teams. The key objective was to build a prototype that could then be accelerated to production. An AWS team of Solutions Architects (SAs) was present on-site to provide technical guidance to the participants. On the final day, a panel of judges comprising senior DHI leadership evaluated the teams to identify the winner. The three use cases the teams worked on:

  • Real-time Employer Analytics Dashboard: Addressing the Streamlining Recruiter Experience theme, this team built a real-time Employer Analytics Dashboard powered by Amazon Bedrock AgentCore and the Strands framework. The solution automates Quarterly Business Review (QBR) reporting for ClearanceJobs’ employer customers, replacing a manual process that currently demands 3+ QBRs per week across 250 customers.
  • Intelligent Candidate Matching: Addressing the Interpreting Job Descriptions Better and Premium Candidate Experience themes, this team built an intelligent candidate matching system with a real-time analytics dashboard. The solution combines Amazon OpenSearch Service for semantic search, Amazon Bedrock for matching intelligence, and Kiro for rapid frontend development.
  • ClearanceJobs MCP Server + AgileATS: Addressing the AgileATS Network and Streamlining Recruiter Experience themes, this team built a unified talent marketplace that connects ClearanceJobs and AgileATS through an agentic AI layer. By creating a single intelligent interface spanning both systems, the solution significantly boosts recruiter efficiency.

Phase 4: Path to production

DHI leadership was committed to advancing all three hackathon use cases to production, a strong signal of the value each prototype demonstrated. Building on the hackathon’s momentum, DHI and AWS aligned on a roadmap to harden each solution, address scalability and security requirements, and integrate them into DHI’s existing system.

In the next section, we focus on the winning hackathon use case, ClearanceJobs MCP Server + AgileATS, and dive deeper into the architecture.

ClearanceJobs MCP Server + AgileATS

High-level overview of the ClearanceJobs MCP Server and AgileATS agentic solution

Figure 1: High-level overview of the unified ClearanceJobs and AgileATS solution

The winning team’s solution represents a modern agentic AI architecture pattern that’s broadly applicable to organizations looking to unify disparate systems through intelligent automation. The architecture uses the Model Context Protocol (MCP) to expose system capabilities as tools that an AI agent can orchestrate.

Detailed agentic architecture spanning the AgileATS and ClearanceJobs accounts, with Amazon Bedrock AgentCore orchestrating MCP server tools

Figure 2: Agentic architecture for the unified ClearanceJobs and AgileATS talent marketplace

How it works

The solution creates a unified recruiter experience by exposing ClearanceJobs capabilities through an MCP server, orchestrated by an intelligent agent built on Amazon Bedrock AgentCore. A separate ProfileLookup AWS Lambda function provides GitHub profile enrichment for candidates.

The problem it solves: Recruiters on ClearanceJobs currently lack an intelligent interface that can search candidates, retrieve profiles, and enrich them with external data such as GitHub profiles in a single conversational flow. This gap requires manual cross-referencing across systems.

The solution: The team built a single agentic interface where recruiters can issue natural-language commands, such as “Find top cleared software engineers with strong GitHub profiles and add them to my pipeline.” The agent handles the multi-step orchestration automatically, with session memory preserving context and preferences across interactions.

Architecture components

Amazon Bedrock AgentCore (orchestration layer)
AgentCore provides the full agent infrastructure: Agent Runtime for session management and reasoning loops, Gateway (an MCP gateway with AWS Identity and Access Management (IAM) authentication and semantic search) for tool discovery and routing, and McpBearerToken for secure authentication to downstream MCP servers. An IAM role scopes the agent’s permissions.

MCP Server Lambda (tool layer)
The ClearanceJobs MCP Server Lambda function, deployed in a private subnet within a virtual private cloud (VPC), exposes system capabilities as discrete tools:

  • search_candidates performs candidate search with clearance and skills filtering.
  • get_candidate performs detailed profile retrieval.

The Lambda function connects to the ClearanceJobs pilot environment through a NAT gateway with a WAF-allowlisted egress IP address, making sure only authorized traffic reaches the production APIs. Credentials and base URLs are stored in AWS Systems Manager Parameter Store.

ProfileLookup Lambda (external enrichment)
A separate Lambda function (find_github_profile) enriches candidate data with external GitHub profiles, routed through an internet gateway to the GitHub Users API.

Foundation model (reasoning layer)
Anthropic’s Claude 3.5 Haiku in Amazon Bedrock provides the agent’s reasoning capabilities. It interprets recruiter intent, decomposes complex requests into tool calls, and synthesizes results into actionable responses.

CJRecruiterAgent memory (context layer)
AgentCore memory, a capability of Amazon Bedrock AgentCore, persists session state and recruiter preferences across conversations. This context lets the agent recall past searches, preferred candidates, and workflow patterns.

Security and networking
The architecture spans two AWS accounts:

  • AgileATS account houses the AgentCore components, the foundation model, and a Bedrock Adapter Lambda function that provides an alternate MCP JSON-RPC path for classic Amazon Bedrock agent integration.
  • ClearanceJobs account houses the MCP Server and ProfileLookup Lambda functions within a VPC (with private and public subnets), a NAT gateway for controlled egress, and Amazon CloudWatch Logs for structured observability.

Communication between AgentCore and the ClearanceJobs account uses MCP over HTTPS with bearer authentication and custom headers for tenant identification.

Results

The hackathon delivered measurable outcomes across multiple dimensions:

Technical acceleration

  • Teams delivered functioning agentic AI features using Amazon Bedrock AgentCore and MCP servers in three days, compressing what would typically take more than three months.
  • The teams validated a production-ready architecture during the hackathon itself, which reduced post-event rework.
  • Kiro served as more than a coding assistant, driving both new code creation and deep analysis of existing systems to accelerate development velocity.

Organizational transformation

  • Kiro usage across product and engineering teams increased 84% following the hackathon, with more unique daily users each week and adoption continuing to grow.
  • 33% of developers reported increased interest in the AI-enabled SDLC.
  • Delivery velocity rose across teams that fully adopted the AI-enabled software development lifecycle, marking a sustained step change rather than a short-term spike.
  • As the second successful hackathon with AWS, and with DHI leadership committing to make it an annual event, the engagement reflects a sustained, deepening partnership.
  • Kiro has become ClearanceJobs’ productivity tool of choice, with adoption expanding beyond developers to product managers. This accelerates product development and lets product managers self-serve on code base analysis and feature scoping.

“Participating for the second straight year as a judge, this hackathon only deepened my appreciation for the AWS team’s partnership, the ambition our teams brought, and what AI makes possible when you clear the runway. The problems they tackled were real, the solutions were creative, and the energy was contagious. It’s given us a fresh lens on how we build.”

– Alex Schildt, President of ClearanceJobs, DHI Group, Inc.

“Our second hackathon with AWS was even more successful than the first. We walked away with deeper confidence and more excitement about AI, all backed by hands-on experience with AWS’s latest capabilities. Post-hackathon, it’s been great to see our teams continue to lean into AI to accelerate how we ship. I think the hackathon was a real catalyst for that. I can’t wait to see these features get into the hands of our users.”

– Rose Fan, Sr. Director of Product, DHI Group, Inc.

Lessons learned: Making hackathons production-ready

Based on our experience hosting multiple hackathons with customers like DHI, here are key principles for hackathons that ship:

  1. Set production-grade success criteria upfront: Prototypes must be sprint-ready, not only demo-ready.
  2. Put decision-makers on the judging panel: Production go/no-go decisions happen on the final day of the hackathon, not weeks later.
  3. Invest in pre-enablement: Workshops before the event mean teams build on day 1 instead of spending it learning.
  4. Use cross-functional teams: Product, go-to-market (GTM), and subject matter experts (SMEs) alongside engineering make sure real business problems get solved.
  5. Build relationships: On-site AWS presence helps build relationships that accelerate delivery long after the event.
  6. Make it repeatable: DHI’s second hackathon planned faster and set higher expectations because the first one shipped to production.

Hackathons as a production accelerator

Hackathons are often dismissed as team-building exercises or limited to generating ideas that never ship. When structured correctly, they become a powerful production acceleration mechanism. Here’s why:

Time-boxed intensity drives decisions. A time-bound constraint (typically one to three days) forces teams to make architectural choices quickly, which alleviates analysis paralysis. Teams can’t over-engineer when the clock is ticking.

Cross-functional alignment happens naturally. When engineering, product, sales, and executives work side by side for several days, alignment that typically takes weeks of meetings happens organically.

Executive visibility de-risks production decisions: When leadership sees a working demo, not a slide deck, they can make go/no-go decisions with confidence. At DHI, the President and Head of Product & Engineering served as judges, giving them firsthand visibility into feasibility.

Real code beats theoretical architecture. Hackathon prototypes aren’t wireframes. They’re functioning applications built on production-grade services, making the path to production shorter and more predictable.

Conclusion

DHI Group’s experience across its annual hackathons shows that structured hackathons are one of the fastest paths from generative AI experimentation to deployed workloads. Their first hackathon shipped two features to production. Their second is on track to deliver three more, including an agentic AI system that unifies two systems through MCP servers and Amazon Bedrock AgentCore.

The takeaway is that hackathons aren’t only idea generators. They compress the entire innovation lifecycle (ideation, architecture, prototyping, executive alignment, and production planning) into a single high-intensity event. Paired with proper preparation and a clear path to production, they become a strategic tool for digital transformation and workforce enablement.

If your organization is looking to accelerate generative AI adoption, consider whether a structured hackathon could compress months of planning into days of building. To get started:

About the authors

How United Airlines uses Amazon Redshift and AWS Glue Data Catalog federation to query Databricks-managed data

Post Syndicated from Vaibhav Agrawal original https://aws.amazon.com/blogs/big-data/how-united-airlines-uses-amazon-redshift-and-aws-glue-data-catalog-federation-to-query-databricks-managed-data/

This post was co-written with Ankit Aggarwal and Raja Kalluri from United Airlines.

United Airlines processes billions of events daily across its data platform, which spans Amazon Redshift and Databricks with Unity Catalog. To bridge these platforms without duplicating data, the team turned to AWS Glue Data Catalog federation.

In this post, we walk through how to configure AWS Glue Data Catalog federation to connect with Databricks Unity Catalog, so you can run live SQL queries from Amazon Redshift without moving or duplicating data.

Why United Airlines needed catalog federation

United Airlines curates petabytes of data through a medallion architecture (bronze to silver to gold) on Amazon Simple Storage Service (Amazon S3). The airline user interaction data layer alone is several double-digit terabytes of near real-time streamed data. Teams use it to measure customer engagement patterns, feature adoption, and conversion behavior across web and mobile touchpoints. Analysts need to query this curated data through Amazon Redshift Serverless. As part of the existing data platform architecture these data tables are cataloged in Databricks Unity Catalog, not in the AWS Glue Data Catalog. As a result, Amazon Redshift has no native visibility into them. Without catalog federation, the only way to make this data queryable from Amazon Redshift would have been to duplicate it into Amazon Redshift Managed Storage (RMS) and build pipelines to keep it in sync.

AWS Glue Data Catalog federation removed this need. Amazon Redshift users now query the gold layer stored in Amazon S3 directly, with Iceberg metadata resolved from Unity Catalog at query time and no data movement. AWS Glue Data Catalog federation connects Amazon Redshift to external catalogs like Unity Catalog, so analysts query cross-platform data without building sync pipelines or duplicating storage.

Amazon Redshift Serverless is powered by the same Graviton-based query engine used in the new RG instance family, which delivers up to 2x faster data lake query performance compared to prior generations. This engine is purpose-built for reading Apache Iceberg tables directly from Amazon S3, making it well-suited for such federated query workloads.

United Airlines is taking a phased approach to adopting AWS Glue Data Catalog federation across its data platform. The initial focus is the most heavily used user interaction data tables, with 30 tables currently federated in production and 70 more in active rollout. Several hundred additional tables across different business domains are planned for production in the coming months.

Solution overview

AWS Glue Data Catalog federation bridges these platforms at the metadata layer. Here’s how the architecture works.

The architecture follows a four-layer federation chain:

  • Databricks Unity Catalog exposes tables through its Iceberg REST API endpoint. For Delta tables, you can turn on UniForm format to make them Iceberg compatible.
  • AWS Glue Data Catalog creates a federated catalog that connects to Databricks Unity Catalog, making metadata visible within AWS without data movement.
  • A resource link database in the default AWS Glue catalog acts as a bridge, pointing to the federated catalog database. This is required for Amazon Redshift compute.
  • Amazon Redshift Serverless references the resource link database through an external schema. When a query runs, Amazon Redshift traverses the link, calls AWS Glue Federation, and reads the Iceberg data through the Databricks Unity Catalog REST API. AWS Lake Formation governs permissions throughout this chain.

Key services or service features used in this solution:

Figure 1: Federation chain from Databricks Unity Catalog to Amazon Redshift Serverless through AWS Glue and Lake Formation

The architecture follows a six-step flow:

  1. A SQL analyst submits a query to Amazon Redshift Serverless.
  2. Amazon Redshift resolves the external schema through the AWS Glue Data Catalog (resource link to federated catalog).
  3. The AWS Glue federated catalog calls the Databricks Unity Catalog Iceberg REST API to retrieve current table metadata.
  4. The namespace IAM role calls AWS Lake Formation GetDataAccess to obtain scoped, temporary S3 credentials.
  5. Lake Formation evaluates fine-grained access policies and vends credentials for the authorized data files.
  6. Amazon Redshift Serverless reads the Iceberg data files directly from S3 and returns results to the analyst.

Prerequisites

Before you begin, make sure the following are in place:

  • A Databricks workspace with Unity Catalog enabled and at least one catalog, schema, and table. Databricks uses UniForm to generate Iceberg metadata on Delta Lake tables on Amazon S3.
  • An AWS account with permissions to manage AWS Glue, AWS Lake Formation, Amazon Redshift Serverless, and IAM.
  • An Amazon Redshift Serverless workgroup and namespace already provisioned.
  • AWS Lake Formation set up with a data lake administrator.
  • AWS Command Line Interface (AWS CLI) configured with appropriate credentials.
  • Familiarity with Amazon Redshift Query Editor v2 or a SQL client.

Note: For setting up the Databricks Unity Catalog side (Phase 1), follow the steps in the AWS blog post Access Databricks Unity Catalog data using catalog federation in the AWS Glue Data Catalog. This walkthrough picks up after the federated catalog has been created in AWS Glue.

Solution walkthrough

The walkthrough is organized into six steps covering Lake Formation configuration, the resource link pattern, IAM role setup, and querying Databricks tables from Amazon Redshift.

Step 1: Configure AWS Lake Formation

1a. Add a data lake administrator

  • In Lake Formation, choose Administration, then choose Administrators and add your admin IAM user or role.

1b. Confirm the federated catalog is registered

  • Choose Data Catalog, then Catalogs and verify that databricks-federated-catalog is visible and registered.

This step is the key architectural detail in the walkthrough. Amazon Redshift resolves CREATE EXTERNAL SCHEMA only against the default AWS Glue Data Catalog. The federated catalog (databricks-federated-catalog) is a separate, non-default catalog object. To give Amazon Redshift a path to the federated data, you create a resource link database in the default catalog that points to the federated catalog’s database.

A resource link does not copy data or metadata. It’s a pointer that Lake Formation resolves at query time.

To create the resource link in the Lake Formation console:

  • Choose Data Catalog, Databases, Create database. Then select Resource link.
  • For Resource link name, enter databricks_federated_db_link.
  • For Target catalog, enter databricks-federated-catalog.
  • For Target database, enter the database name that was discovered by the AWS Glue crawler (for example, databricks_federated_db).

Alternatively, use the AWS CLI:

aws glue create-database \
  --database-input '{
    "Name": "databricks_federated_db_link",
    "TargetDatabase": {
      "CatalogId": "<account-id>:databricks-federated-catalog",
      "DatabaseName": "databricks_federated_db"
    }
  }'

Step 3: Configure the Amazon Redshift Serverless namespace IAM role

When Amazon Redshift queries through the resource link, it uses the IAM role attached to the Amazon Redshift Serverless namespace to call the Lake Formation GetDataAccess API. Lake Formation permissions must be granted to this namespace role.

Choose one of these two approaches:

  • Option A – Update your existing namespace role by adding the following policy inline.
  • Option B – Create a new dedicated role (named RedshiftServerlessNamespaceRole) and attach it to the namespace alongside existing roles.

Attach the following IAM policy to the role:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "glue:GetDatabase",
        "glue:GetDatabases",
        "glue:GetTable",
        "glue:GetTables",
        "glue:GetPartitions",
        "glue:GetCatalog",
        "glue:GetCatalogs"
      ],
      "Resource": "*"
    },
    {
      "Effect": "Allow",
      "Action": "lakeformation:GetDataAccess",
      "Resource": "*"
    }
  ]
}

Note: The Resource: “*” in this policy is shown for simplicity. In production, scope resources to specific AWS Glue catalog ARNs, database ARNs, and table ARNs based on your use case.*

After creating or updating the role, associate it with your Amazon Redshift Serverless namespace:

  • In the Amazon Redshift Serverless console, choose Namespaces, select [your namespace], then choose Security and encryption, then Manage IAM roles.
  • If you use Option A, the existing role already has the new permissions, so no change is needed.
  • If you use Option B, add the new role alongside the existing roles.

Step 4: Grant Lake Formation permissions to the Amazon Redshift namespace role

4a. Grant DESCRIBE on the resource link database (default catalog)

  • In Lake Formation, choose Permissions, Data lake permissions, then Grant.
  • Principal: RedshiftServerlessNamespaceRole.
  • Resources: Named Data Catalog resources, Default catalog, databricks_federated_db_link (resouce link).
  • Database permissions: DESCRIBE.

4b. Grant SELECT and DESCRIBE on the target tables (Grant on Target)

Resource links permit only DESCRIBE and DROP permissions on the link itself. To allow Amazon Redshift to actually read data, you must separately grant SELECT on the target tables in the federated catalog. This is the Lake Formation Grant on Target pattern.

  • Principal: RedshiftServerlessNamespaceRole.
  • Resources: Named Data Catalog resources, databricks-federated-catalog, databricks_federated_db, then Tables.
  • Table permissions: SELECT, DESCRIBE.
  • Catalog permission: DESCRIBE.

Important: SELECT must be granted on the TARGET tables in the federated catalog, not on the resource link. Granting SELECT only on the resource link won’t work. This is a common configuration error.

Step 5: Create an external schema in Amazon Redshift

With the resource link in place and permissions granted, you can now create an external schema in Amazon Redshift that points to the resource link database. The external schema is the query interface. When a user runs SQL against it, Amazon Redshift traverses the link to the federated catalog and retrieves metadata and data from Databricks Unity Catalog.

The DATABASE parameter must reference the resource link database name in the default AWS Glue catalog (databricks_federated_db_link), not the federated catalog name directly. The CATALOG_ARN parameter isn’t required here because the resource link lives in the default catalog and Amazon Redshift resolves it automatically.

Connect to your Amazon Redshift cluster as a superuser (for example, using Amazon Redshift Query Editor v2) and run:

CREATE EXTERNAL SCHEMA databricks_schema
FROM DATA CATALOG
DATABASE 'databricks_federated_db_link'
IAM_ROLE '<iam-role-arn>'
REGION '<region>';

A key design principle in this architecture is the clear separation between data physically stored in Amazon Redshift and data accessed externally through federation. External schemas provide a transparent abstraction layer, so Amazon Redshift users can query data stored in S3 without ingestion. For consistency and clarity, United Airlines follows a standard naming convention for all federated schemas in Amazon Redshift: {domain}_iceberg. This convention makes it immediately clear that the data isn’t natively stored within Amazon Redshift but is accessed by using federation through AWS Glue and Lake Formation. This distinction is critical for analysts and engineers, because it improves discoverability, avoids ambiguity between storage layers, and reinforces architectural discipline when working across hybrid data environments.

The User Interactions domain exposes curated datasets representing customer interaction activity, engagement behavior, and channel usage patterns. Operational datasets follow the same pattern, providing governed access to supporting business events and reference information through a common federation framework.

You create a view layer over each external schema using WITH NO SCHEMA BINDING, so that analysts always resolve the freshest schema on each query execution. For example:

CREATE VIEW analytics.clickstream_events AS
SELECT * FROM {domain}_iceberg.interaction_events
WITH NO SCHEMA BINDING;

Step 6: Verify and query Databricks tables from Amazon Redshift

After creating the external schema, verify that the Databricks tables are visible and run a test query.

Verify table visibility

-- Confirm federated tables are visible in Redshift
SELECT * FROM SVV_EXTERNAL_TABLES
WHERE schemaname = 'databricks_schema';

Query a Databricks Unity Catalog table

-- Query a Databricks Unity Catalog table via the federated catalog
SELECT *
FROM databricks_schema.<table_name>
LIMIT 10;

When a query runs, Amazon Redshift calls Lake Formation GetDataAccess using the namespace IAM role to obtain temporary credentials. It then contacts the AWS Glue federated catalog, which in turn calls the Databricks Unity Catalog Iceberg REST API to retrieve metadata and read table data. The result is returned to the Amazon Redshift user transparently.

For SAML-authenticated users, connect using your IdP JDBC plugin:

jdbc:redshift:iam://<workgroup-name>.<account-id>.<region>.redshift-serverless.amazonaws.com:5439/<database>
?plugin_name=com.amazon.redshift.plugin.<YourIdPPlugin>
&idp_host=<your-idp-host>
&preferred_role=arn:aws:iam::<account-id>:role/RedshiftSAMLUserRole
&ssl=true

The Amazon Redshift JDBC driver handles authentication automatically. It authenticates with your IdP, receives a SAML assertion, and calls sts:AssumeRoleWithSAML for temporary IAM credentials. It then calls redshift-serverless:GetCredentials to connect as the mapped database user.

Business impact

AWS Glue Data Catalog federation delivered measurable architectural and operational improvements for United Airlines:

Area Before After Impact
Data access Delta Lake and Amazon Redshift data were completely siloed, so Amazon Redshift users had no access to curated datasets on Databricks-managed S3 data Amazon Redshift users get real-time access to Databricks-managed data through AWS Glue Data Catalog federation ~100 analysts gained access to user interaction data tables in the first phase without adding new pipelines.
Disaster recovery Cross-Region DR relied on Amazon Redshift snapshots every 3 hours (recovery point objective, or RPO, of 3 hours or more) Amazon S3 cross-Region replication on the Delta Lake provides a near-continuous RPO. A new Amazon Redshift Serverless workgroup in the DR Region can federate to the same S3 data More resilient architecture. Reduces cost for Amazon Redshift snapshot and copy maintenance across Regions
Architecture simplification Data processing happened in both Databricks and Amazon Redshift, requiring manual catalog synchronization between the two platforms which was operationally expensive and prone to drift With the federated architecture, data processing is consolidated in Databricks, and Amazon Redshift acts solely as a query engine powering user queries and dashboards through catalog federation Single processing platform, zero sync pipelines, single source of truth
Infrastructure cost Running dedicated Amazon Redshift ETL cluster with RMS storage, snapshots, and compute for data processing For this use case with federation, Amazon Redshift is not needed for ETL but only as a query engine. No RMS storage duplication, no snapshot replication required ~$30K/month in redundant ETL infrastructure cost reduced

Security considerations

At United Airlines, identity governance is unified through Azure Active Directory groups. On the AWS consumption side, users authenticate to Amazon Redshift Serverless through SAML federation. AD group membership determines database-level access to federated schemas. On the Databricks side, the same AD groups govern access to Unity Catalog schemas. This single-identity model provides consistent access control across both platforms without requiring separate user provisioning. Lake Formation handles credential vending for S3 data access during federated queries, while schema-level access decisions are managed through the AD group mappings on each platform.

The architecture also provides multiple layers of security controls built into the federation chain:

  • AWS Lake Formation governs fine-grained access control throughout the federation chain, so that principals can only access authorized databases, tables, and columns.
  • IAM roles follow least-privilege principles. The Amazon Redshift namespace role is scoped only to AWS Glue metadata operations and Lake Formation GetDataAccess.
  • SAML-based authentication integrates enterprise identity providers, so that users authenticate through existing SSO infrastructure before accessing federated data.
  • All Amazon Redshift connections enforce TLS encryption (ssl=true), protecting data in transit between clients and the Amazon Redshift endpoint.
  • Lake Formation permission vending issues short-lived, scoped credentials for each query execution rather than long-lived static credentials.

Other considerations

Review the catalog federation service limitations before deploying. Key requirements:

  • Delta Lake tables must have UniForm enabled to expose Iceberg-compatible metadata.
  • We recommend that source tables be well-partitioned and regularly compacted, because the federated query performance reflects how efficiently the data is organized at write time.

Clean up

To avoid ongoing charges for resources created in this walkthrough, remove them in the following order. This teardown doesn’t affect Databricks metadata or your underlying data stored in Amazon S3.

  • Drop the external schema in Amazon Redshift: DROP SCHEMA databricks_schema;.
  • Delete the resource link database in the default AWS Glue catalog (databricks_federated_db_link).
  • Revoke Lake Formation permissions granted to the Amazon Redshift namespace role on both the resource link database and the target tables in the federated catalog.
  • Delete the federated catalog in AWS Glue (databricks-federated-catalog).
  • Deregister the AWS Glue connection for the Databricks Unity Catalog if no longer needed.
  • Optionally, remove the IAM role (RedshiftServerlessNamespaceRole) if it was created solely for this walkthrough.

Conclusion

In this post, we showed how United Airlines uses AWS Glue Data Catalog federation to give Amazon Redshift Serverless analysts real-time access to double-digit terabytes of curated user interaction data on Amazon S3, without duplicating a single byte or building sync pipelines.

The architecture uses the Iceberg REST API, resource link databases, and Lake Formation credential vending to create a governed query path between Amazon Redshift and Unity Catalog. For United Airlines, this eliminated redundant ETL infrastructure costs, removed the need for catalog synchronization, and turned Amazon Redshift Serverless into a dedicated high-performance query engine for analysts and dashboards.

For questions or feedback, leave a comment on this post.


About the authors

Vaibhav Agrawal

Vaibhav Agrawal

Vaibhav Agrawal is a Senior Analytics Specialist Solutions Architect at AWS, focused on helping enterprise customers design and implement modern data architectures using AWS Analytics services.

Ankit Aggarwal

Ankit Aggarwal

Ankit Aggarwal is a Principal Enterprise Architect at United Airlines, where he leads the United Data Hub (UDH) platform architecture—a petabyte-scale data platform built on AWS and Databricks. He brings over 15 years of experience in data engineering and enterprise architecture.

Raja Kalluri

Raja Kalluri is a Principal Architect at United Airlines, where he leads enterprise-scale data architecture and modernization initiatives. He specializes in building cloud-native data platforms, enabling real-time analytics and AI, and transforming legacy ecosystems.

How Sony LIV built real-time video streaming analytics with AWS

Post Syndicated from Rahul Sureka original https://aws.amazon.com/blogs/big-data/how-sony-liv-built-real-time-video-streaming-analytics-with-aws/

This guest post was co-written with Mukund Acharya from Sony LIV.

Real-time data analytics is transforming how streaming applications understand and serve their audiences. In this post, we share how Sony LIV used Amazon Kinesis Data Streams for sub-second processing and AWS services to build a comprehensive streaming analytics solution on AWS. In this post, we share how SonyLIV built a comprehensive streaming analytics solution on AWS using Amazon Kinesis Data Streams for sub-second event ingestion, Amazon Data Firehose for reliable delivery to storage, Amazon EMR with Apache Spark for scalable batch and micro-batch processing, and Apache Iceberg on Amazon S3 for ACID-compliant, queryable data lake tables. Together, these services enable SonyLIV to capture millions of concurrent viewer events, process them cost-efficiently at scale, and surface actionable insights from real-time engagement metrics to historical trend analysis — all within a fully managed, serverless-friendly architecture.

About Sony LIV

Sony LIV, operated by Sony Pictures Networks India, is one of India’s leading over-the-top (OTT) streaming applications offering premium content including live sports, original series, movies, and TV shows to millions of users across mobile, web, and Smart TV devices.

As part of the latest game broadcast rights for Asia Cup, Sony LIV built a streaming analytics solution using AWS services that successfully processed millions of concurrent sessions in real-time.

The challenge of batch processing

As Sony LIV’s audience grew and live sports events attracted increasingly large viewership, the team identified an opportunity to move from batch-based analytics to a real-time data application. The existing architecture processed data reliably in scheduled batches, but the growing scale of live events called for faster, more granular insights. The team set out to address three key areas:

  1. Real-Time Visibility: Track frontend user events, journeys, and conversion funnels in near real-time, especially during high-stakes live sports streams, to see peak concurrent viewership and respond to engagement patterns as they happen.
  2. Comprehensive Quality Monitoring: Implement end-to-end monitoring of video quality key performance indicators (KPIs) including buffering rates, playback failures, video start time, and rebuffering events to proactively enhance the viewer experience.
  3. Unified Data Architecture: Consolidate data from multiple sources into a unified architecture to enable comprehensive customer views, supporting future personalization and machine learning (ML)-driven recommendations.

Sony LIV needed to evolve from reactive batch processing to a real-time data application to deliver actionable engagement insights at scale and establish the foundation for unified customer profiles.

Streaming analytics architecture

Streaming analytics architecture showing data from mobile, web, and Smart TV apps flowing through Amazon EKS into real-time and batch processing paths on AWS

Figure 1: Sony LIV streaming analytics architecture on AWS

Data from mobile, web, and Smart TV applications flows through Application Load Balancers to Amazon Elastic Kubernetes Service (Amazon EKS) pods for validation and preprocessing. The architecture implements parallel processing paths to balance speed and cost efficiency.

Real Time Path: Priority events requiring immediate action such as playback failures, user interactions, and live sports engagement signals stream through Amazon Kinesis Data Streams for sub-second processing. These events flow directly to ClickHouse using the ClickHouse connector, enabling real-time analytics with minimal latency.

Batch Path: High-volume batch events are ingested via Amazon Data Firehose into a raw landing zone on Amazon Simple Storage Service (Amazon S3), partitioned by event type and time. Amazon EMR running Apache Spark then processes this raw data performing schema validation, deduplication, and transformations — and writes the curated output as Apache Iceberg tables back to S3

Unified Data Layer: AWS Glue catalogs data across streaming and batch sources, creating a unified schema that powers the customer data application. This unified data layer serves as the foundation for building comprehensive customer profiles and enabling personalized recommendations.

Analytics and Monitoring: ClickHouse serves as the low-latency analytics engine, powering real-time dashboards and video quality monitoring. Operations teams track critical KPIs including peak concurrent viewership, buffering rates, video start time, and playback failures, enabling rapid response to issues during live events.

Custom dashboards and Datadog provide visualization for business metrics and application performance monitoring, while Amazon CloudWatch tracks infrastructure health, delivering end-to-end visibility across the application.

Results and business impact

The results: product, analytics and engineering teams now operate from a single unified dashboard, making real-time decisions with data that is seconds, not hours old. Auto scaling and serverless design have optimized costs while providing reliability at peak load. At the foundation, a centralized data catalog now serves as a single source of truth for business intelligence, machine learning workloads and operational monitoring.

The transformation helped Sony LIV process millions of concurrent sessions in real time, particularly during high-profile events such as the Asia Cup 2025.

  1. Accelerated insights: Data latency dropped from hours to seconds, enabling instant insights into viewer behavior, content performance, and streaming health.
  2. Unified decision-making: Product, analytics, and engineering teams now use unified dashboards for real-time decision-making.
  3. Optimized operations: The solution auto scaling and serverless design optimized costs while supporting reliability and fault tolerance.

Beyond performance gains, the new architecture established a strong foundation for personalized content recommendation and operational agility. By integrating a centralized data catalog and scalable analytics layer, the solution now provides a single source of truth for business intelligence, machine learning workloads, and monitoring.

Future innovations

Sony LIV plans to expand its analytics capabilities by integrating advanced ML models for real-time recommendations, user churn prediction, and anomaly detection. The team will also focus on building unified customer profiles to enable hyper-personalized experiences.

The AWS based architecture provides a strong foundation for future growth and innovation, enabling Sony LIV to deliver highly personalized experiences to an expanding viewer base.

Conclusion

By adopting AWS, Sony LIV transformed its analytics architecture into a real-time, scalable, and insight-driven solution. The solution reduced data latency from hours to seconds, enabled processing of millions of concurrent sessions during peak events, and positioned Sony LIV as a leader in streaming analytics innovation in India.

To learn more about how AWS can help your media organization implement real-time streaming analytics solutions, explore the following resources:


About the authors

Rahul Sureka

Rahul Sureka

Rahul is an Enterprise Solutions Architect at AWS, helping media and cross-industry customers design scalable, real-time streaming architectures on the cloud. With over 25 years of experience architecting and leading large-scale business transformation programs, Rahul specializes in streaming applications, data and analytics, AI/ML, and Agentic AI.

Umesh Chaudhari

Umesh Chaudhari

Umesh is a Sr. Analytics Solutions Architect at AWS, helping Energy, Industrial, and Semiconductor customers shape their data and AI strategies. With over a decade of experience in distributed data systems and analytics Services, he turns complex data challenges into scalable, business-driving solutions. A recognized public speaker, he’s passionate about enabling innovation through modern, cloud-native architecture.

Maheshwaran G

Maheshwaran is a Principal Solution Architect in Media and Entertainment, dedicated to empowering Indian and SAARC media organizations to accelerate growth through cutting-edge cloud technologies that reimagine workflows, enhance scalability, and unlock new business opportunities — anchored in innovation with 17 granted patents spanning USPTO and IPO across diversified fields.

Varsha Palepu

Varsha Palepu

Varsha is a Solutions Architect at AWS and an analytics specialist on the AWS streaming team. She helps small and medium businesses innovate on AWS and creates technical streaming content to empower customers in their cloud journey.

How Moovit achieved 33% cost optimization through architectural modernization

Post Syndicated from Saar Porat original https://aws.amazon.com/blogs/big-data/how-moovit-achieved-33-cost-optimization-through-architectural-modernization/

Moovit, part of Mobileye (Nasdaq: MBLY), is a leading Mobility-as-a-Service (MaaS) solutions provider and the creator of a leading urban mobility app. Moovit’s iOS, Android, and web apps offer users a smart mobility experience to get to their destination using any mode of public and shared transportation. Transit riders can benefit from mobile ticketing to plan, pay, and ride with transit services. Introduced in 2012, Moovit now serves over 1.7 billion users in more than 3,500 cities across 112 countries, in 45 languages.

Behind these user-facing experiences is a data platform that processes large volumes of mobility, application, and operational data to support product analytics, business intelligence (BI), monitoring, and data science. As the platform grew, Moovit needed to keep analytical workloads reliable and cost-efficient without slowing down teams that depend on fresh data every day.

Over several years, Moovit’s Amazon Redshift cluster grew continuously. It started with an expanding fleet of DC2 nodes, migrated to RA3 nodes, and scaled multiple times to keep pace with growing data demands, ultimately becoming the backbone of their entire data platform.

To address this growth, Moovit transformed their data architecture by building an optimal multi-engine lakehouse architecture and assigning each workload to the most suitable option. This modernization reduced their Amazon Redshift cluster by 50 percent, while establishing a flexible, multi-engine architecture ready for future use cases.

In this post, we share how Moovit gained visibility into workload patterns, cleaned up unnecessary load, selected candidates for offloading, and ran a successful proof of concept (POC) on Amazon EMR Serverless. Moovit ultimately divided the workload between multiple engines, building a modern and cost-optimized data platform that combines provisioned Amazon Redshift, Amazon Redshift Serverless, and Amazon EMR.

The challenge: Outgrowing a single-engine data platform

The Amazon Redshift engine handled a wide variety of workloads, including:

  • Heavy ETL processing: Raw data ingestion from Amazon Simple Storage Service (Amazon S3) followed by complex aggregation pipelines (daily user-aggregation running once per day with a 3-day lookback, and weekly 10-day-lookback jobs).
  • Near-real-time operational monitoring: Queries executing every 20 minutes against raw data for system-health dashboards.
  • Business-intelligence reporting: Tableau extracts and live dashboards.
  • Data-science workloads: Exploratory analysis and model-feature engineering.
  • Ad-hoc analysis: Non-recurring queries done by analysts and engineers.

With business growth, storage grew by orders of magnitude over the past decade as the platform expanded. All these varied workloads competed for the same engine and pushed it to its limits. Jobs experienced increasing queue times, service level agreements (SLAs) were at risk, and adding nodes provided minimal performance gains, creating a need to isolate workloads.

Gaining visibility: Measuring workload impact

Moovit’s first modernization milestone was to create a trusted measurement foundation before changing any workloads. Instead of treating warehouse activity as a single opaque stream, the team implemented automated query attribution that continuously classified each query by workload owner and execution context. The classification combined multiple signals: who executed the query (user or service account), recognizable query-signature patterns, and metadata emitted by orchestration frameworks and scheduled processes.

This produced a historical, query-level map of platform usage that answered three critical questions: who is generating load, what kind of workload is running, and how expensive each workload is in runtime and resource terms. With that baseline in place, the team made offload decisions from evidence rather than assumptions. This approach prioritized the largest and most stable optimization opportunities first and reduced the risk of moving business-critical workloads without visibility.

These classifications and workload metrics were reflected in a Tableau report that aggregated query activity by classification label and execution context. The view exposed operational dimensions such as classification, time granularity, service class, execution-time bucket, unload flags, and sample-query context, supporting both trend monitoring and root-cause drill-down.

The worksheet was parameterized to support multiple measurement modes over the same grouped workload population: total execution time, execution plus queue time, total CPU time, average execution time per query, and ratio-based efficiency views (execution/CPU and CPU/execution). This let the team compare “heavy by volume” workloads against “inefficient by behavior” workloads without creating separate artifacts.

For decision-making, CPU time was used as the primary impact metric because it best represented sustained compute pressure. Execution time, queue time, query-count normalization, and workload-management segmentation were treated as secondary evidence to distinguish:

  • compute-heavy but healthy workloads
  • queue-constrained workloads
  • high-frequency/low-cost workloads
  • noisy or weakly classified workloads that required attribution cleanup first

Using this framework, prioritization became systematic: first improve classification coverage, then rank workloads by CPU contribution, then validate with queue and workload management (WLM) signals, and finally choose the action path per workload (optimize SQL, reschedule, isolate, retire, or move to another engine).

The following figure shows an example of one of the dashboard widgets (CPU time by query).

Dashboard widget showing CPU time consumed by each query

Figure 1: CPU time by query, highlighting the most resource-intensive queries and their usage patterns

Cleanup: Reducing unnecessary data warehouse load

With a long-running data platform, in most cases the workloads will start accumulating, some of which become irrelevant at some point. For example, a report which was created and scheduled, yet it became irrelevant after a few years, but still running since no one disabled it. It’s important to indicate these workloads in general to reduce unnecessary load, yet even more critical before doing any significant architectural changes or migrations. Before migrating any workloads, Moovit first reduced unnecessary warehouse load.

The team:

  • Removed unused processes that were still consuming cluster resources.
  • Reduced unnecessary frequency where possible: some jobs ran more often than downstream consumers needed.
  • Reviewed workload-management guardrails to verify resource allocation matched actual priorities.

This cleanup phase was a prerequisite to migration. By removing waste first, the team verified that the workloads eventually selected for offloading were genuinely heavy rather than simply unoptimized or unnecessary.

The no-longer-relevant processes consumed around 7 percent of overall CPU time and were removed before the optimization work began.

Workload selection: Choosing what to offload

With a clear picture of workload patterns, Moovit faced a common decision point: continue scaling the existing Redshift cluster, or re-architect towards a multi-engine approach. The team evaluated two main paths:

  1. Re-architect with Redshift multi-cluster and data sharing: Identify workloads that could benefit from resource isolation, then redistribute processing and queries between multiple Redshift clusters, combining both serverless and provisioned options. This would redistribute load across use-case-optimized clusters and potentially save costs through better resource use.
  2. Re-architect with purpose-built engines: Identify workloads that could benefit from alternative processing frameworks and offload them to more suitable engines. This would reduce pressure on Amazon Redshift while building a more flexible, cost-efficient architecture.

Moovit decided to do both, because while some workloads benefited from being offloaded, others benefited from isolated Amazon Redshift compute.

The measurement data revealed a primary candidate for offloading: raw-data aggregation pipelines. This workload loaded raw data into Amazon Redshift from Amazon S3, then performed heavy sessionization and aggregation transformations. Raw tables were still used for ad-hoc and exploratory analysis, but recurring production consumers primarily depended on aggregated outputs, making these transformations strong candidates for offloading.

Proof of concept: Offloading to EMR Serverless with Spark SQL

With target workload identified, Moovit initiated a POC using Amazon EMR Serverless with Spark SQL. The choice of EMR Serverless was driven by several factors:

  • Spark SQL compatibility: The existing Redshift SQL logic could be ported with minimal changes to Spark SQL syntax.
  • Serverless simplicity: No cluster-management overhead during the evaluation phase.
  • Data-lake native: Processing could occur directly on data in Amazon S3.

The POC defined quantified success criteria measured over five or more consecutive runs:

  • Runtime reduction: Greater than or equal to 40 percent reduction for the transform portion of selected pipelines.
  • Amazon Redshift cost reduction: Greater than 30 percent reduction in Redshift RA3 compute with no performance degradation for remaining workloads.
  • Data-quality parity: Exact match between Spark and Amazon Redshift outputs on row counts, distinct users, and all published metrics over a frozen parity window.

Overcoming initial performance challenges

The first POC attempts exposed significant challenges. Early Spark jobs with 100 executors took approximately 4 hours, far exceeding the 30–40-minute baseline on Amazon Redshift. Beyond raw performance, the team encountered memory pressure, data-parity gaps between Spark and Amazon Redshift outputs, and subtle SQL behavior differences between the two engines.

The team systematically diagnosed and resolved these issues:

  1. Execution-plan analysis: Reviewing the Spark execution plan revealed suboptimal query patterns that generated excessive data shuffles.
  2. Query rewrites: Rewriting specific SQL constructs to align with Spark’s distributed processing model, including splitting large monolithic logic into staged transformations.
  3. Reducing or rewriting expensive DISTINCT patterns: Identifying and eliminating unnecessary DISTINCT operations that created heavy shuffle pressure.

After applying these optimizations, execution time dropped from 4 hours to approximately 10 minutes, and the required executors dropped to fewer than 50, surpassing the original performance.

Validation: Ensuring data parity before cutover

Before transitioning any workload to production, Moovit implemented a rigorous validation process. The new Spark output was compared with the previous Amazon Redshift output using multiple dimensions:

  • Row counts: ensuring no data was lost or duplicated.
  • Distinct users: verifying entity-level completeness.
  • Metric parity: all published business metrics matched.
  • Daily trends: time-series patterns remained consistent.
  • Row-level checks: spot-checking individual records for correctness.

Only after all validation checks passed consistently over multiple consecutive runs did the team proceed with cutover for each workload.

Moving to production: Expanding workload offloading

With a successful POC demonstrating both performance gains and cost savings, Moovit progressively moved additional workloads from Amazon Redshift to EMR:

  • Heavy-aggregation jobs: The primary daily and weekly aggregation pipelines transitioned fully to EMR.
  • Data-transformation stages: Preprocessing steps that previously consumed Redshift compute moved to Spark, with only final aggregated results loaded back into Amazon Redshift for BI consumption.
  • Weekly batch workloads: Large batch jobs that previously created resource contention during weekend processing windows.

The transition used a measured approach: each workload was migrated individually, with data-quality validation confirming parity before decommissioning the equivalent jobs which were running on Redshift.

Additional optimizations: Redshift Serverless, workload isolation, and Amazon EMR on Amazon EC2

Beyond EMR offloading, Moovit implemented further architectural improvements to isolate workloads and optimize costs.

Amazon Redshift rightsizing: Iterative cluster optimization

With heavy workloads successfully offloaded and isolated, Moovit proceeded to right-size the Redshift cluster. Rather than a single resize, the team reduced the cluster incrementally, two nodes at a time, using elastic resize. At each step, they validated that:

  • Existing BI workloads maintained acceptable performance.
  • Queue wait times remained within SLA thresholds.
  • No workload degradation was observed under peak loads.

This iterative approach minimized risk and allowed the team to find the optimal cluster size with confidence.

Workload isolation with Redshift Serverless

Amazon Redshift persisted as the engine of choice for serving curated BI data. However, not all Amazon Redshift workloads needed provisioned capacity:

  • Ad-hoc analyst queries: Moved to Redshift Serverless, isolating unpredictable workloads from the provisioned cluster through data sharing.
  • Data-science workloads: Transitioned to Redshift Serverless for flexible exploration without impacting production.

This workload isolation through Redshift Serverless provided resource separation without requiring additional provisioned capacity. The architecture now used data sharing to provide a unified view across provisioned and serverless clusters.

Operational isolation refinements

Moovit also refined workload isolation by rebalancing WLM priorities on the provisioned cluster. Because the ETL queue mainly handled raw data loading from Amazon S3 (which was not the bottleneck after heavy aggregations moved to Spark), its priority was reduced. At the same time, with most human users moved to Redshift Serverless, Tableau serving workloads on provisioned Redshift were prioritized higher to keep dashboard performance predictable. The final result: a 50% reduction in provisioned Redshift capacity.

Transitioning to EMR on EC2

EMR Serverless proved efficient for the POC phase: it allowed fast iteration without cluster management overhead. However, for longer-term recurring production workloads, Moovit moved to EMR on EC2 to better fit their production cost and infrastructure model, using existing compute reservations.

The transition between EMR deployment options required zero application code changes, demonstrating the flexibility of the EMR deployment options.

AI-assisted SQL translation

Additionally, Moovit used AI-assisted development tools, Claude Code and Cursor, to accelerate parts of the SQL transition process. These tools helped engineers identify Redshift SQL and Spark SQL syntax differences, suggest rewrites, and debug migration issues, while validation and production approval remained under engineer review.

Results: A modern multi-engine architecture

The architectural modernization delivered measurable outcomes:

  • Cluster size reduction: Redshift cluster size reduced to 50 percent of the initial capacity.
  • Performance improvement: Key aggregation jobs ran faster and more consistently on EMR (50 percent execution time reduction for p90).
  • Workload isolation: No single workload type could impact others through resource contention.
  • 33 percent overall data pipeline cost reduction: Combined savings from cluster reduction, transition to EMR, and efficient serverless usage.
  • Future flexibility: The multi-engine architecture provided pathways for additional use cases without architectural changes.

The following figures compare aggregation-job performance before and after the transition.

Chart comparing aggregation-job execution times before and after the transition, with longer, inconsistent runtimes before and shorter, stable runtimes after

Figure 2: Aggregation-job execution times before and after the transition

Chart comparing wall-clock time for job executions across percentiles, with p90 at 5.48 hours before the transition and 2.77 hours after

Figure 3: Wall-clock time for job executions by percentile, before and after the transition

The resulting architecture assigned each workload to the engine that fits it best:

Workload type Engine Rationale
Heavy ETL and aggregation Amazon EMR (Spark SQL) Distributed processing on Amazon S3. No data warehouse load required
Ongoing processing and BI reporting Amazon Redshift provisioned 24/7 running processes
Ad-hoc queries Amazon Redshift Serverless Burst capacity with workload isolation
Data science Amazon Redshift Serverless Flexible exploration without impacting production

Lessons learned

The Moovit modernization journey produced several key insights applicable to similar architectural transitions:

  1. Measure before you move: Establishing baseline metrics and automated classification was essential for identifying true offloading candidates. Without granular workload-level measurements, the team would not have identified which specific processes were exhausting the cluster.
  2. Clean up before you migrate: Reducing unnecessary load first verified that migration efforts targeted genuinely heavy workloads rather than simply unoptimized or unused processes.
  3. Small SQL changes, big impact: Moving from Redshift SQL to Spark SQL required relatively minor syntax adjustments. The core business logic remained intact, and most transformations translated directly with minimal refactoring.
  4. Optimize for the engine: Porting SQL queries to Spark without optimization produced initially poor results for some workloads. Understanding Spark’s distributed execution model and optimizing for it was critical for achieving target performance.
  5. Validate rigorously: Multi-dimensional data-parity checks (row counts, distinct users, metrics, daily trends, and row-level spot checks) gave the team confidence to cut over without data-quality regressions.
  6. Moving between EMR options is straightforward: EMR Serverless proved very efficient for starting fast and evaluating Spark. When Moovit needed to move to EMR on EC2 to use existing reservations, the transition required no application code changes.
  7. Iterative cluster rightsizing: Rather than a single resize, Moovit reduced the Redshift cluster incrementally (two nodes at a time) using elastic resize, validating performance at each step before proceeding further.

Conclusion

Looking ahead, as another potential optimization, Moovit will be evaluating the new Amazon Redshift RG instances for provisioned clusters, providing up to 2.2x better price performance and priced 30% lower than RA3, powered by AWS Graviton.

The broader takeaway is that AWS provides multiple purpose-built engines that can be used in a single data platform. In Moovit’s case, the biggest improvement came from assigning each workload to the engine that fit it best: Amazon Redshift for curated analytical serving, Redshift Serverless for isolated exploratory workloads, and Amazon EMR for large-scale transformations over data in Amazon S3. This architecture gives Moovit a foundation for future optimization and flexibility as data volumes grow and new analytical use cases emerge.

 


About the authors

Saar Porat

Saar Porat

Saar is the Director of BI & Data Engineering at Moovit, where he has spent more than a decade building and scaling the company’s data engineering capabilities. With nearly 20 years of experience in BI, analytics, and data platforms, he focuses on designing reliable, maintainable, and cost-efficient systems that translate complex data into meaningful business impact. Saar led Moovit’s initiative to migrate major workloads from Amazon Redshift to Apache Spark, improving scalability, performance, and infrastructure efficiency while expanding the team’s engineering capabilities beyond SQL-based processing.

Vova Nevski

Vova Nevski

Vova is a Senior Analytics Specialist Solutions Architect at AWS with more than 15 years of experience in the big data and analytics domain, including data lakes, batch and stream processing, both on premises and in the cloud. He partners with AWS customers to design and build solutions best suited to their unique needs.

Razor Group’s journey to a modern data lakehouse on AWS

Post Syndicated from Yaswanth Kothainti original https://aws.amazon.com/blogs/big-data/razor-groups-journey-to-a-modern-data-lakehouse-on-aws/

Razor Group is one of Europe’s leading ecommerce aggregators, operating 250+ brands across multiple global marketplaces. With a portfolio exceeding $400M in revenue, the company relies on data to power every critical business decision, from dynamic pricing and inventory optimization to advertising spend and supply chain orchestration.

At the heart of this operation sits the Razor Operating System (ROS), a proprietary platform that processes 370M+ API calls monthly through 9,300+ data pipelines, transforming marketplace signals into automated actions at scale.

In this post, we share how Razor Group optimized their data platform by implementing a lakehouse architecture on AWS. We cover the architectural decisions, the phased migration approach, and the measurable business outcomes. Whether you’re looking to optimize workload performance, reduce infrastructure costs, or unlock multi-engine flexibility for your analytics, this blueprint provides actionable insights you can adapt for your organization.

The business challenge: Scaling data infrastructure for hypergrowth

As Razor Group’s brand portfolio expanded rapidly, the demands on their data platform grew significantly. The company needed their analytics infrastructure to keep pace with the speed of ecommerce, where pricing decisions, stock replenishment, and advertising bids happen in near real time.

Their existing architecture, built on Amazon Redshift provisioned clusters, had served them well during earlier growth stages. As workloads diversified and data volumes surged, several optimization opportunities emerged:

Razor Operating System data architecture before the migration: signal sources such as Amazon Selling Partner API, Shopify, NetSuite, Walmart, and Target ingested through AWS Lambda and Amazon MSK, stored in Amazon S3 and Amazon DynamoDB, modeled in Amazon Redshift, and consumed by ML notebooks, ML jobs on AWS Batch, and Tableau dashboards, orchestrated by Apache Airflow

Figure 1: The Razor Operating System data architecture before the migration

  • Workload contention: Over 1,000 SQL models for ETL, transformation, and analytics competed for the same compute resources, creating resource contention during peak processing windows.
  • Cost-to-utilization mismatch: Always-on clusters ran 24/7, but workload analysis revealed that 98% of compute demand came from batch ETL rather than interactive analytics, which resulted in significant idle capacity during off-peak hours.
  • Data freshness gaps: Batch-oriented pipelines delivered data with 4–6 hour latency, limiting the team’s ability to react to fast-moving marketplace dynamics.
  • Scaling constraints: As concurrent users and pipeline complexity grew, vertical scaling alone couldn’t address the need for workload isolation and elastic capacity.

These weren’t failures of any single service. They were signals that the architecture needed to evolve to match the scale and diversity of Razor Group’s workloads.

Why a lakehouse architecture?

Rather than replacing their existing investments, Razor Group recognized the opportunity to optimize workload placement by adopting a modern lakehouse architecture. The core principles driving this decision:

  • Open table formats: Apache Iceberg provides ACID transactions, time travel, and schema evolution. Data is stored once and accessed by any compatible engine without duplication.
  • Elastic, per-workload scaling: With data persisted on Amazon Simple Storage Service (Amazon S3), each engine independently scales compute to match its workload. Each engine spins up for peak processing and scales to zero when idle, without over-provisioning shared infrastructure.
  • Multi-engine flexibility: Different workloads have different requirements. Heavy ETL benefits from distributed Spark processing, ad hoc exploration from serverless queries, and business intelligence (BI) dashboards from high-performance warehouse engines, each optimized for its purpose.

This approach allowed Razor Group to right-size each workload to the best-fit engine while maintaining a single, governed copy of data accessible across the entire platform.

Solution overview

Razor Group partnered with AWS to implement a comprehensive lakehouse architecture that brings together multiple AWS services, each playing a complementary role:

New lakehouse architecture on AWS: the same signal sources ingested through AWS Lambda and Amazon MSK, stored and modeled as Bronze, Silver, and Gold Apache Iceberg tables using Apache Spark Connect on Amazon EC2 with AWS Lake Formation and AWS Glue Data Catalog, served through Amazon Redshift, and consumed by ML notebooks, ML jobs on AWS Batch, and Tableau dashboards

Figure 2: End-to-end lakehouse architecture on AWS

Designing for scale: The lakehouse vision

The core insight driving Razor Group’s new architecture was simple: build a single, open format data lake that any engine can query. In the old model, each tool maintained its own copy of the data. In the new model, a single open-format data lake on Amazon S3 serves as the source of truth, and multiple purpose-built compute engines read from it based on the workload at hand.

This shift, commonly called a lakehouse architecture, combines the cost economics and scalability of a data lake with the query performance and governance of a data warehouse. Its open table format, Apache Iceberg, provides ACID transactions, schema evolution, time travel, and no vendor lock-in.

Storage and governance: The open data foundation

  • Amazon S3 Tables (a capability of Amazon S3) with Apache Iceberg — The primary storage layer, providing open-format tables with ACID transactions, partition evolution, and time travel. Data is stored once and accessible by any Iceberg-compatible engine.
  • AWS Glue Data Catalog — A unified metadata repository for consistent data discovery across all compute engines.
  • AWS Lake Formation — Fine-grained access control with column-level and row-level security so that governance scales with the platform.

Compute: Right engine for the right workload

  • Apache Spark on Amazon Elastic Compute Cloud (Amazon EC2) — Elastic, distributed compute for heavy ETL and transformation workloads. It uses AWS Graviton instances and Amazon EC2 Spot Instances for cost optimization.
  • Amazon Athena — Serverless SQL for ad hoc exploration and lightweight queries directly on Iceberg tables, with no infrastructure to manage.
  • Amazon Redshift Serverless — High-performance serving layer for BI dashboards, Tableau workloads, and interactive analytics. Amazon Redshift Serverless automatically scales to meet demand and pauses when idle, so it stays cost-efficient for the analytics workloads it serves best.

Orchestration and observability

  • Apache Airflow — Pipeline orchestration that manages 9,300+ data pipelines with dependency tracking and service level agreement (SLA) monitoring.
  • Comprehensive observability stack — Cost attribution, pipeline health monitoring, and data quality checks across all layers.

Note: When the architecture was originally designed, Amazon Redshift lacked Iceberg write support, making self-managed Spark the only viable ingestion path. This constraint has since been removed. Amazon Redshift now supports full Apache Iceberg DML (UPDATE, DELETE, MERGE), complementing its earlier CREATE/INSERT capabilities and AWS Glue Iceberg materialized views. This makes it a complete read/write Iceberg engine.

Migration approach

Rather than a risky big-bang cutover, Razor Group adopted a phased migration of five stages, each delivering standalone value while building the foundation for the next. Both Amazon Redshift and Spark pipelines ran in parallel during the transition, which maintained business continuity and let the team compare outputs with confidence. At no point was a production pipeline paused or a dashboard unavailable.

The migration journey: Five phases

The migration unfolded across five structured phases, each building on the previous one and delivering incremental value before the next began.

Phase 1: Establish the lakehouse foundation

Before migrating a single query, Razor Group needed to answer three questions: where does the data live, how is it managed, and how do we query it?

Why S3 Tables over self-managed Iceberg

Razor Group had already committed to Apache Iceberg as the table format: open, engine-agnostic, and equipped with ACID transactions and time travel. The question was whether to self-manage Iceberg on standard S3 buckets or use Amazon S3 Tables.

Self-managed Iceberg is powerful but operationally expensive. Someone has to run compaction jobs to prevent small-file proliferation. Someone has to expire old snapshots before metadata bloat degrades query planning. Someone has to clean up orphaned data files after interrupted writes. With 700+ models running across 40+ schemas, many of them materializing multiple times per day, that maintenance burden would scale with the platform rather than shrink.

S3 Tables eliminated this entire category of work. Compaction, snapshot management, and unreferenced file removal run continuously and automatically. The integrated Iceberg REST Catalog API means any compatible engine, such as Spark, Trino, Athena, Amazon Redshift, and Flink, can discover and query tables without maintaining a separate metastore. Discovery is unified through AWS Glue Data Catalog, which now exposes the Iceberg REST Catalog protocol as its access interface. Because tables are first-class AWS resources, access control, encryption, and lifecycle policies operate at the table level rather than through complex S3 bucket policies layered on top of file-path conventions.

For a company that didn’t want the operational burden of self-managing open table format maintenance, this was the deciding factor.

AWS Glue Data Catalog provides unified metadata discovery across all tiers. Lake Formation handles column- and table-level access control, with AWS Identity and Access Management (IAM) roles that follow least-privilege principles and AWS CloudTrail turned on for a full audit trail.

Choosing the query protocol

Prior to the rearchitecture, the Amazon Redshift cluster was 98% ETL, and only a fraction of compute hours were analyst SELECT queries. The replacement engine needed to handle both heavy batch transformations and interactive ad hoc queries.

Traditional Spark (spark-submit) handles batch ETL well, but couples clients to the cluster. Every job requires packaging driver JARs, managing classpaths, and submitting from within the cluster. For a platform running 200+ production directed acyclic graphs (DAGs) that process massive data volumes daily, this operational friction was a non-starter.

Spark Connect is the gRPC-based client-server protocol introduced in Spark 3.4, and it solved the coupling problem entirely. The cluster runs a persistent gRPC endpoint. Clients connect remotely and submit queries over the wire. Airflow operators become thin clients: they open a session, submit SQL, and get results, with success and failure mapping directly to task states. There are no driver JARs and no polling. Multiple consumers, including pipeline orchestrators, the web application, and developer notebooks, share one cluster without any of them needing Spark installed locally.

Deploying Spark Connect

Razor Group deployed a self-hosted Spark cluster on Amazon EC2: an on-demand AWS Graviton leader node, Spot workers at about 70% cost savings, and the Spark Connect endpoint exposed through an internal Network Load Balancer. Custom Amazon Machine Images (AMIs) bake in the full Spark, Iceberg, and S3 Tables stack, so private-subnet nodes have everything they need without internet access at runtime.

This phase produced no immediate business value, but it made everything that followed possible.

Phase 2: Migrate data ingestion

Razor Group’s ingestion layer pulls data from Amazon Selling Partner API, Seller Central portals, NetSuite ERP, and custom web scrapers. In the previous architecture, all of this landed in Amazon Redshift through COPY commands, which meant data freshness was dictated by batch job schedules and competed for resources on the same cluster that served analytical queries.

Razor Group migrated these pipelines to AWS Lambda functions orchestrated by Apache Airflow, writing data directly to S3 Tables in Iceberg format. The shift from schedule-driven to event-driven significantly improved freshness. Lambda functions spin up only when there’s data to process, and Airflow sensors trigger downstream transformations the moment new data lands. This replaced rigid hourly batch windows with data freshness measured in minutes.

The orchestration layer manages 200+ DAGs across 90+ flows and processes data from dozens of sources at scale. The migration required rewiring destinations from Amazon Redshift COPY to Iceberg writes, but the orchestration logic itself carried over with minimal changes.

This phase alone eliminated roughly 40% of compute costs by severing the always-on cluster dependency for ingestion.

Phase 3: Transform processing pipelines

This was the most technically demanding phase, and where Razor Group learned the most. The team migrated 1,000+ SQL models from Amazon Redshift to Apache Spark, working incrementally up the dependency chain across 40+ schemas. The models moved through a medallion structure: Bronze for raw ingested data, Silver for cleaned and conformed data, and Gold for business-ready aggregates.

Razor Group built automated conversion tooling and a validation framework that ran both Amazon Redshift and Spark outputs in parallel, comparing results row-by-row before decommissioning anything. Several categories of transformation pushed the limits of what automation could handle:

  • Window functions: The QUALIFY clause in Amazon Redshift has no Spark equivalent. Each instance required wrapping in a subquery with explicit row numbering, which affected dozens of models in the inventory schema alone.
  • JSON serialization: The most time-consuming category. Complex columns stored as JSON STRING in Amazon Redshift needed from_json() with hand-written STRUCT definitions in Spark. Every nested payload column across ads, orders, and transaction pipelines required schema introspection, with no shortcuts.
  • Function dialect: More than 20 function-level conversions, including NVL to COALESCE, DATEADD to interval arithmetic, and LISTAGG to ARRAY_JOIN(COLLECT_LIST()).
  • Snapshot elimination: The single biggest hidden cost. Full table copies that ran multiple times daily only to preserve point-in-time state consumed more than 35 hours of weekly Amazon Redshift compute. With Iceberg’s native time travel, these became zero-cost operations overnight.

When migrating 1,000+ SQL models, automated tooling handles the mechanical syntax conversions well. But roughly 30% of the models required human judgment: those with complex JSON payloads, deeply nested window functions, or cross-schema snapshot dependencies. These models consumed 70% of the migration effort.

Razor Group built a structured migration workflow that used Claude to accelerate this work: read source SQL, identify dependencies, convert syntax, resolve missing base tables, add JSON parsing, validate outputs, and write to the lakehouse. The system did more than translate SQL. It applied schema context, traced cross-model dependencies, and flagged edge cases that would have taken engineers hours to find manually. What could have been a multi-year effort became a systematic, repeatable process measured in weeks. This approach fundamentally changed the speed of migration.

Phase 4: Unify the serving layer

With data flowing through Iceberg tables, Razor Group collapsed the serving layer. End users query Gold-layer Iceberg tables through Amazon Redshift Serverless, and internal exploration and machine learning (ML) workloads read the same tables through Spark Connect. This removed the need to maintain separate data copies, materialized views, or extract jobs for different consumers.

This is the strategic payoff of an open table format. Iceberg tables on S3 are engine-agnostic: Spark for batch transforms today, Trino for interactive queries tomorrow, Flink for streaming next quarter. Any engine that speaks Iceberg can read the data without conversion or migration. Razor Group went from being locked into a single vendor’s SQL dialect to having the freedom to adopt new engines without touching the storage layer.

Phase 5: Operationalize and observe

The final phase made the lakehouse production-grade. Razor Group built a comprehensive observability stack that aggregates metrics, traces, and logs from every pipeline component into a unified view. This view supports centralized log search, anomaly detection, and automated alerting that correlates failures across the entire data platform.

This observability layer did more than provide visibility. It gave the team confidence. When you’re running thousands of pipeline executions daily, you need to know within minutes when something breaks, what caused it, and which downstream consumers are affected. That’s the difference between reactive firefighting and proactive operations.

Pipeline orchestration consolidated around three patterns: a daily pipeline (ingestion to materialization to export to AI agent analysis), an operations worker polling every 15 minutes, and weekly scraper jobs.

The cutover was zero-downtime by design: both schedulers ran in parallel for two weeks. Automated comparison checks validated that every pipeline produced identical outputs before the prior architecture system was disabled.

Results and business impact

The lakehouse architecture delivered measurable improvements across every dimension:

Metric Before After Improvement
P95 query runtime 180 seconds 63 seconds 65% faster
Infrastructure cost Always-on provisioned clusters Elastic, workload-optimized 63% reduction
Data freshness 4–6 hour batch cycles Event-driven pipelines 15-minute freshness
Concurrent capacity Limited by cluster size Elastic, independent scaling Unlimited
Engine flexibility Single engine Multi-engine (Spark, Athena, Amazon Redshift) Open format portability

The 63% reduction compares the lakehouse run-rate (January–March 2026) with the pre-rearchitecture run-rate (October–December 2025), the trailing three months before the rearchitecture. The figure is an apples-to-apples blended infrastructure number that includes compute and storage across both architectures. The before column covers Amazon Redshift cluster compute and managed storage. The after column covers Amazon EC2 (Spark workers, both on-demand and Spot), AWS Lambda, AWS Glue, Amazon Athena, Amazon Redshift Serverless, and S3 Tables storage. Data-transfer and ancillary services are excluded because they were not materially different between the two periods. Workload mix (the number of pipelines, models, and end-user query volume) was held broadly comparable across the two windows.

Lessons learned

Start with the decision loops, not the tools, and know your workload before you replace your warehouse.

The most valuable activity of the entire migration wasn’t writing a line of code. It was the Amazon Redshift workload analysis we ran before making any architectural decisions. Discovering that 98% of compute was ETL, with only a sliver going to analyst queries, validated the move to on-demand Spark. It also prevented us from over-provisioning the replacement infrastructure for interactive workloads that barely existed. Architecture decisions should always trace back to core business requirements: pricing accuracy, promotional responsiveness, intraday P&L visibility. Start there, not with the technology.

Design for multiple compute engines, and choose the right engine per workload.

One of the clearest lessons from running a single-engine architecture is what you give up. Avoid locking yourself into one compute layer for BI, ingestion, backfills, and ML alike, because they have fundamentally different cost and performance profiles. Iceberg, Spark, and S3 Tables work well together out of the box once you make the shift. The technology isn’t the hard part. The hard part is mapping 1,000+ models across 40+ schemas, tracing dependencies through 200+ DAGs, and discovering that a column is actually a JSON string silently serialized differently between two engines. Migration is as much an excavation project as an engineering one.

Automate conversion, but budget for the 30%.

Automated tooling handles mechanical syntax conversions well, and it should be the first tool you reach for. But models with complex JSON payloads, deeply nested window functions, or cross-schema snapshot dependencies require human judgment, and that work doesn’t compress. Roughly 30% of our models needed significant manual intervention, and those models consumed 70% of the total migration effort. Plan for it honestly from the start.

Observability must include cost attribution, and watch out for hidden cost bombs.

Snapshot operations were our biggest surprise. Full table copies that ran multiple times daily to preserve point-in-time state were costing more than 35 hours of weekly compute, and nobody questioned it because “that’s how snapshots work.” Iceberg’s time-travel capability eliminated their cost, and that single feature justified a meaningful portion of the migration on its own. More broadly, you cannot optimize what you cannot see, so track query-level usage and attribute it to teams and functions. Cost observability is not a nice-to-have. It’s foundational.

Governance isn’t optional. Build it into the foundation, and align stakeholders from day one.

Catalog and access control need to come first, before you scale adoption, not after. The same principle applies to people: migration is a cross-functional program, not an infrastructure project. Our two-week parallel run caught edge cases that row-level validation missed entirely: time zone differences between Amazon Redshift and Spark, partition pruning behavior under concurrent writes, and subtle ordering differences in non-deterministic window functions. That parallel run wasn’t a safety net. It was where the migration actually proved itself. None of it works without the right stakeholders involved and aligned from the very beginning.

Conclusion

Razor Group’s journey offers valuable lessons for organizations looking to optimize their data architectures:

  1. Analyze your workload mix first. Understanding that 98% of compute was ETL rather than interactive queries guided the decision to offload heavy processing to elastic Spark, while preserving Amazon Redshift Serverless for the interactive analytics it handles best.
  2. Design for multi-engine flexibility. Open table formats like Apache Iceberg eliminate the need to choose a single engine. Each workload runs on the engine best suited to its access pattern, cost profile, and performance requirements.
  3. Automate migration, but budget for complexity. Automated transpilation handled 70% of SQL models, but the remaining 30% consumed 70% of engineering effort. Plan accordingly.
  4. Observability must include cost attribution. Without per-workload cost visibility, optimization is guesswork. Razor Group discovered that Iceberg snapshot maintenance alone consumed more than 35 hours of compute weekly, a hidden cost that observability surfaced and automation resolved.
  5. Build governance into the foundation. AWS Lake Formation and AWS Glue Data Catalog provided fine-grained access control from day one, not retrofitted after the migration.
  6. Validate with parallel systems. A two-week parallel run between old and new architectures caught edge cases that automated testing missed, which supported a confident production cutover.

The road ahead

With the lakehouse foundation in place, Razor Group is positioned to accelerate innovation, from real-time pricing models to AI-driven inventory optimization, all powered by a unified, open, and governed data platform on AWS.

The company’s transformation demonstrates that modern data architectures aren’t about choosing between services. They’re about placing each workload where it performs best, using open formats to eliminate silos, and scaling each layer independently as the business grows.

To learn how other organizations are implementing similar lakehouse architectures on AWS, see How BigBasket uses the Iceberg-based lakehouse architecture on AWS to power lightning-fast grocery delivery across India.


About the authors

Yaswanth Kothainti

Yaswanth is VP of Data Engineering & Platform at Razor Group, a $400M+ ecommerce enterprise, where he built the company’s data platform from the ground up and leads a 65-member global engineering organization. His core expertise spans enterprise data platforms, data governance, FinOps, and agentic AI systems, with a track record of translating complex platform investments into measurable business outcomes.

Shubham Purwar

Shubham Purwar

Shubham is an Analytics Specialist Solutions Architect at AWS. He helps organizations unlock the full potential of their data by designing and implementing scalable, secure, and high-performance analytics solutions on AWS. In his free time, Shubham loves to spend time with his family and travel around the world.

Ravi Kompella

Ravi Kompella

Ravi is Principal Analytics Specialist with experience in driving adoption of modern data architectures, enterprise data lakehouses, and real-time data systems across multiple industry verticals in India across all segments including startups and SaaS providers.

How Picnic configured multiple OAuth providers for Amazon MQ

Post Syndicated from Oscar Mapfumo Sibanda original https://aws.amazon.com/blogs/big-data/how-picnic-configured-multiple-oauth-providers-for-amazon-mq/

This post is co-written with Oscar Mapfumo Sibanda from Picnic.

Picnic is an Amsterdam-based tech scale-up that reinvents how people buy food. It isn’t a supermarket with a digital layer but a tech company that happens to deliver groceries. Picnic is engineered in-house: the customer app, the fulfillment platform, the supply chain, and the routing technology that guides a fleet of thousands of electric vehicles through the Netherlands, Germany, and France. Software doesn’t merely support business. Software is a business.

At the center of this system, RabbitMQ is the core component. It’s the communication backbone connecting hundreds of microservices across the entire business lifecycle, from ordering and logistics to delivery and finance. At peak, Picnic’s platform processes close to one million messages per second. At this scale, messaging is no longer only background infrastructure. It becomes part of the company’s operational nervous system. To keep that system highly available, scalable, and resilient as Picnic grows, the company decided to use Amazon MQ as a managed service.

The next challenge was identity. Picnic’s authentication strategy clearly distinguishes between people and services. Operators sign in through Keycloak, the company’s single sign-on provider, while Picnic’s Amazon Elastic Kubernetes Service (Amazon EKS) workloads are adopting AWS Identity and Access Management (IAM) authentication to eliminate static credentials. A single broker therefore must trust two identity providers at once. The Amazon MQ documentation covers configuring OAuth 2.0 with a single provider. This post extends that guidance to a multi-provider setup on the same broker.

In this post, we show how Picnic solved that problem. You will learn how to configure an Amazon MQ for RabbitMQ broker to accept tokens from multiple OAuth 2.0 identity providers, using Keycloak and IAM as the working example. You will also see how to map each provider’s scopes to RabbitMQ permissions and how to roll the change out on a running broker without disrupting connected users.

Background and prerequisites

Amazon MQ for RabbitMQ supports OAuth 2.0 authentication and authorization, where broker users and their permissions are managed by an external identity provider. User authentication and resource permissions for vhosts, exchanges, queues, and topics are centralized through the OAuth 2.0 provider’s scope system.

RabbitMQ’s OAuth 2.0 plugin supports multiple resource servers and audiences, allowing different OAuth 2.0 providers to issue tokens that a single broker can validate. This capability is essential if you operate in multiple environments or have teams registered with separate identity providers.

Prerequisites

To follow along with this post, you need:

  1. An active AWS account.
  2. An Amazon MQ for RabbitMQ broker with OAuth 2.0 configured for at least one identity provider (see Using OAuth 2.0 authentication and authorization for Amazon MQ for RabbitMQ).
  3. A second OAuth 2.0 identity provider configured and operational.
  4. Outbound web identity federation enabled for your AWS account (if using IAM as a provider).
  5. Basic familiarity with RabbitMQ configuration and OAuth 2.0 concepts.
  6. AWS Command Line Interface (AWS CLI) version 2.27 or later (required for the get-web-identity-token command used in the testing section).

Note: The information in this post reflects Amazon MQ for RabbitMQ features and behavior at the time of publication. We recommend checking the Amazon MQ documentation, release notes and best practices before implementation.

Solution architecture

The design rests on a single idea: a RabbitMQ broker can trust more than one identity provider at the same time, and it decides which one to apply per token rather than per broker. RabbitMQ does this by reading the aud (audience) claim of each incoming token and matching it against a configured resource server. Each resource server is bound to one OAuth 2.0 provider, so the audience determines both which signing keys validate the token and which permission rules apply.

Architecture diagram

The following two diagrams show how services and operators authenticate the broker through their respective identity providers.

Services authentication flow from an Amazon EKS workload through AWS STS to the Amazon MQ for RabbitMQ broker over AMQPS

Figure 1: Services (IAM) flow

Services authenticate through IAM. An Amazon EKS workload assumes an IAM role and generates a web identity token (1), which AWS Security Token Service (AWS STS) issues with an audience of rabbitmq-iam (2). The workload presents that token as its password when it connects to the broker over Advanced Message Queuing Protocol (AMQPS) on port 5671 (3). The broker selects the matching resource server, verifies the token’s signature against the AWS STS signing keys (4), and maps the role’s Amazon Resource Name (ARN) to the permissions the workload needs (5).

Operators authentication flow from the RabbitMQ management console through Keycloak single sign-on to the broker

Figure 2: Operators (Keycloak) flow

Operators authenticate through Keycloak. An operator opens the RabbitMQ management console and initiates login (1), and the console redirects the browser to Keycloak (2). Keycloak authenticates the operator and issues a token whose audience targets the console resource server, rabbitmq-keycloak (3). The browser presents that token to the broker (4), which verifies the signature against Keycloak’s signing keys (5). The broker then reads the operator’s group membership and grants access (6): the Operator group receives read-only permissions, while the Administrator group receives full control.

Two constraints follow this design. First, audience verification is a single broker-wide setting that applies to every provider at once, so each provider must issue tokens carrying the exact audience its resource server expects. Second, the broker is private, deployed inside an Amazon Virtual Private Cloud (Amazon VPC) with no public exposure. Both providers’ endpoints must be resolvable, either through publicly addressable JSON Web Key Set (JWKS) endpoints or through private networking, because the broker fetches signing keys from those endpoints.

Implementation walkthrough

This walkthrough configures one Amazon MQ for RabbitMQ broker to trust two identity providers: Keycloak for operators and AWS IAM for services. The steps assume you already have a running broker, a Keycloak realm, and outbound web identity federation enabled for your AWS account. All configurations are applied through a RabbitMQ configuration revision using the AWS Command Line Interface (AWS CLI).

Enable OAuth 2.0 on the broker

The first block activates the OAuth 2.0 backend and keeps the internal backend in place. Internal authentication remains active deliberately: Amazon MQ creates an administrator user when the broker is provisioned, and that user is needed for break-glass access.

auth_backends.1 = oauth2
auth_backends.2 = internal
auth_oauth2.verify_aud = true

Setting verify_aud = true tells RabbitMQ to reject any token whose aud claim does not match a configured resource server. This single broker-wide setting governs every provider you add.

Add the first identity provider (Keycloak)

A resource server binds an audience value to a provider and a set of permission rules. The Keycloak resource server uses the id rabbitmq-keycloak, which is the audience the realm must place in its tokens. RabbitMQ reads the operator’s group membership from the group_membership claim and resolves it through scope aliases.

auth_oauth2.resource_servers.1.id = rabbitmq-keycloak
auth_oauth2.resource_servers.1.oauth_provider_id = keycloak
auth_oauth2.resource_servers.1.scope_prefix = rabbitmq.
auth_oauth2.resource_servers.1.additional_scopes_key = group_membership
auth_oauth2.resource_servers.1.preferred_username_claims.1 = email
auth_oauth2.resource_servers.1.scope_aliases.1.alias = Operator
auth_oauth2.resource_servers.1.scope_aliases.1.scope = rabbitmq.read:*/* rabbitmq.write:^$ rabbitmq.configure:^$ rabbitmq.tag:monitoring
auth_oauth2.resource_servers.1.scope_aliases.2.alias = Administrator
auth_oauth2.resource_servers.1.scope_aliases.2.scope = rabbitmq.read:*/* rabbitmq.write:*/* rabbitmq.configure:*/* rabbitmq.tag:administrator

The Operator group is read-only: it can read any resource and view the management UI through the monitoring tag. The Administrator group receives full permissions plus the administrator tag. This role is least privilege by design.

The provider is configured with its issuer and JWKS endpoint:

auth_oauth2.oauth_providers.keycloak.https.hostname_verification = wildcard
auth_oauth2.oauth_providers.keycloak.issuer = https://keycloak.example.com/auth/realms/test
auth_oauth2.oauth_providers.keycloak.jwks_uri = https://keycloak.example.com/auth/realms/test/protocol/openid-connect/certs

To let operators sign in from the management console, expose Keycloak as a management resource:

management.oauth_enabled = true
management.oauth_disable_basic_auth = false
management.oauth_scopes = openid email profile
management.oauth_resource_servers.1.id = rabbitmq-keycloak
management.oauth_resource_servers.1.oauth_client_id = rabbitmq-keycloak
management.oauth_resource_servers.1.label = Keycloak SSO

Add the second identity provider (AWS IAM)

Adding a second provider means adding a second resource server and a second entry under oauth_providers. The IAM resource server uses the id rabbitmq-iam because the audience is set when minting the token.

auth_oauth2.resource_servers.2.id = rabbitmq-iam
auth_oauth2.resource_servers.2.oauth_provider_id = aws_iam
auth_oauth2.resource_servers.2.scope_prefix = rabbitmq/
auth_oauth2.resource_servers.2.additional_scopes_key = sub
auth_oauth2.resource_servers.2.scope_aliases.1.alias = arn:aws:iam::123456789012:role/EKSWorkloadRole
auth_oauth2.resource_servers.2.scope_aliases.1.scope = rabbitmq/read:*/* rabbitmq/write:*/* rabbitmq/configure:*/* rabbitmq/tag:policymaker
auth_oauth2.oauth_providers.aws_iam.https.hostname_verification = wildcard
auth_oauth2.oauth_providers.aws_iam.issuer =
auth_oauth2.oauth_providers.aws_iam.jwks_uri =

The IAM workload receives the policymaker tag. It can publish, consume, and manage policies but does not receive the administrator tag.

Apply the configuration and restart the broker:

CONFIG_ID=$(aws mq describe-broker --broker-id $BROKER_ID \
  --query 'Configurations.Current.Id' --output text)
REVISION=$(aws mq update-configuration --configuration-id $CONFIG_ID \
  --data "$(cat rabbitmq.conf | base64 | tr -d '\n')" \
  --query 'LatestRevision.Revision' --output text)
aws mq update-broker --broker-id $BROKER_ID \
  --configuration Id=$CONFIG_ID,Revision=$REVISION
aws mq reboot-broker --broker-id $BROKER_ID

Note: The base64 command syntax differs between Linux and macOS. The preceding command (cat file | base64 | tr -d '\n') is portable on both operating systems. If running exclusively on Linux, you can also use base64 --wrap=0 rabbitmq.conf. On macOS, the equivalent command is base64 -i rabbitmq.conf.

Testing and validation

Validate each provider independently. For IAM, assume the role and request a token from AWS STS, then present it as the AMQP password:

TOKEN=$(aws sts get-web-identity-token \
  --audience "rabbitmq-iam" \
  --signing-algorithm ES384 \
  --duration-seconds 300 \
  --query 'WebIdentityToken' --output text)
# Username is empty (ignored by the OAuth plugin); the token is passed as the password
curl -u ":$TOKEN" https://<broker-id>.mq.<region>.on.aws/api/overview

Note: The get-web-identity-token API requires outbound web identity federation to be enabled on your AWS account and AWS CLI version 2.27 or later.

A successful response confirms the IAM resource server accepted the token. For Keycloak, open the management console, choose Keycloak SSO, and sign in as an operator.

When a login fails, decode the JSON Web Token (JWT) and check two claims. The aud claim must exactly match a resource server id. With verify_aud = true, a missing or mismatched audience is the most common cause of rejection. If the audience is correct but permissions are missing, verify the scope_prefix is set correctly.

Operational considerations

A few points deserve attention before you run this pattern in production.

  1. Key rotation: When rotating signing keys at a provider, publish the new key in the JWKS endpoint before revoking the old one. The broker caches keys, so overlapping both during the transition window prevents authentication failures while the cache refreshes.
  2. Audience validation: Audience remains the linchpin. With verify_aud enabled, every provider must issue tokens carrying the audience its resource server expects, so confirm this whenever you onboard a new one. Don’t disable audience validation in production. The RabbitMQ OAuth 2.0 plugin does not perform token revocation checks, which makes audience binding a critical control that prevents tokens issued for other services from granting access.
  3. Scope prefix: scope_prefix values are optional. They’re needed only if the tokens don’t follow the default format. RabbitMQ only reads scopes carrying the expected prefix, so a token can authenticate yet grant nothing if the prefix is missing. Map each provider to the least privilege its principals need. For example, prefer narrow scopes like read:orders over blanket read:all to limit the scope of impact if a single provider’s credentials are compromised.
  4. Token lifetime: Because the plugin does not support token revocation, token lifetime is your primary control over leaked credentials. Issue short-lived access tokens and have your client applications refresh them proactively at approximately 75 percent of the token’s lifetime to avoid connection disruptions when a token expires mid-session.
  5. Monitoring: Authentication failures and refused tokens are recorded in the broker’s connection log group in Amazon CloudWatch, which can be reached through the Amazon CloudWatch Logs link on the broker’s page in the Amazon MQ console. Beyond logs, set up CloudWatch alarms on RabbitMQMemUsed, RabbitMQDiskFree, and ConnectionCount. An unexpected spike in failed connections is often the first sign of a token or audience misconfiguration. For unaggregated, per-node visibility, consider enabling the Prometheus metrics endpoint: metrics such as rabbitmq_auth_attempts_failed_total surface OAuth rejections faster than the CloudWatch one-minute polling interval.
  6. Network controls: Enforce defense in depth by restricting broker access using security groups so that only authorized VPCs and IP ranges can reach the AMQPS and management endpoints. This matters especially in an OAuth setup because, once a token has been issued, the broker cannot revoke it before it expires.

Cleanup

To avoid incurring future costs, delete the resources created during this walkthrough if you no longer need them:

  1. Delete Amazon MQ broker and configurations.
  2. Remove test OAuth application registrations from your identity providers.
  3. Delete any IAM roles created for testing.

Conclusion

In this post, we demonstrated how Picnic configured an Amazon MQ for RabbitMQ broker to authenticate tokens from two OAuth 2.0 identity providers: Keycloak for human operators and AWS IAM for machine-to-machine services on a single broker instance. The key mechanism is RabbitMQ’s support for multiple resource servers, where the audience claim in each token determines which provider’s signing keys and permission rules apply.

With this approach, the Picnic team was able to cleanly separate human and machine authentication without the operational overhead of running separate brokers, while retaining fine-grained access control for both token issuers.

This pattern works with any combination of OAuth 2.0 providers and is particularly valuable for organizations looking to consolidate messaging infrastructure while maintaining distinct identity boundaries.

To learn more about Amazon MQ for RabbitMQ and OAuth 2.0 authentication, see Authentication and authorization for Amazon MQ. For a hands-on walkthrough of configuring OAuth 2.0 with Amazon MQ for RabbitMQ, see Using OAuth 2.0 authentication and authorization for Amazon MQ for RabbitMQ. The configuration examples in this post are broker-level settings applied through the Amazon MQ API. No standalone code repository is required.


About the authors

Oscar Mapfumo Sibanda

Oscar Mapfumo Sibanda

Oscar is a Senior Site Reliability Engineer at Picnic Technologies in the Netherlands. He builds infrastructure that supports rapid scaling, empowers engineering teams to move independently, and strengthens the security posture across the organization. Outside of work he paints and takes photographs; he is a technology enthusiast in the pursuit of happiness.

Ayush Kumar

Ayush Kumar

Ayush is a Technical Account Manager at Amazon Web Services based in the Netherlands. He works with enterprise customers to optimize their cloud architectures and accelerate innovation on AWS. You’ll find him experimenting in the kitchen in his spare time.

Amit Singh

Amit Singh

Amit is a Senior Solutions Architect at AWS, working with enterprise retail customers in the Benelux region. He helps customers design cloud-native architectures, navigate complex modernization journeys, and adopt AI/ML capabilities at scale. Outside of work, he enjoys exploring new places and chasing the perfect shot, whether through a camera lens or on a running trail.

Gallup scales real-time coaching for thousands with Amazon Bedrock

Post Syndicated from Tamil Sambasivam original https://aws.amazon.com/blogs/architecture/gallup-delivers-real-time-workplace-coaching-to-thousands-of-leaders-with-amazon-bedrock/

How Gallup turned 90 years of workplace science into an AI assistant that gives leaders personalized guidance in seconds, powered by Amazon Bedrock.

Gallup delivers analytics and advice to help leaders and organizations solve their most pressing problems. With more than 90 years of experience and a global reach, Gallup has developed a uniquely deep understanding of workplace behavior and performance.

However, this knowledge wasn’t centralized or delivered in context. Leaders had to navigate multiple resources to find relevant insights and then translate them into action without guidance. The lack of real-time, personalized recommendations meant workplace challenges were often handled reactively instead of proactively.

Gallup needed to transform decades of proprietary research into real-time, personalized guidance that leaders can access instantly within their existing workflow.

In this post, we show how Gallup built Gallup AI, a generative AI assistant powered by Amazon Bedrock. It transforms decades of proprietary workplace research into real-time, personalized coaching delivered directly within the Gallup Access application.

Why Amazon Bedrock

Gallup evaluated multiple approaches to building a generative AI assistant. The team chose Amazon Bedrock for three reasons:

  1. Access to leading foundation models like Anthropic’s Claude without managing infrastructure.
  2. Built-in retrieval augmented generation (RAG) through Amazon Bedrock Knowledge Bases, a fully managed RAG capability, grounds responses in verified research.
  3. Native guardrails to enforce content safety at scale.

This combination allowed Gallup to move from prototype to production in weeks rather than months, without hiring a dedicated machine learning (ML) operations team.

Note: Anthropic’s Claude models on Amazon Bedrock are available in select AWS Regions. For current model and Region availability, see Supported models by Region in Amazon Bedrock.

The approach: Building intelligence into daily workflow

Gallup built Gallup AI, a generative AI-powered assistant integrated directly into Gallup Access. The unified application lets managers review engagement results, build action plans, explore CliftonStrengths insights, and access curated content to better support their teams.

The solution uses Amazon Bedrock with Anthropic’s Claude models to deliver conversational insights grounded in Gallup’s proprietary research. Amazon Bedrock Knowledge Bases and Amazon Kendra retrieve relevant research and organizational data. This is designed to ground responses in verified workplace science. Amazon Bedrock Guardrails enforce content safety policies, while AWS Lambda with FastAPI delivers real-time streaming responses that feel natural and immediate.

The architecture follows a serverless design and supports multiple organizations simultaneously. Amazon ElastiCache Serverless provides sub-millisecond response times for conversation history. Amazon Relational Database Service (Amazon RDS) for MySQL serves as the durable system of record. Amazon Data Firehose streams usage metrics to Amazon Simple Storage Service (Amazon S3) for cost management and performance optimization.

How the solution works

The following diagram shows the Gallup Access AI application architecture.

Architecture diagram of the Gallup Access AI application showing request flow through AWS Lambda to Amazon Bedrock, with Amazon Bedrock Knowledge Bases and Amazon Kendra for retrieval, Amazon ElastiCache Serverless and Amazon RDS for storage, and Amazon Data Firehose streaming metrics to Amazon S3

Figure 1: Gallup Access AI application architecture

The architecture processes requests through the following stages:

Gallup’s proprietary workplace research covers decades of employee engagement studies, performance data, and organizational insights. The content is stored in Amazon S3 and ingested into Amazon Bedrock Knowledge Bases. The application also continuously crawls the Gallup website to capture the latest research publications, articles, and insights, indexing this content in Amazon Kendra for instant retrieval. This dual approach gives the AI assistant access to both historical research archives and current workplace science, delivering responses grounded in verified, up-to-date knowledge rather than generic advice.

When a leader asks Gallup AI a question, the system retrieves relevant research from both Amazon Bedrock Knowledge Bases and Amazon Kendra. The system scores documents based on confidence thresholds, filters them, and consolidates them before sending them to Claude models in Amazon Bedrock.

The conversation flows through AWS Lambda handlers that manage both real-time streaming (for web clients) and synchronous requests (for backend services). Amazon ElastiCache Serverless caches recent conversation history for instant retrieval, while Amazon RDS for MySQL serves as the durable storage layer with organized records of conversations, prompts, responses, and source citations.

Amazon Bedrock Guardrails apply content safety policies during generation, with the ability to intervene mid-stream if policy violations are detected. Interactions persist before streaming begins, preserving transactional integrity even if connections are interrupted.

AWS Systems Manager Parameter Store serves as the application’s centralized configuration hub, managing AI model settings, content safety policies, and performance thresholds. This allows the team to adjust application behavior instantly, without redeploying code or interrupting service for users.

Amazon DynamoDB provides fast, flexible storage for product-specific insights and contextual data, so the application delivers personalized experiences tailored to each user’s role and workflow.

Comprehensive metrics, including I/O tokens, cached tokens, time-to-first byte, and stop reasons, flow through Amazon Data Firehose to Amazon S3, providing visibility into cost, performance, and usage patterns across the application.

What Gallup has achieved

Gallup has transformed decades of workplace research into an intelligent assistant that delivers measurable value across thousands of organizations. Tasks that previously required navigating reports, articles, and tools now resolve through a single conversational interaction. Time to insight dropped from manual research to real-time, AI-delivered guidance within seconds. The application processes billions of tokens through production interactions, with responses grounded in verified workplace science.

Since launching in June 2024, adoption and engagement have grown rapidly:

Metric Result
Prompts Increased ~7x
Conversations Increased ~4.5x
Active users Increased ~5.5x
Engagement depth Average prompts per conversation increased ~55%, indicating sustained, multi-turn interactions
Response latency Sub-second time-to-first byte (TTFB) for streaming responses. Sub-millisecond session retrieval via Amazon ElastiCache Serverless

What the customer said

Gallup’s Director of Product reflects on what this shift means for how leaders access workplace science:

“Gallup AI represents a fundamental shift in how leaders access workplace science. For decades, our research helped organizations make better decisions, but it often required leaders to search, interpret, and apply those insights themselves. By building on Amazon Bedrock, we’re embedding scientifically grounded guidance directly into the flow of work, giving managers real-time support that is both personalized and actionable.”

— Andrew Bridger, Director of Product, Gallup

With this foundation in place, Gallup is focused on expanding what the application can do next.

What’s next

Gallup’s roadmap focuses on making its expertise more accessible, actionable, and embedded into everyday workflows. A key initiative is the development of an AI-curated prompt library that captures the most common questions managers and leaders ask. This library will help users quickly engage with Gallup AI through proven, high-value prompts grounded in workplace research.

In addition, Gallup is introducing guided coaching experiences built around structured conversation flows. These guided prompts walk managers through well-defined coaching scenarios, such as improving engagement, addressing team challenges, or developing employees, by sequencing prompts and responses into purposeful, outcome-driven interactions.

Gallup is building an agent-based foundation using Amazon Bedrock AgentCore. This positions Gallup AI to move beyond a user-facing assistant. By surfacing tools, workflows, and proprietary knowledge programmatically, the system can support not only end users but also other systems and integrations across the application.

Conclusion

By combining the generative AI capabilities of Amazon Bedrock with Gallup’s proprietary workplace research, leaders now have instant access to scientifically grounded guidance exactly when they need it. The serverless architecture enables the application to scale reliably while delivering low-latency streaming responses and comprehensive observability.

To build your own generative AI application, get started with Amazon Bedrock. To learn more about grounding responses in your own data, explore Amazon Bedrock Knowledge Bases.

Further reading


About the authors

How a global payment processor preserved AWS RAM shares and Lake Formation permissions during an AWS Organizations migration

Post Syndicated from Sam Mukherjee original https://aws.amazon.com/blogs/architecture/how-a-global-payment-processor-preserved-aws-ram-shares-and-lake-formation-permissions-during-an-aws-organizations-migration/

Accounts move between organizations in AWS Organizations whenever a business changes shape. A merger folds one estate into another. A divestiture carves one out, and some companies run more than one organization by design.

The moves take more care when AWS Resource Access Manager (AWS RAM) resource shares are involved. An organization-bound share trusts an account through its organization membership, so when the account leaves, AWS RAM removes that association. Anything in production that depends on a shared resource needs a continuity plan before the first account moves.

A leading worldwide provider of payment technology and software solutions, based out of the United States, used temporary AWS RAM resource shares to preserve AWS Lake Formation permissions during an AWS Organizations migration of 382 AWS accounts. The payment processor serves merchants and financial institutions around the world. The program separated them from their former parent company, a US-headquartered financial technology provider serving banking and capital markets clients globally, before a Transitional Service Agreement (TSA) expired in April 2026.

The migration unfolded alongside wider corporate change. In January 2026, the payment processor was acquired by a leading payments technology company headquartered in the United States. The transaction transforms the acquirer into a pure-play commerce solutions provider, serving the full spectrum of clients from small businesses to global enterprises worldwide. The migrating estate also included the payment processor’s embedded payments platform, a US-based provider of embedded payment and automated onboarding tools for software-as-a-service (SaaS) platforms.

Most workloads kept running when the original organization-bound shares broke, but the control plane lost access. They needed a migration pattern that preserved service continuity without leaving temporary permissions behind.

AWS partnered with them to design and validate that pattern in two weeks. It uses retained bridge shares for the move, then restores the original shares as the durable permission objects.

In this post, we explain why the bridge works, why the original share must return, and how the company applied the pattern at enterprise scale. For command-level implementation, see Transfer AWS accounts between AWS Organizations while preserving AWS Lake Formation permissions and aws-samples/sample-aws-ram-org-migration.

Solution overview

An AWS account can consume a resource to share automatically while it belongs to the same organization as the producer. When the account leaves, AWS RAM removes that organization-bound principal association. Creating a retained bridge share as an external association before the move preserves access through the organization’s change.

After the move, the company restores the migrated account to the original share, verifies access, and removes the bridge. The original share remains the AWS Lake Formation-managed source of truth. New grants and resource changes continue to attach to it, not to the point-in-time bridge copy. Keeping both would create duplicate permission state and drift.

The following diagram shows the migration wave structure.

Migration wave structure. Stage one covers fourteen non-production waves across eight months, none of which crossed an organization boundary. Stage two covers sixteen production waves: one pilot wave, twelve scheduled waves on a weekly cadence, and three contingency waves. The transitional service agreement expires inside the contingency window, leaving only the first contingency week usable.

The company’s cloud engineering team ran 14 non-production waves, a production pilot, 12 weekly production waves, and three contingency waves. The TSA expired inside the contingency window, leaving about one week of usable slack.

Challenge

The risk surfaced in a production wave in February 2026. A terraform apply against a shared AWS Transit Gateway failed with a permission error, although traffic kept flowing and no alarm fired.

This was a control plane failure. Existing attachments, DNS paths, and certificates continued to work, but engineers could not change shared resources. The dedicated AWS Transit Gateway, Amazon Route 53 Resolver, and the embedded payments platform’s AWS Glue Data Catalog waves were still ahead.

Why the issue stayed hidden

Most services retain their data plane when an AWS RAM association breaks. For example, an Amazon Elastic Compute Cloud (Amazon EC2) instance in a shared Amazon Virtual Private Cloud (Amazon VPC) keeps running, but the account cannot launch a new instance. Infrastructure-as-code exposed the problem because it needed control-plane access.

The following table summarizes the affected resource types confirmed by the customer.

Resource AWS service Effect of losing the share
AWS Transit Gateway ec2:TransitGateway Loses control plane access, keeps the data plane
Amazon Route 53 Resolver rules route53resolver:ResolverRule Risk of DNS resolution disruption
AWS Private Certificate Authority (AWS Private CA) acm-pca:CertificateAuthority Loses the share, issued certificates keep working
Amazon EC2 prefix lists ec2:PrefixList Keeps the data plane, blocks new resource creation
AWS Glue Data Catalog databases and tables AWS Glue and AWS Lake Formation Requires bridge-share validation before migration

Some services require additional handling. Depending on its resource-cleanup configuration, an AWS Firewall Manager policy can remove AWS Network Firewall rules, which you must then redeploy, and organization-integrated AWS CloudFormation StackSets can delete stacks unless you set them to retain.

The embedded payments platform initially entered the estate through an acquisition by the former parent company. Its embedded payment and automated onboarding capabilities were subsequently integrated into the payment processor’s ecosystem to support a broader platform-focused offering for SaaS providers. The platform shared databases and tables across 10 accounts with account IDs as principals, and the workstream was paused rather than testing the migration against production.

Why the original shares broke

The company had enabled sharing with AWS Organizations in the producer account. AWS RAM therefore trusted each in-organization principal through organization membership, even when a share named an account ID. When an account left the source organization, AWS RAM removed that organization-bound association.

A share created for a principal outside the organization behaves differently. AWS RAM sends an invitation, and the accepted association is external. Because it does not depend on organization membership, it survives the account move. The bridge-share pattern uses this behavior.

Why non-production testing missed it

The company’s non-production accounts were already in a separate organization. They never crossed the boundary that caused production associations to break.

Fourteen clean waves validated the migration process but not the production-only condition. Each validation environment must cross the same trust boundary as production.

Applying the bridge-share pattern

This failure was raised with the AWS account team, which brought AWS RAM, AWS Glue, AWS Lake Formation, and AWS Organizations service teams into the response. The company first used a manual recovery path for 21 accounts while the teams automated a scalable approach.

On February 27, 2026, AWS released RetainSharingOnAccountLeaveOrganization for new AWS RAM resource shares. The setting marks principals as external after they accept the invitation. The customer confirmed that the setting does not retrofit existing shares, so those shares needed a temporary parallel share.

Retaining access during the move

AWS RAM allows a resource to belong to more than one resource share, so a parallel share can exist alongside the original. A second retained share was created alongside each original, targeting the same consumer account.

The consumer accepted the invitation before migration, creating an external association. During the move, AWS RAM removed the original organization-bound association while the bridge continued to grant access. AWS Organizations also support transferring an account directly between organizations, so the move itself does not require an intermediate standalone period.

Why restore the original share?

The bridge is a migration-only continuity copy. The original AWS Lake Formation-created share remains the durable, service-managed permission object. If a team adds a grant or changes a shared resource during the migration window, that change applies to the original share, not automatically to the bridge.

The company therefore restored migrated principals to the original share before deleting the bridge. Leaving both in place would create two permission paths that can diverge, complicate audits, and conceal which share is authoritative.

Automation deletes a bridge only after it finds a non-bridge original whose resources, principals, and permissions cover the bridge and whose associations are all ASSOCIATED. This check confirms the original share’s associations are active before removing the temporary path. Access and connectivity were validated separately, as described in the Outcome section.

Validated workflow

AWS validated the pattern across three test accounts and two organizations before it was used in production. Testing confirmed that allowExternalPrincipals alone was not enough. The bridge also required retainSharingOnAccountLeaveOrganization.

The following diagram shows the bridge before and after the account move.

Bridge share behavior before and after an account moves between AWS Organizations. Before the move, the original organization-scoped share and the accepted bridge share both grant access. After the move, the original share is revoked and the accepted bridge share continues to grant access.

Each production wave used five steps:

  1. Inventory. Map each original share, resource, principal, permission, and Region. AWS RAM is Regional, so repeat the inventory in every in-scope Region.
  2. Create and accept bridges. Create a retained share for the same resource and principals, then accept its invitation from each consumer account before migration.
  3. Migrate. Move the account. AWS RAM removes the organization-bound association, while the accepted bridge keeps access active.
  4. Restore originals. Add the migrated account IDs back to the original shares as external principals. This reactivates the durable shares and includes grants created during the migration window.
  5. Validate and remove bridges. Confirm resources, principals, permissions, and association status, then delete only bridge shares fully covered by active originals.

This workflow ran for every remaining production wave and kept the weekly cadence.

Validating AWS Glue and Lake Formation permissions

The embedded payments platform shared AWS Glue Data Catalog databases and tables across 10 accounts, with account IDs as principals. The configuration was reproduced in disposable accounts, and the resource policy was recorded through a cross-organization move.

The validated automation records principal-to-share mappings, supports dry-run and execute modes, restores principals to the original shares, and deletes bridges only after validation. The embedded payments platform completed its production migration on July 21, 2026.

Outcome

  • The global payment processor migrated 378 of the 382 accounts into its landing zone. The final four awaited approvals from external stakeholders.

The TSA with the former parent company ended on schedule in April 2026. No customer-facing workload lost availability, and the company recorded no network drops across the production waves. The company and AWS moved from discovery to a validated bridge-share pattern in two weeks.

After each migration, the cloud engineering team restored the original shares, verified access and connectivity, and removed the bridges. Deleting the temporary copies confirmed that the original service-managed permission path was active and authoritative.

Production, non-production, and the embedded payments platform now run in one landing zone. They control their own guardrails, security posture, provisioning, and change process.

AWS has published the validated pattern and automation, so other organizations can start with a tested procedure.

Lessons learned

This program produced three lessons for organizations planning similar migrations.

Match validation boundaries to production

A test organization cannot expose this failure unless it crosses the same organization boundary as production. Map each production risk to an environment that can reproduce it before the first wave.

Monitor control plane changes

AWS RAM emits resource share state-change events directly to Amazon EventBridge, and AWS CloudTrail records DisassociateResourceShare API calls for audit. A weekly post-migration sweep provided a periodic reconciliation check to catch stale shares.

Inventory dependencies and destination guardrails

The Account Assessment for AWS Organizations tool inventories AWS RAM dependencies before teams set the wave plan. The company’s cloud engineering team reviewed destination guardrails at the same time. A service control policy that blocked ram:AcceptResourceShareInvitation during migration windows was temporarily adjusted.

Conclusion

The migration of this leading payment technology and software company shows how a retained bridge share can protect access while an account moves between AWS Organizations. The bridge is temporary: restoring the original share keeps AWS Lake Formation permissions aligned with future grants and avoids two sources of permission state. Inventory, dry-run-first automation, and post-move validation helped the company meet its deadline without customer disruption.

Next steps

To apply this pattern, read Transfer AWS accounts between AWS Organizations while preserving AWS Lake Formation permissions and review aws-samples/sample-aws-ram-org-migration. Run the scripts in dry-run mode, validate each Region and account, and engage your AWS account team early when AWS Glue Data Catalog or AWS Lake Formation resources are in scope.


About the authors

How AgentFlo built AI sales agents with Amazon Bedrock AgentCore – Part 2

Post Syndicated from Muhammad Musab Iqbal original https://aws.amazon.com/blogs/architecture/how-agentflo-built-ai-sales-agents-with-amazon-bedrock-agentcore-part-2/

If you’re building AI agents for commerce at scale, you face two critical challenges: handling unpredictable traffic spikes and ensuring your agents can be trusted with real customer transactions.

This post shows how AgentFlo solved these challenges using Amazon Bedrock AgentCore and AWS serverless architecture. You learn the architectural patterns behind their reliability and trust frameworks, see the measurable business results (including +12% net revenue uplift based on early deployment data), and explore their roadmap for voice agents and server-side tool execution.

This is Part 2 of a two-part series. Part 1 covers velocity, standardization, and scalability.

Pillar 4: Trust: guardrails for autonomous commercial action and real-time visibility into agent operations

AgentFlo enforces trust at every layer of the stack, from pre-request filtering to post-response privacy controls, so merchants can deploy autonomous agents with confidence.

The challenge

Enterprise customers won’t deploy autonomous agents unless they can trust them. An agent in production can’t expose sensitive data, offer unauthorized discounts, or access another customer’s information. The system also prevents price hallucination, unauthorized tool calls, opt-out violations, and credential exposure.

Merchants also need fine-grained control over who can interact with their AI agents and what data each segment can access. Enterprise customers require restricted access; B2C businesses need open access for broader reach. Without identity-based controls, deploying customer-facing AI is a non-starter.

Defense in depth

In AgentFlo, trust isn’t only about safe responses. It’s about safe action. Agents can create carts, place orders, apply discounts, access customer data, and interact with backend systems, so policy enforcement must sit outside the model’s reasoning loop. The model proposes. Deterministic policy decides.

AgentFlo applies trust controls across the full agent lifecycle: before the model sees the request, during tool execution, and after the model generates a response.

Three-layer guardrails

When users deploy an agent from the AgentFlo Portal, security enforcement happens at three stages.

First, the AWS Fargate layer detects prompt injection and handles opt-outs before requests reach the agent. WhatsApp messages are authenticated using phone numbers as unique identifiers. Enterprise customers like EBM restrict access to authorized users, while restaurant deployments stay open for broader reach.

Next, the AgentCore layer verifies identity and enforces order locks during tool execution. AgentCore Gateway, a capability of Amazon Bedrock AgentCore, enforces policies that prevent sales agents from accessing customer support tools. Cedar policies enforce business rules like maximum discount percentages independently of the model’s reasoning. Cedar is an open-source policy language developed by AWS for fine-grained, verifiable authorization decisions. Additionally, Policy in Amazon Bedrock AgentCore integrates with Amazon Bedrock Guardrails, so Cedar policies can invoke configurable safeguards for prompt attack detection, content filtering, and sensitive information blocking directly at the gateway boundary.

Finally, post-turn privacy filters screen outputs to block inadvertent token disclosure and unverified price claims before customers see responses. Secrets are managed through AWS Secrets Manager with OIDC (OpenID Connect)-authenticated continuous integration and continuous delivery (CI/CD) pipelines. No credentials are stored in agent code.

Infrastructure security

Trust at the application layer requires sound underlying infrastructure. AgentFlo combines the built-in isolation of AgentCore with application-level controls:

  • Session isolation: Session isolation through AgentCore runtime, a capability of Amazon Bedrock AgentCore, which provides complete separation between merchants’ agent sessions through dedicated microVMs.
  • AgentCore Gateway policies: Fine-grained Cedar policies control which agents can access which tools and data, enforced deterministically regardless of model reasoning.
  • AWS Identity and Access Management (IAM)-based access control: Fine-grained permissions for agent-to-service communication.
  • Amazon Virtual Private Cloud (Amazon VPC) integration: Agent sessions operate within AgentFlo’s VPC with domain-level network restrictions, so agents only communicate with approved endpoints.
  • Compliance: Data residency controls and audit trails for regulatory requirements across multiple jurisdictions.

Observability

Trust requires visibility. Merchants need to see what their agents are doing in real time, not only after something goes wrong. AgentFlo uses AgentCore Observability, a capability of Amazon Bedrock AgentCore, to provide end-to-end tracing of every agent interaction, from initial request through tool execution to final response.

AgentCore Observability captures structured traces for each agent turn, including model latency, tool invocation sequences, token usage, and error rates. These traces flow into Amazon CloudWatch, where AgentFlo builds dashboards showing active sessions, response times, and tool call patterns. Full request-to-response traces enable trace-level debugging. The system tracks P50/P95 latencies and throughput across agent types, alerting merchants when behavior deviates from baselines. Cost attribution provides per-merchant, per-agent breakdowns tied to specific conversations.

Results

Together, observability and trust give enterprises confidence to deploy autonomous agents at scale. Merchants benefit from safer execution, stronger compliance across jurisdictions, and controlled access to tools and data at the session level.

Pillar 5: Reliability: a data foundation that keeps agents grounded in reliable data

Reliable agents need reliable data. AgentFlo grounds every agent action in verified, current information through stateful sessions, merchant knowledge bases, and semantic product discovery.

The challenge

Enterprise customers won’t trust AI agents that forget context, hallucinate product details, or operate on stale data. An agent that quotes the wrong price, forgets a customer’s earlier request, or recommends discontinued products destroys confidence instantly. The system must make sure every agent action is grounded in verified, current information, from product catalogs and pricing to conversation history and business rules.

Merchants also need their agents to maintain continuity across long customer journeys, access up-to-date business-specific knowledge, and surface products through natural language, all without manual intervention or prompt engineering.

Data architecture overview

In AgentFlo, reliability isn’t only about accurate responses. It’s about accurate action grounded in verified data. Agents retrieve product information, manage carts, and complete transactions, so every data source must be authoritative and current. The model reasons. Structured data decides.

AgentFlo applies data reliability controls across three layers: stateful conversation management through Amazon DynamoDB, merchant-specific knowledge through Amazon Bedrock Knowledge Bases, and semantic product discovery through vector embeddings in Amazon S3 Vector.

State management architecture

A real sales journey can span 8 hours or 3 days. A customer might ask about a product in the morning, compare options at lunch, and complete the purchase that evening. The agent must remember context across all turns and maintain state, so customers don’t need to start over.

AgentCore runtime and Amazon DynamoDB manage this context storage. The agent replays relevant history, loads context based on intent, and continues transactions safely.

Each agent needs to be stateful (remembering earlier interactions), autonomous (deciding next steps without human intervention), and safe (operating within business rules without exposing sensitive data).

The two-table DynamoDB design covers all three requirements: session continuity through conversation replay, autonomous context loading based on detected intent, and data integrity through structured ground-truth storage that the model can’t hallucinate over.

Diagram of the per-message conversation flow between AgentCore runtime and the DynamoDB session and cart tables

Figure 1: Per-message conversation flow. AgentCore runtime loads the last 15 messages from the DynamoDB Session Table at the start of each turn. The Cart Table is loaded on demand only when intent detection flags the request as cart-related, preventing the model from generating incorrect prices and quantities.

Knowledge base system

Merchants upload business-specific data (restaurant menus, clinic policies, product specifications, promotion calendars) into Amazon Bedrock Knowledge Bases backed by Amazon Simple Storage Service (Amazon S3). Agents automatically retrieve and reason over this merchant-specific content. Responses stay grounded in accurate, up-to-date business information without merchants writing a single prompt.

Semantic search and vector retrieval

AgentFlo improves product discovery by giving every product in its Amazon Aurora database a lightweight vector embedding. Customers can find products by name or description.

AgentFlo also generates extra searchable tags for each product automatically, and merchants can add their own. The result: phrases like “the pink one,” “the smallest one,” “the new one,” or “the chocolate with the golden wrapper” all map to the right product. Customers ask for things naturally, and the platform finds what they’re looking for.

Observability and billing

The system captures all message interactions through Amazon Data Firehose to Amazon S3, so merchants can track cost per conversation and compare those costs against sales revenue. This pipeline shows merchants the return on investment (ROI) of their agent deployments and provides the data foundation for continuous agent improvement.

Results

The data architecture helps minimize context loss in conversations across multi-day customer journeys, with grounded responses that eliminate price and product hallucination. Merchants benefit from natural language product discovery without keyword dependency and merchant-specific knowledge retrieval without prompt engineering. Full cost visibility and ROI attribution per agent deployment give merchants clear measurement of platform value.

Business impact: measurable results across the customer lifecycle

Salesflo’s solution, powered by Strands Agents SDK and Amazon Bedrock AgentCore, delivered measurable improvements across the customer lifecycle:

Metric Improvement
Net revenue uplift +12%
Customer engagement +40%
Conversion rate +15%
Average order value +8%
Customer reactivation +20%

AgentFlo generated these results by comparing agent-assisted customer journeys against a control group over a 90-day early deployment period.

Note: The foundation models referenced in this post are available in select AWS Regions. For the latest information on model availability, see Supported Regions and models for Amazon Bedrock.

The compounding effect at scale

At AgentFlo’s scale, the impact is significant. With $300 billion in annual transacted value flowing through the Salesflo solution (based on platform transaction data), even single-digit percentage improvements translate to billions in incremental revenue for merchants.

Operational improvements

Beyond the metrics, merchants see operational improvements:

  • 24/7 coverage — AI agents engage customers at any hour, in any time zone, across any channel.
  • Consistent quality — Every customer interaction follows best-practice selling methodologies without the variability of human agents.
  • Scalable personalization — Thousands of concurrent agent sessions, each maintaining unique context per customer.
  • Rapid merchant onboarding — New merchants go live with customized AI agents in days, not months, through the self-service configuration platform.

What’s next: voice, server-side execution, and integration expansion

AgentFlo is actively extending the platform along three directions, each at a different stage of maturity.

Real-time voice agents (in pilot)

AgentFlo already supports voice within WhatsApp. Incoming voice notes go through a two-pass transcription process: a raw initial pass, followed by domain-aware correction that fuzzy-matches text against the live product catalog. The second pass catches brand names and SKUs even when partially misheard, across more than 90 supported languages. For outbound audio, merchants pick from multiple Text-to-Speech (TTS) providers per deployment.

Because Speech-to-Text (STT), reasoning on AgentCore runtime, and TTS are fully independent components, any one can be swapped without disrupting the rest of the pipeline.

The next step: BidiAgent. The next evolution is real-time voice agents built on the Strands SDK BidiAgent and Amazon Bedrock AgentCore WebRTC support. BidiAgent supports bidirectional audio streaming, natural interruptions, and concurrent tool execution. The agent can check inventory or apply a discount while continuing to listen and respond to the customer in the same call.

The AgentCore WebRTC protocol and Amazon Kinesis Video Streams handle peer-to-peer transport for mobile and browser interactions without requiring relay infrastructure. This pushes AgentFlo beyond text messaging into proactive outbound calls and high-value B2B sales, where real-time conversation is essential for building trust.

Server-side tool execution (under development)

AgentFlo is experimenting with server-side tool execution in Amazon Bedrock, which removes client-side orchestration entirely.

Traditional orchestration loops between model and tools repeatedly (model → execute → send result → repeat). With server-side execution, the agent makes a single API call to the Amazon Bedrock Responses API with an AgentCore Gateway Amazon Resource Name (ARN). The model then autonomously discovers, invokes, and processes tools through the Gateway Model Context Protocol (MCP) interface, all inside AWS infrastructure with no roundtrips back to the client.

Here’s what that one API call looks like:

from openai import OpenAI

# OPENAI_BASE_URL = https://bedrock-mantle.us-west-2.api.aws/v1
# OPENAI_API_KEY = <Amazon Bedrock API key>
client = OpenAI()

response = client.responses.create(
    model="openai.gpt-oss-120b",
    stream=True,
    background=False,
    store=False,
    input=[
        {
            "type": "message",
            "role": "user",
            "content": [{"type": "input_text", "text": user_message}],
        }
    ],
    tools=[
        {
            "type": "mcp",
            "server_label": "agentflo_gateway",
            "connector_id": GATEWAY_ARN,  # arn:aws:bedrock-agentcore:...:gateway/...
            "server_description": "AgentFlo commerce tools (cart, catalog, knowledge base)",
            "require_approval": "never",
        },
    ],
)

Note: Model availability varies by AWS Region. The model and endpoint shown in this example may not be available in all Regions. See Supported Regions and models for Amazon Bedrock for current availability.

One request handles the whole loop. The Gateway ARN goes in as an MCP connector, and Bedrock takes it from there; pulling the tool list, picking the right one, invoking it, and feeding the result back to the model without anything leaving AWS. The client never sees credentials, tool schemas, or the intermediate turns. See ShopAssist: E-Commerce Agent Demo for more details.

Early results: For specialist agents with short, focused tool loops, early measurements show approximately 30% lower latency. A sales agent calling three tools in sequence (check inventory, apply discount, update cart) collapses 50+ lines of orchestration into a single API call. All credentials stay server-side, and the simplified architecture makes onboarding new agent developers faster.

Integration ecosystem expansion (ongoing)

AgentFlo adds integrations based on merchant demand. Each new platform (payment processors, shipping providers, loyalty systems, or vertical-specific ERPs) becomes another MCP server connector in AgentCore Gateway. This pattern keeps expansion modular and quick.

Key takeaways

  1. Pick your pillars first. The right AWS stack follows. AgentFlo first defined what velocity, standardization, scalability, and trust meant for the production system. With clear requirements, the AWS stack choices became obvious: Strands for the agent layer, AgentCore runtime for stateful sessions, AgentCore Gateway for tool routing and policy, and AWS Fargate for message ingestion.
  2. Specialized agent recipes beat multi-agent complexity. Deploying a single, well-configured agent with domain-specific tools and knowledge outperforms multi-agent orchestration for most customer interactions. Multi-agent handoffs are reserved for cross-domain transitions where context boundaries are clearly defined.
  3. Commerce is stateful. Plan for that on day one. A real sales journey can span 8 hours or 3 days. Stateless chatbots can’t maintain context that long. AgentCore runtime stateful sessions and DynamoDB-backed context were chosen on day one for exactly this reason. Retrofitting state onto a stateless agent later is much more painful than designing for it up front.
  4. Three-layer security builds trust. Pre-turn guards on Fargate, per-tool guards through AgentCore Gateway policies with Cedar, and post-turn output filters work together to support safe deployment of autonomous agents handling commercial transactions.
  5. System design supports scale. By building agent customization as a software as a service (SaaS) layer on top of AgentCore infrastructure, AgentFlo serves hundreds of merchants on shared infrastructure while still delivering personalized agent behavior for each one.
  6. Serverless + AgentCore = elastic commerce. The combination of Fargate for message handling and AgentCore for agent execution means AgentFlo scales from normal traffic to 50x flash-sale spikes without pre-provisioning or capacity planning.
  7. The feedback loop is the product. Real customer conversations teach the agents how to close. Where customers hesitate, what language converts, when they want a human handoff — all of it feeds back into recipes, prompts, and tool definitions. A new merchant deployment benefits from every conversation that ran on the platform before it. That compounding loop is harder to copy than any single piece of the architecture.

Conclusion

Building production-scale AI agents for commerce requires scalability and trust from day one. By combining Amazon Bedrock AgentCore stateful sessions and microVM isolation with AWS serverless infrastructure, AgentFlo delivers autonomous agents that handle unpredictable traffic while maintaining the security controls enterprises require.

Next steps

For questions about implementing similar architectures, visit the AWS Architecture Center or contact your AWS account team. To start building, open the Amazon Bedrock console or explore the Amazon Bedrock service detail page.

We’d love to hear how you’re building agentic AI systems. Share your experiences in the comments.


About the authors

How AgentFlo built AI sales agents with Amazon Bedrock AgentCore – Part 1

Post Syndicated from Muhammad Musab Iqbal original https://aws.amazon.com/blogs/architecture/how-agentflo-built-ai-sales-agents-with-amazon-bedrock-agentcore-part-1/

In this post, you learn how AgentFlo built intelligent sales agents that convert conversations into completed purchases. We show you how AgentFlo improved revenue performance in early deployments using Amazon Bedrock AgentCore and the Strands Agents SDK.

AgentFlo, the agentic commerce service by Salesflo, helps merchants deploy always-on AI sales, support, and ordering agents across channels like WhatsApp. These agents understand intent, connect to commerce systems, recommend products, create carts, and convert conversations into completed transactions. Today, AgentFlo serves eCommerce merchants managing over $300 billion in annual transacted value, according to Salesflo, across services including Shopify, WooCommerce, Magento, and SAP.

This is Part 1 of a two-part series covering the five pillars of production-grade AI agents. Part 1 covers Velocity, Standardization, and Scalability. Part 2 covers Trust, Reliability, and Business results.

The challenge: customer intent without assistance

Cart abandonment hovers around 70% industry-wide, representing trillions in unrealized revenue annually. For merchants operating on messaging platforms like WhatsApp, the gap widens further:

  • Cart abandonment: Customers abandon carts because a single question goes unanswered.
  • Generic product discovery: Ranked listings replace recommendations tailored to each customer.
  • Missed messaging conversations: Inbound chat volume exceeds staffing capacity across time zones and languages.
  • No personalized guidance: Most merchants can’t afford 1:1 assistance for every interaction.
  • Limited outbound engagement: Teams lack bandwidth for proactive sales motions.
  • Peak traffic spikes: Flash sales and seasonal campaigns can spike traffic 10–50x beyond normal capacity.

Rule-based chatbots can’t handle nuanced sales conversations. Human agents can’t scale across geographies, languages, and time zones. Merchants need specialized AI sales agents that understand customer context, run complex workflows, and operate autonomously 24/7.

What makes an agent?

At its simplest, an agent combines a model, instructions, tools, context, and memory. The model reasons over a user’s request. The instructions define the agent’s role and the limits of what it should do. Tools let the agent take action in the real world. Context grounds it in business-specific data. Memory keeps the conversation coherent across turns.

In AgentFlo, those abstract components map to concrete pieces of the platform:

Component What it does in AgentFlo
Model Understands user intent and decides what to do next
System prompt / persona Defines whether the agent behaves like a sales agent, restaurant agent, support agent, or receptionist
Tools Allow the agent to search products, check inventory, create carts, place orders, raise tickets, or trigger follow-ups
Knowledge Grounds responses in merchant-specific data such as product catalogs, menus, policies, promotions, and FAQs
Memory / state Maintains conversation history, cart state, customer preferences, and previous actions
Channels Connects the agent to WhatsApp, SMS, RCS, web chat, and voice
Guardrails / Policy Prevents unsafe, unauthorized, or incorrect actions
Observability Tracks cost, performance, conversions, and conversation quality

Building a demo agent is straightforward. Building one that runs a business safely, repeatably, and at scale requires a different approach. AgentFlo organizes this approach around five pillars.

Five pillars of production-grade AI agents

AgentFlo’s architecture centers on five pillars: Velocity, Standardization, Scalability, Trust, and Reliable. Each addresses a specific production challenge. These challenges influenced how the team chose AWS services, including Strands Agents SDK, Amazon Bedrock AgentCore, Amazon Bedrock, AWS Fargate, Amazon DynamoDB, Amazon Aurora, Amazon Kinesis, and Amazon Simple Storage Service (Amazon S3).

Architecture overview

AgentFlo production architecture on AWS spanning the messaging, agent runtime, tool gateway, data, and observability layers

Figure 1: AgentFlo’s production architecture on AWS.

Customer messages arrive through WhatsApp Graph API or web/mobile channels and pass through an Application Load Balancer into the AWS Fargate messaging layer. It handles authentication, image optical character recognition (OCR), speech-to-text/text-to-speech, pre-turn guardrails, and prompt injection detection. Validated requests flow into AgentCore runtime, a capability of Amazon Bedrock AgentCore, where the Strands Agents SDK orchestrates an agent that streams model inference to an external large language model (LLM). AgentCore Gateway, a capability of Amazon Bedrock AgentCore, brokers tool calls, with IAM-based authorization, to an API layer of AWS Lambda functions (Cart, Product, and Knowledge Base). It persists state across a data layer comprising Amazon DynamoDB session and cart tables, Amazon Aurora order tables, and an Amazon Bedrock Knowledge Base backed by Amazon S3. Policy in Amazon Bedrock AgentCore enforces deterministic access control independently of model reasoning. Amazon Bedrock Guardrails can also be embedded in Policy to filter prompt attacks, harmful content, and sensitive information on both requests and responses. On the observability side, logs and traces feed into AgentCore Observability, a capability of Amazon Bedrock AgentCore, while Amazon Data Firehose captures every interaction into Amazon S3 for cost and revenue analytics.

AgentFlo evaluated several hosting options before selecting Amazon Bedrock AgentCore. Three capabilities made the difference:

  • Stateful sessions for long-running commerce conversations.
  • Agent runtime: each agent session runs in its own lightweight virtual machine, providing hardware-level security boundaries between tenants.
  • Native MCP integration: Model Context Protocol (MCP) is an open standard that allows AI agents to connect securely to external data sources and tools through a unified interface. AgentCore Gateway supports MCP natively for standardized tool connectivity.

Pillar 1: Velocity: from merchant idea to live agent in minutes

Speed to market determines whether merchants can capture emerging opportunities. AgentFlo addresses this with a streamlined deployment model.

The challenge

Merchants want to launch agents quickly, but each has unique workflows, tone, tools, languages, products, and business rules. Generic chatbot templates are too shallow. Custom-building each agent doesn’t scale.

Recipe-based deployment

AgentFlo uses a recipe-based agent deployment model. Merchants select from pre-configured recipes, each shipping with persona, language, tone, tool sets, prompt templates, knowledge sources, response packs, and business rules. Available recipes include:

  • Sales agent.
  • Restaurant ordering agent.
  • Clinic receptionist.
  • Support agent.
  • B2B reorder agent.
  • Cart recovery agent.

Merchants fine-tune a few choices in the AgentFlo Portal. The rest is automated.

How it works

AgentFlo chose Strands Agents SDK as its agent framework. Strands Agents SDK uses a model-driven architecture: you define tools as Python functions, write a system prompt, and let the model handle orchestration. No rigid workflow graphs or hand-coded state machines.

From the merchant’s perspective, agent creation is entirely no-code. Here’s an example of the agent customization flow:

AgentFlo Portal agent customization workflow

Figure 2: Agent customization workflow

This approach makes the recipe model work. Adding a new capability (like loyalty program enrollment) means writing a new tool function and updating the system prompt. No orchestration layer rewiring is needed.

Behind the scenes, one selection triggers an automated pipeline:

  1. The portal generates a Strands agent configuration from the recipe template.
  2. GitHub Actions packages the agent (tools, prompts, context) into a container.
  3. The container deploys to AgentCore runtime with appropriate Gateway policies.
  4. The agent goes live on WhatsApp within minutes.
Automated pipeline that builds and deploys a customized agent

Figure 3: How to build a customized agent under hook workflow

Each agent is a Strands Agent instance with tool definitions mapped to AgentCore Gateway endpoints. Extending an agent’s capabilities is a code change, not an architectural one.

Results

This recipe-based approach delivers faster merchant onboarding, rapid experimentation with new agent behaviors, and quick addition of new capabilities platform-wide. It also means lower engineering effort per deployment and a tighter feedback loop between customer conversations and product iteration.

Pillar 2: Standardization: reusable recipes, tools, and commerce workflows

Consistency across deployments speeds iteration and reduces maintenance burden. AgentFlo achieves this through shared building blocks.

The challenge

As AgentFlo expanded across industries, fragmentation threatened to slow the team down. Every merchant has unique catalog structures, ERP setups, system configurations (Shopify, WooCommerce, Magento), pricing rules, languages, promotions, and support processes. Without standardization, every deployment becomes a custom project, and custom projects don’t scale to hundreds of merchants.

Repeatable building blocks

AgentFlo standardizes around several core components: agent recipes (domain-specific templates), a tool marketplace (reusable capabilities), MCP-based connectors (standardized integrations), integration contracts (consistent interfaces). Additional components include prompt and context packs (reusable templates), conversation review loops (continuous improvement), and a shared semantic layer (unified product understanding). Merchant-specific complexity is pushed to the system edges. The core remains consistent.

Single-agent architecture with domain expertise

We learned early on that a single agent with domain-specific knowledge and a curated tool set outperforms multi-agent architectures for most customer interactions. Each AgentFlo deployment configures a distinct persona, voice, language, and specialized tool set tailored to the business domain (sales agent, restaurant agent, clinic receptionist). It ships with curated contexts, prompt templates, and response packs. Merchants deploy these domain-specific agents through the self-service portal.

A single focused agent maintaining unified context converts better than multiple generalists coordinating with each other. Multi-agent capabilities remain available where valuable. For example, when conversations transition from sales to support, context hands off cleanly to the specialist agent.

Tool routing through AgentCore Gateway

Centralized tool management routes agent requests to dedicated AWS Lambda functions. When a customer asks about product availability, the agent queries the product catalog. When they’re ready to buy, it handles cart operations. For personalized recommendations, it retrieves data from Amazon Bedrock Knowledge Bases, the fully managed Retrieval Augmented Generation (RAG) capability, backed by merchant data stored in Amazon S3. Sales intelligence APIs provide additional context for each interaction.

OAuth tokens and platform credentials live in the Gateway, not in agent sessions. Policy in AgentCore sits alongside the Gateway, enforcing fine-grained, Cedar-based access control rules that operate independently of model reasoning. Cedar is an open-source policy language developed by AWS that allows fine-grained, verifiable authorization decisions.

Policies define which agent sessions can invoke which tools. For example, a sales agent can’t call customer-support-only APIs.

For standardization, this means every new tool added to the platform (payment integration, shipping provider, loyalty system) becomes available to every applicable recipe through the same mechanism. Tools aren’t re-implemented per merchant.

AgentCore Gateway as the integration backbone

AgentFlo integrates with dozens of eCommerce services: Shopify, WooCommerce, Magento, SAP, payment processors, shipping providers, and loyalty systems. Each integration is defined as an MCP server connector, with Gateway handling discovery, authentication, and routing. Furthermore, AgentFlo has many different agent recipes, each with their own specialized tool packs to provide that functionality. To make these connections modular and efficient, AgentFlo uses AgentCore Gateway.

From the agent’s perspective, the full Gateway tool surface is reachable in a few lines:

from strands import Agent
from strands.models import BedrockModel
from strands.tools.mcp.mcp_client import MCPClient
from mcp.client.streamable_http import streamablehttp_client

def create_transport():
    return streamablehttp_client(
        GATEWAY_URL,
        headers={"Authorization": f"Bearer {access_token}"},
    )

mcp_client = MCPClient(create_transport)

with mcp_client:
    # Discover every tool registered on the Gateway in one call.
    # cart, product catalog, knowledge base, shipping, loyalty, etc.
    tools = mcp_client.list_tools_sync()

    agent = Agent(
        model=BedrockModel(model_id="us.anthropic.claude-sonnet-5-20260630"),
        tools=tools,
        system_prompt=SALES_AGENT_PROMPT,
    )

    response = agent("Do you have the red leather wallet in stock?")

Connecting a Strands agent to AgentCore Gateway. A single list_tools_sync() call gives the agent every integration registered on the Gateway: cart, product catalog, knowledge base, shipping, loyalty. Onboarding a new service for a merchant is a Gateway change, not an agent change.

Key capabilities:

MCP-native tool connectivity: Each platform or tool set integration is a standard MCP server connector.

OAuth and credential management: Platform API keys and OAuth tokens are managed centrally in Gateway, never exposed to individual agent sessions.

Code simplicity: The code is cleaner, shorter, and more modular, which simplifies configuration for scale. The alternative is extensive local code for each connection or tool set.

Because each new service integration and recipe-specific tool set is defined as an MCP server connector in AgentCore Gateway, expansion is modular and quick. Adding a new service or tool set requires a connector definition, not a re-architecture.

Conversation reviews as a standardization loop

Standardization also comes from learning. AgentFlo continuously reviews real conversations to understand how customers ask for products, where they drop off, which recommendations convert, when handoff is needed, and how local language affects buying behavior.

These reviews feed back into recipes, prompts, and tool definitions. Standardization is something the platform earns over time, not something declared at launch.

Results

Every deployment improves future deployments, and new integrations become reusable across the merchant base. Agent behavior stays consistent across recipes and merchants, and workflows become repeatable across industries. The service becomes harder to replicate because it learns from real commerce behavior, not generic templates.

Pillar 3: Scalability: elastic, stateful commerce conversations

Commerce conversations are unpredictable in volume and duration. AgentFlo’s architecture handles both dimensions without manual intervention.

The challenge

During flash sales, product launches, or restaurant rush hours, customer conversations can spike 10–50x. Human teams can’t scale that fast. AgentFlo must also support many merchants and concurrent customer sessions simultaneously. One merchant’s surge can’t affect another’s experience.

AgentFlo’s serverless architecture

AgentFlo uses a serverless architecture. Message ingestion, agent execution, tool execution, state, analytics, and billing each scale independently. Each layer absorbs its own spikes without requiring the rest of the system to over-provision.

AgentCore runtime properties for scale

Several AgentCore runtime properties specifically support scale:

Isolated microVM execution: Each agent session runs in its own environment with dedicated CPU, memory, and filesystem. The environment is sanitized on termination. One merchant’s sessions never interfere with another’s.

Stateful sessions up to eight hours: Long-running conversations don’t lose context. A customer browsing in the morning can continue the same assisted session that evening.

Framework-agnostic: AgentCore runs Strands Agents natively but also supports any containerized agent framework, giving AgentFlo flexibility to evolve the agent architecture over time.

How it works

The architecture is built end-to-end on AWS:

Messaging layer (AWS Fargate): An Application Load Balancer routes incoming WhatsApp messages to a Fargate application that handles authentication, voice message conversion (Opus OGG to MP3 transcription with fuzzy matching for product name recognition). The application also runs pre-turn security guards. AWS End User Messaging provides an alternative channel option for broader reach.

Agent orchestration (Amazon Bedrock AgentCore runtime): Each customer session spawns an isolated agent instance running the Strands Agents SDK. Sessions are stateful for up to eight hours and isolated through microVM architecture, where each session runs in its own lightweight virtual machine. They are persistent, with filesystem access for intermediate results and cached product catalogs. microVM isolation keeps merchants completely separated.

This is the entire bridge between Strands SDK agent code and a production-ready endpoint on AWS:

from strands import Agent
from bedrock_agentcore.runtime import BedrockAgentCoreApp

app = BedrockAgentCoreApp()

agent = Agent(
    tools=tools,  # from the Gateway, per snippet above
    system_prompt=SALES_AGENT_PROMPT,
)

@app.entrypoint
def invoke(payload, context):
    """One AgentCore session per customer conversation."""
    user_message = payload.get("prompt")
    session_id = getattr(context, "session_id", None)  # stable for up to 8 hours
    result = agent(user_message)
    return {"result": result.message}

if __name__ == "__main__":
    app.run()

This is all the glue between a Strands agent and AgentCore Runtime. BedrockAgentCoreApp wraps the agent in the standard /invocations contract, and AgentCore handles microVM provisioning, session isolation, scaling, and stateful sessions up to eight hours. Two CLI commands take it from a local file to a live endpoint on AWS. No Dockerfile, no API routing, no web framework to maintain.

agentcore configure --entrypoint agent.py
agentcore launch

Tool execution (Amazon Bedrock AgentCore Gateway): Tool calls scale separately from agent reasoning, so a sudden burst of cart operations doesn’t slow down the agent loop itself.

Results

The architecture handles peak traffic without pre-provisioning capacity and supports long-running conversations that survive across visits. It provides strong multi-merchant isolation with lower operational overhead than traditional always-on infrastructure, resulting in a better customer experience during high-intent moments like product launches or flash sales.

What’s next

In Part 2 of this series, we explore:

Pillar 4: Trust. Guardrails for autonomous commercial action and real-time visibility into agent operations.

Pillar 5: Reliable. Data foundation that ensures agents act on reliable, up-to-date information to complete tasks with precision.

Business results: Measurable impact across the customer lifecycle.

Future roadmap: Voice agents, server-side tool execution, and integration expansion.

Summary

In this post, we explored three of the five pillars for building production-grade AI agents:

Velocity: How recipe-based deployment allows merchants to launch AI sales agents in minutes using the model-driven architecture of the Strands Agents SDK.

Standardization: How reusable building blocks and centralized tool management through Amazon Bedrock AgentCore create consistency across hundreds of deployments.

Scalability: How AgentFlo handles elastic, stateful commerce conversations at scale through Amazon Bedrock AgentCore and AWS Fargate.

Next steps

We’d love to hear how you’re building agentic AI systems. Share your experiences in the comments.


About the authors

How Clario technology detects PHI/PII in DICOM images using Amazon Bedrock

Post Syndicated from Alex Boudreau original https://aws.amazon.com/blogs/architecture/how-clario-automates-phi-pii-detection-in-dicom-images-using-amazon-bedrock/

Clario, part of Thermo Fisher Scientific, uses Amazon Bedrock to automate PHI (Protected Health Information) and PII (Personally Identifiable Information) detection across thousands of DICOM (Digital Imaging and Communications in Medicine) image slices in clinical trials. Each image slice may carry PII or PHI hidden in metadata tags, in custom vendor fields, or burned directly into the pixels. Across imaging sites, central labs, sponsors, and CROs (Contract Research Organizations), every one of those slices must be cleared of PII and PHI before the image moves downstream.

DICOM is the universal standard for storing, transmitting, and managing medical imaging data across healthcare systems. In clinical trials, DICOM images play a critical role by providing objective, quantifiable evidence of a patient’s medical condition throughout the study lifecycle. From baseline imaging to follow-up scans, modalities such as MRI, CT, PET, and X-ray generate DICOM files. Radiologists, clinicians, and sponsors use these files to assess treatment efficacy, monitor disease progression, and support regulatory submissions. These images serve as a core component of the clinical evidence package, making their accurate management and standardized handling essential to trial integrity.

In this post, we share how the Clario team designed an automated PHI and PII detection solution on AWS for DICOM imaging data, the key design decisions behind the architecture, and the lessons the team learned along the way.

About Clario, part of Thermo Fisher Scientific

Clario science and endpoint solutions support the clinical trials industry through the systematic collection, management, and analysis of specific, predefined outcomes (endpoints) to evaluate a treatment’s safety and effectiveness. For more than 50 years, Clario endpoint solutions have been deployed more than 30,000 times, and since 2015, they have supported more than 700 FDA and EMA new drug approvals.

Business challenge

Clearing PII and PHI from every image slice in the clinical trial is only part of the problem. The imaging workflow around this clearing process must be just as rigorous. A well-structured imaging workflow supports every DICOM file captured across globally distributed trial sites. Files are ingested automatically, consistently standardized, and rigorously validated at every step of the journey. Enforcing standardized image acquisition protocols across sites and geographies alleviates inconsistencies. These inconsistencies could otherwise impact data quality or delay regulatory submissions. A centralized imaging infrastructure that maintains complete metadata traceability, including acquisition parameters, imaging equipment details, and timestamps, supports a fully auditable workflow aligned with GCP (Good Clinical Practice) requirements. This empowers sponsors and CROs to move faster with greater confidence and significantly reduces the risk of data queries or compliance gaps.

An equally important aspect of managing DICOM imaging data in clinical trials is embedding intelligent, automated PHI and PII protection directly into the data management process. DICOM files carry more than images. They include metadata and tags, which can contain sensitive information such as patient names, dates of birth, medical record numbers, and facility identifiers. This sensitive information must be carefully managed before sponsors, CROs, or third-party stakeholders receive the data. Proactively verifying that PII and PHI are accurately identified and de-identified at the source is a critical best practice that safeguards patient privacy in compliance with HIPAA, GDPR, and ICH E6 guidelines. Automated de-identification tools that adhere to DICOM Supplement 142 and NEMA (National Electrical Manufacturers Association) standards reinforce data security and regulatory trust. They also preserve the full clinical and scientific value of imaging data, so trial teams can support confident, high-quality regulatory submissions.

To address these challenges, the Clario team built a comprehensive PHI/PII detection solution on AWS using Amazon Bedrock (Anthropic’s Claude Sonnet 4.5 on Amazon Bedrock) that combines automation, accuracy, and security throughout the clinical trial imaging workflow.

Why Amazon Bedrock and Amazon Textract

When evaluating options for building the solution, the Clario team chose to standardize on Amazon Bedrock and Amazon Textract for several key reasons:

  • Scalability without re-architecting: Amazon Bedrock and Amazon Textract provide scalability, reliability, and strong performance. The solution architecture can scale from a handful of documents to millions without re-architecting the solution or managing additional infrastructure. AWS manages the underlying capacity, so you can focus on building features instead of tuning servers or models.
  • Security and compliance: Keeping customer data secure is non-negotiable. By using Amazon Bedrock and Amazon Textract within Clario managed AWS accounts, processing remains inside a hardened AWS environment, taking advantage of AWS Identity and Access Management (IAM), Amazon Virtual Private Cloud controls, and encryption at rest and in transit. The Clario team can align closely with its organization’s security, compliance, and data residency requirements.
  • Managed foundation models: With Amazon Bedrock, the Clario team can access a range of high-quality foundation models as a fully managed service, without having to manage model training, hosting, or updates. This shortens the time to market, and you can iterate quickly as new models and capabilities become available in Amazon Bedrock.
  • Purpose-built OCR and document processing: Amazon Textract provides purpose‑built optical character recognition (OCR) and intelligent document processing, which significantly improves accuracy over traditional OCR engines. Its ability to automatically detect and extract text, tables, and key‑value pairs from complex documents and images reduces the amount of custom parsing logic that must be maintained.
  • End-to-end observability: Running on AWS provides end-to-end observability across the logs, metrics, and traces through services like Amazon CloudWatch and AWS CloudTrail. The Clario team can enforce governance policies, audit permissions, and track model and document processing usage centrally.
  • Extensible AI foundation: Because the solution builds on Amazon Bedrock and Amazon Textract, the Clario team can adopt new models and document processing features as they become available, without re-architecting.

Solution overview

The solution is built entirely on AWS, designed to bring greater efficiency, accuracy, and security to the detection of PHI and PII embedded within DICOM files. Accessible through Amazon API Gateway with TLS encryption in transit, IAM-backed authorization, and rate limiting, the detection workflow is readily consumable by multiple downstream systems with minimal integration effort.

Clinical trial sites store their DICOM images in Amazon Simple Storage Service (Amazon S3). The detection workflow retrieves each file from that bucket and processes it through the detection pipeline, so every ingestion step is logged and auditable for clinical trial security and compliance reviews. The workflow scans both standard and custom private DICOM metadata tags for PHI and PII. This covers the vendor-specific and non-standard tags where sensitive information often hides. Supporting both DICOM (.dcm) and PDF file formats, the solution is well-positioned to address PHI detection needs across the most used file types in clinical trial workflows.

The Clario AI team made a few deliberate design decisions early on. They ran the backend on Amazon Elastic Kubernetes Service (Amazon EKS) because a single DICOM series can span thousands of slices, and the detection workload is long-running and memory-intensive. They chose Amazon Relational Database Service (Amazon RDS) for PostgreSQL to persist processing metadata because the audit trail needs relational queries and strong consistency for compliance reporting. And they put the service behind Amazon API Gateway so that authentication, API-key management, and rate limiting are handled at the edge, keeping the backend focused on detection.

The following diagram and steps show how a DICOM document moves from upload through detection to structured output:

Architecture diagram showing DICOM images uploaded to Amazon S3, requests routed through Amazon API Gateway to detection on Amazon EKS using Amazon Textract and Amazon Bedrock, with metadata stored in Amazon RDS

Figure 1: Solution architecture for DICOM image ingestion, detection pipeline, and data retention workflow

The following steps describe the data flow through the solution, as shown in the architecture diagram:

  1. A consumer, such as an upstream imaging application, first uploads the DICOM image document to an Amazon S3 bucket location that is accessible to the solution.
  2. The consumer then calls the Clario Internal API running on Amazon API Gateway, providing their consumer-specific API key and the location of the document. This call initiates the DICOM image analysis workflow.
  3. Amazon API Gateway fronts the API and receives the incoming request. API Gateway validates the API key and, on success, forwards the request to the detection backend endpoint running on Amazon EKS to initiate processing.
  4. The solution performs initial checks on the file location (for example, URL format, access, and basic metadata) and then begins the PII/PHI identification process. The pipeline retrieves the file from the source S3 bucket and ingests it into the detection workflow.
  5. The file is stored in an internal Amazon S3 bucket and relevant metadata persisted in a PostgreSQL database on Amazon RDS to support downstream processing and auditability.
  6. The workflow invokes Amazon Textract to perform OCR and intelligent document parsing. Textract extracts text, tables, and form fields from the uploaded document, returning a structured representation of the content.
  7. The OCR output is then passed to a large language model (Anthropic’s Claude Sonnet 4.5 on Amazon Bedrock) that is configured to analyze the extracted text and identify potential PII/PHI elements. The model evaluates the content and associates each detected sensitive element with its position in the document.
  8. Once analysis is complete, the detection workflow returns a structured response to the consumer, containing the coordinates and related metadata for each piece of sensitive data identified in the document. The consumer can take follow‑up actions, such as redaction or masking.
  9. To minimize data exposure and support compliance requirements, the ingested files and associated records are retained only for a limited window. The document stored in Amazon S3 is automatically deleted based on an Amazon S3 lifecycle retention policy, and corresponding records in Amazon RDS are removed via a scheduled cleanup job.

Deep image analysis and detection workflow

Beyond metadata, the solution uses Anthropic’s Claude Sonnet 4.5 on Amazon Bedrock to perform a deep scan of the actual image pixel content, detecting PHI or PII that may be physically burned into the image itself. This includes patient names, dates of birth, and patient IDs across every individual slice within a DICOM series that can span thousands of images.

When PHI or PII is identified, the solution precisely captures the spatial coordinates and type of sensitive information detected, passing these bounding box details downstream to integrated systems responsible for the actual pixel-level redaction. Separating detection from masking was a deliberate design decision. It preserves flexibility, supports full auditability, and a human-in-the-loop can review the flagged findings before redaction is applied.

Diagram separating AI-powered PHI and PII detection from human-supervised quality control review and pixel-level redaction

Figure 2: Deep image analysis and detection workflow showing the separation between AI-powered detection and human-supervised redaction

The detection solution returns structured coordinates for the identified PHI and PII, spanning burnt-in pixel text, standard DICOM header fields, and non-standard custom tags. The solution then hands these results off to two downstream processes. In the quality control (QC) flow, qualified reviewers validate the flagged findings and confirm which items require remediation. In the redaction flow, the system executes the appropriate action for each type of PHI identified: masking or overwriting burnt-in text within the image pixel data, or stripping and zeroing out sensitive DICOM metadata tags.

This separation of detection and redaction is intentional. The AI-powered detection solution focuses on comprehensive, high-recall identification across thousands of image slices and metadata fields. The redaction flow retains human oversight over the irreversible act of modifying clinical data, confirming that no PHI is left exposed and no clinically relevant information is removed.

Sample DICOM image with detected PHI regions marked by bounding boxes

Figure 3: Sample DICOM image with detection

With the detection pipeline in place, the next step was to measure how accurately it identifies PHI and PII across real-world clinical documents.

Evaluation methodology

The team validated the solution in three stages: building a representative test dataset, creating ground truth annotations, and running an automated evaluation pipeline.

Building a representative test dataset

The Clario generative AI team partnered with internal stakeholders to assemble a diverse dataset, including:

  • DICOM images with burned‑in annotations, overlays, and metadata.
  • PDF documents such as reports and clinical summaries.

The dataset intentionally included documents that do and do not contain sensitive PII/PHI, allowing the team to measure both the model’s ability to detect sensitive information and its ability to avoid false alarms.

Creating ground truth annotations

For each document in the dataset, ground truth labels were generated that capture:

  • The exact text corresponding to each PII/PHI element.
  • The bounding box coordinates for each element on the page where available.

These annotations form the “gold standard” that the Clario team can use to compare the output from the production pipeline.

Automated evaluation pipeline

The Clario team implemented a set of evaluation scripts that:

  1. Run the full solution on each test DICOM or PDF, using the same workflow that powers the production system.
  2. Collect the model predictions, including the detected PII/PHI text, associated labels (for example, name, date of birth, medical record number), and coordinates where available.
  3. Compare predictions against ground truth using the following matching strategy:
    1. For PDFs and DICOM images where coordinates are available, a match is performed by spatial proximity, treating a predicted bounding box as a correct match if it falls within a configurable tolerance (by default, 3 pixels for each element of the bounding box).
    2. For DICOM metadata where coordinates are not available, a match is performed by a structured path.

Furthermore, the solution includes automated accuracy and performance (run time) checks, to improve system reliability across deployments. After validating the solution’s accuracy, the team assessed how it improves detection coverage across Clario clinical trial imaging workflows.

Results and benefits

The automated evaluation pipeline measured the solution’s detection performance against the manually annotated ground truth dataset across all three detection surfaces:

Detection surface Detection F1 Label accuracy
PDF text 0.9775 98.12%
DICOM burned-in image text 0.9750 96.15%
DICOM metadata tags 0.9951 99.60%

Detection F1 measures how accurately the solution identifies PHI/PII instances. Label accuracy measures how correctly it classifies the type of identified PHI/PII (for example, person_name, date_of_birth, or gender).

These results demonstrate consistently high detection performance across all three data surfaces, with metadata tag detection achieving near-perfect accuracy. The solution meets the Clario generative AI team’s production-readiness bar for deployment in clinical trial workflows where compliance accuracy is non-negotiable.

Complete detection coverage

Manual QC reviewers bring deep domain expertise to PHI identification. But modern clinical trials generate an enormous volume of data: thousands of image slices per series, each with dozens of metadata tags, including non-standard vendor-specific fields. This volume makes exhaustive manual review impractical at scale. The automated solution extends that human expertise across the full dataset.

In internal testing conducted by the Clario team, the solution scanned 100% of image slices, standard DICOM header fields, and custom private tags in the test dataset. This comprehensive coverage complements the existing QC process by surfacing PHI occurrences that might otherwise require additional review passes, particularly in non-standard private tags and burned-in pixel text where sensitive data is less predictable.

Risk and compliance impact

By automating PII and PHI detection across metadata tags and image slices, the solution can strengthen an organization’s compliance posture against HIPAA, GDPR, and ICH E6 requirements.

Beyond the measurable results, the project surfaced several insights that can guide other organizations building similar solutions.

The AWS collaboration

The AWS Solutions Architecture team partnered with the Clario AI team throughout the design and optimization of the detection solution. Key areas of collaboration included:

  • Scaling and throughput optimization: Provided prescriptive guidance on Amazon EKS pod scaling strategy to handle DICOM series with thousands of slices per request without timeout or memory pressure and tuned concurrent invocations to Amazon Bedrock to maximize throughput within account-level quotas.
  • Cost-efficient inference architecture: Recommended batching strategies for Amazon Textract API calls and optimized prompt token usage for Claude Sonnet on Amazon Bedrock to reduce per-document inference cost at scale.
  • Data retention and security controls: Recommended auto-deletion workflow using Amazon S3 Lifecycle policies and Amazon RDS scheduled jobs to meet HIPAA and GDPR data minimization requirements.

Lessons learned and best practices

Throughout the development and deployment of this solution, several valuable insights emerged that can benefit other organizations implementing similar AI-powered PHI detection systems for clinical trial imaging data.

Evaluate models against production-representative data

The Clario AI team adopted a rigorous model evaluation process early in development. Many open-source frameworks and off-the-shelf detection models demonstrated acceptable performance on curated test samples but experienced significant accuracy degradation when exposed to the full variability of production data. This variability includes diverse imaging modalities, vendor-specific private tags, and inconsistent burned-in text formatting across globally distributed trial sites. This reinforced the importance of evaluating any AI model at realistic, production-level data volumes before adoption. The solution that proved most effective was a carefully tuned pipeline where Amazon Textract handles text extraction and Claude Sonnet on Amazon Bedrock performs PHI/PII classification, with prompt engineering optimized for the specific patterns found in clinical trial DICOM data.

Ground truth data is non-negotiable

Building a reliable, automated evaluation pipeline required the manual creation of a ground truth dataset. The team acknowledges this process is time-consuming but necessary. This highlighted a best practice that is frequently underestimated: investing in high-quality, manually validated ground truth data is a prerequisite for developing and maintaining a trustworthy automated detection system. Attempting to shortcut this step risks deploying a solution whose real-world accuracy remains unknown, an unacceptable risk in the context of clinical trial compliance.

Separating detection from masking improves flexibility and auditability

The Clario team deliberately separated the PHI/PII detection function from the actual pixel-level redaction. Rather than performing masking directly, the solution identifies the precise coordinates and type of PHI/PII detected, passing this structured output downstream to integrated systems responsible for redaction. This separation proved to be a sound best practice. It preserves workflow flexibility, a human expert can review the findings before anyone makes irreversible changes to the image data, and keeps human accountability and auditability clear at every step.

Human-in-the-loop review remains an essential safeguard

Automation accelerates the detection and flagging process, but a key lesson learned is that human oversight should remain an integral part of the workflow. Incorporating a human review step for flagged findings before masking makes sure that edge cases and model uncertainties are appropriately handled. In the context of clinical trial data, where accuracy and regulatory accountability are paramount, this human-in-the-loop approach provides an essential layer of quality assurance that purely automated systems alone cannot fully replace.

Conclusion

The Clario automated PHI/PII detection solution demonstrates how AWS services can transform clinical trial imaging workflows by combining speed, accuracy, and compliance. By replacing manual spot-checks with automated scanning of every slice and metadata tag, the solution delivers complete PHI/PII detection coverage, reducing the risk of missed detections while strengthening compliance with HIPAA, GDPR, and ICH E6 requirements.

The key architectural decisions, comprehensive coverage of custom private tags, separation of detection from redaction, and human-in-the-loop validation provide a blueprint for other organizations managing sensitive imaging data in regulated environments. These lessons learned highlight that successful automation in clinical trials requires not just advanced technology, but thoughtful design that balances efficiency with the rigorous quality standards that patient safety and regulatory compliance demand.

Next steps

Organizations looking to implement similar PHI/PII detection capabilities for clinical trial imaging can start by:

  • Evaluate current manual review processes: Map where reviewers spend the most time and where missed PHI poses the greatest compliance risk.
  • Assess custom private DICOM tags: Catalog vendor-specific and site-specific tags across your imaging network to define the full detection scope.
  • Build ground truth datasets: Annotate a representative sample with precise PHI labels to benchmark automated detection accuracy.
  • Design human-in-the-loop workflows: Define review checkpoints where qualified personnel validate flagged findings before redaction.

About the authors