Tag Archives: Customer Solutions

How CyberArk uses Apache Iceberg and Amazon Bedrock to deliver up to 4x support productivity

Post Syndicated from Moshiko Ben Abu original https://aws.amazon.com/blogs/big-data/how-cyberark-uses-apache-iceberg-and-amazon-bedrock-to-deliver-up-to-4x-support-productivity/

This post is co-written with Moshiko Ben Abu, Software Engineer at CyberArk.

CyberArk achieved up to 95% reduction in case resolution time using Amazon Bedrock and Apache Iceberg.

This improvement addresses a challenge in technical support workflow: when a support engineer receives a new customer case, the biggest bottleneck is often not diagnosing the problem but preparing the data. Customer logs arrive in different formats from multiple vendors, and each new log format typically requires manual integration and correlation before an investigation can begin. For simple cases, this process can take hours. For more complex investigations, it can take days, slowing resolution and reducing overall engineer productivity.

CyberArk is a global leader in identity security. Centered on intelligent privilege controls, it provides comprehensive security for human, machine, and AI identities across business applications, distributed workforces, and hybrid cloud environments.

In this post, we show you how CyberArk redesigned their support operations by combining Iceberg’s intelligent metadata management with AI-powered automation from Amazon Bedrock. You’ll learn how to simplify data processing flows, automate log parsing for diverse formats, and build autonomous investigation workflows that scale automatically.

To achieve these results, CyberArk needed a solution that could ingest customer logs, automatically structure them, establish relationships between related events, and make everything queryable in minutes, not days. The architecture had to be serverless to handle unpredictable support volumes, secure enough to protect customer Personally Identifiable Information (PII), and fast enough to allow same day case resolution.

The legacy architecture: Bottlenecks and manual workflows

When support engineers received customer cases, they would upload log files to the data lake stored in Amazon Simple Storage Service (Amazon S3). The original design then suffered from the complexity of multi-step raw data processing.

First, CyberArk’s custom parsing logic running on AWS Fargate would parse these uploaded log files and transform the raw data. During this stage, the system also had to scan for PII and mask sensitive data to protect customer privacy.

Next, a separate process converted the processed data into Parquet format.

Finally, AWS Glue crawlers were required to discover new partitions and update table metadata for processed Parquet files. This dependency became the most complex and time-consuming part of the pipeline. Crawlers ran as asynchronous batch jobs rather than in real time, often introducing delays of minutes to hours before support engineers could query the data.

But the inefficiency went deeper than just architectural complexity. CyberArk supports customers running diverse product environments across multiple vendors. Each vendor and product produces logs in different formats with unique schemas, field names, and structures. Adding support for a new vendor meant days of integration work to understand their log format and build custom parsers.

CyberArk Legacy Logs Ingestion Flow

Figure 1: Legacy log ingestion architecture diagram showing the flow from S3 upload through AWS Fargate processing with AWS Glue Crawler

Beyond ingestion, the investigation process itself was manual and time consuming. Support engineers would manually query data, correlate events across different log sources, search through product documentation, and piece together root cause analysis through trial and error. This process required deep product expertise and could take hours or days depending on issue complexity. The new architecture addresses these inefficiencies through three key innovations:

  1. Single stage serverless processing: AWS Fargate with PyIceberg directly creates Iceberg tables from raw logs in one pass, removing intermediate processing steps and crawler dependencies entirely.
  2. AI powered dynamic parsing: Amazon Bedrock automatically generates grok patterns for log parsing by analyzing file schemas, transforming what was once a manual, time consuming process into a fully automated workflow.
  3. Autonomous investigation with AI Agents: AI Agents autonomously perform complete root cause analysis by querying log data, analyzing product knowledge bases, identifying event flows, and recommending solutions, transforming hours of manual investigation into minutes of automated intelligence.

The solution: AI-powered automation meets single-stage Iceberg processing

The new system delivers zero touch log processing from upload to query. Support engineers simply upload customer log ZIP files to the system. Here’s where the transformation happens: CyberArk’s custom processing logic still runs on AWS Fargate, but now it uses Amazon Bedrock to intelligently understand the data.

Zero-touch log processing workflow

The system extracts sample log entries from the uploaded log files and sends them to Amazon Bedrock along with context about the log source and table schema from AWS Glue Data Catalog. Amazon Bedrock analyzes the samples, understands the structure, and automatically generates grok patterns optimized for the specific log format.

Grok patterns are structured expressions that define how to extract meaningful fields from unstructured log text. For example, the following grok pattern specifies that a timestamp appears first, followed by a severity level, then a message body %{TIMESTAMP_ISO8601:timestamp} %{LOGLEVEL:severity} %{GREEDYDATA:message}

The system validates these grok patterns against additional samples to verify accuracy before applying them to parse the complete log file. Successfully validated grok patterns are stored in Amazon DynamoDB, creating a repository of known patterns. When the system encounters similar log formats in future uploads, it can retrieve these patterns directly from Amazon DynamoDB, avoiding redundant grok pattern generation. Amazon Bedrock processes log samples in real-time without retaining customer data or using it for model training, maintaining data privacy.

This entire process invokes Claude 3.7 Sonnet model from Amazon Bedrock and is orchestrated by AWS Fargate tasks with retry logic for reliability. The processing uses these AI-generated grok patterns to parse the logs and create or update Iceberg tables using PyIceberg APIs without human intervention.

This automation reduced logs onboarding time from days to minutes, enabling CyberArk to handle diverse customer environments without manual intervention.

Figure 2: Log ingestion architecture diagram showing the flow from S3 upload through AWS Fargate processing with Amazon Bedrock integration to Iceberg table creation

Figure 2: Log ingestion architecture diagram showing the flow from S3 upload through AWS Fargate processing with Amazon Bedrock integration to Iceberg table creation

Apache Iceberg: Simplified architecture, faster queries

Iceberg simplified and improved CyberArk’s data lake architecture by addressing the two primary bottlenecks in the legacy system: slow schema management and inefficient query performance.

Built-in schema evolution removes crawler dependency

In the legacy architecture, AWS Glue crawlers became a source of operational overhead and latency. Even when triggered on demand, crawlers ran as batch jobs over S3 prefixes to discover partitions and update metadata. As data volumes grew and datasets diversified across vendors and schemas, teams had to manage and operate a growing number of crawler jobs. The resulting delays, often ranging from minutes to hours, slowed data availability and downstream investigation workflows.

Iceberg removes this entire layer of complexity. Iceberg’s intelligent metadata layer automatically tracks table structure, schema changes, and partition information as data is written. When CyberArk’s processing creates or updates Iceberg tables through PyIceberg, the metadata is updated instantly and atomically. There’s no waiting for crawlers jobs to complete, and no risk of stale metadata. The moment data is written, it’s immediately queryable in Amazon Athena.

PyIceberg: Making Iceberg accessible beyond Apache Spark

Working with Iceberg usually involved Apache Spark and the complexity of distributed data processing. PyIceberg changed that by letting CyberArk create and manage Iceberg tables using a simple Python library. CyberArk’s data engineers could write straightforward Python code running on AWS Fargate to create Iceberg tables directly from parsed logs, without spinning up Spark clusters.

This accessibility was essential for CyberArk’s serverless architecture. PyIceberg enabled single stage processing where AWS Fargate tasks could parse logs, apply PII masking, and create Iceberg tables in one pass. The result was simpler code and lower operational overhead.

Metadata-driven query optimization delivers speed

In addition to removing crawlers, Iceberg significantly improved query performance through its intelligent metadata architecture. Iceberg maintains detailed statistics about data files, including min/max values, null counts, and partition information. When support engineers query data in Athena, Iceberg’s metadata layer supports partition pruning and file skipping, making sure queries only read the specific files containing relevant data. For CyberArk’s use case, where tables are partitioned by case ID, this means a query for a specific support case only reads the files for that case, ignoring potentially thousands of irrelevant files. This metadata driven optimization reduced query execution time from minutes to seconds, allowing support engineers to interactively explore data rather than waiting for results.

ACID transactions maintain data consistency

In a multi user support environment where multiple engineers may be analyzing overlapping cases or uploading logs simultaneously, data consistency is essential. Iceberg’s ACID transaction support helps verify that concurrent writes do not corrupt data or create inconsistent states. Each table update is atomic, isolated, and durable, providing the reliability CyberArk needed for production support operations.

Time travel enables historical analysis

Iceberg’s built-in versioning allows support engineers to query historical states of data, essential for understanding how customer issues evolved over time. If an engineer needs to see what the logs looked like when a case was first opened versus after a customer applied a patch, Iceberg’s time travel capabilities make this straightforward. This feature proved essential for complex troubleshooting scenarios where understanding the timeline of events was critical to resolution.

Automated table optimization with AWS Glue

Iceberg tables require periodic maintenance to maintain query performance.

CyberArk enabled AWS Glue automatic table optimization for their Iceberg tables, which handles compaction and expired snapshot cleanup in the background.

For CyberArk’s continuous upload workflow, this automation avoids performance degradation over time. Tables stay optimized without manual intervention from the engineering team.

AI Agents: Autonomous investigation workflow

While the Claude 3.7 Sonnet model from Amazon Bedrock automates grok pattern generation for log ingestion, the more advanced use of Amazon Bedrock comes in the investigation workflow. We use AI agents with Bedrock models to change how support engineers analyze and resolve customer issues.

From manual analysis to AI powered investigation

In the legacy workflow, support engineers would manually query data, correlate events across different log sources, search through product documentation, and piece together root cause analysis through trial and error. This process required deep product expertise and could take hours or days depending on issue complexity. AI Agents automate this entire investigation process. Support engineers use an internal portal to ask questions in natural language about customer issues, questions like
“Show me authentication errors for case 12345 in the last 24 hours”, “What were the most common errors across cases opened this week?” or “Compare the error patterns between case 12345 and case 12346.”

Behind the scenes, the system fires specialized AI Agents that autonomously perform thorough analysis.

How support agents work

Each AI Agent operates as an intelligent investigator with a clear mission: understand what happened, determine why it happened, and recommend how to fix it. When a support engineer asks a question, the agent collects relevant data by querying Athena to retrieve log data from Iceberg tables, filtering for the specific case and time period relevant to the investigation. The agent then accesses CyberArk’s internal knowledge base for the specific product involved, understanding known issues, common error patterns, and documented solutions. The agent then performs the following analysis:

  • Flow identification: Analyzes the sequence of events in the logs to understand what actually happened during the customer’s issue
  • Root cause determination: Correlates log events with product knowledge to identify the underlying cause of the problem
  • Solution recommendations: Suggests specific remediation steps based on the root cause analysis and known resolution patterns

This entire process happens in minutes, delivering advanced analysis that would have taken support engineers hours to perform manually.

For complex cases where a solution is not found, the support agent escalates to another, specialized agent that interacts with service engineers to collect additional inputs and expertise. This human-in-the-loop approach makes sure that even the most challenging cases receive appropriate attention while still benefiting from the automated investigation workflow. The insights gathered from these escalated cases are automatically fed back into CyberArk’s knowledge base, continuously improving the system’s ability to handle similar issues autonomously in the future.

Amazon Bedrock never shares customer data with model providers or uses it to train foundation models, case data and investigation insights remain within CyberArk’s environment.

Concurrent agent execution at scale

When multiple support engineers investigate different cases simultaneously, the solution runs specialized agents concurrently. CyberArk currently uses Claude 3.7 Sonnet as the foundation model for these agents. Each agent works independently on its assigned investigation, operating in parallel without resource contention. This concurrent execution allows the investigation workflow to scale automatically with support volume, handling peak loads without performance degradation.

AI-powered investigation advantage

This AI-powered investigation workflow delivers two key advantages.

Investigations that took hours now complete in minutes, enabling support engineers to resolve up to 4x more cases per day.

The system also creates a continuous learning feedback loop. When cases require manual resolution by engineers, these resolutions are automatically recorded and fed back into the knowledge base. Future investigations benefit from this accumulated expertise, with agents applying lessons learned from previous manual resolutions to similar cases. Amazon Bedrock doesn’t use customer data to train foundation models. Case data and investigation insights remain within CyberArk’s environment.
This automated feedback mechanism means the investigation workflow becomes more effective over time, continuously improving resolution accuracy and speed.

CyberArk - AI Powered Logs Investigation Flow

Figure 3: Investigation workflow diagram showing natural language query through AI Agents to Athena queries and knowledge base analysis

Scaling without proportional engineering growth

The business impact of this AI automation is significant. CyberArk can expand its vendor coverage and product portfolio without adding data engineering headcount. The same system that handles today’s log types will automatically handle tomorrow’s additions, whether that’s ten new formats or thousands, significantly reducing time to market for new product and vendor integrations.

The results: Significant improvements in resolution time and productivity

The transformation delivered measurable improvements across every key metric.

Resolution time: CyberArk achieved up to 95% reduction in time from case assignment to resolution. Simple cases that used to take 4 to 6 hours now take just 15 to 30 minutes. Complex cases that previously took up to 15 days are now completed in 2 to 4 hours.

Engineer productivity: Support engineers now handle 8 to 12 cases per day, compared to just 2 to 3 cases before. This means each engineer is helping up to 4x more customers.

Data availability: Logs are queryable within minutes of upload instead of waiting hours or days. Support engineers can start investigating issues almost immediately after receiving customer data.

Operational efficiency: The system requires zero manual intervention for new log formats or schema changes. Cases that used to require days of data engineering work now happen automatically.

Cost optimization: The serverless architecture alleviated idle infrastructure costs while scaling automatically with demand. CyberArk only pays for what they use, when they use it.

Customer satisfaction: Faster resolution times and proactive issue identification significantly improved the customer experience. Problems get solved in hours instead of days, and customers spend less time waiting for answers.

What’s next?

While AWS continues to innovate across both data lake management and agentic AI infrastructure, the following capabilities align well with CyberArk’s architecture and may offer additional operational benefits as the system scale.

Agent infrastructure maturity

As the agent-based architecture scales to handle thousands of concurrent investigations, CyberArk is transitioning to Amazon Bedrock AgentCore for future agent deployments. AgentCore provides a managed runtime for production AI agents with enhanced observability through AWS X-Ray integration, intelligent memory for context retention across sessions, and streamlined operational workflows. While the current AI Agents implementation delivers the performance and reliability CyberArk needs today, AgentCore represents a natural evolution path as operational requirements grow, offering framework-agnostic deployment, automatic scaling, and comprehensive monitoring capabilities without infrastructure management overhead.

Amazon S3 Tables

CyberArk’s current architecture uses Iceberg tables stored in Amazon S3 buckets. Amazon S3 Tables offers fully managed Iceberg tables with built-in optimization.

As CyberArk continue to scale with hundreds of Iceberg tables and rapid data growth, CyberArk is exploring a migration to Amazon S3 Tables to further reduce operational overhead.

S3 Tables remove the need to set up and monitor AWS Glue maintenance jobs. It automatically performs maintenance to enhance the performance of Iceberg tables, including unreferenced file removal, file compaction, and snapshot management. Additionally, S3 Tables provides Intelligent-Tiering that automatically moves data between storage classes based on access patterns, optimizing storage costs without manual intervention.

Because S3 Tables uses Iceberg open table format, migration would not require changes to existing Athena queries and PyIceberg code. This flexibility allows CyberArk to evaluate and adopt S3 Tables when the operational and cost benefits align with their business needs.

Conclusion

CyberArk’s transformation demonstrates how combining modern data lake architecture with AI automation can significantly change operational economics. By combining Iceberg’s intelligent metadata management with AI-powered automation from Amazon Bedrock, CyberArk transformed case resolution from days to minutes while enabling support operations to scale automatically with business growth. Support engineers now spend their time solving customer problems instead of wrangling data, customers receive faster resolutions, and the system scales automatically with the business.

To learn more about Iceberg on AWS, refer to Working with Amazon S3 Tables and table buckets and Using Apache Iceberg on AWS. To learn more about Amazon Bedrock AgentCore, refer to Amazon Bedrock AgentCore.


About the authors

Moshiko Ben Abu

Moshiko Ben Abu

Moshiko is a Software Engineer at CyberArk, specializing in architecting cloud-native applications and building AI-powered solutions. Moshiko advocates for a shift-left approach where security is built in from day one. His drive for innovation has been recognized across the company, earning him the Innovator culture award at CyberArk’s Global Kickoff.

Riki Nizri

Riki Nizri

Riki is a Solutions Architect at AWS. Collaborating with AWS ISV customers, Riki helps them leverage AWS services to build modern, efficient solutions that drive measurable business outcomes.

Sofia Zilberman

Sofia Zilberman

Sofia works as a Senior Streaming Solutions Architect at AWS, helping customers design and optimize real-time data pipelines using open-source technologies like Apache Flink, Kafka, and Apache Iceberg. With experience in both streaming and batch data processing, she focuses on making data workflows efficient, observable, and high-performing.

Verisk cuts processing time and storage costs with Amazon Redshift and lakehouse

Post Syndicated from Karthick Shanmugam, Srinivasa Are original https://aws.amazon.com/blogs/big-data/verisk-cuts-processing-time-and-storage-costs-with-amazon-redshift-and-lakehouse/

This post is co-written with Srinivasa Are, Principal Cloud Architect, and Karthick Shanmugam, Head of Architecture Verisk EES (Extreme Event Solutions).

Verisk, a catastrophe modeling SaaS provider serving insurance and reinsurance companies worldwide, cut processing time from hours to minutes-level aggregations while reducing storage costs by implementing a lakehouse architecture with Amazon Redshift and Apache Iceberg. If you’re managing billions of catastrophe modeling records across hurricanes, earthquakes, and wildfires, this approach eliminates the traditional compute-versus-cost trade-off by separating storage from processing power.

In this post, we examine Verisk’s lakehouse implementation, focusing on four architectural decisions that delivered measurable improvements:

  • Execution performance: Sub-hour aggregations across billions of records replaced long batch process
  • Storage efficiency: Columnar Parquet compression reduced costs without sacrificing response time
  • Multi-tenant security: Schema-level isolation enforced complete data separation between insurance clients
  • Schema flexibility: Apache Iceberg support column additions and historical data access without downtime

The architecture separates compute (Amazon Redshift) from storage (Amazon S3), demonstrating how to scale from billions to trillions of records without proportional cost increases.

Current state and challenges

In Verisk’s world of risk analytics, data volumes grow at exponential rates. Every day, risk modeling systems generate billions of rows of structured and semi-structured data. Each record captures a micro-slice of exposure, event probability, or loss correlation. To convert this raw information into actionable insights at scale, experts need a data engine designed for high-volume analytical workloads.

Each Verisk model run produces detailed, high-granularity outputs that include billions of simulated risk factors and event-level results, multi-year loss projections across thousands of perils, and deep relational joins across exposure, policy, and claims datasets.

Running meaningful aggregations (such as, loss by region, peril, or occupancy type) over such high volumes created performance challenges.

Verisk needed to build a SQL service that could aggregate at scale in the fastest time possible and integrate into their broader AWS solutions, requiring a serverless, open, and performant SQL engine capable of handling billions of records efficiently.

Prior to this cloud-based release, Verisk’s risk analytics infrastructure operated on an on-premises architecture centered around relational database clusters. Processing nodes shared access to centralized storage volumes through dedicated interconnect networks. This architecture required capital investment in server hardware, storage arrays, and networking equipment. The deployment model required manual capacity planning and provisioning cycles, limiting the organization’s ability to respond to fluctuating workload demands. Database operations depended on batch-oriented processing windows, with analytical queries competing for shared compute resources.

Amazon Redshift and lakehouse architecture

Lakehouse architecture on AWS combines data lake storage scalability with data warehouse analytical performance in a unified architecture. This architecture stores vast amounts of structured and semi-structured data in cost-effective Amazon S3 storage while maintaining Amazon Redshift’s massively parallel SQL analytics.

Amazon Redshift is a fully managed, petabyte-scale cloud data warehouse service that delivers fast query performance using massively parallel processing (MPP) and columnar storage. Amazon Redshift eliminates the complexity of provisioning hardware, installing software, and managing infrastructure, keeping focus on deriving insights from their data rather than maintaining systems.

To meet their challenge, Verisk designed a hybrid data lakehouse architecture that combines the storage scalability of Amazon S3 with the compute power of Amazon Redshift. The following diagram shows the foundational compute and storage architecture that powers Verisk’s analytical solution.

Compute and Storage Layer

Architecture Overview

The architecture processes risk and loss data through three distinct stages within the lakehouse architecture, with comprehensive multi-tenant delivery capabilities to maintain isolation between insurance clients.

Amazon Redshift allows retrieving data directly from S3 using standard SQL for background processing. This solution collects detailed result outputs, join them with internal reference data, and executes aggregations over billions of rows. Concurrency scaling guarantees that hundreds of background analyses using multiple serverless clusters can run simultaneous aggregation queries.

The following diagram shows the architecture designed by Verisk

Architecture Design used by Verisk

Data ingestion and storage foundation

Verisk stores risk model outputs, location level losses, exposure tables, and model data in columnar Parquet format within Amazon S3. An AWS Glue crawler extracts metadata from S3 and feeds it into the lakehouse processing pipeline.

For versioned datasets like exposure tables, Verisk adopted Apache Iceberg, an open table format that addresses schema evolution and historical versioning requirements. Apache Iceberg provides transactional consistency through atomicity, consistency, isolation, durability ACID-compliant operations that maintain consistent snapshots during concurrent updates. Snapshot-based time travel allows data retrieval at previous points in time for regulatory compliance, audit trails, and model comparison with rollback capabilities. Schema evolution supports adding, dropping, or renaming columns without downtime or dataset rewrites. Incremental processing uses metadata tracking to process only changed data, reducing refresh times. Hidden partitioning and file-level statistics reduce I/O operations, improving aggregation performance. Engine interoperability allows accessing the same tables across Amazon Redshift, Amazon Athena, Spark, and other engines without data duplication.

Verisk built a foundation that combines S3’s cost-effectiveness with data management by adopting Apache Iceberg as open table format for this solution.

Three-stage processing pipeline

This pipeline orchestrates data flow from raw inputs to analytical outputs through three sequential stages. Pre-processing prepares and cleanses data, modeling applies risk calculations and analytics, and post-processing aggregates results for delivery.

  • Stage 1: Pre-processing transforms raw data into structured formats using Iceberg Tables and Parquet files, then processes it through Amazon Redshift Serverless for initial data cleaning and transformation.
  • Stage 2: Modeling takes place with a process built on AWS Batch the pre-processed data and applies advanced analytics and feature engineering. Results are stored in Iceberg Tables and Parquet files.
  • Stage 3: Aggregated Results are obtained during post-processing using Amazon Redshift Serverless, it produces the final analytical outputs in Parquet files, ready for consumption by end users.

Multi-tenant delivery system

The architecture delivers results to multiple insurance clients (tenants) through a secure, isolated delivery system that includes:

  • Amazon Quick Sight dashboards for visualization and business intelligence
  • Amazon Redshift as the data warehouse for querying aggregated results
  • AWS Batch for modelling processing.
  • AWS Secrets Manager to manage tenant-specific credentials
  • Tenant Roles implementing role-based access control to provide data isolation between clients

Summarized results are exposed through Amazon Quick Sight dashboards or downstream APIs to underwriting teams.

Multi-tenant security architecture

A critical requirement for Verisk’s SaaS solution was supporting comprehensive data and compute isolation between different insurance and reinsurance clients. Verisk implemented a comprehensive multi-tenant security model that provides isolation while maintaining operational efficiency.

Our solution implements an isolation strategy in two layers combining logical and physical separation. At the logical layer, each client’s data resides in dedicated schemas with access controls that prevent cross-tenant operations. Amazon Redshift Metadata security restricts tenants from discovering or accessing other clients’ schemas, tables, or database objects through system catalogs. At the physical layer, for larger deployments, dedicated Amazon Redshift clusters provide workload separation at the compute level, preventing one tenant’s analytical operations from impacting another’s performance. This dual approach meets regulatory requirements for data isolation in the insurance industry through schema-level isolation within clusters for standard deployments and complete compute separation across dedicated clusters for larger-scale implementations.

The implementation uses stored procedures to automate security configuration, maintaining consistent application of access controls across tenants. This defense-in-depth approach combines schema-level isolation, system catalog lockdown, and selective permission grants to create a security model.

For data architects interested in implementing similar multi-tenant architectures, review Implementing Metadata Security for Multi-Tenant Amazon Redshift Environment.

Implementation considerations

Verisk’s architecture reveals three decision points for companies building similar systems.

When to adopt open table formats

Apache Iceberg proved essential for datasets requiring schema evolution and historical versioning. Data engineers should evaluate open table formats when analytical workloads span multiple engines (Amazon Redshift, Amazon Athena, Spark) or when regulatory requirements demand point-in-time data reconstruction.

Multi-tenant isolation strategy

Schema-level separation combined with metadata security prevented cross-tenant data discovery without performance overhead. This approach scales more efficiently than database-per-tenant architectures while meeting insurance industry compliance requirements. Security experts should implement isolation controls during initial deployment rather than retrofitting them later.

Stored procedures or application logic

Redshift stored procedures standardized aggregation calculations across teams and constructed dynamic SQL queries. This approach works best when business logic changes frequently or when multiple teams need different aggregation dimensions on the same datasets.

Conclusion

Verisk’s implementation of Amazon Redshift Serverless with Apache Iceberg and lakehouse architecture shows how separating compute from storage addresses enterprise analytics challenges at billion-record scale. By combining cost-effective Amazon S3 storage with Redshift’s massively parallel SQL compute, Verisk achieved aggregations across billions of catastrophe modeling records, reduced storage costs through efficient parquet compression, and eliminated ingestion delays. Now underwriting teams can run ad-hoc analyses during business hours rather than waiting for long-running batch jobs. The combination of open standards like Apache Iceberg, serverless compute with Amazon Redshift, and multi-tenant security provides the scalability, performance, and cost efficiency needed for modern analytics workloads.

Verisk’s journey has positioned them to scale confidently into the future, processing not just billions, but potentially trillions of records as their model resolution increases.


About the authors

Karthick Shanmugam

Karthick Shanmugam

Karthick is Head of Architecture at Verisk EES. Focused on scalability, security, and innovation, he drives the development of architectural blueprints that align technology direction with business objectives. He is dedicated to building a modern, adaptable foundation that accelerates Verisk’s digital transformation and enhances value delivery across global platforms.

Srinivasa Are

Srinivasa Are

Srinivasa is a Principal Data Architect at Verisk EES, with extensive experience driving cloud transformation and data modernization across global enterprises. Known for combining deep technical expertise with strategic vision, Srini helps organizations unlock the full potential of their data through scalable cost-optimized architectures on AWS—bridging innovation, efficiency, and meaningful business outcomes.

Raks Khare

Raks Khare

Raks is a Senior Analytics Specialist Solutions Architect at AWS based out of Pennsylvania. He helps customers across varying industries and regions architect data analytics solutions at scale on the AWS platform. Outside of work, he likes exploring new travel and food destinations and spending quality time with his family.

Duvan Segura-Camelo

Duvan Segura-Camelo

Duvan is a Senior Analytics & AI Solutions Architect at AWS based out of Michigan, he helps customers architect scalable data analytics and AI solutions. With over two decades of experience in Analytics, Big Data and AI, Duvan is passionate about helping organizations build advanced, highly scalable solutions on AWS. Outside of work, he enjoys spending time with his family, staying active, reading, and playing the guitar.

Ashish Agrawal

Ashish Agrawal

Ashish is a Principal Product Manager with Amazon Redshift, building cloud-based data warehouses and analytics cloud services. Ashish has over 25 years of experience in IT. Ashish has expertise in data warehouses, data lakes, and platform as a service. Ashish has been a speaker at worldwide technical conferences.

How Zalando innovates their Fast-Serving layer by migrating to Amazon Redshift

Post Syndicated from original https://aws.amazon.com/blogs/big-data/how-zalando-innovates-their-fast-serving-layer-by-migrating-to-amazon-redshift/

While Zalando is now one of Europe’s leading online fashion destination, it began in 2008 as a Berlin-based startup selling shoes online. What started with just a few brands and a single country quickly grew into a pan-European business, operating in 27 markets and serving more than 52 million active customers.

Fast forward to today, and Zalando isn’t just an online retailer—it’s a tech company at its core. With more than €14 billion in annual gross merchandise volume (GMV), the company realized that to serve fashion at scale, it needed to rely on more than just logistics and inventory. It needed data. And not just to support the business—but to drive it.

In this post, we show how Zalando migrated their fast-serving layer data warehouse to Amazon Redshift to achieve better price-performance and scalability.

The scale and scope of Zalando’s data operations

From personalized size recommendations that reduce returns to dynamic pricing, demand forecasting, targeted marketing, and fraud detection, data and AI are embedded across the organization.

Zalando’s data platform operates at an impressive scale, managing over 20 petabytes of data in its lake supporting various analytics and machine learning applications. The data platform hosts more than 5,000 data products maintained by 350 decentralized teams, serving 6,000 monthly users, representing 80% of Zalando’s corporate workforce. As a fully self-service data platform, it provides SQL analytics, orchestration, data discovery, and quality monitoring, empowering teams to build and manage data products independently.

This scale only made the need for modernization more urgent. It was clear that efficient data loading, dynamic compute scaling, and future-ready infrastructure were essential.

Challenges with the existing Fast-Serving Layer (data warehouse)

To enable decisions across analytics, dashboards, and machine learning, Zalando uses a data warehouse that acts as a fast-serving layer and backbone for critical data/reporting use cases. This layer holds about 5,000 curated tables and views, optimized for quick, read-heavy workloads. Every week, more than 3,000 users—including analysts, data scientists, and business stakeholders—rely on this layer for instant insights.

But the incumbent data warehouse wasn’t future proof. It was based on a monolithic cluster setup optimized for peak loads, like Monday mornings, when weekly and daily jobs pile up. As a result, 80% of the time, the system sat underutilized, burning compute and leading to substantial “slack costs” from over-provisioned capacity, with potential monthly savings of over $30,000 if dynamic scaling were possible. Concurrency limitations resulted in high latency and disrupted business-critical reporting processes. The system’s lack of elasticity led to poor cost-to-utilization ratios, while the absence of workload isolation between teams frequently caused operational incidents. Maintenance and scaling required constant vendor support, making it difficult to manage peak periods like CyberWeek due to instance scarcity. Additionally, the platform lacked modern features such as online query editors and proper auto scaling capabilities, while its slow feature development and limited community support further hindered Zalando’s ability to innovate.

Solving for scale: Zalando’s journey to a modern fast serving layer

Zalando was looking for a solution that demonstrated capabilities which could meet their cost and performance targets through a “simple lift and shift” approach. Amazon Redshift was selected for the POC to address autoscaling and concurrency needs, while simultaneously reducing operational efforts as well as its ability to integrate with Zalando’s existing data platform and align with their overall data strategy.

The overall evaluation scope for the Redshift assessment covered following key areas.

Performance and cost

The evaluation of Amazon Redshift demonstrated substantial performance improvements and cost benefits compared to the old data warehousing platform.

  • Redshift offered 3-5 times faster query execution time.
  • Approximately 86% of distinct queries ran faster on Redshift.
  • In a “Monday morning scenario”, Redshift demonstrated 3 times faster accumulated execution time compared to the existing platform
  • For short queries, Redshift achieved 100% SLA compliance for queries in the 80-480 second range. For queries up to 80 seconds, 90% met SLA.
  • Redshift demonstrated 5x faster parallel query execution, handling significantly higher concurrent queries than the current data warehouse’s maximum parallelism.
  • For Interactive Usage use cases, Redshift demonstrated strong performance, which is essential for BI tool users, especially in parallel executions scenario.
  • Redshift features such as Automatic Table Optimizations and Automated Materialized views eliminated the need for data producing teams to manually optimize the design of tables, making it highly suitable for a central service offering.

Architecture

Redshift successfully demonstrated workload isolation such as separating transformations(ETL) from serving (BI, Ad-hoc etc.) workload using Amazon Redshift data sharing. It also proved its versatility through integration with Spark and common file formats was also proven.

Security

Amazon Redshift successfully demonstrated end-to-end encryption, auditing capabilities, and comprehensive access controls with Row-Level and Column-Level Security as part of the proof of concept.

Developer productivity

The evaluation demonstrated significant improvements in developer efficiency. A baseline concept for central deployment template authoring and distribution via AWS Service Catalog was successfully implemented. Additionally, Redshift showed impressive agility with its ability to deploy Redshift Serverless endpoints in minutes for ad-hoc analytics, enhancing the team’s ability to quickly respond to analytical needs.

Amazon Redshift migration strategy

This section outlines the approach Zalando took to migrate the fast-serving layer to Amazon Redshift.

From monolith to modular: Redesigning with Redshift

The migration strategy involved a complete re-architecture of the fast-serving layer, moving to Amazon Redshift with a multi-warehouse model that separates data producers from data consumers.Key components and principles of the target architecture include:

  1. Workload Isolation: Use cases are isolated by instance or environment, with data shares facilitating data exchange between them. Data shares enable an “easy fan out” of data from the Producer warehouse to various Consumer warehouses. The producer and consumer warehouses can be either Provisioned (such as for BI Tools) or Serverless (such as for Analysts). This allows for data sharing between separate legal entities.
  2. Standardized Data Loading: A Data Loading API (proprietary to Zalando) was built to standardize data loading processes. This API supports incremental loading and performance optimizations. Implemented with AWS Step Functions and AWS Lambda, it detects changed Parquet files from Delta lake metadata and uses Redshift spectrum for loading data into the Redshift Producer warehouse.
  3. Using Redshift Serverless: Zalando aims to use Redshift Serverless wherever possible. Redshift Serverless offers flexibility, cost efficiency, and improved performance, particularly for the lightweight queries prevalent in BI dashboards. It also enables the deployment of Redshift serverless endpoints in minutes for ad-hoc analytics, enhancing developer productivity.

The following diagram depicts Zalando’s end-to-end Amazon Redshift multi-warehouse architecture, highlighting the producer-consumer model:

Architecture Diagram

The core strategy of migration was “lift-and-shift” in terms of code to avoid complex refactoring and meet deadlines.

The main principles used were:

  • Run tasks in parallel whenever possible.
  • Minimize the workload for internal data teams.
  • Decouple tasks to allow teams to schedule work flexibly.
  • Maximize the work done by centrally managed partners.

Three-stage migration approach

The migration is broken down into three distinct stages to manage the transition effectively.

Stage 1: Data replication

Zalando’s priority was creating a complete, synchronized copy of all target data tables from the old data warehouse to Redshift. An automated process was implemented using Changehub, an internal tool built on Amazon Managed Workflows for Apache Airflow (MWAA), that monitors the old system’s logs and syncs data updates to Redshift approximately every 5-10 minutes, establishing the new data foundation without disrupting existing workflows.

Stage 2: Workload migration

The second stage focused on moving business logic (ETL) and MicroStrategy reporting to Redshift to significantly reduce the load on the legacy system. For ETL migration, semi-automated approach was implemented using Migvisor code convertor to convert the scripts. MicroStrategy reporting was migrated by leveraging MSTR’s capability to automatically generate Redshift-compatible queries based on the semantic layer.

Stage 3: Finalization and decommissioning

The final stage completes the transition by migrating all remaining data consumers and ingestion processes, leading to the full shutdown of the old data warehouse. During this phase, all data pipelines are being rerouted to feed directly into Redshift, and long-term ownership of processes is being transitioned to the respective teams before the old system is fully decommissioned.

Benefits and Results

A major infrastructure change at Zalando occurred on October 30, 2024, switching 80% of analytics reporting from the old data warehouse solution to Redshift. The migration of 80% of analytics reporting to Redshift successfully reduced operational risk for the critical Cyber Week period and enabled the decommissioning of the old data warehouse to avoid significant license fees.

The project resulted in substantial performance and stability improvements across the board.

Performance Improvements

Key performance metrics demonstrate substantial improvements across multiple dimensions:

  • Faster Query Execution: 75% of all queries now execute faster on Redshift.
  • Improved Reporting Speed: High-priority reporting queries are significantly faster, with a 13% reduction in P90 execution time and a 23% reduction in P99 execution time.
  • Drastic Reduction in System Load: The overall processing time for MicroStrategy (MSTR) reports has dramatically decreased. Peak Monday morning execution time dropped from 130 minutes to 52 minutes. In the first four
  • weeks, the total MSTR job duration was reduced by over 19,000 hours (equivalent to 2.2 years of compute time) compared to the previous system. This has led to far more consistent and reliable performance.

The following graph shows one of the critical Monday Morning Workload elapsed duration on old-data warehouse as well as Amazon Redshift.

Critical Monday Morning Workload elapsed duration on old-data warehouse as well as Amazon Redshift

Operational stability

Amazon Redshift has proven to be significantly more stable and reliable, successfully meeting the key objective of reducing operational risk.

  • Report Timeouts: Report timeouts, a primary concern, have been virtually eliminated.
  • Critical Business Period Performance: Redshift performed exceptionally well during the high-stress Cyber Week 2024. This is a stark contrast to the old system, which suffered critical, financially impactful failures during the same period in 2022 and 2023.
  • Data Loading: For data producers, the consistency of data loading is critical, as delays can hold up numerous reports and cause direct business impact. The system relied on an “ETL Ready” event, which triggers report processing only after all required datasets have been loaded. Since the migration to Redshift, the timing of this event has become significantly more consistent, improving the reliability of the entire data pipeline.

The following diagram shows consistency in ETL Ready event, after migrating to Amazon Redshift

ETL Ready Event Execution times

End user experience

The reduction in total execution time of Monday morning loads has resulted in dramatically improved end-user productivity. This is the time needed to process the full batch of scheduled reports (peak load), which directly translates to wait times and productivity for end users, since this is when most users need their weekly reports for their business. The following graphs shows typical Mondays before and after the switch and how Amazon Redshift handles the MSTR queue providing much better end user experience.

MSTR queue on 28/10/2024 (before switch)MSTR queue on 28/10/2024 (before switch)

MSTR queue on 02/12/25 (after switch)MSTR queue on 02/12/25 (after switch)

Learnings and unforeseen challenges

Navigating automatic optimization in a multi-warehouse architecture

One of the most significant challenges Zalando encountered during migration involves Redshift’s multi-warehouse architecture and its interaction with automatic table maintenance. The Redshift architecture is designed for workload isolation: a central producer warehouse for data loading, and multiple consumer warehouses for analytical queries. Data and associated objects reside only on the producer and are shared via Redshift Datashare.

The core issue: Redshift’s Automatic Table Optimization (ATO) operates exclusively on the producer warehouse. This extends to other performance features like Automatic Materialized Views and automatic query rewriting. Consequently, these optimization processes were unaware of query patterns and workloads on consumer warehouses. For instance, MicroStrategy reports running heavy analytical queries on the consumer side were outside the scope of these automated features. This led to suboptimal data models and significant performance impacts, particularly for tables with AUTO-set distribution and sort keys.

To address this, two-pronged approach was implemented:

1. Collaborative manual tuning: Zalando worked closely with the AWS Database Engineering team, who provide holistic performance checks and tailored recommendations for distribution and sort keys across all warehouses.

2. Scheduled table maintenance: Zalando implemented a daily VACUUM process for tables with over 5% unsorted data, ensuring data organization and query performance.

Additionally, following data distribution strategy was implemented:

  1. KEY Distribution: Explicitly defined DISTKEY for tables with clear JOIN conditions.
  2. EVEN Distribution: Used for large fact tables without clear join keys.
  3. ALL Distribution: Applied to smaller dimension tables (under 4 million rows).

This proactive approach has given better control over cluster performance and mitigated data skew issues. Zalando is encouraged that AWS is working to include cross-cluster workload awareness in a future Redshift release, which should further optimize multi-warehouse setup.

CTEs and execution plans

Common Table Expressions (CTEs) are a powerful tool for structuring complex queries by breaking them down into logical, readable steps. Analysis of query performance identified optimization opportunities in CTE usage patterns.

Performance monitoring revealed that Redshift’s query engine would sometimes recompute the logic for a nested or repeatedly referenced CTE from scratch every time it was called within the same SQL statement instead of writing the CTE’s result to an in-memory temporary table for reuse.

Two strategies proved effective in addressing this challenge:

  • Convert to a materialized view: CTEs used frequently across multiple queries or with particularly complex logic were converted into materialized views (MVs). This pre-compute the result, making the data readily available without re-running the underlying logic.
  • Use explicit temporary tables: For CTEs used multiple times within a single, complex query, the CTE’s result was explicitly written into a temporary table at the beginning of the transaction. For example, within MicroStrategy, the “intermediate table type” setting was changed from the default CTE to “Temporary table.”

Implementation of either materialized views or temporary tables ensures the complex logic is computed only once. This approach eliminated the recomputation issue and significantly improved the performance of multi-layered SQL queries.

Optimizing memory usage by right-sizing VARCHAR columns

It may seem like a minor detail, but defining the appropriate length for VARCHAR columns can have a surprising and significant impact on query performance. This was discovered firsthand while investigating the root cause of slow queries that were showing high amounts of disk spill.

The issue stemmed from data loading API tool, which is responsible for syncing data from Delta Lake tables into Redshift. Because Delta Lake’s StringType datatype does not have a defined length, the tool defaulted to creating Redshift columns with a very high VARCHAR length (such as VARCHAR(16384)).

When a query is executed, the Redshift query engine allocates memory for in-transit data based on the column’s defined size, not the actual size of the data it contains. This meant that for a column containing strings of only 50 characters but defined as VARCHAR(16384), the engine would reserve a vastly oversized block of memory. This excessive memory allocation led directly to high disk spill, where intermediate query results overflowed from memory to disk, drastically slowing down execution.

To resolve this, a new process was implemented requiring data teams to explicitly define appropriate column lengths during object deployment. nalyzing the actual data and setting realistic VARCHAR sizes (such as VARCHAR(100) instead of VARCHAR(16384)), significantly improved memory usage, reduced disk spill, and boosted overall query speed. This change underscores the importance of precision in data definition for an optimized Redshift environment.

Future outlook

Central to Zalando strategy is the shift to a serverless-based warehouse topology. This move enables automatic scaling to meet fluctuating analytical demands, from seasonal sales peaks to new team projects, all without manual intervention. The approach allows data teams to focus entirely on generating insights that drive innovation, ensuring platform performance aligns with business growth.

As the platform scales, responsible management is paramount. The integration of AWS Lake Formation create a centralized governance model for secure, fine-grained data access, enabling safe data democratization across the organization. Simultaneously, Zalando is embedding a strong FinOps culture by establishing unified cost management processes. This provides data owners with a comprehensive, 360-degree view of their costs across Redshift’s services, empowering them with actionable insights to optimize spending and align it with business value. Ultimately, the goal is to ensure every investment in Zalando’s data platform is maximized for business impact.

Conclusion

In this post, we showed how Zalando’s migration to Amazon Redshift has successfully transformed its data platform, making it a more data-driven fashion tech leader. This move has delivered significant improvements across key areas including enhanced performance, increased stability, reduced operational costs, and improved data consistency. Moving forward, a serverless-based architecture, centralized governance with AWS Lake Formation, and a strong FinOps culture will continue to drive innovation and maximize business impact.

If you’re interested in learning more about Amazon Redshift capabilities, we recommend watching the most recent What’s new with Amazon Redshift session in the AWS Events channel to get an overview of the features recently added to the service. You can also explore the self-service, hands-on Amazon Redshift labs to experiment with key Amazon Redshift functionalities in a guided manner.

Contact your AWS account team to learn how we can help you modernize your data warehouse infrastructure.


About the authors

Srinivasan Molkuva

Srinivasan Molkuva

Srinivasan is an Engineering Manager at Zalando with over a decade and a half of expertise in the data domain. He currently leads the Fast Serving Layer team, having successfully managed the transition of critical systems that support the company’s entire reporting and analytical landscape.

Sabri Ömür Yıldırmaz

Sabri Ömür Yıldırmaz

Ömür is a Senior Software Engineer at Zalando, based in Berlin, Germany. Passionate about solving complex challenges across backend applications and cloud infrastructure, he specializes in the end-to-end lifecycle of critical data platforms, driving architectural decisions to ensure robustness, high performance, scalability, and cost-efficiency.

Prasanna Sudhindrakumar

Prasanna Sudhindrakumar

Prasanna is a Senior Software Engineer at Zalando, based in Berlin, Germany. Brings years of experience building scalable data pipelines and serverless applications on AWS. Passionate about designing distributed systems with a strong focus on cost efficiency and performance, with a keen interest in solving complex architectural and platform-level challenges.

Paritosh Kumar Pramanick

Paritosh Kumar Pramanick

Paritosh is a Senior Data Engineer at Zalando, based in Berlin, Germany. He has over a decade of experience spearheading data warehousing initiatives for multinational corporations. Expert in transitioning legacy systems to modern, cloud-native architectures, ensuring high performance, data integrity, and seamless integration across global business units.

Saman Irfan

Saman Irfan

Saman is a Senior Specialist Solutions Architect at Amazon Web Services, based in Berlin, Germany. Saman is passionate about helping organizations modernize their data architectures to drive innovation and business transformation.

Werner Gunter

Werner Gunter

Werner is a Principal Specialist Solutions Architect at Amazon Web Services, based in Berlin, Germany. As a seasoned data professional, he has helped large enterprises worldwide over the past 2 decades, to modernize their data analytics estates.

How Convera built fine-grained API authorization with Amazon Verified Permissions

Post Syndicated from Santhosh Veeraraman original https://aws.amazon.com/blogs/architecture/how-convera-built-fine-grained-api-authorization-with-amazon-verified-permissions/

Convera processes billions in cross-border payment volume yearly for businesses and financial institutions worldwide. As their platform grew, they needed a robust authorization system that could protect sensitive financial data while maintaining operational efficiency across their global network.

In this post, we share how Convera used Amazon Verified Permissions to build a fine-grained authorization model for their API platform.

Background

As Convera’s service offerings expanded, they needed a scalable, secure, and auditable way to enforce role-based and attribute-based access control. Their goal was to make sure users, both internal and external, had access only to the resources and actions they were explicitly authorized for, while maintaining flexibility to adapt to evolving business needs. Initially, Convera explored building an in-house access control solution. However, they realized that implementing policy management, real-time authorization, logging, and auditing from scratch would require significant engineering effort and ongoing maintenance, diverting resources from their core business priorities. Convera chose Verified Permissions for implementing fine-grained authorization for their payment APIs. This choice was driven by the following factors:

  • Direct integration with AWS services like Amazon Cognito and Amazon API Gateway
  • Cedar policy language’s flexibility in defining complex authorization rules
  • Ability to evaluate multiple attributes like user roles, transaction amounts, and geographic locations
  • High-performance characteristics with millisecond-level authorization decisions

Given its flexibility and scalability, Verified Permissions became the foundational reference architecture for managing access control across two main scenarios:

  • Fine-grained access control – Convera’s Payment platform serves diverse users including customers, internal staff, and machine-to-machine communications, each requiring specific entitlements based on their roles, organizational hierarchy, and context.
  • Multi-tenancy controls – One of Convera’s most complex requirements was enabling multi-tenant access control while enforcing strict data isolation. Verified Permissions make it possible to define policies that dynamically evaluated tenant ownership, user roles, and contextual attributes.

The following analysis breaks down Convera’s implementation approach across their major use cases

Fine-grained access control

Fine-grained access control is a critical aspect of application security that makes sure users have precisely defined permissions, granting access only to specific resources or actions within an application. Using Verified Permissions, you can define a schema in terms of entity type, including attributes relevant to the authorization model and the valid combinations of principal types, resource types, and actions. Verified Permissions uses this schema to validate that a static policy or policy template is consistent with the application’s authorization model.

Convera implemented Verified Permissions for fine-grained access control across multiple user types and interaction patterns.

Customer access management

At the UI level, Convera used Verified Permissions to manage API access based on specific user characteristics. For example, in their financial applications, the visibility of transaction features like the modify payment parameters is dynamically controlled based on Verified Permissions policies.

The following diagram illustrates the user authentication decision flow for financial transactions.

User Authorization Decision for financial transactions

Figure 1: User Authorization Decision for financial transactions

The following is an example Cedar policy that can be used in conjunction with Verified Permissions to achieve this use case. This policy makes sure only authorized users with specific roles can see sensitive financial controls. The same policy must be evaluated at the API level when the actual transfer request is made. At the service level, the policy is designed to provide fine-grained access controls.

permit (
    principal,
    action in [MyApp::Action::"ViewTransferButton"],
    resource
) when {
    principal.role == "PAYMENT_INITIATOR" &&
    resource.accountType == "BUSINESS" &&
    resource.status == "ACTIVE"     
};

You can integrate the application with Verified Permissions through the API to authorize user access requests. For each authorization request, the service retrieves the relevant policies and evaluates those policies to determine whether a user is permitted to take an action on a resource given context input such as users, roles, group membership, and attributes.

The following figure illustrates the end-to-end architectural diagram of this implementation.

Fine-grained User Authorization control with Verified PermissionsPermissions

Figure 2: Fine-grained User Authorization control with Verified Permissions

The workflow consists of the following steps:

  1. Users initiate login through the client application.
  2. The client authenticates with Amazon Cognito.
  3. Amazon Cognito triggers a pre-token generation AWS Lambda function to get user roles.
  4. The Lambda function fetches user roles from Amazon Relational Database Service (Amazon RDS).
  5. The Lambda function enriches a JSON Web Token (JWT) with user role information.
  6. The enriched JWT with user roles is returned to the client application.
  7. The client application makes an API call, sending an authorization request to API Gateway through the enriched JWT.
  8. A Lambda authorizer validates the JWT and the role permissions from the JWT and makes a call to Verified Permissions.
  9. Verified Permissions reads access policies stored as Cedar policies and makes an authorization decision.
  10. Verified Permissions returns the authorization result to the Lambda authorizer.
  11. The Lambda authorizer, based on the authorization result, sends an AWS Identity and Access Management (IAM) policy that allows or denies the request to API Gateway.
  12. API Gateway either allows or denies the request to the client application.
  13. API Gateway caches the IAM policy.

The policy governance is owned by Convera’s infosec team through a strictly regulated IAM role. The changes to the Cedar policies are captured using Amazon DynamoDB Streams and continuously synced with Verified Permissions.

To improve speed, Convera created a two-level cache system, using the API Gateway built-in cache for authorization decisions and application-level caching for Amazon Cognito tokens. The Lambda function invokes Verified Permissions to authorize the request. If Verified Permissions returns deny, the request is rejected, and an HTTP unauthorized response (403) is sent back. If Verified Permissions returns allow, the request is moved forward. This multi-level caching approach successfully delivers sub-millisecond response times while reducing operational costs and maintaining security controls.

Internal customer connect applications

Convera was able to reuse the same architecture for their internal user access as well, such as customer service associates who need quick access to client information to provide efficient service while protecting access to sensitive data. Using Verified Permissions, a role-specific Cedar policy was created to achieve the following:

  • Enable view-only access to basic customer profiles, including contact information and service history
  • Restrict edit capabilities to specific fields, such as updating contact preferences or logging support interactions
  • Block access to sensitive financial data or internal business metrics

In this flow, internal users authenticate through their enterprise identity provider (IdP), in this case Okta, through the Convera Connect App and obtain ID and access tokens from Amazon Cognito. Amazon Cognito, using a pre-token generation hook, customizes the access token with user attributes stored in Amazon DynamoDB. Although the Cedar policies are tailored for internal roles and responsibilities (different from customer-facing policies), the fundamental flow involving API Gateway, a Lambda authorizer, Verified Permissions policy evaluation, and decision caching remains identical. This architectural reuse meant Convera didn’t need to rebuild their authorization infrastructure, so they can use the same performance optimizations, security controls, and operational processes across both customer and internal user access patterns.

Extending the model to service communication

After successfully implementing Verified Permissions for customer and internal user access control, Convera recognized they could use the same architecture for securing service-to-service communications. Similar to how they manage user authentication through Amazon Cognito user pools, each client service is registered in their client configuration system, with Verified Permissions creating a dedicated policy store for service-specific permissions. Services authenticate through Amazon Cognito using client credentials (instead of user credentials) to obtain access tokens that carry service-specific attributes such as service identifier, tier, allowed operations, and rate limits.

The following diagram illustrates the machine-to-machine architecture for internal and external partner integration.

Machine to machine architecture for internal and external partner integration

Figure 3: Machine to machine architecture for internal and external partner integration

The workflow consists of the following steps:

  1. Service A sends an API request to the authentication token endpoint in API Gateway, including its access token.
  2. The request is forwarded to Amazon Cognito, which validates the token from the machine-to-machine user pool.
  3. Service A, now authenticated, makes a request to the business API (Service B) through API Gateway.
  4. API Gateway forwards the request to the Lambda authorizer, which processes the incoming request and extracts the service context from the token.
  5. The Lambda authorizer sends the authorization request to Verified Permissions to evaluate it against stored Cedar policies. The evaluation considers:
    1. Service identity
    2. Requested operation
    3. Resource context
    4. Environmental factors
  6. Verified Permissions returns an allow or deny decision. If allowed:
    1. The Lambda authorizer generates an appropriate IAM policy.
    2. API Gateway caches the authorization decision.
    3. The request is forwarded to Service B.
    4. Future similar requests can use the cached decision.

Multi-tenancy controls

As Convera expanded to support multi-tenant software as a service (SaaS) integrations, they needed a way to implement tenant-specific access controls and data isolation. The challenge was to make sure each tenant’s users could only access their authorized resources while allowing tenant administrators to manage their own access policies. Convera used their existing Verified Permissions architecture with a per-tenant policy store approach to address these requirements.

Verified Permissions per-tenant policy store

Convera decided to use per-tenant policy store approach for the following reasons:

  • Low-effort tenant policies isolation
  • The ability to customize templates and schema per tenant
  • Low-effort tenant onboarding and offboarding
  • Per-tenant policy store resource quotas

The following figure shows the process of implementing fine-grained authorization control using Verified Permissions with a per-tenant policy store.

Fine-grained authorization control using Verified Permissions with per-tenant policy store

Figure 4: Fine-grained authorization control using Verified Permissions with per-tenant policy store

The end-to-end process flow consists of the following steps:

  1. The tenant-specific Amazon Cognito pool is created with a custom attribute called tenant_id. The user logs in to the pool with user claims (for example, user_id).
  2. Amazon Cognito uses a pre-token generation Lambda function that looks up the user_id from a user tenant mapping DynamoDB table.
  3. A DynamoDB table is maintained to map user and tenant configuration.
  4. The pre-token generation Lambda hook gets the tenant_id back from DynamoDB, and adds to the custom tenant_id attribute in the access token.
  5. The user makes an API call to API Gateway with the enriched JWT.
  6. API Gateway validates the token and forwards the request to a custom Lambda authorizer.
  7. The Lambda authorizer function reads the tenant_id from the JWT and looks up the associated Verified Permissions policy-store-id from a DynamoDB table.
  8. The Lambda authorizer verifies the JWT for validity and claims with the Verified Permissions associated Amazon Cognito pool taken from the access token.
  9. If authentication is successful, it calls Verified Permissions to verify that the user is permitted to do the requested action.
  10. If allowed, Verified Permissions returns an IAM policy with Allow access (Deny-by-Default) and forwards the request to backend Kubernetes pods with tenant_id in a custom header.
  11. Backend services receive the tenant_id and validate with Verified Permissions again (for zero-trust policy), creates a tenant context, and forwards to Amazon RDS. Amazon RDS is configured to accept only requests with specific tenant context and returns data specific to the requested tenant_id.

The following are some examples of Cedar policies used for multi-tenant isolation:

permit (
    principal in
        convera_connect_authz::userGroup::"ConveraConnect-PAYEE_MGMT",
    action in [convera_connect_authz::Action::"PUT /customer/user/{id}"],
    resource
);
 
permit (
    principal,
    action in [convera_connect_authz::Action::"EDIT"],
    resource
)
when
{
    principal.role.contains("UPDATE_USER_STATUS") &&
    resource.type == "PUT" &&
    resource.path == "/customers/user"
};

Conclusion

In this post, we explored how Convera used Verified Permissions to build a sophisticated, fine-grained authorization model for their API platform. We discussed about how Convera was able to implement fine-grained access for their customers, multi-tenant SaaS integrations, machine-to-machine communication scenarios, and internal customer connect applications with the help of Verified Permissions. With Verified Permissions, Convera was able to achieve the following:

  • Implement fine-grained access control across multiple use cases
  • Enhance security with attribute-based access control across multi-tenant environments
  • Improve scalability, handling over thousands of authorization requests per second with submillisecond latency.
  • Increase operational efficiency, reducing time spent on access management tasks by 60%.
  • Future-proof their authorization framework to adapt to evolving business needs

To learn more about implementing these patterns and best practices, refer to the Verified Permissions User Guide. For hands-on experience, we recommend exploring the Verified Permissions workshop, which provides practical examples and guided exercises.


About the authors

How Artera enhances prostate cancer diagnostics using AWS

Post Syndicated from Hariharan Ananthakrishnan original https://aws.amazon.com/blogs/architecture/how-artera-enhances-prostate-cancer-diagnostics-using-aws/

This post was co-written with Hariharan Ananthakrishnan from Artera.

Artificial intelligence (AI) and machine learning (ML) are transforming cancer diagnosis and treatment, enabling faster and more accurate decisions for patients. One company at the forefront of this transformation is Artera, a precision medicine company developing an AI-powered platform for cancer treatment planning. The U.S. Food and Drug Administration (FDA) has granted De Novo authorization for the ArteraAI Prostate, establishing it as the first and only AI-powered software authorized to prognosticate long-term outcomes for patients with nonmetastatic prostate cancer. The ArteraAI Prostate is now recognized as an FDA-regulated software as a medical device (SaMD). In this post, we explore how Artera used Amazon Web Services (AWS) to develop and scale their AI-powered prostate cancer test, accelerating time to results and enabling personalized treatment recommendations for patients.

Customer overview

Artera offers AI-enabled predictive and prognostic cancer tests, including the ArteraAI Prostate Test. This innovative test analyzes images of a patient’s biopsy to accurately predict the risk of localized cancer spreading as well as the likelihood a patient will benefit from specific therapies. This is the first test that can predict therapeutic benefit for patients with localized prostate cancer, and physicians can use it to make treatment decisions with more confidence, ultimately improving patient outcomes.

Artera is making significant strides in the field of precision medicine, operating in multiple regions. Recently, the FDA granted De Novo authorization for the ArteraAI Prostate platform, highlighting its potential to address unmet needs in cancer care. Since 2024, the ArteraAI Prostate Test has been considered the standard of care for localized prostate cancer, being included in the National Comprehensive Cancer Network Clinical Practice Guidelines in Oncology. The technology’s De Novo authorization establishes a new product code category for future AI-powered digital pathology risk-stratification tools, and it enables its implementation at the point of diagnosis at qualified pathology labs across multiple countries. This capability addresses a critical gap in prostate cancer care by reducing delays in delivering actionable insights at diagnosis, helping clinicians and patients make informed treatment decisions with greater confidence.

The challenge of matching treatment to patient

When patients are diagnosed with cancer, their next step is to determine the course of therapy that will yield the best outcome. Typically, more aggressive cancers require more aggressive therapy. However, it’s not always clear how aggressively the cancer may progress. Furthermore, patients respond differently to the same therapy based on their unique biological makeup. As a consequence, some patients with less aggressive disease are inadvertently overtreated, receiving unnecessary therapies involving a host of side effects, while others with more aggressive cancers are undertreated, leading to potentially worse outcomes.

Before Artera’s solution, there were no AI-based tools to help physicians and cancer patients make personalized, timely treatment decisions. Instead, physicians submitted a patient’s biopsy tissue sample to a lab, where a chemical assay measured the expression levels of a small set of genes. The RNA expression of these genes was then used to assess a patient’s risk level. These tests have several limitations:

  • The entire process can take 6 weeks—a long time to wait when making a high-stress decision about cancer therapy.
  • These tests typically only identify a small number of key genes (as science continues to advance faster than the diagnostic tests can keep up) linked to cancer risk.
  • These tests consume the original tissue samples, limiting the physician’s ability to order additional tests, as well as the patient’s ability to enroll in future clinical trials or participate in long-term monitoring

Developing an AI-powered diagnostic tool for cancer treatment presents unique technical challenges. Artera had to manage and process a large volume of high-resolution biopsy image files to power their AI-driven cancer diagnostics. These images are enormous, sometimes reaching 8 GB, and they need to be broken down into tens of thousands of smaller patches for the model to handle. Training Artera’s foundation models (FMs) requires serving millions of image patches at high volume to AWS servers.

Additionally, as a healthcare company handling sensitive patient data, Artera needed to ensure compliance and data residency and regulatory requirements across multiple countries, including the Health Insurance Portability and Accountability Act (HIPAA) in the United States. They needed a robust, scalable storage solution that would enable their ML engineers to focus on the core cancer research rather than infrastructure management.

Modern, scalable design delivers fast results

Artera implemented a comprehensive AWS based solution to address their challenges. The architecture follows a modern, scalable design that enables secure processing of sensitive medical data while delivering fast results to healthcare providers. Their solution starts with training AI models, advanced workflow orchestration, and data locality principles that are critical for global deployment of clinical AI models.

“Artera was founded with the belief that there were a lot of signals in the histopathology image data that were not being used, but if an AI algorithm could be specifically developed with this in mind, you could radically change cancer patient care,”

– Nathan Silberman, Chief Technology Officer of Artera.

The following architecture diagram illustrates how Artera has built a secure, scalable solution on AWS. At its core, Artera’s AI products are composed of many individual steps in a complex workflow, often involving multiple AI models that perform different specialized tasks. This sophisticated workflow orchestration helps them move faster and abstract away complexity as they build their compound AI system.

AWS architecture diagram showing medical professionals accessing ArteraAI portal through AWS Global Accelerator, WAF, load balancer, with ECS web portal and EKS AI inference cluster in a VPC, connected to data storage services and comprehensive security monitoring.

Comprehensive AWS architecture diagram showing the integration of cloud services for a medical professionals’ portal with AI inference capabilities, including data flow from end users through global acceleration services to compute, storage, and security infrastructure in a VPC within Region A.

Medical professionals access the Artera Portal, which serves as the interface for uploading biopsy images and receiving diagnostic results. AWS Global Accelerator sits in front of the Application Load Balancer, providing improved availability and performance by directing traffic through the AWS global network. Amazon CloudFront provides a fast, secure content delivery network for the portal’s static assets, providing low-latency access globally

Within a virtual private cloud (VPC), Elastic Load Balancing distributes incoming traffic across the application servers. Amazon Elastic Container Service (Amazon ECS) hosts the web portal containers, providing the user interface for healthcare professionals. An Amazon Elastic Kubernetes Service (Amazon EKS) cluster runs the AI/ML inference workloads that analyze biopsy images using computer vision models.

Amazon Elastic File System (Amazon EFS) provides shared file storage, accessible by both Amazon ECS and Amazon EKS for storing and processing biopsy images. Amazon Relational Database Service (Amazon RDS) delivers a managed relational database for patient records, diagnostic results, and application data with high availability. Amazon ElastiCache provides in-memory caching to improve application performance and reduce latency for frequently accessed data.

AWS Identity and Access Management (IAM) provides proper access controls and permissions. AWS Key Management Service (AWS KMS) manages encryption keys for sensitive patient data. Amazon CloudWatch monitors the entire infrastructure for performance and health. Amazon Simple Storage Service (Amazon S3) provides durable, secure storage for biopsy images and analysis results.

This architecture enables a complete workflow:

  1. Data ingestion – Biopsy images are securely uploaded through the portal and stored in Amazon S3.
  2. Processing pipeline – The EKS cluster orchestrates containerized preprocessing applications that prepare images for analysis.
  3. ML model training and execution – The AI models are trained and deployed on Amazon EKS and access the preprocessed images from Amazon EFS, then run Artera’s proprietary ML algorithms, with metadata and results stored in Amazon RDS. The company’s ML teams use EKS to train their massive pan-tumor FM, which is capable of assessing patient risk and therapy benefit across any cancer sample.
  4. Results storage and delivery – Analysis results are stored in Amazon S3 and made available to healthcare providers through the secure web portal.

Data locality and global scalability

One of the key challenges Artera faced was maintaining data locality while serving AI globally. The company uses multiple AWS services to create a comprehensive solution that addresses both performance and compliance requirements.

AWS global infrastructure enables Artera to deploy Region-specific resources that keep sensitive patient data within appropriate jurisdictional boundaries. Amazon S3 provides secure, Region-specific storage buckets, and Amazon EKS allows for containerized workloads to run locally in each Region.

“One of the nice things about Amazon EFS is that it’s very simple to achieve data locality,” says Silberman. “We can mount file systems in the same AWS Region as our applications, ensuring data stays close to where it’s processed.”The combination of Amazon S3, Amazon EKS, Amazon EFS, and other AWS networking services creates a robust foundation for Artera’s global operations. This integrated approach helps Artera accelerate time to market in new regions while maintaining the highest standards of data security and compliance with regional regulations.

To learn more about how Artera uses Amazon EFS, visit the case study, Artera Shapes the Future of Cancer Treatment Using Machine Learning on AWS.

Results and patient impact

By using AWS Cloud services, Artera has transformed cancer diagnostics with tangible benefits for patients:

  • Accelerated results – Patients receive personalized treatment recommendations in only 1–2 days, compared to 6 weeks for traditional genomic tests—dramatically reducing the waiting period for critical treatment decisions.
  • Improved clinical decisions – The speed and accuracy of Artera’s AI-powered diagnostics help physicians make more informed treatment decisions, potentially improving outcomes for prostate cancer patients.
  • Tissue preservation – Unlike traditional tests that destroy tissue samples through chemical assays, the ArteraAI Prostate Test uses only digital imagery, preserving the original tissue for additional tests or clinical trials.

In 2024, almost 300,000 Americans were diagnosed with prostate cancer. For these patients, timely and accurate diagnostics are essential.

“Imagine a patient getting the worst news they’ve ever had and having to sit on that for 6 weeks to determine what the treatment plan is,” says Silberman. “Instead, Artera provides custom-tailored, personalized results within days.”

There are over 3.5 million prostate cancer survivors in the United States. By recommending personalized treatment plans, Artera is helping patients determine the best therapeutic options to achieve progression-free survival while minimizing unnecessary side effects.

“We’ve heard from patients who have said that because of our test, they were able to avoid unnecessary treatments with a lot of side effects,” says Silberman. “That’s why all of us at Artera are here, giving clinicians as many data-backed insights as possible to inform the patient and make the best possible choice for their care.”

Operational benefits

Using AWS services has meant that Artera has achieved significant operational advantages:

  • Enhanced focus on innovation – With AWS managing the infrastructure, Artera’s engineers can dedicate more time to refining their ML algorithms and expanding diagnostic capabilities.

“Using AWS, we can focus on the histopathology problems, rather than on maintenance and monitoring,” says Silberman.

  • Global scalability – Artera has successfully expanded operations while maintaining compliance with regional data regulations across multiple countries.
  • Efficient processing – The test processes tens of thousands of image files through ML workflows per biopsy slide, completing in hours instead of weeks. This efficiency comes from Artera’s sophisticated workflow orchestration that breaks up large input images (sometimes reaching 8 GB) into many small patches processed in parallel across EKS clusters.

The FDA’s De Novo authorization for the ArteraAI Prostate Test underscores the potential impact of this technology on cancer care. With AWS powering their infrastructure, Artera is well-positioned to continue revolutionizing how cancer is diagnosed and treated.

Future innovations

As Artera continues to innovate in the field of AI-powered cancer diagnostics, their AWS based infrastructure provides the foundation for future growth. The company’s ultimate goal is a massive pan-tumor FM capable of assessing patient risk and therapy benefit across any cancer sample. Using elastic, scalable solutions on AWS, Artera has a solid foundation for developing ML models for additional cancer tests. The company has announced plans for a breast cancer product, with several more products close behind.

“What we have coming up is a rapid acceleration across different areas of cancer,” says Silberman. “As proud as we are of the work that we’ve done in the prostate cancer space, we’re just getting started.”

Artera plans to expand their AI capabilities in several ways:

  • Analyze additional biomarkers
  • Integrate genomic data with imaging analysis
  • Create more comprehensive diagnostic tools
  • Partner with major healthcare systems to integrate diagnostic tools directly into clinical workflows

With the scalability of AWS services, Artera is positioned to handle the increasing data demands as they expand to new cancer types and regions globally.

Conclusion

Artera’s journey demonstrates how AWS Cloud services can empower healthcare innovators to develop and scale life-changing technologies. By using Amazon EKS, Amazon ECS, Amazon EFS, Amazon RDS, Amazon S3, AWS Global Accelerator, and Amazon ElastiCache, Artera built a robust, scalable infrastructure they use to keep their focus on their core mission: improving cancer treatment through AI-powered diagnostics. To learn more about how AWS can help your healthcare organization implement AI and ML solutions, visit AWS for Healthcare.

To learn more about Artera and their innovative cancer diagnostics, visit Artera.ai.


About the authors

How Tipico democratized data transformations using Amazon Managed Workflows for Apache Airflow and AWS Batch

Post Syndicated from Jake J. Dalli original https://aws.amazon.com/blogs/big-data/how-tipico-democratized-data-transformations-using-amazon-managed-workflows-for-apache-airflow-and-aws-batch/

This is a guest post by Jake J. Dalli, Data Platform Team Lead at Tipico, in partnership with AWS.

Tipico is the number one name in sports betting in Germany. Every day, we connect millions of fans to the thrill of sport, combining technology, passion, and trust to deliver fast, secure, and exciting betting, both online and in more than a thousand retail shops across Germany. We also bring this experience to Austria, where we proudly operate a strong sports betting business.

In this post, we show how Tipico built a unified data transformation platform using Amazon Managed Workflows for Apache Airflow (Amazon MWAA) and AWS Batch.

Solution overview

To support critical needs such as product monitoring, customer insights, and revenue assurance, our central data function needed to provide the tools for several cross-functional analytics and data science teams to run scalable batch workloads on the existing data warehouse, powered by Amazon Redshift. The workloads of Tipico’s data community included extract, transform, and load (ELT), statistical modeling, machine learning (ML) training, and reporting across diverse frameworks and languages.

In the past, analytics teams operated in isolation, distinct from each other and the central data function. Different teams maintained their own set of tools, often performing the same function and creating data silos. Lack of visibility meant a lack of standardization. This siloed approach slowed down the delivery of insights and prevented the company from achieving a unified data strategy that ensured availability and scalability.

The need to introduce a single, unified platform that promoted visibility and collaboration became clear. However, the diversity of workloads brought another layer of complexity. Teams needed to tackle different types of problems and brought distinct skillsets and preferences in tooling. Analysts might rely heavily on SQL and business intelligence (BI) platforms, whereas data scientists preferred Python or R, and engineers leaned on containerized workflows or orchestration frameworks.

Our goal was to architect a new system that supports diversity while maintaining operational control, delivering an open orchestration platform with built-in security isolation, scheduling, retry mechanisms, fine-grained role-based access control (RBAC), and governance features such as two-person approval for production workflows. We achieved this by designing a system with the following principles:

  1. Bring Your Own Container (BYOC) – Teams are given the flexibility to package their workloads as containers and are free to choose dependencies, libraries, or runtime environments. For teams with highly specialized workloads, this meant that they could work in a setup tailored to their needs while also operating within a harmonized platform. On the other hand, teams that didn’t require fully customized environments could redesign their workloads to align with existing workloads.
  2. Centralized orchestration for full transparency – All teams can see all workflows and build interdependencies between them
  3. Shared orchestration, isolated compute – Workloads run in team-specific Docker containers within a unified compute environment, providing scalability while keeping execution traceable to each team.
  4. Standardized interfaces, flexible execution – Common patterns (operators, hooks, logging, or monitoring) reduce complexity, and teams retain freedom to innovate within their containers.
  5. Cross-team approvals for critical workflows stored inside version control – Changes follow a four-eye principle, requiring review and approval from another team before execution, providing accountability and reducing risk. This allowed our core data function to monitor and contribute suggestions to work across different analytics teams.

We devised a system wherein orchestration and execution of tasks operate on shared infrastructure, which teams interact with through domain-specific infrastructure. In Tipico’s case, each team pushes images to team-owned container instances. Such containers provide code for workflows, including execution of ELT pipelines or transformations on top of domain-specific data lakes.

The following diagram shows the solution architecture.

The technical challenge was to architect a flexible and high-performance orchestration layer that could scale reliably while also remaining framework-agnostic, integrating seamlessly with existing infrastructure.

When designing our system, we were aware of the several container orchestration solutions offered by Amazon Web Services (AWS), including Amazon Elastic Kubernetes Service (Amazon EKS), Amazon Elastic Container Service (Amazon ECS), and AWS Batch, among others. In the end, the team selected AWS Batch because it abstracts away cluster management, provides elastic scaling, and inherently supports batch workloads as a design feature.

Solution details

Before adopting the current solution, Tipico experimented with operating a self-managed Apache Airflow setup. Although it was functional, it became increasingly burdensome to maintain. The shift toward a managed and scalable solution was driven by the need to focus more on empowering teams to deliver rather than maintaining the infrastructure. Tipico replatformed the central orchestration solution using Amazon MWAA and AWS Batch.

Amazon MWAA is a fully managed service that simplifies running open source Apache Airflow on AWS. Users can build and execute data processing workflows while integrating seamlessly with various AWS services, which means developers and data engineers can concentrate on building workflows rather than managing infrastructure.

AWS Batch is a fully managed service that simplifies batch computing in the cloud so users can run batch jobs without needing to provision, manage, or maintain clusters. It automates resource provisioning and workload distribution, with users only paying for the underlying AWS resources consumed.

The new design provides a unified framework where analytics workloads are containerized, orchestrated, and executed on scalable compute and integrated with persistent storage:

  1. Containerization – Analytics workloads are packaged into Docker containers, with dependencies bundled to provide reproducibility. These images are versioned and stored in Amazon Elastic Container Registry (Amazon ECR). This approach decouples execution from infrastructure and enables consistent behavior across environments.
  2. Workflow orchestration – Airflow Directed Acyclic Graphs (DAGs) are version-controlled in Git and deployed to Amazon MWAA using a continuous integration and continuous delivery (CI/CD) pipeline. Amazon MWAA schedules and orchestrates tasks, triggering AWS Batch jobs using custom operators. Logs and metrics are streamed to Amazon CloudWatch, enabling real-time observability and alerting.
  3. Data persistence – Workflows interact with Amazon Simple Storage Service (Amazon S3) for durable storage of inputs, outputs, and intermediate artifacts. Amazon Elastic File System (Amazon EFS) is mounted to Amazon MWAA for fast access to shared code and configuration files, synchronized continuously from the Git repository.
  4. Scalable compute – Amazon MWAA triggers AWS Batch jobs using standardized job definitions. These jobs run in elastic compute environments such as Amazon Elastic Compute Cloud (Amazon EC2) or AWS Fargate, with secrets securely injected using AWS Secrets Manager. AWS Batch environments auto scale based on workload demand, optimizing cost and performance.
  5. Security and governance – AWS Identity and Access Management (IAM) roles are scoped per team and workload, providing least-privilege access. Job executions are logged and auditable, with fine-grained access control enforced across Amazon S3, Amazon ECR, and AWS Batch.

Common operators

To streamline the execution of batch jobs across teams, we developed a shared operator that wraps the built-in Airflow AWS Batch operator. This abstraction simplifies the execution of containerized workloads by encapsulating common logic such as:

  1. Job definition selection
  2. Job queue targeting
  3. Environment variable injection
  4. Secrets resolution
  5. Retry policies and logging configuration

Parameterization is handled using Airflow Variables and XComs, enabling dynamic behavior across DAG runs. The operator is maintained in a shared Git repository, versioned and centrally governed, but accessible to all teams.

To further accelerate development, some teams use a DAG Factory pattern, which programmatically generates DAGs from configuration files. This reduces boilerplate and enforces consistency so teams can define new workflows declaratively.

By standardizing this operator and supporting patterns, Tipico reduces onboarding friction, promotes reuse, and provides consistent observability and error handling across the analytics ecosystem.

Governance

Governance is enforced through a combination of fine-grained IAM roles, AWS IAM Identity Center and automated role mapping. Each team is assigned a dedicated IAM role, which governs access to AWS services such as Amazon S3, Amazon ECR, AWS Batch and Secrets Manager. These roles are tightly scoped to minimize the extent of damage and provide traceability.

Given that the airflow environment runs version 2.9.2, which doesn’t support multi-tenant access, Tipico developed a custom component that dynamically maps AWS IAM roles to Airflow roles. The component, which executes periodically using Airflow itself, dynamically syncs IAM role assignments with Airflow’s internal RBAC model. Airflow tags are used to govern access to different DAGs, governing which teams have access to execute or modify the settings on the DAG. This aligns access permissions remain with organizational structure and team responsibilities.

Adoption

The shift toward a managed, scalable solution was driven by the need for greater team autonomy, standardization, and scalability. The journey began with a single analytics team validating the new approach. When it was successful, the platform team generalized the solution and rolled it out incrementally to other teams, refining it with each iteration.One of the biggest challenges was migrating legacy code, which often included outdated logic and undocumented dependencies. To support adoption, Tipico introduced a structured onboarding process with hands-on training, real use cases, and internal champions. In some cases, teams also had to adopt Git for the first time—marking a broader shift toward modern engineering practices within the analytics organization.

Key benefits

One of the most valuable outcomes of our new architecture that is primarily built around Amazon MWAA and AWS Batch is to accelerate analytics teams’ time to value. Analysts can now focus on building transformation logic and workloads without worrying about the underlying infrastructure. With this system, analysts can rely on preprepared integrations and analytics patterns used across different teams, supported by standard interfaces developed by the core data team.

Aside from building analytics on Amazon Redshift, the orchestration solution also interfaces with several other analytics services such as Amazon Athena and AWS Glue ETL, providing maximum flexibility on the type of workloads being delivered. Teams within the organization have also shared practices in using different frameworks, such as dbt Labs, to reuse custom developments to carry out standard processes.

Another valuable outcome is the ability to clearly segregate costs across teams. Within the architecture, Airflow delegates heavy lifting to AWS Batch, providing task isolation that spans beyond Airflow’s built-in workers. Through this, we gain granular visibility into resource usage and accurate cost attribution, promoting financial accountability across the organization.

Finally, the platform also provides embedded governance and security, with RBAC and standardized secrets management providing an operationalized model for securing and governing working flows across different teams.

Teams can now focus on building and iterating quickly, knowing that the surrounding structures provide full transparency and are coherent with the organization’s governance, architecture, and FinOps goals. At the same time, centralized orchestration fosters a collaborative environment where teams can discover, reuse, and build upon each other’s workflows, driving innovation and reducing duplication across the data landscape.

Conclusion

By reimagining our orchestration layer with Amazon MWAA and AWS Batch, Tipico has unlocked a new level of agility and transparency across its data workflows.

Previously, analytics teams faced long lead times, often stretching into weeks, to implement new reporting use cases. Much of this time was spent identifying datasets, aligning transformation logic, discovering integration options, and navigating inconsistent quality assurance processes. Today, that has changed. Analysts can now develop and deploy a use case within a single business day, shifting their focus from groundwork to action.

The modern architecture empowers teams to move faster and more independently within a secure, governed, and scalable framework. The result is a collaborative data ecosystem where experimentation is encouraged, operational overhead is reduced, and insights are delivered at speed.

To start building your own orchestrated data platform, explore the Get started with Amazon Managed Workflows for Apache Airflow and AWS Batch User Guide. These services can help you achieve similar results in democratizing data transformations across your organization. For hands-on experience with these solutions, try our Amazon MWAA for Analytics Workshop or contact your AWS account team to learn more.


About the authors

Jake J. Dalli

Jake J. Dalli

Jake is the Data Platform Team Lead at Tipico, where he is engaged in architecting and scaling data platforms that enable reliable analytics and informed decision-making across the organization. He’s passionate about empowering analysts to deliver faster insights by simplifying complex systems and accelerating time to value.

David Greenshtein

David Greenshtein

David is a Senior Specialist Solutions Architect for Analytics at AWS, with a passion for building distributed data platforms aligned with governance requirements. He works with customers to design and implement scalable, governed analytics solutions to turn data into actionable insights and measurable business outcomes.

Hugo Mineiro

Hugo Mineiro

Hugo is a Senior Analytics Specialist Solutions Architect based in Geneva. He focuses on helping customers across various industries build scalable and high-performing analytics solutions. He loves playing football and spending time with friends.

Get started faster with one-click onboarding, serverless notebooks, and AI agents in Amazon SageMaker Unified Studio

Post Syndicated from Siddharth Gupta original https://aws.amazon.com/blogs/big-data/get-started-faster-with-one-click-onboarding-serverless-notebooks-and-ai-agents-in-amazon-sagemaker-unified-studio/

Data teams today struggle with fragmented tools, complex infrastructure provisioning, and hours spent writing boilerplate code to connect to data sources. This forces analysts, data scientists, and engineers to work in separate environments, which slows collaboration and time to insight. Since our launch of Amazon SageMaker Unified Studio in March 2025, leading companies such as Bayer, NatWest, and Carrier have adopted it to bring their data teams into one collaborative workspace with unified tools, straightforward infrastructure provisioning, and fast connections to data sources.

Continuing our mission to provide faster time-to-value for customers, in November 2025, we announced Amazon SageMaker notebooks, a serverless workspace with a built-in AI agent in Amazon SageMaker Unified Studio. You can now launch a notebook in seconds, generate code from natural language prompts, and connect automatically to data across Amazon Simple Storage Service (Amazon S3), Amazon Redshift, third-party databases, and more from a single environment without needing to pre-provision or tune data processing infrastructure. Inside these serverless notebooks, analysts can perform SQL queries, data scientists can execute Python code, and data engineers can process large-scale data jobs in Spark within a single workspace. Together with the new one-click onboarding available for SageMaker Unified Studio, customers can go from their existing AWS data to running analytics and machine learning workloads much faster, spending their time on analysis rather than setup and configuration.

In this post, we walk you through how these new capabilities in SageMaker Unified Studio can help you consolidate your fragmented data tools, reduce time to insight, and collaborate across your data teams. Here’s a short demo of the new capabilities:

One-click onboarding of existing AWS datasets

Get started exploring your data with one-click onboarding that provisions and configures environments in minutes instead of weeks. The new onboarding experience can reuse existing AWS Identity and Access Management (IAM) roles to provide access to SageMaker Unified Studio, automatically connecting to data sources across S3 buckets, S3 Tables, AWS Glue Data Catalog, and AWS Lake Formation policies, removing the need for additional data permission setup. Under the covers, a new IAM-based domain and project are created with default notebook and compute resources preconfigured. When complete, you enter SageMaker Unified Studio with all your tools available in the left-side navigation along with built-in samples to accelerate first use, as seen in the following screenshot.

“New features with Amazon Sagemaker will unlock a new paradigm of innovation, allowing Codex to significantly accelerate time-to-value for our customers, and transform them from aging to agentic in weeks, not months.“

– Abhinav Sharma, Chief Data Officer, Codex

You can start directly from Amazon SageMaker, Amazon Athena, Amazon Redshift, or Amazon S3 Tables, giving them a fast path from their existing tools and data to the unified experience in SageMaker Unified Studio. After you choose Get Started and specify an IAM role, SageMaker automatically creates a project with the existing data permissions intact from Data Catalog, Lake Formation, and Amazon S3. As a result, teams can immediately discover and act on their data using the existing data permissions and infrastructure.

For more information, see New one-click onboarding and notebooks with a built-in AI agent in Amazon SageMaker Unified Studio

Serverless SageMaker notebooks

The fully managed, web-based notebooks in SageMaker Unified Studio support multiple programming languages, letting you write Python, SQL, and Spark code in the same notebook. The infrastructure adjusts automatically based on your workload, while built-in libraries create charts and insights directly in your workflow. When your analysis scales beyond interactive queries to large-scale data processing, Amazon Athena for Apache Spark engine delivers optimized performance, integrating with the serverless notebook experience to execute analytical workloads efficiently. This serverless approach eliminates the need to provision clusters or maintain servers, reducing the time from question to insight.

“The new SageMaker interface brings clarity and speed to the entire ML lifecycle. Its developer-friendly design has made our experimentation and delivery significantly faster,“

– Sachin Mittal, Product Manager at Deloitte.

As shown in the preceding image, the notebook gives data engineers, analysts, and data scientists one place to perform SQL queries, execute Python code, process large-scale data jobs, run machine learning workloads, and create visualizations without having to switch between tools.

AI-assisted development with Data Agent

To accelerate development further, the new SageMaker Data Agent helps create SQL, Python, or Spark code using natural language prompts. Instead of spending hours writing boilerplate code to connect to your data sources and understand schemas, you can describe what you want to accomplish. The agent analyzes data catalog metadata about your available datasets, schemas, and relationships to provide context-aware assistance.

In the preceding example image, if you prompt Build and analyze a complete sales forecast based on the sample retail data, the agent helps identify the relevant tables and suggests the appropriate joins and analysis approach, transforming what might take hours into minutes. To try this yourself, navigate to the Overview tab in your SageMaker Studio environment and look for the Retail Sales Forecasting with SageMaker XGBoost notebook in the sample notebooks collection—these examples are automatically available when you first set up SageMaker Studio. The agent breaks down complex analytical workflows into manageable, executable steps, so you can move from question to insight faster.

Learn more about SageMaker

In this post, we focused on three new SageMaker Unified Studio capabilities recently made available, but they’re a fraction of the more than 40 launches last year. Here’s a list of videos of re:Invent sessions and the measurable results from leading organizations adopting SageMaker Unified Studio, including:

  • Summary of 2025 launches: What’s new with Amazon SageMaker in the era of unified data and AI (ANT216)
  • NatWest Group plans to scale to 72,000 employees having federated data access using SageMaker Unified Studio. Watch their presentation.
  • Commonwealth Bank of Australia migrated 10 petabytes and 61,000 pipelines into AWS and has setup SageMaker Unified Studio to provide unified access to 40 different lines of business in their ongoing data transformation journey. Watch their presentation.
  • Carrier Global Corporation improved natural language to SQL agent accuracy by 38% through the SageMaker Catalog’s governed metadata and business glossary. Watch their presentation.
  • Bayer is now positioned to onboard over 300 TB of biomarker data and integrate siloed omics, clinical, and chemistry data repositories into a cohesive environment built on Amazon SageMaker. Read their story.

Conclusion

Using Amazon SageMaker Unified Studio serverless notebooks, AI-assisted development, and unified governance, you can speed up your data and AI workflows across data team functions while maintaining security and compliance. To learn more visit the SageMaker product page or get started in the SageMaker console.


About the authors

Siddharth Gupta

Siddharth Gupta

Siddharth is heading Generative AI within SageMaker’s Unified Experiences. His focus is on driving agentic experiences, where AI systems act autonomously on behalf of users to accomplish complex tasks. An alumnus of the University of Illinois at Urbana-Champaign, he brings extensive experience from his roles at Yahoo, Glassdoor, and Twitch.

Matt David

Matt David

Matt is a Product Marketing Manager at AWS, specializing in helping data teams with AI-powered analytics. His areas of interest include self-service analytics, data democratization, and preparing organizations for the age of AI agents. He brings extensive experience from his roles at Atlassian, Hex, and DataCamp.

Sean Ma

Sean Ma

Sean is a leader on Amazon SageMaker and an AWS Principal Product Manager. He is passionate about delivering products that Data and AI professionals love through user experience focused product design. Sean’s track record of innovation with successful products includes AWS Glue, Google Cloud Data Analytics, Informatica and Alteryx (Trifacta).

How Bazaarvoice modernized their Apache Kafka infrastructure with Amazon MSK

Post Syndicated from Oleh Khoruzhenko original https://aws.amazon.com/blogs/big-data/how-bazaarvoice-modernized-their-apache-kafka-infrastructure-with-amazon-msk/

This is a guest post by Oleh Khoruzhenko, Senior Staff DevOps Engineer at Bazaarvoice, in partnership with AWS.

Bazaarvoice is an Austin-based company powering a world-leading reviews and ratings platform. Our system processes billions of consumer interactions through ratings, reviews, images, and videos, helping brands and retailers build shopper confidence and drive sales by using authentic user-generated content (UGC) across the customer journey. The Bazaarvoice Trust Mark is the gold standard in authenticity.

Apache Kafka is one of the core components of our infrastructure, enabling real-time data streaming for the global review platform. Although Kafka’s distributed architecture met our needs for high-throughput, fault-tolerant streaming, self-managing this complex system diverted critical engineering resources away from our core product development. Each component of our Kafka infrastructure required specialized expertise, ranging from configuring low-level parameters to maintaining the complex distributed systems that our customers rely on. The dynamic nature of our environment demanded continuous care and investment in automation. We found ourselves constantly managing upgrades, applying security patches, implementing fixes, and addressing scaling needs as our data volumes grew.

In this post, we show you the steps we took to migrate our workloads from self-hosted Kafka to Amazon Managed Streaming for Apache Kafka (Amazon MSK). We walk you through our migration process and highlight the improvements we achieved after this transition. We show how we minimized operational overhead, enhanced our security and compliance posture, automated key processes, and built a more resilient platform while maintaining the high performance our global customer base expects.

The need for modernization

As our platform grew to process billions of daily consumer interactions, we needed to find a way to scale our Kafka clusters efficiently while maintaining a small team to manage the infrastructure. The limitations of self-managed Kafka clusters manifested in several key areas:

  • Scaling operations – Although scaling our self-hosted Kafka clusters wasn’t inherently complex, it required careful planning and execution. Each time we needed to add new brokers to handle increased workload, our team faced a multi-step process involving capacity planning, infrastructure provisioning, and configuration updates.
  • Configuration complexity – Kafka offers hundreds of configuration parameters. Although we didn’t actively manage all of these, understanding their impact was important. Key settings like I/O threads, memory buffers, and retention policies needed ongoing attention as we scaled. Even minor adjustments could have significant downstream effects, requiring our team to maintain deep expertise in these parameters and their interactions to ensure optimal performance and stability.
  • Infrastructure management and capacity planning – Self-hosting Kafka required us to manage multiple scaling dimensions, including compute, memory, network throughput, storage throughput, and storage volume. We needed to carefully plan capacity for all these components, often making complex trade-offs. Beyond capacity planning, we were responsible for real-time management of our Kafka infrastructure. This included promptly detecting and addressing component failures and performance issues. Our team needed to be highly responsive to alerts, often requiring immediate action to maintain system stability.
  • Specialized expertise requirements – Operating Kafka at scale demanded deep technical expertise across multiple domains. The team needed to:
    • Monitor and analyze hundreds of performance metrics
    • Conduct complex root cause analysis for performance issues
    • Manage ZooKeeper ensemble coordination
    • Execute rolling updates for zero-downtime upgrades and security patches

These challenges were compounded during peak business periods, such as Black Friday and Cyber Monday, when maintaining optimal performance was essential for Bazaarvoice’s retail customers.

Choosing Amazon MSK

After evaluating various options, we selected Amazon MSK as our modernization solution. The decision was driven by the service’s ability to minimize operational overhead, provide high availability out of the box with its three Availability Zone architecture, and offer seamless integration with our existing AWS infrastructure.

Key capabilities that made Amazon MSK the clear choice:

  • AWS integration – We already used AWS services for data processing and analytics. Amazon MSK connected directly with these services, alleviating the need to build and maintain custom integrations. This meant our existing data pipelines would continue working with minimal changes.
  • Automated operations management – Amazon MSK automated our most time-consuming tasks. We no longer need to manually monitor instances and storage for failures or respond to these issues ourselves.
  • Enterprise-grade reliability – The platform’s architecture matched our reliability requirements out of the box. Multi-AZ distribution and built-in replication gave us the same fault tolerance we’d carefully built into our self-hosted system, now backed by AWS’s service guarantees.
  • Simplified upgrade process – Before Amazon MSK, version upgrades for our Kafka clusters required careful planning and execution. The process was complex, involving multiple steps and risks. Amazon MSK simplified our upgrade operations. We now use automated upgrades for dev and test workloads and maintain control over production environments. This shift reduced the need for extensive planning sessions and multiple engineers. As a result, we stay current with the latest Kafka versions and security patches, improving our system reliability and performance.
  • Enhanced security controls – Our platform required ISO 27001 compliance, which typically involved months of documentation and security controls implementation. Amazon MSK came with this certification built-in, alleviating the need for separate compliance work. Amazon MSK encrypted our data, controlled network access, and integrated with our existing security tools.

With Amazon MSK selected as our target platform, we began planning the complex task of migrating our critical streaming infrastructure without disrupting the billions of consumer interactions flowing through our system.

Bazaarvoice’s migration journey

Moving our complex Kafka infrastructure to Amazon MSK required careful planning and precise execution. Our platform processes data through two main components: an Apache Kafka Streams pipeline that handles data processing and augmentation, and client applications that move this enriched data to downstream systems. With 40 TB of state across 250 internal topics, this migration demanded a methodical approach.

Planning phase

Working with AWS Solutions Architects proved critical for validating our migration strategy. Our platform’s unique characteristics required special consideration:

  • Multi-Region deployment across the US and EU
  • Complex stateful applications with strict data consistency needs
  • Vital business services requiring zero downtime
  • Diverse consumer ecosystem with different migration requirements

Migration challenges

The biggest hurdle was migrating our stateful Kafka Streams applications. Our data processing runs as a directed acyclic graph (DAG) of applications across regions, using static group membership to prevent disruptive rebalancing. It’s important to note that Kafka Streams keeps its state in internal Kafka topics. For applications to recover properly, replicating this state accurately is crucial. This characteristic of Kafka Streams added complexity to our migration process. Initially, we considered MirrorMaker2, the standard tool for Kafka migrations. However, two fundamental limitations made it challenging:

  • Risk of losing state or incorrectly replicating state across our applications.
  • Inability to run two instances of our applications simultaneously, which meant we needed to shut down the main application and wait for it to recover from the state in the MSK cluster. Given the size of our state, this recovery process exceeded our 30-minute SLA for downtime.

Our solution

We decided to deploy a parallel stack of Kafka Streams applications reading and writing data from Amazon MSK. This approach gave us sufficient time for testing and verification, and enabled the applications to hydrate their state before we delivered the output to our data warehouse for analytics. We used MirrorMaker2 for input topic replication, while our solution offered several advantages:

  • Simplified monitoring of the replication process
  • Avoided consistency issues between state stores and internal topics
  • Allowed for gradual, controlled migration of consumers
  • Enabled thorough validation before cutover
  • Required a coordinated transition plan for all consumers, because we couldn’t transfer consumer offsets across clusters

Consumer migration strategy

Each consumer type required a carefully tailored approach:

  • Standard consumers – For applications supporting Kafka Consumer Group protocol, we implemented a four-step migration. This approach risked some duplicate processing, but our applications were designed to handle this scenario. The steps were as follows:
    • Configure consumers with auto.offset.reset: latest.
    • Stop all DAG producers.
    • Wait for existing consumers to process remaining messages.
    • Cut over consumer applications to Amazon MSK.
  • Apache Kafka Connect Sinks – Our sink connectors served two critical databases:
    • A distributed search and analytics engine – Document versioning depended on Kafka record offsets, making direct migration impossible. To address this, we implemented a solution that involved building new search engine clusters from scratch.
    • A document-oriented NoSQL database – This supported direct migration without requiring new database instances, simplifying the process significantly.
  • Apache Spark and Flink applications – These presented unique challenges due to their internal checkpointing mechanisms:
    • Offsets managed outside Kafka’s consumer groups
    • Checkpoints incompatible between source and target clusters
    • Required complete data reprocessing from the beginning

We scheduled these migrations during off-peak hours to minimize impact.

Technical benefits and improvements

Moving to Amazon MSK fundamentally changed how we manage our Kafka infrastructure. The transformation is best illustrated by comparing key operational tasks before and after the migration, summarized in the following table.

Activity Before: Self-Hosted Kafka After: Amazon MSK
Security patching Required dedicated team time for Kafka and OS updates Fully automated
Broker recovery Needed manual monitoring and intervention Fully automated
Client authentication Complex password rotation procedures AWS Identity and Access Management (IAM)
Version upgrades Complex procedure requiring extensive planning Fully automated

The details of the tasks are as follows:

  • Security patching – Previously, our team spent 8 hours monthly applying Kafka and operating system (OS) security patches across our broker fleet. Amazon MSK now handles these updates automatically, maintaining our security posture without engineering intervention.
  • Broker recovery – Although our self-hosted Kafka had automatic recovery capabilities, each incident required careful monitoring and occasional manual intervention. With Amazon MSK, node failures and storage degradation issues such as Amazon Elastic Block Store (Amazon EBS) slowdowns are handled entirely by AWS and resolved within minutes without our involvement.
  • Authentication management – Our self-hosted implementation required password rotations for SASL/SCRAM authentication, a process that took two engineers several days to coordinate. The direct integration between Amazon MSK and AWS Identity and Access Management (IAM) minimized this overhead while strengthening our security controls.
  • Version upgrades – Kafka version upgrades in our self-hosted environment required weeks of planning and testing as well as weekend maintenance windows. Amazon MSK manages these upgrades automatically during off-peak hours, maintaining our SLAs without disruption.

These improvements proved especially valuable during high-traffic periods like Black Friday, when our team previously needed extensive operational readiness plans. Now, the built-in resiliency of Amazon MSK provides us with reliable Kafka clusters that serve as mission-critical infrastructure for our business. The migration made it possible to break our monolithic clusters into smaller, dedicated MSK clusters. This improved our data isolation, provided better resource allocation, and enhanced performance predictability for high-priority workloads.

Lessons learned

Our migration to Amazon MSK revealed several key insights that can help other organizations modernize their Kafka infrastructure:

  • Expert validation – Working with AWS Solutions Architects to validate our migration strategy caught several critical issues early. Although our team knew our applications well, external Kafka experts identified potential problems with state management and consumer offset handling that we hadn’t considered. This validation prevented costly missteps during the migration.
  • Data verification – Comparing data across Kafka clusters proved challenging. We built tools to capture topic snapshots in Parquet format on Amazon Simple Storage Service (Amazon S3), enabling quick comparisons using Amazon Athena queries. This approach gave us confidence that data remained consistent throughout the migration.
  • Start small – Beginning with our smallest data universe in QA helped us refine our process. Each subsequent migration went smoother as we applied lessons from previous iterations. This gradual approach helped us maintain system stability while building team confidence.
  • Detailed planning – We created specific migration plans with each team, considering their unique requirements and constraints. For example, our machine learning pipeline needed special handling due to strict offset management requirements. This granular planning prevented downstream disruptions.
  • Performance optimization – We found that utilizing Amazon MSK provisioned throughput offered clear cost advantages when storage throughput became a bottleneck. This feature made it possible to improve cluster performance without scaling instance sizes or adding brokers, providing a more efficient solution to our throughput challenges.
  • Documentation – Maintaining detailed migration runbooks proved invaluable. When we encountered similar issues across different migrations, having documented solutions saved significant troubleshooting time.

Conclusion

In this post, we showed you how we modernized our Kafka infrastructure by migrating to Amazon MSK. We walked through our decision-making process, challenges faced, and strategies employed. Our journey transformed Kafka operations from a resource-intensive, self-managed infrastructure to a streamlined, managed service, improving operational efficiency, platform reliability, and team productivity. For enterprises managing self-hosted Kafka infrastructure, our experience demonstrates that successful transformation is achievable with proper planning and execution. As data streaming needs grow, modernizing infrastructure becomes a strategic imperative for maintaining competitive advantage.

For more information, visit the Amazon MSK product page, and explore the comprehensive Developer Guide to learn about the features available to help you build scalable and reliable streaming data applications on AWS.

About the authors

Oleh Khoruzhenko

Oleh Khoruzhenko

Oleh is a Senior Staff DevOps Engineer at Bazaarvoice Inc, specializing in architecting and optimizing high-throughput data streaming solutions. He is an expert in the Apache Kafka ecosystem, utilizing Apache Kafka Streams for complex event processing and Apache Spark for large-scale data ingestion and transformation.

Christian Silva

Christian Silva

Christian is a Sr. Solutions Architect at AWS based in Houston, TX. He works with independent software vendor customers, helping them build and optimize their solutions on AWS. With a background in cloud architecture, Christian is passionate about networking and security, guiding customers to implement robust and efficient cloud infrastructures. Outside of work, he enjoys spending time with his kids, playing soccer, and fishing.

Aravind Marthineni

Aravind Marthineni

Aravind is a Technical Account Manager at AWS based in Austin, TX. He supports SaaS customers in ecommerce and social media analytics on migrations and modernizations using cloud-based architectures. He specializes in cloud governance and is passionate about helping customers operate efficiently on the cloud using generative AI. When not working, he loves playing and teaching cricket to his 2-year-old and cooking Indian delicacies.

How Slack achieved operational excellence for Spark on Amazon EMR using generative AI

Post Syndicated from Avijit Goswami original https://aws.amazon.com/blogs/big-data/how-slack-achieved-operational-excellence-for-spark-on-amazon-emr-using-generative-ai/

At Slack, our data platform processes terabytes of data each day using Apache Spark on Amazon EMR on Amazon Elastic Compute Cloud (Amazon EC2), powering the insights that drive strategic decision-making across the organization.

As our data volume expanded, so did our performance challenges. With traditional monitoring tools, we couldn’t effectively manage our systems when Spark jobs slowed down or costs spiraled out of control. We were stuck searching through cryptic logs, making educated guesses about resource allocation, and watching our engineering teams spend hours on manual tuning that should have been automated. That’s why we built something better: a detailed metrics framework designed specifically for Spark’s unique challenges. This is a visibility system that gives us granular insights into application behavior, resource usage, and job-level performance patterns we never had before. We’ve achieved 30–50% cost reductions and 40–60% faster job completion times. This is real operational efficiency that directly translates to better service for our users and significant savings for our infrastructure budget. In this post, we walk you through exactly how we built this framework, the key metrics that made the difference, and how your team can implement similar monitoring to transform your own Spark operations.

Why comprehensive Spark monitoring matters

In enterprise environments, poorly optimized Spark jobs can waste thousands of dollars in cloud compute costs, block critical data pipelines affecting downstream business processes, create cascading failures across interconnected data workflows, and impact service level agreement (SLA) compliance for time-sensitive analytics.

The monitoring framework we’re examining captures over 40 distinct metrics across five key categories, providing the granular insights needed to prevent these issues.

How we ingest, process, and act on Spark metrics

To address the challenges of managing Spark at scale, we developed a custom monitoring and optimization pipeline—from metric collection to AI-assisted tuning. It begins with our in-house Spark listener framework, which captures over 40 metrics in real time across Spark applications, jobs, stages, and tasks while pulling critical operational context from tools such as Apache Airflow and Apache Hadoop YARN.

An Apache Airflow-orchestrated Spark SQL pipeline transforms this data into actionable insights, surfacing performance bottlenecks and failure points. To integrate these metrics into the developer tuning workflow, we expose a metrics tool and a custom prompt through our internal analytics model context protocol (MCP) server. This enables seamless integration with AI-assisted coding tools such as Cursor or Claude Code.

The following is the list of tools used for our Spark monitoring solution, which includes metric collection to AI-assisted tuning:

The result is fast, reliable, deterministic Spark tuning without the guesswork. Developers get environment-aware recommendations, automated configuration updates, and ready-to-review pull requests.

Deep dive into Spark metrics collection

At the center of our real-time monitoring solution lies a custom Spark listener framework that captures thorough telemetry across the Spark lifecycle. Spark’s built-in metrics are often coarse, short‑lived, and scattered across the user interface (UI) and logs, which leaves four critical gaps:

  1. Consistent historical record
  2. Weak linkage from applications to jobs to stages to tasks
  3. Limited context (user, cluster, team)
  4. Poor visibility into patterns such as skew, spill, and retries

Our expanded listener framework closes these gaps by unifying and enriching telemetry with environment and configuration tags, building a durable, queryable history, and correlating events across the execution graph. It explains why tasks fail, pinpoints where memory or CPU pressure occurs, compares intended configurations to actual usage, and produces clear, repeatable tuning recommendations so teams can baseline behavior, minimize waste, and resolve issues faster. The following architecture diagram illustrates the flow of the Spark metrics collection pipeline.

Spark metrics ingestion architecture diagram

Spark listener

Our listener framework captures Spark metrics at four distinct levels:

  1. Application metrics: Overall application success/failure rates, total runtime, and resource allocation
  2. Job-level metrics: Individual job duration and status tracking within an application
  3. Stage-level metrics: Stage execution details, shuffle operations, and memory usage per stage
  4. Task-level metrics: Individual task performance for deep debugging scenarios

The following Scala example code shows the SparkTaskListener extends the class SparkListener to capture detailed task-level metrics:

class SparkTaskListener(conf: SparkConf) extends SparkListener {
 val taskToStageId = new mutable.HashMap[Long, Int]()
 val stageToJobID = new mutable.HashMap[Int, Int]()
 private val emitter: Emitter = getEmitter(conf)
  override def onTaskStart(taskStart: SparkListenerTaskStart): Unit = {
   taskToStageId += taskStart.taskInfo.taskId -> taskStart.stageId 
 }
 override def onTaskEnd(taskEnd: SparkListenerTaskEnd): Unit = {
   val taskInfo = taskEnd.taskInfo
   val taskMetrics = taskEnd.taskMetrics
   val jobId = stageToJobID.apply(taskToStageId.apply(taskInfo.taskId))
   val metrics = Map[String, Any](
     "event_type" -> "task_metric",
     "job_id" -> jobId,
     "task_id" -> taskInfo.taskId,
     "duration" -> taskInfo.duration,
     "executor_run_time" -> taskMetrics.executorRunTime,
     "memory_bytes_spilled" -> taskMetrics.memoryBytesSpilled,
     "bytes_read" -> taskMetrics.inputMetrics.bytesRead,
     "records_read" -> taskMetrics.inputMetrics.recordsRead
     // additional metrics.....
   )
   emitter.report(convertToJson(metrics))
 }
}

Real-time streaming to Kafka

These metrics are streamed in real time to Kafka as JSON-formatted telemetry using a flexible emitter system:

class KafkaEmitter(conf: SparkConf) extends Emitter {
     private val broker = conf.get("spark.custom.listener.kafkaBroker", "<broker_address>")
     private val topic = conf.get("spark.custom.listener.kafkaTopic", "<topic_name>")
     private var producer: Producer[String, Array[Byte]] = _
     override def report(str: String): Unit = {
         val message = str.getBytes(StandardCharsets.UTF_8)
         producer.send(new ProducerRecord[String, Array[Byte]](topic, message))
     }
}

From Kafka, a downstream pipeline ingests these records into an Apache Iceberg table.

Context-rich observability

Beyond standard Spark metrics, our framework captures essential operational context:

  • Airflow integration: DAG metadata, task IDs, and execution timestamps
  • Resource tracking: Configurable executor metrics (heap usage, execution memory)
  • Environment context: Cluster identification, user tracking, and Spark configurations
  • Failure analysis: Detailed error messages and task failure root causes

The combination of thorough metrics collection and real-time streaming has redefined Spark monitoring at scale, laying the groundwork for powerful insights.

Deep dive into Spark metrics processing

When raw metrics—often containing millions of records—are ingested from various sources, a Spark SQL pipeline transforms this high-volume data into actionable insights. It aggregates the data into a single row per application ID, significantly reducing complexity while preserving key performance signals.

For consistency in how teams interpret and act on this data, we apply the Five Pillars of Spark Monitoring, a structured framework that turns raw telemetry into clear diagnostics and repeatable optimization strategies, as shown in the following table.

Pillar Metrics Key purpose/insight Driving event
Application metadata and orchestration details
  • YARN metadata (app, attempt, allocated memory, compute cluster, final job status, run duration)
  • Airflow metadata (DAG, task, owner)
Correlate performance patterns with teams and infrastructure to identify inefficiencies and ownership.
  • Airflow metadata
  • YARN metadata on Amazon EMR on EC2
User-specified configuration
  • Given memory (driver, executor)
  • Dynamic allocation (min/max/initial executor count)
  • Cores per executor
  • Shuffle partitions
Compare configuration as opposed to actual performance to detect over- and under-provisioning and optimizing costs. This is where significant cost savings often hide. Spark event:

  • app_metric
Performance insights
  • Maximum skew ratio (75th percentile as opposed to max shuffle_total_bytes_read by Spark tasks per stage)
  • Total spill
  • Spark stage/task retry/failure
This is where the real diagnostic power lies. These metrics identify the three primary stoppers of Spark performance: skew, spill, and failures. Spark event:

  • task_metric
  • stage_metric
Execution insights
  • Spark job/stage/task count
  • Spark job/stage/task duration
Understand runtime distribution, identify bottlenecks, and highlight execution outliers. Spark event:

  • task_metric
  • stage_metric
  • job_metric
Resource usage and system health
  • Peak JVM heap memory
  • Max GC overhead %
Reveal memory inefficiencies and JVM-related pressure for cost and stability improvements. Comparing these against given configs helps identify waste and optimize resources. Spark event:

  • task_metric
  • stage_metric
  • executor_metric

AI-powered Spark tuning

The following architecture diagram illustrates the use of agentic AI tools to analyze the aggregated Spark metrics.

AI-powered Spark tuning diagram

To integrate these metrics into a developer’s tuning workflow, we build a custom Spark metrics tool and a custom prompt that any agent can use. We use our existing analytics service, a homegrown web application that users can query our data warehouse with, build dashboards, and share insights. The backend is written in Python using FastAPI, and we expose an MCP server from the same service by using FastMCP. By exposing the Spark metrics tool and custom prompt through the MCP server, we make it possible for developers to connect their preferred assisted coding tools (Cursor, Claude Code, and more) and use data to guide their tuning.

Because the data exposed by the analytics MCP server might be sensitive, we use Amazon Bedrock in our Amazon Web Services (AWS) account to provide the foundation models to our MCP clients. This keeps our data more secure and facilitates compliance because it never leaves our AWS environment.

Custom prompt

To create our custom prompt for AI-driven Spark tuning, we design a structured, rule-based format that encourages more deterministic and standardized output. The prompt defines the required sections (application overview, current Spark configuration, job health summary, resource recommendations, and summary) for consistency across analyses. We include detailed formatting rules, such as wrapping values in backticks, avoiding line breaks, and enforcing strict table structures to maintain clarity and machine readability. The prompt also embeds explicit guidance for interpreting Spark metrics and mapping them to recommended tuning actions based on best practices, with clear criteria for status flags and impact explanations. The prompt means that the AI’s recommendations can be traced, reproduced, and actioned based on the provided data by tightly controlling the input-output flow and attempting to prevent hallucinations.

Final results

The screenshots in this section show how our tool performed the analysis and provided recommendations. The following is a performance analysis for an existing application.

performance analysis for an existing application

The following is a recommendation to reduce resource waste.

recommendation to reduce resource waste

The impact

Our AI-powered framework has fundamentally changed how Spark is monitored and managed at Slack. We’ve transformed Spark tuning from a high-expertise, trial-and-error process into an automated, data-backed standard by moving beyond traditional log-diving and embracing a structured, AI-driven approach. The results speak for themselves, as shown in the following table.

Metric Before After Improvement
Compute cost Non-deterministic Optimized resource use Up to 50% lower
Job completion time Non-deterministic Optimized Over 40% faster
Developer time on tuning Hours per week Minutes per week >90% reduction
Configuration waste Frequent over-provisioning Precise resource allocation Near-zero waste

Conclusion

At Slack, our experience with Spark monitoring shows that you don’t need to be a performance expert to achieve exceptional results. We’ve shifted from reacting to performance issues to preventing them by systematically applying five key metric categories.

The numbers speak for themselves: 30–50% cost reductions and 40–60% faster job completion times represent operational efficiency that directly impacts our ability to serve millions of users worldwide. These improvements compound over time as teams build confidence in their data infrastructure and can focus on innovation rather than troubleshooting.

Your organization can achieve similar outcomes. Start with the basics: implement comprehensive monitoring, establish baseline metrics, and commit to continuous optimization. Spark performance doesn’t require expertise in every parameter, but it does require a strong monitoring foundation and a disciplined approach to analysis.

Acknowledgments

We want to give our thanks to all the people who have contributed to this incredible journey: Johnny Cao, Nav Shergill, Yi Chen, Lakshmi Mohan, Apun Hiran, and Ricardo Bion.


About the authors

Nilanjana Mukherjee

Nilanjana Mukherjee

Nilanjana is a staff software engineer at Slack, bringing deep technical expertise and engineering leadership to complex software challenges. She specializes in building high-performance data systems, focusing on data pipeline architecture, query optimization, and scalable data processing solutions.

Tayven Taylor

Tayven Taylor

Tayven is a software engineer I on Slack’s Data Foundations team, where he helps maintain and optimize large-scale data systems. His work focuses on Spark and Amazon EMR performance, cost optimization, and reliability improvements that keep Slack’s data platform efficient and scalable. He’s passionate about creating tools and systems that make working with data faster, smarter, and more cost-effective.

Mimi Wang

Mimi Wang

Mimi is a staff software engineer on Slack’s Data Platform team, where she builds tools to facilitate data-driven decision-making at Slack. Recently she has been focusing on using AI to lower the barrier to entry for non-technical users to derive value out of data. Previously, she was on the Slack Security team focusing on a customer-facing real-time anomaly detection pipeline.

Rahul Gidwani

Rahul Gidwani

Rahul is a senior staff software engineer at Salesforce specializing in search infrastructure. He works on Slack’s data lake development and processing pipelines and contributing to open-source projects such as Apache HBase and Druid. Outside of work, Rahul enjoys rock climbing.

Prateek Kakirwar

Prateek Kakirwar

Prateek is a senior engineering manager at Slack leading the AI-first transformation of data engineering and analytics. With over 20 years of experience building large-scale data platforms, AI systems, and metrics frameworks, he focuses on scalable architectures that enable trusted, self-service analytics across the organization. He holds a master’s degree from the University of California, Berkeley.

Avijit Goswami

Avijit Goswami

Avijit is a principal specialist solutions architect at AWS specializing in data and analytics. He helps customers design and implement robust data lake solutions. Outside the office, you can find Avijit exploring new trails, discovering new destinations, cheering on his favorite teams, enjoying music, or testing out new recipes in the kitchen.

How Salesforce migrated from Cluster Autoscaler to Karpenter across their fleet of 1,000 EKS clusters

Post Syndicated from Sana Jawad original https://aws.amazon.com/blogs/architecture/how-salesforce-migrated-from-cluster-autoscaler-to-karpenter-across-their-fleet-of-1000-eks-clusters/

As organizations scale their Kubernetes deployments, Kubernetes cluster scaling has traditionally been complex and slow, requiring careful management of node groups and auto scaling configurations. Karpenter, an open source node provisioning project for Kubernetes, can help transform this approach by directly provisioning right-sized nodes based on real-time workload demands. A recent Datadog report reveals that the percentage of nodes provisioned by Karpenter rose by 22% in the last 2 years as organizations migrate from traditional auto scaling approaches. This growth underscores Amazon Web Services (AWS) leadership in cloud-based innovation and the container ecosystem’s recognition of Karpenter’s strong performance and cost efficiency benefits. The following post examines how Salesforce, operating one of the world’s largest Kubernetes deployments, successfully migrated from Cluster Autoscaler to Karpenter across their fleet of 1,000 plus Amazon Elastic Kubernetes Service (Amazon EKS) clusters.

Salesforce operates one of the world’s most complex Kubernetes platforms, managing over 1,000 EKS clusters that serve thousands of internal tenants across the company. These clusters power a wide range of applications, from mission-critical services to experimental projects, and demand a high degree of scalability, reliability, and operational efficiency.

As the platform grew, Salesforce’s Kubernetes platform team began to face major hurdles with its traditional auto scaling approach based on AWS Auto Scaling groups and the Kubernetes Cluster Autoscaler. These limitations hampered the team’s ability to respond to application demands quickly, optimize compute resources, and empower internal developers to self-serve infrastructure needs.

To address these challenges, Salesforce undertook a large-scale migration to Karpenter, an open source Kubernetes [1] auto scaler built by AWS. This blog post details the motivation behind the transition, the implementation strategy, the challenges encountered along the way, and the impact it had on cost, performance, and operational complexity.

Opportunity for operational transformation

At Salesforce’s massive scale, the traditional Kubernetes infrastructure faced several critical challenges. The need to accommodate diverse workload requirements led to a proliferation of thousands of node groups and Auto Scaling groups, creating operational bottlenecks and slowing innovation. This architectural complexity was compounded by significant scaling performance issues, where the Auto Scaling group-dependent Cluster Autoscaler struggled to handle dynamic workloads, often resulting in multi-minute delays during demand spikes and degraded user experience. Resource utilization suffered as well, with inefficient bin-packing and conservative scale-down strategies leading to stranded resources and underutilized infrastructure—a particular concern given Salesforce’s focus on cost-to-serve and sustainability goals. These challenges were further exacerbated by structural limitations in the Auto Scaling group–based architecture, including poor Availability Zone balance and performance bottlenecks in large clusters, particularly for memory-intensive workloads. The combination of these factors made it clear that a more modern, flexible auto scaling solution was essential for maintaining Salesforce’s competitive edge and operational efficiency.

Solution overview

To migrate over 1,000 production clusters, without disruption, Salesforce engineered a highly automated, risk-mitigated transition process centered on Karpenter. Here’s how the migration was executed.

At this scale, a manual migration was infeasible. The team developed an in-house Karpenter transition tool to orchestrate the switch-over safely and consistently, and a Karpenter patching check tool. Karpenter transition tool and Karpenter patching check tool provide a comprehensive solution for migrating Kubernetes clusters to and from Karpenter node management while maintaining operational continuity through automated node rotation, Amazon Machine Image (AMI) validation, and graceful pod eviction handling.

Key design principles included:

  • Zero disruption – The tool cordoned and drained legacy nodes with full respect for pod disruption budgets (PDBs), maintaining workload safety
  • Rollback support – A reverse transition capability allowed fast recovery to Auto Scaling group–based auto scaling if needed
  • Continuous integration and continuous delivery (CI/CD) integration – The tool was embedded in the core infrastructure provisioning pipeline, standardizing the migration across services.

This foundation enabled repeatability across thousands of clusters and node pools, inspiring confidence in Salesforce developers.

Automated configuration mapping

To convert existing Auto Scaling group configurations to Karpenter-based definitions, the team automated the mapping logic between legacy and modern configurations. For example:

  • Auto Scaling group instance types → EC2NodeClass instance types
  • Root volume sizes → Storage parameters in Karpenter config
  • Node labels → Applied in both NodePool and EC2NodeClass

With over 1,180 node pools containing highly diverse configurations, automation was essential to minimize errors and reduce manual toil.

Example:

metadata:
 name: m5.8xlarge-min-300-max-2500
data:
 k8s_instance_type: m6i.8xlarge
 k8s_root_volume_size: '100'
 k8s_root_volume_iops: '3000'
 k8s_root_volume_type: 'gp3'
 k8s_root_volume_throughput: '125'
 k8s_min_node_number: '300'
 k8s_max_node_number: '2500'
 multi_az_provisioned_workers: 'false'
 asg_launch_type: 'launch_template'
 gpu_enabled: 'false'

A deliberate, phased rollout strategy was adopted:

  • Mid-2025 to Early 2026 – A multistage migration across internal environments with soak times between stages
  • Start with lower-risk environments – Less critical workloads were migrated first to validate tooling and operational processes
  • Risk-based sequencing – High-stakes production environments continue to be migrated last after testing the process

By using this approach Salesforce, continuously learned and adapted, avoiding large-scale regressions.

Key insights from the migration

During this migration journey, the Salesforce team gained valuable insights and best practices that we’ll share to help guide your own transformation initiatives.

Managing application availability during nude Updates

PDBs emerged as a critical consideration during the migration because several services had overly restrictive or misconfigured PDBs that blocked node replacements. The team addressed this by identifying problematic configurations, partnering with application owners on remediation, and implementing Open Policy Agent (OPA) policies for proactive PDB validation. This experience highlighted how proper PDB configuration is essential for safe auto scaling and helped establish stronger governance practices.

Optimizing node maintenance workflows

The initial migration approach of cordoning Karpenter nodes in parallel led to unexpected cluster health issues. To address this, the team refined their strategy by implementing sequential node cordoning, adding manual verification checkpoints with rollback capabilities, and deploying enhanced monitoring for early detection of cluster instability. This experience reinforced that even with modern infrastructure tooling, careful orchestration of node maintenance remains crucial for system reliability.

Understanding Kubernetes label constraints

During the migration, the team discovered that Salesforce’s human-friendly legacy naming conventions often exceeded Kubernetes’s 63-character label length limit, creating challenges with Karpenter’s label-dependent operations. The team resolved this by refactoring naming conventions across node pools to comply with Kubernetes standards. This experience highlighted how seemingly minor technical constraints, such as label length limits, can become significant blockers in automated infrastructure management if not properly addressed early in the migration process.

For example, the following name is 67 characters long:

analytics-bigdata-spark-executor-pool-m6a-32xlarge-az-a-b-c

It produced the result:

error: metadata.labels: Invalid value: must be no more than 63 characters

Protecting single-instance applications

The team discovered that Karpenter’s efficient bin-packing and consolidation features could unexpectedly impact applications running single-replica pods, leading to service disruptions in critical scenarios. To address this, we began implementing guaranteed pod lifetime features and workload-aware disruption policies to safeguard these singleton workloads. This experience demonstrated that effective auto scaling solutions must balance infrastructure efficiency with application availability requirements, particularly for mission-critical services.

Managing storage requirements in node migrations

The migration revealed that certain workloads failed to schedule due to incomplete ephemeral storage configurations. The team resolved this by implementing precise 1:1 mappings between the original Auto Scaling group–defined volume settings and Karpenter’s EC2NodeClass parameters. This experience emphasized the importance of carefully translating storage requirements during infrastructure migrations, particularly for I/O-intensive applications.

Realized value

The transition to Karpenter delivered measurable impact across multiple dimensions—performance, cost, and developer experience.

Operational efficiency

Salesforce eliminated thousands of node groups, significantly simplifying infrastructure management across its Kubernetes platform. Manual operational overhead was reduced by 80% through automation and the introduction of self-service capabilities. Developers can now define their own node pool requirements without waiting for centralized approvals, resulting in faster onboarding and greater agility.

Performance gains

With Karpenter, scaling latency was reduced from minutes to seconds by provisioning nodes based on actual pending pods, effectively bypassing delays associated with Auto Scaling groups. Node utilization improved significantly due to advanced bin-packing algorithms, resulting in fewer stranded resources and better efficiency. The migration eliminated Auto Scaling group thrashing, leading to more stable workloads and fewer scaling events during traffic spikes.

Cost optimization

Salesforce achieved 5% in cost savings in FY2026 by improving bin-packing efficiency and reducing idle capacity across its Kubernetes clusters. With the Karpenter rollout still in progress, an additional 5–10% in savings is projected for FY2027. The migration also lowered the overall cost-to-serve (CTS) by reducing the number of required nodes and improving multi-instance handling.

Enhanced developer and customer experience

The migration to Karpenter introduced true self-service infrastructure, allowing developers to define their capacity needs through straightforward node pool declarations. It also enabled greater flexibility by supporting heterogeneous instance types, including GPU, ARM, and x86, within a single node pool. Karpenter further improved IP efficiency by decoupling node provisioning from specific subnets, helping reduce IP fragmentation and exhaustion across the platform.

Conclusion

The migration to Karpenter represents a fundamental shift in how Salesforce manages Kubernetes infrastructure at scale. By addressing the limitations of traditional auto scaling approaches, we’ve achieved significant improvements in operational efficiency, cost optimization, and customer experience.

The key to our success was a combination of careful planning, custom tooling, and a phased approach that prioritized stability and zero-disruption migration. The results demonstrate that modern Kubernetes auto scaling solutions like Karpenter can transform platform operations while maintaining the reliability required for enterprise-scale deployments.

Salesforce’s success with Amazon EKS and Karpenter demonstrates how AWS continues to innovate alongside its largest enterprise customers, delivering solutions that scale from hundreds to thousands of clusters while reducing costs and operational complexity. This partnership showcases the power of combining AWS managed Kubernetes service with open source innovations like Karpenter to solve real-world challenges at unprecedented scale. To learn more, refer to the Karpenter Best Practices Guide in the Amazon EKS documentation.


About the Authors

How Taxbit achieved cost savings and faster processing times using Amazon S3 Tables

Post Syndicated from Larry Christensen original https://aws.amazon.com/blogs/big-data/how-taxbit-achieved-cost-savings-and-faster-processing-times-using-amazon-s3-tables/

In this post, we discuss how Taxbit partnered with Amazon Web Services (AWS) to streamline their crypto tax analytics solution using Amazon S3 Tables, achieving 82% cost savings and five times faster processing times.

Taxbit is a leading tax compliance suite serving cryptocurrency exchanges, digital platforms, and government agencies, generating more than 100 million forms for users and reconciling more than 500 billion digital asset transactions. The suite powers a complex environment that handles real-time pricing data from 29 cryptocurrency exchanges covering over 10,000 digital assets.

Recently, Taxbit experienced challenges with their pricing data infrastructure. As data volumes continued to expand, infrastructure costs rose sharply, putting pressure on operational budgets. At the same time, the system struggled to efficiently ingest the growing number of pricing data points, creating persistent bottlenecks in their data pipeline. These technical limitations led to customers missing data and experiencing slow processing times, leading to dissatisfaction. In addition to these operational challenges, Taxbit has strict regulatory compliance requirements to be considered when designing solutions. This combination of issues led Taxbit to modernize their pricing data infrastructure with a focus on helping to meet regulatory standards.

“During peak workloads, our solutions process hundreds of millions of digital asset transactions across blockchain and cryptocurrency exchanges,”

– says Clark Roberts, CTO at Taxbit.

“Our legacy database architecture was becoming a bottleneck, leading to increased costs and slower response times for our enterprise and government customers.”

Solution overview

Taxbit’s modernized architecture uses Amazon S3 Tables with Apache Iceberg as the foundation, combined with purpose-built AWS services for data ingestion, processing, and analytics. The solution processes real-time pricing data from 29 cryptocurrency exchanges including over 10,000 digital assets. This architecture is shown in the following diagram.

This AWS cloud architecture diagram illustrates a comprehensive data pipeline for processing digital assest market data.

The data pipeline architecture uses AWS services to deliver a comprehensive solution. At its foundation, Amazon S3 Tables provides the scalable storage infrastructure necessary for managing large volumes of pricing data. For data processing and transformation, the solution combines Amazon EMR and AWS Glue, handling both extract, transform, and load (ETL) operations and asynchronous API requirements efficiently.

Real-time data handling is managed through Amazon Kinesis, enabling streaming of pricing updates. AWS Lambda functions perform multiple tasks, including periodic polling of vendor APIs, transformation of streaming data, and data enrichment. The orchestration of these components is managed by AWS Step Functions, helping to ensure coordination of data workflows. Completing the architecture, Amazon Athena provides query capabilities, supporting both synchronous APIs and one-time analytical queries. This approach creates a scalable system built to handle both real-time and batch processing workflows while maintaining high performance and reliability.

Data ingestion layer

The ingestion layer operates through two key components: API integration and stream processing. The API integration uses Lambda functions to systematically poll multiple external APIs. These polling operations are orchestrated by Amazon EventBridge, which manages the scheduled data collection tasks. Additionally, WebSocket listeners maintain continuous connections to capture real-time price updates as they occur.

On the stream processing side, Amazon Kinesis Data Streams serves as the backbone for handling real-time data ingestion at scale. As data flows in, Lambda functions perform transformations and enrichment operations to prepare the data for downstream use. Throughout this process, custom validation checks are applied to help ensure the quality and completeness of the data, helping to maintain the integrity of the pricing information pipeline.

Data storage layer

At the storage layer, Taxbit uses Amazon S3 Tables because of its optimized storage format designed for analytical queries. Amazon S3 Tables is designed to automatically handle table optimization and compaction, helping to streamline data management processes. The system also incorporates time-travel capabilities, allowing Taxbit to meet audit requirements and their need for historical data analysis.

The data organization strategy is designed to maximize efficiency and accessibility. Data is systematically partitioned by date and exchange, allowing for targeted data retrieval and improved query performance. The implementation of columnar storage further enhances query efficiency by minimizing unnecessary data scans. Additionally, version control mechanisms are in place to maintain clear data lineage, enabling precise tracking of data changes and transformations over time.

Analytics layer

At the analytics layer, the query engine forms the foundation, using Amazon Athena to facilitate flexible ad-hoc analysis of the pricing data. This is complemented by Presto-based queries that handle complex aggregations efficiently. The system includes carefully crafted execution plans optimized for common query patterns, designed to provide consistent and reliable performance.

To maximize efficiency, the analytics layer incorporates several key performance optimizations. The system uses an Athena reuse query result to minimize redundant processing and parallel query execution capabilities to handle multiple simultaneous requests effectively.

Security and compliance

The data protection strategy implements multiple layers of security, starting with AWS Key Management Service (AWS KMS) encryption for all data at rest. This is complemented by TLS encryption for data in transit, helping to secure data movement throughout the system. Access to data and resources is controlled through AWS Identity and Access Management (IAM), providing fine-grained permissions that enforce the principle of least privilege.

The audit trail component provides comprehensive monitoring and compliance capabilities. AWS CloudTrail logging captures detailed records of system activities, enabling thorough security analysis and incident investigation. Data lineage tracking maintains clear records of data movement and transformations throughout the pipeline. These features are augmented by robust compliance reporting capabilities, helping the system demonstrate adherence to regulatory requirements and internal governance policies. Together, these security controls create an environment that protects sensitive data, maintains transparency, and provides accountability.

Business impact

Most notably, Taxbit achieved an 82% reduction in storage infrastructure costs, while simultaneously delivering processing speeds five times faster than their previous architecture. Data completeness for calculations achieved approximately 99.99% accuracy and the workload can now successfully support over 10,000 digital assets.The benefits extended beyond these quantitative improvements. Customer experience has improved, with transaction pricing times shrinking from hours to minutes. Higher throughput capabilities increased operational efficiency, enabling faster data loading while reducing compute costs. The new architecture also established a scalable foundation that provides faster data access and the flexibility to expand into new markets. The modern infrastructure has also enabled Taxbit to pursue new product offerings by supporting advanced analytics and real-time insights that were previously unattainable. These capabilities created new business opportunities and revenue streams that weren’t possible under the constraints of the legacy system.

Conclusion

Taxbit’s implementation of Amazon S3 Tables has transformed their cryptocurrency tax compliance solutions, delivering 82% cost savings and five times faster processing speeds. The modernized architecture, combining Amazon EMR, AWS Glue, Amazon Kinesis, and Lambda, now processes transactions in minutes instead of hours. Additionally, the architecture has helped Taxbit maintain approximately 99.99% data accuracy across more than 10,000 digital assets. Beyond operational improvements, this transformation has enabled new product offerings and real-time analytics capabilities. By partnering with AWS, Taxbit addressed their scaling challenges and built a foundation for continued innovation in the digital asset space.

For more information, see Amazon S3 Tables.


About the authors

Larry Christensen

Larry Christensen

Larry is a Principal Engineer at Taxbit based in the Salt Lake City area. He’s spearheaded many architectural, big data, and AI transformations across Taxbit.

Washim Nawaz

Washim Nawaz

Washim is an Analytics Specialist Solutions Architect at AWS with extensive professional experience building and tuning data warehouse and data lake solutions. He is passionate about helping customers modernize their data platforms with efficient, performant, and scalable analytics solutions. Outside of work, he enjoys watching sports and traveling.

Derek Ziehl

Derek Ziehl

Derek is a Senior Technical Account Manager (TAM) at AWS. He has a background designing large-scale network systems and managing cloud migrations. As a TAM he enjoys enabling customers to run resilient, optimized workloads on AWS.

Pranjal Gururani

Pranjal Gururani

Pranjal is a Solutions Architect at AWS based out of Seattle. Pranjal works with various customers to architect cloud solutions that address their business challenges. He enjoys hiking, kayaking, skydiving, and spending time with family during his spare time.

How Socure achieved 50% cost reduction by migrating from self-managed Spark to Amazon EMR Serverless

Post Syndicated from Junaid Effendi, Pengyu Wang original https://aws.amazon.com/blogs/big-data/how-socure-achieved-50-cost-reduction-by-migrating-from-self-managed-spark-to-amazon-emr-serverless/

Socure is one of the leading providers of digital identity verification and fraud solutions. Its predictive analytics platform applies artificial intelligence (AI) and machine learning (ML) techniques to process both online and offline intelligence, including government-issued documents, contact information (email, phone, address), personal identifiers (DOB, SSN), and device or network data (IP, velocity) to verify identities accurately and in real time.

Socure ID+ is an identity verification platform that uses multiple Socure offerings such as KYC, SIGMA, eCBSV. Phone Risk and more. It has two environments focused on proof of concept (POC) and live customers. The Data Science (DS) environment is designed for the POC or proof of value (POV) stage. In this environment, customers provide datasets via SFTP, which are processed by Socure’s data scientists through an internal endpoint. The data undergoes ML-based scoring and other intelligence calculations depending on the selected modules and processed results are stored in Amazon Simple Storage Service (Amazon S3) in delta open table format . In the Production (Prod) environment, customers can verify identities either in real time through live endpoints or via a batch processing interface.

Socure’s data science environment includes a streaming pipeline called Transaction ETL (TETL), built on OSS Apache Spark running on Amazon EKS. TETL ingests and processes data volumes ranging from small to large datasets while maintaining high-throughput performance.

The primary purpose of this pipeline is to give data scientists a flexible environment to run POC workloads for customers.

Data scientists…

  • trigger ingestion of POC datasets, ranging from small batches to large-scale volumes.
  • consume the processed outputs written by the pipeline for analysis and model development.
  • share the results with Socure’s customers.

The following diagram shows the Transaction ETL (TETL) architecture.

Transaction ETL architecture

This pipeline directly supports customer POCs, ensuring that the right data is available for experimentation, validation, and demonstration. As such, it is a critical link between raw data and customer-facing outcomes, making its reliability and performance essential for delivering value. In this post, we show how Socure was able to achieve 50% cost reduction by migrating the TETL streaming pipeline from self-managed spark to Amazon EMR serverless.

Motivation

As data volumes have scaled by 10x, several challenges like latency and data reliability have emerged that directly impact the customer experience:

  • Performance issues due to inefficient autoscaling leading to increase in latency up to 5x
  • High operational cost of maintaining an OSS Spark environment on EKS

Additionally, we have identified other important issues:

  • Resource constraints due to instance provisioning limits, forcing the use of smaller nodes. This leads to frequent spark executor out of memory (OOM) failures under heavy loads, increasing job latency and delaying data availability.
  • Performance bottlenecks with Delta Lake, where large batch operations such as OPTIMIZE compete for resources and slow down streaming workloads.

During this migration, we also took the opportunity to transition to AWS Graviton, enabling additional cost efficiencies as explained in this post.

With these two primary drivers we began exploring alternative architecture using Amazon EMR. We already dd extensive benchmarking on several identity verification related batch workloads on different EMR platforms and came to the conclusion that Amazon EMR Serverless (EMR-S) offers a path to reduce operational cost, improve reliability, and better handle large-scale batch and streaming workloads; tackling both customer-facing issues and platform-level inefficiencies.

The new pipeline architecture

The data processing pipeline follows a two-stage architecture where streaming data from Amazon Kinesis Data Stream first flows into the raw layer, which parses incoming data into large JSON blobs, applies encryption, and stores the results in append-only Delta Tables. The processed layer consumes data from these raw Delta tables, performs decryption, transforms the data into a flattened and wide structure with proper field parsing, applies individual encryption to personally identifiable information (PII) fields, and writes the refined data to separate append-only Delta Tables for downstream consumption.

The following diagram shows the TETL before/after architecture we implemented, transitioning from OSS Spark on EKS to Spark on EMR Serverless.

Transaction ETL architecture

Benchmarking

We benchmarked end-to-end pipeline performance across OSS Spark on EKS and EMR Serverless. The evaluation focused on latency and cost under comparable resource configurations.

Resource Configuration

EKS (OSS Spark):

  • Min 30 executors
  • Max 90 executors
  • 14 GB memory / 2 cores per executor

EMR Serverless:

  • Min 10 executors
  • Max 30 executors
  • 27 GB memory / 4 cores per executor
  • Effectively ~60 executors when normalized for 2x memory and cores, designed to mitigate the OOM issues described earlier.

Observations

  • Autoscaling Efficiency: EMR Serverless scaled down effectively to 20 workers on average over the weekend (low traffic day), resulting in lower costs up to 12% compared to weekday.
  • Executor Sizing: Larger executors on EMR Serverless prevented OOM failures and improved stability under load.

Definitions

  • Cost: It is the service cost for both raw & processed jobs from the AWS Cost Explorer.
  • Latency: End-to-end latency measures the time from Socure ID+ event generation until data arrives in the processed delta table, calculated as Inserted Date minus Event Date.

Results

The values in the following table represent percentage improvements observed when running on EMR compared to EKS.

Low Traffic (Weekend) Regular Traffic (Weekday)
Records Count ~1M ~5M
Min Latency (best case) 73.3% 69.2%
Avg Latency (representative workload) 51.0% 47.9%

Max Latency

(worst case)

12.3% 34.7%
Total Cost 57.1% 45.2%

Note: Even with a conservative 40% cost reduction applied to the EKS environment to account for Graviton, EMR-S remains approximately 15% cheaper.

Performance improvement graph

The benchmarking results clearly demonstrate that EMR Serverless outperforms OSS Spark on EKS for our end-to-end pipeline workloads. By moving to EMR Serverless, we achieved:

  • Improved performance: Average latency reduced by more than 50%, with consistently lower min and max latencies.
  • Cost efficiency: Overall pipeline execution costs dropped by more than half.
  • Scalability: Autoscaling optimized resource usage, further lowering cost during off-peak periods.
  • Operational overhead: EMR-S fully managed and serverless nature eliminates the need to maintain EKS and OSS Spark.

Conclusion

In this post, we showed how Socure transitioning to EMR Serverless not only resolved critical issues around cost, reliability, and latency, but also provided a more scalable and sustainable architecture for serving customer POCs effectively, enabling us to deliver results to customers faster and strengthen our position for potential custom contracts.


About the authors

Junaid Effendi

Junaid Effendi

Junaid is a Senior Data Engineer at Socure. He designs and builds data infrastructure, pipelines, and services for both batch and streaming workloads, enabling data-driven insights that power identity verification. In his free time, he enjoys writing tech blogs and playing soccer.

Pengyu Wang

Pengyu Wang

Pengyu is a Senior Manager of Data Engineering at Socure. He leads teams that design and build scalable data platforms and pipelines, driving high-quality data solutions that power identity verification and analytics. In his free time, he enjoys skiing in the winter and exploring new technologies.

Raj Ramasubbu

Raj Ramasubbu

Raj is a Senior Analytics Specialist Solutions Architect focused on big data and analytics and AI/ML with Amazon Web Services. He helps customers architect and build highly scalable, performant, and secure cloud-based solutions on AWS. Raj provided technical expertise and leadership in building data engineering, big data analytics, business intelligence, and data science solutions prior to joining AWS. He helped customers in various industries like healthcare, medical devices, life science, retail, asset management, car insurance, residential REIT, agriculture, title insurance, supply chain, document management, and real estate.

How Bayer transforms Pharma R&D with a cloud-based data science ecosystem using Amazon SageMaker

Post Syndicated from Avinash Erupaka original https://aws.amazon.com/blogs/big-data/how-bayer-transforms-pharma-rd-with-a-cloud-based-data-science-ecosystem-using-amazon-sagemaker/

This post was written with Avinash Erupaka from Bayer (IT PH, Drug Innovation platform)

How can pharmaceutical companies unlock the full potential of their data to drive breakthrough innovations? Bayer, a global leader in health and nutrition, is dedicated to tackling the pressing challenges of our time, including a growing and aging population and the strain on our planet’s ecosystems. Its mission of “Health for All, Hunger for None” drives its commitment to addressing societal and environmental needs through groundbreaking research. Bayer is focused on developing innovative solutions that make a tangible difference in the world and value for its customers, employees, and stakeholders. Headquartered in Leverkusen, Germany, Bayer operates across 80 countries and is pioneering a data science ecosystem that transforms how research teams access, analyze, and derive insights from complex scientific data.

By harnessing the power of data, analytics, artificial intelligence and machine learning (AI/ML), and generative AI, Bayer is creating a cloud-based Pharma R&D Data Science Ecosystem (DSE) on AWS that powers cutting-edge technologies and concepts with robust data management. In doing so, R&D teams can fully realize the potential of unified data and analytics.

In this post, we discuss how Bayer used the next generation of SageMaker to build a solution that unified data ingestion, storage, analytics, and AI/ML workflows. Built on data mesh principles, Bayer’s DSE integrates advanced data ingestion, storage, analytics, and ML workflows to enable agile experimentation and scalable insight generation. It democratizes access to analytics, fosters cross-Region collaboration, and provides flexible integration of structured, semi-structured, and unstructured data.

Challenges in pharmaceutical research

In pharmaceutical research, data has become the most critical asset for driving innovation. However, managing this data effectively presents unprecedented challenges and traditional data management approaches are becoming increasingly inadequate for complex, global research initiatives. Many pharma R&D organization face a complex ecosystem of data and analytics related obstacles that hinder scientific discovery and operational efficiency:

  • Siloed datasets – Research datasets are siloed across domains, limiting reuse and slowing discovery.
  • Multiple data modalities – Clinical trial data (structured), real-world evidence (semi-structured), and genomic files (unstructured) existed in isolation, complicating integration and analysis.
  • Inflexible ingestion capabilities – Systems that support batch processing (such as trial data), real-time data streams (for example, from lab equipment), and event-driven ingestion (such as regulatory updates).
  • Rising R&D costs – Disparate technologies and disconnected systems create operational inefficiencies and increased licensing and maintenance costs.
  • Inconsistent landscape to fully use ML – The absence of a unified data architecture and standardized, domain-agnostic MLOps workflows mean that data and analytics innovation is often ad hoc and non-repeatable. Teams lack a streamlined way to scale successful patterns, resulting in redundant efforts, longer development cycles, and missed opportunities for cross-domain synergy.
  • Disconnected architectures – Software solutions are not integrated into the wider unified ecosystem, resulting in silos, redundancies, and inefficiencies.

Recognizing these systemic challenges, Bayer embarked on a transformative journey. DSE is not just a technological solution, but a strategic reimagining of how research data and analytics could be used across a global organization. By bringing together cutting-edge technologies, standardized frameworks, a collaborative data mesh, and lakehouse architecture, Bayer set out to help researchers and engineers accelerate pharmaceutical innovation.

Finding a solution with the next generation of SageMaker

Bayer envisioned a unified data science ecosystem that would provide the following:

  • A unified collaborative development experience for all data scientists regardless of their location or specialization
  • Seamless access to both structured and unstructured data through a consistent interface
  • Built-in governance and compliance controls appropriate for pharmaceutical research
  • Scalable compute resources to handle the most complex analytical workloads

Bayer conducted a comprehensive evaluation of various solutions before selecting the next generation of SageMaker as the cornerstone of their new data science ecosystem. Although other options had merits, Bayer prioritized the following capabilities:

  • Access to multimodal data – Essential for genomics, proteomics, and advanced biomarker research
  • Centralized asset marketplace – Central hub to discover and reuse data, features, models, and other enterprise assets
  • Integrated tooling ecosystem – Streamlined access to key tools like Git, ETL, MLflow, and generative AI application builders in one place
  • Multi-domain and cross-Region support – Critical for global research collaboration
  • Price-performance – Necessary for sustainable, long-term scaling

The capabilities of Amazon SageMaker Unified Studio and Amazon SageMaker Catalog aligned with Bayer’s vision of decentralized mesh execution combined with centralized discovery and governance. They enabled teams to work with their preferred tools, such as Jupyter Notebooks or workflow builders, while maintaining discoverability and reusability of assets.

Solution overview

This section describes the key features and architecture of Bayer’s DSE built on SageMaker. The DSE solution addresses the identified challenges through a multi-layered architecture:

  • Breaking down data silos – Multimodal data ingestion capabilities of the solution break down data silos by enabling unified storage, processing of structured, semi-structured, and unstructured data through batch, streaming, and event-driven pipelines.
  • Handling diverse data modalities – A hybrid lakehouse architecture, built on Amazon Simple Storage Service (Amazon S3), Apache Iceberg, and Amazon Redshift, provides a flexible foundation for handling diverse data modalities and maturities while providing data consistency and accessibility.
  • Reducing costs through standardization – To address rising R&D costs and operational inefficiencies, pre-wired analytical workbenches offer standardized templates and integrated development environments (IDEs) that reduce redundancy and accelerate workflow development.
  • Unlocking AI/ML with Amazon SageMaker AI and Amazon Bedrock – Advanced AI/ML capabilities, powered by Amazon SageMaker AI and Amazon Bedrock, create a standardized, domain-agnostic MLOps environment that enables repeatable innovation and cross-domain synergy.
  • Managing tools ecosystem with end-to-end observability – Robust governance and observability features provide compliance and system reliability while integrating previously disconnected tools into a unified, well-monitored ecosystem that breaks down architectural silos and promotes efficient resource utilization.

The DSE architecture implements data mesh principles where data domains (omics, regulatory, clinical trials) are treated as products, with ownership and management responsibilities assigned to domain experts. These domains are decentralized for execution but remain discoverable and reusable through SageMaker Catalog. At the core of the architecture is a hybrid mesh lakehouse architecture that combines Amazon S3 and Iceberg, providing the flexibility to handle both structured and unstructured data efficiently. SageMaker Unified Studio provides an analytical layer where researchers can access the full suite of tools needed for their work. The following diagram illustrates this architecture.

architecture diagram showing Bayer's data science ecosystem

Impact

The first phase of Bayer’s DSE confirmed the next generation of SageMaker as a powerful foundation for their R&D DSE—designed to balance decentralized innovation with centralized governance through a scalable data mesh architecture. With this solution, Bayer can catalog and manage multimodal data assets—including structured and unstructured data, ML features, models, and custom scientific assets—with context-rich metadata across diverse Pharma R&D domains. Bayer is now positioned to onboard over 300 TB of biomarker data and integrate siloed omics, clinical, and chemistry data repositories into a cohesive environment. With integrated tools like JupyterLab Spaces, MLflow, and SageMaker AI Studio, the DSE platform is laying the groundwork for a comprehensive, GxP-aware ML workbench—paving the way to operationalize over 25 high-value ML use cases and support more than 100 data scientists across the organization.

“The Data Science Ecosystem is vital for developing our medicines,” says Daniel Gusenleitner, Mission Lead for the R&D Data Science Ecosystem. “It enhances our business workflows with advanced analytics, helping us accelerate the search for new treatments. By integrating data from the entire research and development process, we improve the chances of technical success and ensure our efforts are efficient. Unlocking our data also facilitates target discovery, leading to groundbreaking advancements in patient care.”

Next steps

Bayer has successfully begun their Data Science Ecosystem on the next generation of Amazon SageMaker and is working to onboard the first use case of advanced biomarker research. Building on the strong foundation, Bayer is also accelerating the evolution of the DSE solution with the following key enhancements:

  • Federated catalogs and cross-domain integration – Enabling search and reuse of data assets across therapeutic areas and business units
  • Advanced ontology and semantic layer – Enriching metadata with domain knowledge to support AI-based search, discovery, and reasoning
  • Adoption of generative and agentic AI workflows – Driving novel drug discovery and accelerating hypothesis generation

Conclusion

By leveraging the next generation of Amazon SageMaker to build their cloud-based Data Science Ecosystem, Bayer is creating a foundation for faster, more efficient research and discovery. Amazon SageMaker is unifying diverse data types, enabling global collaboration, and standardizing ML workflows to help position Bayer at the forefront of data-driven innovation.

To learn more and get started with the next generation of SageMaker, refer to Amazon SageMaker or the AWS console.


About the Authors

Avinash Erupaka

Avinash Erupaka

Avinash is a Principal Engineering Lead at Bayer’s Drug Innovation platform. With deep experience across pharmaceuticals, crop science, and consumer health, he has led large-scale transformations spanning cloud platforms, AI/ML, and data infrastructure. Avinash brings a unique blend of technical depth and business acumen, having worked across the life sciences value chain—from research to manufacturing. He holds a Master’s in Engineering and an Executive MBA, and is passionate about building scalable, reusable solutions to accelerate scientific discovery.

Modood Alvi

Modood Alvi

Modood was a Senior Solutions Architect at AWS. Modood is passionate about digital transformation and is committed to helping large enterprise customers across the globe accelerate their adoption of and migration to the cloud. Modood brings more than a decade of experience in software development, having held a variety of technical roles within companies like SAP and Porsche Digital. Modood earned his Diploma in Computer Science from the University of Stuttgart.

Radhika Kashyap

Radhika Kashyap

Radhika is a Senior Customer Solutions Manager at AWS. Radhika brings over a decade of experience in technical program management and works with AWS customers to accelerate their journey to the cloud. She holds a master’s degree in management information systems and a bachelor’s degree in information technology.

Building zero trust generative AI applications in healthcare with AWS Nitro Enclaves

Post Syndicated from Nathan Pogue original https://aws.amazon.com/blogs/compute/building-zero-trust-generative-ai-applications-in-healthcare-with-aws-nitro-enclaves/

In healthcare, generative AI is transforming how medical professionals analyze data, summarize clinical notes, and generate insights to improve patient outcomes. From automating medical documentation to assisting in diagnostic reasoning, large language models (LLMs) have the potential to augment clinical workflows and accelerate research. However, these innovations also introduce significant privacy, security, and intellectual property challenges.

Healthcare data often contains Protected Health Information (PHI), which is governed by strict regulations and compliance frameworks. At the same time, organizations or researchers who have invested substantial time and compute resources into training medical LLMs must protect their proprietary model architectures, weights, and fine-tuned datasets. Traditional deployment models necessitate mutual trust between the model publisher and the healthcare data provider — trust that sensitive data won’t be leaked, and that the model itself won’t be copied, tampered with, or exfiltrated. The absence of a secure and verifiable trust model between model publishers and consumers remains one of the main barriers to scaling generative AI in regulated medical environments.

To address this concern, both parties need a secure environment to publish and consume models without exposing data or intellectual property. Amazon Web Services (AWS) Nitro Enclaves provide isolated, attested, and cryptographically verified compute environments that help protect sensitive workloads. Model owners can encrypt their LLMs with AWS Key Management Service (AWS KMS) and allow only verified Nitro Enclaves to decrypt and run them, making sure that the model can’t be accessed outside the Nitro Enclave. Healthcare organizations and consumers can use this to process sensitive data within their own AWS environment entirely within the Nitro Enclave, helping keep PHI private and contained. Hardware-based attestation provides proof that the Nitro Enclave is running trusted code, so that both sides can exchange information with confidence.

In this post, we demonstrate how to deploy a publicly available foundational model (FM) using Nitro Enclaves for isolated, more secure compute, AWS KMS for model encryption, Amazon Simple Storage Service (Amazon S3) for storing model artifacts and images, and Amazon Simple Queue Service (Amazon SQS) for securely delivering queries, enabling private, privacy-preserving inferences while helping protect both model intellectual property and sensitive patient data.

Solution overview

This solution outlines how to build a more secure end-to-end pipeline that enables zero trust medical LLM publication and inference with Nitro Enclaves. This post demonstrates a guide for setting up an Amazon Elastic Compute Cloud (Amazon EC2) instance with Nitro Enclaves enabled, downloading and encrypting a publicly available FM to an S3 bucket with an AWS KMS key, sending medical text and image-based queries to an SQS queue for processing, and storing results in an Amazon DynamoDB table.

This project is intended solely for educational and demonstration purposes and isn’t suitable for production or clinical use. Its outputs aren’t validated for clinical accuracy and must not be used for patient care or medical decision-making. Before any real-world deployment, make sure that you implement comprehensive security, privacy, and compliance safeguards. These include health data protection controls, secrets management, and regulatory validation. Furthermore, you must consult the appropriate clinical, legal, and security experts.

For demonstration purposes, this solution is deployed in a single AWS account. Ideally, in production, it would be deployed across separate AWS accounts: one for the model owner and one for the model consumer. The model owner can use cross-account AWS Identity and Access Management (IAM) permissions and encrypted model sharing through AWS KMS to securely provide access to their model without exposing the underlying weights or logic. At the same time, the consumer can run sensitive inferences within their own environment, maintaining strict data privacy and zero trust principles. In a real-world implementation, the model provider should also establish a robust entitlement and licensing framework to manage customer access, enabling fine-grained control over who can invoke the model, track usage, and support license revocation to immediately remove permissions from specific customers when necessary.

The following diagram shows the solution architecture:

Scope of solution

The steps of the solution include:

  1. Amazon EC2 setup: An EC2 instance is launched with Nitro Enclaves and Trusted Platform Module (TPM) enabled. For this project, a c7i.12xlarge instance with a 150 GB Amazon EBS volume is used to provide the necessary compute resources for running LLMs.
  2. Public FM download: A publicly available FM is retrieved from Hugging Face and stored in an S3 bucket within the model consumer’s AWS account.
  3. Model encryption: The model is encrypted using AWS KMS envelope encryption. Only a Nitro Enclave presenting a valid attestation document can request the decryption key from AWS KMS, which helps prevent unauthorized access to the model weights outside the Nitro Enclave.
  4. Nitro Enclave setup: A Docker image containing the llama.cpp inference runtime is built and deployed inside the Nitro Enclave.
  5. Model decryption and setup: When the Nitro Enclave launched, it requests decryption of the model artifacts using its attestation credentials. Then, the model can be securely decrypted inside the memory of the Nitro Enclave and loaded by the llama.cpp server. This means that the decrypted model weights aren’t visible outside of the Nitro Enclave boundary.
  6. Medical query: Users can submit either text or image-based queries to the model. Queries are sent through vsock, a secure communication channel from the client application to the model server inside the Nitro Enclave. Image queries necessitate that users upload images to an S3 bucket. The upload event triggers an SQS queue, which signals the Amazon EC2 parent to fetch and send the image to the Nitro Enclave image for the medical LLM to process with its multimodal capabilities.
  7. Message history: Each interaction, including the user’s prompt and the model’s response, is logged to a DynamoDB table. This provides a persistent conversation history that enables traceability and auditing while keeping PHI securely stored within the consumer’s account. If necessary, the DynamoDB table can be encrypted and sealed for another layer of security and privacy.

About Google MedGemma 4B

Google MedGemma is a family of medically-optimized LLMs built on Gemma 3, with 4B and 27B parameter variants supporting both text and multimodal versions for medical image inputs. The 4B model offers efficiency and strong performance for multimodal tasks such as report generation and medical Q&A, while the 27B models excel at more demanding scenarios, such as electronic health record interpretation and complex longitudinal data analysis.

MedGemma models are well-suited for automated radiology report generation, clinical triage and documentation, patient education, medical image pre-interpretation, and medical education systems. The 4B model is ideal for portable or resource-constrained deployments, whereas the 27B multimodal delivers maximal performance.

In this project, MedGemma 4B serves as a reference medical LLM, showing how domain-adapted fine-tuning can enhance a model’s ability to interpret, reason about, and respond to complex medical queries. It also provides a foundation for exploring the safe and effective use of LLMs in healthcare applications, while being securely deployed within a Nitro Enclave. However, you can choose to deploy your own medical FM if needed. This is a deeper overview on the 4B model.

Prerequisites

To implement the proposed solution, make sure that you have the following:

  • The AWS Command Line Interface (AWS CLI) installed on your machine to create the EC2 instance.
  • AWS permissions with access to EC2 c7i.12xlarge instances and Nitro Enclaves.
  • Knowledge of Amazon S3, AWS KMS, Amazon SQS, AWS Lambda, and DynamoDB.
  • Basic knowledge of Nitro Enclaves and healthcare data security.
  • The GitHub repository cloned to your local machine.

Environment setup

The following sections outline how to set up your environment for this solution.

Create S3 buckets

In this solution, you create two S3 buckets: one for the model artifacts and one for the image inputs.

To create the S3 buckets

  1. Sign in to the Amazon S3 console, choose Create bucket, and follow the prompts to create a new S3 bucket.
  2. For the model artifact bucket, give it a unique name (for example AWSACCOUNTNUMBER-medgemma-model) in the same Region you use for the other project resources.
  3. Repeat the same process for the image bucket (for example AWSACCOUNTNUMBER-medgemma-image-inputs).
  4. Update the S3_BUCKET_NAME variable with your model bucket name in envelope_encrypt_model.sh and run.sh.

Create an SQS queue

When images are uploaded to the S3 image bucket, they are sent to an SQS queue for processing in sequential order by the model running in the Nitro Enclave.

To create an SQS queue

  1. Sign in to the Amazon SQS console, choose Create queue, and follow the prompts to create a new SQS queue.
  2. Choose Standard Queue, provide a name, leave the rest as default, and choose Create queue.
  3. Replace the SQS_QUEUE_URL variable in image_processor.py and lambda_function.py (in the client and assets folder, respectively) with your URL.

Create a Lambda function

For image-based queries, MedGemma 4B expects images encoded in base64 format to be passed in the prompt. To convert the images to this format, a Lambda function is invoked using an Amazon S3 trigger when an image is uploaded to the bucket.

To create a Lambda function

  1. Sign in to the Lambda console, choose Create function, and follow the prompts to create a new Lambda function from scratch.
  2. Choose a name, choose a Python runtime (for example Python 3.13), and paste in the Lambda function code from the assets folder.
  3. Next, update the Lambda function’s IAM role in Permissions under the Configuration tab with access to your S3 image bucket and the SQS queue that you created with inline policy permissions. Attach the following policies:
    1. Amazon S3 policy:
      {
       "Version": "2012-10-17",
       "Statement": [
           {
               "Sid": "Statement1",
               "Effect": "Allow",
               "Action": [
                   "s3:*"
               ],
               "Resource": [
                   "arn:aws:s3:::<IMAGE_BUCKET_NAME>",
                   "arn:aws:s3:::<IMAGE_BUCKET_NAME>/*"
               ]
           }
       ]
      }

    2. Amazon SQS policy:
      {
       "Version": "2012-10-17",
       "Statement": [
           {
               "Sid": "VisualEditor0",
               "Effect": "Allow",
               "Action": "sqs:ListQueues",
               "Resource": "*"
           },
           {
               "Sid": "VisualEditor1",
               "Effect": "Allow",
               "Action": "sqs:*",
              "Resource": "arn:aws:sqs:<REGION>:<ACCOUNT_NUMBER>:<QUEUE_NAME>"
           }
       ]
      }

  4. Finally, within the Lambda Designer, add a trigger, choose Amazon S3, and choose your image bucket. You should see the following example when the trigger is enabled.

AWS Lambda configuration interface showing S3 bucket trigger setup for medical image processing workflow

Create a DynamoDB table

When the queries have been processed by the model for inference, the prompts and responses are logged to a DynamoDB table for auditing and message history purposes.

To create a DynamoDB table

  1. Sign in to the DynamoDB console, choose Create table, and follow the prompts to create a new DynamoDB table.
  2. Give it a partition key named ID as a String type.
  3. Replace the TABLE_NAME variable with the table name and REGION variable with your AWS Region in direct_query.py and image_processor.py files.

Create an AWS KMS key

An AWS KMS key is used to envelope-encrypt the model artifacts before they are uploaded to the S3 model bucket. During encryption, the AWS KMS key policy is configured with conditions that restrict decryption to only those Nitro Enclaves presenting a valid attestation document. This attestation includes platform configuration registers (PCR) hashes that represent the measured state of the Nitro Enclave, which covers the signed Nitro Enclave image, runtime, and configuration. When the Nitro Enclave is launched, it generates an attestation document signed by the Amazon EC2 Nitro hypervisor, proving that its PCR values match the expected trusted measurements defined in the AWS KMS key policy. The key is released only if these PCR hashes align and the attestation is verified by AWS KMS, allowing the Nitro Enclave to decrypt and load the model securely in memory.

To create an AWS KMS key

  1. Sign in to the AWS KMS console, choose Create key, and follow the prompts to create a new AWS KMS key.
  2. Choose Symmetric as the key type and Encrypt and decrypt for the key usage. Make the alias AppKmsKey. Leave the default settings and choose Finish.
  3. Replace the REGION variable in vsock-proxy.yaml in the client folder with your AWS Region.

Create an EC2 instance

Now that the necessary resources are set up, you can proceed to launch the EC2 instance and create the Nitro Enclave image. For this solution, a c7i.12xlarge instance with a 150 GB EBS volume is provisioned.

To launch an EC2 instance with Nitro Enclaves enabled

  1. Within the GitHub repository on your local machine, run ./create_ec2.sh to create the EC2 instance.
cd scripts 
chmod +x create_ec2.sh 
./create_ec2.sh

  1. The script launches an EC2 instance called MedGemmaNitroEnclaveDemo. When the instance is running, you must create an IAM policy and add it to the Amazon EC2 IAM role with necessary permissions to the resources created previously.
  2. Sign in to the IAM console and navigate to Policies, choose Create policy, choose JSON, and paste the following policy, making sure that you update the bucket, queue URL, AWS Region, account number, and table variables:
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "s3modelbucket",
      "Effect": "Allow",
      "Action": [
        "s3:*"
      ],
      "Resource": [
        "arn:aws:s3:::<MODEL_BUCKET_NAME>",
        "arn:aws:s3:::<MODEL_BUCKET_NAME>/*"
      ]
    },
    {
      "Sid": "dynamodb",
      "Effect": "Allow",
      "Action": [
        "dynamodb:*"
      ],
      "Resource": [
        "arn:aws:dynamodb:<REGION>:<ACCOUNT_NUMBER>:table/<TABLE_NAME>"
      ]
    },
    {
      "Sid": "sqslist",
      "Effect": "Allow",
      "Action": "sqs:ListQueues",
      "Resource": "*"
    },
    {
      "Sid": "sqsqueue",
      "Effect": "Allow",
      "Action": "sqs:*",
      "Resource": "arn:aws:sqs:<REGION>:<ACCOUNT_NUMBER>:<QUEUE_NAME>"
    }
  ]
}
  1. Give it a name (for example enclave-permissions) and choose Create policy.
  2. Navigate to Roles, choose Create role, choose EC2 as the AWS service for the Trusted entity type, then choose your policy that you created under Add permissions.
  3. Update your EC2 instance to use the role by going to the Security setting under the Actions dropdown, then modifying its IAM role.
  4. You can upload the modified repository to your EC2 instance using SCP. Alternatively, you can transfer the repository through rsync
rsync -avz -e "ssh -i /path/to/directory/sample-for-secure-medical-llm-inference-with-nitro-enclaves/nitro-enclave-key.pem" --exclude='*.pem' /path/to/directory/sample-for-secure-medical-llm-inference-with-nitro-enclaves/ ec2-user@<PUBLIC_IP>:~ 
cd sample-for-secure-medical-llm-inference-with-nitro-enclaves
  1. Make the scripts executable (in client, server and scripts).
chmod +x *.sh

Nitro Enclave setup

With Amazon EC2 loaded with the necessary scripts, you can begin building the Nitro Enclave image. During this process, the Docker container is converted into an Enclave Image File (EIF), which generates cryptographic measurements (PCR hashes) that uniquely identify the code and configuration of the enclave. These measurements are embedded into the AWS KMS key policy, creating a hardware-attested trust boundary that makes sure only this specific, unmodified Nitro Enclave can decrypt and access the model weights.

  1. Run the complete setup script, which sets up the client on the EC2 parent instance and the server running within the Nitro Enclave. You can observe the different scripts in the client, server, and scripts folders.
sudo ./run_complete_setup.sh
  1. The various scripts run to download the MedGemma 4B model, encrypt the model with the AWS KMS key, build a Docker image to run a llama.cpp server, start a Nitro Enclave, and decrypt and run the model. This process takes approximately 10 minutes.
  2. When the Nitro Enclave is running, it runs in debug mode so that you can observe the various startup logs outputted. Wait until llama.cpp server logs are outputted that indicate the server is ready and listening.

Inference examples

When the model is decrypted and running on the llama.cpp server within the Nitro Enclave, you can begin to invoke the model with either image or text-based queries. Open a new terminal session in your EC2 instance. You can navigate to the client folder to run the scripts for queries.

Image-based queries

For inference on medical images, upload an image to your Amazon S3 image bucket. When it is uploaded, run python3 image_processor.py to pass the image from the SQS queue to the Nitro Enclave for processing. The following are examples of image inputs and model outputs.

Brain CT scan:

Medical brain CT scan image showing anatomical structures in grayscale

Case courtesy of Dr Henry Knipe, Radiopaedia.org, rID: 46289

Model response:Complete processing log showing secure medical image analysis pipeline from upload through diagnostic outputChest X-ray:

Chest X-ray image in AP sitting position

Model response:Model inference logs showing automated chest X-ray analysis with detailed radiological findings and considerationsText-based queries

For inference on text-queries, run python3 direct_query.py "<YOUR_MEDICAL_QUERY>" to invoke the model. The following are examples of text-based inputs and model outputs.

Basic usage:

python3 direct_query.py "What are the symptoms of pneumonia?"

Model response:Model inference output describing pneumonia symptoms, diagnosis, and treatment optionsLab result interpretation:

python3 direct_query.py "Patient has elevated troponin levels (15.2 ng/mL), elevated CK-MB, and ST elevation in leads II, III, aVF. What does this suggest?"

Model response:Model inference output analyzing cardiac lab results and ECG findings

Cleaning up

To avoid incurring future charges, delete the resources used in this solution:

  1. Stop and terminate the EC2 instance.
  2. Empty and delete the S3 buckets.
  3. Delete the DynamoDB table.
  4. Delete the SQS queue.
  5. Delete the AWS KMS Key.
  6. Delete the Lambda function.

Conclusion

You can combine the isolation and attestation capabilities of AWS Nitro Enclaves, the encryption controls of AWS KMS, and the scalability of services such as Amazon S3, Amazon SQS, and Amazon DynamoDB to build a more secure, zero trust pipeline for deploying generative AI models in healthcare. Using Google MedGemma 4B as your reference medical LLM, you can enable privacy-preserving inference where both PHI and model intellectual property remain protected. For more information, consult the following resources:

Meet digital sovereignty needs with AWS Dedicated Local Zones expanded services

Post Syndicated from Max Peterson original https://aws.amazon.com/blogs/security/meet-digital-sovereignty-needs-with-aws-dedicated-local-zones-expanded-services/

At Amazon Web Services (AWS), we continue to invest in and deliver digital sovereignty solutions to help customers meet their most sensitive workload requirements. To address the regulatory and digital sovereignty needs of public sector and regulated industry customers, we launched AWS Dedicated Local Zones in 2023, with the Government Technology Agency of Singapore (GovTech Singapore) as our first customer.

Today, we’re excited to announce expanded service availability for Dedicated Local Zones, giving customers more choice and control without compromise. In addition to the data residency, sovereignty, and data isolation benefits they already enjoy, the expanded service list gives customers additional options for compute, storage, backup, and recovery.

Dedicated Local Zones are AWS infrastructure fully managed by AWS, built for exclusive use by a customer or community, and placed in a customer-specified location or data center. They help customers across the public sector and regulated industries meet security and compliance requirements for sensitive data and applications through a private infrastructure solution configured to meet their needs. Dedicated Local Zones can be operated by local AWS personnel and offer the same benefits of AWS Local Zones, such as elasticity, scalability, and pay-as-you-go pricing, with added security and governance features.

Since being launched, Dedicated Local Zones have supported a core set of compute, storage, database, containers, and other services and features for local processing. We continue to innovate and expand our offerings based on what we hear from customers to help meet their unique needs.

More choice and control without compromise

The following new services and capabilities deliver greater flexibility for customers to run their most critical workloads while maintaining strict data residency and sovereignty requirements.

New generation instance types

To support complex workloads in AI and high-performance computing, customers can now use newer generation instance types, including Amazon Elastic Compute Cloud (Amazon EC2) generation 7 with accelerated computing capabilities.

AWS storage options

AWS storage options provide two storage classes including Amazon Simple Storage Service (Amazon S3) Express One Zone, which offers high-performance storage for customers’ most frequently accessed data, and Amazon S3 One Zone-Infrequent Access, which is designed for data that is accessed less frequently and is ideal for backups.

Advanced block storage capabilities are delivered through Amazon Elastic Block Store (Amazon EBS) gp3 and io1 volumes, which customers can use to store data within a specific perimeter to support critical data isolation and residency requirements. By using the latest AWS general purpose SSD volumes (gp3), customers can provision performance independently of storage capacity with an up to 20% lower price per gigabyte than existing gp2 volumes. For intensive, latency-sensitive transactional workloads, such as enterprise databases, provisioned IOPS SSD (io1) volumes provide the necessary performance and reliability.

Backup and recovery capabilities

We have added backup and recovery capabilities through Amazon EBS Local Snapshots, which provides robust support for disaster recovery, data migration, and compliance. Customers can create backups within the same geographical boundary as EBS volumes, helping meet data isolation requirements. Customers can also create AWS Identity and Access Management (IAM) policies for their accounts to enable storing snapshots within the Dedicated Local Zone. To automate the creation and retention of local snapshots, customers can use Amazon Data Lifecycle Manager (DLM).

Customers can use local Amazon Machine Images (AMIs) to create and register AMIs while maintaining underlying local EBS snapshots within Dedicated Local Zones, helping achieve adherence to data residency requirements. By creating AMIs from EC2 instances or registering AMIs using locally stored snapshots, customers maintain complete control over their data’s geographical location.

Dedicated Local Zones meet the same high AWS security standards and sovereign-by-design principles that apply to AWS Regions and Local Zones. For instance, the AWS Nitro System provides the foundation with hardware- and software-level security. This is complemented by AWS Key Management Service (AWS KMS) and AWS Certificate Manager (ACM) for encryption management, Amazon Inspector, Amazon GuardDuty, and AWS Shield to help protect workloads, and AWS CloudTrail for audit logging of user and API activity across AWS accounts.

Continued innovation with GovTech Singapore

One of GovTech Singapore’s key focuses is on the nation’s digital government transformation and enhancing the public sector’s engineering capabilities. Our collaboration with GovTech Singapore involved configuring their Dedicated Local Zones with specific services and capabilities to support their workloads and meet stringent regulatory requirements. This architecture addresses data isolation and security requirements and ensures consistency and efficiency across Singapore Government cloud environments.

With the availability of the new AWS services with Dedicated Local Zones, government agencies can simplify operations and meet their digital sovereignty requirements more effectively. For instance, agencies can use Amazon Relational Database Service (Amazon RDS) to create new databases rapidly. Amazon RDS in Dedicated Local Zones helps simplify database management by automating tasks such as provisioning, configuring, backing up, and patching. This collaboration is just one example of how AWS innovates to meet customer needs and configures Dedicated Local Zones based on specific requirements.

Chua Khi Ann, Director of GovTech Singapore’s Government Digital Products division, who oversees the Cloud Programme, shared:
“The deployment of Dedicated Local Zones by our Government on Commercial Cloud (GCC) team, in collaboration with AWS, now enables Singapore government agencies to host systems with confidential data in the cloud. By leveraging cloud-native services like advanced storage and compute, we can achieve better availability, resilience, and security of our systems, while reducing operational costs compared to on-premises infrastructure.”

Get started with Dedicated Local Zones

AWS understands that every customer has unique digital sovereignty needs, and we remain committed to offering customers the most advanced set of sovereignty controls and security features available in the cloud. Dedicated Local Zones are designed to be customizable, resilient, and scalable across different regulatory environments, so that customers can drive ongoing innovation while meeting their specific requirements.

Ready to explore how Dedicated Local Zones can support your organization’s digital sovereignty journey? Visit AWS Dedicated Local Zones to learn more.

TAGS: AWS Digital Sovereignty Pledge, Digital Sovereignty, Security Blog, Sovereign-by-design, Public Sector, Singapore, AWS Dedicated Local Zones

Max Peterson
Max Peterson

Max is the Vice President of AWS Sovereign Cloud. He leads efforts to help public sector organizations modernize their missions with the cloud while meeting necessary digital sovereignty requirements. Max previously oversaw broader digital sovereignty efforts at AWS and served as the VP of AWS Worldwide Public Sector with a focus on empowering government, education, healthcare, and nonprofit organizations to drive rapid innovation.
Stéphane Israël
Stéphane Israël

Stéphane is the Managing Director of the AWS European Sovereign Cloud and Digital Sovereignty. He is responsible for the management and operations of the AWS European Sovereign Cloud GmbH, including infrastructure, technology, and services, and leads broader worldwide digital sovereignty efforts at AWS. Prior to AWS, he was the CEO of Arianespace, where he oversaw numerous successful space missions, including the launch of the James Webb Space Telescope.

How BASF’s Agriculture Solutions drives traceability and climate action by tokenizing cotton value chains using Amazon Managed Blockchain

Post Syndicated from Kevin S. Ridolfi original https://aws.amazon.com/blogs/architecture/how-basfs-agriculture-solutions-drives-traceability-and-climate-action-by-tokenizing-cotton-value-chains-using-amazon-managed-blockchain/

BASF Agricultural Solutions combines innovative products and digital tools with practical farmer knowledge. With over a century of experience, BASF offers a broad portfolio spanning seeds, crop protection, soil management, plant health, and digital agriculture solutions. Through collaboration with farmers, scientists, and partners, BASF strives to meet societal needs sustainably while creating a lasting agricultural legacy. Infosys is a global premier consulting and managed services partner of Amazon Web Services (AWS). Through this unique partnership, AWS helps customers integrate software, services, and processes to accelerate business transformation. This post explores the commitment of this partnership to driving positive change in the agricultural industry by using Amazon Managed Blockchain to tokenize food and cotton value chains for traceability, climate action, and circularity.

Global challenges and the agricultural industry

The world’s population is growing, with the UN projecting an estimated world population of 8.5 billion in 2030 and 10.4 billion by the end of the century. Along with this growth, as well as a global increase in standards of living, comes a rising demand for agricultural products such as fiber and food crops. At the same time, as society becomes more aware of the ecological impact of agriculture, both local communities and farmers are placing larger focus on a sustainable management of natural resources. The agricultural industry is uniquely positioned at the intersection of these two trends.

The agricultural industry faces numerous complex challenges that span both business and technical domains. From a business perspective, today’s agricultural supply chains have become incredibly complex, often involving multiple intermediaries across different countries. This complexity makes it difficult to ensure fair pricing and adequate compensation for farmers, who are often at the bottom of the value chain. Furthermore, verifying sustainable farming practices and organic certifications has become increasingly challenging, even as consumer demand for product authenticity and sustainability information continues to grow. Adding to these pressures, agricultural businesses must navigate increasing regulatory requirements for environmental regulation compliance and reporting, along with complex international trade regulations and documentation.

On the technical front, the industry struggles with limited digital infrastructure in rural farming areas, where internet connectivity and technology adoption remain significant hurdles. Data collection methods vary widely across different farms and regions, making it difficult to establish consistent metrics and reporting standards. Many agricultural businesses still operate with legacy systems that resist integration with modern tracking solutions, and the lack of standardization in agricultural data formats creates additional complications. Maintaining data integrity across multiple stakeholders has proven particularly challenging, as has the implementation of real-time tracking and tracing capabilities.

Cotton and fast fashion: Industry`s challenges

Cotton is the world’s most important natural fiber crop, with a yearly production of 126.5 million bales in the 2022–2023 season, enough to produce 25.3 billion pairs of jeans or 151.8 billion T-shirts. It also plays a major role in the fast fashion industry, where garments and clothing undergo a fast production and disposal cycle to quickly address customer attention and the latest fashion trends, with around 30% of clothing sold in the US being made with cotton. This accelerated production and disposal cycle comes at the expense of considerable environmental impact, with the fast fashion industry accounting for approximately 20% of the world’s water consumption and 10% of the world’s total CO2 emissions.

The cotton industry faces its own set of distinct challenges. Water usage stands as one of the most pressing concerns, with a single cotton T-shirt requiring approximately 2,700 liters of water to produce. Chemical usage tracking presents another significant challenge, as stakeholders must carefully monitor pesticide and fertilizer application throughout the growing process. Labor practices verification has become increasingly important, with brands and consumers demanding assurance of ethical working conditions throughout the supply chain.

Quality verification poses another crucial challenge, given that maintaining accurate documentation of cotton grade and characteristics is essential for pricing and processing. The industry’s global nature creates additional complexities in cross-border logistics, requiring careful management of international shipping and customs processes. Furthermore, the growing importance of sustainability certification has created new pressures to validate organic and sustainable farming practices with reliable, transparent documentation.

As consumer expectation of guaranteed fair practices, lower carbon emissions, and sustainable use of natural resources grows, so does the demand for traceability systems that can provide near real-time visibility into each step of the value chain by tracking sustainability information such as water consumption and CO2 emissions.

The potential for a blockchain-based solution

To address this demand, BASF identified blockchain as a foundational technology for a digital solution to deliver transparency along the value chain, targeting specific customer requirements for digital assets backed by information, validation, certificates, and know your business (KYB) policies for value chain partners.

Blockchain technology emerges as a particularly powerful solution to these challenges, offering unique capabilities that directly address many of the industry’s pain points. At its core, blockchain provides immutable record-keeping, creating permanent, tamper-proof records of transactions and events that ensure data integrity throughout the supply chain. This feature proves especially valuable in preventing fraudulent modification of sustainability certificates and maintaining the credibility of organic farming claims.

Smart contracts, a key feature of blockchain technology, enable the automation of compliance with agricultural standards and facilitate automatic payment execution based on predefined conditions. This automation significantly reduces administrative overhead in supply chain management and helps ensure fair compensation for farmers.

The technology’s traceability capabilities provide end-to-end visibility of cotton from seed to garment, enabling real-time tracking of sustainability metrics and creating transparent audit trails for certification purposes. This transparency helps brands and consumers verify the authenticity and sustainability of their cotton products while enabling farmers to demonstrate their commitment to sustainable practices.

Blockchain’s decentralized data management allows multiple stakeholders to maintain shared records without requiring a central authority, eliminating single points of failure in data storage and reducing dependency on central authorities. This decentralized approach proves particularly valuable in agricultural supply chains, where numerous parties need to access and verify information.

The implementation of token economics through blockchain creates new opportunities for incentivizing sustainable farming practices. Through tokenization, farmers can access new revenue streams, including carbon credits, while establishing more direct relationships with buyers. Additionally, blockchain’s digital identity capabilities provide secure authentication for supply chain participants, enabling granular access control to sensitive data and facilitating compliance with know your customer (KYC) and KYB requirements.

Solution overview

Using a permissioned blockchain based on open-source systems, BASF Agricultural Solutions has developed a novel way to promote data democratization and address the challenges of data recording, off-chain processes, and on-chain activities at scale. The solution enables value chain players to independently verify activities progressively, and an organizational structure within chain and off-chain monitors key performance indicators (KPIs) through a DAO (Distributed Autonomous Organization) interface.

To focus on building such a system rather than managing the underlying blockchain infrastructure, BASF selected Amazon Managed Blockchain alongside additional AWS services. Amazon Managed Blockchain simplifies BASF’s approach because it brings a suite of offerings that can be configured to build this solution without the need to add more layers and external or internal sources.

As a foundational system, Amazon Managed Blockchain augments the solution’s ability to generate smart certificates along with off-chain opportunities to further expand the offering as a platform, such as with AI and AWS Lambda. This fits into BASF’s vision to deliver best-in-class solutions for the farming community and deliver trusted information to communities that want to drive a positive impact for the planet.

The following are the key structural components of the solution:

  • Peers – These blockchain nodes run smart contracts (chain code) and maintain the ledger.
  • Ordering service – The ordering service makes sure a transaction meets the consensus requirements based on configured channel and endorsement policies for the installed chain code.
  • Fabric certificate authority (CA) – This component enrolls and generates blockchain identities needed to sign transactions.
  • AWS services – The solution uses various AWS services to perform operations on the blockchain efficiently. These services include:
    • Amazon Cognito – We use Amazon Cognito to onboard external users and clients to the platform.
    • AWS Fargate – A block listener is a custom service that listens to every block event from the blockchain and updates the off-chain storage accordingly. It’s hosted as a container on Fargate. Running a container using Fargate is more straightforward than other Kubernetes services because you don’t have to manage servers or clusters of Amazon Elastic Compute Cloud (Amazon EC2) instances. With Fargate, we no longer have to provision, configure, or scale clusters of virtual machines to run containers.
    • AWS Lambda – Middleware services are hosted as Lambda functions, which makes sure the services are automatically scalable by default and cost-efficient. This is important because we’re charged based on the number of requests for the function and the time it takes for the code to run.
    • Amazon OpenSearch Service – We use OpenSearch Service as an off-chain data store because the solution requires complex queries to aggregate the ledger data. The off-chain storage is kept in sync with the ledger and is restricted for direct updates. It can be updated only by an authorized application, based on ledger events.
    • AWS Secrets Manager – We use Secrets Manager to manage blockchain identities.
    • Amazon Simple Notification Service (Amazon SNS) – We use Amazon SNS to connect various services asynchronously.

The following diagram illustrates the solution architecture.

The solution architecture is extensible and scalable to meet the dynamic load requirements. It can seamlessly connect various data sources with appropriate connectors such as Salesforce, mobility platforms, third-party services, and more.

External users such as value chain players, retailers, and others who could benefit from tokens can access the platform through different methods. Generally, access of DAOs is done through business-to-customer (B2C) login, and API streams can be subscribed by end retailers for checkouts, point of sale (POS), and so on. Additionally, we provide internal access for admins and auditors to visualize the product flows.

Conclusion

Climate challenges are quite complex and require a joint approach between technology, the custodians of our planet (namely the farmers), and public chains that deliver the right protocols. BASF Agriculture Solutions represents the farming needs and the link to the right communities and crops on the ground, AWS brings in the right infrastructure and support of the cloud and scale, and Infosys brings in development support as a partner to both AWS and BASF.

BASF is connected to millions of farmers. BASF considers farming to be the biggest job on earth. Sustainable farming means bringing back lost biodiversity and increasing carbon capture within the soil. And sustainability overall requires additional effort by the farmers. Additionally, consumers like us make choices daily when it comes to our own purchase decisions, such as to buy sustainable products or take action that brings positive impact to the climate.

The solution outlined in this post creates a solution using blockchain as the base technology to enable a secure and reliable method for information sharing across all stakeholders. It’s the baseline to onboard use cases in the agriculture industry to enable end-to-end traceability with a 360-degree view. Smart contracts incentivize farmers and other stakeholders to follow the sustainable measures based on the information in the system, which is reviewed and authorized by validators. All the actions in the system are monitored and logged as immutable records, which enforce the information trust by default. This acts as a baseline for 100% traceability, tokenization for sustainable measures, and digital assets that can be exchanged and create a positive economy around sustainability. The design discussed in this post is flexible to onboard different use cases and can auto scale to meet dynamic data volumes.

We encourage you to join BASF, Infosys, and AWS in driving sustainability through trusted value chains that incentivize farmers, empower consumers, and create a positive economy around climate action. If you want to dive deep into topics surrounding sustainability and AWS architecture, we suggest visiting the AWS Architecture Blog.


About the Authors

AWS launches AI-enhanced security innovations at re:Invent 2025

Post Syndicated from Lise Feng original https://aws.amazon.com/blogs/security/aws-launches-ai-enhanced-security-innovations-at-reinvent-2025/

At re:Invent 2025, AWS unveiled its latest AI- and automation-enabled innovations to strengthen cloud security for customers to grow their business. Organizations are likely to increase security spending from
$213 billion in 2025 to $377 billion by 2028 as they adopt generative AI. This 77% increase highlights the importance organizations place on securing their AI investments as they expand their digital footprints.

AWS uses artificial intelligence, machine learning, and automation to help you secure your environments proactively. These advancements include AI security agents, machine-learning and automation-driven threat detection, and agent-centric identity and access management. Together, they unify defense-in-depth across the application, infrastructure, network, and data layers to protect organizations from a wide spectrum of threats, vulnerabilities, and misconfigurations that could disrupt business operations.

AI security agents

AWS is embedding AI agents directly into security workflows to perform code reviews, collate incident response signals, and secure agentic access.

  • AWS Security Agent is a frontier agent that proactively secures applications throughout the development lifecycle. It conducts automated security reviews tailored to organizational requirements and delivers context-aware penetration testing on demand. By continuously validating security from design to deployment, it helps prevent vulnerabilities early in development.
  • AWS Security Incident Response delivers agentic AI-powered investigation capabilities designed to help enhance and accelerate security event response and recovery.
  • AgentCore Identity now offers authentication that provides enhanced access controls for AI agents, which restricts their interactions to authorized services and data based on specific user permissions and attributes. Enabling granular boundaries for how AI agents interact with enterprise applications reduces the risk of unauthorized access or data exposure.

ML and automation-driven threat detection

Machine learning models and automation now accelerate threat detection across more AWS environments, surfacing otherwise hard to see correlations, such as for sophisticated multistage attacks, at scale. These latest advancements save time by automatically correlating signals into consolidated sequences.

Agent-centric identity and access management

Intelligent access controls are redefining how organizations manage identities and permissions. These controls automate policy generation and improve your zero trust maturity level, making it easier for you to use AWS services.

  • IAM policy autopilot helps AI coding assistants quickly create baseline IAM policies that teams can refine as the application evolves, so organizations can build faster.
  • Outbound identity Federation helps IAM customers to securely federate their AWS identities to external services, making it easy to authenticate AWS workloads with cloud providers, SaaS platforms, and self-hosted applications.
  • Private access sign-in routes 100% of console traffic through VPC endpoints instead of public internet, using intelligent routing to maintain security without compromising performance.
  • Login for AWS local development lets developers use their existing console credentials to programmatically access AWS.

Transforming security through AI

These AI and ML advancements transform security from reactive manual processes to proactive, scalable protection. You can use them to operationalize threat hunting and advance your security posture, even as you grow your digital real estate.

The confidence organizations place in cloud-native security validates this approach. The AWS-sponsored report of 2,800 IT and security decision makers and practitioners revealed that 81% agree that their primary cloud provider’s native security and compliance capabilities exceed what their team could deliver independently. Additionally, 56% responded that the public cloud was better positioned to deliver security as opposed to 37% that selected on-premises, and 51% believe the public cloud is better positioned to meet regulations versus 41% that responded on-premises.

Cloud is the foundation on which customers build their businesses, and AWS continues to deliver security innovations that reinforce that foundation.

If you have feedback about this post, submit comments in the Comments section below.

Lise Feng

Lise Feng

Lise is a Seattle-based PR Manager focused on AWS security services and customers. Outside of work, she enjoys cooking and watching most contact sports.

Medidata’s journey to a modern lakehouse architecture on AWS

Post Syndicated from Mike Araujo original https://aws.amazon.com/blogs/big-data/medidatas-journey-to-a-modern-lakehouse-architecture-on-aws/

This post was co-authored by Mike Araujo Principal Engineer at Medidata Solutions.

The life sciences industry is transitioning from fragmented, standalone tools towards integrated, platform-based solutions. Medidata, a Dassault Systèmes company, is building a next-generation data platform that addresses the complex challenges of modern clinical research. In this post, we show you how Medidata created a unified, scalable, real-time data platform that serves thousands of clinical trials worldwide with AWS services, Apache Iceberg, and a modern lakehouse architecture.

Challenges with legacy architecture

As the Medidata clinical data repository expanded, the team recognized the shortcomings of the legacy data solution to provide quality data products to their customers across their growing portfolio of data offerings. Several data tenants began to erode. The following diagram shows Medidata’s legacy extract, transform, and load (ETL) architecture.

Built upon a series of scheduled batch jobs, the legacy system proved ill-equipped to provide a unified view of the data across the entire ecosystem. Batch jobs ran at different intervals, often requiring a sufficient degree of scheduling buffer to make sure upstream jobs completed within the expected window. As the data volume expanded, the jobs and their schedules continued to inflate, introducing a latency window between ingestion and processing for dependent consumers. Different consumers operating from various underlying data services further magnified the problem as pipelines had to be continuously built across a variety of data delivery stacks.

The expanding portfolio of pipelines began to overwhelm existing maintenance operations. With more operations, the opportunity for failure expanded and recovery efforts further complicated. Existing observability systems were inundated with operational data, and identifying the root cause of data quality issues became a multi-day endeavor. Increases in the data volume required scaling considerations across the entire data estate.

Additionally, the proliferation of data pipelines and copies of the data in different technologies and storage systems necessitated expanding access controls with enhanced security features to make sure only the correct users had access to the subset of data to which they were permitted. Making sure access control changes were correctly propagated across all systems added a further layer of complexity to consumers and producers.

Solution overview

With the advent of Clinical Data Studio (Medidata’s unified data management and analytics solution for clinical trials) and Data Connect (Medidata’s data solution for acquiring, transforming, and exchanging electronic health record (EHR) data across healthcare organizations), Medidata introduced a new world of data discovery, analysis, and integration to the life sciences industry powered by open source technologies and hosted on AWS. The following diagram illustrates the solution architecture.

Fragmented batch ETL jobs were replaced by real-time Apache Flink streaming pipelines, an open source, distributed engine for stateful processing, and powered by Amazon Elastic Kubernetes Service (Amazon EKS), a fully managed Kubernetes service. The Flink jobs write to Apache Kafka running in Amazon Managed Apache Kafka (Amazon MSK), a streaming data service that manages Kafka infrastructure and operations, before landing in Iceberg tables backed by the AWS Glue Data Catalog, a centralized metadata repository for data assets. From this collection of Iceberg tables, a central, single source of data is now accessible from a variety of consumers without additional downstream processing, alleviating the need for custom pipelines to satisfy the requirements of downstream consumers. Through these fundamental architectural changes, the team at Medidata solved the issues presented by the legacy solution.

Data availability and consistency

With the introduction of the Flink jobs and Iceberg tables, the team was able to deliver a consistent view of their data across the Medidata data experience. Pipeline latency was reduced from days to minutes, helping Medidata customers realize a 99% performance gain from the data ingestion to the data analytics layers. Due to Iceberg’s interoperability, Medidata users saw the same view of the data regardless of where they viewed that data, minimizing the need for consumer-driven custom pipelines because Iceberg could plug into existing consumers.

Maintenance and durability

Iceberg’s interoperability provided a single copy of the data to satisfy their use cases, so the Medidata team could focus its observation and maintenance efforts on a five-times smaller subset of operations than previously required. Observability was enhanced by tapping into the various metadata components and metrics exposed by Iceberg and the Data Catalog. Quality management transformed from cross-system traces and queries to a single analysis of unified pipelines, with an added benefit of point in time data queries thanks to the Iceberg snapshot feature. Data volume increases are handled with out-of-box scaling supported by the entire infrastructure stack and AWS Glue Iceberg optimization features that include compaction, snapshot retention, and orphan file deletion, which provide a set-and-forget experience for solving a number of common Iceberg frustrations, such as the small file problem, orphan file retention, and query performance.

Security

With Iceberg at the center of its solution architecture, the Medidata team no longer had to spend the time building custom access control layers with enhanced security features at each data integration point. Iceberg on AWS centralizes the authorization layer using familiar systems such as AWS Identity and Access Management (IAM), providing a single and durable control for data access. The data also stays entirely within the Medidata virtual private cloud (VPC), further reducing the opportunity for unintended disclosures.

Conclusion

In this post, we demonstrated how legacy universe of consumer-driven custom ETL pipelines can be replaced with a scalable, high-performant streaming lakehouses. By putting Iceberg on AWS at the center of data operations, you can have a single source of data for your consumers.

To learn more about Iceberg on AWS, refer to Optimizing Iceberg tables and Using Apache Iceberg on AWS.


About the authors

Mike Araujo

Mike is a Principal Engineer at Medidata Solutions, working on building a next generation data and AI platform for clinical data and trials. By using the power of open source technologies such as Apache Kafka, Apache Flink, and Apache Iceberg, Mike and his team have enabled the delivery of billions of clinical events and data transformations in near real time to downstream consumers, applications, and AI agents. His core skills focus on architecting and building big data and ETL solutions at scale as well as their integration in agentic workflows.

Sandeep Adwankar

Sandeep is a Senior Product Manager at AWS, who has driven feature launches across Amazon SageMaker, AWS Glue, and AWS Lake Formation. He has led initiatives in Amazon S3 Tables analytics, Iceberg compaction strategies, and AWS Glue Iceberg optimizations. His recent work focuses on generative AI and autonomous systems, including the AWS Glue Data Catalog model context protocol and Amazon Bedrock structured knowledge bases. Based in the California Bay Area, he works with customers around the globe to translate business and technical requirements into products that accelerate their business outcomes.

Ian Beatty

Ian is a Technical Account Manager at AWS, where he specializes in supporting independent software vendor (ISV) customers in the healthcare and life sciences (HCLS) and financial services industry (FSI) sectors. Based in the Rochester, NY area, Ian helps ISV customers navigate their cloud journey by maintaining resilient and optimized workloads on AWS. With over a decade of experience building on AWS since 2014, he brings deep technical expertise from his previous roles as an AWS Architect and DevSecOps team lead for SaaS ISVs before joining AWS more than 3 years ago.

Ashley Chen

Ashley is a Solutions Architect at AWS based in Washington D.C. She supports independent software vendor (ISV) customers in the healthcare and life sciences industries, focusing on customer enablement, generative AI applications, and container workloads.

How Octus achieved 85% infrastructure cost reduction with zero downtime migration to Amazon OpenSearch Service

Post Syndicated from Vaibhav Sabharwal original https://aws.amazon.com/blogs/big-data/how-octus-achieved-85-infrastructure-cost-reduction-with-zero-downtime-migration-to-amazon-opensearch-service/

As data volumes continue to grow exponentially, there is increasing pressure to optimize search infrastructure costs while maintaining the high performance and reliability that mission-critical workloads demand. Many companies find themselves managing complex, expensive search systems that require significant operational overhead and limit their ability to scale efficiently. The challenge becomes even more acute when organizations need to migrate between search systems, a process that traditionally involves substantial downtime, complex data synchronization, and significant impact on business operations. Enterprise applications cannot afford service interruptions that could impact customer experiences, business intelligence, or operational continuity. Migration strategies need to deliver cost optimization and operational improvements while maintaining zero downtime and facilitating complete data integrity throughout the transition process.

Founded in 2013, Octus, formerly Reorg, is the essential credit intelligence and data provider for the world’s leading buy side firms, investment banks, law firms and advisory firms. By surrounding unparalleled human expertise with proven technology, data and AI tools, Octus unlocks powerful truths that fuel decisive action across financial industries.

This post highlights how Octus migrated its Elasticsearch workloads running on Elastic Cloud to Amazon OpenSearch Service. The journey traces Octus’s shift from managing multiple systems to adopting a cost-efficient solution powered by OpenSearch Service. Along the way, we share the architecture choices and implementation strategies that made the migration successful. The result is uninterrupted service availability throughout migration, with improved performance and greater cost efficiency.

Strategic requirements

We identified several requirements that made Amazon OpenSearch Service the right choice for their migration:

  • Cost efficiency: The OpenSearch Service pricing model enabled us to optimize cloud spend without compromising performance.
  • Responsive support: AWS provided dependable, high-quality support to accelerate issue resolution and instill confidence.
  • Consistent reliability: OpenSearch Service provides an SLA up to 99.99% offering the reliability required for Octus’s mission-critical workloads.
  • Seamless migration with no query downtime: Migration Assistant for Amazon OpenSearch Service provided Octus with a migration path while maintaining uninterrupted query availability during the migration, facilitating business continuity.
  • Operational simplification: Consolidating onto AWS reduced infrastructure complexity while maintaining high security standards.

Solution overview

The Migration Assistant for Amazon OpenSearch Service provides a suite of tools to aid in Elasticsearch to OpenSearch Service migrations. Octus use the following capabilities for their migration:

  • Metadata migration: The tool enabled Octus to migrate dozens of indices with diverse mappings and settings. When a backward incompatibility was identified with timestamp metadata, a custom JavaScript transformation, integrated directly into the Migration Assistant tooling, was applied to automatically adjust the mappings across the indices and facilitate compatibility.
  • Historical data migration: Octus used Reindex-from-Snapshot to migrate the historical documents from a point-in-time snapshot of the source cluster, scaling this process without impacting the source cluster since the snapshot was stored in Amazon Simple Storage Service (Amazon S3). Reindex-from-Snapshot also enabled Octus to adjust the sharding scheme during migration, helping to optimize cluster performance on the target.
  • Live Traffic Replay: Once backfill was complete, Octus used Migration Assistant’s Traffic Replayer to send the captured live traffic (from the Traffic Capture Proxy) to the target cluster with required request transformations for OpenSearch Service compatibility, resulting in the target cluster containing the documents from the source cluster with updates being performed in real time.

The following diagram illustrates the implementation architecture diagram for this migration.


Figure 1 – Migration Assistant architecture with migration steps

For more information about the Migration Assistant for Amazon OpenSearch Service, visit the AWS Solutions home page.

Each node in the diagram correlates to the following steps in the migration process:

  1. Client traffic is directed to the existing cluster.
  2. An Application Load Balancer with capture proxies relays traffic to a source while replicating data to Amazon Managed Streaming for Apache Kafka (Amazon MSK).
  3. Using the migration console, a point-in-time snapshot is taken. Once the snapshot completes, the Metadata Migration Tool is used to establish indexes, templates, component templates, and aliases on the target cluster. With continuous traffic capture in place, Reindex-from-Snapshot, migrates data from the source.
  4. Once Reindex-from-Snapshot is complete, captured traffic is replayed from Amazon Managed Streaming for Apache Kafka (Amazon MSK) to the target cluster by Traffic Replayer.
  5. Performance and behavior of traffic sent to the source and target clusters are compared by reviewing logs and metrics.
  6. After confirming that the target cluster’s functionality meets expectations, clients are redirected to the new target.

Complete migration and optimization journey

Octus’s migration from Elastic Cloud to Amazon OpenSearch Service encompassed both the core migration effort and subsequent optimization phases. The goal was to successfully migrate the search infrastructure, applications, and data from Elastic Cloud to a new OpenSearch Service domain with minimal disruption, while continuously optimizing performance and costs based on real-world usage data.

Octus used their in-house custom infrastructure frameworks (their internal tooling for infrastructure automation) to build, deploy and monitor the target OpenSearch Service 1.3 domain, establishing a solid foundation for the migration. This approach used familiar internal processes while moving to the fully managed AWS service. Refer to AWS documentation to implement security best practices when using OpenSearch Service.

Pre-migration optimization

Prior to initiating the migration, Octus conducted optimization activities on the source Elasticsearch cluster to streamline the migration process. This included removing unused indexes that had accumulated over time and removing large documents that would unnecessarily extend migration duration and increase storage transfer costs. These preparatory steps significantly reduced the data volume requiring migration and minimized the overall migration complexity, enabling more efficient use of the Migration Assistant tools.

Technical constraints and version considerations

The migration involved specific version compatibility challenges that influenced the technical approach. The source Elasticsearch cluster was running version 7.17, and the Python client applications were also constrained to Elasticsearch 7.17 compatibility. To support the transition, the team used Reindex-from-Snapshot, which enables cross-system migrations by reindexing data from existing snapshots into a new OpenSearch Service cluster. RFS also rewrites indices created on older versions of Lucene, simplifying future upgrades to the latest version of OpenSearch Service. While evaluating a move to OpenSearch 1 or 2, Octus selected OpenSearch 1.3 as the target to minimize client-side changes and reduce migration complexity, while positioning themselves for simpler upgrades later.

The version selection particularly impacted the R application environment, as R language (an open-source programming language for statistical computing and data analysis) lacked native OpenSearch 1.3 client support. This constraint required Octus to develop a custom client solution using the ropensci/elastic library to integrate with the new OpenSearch Service domain. The Python environment presented similar challenges, where the Elasticsearch 7.17 client constraints necessitated careful consideration of the migration approach. These client compatibility concerns were among the factors that influenced the choice of Migration Assistant tools over traditional snapshot-based methods, as the Migration Assistant provided better support for managing version-specific client interactions during the transition.

Looking forward, Octus plans to upgrade to newer OpenSearch versions as their application stack evolves and client library support matures, so that they can leverage the latest features and performance improvements while maintaining the stability achieved through this migration.

Application modernization across multiple languages

The application changes represented a significant technical undertaking across multiple programming environments:

  • Legacy PHP systems (5.6 and Laravel 4.2): Octus handled mapping type deprecation on OpenSearch requests as specifying these mapping types are not supported, while continuing to use the elasticsearch connector library with username/password authentication.
  • Modern PHP applications (8.1 and Laravel 9): These underwent more comprehensive changes, replacing the elasticsearch/elasticsearch library with the opensearch-project/opensearch-php client and leveraging IAM authentication to connect to the clusters.
  • Python environment: Applications spanning versions 3.8, 3.10, 3.11, and 3.13 with Django frameworks 2.1, 3.2, and 5.2 required replacing the elasticsearch library with opensearch-py and transitioning to IAM authentication.
  • R applications: For R 4.5.1 applications, Octus utilized a custom library ropensci/elastic to facilitate compatibility.

Traffic routing and enhanced monitoring

To facilitate the migration, Octus redirected their existing clients to route requests to the source cluster through Migration Assistant’s Traffic Capture Proxy, migrating the data from live traffic to their target cluster.

The monitoring infrastructure underwent significant enhancement during this process. Octus’s observability infrastructure monitors the overall health of OpenSearch Service clusters which includes cluster manager and data nodes, network, data storage, security and IAM access. It also monitors the indexing and search performance of their applications. This alleviated the need for a separate monitoring cluster as logs and metrics were shipped directly to Datadog, significantly improving observability. The Datadog monitors were defined using Infrastructure-as-Code and integrated seamlessly into their infrastructure frameworks.

Cutover and initial results

The Site Reliability Engineering team meticulously planned the release, achieving a successful migration from Elasticsearch to OpenSearch Service and cutover of the Elasticsearch client to the OpenSearch Service clients with no downtime for the system application and zero data loss. The initial migration phase resulted in a 52% cost reduction while achieving operational benefits including zero downtime for the system app, no data loss, full Infrastructure-as-Code implementation for infrastructure and monitoring, and enhanced observability.

Post-migration optimization

Following the migration, Octus conducted comprehensive optimization based on operational data from production and other environments in the new OpenSearch Service setup. This real-world usage data provided valuable insights into actual resource consumption, enabling informed decisions regarding further cluster resizing.

Through usage metric analysis and strategic resizing, Octus aligned cluster size more precisely with operational needs, facilitating continued performance while minimizing expenditure. This optimization phase delivered an additional 33% cost reduction compared to the original Elastic Cloud costs, bringing the total reduction to 85% while maintaining consistent and optimal performance.

Operational monitoring

Octus uses Datadog to monitor both search and indexing latency providing real-time visibility into Amazon OpenSearch Service cluster performance. The following screenshot showcases how custom Datadog dashboards provide a live view of the OpenSearch Service clusters. This visualization offers both a high-level overview and detailed insights into the ingestion process, helping us understand the storage and document count. The bottom half of the dashboard presents a time-series view of individual node health and performance metrics like read and write latency, throughput and IOPS.


Figure 2 – DataDog dashboards

Migration observability

Migration Assistant for Amazon OpenSearch Service provides several dashboards to observe and validate the progress of a migration. By using these observability features customers can track both backfill and live capture and replay progress, facilitating confidence before switching production workloads to the target cluster.The following graphs are an example from Octus’s migration, where approximately 4TB of data was migrated in about 9 hours (from 08:00 to 17:00).


Figure 3 – Backfill progress by disk usage


Figure 4 – Backfill progress by searchable documents

Once the backfill is complete, the captured traffic is replayed to synchronize ongoing activity between the source and target clusters.

At the time the backfill finished (around 17:00), the target cluster was approximately 467 minutes behind the source. The replay process rapidly reduced this lag by processing captured traffic at a faster rate than it was originally ingested at the source.


Figure 5 – Replay lag after backfill completion

When the lag time reached 0, the target cluster was fully in sync and production traffic could safely be rerouted. Octus chose to observe replayed traffic on the target for several days before making the final switchover.

Achieving excellence

Octus’s migration to Amazon OpenSearch Service has yielded remarkable results:

  • Scalability – Octus has almost doubled the number of documents available for Q&A across three environments in days instead of weeks. Their use of Amazon Elastic Container Service (Amazon ECS) with AWS Fargate with auto scaling rules and controls gives them elastic scalability for their services during peak usage hours.
  • Cost reduction – By moving away from Elastic Cloud to OpenSearch Service, Octus’s monthly infrastructure costs are now 85% lower.
  • Enhanced search performance – Octus maintained consistent response times throughout the migration with no negative impact on latency, while achieving a 20% improvement in query throughput and overall search performance.
  • Zero downtime – Octus experienced zero downtime during migration and 100% uptime overall for the whole application.
  • Reduced operational overhead – Post-migration, Octus’s DevOps and SRE teams see 30% less maintenance burden and overheads. Supporting SOC2 compliance is also straightforward now that they’re using one system.
  • Accelerated timeline delivery – The entire migration was completed ahead of schedule, moving from planning to full completion in under one quarter.

“Moving from Elastic Cloud to Amazon OpenSearch Service was a key component of our broader strategy to minimize third-party dependencies and strengthen the reliability of Octus’ system infrastructure. Migration Assistant for Amazon OpenSearch Service enabled us to execute a seamless transition with zero data loss and virtually no downtime for our users.” – Vishal Saxena, CTO, Octus

Conclusion

In this post, we showed you how Octus successfully migrated their Elasticsearch workloads from Elastic Cloud to Amazon OpenSearch Service using the Migration Assistant for OpenSearch Service, achieving zero downtime and significant operational improvements.

The Migration Assistant for OpenSearch Service supported this complex migration through its comprehensive suite of tools. The Metadata Migration capability migrated dozens of indices with diverse mappings and settings, with custom JavaScript transformations handling backward incompatibilities. Reindex-from-Snapshot migrated the historical documents from point-in-time snapshots without impacting the source cluster, while also optimizing the sharding scheme for improved performance. Live Traffic Replay made sure the target cluster remained synchronized with real-time updates throughout the migration process.

The migration delivered substantial results across the dimensions. Octus achieved an 85% reduction in monthly infrastructure costs while nearly doubling the number of documents available for search across three environments. Search performance improved by 20% in query throughput with consistent response times and no negative impact on latency. The migration maintained zero downtime and 100% uptime for the entire application, with DevOps and SRE teams experiencing 30% less maintenance burden and operational overhead. The entire migration was completed ahead of schedule in under one quarter.

To learn more about the Migration Assistant for OpenSearch Service and how it can help you achieve similar results, visit the AWS Solutions home page.

Visit Octus to learn how we deliver rigorously verified intelligence at speed and create a complete picture for professionals across the entire credit lifecycle. Follow Octus on LinkedIn and X.


About the Authors

Harmandeep Sethi

Harmandeep Sethi

Harmandeep is Head of SRE Engineering and Infrastructure Frameworks at Octus. with nearly 10 years of experience leading high-performing teams in the implementation of large-scale systems. He has played a pivotal role in transforming and modernizing Octus’s Search Engine infrastructure and services by driving best practices in observability, resilience engineering, and the automation of operational processes through Infrastructure Frameworks.

Serhii Shevchenko

Serhii Shevchenko

Serhii is a Site Reliability Engineer at Octus. With 9 years of combined experience in software development and site reliability engineering, his expertise focuses on enhancing system reliability and performance. He was a key developer on the application side for the company’s critical migration from Elasticsearch Cloud to AWS OpenSearch. His planning was instrumental in executing the transition with zero client-facing downtime.

Govind Bajaj

Govind Bajaj

Govind is a Senior Site Reliability Engineer at Octus, specializing in architecting and implementing scalable infrastructure that supports high-performing engineering teams and critical systems. With over 8 years of experience, he excels at breaking down complex problems and turning them into practical, well-designed solutions, with a strong focus on building secure, observable, and resilient platforms.

Virendra Shinde

Virendra Shinde

Virendra is the Head of Platform at Octus, where he oversees cloud infrastructure, site reliability, and the core frameworks that power the Octus product suite. Before joining Octus, he spent two years at Grayscale Investments building an investor portal and data APIs from the ground up. Prior to that, he spent eight years at Blackstone leading multiple development teams. He holds a Master’s degree in Information Management from the University of Maryland.

Brian Presley

Brian Presley

Brian is a Software Development Manager at OpenSearch, leading teams behind OpenSearch Migrations and OpenSearch Serverless to build scalable, high-impact search and analytics solutions.

Andre Kurait

Andre Kurait

Andre is a Software Development Engineer II at AWS, based in Austin, Texas. He is currently working on Migration Assistant for Amazon OpenSearch Service. Prior to joining Amazon OpenSearch, Andre worked within Amazon Health Services. In his free time, Andre enjoys traveling, cooking, and playing in his church sport leagues. Andre holds Bachelor of the Science degrees from the University of Kansas in Computer Science and Mathematics.

Vaibhav Sabharwal

Vaibhav Sabharwal

Vaibhav is a Senior Solutions Architect at AWS based out of New York. He is passionate about learning new cloud technologies and assisting customers in building cloud adoption strategies, designing innovative solutions, and driving operational excellence. As a member of the Financial Services and Storage Technical Field Communities at AWS, he actively contributes to the collaborative efforts within the industry.

Announcing CloudFormation IDE Experience: End-to-End Development in Your IDE

Post Syndicated from Damola Oluyemo original https://aws.amazon.com/blogs/devops/announcing-cloudformation-ide-experience-end-to-end-development-in-your-ide/

If you’ve developed AWS CloudFormation templates, you know the drill; write YAML(YAML Ain’t Markup Language) in your IDE(Integrated Development Environment), switch to the AWS Management Console to validate, jump to documentation to verify property names. Then run CFN Lint(Cloudformation Linter) in your terminal, deploy and wait, then troubleshoot failures back in the console. This constant context switching between your IDE, AWS Console, documentation pages, and validation tools fragments your workflow and kills productivity. What should take 30 minutes often stretches into hours of iteration cycles.

Today, we’re excited to introduce the CloudFormation IDE Experience, a comprehensive solution that brings the entire CloudFormation development lifecycle into your IDE. No more context switching. No more fragmented workflows. Just one unified, intelligent development experience from authoring to deployment.

In this post, you’ll learn how the Cloudformation IDE Experience transforms your workflow with intelligent authoring, real-time validation, AWS integration, and more.

What is the CloudFormation IDE Experience?

The CloudFormation IDE Experience reimagines how you build infrastructure as code by creating an end-to-end development loop entirely within your IDE. Unlike generic YAML or JSON editors, this is a CloudFormation-first solution built specifically for infrastructure developers.

This solution covers the complete lifecycle; from intelligent authoring with smart code completion and navigation that understands CloudFormation semantics, to real-time multi-layer validation that catches issues before deployment. It provides direct AWS integration for seamless resource imports and stack visibility, monitors configuration drift between your templates and deployed resources, and includes server-side pre-deployment checks that prevent common deployment failures.
The result? A development environment that understands your infrastructure code as deeply as your IDE understands your application code.

Core Features

Quick Project Setup with CFN Init

CFN Init streamlines project setup by creating a structured CloudFormation project with environment configurations in seconds. Run “CFN Init: Initialize Project” from the Command Palette, configure your environments (dev, staging, production), and associate each with an AWS profile.

The CloudFormation Explorer displays your environments, letting you switch between them with a single click. Each environment maintains its own deployment settings and parameter values, eliminating manual configuration and ensuring consistent deployments across your infrastructure lifecycle.

Intelligent Authoring with Intelligent Code Completion

The IDE understands CloudFormation semantics and provides context-aware suggestions as you type. Only required properties appear automatically, while optional properties surface on hover, so when you add a Properties section to an EC2 VPC resource, nothing appears because it has no required properties. Create a subnet, however, and VpcId appears immediately because it’s required.

When you use !GetAtt or !Ref, the IDE knows exactly which attributes and resources are available. Navigation features like go-to-definition for logical IDs and hover tooltips let you explore complex templates without losing context. The IDE also provides full support for CloudFormation intrinsic functions and pseudo parameters.

Multi-Layer Validation System

The IDE provides comprehensive validation at multiple levels:

Static Validation (Real-time)

  • CloudFormation Guard Integration: Security and compliance checks using AWS Security pillar rules. For example, it automatically flags insecure configurations like MapPublicIpOnLaunch: true on subnets
  • CFN Lint Integration: Advanced syntax and logic validation, including overlapping CIDR block detection, resource dependency validation, and property checks beyond basic schema validation

Interactive Error Resolution
When errors occur, the IDE doesn’t just highlight them, it helps you fix them. Contextual error messages explain what’s wrong and why it matters, while one-click quick fixes automatically correct common issues like missing required properties or invalid reference formats. If you reference a non-existent resource, the IDE suggests valid alternatives from your template. Reference an invalid attribute with !GetAtt, the IDE immediately shows which attributes are actually available for that resource type.

AWS Resource Integration (CCAPI)

Import existing AWS resources directly into your templates using the Cloud Control API (CCAPI). Browse live resources and view all CloudFormation stacks in your AWS account from within the IDE. Pull resource configurations directly into your template with one click, complete with accurate property values. This transforms existing infrastructure into Infrastructure-as-Code without manual reconstruction or switching to the console to look up property values.

Server-Side Validation

Before you deploy, the IDE performs comprehensive server-side validation through AWS’s intelligent validation service that analyzes your CloudFormation templates against real-world deployment patterns and catches issues static analysis can’t detect.

The AWS’s intelligent validation service uses AWS-managed hooks to analyze your change sets before execution across three categories. Enhanced template validation covers CFN Lint blind spots like transforms and parameter values. Primary identifier conflict detection finds existing resources with the same identifiers before you attempt deployment. Resource state validation checks resource readiness ensuring, for example, that Amazon Simple Storage Service(S3) buckets are empty before deletion attempts.

This validation is based on analysis of the top CloudFormation failure patterns, helping you catch issues before they cause rollbacks or failed states.

Getting Started

Getting started with the CloudFormation IDE Experience is straightforward:

Prerequisite:

  1. Install an IDE that supports the CloudFormation extension, such as Visual Studio Code, Kiro
  2. Download the CloudFormation extension for your platform (available through the AWS Toolkit)
  3. Install the extension following the standard VS Code extension installation process

No complex dependency management or schema updates required—all configuration and updates are handled automatically.

Let’s See How It Works

Let’s walk through a practical example that demonstrates the IDE experience in action. We’ll build a simple Amazon Virtual Private Cloud (Amazon VPC) infrastructure with subnets and an S3 bucket.

Setting Up Your Project

Start by initializing a new CloudFormation project. Open the Command Palette, run “CFN Init: Initialize Project”, choose your project location, and set up environments. For this example, create a “beta” environment and associate it with your AWS development profile. The IDE creates your project structure with configuration files ready to use. You can now select your “beta” environment from the CloudFormation Explorer to ensure all deployments use the correct settings.

Figure 1: Initializing a CloudFormation project with environment configuration

Starting with Intelligent Authoring

Create a new CloudFormation template and start typing AWS::EC2::VPC. The IDE provides intelligent completions as you type.

Cloudformation IDE extension intelligent completion

Figure 2.0: Resource type auto-completion with CloudFormation-aware IntelliSense

When you add the Properties section, notice something interesting: nothing appears automatically. That’s because Amazon Elastic Compute Cloud (Amazon EC2) VPC has no required properties.

Cloudformation IDE extension doesn't suggest optional properties
Figure 2.1: No automatic suggestions for VPC properties since none are required

Hover over Properties to see all available options with their types and documentation links.

Hover information displaying optional properties and their documentation

Figure 2.2: Hover information displaying optional properties and their documentation

Add a CIDR block, then create a subnet. This time, when you type Properties, VpcId appears immediately because it’s required.

Required properties VpcID automatically suggested for EC2 Subnet
Figure 2.3: Required properties VpcID automatically suggested for EC2 Subnet

The IDE provides the resource names in your template, and when you use !GetAtt or !Ref, it knows which attributes are available for each resource type.

Type-aware completions for intrinsic functions like !GetAtt & !Ref

Figure 2.4: Type-aware completions for intrinsic functions like !GetAtt & !Ref

Real-Time Validation in Action

As you continue building, add MapPublicIpOnLaunch: true to make a public subnet. Immediately, a blue squiggly line appears.

CloudFormation Guard warning highlighted in real-time

Figure 3: CloudFormation Guard warning highlighted in real-time

Hovering reveals a CloudFormation Guard warning from the AWS Security pillar rules: this configuration isn’t recommended for security compliance.

Security compliance warning with detailed explanation

Figure 3.1: Security compliance warning with detailed explanation

Create a second subnet by copying the first, but now red squiggly lines appear. CFN Lint has detected overlapping CIDR blocks between your two subnets – an issue that would fail during deployment. You can fix it immediately with the contextual information provided.

CFN Lint error detection for overlapping CIDR blocks providing detailed error information helping you resolve the issue quickly
Figure 3.2: CFN Lint error detection for overlapping CIDR blocks providing detailed error information helping you resolve the issue quickly

Importing Existing Resources

Now you need an S3 bucket. Instead of writing it from scratch, open the Resource Explorer panel on the left. Using CCAPI integration, you can see all your existing AWS resources. Select an S3 bucket and click “Import resource state”. The IDE pulls in the complete resource configuration with all properties already set. You can now iterate on this resource without needing to remember or look up all the configuration details.

Automatically imported resource configuration from live AWS resources

Figure 4: Automatically imported resource configuration from live AWS resources

Developer Experience Benefits

The CloudFormation IDE Experience delivers measurable improvements across productivity and quality:

Productivity Gains:

  • Reduced context switching: Keep your entire workflow in one place
  • Faster iteration cycles: Catch and fix issues in seconds, not minutes or hours
  • Shift-left validation: Identify problems before deployment, not after
  • Intelligent assistance: Spend less time in documentation, more time building

Quality Improvements:

  • Proactive error prevention: Multi-layer validation catches issues early
  • Security by default: Built-in compliance checks from CloudFormation Guard
  • Best practice enforcement: Automated guidance aligned with AWS recommendations
  • Deployment confidence: Pre-deployment validation reduces rollback scenarios

What previously took hours of troubleshooting and multiple deployment attempts now becomes a confident 30-minute development cycle.

“I will definitely use these features; they help to reduce the feedback loop and speed up the development of IaC templates.” – AWS Community Builder

Things to Know

Platform Support

The CloudFormation IDE Experience is available for:

  • Visual Studio Code: Full feature support
  • Kiro: Full feature support
  • Cursor: Full feature support
  • JetBrains IDEs: Complete integration across the IntelliJ family (Fast Follow)
  • Operating Systems: macOS (ARM), Linux (x64) and Windows(…)

Conclusion

The CloudFormation IDE Experience eliminates the context switching that fragments your workflow. Write, validate, and deploy all from one environment. What used to take hours of iteration now takes minutes.

Ready to get started? Install the CloudFormation extension from the AWS Toolkit for VS Code and experience the difference. For detailed setup instructions and feature documentation, see the CloudFormation IDE Experience guide.

About the Authors:

Damola Oluyemo

Damola Oluyemo is a Solutions Architect at Amazon Web Services focused on Enterprise customers. He helps customers design cloud solutions while exploring the potential of Infrastructure as Code and generative AI in software development.

Jehu Gray

Jehu Gray is a Prototyping Architect at Amazon Web Services where he helps customers design solutions that fits their needs. He enjoys exploring what’s possible with IaC.