All posts by Deepti Tirumala

How MHK built a HIPAA-eligible agentic AI solution on Amazon Bedrock

Post Syndicated from Deepti Tirumala original https://aws.amazon.com/blogs/architecture/how-mhk-built-a-hipaa-eligible-agentic-ai-solution-on-amazon-bedrock/

Healthcare organizations face an increasingly complex challenge: processing vast volumes of medical documents, including clinical records, claims, prior authorizations, appeals, and pharmacy data, while maintaining strict HIPAA compliance and security standards. Traditional approaches require dedicated engineering teams to build individual AI systems for each use case, each needing its own compliance infrastructure, audit trails, and security controls. Agentic frameworks that can scale across use cases can reduce lengthy development cycles and high operational overhead.

MHK, a Hearst Health company ranked #1 in payer care management solutions in the 2024 Best in KLAS: Software & Services Report, faced this exact challenge. Their medical management solution serves health plans across multiple workflows (medical, pharmacy, grievance, and appeals), each requiring intelligent document processing and decision support. Building separate AI systems for each workflow was unsustainable as demand grew.

To solve this, they developed the SmartProminence AI Orchestrator, a HIPAA-eligible agentic workflow framework built on AWS that reduced manual medical review effort by 90%. New AI features that previously took 3+ months to deploy now ship in 2 weeks.

In this post, we walk through how MHK architected this solution using Amazon Bedrock, Amazon Elastic Container Service (Amazon ECS), and event-driven patterns to create a reusable, multi-tenant orchestrator for healthcare AI.

MHK uses Amazon Bedrock exclusively for foundation model inference. They built their own orchestration, retrieval, and validation layers because healthcare workflows require domain-specific controls: DAG-based multi-step execution, clinical document retrieval tied to case context, and HIPAA-specific validation logic that goes beyond general-purpose guardrails. This approach keeps Bedrock focused on scalable model access while MHK retains full control over workflow behavior and compliance enforcement.

Background

MHK provides healthcare cost management and compliance solutions to health plans across the United States. Their medical management system supports the full lifecycle of care decisions, from the moment a provider submits a request for service through final resolution. This includes prior authorization, claims adjudication, appeals processing, pharmacy benefit verification, and medical director reviews.

Each workflow involves analyzing unstructured medical documents such as clinical notes, lab results, imaging reports, and multi-page faxed records, against structured policy criteria. Before MHK’s SmartProminence AI Orchestrator, case managers spent 5 to 10 minutes manually processing each incoming document, while medical directors spent longer reviewing complex cases that required policy adherence determinations.

MHK needed to automate this research while maintaining healthcare’s audit trail and compliance requirements, and to do so across all product modules without building separate AI infrastructure for each one.

Business challenge

As MHK evaluated how to bring AI capabilities across their entire product suite, three core challenges emerged.

  • Fragmented AI infrastructure. Each AI-powered feature would require its own deployment pipeline, HIPAA compliance certification, security controls, and monitoring. Every new AI roadmap item meant a new cluster, a new compliance engagement, and a new operational burden. For a company serving multiple health plans across multiple modules, this approach could not scale.
  • Lengthy development cycles. Deploying a new AI workflow through traditional engineering took 3 to 6 months, not including requirements gathering. The engineering team could not keep pace with the product roadmap.
  • Manual effort at premium cost. Case managers, nurses, pharmacists, and medical directors spent hours per case manually searching through patient records. The cost was especially acute for medical directors (physicians) and pharmacists, whose hourly rates make even small-time savings translate into significant ROI.

Solution overview: SmartProminence AI Orchestrator

MHK built the SmartProminence AI Orchestrator, a multi-tenant, agentic workflow framework running entirely on AWS. Rather than building separate AI systems for each use case, MHK created a single orchestrator where various AI workflows can be deployed through configuration. Define your prompts, specify your input/output schemas, and register the agent. The solution handles everything else including HIPAA compliance, encryption, audit trails, scaling, and orchestration.

The solution is architected around a controller-agent pattern in which a Workflow Engine Controller resolves workflow dependencies and dispatches individual steps to LLM processing agents. The entire system is stateless, event-driven, and independently scalable.

The following diagram illustrates the high-level architecture of the solution.

Architecture of the MHK SmartProminence AI Orchestrator on AWS, showing the orchestration core, workflow controllers, and processing agents communicating through Amazon SQS queues

Figure 1: MHK SmartProminence AI Orchestrator architecture on AWS

At the core of the architecture, the Agent Orchestration Core serves as the central nervous system. Built on Spring Boot and running on AWS Fargate, it exposes a REST API that handles job submission, workflow management, LLM proxying, and token management. Critically, it is the only component that directly accesses the database: controllers and agents interact exclusively through the orchestration core’s API, enforcing strict data access boundaries.

Architecture overview

This section examines the key architectural patterns the orchestrator uses to process diverse healthcare workflows at scale.

Controller-agent pattern with Amazon Bedrock

MHK selected Amazon Bedrock for its multi-model access through a single API, letting them choose the best model per workflow step without separate integrations. As a managed AWS service, Bedrock inherits existing AWS Identity and Access Management (IAM), Amazon Virtual Private Cloud (Amazon VPC), and encryption controls, avoiding a new trust boundary. Built-in content filtering and invocation logging satisfy healthcare auditability requirements, and its model-agnostic architecture lets MHK adopt newer models without rearchitecting the solution.

The orchestrator enforces a strict separation between workflow orchestration and LLM processing. The Workflow Engine Controller determines what needs to happen and in what order, while LLM processing agents execute individual steps. This separation lets agent processing scale independently from workflow logic, and it makes the workflow the single source of truth while agents operate only on specific, actionable steps.

When a job arrives, the controller loads the version-pinned workflow definition, resolves step dependencies into a DAG using Kahn’s algorithm, and pre-creates step executions in a WAITING state. It uses conditional Spring Expression Language (SpEL) expressions to decide which steps to run versus skip, then dispatches agents layer by layer. Steps at the same depth run in parallel, and the controller polls for completion before advancing to the next depth.

Each agent runs a standardized pipeline: input binding (resolving expressions to gather prior step results), optional vision processing for scanned documents, prompt assembly with enriched context, LLM invocation to Amazon Bedrock (Claude), and post-processing for field extraction, type coercion, and structured output.

For parallel workloads within a single step, agents use Java virtual threads for each execution. This lets them process multiple items concurrently, such as extracting data from each page of a multi-page document simultaneously.

Event-driven orchestration with Amazon SQS

Communication between the orchestration core, controllers, and agents flows through Amazon Simple Queue Service (Amazon SQS) queues. To trigger a workflow, the orchestration core places a ControllerTaskMessage on the Controller Invoke Queue. To dispatch an individual step, it places an AgentTaskMessage on the Agent Invoke Queue. Each message is secured with a capability token scoped to only that operation’s data.

This design delivers four properties. Stateless processing means available instances can pick up pending messages. Independent scaling lets agents scale horizontally through ECS Fargate. Fault isolation keeps a failed task from blocking parallel steps, and dead letter queues capture failures. Decoupled deployment ships new agent versions without system-wide restarts.

The only blocking call in the pipeline is the LLM invocation to Amazon Bedrock. Everything else is asynchronous and event-driven, so the system can process hundreds of concurrent jobs without resource contention.

DAG-based parallel execution

The workflow engine uses depth-based parallel execution to maximize throughput. Consider a medical policy review workflow: at Depth 0, agents simultaneously extract patient demographics and pull claims history. At Depth 1, once both are complete, a policy lookup agent identifies the relevant criteria. At Depth 2, an evidence-gathering agent searches through the patient’s clinical history for documentation that satisfies each policy criterion. The controller only advances to the next depth when steps at the current depth have completed.

Conditional expressions can dynamically skip steps based on upstream results. For example, if the initial classification step determines that a case does not involve prescription drugs, the pharmacy verification step at the next depth is automatically skipped, saving both time and token costs. This conditional logic is evaluated by the controller using Spring Expression Language (SpEL) against the structured outputs of completed steps.

Dynamic agent registry

When MHK needs a new agent type, whether for a new medical management module or a new kind of analysis, the process is configuration-driven rather than engineering-driven.

A developer defines the agent configuration (prompt templates, input/output schemas, and model selection), then registers it through the orchestration core’s API. Terraform automatically provisions the supporting infrastructure: SQS queues, IAM roles, and ECS task definitions. The agent immediately becomes available for workflow step assignments, with no new compliance certification needed, since it runs within the already certified orchestrator.

This transformed MHK’s development velocity. The engineering team focuses on prompt design and workflow logic rather than infrastructure scaffolding.

Conversational memory and case association

The orchestrator maintains context across multiple workflow executions for the same patient case. Each execution returns a job ID that the upstream system associates with the case record. Over a case’s lifetime there may be three or more executions (initial intake, policy review, and appeal processing), each producing structured outputs that stay available for later executions.

Once a 30-page clinical record has been analyzed, its structured output is available for future queries on that case without re-running ingestion. When a medical director reviews an appeal weeks later, the patient’s history is already organized and searchable.

Prior context remains in Amazon Simple Storage Service (Amazon S3), encrypted with the client’s dedicated AWS Key Management Service (AWS KMS) key.

Responsible AI controls

MHK enforces safe LLM outputs through application-layer validation built into each agent’s processing pipeline. Every agent post-processes model responses against expected output schemas, cross-references extracted data with source documents to detect hallucinations, and rejects responses that fail confidence thresholds. Domain-specific checks verify that outputs reference only the patient’s own clinical records and match policy-specific medical criteria. LLM inputs and outputs are logged with full audit trails, which supports compliance review and reproducibility for every AI-assisted decision.

AWS services used

The following table summarizes the AWS services that compose the SmartProminence AI Orchestrator and the role each plays in the architecture.

Service Role in architecture
Amazon Bedrock Foundation model inference with IAM role-based authentication
Amazon ECS (Fargate) Containerized orchestration core, workflow controllers, and processing agents
Amazon SQS Event-driven inter-component communication with dead letter queues for fault tolerance
Amazon RDS (MySQL 8.4) Workflow definitions, execution state tracking, multi-AZ for high availability
Amazon S3 Job artifacts, document storage, immutable workflow configurations (KMS encrypted)
AWS KMS

Capability token signing and validation

Per-client encryption keys for multi-tenant data isolation

Amazon Cognito OAuth2/JWT authentication for API access and user identity
Amazon CloudWatch Logging, metrics, token usage tracking, and alerting (no PHI)
Elastic Load Balancing TLS 1.3-terminated application load balancer
Amazon VPC Network isolation with private subnets, VPC endpoints for service access

Security and compliance

Healthcare data demands the highest security standards, and MHK’s architecture implements defense-in-depth across every layer. The orchestrator processes protected health information (PHI) for multiple health plan clients simultaneously, making multi-tenant data isolation a foundational feature.

  • Per-client encryption. Every client has their own AWS KMS key. Documents stored in Amazon S3 are double-encrypted: S3 server-side encryption plus client-specific KMS encryption. Even if a job were somehow misrouted (which the token system helps prevent), the receiving agent could not decrypt another client’s data because it would not have access to that client’s KMS key. The database layer adds row-level encryption on top of Amazon Relational Database Service (Amazon RDS) storage-level encryption, providing defense-in-depth for data at rest.
  • Capability token model. A least-privilege token system limits what each component can access. A controller-scoped token can read workflow definitions and job data, dispatch agent tasks, and create step executions. An agent-scoped token can only read its step’s input, write its own result, call the LLM through the proxy, and upload artifacts. Tokens are generated per-dispatch through KMS, so even a compromised agent cannot reach data from other steps, workflows, or clients.
  • Network isolation. The database subnets have no internet access. AWS service communication (Amazon S3, Amazon SQS, AWS KMS, AWS Secrets Manager, Amazon CloudWatch, Amazon Elastic Container Registry (Amazon ECR)) flows through VPC endpoints, meaning no data ever traverses the public internet. Connections use TLS 1.3 for encryption in transit.
  • Compliance controls. LLM request and response bodies are not logged. Only token counts and content hashes are recorded. Workflow configurations are stored immutably in S3 for complete version history. Agents receive only the minimum context needed for their step, following the principle of data minimization.

Results and impact

The SmartProminence AI Orchestrator delivered measurable business outcomes across both MHK’s internal operations and their health plan clients.

90% reduction in manual review effort. For document intake workflows, processing time dropped from 5–10 minutes per document (manual) to under 1 minute (automated with human-in-the-loop verification). For complex medical director reviews, the system pre-gathers the relevant evidence and presents a structured summary, reducing the physician’s task from hours of document searching to a 30-second approval or denial decision.

85% faster AI feature deployment. New AI capabilities that previously required a full 3–6 month engineering release cycle now deploy in approximately 2 weeks. The engineering team defines workflow configuration and prompt logic without building custom infrastructure, compliance pipelines, or security controls for each feature.

Unified compliance posture. Instead of attesting each AI feature independently, MHK maintains a single orchestrator-level HIPAA and SOC 2 attestation that covers the agents. New agents inherit the orchestrator’s security controls automatically: per-client encryption, audit logging, token-based access, and data minimization.

Multi-tenant extensibility. Health plan clients can run AI workflows through the framework without building their own HIPAA-eligible infrastructure. Because the orchestrator is configuration-driven, MHK can onboard new use cases for existing clients or deploy entirely new health plan customers with minimal engineering effort.

Conclusion

The orchestrator’s controller-agent architecture provides a blueprint for organizations that need to scale AI capabilities across multiple use cases without multiplying their compliance burden. The key insight is that compliance infrastructure should be an orchestrator-level concern, not a per-feature concern, and that agentic orchestration patterns can be both powerful and auditable when designed with healthcare-grade security from the ground up.

Looking ahead, MHK is extending the orchestrator with conversational interfaces so case managers and medical directors can interactively query case data, using the same workflow memory and security infrastructure. The dynamic agent registry continues to grow as new medical management modules adopt AI-powered decision support. MHK is also exploring AWS Marketplace as a distribution channel to bring their HIPAA-eligible agentic framework to organizations beyond healthcare that require similar compliance thresholds.

Share your experience building HIPAA-eligible AI workflows in the comments or reach out if you’re exploring agentic architectures for regulated industries.

To learn more, get started with Amazon Bedrock and explore the Amazon Bedrock code samples to build your own agentic AI solutions on AWS.

 


About the authors