Post Syndicated from Nidhi Ramakant original https://aws.amazon.com/blogs/security/building-your-ai-vulnerability-harness-part-1/
Vulnerability scanners produce findings faster than manual triage can process them. Your developers ship more code with more dependencies, and the volume of candidate findings grows with it.
Many findings a scanner produces are unlikely to be exploited. The ones that matter need to reach an engineer fast, with enough evidence that they can act immediately rather than repeat the analysis. The challenge is separating signal from noise at the speed your pipeline demands.
This post shows you how to close the gap between detection and action. We built a three-layer pipeline that takes raw scanner findings and narrows them to a small, prioritized set with documented evidence of exploitability. The companion post, Configuring your AI vulnerability harness, will cover the steering file that drives the model’s behavior inside this pipeline. This post covers the layers around it.
Because the pipeline’s value comes from its filtering logic and evidence standards—not from any specific tool—you can substitute your own scanners and models without changing the architecture. Your AI provider, your infrastructure stack, and your scanner ensemble will differ from ours. The architectural decisions—what to filter at each layer, what counts as evidence, where to stop trusting the model—transfer regardless.
What a test harness is
A test harness is an automated framework that subjects a system to controlled inputs and observes whether outputs match expected behavior. Applied to vulnerability detection, the harness takes each candidate finding, constructs a test to evaluate exploitability, executes that test in a controlled environment, and records the evidence. The harness is the testing infrastructure, not the results it produces.
Traditional static application security testing (SAST) tools rely on pattern matching, with data-flow analysis varying by tool and language. They flag known-shape sinks well but struggle with multi-hop chains, business-logic flaws, context-dependent sanitization, and whether a sink is reachable from attacker-controlled input. AI-augmented analysis addresses that reasoning gap: tracing across files, inferring intent, and weighing deployment context.
The key conceptual distinction is that we don’t run a model against the codebase. The codebase is context passed to a model with a specific prompt. The model receives code and applies reasoning about how data flows, where trust boundaries exist, and whether security-relevant patterns are present.
Context scoping is one of the harder parts of this architecture because too much context can overwhelm the model’s reasoning window, while too little can produce analysis gaps that generate false negatives. For a basic AWS Lambda function, the relevant context might be a single file plus its AWS Identity and Access Management (IAM) role. For a complex microservice with shared libraries and layered infrastructure, deciding what to include requires understanding the application’s dependency graph. The model can reason about your deployment topology if you provide it; for example, by passing AWS Cloud Development Kit (AWS CDK) constructs or AWS CloudFormation templates as additional context. Starting narrow and expanding only when the model’s initial assessment is inconclusive generally produces better results than passing everything at once.
Building the pipeline
Security scanners can produce high volumes of findings each time they run. Many of these are false positives that look like vulnerabilities but aren’t real issues. To separate real problems from noise, the pipeline filters findings through three layers. Each layer applies a different type of evidence. By the end, a smaller number of prioritized issues backed by documented evidence remains.
Layer 1: Multi-scanner agreement
Plausibility comes first. Run multiple scanners independently against the same codebase: when two or more converge on the same finding, confidence increases. When only one scanner reports an issue and no other tool agrees, the pipeline flags that finding for extra scrutiny.
This approach—called multi-scanner agreement—is a low-cost way to increase confidence. It works with your existing tools and requires no new infrastructure.
Layer 2: Structural verification
The second layer verifies structure. AI-powered scanners describe how they think a vulnerability works: data enters through one function, passes through another, and reaches a point where it could cause harm. However, these descriptions can be wrong.
Before a finding moves forward, the pipeline verifies it against the actual code. It checks to see if the function that the scanner mentioned exists, that the file exists, and if data can flow along the path described by the scanner. If not, the pipeline rejects the finding.
This verification step uses the code’s abstract syntax tree (AST): a structured map of how functions, files, and data flows connect in your codebase. This step helps protect the credibility of your pipeline’s output. If unverified findings reach human reviewers, teams quickly learn to distrust the results, which reduces the security value of the tool.
Layer 3: Deployment context
The final layer adds deployment context. Vulnerabilities exist within the context of an architecture, which might include protective controls. These controls can include AWS WAF rules, network isolation, authentication requirements, and input validation.
This layer reads your infrastructure-as-code (IaC) templates—such as CloudFormation or Terraform files that define your cloud architecture—alongside your application source code. It then checks to see if an attacker can exploit a vulnerability by avoiding the controls that are in place.
Not all controls provide the same level of protection. A web application firewall rule that directly blocks the relevant attack technique provides stronger mitigation than one that addresses a different type of vulnerability. The pipeline treats control effectiveness as a spectrum, not a yes-or-no checkbox.
The result
The result is a short list of findings. Each one has cleared all three layers: it passed the plausibility check, its structure checks out against the code, and its deployment context indicates it can be taken advantage of. The number of findings decreases at each layer, and confidence in the remaining findings increases.
Principles
Four principles emerged from building this pipeline:
- Progressive filtering with escalating evidence – Each layer demands a different kind of proof. The first layer checks that the finding is plausible. The second determines if it’s structurally real. The third analyzes if it’s exploitable in context. Using three focused layers works better than using one comprehensive layer, because each layer catches a different type of error.
- Infrastructure context transforms prioritization – Without deployment context, you might assume that a critical-severity finding is more urgent than a medium-severity finding. However, a critical finding behind strong protective controls might be less urgent than a medium finding that’s directly exposed to the internet. To make this assessment, the pipeline needs access to your IaC templates, not only your application source code.
- Independent agreement outperforms single-tool confidence – No single scanner reliably detects the full range of vulnerabilities. No single model produces correct results across all situations. When multiple independent tools agree on a finding, that agreement provides a stronger signal than one tool’s self-reported confidence score. This principle generally holds across tool choices.
- AI-generated claims require verification before human review – Language models can produce outputs that sound plausible but might not be correct. Verifying AI-generated findings against the actual code structure—before a human reviews them—is what separates a useful system from one that erodes trust through confident-sounding errors.
How to apply this to your stack
You don’t need to build everything at once. Adopt them in order of immediate return:
- Start with indexing – Build (or use) an AST-based index of your codebase that surfaces entry points, sinks, and authentication configuration. This costs nothing to run and gives you a map of your attack surface before any AI is involved.
- Add multi-scanner triage – Take the output you already get from
Semgrep,CodeQL, or equivalent tooling, and run it through the AST gate plus a large language model (LLM) exploitability classification. This delivers substantial noise reduction for a modest investment: you’re filtering output from tools you already run. - Add hypothesis generation – After triage is stable, add AI-driven discovery of vulnerabilities scanners miss: multi-hop chains, business-logic flaws, and cross-package data flows. This is the step where steering quality matters a great deal; see Configuring your AI vulnerability harness.
- Add infrastructure context – Parse your CDK or CloudFormation, map controls to vulnerability classes, and apply the multipliers. This is where your prioritization becomes more targeted.
- Add live verification – Run proofs of concept (PoCs) against a deployed pre-production target with structured success criteria. This step requires the most operational overhead, so run it last and only for findings that have already passed your confidence threshold.
Treat the pipeline as something you tune over time. Validate it against known true and false positives, adjust the weights, and rerun. Your results will change.
What this doesn’t solve
The harness is detection infrastructure. It helps you see which findings are likely real and which deserve attention first. It doesn’t:
- Fix the code – Remediation is a separate problem. The harness produces findings with enough evidence that a human or agent can act on them; the act of remediation requires its own tooling and process.
- Replace human review for final action – Even after three layers of filtering, the output is a prioritized list of likely-exploitable findings, not proven exploits. Priority 0 (P0) findings still warrant review by an engineer before they drive code changes.
- Catch vulnerability classes outside the scanner ensemble’s detection profile – If none of your scanners look for a particular vulnerability category, the harness won’t surface it. AI-augmented hypothesis generation closes some of this gap, but novel or business-logic flaws still depend on what the model can reason about from code alone.
- Verify live infrastructure state by default – Parsing IaC tells you what was defined. It doesn’t tell you whether a WAF rule got disabled last week, or whether a security group was modified after deployment. Live verification is an additional integration, not a property of the layers themselves.
Each of these limits points to where the architecture extends, not where it breaks. The harness is the first floor of a multi-story system, not the whole building.
What comes next
In this post, we showed how a three-layer pipeline—multi-scanner agreement, structural verification, and deployment context—narrows scanner output to a small, higher-confidence set of findings that your team can act on. The architecture is tool-agnostic; what transfers is where to filter, what counts as evidence, and when to stop trusting the model.
The architecture produces consistent results when the model receives consistent instructions. Our companion post, Configuring your AI vulnerability harness, covers the steering file that encodes this methodology as machine-executable instructions, including the confidence formula, the infrastructure control multipliers, and the verification gates that help prevent hallucinated findings from reaching your engineers. The same model that produces rigorous, evidence-grounded assessments with well-crafted instructions can produce inflated severity ratings and fabricated attack chains without those instructions. Configuration is what turns capability into judgment.
The window between disclosure and exploitation is narrowing. Embedding automated detection into your development workflows helps your team keep pace. We built this pipeline through trial and error and are sharing it so you can skip some of the mistakes we made.
For more information about the services mentioned in this post, see the AWS Lambda, AWS WAF, and AWS CloudFormation documentation.
If you have feedback about this post, leave a comment in the Comments section below.