All posts by Vikram Grover

We tested our own WAF with frontier AI models. Here’s what we found

Post Syndicated from Vikram Grover original https://blog.cloudflare.com/adaptive-ai-waf-testing/

“Is your WAF ready for frontier AI models?” We keep hearing this question from our customers, so we decided to find out.

When it comes to exploiting applications, what LLMs are really good at is iterating and mutating attack payloads faster than any human hacker could do. LLMs can use real-time responses to iterate and change their techniques by, for example, testing different encodings, sending the payload in a different part of the HTTP request, or moving to the next vulnerability to test.

Even before LLMs were around, security engineers used two common approaches to test applications: static and dynamic application security testing. The former analyzes code without executing it to identify vulnerabilities, while the latter probes running applications to find runtime flaws. There are plenty of works scanning code with frontier AI models, including details on how to build your own harness.

For the project described in this blog post, we took a dynamic approach: making the LLM act as if it was a hacker to evaluate whether a WAF is doing its job. The LLM had no visibility into source code, no view of the WAF's rules, and could only see selected HTTP response data.

We built a WAF tester that starts from known exploits and then iterates by changing how it is encoded or delivered, sends it again, and uses the response to choose the next variation. A request that was not blocked became a lead for human review, not a confirmed exploit.

We ran the tester against an authorized customer staging environment across six attack categories and recorded 1,107 attempts. After reviewing the non-blocked requests and removing malformed, benign, duplicate, and out-of-scope observations, the vast majority of the attacks were blocked by the Cloudflare WAF. The requests that got through helped us create new detections to harden our security to benefit all Cloudflare customers.

Here we will explain how we set up the system, the types of attacks we tested, which attack vectors bypassed the WAF more easily, and how we fixed it. Most importantly, we share what we learned from this process and how this exercise is becoming a foundational building block of our WAF development lifecycle.

Finally, we offer guidance to help you correctly deploy your WAF in front of your application and, most importantly, patch your software. A payload that bypasses the WAF still needs an exploitable application to succeed, so keeping your stack up-to-date remains one of the strongest defenses against attackers.

How the adaptive loop works

To test our WAF with frontier models, we built a system that iterates over multiple scenarios. A scenario means choosing one attack category, placing the input in a specific part of the request, starting with a version the WAF already blocked, and giving the tester a fixed number of attempts to try other variations. The loop runs LLM models twice: the first is the proposal call, the second is the review call.

The first call receives the starting request, the context, a short history of earlier results, and suggests the next variation, then the code builds and sends the request. The review call receives the request context, response status, selected headers, and the response body. The loop stops when mutations stop producing useful variations or when a hard coded attempt limit has been reached.

Both model calls work without access to WAF internal information. Neither receives rule expressions, rule IDs, WAF Attack Score details, or the identity of the security layer that acted. We implemented the system in Python rather than wrapping an existing penetration-testing tool. It handles HTTP replay, scenario orchestration, state tracking, and result collection.

In the current implementation, the models do not send requests directly — code controls what happens at each step. Before each request, it checks the target hostname against an allowlist, disables redirects, records the attempt, and enforces the attempt limit. After each request, it records the response and uses the model's review to choose the next predefined step. Response text may appear in a later prompt, so the tester treats it as untrusted input. Neither model call can deploy a rule nor change enforcement.

The system records structured evidence for each attempt.

Six attack categories against one WAF configuration

The main run targeted an authorized customer staging environment protected by Cloudflare’s WAF. We used an allowlisted test User-Agent so the customer’s automated-traffic controls would not stop the test before requests reached the WAF.

We ran 45 scenarios. For each, we looked for ways to deliver the same attack differently: different encoding, different part of the request, or the same destination written another way. Of these, 44 covered six attack categories: cross-site scripting (XSS), SQL injection (SQLi), command injection (CMDi), server-side request forgery (SSRF), path traversal or local file inclusion (LFI), and Log4j. The remaining scenario covered log injection, reported separately.

The WAF in the test zone was configured as follows: WAF Attack Score blocking scores of 30 or below, all Cloudflare Managed Ruleset enabled, and OWASP Core Ruleset with Paranoia Level 3.

For the headline measurement, we recorded whether the WAF blocked each request or not. The results describe the configured WAF boundary as a whole, not the performance of any individual rule or detection mechanism.

What adaptation looked like in one recorded session

Here is an example of how the LLM adapts a Server-Side Request Forgery (SSRF) attack during the test.

Cloud metadata services can expose temporary credentials to workloads. An SSRF vulnerability can let an application fetch that data on an attacker's behalf. A WAF can help stop the malicious request before it reaches the application, but it is only one layer of protection.

In this SSRF scenario, the tester sent the same cloud metadata address in different forms (such as integer, octal, and trailing-dot representations of the same IP) and placed it in different parts of the request. The WAF blocked all of them except one. At attempt 18, the model kept the same request structure as the previous blocked attempt and switched to the trailing-dot form. The client encountered a redirect rather than a WAF block.

The table below shows selected moments from the session. The hypothesis column summarizes what the model said it was trying before each move. It is not a verbatim transcript, and it is not proof that the explanation was correct.

Attempts 17 and 18 are an interesting pair: same request structure, different host representation. One was blocked, one was not. That gave us a specific question: does the trailing dot change how the WAF reads the destination? It was a lead to investigate, but not proof that metadata was accessed.

This was one selected trajectory among 45 scenarios. The next section shows how we counted and triaged the full run.

What we found

Our tester generated 1,107 attempts and the overall result was strong with XSS, LFI, SQLi, and Log4j having near full coverage. While the run produced useful findings, it also produced noise. After human review, we were left with 49 findings worth investigating, 48 of them belonging to CMDi and SSRF. 

Here is how they break down:

Metric

Value

What it means

Recorded mutation attempts

1,107

Model iterations across 45 active scenarios; not all produced a usable result

Post-triage result set

607

The 558 blocked requests plus 49 documented WAF-relevant findings

Blocked requests

558

The WAF stopped these before they reached the application

WAF-relevant findings

49

Documented for remediation analysis after human review

The rest did not produce a result worth counting as the model failed to generate a usable HTTP request, some failed before reaching the target, or the payload generated was benign.

When a request was not blocked, we worked through five questions before counting it as a finding:

Question

Why it matters

Did the tester actually send a valid request?

If the model failed or the request never reached the target, the result tells us nothing about the WAF.

Was the request clearly not blocked?

An ambiguous response is not enough to count.

Was the request still malicious?

Changing a request to get it past the WAF can also make it harmless.

Did the behavior belong to the WAF?

Some attacks only work through DNS or network paths the WAF cannot stop at request time.

Could engineers reproduce it safely?

A fix needs a stable test case with a clear expected result.

We removed anything that failed those checks and combined duplicate cases. What remained became the input for rule, normalization, and mitigation work.

Findings became detections

Not every finding needed a new rule. Some pointed to gaps in existing Managed Rules coverage. Others pointed to how the WAF normalized the request or belonged to another security control. We replayed each case and decided where the change should happen.

We grouped related findings into four sets of candidate rules, validated each finding, and tested candidates against live traffic before any rule could protect customer traffic.

Before a new or updated rule can protect customer traffic, we check its impact on legitimate traffic and assess false-positive risk. Some of the issues we find when evaluating a new rule candidate include:

Issue

Next step

Missing or narrow detection

Review whether existing rules cover the finding

Equivalent inputs interpreted differently

Engine or normalization review

False-positive risk is too high

Revise or reject the candidate

This work contributed to three changes in Cloudflare's Managed Ruleset: new detections for SSRF – Obfuscated Host and SSRF – Restricted Protocol in the July 21 release, and improvement of the existing SSRF – Cloud rule. The SSRF – Obfuscated Host detection came directly from requests that encoded internal addresses in non-standard numeric forms.

What we learned

The model was only one part of the test. We ran the same scenarios with two versions of the same model family. They produced different variations – and the same underlying issues appeared in both. Because request replay and evidence capture stayed consistent, we could compare the runs without treating either model's output as ground truth.

More attempts within one scenario did not always find more. Some scenarios started repeating earlier ideas near the end of the 25-attempt limit. We got broader coverage by testing more starting requests, attack categories, and input locations instead of extending one sequence.

The model generated requests. We decided which ones mattered. A request that was not blocked still needed replay and human review before it could become a finding, a mitigation, or a regression test. Without that review, there were no findings.

What customers can do now

WAF is just one layer of detections you can deploy. When you deploy all available protections you increase the effectiveness of your overall stack. 

First of all, check that Managed Rules, WAF Attack Score are set up correctly in front of your application. Other tools you can deploy include API Security, Bots and Fraud detection, and Threat Intelligence to strengthen your posture even further. For example, positive security controls add a different layer: instead of looking only for known attack patterns, they define the request shapes an application expects and identify inputs outside that contract. This drastically reduces your attack surface area. 

Customers do not need to reproduce this experiment. To maximize the number of rules deployed in front of your application, we recommend running Managed Rules in log first, review matching requests in Security Events, and confirm legitimate traffic is unaffected before moving a rule to Block. Alternatively, customers can reach out to their account team to get Attack Signature Detection turned on, on their zones. This new feature simplifies how to review matched traffic and how to deploy signature detections. If you already perform application security testing, run those tests against a staging hostname protected by the same Cloudflare controls as production.

Next steps

By combining adaptive AI-driven testing with human triage and validation, we found detection gaps that fixed tests might miss and turned those findings into stronger WAF protections, improving our block rate. In a future post, we will share results from further testing using a white-box approach, where the model knows both the application’s vulnerabilities and the WAF rules protecting it.

Improving the accuracy of our machine learning WAF using data augmentation and sampling

Post Syndicated from Vikram Grover original https://blog.cloudflare.com/data-generation-and-sampling-strategies/

Improving the accuracy of our machine learning WAF using data augmentation and sampling

Improving the accuracy of our machine learning WAF using data augmentation and sampling

At Cloudflare, we are always looking for ways to make our customers’ faster and more secure. A key part of that commitment is our ongoing investment in research and development of new technologies, such as the work on our machine learning based Web Application Firewall (WAF) solution we announced during security week.

In this blog, we’ll be discussing some of the data challenges we encountered during the machine learning development process, and how we addressed them with a combination of data augmentation and generation techniques.

Let’s jump right in!

Introduction

The purpose of a WAF is to analyze the characteristics of a HTTP request and determine whether the request contains any data which may cause damage to destination server systems, or was generated by an entity with malicious intent. A WAF typically protects applications from common attack vectors such as cross-site-scripting (XSS), file inclusion and SQL injection, to name a few. These attacks can result in the loss of sensitive user data and damage to critical software infrastructure, leading to monetary loss and reputation risk, along with direct harm to customers.

How do we use machine learning for the WAF?

The Cloudflare ML solution, at a high level, trains a classifier to distinguish between various traffic types and attack vectors, such as SQLi, XSS, Command Injection, etc. based on structural or statistical properties of the content. This is achieved by performing the following operations:

  1. We inspect the raw HTTP input and perform some number of transformations on it such as normalization, content substitutions, or de-duplication.
  2. Decompose or partition it via some process of tokenization, generate statistical information about the content, or extract structural data.
  3. Compute optimal internal numerical representations of the inputs via the process of training the model. The nature of these internal representations depends on the class of model and architecture.
  4. Learn to map internal content representations against classes (XSS, SQLi or others), scores or some other target of interest.
  5. At run-time, use previously learned representations and mappings to analyze a new input and provide the most likely label or score for it. The score ranges from 1 to 99, with 1 indicating that the request is almost certainly malicious and 99 indicating that the request is probably clean.
Improving the accuracy of our machine learning WAF using data augmentation and sampling

This reasonable starting point stumbles immediately upon a critical challenge right from the start: we need high quality labeled data, and lots of it as that has the biggest impact on model performance. Contrary to well-researched fields like image recognition, text sentiment analysis, or classification, large datasets of HTTP requests with malicious payloads embedded are difficult to get.

To make matters even harder, strict implementation requirements for a production-quality WAF restrict the complexity of our potential ML models or architectures to ones that are relatively simple and light-weight, implying that we cannot simply pave over shortcomings of the data.

Data and challenges

The selection of a dataset is likely the most difficult of all the aspects that contribute to the final set of attributes of a machine learning model. In most cases, the model is tasked with learning the distribution of the data in some statistical sense, thus choosing and curating the dataset to ensure that the desired properties of the final solution are even possible to learn is incredibly crucial! ML models are only as reliable as the data used to train them. If we train an ML model on an incomplete dataset, or on data that doesn’t accurately represent the population, predictions might be inaccurate as they will be a direct reflection of the data.

To build a strong ML WAF, a good dataset must have large volumes of heterogeneous data covering malicious samples for all attack categories, a diverse set of negative/benign samples, and samples representing a broad spectrum of obfuscation techniques.

Due to those constraints, creating a solid dataset has a number of challenges:

Privacy

Privacy requirements limit data availability and how it can be used. Cloudflare has strict privacy guidelines and does not keep all request data – it simply isn’t available, and what is available must be carefully selected, anonymised, and stripped of sensitive information.

Heterogeneity of samples

Due to the wide assortment of potential request content types and forms, finding enough benign samples is difficult. Furthermore, it is challenging to collect data that represents requests with various charsets and content-encodings. Covering all attack configurations is also important because some attacks can be inserted into essentially any kind of request (e.g. five bytes in a huge “regular” request)

Sample difficulty

We want a dataset with a good mix of attack techniques and isn’t dominated by the ones that are easily generated by tools which simply swap out constants, transform expressions through invariants, and so on (sqli-fuzzer). Additionally, the vast majority of freely available samples in the wild are fairly trivial auto-generated payloads as part of indiscriminate scanning and discovery tools. They have very similar structural and statistical characteristics. Some of them are fairly old as well and do not reflect the current software landscape. How to “grade” the sample difficulty is not immediately obvious! What’s easy to a human may not be easy for a particular preprocessor/model, and vice-versa.

Noisy labels

Label noise affects results a lot, especially when it comes to esoteric, specific, or unusual attacks which are likely to be classified as benign by rules WAF.

What’s the strategy to overcome this?

Data augmentation

In simple terms, Data Augmentation is a process of generating artificial (but realistic) data to increase the diversity of our data by studying statistical distribution of existing real-world data.

This is crucial for us because one of the biggest concerns with rules-based WAFs is false positives. False positives are a serious challenge for WAFs because the risk of accidentally filtering legitimate traffic deters users from employing very strict rulesets. Data augmentation is used to build a solution that does not rely on observing specific high-risk keywords or character sequences, but instead uses a more holistic analysis of content and context, making it considerably less likely to block legitimate requests.

There are many sequences of characters which appear almost exclusively in payloads, but are themselves not dangerous. In order to reduce false positives and improve overall performance, we focussed on generating a lot of heterogeneous negative samples to force the model to consider the structural, semantic, and statistical properties of the content when making a classification decision.

In the context of our data and use cases, data augmentation means that we mutate benign content in a variety of ways as the content will remain benign (this isn’t going to accidentally turn it into a valid payload, with probability 1). For instance, we can add random character noise, permute keywords, merge benign content together from multiple sources, and so on. Alternatively, we can seed benign content with ‘dangerous’ keywords or ngrams frequently occuring in payloads – this results in a benign sample, but ideally will teach the model not to be too sensitive to the presence of malicious tokens lacking the proper semantics and structure.

Benign content

First and foremost, generating benign content is way easier. Mutating a malicious block of content into different malicious blocks is difficult because malicious payloads have a stricter grammar and syntax than general HTTP content due to the fact that it has code, therefore they must be manipulated in a specific manner.

However, there are a few options  if we want to do this in the future. Tools like sqli-fuzzer,  automates the process of fuzzing a given payload by applying transformations which preserve the underlying semantics while changing the representation or adding obfuscation. Outside existing third-party tools, it’s possible to generate our own malicious payloads using various “append malicious content to non-malicious content” techniques, with the trade off that this doesn’t actually generate *new* malicious content, just puts it into a different context.

Pseudo-random noise samples

A useful approach we identified for bolstering the number of negative training samples was to generate large quantities of pseudo-random strings of increasing complexity.

The probability of any pseudo-random string (drawn from essentially any token distribution) being a valid payload or malicious attack is essentially zero, but we can build a series of token sampling distributions that make it increasingly difficult for the model to distinguish them from a real payload, and we discovered that this resulted in dramatically better performance in terms of false positive rate, robustness, and overall model properties.

This approach works by taking a collection of tokens and a probability distribution over these tokens, and independently sampling a stream of tokens from it to create our ‘sample’. Each sample length is selected from a separate discrete sample length distribution.

For an extremely simple example, we could take a token collection consisting of ASCII characters and a uniform sampling distribution:

['a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i', 'j', 'k', 'l', 'm', 'n', 'o', 'p', 'q', 'r', 's', 't', 'u', 'v', 'w', 'x', 'y', 'z', '0', '1', '2', '3', '4', '5', '6', '7', '8', '9']

We sample random strings of length 0-32 from this to get some (uninteresting) negative samples:

8hwk1d740hfstbb4aogbpi4qayppvdl41b6blornuzktp4yl

1deq7rug1zftmn9tjr73yttjnye99zh2140z2x9lr8n6sxhucdgn6bmqvfv7auw8fwbkrtxilk45ht-

We wouldn’t expect even a very simple model to struggle to learn that these samples are benign,  but as we increase the complexity of the token collections, we can move towards much more ‘difficult’ noise examples, including elements such as: fragments of valid URIs, user agents, XML/XSLT content or even restricted language identifiers, or keywords.

Here are some examples of more complex token collections and the kinds of random strings they produce as our negative samples:

Ascii_script: alphanumeric characters plus  ‘<‘, ‘>’, ‘/’, ‘</’, ‘-‘, ‘+’, ‘=’, ‘< ‘, ‘ >’, ‘ ‘, ‘ />’

Improving the accuracy of our machine learning WAF using data augmentation and sampling

alphanumerics, plus special characters, plus a variant of full javascript or sql keywords and (multi-character) sub-token fragments

Improving the accuracy of our machine learning WAF using data augmentation and sampling

It’s fairly straightforward to construct a suite of these noise generators of varying complexity, and targeting different types of content: JSON, XML, URIs with SQL-esque ‘noise’, and so on. As the strings get sufficiently long, the probability that they will contain at least some dangerous looking subsequences grows, so it’s also an excellent test of model robustness.

We make extensive use of noise strings to enhance the core dataset used for training and testing the model by directly training the model on increasingly difficult noise before fine-tuning on exclusively real data, appending noise of varying complexity to malicious(real) samples or benign samples to both induce and test for model robustness for padding attacks, and estimating false positive rate for certain classes of benign content.

Beyond independent sampling of random strings?

A natural extension to the above method for generating pseudo-random strings is to drop the ‘independence’ assumption for sampling tokens. This means that we’re starting to emulate the process by which real data is generated, to some extent, yielding samples with increasingly realistic local (and eventually global) structure. Some approaches for this might include a simple Markov chain, and extend all the way to state-of-the-art Large Language Models.

We experimented with using contemporary autoregressive language models trained on our corpus of real malicious payloads and found it extremely effective at generating novel payloads, as well as transforming payloads into sophisticated obfuscated representations. As the language models approached convergence on the data the likelihood of each sample being a valid payload approached 100%, allowing us to use early samples as ‘extremely strong negatives’ and the later samples as positive samples. The success of this work has suggested that deeper investigation into the use of language models for security analysis may be fruitful, not only for training classifiers, but also for creating powerful adversarial pen-testing agents.

Results summary

Let’s see a comparative summary of results and improvements, before and after the augmentation:

Model performance on evaluation metrics

The effectiveness of machine learning models for classification problems can be evaluated using a wide range of metrics, including accuracy, precision, recall, F1 Score, and others. It is important to note that in addition to using quantitative metrics, we also consider the model’s general properties and behavioral constraints. This criteria and metrics-based approach is especially important in our domain where data is inherently noisy, labels are not trustworthy, the domain of the inputs is extremely large, and hard to cover with samples.

For this post, we will concentrate on key quantitative metrics like F1 score even though we examine a variety of metrics to assess the model performance. F1 score is the weighted average (harmonic mean) of precision and recall. We can represent the F1 score with the formula:

Improving the accuracy of our machine learning WAF using data augmentation and sampling

Where,

True Positives (TP): malicious content classified correctly by the model

False Positives (FP): benign content that the model classified as malicious

True Negatives (TN): benign content classified correctly by the model

False Negatives (FN): malicious content that the model classified as benign

Since this formula takes false positives and false negatives into consideration, this score is more reliable than other metrics. There are a few methods to calculate this for multi-class problems, like Macro F1 Score, Micro F1 Score and Weighted F1 Score. Although each method has advantages and disadvantages, we obtained nearly identical results with all three methods. Below are the numbers:

Without Augmentation With Augmentation
Class Precision Recall F1 Score Precision Recall F1 Score
Benign 0.69 0.17 0.27 0.98 1.00 0.99
SQLi 0.77 0.96 0.85 1.00 1.00 1.00
XSS 0.56 0.94 0.70 1.00 0.98 0.99
Total(Micro Average) 0.67 0.99
Total(Macro Average) 0.67 0.69 0.61 0.99 0.99 0.99
Total(Weighted Average) 0.68 0.67 0.60 0.99 0.99 0.99

The important takeaway is that the range of this F1 score is best at 1 and worst at 0.

The model after augmentation appears to have similar precision and recall with good overall performance, as indicated by a value of 0.99 after augmentation, compared to 0.61 for Macro F1.

So far in the results summary, we’ve only discussed F1 Score; however, there are other improvements in characteristics that we’ve observed in the model that are listed below:

False positive characteristics

  • Estimated false positive rate reduced by approximately 80% on test data sets. There are significantly fewer false positives involving PromQL and other SQL-structured analogues. PromQL examples result in high scores and are classified correctly:
Improving the accuracy of our machine learning WAF using data augmentation and sampling

Today, the only major category of false positives are literal SQL or JavaScript files.

  • General false positive rate on noise from JSON-esque, XML/SOAP-esque, and SQL-esque content-generators reduced to about a 1/100,000 rate from about 1/50 to 1/1.

True positive characteristics

  • True positive rate for highly fuzzed content is vastly improved. Models trained solely on real data were easily bypassed by advanced fuzzing tools, whereas models trained on real plus augmented data are extremely resistant, with many payloads receiving higher risk scores as fuzzing increases. Examples:
Improving the accuracy of our machine learning WAF using data augmentation and sampling

These yield approximately same scores as they are a result of only a few byte   alterations

  • Proportion of client-provided test sets that primarily contain payloads not blocked by rules-waf for XSS/SQLi successfully classified is about 97.5% (with the remaining 2.5% being arguable) up from about 91%.
  • Padding a payload with almost any amount of ASCII, JSON-esque, special-characters, or other content will not reduce the risk score substantially. Due to the addition of hard noise long length augmented training samples, even a six byte payload in a 100 kilobyte string will be caught. Examples:
Improving the accuracy of our machine learning WAF using data augmentation and sampling

They both generate similar scores even though the latter has junk padding around the payload.

Execution performance

  • Runtime characteristics are unchanged for inference.

On top of that, we validated the model against the Cloudflare’s highly mature signature-based WAF and confirmed that machine learning WAF performs comparable to signature WAF, with the ML WAF demonstrating its strength particularly in cases of correctly handling highly obfuscated or irregularly fuzzed content (as well as avoiding some rules-based engine false positives). ​​Finally, we were able to conclude that augmentation helps in improving the model performance and induce the right set of properties.

Conclusion

We built a machine learning powered WAF, with the substantial challenge to gather a diversified training set, given constraints to avoid sensitive real customer data for privacy and regulatory considerations. To create a broader and diversified dataset without requiring vast amounts of sensitive data, we used techniques such as fuzzing, data augmentation, and synthetic data generation. This allowed us to improve the solution’s false positive robustness and overall model performance.

Furthermore, these techniques reduced the time complexity required to retrieve/clean real data, and helped induce the correct model behavior. In the future, we intend to investigate autoregressive language models to generate synthetic pseudo-valid payloads.