Lifestyle

Cloudflare Used Frontier AI Models to Attack Its Own WAF: Most Attempts Were Blocked, and the Gaps Led to New SSRF Rules

On September 29, 2026, Cloudflare said it had advanced AI models act as attackers against its own web application firewall (WAF). According to Cloudflare, most of the 1,107 attempts were blocked. Human review left 49 findings worth investigating, and these led to SSRF rule updates that Cloudflare says benefit all its customers.

About 8 min read

Cloudflare Used Frontier AI Models to Attack Its Own WAF: Most Attempts Were Blocked, and the Gaps Led to New SSRF Rules
Image: Mokaair (Original editorial artwork)

What happened: Cloudflare had AI play the attacker against its own firewall

On September 29, 2026, Cloudflare published a post on its official blog written by Vikram Grover, Daniele Molteni and Kuber Nandwani. Cloudflare said customers keep asking, "Is your WAF ready for frontier AI models?", so it decided to find out. A WAF (web application firewall) is a layer of protection that sits in front of a website and filters out malicious requests. Frontier AI models are the most advanced AI models currently available.

Cloudflare says large language models are strong attackers because they can change attack payloads faster than any human hacker and adjust their tactics based on real-time responses. To examine this, Cloudflare used dynamic testing, which means probing a running website from the outside. The models could not see the source code or the WAF rules, and could see only part of the HTTP response data. Cloudflare stressed that an unblocked request was only a lead for human review, not a confirmed exploit.

Cloudflare Used Frontier AI Models to Attack Its Own WAF: Most Attempts Were Blocked, and the Gaps Led to New SSRF Rules
Mokaair editorial verification flow · Image: Mokaair (Original editorial artwork)
Read the full description

Sources are collected, independently checked, then reviewed by Jev.

How the test worked: an AI testing loop controlled by code

According to Cloudflare, the test target was an authorized customer staging environment, a test copy of a site used before changes go live. Cloudflare ran 45 scenarios. Of these, 44 covered six common attack types: cross-site scripting (XSS), SQL injection, command injection (sneaking system commands into a request), server-side request forgery (SSRF), path traversal or local file inclusion (LFI), and Log4j. The remaining scenario covered log injection and was reported separately.

Cloudflare said the test system was written in Python and called the model twice per loop: once to propose the next variation and once to review the response. The model did not send requests directly. Instead, code checked that the target host was on an allowlist, disabled redirects, logged every attempt and enforced the attempt limit. The model had no access to rule contents, rule IDs or scoring details, and it could not deploy rules or change protection settings.

Cloudflare described the WAF configuration of the test zone as follows. WAF Attack Score blocked requests scored 30 or below. The Cloudflare Managed Ruleset, a set of detection rules that Cloudflare maintains, was fully enabled. The OWASP Core Ruleset ran at Paranoia Level 3. Cloudflare noted that the results reflect the overall configuration as a whole, not the performance of any single rule.

The results: the numbers Cloudflare published

Key figures from Cloudflare's AI WAF test (source: Cloudflare official blog, not independently verified)
ItemNumberCloudflare's explanation
Attempts logged1,107All attempts generated by the models across 45 scenarios; not all produced usable results
Filtered result set607558 blocked requests plus 49 findings
Blocked requests558Stopped by the WAF before reaching the application
WAF-related findings49Included in remediation analysis after human review; 48 were in command injection and SSRF
Attempt limit per scenario25Some scenarios began repeating earlier ideas as they neared the limit

Cloudflare said the overall results were strong, with near-complete coverage of XSS, LFI, SQL injection and Log4j. Some attempts were not counted. In those cases the model did not produce a usable request, the request did not reach the target, or the generated content was harmless.

Cloudflare also described one SSRF scenario. Cloud metadata services can hand out temporary credentials, and an SSRF flaw can let an attacker make an application fetch that data for them. The model wrote the same cloud metadata address in several different forms, and the WAF blocked all of them except one. That single unblocked attempt produced only a redirect. Cloudflare noted that there is no evidence the application actually fetched metadata, so this was only a lead worth further investigation.

From findings to protection: human review and new rules

Cloudflare said every unblocked request had to pass a five-question review before it was counted as a finding:

  1. Did the test tool actually send a valid request?
  2. Was the request clearly not blocked?
  3. Is the mutated request still malicious?
  4. Is this behavior within the scope of what a WAF can address?
  5. Can engineers safely reproduce it?

According to Cloudflare, the findings were grouped into four sets of candidate rules. These were tested against live traffic before they could protect customers, to check the risk of false positives, which means wrongly blocking legitimate visitors. This work led to three changes to the Managed Ruleset. The July 21 release added the "SSRF - Obfuscated Host" and "SSRF - Restricted Protocol" detections and improved the existing "SSRF - Cloud" rule.

What it means for general readers and site administrators

For general readers, the report shows AI models being used to attack systems in order to test their defenses. Cloudflare had AI try large numbers of variations, and then humans judged which ones truly needed fixing. In Cloudflare's words, the model generated the requests, but people decided what mattered; without human review, there were no findings. Cloudflare also said that two versions from the same model family produced different variations but found the same underlying issues.

For site administrators, Cloudflare's first recommendation is to keep software up to date, because a request that bypasses the WAF still needs an exploitable application vulnerability to succeed. It also advises making sure Managed Rules and WAF Attack Score are configured correctly. Cloudflare said customers do not need to reproduce this experiment. It recommends first running Managed Rules in log mode, which records matches without blocking them. Administrators can then review matched requests in Security Events and switch rules to block once legitimate traffic is confirmed to be unaffected.

Cloudflare also said it will share results from white-box testing in a future post. In that kind of test, the model knows both the application's vulnerabilities and the WAF rules protecting it.

Frequently asked questions

What is the WAF Cloudflare tested?

A WAF is a web application firewall that sits in front of a website and filters malicious requests. Cloudflare tested its own WAF, configured with WAF Attack Score, the Cloudflare Managed Ruleset and the OWASP Core Ruleset.

Did the AI models successfully break into the site?

According to Cloudflare, there is no evidence of that. Unblocked requests were treated only as leads for human review. In the SSRF scenario it described, Cloudflare explicitly said there is no evidence the application fetched metadata.

What is SSRF, and why does it matter here?

Server-side request forgery (SSRF) is a flaw that can let an attacker make an application fetch data on their behalf, for example from cloud metadata services that can hand out temporary credentials. According to Cloudflare, 48 of the 49 findings involved command injection and SSRF, and all three rule changes concern SSRF.

What changed as a result of the test?

Cloudflare said the July 21 release of its Managed Ruleset added the "SSRF - Obfuscated Host" and "SSRF - Restricted Protocol" detections and improved the existing "SSRF - Cloud" rule.

I'm a Cloudflare customer. Do I need to run a similar test myself?

Cloudflare said customers do not need to reproduce this experiment. It recommends confirming that Managed Rules and WAF Attack Score are configured correctly, watching in log mode before switching to block, and keeping software up to date. Teams already doing application security testing can run their tests against a staging hostname protected by the same Cloudflare controls as production.

Does this mean a WAF is enough to protect a website?

No. Cloudflare itself stresses that a WAF is only one layer of protection. A request that bypasses the WAF still needs an exploitable application vulnerability to succeed, so patching software remains one of the strongest defenses.

Browse the latest news in this topic

Latest travel guides

Sources

Lifestyle