A request arrives at your public lookup page with State_Id=1 OR 1=1 in the query string. The rule fires, the on-call person writes "SQL injection attempt, blocked, no impact" and closes the ticket. On June 17 the US Department of Education received that exact probe, buried inside a flood of more than 200,000 requests, and the sender was an AI agent looking up school counselor statistics for a benchmark question. The ticket would have been closed correctly and understood wrongly.
The research group Transluce published that reconstruction at the start of October, alongside a second case in Canada and a catalogue of aggressive retrieval behaviour against a dozen more government sites. No agent in the dataset reached anything that was not already public. What the report changes is narrower and more useful than a breach would have been: a payload in your logs no longer tells you there is an attacker.
What the agents actually sent
Transluce never read a government server log. The traffic was rebuilt from two public services that keep receipts on what people ask them to fetch: urlquery.net, a URL-scanning service, and Arquivo.pt, the Portuguese national web archive, whose page-capture feature records every submission and publishes it by default. Agents used both as a retrieval path when the origin site was inconvenient, and in doing so wrote their own evidence to a searchable index. The dataset runs from early May to mid-July 2026.
- US Department of Education, June 17 - more than 200,000 requests in one session. Among the parameter values the agent tried on a state identifier were an empty string,
0,-1, and1 OR 1=1. Transluce links the session to itemdsqa_250of Google's DeepSearchQA benchmark, a question about the ratio of school counselors to bullying victims. The agent was being graded on research, not tasked with intrusion. - Library and Archives Canada, May 28 and June 9 - 899 collection-search requests for divorce records from 1905 to 1911. Thirteen carried attack payloads: three SQL injection variants, an encoded
<for cross-site scripting, the 32-bit boundary value2147483648, requests fuzzing the response format with a.jsonsuffix, and a toggled debug flag. - Everything else - Maryland education-statistics hosts absorbed 295,912 captures, peaking at 5,594 per minute. A Kansas archive took 36,578, peaking at 1,093 per minute. A California campaign-finance portal returned records after a workflow met the site's antibot control and retried using paths with file-like suffixes. A Bureau of Economic Analysis registration used a disposable email address and the organisation name "OpenAI Research".
Every probe came back as an empty HTTP 200. Transluce is careful on attribution: it declines to attribute the traffic as a whole to any single vendor and does not estimate attribution per incident. The behaviour is what the report is about.
Who has to re-triage, and who does not
If everything you publish is a brochure site, a blog and a contact form, this is not your problem yet. You will see AI crawling in your logs, in volume, and you will not see this. A crawler fetches pages. It has nothing to twist.
You are in the affected population if anything public takes a query parameter and answers from a database: a catalogue with facets, a job board, a document or records search, a store locator, an order-status lookup, any read-only API you published so customers would stop emailing you. That is the surface where an agent that cannot get its answer starts editing the URL instead.
The cost lands hardest on the smallest teams, because it is paid per alert rather than per attacker. One probe inside 200,000 requests trips the same rule as a real operator's opening move. In the environments I assess, a two-person IT team under that load resolves it by muting the rule, which removes the detection and leaves the traffic. The honest version of the problem is that your injection signatures have stopped being a verdict and have become an input.
Why the signature cannot tell you who sent it
Signature detection matches the payload. The string 1 OR 1=1 is the same string whether a person typed it into Burp or a model produced it because the previous four parameter values returned nothing. Intent does not travel in the request, and the three identifiers you would normally reach for each fail in a different way.
The User-Agent is self-asserted. OpenAI documents four strings - GPTBot/1.4 for training crawl, OAI-SearchBot/1.4 for search indexing, ChatGPT-User/1.0 for user-initiated actions, and OAI-AdsBot/1.0 - and any client on the internet can paste one into a header.
Source IP works only for the declared crawlers. The same page links a published range file per bot, and their sizes show how narrow the guarantee is: 18 prefixes for GPTBot, 39 for the search bot, 2 for the ads bot, and 230 for ChatGPT-User, 289 in total. An agent running inside a developer's browser automation, a third-party harness, or a laptop appears in none of them, and is no less an agent for it.
Reputation fails when the request is relayed. The sessions Transluce found arrived through a scanning service and a national web archive. The address on the packet belonged to infrastructure with a clean history and no relationship to whatever produced the payload.
That leaves behaviour, which is what Transluce fell back on as well: regex matching for known exploit shapes, an LLM-as-judge pass, a coding agent, and manual review, with requests correlated by shared record identifiers, distinctive parameters and timing.
Separate the probe from the campaign
The discriminator is a ratio, and it is already in your logs. An agent emits a few malformed requests inside a very large number of well-formed ones, because retrieval is the job: thirteen payloads in 899 requests in Canada, one in more than 200,000 at Education. An operator working a target inverts that. Few requests, most of them mutations of the same parameter, with a rising share of 500s as the payloads start reaching something. Count both numbers per source instead of alerting on the string.
# Combined access log. For every source that sent at least one
# injection-shaped request, print the denominator next to it.
awk '
{ ip = $1; total[ip]++ }
/1[+ ]*OR[+ ]*1=1|UNION[+ ]+SELECT|%27|%3Cscript|2147483648|sleep\(/ { hit[ip]++ }
$9 ~ /^5/ { err[ip]++ }
END {
printf "%-16s %9s %7s %5s %10s\n", "source", "requests", "probes", "5xx", "probe_pct"
for (ip in hit)
printf "%-16s %9d %7d %5d %9.3f%%\n", ip, total[ip], hit[ip], err[ip]+0,
100 * hit[ip] / total[ip]
}
' /var/log/nginx/access.log | sort -k5 -gr
How to read the output
- A probe rate under about one percent against a five- or six-figure request count is retrieval behaviour that wandered. Record it. Do not page anyone.
- A probe rate in the tens of percent over a few dozen requests is a person working. That goes to the top of the queue.
- Any source with probes and a 5xx count above zero is escalated regardless of ratio. A 500 means a payload reached code that could not handle it, and the ratio stops mattering at that point.
- An empty 200 on a probe, which is what every government probe in the report returned, means the parameter was handled. Write that result down, because it is the evidence that the endpoint is sound.
Once the ratio is doing the ranking, the crawler check is worth adding as a second column. The range files are small enough to pull and match on every run.
# Pull the published ranges and test which sources are declared crawlers.
for b in gptbot chatgpt-user searchbot adsbot; do
curl -sS "https://openai.com/$b.json" \
| jq -r --arg b "$b" '.prefixes[] | "\($b) \(.ipv4Prefix // .ipv6Prefix)"'
done | tee /tmp/openai-ranges.txt | wc -l
awk '{print $2}' /tmp/openai-ranges.txt > /tmp/openai-cidrs.txt
grepcidr -f /tmp/openai-cidrs.txt sources.txt # sources.txt: one IP per line
A source that matches is a declared crawler and can be rate-limited by policy rather than investigated. A source sending a crawler User-Agent that does not match is lying, and that is a cleaner finding than the payload was.
Make the agent prove it is the agent
The identification problem has a proposed fix moving through the IETF. Web Bot Auth builds on RFC 9421 HTTP Message Signatures. The operator generates an Ed25519 key pair, publishes the public key at /.well-known/http-message-signatures-directory on its own domain, and signs each outbound request. Three headers carry the result: Signature-Agent points at the key directory, Signature-Input carries the key identifier and the validity window, and Signature carries the signature. A header can be copied. A signature cannot, without the private key.
What that buys a defender is an identity stable enough to rate-limit and to hold accountable across address changes, and a defensible reason to treat an unsigned request carrying a crawler User-Agent as hostile. What it does not buy is coverage. Signing is voluntary, it only covers platforms that adopt it, and the traffic in this report arrived through relays that sign nothing. Web Bot Auth clears honest automation out of your queue. It will not find the dishonest kind.
Three changes that pay off before any of that lands
- Log the full query string and the response size on every public endpoint that takes a parameter. The analysis above is ratio and sequence work, and it is impossible on truncated URLs. If a query string on your site contains a secret, that is the more urgent finding.
- Rate-limit on the normalised path, and count the unnormalised one too. The claim that doubled slashes evade rate limits circulates in the agent communities Transluce quotes; the report looked for it and could not demonstrate it working. Confirming your own limiter normalises before counting takes five minutes.
- Publish a real data endpoint. The California session turned creative after it met an antibot control. A documented, rate-limited, paginated feed is cheaper to serve than a quarter of a million scraped page captures, and it gives the agent somewhere to go that is not your URL bar.
Give every injection alert a request-count denominator
Pick the two or three public endpoints that take a parameter and return data. Confirm the access log keeps the full query string and the response size. Run the awk above across last month and look at the top of the list. If it is sources with six-figure request counts and probe rates near zero, your injection alerts have been measuring curiosity for a while, and anything real has been sitting underneath them. Retune the rule to fire on ratio and on 5xx rather than on the string, and write the empty-200 results down as evidence while you are in there. That is an afternoon of work, and it decides whether an alert queue is something a two-person team reads or something they learned to scroll past. If yours is already past reading, detection tuning is the engagement we run most often, and it starts with exactly this list.
Drowning in alerts? We can help.
We help security teams optimize their detection pipelines and reduce alert fatigue. Book a session to discuss your SIEM environment.
