Contents

09/27/2026

Fake AI Crawler User Agents in Nine Days of Server Logs

Anyone can put ClaudeBot in a request header. I said so in my AI bot census, and then moved on, because one week of logs didn't give me a clean example. Nine days of daily summaries from my Site Ops dashboard did.

The summaries cover September 17 to 26, 2026, on five small sites. They keep request counts per crawler name and per response status. They don't keep IP addresses, on purpose, so what follows is what a user agent string alone can and can't tell you.

What a user agent can tell you

  • What the sender claims

    The user agent is a header the client writes itself. It costs nothing to copy a well-known crawler's name.

  • What the server answered

    The status code is the server's side of the story, so it can't be faked by the client.

  • Not who actually asked

    Only the source address, checked against the vendor's published ranges, says whether a request really came from that vendor.

The day the rare names showed up

September 24 was an ordinary day on nextdayvideos.com: 3,073 requests in all. Mixed in were a few crawler names that never appeared on that site on any other day of the ten:

Name in the user agent Requests Got a 200
CCBot 19 0
PerplexityBot 16 0
Google-Extended 15 0
MistralAI-User 10 0
DuckAssistBot 6 0

Every one of those requests got a redirect or an error. For comparison, ClaudeBot made 44 requests to the same site that day and 10 of them got a 200.

One of those names can't be real

Google's list of common crawlers is explicit: Google-Extended has no user agent string of its own. It's a robots.txt token Google uses to control how content is used, and the crawling is done under Google's ordinary user agents. On September 24, 2026, nextdayvideos.com logged 15 requests whose user agent claimed to be Google-Extended, a name Google says it never sends as a user agent, and none of them got a 200 response. Whoever sent those, by Google's own documentation it wasn't Google.

The rest might be real, and I can't tell

On that same day, nextdayvideos.com also logged 66 requests claiming to be CCBot, PerplexityBot, MistralAI-User or DuckAssistBot, none of those names appeared there on any other day from September 17 to 26, and none of the 66 got a 200 response. Showing up together, on one day, with nothing but redirects and errors, looks more like one scanner wearing several name tags than four crawlers that happened to coincide. But that's a guess. Without source addresses I can't prove it either way.

Two quick answers

Which rare AI crawler names appeared on nextdayvideos.com on September 24, 2026?

On that same day, nextdayvideos.com also logged 66 requests claiming to be CCBot, PerplexityBot, MistralAI-User or DuckAssistBot, none of those names appeared there on any other day from September 17 to 26, and none of the 66 got a 200 response.

Did Google-Extended show up as a user agent on nextdayvideos.com?

On September 24, 2026, nextdayvideos.com logged 15 requests whose user agent claimed to be Google-Extended, a name Google says it never sends as a user agent, and none of them got a 200 response.

How to check your own

  • Treat a crawler name as a claim. Before blocking, allowing or celebrating a bot, check the source address. Google explains how to verify a request really came from Google, and OpenAI's crawler overview links the IP lists for its agents.
  • Watch for names that only exist in robots.txt. Google-Extended is one. A request carrying it is a free tell.
  • Look at the status codes. A crawler that only ever gets redirects and errors is either working from a stale list or probing for something.

This post is also half of a small test I'm running on whether structured data changes what AI answer engines repeat about a page. I'll write up what happened once there's enough to say.

If you'd like help reading your own site's logs and deciding what to do about who shows up, that's what my one-on-one AI guidance sessions are for. Let my Claude help your Claude. Book a free 30-minute consult.

Filed under