Contents

09/24/2026

Who's Crawling Us? An AI Bot Census of Five Sites' Logs

Everybody has an opinion about AI crawlers. Very few of those opinions come with a log file attached. So I pulled one week of raw Apache access logs from the five sites I run and counted who actually showed up: ClaudeBot, GPTBot, Meta's crawler, Bytespider and the rest. I also checked whether they read robots.txt, and whether any of them fetched llms.txt.

The five are belchamber.us, brandager.com, nextdayvideos.com, prolificfutility.com and this site, tools.belchamber.us. They're all small static sites on the same DreamHost plan. The window runs from 00:00 on September 17, 2026 to about 03:00 on September 24, Pacific time, the server's own clock. That's roughly seven days and three hours, which is all the rotation had kept.

The counting was done by Claude Code over a read-only SSH session while I directed it. Nothing on the server was changed. The method is below, and so are the two ways it went wrong before it went right.

Two ways it went wrong

  • Gotcha 1: the log folder that looks right and isn't

    The live logs are under https. If your numbers look oddly quiet, check the dates on the files before you trust the counts.

  • Gotcha 2: access.log.0 counts a day twice

    A glob like access.log* reads yesterday twice. The fix is to read regular files only (find -type f), which skips the symlink.

The short version

  • Named AI crawlers made about 5.4 times as many requests as Googlebot and Bingbot combined: 5,819 against 1,082, across all five sites.
  • Meta's crawler was the busiest AI agent by a wide margin, with 1,801 requests on this site alone. Of those, 736 hit URLs from the WordPress era that now answer with a 301 redirect.
  • Nobody broke a robots.txt rule. That's a weaker result than it sounds, because my robots.txt files barely say anything. More on that below.
  • No named AI crawler fetched llms.txt. It was requested 330 times, and most of those requests look like my own monitoring and deploy checks.

Gotcha 1: the log folder that looks right and isn't

On DreamHost, each domain has a folder under ~/logs/<domain>/ holding two symlinks, http and https. The http link points at a real folder full of real, rotated, gzipped access logs. For these five sites, though, its newest file was anywhere from early September back to 2024. Every site now answers over HTTPS, so the plain-HTTP folder stopped getting traffic, and nothing tells you it's stale.

The live logs are under https. If your numbers look oddly quiet, check the dates on the files before you trust the counts.

Gotcha 2: access.log.0 counts a day twice

Inside the https folder, access.log is today's file, older days are access.log.YYYY-MM-DD (gzipped after a day or two), and access.log.0 is a symlink to yesterday's file. A glob like access.log* reads yesterday twice.

My first pass fell right into it. It reported 18,684 lines for belchamber.us, and the correct figure is 16,077. The gap is almost exactly one extra copy of the September 23 file. Every per-bot count was inflated by that day's traffic too. The fix is to read regular files only (find -type f), which skips the symlink.

Method, and what it can't tell you

The logs are in Apache's combined log format, where the user agent is the third quoted field. Splitting each line on " puts the user agent in $6 and the request line in $2. A request counts for a bot when the bot's name appears anywhere in the user-agent field, compared case-insensitively.

Caveats, stated plainly:

  • A user agent is a claim, not an identity. Anyone can send ClaudeBot in a header. I didn't verify source addresses against the vendors' published ranges. OpenAI's crawler documentation and Perplexity's crawler documentation both publish IP lists, and Google explains how to verify that a request really came from Google. If a decision depends on these numbers, do that step.
  • Requests, not pages. A crawler fetching a page, its redirect target and its images shows up as several requests.
  • One week is a small sample. Within a minute of my extraction closing, GPTBot made 44 requests to this site. All week it had made 24.
  • An escaped quote inside a request line can throw off the field split. That's rare, but it happens.

As a cross-check, the first pass matched names anywhere on the line, which included URLs and referrers. After I removed the duplicated day, the two methods agreed to within a few hits per bot. So the substring shortcut wasn't the problem. The symlink was. A second, independent tally from my Site Ops dashboard, which stores a summary per site per day and so stops at the end of September 23, comes within about 1% of these totals for every bot it tracks.

The numbers

Requests by user agent, from 00:00 on September 17 to about 03:00 on September 24, 2026, Pacific time:

Agent belchamber.us brandager.com nextdayvideos.com prolificfutility.com tools.belchamber.us Total
meta-externalagent (Meta) 78 211 80 16 1,801 2,186
ClaudeBot (Anthropic) 179 192 238 139 394 1,142
Bytespider (ByteDance) 109 213 219 11 561 1,113
OAI-SearchBot (OpenAI) 72 85 42 42 137 378
Amazonbot 26 51 15 3 266 361
PerplexityBot 0 0 0 0 262 262
GPTBot (OpenAI) 32 80 70 20 24 226
YouBot (You.com) 10 66 6 0 39 121
ChatGPT-User 2 6 1 0 10 19
Claude-User 0 2 0 0 8 10
DuckAssistBot 0 1 0 0 0 1
For comparison: Googlebot 102 161 69 38 220 590
For comparison: bingbot 61 188 119 5 119 492
For comparison: PetalBot (Petal Search) 332 604 211 38 1,551 2,736
All requests 16,077 17,323 12,555 5,358 71,715 123,028

A few things stand out.

Meta is busy, but a lot of it is history. Meta's crawler documentation says meta-externalagent crawls for uses "such as training foundation AI models." On this site, 736 of its 1,801 requests landed on a 301: old WordPress tag archives and image attachment pages that the static rebuild redirects. It's working through a list of URLs that stopped existing weeks ago. So "Meta crawled us 1,801 times" and "Meta read 1,801 pages" are two very different claims.

PerplexityBot only visited one site. It made 262 requests here and none on the other four.

The user-triggered agents are the interesting small numbers. ChatGPT-User and Claude-User fetch a page because a person asked an assistant something (Anthropic describes Claude-User this way). Together they made 29 requests all week. That's small, but it's the closest thing in these logs to a human reading via AI.

PetalBot out-crawled everyone. It isn't an AI crawler by name. I'm including it because it's a reminder that "who is hammering my site" and "which AI is reading my site" are separate questions.

robots.txt: who read it, and who obeyed it

First, what the files say. All five robots.txt files allow everything. Three also disallow leftover paths: /_site/ on nextdayvideos.com, /_site/ and /cms/ on prolificfutility.com, and /cms/ here. None of them names an AI crawler. According to each site's git history, none of the files changed during the window. The robots.txt standard itself is RFC 9309.

Agent robots.txt fetches (all 5 sites) Requests to a disallowed path
ClaudeBot 345 0
OAI-SearchBot 270 0
facebookexternalhit 110 0
YouBot 92 0
Googlebot 70 0
bingbot 69 0
Bytespider 57 0
PerplexityBot 4 0
Claude-User 4 0
GPTBot 0 0
meta-externalagent 0 0
Amazonbot 0 0
ChatGPT-User 0 0

Zero violations, and I wouldn't lean on that. With only three narrow Disallow rules on paths nothing links to, a crawler had almost no reason to go there. Across the week, 208 requests hit a disallowed path. None came from a named AI crawler. The biggest single source was WordPress.com's Photon image service, still asking for images from an install that no longer exists. A real compliance test needs a rule that a crawler actually wants to break. I'll run that as a follow-up, with a Disallow aimed at one named bot on a path it's already crawling.

Some bots mostly just check the rules. On three of the five sites, every OAI-SearchBot request was for robots.txt: 85 of 85, 42 of 42 and 42 of 42. ClaudeBot re-read each site's robots.txt about ten times a day.

Three busy crawlers never asked for robots.txt under their own names: GPTBot, meta-externalagent and Amazonbot. That doesn't prove they ignore it. A company can fetch robots.txt under another agent name or cache it for longer than a week. (One thing it isn't: OpenAI says it may add a robots.txt marker to a crawler's user agent when fetching the file, but a GPTBot request marked that way still contains GPTBot, so it would have been counted here.) But if you're relying on a rule for one of these, check your logs and don't assume.

Remember that user-triggered fetchers play by different rules. OpenAI says that for ChatGPT-User, "robots.txt rules may not apply," and Perplexity says Perplexity-User "generally ignores robots.txt rules." Amazon's Amazonbot page says something similar for Amzn-User. Blocking the training crawler doesn't block the assistant a reader is using.

llms.txt: who actually fetched it

All five sites publish an llms.txt. It was requested 330 times during the week, and not once by a named AI crawler.

Where the requests came from:

  • 207 from python-requests/2.34.2. That's the same library version my Site Ops health checks run with, and those checks request robots.txt and llms.txt from every site. I'm confident most of these are me checking on myself, but I haven't matched them request by request.
  • 60 from a user agent that's just node. These are most likely mine too. After every deploy, my sites' release checks fetch each site's llms.txt with Node's built-in fetch, which identifies itself as exactly node, and the sites were deployed many times that week. As with the Python requests, I haven't matched them one by one.
  • 23 from Lighthouse, which reads llms.txt in its newer agent-readiness audit. I wrote that audit up in Lighthouse now grades your site for AI agents.
  • The rest were a scattering of SEO tools and small crawlers.

If you added llms.txt hoping ClaudeBot or GPTBot would read it, one week on five sites says they don't, at least not under those names. It costs nothing to publish, so I'm keeping it. I just wouldn't expect the big crawlers to read it yet.

A reusable census script

This is a tidied version of the script behind the per-bot counts. Its list covers the main agents; add YouBot, DuckAssistBot, facebookexternalhit or any other name you want counted. It's read-only: it takes a log directory, reads regular files only (skipping the access.log.0 symlink) and uses zcat -f so gzipped and plain files work the same way.

What the script does

  1. Takes a log directoryThe live logs are under https
  2. Reads regular files onlySkipping the access.log.0 symlink
  3. zcat -fGzipped and plain files work the same way
  4. Splitting each line on "Puts the user agent in $6 and the request line in $2
  5. A request counts for a botWhen the bot's name appears anywhere in the user-agent field
#!/usr/bin/env bash
# bot-census.sh -- count AI and search crawler requests in Apache combined logs.
# Usage: bot-census.sh ~/logs/example.com/https
# The trailing slash on "$dir/" lets find descend a symlinked log directory.
dir="${1:?usage: bot-census.sh LOG_DIR}"
find "$dir/" -maxdepth 1 -type f -name 'access.log*' -print0 |
  xargs -0 zcat -f |
  awk -F'"' '
    BEGIN {
      n = split("ClaudeBot Claude-User Claude-SearchBot GPTBot OAI-SearchBot ChatGPT-User " \
                "PerplexityBot Perplexity-User meta-externalagent Bytespider Amazonbot " \
                "PetalBot Googlebot bingbot", bots, " ")
    }
    {
      ua = tolower($6)            # user agent: third quoted field
      split($2, req, " ")         # "GET /path HTTP/1.1"
      path = req[2]
      for (i = 1; i <= n; i++) {
        if (index(ua, tolower(bots[i]))) {
          hits[i]++
          if (path == "/robots.txt") robots[i]++
          if (path ~ /^\/llms(-full)?\.txt/) llms[i]++
          break
        }
      }
    }
    END {
      printf "%-20s %8s %8s %8s\n", "agent", "hits", "robots", "llms"
      for (i = 1; i <= n; i++)
        printf "%-20s %8d %8d %8d\n", bots[i], hits[i], robots[i], llms[i]
    }'

To check compliance, add one line per Disallow prefix inside the loop, for example if (path ~ /^\/cms\//) blocked[i]++, and print it with the rest. Run it against a single day's file first and compare the result to a quick grep -ci claudebot on the same file. My first pass double-counted a day, so review AI-written scripts like this one before you trust their numbers.

What I'd do with this

  • Decide per bot, on purpose. Training crawlers, search crawlers and user-triggered fetchers are separate agents with separate robots.txt names. Each vendor's own page lists them.
  • Look at where the crawlers land, not just the totals. A big count can be mostly redirects from a URL list that went stale months ago.
  • Keep the logs longer than a week. One burst from a single crawler can change the story, and DreamHost's rotation won't keep the history for you.

Want this for your own sites?

This whole census took one conversation: I asked, Claude Code did the digging over a read-only connection, and I checked the numbers. If you run a business website and want to know who's really reading it, human or bot, and what to do about it, that's the kind of thing my one-on-one AI guidance sessions cover. Let my Claude help your Claude. Book a free 30-minute consult.

Filed under