AI crawlers spend real budget on dead URLs
Since 2026-07-31, one large content site in the measured fleet has returned 404 to 12,680 of 28,471 GPTBot requests. On this site, 109 of 135 ChatGPT-User requests did the same.
An answer engine can only cite a page it can fetch. Every 404 a crawler meets is a page it cannot read, index, or quote, and at volume dead URLs eat the crawl budget that should have reached canonical content. On one large content site in the fleet we measure, 12,680 of 28,471 GPTBot requests since 2026-07-31 have returned 404, with another 476 redirects on top — 46% of that crawler’s requests went nowhere.
This site is not exempt. In the same window, 109 of 135 ChatGPT-User requests to sorcerai.co returned 404, mostly requests for retired paths from before the rebuild that are still being probed weeks later. Crawler memory outlives redesigns.
The fix pattern is boring: enumerate the 404 sources, repair or redirect each one, confirm the canonical paths return 200 to AI user agents, then re-measure on the same request-level signal. Every number here is request-level first-party edge measurement with scanner probes excluded, window 2026-07-31 through 2026-08-07.