{"slug":"crawler-access-probe","name":"AI Crawler Access Reader","version":"1.0.0","updated_at":"2026-10-08T14:30:21.379Z","use_when":"Reads the results of probing a website with AI-crawler user agents (status codes per bot, a control request, sometimes robots.txt or the site's settings) and says what they really show as JSON - whether training bots, search bots and user-triggered agents are open, blocked, mixed or unknown, what is behind a refusal, and whether the robots.txt that answered is the site's own file. Knows the traps - a bare bot name proves nothing and lies differently for each bot, a blocked training bot is a choice while a blocked search bot is a defect, Python-urllib and libwww-perl are refused by Browser Integrity Check and not by the site's policy, since 15 September 2026 Cloudflare refuses training bots on zones nobody touched, and a robots.txt of only comments is a stand-in from the edge. Use when asked whether a site blocks ChatGPT, Claude, Perplexity or Google AI, how to read curl results with bot user agents, or why an AI crawler gets a 403.","not_for":"Running the probe or fetching any site: it reads only the results you paste. Not for writing a robots.txt file, which is another skill, and not for judging content quality, rankings or whether a bot ought to be allowed. It does not read your Cloudflare settings, because an API token cannot; it judges by how the site actually answered.","languages":["any"],"tags":["ai-crawlers","gptbot","claudebot","cloudflare","robots-txt","technical-seo","geo"],"category":"seo","category_url":"https://aiskills402.com/categories/seo","keywords":["Browser Integrity Check","Python-urllib","Perplexity"],"faq":[{"q":"What comes back for a set of probe results?","a":"A single JSON object with fixed fields that code can switch on: training, search and agent, each open, blocked, mixed or unknown; a cause such as cloudflare_ai_setting or browser_integrity_check; whether the robots.txt that answered is the site's own file or a stand-in; and a short reason in plain words. Mixed results, an unprobed bot, or a failed control are reported as mixed or unknown instead of guessed, which is what keeps an audit from publishing a claim the numbers never supported."},{"q":"Why not simply read the status codes myself?","a":"Because a probe with a bare bot name lies, and lies differently for each bot. On one site bare ClaudeBot got 403 and bare GPTBot got 200, which looked like a policy against Anthropic; with full user agents both were refused alike. The skill discards such results, insists on a control request, and pairs each company's training bot with its search bot, since a refused search bot means lost visibility while a refused training bot is usually a choice. Statuses alone cannot show this."},{"q":"Does it know about the Cloudflare change of September 2026?","a":"Yes. On 15 September 2026 Cloudflare moved zones to separate Search, Training and Agent settings, so training bots receive 403 on sites nobody touched. The skill tells that default apart from a rule the owner wrote, and from Browser Integrity Check, which turns away Python-urllib and libwww-perl with an HTML page and leaves no line in the Worker log. It also knows Google-Extended cannot be probed at all. Whoever audits many sites for visibility to assistants gets the same reading every time."},{"q":"Does it help Claude Sonnet?","a":"Somewhat, on specific points. Over 23 probe situations Sonnet got 19 right alone and all 23 with this file; Haiku rose from 17 to 21. The real gains: Sonnet treated bare bot names as proof, blamed an owner rule on an untouched site, and marked a Google-Extended probe open. Pairs, Browser Integrity Check and robots.txt stand-ins it already handled. The facts were read on 8 October 2026; Anthropic publishes only tokens, so copy Claude strings from your logs."}],"examples":[{"lang":"en","model":"claude-sonnet-5-5","input_excerpt":"We tested https://shop.example/ with curl, one request per user agent, all within one minute. -A \"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/141.0.0.0 Safari/537.36\" gave 200. -A \"GPTBot\" gave 200. -A \"ClaudeBot\" gave 403. Does the site treat OpenAI's and Anthropic's training crawlers differently?","output_excerpt":"{\"training\": \"unknown\", \"search\": \"unknown\", \"agent\": \"unknown\", \"cause\": \"probe_error\", \"robots_txt\": \"unknown\", \"why\": \"The control returned 200, but GPTBot and ClaudeBot were sent as bare names, and Cloudflare matches each bot by its full user agent string, so a 200 for one and a 403 for the other shows nothing about how the site treats OpenAI versus Anthropic.…"}],"page_url":"https://aiskills402.com/skills/crawler-access-probe","markdown_url":"https://aiskills402.com/skills/crawler-access-probe.md","image_url":"https://cdn.aiskills402.com/og/skills/crawler-access-probe/16728e6e.png","related_url":"https://api.aiskills402.com/v1/skills/crawler-access-probe/related","purchases_count":null,"tested":{"date":"2026-10-08","strong":{"model":"claude-sonnet-5-5 (Claude Code alias \"sonnet\")","verdict":"Read all 23 probe situations correctly: bare bot names were set aside as no evidence (training unknown, cause probe error) and a control request that was refused too made every result unknown; the full GPTBot refusal beat the bare 200; the refusal of a training bot on an untouched zone was put down to the Cloudflare default, and an agent refused only on pages with ads was called mixed; a refused search bot was named as a defect; Python-urllib and the Perl and Java clients were put down to Browser Integrity Check, and a probe of Google-Extended was left unknown; a comments-only robots.txt with the 404's headers was called a stand-in, and rules in the live file that the repository lacks were traced to a Cloudflare setting."},"weak":{"model":"claude-haiku-5-5 (Claude Code alias \"haiku\")","verdict":"21 of 23 with the skill, up from 17, and the same readings as Sonnet on bare names, the failed control, the ads pages and Google-Extended. But one answer was cut off and is not valid JSON (the mixed training case), and for the rules injected above a robots.txt it said the cause was probably a Cloudflare setting in its text yet wrote \"unknown\" in the field."},"note":"Twenty-three probe situations written by us the way an auditor pastes curl results: 15 where the obvious reading is wrong and 8 where it is right, so any harm from the skill would show. The answer is JSON with three bot kinds, a cause and a robots.txt state, scored on the fields that have one defensible answer. The facts come from the vendors' crawler pages read on 8 October 2026 and from our own measurements on live sites in September 2026, named in the skill's notes. One run per model and situation.","baseline":{"date":"2026-10-08","rows":[{"label":"Right reading of the crawler probe results (23 situations)","better":"higher","strong":{"with":{"n":23,"of":23},"without":{"n":19,"of":23}},"weak":{"with":{"n":21,"of":23},"without":{"n":17,"of":23}}}],"note":"The same request on both sides: the same fields with the same neutral one-line meanings. Read by hand, Sonnet's four misses without the skill are real. Two read bare bot names as evidence: it called training \"mixed\" because bare ClaudeBot got 403 and bare GPTBot got 200, and called training \"open\" because three bare names got 200. For a zone nobody had touched since August it named an owner rule, against the facts it was given. For a Google-Extended probe it wrote \"open\" while its own reason said the probe proves nothing; that one is a mislabelled field, not a gap in knowledge. Haiku missed the same four plus two more."},"report_url":null},"price_usd":"0.05","price_micro":50000,"size_bytes":9721,"sha256":"5dd2aaa4e1022d05705d3f59f2f775e0284bdb99b41198d070e6b2c1ac63250a","outline":["The answer","The three kinds of bot","What decides the reading","Three readings, worked through","Where the obvious reading is right","Work in this order"],"license":{"summary":"Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing","holder":"Georgi Kalchev, aiskills402.com","url":"https://aiskills402.com/docs#license"},"buy_url":"https://api.aiskills402.com/v1/skills/crawler-access-probe/file","redownload_url_template":"https://api.aiskills402.com/v1/purchases/{token}","mcp_tool":null,"payment":{"protocol":"x402","scheme":"exact","asset":"USDC","selling":true,"network":"base","network_caip2":"eip155:8453","pay_to":"0x8e37022edcf0f21cf3c9f93fee9d4d32519f36f4","facilitator":"cdp"},"seo_title":"AI Crawler Access Reader Skill","seo_description":"Paste curl results made with AI crawler user agents and get what they prove as JSON: training, search and agent bots, the cause, the robots.txt. Bought once.","versions":[{"version":"1.0.0","date":"2026-10-08","changelog":"# Changelog\n\n## 1.0.0 — 2026-10-08\n\nFirst release: reads the results of probing a site with AI-crawler user agents and returns JSON with the state of training, search and user-triggered bots (open, blocked, mixed or unknown), the cause of a refusal (a Cloudflare AI setting, Browser Integrity Check, a rule the owner wrote, or an error in the probe itself) and whether the robots.txt that answered is the site's own or a stand-in from the platform. Discards results made with bare bot names, requires a working control request, pairs each company's training and search bot, and knows that Google-Extended cannot be probed.\nPrice: 0.05 USD.\n"}]}