{"slug":"robots-txt-policy","name":"Robots.txt Policy Writer","version":"1.0.0","updated_at":"2026-10-08T10:45:10.883Z","use_when":"Turns a crawler policy written in words into a correct robots.txt, answered as the file only. Knows the mistakes that make a file do the opposite of what was meant - a crawler obeys only the one group that names it, so a GPTBot group silently drops every rule written for the star group and the private paths must be repeated; longest match wins and Allow wins a tie; Google ignores Crawl-delay; Host is a Yandex-only line; Sitemap must be an absolute address; Google-Extended is a control token for Gemini training and does not touch Google Search; training, search and user-triggered AI bots are separate tokens. Says in a NOTE comment what robots.txt cannot do, such as removing a page from Google. Use when asked to write, check or fix a robots.txt, to allow or block search engines, AI crawlers or AI training bots, or to keep private paths out of crawlers.","not_for":"Checking a live site's file or its indexing in Search Console, removing pages from Google, or blocking bad bots by IP. It writes the file from the policy you describe; it cannot see what the site serves today.","languages":["any"],"tags":["robots-txt","crawlers","ai-bots","seo","indexing","sitemap"],"category":"seo","category_url":"https://aiskills402.com/categories/seo","keywords":["crawler policy","AI training bots","Google-Extended"],"faq":[{"q":"Is the reply only the file, or do I get an explanation too?","a":"Only the robots.txt, ready to save at the root of the host. Where your policy asks for something a robots.txt cannot do, such as taking a page out of Google, the file carries a NOTE comment that says so and names the right tool."},{"q":"What is the most common mistake it prevents?","a":"Giving one bot its own group. A crawler follows only the group that names it, so a new GPTBot group quietly cancels every rule written for all crawlers, and the private folders become open to that bot. The skill repeats the private paths in each named group."},{"q":"Can I block AI training but stay in AI search?","a":"Yes. Training, search and user-triggered bots use separate names, so GPTBot can be blocked while OAI-SearchBot stays allowed. Google-Extended is only a control token for Gemini training and does not change Google Search; blocking Googlebot would."},{"q":"Is it worth buying for a strong model?","a":"Only a little. Across 22 policies Sonnet got 21 files right unaided and 22 with the skill; the gain was one Host line. Haiku went from 18 to 20 and still wrote a Host line once. The skill helps most a smaller model, or an agent that should explain in comments what robots.txt cannot do."}],"examples":[{"lang":"en","model":"claude-sonnet-5-5","input_excerpt":"Every crawler may crawl the whole site except /admin/ and /account/. In addition GPTBot must stay out of /pricing/.","output_excerpt":"User-agent: *\nDisallow: /admin/\nDisallow: /account/\n\nUser-agent: GPTBot\nDisallow: /admin/\nDisallow: /account/\nDisallow: /pricing/"}],"page_url":"https://aiskills402.com/skills/robots-txt-policy","markdown_url":"https://aiskills402.com/skills/robots-txt-policy.md","image_url":"https://cdn.aiskills402.com/og/skills/robots-txt-policy/eea77783.png","related_url":"https://api.aiskills402.com/v1/skills/robots-txt-policy/related","purchases_count":null,"tested":{"date":"2026-10-08","strong":{"model":"claude-sonnet-5-5 (Claude Code alias \"sonnet\")","verdict":"Wrote a correct robots.txt in all 22 policies: the private paths repeated in every group that names a bot, GPTBot blocked while OAI-SearchBot stays allowed, Google-Extended blocked without touching Googlebot, no Host line, absolute sitemaps, longest-match exceptions for a search folder, PDFs and WordPress, and a comment instead of a pretend rule when the policy asked to remove a page from Google, to slow Google down, or to cover a staging host and a Cloudflare setting."},"weak":{"model":"claude-haiku-5-5 (Claude Code alias \"haiku\")","verdict":"Got 20 of 22 right. It wrote a Host line when the policy asked for a preferred host, and for a page the policy asked to block but keep out of Google results it left the page open and explained noindex in a comment instead of writing the Disallow. The group rule, the AI bot tokens, the sitemaps and the pattern cases were right."},"note":"Twenty-two crawler policies written by us, all in English: about two thirds with a trap (a bot group that must repeat the private paths, training blocked while AI search stays open, Google-Extended against Googlebot, a Host request, a relative sitemap, a request to remove a page from Google, a staging host, longest-match exceptions) and the rest plain policies where the obvious file is right. Each answer was parsed as a robots.txt by the RFC 9309 rules (the most specific group replaces the star group, longest match wins, Allow wins a tie, * and $) and every named bot was asked whether it may fetch the stated paths; sitemaps had to be absolute, a Host line was refused, and cases that ask for something robots.txt cannot do also needed a comment that says so. The rules and bot names in the skill were read on the vendors' own pages on 8 October 2026. One run per model and policy.","baseline":{"date":"2026-10-08","rows":[{"label":"Files that do what the policy says (22 policies)","better":"higher","strong":{"with":{"n":22,"of":22},"without":{"n":21,"of":22}},"weak":{"with":{"n":20,"of":22},"without":{"n":18,"of":22}}}],"note":"The same request on both sides, asking for the robots.txt only; a fence around the answer is removed first, and the comment check looks for a comment that says the thing, not a fixed label. Sonnet already writes nearly every policy right alone, group rule included, so its measurable gain is one line: without the skill it wrote a Host line in the one case that asked for a preferred host. Haiku without the skill also wrote Host, blocked a page the policy wanted removed from Google, opened the private folder to the user-triggered agents, and gave no noindex comment. With the skill the first two were fixed and the comment was written, yet it wrote Host again and left one page open that the policy asked to block. A strong model gains little here; the smaller one gains more."},"report_url":null},"price_usd":"0.03","price_micro":30000,"size_bytes":8188,"sha256":"53cad17f2c6a5ec28656fcac5ae8e11c5a80e99920ef5b364aeeb7aa61ff4bdf","outline":["The answer","The rules that decide what a file means","AI crawlers: three purposes, separate tokens","A setting outside the file","Work in this order","Short examples"],"license":{"summary":"Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing","holder":"Georgi Kalchev, aiskills402.com","url":"https://aiskills402.com/docs#license"},"buy_url":"https://api.aiskills402.com/v1/skills/robots-txt-policy/file","redownload_url_template":"https://api.aiskills402.com/v1/purchases/{token}","mcp_tool":null,"payment":{"protocol":"x402","scheme":"exact","asset":"USDC","selling":true,"network":"base","network_caip2":"eip155:8453","pay_to":"0x8e37022edcf0f21cf3c9f93fee9d4d32519f36f4","facilitator":"cdp"},"seo_title":"Robots.txt Policy Writer for AI Agents","seo_description":"Turn a crawler policy in words into a correct robots.txt: AI bots, search engines, private paths, sitemap. Notes what a file cannot do. $0.03 once.","versions":[{"version":"1.0.0","date":"2026-10-08","changelog":"# Changelog\n\n## 1.0.0 — 2026-10-08\n\nFirst release: turns a crawler policy in words into a robots.txt, answered as the file only. Covers the group rule (a crawler obeys only the group that names it, so private paths are repeated in every named group), longest match with Allow winning ties, wildcards, Google ignoring Crawl-delay, no Host line, absolute Sitemap addresses, one file per host, the split between blocking crawling and indexing (a NOTE comment instead of a fake fix), AI crawlers by purpose (training, search, user-triggered) with Google-Extended as a control token, and a check of Cloudflare's managed robots.txt setting.\n"}]}