What is robots.txt
robots.txt is a plain text file at the root of your website that provides instructions to web crawlers about which parts of your site they should and shouldn't access.
Why it matters
- Express crawler preferences — publish path rules that conforming crawlers may follow
- Avoid accidental exposure assumptions — robots.txt is public and is not access control; sensitive content still needs authentication/authorisation
- Declare sitemaps — the
Sitemap:directive tells crawlers where to find your XML sitemap - Declare crawler-specific policy intent — some publishers name AI-related crawler products, but robots.txt cannot guarantee compliance, use or training outcomes
Basic format
User-agent: *
Allow: /
Disallow: /admin/
Sitemap: https://example.com/sitemap.xml
When to use it
Decide which crawler groups and paths should be crawlable. A broad Disallow can affect many URLs; a narrow rule should reflect the actual path and intended crawler.
A rule disallowing /admin/ asks matching crawlers not to crawl that path. Protect an administration area with proper access controls as well.
Location and responsibility
Use the root /robots.txt address and confirm which CMS, plugin or file serves it. robots.txt is public crawl guidance, not authentication or a way to keep confidential content private. Blocking crawling is also different from requesting noindex.
Refer to Robots Exclusion Protocol (RFC 9309) for the external specification or convention. Use Flowpane’s check details to understand what its own assessment covers; a product score is not a certification.