A robots.txt file publishes site-level instructions about which URL paths automated crawlers may request. It is a plain-text file at the origin root, such as https://example.com/robots.txt, and uses records that target crawler user agents.

The common Disallow and Allow fields describe path rules for a matching user agent. A Sitemap field can point to an XML sitemap. Before crawling a host, a crawler should fetch and parse the file for its declared user agent, apply the most specific matching group, and handle temporary fetch failures conservatively.

The file is a directive mechanism, not an access-control system. It does not stop a client from requesting a URL, make private material safe to publish, or grant permission to collect content that other rules prohibit. A disallowed URL can still be discovered through links or indexed from outside references. Sensitive resources need authentication and server-side authorization rather than a robots rule.

Crawler implementations differ in how they handle nonstandard fields and unusual status responses. Following the documented standard and respecting site terms reduces ambiguity. Applicable law and contractual restrictions remain separate from robots.txt.

Related terms: web crawler and web scraping.

Browse all glossary terms

Use residential proxies with your own workflow.