robots.txt
robots.txt is a file at a website's root that tells cooperating crawlers which parts of the site they may request.
In full
The robots.txt file sits at the root of a domain and states which paths a given crawler may or may not request, identified by user-agent. It is a convention rather than a security mechanism: well-behaved operators honour it, and one that ignores it will keep ignoring it. That distinction matters because robots.txt is frequently mistaken for access control. It is a request, and anything that genuinely must be prevented has to be enforced at the network or application layer. A robots.txt file also commonly declares the location of the site's XML sitemap.
Why it matters
A misconfigured robots.txt is one of the fastest ways to make a website disappear from search. A staging-site block that ships to production will deindex an entire site quietly.
A common misconception
That disallowing a page in robots.txt removes it from search results. Blocking crawling can prevent an engine from seeing a noindex directive, occasionally leaving a URL listed with no description. To remove a page, allow crawling and use noindex.