Glossary

robots.txt

robots.txt is a file at the root of your site that tells search and AI bots which parts they may read. It controls reading, not whether a page appears.

By SupaVisible Editorial

robots.txt is a plain text file at yourdomain.com/robots.txt that tells bots (the programs search engines and AI tools use to read the web) which parts of your site they may read. Google says it is used mainly to manage how much bots fetch from your site.

It does not hide pages. Google's own guide says robots.txt "is not a mechanism for keeping a web page out of Google": a blocked page can still appear in results if other sites link to it. To keep a page out, use noindex.

A robots.txt rule left over from a staging site is a common reason a new site never appears. URL Inspection in Search Console shows whether one is blocking a page; see how long until a new site shows on Google.

The file also decides which AI bots can read your site. Search bots such as OAI-SearchBot fetch pages to quote in answers; training bots such as GPTBot collect training data. More in answer engine optimization.

Quick answers

Does robots.txt stop a page appearing in Google?
No. Google says robots.txt is not a way to keep a page out of Google: a blocked page can still be indexed if other sites link to it. To keep a page out, use a noindex tag and leave the page readable so Google can see it.

Try it on your own site

Find the searches you almost rank for.

Get in touch
All terms