Robots.txt: what is this file for?
The robots.txt file is a text file placed at the root of a site that tells crawlers (Googlebot, Bingbot…) which areas they're allowed to explore and which they aren't.
Why this file matters for an online store
An e-commerce site often generates thousands of technical URLs: cart pages, checkout pages, customer account pages, or faceted filter results (colour, size, price). If crawlers explore all these URLs indiscriminately, they waste part of the "crawl budget" allocated to the site and can index duplicate content generated by the filters.
The robots.txt file directs crawlers to the pages that actually matter: categories, product pages, editorial content, while excluding the cart, the back office, or the parameters generated by filters.
A concrete example
The file is accessible at https://example.com/robots.txt and contains directives such as Disallow: /cart or Disallow: /admin, usually followed by a Sitemap: line pointing to the XML sitemap, which helps search engines discover the important pages.
Frequently asked questions
Is the robots.txt file enough on its own to stop a page being indexed?
Where should the robots.txt file be placed?
Should the back office be blocked in robots.txt?
Describe your need in one minute
A few targeted questions so I can reply with an estimate rather than another questionnaire.