Shopify AI crawlers: search access is not training consent
Understand OAI-SearchBot, GPTBot, Google Search and Shopify Catalog controls before editing robots.txt for AI discovery.
- PUBLISHED
- FILED UNDER
- GEO
- READING TIME
- 3 minutes
- WORDS BY
- Shopify GEO SEO Agency

‘Allow AI’ is too vague to be a useful setting. A merchant may want shoppers to discover products through an assistant while making a different decision about training access. Begin by writing down that intention. Otherwise a broad block or allow rule can achieve the opposite of what the business wanted.
Identify the purpose of each control
OpenAI identifies OAI-SearchBot as a search crawler and GPTBot as a crawler for potential model-training use. Those choices are independent. ChatGPT-User handles certain user-initiated actions and is not the control for automatic search crawling; robots.txt rules may not apply to those user-triggered requests. OpenAI: crawler purposes and controls ↗
| Surface | Question to answer |
|---|---|
| Open-web search crawling | Can a search system retrieve the public pages we want discovered? |
| Training crawling | What is the business's policy on this distinct use? |
| User-triggered retrieval | Can a user-requested visit reach the useful public information? |
| Product-catalogue distribution | Which enabled channels receive our submitted catalogue data? |
Keep the decision record short: provider, intended use, desired access, relevant documentation, owner and review date. Use the provider's actual crawler names rather than a guessed umbrella name. Revisit the record when the service changes its documentation or when you add a new sales channel.
Inspect Shopify's generated output first
Open /robots.txt on the live storefront and compare it with any robots.txt.liquid customisation in the published theme. Identify old snippets, broad rules and specific user-agent groups. Preserve generated defaults rather than replacing the entire file with a static list copied from another store.
Shopify explains that product syndication through Shopify Catalog is independent of robots.txt. Blocking an open-web crawler does not switch off an activated catalogue channel. Use the Agentic channel settings for catalogue distribution decisions. Shopify: editing robots.txt.liquid ↗
For that separate catalogue route, read how Shopify product discovery works in ChatGPT.
Test public information, not private areas
- List the product, collection and guide URLs intended for discovery.
- Evaluate the relevant crawler rules against that sample, including market subfolders.
- Check that public pages return useful content without an unintended login or challenge.
- Keep account, cart and checkout controls appropriate to their purpose; discovery does not require exposing private information.
- Save the previous template and record the exact change so a problem can be reversed.
On a hosted Shopify storefront, work within supported platform controls. If you use a headless storefront with your own hosting or bot protection, inspect that additional layer too. A permissive robots file does not help if the infrastructure returns a challenge instead of the page.
Keep Google Search controls distinct
Google says Googlebot access governs its Search AI features. Its Google-Extended control concerns other specified uses; it is not a separate switch for appearing in AI Overviews. Follow the Search documentation for indexing and snippet controls. Google: AI features and your website ↗
Confirm access without promising visibility
After publishing, retrieve the generated file again and test the intended paths. Record when the change went live; a provider may take time to refresh its rules or revisit a page. Then monitor real retrieval, indexed pages and observed referrals where available. A crawler visit demonstrates access. It does not establish that the product was recommended, that the answer was accurate or that a customer bought anything.
Put this into practice with Shopify SEO and Shopify GEO.