FoundGPTFoundGPT
Published 2026-08-09

We Generate llms.txt for 100+ Stores. Here's the Honest Truth About What It Does.

FoundGPT ships llms.txt — and our own 207k-citation dataset says AI engines cited it zero times. The honest truth about what llms.txt does, and why we still generate it.

FoundGPT — AI visibility for Shopify stores

FoundGPT ships llms.txt files. So you'd expect us to tell you llms.txt is the secret to AI visibility. Our own 207,297-citation dataset says otherwise — and we think telling you that matters more than the marketing.


llms.txt is a proposed standard: a Markdown file at your domain root that gives AI systems a clean, curated summary of your site — your products, your policies, your story — without the HTML noise. We generate one for every store that installs FoundGPT, and over a hundred Shopify stores in our fleet have one live right now. (New to it? Here's what llms.txt is and how to generate one.)

So this is us auditing our own product feature in public.

What our data says

We track which URLs AI engines actually cite when answering shopping questions. Across 207,297 citations collected for 107 stores over four months (April–August 2026):

  • Citations of an llms.txt file: 0
  • Citations of any .md URL: 0
  • Citations of /products/ pages: 11,942
  • Citations of blog articles: 8,000
  • Citations of /collections/ pages: 3,679

Not one engine — ChatGPT, Gemini, Perplexity, Google AI Mode, Claude — cited an llms.txt or Markdown URL even once. Meanwhile they cited ordinary HTML product pages twelve thousand times.

We're not alone in finding this. OtterlyAI ran controlled experiments on their own site: their llms.txt received 0.1% of AI bot traffic over 90 days, and their 14-day Markdown-mirror experiment recorded zero AI crawler visits and zero citations to .md versions of pages while the HTML twins collected both. Different site, different vertical, same shape: AI search runs on standard web infrastructure. Answers cite HTML.

So why do we still generate llms.txt?

Fair question. Four honest reasons:

  1. The cost is near zero. Generation is automatic, the file is static, and it can't hurt you. This is not true of most GEO tactics.
  2. It's insurance on an emerging standard. llms.txt is young. If any major engine starts consuming it for grounding, the stores that already have one are ahead by a deploy cycle. We'd rather our merchants hold a free option than not.
  3. On-demand agents are not crawlers. Citation data measures answer engines' retrieval. Purpose-built agents (a user's personal shopping agent, a developer's tool) can be pointed at llms.txt and read it happily — a real, if currently niche, consumer.
  4. The curation is valuable even if the file isn't. Building a store's llms.txt forces the question "what are this store's actual facts?" — the same distillation that improves the surfaces engines do read.

What we won't do is tell you llms.txt is why you'll show up in ChatGPT. It isn't. Our own numbers say so.

Where the citations actually come from

If 0 citations went to llms.txt, where did the 207k go? To the boring, load-bearing web:

  • Retailer and brand product pages — the majority on every engine
  • Buying guides and blog articles — 8,000 citations to blog paths; comparison content on your own domain works
  • Editorial roundups, YouTube reviews, specialist magazines — the earned-media layer

The uncomfortable, liberating conclusion of every experiment we and others have run: there is no side door. No special file, no alternative format, no metadata trick substitutes for HTML pages that deserve to be cited — structured data that's correct, descriptions with substance, guides that answer real questions.

That's also why it's winnable. The stores beating Amazon in their niches inside ChatGPT answers (we've measured it — by 10x) did it with excellent ordinary pages, not exotic files.

The honest checklist

  1. Have an llms.txt (free insurance) — but spend zero ongoing effort on it.
  2. Spend that effort where citations live: product page schema, description depth, buying guides.
  3. Distrust any GEO pitch built on alternative file formats. Ask for citation evidence.
  4. Measure outcomes, not artifacts: mentions, citations, and AI-referred sessions — not "files deployed."

Where does your store stand on the things that measurably matter? foundgpt.app/scan — free, 14 checks, no login.


Methodology: URL-pattern analysis over 207,297 citation instances (April 19 – August 9, 2026, 107 Shopify stores, 3,601 buyer prompts, five engines). External results referenced: OtterlyAI's llms.txt (90-day) and Markdown-vs-HTML (14-day) experiments, which reached the same conclusion on non-ecommerce data. We'll publish our own controlled .md experiment next — running it now.


About the author

Rahul — Founder, FoundGPT

Rahul built FoundGPT and ran the analysis behind this study. Across 111 Shopify stores he has run 23,000+ real AI-search checks on ChatGPT, Gemini, Google AI Mode and Perplexity, and applied ~99,000 structured-data fixes.

About FoundGPT

Is your store AI-ready?

Run a free scan of the structured data ChatGPT, Gemini & Perplexity read to answer shoppers' questions — get your readiness score and the fixes. No login.

On Shopify? Get the app free →·★★★★★ 5.0 on the Shopify App Store