Product pages take 6% of the sources AI assistants cite on buyer queries, across the 40 Shopify stores we track, and 19.3% of brand-domain citations in AirOps' data. They are never the page that wins the recommendation. This post is about the moment your page does matter: when the assistant opens it to verify what it is about to recommend.
Most product pages fail that moment on mechanics, not marketing. The facts sit in a JavaScript widget no AI crawler executes, or in the bottom half of the page that 44.2% of citations never reach, or in a spec sheet shipped as a JPEG. The fixes are cheap, and the evidence says they pay: in Princeton's GEO study, modest content changes lifted visibility in generative-engine answers by up to 40%, with the largest gains going to sites ranked lower in classic search. For a small store, page mechanics are a levelling opportunity.
A citation, as throughout this series, means the source an assistant names behind its answer. The numbers this post is built on:
| What the data says | Number | Source |
|---|---|---|
| Major AI crawlers that execute JavaScript | 0 | Vercel, 500M+ fetches analysed |
| ChatGPT citations drawn from the first 30% of a page | 44.2% | Kevin Indig, 18,012 citations |
| Citation change after 1,885 pages added schema (vs 4,000 controls) | None | Ahrefs |
| Visibility lift from adding statistics, quotations and sourced claims | Up to 40% | Princeton, KDD 2024 |
| The typical ChatGPT-cited page | 941 words, 4 H2s, 15 links | Evertune, 40,000 URLs |
| Product pages' share of cited sources on buyer queries | 6% | CitoRank cohort, May 2026 |
Each number is unpacked, with method and caveats, in its section below.
Can the assistant read your page at all?
Start with the blunt finding. Vercel analysed how AI crawlers process the web, including more than 500 million GPTBot fetches, and found that none of the major AI crawlers render JavaScript. GPTBot, ClaudeBot and PerplexityBot download your script files; none of them run them. Anything your page draws in after load (a reviews widget, a tabbed spec panel, an app-injected size guide) does not exist for ChatGPT or Claude.
The exception is Google. Googlebot renders JavaScript, and Google's AI surfaces ride its index, so a JS-dependent page is visible to Gemini and AI Overviews while staying invisible to everyone else. That is the same per-engine split the engine teardown found in citation sources, now operating at the rendering layer.
For a standard Shopify store the default position is good: Liquid themes are server-rendered, so your catalogue ships as real HTML. The failure modes are the additions. Review apps that inject their content client-side put your social proof where only Google can see it. Headless builds without server rendering hide the whole catalogue. Bot-protection rules (a firewall app, a Cloudflare setting) can challenge AI user agents even when robots.txt welcomes them. And slow responses get dropped: the four-week sequence in the hub post starts with getting server response under 200ms for exactly this reason.
The one-minute test: open view-source on your best-selling product page, or curl it. If the spec table and the reviews are not in that raw HTML, they are not in any ChatGPT or Claude answer.
The first 30% decides what gets quoted
Kevin Indig analysed 3 million ChatGPT responses and 30 million citations, isolating 18,012 verified citations to study where in a page the quoted material lives. The distribution is steep: 44.2% of citations come from the first 30% of the content. The middle 40% of the page supplies 31.1%. The final third supplies 24.7%, with a sharp drop near the footer.
Share of citations by position in the page.
Source: Kevin Indig, 18,012 verified ChatGPT citations, as reported by Search Engine Land, February 2026.
The paragraph-level finding is the useful surprise: 53% of citations come from the middle of paragraphs, not the openers, and plain short sentences outperform dense prose. An assistant reads the way a skimming buyer does. It rewards pages that state facts early and plainly.
On a product page, that means the first 100 words carry the load. What the product is, the three numbers that matter (capacity, runtime, weight, whatever your category trades in) and who it fits. "Crafted for adventurers who demand more" extracts to nothing. "40-litre compressor fridge, 45 W draw, runs 18 hours on a 100Ah leisure battery" hands the model three facts it can quote. The same logic kills the spec-sheet JPEG: an image of a table is invisible to a text crawler. Real table, real text.
One scoping note from Evertune's data: the typical ChatGPT-cited page runs 941 words with four H2 sections and 15 external links. That is a buying guide, not a product listing. Don't inflate your product page to chase it. The product page's job is verification facts up top plus clean structured data; the citable long-form lives on your guides, comparisons and video transcripts.
Schema won't earn the citation. Do it anyway
The honest evidence first. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026, matched them against 4,000 control pages, and measured citations across AI Overviews, AI Mode and ChatGPT. Adding schema produced no citation lift on any platform. Cyrus Shepard's synthesis of 54 studies scores structured data 5.6 out of 10 as a citation factor. Anyone selling schema as the way to get cited is selling against the data.
Schema earns its place doing three other jobs. It feeds the shopping surfaces ChatGPT now attaches to 87% of product questions, which are built from product feeds and structured data rather than prose. It makes the page unambiguous when an assistant does read it: Product, Offer, AggregateRating, FAQ and Breadcrumb markup state your price, stock and rating in a form that can't be misread. And it keeps your entity consistent across every page that mentions you. Necessary plumbing; not a citation lever.
The same limit applies to structured product data you hand to a shopping surface: clean, machine-readable fields decide whether you can be listed, not whether you get named. Shopify gets your products into AI, getting picked is a different job works through that distinction.
This is also the part of the audit we automate. CitoRank scans your live product pages, scores the JSON-LD you have today, and flags exactly which of those fields are missing or broken, alongside the citation tracking. How the audit runs covers both halves.
What Princeton measured: facts beat adjectives
The Princeton GEO study (KDD 2024) tested content changes against generative engines and measured the visibility effect. Adding statistics, quotations and citations from credible sources lifted visibility by up to 40%. Keyword stuffing, the reflex two decades of SEO trained, did not help. And the authors showed the gains skew toward sites ranked lower in classic search: the AI answer layer rewards content quality over accumulated domain authority more than the blue links ever did.
The merchant translation is one sentence: a sourced number beats an adjective, every time. "IP44-rated, 45 W measured draw, kept contents at 4°C for 18 hours in Outdoorsmagic's bench test" is a quotable claim with a named source. "Rugged, efficient and reliable" is filler the model skips. Write the product page the way this series quotes studies, and the assistants have something to work with.
Rebuild one product page this week
Pick the product your margin depends on and run the five steps:
- Fetch the page the way a crawler does. View-source or curl. Confirm the specs and the reviews are in the raw HTML. If they aren't, that's the week's first fix.
- Rewrite the first 100 words answer-first. What it is, the three numbers that matter, who it fits. 44.2% of citations come from the first 30% of the page.
- Move the spec sheet into an HTML table. Delete the JPEG version, or keep it for humans and mirror it in text.
- Fix the JSON-LD and the feed. Product, Offer, AggregateRating, FAQ, Breadcrumb, every field populated, and the same data in your ChatGPT merchant feed.
- Add a visible updated date and dateModified. Then refresh prices, stock and FAQs quarterly so the page reads as current.
Then measure whether it moved anything. CitoRank runs your buyer-intent queries against ChatGPT, Claude and Gemini, scores your product pages' schema, and names every source behind every answer, so a page rebuild shows up as a rank change you can date. See how the audit runs or compare plans.
FAQ
Do AI assistants run JavaScript when they read my store?
No. Vercel's analysis, covering more than 500 million GPTBot fetches, found none of the major AI crawlers execute JavaScript: GPTBot, ClaudeBot and PerplexityBot download script files but never run them. Googlebot is the exception, so Gemini and AI Overviews can see JS-rendered content while ChatGPT and Claude cannot. The only version of your page every engine sees is the server-rendered HTML.
Does Shopify block AI crawlers by default?
No. A standard Shopify robots.txt allows them, and Liquid themes ship server-rendered HTML, which is the right default. The risks come from additions: a customised robots.txt.liquid, a bot-protection or firewall app, or Cloudflare rules that challenge unfamiliar user agents. Check two things: your live robots.txt for GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot and PerplexityBot, and your firewall logs for blocked requests from those agents.
Does adding Product schema get me cited by ChatGPT?
On the evidence, no. Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 control pages and measured no citation lift on AI Overviews, AI Mode or ChatGPT. What schema does do: it feeds the shopping surfaces ChatGPT attaches to 87% of product questions, it states your price, stock and rating unambiguously when a page is read, and it keeps your brand entity consistent. Plumbing, not a citation lever, and still worth doing properly.
How long should a product page be for AI search?
Length is the wrong target. Evertune's typical ChatGPT-cited page runs 941 words with four H2s and 15 external links, but those pages are buying guides and round-ups, not product listings. Your product page needs the verification facts in its first 100 words and clean structured data below; padding it to 941 words helps nobody. Put the long-form on a buying guide or comparison page, where the citation gets earned.
Where on the page should the important facts go?
Early, and mid-paragraph. In Kevin Indig's analysis of 18,012 verified ChatGPT citations, 44.2% came from the first 30% of the content and citation rates dropped sharply near the footer. At paragraph level, 53% of citations came from the middle of paragraphs, and short plain sentences outperformed dense prose. Lead the page with the facts, keep sentences short, and don't bury the one number that wins the comparison in the last section.
What did the Princeton GEO study find works?
Adding statistics, quotations and citations from credible sources lifted visibility in generative-engine answers by up to 40%. Keyword stuffing did not help. The study (Aggarwal et al., KDD 2024) also showed the gains concentrate in sites ranked lower in traditional search, which makes the AI answer layer one of the few places a small store can outperform a bigger competitor's domain authority with better-evidenced content.
Does CitoRank check my product schema?
Yes. The audit scans your live product pages, scores the JSON-LD you have today, and tells you which fields are missing or broken across Product, Offer, AggregateRating, FAQ and Breadcrumb. It runs alongside the citation tracking, so when you fix a page you can watch whether your rank in ChatGPT, Claude or Gemini moves in the weeks that follow, and which sources moved it.
Find out what the assistants can read on your store today
Same-day first audit. Cancel anytime in Shopify.