top of page

How ChatGPT and Perplexity Decide What to Recommend

  • Jul 15
  • 8 min read
Businessperson using laptop with glowing AI play icon and media symbols, suggesting tech and digital content work

Quick Answer: AI engines decide what to recommend through a mix of live web retrieval, training-time associations, source authority signals, and citation extractability. How ChatGPT and Perplexity decide what to recommend differs slightly from Gemini and Claude, but all four weigh these signals differently. The dominant patterns are authoritative domains, recent and well-structured content, third-party mentions, and Reddit or forum discussion signals.


The shortest honest answer: nobody outside OpenAI, Anthropic, Perplexity, and Google knows the exact ranking formulas, and the formulas change. What's public knowledge is the architecture (how the engines work mechanically), the training data they used (which influences what they "know" about your industry and your competitors), the live retrieval providers they query, and the patterns that emerge when you test thousands of prompts across each engine.


Princeton researchers Aggarwal et al. published the foundational GEO paper in 2024, testing 9 different content optimization strategies across 10,000 queries. Their findings established the public baseline: citation density, named statistics, and authoritative quotations lifted source visibility by 30% to 40%, while classic SEO levers (keyword stuffing, link volume) had little effect on generative responses. Subsequent work from independent practitioners has filled in engine-specific patterns.


This article covers what's actually known about each major engine's behaviour, the cross-engine signals that hold up consistently, the role of retrieval-augmented generation in 2026, what triggers a citation vs a passing mention, and what this means for how Calgary businesses should structure their content.


At a Glance


Quick Facts:

  • Perplexity: Heaviest reliance on live web retrieval; cites 5 to 10 sources per answer; updates fastest when you change your content

  • ChatGPT Search: Hybrid retrieval + training data; cites 3 to 6 sources; weights authoritative domains and recent publications

  • Gemini: Direct integration with Google Search index; weights schema markup and Google E-E-A-T signals heavily

  • Claude: Training-data-weighted; less frequent updates; rewards established brand and citation presence built over the years

  • Cross-engine pattern: Citation-rich, statistic-dense, well-structured content outperforms keyword-optimized content by 30 to 40% (Aggarwal et al., 2024)

  • Reddit and forum signals: Treated as a proxy for genuine user opinion; routinely surface in answers for "best of" and recommendation queries


How Retrieval-Augmented Generation Actually Works

Most modern AI engines use retrieval-augmented generation (RAG) for queries that involve facts, recommendations, or current information. The process: the engine takes the user's query, runs a search against a live or indexed corpus, retrieves a candidate set of relevant documents, and then uses the retrieved content as context for generating the answer. Citations point back to the documents used.


The implication for GEO is direct. To be cited, you have to be retrieved first. To be retrieved, you need to (a) be indexed by the search provider the engine uses and (b) match the semantic intent of the query well enough to surface in the top retrieval results. Then, to be cited from those results, your content has to contain an extractable, attributable answer.


Each engine uses different retrieval providers. Perplexity built its own search infrastructure and operates more like a search engine with a chat layer. ChatGPT Search uses Bing as its primary retrieval backbone. Gemini uses Google Search directly. Claude (when web access is enabled) uses Brave Search. That mix means: if you're well-indexed on Google and Bing, you're already covered for retrieval into Gemini and ChatGPT, respectively. Brave and Perplexity's indexes have less overlap and reward businesses that ensure their content is technically clean and crawlable.


What ChatGPT Specifically Looks For

ChatGPT (with browsing and search enabled) blends live retrieval with training-time associations. For recommendation queries, the engine tends to surface 3 to 6 named sources. The cited sources skew toward established domains: well-known publications, manufacturer sites, large directories, and sites with strong domain authority signals. Smaller business sites can earn citations, but typically through one of three paths: appearing in authoritative roundups, being mentioned in trusted third-party reviews, or hosting genuinely best-in-class content on a specific topic.


What lifts ChatGPT citation probability:

  • Recent publish date or last-updated timestamp (visible in the page metadata and schema)

  • Clear, attributable claims (named statistics, named experts, distinct paragraphs)

  • Domain authority signals (the kind of trust that translates from classic SEO foundations)

  • Schema markup (Article, FAQ, Organization help ChatGPT identify the content type)

  • Distinct, self-contained paragraphs that read well as standalone excerpts


What suppresses citation: paywalled content, JavaScript-heavy sites that don't render server-side, thin content, and pages that bury the answer below a long preamble.


Hands typing on a laptop keyboard beside a glowing code screen in a dark, moody blue-lit setting.

What Perplexity Specifically Looks For

Perplexity is the most retrieval-heavy of the major engines and the most responsive to content changes. New or updated content can appear in Perplexity citations within days, where ChatGPT and Claude often take months. That makes Perplexity the highest-leverage starting point for Calgary businesses beginning GEO work; the feedback loop is short enough to test what works.


Perplexity's citation behaviour is shaped by three patterns. It cites more sources per answer (often 5 to 10), it ranks freshness highly for time-sensitive queries, and it pulls from Reddit and community discussions aggressively when the query involves recommendations or opinions. The "best Calgary [service]" type queries routinely include Reddit threads in the cited sources.


The practical implications: businesses that update their content regularly, structure it for citation extraction, and have at least some presence in Reddit discussions where their industry is debated will be cited more often in Perplexity than competitors with stronger Google rankings but weaker community presence.


What Gemini and Claude Each Do Differently

Gemini integrates directly with Google Search, which means everything that helps you rank on Google (E-E-A-T, schema, technical SEO, content quality) directly helps in Gemini. The flip side: Gemini answers often reflect Google's existing biases (the same authoritative domains that dominate Google SERPs also dominate Gemini citations). For Calgary businesses with strong existing SEO, Gemini is the most familiar terrain.


Claude weighs training-data associations heavily and updates its model knowledge less frequently than ChatGPT. That means for Claude, the brand and citation footprint built over the years matters more than recent content changes. Businesses that have been consistently mentioned in industry publications, cited in authoritative articles, and built brand awareness over time tend to appear more often in Claude's recommendations than businesses that recently spun up high-quality content.


Anthropic has been transparent about Claude's design philosophy (constitutional AI, Acceptable Use Policy disclosure), which is reflected in citation behaviour: Claude is more conservative about recommending specific businesses without strong third-party validation. Earning Claude citations is the slowest of the four engines, but also the most defensive once established.


Cross-Engine Signals That Hold Up Consistently

While each engine has quirks, the signals that lift visibility across all four are consistent enough to act on. The Princeton GEO study identified the strongest patterns, and subsequent practitioner testing has confirmed them at scale.


The cross-engine signals that work:

  • Citation density. Pages that include 3 to 8 named, attributed citations to authoritative sources earn higher visibility than pages with the same content but no citations.

  • Named statistics. Specific numbers attributed to named sources (Statista, Bank of Canada, industry research) are extracted more often than vague claims.

  • Quotation-friendly structure. Distinct paragraphs that read well as standalone excerpts get cited more than long, flowing paragraphs.

  • Authoritative third-party mentions. The brand being named in trusted external sources (industry publications, reputable directories, news media) lifts citation likelihood for all four engines.

  • Schema markup. Article, FAQ, Organization, and LocalBusiness schema help all engines understand context and entity relationships.

  • E-E-A-T signals. Author credentials, experience markers, and trust signals translate across SEO and GEO.


What does not move the needle: keyword density, internal link volume, exact-match anchor text, and most of the historical SEO levers that were already losing weight under Google's E-E-A-T era.


What Triggers a Citation vs a Passing Mention

There's a meaningful distinction between a citation (your site or brand named as a source the user can click) and a passing mention (your business referenced in the answer text without a clickable citation). Citations drive referral traffic; passing mentions drive brand awareness and brOff-Site Authority Buildinganded search lift. Both matter, but they're earned differently.


Citations typically require: an indexed, retrievable web page that contains the specific answer the AI engine pulled into its response. The page must be accessible to the retrieval system and contain the claim or fact being attributed.


Passing mentions typically require: the brand being strongly associated with the topic in the training data and across third-party sources. A business that's been mentioned in dozens of Calgary marketing roundups, reviewed on Google and Trustpilot, discussed on Reddit, and referenced in industry publications will be named in AI responses about Calgary marketing agencies even when no specific page from the business is cited.


The practical takeaway: invest in both. On-site citation-friendly content earns clicks. Off-site brand presence earns the persistent mentions that compound over the years.


Laptop with glowing CONTENT text and cloud, lock, search, and device icons on a tech-style background, suggesting digital content management

What This Means for Your Content Strategy

Three implications follow directly from how the engines actually work. First, prioritize Perplexity for fast feedback: it's the engine where you can see whether your content changes are working within weeks rather than months. Understanding how ChatGPT and Perplexity decide what to recommend also reinforces the importance of publishing authoritative, well-structured content that is easy for AI engines to retrieve and cite. Second, build for retrieval and extractability simultaneously: technical accessibility, schema, and citation-friendly structure together determine whether your content can be both found and cited. Third, treat off-site authority building as half the GEO program: reviews, Reddit, directory listings, and PR mentions are not optional supplements; they directly influence citation rates.


The companion article on structuring content for generative AI engines covers the on-page tactics in detail. The article on the role of reviews and Reddit covers the off-site authority side. Together, they form the working playbook for Calgary businesses adapting to AI-driven discovery.


Frequently Asked Questions


Do AI engines penalize content the way Google penalizes spam?

Not in the same explicit way, but they functionally suppress low-quality content because the citation extraction logic prioritizes clear, attributable, well-structured content. Thin pages, keyword-stuffed pages, and content that buries the answer get retrieved less often and cited even less often, which produces the same outcome as a penalty without the explicit ranking action.

You can, but the signals overlap heavily across engines, which means optimizing for ChatGPT alone leaves Perplexity, Gemini, and Claude visibility on the table for marginal additional effort. A unified GEO program covers all four with mostly shared work and engine-specific refinements where they matter.

Perplexity refreshes nearly in real time for retrievable web content. ChatGPT updates its retrieval index continuously, but its training-time associations update with each major model release (every few months). Gemini reflects Google index changes within days. Claude updates training knowledge with model releases (typically every 6 to 12 months for major updates), though web-enabled queries reflect current information.

Yes, the underlying signals overlap heavily. Author credentials, demonstrated experience, expertise markers, and trust signals lift visibility across both Google SERPs and AI citations. The investment in E-E-A-T pays double in 2026.

It happens, though less often than it did in 2023 to 2024. The defences are: clear, consistent, accurate information across your owned channels (website, Google Business Profile, directories), regular monitoring through manual prompt testing, and earning enough authoritative third-party mentions that the engines have a reliable corpus to draw from instead of guessing.


Stylized black and teal circular logo with LTL above the word CREATIVE on a white background.

About LTL Creative: LTL Creative is a Calgary digital marketing agency providing Calgary generative engine optimization services for ambitious local businesses, specializing in AI engine analysis, citation strategy, and integrated SEO/AEO/GEO programs, delivered through Google Partner and CXL-certified specialists for owners and marketing leaders requiring measurable, trusted results.


Ready to Drive Results Today by understanding exactly how AI engines find and cite businesses like yours? LTL Creative helps Calgary businesses earn AI citations across ChatGPT, Perplexity, Gemini, and Claude, backed by Google Partner, Meta-certified, and CXL-trained specialists.


Connect with LTL Creative today to discuss your Calgary generative engine optimization strategy.


Disclaimer: Results vary by business, industry, and market conditions. Statistics, platform data, and pricing referenced reflect current industry benchmarks and are subject to change.

Comments


bottom of page