Strategy · GEO

More people every day search for products and services directly in ChatGPT or Gemini before they go to Google. Not to browse — to get a recommendation. “What’s the best accounting software for freelancers?” “Which legal contract tool do law firms actually use?” “What business insurance covers remote workers?” The answer arrives in seconds. And if your company isn’t in that answer, that potential customer doesn’t know you exist.
What determines who appears and who doesn’t in those answers is not the same mechanism that determines who ranks on Google. It’s a different logic, with different criteria and different metrics. The discipline that works on those criteria is called GEO — Generative Engine Optimization. This guide explains what it is, how it works, and what separates companies that appear from those that don’t. It draws on academic research from Princeton and Georgia Tech published in 2024, and on original experiments we ran across four sectors — Spanish accounting services, legal tech, private health insurance in Spain, and private health insurance in the UK.
What GEO is and why the term matters
GEO (Generative Engine Optimization) is the discipline that optimizes a brand’s presence in the responses of generative language models — ChatGPT, Gemini, Perplexity, Claude — so that those models select it as a reliable source when answering their users’ questions.
The term was formally introduced in November 2023 in a paper by researchers from Princeton, Georgia Tech, the Allen Institute for AI and IIT Delhi, published at the ACM KDD 2024 conference — one of the most prestigious data science conferences in the world. It was the first rigorous academic framework for this discipline, and demonstrated that applying GEO techniques can increase content visibility in generative responses by up to 40%.
In practice, GEO goes by several names depending on the context: AEO (Answer Engine Optimization) and LLMO (Large Language Model Optimization) are technical variants used by some practitioners. GEO is the globally dominant term with academic backing — it’s the one we use at Deslumbra IA as the standard reference.
The difference between SEO and GEO is not one of degree — it’s one of logic:
| Dimension | SEO | GEO |
|---|---|---|
| Objective | Position in a ranked list of results | Citation in a generative response |
| Unit of success | Rankings, clicks, CTR | Share of answer, citation rate |
| Primary lever | Domain authority + keyword relevance | Content structure + brand entity signals |
| Time horizon | Weeks or months | Days (RAG) or quarters (training data) |
Key dimensions comparing SEO and GEO. Source: own elaboration.
Why this changes now — the shift that isn’t reversing
This isn’t a forecast. It’s a measurable reality.
In 2024, SparkToro and Datos — Semrush’s data division — published the most rigorous study to date on Google search behaviour. The finding: of every 1,000 searches on Google in the United States, only 360 end with a click to the open web. Nearly 60% of searches generate no click at all — the user gets the answer directly on the results page and doesn’t need to go anywhere else. In Europe the figure is almost identical: 374 clicks per 1,000 searches. The trend has been worsening for years, and the arrival of Google’s AI Overviews in 2024 has accelerated it.
What this means for a business is not that organic traffic is dead — it’s that its nature has changed. The user who does click arrives with more information, more context and more purchase intent than before. Volume falls, quality rises. But that only benefits whoever appears in the generative answer that precedes that click.
Research shows that AI-triggered searches generate zero-click rates above 80% — users get a complete answer without visiting any external site. The businesses that benefit from the traffic that remains are the ones already inside the answer.
How AI decides which companies to recommend
The distinction no one explains clearly: training data vs RAG
Language models operate with two types of knowledge, and understanding the difference is fundamental for knowing what to optimise and when to expect results. The first is training knowledge: information the model incorporated during its training process, processing vast volumes of text from the internet, books and other sources. This knowledge is static — it doesn’t change until the model is retrained, which happens every several months with each new version. If your company didn’t exist in that corpus, the model doesn’t know it. If it existed with outdated information, the model works with that outdated version.
The second is RAG (Retrieval-Augmented Generation): an architecture that allows the model to retrieve information in real time from external sources before generating its response. Perplexity uses RAG extensively — it’s essentially a search engine with language generation on top. ChatGPT uses it for real-time searches. This mechanism means content published today can appear in a response tomorrow, because the model doesn’t rely exclusively on its training knowledge.
The practical implication: optimising for training knowledge takes time — weeks or months until crawlers index content and that content enters the next training cycle. Optimising for RAG can have effect within days, if the content is well structured and AI crawlers can access it.
What signals each model uses to select sources
Models don’t behave the same way. Each platform has its own criteria for weighting the sources it includes in its responses. According to the State of AI Search 2026 report by AthenaHQ, based on analysis of 8 million AI responses, each model cites an average of 12 distinct sources per response — but with significant variation: Grok reaches 16 sources per response, while Gemini stays below 10.
What our own experiments confirm is that this variation is not random. In the UK private health insurance experiment, four different platforms consulted the same external sources about Bupa, AXA Health and Vitality — and reached different conclusions about Vitality because each weighted those sources differently. The same input, four different outputs. That’s what makes AI visibility not a technical problem with a single solution: it’s a multi-channel presence problem where each channel applies its own criteria.
Why corporate websites almost never appear as sources
This is the most counterintuitive finding for most businesses. According to AthenaHQ’s analysis, only 14.98% of AI responses cite a brand’s own domain as a source. More than 85% of responses that mention a company include no link to that company’s own website.
The reason is structural. A corporate website talks about the company: who we are, what we do, why we’re good. An AI model, when building a response to a purchase decision question, looks for content that answers that question with explicit criteria, comparison between options and verifiable data. Comparison platforms, specialist publications and user forums produce exactly that type of content. Corporate websites, in general, don’t.
We documented this in the UK private health insurance experiment: not one corporate website for Bupa, AXA Health or Vitality appeared as a cited source across any of the platforms analysed. The companies controlling what models said about those insurers were WeCovr, Drewberry, MoneySavingExpert and Which?. The insurers themselves did not participate in the narrative their potential customers received through AI.
The five factors that determine whether AI cites you
There is a separate article on this blog that covers in detail the technical factors that prevent AI from finding a company — robots.txt, llms.txt, Schema markup, brand entity. If you haven’t read it, that is the starting point before this one. See the technical factors guide. What follows here are the five strategic factors that determine whether, once models can find you, they decide to cite you.
1. Content that answers questions, not content that describes products
The Princeton and Georgia Tech study published at ACM KDD 2024 measured the effect of different optimisation techniques on visibility in generative responses. The clearest finding: adding verifiable statistics to content improves visibility by up to 41%. Citing authoritative external sources improves it by up to 115% for lower-ranked content. Adding persuasive phrasing without data had no significant effect — or made things worse.
The practical translation: content that answers specific questions with data, comparisons and explicit criteria is orders of magnitude more likely to be cited than content that describes a product or service in positive terms. A decision guide with original data is worth more to models than ten pages of corporate copy.
We saw this in the Spanish accounting services experiment: in February 2026 we asked four AI platforms the same question about the best accounting services for self-employed workers in Spain. The top result on Gemini was Xolo — a company founded in Estonia, with no local team in Spain — ranked above every domestic competitor. Not because it had the best product, but because it had built the most complete content architecture for that question: specific pages for different self-employed profiles, articles answering exactly the questions those users ask, and verifiable published pricing. The models found that content and cited it. Spanish competitors with equivalent products but without that architecture didn’t appear.
2. A recognisable brand entity for the models
For a model to cite you with confidence, it needs enough external signals confirming that your company exists, does what it claims and is considered relevant by others in its category. That includes Schema Organization markup on the website, a verified Google Business Profile, mentions in relevant sector directories and consistency between the company name and its description across all those sources.
The Abaq case illustrates the limit of this logic: 72 Google reviews with a 4.9 out of 5 rating, competitive product, years in the market — and invisible across all four AI platforms we tested. Reviews are a trust signal for users, but they are not a direct citation mechanism for models. What was missing wasn’t reputation — it was content architecture that models could extract and use to answer decision questions.
3. Presence in the sources each model already cites in your sector
The lever isn’t always publishing more content on the owned domain. Sometimes — especially in sectors where comparison platforms dominate — the lever sits entirely in the sources models already consult.
Our own experiments show this playing out very differently by sector. In online accounting services for self-employed workers, publishing well-structured content on the owned domain has a measurable effect on citations — publishing on your own domain works. In private health insurance, as the UK experiment documented, comparison platforms dominate so completely that publishing more on the corporate website moves nothing — the right strategy involves understanding what the dominant comparison sources say about your company and how to update that narrative. Two sectors, the same starting question, two completely different answers.
4. Authority built over time
Models incorporate new information into their base knowledge in cycles of months, not weeks. Whoever starts building AI presence today is working ground their competitors haven’t touched — and that advantage compounds with each new model version.
The Princeton finding is relevant here: content in a low search position benefits 115% more from GEO optimisation than content already in position 1. The opportunity is greatest for those starting from zero or from a weak position. Companies that have spent years investing in SEO and hold consolidated positions on Google have less relative improvement margin in GEO than a well-structured company starting from scratch.
5. Coherence between what you say and what others say about you
If the content about your company online is produced primarily by comparison platforms, directories or aggregators, the narrative models build about you isn’t yours. It belongs to those sources, with their editorial criteria, their data — which may be out of date — and their commercial interests. The UK private health insurance experiment showed the extreme consequence of this factor: Bupa, AXA Health and Vitality did not appear as a source on any platform, and the result was that different models gave contradictory assessments of Vitality based on how each weighted the same external sources. The companies didn’t participate in any of those assessments. Not from lack of digital presence — from lack of presence in the specific sources models consult.
The GEO metrics that replace rankings
SEO has a mature measurement system: Google positions, impressions, clicks, CTR. GEO is building its own. These are the six metrics marketing teams need to start tracking.
Share of answer — the percentage of a model’s responses that mention your brand for a defined set of target prompts. The GEO equivalent of share of voice in traditional media, measured inside generative responses.
Citation rate — how frequently the owned domain appears as an attributed source in AI responses. Being mentioned isn’t enough — being cited as a source carries greater weight because it signals the model considers your content a primary reference.
Mentions per prompt — the consistency of appearance across different formulations of the same question. A model may cite you when a question is framed one way and not when it’s framed another. Measuring this reveals coverage gaps.
Daily citations — the number of times AI engines cite your content daily. A volume metric that allows detecting changes in visibility before they translate into traffic changes.
AI-sourced leads — conversions whose origin is discovery in a generative response. Difficult to measure directly, but partly trackable through traffic with no known referrer and through monitoring which channels cite blog content as the source of a contact.
ROI per prompt — the business impact generated by appearing in specific high-value queries. Allows prioritising which questions to optimise first based on their potential business impact.
These metrics don’t replace SEO — they complement it. A company that isn’t indexed by Google won’t easily appear in models that use RAG with search engines as a base. SEO remains the infrastructure. GEO is the visibility layer that determines what happens inside generative responses.
Does GEO work the same way in every sector?
No. And that difference has concrete implications for strategy. Our experiments make this visible: in online accounting services, publishing well-structured content on the owned domain moves the needle. In private health insurance, comparison platforms dominate so completely that the lever sits entirely in external sources. Two sectors, the same starting question, two completely different answers. According to the AthenaHQ State of AI Search 2026 report, in technology and software the /blog accounts for 59.89% of AI entries into brand websites; in logistics and mobility the /home dominates at 48.15%. The content architecture that works depends on the sector.
GEO is not a discipline that can be applied generically. What works in accounting doesn’t work in insurance. What moves the needle in legal tech is not what moves it in healthcare. The concepts in this guide are the framework — but the framework only makes sense when applied to a specific sector with real data.
The most direct way to understand how GEO operates in practice is not to read more about it. It’s to see what happened when we measured it. The Spanish accounting services sector is the most documented in the blog: three experiments covering general services for self-employed workers, an in-depth case study on Abaq, and a niche experiment in creator economy accounting after Spain’s 2026 Tax Authority enforcement plan. The first experiment is the most direct entry point to what GEO means when it stops being theory: who appears, who doesn’t, and why product quality isn’t the variable that explains it.
Before building a GEO strategy, the starting point is knowing how AI sees you today. At Deslumbra IA we run that diagnostic with real data: prompts executed across the main platforms, sources cited by each model, and a clear read on which levers make the most sense for your specific sector.