Strategy · GEO
In February 2026, a BBC senior technology reporter named Thomas Germain decided to test how easy it was to lie to the world’s most-used AI tools. He wrote one article on his personal website claiming he was a competitive hot-dog-eating champion. The next day, ChatGPT, Gemini and Google’s AI Overviews were citing him as a “heavy hitter” in the 2026 South Dakota International Hot Dog Eating Championship. The whole experiment took twenty minutes. The article that documented it was published in BBC Future on 20 May 2026 and updated the following day.
The joke landed because the mechanism is real. The same technical reality that lets one fabricated post poison Gemini is the reality that explains why the gestoría with 72 five-star reviews never gets recommended, why the insurer with the better policy is invisible behind a comparison site, and why the journalist writing a satirical post becomes, for a brief window, a hot-dog champion in the eyes of 2.5 billion AI Overviews users. Three faces of one technical fact: language models do not parse structure. They read text.
What Germain actually did, and why it worked in twenty minutes
The setup was elementary. Germain published a single article on his personal site claiming he had taken first place in the news division at a fictional eating contest. He did not buy backlinks. He did not stuff schema. He wrote a post in the format AI models reward: a clear claim, a clear context, a verifiable-sounding number. By the next day, queries about competitive hot-dog eating in 2026 were returning him as a top result in Gemini, ChatGPT and Google’s AI Overviews.
The investigation Germain published alongside the experiment documented the same trick being used non-satirically. AI was dismissing safety concerns about medical supplements based on a single sponsored blog post. AI was steering people toward retirement strategies authored by entities with undisclosed commercial interests. Lily Ray, founder of the SEO and AI search consultancy Algorythmic, told the BBC the line that organises the rest of this article: “You should assume that you’re being manipulated until they have better systems in place.”
Why does the trick work? Because AI tools often build a response from a single web page or social post rather than from corpus consensus. The model is not weighing thousands of sources and converging on the most defended claim. It is finding the cleanest extractable answer to the question it was given. If your post is the cleanest extractable answer, you are the answer — whether you are telling the truth or not.
The same mechanism explains both manipulation and invisibility
This is where the article we already wrote on FAQ schema connects to the article you are reading now. When Google deprecated the FAQ Rich Result in May 2026, the GEO industry insisted that schema mattered more than ever for AI citation. A controlled experiment by Mark Williams-Cook showed the opposite: large language models do not parse JSON-LD as structured data. They tokenise it as text. The tag @type: Organization is fragmented into separate tokens — @type and Organization — and the semantic frame around them disappears.
That same fact cuts in two directions at once.
In the direction of invisibility: a business with strong services but content buried inside JSON-LD, PDFs, or visually complex layouts is illegible to the model. The Spanish gestoría Abaq has 72 Google reviews at 4.9 stars and a documented edge in international billing. AI tools have never heard of it. The information about Abaq exists. The model cannot extract it cleanly.
In the direction of manipulation: a fabricated claim written in extractable prose is, from the model’s perspective, indistinguishable from a verified claim written in extractable prose. The model is not checking authority. It is checking parseability. A journalist who understands this can spend twenty minutes writing a satirical post and end up cited as a competitive eater. A bad actor who understands this can spend the same twenty minutes on something that matters more. The mechanism is the same. The consequences are not.
Google’s 15 May policy update is closing the front door. The back doors are wide open.
On 15 May 2026, Google updated its Search spam policies to explicitly cover “attempting to manipulate generative AI responses in Google Search.” The new language captures what the industry has been quietly calling generative engine optimisation when sold to clients, and recommendation poisoning when described by regulators. The penalisable tactics now include biased ranking listicles, prompt injection attempts, mass production of low-value AI-generated pages, cloaking, and abuse of expired trusted domains. Penalties range from demotion to full removal from search.
Germain reports a striking detail in his BBC piece. Lily Ray ran a variant of his experiment in the week after the policy update was published. She published a post claiming a fellow SEO specialist was an expert at building sandcastles. Google’s AI cited it. The policy update had been in force for less than seven days, and the front door was already being walked around.
Harpreet Chatha, who runs the consultancy Harps Digital, framed it directly for the BBC: “Google is playing whack-a-mole. They’re announcing the policy update to deter people, but the tactics will just move.” Chatha’s example is concrete. As Google penalises manipulative blog posts, operators are paying YouTube influencers to make the same claims in video. Google’s AI is now citing YouTube videos.
The policy update is a signal. It is not a fix. The marketing manager who reads about it and concludes the problem is solved is reading the wrong signal.
Sycophancy plus external manipulation: a double diagnostic failure
We have written before on this blog about why asking an AI model to diagnose your own AI visibility is structurally unreliable. The Cheng et al. paper published in Science in March 2026 documented that AI models have a measurable structural bias toward affirmation. They tell users what users seem to want to hear. If you ask ChatGPT why your business is not appearing in AI responses, the model will reach for explanations that fit your priors. That is the first diagnostic failure mode.
Germain’s investigation reveals the second. If you ask ChatGPT whether a competitor’s prominent placement is the result of manipulation, the model has no privileged access to its own training pipeline. It cannot tell you whether the post that established that competitor as authoritative was a verified case study or a satirical lie published the day before. It will give you a confident answer either way.
Stacking the two: the marketing manager using ChatGPT to diagnose why their visibility is poor is being told a flattering story about their own content — by a model that may itself have been the downstream victim of upstream poisoning. The model is biased toward affirming the user and blind to its own source quality. Self-diagnosis through the same channel produces a doubly distorted picture.
This is not a methodology problem that the AI tool itself can fix by being asked the right question. It is a methodology problem that requires an external frame.
What this means for the marketing manager who has been told to “optimise for AI”
The operational consequences are three, and worth stating without softening.
First, the questions worth asking about your AI visibility are not about your content. They are about the sources AI models already cite in your sector. When we ran the UK private health insurance experiment, four AI platforms answered questions about Bupa, AXA Health and Vitality citing exclusively WeCovr, Drewberry, MoneySavingExpert and Which?. None of the three insurers’ own corporate websites appeared as sources in any response. The narrative the models built about each insurer was constructed entirely from third-party assessments. Influencing those third-party sources is a distribution problem, not a content problem.
Second, being cited as a source and being recommended as a solution are not the same thing, and the gap is widening. Lily Ray notes in the BBC piece that Google and ChatGPT appear to be quietly removing companies from their AI recommendations when self-promotion is detected, while continuing to cite the underlying article. That distinction is operationally critical. An audit that counts citations of your domain but does not count whether the citation is accompanied by a recommendation is measuring the wrong thing.
Third, the only methodologically honest path to diagnosing AI visibility is external. It is not a conversation with the same model you are trying to appear in. It involves running queries from clean browser environments, documenting which sources are cited, comparing how the same query is answered across platforms, and tracking whether the response changes when the inputs change. None of that is something the model can do for itself, because the model is not a neutral observer of its own behaviour.
Germain ended his piece with a sentence worth noting: just because it looks like a giant tech company is speaking to you instead of some random website does not mean you should have faith. The reverse is also true. Just because a giant tech company is failing to mention your business does not mean your business has failed. It may mean the model cannot read you, the model has been fed a cleaner story about someone else, or the model is removing recommendations it suspects are self-promotional while still citing the underlying page. Three different problems. Three different fixes. If you have been sold a single GEO product that promises to solve all three at once, you have been sold a story that does not survive contact with the methodology.