Research brief
- What this research adds
- A first-party citation capture of one commercial phrase across Perplexity, Claude, Gemini and Google AI Overview on a single day, including a within-engine repeat test that measured how much two identical runs disagree, plus an independent schema and freshness audit of all 11 cited pages that corrects my own earlier finding.
- Research question
- When someone asks an AI engine about "Claude Cowork for marketers", which sources get cited, do the engines agree with each other, and does the same engine agree with itself?
- Method
- On 6 August 2026 from a Copenhagen IP, I put the identical prompt to Perplexity (Kimi K3, Computer mode), Claude (Fable 5, effort high, web search on), and Gemini (3.6 Thinking), and captured the Google AI Overview for the same phrase. Every cited URL, its position, and Perplexity's full 30-item retrieved set were recorded from the rendered DOM. Claude was queried twice, 17 minutes apart, under identical settings, as a repeat test. I then fetched all 11 cited pages plus 2 retrieved-but-uncited pages directly over HTTP and inspected their raw JSON-LD and datePublished values myself. Limits: one capture day, one geo, one account, one run per engine except Claude. ChatGPT failed twice and returned no data, so all cross-engine counts are out of 3, not 4. No Perplexity stream capture was performed, so there is no classifier, trust-object or index-timestamp data in this run.
- Confidence
- Medium
- Evidence
- first-party test, LLM citation capture, primary documentation, independent schema audit
- Next verification
TL;DR
I asked Perplexity, Claude and Gemini the same question about Claude Cowork for marketers on 6 August 2026 and logged every source they cited. Anthropic's own domains were the only publisher all three agreed on. Of the 20 distinct third-party pages cited across the engines, 16 were cited by exactly one engine and none were cited by all three. Then I ran Claude twice, 17 minutes apart, same model, same prompt, same account: the two runs cited 13 distinct URLs between them and shared exactly one. One cited source returned a 404. Another, cited by Gemini, described a product state Anthropic had already replaced four months earlier. The practical consequence: a single citation check is a sample, not a scoreboard, and if you are optimising for one engine's citation list you are optimising for noise.
Evidence ledger
| Claim | Evidence | Source type | Verified | Confidence | Caveat |
|---|---|---|---|---|---|
| Anthropic's own domains were the only publisher cited by all three answering engines | Perplexity cited claude.com/product/cowork and the Cowork product guide; Claude cited claude.com/plugins/marketing and support.claude.com get-started; Gemini cited the same support.claude.com get-started article | LLM citation capture | 2026-08-06 | High | Counts claude.com, support.claude.com, anthropic.com and anthropic.skilljar.com as one publisher. Treated separately, no single URL reached 3/3 |
| No third-party page was cited by all three engines | Highest third-party score was 2 of 3, reached by Coupler.io, techsy.io, Customer.io and CNBC | LLM citation capture | 2026-08-06 | High | Three engines, not four. ChatGPT returned no data |
| 16 of 20 distinct third-party pages cited were cited by exactly one engine | Full per-engine cited sets recorded in the run manifest; union across Perplexity, Claude, Gemini and Google AI Overview totals 20 third-party URLs, of which 4 appear in two engines and 16 in one | LLM citation capture | 2026-08-06 | High | Google AI Overview counted inside the Gemini family, matching how the two surfaces share an index |
| Two identical Claude runs 17 minutes apart shared 1 of the 13 URLs they cited | Run 1 (20:56 CEST) cited 5 URLs, run 2 (21:13 CEST) cited 9, both on Fable 5 at high effort with web search on; only Coupler.io appeared in both | First-party test | 2026-08-06 | Medium | Two runs is a small sample. It establishes that instability exists, not its exact magnitude |
| Perplexity retrieved 30 sources and cited 7 of them | Perplexity's expanded Sources panel listed 30 entries against 6 inline citation chips plus 1 confirmed bundled URL, Engadget | LLM citation capture | 2026-08-06 | High | Three grouped chips had an unverifiable second bundled source each, so the true cited count is 7 to 10 |
| A source cited by Claude returns HTTP 404 | Claude cited an AOL article URL; direct request returns final_status=404 and AOL's own "page does not exist" body | First-party test | 2026-08-07 | High | Confirmed twice, once by browser navigation and once by direct HTTP request. Not a bot wall |
| A source cited by Gemini described a superseded product state | gend.co presents Cowork as a Max-only macOS research preview with a waitlist, while Anthropic's own product page documents general availability across paid plans on macOS and Windows | Primary documentation | 2026-08-07 | High | The page carries BlogPosting and FAQPage JSON-LD but publishes no datePublished value at all |
| Schema markup did not separate cited pages from uncited ones | 7 of the 11 cited pages I fetched carry Article or BlogPosting JSON-LD; so does Adspirer, which Perplexity retrieved and never cited | Independent schema audit | 2026-08-07 | High | Raw HTML inspection of head-injected JSON-LD. Corrects my own earlier extraction pass, which reported zero schema across the set |
| Only 3 of the 11 cited pages carry FAQPage schema | FAQPage present on techsy's Cowork guide, techsy's Marketing Ops piece and gend.co only | Independent schema audit | 2026-08-07 | High | claude.com/product/cowork emits Question and Answer entities without wrapping them in FAQPage |
| The most-cited third-party page is not the newest | Coupler.io publishes datePublished of 2026-04-17 and was cited by 2 engines; The Ad Spend is dated 2026-07-12, carries no JSON-LD, and was retrieved by Perplexity without being cited | Independent schema audit | 2026-08-07 | Medium | Single phrase, single day. Freshness may still matter on other query classes |
| Video holds real citation surface in Google's AI Overview but not in Perplexity | 2 of the 8 Google AI Overview source cards were Grace Leung YouTube videos (one, two); Perplexity retrieved 4 YouTube URLs and cited none | LLM citation capture | 2026-08-06 | Medium | Perplexity was running in Computer mode, which may weight video differently from its default search mode |
| The named-expert slot behaves differently from the guide slot | Perplexity cited Simon Willison's first-impressions post in second position despite it carrying no JSON-LD, no FAQ and roughly 1,800 words | LLM citation capture | 2026-08-06 | Medium | One observation on one engine. Not enough to generalise about author authority |
Where the evidence conflicts
My own two passes disagreed about schema, and the first one was wrong. My initial extraction pass over the ten top-cited pages reported zero schema markup across the entire set. That finding was tidy, dramatic, and false. When I fetched the same pages directly and grepped the raw HTML, seven of eleven cited pages carried Article or BlogPosting JSON-LD, several with ten or more separate ld+json blocks. The extraction tool had missed markup injected into the document head. I am publishing the correction rather than the original claim because the original claim would have sent readers off to add schema they mostly already have.
Engine self-reports do not match engine behaviour. Perplexity's rendered Sources panel showed 30 entries. Only 6 URLs appeared as inline citation chips. Four of those chips were grouped pills showing "+1", meaning each bundles a second source that the interface never exposes. I could confirm one bundled URL and not the other three. So the honest cited count for that run is a range, 7 to 10, not a number. Any tool that reports a precise citation count off a rendered Perplexity answer is reporting a number it cannot see.
Cited does not mean current, and the engines have no visible freshness floor. gend.co is cited by Gemini while describing a waitlist that no longer exists. It publishes no date. Meanwhile the freshest page in the whole set, dated 2026-07-12, was retrieved and dropped. I cannot tell from this run whether the engines lack a freshness signal, weight it below other factors, or are reading stale index copies. Distinguishing those three would need Perplexity stream capture with index timestamps, which this run did not perform.
Retrieved and absent are not the same failure, and I can only separate them for one engine. Perplexity exposed its full 30-item retrieved set, so I can say that Wired, Fortune, DataCamp, Reddit and four YouTube videos made the candidate pool and lost at synthesis. Claude, Gemini and the AI Overview expose no equivalent list. For those, a missing page could be a ranking failure or an extraction failure and there is no way to tell from outside.
Unverifiable in this run: every classifier probability, trust object, index timestamp and fan-out query. Those live in Perplexity's event stream and I did not instrument it. Anything you read about Perplexity's intent thresholds, including in my own earlier notes, is not supported by this capture.
What I tested
Date: 6 August 2026, 18:42 to 21:13 CEST. Location: Copenhagen, Denmark, residential IP, logged-in paid accounts on all four services.
Prompt, identical across every engine and typed rather than pasted from a template:
I'm trying to understand Claude Cowork for marketers. Give me a thorough, source-backed answer with citations.
Engines and models, each verified in the live model picker at run time rather than assumed:
| Engine | Model | Mode | Result |
|---|---|---|---|
| Perplexity | Kimi K3 | Computer, 9 steps, 1m46s | 30 sources retrieved, 6 inline chips |
| Claude | Fable 5, effort high | Web search on | 5 URLs cited, run 1 |
| Claude | Fable 5, effort high | Web search on, +17 min | 9 URLs cited, run 2 |
| Gemini | 3.6 Thinking | Default | 5 URLs, 14 citation instances |
| Google AI Overview | n/a | Same phrase in Search | 8 source cards |
| ChatGPT | n/a | Two attempts | No data |
Procedure: for each engine I recorded the selected model before submitting, submitted the prompt once, waited for generation to finish, expanded every collapsed source list, and read the URLs out of the DOM rather than out of the visible answer text. Grouped citation chips were clicked individually where they could be expanded without navigating away. Where a bundled source could not be revealed, I logged it as unverified instead of guessing.
Then, on 7 August, I fetched all 11 cited pages and 2 retrieved-but-uncited pages over plain HTTP with a desktop user agent, and inspected application/ld+json blocks, @type values and datePublished fields in the raw response body.
The repeat test. The second Claude run was an accident. My ChatGPT session failed to connect on the first attempt, and on the retry the browser landed on claude.ai instead and re-ran the same prompt. That mistake produced the most useful number in this dossier. Two runs, same model, same effort setting, same account, same prompt, 17 minutes apart. Run 1 cited 5 URLs. Run 2 cited 9. Between them they named 13 distinct URLs and agreed on exactly one, Coupler.io. Group by publisher and treat Anthropic's four domains as one, and the agreement rises to two publishers out of roughly eleven. Either way it is a minority.
Expected result: meaningful overlap between engines, and near-total overlap between two runs of the same engine minutes apart.
Observed result: neither. Cross-engine consensus existed only for the vendor's own documentation. Within-engine consensus was 1 URL in 13.
Reproducibility limits. One day, one geo, one account, one run per engine except Claude. Model routing, personalisation and index state all vary. Anyone repeating this will get different URLs. That is the finding, not a defect in it. If you want to try, keep the prompt identical, record the model from the picker rather than assuming it, and expand every source list before reading anything.
What ChatGPT's absence costs this page. Every cross-engine count here is out of three. A 4/4 consensus source might exist and I would not have seen it. I would rather publish a three-engine result labelled as one than pad it to four.
Change log
| Date | Change found | Evidence affected | Conclusion changed? |
|---|---|---|---|
| 2026-08-06 | Initial capture across Perplexity, Claude, Gemini and Google AI Overview | All citation claims | Initial publication |
| 2026-08-07 | Direct HTTP audit found Article or BlogPosting JSON-LD on 7 of 11 cited pages, contradicting the earlier extraction pass that reported zero | Schema claims | Yes. "Nobody has schema" was replaced with "schema did not separate cited from uncited" |
| 2026-08-07 | AOL URL cited by Claude confirmed returning HTTP 404 on direct request | Dead-link claim | No, strengthened |
My judgment
My conclusion
Stop treating a citation check as a scoreboard. On this phrase, the same engine disagreed with itself 17 minutes later on 12 of 13 URLs, and three different engines agreed on nothing except the vendor's own documentation. Any single-run citation audit, including the kind sold as a GEO dashboard reading, is one sample from a distribution nobody has published the width of. The one durable pattern in the whole capture is that the vendor's first-party pages win the definitional layer outright, and everything above that layer is close to a lottery among competent pages.
What builders should do
Run the capture at least three times per phrase and only act on sources that appear in two or more runs. Everything else is noise dressed as a finding, and chasing it will have you rewriting pages against a target that moved while you read it.
Concede the definitional layer. If the question has a vendor, the vendor's docs own "what is it" and "what does it cost" across every engine. Link to them and compete one level up, on application, judgment and limits, where I found zero cross-engine consensus and therefore an open field.
Ignore the schema panic. Seven of eleven cited pages already had Article or BlogPosting markup, and so did a page that got retrieved and dropped. Schema is table stakes here. Ship it, then stop optimising it and go make the page worth extracting instead. I keep coming back to this in my Cowork plugins and memory guide: the structural work is necessary and it is not the thing that wins.
Check what your cited competitors actually say before you copy their angle. One of them is a 404. Another is describing a waitlist that closed months ago and is still being cited today.
What I would not trust yet
The exact instability figure. Two Claude runs is enough to prove instability exists and nowhere near enough to size it. The 1-in-13 number is a data point, not a rate.
Any conclusion about ChatGPT. It failed twice and contributed nothing. Every count here is out of three engines.
The freshness reading. I have one phrase on one day showing an April page beating a July page. That is suggestive and it is not a rule.
Anything about Perplexity's classifier, trust levels or index timestamps. I did not capture the event stream, so I have no data on any of it, and neither does anyone quoting thresholds captured on a different account in a different country.
The video finding. Perplexity ran in Computer mode, which is not its default search path, so its total indifference to four retrieved YouTube URLs may be a mode artifact rather than a ranking signal.
What would change my mind
Ten repeat runs per engine on the same phrase producing a stable core of sources across 8 or more of them. That would mean the noise I measured is a small-sample artifact and single-run audits are defensible after all.
A stream capture showing the engines carry index timestamps that correlate with citation choice. That would convert my freshness shrug into an actual mechanism.
A third-party page reaching 3-of-3 cross-engine consensus on a vendor-owned phrase. It did not happen here, and if it happens reproducibly elsewhere then the vendor lock on the definitional layer is weaker than this capture suggests. I run the same kind of capture across Perplexity's own surfaces and would expect the pattern to hold, so a clean counterexample would be genuinely interesting.
FAQ
Does this mean GEO optimisation is pointless?
No, it means single-run measurement is pointless. The vendor-documentation pattern held across every engine, and that is actionable: do not spend effort competing for definitional queries a vendor owns. What you cannot do is read one citation list and conclude your page lost.
Why did ChatGPT return nothing?
Two attempts, two different failures. The first never established a browser connection. The second was redirected to claude.ai and re-ran the prompt there, which is where the accidental repeat test came from. I logged it as no data rather than substituting a result.
Is 17 minutes long enough for citations to legitimately change?
That is the point. Nothing about the underlying web changed in 17 minutes. The index did not meaningfully move, no new page was published, and the model and account were identical. The variation came from the retrieval and synthesis path, which is exactly the thing a single-run audit assumes is stable.
How is this different from a rank tracker?
A rank tracker reads one ordered list from one surface. This reads which sources an engine chose to quote inside a generated answer, which is a different and noisier decision. Perplexity retrieved 30 sources here and cited around 7 of them, so ranking into the candidate pool and getting quoted are separate outcomes with separate failure modes.
Can I reproduce this without paid accounts on all four services?
Partly. Google's AI Overview needs no account and Perplexity and Gemini both answer without one, though results differ from a logged-in session. Claude's web search and Perplexity's Computer mode both need a paid plan. Expect different URLs regardless, since geo and account history both affect retrieval.