Meet Cortex - AI Powered, Expertise Refined Decision EngineYour AI Optimization Engine
GEOAug 18, 2026·13 min read

GEO for News Publishers: Citation Economics When the Click Never Comes

TL;DR

The publisher problem is no longer traffic acquisition, it is attribution. Referral volume is structurally declining and no amount of technical SEO reverses that, so the strategic question becomes whether you are named as the originating source when a story you broke gets summarised. That is winnable, and it turns on three things: publishing what only you have, structuring reporting so the attribution travels with the fact, and making a deliberate decision about which crawlers get access rather than defaulting to all or none. Blocking everything protects nothing if the story reaches the engine through a syndication partner anyway.

Audience

Audience development leads, SEO directors, and editorial leadership at news organisations facing structural referral decline.

Cortex

Cortex is modern marketing. Old marketing waited on people. Modern marketing fuses the efficiency of AI with the experience of experts. Meet your optimization engine.

Get Cortex

Effective

Google's guidance on creating helpful content asks whether content provides original information, reporting, research or analysis, which is the standard that separates primary reporting from aggregation. [src]

Impact

Google's article structured data documentation specifies the NewsArticle type and the author, publisher and date properties that carry attribution. [src]

Action

Google publishes dedicated guidance for appearing in Discover, which has become a larger share of publisher referral traffic as search referrals decline. [src]

Platform

Google documents that Google-Extended controls Gemini training and grounding without affecting inclusion or ranking in Google Search, which is the mechanism behind a partial publisher opt-out. [src]

Methodology

Cortex built this post from AI answer sets across 30 breaking and developing news queries, tracking which outlets were named as originating source versus merely cited, and cross-referenced the results against each outlet's crawler access policy and structured data.

Publishers have spent three years trying to fix a traffic problem that is not fixable. Search referrals are declining structurally because the queries that drove them are now answered in the result, and no technical SEO programme reverses that.

The winnable problem is different. When an engine summarises a story your newsroom broke, does it name you? That question has a real answer, the answer varies enormously between outlets, and the variance is driven by things a publisher controls.

This guide covers what determines whether you are named as the originating source, why blocking every crawler often fails to protect anything, and which numbers to report once sessions stop being the goal. Read it alongside our guide to SEO for news publishers, which covers the traffic side, and our data on zero-click search.

The Shift from Traffic to Attribution

Accept the premise before optimising against it, because half the industry's effort is still aimed at the wrong outcome.

Informational news queries resolve inside the answer now. A user asking what happened in a story gets a synthesis of several outlets rather than a list of ten headlines. That is not a ranking failure and it is not recoverable by better markup.

What remains contested is the attribution layer. Every synthesised answer names sources, usually 2 to 7 of them out of the dozens covering the story, and being in that set has 3 distinct values.

Brand impression at the moment of information consumption, which is closer to how television news was valued than to how a session is valued.

Authority reinforcement, since being repeatedly named as the source on a beat compounds. Engines that have learned an outlet is the primary source on a subject reach for it again.

Residual click share, because a proportion of readers do click through on stories they care about, and that proportion is higher on original reporting than on commodity coverage.

The strategic consequence is that a publisher should stop competing on stories everybody has and concentrate on stories only they have, because attribution is the currency and originality is what earns it.

What Only You Have

Original reporting is the only durable position, and the reason is structural rather than editorial. A summarisation system can rewrite an explainer. It cannot manufacture a primary source.

Google's helpful content guidance asks whether content provides original information, reporting, research or analysis, and the categories that qualify for a newsroom are specific.

  • Documents nobody else obtained. A records request, a leaked memo, a court filing pulled before the wires noticed.
  • Interviews with named people. A direct quote is unrepeatable. Whoever else covers the story has to cite the outlet that got it.
  • Proprietary data. A survey, an analysis of a public dataset, a tally the newsroom maintained.
  • Physical presence. Being in the room, at the scene, in the courtroom. Detail that cannot be sourced remotely.
  • Beat expertise that produces a story before the event. The piece explaining what the vote will mean, published before the vote.

Each one produces a fact that only exists in your reporting, and a fact that only exists in one place is a fact that must be attributed. That is the entire mechanism.

The corollary is worth stating bluntly. Commodity coverage of wire stories now has close to zero attribution value. Ten outlets rewriting the same agency copy give an engine 10 interchangeable sources for 2 to 7 slots, and it will name whichever it likes. Resource allocated there is resource not spent on the reporting that would have been cited.

Writing So Attribution Travels

Original reporting still loses attribution when the sourcing is separated from the fact. This is a craft problem and it is fixable.

Retrieval works on passages. A system extracting a paragraph gets whatever is in that paragraph and nothing from three paragraphs earlier. So a story that establishes its sourcing in the second paragraph and states its exclusive fact in the ninth has produced a highly extractable fact with no attribution attached.

Five habits fix it, and none of them harms the reading experience.

  • Put the sourcing in the sentence with the claim. According to the filing obtained by this outlet, rather than establishing the provenance earlier and relying on the reader to carry it.
  • Name the outlet in the passage where the exclusive detail sits. Self-reference reads oddly to editors and it is what survives extraction.
  • Attribute quotes in the same paragraph they appear in. A quote in one paragraph and the speaker's identification in the next can arrive as an unattributed quotation.
  • State what is exclusive explicitly. First reported here, obtained by this newsroom, confirmed independently. These phrases are signals to a retrieval system as much as to a reader.
  • Front-load the new information. The lede is the most extracted passage on the page, so the thing only you have belongs in it.

Our post on content chunking covers the underlying retrieval mechanics.

The Syndication Leak

Here is the failure that undermines most publisher blocking strategies, and it is rarely modelled before the decision is made.

An outlet blocks every AI crawler to protect its journalism. The same outlet syndicates to wire services, partner sites, and aggregators, all of which republish the story. Those partners have not blocked anything. The story reaches every engine through them, stripped of the originating attribution, and the outlet that did the reporting is absent from the answer while its content is fully represented in it.

The block achieved the opposite of its purpose. It removed the publisher from the citation without removing the content from the corpus.

Three things to audit before deciding anything about crawler access.

Map where your content legitimately appears. Wire partners, syndication deals, aggregator feeds, MSN-style portals, and licensed republication all count.

Check whether those partners carry canonical or source attribution back to you. Many do not, or do so in a form no engine will resolve.

Test the actual outcome. Ask an engine about a story you syndicated and see which outlet it names. That single check tells you more about your exposure than any policy discussion.

If the leak is wide, the negotiation that matters is with your syndication partners about attribution requirements, not with the AI companies about access.

Block, License, or Allow

There are three coherent positions and one incoherent one, and most publishers are in the fourth.

The incoherent position is blocking selectively without understanding the syndication picture, which produces the leak above.

Allowing broadly is coherent when your value is attribution and reach. You accept that content is used, you optimise to be the named source, and you treat citation as brand exposure. This suits outlets whose revenue is advertising and brand-driven rather than subscription-driven.

Blocking broadly is coherent when your content is genuinely exclusive and your revenue is subscription. If a reader can get the substance from an assistant, the subscription case weakens. This requires the syndication side to be tight, or the block is decorative.

Licensing is coherent and is where the larger outlets have landed. Access in exchange for payment and attribution terms, which converts an uncontrolled use into a revenue line. It requires leverage, which means exclusivity, which brings the argument back to original reporting.

One mechanism worth knowing for a partial position. Google documents that Google-Extended controls Gemini training and grounding without affecting inclusion or ranking in Google Search. That lets a publisher decline Gemini training while keeping Search and Discover intact, which is a genuinely available middle path. Our post on the difference between AI training and AI retrieval covers the distinction that makes these decisions separable.

Structured Data for News

Markup does not create attribution but it removes ambiguity about who published what and when, which matters on a fast-moving story.

Google's article structured data documentation covers NewsArticle. Five properties carry disproportionate weight for a publisher.

datePublished with a precise timestamp including timezone. On a developing story where 40 outlets publish inside 90 minutes, the first has a claim to primacy, and a date-only value cannot establish it.

dateModified, kept honest. Updating it without changing anything substantive is a pattern engines discount.

author as a Person with a profile URL, not the newsroom. A named reporter with a beat is a resolvable entity and an institutional byline is not. Our guide to ProfilePage schema covers building those entities.

publisher resolving to a single canonical Organization node with a logo.

citation where the story rests on a document, linking the filing, dataset, or judgement. This is underused and it is a direct signal of primary sourcing.

Two more notes. Keep NewsArticle for actual news and use Article for explainers and features, because misapplying the type on evergreen content is a mismatch engines notice. And with search referrals declining, Discover has become a larger share of publisher traffic, so the image and titling requirements there deserve attention proportional to that shift.

What to Measure Instead of Sessions

Reporting session decline every month tells leadership nothing they can act on. Four numbers replace it.

Citation share on originated stories is the headline metric. Take the stories you broke, query the engines about them, and record how often you are named. This is directly attributable to editorial decisions and it moves when the work improves.

Citation share versus named competitors on your core beats. The absolute number matters less than whether you are gaining on the outlets you compete with.

Attribution accuracy, meaning how often you are named as originating source rather than merely cited alongside outlets that followed your story. This is the number the syndication leak destroys, so it doubles as a diagnostic.

Referral quality rather than volume. Sessions from AI referrals are fewer and typically more engaged, so measure subscription conversion rate on that traffic rather than its size.

Our guide to measuring GEO performance covers the instrumentation, and our post on SEO reporting covers presenting it.

Common Mistakes

  • Blocking crawlers while syndicating freely. Removes you from the citation without removing the content from the corpus.
  • Competing on wire rewrites. Ten interchangeable versions of the same story give an engine no reason to name yours.
  • Sourcing established early and exclusives buried late. Extraction takes the passage, not the article, so attribution has to sit with the fact.
  • Institutional bylines. A newsroom name is not a resolvable author entity and forfeits the reporter's beat authority.
  • Date-only timestamps on developing stories. Primacy cannot be established without a precise time.
  • Reporting session decline as the metric. It is structural, unactionable, and it crowds out the numbers that are neither.
  • Refreshing dateModified with no substantive change. A recognised pattern that engines discount.

Implementation Sequence

  1. Audit the syndication footprint and test which outlet engines name for stories you originated. Do this before any access policy discussion.
  2. Fix attribution requirements with syndication partners, since a leak there undermines everything downstream.
  3. Move sourcing into the sentence carrying the claim, and get exclusives into the lede. This is a style guide change.
  4. Replace institutional bylines with named reporters carrying profile pages and beat descriptions.
  5. Get NewsArticle markup right: precise timestamps, honest dateModified, Person authors, and citation on document-based stories.
  6. Take a deliberate access position per crawler, using Google-Extended for a partial stance if that fits.
  7. Report citation share on originated stories monthly, alongside competitive citation share on core beats.

Frequently Asked Questions

Should news publishers block AI crawlers?

Only after auditing syndication. If your stories reach engines through wire partners and aggregators who have not blocked anything, a block removes you from the citation while the content circulates anyway. Fix the attribution terms with partners first.

Can we decline AI training without losing Google Search?

Yes for Gemini. Google documents Google-Extended as governing training and grounding without affecting inclusion or ranking in Search, so it offers a genuine partial position. Other vendors draw the line differently, so the policy has to be written per crawler.

What actually determines whether we get named as the source?

Having something only you have. A document you obtained, a quote you got, data you compiled, or presence somewhere nobody else was. A fact existing in one place must be attributed, and that is the whole mechanism.

Does structured data improve attribution?

It removes ambiguity rather than creating attribution. Precise timestamps support a primacy claim on developing stories, named Person authors build resolvable beat authority, and citation signals primary sourcing. None of it substitutes for original reporting.

What should we report to leadership instead of sessions?

Citation share on stories you originated, competitive citation share on core beats, attribution accuracy as originating source, and subscription conversion rate on AI referral traffic. Those move with editorial decisions in a way session decline does not.

Key Takeaways

  • -Referral decline is structural, so the goal shifts from clicks to being named as the originating source.
  • -Original reporting is the only durable moat, because summarisation cannot manufacture a primary source.
  • -Attribution travels with a fact only if the sourcing sits in the same sentence as the claim.
  • -Blocking every crawler fails when your wire partners republish the same story without the block.
  • -The metric that replaces sessions is citation share on stories you originated.

Ready to optimize for the AI era?

Get a free AEO audit and discover how your brand shows up in AI-powered search.

Get Your Free Audit