← All posts

Embeddings for SEO: six ways cosine lies on a live corpus

· 6 min read · By Mikhail Kuzmitskii

embeddingssemantic SEOinternal linkingcannibalizationcontent ops

View as Markdown ↗

String matching has a hard ceiling. Two pages can cover the same topic with almost no shared vocabulary — “cheapest eSIM for Japan” and “prepaid data plans for Tokyo travellers” overlap on roughly nothing — and every duplicate-detection tool built on word overlap will call them unrelated. Embeddings remove that ceiling: a model turns each text into a vector positioned by meaning, and closeness becomes arithmetic.

That part is well covered. What is not covered is the next part — the six specific ways cosine similarity misleads you once you point it at a real corpus instead of a demo. These are field notes, each one from a run that shipped something wrong before we caught it.

What embeddings actually buy you

Three corpus questions that are unanswerable with string matching, and trivially answerable with vectors:

Embeddings for SEO: six ways cosine lies on a live corpus
  • Near-duplicates that share no words. Paraphrase, translation drift, two writers briefed from the same outline.
  • Latent cannibalization. Page A’s body covers page B’s topic while A claims to be about something else — see the cannibalization Ahrefs can’t see for how direction, not similarity, is the signal there.
  • Internal-link candidates and orphans. Which paragraph on page A is genuinely the closest thing to page B’s topic — and which pages sit far from everything, semantically orphaned even when they have inbound links.

All three ride on one primitive: cosine similarity between vectors. Which is exactly why its failure modes matter.

1. Boilerplate pairs with itself at cosine 1.0

The most expensive lesson. A site-wide paragraph — a YMYL disclaimer, an exchange-rate note, a CTA block — is byte-identical on every page. Identical text embeds to an identical vector. So every article is maximally similar to every other article through the block that carries no topic at all, and an internal linker faithfully links on the disclaimer.

We watched this happen on a 52-page corpus: link candidates ranked by similarity, and the top of the ranking was the same legal paragraph pairing with itself, over and over.

2. There is no universal “duplicate” threshold

Every tool that quotes a magic cosine number without naming its model and its unit of text is guessing. The number moves with both.

On our corpora, paragraph-level paraphrase lands around 0.92 and up, so the near-duplicate cut sits at 0.90 — deliberately conservative. Internal-link candidates use 0.65. Same math, different numbers, because the two questions have opposite error costs: a false duplicate verdict merges or 301s a page that should have lived, while a weak link suggestion costs a human three seconds to reject. Pick your threshold from what a mistake costs, then calibrate it on your own corpus.

3. Length skews the number

A 1,800-word body and a five-word target keyword never reach a high cosine, no matter how perfectly the page is about that keyword. Long text averages toward the centre of the space; short text sits at the edges. Compare like with like — body to body, keyword-phrase to keyword-phrase — or compare gaps between two similarities rather than the absolute value of either.

This is the single most common way a home-made semantic audit produces nonsense: it compares a page body to a query string, sees 0.58, and concludes the page is off-topic.

4. Cosine does not know direction

“A is similar to B” is symmetric. Real content relationships are not. A comprehensive guide that covers a subtopic is not the same situation as a thin page that duplicates that guide’s topic — one deserves a link, the other deserves a merge, and a single similarity score cannot tell them apart. You need a second measurement (what each page claims to be about, versus what its body covers) and the difference between the two.

5. A cached vector describes the page you no longer have

Caching embeddings is mandatory — a re-run over an unchanged corpus should cost nothing. But the cache key has to be the text, not the URL or the slug. Key it by URL and every audit after an edit reasons about the previous version of the page, silently and forever. The symptom is nasty precisely because there isn’t one: no error, no warning, just an audit that keeps agreeing with itself.

6. The spend is invisible, not large

Embedding a normal site costs cents. The problem isn’t the amount — it’s that the call hides inside an audit that feels free, so it never appears on a cost line until someone specifically goes hunting. We shipped three audits that embedded paid vectors without a single metering line, and the cost report happily printed a total that excluded them.

Meter at the shared helper where the call is actually made, never in each caller — the next audit someone adds will forget, and the guard has to make forgetting impossible.

The practical setup

If you’re adding embeddings to a content workflow, the order that survives contact with a real corpus:

  1. Strip repeats first. Verbatim paragraphs on 2+ pages never reach the model.
  2. Cache by text hash. Unchanged text, zero cost, correct vector.
  3. Two thresholds, not one. Conservative for destructive verdicts (merge, 301), permissive for suggestions a human reviews.
  4. Compare like with like. Body↔body, claim↔claim. Use gaps, not absolutes, when the lengths differ.
  5. Meter at the call site. One place, so it can’t be forgotten.
  6. Positive-control your zero. “No duplicates found” and “the scan silently measured nothing” look identical in a report. Plant a known duplicate and confirm the audit catches it — otherwise a clean result proves nothing.

That last one is the habit that generalises beyond embeddings entirely. Silence is not the same as cleanliness, and any audit that can’t tell you which it just produced is not finished.

Our own semantic audits — duplicate scanning, entity shadowing, link-equity spread — run on exactly this pipeline, and each of the six rules above exists because breaking it shipped something wrong first. If you want the shape of the whole system rather than the primitive, topical authority and internal-linking silos covers where these scores land in a site architecture.

FAQ

What is an embedding, in SEO terms?

A model reads a piece of text and returns a list of numbers — a vector — positioned so that texts about the same thing land near each other. Once every page is a vector, 'are these two pages about the same thing?' becomes an arithmetic question (cosine similarity) instead of a word-overlap question. That is the whole trick: it catches pages that share a topic while sharing almost no vocabulary.

What cosine threshold means 'duplicate'?

There is no universal number, and any tool quoting one without naming its model and its unit of text is guessing. On our corpora, paragraph-level paraphrase sits around 0.92 and above, so we cut near-duplicates at 0.90 — deliberately conservative, because a false MERGE costs a page. Internal-link candidates use 0.65, because there the cost of a miss is higher than the cost of a bad suggestion a human rejects.

Why did our internal linker start linking on disclaimers?

Because a site-wide boilerplate paragraph — a YMYL disclaimer, an FX-rate note, a CTA — is the same text on every page, so it embeds to the same vector on every page and pairs with itself at cosine ~1.0. Every article becomes maximally similar to every other article, through the block that carries no topic at all. The fix is to drop paragraphs repeated verbatim across two or more articles before embedding, which also stops you paying to embed the same paragraph N times.

Do embeddings replace keyword research?

No. They answer a different question. Keywords tell you what people type and how much demand sits behind it; embeddings tell you how your own pages relate to each other in meaning. Use embeddings for corpus questions — duplicates, cannibalization, link graph, topical spread — and demand data for what to write next.

How much does embedding a corpus cost?

Cents for a normal site, but the real risk is that the spend is invisible rather than large. Embedding calls sit inside audits that feel free, so they never show up on a cost line until someone goes looking. Meter the call where it is made — in the shared embed helper, not in each caller — and cache vectors keyed by the text, so a re-run of an unchanged corpus costs nothing.

Was this helpful?
Launch SEO/GEO with seo·matrix → or a request without Telegram →

Comments

  • …

Comments are moderated.