Why Link Correlation Studies Prove Less Than They Claim
A link correlation study takes a large sample of search results, measures a link metric for each ranking URL, and reports a statistical association between the metric and the position. That association is usually real. What it cannot establish is that links caused the position, that more links would improve a position, or that the metric measures what its name suggests.
Those three gaps are not nitpicks. They are the difference between “higher-ranking pages tend to have more referring domains” — which is well-established and unsurprising — and every strategic claim built on top of it.
The four structural problems
Confounding, unavoidably. Pages that rank well tend to be good pages on established sites. Good pages on established sites get cited. So links and rankings share upstream causes: quality, age, brand, distribution, budget. No observational study of live search results can separate these, because you cannot hold them constant. There is no control group of identical pages differing only in link count.
Reverse causation is live. Ranking well makes a page more discoverable, which makes it more likely to be cited. Some portion of the association runs from position to links rather than the other way. Nobody has a credible estimate of the split, and the direction problem is intrinsic to cross-sectional data.
The metric is not the construct. A study that correlates Domain Rating with position is correlating one vendor’s index-dependent estimate with position. Whatever the true underlying link signal is, the metric is a proxy for it with unknown error — see Domain Rating versus Domain Authority. Measurement error attenuates correlations, which means these studies both overstate their causal reach and understate their own uncertainty about the construct.
Aggregation across queries. Search results for a medical query, a local service query, and a software comparison query are produced under different conditions. Pooling thousands of queries produces an average association that may not describe any individual SERP, and the variance is usually far more interesting than the mean.
Why the correlation coefficient is the least useful number
Published studies often report a correlation figure. Three reasons not to lean on it.
Position is ordinal and truncated. Ranks 1–10 are ordered categories from a distribution that has been cut at 10. Correlations computed on truncated ranked data behave badly and are not comparable across studies with different cutoffs.
Link metrics are compressed scales. Correlating a logarithmic-ish score with a rank position gives a number whose magnitude depends heavily on transformation choices most write-ups don’t state.
The coefficient can’t be converted into anything actionable. There is no arithmetic that turns “moderate positive correlation” into “N links moves you M positions.” Any write-up that makes that leap has added an assumption it didn’t test.
What a genuinely informative study would need
This is worth stating positively, because it clarifies why the good version is rare.
- Random assignment, which is impossible: you cannot randomly assign citations to pages at scale, and if you manufactured them you’d be studying manufactured links.
- A natural experiment — some exogenous shock that changes links without changing quality, brand, or intent match. These exist occasionally and are hard to find and harder to generalise from.
- Longitudinal data with per-URL controls, tracking the same URLs over time so each page acts as its own control. This is the strongest available design and it still can’t rule out simultaneous changes.
- Reported variance across query types, not a single pooled coefficient.
- A stated index and date for every metric used, since the metric is a function of a crawl.
Very little published link research meets even half of this, and that is not primarily a criticism of the researchers. The clean design isn’t available.
How to read a study you’re handed
Five questions, in order.
- What exactly was measured, from which index, on what date? If the write-up doesn’t say, the numbers aren’t reproducible.
- What was the sample? Which queries, from where, in what language, and how selected. A sample of high-volume commercial English queries generalises to high-volume commercial English queries.
- Does the language stay correlational? Look for “associated with” holding all the way to the conclusion. If the abstract says “correlated” and the conclusion says “drives,” the study didn’t change — the write-up did.
- Is variance reported? A mean without a spread is a headline, not a result.
- Who benefits from the finding? Vendor research is not automatically compromised — vendors have the only datasets large enough — but a study whose conclusion is “buy more of what we sell” deserves the same scrutiny as any other interested party’s.
The claims to refuse outright
“Links are N% of the algorithm.” Not published, not knowable externally, and the framing assumes a decomposition that may not exist.
“Sites with X referring domains rank Y times better.” Ratio claims on rank positions are not meaningful quantities.
“We tested this on N sites and saw Z% lift.” Without a control group this is a description of what happened to those sites during a period when many things happened. See how to tell whether a link did anything.
Any specific decay figure — for redirects, for nofollow, for link age. These circulate as folklore with no source.
What to do with the honest residue
After all that, something survives, and it’s worth saying plainly: the association between link metrics and rankings is one of the more consistently observed patterns in SEO measurement, across many independent datasets, over many years. Links matter. Search engines say so, and the observational data agrees.
What no study establishes is how much, for which queries, with what diminishing returns, or whether a given link will do anything for you. So the defensible use of correlation research is as background — it justifies caring about links at all — and not as a source of targets, forecasts, or attribution.
If someone asks for a number anyway, the useful move is the bounded, labelled competitor comparison rather than a study coefficient. That method, and its limits, is in how many backlinks do you need to rank.