Counting is not analysis
Almost every link profile review I am shown begins the same way: somebody exported a backlink report, looked at the total, compared it with a competitor's total, and formed a conclusion. That number is close to meaningless, and the reasons why are the whole subject.
A raw export from any link index is a list of records, not a list of links. The same link appears repeatedly because it was found at several URL variants of the same page. Links appear that were removed years ago and have not been recrawled. Links appear from pages that exist only as a redirect. Site-wide footer and sidebar links inflate a count by tens of thousands while representing a single editorial decision made once. And a considerable share of any large profile is scraped content, aggregators and automatically generated pages that no human ever chose to publish.
I gave a talk on this at BrightonSEO in September 2023 called “Advanced link profile analysis in 2023,” and the argument was simple: the work is not gathering more data, it is reducing what you already have to the set of links that are actually real, actually live, and actually mean something. Everything useful happens after that reduction.
Getting the raw data right
No single link index sees the whole web. Each maintains its own crawl, its own retention rules and its own definition of when a link is dropped from the set, which is why two tools reporting on the same site disagree by margins that would be alarming anywhere else. For anything consequential — a penalty investigation, an acquisition, an analysis that will be shown to someone who will argue with it — pull from more than one index and treat the union as your working set.
Export at the right granularity. You want, at minimum, source URL, target URL, anchor text, first seen date, last seen date, link type, whether it carries a nofollow or sponsored attribute, and whatever host-level metric the tool provides. Reports that give you referring domains only have already thrown away the information you need.
Add the site's own record to the pile. Search Console's link report is the one source drawn from the search engine's own index rather than a third-party crawl, and while it is sampled and its anchor reporting is limited, a link that appears there is a link the search engine knows about. That is a different and more interesting fact than a link a commercial crawler found.
One caution about dates: first seen means when the tool first observed it, not when it was published. For a link that existed before the tool crawled that page, the date is wrong, and in any matter where timing is contested that distinction needs stating explicitly.
A row is not a link
Deduplication is unglamorous and it is where most of the value is created. Work through it in this order.
- Normalize the source URLs. Lowercase the host, strip the protocol distinction, unify www and non-www, remove trailing slash inconsistencies, and strip tracking and session parameters. A meaningful fraction of any export collapses at this step alone.
- Collapse to one record per source page and target. Several links from one article to one destination are one editorial decision, not several endorsements.
- Identify template links. If a domain links to you from thousands of pages with the same anchor in the same position, that is a site-wide element. Count it as one link with a note about its scale, not as thousands.
- Roll up to referring domain, then to registrable domain. Subdomains of the same host are frequently the same publisher, and free subdomain platforms are the opposite — thousands of unrelated publishers on one registrable name.
- Separate the link types. Editorial in-content links, navigation and footer links, user-generated links from forums and comments, directory listings, syndicated copies of a single original, and redirect-only records each deserve their own bucket. Averaging them together produces a number that describes nothing.
What comes out the other side is usually between a fifth and a half the size of what went in. That smaller set is the profile. The original export was inventory.
Verify that the links exist
A link index tells you what a crawler saw on the date it last visited. It does not tell you what is on the page now. So fetch the source pages yourself and check, at scale, whether the link is present in the response, whether it appears in the raw HTML or only after JavaScript executes, whether it carries a nofollow, sponsored or user-generated attribute, and what the page's own indexability status is.
Several categories fall out immediately. Links on pages that now return a 404 or a redirect. Links on pages carrying a noindex directive, which cannot pass anything because the page is not in the index. Links that exist in a rendered document but not in the served HTML. Links inside markup that was never crawlable in the first place. And links on pages that have been rewritten since the link was placed, where the surrounding context no longer has anything to do with your site.
Do the same check on the destination. A link pointing at a URL that redirects twice, or at an old HTTP address, or at a page that has since been removed, is worth substantially less than the export implies — and reclaiming those is usually the highest-return work in the entire exercise, because the link already exists and somebody already agreed to it.
Read the anchor text distribution
Anchor text is the clearest available signal of whether a profile grew or was built. Bucket every anchor into branded, naked URL, generic, partial-match, exact-match commercial, and image or empty anchors, then look at the shape.
A profile that accumulated naturally is dominated by brand mentions, bare URLs and generic phrases, because that is how people actually link. The commercial phrases exist but sit in the tail. When the distribution inverts — when a competitive money phrase is the most common anchor on the site — that did not happen by itself.
Then add the dimension people forget: time. Plot anchors by first seen date. Manufactured profiles show bursts. A hundred links with near-identical commercial anchors appearing across three weeks and then stopping is a campaign, and it is visible in a chart long before it is visible in a total. Natural acquisition looks lumpy but continuous, with the bursts tied to identifiable events — a launch, a piece of coverage, a conference.
Anchor variation matters as much as anchor type. Spun anchors are a recognizable pattern: the same phrase with word order shuffled and a filler word swapped, over and over, across unrelated hosts. Nobody writes that way.
Finding the network behind the links
Link networks are built by people reusing their own infrastructure, and the work is finding the reuse. Resolve every referring domain and cluster on the signals, roughly in descending order of strength:
- Reused analytics and advertising identifiers. The strongest available signal, because a human had to paste the same identifier into multiple properties. Historical page captures often preserve identifiers that were later removed.
- Registration data. Shared registrant email addresses or organization names where they are visible, common registrars, and creation dates clustered into the same short window.
- Nameservers. A cluster of otherwise unrelated sites on the same unusual nameserver pair is a stronger indicator than shared hosting, because it reflects a management decision rather than a purchase.
- IP addresses and subnets. Useful, but the weakest of these and the most abused in reporting. Large hosts place tens of thousands of unrelated sites behind one address, so co-location is a lead to check, never a finding to report. What is genuinely interesting is a set of sites spread across many addresses that all fall within a small number of adjacent subnets — that is a deliberate attempt to look diverse.
- Certificate records. Public certificate transparency logs record every issued certificate, and a certificate covering several unrelated names ties them together at a moment in time.
- Content and template fingerprints. Identical boilerplate, matching favicons, the same theme with the same customizations, recycled stock imagery, reused phone numbers and addresses.
A cluster built from several independent signals is persuasive. A cluster built from shared hosting alone is the kind of exhibit that gets a whole analysis discounted, and I have seen that happen.
Relevance is the judgment that remains
After the cleaning, the verification and the clustering, what is left requires a human decision: is this a site whose readers would plausibly care about the destination?
That question resists metrics. Host-level authority scores are third-party estimates, they are gameable, and a high score on a site with no topical connection and no real audience is exactly what an expired-domain network is engineered to produce. I would rather see a link from a small trade publication that genuinely covers the industry than from a general-interest site with an impressive score and no readers.
The practical tests: does the linking page have an author and a discernible editorial purpose; is the surrounding content about the same subject or is your link the only commercial mention on an otherwise unrelated page; does the site rank for anything itself, which tells you whether the search engine takes it seriously; and would you be comfortable showing the page to the client. That last one is not a joke. It correlates better with outcomes than most scores do.
Segment the surviving set by topic and by intent, and you can finally answer the question the exercise was for: is this profile concentrated where the business competes, or is it spread across places that happen to have been available?
What to do with the answer
Restraint matters here. The disavow file is a blunt instrument and it is overused. Search engines have become far better at ignoring manipulative links than at penalizing them, which means most bad links are already discounted and disavowing them changes nothing while risking the removal of links that were fine. I disavow when there is a manual action to clear, when a profile shows an unmistakable purchased or networked pattern at scale, or when a domain was acquired with an inherited profile built for a purpose that has nothing to do with its new owner. Otherwise I leave it alone.
The more productive outputs are usually reclamation and comparison. Fix the links pointing at broken or redirecting URLs. Chase unlinked brand mentions. Then run the same cleaning process against two or three competitors and compare the cleaned sets rather than the raw totals, which is the only way that comparison means anything.
The slides from my BrightonSEO talk are public on Speaker Deck if you want the compressed version. Consulting engagements, including link audits and cleanups, are handled through Hartzer Consulting.