Skip to main content
StackraStackra
Verified
Case File

We Red-Team Our Own Benchmarks

10 min readAugust 9, 2026Verified against codebase

Every number in your Stackra report compares you to real businesses. That only helps if the comparison is honest.

So we treat our own data the way a good auditor treats a set of books: guilty until the numbers reconcile. Here is how we keep our benchmarks trustworthy, and why the fixes we built under pressure turned out to be textbook methods.

AIndustry
Website and SEO benchmarking
BStack
Common Crawl web graph, Chrome UX Report, BigQuery, a 1.1 million-site US corpus
COutcome
A set of guardrails that keep our published benchmarks honest, most of which turned out to be established statistical methods we arrived at from the data rather than from a textbook.

At a glance

A benchmark is only as trustworthy as the data behind it, and web data is noisy, lopsided, and easy to misread.

20+
Guardrails on our numbers
each built to stop a specific way a benchmark can mislead
11
That map to established methods
reverse-engineered from real problems, then matched to the textbook
Same as yours
Standard for our own studies
we red-team our published numbers with fresh data
0
Manufactured urgency
no scare tactics, ever; it is a deliberate rule

Three pain points, by site type

Tap the one that sounds like your site to read the full story.

1

An average makes the typical site look richer than it is

Affects: Any benchmark that reports a mean over web data: backlinks, referring domains, traffic, page counts.

If this is you

If a report tells you the average site in your industry has 45,000 referring domains and you have 40, you feel broken. You are not.

What it looks like

Web numbers are ruled by a handful of giants. A few enterprise sites drag the average into a range no normal business lives in, so the 'average' describes nobody.

What worked

We report the geometric mean, the honest center of a lopsided distribution. It answers what a typical site actually looks like, not what happens when you average in the outliers.

Result

A category whose plain average is 45,000 referring domains has a typical value nearer 128. One of those is a number you can act on. The other is discouraging and useless.

2

A gap between two dates looks like missing data

Affects: Any system that keeps a stable snapshot for comparison next to a live, constantly updating feed.

If this is you

You want your score to mean the same thing month over month, so we freeze the benchmark you are measured against. That frozen copy sits next to a live feed of the latest web crawl.

What it looks like

Join the frozen snapshot to the live feed the naive way and it looks like 41% of the data disappeared. It is the kind of number that starts a fire drill.

What worked

Nothing disappeared. The two sides were just taken on different dates. We compare like-to-like, the same snapshot to the same snapshot, before we ever call something missing.

Result

The supposed 41% gap was 90% still present. We caught it because we assume a gap is our own mistake until the data proves otherwise.

3

Showing up in a crawl is not the same as using a tool

Affects: Any trend measured from what a web crawler happened to capture in a given month.

If this is you

If we told you a platform's adoption tripled in a year, you would expect that to mean real businesses switched to it.

What it looks like

One vendor's sites did appear to triple. It looked like explosive growth worth writing up.

What worked

Almost none of it was real. The crawler had simply started reaching sites that already existed the year before. We measure adoption from the link graph instead, which does not depend on what a crawler captured that month.

Result

The tripling was a coverage artifact, not growth. Nearly 89% of the 'new' sites were online and linked a year earlier. Reporting it as adoption would have put a made-up trend in front of you.

Your benchmark is a claim about other people's websites

Every number in your Stackra report compares you to real businesses. The score, the industry benchmarks, the analytics cohorts: all of it rests on measurements of other sites. If those measurements are careless, your report inherits the carelessness. So we treat our own data the way a good auditor treats a set of books, guilty until the numbers reconcile.

We built the fixes by hand, then found out they were textbook

None of our guardrails started in a statistics book. Each one came from a specific number that looked right and turned out wrong: a benchmark that put every small business three orders of magnitude away from 'typical,' a trend that swung wildly for no real reason, a coverage gap that was not a gap. We reverse-engineered a fix each time. Later, comparing notes against the academic literature, most of them turned out to be established methods with real names, from capture-recapture to control charts to robust statistics for lopsided data.

The strongest validation of an approach, short of a proof, is discovering you independently rebuilt the one the field already trusts.

A number that swings 2 to 4x release to release still means something, in aggregate

The underlying web graph is noisy. The same site's raw link count can swing 2 to 4 times from one crawl release to the next, purely from what a crawler happened to reach that month. Print that as a delta and you get a chart that looks like your links collapsed when nothing about the site changed. We stopped printing raw per-release deltas and moved to percentile bands instead, the same principle behind why PageRank is described as stable in aggregate even though any single page's raw score jumps around release to release. A trend only counts if it moves the band, not just the point estimate.

A vendor's numbers can drift 2 to 3% and mean nothing

When we measured how fast a hosting platform or CMS was gaining or losing sites, healthy vendors with zero real change still moved 1.9% to 2.8% between measurements from crawl sampling noise alone. Read that as real churn and you would publish a decline story every quarter for a vendor that never lost a customer. So a move has to clear a floor, roughly 3% of the base or 60 sites, whichever is larger, before we call it a trend. It is the same idea as a control limit on a factory floor: not every wobble is a signal.

Present at two different times is not proof something persists

We wanted to know whether a technology, once adopted, tends to stick around. The naive check is to sample a site and see if the same vendor tag shows up a year later, but that check is circular if the later sample is what built the earlier list in the first place, so it will look like nothing ever drops off. We fixed it by only crediting persistence to sites we can independently confirm existed and were tagged in both windows, the same discipline as a mark-and-recapture study estimating a wildlife population: you only trust the overlap you can actually observe twice.

We own the real population, so we can correct for what the crawl missed

A web crawl does not see every site with equal frequency. Small, low-traffic businesses get revisited less often than large ones, which is missing data, not random noise. We also maintain our own list of every business we know is real and active, independent of any single crawl. Where the two disagree, we reweight the crawled sample back toward that known population instead of trusting the raw crawl distribution as-is, the same move a survey statistician makes when calibrating a sample back to known totals.

Overlapping time windows fake a smoother trend than the real one

Comparing two three-month windows sounds safe until you notice they share two of the same three months. That overlap makes a trend look steadier than it is, because most of the underlying data literally did not change between the two windows being compared. We only difference windows with zero shared months, the same fix that corrects for autocorrelation inflating confidence in a trend that has not actually moved yet.

One signal that disagrees with itself is a coin flip, not a finding

We tested letting a single text signal override a site's category label on its own. It agreed with the existing label 36.2% of the time and disagreed 37.7% of the time, a coin flip with extra steps. We now require a second, independent signal to agree before anything overrides a label. A vote of one is not corroboration, it is an ensemble of one.

When honest accuracy is 25 to 65%, the right answer is 'not sure,' not a guess

Our finest-grained classification tiers only classify correctly 25% to 65% of the time on the vaguest clusters. Forcing a specific answer there means being wrong more often than right, so the system is allowed to abstain and roll a site up to a coarser, still-accurate label instead of guessing at the most specific one. It is the same trade-off any statistical classifier makes between precision and coverage: refusing to answer with more confidence than you have is not a workaround, it is a deliberate choice.

We red-team our own studies, not just your site

It is easy to publish a chart that flatters your own thesis. So before we trust a study, we attack it. One group tries to break the finding with fresh data, and a second group defends it. A number survives only if the defense holds up with evidence, not with a show of hands, the same posture behind multi-agent debate as a check against a single confident answer being wrong. We run this on our own published research, including studies with our name on them, and when something does not hold up we correct it. We would rather catch that ourselves than have you catch it.

What this means for your report

You do not need to care about any of this to use Stackra, and that is the point. It means that when your report says a typical site in your industry loads faster than yours, that 'typical' is a number a real business could hit, not an average warped by giants. It means a benchmark will not flip on you because two snapshots were captured on different days. If you want to see the methods in the open, our one-million-site audit and our piece on what backlink tools miss both show the work. Or you can run a free audit and read the numbers knowing they were pressure-tested first.

Frequently asked questions

What is the geometric mean, and why do you use it for benchmarks?
The geometric mean is the average of a set of numbers computed on a log scale, then converted back to normal scale. It resists being skewed by a handful of extreme outliers, which is why we use it for lopsided web metrics like backlinks and referring domains. Where a plain average might put a category's 'typical' referring-domain count at 45,000, the geometric mean puts it closer to 128, a number a typical business can actually recognize as itself.
What is a control limit, and how do you use it in benchmarking?
A control limit is a statistical boundary borrowed from process control charts, used to decide whether a variation is a real signal or ordinary noise. In our benchmarks, healthy vendors with no real change still drift 1.9% to 2.8% from crawl sampling noise alone, so a move has to clear roughly 3% of the base, or 60 sites, whichever is larger, before we report it as a trend.
What is capture-recapture, and how does it apply to web crawl data?
Capture-recapture, also called mark-and-recapture, is a statistical method, originally from wildlife ecology, for estimating a population by comparing overlapping samples taken at two different times. We apply the same discipline to crawl data: we only credit a technology with persisting on a site if we can independently confirm it was tagged in both time windows, not just the most recent one.
What is inverse probability weighting, and why do you reweight crawl data?
Inverse probability weighting is a technique for correcting a sample that under-represents part of a population by giving underrepresented members more weight. Web crawls revisit small, low-traffic sites less often than large ones, so we reweight our crawled sample back toward our own independently maintained list of known active businesses before drawing conclusions from it.
Why do you avoid overlapping time windows when measuring trends?
Comparing two time windows that share months in common creates autocorrelation: the shared data makes the two windows look artificially similar, which understates real volatility and can manufacture a trend that is not there. We only compare windows with zero shared months, so a measured change reflects genuinely new data.
What is PageRank stability, and how does it relate to your benchmarks?
PageRank is designed to be stable in aggregate even though any single page's raw score can swing significantly between crawl releases. We apply the same principle to our link and authority benchmarks: instead of charting the raw number release to release, we track percentile bands, so a real change has to move the band, not just a noisy point estimate.
What is an ensemble, and why do you require two signals to agree?
An ensemble combines multiple independent signals to produce a more reliable answer than any single signal alone. When we tested letting one text signal override a site's category label by itself, it was right about as often as it was wrong. We now require a second, independent signal to corroborate before anything overrides a label.
What is a reject option in classification, and when do you use one?
A reject option, also called abstention, lets a classifier decline to answer rather than force a low-confidence guess. Our most granular classification tiers are only 25% to 65% accurate on the vaguest categories, so we let the system abstain and report a coarser, more accurate label instead of guessing at the most specific one.
What is multi-agent debate, and how do you use it to check your own research?
Multi-agent debate is a validation method where one process tries to break a claim with fresh evidence while another defends it, and the claim survives only if the defense holds up. We run this against our own published research, not just customer-facing scores, and correct anything that does not survive the challenge.
LB
Luke Beck
Founder, Stackra
Last verified August 9, 2026

From the founder

Real tips. Occasional rants.

Dispatches from a founder driving a car while it's being built. No polish, no content calendar. Just what's actually working.

No spam. Unsubscribe any time.

Verdict

Want to see what bots see on your site?