What the US business web is made of
Start with the shape of it. Every classifiable site in the corpus is sorted into one of 25 business categories, and the distribution is lopsided in a way that says a lot about who actually has a website. Retail and restaurants alone are a quarter of the entire US business web. Add healthcare, nonprofits, education, and B2B services and six categories carry more than half of everything. The digital-native categories that dominate tech coverage barely register down here: SaaS is 1.3% of the web, portfolios 0.4%. The web that real businesses live on is local, physical, and service-driven, which is exactly the slice most benchmarking tools underweight. Not every site fits a category honestly: 3.8% (43,221 sites) is left genuinely unclassified rather than forced into the nearest wrong bucket -- see the note below the table.
Retail and restaurants (highlighted) are 25% of the corpus between them. By contrast SaaS is 1.3% and portfolios 0.4%. The full 25-category breakdown is in the table below.
| Category | Sites | Share |
|---|---|---|
| Retail | 153,548 | 13.4% |
| Restaurants | 137,829 | 12.0% |
| Healthcare | 93,686 | 8.2% |
| Nonprofits | 77,212 | 6.7% |
| Education | 68,361 | 6.0% |
| B2B services | 63,064 | 5.5% |
| Automotive | 56,488 | 4.9% |
| Real estate | 53,103 | 4.6% |
| Home services | 48,684 | 4.2% |
| Entertainment | 46,212 | 4.0% |
| Wellness and beauty | 37,654 | 3.3% |
| Content and media | 37,546 | 3.3% |
| Professional services | 37,318 | 3.3% |
| Hospitality | 36,618 | 3.2% |
| Recreation | 33,315 | 2.9% |
| Religious | 30,680 | 2.7% |
| Government | 29,344 | 2.6% |
| SaaS | 14,731 | 1.3% |
| Financial services | 14,145 | 1.2% |
| Funeral services | 9,975 | 0.9% |
| Transportation | 6,603 | 0.6% |
| Senior care | 6,164 | 0.5% |
| Portfolio | 5,105 | 0.4% |
| Utilities | 4,664 | 0.4% |
| Storage | 1,883 | 0.2% |
These 25 categories sum to 96.2% of the corpus (1,102,598 sites). The remaining 3.8% (43,123 sites) is honestly unclassifiable -- no confident category signal, not force-fit into the nearest wrong bucket. Counts are from stackra_us_corpus, the frozen US snapshot, re-verified 2026-08-09.
25 categories, 469 subcategories
Category is the top level. Beneath it, every classified site also carries one of 469 subcategories, so a restaurant is not just a restaurant but a sushi bar, a brewery, or a taqueria, and a retailer is a clothing store or a furniture showroom. That second level is what makes cohort benchmarking specific: your site gets compared against others doing the same thing, not against the whole 1.1 million. There are far too many subcategories to chart, but they are the reason a benchmark can say 'sites like yours' and actually mean it.
Where the sites come from
HTTP Archive's May 2026 crawl (mobile client, root pages only), queried through BigQuery, with a content floor of at least 1,000 bytes of HTML and 20 DOM elements to drop parked domains and empty shells. Excluded on purpose: platform freebie subdomains (WordPress.com, Wix.com), throwaway preview hosts with no custom domain (bare vercel.app, netlify.app), personal publishing hosts (Blogspot, Medium, Tumblr, Substack), institutional domains (.gov, .edu, .mil), and bare IP-address hosts. What is left is real, independently-hosted US business sites.
How a site counts as US
Not by domain ending or a locale guess: we tested the TLD-proxy approach and it caught only 32% of confirmed-real sites. Two real signals instead, neither treated as a ceiling on its own:
| Signal | Sites confirmed | Notes |
|---|---|---|
| Overture Maps address match | ~744,500 | Physical US street address on file, open places dataset |
| US city/state text signal | +442,808 net-new | Site's own text, screened for false-positive city names (e.g. 'reading', 'paradise') |
How we know the labels are right
A category map is only as good as the labels under it, so the labels got two rounds of checking. First, we removed every borrowed shortcut that looked authoritative but wasn't: a Google-Maps-embed rule, a schema.org Person tag, and a Squarespace-means-portfolio assumption between them mislabeled roughly 1.57 million sites before any AI pass ran, so we measured each one's real error rate by direct count and dropped all three. The categories the model kept returning off-list (financial services, funeral care, recreation, storage, transportation) were promoted to real types rather than forced into the nearest wrong bucket. Then we checked the finished labels two independent ways: a 504-row blind audit where a separate model graded raw page content with no category hints, and a 38-row hand-read of the actual pages behind the biggest label changes, which came back 33 to 35 out of 38 clean. Where both blind methods classified the same row, they agreed 97.4% of the time.
| Signal | Assumption | False positives | Reality |
|---|---|---|---|
| Google Maps embed | → local business | 600,000 | 84,000 real vs 600K wrong |
| schema.org Person tag | → personal portfolio | 738,000 | removed outright |
| Squarespace CMS | → portfolio | 233,000 | 94% of Squarespace sites are actually retail |
~1.57M combined false positives removed before any AI reclassification pass ran.
Where this stands, and what you get from it
This is the same corpus Stackra's category benchmarks and the link-authority research draw from, with every row keeping its prior label (business_type_previous) for audit. It is a reference dataset, not a live scan feature on its own: when you run a Stackra audit, your site is classified the same way and then compared against the category and subcategory cohorts measured here. That is how a benchmark tells you whether your site is ahead of or behind the specific slice of the web you actually compete in.