Skip to main content
StackraStackra
Verified
Case File

We Sorted 1.1 Million US Business Websites by Category. Here's the Map.

5 min readJune 25, 2026Updated August 9, 2026Verified against codebase

The shape of the US business web, measured. Real business sites broken down into 25 categories and 469 subcategories.

Retail and restaurants are more than a quarter of it, SaaS barely registers, and the whole thing skews local and physical. Here is the full breakdown, how a site earns a US label, and how we checked the labels before trusting them.

AIndustry
Internal data research, US business web corpus
BStack
BigQuery, HTTP Archive, Overture Maps, Vertex AI / Gemini, TypeScript
COutcome
1.1 million US business sites sorted into 25 categories and 469 subcategories, every site labeled, with the three biggest false-positive shortcuts removed and the labels cross-checked two independent ways before publishing.

At a glance

The data comes from a May 2026 HTML archives crawl. 15.4 million websites. Removing and classifying the ones I could verify left this sample.

1,146,202
Sites classified
real US business websites; 96.2% carry a category label, the rest are honestly unclassifiable
25
Business categories
every classifiable site sorted into exactly one
469
Subcategories
the finer cohort beneath the 25 categories
25%
Retail + restaurants
the two biggest categories are a quarter of the entire business web
1.3%
SaaS share of the web
14,731 sites; the digital-native categories are rare
97.4%
Label agreement, blind check
where two independent methods both classified the same row

What the US business web is made of

Start with the shape of it. Every classifiable site in the corpus is sorted into one of 25 business categories, and the distribution is lopsided in a way that says a lot about who actually has a website. Retail and restaurants alone are a quarter of the entire US business web. Add healthcare, nonprofits, education, and B2B services and six categories carry more than half of everything. The digital-native categories that dominate tech coverage barely register down here: SaaS is 1.3% of the web, portfolios 0.4%. The web that real businesses live on is local, physical, and service-driven, which is exactly the slice most benchmarking tools underweight. Not every site fits a category honestly: 3.8% (43,221 sites) is left genuinely unclassified rather than forced into the nearest wrong bucket -- see the note below the table.

Share of the US business web, by category (top 8 of 25)
RetailRestaur.HealthNonprofitEduc.B2B svcsAutoReal est.0%4%8%12%16%

Retail and restaurants (highlighted) are 25% of the corpus between them. By contrast SaaS is 1.3% and portfolios 0.4%. The full 25-category breakdown is in the table below.

All 25 business categories, by share of the 1.1M-site corpus
CategorySitesShare
Retail153,54813.4%
Restaurants137,82912.0%
Healthcare93,6868.2%
Nonprofits77,2126.7%
Education68,3616.0%
B2B services63,0645.5%
Automotive56,4884.9%
Real estate53,1034.6%
Home services48,6844.2%
Entertainment46,2124.0%
Wellness and beauty37,6543.3%
Content and media37,5463.3%
Professional services37,3183.3%
Hospitality36,6183.2%
Recreation33,3152.9%
Religious30,6802.7%
Government29,3442.6%
SaaS14,7311.3%
Financial services14,1451.2%
Funeral services9,9750.9%
Transportation6,6030.6%
Senior care6,1640.5%
Portfolio5,1050.4%
Utilities4,6640.4%
Storage1,8830.2%

These 25 categories sum to 96.2% of the corpus (1,102,598 sites). The remaining 3.8% (43,123 sites) is honestly unclassifiable -- no confident category signal, not force-fit into the nearest wrong bucket. Counts are from stackra_us_corpus, the frozen US snapshot, re-verified 2026-08-09.

25 categories, 469 subcategories

Category is the top level. Beneath it, every classified site also carries one of 469 subcategories, so a restaurant is not just a restaurant but a sushi bar, a brewery, or a taqueria, and a retailer is a clothing store or a furniture showroom. That second level is what makes cohort benchmarking specific: your site gets compared against others doing the same thing, not against the whole 1.1 million. There are far too many subcategories to chart, but they are the reason a benchmark can say 'sites like yours' and actually mean it.

Where the sites come from

HTTP Archive's May 2026 crawl (mobile client, root pages only), queried through BigQuery, with a content floor of at least 1,000 bytes of HTML and 20 DOM elements to drop parked domains and empty shells. Excluded on purpose: platform freebie subdomains (WordPress.com, Wix.com), throwaway preview hosts with no custom domain (bare vercel.app, netlify.app), personal publishing hosts (Blogspot, Medium, Tumblr, Substack), institutional domains (.gov, .edu, .mil), and bare IP-address hosts. What is left is real, independently-hosted US business sites.

How a site counts as US

Not by domain ending or a locale guess: we tested the TLD-proxy approach and it caught only 32% of confirmed-real sites. Two real signals instead, neither treated as a ceiling on its own:

US-identification signals
SignalSites confirmedNotes
Overture Maps address match~744,500Physical US street address on file, open places dataset
US city/state text signal+442,808 net-newSite's own text, screened for false-positive city names (e.g. 'reading', 'paradise')

How we know the labels are right

A category map is only as good as the labels under it, so the labels got two rounds of checking. First, we removed every borrowed shortcut that looked authoritative but wasn't: a Google-Maps-embed rule, a schema.org Person tag, and a Squarespace-means-portfolio assumption between them mislabeled roughly 1.57 million sites before any AI pass ran, so we measured each one's real error rate by direct count and dropped all three. The categories the model kept returning off-list (financial services, funeral care, recreation, storage, transportation) were promoted to real types rather than forced into the nearest wrong bucket. Then we checked the finished labels two independent ways: a 504-row blind audit where a separate model graded raw page content with no category hints, and a 38-row hand-read of the actual pages behind the biggest label changes, which came back 33 to 35 out of 38 clean. Where both blind methods classified the same row, they agreed 97.4% of the time.

The three borrowed shortcuts we measured and removed
SignalAssumptionFalse positivesReality
Google Maps embed→ local business600,00084,000 real vs 600K wrong
schema.org Person tag→ personal portfolio738,000removed outright
Squarespace CMS→ portfolio233,00094% of Squarespace sites are actually retail

~1.57M combined false positives removed before any AI reclassification pass ran.

Where this stands, and what you get from it

This is the same corpus Stackra's category benchmarks and the link-authority research draw from, with every row keeping its prior label (business_type_previous) for audit. It is a reference dataset, not a live scan feature on its own: when you run a Stackra audit, your site is classified the same way and then compared against the category and subcategory cohorts measured here. That is how a benchmark tells you whether your site is ahead of or behind the specific slice of the web you actually compete in.

Frequently asked questions

How many US business websites are there, and how many did this study classify?
The frozen corpus behind this study holds 1,146,202 real, independently-hosted US business websites, drawn from HTTP Archive's May 2026 crawl after removing parked domains, empty shells, platform freebie subdomains, and institutional (.gov/.edu) hosts. Of those, 1,102,598 (96.2%) carry a confident business-category label; the remaining 43,123 are left honestly unclassified rather than force-fit.
What are the 25 business categories used to classify US websites?
Every classifiable US business site is sorted into exactly one of 25 categories: retail, restaurants, healthcare, nonprofits, education, B2B services, automotive, real estate, home services, entertainment, wellness and beauty, content and media, professional services, hospitality, recreation, religious, government, SaaS, financial services, funeral services, transportation, senior care, portfolio, utilities, and storage. Retail and restaurants alone make up a quarter of the corpus.
What is a subcategory, and why does it matter for benchmarking a website?
A subcategory is the finer classification beneath the 25 top-level categories, of which there are 469 total, such as sushi bar or brewery beneath restaurants, or clothing store or furniture showroom beneath retail. Comparing a site against its subcategory cohort, not the whole 1.1 million or even its whole category, is what lets a benchmark credibly say 'sites like yours.'
How do you verify that a US business is actually based in the US, without relying on the domain ending?
A .com/.org domain ending alone identified only 32% of confirmed-real US sites, so this study used two independent signals instead: a physical US street address matched against the Overture Maps open places dataset (about 744,500 sites), and a US city/state text mention on the site's own pages, screened for false-positive city names, adding 442,808 net-new confirmed sites.
What is a false-positive classification shortcut, and which ones did this study remove?
A false-positive shortcut is a rule that looks authoritative but mislabels sites at scale. This study measured and removed three: treating any Google Maps embed as proof of a local business (600,000 false positives), treating a schema.org Person tag as proof of a personal portfolio (738,000 false positives), and treating the Squarespace CMS as proof of a portfolio site when 94% of Squarespace sites in the corpus are actually retail (233,000 false positives). Together those removed about 1.57 million mislabels before any AI reclassification ran.
How was the accuracy of the business-category labels checked?
Two independent checks. First, a 504-row blind audit where a separate model graded raw page content with no category hints, agreeing with the production labels 97.4% of the time on rows both methods classified. Second, a 38-row hand-read of the actual pages behind the biggest label changes, which came back clean on 33 to 35 of 38.
LB
Luke Beck
Founder, Stackra
Last verified August 9, 2026

From the founder

Real tips. Occasional rants.

Dispatches from a founder driving a car while it's being built. No polish, no content calendar. Just what's actually working.

No spam. Unsubscribe any time.

Verdict

Want to see what bots see on your site?