Market research firms use Cookieless Audience as a census-style input: a pre-computed audience profile for each of 102 million domains, coded in fixed, versioned vocabularies aligned with IAB Audience Taxonomy 1.1. Instead of extrapolating a media landscape from a panel that only resolves the largest sites, you aggregate coded demographics, 285 sub-interests, 283 purchase-intent segments and 1,667 deterministic personas across every domain in a category — and every number in the deliverable traces back to a versioned code a client can audit.
The standard inputs for a media landscape study — consumer panels, syndicated audience measurement, survey recall — share one structural limit: they only resolve the head of the web. A panel of tens of thousands of people produces stable audience estimates for a few thousand large sites; below that, the sample thins out and the long tail of a category becomes statistically invisible. Yet in most categories the long tail is precisely where the interesting structure lives: the specialist publishers, the niche communities, the emerging players a market map is supposed to surface.
Tracking-based alternatives have their own problem: they depend on third-party cookies and device identifiers, and roughly 40%+ of web traffic — Safari, Firefox and iOS — is already cookieless (Chrome still supports third-party cookies, but privacy regulation keeps adding pressure). A measurement approach anchored to tracking under-observes an increasingly large slice of exactly the audiences a study is meant to describe.
Content-derived audience classification takes the opposite route. Every domain is profiled from what it publishes, using the same fixed vocabularies — 8 age brackets, a 5-point gender skew, 6 income bands, 7 education levels, 14 life stages, household composition, employment, home ownership, urbanicity, plus coded interests (INT.*), purchase intent (PI.*) and B2B firmographics. Coverage is uniform across the head and the tail, no panelists are involved, and no PII enters the pipeline at any point. For a research firm that means one thing above all: the denominator of a category study can finally be the whole category. The full field list is documented on the audience segmentation taxonomy page.
The dataset behaves like a structured census of web properties — so it slots into the study types you already sell.
Enumerate every domain serving a category, then cluster by persona, interest mix and demographic profile to show who competes for which audience — head to tail, not just the top 50 sites a panel can see.
Count and segment the supply side of a content category: how many properties address a given intent segment, at what confidence, with what demographic tilt — a defensible structural size for the media landscape.
Describe how a market's attention is distributed: which audience segments are over-served or under-served by existing publishers, and where white space exists for a client's product or content play.
Quarterly refreshes on stable v1.0 vocabularies mean the same query re-run next quarter measures the same thing — category drift becomes observable instead of being an artifact of methodology churn.
A landscape study on this data is a sequence of reproducible queries, not a sequence of judgment calls. Four steps, whether the universe is a vertical slice or the full corpus.
Select domains by coded attribute — e.g. every domain carrying INT.home_garden.home_improvement or PI.home_garden_services.home_improvement_and_repair — optionally filtered by country slice. The selection rule is itself part of the methodology section.
Cross-tabulate the universe on any field: age_bracket, income_level, life_stage, urbanicity, persona. Confidence bands (low / medium / high) let you report a strict cut and a broad cut side by side.
Cluster domains by attribute similarity to reveal sub-markets, position named players on the map, and quantify white space — audience cells with demand signals but few dedicated properties.
Every figure cites its vocabulary code and version. A client — or a peer reviewer — can re-run the query on the same file and get the same number. Deep-dive individual properties with the real-time API where a page-level view is needed.
Illustrative scenario: a firm is mapping the home-improvement media landscape for a tools manufacturer. Left, the aggregate distribution across the selected universe; right, one domain from the map, as it appears in the data.
home_ownership: owner urbanicity: suburban income_level: upper_middle life_stage: family_young_children vocab: v1.0
This property sits in the category's densest cell — suburban owner-families, 35–44. The map's white space is elsewhere: the aggregate shows an 18% share of domains skewing 25–34, mostly renters (home_ownership: renter), yet few properties serve that cell with dedicated rental-friendly improvement content — a findable, codeable gap the client can act on.
The point of the exercise is not any single profile — it is that both panels of this example come from the same file, same vocabulary, same version. The aggregate and the property-level view reconcile by construction, which is what makes the finding citable rather than anecdotal.
Domain-level audience data does not replace panels or surveys — it covers the questions they structurally cannot answer, and it is the only one of the four that scales to the full web.
| Input | Strength | Structural limit | Role in a landscape study |
|---|---|---|---|
| Consumer panels | People-level behavior on large sites | Sample thins below the head; long tail invisible | Reach validation for the largest mapped properties |
| Surveys | Attitudes, awareness, stated preference | Recall-based; cannot enumerate a media landscape | Demand-side context layered onto the map |
| Clickstream / traffic estimates | Volume ranking of sites | Says how many, not who; tracking-dependent | Weighting the map by scale |
| Coded domain-level audience data | Uniform audience attributes across 102M domains, versioned and reproducible | Describes properties and their audiences, not individuals | The census frame: universe definition, segmentation, white-space analysis |
Licensing follows study scope: the top-100k file ($490) covers head-of-market studies, the top-1M file ($1,990, instant checkout) covers most national landscapes, vertical and country slices run $190–$490, and multi-country or full-corpus programs — 5M up to all 102M domains — are quoted individually via contact. Quarterly refreshes keep trackers on a consistent baseline.
It answers a different question with different machinery. Panels observe a sample of people and estimate audiences for the sites large enough to register in that sample. This dataset profiles the properties themselves — every one of 102 million domains gets the same coded attribute set derived from its content — so coverage is uniform from the head to the long tail. Most research teams use both: the coded corpus as the census frame that defines and segments the universe, panel data to validate reach on the largest properties within it.
Yes, and it is built for exactly that. Every attribute value comes from a fixed, versioned vocabulary (v1.0, aligned with IAB Audience Taxonomy 1.1), every attribute carries a confidence band, and classification is deterministic — the same input yields the same output. A methodology section can state the universe-selection rule, the vocabulary version and the confidence threshold, and anyone with the same file can reproduce every figure. The full vocabulary is public on the taxonomy page.
A single-category study is usually covered by a vertical slice ($190–$490) or the top-1M file ($1,990, instant download), which resolves most nationally relevant properties. Multi-market programs, syndicated products or trackers that need the deeper tail license larger extracts — 5M domains up to the full 102M corpus — under custom terms, quoted individually. Quarterly refresh options ($190–$590 depending on tier) keep longitudinal studies on a stable baseline.
No. Profiles are derived from the published content of each domain, not from tracking individuals — no cookies, no device IDs, no panelists, no PII anywhere in the pipeline. That also means coverage does not degrade in cookieless environments (Safari, Firefox, iOS — roughly 40%+ of traffic), which increasingly distort tracking-based inputs. For research firms this simplifies both procurement review and the privacy statement in the final report.
The same coded corpus supports the neighboring study types — and the industries that commission them.
Open the demo and read the coded profile of any domain in your category — then license the slice that covers your universe and run the whole map in one pass.