Strategy and diligence engagements run on evidence that has to hold up — to a partner review, to a client's board, to the other side's advisors. Cookieless Audience gives consulting teams a deterministic audience profile for any of 102 million domains: coded demographics, 285 sub-interests, 283 purchase-intent segments, B2B firmographics and 1,667 personas, every value drawn from fixed v1.0 vocabularies aligned with IAB Audience Taxonomy 1.1. The same query re-run by anyone, on the same file, returns the same figure — which is what “auditable” actually means. And it is available by instant download, on engagement timelines, not procurement timelines.
When an engagement touches a media property, a content business or anything whose value rests on “who the audience is”, the evidence base gets thin fast. Management presentations assert an audience (“affluent, professional, decision-makers”) that was itself assembled for persuasion. Expert calls give directional color that cannot be quantified or reproduced. Traffic estimates say how many but never who. And survey work — the rigorous option — takes longer than the exclusivity window of most deals and the sprint cadence of most strategy studies.
The structural issue is reproducibility. A diligence finding is only as strong as the answer to “how do you know, and could we check?” Probabilistic audience estimates from tracking-based vendors change between pulls, depend on opaque models, and degrade precisely where measurement is weakest — the roughly 40%+ of traffic on Safari, Firefox and iOS where third-party cookies are already blocked (they remain supported on Chrome). An analysis a client cannot re-run is an opinion with a chart.
Content-derived, deterministic classification changes the character of the evidence. Each domain's profile is computed from what it publishes, every attribute takes a value from a fixed vocabulary — documented in full on the taxonomy page — and carries a banded confidence level (low / medium / high). The workpaper can state the exact rule applied; the reviewer can apply it again and reconcile to the digit.
Any workstream that needs a defensible answer to “who does this property, partner or market actually reach?”
Independently verify a target's audience claims — publisher, content network, affiliate portfolio, community — against coded profiles rather than the CIM's own narrative.
Map a category's media landscape — who serves which audience segment, where the white space is — from a census-style corpus instead of a hand-built site list.
Audit where a client's spend and partnerships actually land: profile the domains on the media plan and score them against the client's own target-audience definition.
Re-run the diligence queries each quarter on refreshed data to track whether the audience thesis is playing out — same vocabulary, same rules, comparable numbers.
Built to fit inside a two-to-six-week workstream. No integration project, no panel fieldwork — a file, a join, and a rules pass.
Translate the claim under test into vocabulary terms. “Affluent professional audience” becomes income_level: high|affluent, life_stage: established_professional, b2b_seniority: director+ — testable, not rhetorical.
Instant-download the top-1M file ($1,990) or a vertical slice ($190–$490) the day the engagement starts. Spot-check specific properties and sections with the real-time API where page-level granularity matters.
Join the target's domains (or the market's) to the file and apply the coded rule at a stated confidence threshold. Report the strict cut (high confidence only) and the broad cut (medium+) side by side.
Deliverable cites vocabulary version, selection rule and threshold. Anyone with the same file reproduces every exhibit — the finding survives partner review, client challenge and the other side's advisors.
Illustrative scenario: a client is acquiring a personal-finance content group. The information memorandum claims “an affluent, investment-active professional audience.” Left, the claim coded as a testable rule; right, the flagship property's profile from the data.
income_level: high | affluent life_stage: established_professional PI.finance_insurance.stocks_and_investments threshold: medium+ vocab: v1.0
Applied to all 14 domains in the target portfolio; exhibit reports the share of properties satisfying the full rule at the stated threshold, plus each property individually.
Partially substantiated. Investment intent and professional profile confirm at high/medium confidence — but income resolves to upper_middle, not high | affluent as claimed. Across the portfolio, 9 of 14 domains pass the full rule. Material for pricing the advertising-revenue thesis; the sell-side's “affluent” language overstates the modal audience.
Note what the exhibit does not depend on: no panel extrapolation, no vendor black box, no expert's impression. The client's own analysts can re-run the rule after close — and each quarterly refresh turns the same query into a value-creation tracking metric.
Each source answers a different question. The coded corpus is the only one that is simultaneously fast, quantitative and reproducible.
| Source | Speed | Reproducible? | Best used for |
|---|---|---|---|
| Management materials / data room | Immediate | Asserted, not testable | The claims to be tested, not the test |
| Expert interviews | Days | Directional, unquantified | Context, hypotheses, sanity checks |
| Commissioned surveys | Weeks–months | Partially (with full methodology) | Demand-side questions when the timeline allows |
| Traffic / clickstream estimates | Immediate | Model-dependent, varies between pulls | Scale ranking; says how many, not who |
| Coded domain-level audience data | Instant download | Deterministic, versioned, confidence-banded | Testing audience claims; mapping markets; quarterly tracking |
For firm-wide use — multiple engagement teams querying the corpus, larger extracts from 5M up to the full 102M domains, or feeds into internal benchmarking tools — custom licensing is available (from $15,000/year); contact us for a quote. Single engagements are usually covered by self-serve tiers.
Yes — the top-1M file ($1,990) and the top-100k file ($490) are instant card checkout with immediate download, and vertical or country slices ($190–$490) ship the same way. A team can be joining target domains against coded profiles on day one of the workstream. Only large custom extracts (5M+ domains) go through a quoted process. Individual properties can be checked immediately, for free, in the live demo.
That is the design goal. Classification is deterministic — the same domain and data version always yield the same profile — and every attribute value comes from a fixed, versioned vocabulary (v1.0, aligned with IAB Audience Taxonomy 1.1) with a banded confidence level. Your workpaper states the selection rule, threshold and vocabulary version; any party with the same file reproduces the exhibit exactly. Disagreement then has to be about the rule, which is a substantive argument, not a data-quality one.
No. The database is a flat, documented file keyed on domain — a join and a filter in Excel, SQL, R or Python, well within a normal analyst's toolkit. The coded fields (age_bracket, income_level, INT.*, PI.*, personas, firmographics) are enumerated on the taxonomy page. The real-time API (plans from $99/month for 10,000 credits) is there when a workstream needs page-level profiles of specific URLs.
Profiles are derived from each domain's published content, not from tracking users, so they are identical for Chrome, Safari, Firefox and iOS audiences. Tracking-based estimates under-observe the roughly 40%+ of traffic where third-party cookies are already blocked (Safari, Firefox, iOS — Chrome still supports them); a content-derived profile has no such blind spot, and no PII is processed at any point — which shortens the client's own privacy and procurement review.
The same corpus powers the neighboring diligence and research workflows.
Check any target domain in the live demo now — then license the file that covers the workstream and hand the client an analysis they can re-run themselves.