Ask how the number was built before you spend against it. Martech Futurist | August 8, 2026

Tuesday's edition traced how the EU AI Act moved part of the verification cost off the internal decision sheet and attached it to the asset as a legal obligation. The cost stopped being a choice.

This post extends that in a direction with more immediate budget consequence. The systems generating the numbers marketers act on increasingly sit outside the enterprise: advertising platforms whose models set price and delivery, AI assistants whose answers decide whether a brand surfaces at all, and vendors selling measurement of those answers. Four items published between August 3 and August 7 converge on the same operational question. What has to be true about a number before you spend against it.

Nobody proposed opening the systems. They proposed grading the evidence.

Ask what evidence grade a number carries before you move budget behind it

The Interactive Advertising Bureau published Measuring Visibility in the AI Era on August 3, a 36-page framework covering brand and publisher visibility inside AI-generated answers. It arrives into a market IAB describes plainly: more than 20 companies sell AI visibility measurement tools, each using different query sets, platform coverage, and scoring rubrics, and two of them can measure the same brand in the same category during the same week and return materially different results. Sixteen percent of brands systematically track AI visibility today.

The framework's structural contribution is a two-tier split. Directional measurement supports trend monitoring, early signal detection, and internal briefings. IAB states directly that it does not clear the bar for budget allocation, provider selection, or executive strategy. Decision-grade measurement clears a higher standard across nine dimensions including query volume, sample size, prompt-type coverage, testing cadence, reproducibility, and platform coverage. One threshold is numeric: fewer than 50 queries per program counts as exploratory, below even directional.

Two requirements carry the most weight for a marketing leader.

  • Providers must report hallucinated mentions and factually inaccurate mentions separately, per platform, and must surface flagged mentions to clients. A vendor that quietly filters hallucinations out of a mention rate produces a cleaner number and hides a reputational exposure. IAB instructs buyers to treat a missing detection methodology as a material gap.

  • Where a provider will not disclose against a required item, buyers should read that absence as the answer. That single instruction shifts negotiating leverage without requiring anyone to build methodological expertise in-house.

Underneath both sits the property that makes this hard. AI platforms do not produce deterministic output. IAB's position is that a brand's visibility on a given query is a distribution, and any metric derived from one response per query reflects a sample of one. Similarweb research from November 2025 found citation sets changing by roughly 50% each month with only 11% overlap between major platforms.

This lands squarely on the Trust and Verification layer in the capability framework I use, in its Insights form: is the pattern real, or a confident artifact of how somebody built the query set. The answer now has a documented shape.

It reminds me of what Jennifer Griffin Smith, Chief Market Officer from Acquia, said when I interviewed her on The Agile Brand podcast: "We spent 20 years optimizing for search ranks and the next five years are going to be about winning the answer battle." The battle has a scoreboard now. Read the scoring rules before you accept the score.

Buy independent measurement for the platforms that carry your budget

Gartner published a forecast on August 6 putting more than 70% of global ad spend and 80% of U.S. ad spend through self-serve advertising platforms where AI materially influences media buying, cost, and outcomes by 2028. Eric Schmitt, VP Analyst in the Gartner Marketing practice, draws a distinction worth carrying into planning conversations: this back-office AI, embedded in delivery, audience selection, and pricing, operates separately from the generative AI teams use to make creative.

Schmitt's framing of the economics is the part CMOs should sit with. Better platform economics does not automatically become lower cost for the advertiser. Gartner's recommendations run to concentration and evidence: prioritize the platforms most strategically important to the business, curate a portfolio of additional platforms selectively, and direct investment toward platforms that support transparency and independent evaluation of results.

Schmitt puts the causal link directly: the more influence AI has over advertising decisions, the more important independent measurement becomes.

Read that alongside the IAB framework and a pattern emerges. Two organizations, working on different surfaces, arrived at the same instrument. Neither asked for access to the model. Both asked for evidence a buyer can grade.

It reminds me of what Chris Golec, Founder & CEO from Channel99, said when I interviewed him on The Agile Brand podcast: "Clicks feel good for marketers, like paid search, where you only pay for clicks, but 85% of those are from people or companies that will never buy anything. Clicks are probably 10% of the signal online, so you have to capture not only what's happening on your website but what's happening off your website." The number a platform reports and the business impact a CFO recognizes have never been the same number. AI-influenced buying widens that gap and removes the line of sight teams used to close it.

Check who supplies your measurement and how much of the surface it covers

Cloudflare released its AEO Visibility Dashboard on August 6 in early access, reporting Citation Rate, Mention Rate, Prominence, and Share of Voice for a site within its inferred industry category, alongside per-operator crawl and referral traffic drawn from Cloudflare's own network. Stephanie Cohen, Cloudflare's chief strategy officer, framed the gap it addresses: brand recommendations now happen at scale inside AI answers where site owners cannot see them.

Three things about this launch matter more than the launch.

The vocabulary matches IAB's Presence and Prominence metrics almost exactly, three days after publication. That speed suggests the standard will hold.

The data comes from the network layer, meaning real crawl activity and real referral traffic instead of simulated queries. Under IAB's disclosure framework, that is a distinct data collection architecture with its own validity profile, and outputs from different architectures are not equivalent without disclosure.

The citation probing covers two assistant families, Claude and GPT, according to independent analysis published August 7. Gemini, Perplexity, Copilot, and Grok sit outside it. IAB's decision-grade bar requires platforms collectively representing a substantial majority of consumer AI traffic in the target market, with weighting and market share basis disclosed. ‍

Apply the framework to the product and you get a usable read within a day of launch. That is the framework working as intended, and it works the same way on the vendor you already pay.

Here is the worked version. Take a CMO with $40 million in working media and a $180,000 annual AI visibility contract up for renewal. Under the old process, the renewal conversation covers dashboard features and price. Under the new one, procurement sends four questions before the meeting: how many queries per measurement program, and what is the distribution across informational, comparison, recommendation, and transactional intents; which platforms, weighted how, on what consumer share basis; what variation do you observe across identical repeated queries inside a seven-day window; and do you separate hallucinated mentions from factually inaccurate ones, per platform. A vendor supplying 40 queries against two platforms with no variability baseline is selling directional data. That data still has a job, in competitive awareness and early signal detection. It cannot carry a budget reallocation, and now the CMO has a published standard to say so without relitigating methodology in a renewal meeting.

Set your own productivity baseline before you accept a vendor's number

Three senior media executives working across Asia-Pacific graded their own organizations' AI performance at five or six out of ten on a panel published August 6 from ATS Singapore. Eileen Ooi, APAC president of PHD at Omnicom, put agent coverage at roughly 10 to 20% in overall productivity terms, concentrated in insights reporting. Ganga Chiravurri, president of product and solution development for APAC at dentsu, put combined true-agent and AI-assisted-tool coverage at a good 15 to 20% of work performed.

Set that against the survey backdrop. Research covered in July recorded 87% of marketers believing their organization uses generative AI effectively, against a MiQ survey where 45% felt confident. Three executives with every commercial reason to describe agentic AI as transformative marked themselves at six.

Two operator observations from that panel deserve wider circulation.

Ooi named the organizational risk as outsourcing thinking and taking outputs for granted, and said the concern sharpens in optimization agents handling millions of dollars. She described human sign-off before any campaign goes live as non-negotiable at PHD, on the reasoning that the agents were built by humans and gaps still surface.

Chiravurri drew a definitional line the industry blurs constantly. Putting a process on top of a language model produces an AI-assisted tool. A genuine agent requires independence in the orchestration layer, the ability to hold embeddings, fine-tuning against proprietary documents, and the capacity to advance a workflow without step-by-step handholding. ‍

The panel could not close on cost. Ooi said the industry underdiscusses it and called for a conversation about blended cost, noting that technology costs historically absorbed by the agency now need a different commercial arrangement. Omnicom reported third-party service costs of almost $2.9 billion in the first half of 2026 against $1.7 billion a year earlier. Chiravurri described calibration as poor on both sides, offering the example of a one-sentence question about 500 years of history returning 60 pages of billed output.

That maps to two tiers of my capability framework at once. Compute economics is a Force, external pressure that bends the system from outside. The 10 to 20% coverage figure sits in Orchestration, and it stops at the point where somebody has to check the work.

Agile Brand Principle 7 asks for the humility to acknowledge room for improvement. Three people scoring themselves at six, in public, at a conference, did more for the industry's evidence base this week than another vendor survey reporting 87%.

Featured Insights

IAB | August 3, 2026Measuring Visibility in the AI Era (detailed analysis via PPC Land, August 4) Practitioner takeaway: Insert the nine decision-grade criteria into your next measurement RFP verbatim. The disclosure format converts vendor comparison from a claims exercise into a document exercise, and a provider's refusal to answer becomes usable procurement evidence.

Gartner | August 6, 2026Gartner Predicts More Than 70% of Global Ad Spend Will Flow Through AI-Influenced Self-Serve Advertising Platforms by 2028Practitioner takeaway: Rank your platforms by strategic weight, then fund independent measurement in proportion to that weight. Platform-reported performance and business impact diverge as AI takes more of the buying decision.

Cloudflare | August 6, 2026Cloudflare Adds AEO Visibility Dashboard to Its AEO SuitePractitioner takeaway: Network-layer measurement gives you crawl and referral data your vendors cannot see. Confirm platform coverage before you retire an existing tool, since two assistant families leaves most of the surface unmeasured.

ATS Singapore 2026 panel, via PPC Land | August 6-7, 2026Omnicom's PHD gains just 10-20% productivity from AI in APACPractitioner takeaway: Use 10 to 20% coverage, concentrated in reporting and pre-buy analysis, as your planning benchmark. Then settle who pays for inference and on what basis before you sign the next agency agreement.

Key Takeaways‍ ‍

1. Evidence grade is now a budget control. IAB's directional and decision-grade split gives marketing leaders published language for refusing to reallocate spend against a number built from 40 queries and one platform. Write the criteria into procurement this quarter.

2. Independent measurement scales with AI influence. Gartner puts more than 70% of global ad spend through AI-influenced platforms by 2028. Concentrate spend on platforms that support independent evaluation, and fund verification as a line item against the platforms carrying the most budget.

3. Ask who built the measurement and what it covers. Cloudflare's dashboard draws on real network signals and probes two assistant families. Both facts change what the number can support. Every measurement source has an architecture and a coverage boundary, and both belong in the report.

4. Practitioner-reported returns run well below survey-reported returns. Senior agency executives put agent coverage at 10 to 20% and scored themselves at six out of ten, with human sign-off retained before launch. Plan against the operator number.

A note from Greg

I have sat on both sides of a diligence table. What I learned buying and selling companies is that the number matters less than the process that produced it, and the fastest way to find out whether somebody believes their own figure is to ask how they built it. Most of the time the question itself does the work.

Twelve years of piano lessons taught me the same lesson in a different register. My teacher never asked how the piece sounded. She asked what I had practiced that week, and how many times.

What happened this week is that our industry finally wrote down the questions. The IAB criteria, Gartner's push toward independent evaluation, and three agency leaders scoring themselves honestly at a conference all point at the same discipline. Intelligence tells you what is and what is likely. Deciding what to spend against it stays a human judgment, and it always will.

Ask how the number was built. Then decide.

Previous
Previous

Evidence Expires. Martech Futurist | August 10, 2026

Next
Next

Verification just became a property of the asset. Martech Futurist | August 5, 2026