Naming an owner is easy, but measurement is hard. Change Futurist | September 3, 2026

Last week the answer was to name an owner. This week the question is what that owner gets measured on.

Tina Nunno and her Gartner colleagues published survey results on September 1 showing that roughly 11% of organizations cannot say what their own function spent on AI in 2025. Eighty-five percent of those same functional leaders plan to spend more in 2026.

An executive who owns a program and holds no instrument owns a title.

I have sat in enough quarterly reviews to know how this plays out. The CFO asks what the AI investment returned. The marketing leader describes activity, cites a vendor dashboard, and names a percentage that came from the vendor selling the thing under measurement. The meeting ends without a decision, and the budget survives one more quarter on momentum.

Four items published between September 1 and 2 mark out where the instruments are missing and what installing them costs.

Start by finding out what your function actually spent

Nunno's team surveyed 1,303 respondents at organizations with at least $50 million in fiscal 2025 revenue, fielding between January and April 2026. Only 22% have scaled AI across multiple business units or adopted an AI-first approach. Functional leaders put an average of 12% of their budgets into AI in 2025, and 85% plan to increase that in 2026.

Nunno states the risk directly: without disciplined measurement tied to business outcomes, organizations waste resources and miss expectations.

The finding worth carrying into planning season sits further down. Seventy-five percent of functional leaders named productivity as a target outcome, and productivity commands about 30% of functional AI spend, roughly double the next objective. Then compare what leaders pursue against what pays. The most frequently pursued use cases are cybersecurity threat detection and response at 54%, IT service desk automation at 54%, and automated code generation at 44%. The use cases where the largest share of leaders report positive returns are intelligent IT asset and cost optimization at 40%, synthetic data generation at 28%, and automated code generation at 23%.

One item appears on both lists. The rest diverge completely.

Nunno's guidance follows from that gap: leaders who track every dollar by outcome category defend their investments and reallocate quickly when projects underperform. That instruction assumes a ledger most marketing functions have not built.

Commentary: Nunno's team supplies the number that makes accountability enforceable, because a leader who cannot produce a spend figure by outcome category has nothing to be held to.

Shorten the distance between an insight and a decision

Emily Manock reported research on September 1 from two vendors that measures the same gap from inside the marketing function. Her weekly statistics roundup carries both.

Bitly surveyed marketers and found that 82% lack a clear view of what works in their campaigns. Fifty-six percent often or always decide from instinct or past experience. Seventy-three percent regularly discover that a campaign underperforms after the point where they can fix it. The average time from a new insight surfacing to a marketer acting on it runs five days. Respondents use an average of six tools to analyze performance, and more than a third use seven or more.

Six tools producing a five-day lag describes a measurement estate that costs money and returns opinions.

Validity surveyed 500 B2B and B2C marketers and found the downstream effect. Sixty percent of C-suite marketers and 52% of SVP and VP respondents feel pressure to deploy AI tools immediately while knowing their underlying data cannot support them. Seventy-eight percent of C-suite and 92% of SVP and VP respondents have acted on an AI recommendation they later suspected was wrong because the data feeding it was poor. Two-thirds of organizations handed more decisions to autonomous agents over the past year. Twenty-one percent describe their CRM data as "very well prepared" to support AI. Sixty-two percent report losing revenue directly to poor data quality.

Read those together and the sequence becomes clear. Organizations delegated decisions to agents, kept the data quality that produced wrong recommendations, and added five days to the interval where anyone could catch it.

Decision latency belongs on the same dashboard as accuracy. An agent that acts continuously compounds a five-day blind window into five days of executed decisions.

It reminds me of what Arianna Vogel, Senior Director of Product Marketing at Foursquare, said when I interviewed her on The Agile Brand podcast: "If you're really just thinking about a report card after your campaign, you're missing out on so many opportunities to optimize and drive efficiencies while your campaign is in flight."

Vogel names the design requirement. The instrument has to report on the timeline decisions get made, and a quarterly read cannot govern a system that acts hourly.

Commentary: Both studies come from companies selling measurement and data quality software, so treat the magnitudes as directional, and note that the five-day lag figure describes the same failure the Gartner survey found at the budget level.

Measure what customers do, and stop pricing what they say

Eric Keller published Gartner survey findings on September 2 drawn from 3,566 B2B and B2C customers surveyed in February and March 2026. The headline number: 27% of customers would try a chatbot again after a negative experience.

The number underneath it matters more for planning. Forty-nine percent of customers said they would have used a chatbot if the company had offered one. Seven percent actually used a chatbot or digital assistant in their most recent service interaction. Gartner also found customers roughly three times more likely to reach for a third-party generative AI tool than a company-provided chatbot, and 87% saying access to a human agent is essential when a company uses generative AI in service.

Any team that built its deployment case on the 49% built it on a number that overstates demand by a factor of seven.

Keller draws the operating conclusion: service leaders should prioritize reliability over reach. A bot that resolves a narrow set of issues consistently earns more use than one that attempts everything and fails often. He describes the 27% figure as a leaky bucket, where one bad interaction suppresses future use even as the underlying capability improves.

This lands squarely on the Trust layer in the capability framework I keep returning to. Under Orchestration, trust means an audit trail of what the agent actually did. Here the audit trail includes what the customer did next, and stated intent is the wrong instrument for reading it.

Agile Brand Principle 6 says to listen to customers and stop talking at them. A survey asking whether someone would use a chatbot measures politeness. Behavior after a failed interaction measures the thing you are spending against.

Commentary: Keller gives CX leaders a specific correction to apply to any deployment forecast built on stated willingness, and the seven-point gap between intent and behavior is large enough to invalidate most 2026 business cases.

Put your instrument where your CFO already looks

Amrit Virdi reported on September 1 that brands are plugging creators into marketing mix models and setting flexible KPIs to prove influencer effectiveness. Rightmove CMO Matt Bushby describes influencer as connected into the firm's MMM, showing what he calls brilliant long and short-term effects.

The move is unglamorous and correct. Influencer marketing resisted measurement for a decade because it produced engagement numbers that no finance team recognized. Bushby's team stopped inventing a channel-specific score and routed the channel into the model the business already uses to allocate budget.

Apply the same logic to AI. Most marketing organizations are building AI-specific dashboards that report adoption, prompt volume, assets generated, and hours saved. None of those appear on a P&L. The Gartner respondents who defend their investments track dollars by outcome category, which is a finance construct.

It reminds me of what Jim Sturm, President of North America at Capillary Technologies, said when I interviewed him on The Agile Brand podcast: "Every CFO tends to be a skeptic of loyalty initiatives in the beginning, but as long as the initiatives you've created can be directly attributed to the bottom line in the reports a CFO looks at, they become believers."

The operative phrase is the reports a CFO looks at. Build your AI measurement into the existing model, on the existing reporting cycle, in the existing units.

Commentary: Virdi documents a channel that solved its own credibility problem by entering the finance stack, and AI programs face the identical problem with a shorter runway.

Run the arithmetic on one program

A $1.4B specialty retailer spends $58M on marketing. Applying Gartner's 12% average, the AI line runs about $7M, split across nine initiatives on a single vendor invoice. Nobody meters per initiative.

Take one program. Lifecycle personalization spends $4.2M a year across paid retargeting and email, running 240 campaign cycles at 14 days each. That works out to $17,500 of spend per cycle.

Thirty percent of those cycles underperform target, which is 72 cycles. At the five-day detection lag Bitly measured, five of the 14 days run at a known-bad rate before anyone acts. Five-fourteenths of $17,500 is $6,250 per underperforming cycle. Across 72 cycles, $450,000 of annual spend runs inside a window where nobody can see it.

Cut the lag to one day. Exposure drops to $1,250 per cycle, or $90,000 across the same 72 cycles. Three hundred sixty thousand dollars of spend moves out of the blind window and into a window where a person can act on it.

Now price the instrument. Consolidate six analysis tools to two and fund half an analyst to build a daily read against a defined underperformance threshold. Call it $180,000 all-in.

That is a $180,000 investment producing a $360,000 change in exposure, on one program, inside one of nine initiatives. It is also the first number in this scenario that a CFO can evaluate. The other eight initiatives remain unmeasured, which is the actual finding.

Note what the arithmetic required. Spend per cycle, cycle count, underperformance rate, and detection lag. Four numbers. Most marketing organizations hold the first two and guess at the last two.

Do these four things before Q4 planning closes

Produce a spend figure by outcome category. Take your total AI spend and split it across productivity, revenue growth, risk mitigation, and innovation. If you cannot do that in a week, you have found your first project. Nunno's high performers can do it, and that capability is what lets them reallocate.

Measure detection lag as a named metric. Time the interval from a performance signal appearing to a human changing something. Put the number on the same page as your accuracy metrics. Five days is the benchmark to beat.

Rebuild every deployment forecast on observed behavior. Strip stated-intent figures out of your business cases. Gartner's seven-percent-against-forty-nine-percent gap is the correction factor, and any forecast built on survey willingness needs rerunning before it enters a budget.

Route AI measurement into the model finance already runs. Stop building a parallel AI dashboard. Put the spend and the return into your MMM, your CAC calculation, and your contribution margin reporting, on the reporting cycle your CFO already reads.

You named an owner last month. Give that person four numbers and a reporting line by the end of this quarter, and the ownership becomes real. Leave them with a vendor dashboard and an adoption percentage, and you have assigned a title.

Featured this cycle:

  1. Gartner Survey Finds Only 22% of Organizations Have Successfully Scaled AI Across Multiple Business Units. Tina Nunno, Gartner, September 1, 2026

  2. Inflation, in-game advertising, performance tracking: 5 interesting stats to start your week. Emily Manock, Marketing Week, September 1, 2026

  3. Gartner Survey Finds Only 27% of Customers Would Try a Chatbot Again After a Negative Experience. Eric Keller, Gartner, September 2, 2026

  4. 'Long-term value creation': Marketers on nailing influencer effectiveness. Amrit Virdi, Marketing Week, September 1, 2026

Featured Insights

Gartner, September 1. Only 22% of organizations have scaled AI across business units. Roughly 11% cannot say what their function spent in 2025. The most pursued use cases and the highest-returning use cases overlap on one item out of three. Build the spend-by-outcome ledger this quarter. It is the artifact that converts a named owner into an accountable one.

Marketing Week, September 1. Bitly reports 82% of marketers lack a clear view of campaign performance, with a five-day average lag from insight to action across six analysis tools. Validity reports 78% of C-suite marketers have acted on an AI recommendation they suspected was wrong. Count your performance tools and time your detection lag. Both numbers are available this week and neither requires a vendor.

Gartner, September 2. Forty-nine percent of customers say they would use a chatbot. Seven percent did. Twenty-seven percent would return after a bad experience. Pull every deployment forecast built on stated willingness and rerun it against observed channel usage before it enters your 2027 budget.

Marketing Week, September 1. Rightmove routed influencer marketing into its marketing mix model to establish long and short-term effects. Choose one AI initiative and put its spend and return into your MMM this quarter. A channel that enters the finance model gets funded.

Four Takeaways

  1. Spend visibility precedes accountability. An owner who cannot report spend by outcome category holds a title and nothing enforceable.

  2. Detection lag is a budget line. Five days between signal and action, multiplied across campaign cycles, converts directly into spend running at known-bad rates.

  3. Stated intent overstates behavior by a factor of seven. Rebuild forecasts on what customers did, and treat survey willingness as a ceiling nobody reaches.

  4. Measurement that lives outside the finance model gets defunded. Route AI reporting into the MMM, the CAC calculation, and the margin report your CFO already reads.

A Closing Note

I studied photography in college, before digital cameras were common, and spent a lot of hours in a darkroom. The discipline that stayed with me came from the test strip. Before committing an expensive sheet of paper to the enlarger, you exposed a narrow strip in bands at different durations, developed it, and read the result. Two minutes of work told you what the full print would do.

Photographers who skipped the test strip produced beautiful work occasionally and wasted paper constantly. They also could not explain why any given print succeeded, which meant they could not repeat it.

Marketing organizations are running full prints on nine initiatives with no test strips. The Gartner data says roughly one in nine cannot even report the cost of the paper. My second book holds that Intelligence tells you what is and what is likely, and that deciding what an organization should do stays a human responsibility. That division only functions when the descriptive half produces numbers a person can act on. An agent that scores a workflow and a dashboard that reports adoption both leave the normative decision unsupported.

Pick one program this quarter and build the strip. Four numbers, one reporting line, one name.

Previous
Previous

AI pricing moved from a seat to a meter. Change Futurist | September 7, 2026

Next
Next

Who owns your AI spend? Martech Futurist | August 31, 2026