elmerdata.ai blog

My blog

How Should We Measure AI?

Enterprise AI spending is rising faster than the evidence that it works. Spending, deployments, pilots, users, and tokens are easy to count. Institutional outcomes are not. The gap between the two is not an accounting problem. It is a governance problem.


The clearest picture yet of enterprise AI use arrived in February, and the most striking thing about it is what the respondents could not say. In NBER Working Paper 34836, Firm Data on AI, research teams at the Federal Reserve Bank of Atlanta, the Bank of England, the Deutsche Bundesbank, and Macquarie University surveyed nearly 6,000 senior executives across the United States, the United Kingdom, Germany, and Australia between November 2025 and January 2026. Adoption is no longer in question: 69% of firms currently use some AI technology, rising to 78% in the United States, and 75% expect to be using it within three years. Spending is following the same curve. A separate Atlanta Fed survey conducted in March found that per-employee AI spending rose from $1,358 in 2025 to an anticipated $2,068 in 2026, an increase of roughly 50% on an employment-weighted basis, though that average conceals a wide distribution in which more than half of firms plan to spend no more than $200 per employee while the top 10% plan at least $2,800.

Then comes the finding that should interest anyone who writes AI policy for a living. Asked about the effects of AI at their own firms over the preceding three years, roughly 9 in 10 executives reported no impact on employment and no impact on productivity. Adoption is widespread, spending is accelerating, and roughly 9 in 10 executives report no effect on the two outcomes economists know best how to measure.

There are at least two readings of that result, and the difference between them matters enormously. The first is that AI has not yet delivered, which would be an unremarkable finding about a general-purpose technology in its early years; electrification took decades to show up in productivity statistics. The second is that the effects exist but the organizations have no instrument capable of detecting them. Both are probably true in some proportion, and no one currently knows which proportion, because an institution that cannot demonstrate a gain generally cannot demonstrate its absence either. A measurement that returns nothing is not the same as a measurement that returns zero. The first is a fact about the world and the second is a fact about the instrument, and an organization unable to tell them apart is allocating capital in the dark while calling the result a finding.

National Bureau of Economic Research headquarters in Cambridge, Massachusetts

National Bureau of Economic Research headquarters at 1050 Massachusetts Avenue, Cambridge, Massachusetts, 2022. Photograph by Astrophobe via Wikimedia Commons. Licensed under CC BY SA 4.0.

Activity Is Not Outcome

Organizations are not failing to count things. They are counting the wrong things with considerable diligence. Tools purchased, pilots launched, users enrolled, queries submitted, tokens consumed, models deployed, and training hours completed are all tracked, reported, and presented to boards, and none of them demonstrates value. Dave Russell, who serves as SVP and Head of Strategy at Veeam, has described the pattern as confusing activity for outcomes, and while the phrase comes from a vendor with a commercial interest in the diagnosis, the distinction it draws holds independently of who is drawing it.

A chatbot fielding thousands of questions is active. Whether it shortens response time, improves accuracy, raises satisfaction, reduces workload, or produces better decisions is a separate question with a separate answer, and the first number tells you nothing about the second. Activity measures establish that an AI system exists and is being used. Outcome measures establish whether the institution is better off for it. Substituting the former for the latter produces operational dashboards that look like accountability and function as decoration.

Veeam's Data and AI Trust Gap report, based on a survey of 600 senior executives across five industries conducted in late March and early April, puts numbers on the substitution, with the standard caveat that it is vendor-sponsored research measuring a problem the sponsor sells against. Some 85% of respondents reported significant success from their data initiatives over the past year, and 45% had never formally measured the return. Only 29% included people and culture measures such as adoption, satisfaction, or retention when calculating value, the lowest of any category the survey tracked. The report also finds that just 7% of organizations qualify as genuinely ready by its own three-part standard of ambition, visibility, and governance, and that 97% of those report quantified business outcomes. The correlation is suggestive rather than causal, and Veeam does not disclose how many respondents are its customers, but the shape of it is hard to ignore: the organizations that built the measurement apparatus are the ones with something to report.

What the 85% and the 45% describe together is an institution that has concluded AI works before assembling the evidence that it works. Confidence is not measurement. The gap becomes consequential precisely when confidence is used to justify the next round of investment, which is the moment the Atlanta Fed spending curve describes.

The Missing Counterfactual

Underneath the reporting problem sits a methodological one that no dashboard can solve after the fact. Suppose an organization deploys an AI assistant and observes that employees now complete a process in 20 minutes. Is that good? The question cannot be answered, because the number has no comparison. If the process previously took 30 minutes, the system created substantial value. If it took 18, the system made things worse while feeling faster. If it took 30 but error rates have since doubled, the organization has purchased speed with accuracy and has no way to price the trade.

Without a baseline, nearly any observation can be narrated as success, and organizations under pressure to justify AI spending will narrate accordingly. The uncomfortable implication is that measurement cannot begin at deployment. It has to begin before it, when the current state is still observable and no one has an investment to defend. An organization that has already deployed a system without capturing a baseline has not merely delayed its evaluation; it has forfeited the ability to conduct one.

The same logic applies to the workflow as a whole rather than the AI component within it. An implementation can succeed technically and fail organizationally in ways that never appear in system telemetry. Employees quietly route around the tool, or use it incorrectly, or build parallel workflows that restore the old process alongside the new one. Managers overestimate reliability. Automation relocates work rather than eliminating it, and time saved in one step reappears as review work in the next. A 5-minute AI task that generates 10 minutes of human verification has not saved 5 minutes, and an organization measuring only the 5 minutes will record a gain it did not receive. The 29% figure suggests most organizations are not looking where the loss occurs.

Measurement Is a Data Governance Problem

A governance committee that reviews AI proposals for risk and compliance but not for evidence has done half its job. The questions that belong in that review are not exotic: what problem the system is meant to solve, what the current baseline is, which outcome should change and by how much before anyone calls it success, how the change will be measured, who owns the measurement, how often it will be reviewed, what evidence would justify expanding the system, and what evidence would require modifying or stopping it. Most AI governance frameworks handle the first few competently and stop before the last one.

Answering any of them depends on data the institution may not have in usable form. Veeam found that 95% of executives say data challenges have already slowed their AI progress, which fits a pattern familiar to anyone who has run a data function. Organizations frequently approach AI as a way to vault over unresolved data problems, and AI instead exposes them. Fragmented environments make baselines hard to establish. Inconsistent definitions make metrics incomparable across units. Incomplete records make before-and-after comparison unreliable. Weak ownership means no one is accountable when the numbers disagree. AI governance cannot operate at arm's length from data governance, because the evidence that AI governance requires is manufactured by data governance, and the two functions converge whether or not the organization chart reflects it.

That convergence raises the question of who owns the evidence. Technology leaders own systems, business units own processes, finance calculates return, and risk and legal oversee compliance, but credible measurement depends on definitions, lineage, provenance, baselines, a single source of truth, and metrics that return the same answer twice. The Chief Data Officer need not own every AI system to matter here; someone has to ensure the institution can distinguish an assertion from a finding, and it is difficult to see who else is structurally positioned to do it.

Financial return is only one axis, and for many systems the wrong one. Time saved, error reduction, service quality, decision quality, employee capacity, risk reduction, and accessibility are all legitimate measures, and the appropriate one follows from the purpose of the system rather than from a standard template. A university deploying AI to identify students who may need support should not be evaluated primarily on token costs or seat counts; the relevant measures are identification accuracy, how quickly interventions follow, how many false positives the system generates, what happens to the students it flags, and whether its decisions are equitable and explainable. Good measurement follows the mission, which means it cannot be procured with the software.

The Stop Rule

Autonomy raises the stakes on all of this. Veeam reports that 88% of organizations are already using or piloting AI agents while only 28% are confident they could detect a system operating outside approved parameters. A poorly performing assistant produces a bad answer; a poorly governed agent produces bad actions, repeatedly, at machine speed, which is why these systems need operational telemetry, exception rates, override counts, escalation behavior, and outcome measures reviewed continuously rather than annually. The faster a system acts, the faster governance has to be able to see what it is doing.

I argued in Context Engineering Is Governance at Inference Time that deciding what enters an AI system's context window is a governance act, because it determines which version of the institution is true at the moment the model answers. Measurement is the other half of that argument: context engineering asks what information the system receives, measurement asks what happened because it received it, and an organization that has done the first without the second has designed its input environment carefully and left the output unexamined.

Which brings us to the mechanism most conspicuously absent from AI governance frameworks. Organizations have elaborate criteria for approving AI projects and almost none for ending them. Every significant initiative should carry a stop rule defined before deployment, specifying in advance what would constitute failure: productivity gains that do not materialize, accuracy problems that persist past a defined threshold, costs that exceed projections, adoption that never arrives, risk that proves unacceptable, alternatives that render the system obsolete, or human review requirements that consume the efficiency the system was purchased to create. Writing the rule down in advance changes behavior, because it converts AI from a technology that must be retroactively justified into an investment that continues to earn its position.

Deployment itself is ceasing to be impressive. When 69% of firms use AI, using AI distinguishes no one, and organizations will separate themselves instead by knowing which systems work, under what conditions, and when they have stopped working. The first question executives asked was whether the organization was using AI, the question that followed was how much, and the mature question is what measurable difference it made.

Governance ultimately requires more than controlling what an AI system is permitted to do; it requires knowing what happened after we permitted it. Spending can be counted, usage can be counted, and tokens can be counted, and none of those establishes that an institution became better at anything. An approval rule without a measurement rule and a stop rule produces organizations that cannot distinguish a system that did nothing from a system whose effects they never learned to see, and an organization that cannot measure the consequences of its AI cannot credibly claim to govern it.

Further Reading

My earlier posts

Other sources


AI Assistance Statement ▾
Preparation of this blog entry included drafting assistance from ChatGPT using a GPT-5 series reasoning model. The tool was used to help organize ideas, propose structure, refine language, and accelerate revision. It was also used to assist in identifying image sources and verifying that selected images appear to be released for reuse (for example through public domain or Creative Commons licensing). The author selected the topic, determined the argument, reviewed and edited the text, confirmed image licensing, and takes full responsibility for the final published content.

#AIData #Observations