
Before NASA sends anything to space, they test it relentlessly. The rocket, the heat shield, the communication systems; every component is stress-tested against failure conditions before a single astronaut straps in. The evidence produces confidence in the tool.
Most organizations deployed AI the other way around. The tool went live and the rollout was declared a success. The mission launched, but nobody checked if the heat shield worked.
Victory was declared, but do we know the consequences of victory?
The measurement frameworks most organizations have in place were designed to track activity, not decision quality. That distinction matters because adoption metrics don’t show if AI is improving the decisions that matter, and retrospective evaluation cannot reliably fill that gap. What follows is a case for methodologically sound measurements that companies should establish before their next AI deployment.
Most organizations that have deployed an AI tool in the last few years are measuring success with metrics such as adoption rate, utilization, and time saved per task, irrespective of cost and time to build. These are the metrics that made it into the dashboard and reported upwards each quarter.
They also may be the wrong ones for your company.
Activity metrics tell you whether people showed up. They don't tell you whether anything improved. Tracking AI utilization to evaluate AI value is like counting how many times someone opened the fridge to evaluate their nutrition. Put simply, presence isn't impact and motion isn't inherently progress.
The measurement problem runs deeper than metric selection. The more uncomfortable truth is that most organizations never set up the conditions to measure the right thing in the first place.
Think about how AI tools were deployed at many companies. A vendor made a compelling case, procurement signed off, IT stood up the infrastructure, and a rollout plan was built. Then, what is potentially, a watershed moment in tech, was just simply launched.
What didn't happen: a hypothesis. A baseline. A counterfactual. Any meaningful answer to the question, “compared to what?”
In other domains, this would be considered irresponsible. Clinical trials don't declare drug effective simply because patients reported feeling better. We've built entire disciplines around the simple principle that outcomes need comparison conditions to mean anything.
AI got a pass at too many companies. Companies treated deployment as a launch event, not an experiment. Now CFO’s are trying to evaluate results from an experiment that was never designed.
The reason this is particularly hard to fix is the outcome that matters most is decision quality, and almost no one measures it.
Efficiency metrics are visible and fast. Decision quality is slow, diffuse, and harder to attribute. It shows up months later, in a strategic call that went differently, a forecast that proved accurate, or a risk that got caught early. That kind of value doesn't show up cleanly in a quarterly dashboard.
Even when organizations try to measure AI decision quality, they run into a deeper problem. Namely, human self-assessment of whether AI helped them decide better is systematically unreliable. Furthermore, recent research suggests it may be getting worse, not better. A 2025 study out of Aalto University found that people using AI consistently overestimated how well they performed, and that the most AI-literate participants had the least accurate self-assessments [link]. Familiarity with the tool didn't sharpen judgment, instead it amplified overconfidence. When you ask someone, "Did the AI help you?", you're not measuring decision quality. You're measuring how good they feel about it, if anything at all, which is a different thing entirely.
This is why measurement must be taken seriously from the start.
For many organizations, the foundational AI experiment wasn’t designed the way it should have been, or maybe even at all. That’s not an indictment. It simply reflects how quickly the technology moved and how little guidance existed at the time. But as the great Maya Angelou put it: “do the best you can until you know better. Then when you know better, do better.” Though, something tells me she didn’t realize this sentiment would be paired with rapidly emerging technology.
We know better now. And the next opportunity is already approaching.
Agents, decision automation, real-time recommendations embedded directly into workflows; that next wave of AI capability is already in the vendor pipeline. When that conversation arrives, the instinct will be familiar: build the business case, get procurement aligned, stand it up, launch. The difference this time is that experience is available as a guide. The gaps in that sequence are documented. And the cost of leaving measurement out is no longer theoretical.
The organizations that get the next deployment right won't necessarily have more resources or more sophisticated tools. They'll have made a different decision at the beginning: to define what success looks like in clear terms before anyone touches the technology. That decision is still available. The window hasn't closed. It just requires making the choice before the launch, not after.
Before the next deployment, define what success looks like in outcome terms; not adoption, not utilization, but decision quality. Identify what you'd need to observe, and over what epoch, to know it's working. Establish and advocate for a control group, even a small one. You’re not hamstringing your company; you’re attempting to actually measure this large and increasingly expensive investment your company continues to make. A hypothesis, clear metrics, and something to compare to are a small price to pay to avoid being wrong.
To make this concrete imagine a financial services firm preparing to deploy an AI tool that supports loan officers in credit risk assessment. Before launch, they define success in outcome terms; whether the loans they approve perform better over a 24-month period. They identify a small control group of loan officers who will continue working without the AI input during the pilot. They document the current approval rate, default rate, and average risk score as their baseline. The hypothesis is explicit: AI-assisted loan officers will produce a portfolio with measurably lower default rates over two years, holding loan type and market conditions constant. This allows the organization to actually answer the question they will most certainly be asked, “Was our investment worth it?”.
The organizations that get this right will move beyond knowing whether their AI tools are working. This matters more urgently as AI moves from task assistance toward decision influence. The tools organizations are deploying today are increasingly embedded in consequential calls, business critical decisions like forecasts, resource allocation, customer decisions, risk assessments. That shift changes the measurement stakes fundamentally.
The KPIs you set now don't just evaluate past performance. They feed back into decisions about where to scale the technology next, which use cases to expand, where to invest further. Organizations that are measuring adoption will scale adoption. Organizations that are measuring utilization will scale utilization. Neither of those is the same as scaling good decisions, and the gap between them widens as the technology becomes more deeply embedded in how the business actually runs.
Goodhart's Law is blunt about what happens next - when a measure becomes a target, it ceases to be a good measure. If we apply that logic here, if AI adoption is what gets reported, celebrated, and funded, then AI adoption is what you'll get. The feedback loop will perform exactly as designed. The question you should ask yourself is, is that the right metric to hang your company’s proverbial hat on.
Defining the right KPIs before the next deployment is the mechanism by which an organization ensures that as AI scales its influence over decisions, the feedback it receives is about decision quality. Not activity, not sentiment, not utilization. Without that, you could be accelerating in the wrong direction with increasing confidence that you're right on track.
New technology will keep coming, and that is genuinely exciting. The pace of AI development means that every organization has another opportunity in front of them. The question is whether the enthusiasm that drives adoption also drives the discipline to measure it — because wouldn’t it be better to know the ROI on your ideas?
That discipline doesn’t necessitate complexity. But intention is non-negotiable – a hypothesis, a baseline, something to compare to, defined before anyone touches the technology. That is what turns excitement into evidence and evidence into decisions an organization can iterate on.
Not sure on your next step? We'd love to hear about your business challenges. No pitch. No strings attached.