There are several ways companies can check whether AI is working for them. The 3 most common ones, in order of how often i have seen them around last couple of months are:
asking employees whether the tools made their work easier
measuring how long a piece of work takes from start to finish, and
looking at whether a cost or a revenue figure changed
The first check is definitely the easiest and quickest, but from what i am seeing around, it also tells you the least about org level outcomes.
METR surveyed 350 tech workers earlier this year and asked how much faster did AI make them - the median answer was 3 times. METR’s own staff reported the smallest gains of any group, and METR thinks that this was because the staff knew its earlier research on the difference between reported and measured productivity. In the 2025 version of the same trial the developers said AI made them about 20% faster, and the measured time showed they were 19% slower. METR also noted that surveys have consistently reported larger AI gains than various field experiments.
There are various reasons people may overstate without deliberately lying.
Some tasks are now very cheap to do with AI, and people end up doing stuff they would never have bothered with earlier. i am sure someone in your team has shared a dashboard for a dataset that did not need one… or a fancy Claude artefact, built to impress a Slack group, that nobody uses afterwards. Each of those “tasks” would have taken hours by hand, so people count them as hours saved. Now sum those hours across everyone in the pilot and the total looks huge, even though the work that really mattered for the business may take about the same time.
People also go by how “tiring” a task felt when they guess how long it took. AI takes away most of the boring bits like the first draft of a 2027 planning doc, but you still have to read, review and fix. Because the boring bit is gone the task feels shorter, even when the combined process of creation + verification could take way longer than going ahead and writing it yourself.
People also feel productive when they make more stuff, because most of us remember how much we made but don’t really care (or know) to check if it was any good. That’s also because in our brains, we can see how much we generated, but have a harder time judging how good it was. The person who makes a deck feels good that day, and the problems come up a week later with whoever has to use it, and by then nobody links it to the tool. In a UK government trial of Copilot, slides got made 7 mins faster at 50% the quality score but 72% of the users still said that they were “satisfied”.
The second check measures how long the work took pre/post AI adoption. With pretty much any meaningful process in an org, the “minutes” someone spends working on an item are a small fraction of the “days” that same item spends in the system. That is because the rest of the time goes to queues, handoffs and approvals. So even if every working step gets faster, the item still sits in the queues, and extrapolating the velocity of a step to a process is a misattribution I am seeing everywhere.
Most of the satisfaction surveys ask people about the steps, which get reported satisfactorily, while the delay is usually between them.
The third check is whether the business made or saved money. In McKinsey’s updated 2026 state of AI survey, 80% of respondents said AI had improved their own productivity and 6% said it accounted for 5% or more of EBIT (which is unchanged from 2025 after another year of spending).
Of that 6%, about three quarters had redesigned their workflows and they were twice as likely to have a defined way of measuring what AI had done (a very interesting discourse here on the 6% if you want to explore further)
One of the reasons companies stay stuck at the first check is because of how the tools are priced today. A vendor who could charge for outcomes needs an estimate to baseline against - e.g., cycle time to ship a campaign, or % of exceptions deflected - and companies (buyers) usually do not have that number handy when they sign the contract. So the vendor goes ahead and charges per seat.
And, what’s the easiest way to evaluate seats? Asking people whether they like the tool.
Several things then push the answers up:
the team that bought the licences also runs the survey, and people can tell which answer that team wants
for past 2 years or so, the message from CXOs on social media has been “AI is helping us and that people who cannot get value from it will be left behind”, so saying it did not help sounds like a statement about the person rather than the tool
So, finally, the survey comes back positive and the budget gets renewed, and since the contract never asked for the right estimates, nobody goes and measures the process.
The immediate fix some orgs are trying out is measuring the process before anything is built. Orgs which are doing it well are designating someone closer to the business to partner with the AI technologists to write down how long an outcomes and supporting blocks of work take from start to finish, how much of that time is people working vs queued/ waiting, and which steps across those work streams could be augmented vs rethought of with AI. Once that estimate has been sized, a vendor can be paid on it and the survey doesn’t remain the only quicker proof.
It is definitely hard to get it right, and building the right team is a tedious process with a lot of risks attached, but i do expect more companies adopting this measurement thesis over the next 6 months.
Token spend will only continue to get larger until companies have an understanding of setting up outcome oriented differential access and cost mechanisms, figuring out outcome-based (and not only task based) routing and deployment of different LLMs.
In the meantime, there is going to be increasing pressure from CFOs to assign an owner who can explain what those tokens bought and to whom. That owner needs to have a good hold on 2 key measurements:
how long an outcome and its supporting processes take from start to finish, broken down by before and after the pilot
what a specific change brought by AI did to a figure a specific function (e.g. PLG-to-sales handoff rate) or business (e.g. days sales outstanding) already reports
The first estimate can be made pretty quickly but the second one will take a few quarters.. but I am (a) more than certain that orgs are going to realise very soon that both are needed and (b) it will take a few cycles for the second realisation to occur but it’s best to have one person own both those estimations.
Until that happens, i really hope companies stop taking 'people like it' as all it takes to pick a vendor or plan budgets.


This maps to something I've seen in AI-native teams: the second check, cycle time, only works if you measured it before the tool showed up. Most companies skip that baseline, so six months in they're comparing a vague memory of 'before' to a real number 'after,' and the comparison is basically vibes with a spreadsheet attached. Do you think the lack of pre-AI baselines is why so many teams default to check one, since it doesn't require any historical data at all?