Every AI workflow you deploy either gives you capacity back or it doesn't, and the single variable that decides which one happens is how expensive it is to check the output. That's it. Not the model, not the vendor, not how good the demo looked.
I want to give you a test you can run this week on one workflow. It takes about ninety minutes. I've started calling it the verification test, and I've been running it with clients since a conversation that genuinely rearranged how I think about AI ROI.
Here's the conversation. At the professional services firms I work with, I went looking for time savings and asked the biggest AI champions, not the skeptics, the champions, how much time they were getting back. One power user paused and said, "Honestly? None. Whatever the AI saves me, I spend double-checking its work."
She does the task in a fraction of the time. Then she burns the entire remainder verifying it. Triple-checking. Quadruple-checking. The fear of one bad citation eats the whole dividend.
She is not doing it wrong. In a field where a single hallucination costs you a client, verification isn't waste, it's the job. The workflow is functioning exactly as designed. It just isn't producing the thing everyone assumed it would produce.
So the test is not "is this AI any good." The test is "what does it cost to know."
Step 1: Pick one workflow and time the whole loop, not the generation
Take a single recurring task that somebody on your team already runs through AI. One task. Not a department, not a category.
Now measure three numbers, in minutes, over one week:
T1: how long the task took before AI
T2: how long the AI takes to produce output
T3: how long a human spends checking that output before it's usable
Most people only ever measure T2, because T2 is the number that feels magical. T2 is what shows up in the vendor case study. T3 is where your ROI actually lives.
What good output looks like: T2 plus T3 is meaningfully less than T1. What the power user at the law firm has: T2 plus T3 roughly equals T1. Same result, different sequence, zero capacity gained.
Where this usually breaks: people estimate T3 from memory instead of measuring it. Everybody underestimates checking time by about half, because checking doesn't feel like work in the same way producing does. Make somebody write it down in real time for five days.
Step 2: Classify the verification
This is the step that actually tells you something. For that workflow, answer one question: can a machine tell you when it's wrong, or does a person have to read it?
There are only two answers.
Cheap verification. The tests pass or they don't. The query is faster or it isn't. The number reconciles or it doesn't. Somebody or something other than the original author can confirm correctness in seconds without re-doing the reasoning.
Expensive verification. Correctness requires a qualified human to re-read the whole thing with the same expertise it would have taken to write it. Legal citations. Clinical language. Client-facing analysis where one wrong clause is a relationship.
I watched this distinction play out in the same month, in opposite directions.
On a dev team I work with, one developer's monthly AI bill crossed $1,000. That team shipped an automated documentation system, a support site the AI essentially built itself, and a round of database optimizations that cut query times in half. Cheap verification the whole way down. Tests pass or they don't. And on the documentation project, when the agent hit something ambiguous, it didn't guess. It flagged the item "needs clarification," wrote up exactly what confused it, and moved on. Somebody answers the question, and the next nightly pass picks it up and finishes. The system knowing what it doesn't know is what made the verification cheap.
At the law firms, same dollars, expensive verification, zero throughput gain. The spend is real. The capacity isn't there yet.
Where this usually breaks: teams classify optimistically. If your honest answer is "well, mostly a machine can check it, except for the part that matters," that's expensive verification. Round down.
Step 3: Ask what happens when it's wrong and nobody catches it
Write one sentence describing the worst realistic outcome of an uncaught error in this workflow.
For the nightly documentation agent: a support page is stale for a day. Somebody notices, it gets fixed on the next pass.
For a legal brief with a fabricated citation: you lose a client, and possibly worse.
That sentence tells you how much verification is rational. It also tells you, honestly, whether you're in year one of this workflow or year three. In year one on a high-stakes workflow, the honest ROI answer is that there isn't one yet. You're paying for table stakes, the same way nobody measured ROI on rolling out Word. The payoff shows up later, when trust catches up to capability and verification drops from 100% of outputs to spot-checks.
I keep seeing a six to twelve month lag on that curve. I have not found a way to shorten it, and I've stopped pretending to clients that I have.
Step 4: Decide what you're actually buying
Now put the workflow in one of three buckets, and be honest about it.
Capacity. Cheap verification, T2 plus T3 well under T1. This is throughput you didn't have. Fund it like headcount, not like software. Nobody asks whether a contractor is "too expensive" in the abstract, they ask what got delivered.
Trust-building. Expensive verification, no time savings yet, but the work is real and the stakes justify the checking. Budget it as table stakes, set a review date six months out, and measure one thing: is the percentage of outputs requiring full verification going down? If it's flat after two review cycles, you have a workflow problem, not a patience problem.
Neither. No time savings, no trust curve, no clear worst-case worth the checking. This one gets cut.
The failure this test is designed to prevent
Here's the mistake I made first, and I made it on a call, out loud, before I caught myself.
A founder saw that $1,000 monthly bill and said, "I said we wanted to get this stuff done. I didn't say we wanted to go out of business." My immediate instinct was containment. Per-seat caps. Model tiers. Usage dashboards. Approval thresholds. I spent a solid fifteen minutes being helpful in exactly the wrong direction before I thought to ask what the team had actually shipped that month.
Three real projects. Cheap verification on all three. That spend was the best-converting money on the P&L, and my reflex was to cap it.
A blanket per-seat ceiling would have killed the documentation system, because a nightly agent crawling an entire knowledge base is precisely the usage pattern that trips a limit designed for chat windows. The cap punishes the workflow with cheap verification, which is the only kind that was working, while leaving the expensive-verification workflows untouched because they never spend much anyway.
Caps are what you reach for when you don't have this test. That's the whole reason to run it.
What it looks like when you're done
One page. One workflow. Three numbers, a verification classification, a worst-case sentence, and a bucket. Ninety minutes of somebody's week.
The output isn't a strategy. It's a decision about one thing you're already doing, made with something other than vibes. Run it four more times and you have a map of where your AI dollars convert and where they're buying patience. Those are both fine things to buy. They are not the same thing, and right now most companies are reporting them in the same line item.
Homework
Pick the AI workflow your team uses most. Have the person who runs it log T1, T2, and T3 for five working days, in minutes, written down as it happens rather than recalled on Friday. Then classify the verification as cheap or expensive in one sentence.
If T2 plus T3 is not clearly under T1, you don't have a capacity gain yet, and the next conversation should be about whether verification will get cheaper by year-end or whether this workflow is one you're funding out of hope.
This week on LinkedIn
Monday - The SaaS CEO who isn't sure he adds value anymore
Tuesday - A sales guy wrote a million lines of code in 30 days
Wednesday - I built an AI-first company in 2022, and I don't recommend it
Thursday - No password breaches is now the hiring red flag
Friday - Every AI employee needs a human boss, until you actually try it
Last week, in case you missed it
Two things I'd genuinely like back from you. First, hit Reply and tell me which part of this matches what you're seeing. Second, if there's someone in your world who should be on this list, forward it to them. Thank you!
- Robbie