The short version

Start with a business problem that costs money you can name. Map the workflow and split it into what a machine can do, what it can prepare for a person, and what a person has to own. Price what that costs you today, before you switch anything on. Estimate the future state step by step rather than as a blanket percentage, and say out loud how freed capacity turns into money. Cost the whole investment, not the license. Then give the decision a payback month and a range, and label every soft number as soft before somebody else finds it.

The math isn't hard. The discipline is measuring the before, naming the mechanism, and being willing to write down the number that would prove you wrong.

One rule sits underneath all of it: a business case built with your finance partner gets approved. A business case built at them gets audited.

Sections 1 through 7 and 9 through 13 are the method. Section 8 is the arithmetic worked in full, for people who have to defend it in a room. The appendix is how to measure the number the arithmetic depends on.

Most AI business cases are built backward. Someone sees a tool, gets interested, and goes looking for a problem that justifies buying it. The business case gets written last, usually by the person who already knows what they want the answer to be.

I know because I've written that version. Through most of the last automation wave I spent a good part of my working life building business cases for robotic process automation and walking them into finance. Some got approved. Some came back with the assumptions circled. The ones that came back taught me more, and most of what's in this guide is what I wish I'd known before the first meeting rather than after the third.

This guide runs the other direction. Start with a business problem, work out what it costs you right now, and only then ask what technology is worth spending on it.

It assumes your organization has done little or none of this before, and it covers both situations: you haven't run a pilot yet, or you ran one and now need to know what it was worth.

You don't need a data science team for any of this. You need one workflow, six numbers, and a finance partner willing to sign off on them.

What's different this time. The method here would have worked for RPA. The reason it needs restating is that the technology changed underneath it. An RPA bot is deterministic: it completes the task or it fails, and the failure is visible. A language model is probabilistic: it gives you something usable seven times in ten, and the three misses look like successes until someone reads them. That's why acceptance rate is the number that matters now and didn't exist as a concept in the last wave. It's why review time is a real cost line instead of a rounding error. And it's why keeping a person in the loop is an engineering requirement, not a values statement. If you built cases in the last wave, most of your instincts still hold. The ones that don't are the ones this guide spends the most time on.

This guide goes deep on one workflow on purpose. Deciding which workflow deserves that attention in the first place is the harder problem, and a different one.


Start with the business problem, not the tool

Write the problem as one sentence with a number in it.

"Our sales team takes eleven days to get a proposal out, and we lose deals to competitors who take three" is a problem. "We should be using AI in sales" is a budget line looking for a reason.

Three questions to screen any candidate:

  1. Does this cost us money we can name?
  2. Do we know how much, or could we find out in two weeks?
  3. Would we fix this even if AI didn't exist?

The third question catches the most projects: if the answer is yes, then AI is the method and the problem is real. If the answer is no, you have found a capability you want rather than a problem you have. That doesn't make it a bad project per se, it's just less likely to pass the approval process and should be prioritized lower on the list.

Sequence: revenue first, then what supports revenue

Since you're likely reading this because you're trying to justify an early AI investment, you'll use this use case as the basis to fund future projects.

Sort candidate workflows into three tiers:

Tier 1 is revenue-generating work. Proposal and quote generation. Lead qualification. Renewal outreach. Customer onboarding. These sit closest to the money and have the shortest chain between "we changed this workflow" and "revenue moved." Short chains are what make a return defensible.

Tier 2 is the work that feeds Tier 1. For example, marketing operations for the campaigns that fill the pipeline. This could include marketing operations, content production and data preparation for targeting. The value is real but it arrives through Tier 1, so you measure it in throughput rather than in revenue directly.

Tier 3 is everything else. Finance close, HR service delivery, IT support, internal knowledge search. Genuine value but harder to attribute and easier to fund once you have a proven return to fund it with.

Sequencing this way is a funding strategy, not just a priority list. A defensible Tier 1 return is what pays for the platform, the integration work, the security review, and the governance that everything after it will need. Start in Tier 3 and you are asking for a budget whose ROI you can't justify.

There is a second reason to pair Tier 1 and Tier 2 rather than choosing between them. A Tier 1 win fills the pipeline while a Tier 2 win gives you the capacity to serve what you filled. Do one without the other and you either create demand you cannot service or free up capacity with nothing to put in it. The two work together, and the business case is stronger when it says so.

Draw the workflow before you price it

Map the current state at the task level, not the process level. A process diagram tells you what is supposed to happen but you need what actually happens, including the waiting.

For each step, capture who does it, how long it takes, how often it runs, what triggers it, what it produces, and where it stalls.

Then sort every step into three buckets.

Automate. Retrieval, formatting, drafting, routing, rules that are already written down somewhere.

Assist. A human decides, and AI prepares the decision. Research, first drafts, summarizing, options.

Leave alone. Judgment, relationships, accountability, and anything where being wrong is expensive.

You can claim savings on the first bucket and part of the second. That's it. Business cases that assume an entire workflow disappears are the ones that get rolled back in month nine, when the exceptions and the review time turn out to be real.

One thing to expect when you run this exercise: people sort their own work into the judgment pile, because that's the pile that keeps their job interesting. The tell is when "leave alone" comes out largest. Arguing about it doesn't work. Changing the question does. Instead of "does this step need judgment," ask "what happens if this step is wrong and nobody catches it for a week?" If the answer is "someone fixes it," the step is automate or assist. If the answer involves a customer, a regulator, or a number a senior person has already seen, it's leave alone. People answer the consequence question honestly because it isn't about them. It's about the blast radius.

One thing worth looking for while you map: in revenue workflows, the waiting between steps is usually the biggest cost, and it never appears on an org chart. A proposal that takes eleven days often contains six hours of work and ten days of queue.

Price the current state

This is the baseline, skipping it is the single most common reason a company cannot calculate a return later. Without a before, you don't have a measurement. You have a story.

Six numbers per workflow:

  1. Volume. How many times a month this workflow runs.
  2. Touch time. Hours of human effort per run, broken out by role.
  3. Cycle time. Calendar time from trigger to done. Not the same as touch time, and for revenue workflows it often matters more.
  4. Fully loaded labor rate. Salary plus benefits, payroll taxes, tooling, and overhead. Use the raw salary number and someone in finance will correct you in front of your sponsor, usually by a factor of 1.25 to 1.4.
  5. Rework rate. How often output goes back for a fix, and what the fix costs.
  6. Outside spend. Agencies, contractors, and vendors doing part of this work today.

Because volume is a monthly count, this gives you a monthly figure: current-state monthly cost is (volume × touch time × loaded rate) + rework + outside spend. Multiply by twelve when you need the annual number, and say which one you are showing.

For Tier 1 workflows, add a delay cost: what a shorter cycle time is worth. Win rate against faster competitors, deals that age out of the pipeline, revenue recognized earlier in the quarter. This number needs your finance partner's blessing before it goes in a deck, because it is the number that gets challenged first.

If you have no data at all, you can build a defensible baseline in two weeks. Sample three people over ten working days with a simple time log, and pull whatever your CRM and ticketing systems already timestamp. It won't be perfect but it will be defensible, which is a different and more useful standard than perfect. Label it an estimate and move.

Estimate the future state honestly

Estimate efficiency step by step, not as a blanket percentage. "Thirty percent faster overall" is a number nobody can verify and everybody can dispute. "Drafting drops from ninety minutes to twenty, review adds fifteen" is a claim you can test.

Four adjustments that most business cases leave out:

Ramp. Month one is slower than the old way. People are learning, the prompts are wrong, the integrations are half done. Model a curve rather than a step change, and assume no net gain in the first month. If your case only works when the benefit starts on day one, it doesn't work.

Review time. Human-in-the-loop is a real line item, not zero. Someone reads the output before it goes to a customer. Budget those minutes, every time, at the loaded rate.

Acceptance rate. The share of output usable without meaningful rework. This is the number most likely to be flattering in a vendor demo and disappointing in your environment, because your data is messier than the demo data. Measure it on your own work before you commit to it.

Exception handling. The cases that fall out of the automated path. Some of them will cost more to handle than they did before, because now someone has to figure out why the system didn't take them.

Build three cases: conservative, expected, optimistic. Lead with the conservative one. If the conservative case doesn't clear your investment bar, you don't have a project yet. You have a hypothesis, and the next section of this guide is about how to test it cheaply.

Where efficiency turns into money

This is the section most business cases are missing, and it is the reason so many AI investments produce real efficiency and no visible return.

Time saved is not money saved but it is capacity that becomes money in exactly three ways.

The team absorbs more work. More campaigns, more proposals, more accounts per person. This is the outcome to design for, and it is the one that pairs with your Tier 1 investment. You freed capacity in marketing operations, and you filled it with the campaign volume your new pipeline needs.

You stop paying someone outside. Agency and contractor spend comes down. These are hard dollars, they show up in the general ledger, and they are the easiest kind of return to defend.

You reduce headcount. This is the version that shows up in most vendor decks and the one we would put last for a practical reason rather than a sentimental one.

The quiet cost of the headcount route

The moment a team understands that "efficiency" means fewer of them, they stop telling you where the tool fails.

That reporting is the most valuable thing you own in the first year of an AI investment. It is how you find the exceptions, fix the prompts, catch the errors before a customer does, and learn which steps genuinely needed a human. Buy that information with job security and it costs you nothing. Trade it for a headcount reduction and you have bought a one-time saving with the mechanism that would have made the second year work.

This is why rollbacks happen. Not because the technology failed, but because nobody would say so early enough to fix it.

Whichever of the three you choose, write it down before you build. A business case is only honest when the mechanism is named. "We'll figure out what to do with the time later" is how a real gain becomes an unprovable one.

Cost the investment, not the license

How each billing shape behaves in a forecast

Billing shapes, how each behaves in a forecast, and what to watch for.
ShapeForecasts likeWatch for
Per seatPredictable, scales with headcountCost decoupled from value. You pay for licenses nobody opens, and broad rollout gets punished exactly when adoption is the goal.
Consumption or tokenScales with actual useHard to forecast and spiky. A chatty workflow or a retry loop can multiply cost overnight. Require caps and alerts.
Platform fee plus creditsA floor plus a variableCredit expiry, the overage rate, and what a credit actually buys. This is where the surprise usually lives.
Per action or workflow runClean unit economics, easy to model per transactionVerify what counts as one action. Multi-step agents and retries inflate the count fast.
Outcome-basedSpend tracks valueDefine the outcome and the attribution rule in the contract, or you will be arguing about it in month four.

What sits outside all of them

Integration and data plumbing. Data cleanup. Security and legal review. Model evaluation. Prompt and agent maintenance as your business changes. Monitoring. Training hours. And the internal people who run the thing, whose time is real cost even though no invoice arrives for it.

A working rule for the first year: the license is the smaller half of the bill. Put the internal hours in the business case as a line item with a name against it, because that is the cost most likely to be discovered late.

Five things to settle with a vendor before signing: a usage forecast at your actual volume, a spend cap, the overage rate in writing, a re-price trigger, and what happens to your data and your configured workflows if you leave.

Do the math

Net monthly benefit = (baseline cost − future-state cost) + attributable new revenue − the cost of running it.

Payback = the month your cumulative position crosses zero, counting the setup cost and every month you paid for the thing before it started working.

Lead with a payback month and a range rather than a single ROI percentage. A percentage invites an argument about the denominator. A payback month is a decision.

One thing to settle before you present: year-one payback is the wrong bar, and insisting on it is what produces inflated business cases. Nobody expects an ERP rollout or a warehouse to pay for itself in twelve months. When AI is held to a standard we apply to nothing else of comparable cost, the pressure lands on the assumptions, they become too optimistic, the project misses a number it should never have been given, and it gets pulled. That is a meaningful share of the rollbacks.

A worked example

A mid-sized B2B services firm, roughly 400 people. The figures are illustrative and the arithmetic is checkable, which is the only claim being made for them.

A word on precision before we start. What follows is deliberately simplified so the structure stays visible. Your finance partner will want to substitute their own labor rates, their own margin, and possibly a discount rate, and they should. Organizations differ enormously in the rigor they require: some need a fully modeled NPV, others need a defensible order of magnitude and a named owner. Bring finance the structure and let them set the precision. A business case built with them is approved; a business case built at them is audited.

Practically, that means bringing them the model before it has an answer in it. Ask which labor rate to use before you've multiplied anything. Ask what margin they'd count incremental deals at before you've estimated a single deal. Once finance has supplied an input, the output is partly their number, and people defend their own numbers in front of the CEO. Bring them a finished case and the only thing left for them to do is find the hole.

The baseline. 120 proposals a month, 9 hours of touch time each across a solutions lead, a subject matter expert, and a proposal writer, at a blended loaded rate of $85. That is about $91,800 a month. Cycle time is 11 days.

Caveat to raise yourself: a blended rate hides mix. If AI displaces your cheapest hours and leaves the expensive ones, blending overstates the gain. Finance can rerun it by role in ten minutes.

Split the hours before you touch them. Of the 9 hours, about 5.5 are drafting and assembly and 3.5 are scoping, pricing judgment, and sign-off. Only the first pile is in play. Any estimate that compresses all 9 hours is quietly claiming to automate the part where somebody's name goes on the decision.

That split comes from the three-pile exercise in Section 3, done with the people who do the work. It is worth more attention than anything else in this model, for reasons the sensitivity table below makes plain.

The future state. When the draft is usable, the 5.5 drafting hours become 1.5 and review adds 45 minutes, so 5.75 hours. When it is not, the team writes it themselves and the review time is spent anyway, so 9.75 hours. Blended at a measured 70% acceptance rate that is 6.95 hours, about $70,890 a month, freeing roughly $20,910.

Caution: it is not $35,000, which is what you get by applying 70% to all 9 hours instead of to the part that was ever automatable.

Cycle time, with the queue shown. Nine hours of work inside an 11-day cycle means about 10 days of waiting: two days of intake, four days of drafting (mostly waiting for the SME to have a free block), three days of pricing approval, two days of final review.

AI shrinks the SME's task from "write your section," which needs a two-hour block and gets scheduled for next week, to "correct this draft," which fits between meetings. Smaller tasks queue shorter, which plausibly takes the drafting window from four days to about one and a half. Pricing approval does not move; it is three days of someone else's process. So the defensible claim is 11 days to 7, not 11 to 4. Going below 7 means changing the approval process, which is a different project with a different owner.

The revenue argument, and its weakness. Sales believes a 4-day faster turnaround lifts win rate by a point. Finance halves it to half a point and counts it at a 40% gross margin rather than at revenue: 0.6 additional deals a month, about $9,600 of margin.

Caveat to raise yourself: nobody has yet produced a deal that was lost on speed, and 40% is at the optimistic end for services. Until someone pulls loss reasons from the CRM this is a hypothesis with a dollar sign in front of it. It is also, as you will see, the line the whole case rests on.

The supporting workflow, and why it is here unpriced. Campaign build, list preparation, and content adaptation run 220 hours a month at a $65 rate, split 130 production and 90 judgment. Running the same assumptions as above (a 73% cut to automatable hours, review at 14%, acceptance at 70%) recovers about 48.5 hours a month, or 22% of the workflow.

Both workflows use identical ratios on purpose. A worked example that changes its own assumptions between sections cannot be checked by the reader, and being checkable matters more here than being exact. In a real case you would derive each separately and they would not match.

We are not converting those 48.5 hours into a dollar figure, and that is deliberate. Recovered time arrives in fragments across a small team, and fragments do not staff a campaign unless somebody consolidates them on purpose. Naming a number of additional campaigns would be exactly the kind of claim this guide keeps telling you to distrust.

What the capacity actually buys is insurance. If the proposal workflow works and volume rises, marketing operations is the constraint that appears next. Funding the two together keeps the first gain from evaporating into a bottleneck one function downstream. That argument is worth making on its own terms, without a number attached.

If finance wants it priced, it can be: value the campaigns at their historical contribution to pipeline, or price the capacity as the contractor you would otherwise hire when volume rises. Both are defensible, both need your numbers, and both would double the length of this example. Model it with them, separately, if the case needs strengthening.

Investment. $6,500 a month for platform and seats, $9,000 a month for the 0.4 of a person who runs it, $70,000 one time for integration and data work, $25,000 for training and change. Months one and two are baseline measurement, integration, and training, so the $15,500 monthly is paid with nothing coming back.

Three caveats to raise yourself. The $9,000 is internal labor, not cash out the door, and finance will want cash and notional separated. On cash alone (platform and setup out, deal margin in) this runs about $3,100 a month positive from month four and is still slightly negative at month 36, which is worth knowing before someone else finds it. Second, 0.4 FTE in year one is optimistic; new tools tend to consume closer to a full person until they stabilize. Third, there is no contingency in the $95,000, and integrations rarely come in flat.

The result, with acceptance improving from 70% to 85% as the system is tuned, shown with and without the revenue line so a reader can see what it is carrying:

Cumulative position by month, with and without the revenue line.
Month Capacity freed Cumulative, with revenue Cumulative, cost savings only
1–2$0−$126,000−$126,000
3$20,910−$120,590−$120,590
6$23,970−$69,440−$98,240
9$27,030−$9,110−$66,710
10$27,030$12,020−$55,180
12$27,030$54,280−$32,120
15$27,030$117,670$2,470

Payback is month 10. Remove the revenue column and it is month 15. Those five months are the entire value of a CRM loss-reason query somebody could run this week.

Neither number kills the project. Month 15 on a technology investment of this size is ordinary, and any sponsor who has approved a systems migration knows it. What would kill it is presenting month 5, which is what you get by rounding every assumption toward the answer you wanted.

What the second workflow costs. Most of that $95,000 isn't the price of this workflow. It's the price of entering the category: the integration pattern, the data cleanup, the security review, the vendor relationship, and the first real argument with your team about what a machine should and shouldn't touch. Workflow two inherits nearly all of it.

Run the same model on a second workflow carrying only marginal setup cost and payback arrives substantially sooner. That's the difference between funding a series of experiments and funding a program, and it's worth saying out loud to whoever controls the budget. The first case is the expensive one. It's also the one that buys the right to make the next five cheaply.

What actually moves the answer. Three assumptions carry almost all the uncertainty. Show this table and you will spend the meeting on the right argument:

Sensitivity: the three assumptions that carry almost all the uncertainty.
Assumption Range Payback with revenue Cost savings only
Hour split (automatable / person-only)4.5h / 4.5h → 6.5h / 2.5hmonth 13 → 8month 30 → 11
Deployment length2 → 6 monthsmonth 10 → 17month 15 → 25
Revenue lineincluded → removedmonth 10month 15

Why that first row matters most. It's worth understanding why one hour moves the answer so far, because the reason generalizes.

When you hand an hour of work to the machine, that hour doesn't disappear. It shrinks to roughly fifteen minutes, because someone still has to read the output and fix it. When you keep an hour on the person side, it stays a full hour. So every hour you move across that line buys back about forty-five minutes on the proposals where the draft comes out usable, which is seven in ten of them. Averaged across all 120 proposals, call it half an hour each, or about $5,200 a month.

(The range in the table, plus or minus an hour, is a judgment rather than a derivation. It's roughly how far apart three people doing the three-pile exercise would land on a 9-hour workflow. Use whatever spread your own team actually argues about.)

The uncomfortable part. Now put that $5,200 next to what this project earns.

In month three the workflow frees $20,910 of capacity and costs $15,500 a month to run. The running cost eats 74% of the gain, which leaves a cushion of about $5,400 a month.

One hour moved between the piles is worth $5,200. The entire monthly cushion is $5,400. Move the hour the good way and you roughly double what the project makes. Move it the wrong way and almost nothing is left.

It's the household version of the same problem. A family bringing in $5,000 a month and spending $4,700 has a $300 cushion, and that's fine right until the car needs brakes. A family with a $3,000 cushion doesn't notice the same repair.

So this project isn't fragile because the hour split is uniquely important. It's fragile because it barely makes money, and when the margin is thin every assumption swings the answer. A workflow that freed $60,000 a month against the same $15,500 would shrug off the same disagreement.

That's the argument for choosing a first workflow with real headroom in it rather than the one that irritates people the most. Headroom isn't only about earning more. It's about being able to forecast at all.

And on cost savings alone, this project does not pay back inside a year. The case works because of the revenue line, which is the number nobody has evidence for yet.

That's not a reason to kill it. It's a reason to know exactly what you're betting on, and to spend the first two weeks pulling loss reasons out of the CRM instead of shopping for tools. If speed isn't actually why you lose deals, you'd rather learn that in week two than in month ten.

Notice what carried this case, and what didn't. Not the technology. The hours split before anything was estimated. The queue broken out so the cycle-time claim had a mechanism behind it. The acceptance rate measured rather than assumed. And every soft number labeled as soft, in front of the person who was going to find it anyway.

If you have never run a pilot

Instead of a pilot, start with a baseline and a decision rule.

Most first pilots fail as investments because they were designed to prove that the technology works, and whether the technology works was never really the question. It works. Whether it works in your workflow, at your volume, on your data, with your actual team is the question, and a demo cannot answer it.

The shape recommended:

Two workflows, not one. Six to eight weeks with one Tier 1 and one Tier 2 that supports it.

Measure the baseline first. Two weeks of instrumentation before anything is switched on. Without it you cannot accurately calculate a return afterward.

Write the success threshold before you start, as a number and a decision rule: continue, change, or stop. Put a date on the decision and a name against it.

A named owner in the business who owns the number, not the tool. It shouldn't be IT, and not the vendor.

Buy, don't build. A first investment is not the place for custom work.

Real users, real work, real volume. Don't pick volunteers running hand-picked easy cases in a sandbox. Connect with the folks doing the work and surface the realities they face and you will be a lot less likely to have to rollback.

Cap the spend and pick something you can walk away from.

Instrument six things from day zero: volume, touch time, cycle time, rework rate, acceptance rate, and adoption. That last one, how many people actually used it and how often, will explain more about your results than any of the others.

If you ran a pilot, here is how to calculate the return

Reconstruct the baseline you didn't measure. System timestamps, calendar data, ticket history, and a structured recall session with the team. Label it an estimate. Get finance to sign off on the labor rates specifically, since those are the input most likely to be challenged.

Correct for pilot bias. Every pilot number is optimistic, for the same predictable reasons. Enthusiast users rather than the median user. Clean, hand-picked cases. Vendor support hours you will not get at scale. Free or discounted credits. Plus, the novelty effort fades around month three.

Apply a haircut and state what it is. Reviewers trust a stated haircut far more than an unstated one, and stating it moves the conversation from "do I believe you" to "is 30% the right number," which is more constructive.

Convert pilot cost to run-rate cost. List price, full volume, all seats, plus the internal hours the pilot got for free because three people were excited about it.

Annualize with a ramp, not by multiplying your best pilot week by 52. Document your ramp assumptions.

Then make one of three calls. Scale it, extend it with proper instrumentation, or stop. If you extend, attach a new end date and a new threshold. Otherwise you have created a permanent pilot, which wastes real spend, real attention, with no decision at the end of it.

Before you take the number to finance

Run a sensitivity check. Cut your efficiency assumption in half. Does the case still clear the bar? If it doesn't, say so in the room before someone else finds it.

Fix the attribution rule. Whoever signs off on the labor rates and the revenue attribution should not be the person who wants the project approved.

Don't double count. The same hour cannot be saved by two different projects, and the same deal cannot be credited to two initiatives. This gets caught, usually late and usually publicly.

State the counterfactual. What happens if you do nothing for twelve months. Sometimes the honest answer is "not much," and it is much better to know that before you spend than after.

Include the exit cost. What it takes to unwind this in eighteen months if it doesn't work. A case that has thought about its own failure is more credible, not less.

Know what each person in the room will ask

You'll present this to people whose jobs are different from yours. Walking in with their questions already answered is most of what separates a proposal that gets approved from one that gets deferred for more analysis.

Your CFO will ask: does the integration work capitalize or expense, and which budget does it hit? What does the three-year total look like, not just year one? Which lines are cash leaving the building and which are internal time you're already paying for? What's the contract term, and what happens at renewal?

Have ready: the split between capital and operating spend (integration and build often capitalize, subscription doesn't, and it changes which approval you need), a 36-month view, and a version of the model with only cash in it.

Your CIO will ask: what data leaves our environment, and does the vendor train on it? What does it cost to get out in eighteen months, and what's stranded if we do? Why doesn't the license we already own cover this?

Have ready: a straight answer on data handling, an exit cost sitting in the investment table rather than in a footnote, and an honest comparison against whatever adjacent tool your organization already pays for. That last question stops more projects than any other, and the only bad answer is not having considered it.

Your CEO will ask: why this workflow and not one closer to where the company is going? What happens when we run twelve of these? What do we lose by waiting a year?

Have ready: the counterfactual, and the compounding argument from Section 8. One project at a fifteen-month payback is a decision. A program where each project costs less than the last is a strategy.

Your sponsor will ask: who owns this number when it moves, what would make us stop, and whose team do the integration hours come out of?

Have ready: a name, a written kill criterion with a date on it, and a confirmed commitment for the technical capacity, secured before the number went into a deck rather than after.

That last one has a specific failure mode worth naming. A sponsor's yes is worthless if the person whose team does the integration wasn't in the room. Sponsors estimate other people's capacity generously, always in the same direction. Get the estimate from the engineering or data lead directly, in writing, before the number goes into a deck. If they won't commit to it, that's your answer, and it's better to have it now.

And be ready for the question nobody asks out loud: what would make you recommend not doing this? Having a real answer, and offering it before you're asked, buys more credibility than any table in this document.

What this looks like on one page

Problem, in a sentence with a number in it. Baseline, with the six inputs and where they came from. Assumptions, with the conservative case leading. Investment, license and everything else. Payback month. Decision date. Named owner.

That's the whole thing. The math is not difficult. The discipline is measuring the before, naming how capacity turns into money, and being willing to write down the number that would prove you wrong.

And this was one workflow. The harder discipline, once you can do this, is choosing which workflow deserves it, and being willing to tell a senior person that theirs doesn't.


Appendix: Measuring acceptance rate so it means something

Section 5 tells you acceptance rate is the number that matters. Here is how to get one you can defend.

The problem is that a language model's misses look like hits. In the last wave a failed bot stopped and generated a ticket. A model produces something every time, and the three bad outputs in ten look exactly like the seven good ones until somebody reads closely. So you can't count failures. You have to go looking for them, in three layers.

Layer 1: the review step records a verdict. Every reviewed output gets one of three marks: used as it came, used with edits, redone. Instrument it in the tool, or with a three-button form if the tool won't. If the reviewer fixes the output silently and moves on, you have a review process and no data. End-of-week recall doesn't count; people remember the bad ones and lose count of the rest.

Layer 2: blind audit of what was accepted. Reviewers drift. Acceptance rate rises while quality falls, because tired reviewers start approving on sight, and nothing in the primary measurement can tell those apart. A second reviewer scores a random 5 to 10% of accepted outputs without knowing they were accepted. If acceptance is rising and audit agreement is falling at the same time, you're measuring fatigue, not improvement.

Layer 3: lineage. Some misses pass both layers and surface later as a customer complaint, a bounced compliance check, a campaign carrying last month's offer. Every AI-generated output carries an identifier through the workflow, so a downstream failure can be traced back to the output, the reviewer who passed it, and the prompt version that produced it. Without this, escape rate is a guess.

The acceptance-rate KPI set, and how each one goes wrong.
KPIWhat it tells youWhere it goes wrong
Acceptance rateShare of outputs used without meaningful rework"Meaningful" undefined until after the fact
Edit ratioHow much the reviewer changed, per outputCollapsed into a binary that hides partial value
Review time per unitThe cost line behind human-in-the-loopRecalled, not captured
Escape rateShare of accepted outputs later found wrongUnmeasurable without Layers 2 and 3
Audit concordanceHow often the second reviewer agrees with the firstA fall gets read as model improvement
Exception rateShare that leaves the automated path entirelyLeft out of the cost model
AdoptionShare of eligible work actually routed through the toolExplains the most and is measured least
Cost per accepted outputTotal cost divided by usable unitsThe one that ties the rest to economics

Seven rules for calculating them. Define "accepted" as a rule before day one (under five minutes of edits, no change to named fields, an edit-distance threshold), or the number moves with morale. Use everything generated as the denominator, not everything reviewed, or an adoption problem shows up as a quality win. Segment by case type, because a 70% blend is often 95% on easy cases and 30% on hard ones, and the split feeds straight back into the three-pile exercise in Section 3. Keep the scorer separate from the stakeholder, the same attribution rule Section 11 applies to labor rates. Compare against the team's rework rate from Section 4, not against zero. Treat a rising acceptance rate with falling audit agreement as fatigue. And trend it, because a flat line for three months means the tuning loop isn't running, which is a management finding rather than a model finding.

Acceptance rate is the highest-leverage input in Section 8. An input with that much leverage and no measurement behind it isn't an assumption. It's a hope with a decimal point.

Heather Besner writes about AI enablement and operationalization: how organizations turn AI experiments into capability that survives a budget review.

Not sure your organization is ready to make this decision yet?

The free AI Transformation Decision Snapshot shows where the decision actually sits — who owns it, how consistently problems get defined before a tool is chosen, and how governance is working — before anyone spends money.

All field reports