August 17, 2026 · 26 min read
How to Estimate AI Automation ROI Without Relying on Vendor Benchmarks
A credible AI business case should explain how a specific workflow changes economically. External productivity research can inform the hypothesis, but the ROI assumptions should come from the organization's own process and pilot evidence.
Executive Summary
AI automation business cases are often presented with a familiar pattern: estimate the percentage of work that can be automated, multiply that percentage by labor cost, subtract software expense, and report the remainder as return on investment. The arithmetic is simple, but the model is frequently misleading because it treats heterogeneous knowledge work as though it were a uniform pool of labor hours. In practice, AI changes different tasks in different ways. It can reduce active handling time, improve quality, increase throughput, shift work from execution to review, create new exception categories, or have little effect on the highest-skill portions of a process. Some benefits appear as avoided future hiring rather than immediate cost reduction; others appear as shorter cycle times, better service levels, reduced rework, or increased capacity. A credible ROI model should therefore begin with the workflow and its observed operating data rather than a vendor benchmark.
The empirical literature makes the case for this approach unusually clearly because measured effects vary substantially by task and setting. In the 2025 Quarterly Journal of Economics article Generative AI at Work, Brynjolfsson, Li, and Raymond studied deployment of an AI assistant across 5,172 customer-support agents and found a 15 percent average increase in issues resolved per hour. The lowest-skill quintile experienced a 36 percent productivity increase, while the most skilled workers saw little productivity improvement and some evidence of a small quality decline. This is strong evidence that AI assistance can create material value, but it is equally strong evidence against applying one average productivity percentage to every worker or process.
Noy and Zhang's experiment on professional writing provides another positive result from a different task class. MIT's summary of the study reports that approximately 450 college-educated professionals completing occupation-specific writing assignments finished about 40 percent faster with ChatGPT, while independently rated output quality increased by 18 percent. Yet the tasks were deliberately bounded and did not require full organizational context or high-stakes factual verification. The implication for an ROI model is not that a writing workflow should assume a 40-percent labor saving. It is that bounded language-intensive tasks can exhibit large effects when the model's capabilities align well with the task and human review remains practical.
The 2026 Organization Science article Navigating the Jagged Technological Frontier sharpens the point. In a preregistered experiment with 758 Boston Consulting Group knowledge workers, participants using GPT-4 on 18 tasks designed to fall within the model's capability frontier completed 12.2 percent more tasks and worked 25.1 percent faster, with significantly higher quality. On a complex managerial task intentionally selected outside the frontier, however, AI-assisted participants were 19 percent less likely to produce the correct solution. The same technology can therefore create positive and negative economic effects within superficially similar knowledge work. ROI is conditional on task fit.
Software-development evidence is similarly heterogeneous. METR's 2025 randomized controlled trial recruited 16 experienced open-source developers working on 246 real issues in mature repositories they knew well. Developers expected AI to reduce completion time by 24 percent before the study and, after using the tools, believed AI had accelerated them by 20 percent. Measured performance showed the opposite: allowing early-2025 AI tools increased completion time by 19 percent. METR explicitly warns against generalizing the result to most software development, and its 2026 follow-up says later data are affected by serious selection issues as AI adoption increased. The study is nevertheless economically important because it shows that perceived productivity and measured productivity can diverge sharply.
A separate randomized controlled trial involving 96 full-time Google engineers found approximately a 21-percent reduction in time on a complex enterprise-grade task when developers used three AI features, although the authors emphasize that the confidence interval was wide and that results from internal tooling and a 2024 task should not be assumed to transfer broadly. Taken alongside METR, the study demonstrates why category-level claims such as AI speeds up software development are too coarse for investment analysis. The economically relevant unit is the actual task distribution, workforce, tooling, codebase, and review burden.
Large-scale field experiments also show variation across workflows within the same firm. A 2025 study of generative AI in cross-border online retail reported randomized experiments across seven customer-facing workflows involving millions of users and products; estimated sales effects ranged from 0 percent to 16.3 percent depending on the marginal contribution of the AI relative to existing business practices. The precise result remains setting-specific, but the pattern is informative: where the baseline process is already strong or AI adds little incremental value, returns can be small even at very large scale.
The correct conclusion from this literature is therefore not an average AI productivity rate. It is a methodology. External studies can establish plausible mechanisms—faster drafting, improved novice performance, higher throughput, lower quality outside the capability frontier, additional review burden, or no measurable incremental effect. The numerical business case should then be built from internal workflow volume, touch time, exception mix, quality, capacity, costs, and pilot evidence.
Implementation connection: The Workflow Discovery Template helps establish the process baseline described here, while the Invoice Processing ROI Calculator provides a concrete resource for translating workflow assumptions into an economic model.
1. Define the Economic Unit Before Calculating ROI
An ROI model needs a unit of work. Depending on the process, that may be an invoice, support case, document, proposal question, shipment exception, marketing asset, legal request, order exception, or another transaction. The unit should correspond to something the organization can count consistently and follow from trigger to completion. Without a stable unit, time savings and cost estimates often mix incompatible activities and become difficult to validate.
For each unit, describe the current process in stages. A support case may include intake, classification, account lookup, investigation, drafting, approval, customer communication, and documentation. An invoice may include receipt, extraction, matching, exception review, approval, and ERP posting. The objective is to identify which activities AI might change rather than applying an automation percentage to the whole process.
2. Establish a Baseline From Observed Operations
Start with observed volume over a representative period. Monthly averages may be adequate for a stable workflow; seasonal or rapidly growing processes may require weekly or quarterly distributions. Record both average and peak volumes where capacity matters. A business case based on an annual average can miss the value of handling a seasonal surge without temporary staff or overtime.
Then measure active touch time rather than elapsed cycle time. If a case sits in a queue for two days but an employee actively works on it for twenty minutes, the directly automatable labor base is twenty minutes unless the automation also changes the queueing mechanism. Cycle-time reduction can still have economic value, but it should be modeled separately as a service, working-capital, conversion, or capacity effect rather than mislabeled as labor savings.
Where time data are unavailable, sampling can be more defensible than executive estimates. Observe a representative set of cases, use application activity data where appropriate, or ask operators to record active work by stage for a limited period. The purpose is not false precision. It is to establish an auditable baseline that can later be compared with pilot performance.
3. Decompose the Workflow at the Task Level
The research evidence strongly supports task decomposition. The BCG experiment found significant gains on tasks within the AI capability frontier and a 19-percent decrease in correct solutions on the selected outside-frontier task. The customer-support study found much larger gains for less experienced workers than for the highest-skill group. The software-development studies report effects with opposite signs in different settings. A process-level average obscures precisely the heterogeneity that determines ROI.
For each stage, classify the work as deterministic, interpretive, generative, retrieval-oriented, judgment-sensitive, or action-oriented. Deterministic work may be better automated through ordinary software or rules. Interpretive work may benefit from a model. High-consequence judgment may remain human-controlled even if AI prepares evidence. The business case should value the change in each stage rather than assume that the agent eliminates the entire unit of work.
4. Separate Straight-Through Work, Review, and Hard Exceptions
Most production automation produces at least three operating paths. Some cases complete automatically. Some are prepared by AI but require human review. Others remain hard exceptions that people investigate directly. The relative volume and handling time of these paths determine the economics.
Suppose a workflow currently takes fifteen active minutes per case. If a future system can complete 50 percent automatically, prepares 35 percent for a three-minute review, and leaves 15 percent as fifteen-minute manual exceptions, the expected human touch time becomes 3.3 minutes per case: zero minutes for half the volume, 1.05 weighted minutes for reviewed cases, and 2.25 weighted minutes for exceptions. That is a 78-percent reduction in modeled touch time, but it is not a claim about any KeenSight deployment. It is an illustrative calculation showing why path mix matters more than a blanket automation rate.
The review path deserves special attention. A system may automate execution while creating a large approval queue. If reviewers need to reconstruct context, the apparent automation benefit can disappear. Measure review time, escalation causes, and reviewer skill level during the pilot so that human-control cost is included rather than treated as free.
5. Distinguish Productivity From Cash Savings
A reduction in touch time creates capacity. It does not automatically reduce expenditure. If ten employees each recover five hours per week but staffing does not change, the company has gained capacity rather than cash. That capacity may still be valuable if it absorbs growth, reduces backlog, improves responsiveness, permits additional revenue-generating work, or defers hiring. The business case should identify which economic mechanism is expected rather than translating every hour into payroll savings.
The customer-support field study illustrates this distinction. Its authors estimated that the firm could theoretically handle the same support volume with about 12 percent fewer worker-hours under the observed productivity improvement. They explicitly caution, however, that demand, staffing, training, and longer-run labor responses may change. The correct financial model would therefore ask whether the organization actually plans to reduce overtime, avoid hires, increase service capacity, or redeploy staff—not simply multiply a 15-percent productivity result by salary expense.
6. Model Worker Heterogeneity
Average effects can hide strategically important distributional patterns. In the QJE support study, the lowest-skill quintile experienced a 36-percent increase in resolutions per hour, while the most skilled workers saw little productivity improvement. Newer workers also moved down the experience curve faster. If a proposed AI system primarily affects onboarding, training, or novice performance, the economic case may be strongest in avoided training time and faster time-to-proficiency rather than average labor reduction across the entire workforce.
This also means that using one blended wage and one blended productivity factor can misprice the opportunity. Segment workers or cases when the process materially differs by experience, complexity, geography, product line, or customer tier. The investment may create large gains for one segment and none for another. A smaller targeted deployment can therefore outperform a broad rollout even if the technology is capable of serving both.
7. Include Quality, Rework, and Error Costs
Speed is not the only economic outcome. Noy and Zhang found quality improvements alongside faster completion on bounded writing tasks. The BCG study found higher quality on within-frontier tasks but lower correctness on the outside-frontier task. The support study found some evidence of small quality declines among the highest-skill workers. These results reinforce a practical rule: ROI should include the cost of correct work, rework, and incorrect automation rather than optimizing handling time alone.
Measure current rework rates, corrections, duplicate entry, follow-up contacts, escalations, or downstream defects. Then define how the AI system could affect them. If automation reduces manual copying but introduces a new model-review process, both changes belong in the model. For high-consequence workflows, attach an expected cost to incorrect automated actions using internal incident history or scenario analysis. Even a low failure probability can dominate ROI if the consequence is large enough.
8. Include Implementation Costs That Usually Disappear From Simple Calculators
A production AI project can include workflow discovery, data mapping, integration engineering, identity and access work, security review, model evaluation, interface development, human-review design, testing, change management, training, and deployment. A document assistant connected to one repository has a different implementation cost from an agent orchestrating CRM, ERP, email, and financial approvals. Treat implementation cost as a work breakdown rather than a single arbitrary percentage of software spend.
Also distinguish one-time and recurring engineering. An integration may require initial build effort and ongoing maintenance when vendor APIs change. Evaluation sets must evolve as the workflow changes. Permissions and credentials need ownership. Knowledge sources require maintenance. Operational support and incident response may be small for a narrow assistant and meaningful for an agent that changes business state.
9. Model Recurring AI and Infrastructure Cost Per Successful Unit
Model costs should be expressed in terms of the workload. Estimate model calls, input and output volume, retrieval operations, tool calls, infrastructure, observability, and storage per case. Then incorporate retry and exception behavior. A workflow with long contexts and repeated tool use can have a different unit cost from a single classification call even if they use the same model.
A useful metric is recurring technology cost per successfully completed unit rather than cost per model call. If an agent uses inexpensive calls but fails frequently, retries several times, or requires expensive human intervention, apparent inference savings may be irrelevant. Unit economics should include the entire operating path.
10. Include Reviewer Capacity as an Operating Cost
Review cost is often underestimated because reviewers are existing employees. If 10,000 cases per month generate a 30-percent escalation rate and each escalation requires six minutes of specialist attention, the workflow creates 300 reviewer-hours per month. Whether or not the company hires new people, those hours have an opportunity cost and may create a capacity bottleneck.
Track escalation by reason. Missing data may be fixed through integration. Policy exceptions may remain permanently human-owned. Low-confidence classification may improve with a model or data change. Reviewer edits may reveal systematic instruction problems. The economic model should become more precise as the organization learns which review categories are structural and which can decline.
11. Account for Adoption and Utilization
Potential productivity has no financial effect if the system is rarely used or routinely bypassed. Adoption assumptions should therefore be separate from technical automation assumptions. Measure the share of eligible cases actually processed through the new workflow, the share of recommendations accepted or edited, and whether users create parallel manual processes because they do not trust the system.
The METR developer study is a particularly useful warning about perception. Participants expected a 24-percent speedup before the trial and still believed they had been accelerated by 20 percent afterward, despite measured completion time being 19 percent longer when AI was allowed. The result is specific to the early-2025 tools, experienced developers, and mature repositories in the study, but it demonstrates why survey sentiment should not substitute for operational measurement.
12. Use External Research as a Prior, Not a Forecast
External evidence can help determine which hypotheses are plausible. The support study suggests that lower-experience workers may gain disproportionately in some knowledge-transfer settings. The writing experiment suggests substantial potential for bounded drafting. The BCG experiment suggests that capability fit matters and that the same users can be helped on some tasks and harmed on another. The Google and METR coding RCTs demonstrate that even the direction of a productivity effect can vary with context. The online-retail field experiments show that marginal business impact can vary from effectively zero to double-digit percentages across workflows in one broader operating environment.
These studies should influence what the pilot measures, not populate the financial spreadsheet as default assumptions. A company evaluating support automation can use the QJE study to justify measuring differences by worker tenure. A consulting workflow can test in-frontier and outside-frontier tasks separately. A coding deployment can compare perceived and actual time. Research becomes most useful when it improves the experiment design.
13. Build Conservative, Expected, and Upside Scenarios
Point estimates imply more certainty than early-stage projects deserve. Build at least three scenarios. Vary eligible volume, straight-through rate, review rate, time per review, exception rate, quality effect, adoption, recurring cost, and implementation cost. The conservative case should represent a plausible outcome where technical performance is useful but materially below the team's target. The upside case should not assume perfection; it should reflect evidence that could reasonably be achieved after iteration.
Sensitivity analysis is often more useful than the headline ROI. If the project becomes unattractive when the review rate moves from 15 to 25 percent, review is a critical pilot metric. If the business case remains strong even when model cost doubles, inference pricing is not the main risk. If value depends almost entirely on reducing headcount, the investment thesis is more fragile than one that creates multiple forms of capacity or quality improvement.
14. Use Payback and Unit Economics Before Sophisticated Finance
For many workflow projects, a few transparent measures are sufficient for the first decision. Annual net benefit equals annual monetized benefits minus recurring operating cost. Simple ROI can be expressed as annual net benefit divided by initial implementation cost. Payback period is initial implementation cost divided by monthly net benefit. Cost per completed unit compares the current process with the proposed process after technology and human review are included.
For multi-year programs or investments with material timing differences, discounted cash flow and net present value may be appropriate. But financial sophistication cannot rescue weak operating assumptions. A detailed NPV based on a guessed 70-percent automation rate is less useful than a simple payback model built from measured pilot data.
15. A Practical Net-Value Equation
A useful conceptual model is: net value equals labor capacity value plus quality and rework value plus throughput or revenue value plus avoided future cost, minus implementation cost, recurring technology cost, human-review cost, maintenance cost, and expected risk cost. Each term should be explicit enough that finance, operations, and engineering can challenge the assumption.
Expected risk cost should not be interpreted as an attempt to price every possible AI incident precisely. It is a way to prevent the model from treating failure as economically free. For a low-impact internal drafting workflow, the risk term may be modest. For a workflow capable of financial, legal, access-control, or customer commitments, scenario analysis around incorrect actions can materially affect the acceptable level of autonomy.
16. Design the Pilot to Replace the Largest Assumptions
The pilot should not merely prove that the agent can work. It should produce the evidence needed to update the business case. Measure baseline and AI-assisted touch time by stage, straight-through completion, review time, exception rate, failure categories, quality, rework, adoption, model/tool cost, system latency, and operator behavior. Where possible, use controlled comparisons rather than relying on before-and-after impressions.
The empirical studies cited here demonstrate the value of controlled measurement. Noy and Zhang randomized access to ChatGPT. The BCG experiment used randomized conditions after a performance baseline. METR randomized real repository tasks to AI-allowed and AI-disallowed conditions. Google's study randomized engineers working on an enterprise-grade task. A production pilot may not support a formal RCT, but teams can still create matched samples, staged rollouts, holdouts, or repeated measures that provide stronger evidence than anecdotes.
17. Recalculate After Deployment
ROI is not a one-time approval artifact. Real volumes, model usage, exception rates, reviewer behavior, adoption, and downstream costs will differ from pre-launch assumptions. Maintain the unit-economic model after deployment and replace assumptions with observed data. This makes it possible to decide whether to expand the workflow, change the model, invest in an integration that removes a common exception, or reduce autonomy where review and risk costs are higher than expected.
Recalculation also prevents an early successful pilot from becoming permanent mythology. A model or workflow can change. Volume can shift toward harder cases. Staff may adapt. A new control may add review time. Technology prices may fall while maintenance costs rise. The economic case should evolve with the operating system.
18. What the Research Actually Says About AI ROI
The strongest conclusion from current productivity research is conditional rather than universal. AI can produce large gains in some bounded knowledge tasks and for some worker segments. It can produce smaller gains in enterprise software work. It can produce no incremental effect in particular workflows. It can make expert workers slower in some settings. It can increase throughput while shifting work toward verification. It can improve quality inside a capability frontier and reduce correctness outside it. This is not evidence that AI ROI is unknowable; it is evidence that ROI has to be measured at the workflow and task level.
For executives, that changes the investment question. The decision is not whether generative AI has demonstrated productivity benefits somewhere in the economy. It has. The decision is whether the proposed system changes this organization's operating equation enough to justify implementation and ongoing ownership. The answer requires internal data.
For a worked domain example, continue with AI Invoice Processing ROI. When the organization has real baseline and pilot data, discuss the business case with KeenSight rather than substituting a generic benchmark for the deployment decision.
Conclusion
A defensible AI automation business case begins with observed workflow volume, active touch time, task composition, exception mix, quality, and the economic value of capacity. It separates straight-through work from review and hard exceptions, distinguishes productivity from cash savings, includes implementation and recurring operating cost, and models risk and human oversight rather than treating them as externalities. It uses ranges rather than a single confident percentage and designs the pilot to replace the most consequential assumptions with evidence.
External research remains valuable precisely because it shows why this discipline is necessary. A 15-percent average productivity increase across 5,172 support agents, a 40-percent time reduction in bounded writing tasks, a 25.1-percent speed improvement for consultants on in-frontier tasks, a 19-percent correctness decline on an outside-frontier task, an approximately 21-percent speed improvement in one Google engineering experiment, and a 19-percent slowdown in METR's early-2025 experienced-developer RCT cannot all be reduced to one AI productivity benchmark. They describe different tasks, workers, tools, and environments. The business case should do the same.
Research and Further Reading
Brynjolfsson, Li, and Raymond, Generative AI at Work, Quarterly Journal of Economics, 2025.
MIT summary of Noy and Zhang's professional-writing experiment.
Dell'Acqua et al., Navigating the Jagged Technological Frontier, Organization Science, 2026.
METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity.
METR, 2026 update on developer-productivity experiment design and selection effects.
Fang et al., Generative AI and Firm Productivity: Field Experiments in Online Retail.
What the Empirical Evidence Shows
These studies measure different technologies, tasks, populations, and outcomes. They are evidence for task-contingent effects—not transferable ROI benchmarks.
+15% Average Productivity
QJE study of 5,172 customer-support agents: issues resolved per hour increased 15% on average; the lowest-skill quintile improved 36%.
40% Faster / 18% Higher Quality
Noy and Zhang's bounded professional-writing experiment found about 40% lower completion time and 18% higher independently rated quality.
+12.2% Tasks / 25.1% Faster
In the 758-consultant BCG study, AI users completed 12.2% more in-frontier tasks and worked 25.1% faster.
19% Less Likely Correct
On the BCG experiment's selected outside-frontier managerial task, AI-assisted participants were 19% less likely to produce the correct solution.
19% Slower
METR's early-2025 RCT of 16 experienced developers across 246 real repository tasks found AI-allowed tasks took 19% longer.
~21% Faster
A randomized study with 96 Google engineers estimated about a 21% reduction in time on a complex enterprise-grade task, with a wide confidence interval.
Inputs for a Defensible ROI Model
Workflow Volume
Items processed over a representative period, including seasonality and peak demand where relevant.
Manual Touch Time
Active staff time by major step rather than total elapsed process time.
Path and Exception Mix
The share of work expected to complete automatically, require review, or remain a hard exception.
Quality and Rework
Observed corrections, follow-up, duplicate handling, defect costs, and downstream impact.
Implementation and Adoption
Discovery, engineering, integrations, testing, security, rollout, training, and actual workflow utilization.
Operating Cost
Model usage, infrastructure, monitoring, maintenance, review effort, retries, and expected risk cost.
Build the Business Case From Your Own Workflow
Use external research to design the hypothesis and the pilot, then replace assumptions with your own volume, touch-time, exception, quality, review, adoption, and operating-cost evidence. The existing invoice calculator remains available for invoice workflows; other processes should start with workflow discovery rather than a borrowed benchmark.
Related Analysis
Continue with research and practical guidance on adjacent AI architecture, governance, and operating-model questions.
AI Customer Support ROI: Why Handle Time Alone Is an Incomplete Business Case
A service-operations framework for evaluating AI customer support ROI across productivity, resolution quality, escalation, repeat contacts, workforce learning, adoption, and operating cost.
AI Document Processing ROI: Measuring the Economics of Intake, Classification and Extraction
A technical framework for evaluating document AI ROI using cost per accepted record, classification and extraction quality, review effort, exception handling, privacy, and downstream integration.
The ROI of AI for RFP and Proposal Response: Measuring SME Capacity, Turnaround and Review Economics
A commercial framework for evaluating AI RFP automation ROI using proposal volume, SME time, content reuse, turnaround, review effort, qualification, response capacity, and revenue opportunity.
