How to Tell If an Automation Paid for Itself
By Nick Bryant, Co-Founder and CTO, SMB Investor Network
7 min read
In brief
Feeling faster is not the same as being faster. Measure the task before you build, count review time, and check whether any spending actually changed.
Most owners can't say whether their automation paid for itself, because nobody timed the work before it started. In a controlled study of experienced developers, people predicted AI would make them 24% faster, measured 19% slower, and still believed afterward they'd been 20% faster1. Feeling faster is not evidence.
I build agents, and I'm as prone to that trap as anyone. The fix is boring: a stopwatch before, a count after. Here's the arithmetic that actually tells you whether an automation paid off.
ROI calculators from vendors count every freed hour as cash. Nobody tells the owner to measure before, count review time, or check whether any spending actually changed. That's the gap this post covers, with a worked example you can swap your own numbers into.
Feeling faster is not the same as being faster
METR's randomized trial put 16 experienced open-source developers on 246 real tasks. Allowed to use early-2025 AI tools, the tasks took 19% longer, not shorter. The developers had forecast a 24% speed-up, and afterward they still believed they'd been sped up by 20%1.
Small businesses report the same pattern in softer form. Among small business owners currently using AI, 30% named higher productivity as the impact and 8% named lower operating costs, in a single-choice survey question2. That's a gap worth sitting with: 30% reported higher productivity, while 8% reported lower operating costs2.
In Denmark, a study linking chatbot-adoption surveys to actual payroll records through May 2025 found a precise null effect on earnings and hours, ruling out any effect larger than 2%3. This one didn't ask anyone how they felt. It checked the payroll records directly, and found no effect larger than 2% on earnings or recorded hours3.
Takeaway: self-report overstates the win. Measure the task, don't ask how it felt.
What real gains look like when someone measured them
Real, measured gains exist, and they're narrower than the pitch. A study of 5,179 customer support agents found a generative AI assistant raised issues resolved per hour by 14% on average and 34% for novices, with almost no effect on the most experienced4. St. Louis Fed survey respondents who used generative AI reported saving 5.4% of their work hours, about 2.2 hours a week, and the aggregate share of all US work hours saved by generative AI rose from 1.6% to 2.2% between Q3 2024 and Q2 202656.
Takeaway: measured gains varied, with a 14% average and 34% for novices in the support-agent study4. An illustrative pitch promising 50% deserves a stopwatch, not trust.
The payback math, line by line
Take one invoice-entry automation handling 400 invoices a month. A bookkeeper's loaded rate (BLS median wage grossed up by the June 2026 employer benefit share) runs $34.80 an hour78. Keying an invoice by hand takes 4 minutes; reviewing the automated output takes 1 minute; about 5% need a 4-minute redo; the illustrative Sonnet 5 token cost of about a cent per invoice assumes 3,000 input tokens and 400 output tokens at list prices9.
Same automation, different payback
A calculator for one invoice-entry automation handling 400 invoices a month. Adjusting the rework rate and the number of rule changes routed through a developer each month changes the net monthly saving and the payback period against a $1,544 one-time setup cost. At the default 5% rework rate and no developer-routed changes, the net saving is about $525 a month and payback is about 2.9 months.Gross time replaced
Review time
Rework
Model tokens
Upkeep
Tool fee
Source: Illustrative worked example: 400 invoices a month, bookkeeper loaded rate $34.80/h (BLS May 2025 wage grossed up by the June 2026 ECEC benefit share)78. Sonnet 5 token cost read 2026-09-269. Volumes, minutes, upkeep, tool fee and setup cost are illustrative assumptions, not a measured result.
Figure data
| Invoices that need rework | Monthly labor saved before costs | Net monthly saving | Payback on setup cost |
|---|---|---|---|
| 0% | $928 | $572 | 2.7 months |
| 5% | $928 | $525 | 2.9 months |
| 10% | $928 | $479 | 3.2 months |
| 20% | $928 | $386 | 4.0 months |
Illustrative: 400 invoices a month at $34.8/hour loaded, $1,544 setup, no developer changes.
Run the numbers at the default 5% rework rate: $928 of gross time replaced, minus $232 of review, minus $46 of rework, minus about $5 of tokens, minus $70 of upkeep, minus a $50 tool fee. That leaves about $525 a month, against a one-time setup cost of $1,544. Payback lands around 2.9 months. All of these inputs except the wage, the benefit share and the token price are illustrative.
Takeaway: review time is the single biggest deduction, and it's the one line most vendor ROI calculators leave out entirely.
Hours saved are not cash until something changes
The same automation with no change in staffing, overtime or hiring is a monthly cost, not a saving. Adoption surveys back this up: AI-related headcount decreases happen in only 2% of AI-using firms10, and 98% of small employers using AI say it hasn't changed their workforce size11.
In this illustrative example, if the freed hours don't go anywhere, the cash result of that same automation is negative: about $55 a month, once you count only the tokens and the tool fee against nothing gained. Write down, before you build anything, what the freed hours will actually become: overtime cut, a hire deferred, a backlog cleared, or the owner out of the invoice queue. If you can't name it, the saving won't show up in the bank account.
This isn't a hypothetical trap. For this automation, check whether spending stays fixed even when the task takes less time. A subscription bill arrives every month whether or not anyone acted on the freed capacity. Cash only moves when a person's schedule, a headcount decision or a vendor payment actually changes because of it. Ask the question at the moment you're deciding to build, not six months later when someone asks why the tool bill never went down.
Takeaway: name where the freed hours go before you build. Otherwise the automation is a subscription, not a saving.
Where payback dies: tuning
Rules change. Approval limits, which vendors auto-post, the wording of a reminder, the day a report runs. If every one of those changes means an engineer sitting down with an agent, you pay for it in tokens and hours every single time.
Here's how that breaks, adapted from a lesson I learned building something else entirely: the agent spins up and re-reads the whole codebase, over-plans the change, edits the one line you wanted plus ten you didn't, and runs every test in the suite. Multiply that by every small tuning decision a business makes in a month.
The fix is the same one I use in my own projects: put every rule that could change, dollar limits, vendor lists, schedules, wording, in one plain settings file or sheet the owner can edit directly. The logic reads from it. Nobody opens the code to change a threshold. My own version of this was a game, not a business (ROLL_TIME = 0.62, easy to swap the animation speed), but the principle carries over exactly: separate the knobs from the logic.
Run four rule changes a month through a developer at $100 an hour, and the same automation's net monthly saving drops from $525 to $125, and payback stretches from about 3 months to about 12 [illustrative]. The calculator above shows the same shift: move the rule-changes slider and watch the payback period move with it.
Notice what didn't change in that comparison: the model, the wage, the volume. The only thing that moved was who could touch the rules and how often. That's the lever owners miss when they're shopping for a "better" AI model instead of asking who maintains the one they already have.
Takeaway: separate the rules from the logic so the owner can tune the automation without paying a developer every time.
The four numbers to track weekly
A stopwatch and a settings file are cheap. A dashboard is not required. Once something's live, track four numbers on the weekly scorecard:
- Volume handled.
- Exception rate: the share a person had to redo.
- Review minutes per item.
- The cash line that actually changed.
Under an illustrative twelve-month payback policy, a climbing exception rate or longer payback would trigger a review of whether to turn it off, not an automatic decision.
None of these four numbers require a data team. A shared sheet and a Monday habit is enough. The point isn't precision to the cent, it's catching the automation that quietly stopped paying off before it's been running unexamined for a year. Put the same four numbers next to the ones from the setup week, and the trend line tells you more than any single week's snapshot.
What I'd do
- Before touching anything, time 20 runs of the task by hand and write down the monthly volume. No baseline, no ROI.
- Write one sentence on where the freed hours will go: overtime cut, a hire deferred, a backlog cleared, or the owner out of the queue. If you can't write it, don't expect it in the bank account.
- Build the smallest version. Put every rule that could change in one settings file or sheet the owner can edit; the logic reads from it.
- For four weeks, track volume, exception rate, review minutes per item and the cash line from step 2, every week.
- At week four, run the waterfall: gross minutes minus review, rework, tokens, upkeep and fees, then divide setup cost by the net. For example, under a 12-month payback policy, if payback exceeds that threshold or the exception rate climbs, review whether to turn it off and keep the process map.
- Re-check tool and model prices every quarter. They move more than you'd expect12.
For the piece this leans on most, see what an AI agent actually costs per task: the same review-and-redo math that decides payback here is what prices a single task there. And before building anything, work through what to automate first and the manual work audit to make sure you're timing a task worth timing.
Source notes
- METR, early-2025 AI and experienced developer productivity1.
- NFIB, Small Business and Technology 2025112.
- Humlum and Vestergaard, "Large Language Models, Small Labor Market Effects" (NBER w33777)3.
- Brynjolfsson, Li and Raymond, "Generative AI at Work" (NBER w31161)4.
- St. Louis Fed, "The Impact of Generative AI on Work Productivity"5; FRED Blog update6.
- US Census Bureau, CES Working Paper 26-2510.
- Federal Reserve Board, FEDS Notes on AI adoption13.
- JPMorganChase Institute, understanding AI use among small businesses14.
- BLS wage data and Employer Costs for Employee Compensation78; Anthropic pricing9.
- Goldman Sachs 10KSB15.
Sources
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (2025-07-10, updated Feb 2026) ↑
- NFIB, Small Business and Technology 2025, page 6 ↑
- Humlum and Vestergaard, Large Language Models, Small Labor Market Effects (NBER w33777, revised March 2026) ↑
- Brynjolfsson, Li and Raymond, Generative AI at Work (NBER w31161; Quarterly Journal of Economics 2025) ↑
- St. Louis Fed, The Impact of Generative AI on Work Productivity (Bick, Blandin, Deming; 2025-02-27) ↑
- FRED Blog, St. Louis Fed, Does generative AI save time at work? (August 2026) ↑
- BLS Occupational Outlook Handbook, Bookkeeping, Accounting, and Auditing Clerks ↑
- BLS, Employer Costs for Employee Compensation, June 2026 (released 2026-09-09) ↑
- Anthropic, Claude pricing page (read 2026-09-26) ↑
- US Census Bureau, CES Working Paper 26-25, The Microstructure of AI Diffusion (Bonney, Breaux, Dinlersoz, Foster, Haltiwanger, Pande; April 2026) ↑
- NFIB, Small Business and Technology 2025 (24% read in the PDF; size split and 98% from ASBN summary of the release) ↑
- Google AI for Developers, Gemini API pricing (read 2026-09-26) ↑
- Federal Reserve Board, FEDS Notes, Monitoring AI Adoption in the U.S. Economy (Jeffrey S. Allen, 2026-04-03) ↑
- JPMorganChase Institute, Understanding the use of AI among small businesses (2026-04-14) ↑
- Goldman Sachs 10,000 Small Businesses Voices survey ↑
