What an AI Agent Costs Per Task vs. a Human Hour
By Nick Bryant, Co-Founder and CTO, SMB Investor Network
7 min read
In brief
Price a back-office task both ways: the model bill and the loaded human minutes. The token cost is rarely the number that decides anything.
In an illustrative routine task, keying one vendor invoice costs about a cent in model fees and about $2.32 in bookkeeper time, using published model prices and wage data123. The cent was never the decision.
I build agents, and the token counter is the one number people ask about first. It's also the smallest number in the equation. The cost that decides whether an automation pays off is the human minutes around the model call: the review, the exceptions, and the redo work when a run fails. Here's how to price a task both ways before you buy anything.
An owner shopping for automation usually gets shown a per-seat SaaS price or a listicle about AI replacing staff. Neither one does the arithmetic at the level that matters: one task, priced twice, with real numbers on both sides. That's the gap this post fills.
What a token costs in September 2026
A token is about four characters, or three quarters of a word. A hundred tokens run about seventy five words4. The API prices in the table below are quoted per million tokens, split into input (what you send) and output (what comes back).
| Model | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
| Claude Haiku 4.5 | $1 | $5 |
| Claude Sonnet 5 | $2 | $10 |
| Claude Opus 5.5 | $4 | $20 |
| OpenAI gpt-6-luna | $0.10 | $0.50 |
| OpenAI gpt-6-sol | $2 | $10 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 |
| Gemini 3.8 Flash | $0.75 | $3.75 |
Prices read 2026-09-26156. Batch processing runs about half off across providers, and cached input tokens are cheaper still. None of that is stable. Gemini 3.8 Flash doubles its listed price on January 1, 2027, and inference for GPT-3.5-level output fell more than 280-fold between November 2022 and October 202467. Whatever number you use today, put a date on it and check again next quarter.
That instability cuts both ways. A model that looks expensive this quarter can be a rounding error next quarter, and a "cheap" model picked once can quietly become the expensive option once a provider repriced it. I keep the actual list prices in a settings file, not in a spreadsheet someone built once and never opened again.
Takeaway: model prices are cheap and they move. Budget in ranges, and re-check quarterly rather than trusting last year's number.
Six back-office tasks, priced both ways
Here's the same six tasks priced as loaded human minutes and as model tokens. The human side uses BLS May 2025 median wages for bookkeeping clerks, customer service reps and general office clerks, grossed up by the private-industry wage share of total compensation (70%, per the June 2026 Employer Costs for Employee Compensation release) to get a loaded hourly rate2893. The model side uses Claude Sonnet 5 pricing1. Token counts and minutes are illustrative assumptions; only the unit prices and wages are sourced.
The model line vs the people line
Six back-office tasks plotted on a log dollar scale from a hundredth of a cent to a hundred dollars, comparing loaded human cost against the Sonnet 5 model cost for the same task. Human cost ranges from 58 cents to 52 dollars and 20 cents; the Sonnet 5 cost for the same six tasks ranges from four tenths of a cent to a dollar thirteen. A faint tick marks the cheapest listed model on each row.Categorize a bank transaction
$0.580 human · $0.0040 model
Key a vendor invoice
$2.32 human · $0.012 model
Draft a customer email reply
$3.08 human · $0.017 model
Summarize a 30-minute call
$5.15 human · $0.024 model
Weekly KPI memo from exports
$52.20 human · $0.080 model
25-step agentic reconciliation
$17.40 human · $1.13 model
Filled dot: loaded human cost per task. Hollow dot: Sonnet 5 token cost. Faint tick: cheapest listed model (gpt-6-luna). Log scale: equal spacing is a 100x change.
Source: Token counts and minutes are illustrative assumptions. Prices: Anthropic, OpenAI, read 2026-09-2615. Wages: BLS May 2025 medians, loaded with the June 2026 ECEC benefit share2893.
Figure data
| Task | Loaded human cost | Sonnet 5 cost | Cost on the cheapest listed model (gpt-6-luna) |
|---|---|---|---|
| Categorize a bank transaction | $0.58 | $0.004 | $0.0002 |
| Key a vendor invoice | $2.32 | $0.01 | $0.0006 |
| Draft a customer email reply | $3.08 | $0.02 | $0.0008 |
| Summarize a 30-minute call | $5.15 | $0.02 | $0.0012 |
| Weekly KPI memo from exports | $52.20 | $0.08 | $0.004 |
| 25-step agentic reconciliation | $17.40 | $1.13 | $0.06 |
Per task; token counts and minutes are illustrative assumptions.
Look at the gap in these illustrative tasks. On a single-pass task, categorize, extract, draft, the model is a hundred to a thousand times cheaper than the minutes it replaces. That gap is real in the example, and it's also not the interesting part of this post. A cheap API call was never going to lose to a $30-an-hour clerk.
Takeaway: in this illustrative comparison, the model bill for single-pass tasks (categorize, extract, first-draft) is low relative to the cost of human minutes. The model bill is a rounding error next to the minutes.
Price the completed task, not the call
The interesting part starts once a task needs more than one step. A multi-step agent has to plan, act, check its own work, and sometimes redo it. The formula for what it actually costs:
Token cost + review minutes + (failure rate × human redo)
Take an illustrative 25-step agentic vendor-statement reconciliation, no caching: about $1.13 in Sonnet 5 tokens, against $17.40 in bookkeeper time to do the whole thing by hand, using published model prices and wage data123. Run the formula at three success rates:
What a completed task costs
Three stacked bars build up the expected cost of a completed reconciliation task at 30%, 60% and 90% agent success, each from token cost, review time and redo time, against a dashed reference line at $17.40 for an all-human reconciliation. The 30% bar settles at $14.18, the 60% bar at $9.83, and the 90% bar at $5.48.30% success
60% success
90% success
Bookkeeper does it by hand: $17.40 [illustrative].
Source: 25-step agentic vendor-statement reconciliation, illustrative tokens and minutes; Sonnet 5 pricing read 2026-09-261. Bookkeeper loaded rate from BLS May 2025 wages and the June 2026 ECEC benefit share23. Success-rate anchor: best public agent fully completed 30.3% of 175 simulated-company tasks10.
Figure data
| Agent success rate | Tokens | Review | Redo | Expected cost per completed task |
|---|---|---|---|---|
| 30% | $1.13 | $0.87 | $12.18 | $14.18 |
| 60% | $1.13 | $1.74 | $6.96 | $9.83 |
| 90% | $1.13 | $2.61 | $1.74 | $5.48 |
| All human | — | — | — | $17.40 |
25-step reconciliation task; illustrative tokens and minutes.
At 30% success, the expected cost is $14.18: tokens plus a short review of the ones that worked, plus a full human redo of the seven in ten that didn't10. At 60%, it's $9.83. At 90%, it's $5.48, against $17.40 for a person doing the whole thing. The 30% row isn't a hypothetical low bar. On TheAgentCompany, a public benchmark of 175 tasks in a simulated company, the best agent (Gemini 2.5 Pro) fully completed 30.3% of tasks, averaging 27.2 steps and $4.20 per task, with administrative and finance work among the hardest categories10.
One more number worth knowing: in this illustrative no-cache reconciliation, input tokens are about 89% of the bill on Sonnet 5, using its published token prices1. The agent is spending most of its money reading, not writing. Caching that context (90% of it, in this example) drops the token cost from $1.13 to about $0.32. Caching lowers token costs in this illustrative case.
Takeaway: below about half success, an agent on a multi-step task mostly adds a step before the human does the job anyway.
How the naive version breaks
I've watched agents fail the same handful of ways enough times to write them down plainly:
- Context re-reads on every step. Most of the token bill is the agent reading, not writing.
- Retries and loops when the first attempt doesn't land, quietly multiplying cost with no extra output to show for it.
- A stronger model picked "to be safe," doubling the bill for no measurable quality gain over the cheaper one.
- Nobody owns review, so a wrong answer ships instead of getting caught before it reaches a customer or a ledger.
- Prices change and nobody notices until the invoice is bigger than expected, because the price was never dated in the first place.
- The one case the agent can't handle is usually the expensive one, and it still needs a person, on top of whatever the agent already spent trying.
Takeaway: most of the token bill is the agent re-reading its own context, not producing new output.
Where this leaves a $1 to $20 million business
Adoption is still thin and narrow. Among US firms using AI, 65% use it for three tasks or fewer11. The median small business paying for AI spent about $28 a month in 202512, and adoption in construction and transportation sits under 9% even as it clears 30% in professional services and 39% in information13. This isn't a business running a fleet of autonomous agents. It's a business trying one tool on one task.
The measured wins are in narrow, supported tasks, not autonomous ones. A study of 5,179 customer support agents found a generative AI assistant raised issues resolved per hour by 14% on average and 34% for novice or low-skilled workers, with almost no effect on the most experienced14. That's a person doing the task with help, reviewed as they go, not software running the task alone.
Among 4.6 million Chase small business customers, 17.7% had adopted paid AI services by December 2025, up from 5.2% in January 2023, and 72.5% of adopters paid for a single AI service, not a stack of them12. That matches what the task math above says: start with one task, not a platform.
Takeaway: start with a single-pass task a person already reviews. That's where the public evidence actually points.
What I'd do
I pay these token bills myself, and the arithmetic above is the same one I run before adding anything to an agent's plate.
- Pick one task that repeats at least 100 times a month and already has a person checking it. Invoices, transaction coding and first-draft customer replies are the usual candidates.
- Time 20 of them by hand this week. Write down the minutes, not a guess.
- Run the same 20 through a model with the real documents. Log tokens in and out from the API response, not from a calculator.
- Have the person who does the task today review every output, and time that review. Count how many they had to redo.
- Compute the cost per completed task: tokens plus review minutes plus redo minutes, at the loaded hourly rate. Compare it to step 2.
- If it's a multi-step agent and fewer than half the runs finish clean, stop. Keep the single-pass piece (extract, draft) and let a person do the rest.
- Put the model name, prices and thresholds in one settings file the owner can read. Prices change; one listed price above doubles on January 16. Re-check every quarter.
Agents are good at a narrow task someone already defined. They're bad at deciding what the task is, and administrative and finance tasks were among the hardest categories in a public benchmark set in a simulated software company10. If you want the full automation costs picture, that post covers the training and review time a subscription price never shows. And before you run this math on your own numbers, check whether an automation you already built paid for itself, because the same review-and-redo arithmetic applies there too.
Source notes
- Anthropic, Claude pricing, read 2026-09-261.
- OpenAI, API pricing, read 2026-09-265.
- Google AI for Developers, Gemini API pricing, read 2026-09-266.
- OpenAI Help Center, what tokens are4.
- Stanford HAI, AI Index Report 20257.
- BLS Occupational Outlook Handbook, bookkeeping/customer service/office clerk wages, May 2025289.
- BLS, Employer Costs for Employee Compensation, June 20263.
- Brynjolfsson, Li and Raymond, "Generative AI at Work" (NBER w31161)14.
- Xu et al., TheAgentCompany (CMU, NeurIPS 2025 Datasets and Benchmarks)10.
- JPMorganChase Institute, understanding AI use among small businesses1213.
- US Census Bureau, CES Working Paper 26-2511.
- Census Bureau Business Trends and Outlook Survey15; Goldman Sachs 10KSB16.
Sources
- Anthropic, Claude pricing page (read 2026-09-26) ↑
- BLS Occupational Outlook Handbook, Bookkeeping, Accounting, and Auditing Clerks ↑
- BLS, Employer Costs for Employee Compensation, June 2026 (released 2026-09-09) ↑
- OpenAI Help Center, What are tokens and how to count them ↑
- OpenAI API pricing page (read 2026-09-26) ↑
- Google AI for Developers, Gemini API pricing (read 2026-09-26) ↑
- Stanford HAI, AI Index Report 2025 ↑
- BLS Occupational Outlook Handbook, Customer Service Representatives ↑
- BLS Occupational Outlook Handbook, General Office Clerks ↑
- Xu et al., TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks (CMU; NeurIPS 2025 Datasets and Benchmarks) ↑
- US Census Bureau, CES Working Paper 26-25, The Microstructure of AI Diffusion (Bonney, Breaux, Dinlersoz, Foster, Haltiwanger, Pande; April 2026) ↑
- JPMorganChase Institute, Understanding the use of AI among small businesses (2026-04-14) ↑
- JPMorganChase Institute, Understanding the use of AI among small businesses (2026-04-14) ↑
- Brynjolfsson, Li and Raymond, Generative AI at Work (NBER w31161; Quarterly Journal of Economics 2025) ↑
- US Census Bureau, Business Trends and Outlook Survey (America Counts story) ↑
- Goldman Sachs 10,000 Small Businesses Voices survey ↑
