What changed
Gartner forecasts that AI inference costs per agentic workflow will increase more than fivefold through 2028.
For an agent doing a customer service task in banking, token costs frequently represent just 20 to 25 percent of the variable run costs of an AI agent. Human oversight takes most of the rest.
A Gartner survey of 199 service and support leaders, run in April through May 2026, found spending on AI up 38% while overall service and support budgets grew just 2%.
Why this matters for you
The price per token fell. The number of tokens each workflow burns rose faster, and the expensive part was never the tokens.
If your agentic budget treats tokens as the main cost, our read of McKinsey's banking example is that the split runs the other way.
What you can do
Ask for one number before you approve the next AI spend: fully loaded cost per completed task, with human review time counted in. If nobody can produce it, you have your answer.
Watch out for
Gartner's number is a forecast through 2028, not a measurement. The McKinsey split comes from customer service work in banking, so treat it as a shape rather than your number.
Variable run costs of an AI agent performing a customer service task in banking
McKinsey scopes this to an agent performing a customer service task in banking and puts human oversight at 70 to 75 percent of the variable cost.
McKinsey QuantumBlack, 24 Aug 2026