Why this matters now:Gartner's August 2026 forecast puts inference at about $23.3 billion of $42 billion in AI-optimized infrastructure spend this year, the first time it has passed training, and expects the share to reach 59 percent in 2027. Gartner also reports that only 44 percent of organizations have adopted financial guardrails for AI, as cited by CIO. India's 2,117 GCCs now run AI for parent companies at scale, with more than 1,200 holding embedded AI and machine learning capabilities according to Nasscom and Zinnov, so the question of who pays for each token lands at the GCC's door first.

The Invoice Arrives in the GCC Before Anyone Agrees Who Owns It

A platform team inside a GCC ships an internal assistant and, a few months later, an agent that triages support tickets. Three business units at the parent adopt them. Usage climbs, the cloud bill climbs with it, and finance sees one line labeled AI platform booked to the GCC's cost center. The parent asks why the center's run cost doubled. The GCC answers that the business units are using more. Nobody in the room can say which workflow drove the increase, which team started it, or what the company got in return.

The usual reassurance is that models are getting cheaper. Per-token prices are falling, and Goldman Sachs estimates the decline at 60 to 70 percent a year, as reported by Tech Times. The same report cites Goldman's projection that global token usage grows 24-fold between 2026 and 2030. Volume is winning against price. Agentic design adds to it: Gartner's March 2026 analysis found that agentic models need 5 to 30 times more tokens per task than a standard chatbot query. A cheaper token multiplied by thirty tokens per task is not a smaller bill.

A GCC feels this sooner than most because of where it sits. The parent sets the budget, business units create the demand, and the center runs the shared platform in between. Whatever funding model your center uses, an AI line item that cannot be traced to a consumer turns into a quarterly negotiation. The negotiation is not really about money. It is about the absence of a meter, and a meter is something engineering can build.

"A token price tells you what the model costs. Only a cost per outcome tells you whether the work was worth doing."
55%
Share of 2026 AI-optimized IaaS spend that goes to inference ($23.3B of $42B), the first year it exceeds training. Source: Gartner forecast, August 2026, as reported by Tech Times.
5–30x
Tokens an agentic model needs per task compared with a standard chatbot query. Source: Gartner, March 2026, as reported by Tech Times.
0
Public benchmarks we could find on how GCCs allocate AI inference cost to parent business units. Source: 10decoders Editorial research for this post, October 2026. The models below are engineering judgment, not survey data.

Six Places a GCC's AI Cost Goes Missing Between the Model and the Business Unit

Cost leakWhat goes wrong in practiceSeverity
Shared endpoints with no consumer tagEvery call reaches the model through one gateway service account, so the invoice cannot be split by team, workflow, or environmentCritical
Context re-sent on every agent stepEach step in an agent loop sends the conversation and tool results again. The Stanford Digital Economy Lab put re-sent context at about 62 percent of agent inference cost, as reported by Tech TimesCritical
Idle and forgotten endpointsA dedicated endpoint left running after a pilot ends keeps billing. Tech Times reports hosting costs of $50 to $70 a day for an idle deploymentHigh
Pilot budgets carried into productionA pilot is funded as a finite project, but inference runs every day. The budget has no line for month fourteenHigh
Platform work billed to the platformEvaluation runs, guardrail checks, retrieval, and logging are charged to the center, not to the workflow that triggered themModerate
Currency and period mismatchUsage is billed in dollars daily and recovered from business units in local currency monthly, so forecasts drift between the twoLower

Not sure where your GCC's AI cost is leaking?

10decoders works with GCC platform and finance teams on metering, tagging, and unit economics for AI workloads. We trace one real workflow from invoice to outcome and show you what share of the bill you can attribute today.

Book a Free AI Assessment →

Why Cost per Outcome Beats Cost per Token for a Shared-Service Center

A business unit does not buy tokens. It buys a closed ticket, a reviewed contract, a resolved claim, a screened alert. Cost per outcome is everything it took to produce one accepted result: model calls, retrieval, guardrail checks, evaluation runs, and the minutes a person spent reviewing. Rework belongs in the numerator. An output that was rejected still cost tokens and still cost a reviewer, and ignoring it makes a weak workflow look cheap.

The absence of a per-consumer meter has already caught well-known organizations. Tech Times reports that Uber exhausted its annual AI budget in four months, with per-engineer monthly costs between $500 and $2,000, according to its CTO. We do not know the internal details, but the pattern is familiar: spend grew faster than anyone could see, because nothing connected a person or a workflow to a charge until the bill arrived. The FinOps Foundation's 2026 State of FinOps report, as summarized by BERI, finds 98 percent of practitioners now manage AI spend, up from 63 percent in 2025 and 31 percent in 2024. Managing it is now normal. Attributing it is the part many centers have not built.

Our advice is to start with showback, not chargeback. For one quarter, publish each business unit's AI cost and cost per outcome without billing it. Behavior usually changes once a team sees its own number, and you learn which allocations are disputed before money moves. Chargeback then follows with the disputes already resolved, and it comes with per-workflow ceilings so that a runaway loop stops at a limit rather than at the end of the quarter.

Stage 1
One pooled bill

The AI Line Item

All inference is billed to the center. Cost is explained by anecdote, and every budget review ends with a debate about whose usage grew.

Stage 2
Visible but not billed

Tagged Showback

Calls carry consumer and workflow tags, and each business unit receives a monthly cost report. Ceilings exist on paper, and outcomes are counted for the biggest workloads.

Stage 3
Priced per result

Cost-Per-Outcome Chargeback

Business units are billed per accepted outcome with agreed definitions, rework included. Each workflow has a ceiling, step limits, an alert, and a named owner.

Checklist: Can Your GCC Price One AI Workflow in an Afternoon?

Choose a single production workflow and test each item. If you cannot show it as a tag, a report, or a named person, it is a gap.

What a priceable AI workflow includes
Every inference call carries a consumer tagBusiness unit, workflow ID, and environment travel with each request, and the gateway rejects calls that arrive without them.
Each workflow has a named business ownerOne person on the consumer side answers for the workflow's volume and value, not just a shared mailbox.
Agent steps and context size are meteredYou can see steps per task and tokens per step, so a loop that re-sends a growing context shows up as a line, not a surprise.
Idle endpoints expire by defaultDedicated deployments carry an end date and an owner, and the weekly report lists any that outlived their pilot.
The outcome is defined with the business unitA closed ticket or a reviewed document is written down in plain terms, including what counts as accepted.
Rework is counted in the costRejected outputs and reviewer minutes sit in the same calculation, so the cost per outcome reflects the real process.
Ceilings and alerts exist per workflowA monthly limit, a step limit for agent loops, and a page to a named person when spend reaches the threshold.
Showback is published and variance is explainedA monthly report goes to each business unit, and a forecast miss of more than the agreed margin gets a written reason.
"Charge the business unit for the outcomes it asked for, and hold the GCC accountable for the waste it could have prevented."

What to Do This Week

01Trace last month's AI invoice to workflows

Take last month's AI cost and group it by endpoint, API key, and service account. Then try to assign each group to a workflow and a business unit. Record the share you can attribute and the share you cannot. That unattributed percentage is your baseline, and it is the number to bring to the first conversation with finance.

02 Require tags at the gateway

Add three required fields to every inference call: business unit, workflow ID, and environment. Start in log-only mode for a week to see who is missing them, then reject untagged calls in non-production first. Most of the work is telling the teams that call the model through shared libraries, so update the library once and the tags follow.

03 Compute cost per accepted outcome for one workflow

Pick the workflow with the highest spend. Add model calls, retrieval, guardrails, evaluation runs, and reviewer minutes, then divide by accepted outcomes for the same period. Compare the result with the cost of producing the same outcome manually. If the AI path is not cheaper or faster, you have learned that before scaling it, which is the cheaper time to learn it.

04 Set one ceiling and one alert

For that same workflow, set a monthly ceiling at about 120 percent of forecast, add a step limit to the agent loop, and route an alert to a named owner. Review the list of idle endpoints every Monday and shut down anything without an owner. One working ceiling will teach you more about your gaps than another round of planning.

Let 10decoders Trace Your GCC's AI Spend to the Outcomes It Buys

We review one production workflow end to end: tagging at the gateway, metering of agent steps, idle resource policy, outcome definitions, showback reporting, and ceilings. You leave with a ranked leak list and a chargeback design your finance and business-unit leads can adopt.