The Invoice Arrives in the GCC Before Anyone Agrees Who Owns It
A platform team inside a GCC ships an internal assistant and, a few months later, an agent that triages support tickets. Three business units at the parent adopt them. Usage climbs, the cloud bill climbs with it, and finance sees one line labeled AI platform booked to the GCC's cost center. The parent asks why the center's run cost doubled. The GCC answers that the business units are using more. Nobody in the room can say which workflow drove the increase, which team started it, or what the company got in return.
The usual reassurance is that models are getting cheaper. Per-token prices are falling, and Goldman Sachs estimates the decline at 60 to 70 percent a year, as reported by Tech Times. The same report cites Goldman's projection that global token usage grows 24-fold between 2026 and 2030. Volume is winning against price. Agentic design adds to it: Gartner's March 2026 analysis found that agentic models need 5 to 30 times more tokens per task than a standard chatbot query. A cheaper token multiplied by thirty tokens per task is not a smaller bill.
A GCC feels this sooner than most because of where it sits. The parent sets the budget, business units create the demand, and the center runs the shared platform in between. Whatever funding model your center uses, an AI line item that cannot be traced to a consumer turns into a quarterly negotiation. The negotiation is not really about money. It is about the absence of a meter, and a meter is something engineering can build.
"A token price tells you what the model costs. Only a cost per outcome tells you whether the work was worth doing."
Six Places a GCC's AI Cost Goes Missing Between the Model and the Business Unit
| Cost leak | What goes wrong in practice | Severity |
|---|---|---|
| Shared endpoints with no consumer tag | Every call reaches the model through one gateway service account, so the invoice cannot be split by team, workflow, or environment | Critical |
| Context re-sent on every agent step | Each step in an agent loop sends the conversation and tool results again. The Stanford Digital Economy Lab put re-sent context at about 62 percent of agent inference cost, as reported by Tech Times | Critical |
| Idle and forgotten endpoints | A dedicated endpoint left running after a pilot ends keeps billing. Tech Times reports hosting costs of $50 to $70 a day for an idle deployment | High |
| Pilot budgets carried into production | A pilot is funded as a finite project, but inference runs every day. The budget has no line for month fourteen | High |
| Platform work billed to the platform | Evaluation runs, guardrail checks, retrieval, and logging are charged to the center, not to the workflow that triggered them | Moderate |
| Currency and period mismatch | Usage is billed in dollars daily and recovered from business units in local currency monthly, so forecasts drift between the two | Lower |
Not sure where your GCC's AI cost is leaking?
10decoders works with GCC platform and finance teams on metering, tagging, and unit economics for AI workloads. We trace one real workflow from invoice to outcome and show you what share of the bill you can attribute today.
Book a Free AI Assessment →Why Cost per Outcome Beats Cost per Token for a Shared-Service Center
A business unit does not buy tokens. It buys a closed ticket, a reviewed contract, a resolved claim, a screened alert. Cost per outcome is everything it took to produce one accepted result: model calls, retrieval, guardrail checks, evaluation runs, and the minutes a person spent reviewing. Rework belongs in the numerator. An output that was rejected still cost tokens and still cost a reviewer, and ignoring it makes a weak workflow look cheap.
The absence of a per-consumer meter has already caught well-known organizations. Tech Times reports that Uber exhausted its annual AI budget in four months, with per-engineer monthly costs between $500 and $2,000, according to its CTO. We do not know the internal details, but the pattern is familiar: spend grew faster than anyone could see, because nothing connected a person or a workflow to a charge until the bill arrived. The FinOps Foundation's 2026 State of FinOps report, as summarized by BERI, finds 98 percent of practitioners now manage AI spend, up from 63 percent in 2025 and 31 percent in 2024. Managing it is now normal. Attributing it is the part many centers have not built.
Our advice is to start with showback, not chargeback. For one quarter, publish each business unit's AI cost and cost per outcome without billing it. Behavior usually changes once a team sees its own number, and you learn which allocations are disputed before money moves. Chargeback then follows with the disputes already resolved, and it comes with per-workflow ceilings so that a runaway loop stops at a limit rather than at the end of the quarter.
The AI Line Item
All inference is billed to the center. Cost is explained by anecdote, and every budget review ends with a debate about whose usage grew.
Tagged Showback
Calls carry consumer and workflow tags, and each business unit receives a monthly cost report. Ceilings exist on paper, and outcomes are counted for the biggest workloads.
Cost-Per-Outcome Chargeback
Business units are billed per accepted outcome with agreed definitions, rework included. Each workflow has a ceiling, step limits, an alert, and a named owner.
Checklist: Can Your GCC Price One AI Workflow in an Afternoon?
Choose a single production workflow and test each item. If you cannot show it as a tag, a report, or a named person, it is a gap.
"Charge the business unit for the outcomes it asked for, and hold the GCC accountable for the waste it could have prevented."
What to Do This Week
01Trace last month's AI invoice to workflows
Take last month's AI cost and group it by endpoint, API key, and service account. Then try to assign each group to a workflow and a business unit. Record the share you can attribute and the share you cannot. That unattributed percentage is your baseline, and it is the number to bring to the first conversation with finance.
02 Require tags at the gateway
Add three required fields to every inference call: business unit, workflow ID, and environment. Start in log-only mode for a week to see who is missing them, then reject untagged calls in non-production first. Most of the work is telling the teams that call the model through shared libraries, so update the library once and the tags follow.
03 Compute cost per accepted outcome for one workflow
Pick the workflow with the highest spend. Add model calls, retrieval, guardrails, evaluation runs, and reviewer minutes, then divide by accepted outcomes for the same period. Compare the result with the cost of producing the same outcome manually. If the AI path is not cheaper or faster, you have learned that before scaling it, which is the cheaper time to learn it.
04 Set one ceiling and one alert
For that same workflow, set a monthly ceiling at about 120 percent of forecast, add a step limit to the agent loop, and route an alert to a named owner. Review the list of idle endpoints every Monday and shut down anything without an owner. One working ceiling will teach you more about your gaps than another round of planning.
Let 10decoders Trace Your GCC's AI Spend to the Outcomes It Buys
We review one production workflow end to end: tagging at the gateway, metering of agent steps, idle resource policy, outcome definitions, showback reporting, and ceilings. You leave with a ranked leak list and a chargeback design your finance and business-unit leads can adopt.
