Skip to main content

BLOG | Tokenomics: The AI cost challenge telcos can't afford to ignore

AT&T is running 40 billion tokens on an average day. Six months ago, it was 5 billion. Mark Austin of AT&T sat down with TM Forum's Guy Lupo at DTW Ignite 2026 to talk about what that scale demands — and what the industry needs to do about it.

July 30, 2026

There's a pin on Mark Austin's jacket. It says: "No token cost."

It's a conversation starter. It's also a strategic goal.

Tokenomics is the theme of the industry right now — and at AT&T, with over 40 billion tokens flowing through its systems on an average day, getting a handle on how those tokens are being spent is critical. As recently as September 2025, that number stood at 5 billion. The jump to 40 billion happened in roughly six months — and that explosive growth is exactly why AT&T is so focused on it.

At that rate, tokenomics isn't a future problem. It's a live operational challenge, right now.

More than a gateway

Many operators have an AI gateway. That's the starting point, not the solution.

A gateway gives you routing capability: the ability to direct tokens to the right model, on the right cloud, at the right cost. That matters. When the same model is available across multiple cloud providers, you can route to the cheapest. When one cloud is congested, you can route around it. And beyond the major frontier models, open-source alternatives are worth serious consideration. Plus, they're only around six to ten months behind on capability. AT&T looked at its model inventory and found a significant share were more than a year old, meaning open-source equivalents already existed. Migrating those is a straightforward cost play. Specialized or fine-tuned small language models go further still, capable of cutting costs by as much as 90%. On-premises routing adds another lever for operators with the infrastructure to support it.

So routing is powerful. But routing alone doesn't give you control when control is what tokenomics actually requires.

That's where the gateway needs to become something bigger. On top of routing, you need FinOps: the financial discipline to measure token spend at a use-case level. How much is each use case costing? Which teams or applications are running over budget? Can you set quotas? Can you automatically downgrade to a cheaper model when the task doesn't warrant the premium one? These are operational questions that require observability; the ability to go back through logs, understand what happened, and intervene before costs spiral.

TM Forum's Model as a Service (MODaaS) framework is built around exactly this. It covers the routing layer, but also the governance, FinOps, and observability capabilities that turn a gateway into a genuine AI cost control plane. As Mark puts it: it's more than a gateway — and operators who treat it as just a gateway will find themselves managing a tokenomics problem they can't see clearly enough to solve.

What the data actually showed

AT&T didn't theorize their way to these insights. They built and measured, working through TM Forum's MODaaS framework and a Catalyst project with Databricks and other partners.

When they started measuring and looking at the logs, they found that 85% of the time, developers were defaulting to the most expensive model, when they only actually needed it 8% to 10% of the time. AT&T graded 3,600 tasks and asked: can we get away with Sonnet instead of Opus? Sure enough, Sonnet handled 85 to 90% of them just as well.

"We probably would have never figured that out unless we did the project."

That's the value of building in code first. Real patterns only emerge in production.

There's one more wrinkle worth noting: cache awareness. When routing across models in a coding context, switching models breaks the GPU cache, meaning tokens that should cost a fraction of the usual price end up billed at full rate. Smart routing has to account for this, or it can inadvertently increase cost rather than reduce it.

AI that speaks telco

OTEL models are specialized telco models, built because frontier models like Claude and OpenAI simply don't speak telco — they don't understand the telco lingo. Asked telco-specific questions, they get the answer right just 60 to 70% of the time. AT&T trained specialized models for telco use cases and pushed accuracy to 90%, and at DTW Ignite, opened up access to a newly fine-tuned Google Gemma 4 model hitting 92%.

Now integrated into the MODaaS framework, these OTEL models give the industry a foundation for telco-native AI reasoning that's actually production-ready. They are built for root cause analysis, fault management, and complex telco queries that general-purpose models simply can't handle reliably.

From cost control to opportunity

Mastering tokenomics internally is table stakes. AT&T is using it themselves first, but the next move is offering that capability as a service to others. Some operators at DTW Ignite were already doing exactly that, turning token management infrastructure into a commercial offering. The gap between operators who move and those who don't is widening fast.

For Mark, the thread running through all of it is trust. As a member of TM Forum's Trustworthy AI and Data Mission Board, he sees sovereign AI and tokenomics as directly connected: controlling where your tokens go, which models process them, and under what conditions is a trust question as much as a cost question.

"You can't scale without trust."

Get the tokenomics right. Get the governance right. The rest follows.

Watch the full conversation https://inform.tmforum.org/videos/tokenomics-the-ai-cost-challenge-telcos-cant-afford-to-ignore