Blueprints 7 min read

Blueprint: The Network Value of AI Tokens

Tokens are becoming the currency of AI. Operators can’t yet measure the ones crossing their networks.

By Stephen Douglas, Head of Market Strategy, Spirent (a Keysight company)

A few months ago, I’d have told you AI was just another application riding across the network. Heavy, sure. Bursty, unpredictable, demanding. But still an app, like video or gaming before it, something the network carries rather than something the network is built around.

I’ve changed my mind. AI is becoming primary traffic. It’s starting to dictate how networks behave instead of just consuming what they offer. And if that’s right, it changes the question operators should be asking. The question isn’t “can my network handle AI traffic.” It’s “can I measure and charge for what that traffic is actually worth.”

The unit that answers both is the token.

Why tokens, and why now

Tokens are the model-specific units into which AI systems break input and output text, code, images, audio, or other modalities for processing and billing. For many generative AI services, especially LLM APIs, input and output tokens have become a core pricing unit. With consumption-based billing now becoming standard practice, you are charged based on both input and output tokens consumed.

The token therefore represents the connection point between the amount of raw computing power available to your AI model, and what your customers actually receive.

And this is when things start to get difficult for the operators. While networks forward packets across their infrastructure, the outputs that correspond to model tokens are serialized into application-layer payloads that are ultimately carried across packets, but tokens do not map cleanly to packets. Routers can observe flows, timing, and loss. They generally cannot observe application-level time-to-first-token (TTFT) unless that signal is exported by the application, model gateway, or synthetic test agent. The difference between what the network is measuring and what the customer believes has value is the core issue. It’s also the core opportunity.

If telcos want a share of the AI economy, they have to speak its language, and that language is tokens. Tokens don’t replace packets. They travel inside them. But they add a new layer that defines value, and it’s the operators who can measure that layer who get to charge against it.

A fair amount of skepticism, and some of it lands

I’ll be honest that not everyone buys this. When the idea of “AI token plans” started circulating, plenty of smart people pushed back, and they weren’t wrong to. Some leading analysts have argued operators should sell complete services rather than meter intelligence by the unit, comparing piecemeal token billing to trying to sell content by the pixel. Others have pointed out that telcos keep getting their lunch eaten by cloud providers every time there’s a new monetization play, from APIs to edge compute. Operators have chased mobile edge computing for years, but adoption and monetization have developed more slowly than early expectations.

That history is real, and I’d rather acknowledge it than pretend it away. The lesson I take from it isn’t “don’t bother.” It’s that operators lost those earlier rounds partly because they couldn’t prove what their network was delivering. They had a product and no way to measure its value, so the value accrued to whoever could. Token-level visibility is how you avoid running that play again. You can’t charge a premium for AI service quality if you can’t demonstrate it.

What you’d actually have to measure

The metrics we’ve leaned on for decades (latency, jitter, packet loss, throughput, availability) still matter. They’re just no longer enough on their own for interactive AI. Several ecosystem players are already asking us what the new measurements should look like, and a handful keep coming up.

Time to first token is the big one. How long after a request does the user or device receive the first meaningful chunk of AI output, and how much of that delay is attributable to network path versus model-serving stack? For a chatbot that’s a nice-to-have. For robotics, industrial automation, or assisted autonomy, first-response latency can determine whether an AI-assisted function is useful in the control window. AI value tends to live in the immediate reaction, not the total completion time.

Then there is deterministic reliability, a composite metric providing a service-level view of whether latency variation, loss, and retransmission behavior remain within the application’s operating envelope. Real-time AI and closed loop systems need delivery to be predictable not best effort. Timing variation can degrade streaming interaction, disrupt synchronization between AI components, or push closed-loop systems outside their timing budget, and for safety/automation use cases that’s the difference between working and not.

Energy per token is the one I think gets overlooked. It connects infrastructure efficiency directly to delivered value. If inference traffic shifts closer to users and devices, operators could absorb more AI-related transport and edge-compute demand than they do today. If you cannot measure energy per token, then you can’t manage the economic underpinning of this whole thing.

There are others worth tracking: precise time and location sync for AI that fuses data from many sources, how efficiently compute turns into a real-world action depending on where it runs, and whether the network can re-route a workload when an edge node fails without breaking the session. Each one ties a network behavior to something a customer would actually pay for.

The hard part nobody likes to say out loud

Here’s the catch. Unless the operator owns or brokers the models, they cannot see tokens from the network alone. It resides within the application layer and model layer. AI inference traffic is often carried over application protocols such as HTTPS, gRPC over HTTP/2, or HTTP/3 over QUIC. In most production deployments, TLS or QUIC encryption prevents the network from inspecting payload contents. Data payloads are also encrypted so, while the network may infer metadata such as timing, sizes, directionality, and flow behavior, it generally cannot identify the specific token or semantic content inside encrypted payloads. Additionally, just because one packet may equate to one chunk, one chunk does not necessarily equate to one token. That is why “counting tokens on the wire” breaks down quickly. The practical path is not packet inspection, it is correlation between application-level token telemetry and network-level performance data.

Two things get you most of the way there.

The first is cooperative telemetry. In essence, it treats each token as if it were an application-level event; have your model stack export the timestamp, count and request ID for each token; then correlate these values against network flow and path data. You’re not extracting token information from the wire. You’re correlating application truth with network truth and providing accurate AI KPIs based on the correlation.

The other is active testing with synthetic AI traffic. Here’s where I should admit my bias, since this is the work we do, so weigh it accordingly. Instead of trying to infer every production token through encryption, you generate test sessions whose request patterns and streaming bursts are built to behave like real AI inference. That makes the workload repeatable and measurable. You can run it continuously, across locations, before and after a routing change, and use it to prove what a given route or edge location is capable of delivering. It gives you a baseline: what the network should do under a controlled token-like load.

Passive monitoring tells you what real traffic is doing. Active testing tells you what the network can do under control. Run them together and you can tell whether a slow AI response is the network’s fault or the application’s, which is exactly the argument operators and cloud teams keep having with no evidence to settle it.

Where this goes

I’m not predicting telcos win the AI economy. What I’d say is, you can’t monetize what you can’t measure. The operators who build token-level visibility now, the KPIs, the active and passive testing, the correlation between the two, are the ones who’ll have something to sell when the rest of the market is still describing AI quality in adjectives.

Operators need to decide where they want to play. Operators that control or integrate into the AI service layer can monetize tokens more directly, while operators that only control the network can still establish monetization models from assurance, connectivity, performance guarantees, and infrastructure quality.

The token is becoming the currency of AI. Whether it becomes the currency of the network too depends less on the technology and more on whether operators decide to measure what they’re carrying.

Share this article

Help others discover this reporting.

Explore More

On this day

July 27

From 25 years of the Converge Digest archive.