The Unit of Return
In June 1875, delegates from twenty countries gathered in St. Petersburg to settle the rules of international telegraphy. Among the questions occupying them was a surprisingly difficult one: what counted as a word?
Telegraph tariffs depended on word counts, but languages did not divide information evenly. German could combine several words into one. Businesses used ciphers that compressed complete instructions into a few characters. The delegates responded with limits: no more than fifteen letters to a word in Europe, ten outside it, five characters of cipher counted as one word.
The rules gave telegraph administrations something consistent to charge for, and customers learned to work around them. Businesses compiled codebooks in which a single word could represent a complete commercial instruction. The network counted one word. The recipient received a paragraph. The message and the charge had begun to separate.
AI companies use tokens for a similar purpose. A token may be a word, part of a word, punctuation, or a space, and two models can consume the same number of tokens and produce answers of very different quality. One finishes the task immediately. Another needs a second attempt, a check from another model, or a person to fix what it produced.
Tokens are useful for billing because running models has a real cost. They become less useful when carried into an ROI calculation as though they measure the thing a business receives.
Most AI business cases start with a cost that's easy to count: licenses, model calls, tokens. The return sits opposite it as productivity, better decisions, improved experience, transformation. Those may all be real. They're also too broad to sit across from a number that precise. The calculation puts an exact measure of consumption next to a vague ambition and expects the comparison to mean something.
As capable models get cheaper, the comparison gets harder. More models become viable for the same task, and the path through each one varies — in cost, in speed, in how often it needs a second pass. A model with cheaper tokens can cost more overall when it requires three attempts and a person to finish the job. Token prices describe every charge along the way without saying which path produced the better result.
A useful ROI calculation needs a unit both sides can use. The investment becomes the full cost of producing that unit. The return becomes the value created each time it's delivered.
Every AI proposal should be able to answer three questions.
What useful result will the system produce?
Consider a system handling customer requests. "Answering questions" is too broad to measure. A resolved request has observable conditions: the correct policy was applied, the requested action was completed, and the customer did not have to return the next day. The result could just as easily be an invoice processed, a code change accepted, or a report delivered in time to inform a decision.
The result needs a connection to value. A resolved request may reduce repeat contacts or free up the service team. An accepted code change may shorten delivery time. Naming what changes gives the return somewhere to appear.
That value needs a baseline. If the team already resolves the same request quickly and cheaply, automating it creates less return than the completed result alone suggests. The relevant change might be capacity created, cost avoided, revenue gained or risk reduced.
What evidence will show the result is good enough?
An answer isn't evidence that a request was resolved. The answer has to match the policy, the action has to happen, and sometimes the customer's experience has to be checked afterward. That's what turns an output into an accepted result.
The measure has to stay honest. A support system rewarded only for avoiding repeat contact can learn to end conversations fast while leaving people confused — so the evidence has to include the thing you were actually trying to protect, not just the metric standing in for it.
Some evidence arrives fast. A payment can be matched to an invoice, and code can be checked against tests. A recommendation to enter a new market depends on competitive responses that may take years to play out. How fast and how cheaply the evidence arrives determines how confidently the result can be counted at all.
What will the full path to that result cost?
The model call is one part of the answer. A system may also search, use other software, check its own work, retry or hand off to a person. Building, integration and process change create costs before the first result arrives; maintenance adds more over time. Some costs occur once and others on every request, so the full calculation needs an expected volume over a defined period.
Answer all three, and the two sides of ROI can be expressed in the same unit: value per accepted result and full cost per accepted result. Assumptions remain, but they are now assumptions about something the business can observe.
The common unit also lets you compare across providers. A model that applies the right policy, completes the action, and satisfies the definition of resolution delivers the same unit of work as any other model that does the same thing. The business can move between them and keep getting resolved requests. The work and the machinery that produced it begin to separate.
An evaluator makes that possible. It checks the result and routes failed, unusual or high-stakes work down a different path.
Developers will still count tokens to manage cost, monitor response times and catch runaway loops. Providers may continue charging by usage because different requests require different amounts of work. Businesses can still compare providers by cost per accepted result.
A proposal that can’t yet answer the three questions may still deserve funding as an experiment. The next investment should identify the useful unit, the evidence that proves it and the real cost of producing it, discovering whether an ROI case exists.
The delegates in St. Petersburg needed to know how many words crossed the wire. The sender needed the instruction to be understood and acted upon. Each measure answered a different question.
Leaders evaluating an AI proposal need the same separation. What useful result will the system produce? What evidence will show it's good enough? What will the full path to that result actually cost?
Together, those three answers define the unit of return.