Multidimensional value metrics
On the Monday after launch, Distribuidora Norte’s new dashboard filled the big screen in the meeting room. It showed sales, deliveries, complaints and customers at risk of churning. The data refreshed on its own and the charts matched the approved design. Users, queries and processed alerts could all be counted. At a glance, the project could already be declared a success.
Then the manager asked a less comfortable question: “Are we deciding better?” No one answered right away. The dashboard said how much it had been used, not what had changed thanks to that use. It did not say whether the supervisor got to a complaint sooner, or whether salespeople were preventing churn that would otherwise have happened. Nor did it say whether the drivers felt the tracking helped them organize the route or simply watched them more closely.
That silence is the distance between activity and value. An organization can put models into production, buy licenses, log thousands of conversations with an assistant and improve an algorithm’s accuracy without any relevant decision getting better. It can also obtain an economic benefit and, at the same time, erode autonomy, concentrate knowledge or exclude those who cannot use the new channel. If only the easy things are measured, what gets left out is the part of the change that makes the technology sustainable, or intolerable.
IMIA asks whether the organization has the conditions to move forward. This chapter asks a later and harder question: was the intervention worth it, and for whom? The answer cannot be improvised at closure. It has to start before building, while it is still possible to agree on what change would justify the effort.
The appeal of available figures
Every system produces numbers about itself: response time, active sessions, documents processed, predictions issued, recommendations accepted. They are useful for knowing whether the solution works and whether anyone uses it. The trouble starts when they are presented as proof of value.
A high usage rate may mean the tool helps, or that management made mandatory a step that did not exist before. Many queries to a chatbot may show interest, or show that the answers do not solve the problem and each person has to ask several times. Accuracy of 92% can be excellent in one context and disastrous in another, depending on where the remaining 8% is concentrated and what each error costs.
What this book calls vanity metrics are numbers that describe activity and are used to imply a result they did not measure. They are not necessarily false. They are incomplete and, for that reason, easy to bend toward the story someone wants to tell.
The pressure to show success comes early. A project has sponsors, a budget and a promise to defend. If no one agreed on a baseline, in the end whichever indicator came out best gets chosen. Perhaps the time for a task went down, though rework went up. Perhaps more cases were handled, though resolution got worse. Perhaps employees deliver sooner, but keep working outside the system to do it.
Measuring value requires resisting that retrospective selection. It does not eliminate interpretation or produce an automatic truth, but it forces the questions to be committed before the answers are known.
The distance between deploying and transforming
During 2025 and 2026, evidence at scale appeared on the distance between what is invested in AI and the value obtained. The preliminary MIT report that popularized the term GenAI Divide found that the vast majority of the organizations studied got no return from their generative initiatives, despite the size of the investment. Because it is preliminary, a striking figure should not be turned into a general law. Its explanation is more fertile than the percentage: the obstacle lay not mainly in the models but in an organizational learning gap. The solutions did not retain enough context, nor were they integrated into operations in a way that let them learn from it.
The observation matches the thesis of the bridge. Buying access to a technical capability is not the same as incorporating it into the work. Between the two lie definitions, institutional memory, trust, owners and mechanisms for correction. If that fabric is missing, a more powerful model can improve a demo without changing the organization’s results.
There is also evidence of concrete benefits. Brynjolfsson, Li and Raymond studied the use of generative AI in a support center and published their results in The Quarterly Journal of Economics in 2025. They observed an average 15% increase in cases resolved per hour, with a larger improvement among less experienced workers. How the benefit is distributed matters as much as its size: the tool seemed to bring newcomers part of the knowledge that had been concentrated in the experts.
That finding does not license promising 15% in any activity. It happened in a specific context, task and organization. What it does show is why value has to follow people: the same technology benefits people differently depending on their prior experience, and the average hides that distribution.
McKinsey’s global survey from late 2025 offered another contrast. About 6% of respondents were classified as high performers: organizations that attributed at least 5% of their operating profit (EBIT) to AI. Some 39% reported some impact on that result, a much lower bar. Like any industry survey, it calls for caution. Even so, it helps frame the right question: what distinguishes visible adoption from a sustained capacity to capture value?
International evidence offers mechanisms and orders of magnitude, but it does not replace local measurement. A small business in Mendoza does not work under the same conditions as the support center of a large company. The bridge uses these studies to avoid two opposite errors: denying that there are measurable benefits, or assuming they transfer without looking at the context.
Before measuring an improvement
The baseline is a description of the state before the intervention. It seems an obvious requirement, yet many projects skip it because the problem is presented as urgent or because there is no historical data. Once the change is in place, how work was done before can no longer be reconstructed with precision. What remains is memory, and memory tends to adjust itself to the result.
A baseline does not need to be sophisticated. It can combine times recorded over a few weeks, observed errors, interviews, samples of decisions and a short survey. What matters is that it measures what the intervention says it will change.
If the goal is to respond to complaints sooner, one records when a complaint is considered open, when it is considered resolved and how many cases are reopened. If the aim is better commercial decisions, counting meetings is not enough: one has to observe how long a decision takes, what evidence it uses and how often it has to be corrected. If the promise is to free up repetitive work, measuring the hours of the original task is not enough. The new time spent reviewing results, correcting exceptions and maintaining the system also counts.
“We did not measure that” is a valid answer. It describes a real limitation and can become the project’s first improvement: designing an observation period before committing to a percentage. Inventing a baseline after the fact yields a more comfortable figure and a worse decision.
One also has to decide what variation would be significant. Five fewer minutes can matter in a task that happens thousands of times and be trivial in one that happens once a month. The threshold depends on frequency, cost, risk and the experience of the people involved.
Governing meaning
Before building a dashboard, the organization needs to agree on what its numbers mean. Several areas may use the word “sale” for different things: one records the order, another the invoice and another the payment. No definition is absurd in itself. The problem appears when they are combined without declaring the difference.
Metric anarchy shows itself when two dashboards answer the same question with incompatible results and each area defends its own. An AI project does not resolve that conflict; it can, instead, hide it behind a more complex layer.
Every relevant indicator needs an accessible definition, an identified source and a business owner. The owner is not whoever codes the calculation. It is whoever can explain why the definition represents the phenomenon, accept its limits and authorize a change when the process changes.
The semantic layer, a place where shared definitions and rules are kept, prevents each dashboard from recalculating core concepts on its own. But architecture alone does not govern: a review ritual is needed. An indicator may once have been useful and then lost its connection to a decision. If it stays on screen out of inertia, it becomes a zombie KPI, consuming attention without guiding any action.
A simple question usually brings order to the design: what Monday decision changes if this number goes up or down? If no one can answer, the chart may be interesting, but it is not yet a management metric. If there is an answer, it must be clear who decides, within what time frame and what other information they need.
Quantity matters too. An executive dashboard with dozens of indicators spreads attention until it becomes irrelevant. A small, reliable set is preferable, with detail views for when a signal calls for investigation. Measuring more is not the same as understanding better.
Five dimensions of value
Economic value is indispensable, especially in organizations with scarce resources. An intervention must be able to explain what time it frees, what errors it avoids, what cost it reduces or what revenue it makes possible. But the economic result does not exhaust the transformation. Sometimes it arrives late; other times it improves at the cost of a fragility that ends up destroying it.
That is why the method looks at five dimensions of value: economic, decisional, human, organizational and social. They do not form a universal index, nor are they added up as if one improvement could offset any harm. They are perspectives that force attention onto different consequences.
Economic value
It covers time, cost, revenue and use of resources. Measuring it well means avoiding automatic equivalences. Ten hours “saved” do not necessarily become ten billable hours. Perhaps they are redistributed toward backlogged tasks, reduce overtime or allow better service. That result still has value, but it has to be named precisely.
The total cost also has to be followed. An automation may reduce manual workload and, at the same time, increase licenses, supervision, infrastructure consumption and vendor dependency. The net benefit appears after adding maintenance and transition, not in the comparison between the old task and the demo of the new system.
Decisional value
It asks whether decisions are more timely, consistent, informed and traceable. At Distribuidora Norte it was not enough to know how many alerts had been opened. What mattered was how much time passed from the first sign of a problem to a conversation with the customer, what information the salesperson used and whether the reason for accepting or rejecting the recommendation remained available to learn from.
A faster decision is not always a better one. If speed removes review in a high-risk case, quality can get worse. That is why decisional value looks at timeliness and correctness together, and distinguishes reversible situations from decisions that affect rights or produce costs that are hard to repair.
Human value
It follows what changes in the experience and capability of those who work with the system: autonomy, workload, learning, trust and well-being. Some signals are quantitative, such as overtime, turnover or the frequency of corrections; others require interviews or observation.
This dimension prevents declaring success when productivity rises because work was intensified. It also makes it possible to recognize benefits the economic balance is slow to show: fewer interruptions, more clarity about priorities, the chance to learn a new task or the end of a repetitive emotional burden.
Perception is no minor datum. If the team feels that a tool meant to help them is watching them, that interpretation will change how they use it and, over time, the economic result. The human is not a soft externality: it is part of the mechanism that produces value.
Organizational value
An intervention can leave capabilities that go beyond its initial case: better-defined data, new owners, review practices, integration between areas and experience for evaluating future technology. It can also do the opposite and concentrate knowledge in a vendor or in a single person.
IMIA makes it possible to observe changes in that configuration. An improvement in culture and capabilities, in data quality or in governance can be a result in itself, even if the first pilot delivers a modest economic benefit. Even so, “we gained maturity” should not become an automatic consolation. The new capability needs evidence: a policy that is applied, a process that no longer depends on a personal spreadsheet, a second lead who can take over.
Social value
In the public sector this dimension is usually mandatory. It measures access, equity, the possibility of appeal, quality of service and how benefits and errors are distributed. A digital procedure can reduce costs and, at the same time, exclude those with poor connectivity. A model can improve the average and concentrate false positives on a specific group. The aggregate result does not describe that distribution.
It also applies to private organizations when a decision affects opportunities, prices or conditions of access. In sustainability-related projects, the social dimension has to engage with the environmental cost of the infrastructure. Promises of “intelligence” or “green efficiency” require measuring energy, inclusion and material effects, not just changing the language of the report.
Choosing without reducing
Not every project can measure all five dimensions with the same depth. Trying to can produce a measurement system so costly that no one sustains it. At the start, two or three priority axes are chosen according to the sociotechnical diagnosis, and the others are reviewed to detect serious harm.
In a low-risk administrative automation, economic and human value may be the center: how much time is freed and how the workload changes. In a system that prioritizes public inspections, decisional and social value are inseparable. In an organization that is only now putting its data in order, organizational value may come before economic return.
Prioritizing is not forgetting. It is declaring what will be measured systematically and what will be watched as a guardrail. An improvement on two axes does not license ignoring severe harm on another: no number of hours saved makes an uncorrected discrimination acceptable.
A protocol that travels with the project
Measurement begins during the diagnosis. That is where the baseline is agreed, the dimensions are chosen and the assumptions are recorded. Before building, objectives and thresholds are defined: what change would be valuable, how long it must hold and what consequence would require stopping the pilot.
During implementation, early indicators are followed, not to declare victory but to correct. If usage rises and so do manual workarounds, something needs investigating. If time goes down but reopenings grow, speed is hiding rework.
Closure compares against the baseline and separates three kinds of result: what can be attributed to the intervention with reasonable confidence, what improved but has multiple causes and what could not be measured. That distinction protects against the temptation to credit the project with any favorable change that happened in the same period.
Then sociotechnical hygiene comes in. Some benefits fade when the accompaniment ends; others appear when the team learns to use the tool with more judgment. A review at three or six months says more about adoption than launch week.
Back to the meeting room
At Distribuidora Norte, the final dashboard did not need to start with fifty indicators. It could start with two well-measured questions. The first: does the supervisor act sooner on a relevant complaint? The second: do the drivers feel the new flow helps them resolve the route, or that it only adds surveillance?
The first required agreeing on when a complaint began, when there was an action and how many cases were reopened. The second could not be inferred from usage. It called for listening to the team before and after, observing the workarounds and recording whether corrections made outside the system increased.
Savings in pesos could come later; without a reliable baseline, promising them would have been mere narrative. The improvement in data quality and governance, by contrast, could be observed in the corresponding IMIA dimensions, with the honesty of making clear that the instrument describes configuration and still awaits empirical calibration.
Measuring this way does not produce a perfect figure. It produces something more useful: a debatable and verifiable explanation of what changed, for whom and at what cost. With it, one can recognize a partial success, stop an intervention that does harm or learn from a result different from the one expected.
Value does not live in the model, because the model does not wait for an answer to a complaint, does not make a hard decision, does not learn a trade and does not lose access to a right. All of that happens to people and organizations. Metrics gain meaning when they follow those consequences and refuse to confuse the glow of activity with the transformation of work.
See also: Concepts of our own (value metric) · IMIA — the maturity instrument · The bridge applied to SMEs · Sociotechnical hygiene