Determining the ROI of AI requires data that most companies lack

The widespread push by corporate leadership to scale artificial intelligence initiatives is encountering a significant financial accountability hurdle. With budgets for AI tripling and adoption rates soaring across industries, the critical question posed by CFOs and boards is no longer if AI investments are being made, but which of these initiatives are actually proving profitable. This fundamental question, however, remains largely unanswerable for most organizations, not due to a lack of cost visibility, but because the data provided by AI vendors is fundamentally misaligned with the needs of business analysis.
The lessons learned from managing cloud computing expenditures, which once presented a similar complexity in tracking return on investment (ROI), offer a partial blueprint but fall short when applied to the unique challenges of AI. Cloud providers, like Amazon Web Services (AWS), offer a granular level of detail in their billing statements. This includes resource IDs, account hierarchies, regional data, Stock Keeping Unit (SKU) information, tag metadata, and even usage by the minute. This richness allows mature FinOps (Financial Operations) teams to attribute every dollar spent to specific workloads, teams, or customer segments, provided proper tagging practices are in place. This granular data, when merged with business context such as customer and product mappings, enables a clear focus on cloud spend ROI.
However, artificial intelligence introduces a new layer of complexity that necessitates a more comprehensive data approach. Beyond cost and business data, AI’s ROI calculation critically depends on telemetry – the automatic collection of data from disparate sources that illuminates the "what" and "why" of system behavior. While executives and engineering leads may possess AI invoices detailing token consumption and understand customer revenue, the crucial link between the two is missing. The token count on an AI provider’s invoice, for instance, does not specify which customer initiated a particular call, which feature it supported, or whether the generated output delivered tangible business value. This level of detail simply does not reside within the billing systems of AI providers.
AI Providers’ Limited Scope: Billing vs. Business Value
The inherent limitation stems from the core business model of AI providers. They are in the business of selling computational resources, primarily measured in "tokens," not in the business of attributing an enterprise’s AI costs to its specific customers or internal product lines. Consequently, the granularity they expose in their billing is dictated by their internal accounting needs, not the analytical requirements of a Chief Financial Officer.
Consider the stark contrast between the data provided by a cloud provider and an AI service. An AWS invoice provides a wealth of information allowing for meticulous cost allocation. In contrast, an AI provider’s invoice typically offers token consumption per model, with an optional grouping by API key. This is the extent of the resolution offered. There is no request-level attribution, no direct customer identification, no feature mapping, and no insight into the outcome of a prompt. Even complex, multi-step agentic workflows, where a single user request can trigger a cascade of AI interactions, are collapsed into a single token count. This lack of detail creates significant blind spots. For example, a large financial institution receiving a multi-million dollar monthly AI invoice may have no clear visibility into which specific business units or functions are responsible for particular cost components, rendering accurate internal cost allocation impossible.
Therefore, if an enterprise seeks to understand the precise AI costs associated with specific customers or features, it must proactively capture this data internally, within its applications, before the AI calls are even made to the provider. This requires a deliberate architectural decision to instrument the application layer for detailed event tracking.
The Three Pillars of AI ROI Measurement
To effectively measure the ROI of AI investments, organizations must integrate three distinct data sources into a unified analytical model:
- Cost Data: This encompasses the raw billing information from AI providers, detailing token consumption, model usage, and API calls. While essential, this data is insufficient on its own.
- Business Context Data: This includes information about customers, products, features, revenue streams, and market segments. It provides the framework for understanding the commercial impact of AI applications.
- Telemetry Data: This is the detailed, real-time operational data generated by applications and AI workflows. It captures the specifics of each AI interaction, including which user, which feature, which specific prompt, any retries, and the ultimate outcome of the interaction.
When these three data sources are meticulously stitched together and modeled, they unlock the unit economics that are now indispensable for informed AI investment decisions. This allows for calculations such as cost per customer interaction, margin per feature, profitability per agent workflow, and ROI per model choice. Crucially, none of these vital metrics can be derived from billing data alone, nor from telemetry in isolation. They necessitate the convergence of all three sources, modeled in a way that directly maps AI costs to measurable business outcomes.
The Urgency Amplified by Agentic AI
The challenges in tracking AI costs are amplified significantly by the rise of agentic AI. While single-call inference, where one request maps to one cost and one outcome, presents a relatively straightforward scenario, agentic workflows introduce a new dimension of complexity. An AI agent is designed to decompose a task into multiple discrete steps, with each step potentially invoking a different AI model. These workflows can involve fallback mechanisms to alternative models when initial attempts fail, retries triggered by suboptimal results, and integrations with external tools that themselves incur costs. Consequently, a single user request can result in dozens of inference calls across multiple AI providers, with costs compounding in ways that are opaque to the provider’s invoice.
If application telemetry does not capture the granularity of each agent step, it becomes impossible to discern which specific actions within a workflow are profitable and which are not. Aggregate costs may only surface weeks later on the provider invoice, by which time the workflow could have been running at scale, customers onboarded, and unprofitable operational paths retried thousands of times. As agents become more sophisticated and autonomous, the volume of cost-generating events without attached business context can increase by an order of magnitude. The window for implementing robust instrumentation and attribution before these complex workflows become unmanageable is rapidly closing.
Transforming AI Investment Conversations
The integration of cost, business context, and telemetry data fundamentally alters the conversation around AI investments. Previously equivalent AI capabilities, when evaluated through the lens of unit economics, can reveal dramatic cost divergences – sometimes by a factor of ten – despite similar adoption metrics. This visibility empowers teams to select the approach that delivers comparable business outcomes at a fraction of the cost.
Product teams can design features with an intrinsic understanding of their potential profit margins from the initial architectural phase, rather than relying on post-launch budget reviews. Engineering teams can make informed decisions about model architectures, balancing latency and quality with actual cost-per-outcome data. Leadership can evaluate AI initiatives with the same rigor applied to any other capital allocation decision, based on demonstrable unit economics rather than solely on engagement charts. Aggregated invoices can track the cost per customer interaction, engagement metrics can reveal margin per feature, and the intuitive selection of AI models can be validated against empirical cost-per-outcome data.
Within seconds, stakeholders can identify which AI features are driving profitability, which warrant scaling, and which should be discontinued. This level of insight is precisely what organizations are seeking to optimize the benefits derived from their AI investments.
Navigating the "Build Trap" for AI Cost Management
The compounding nature of AI costs, coupled with the rapid pace of innovation, presents a significant challenge. Boards are unlikely to wait 18 months for an internal project to develop a sophisticated AI cost attribution system to reach production. The temptation to "build it ourselves" can be strong, especially with the proliferation of AI coding tools that empower small engineering teams to ship substantial functionality rapidly. The instrumentation layer might appear tractable, cost normalization a manageable weekend project, and the semantic model seemingly draftable within a sprint.
However, this approach often leads to what can be termed the "build trap," for several critical reasons.
Firstly, the sheer volume of data generated by a production AI footprint is immense. Millions of telemetry events per hour are not uncommon, and this volume escalates dramatically with agentic adoption. Real-time ingestion, correlation, and attribution at this scale represent a fundamentally different and more demanding problem than developing a prototype for afternoon experimentation. It requires a permanent, robust operational system that must function flawlessly, minute by minute, day by day.
Secondly, the vendor landscape is a dynamic and complex ecosystem. Cost data arrives from providers in delayed billing windows, often with non-interoperable schemas that can change without notice. Furthermore, new AI providers emerge monthly, each introducing its own unique taxonomy and metering methods. This means that any internal system built to manage AI costs is not a static solution; it requires continuous maintenance against a moving target that evolves faster than most internal release cycles.
The third critical factor is the cumulative impact of the first two: this is business-critical infrastructure. The financial health and strategic direction of an enterprise will increasingly rely on the data produced by such a system for capital allocation decisions. When schema drift goes unnoticed for weeks, or when an agent telemetry stream ceases to correlate with a vendor that has quietly altered its billing API, the cost of errors is not a minor cleanup effort. It can translate into months of misallocated capital, hindering strategic growth and potentially leading to significant financial repercussions.
The build-versus-buy decision for engineering leaders has thus evolved. The question is no longer simply "Can we build this?" The honest answer is often yes. The more pertinent question is whether the marginal hour of a highly skilled engineer is best spent stitching together disparate cost data, telemetry streams, and business outcomes, or focused on developing the AI products that generate the revenue this cost data is intended to measure.
The capability to accurately track and attribute AI costs is reproducible and can be achieved within weeks. The strategic choice lies in whether an organization dedicates the next 18 months to building this foundational infrastructure, or to leveraging its insights to drive impactful AI product development and business growth. The former risks obsolescence and operational overhead, while the latter promises optimized AI investments and a competitive edge.







