Radar · 17/07/2026 · happened on 14/07/2026 · business

OpenAI releases 'AI scorecard' to measure business ROI: work completed, cost per task, reliability

Sarah Friar, OpenAI’s CFO, has published an operational framework to measure the return on AI investments in business. Not generic benchmarks or claimed productivity percentages: four concrete metrics that apply to real workflows.

The four metrics: useful work completed (end-to-end tasks delivered), cost per successful task (dollars divided by verified outputs), reliability (percentage of outputs passing quality control), and return on compute (value produced per dollar of inference). Each metric is defined with examples from data science teams and sales teams, documented cases with verifiable workflows.

Why it matters to you. If you’re using or planning to use AI in a business workflow, eventually someone will ask: how much does it cost, how much does it produce, how reliable is it. This framework gives you a way to answer with numbers that hold up to audit, not impressions. The logic applies beyond enterprise too: if you’ve automated something with an agent, you can count how many tasks it completed without intervention, what it cost you, and how many outputs you had to discard.

Where to look. OpenAI has documented the framework’s application on two use cases (data science and sales) in an official post, with step-by-step workflows. If you’re evaluating an AI investment or need to justify an existing one, the material is ready and understandable even for non-developers.

In detail

The context.

The problem is familiar to anyone who’s tried to get an AI budget approved at a company: public benchmarks say nothing about your workflows, vendor productivity claims are opaque, and the controller wants numbers they can verify. Until now, many companies measured AI with generic proxies (time saved, license costs) or didn’t measure it at all, accepting the risk of wasting resources on tools that seem useful but produce no measurable value.

Sarah Friar chose to tackle the issue from inside: OpenAI sells AI products, but also needs to justify its own internal investments. The framework comes from that need and applies to any company using AI to automate or enhance operational workflows.

The four metrics, explained.

Useful work completed counts end-to-end tasks delivered without significant human intervention. Not tokens generated, not sessions opened: how many reply emails were sent, how many reports were produced, how many analyses were delivered. It’s the output metric.

Cost per successful task divides total cost (inference, tools, infrastructure) by the number of tasks passing quality control. A failed or discarded task costs the same in compute as a successful one, but produces no value: this metric makes it visible. If the cost per successful task is too high relative to the task’s value, the workflow doesn’t make sense.

Reliability measures the percentage of outputs passing quality control on the first attempt. An agent completing 100 tasks but failing 30 has 70% reliability. Below a certain threshold (depending on the use case), manual correction time erases the automation gain.

Return on compute relates the value produced (measured in dollars saved, revenue generated, or other business metric) to the cost of inference. It’s the metric that ties AI to the bottom line: one dollar spent on compute must produce more than one dollar in value, or the workflow doesn’t scale.

What OpenAI says and doesn’t say.

The post documents the framework’s application on two concrete cases: a data science team using ChatGPT Work to explore datasets and a sales team using it to prepare calls and follow-ups. For each case, OpenAI lists the tasks, shows how to count them, and measures the four metrics. The workflows are verifiable: people working in those roles recognize the tasks and can replicate the measurement on their own workflows.

What the post doesn’t provide are absolute numbers (how many tasks, what reliability, what return on compute) for the two documented cases. That’s not a limitation: those numbers depend on business context and publishing them would only fuel misleading comparisons. The value lies in the method, not the benchmarks.

Practical implications.

For those building or managing AI workflows in business, the framework offers a shared language with finance and operations. Instead of arguing abstractly about AI’s usefulness, you can show costs, outputs, and reliability measured on your workflow. If return on compute is negative, you have the numbers to decide whether to optimize (cheaper model, better prompt, tighter quality control) or abandon the effort.

For those evaluating an AI investment, the four metrics are questions to ask vendors: what’s the cost per task for your tool on my use case? What reliability can I expect? How do I measure useful work completed? A vendor who can’t answer with verifiable numbers is selling hype, not a tool.

The framework works outside enterprise too. If you’ve automated a personal workflow with an agent (email triage, daily digest, document drafts), you can count how many tasks it completed in a week, what it cost, and how many outputs you had to discard or rewrite. Those three pieces of information tell you whether the workflow pays for itself or whether you’re spending time and money delegating something you’d do faster by hand.

The limits.

The framework measures existing workflows, not explores new possibilities. If you don’t yet know what to automate, the four metrics won’t help you choose: you need experiments, not measurements. And the framework assumes you already have quality control in place: without it, reliability and cost per successful task can’t be calculated.

Finally, return on compute requires translating produced value into dollars. For some workflows (sales, customer support) it’s straightforward. For others (research, internal analysis) you need advance agreement on what counts as value, or the metric becomes arbitrary.

Type to search across course, playbooks, skills, papers…