The Digital Transformation Playbook
Kieran Gilmurray is an Internationally acclaimed expert in leadership, AI, strategy and transformation.
He helps boards, executive teams and senior leaders make sense of complex technological change and turn it into practical business value.
Most experts make technology feel more complex. Kieran makes complex ideas simple, useful and actionable.
He has worked with leadership teams across the globe to help them understand AI, use data to make better decisions and apply technology in ways that improve performance.
The outcome is clearer thinking, stronger leadership confidence, better adoption and more measurable business benefit from technology.
Kieran and his team bring the practicality many thought leaders lack, the human clarity large consultancies often miss, and the strategic depth that goes beyond standard AI training.
If your organisation is trying to digitally transform and make AI useful, safe and commercially relevant, then connect.
📅 Book a call: https://calendly.com/kierangilmurray/catch-up
🌎 Website: www.KieranGilmurray.com
📘 Kieran Gilmurray | LinkedIn
🌐 Substack: https://kierangilmurray.substack.com
📕 Amazon https://tinyurl.com/MyBooksOnAmazonUK or Audible https://www.audible.com/search?keywords=kieran+gilmurray
Kieran
The Digital Transformation Playbook
Measuring What Actually Matters: The Value Layer of AI Scale
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
AI adoption is rising fast, yet many organisations still struggle to prove real business value. This episode examines why activity metrics can create confidence without showing whether AI is improving performance.
It explores the Value layer of AI scale.
TLDR / At a Glance
• Activity versus value
• Stronger AI measurement chains
• Output quality and workflow performance
• Business outcomes and economic impact
• Risk adjusted value metrics
• Workflow level evidence
The key takeaway is that AI becomes defensible when leaders can connect usage to measurable performance, financial impact, and controlled risk.
If you are leading your businesses strategic transformation and need greater clarity, stronger execution and measurable results, let’s connect.
🌎 Website: www.KieranGilmurray.com
📅 Book a call: https://calendly.com/kierangilmurray/catch-up
📘 Kieran Gilmurray | LinkedIn
🌐 Substack: https://kierangilmurray.substack.com
📕 Amazon https://tinyurl.com/MyBooksOnAmazonUK
AI Transparency Notice: This podcast uses a hybrid format. When an episode features one of Kieran Gilmurray’s written articles, the narration is generated using a synthetic clone of his voice via ElevenLabs AI (the underlying article text is entirely human-authored). Episode descriptions and summaries are assisted by AI and should be considered unedited by a human unless specified.
Why AI Value Is Hard
SPEAKER_00Measuring what actually matters the value layer of AI scale TLDR slash at a glance Most organizations can measure AI activity more easily than AI value. That creates the appearance of momentum without a dependable basis for prioritization or scale. Seats, usage, prompt volume, and anecdotal productivity are weak signals on their own. They say little about output quality, business outcomes, or financial impact. Serious AI measurement moves through a stronger chain, activity, output quality, workflow performance, business outcomes, and economic impact. Value usually appears at workflow and system level, not just at user or tool level. Faster tasks do not automatically produce better economics. Weak measurement weakens governance, capital allocation, and executive confidence because leaders cannot clearly decide what to stop, scale, or fund. The value layer is where AI moves from visible movement to evidence that boards, executives, and investors can defend. Many organizations can show AI activity, far fewer can show AI value. Dashboards can report licenses, active users, prompts, and assisted hours almost immediately, but those measures are weak proxies for output quality, workflow performance, business outcomes, or economic results. In the first seven articles in this series, we established why AI fails before it scales, defined the human AI operating system, and showed how ownership, workflow design, capability, system building, and governance shape outcomes. This article turns to the value layer and asks a harder question. Even when all those elements improve, how do we know value is being created? The argument is simple. AI activity is easy to count, real value is harder to prove. Until organizations measure what matters, AI will remain difficult to prioritize, govern, and scale with confidence.
Activity Metrics Mislead Leaders
SPEAKER_00Why activity is easier to count than value? One reason AI measurement remains shallow is that activity sits at the top of the funnel. It is easy to instrument because it lives in vendor dashboards, local usage logs, and rollout reports. Leaders can quickly see how many licenses have been issued, how many people are using the tool, how often they return, and whether prompts or assisted hours are rising. Value sits much lower in the logic chain. It depends on baselines, attribution, workflow redesign, ownership, time horizon, and financial interpretation. That makes it slower to capture and harder to simplify. Accenture's research found that AI adoption is rising, but many firms are still struggling to convert it into productivity and revenue gains. The gap is execution, AI is being used, but workflows and operating models are not changing fast enough around it. Other major studies reinforce the same pattern. AI adoption is widespread, while scaled financial impact remains uneven. PWC found that many organizations expect AI to drive growth and efficiency, but fewer are yet seeing clear financial return at enterprise level. Deloitte similarly found that AI investment is rising, while fast payback remains rare. The message is not that AI lacks value, it is that value proof lags activity by a wide margin. That distinction matters because visibility can distort judgment. If leaders confuse movement with impact, they can overfund noise, underinvest in redesign, and misread experimentation as evidence of durable progress.
The Measurement Chain That Matters
SPEAKER_00What the value layer actually measures. In the human AI operating system, the value layer is not a finance add-on. It is the part of the management architecture that determines whether AI is creating measurable performance strong enough to justify further investment, governance, and scale. That requires a more disciplined hierarchy of measurement. At the top set activity metrics such as licenses, active users, prompts, experiments, and assisted hours. These are useful in pilot mode because they show whether the tool is being touched at all. Below that sit output quality metrics such as accuracy, review rate, error rate, repeat work, and fit for purpose. Then come workflow metrics such as cycle time, throughput, first pass resolution, handoffs, and rework. After that come business outcome metrics such as conversion, retention, cost to serve, sales effectiveness, or repeat inquiry rate. Finally, at the bottom of the chain, cite economic impact metrics such as e-bit contribution, margin uplift, cash flow, payback, or return on invested capital. Agentic AI also changes what needs to be measured. If an AI system is taking actions across a workflow, leaders need to measure not only usage or time saved, but action quality, exception rates, human override rates, downstream rework, customer impact, and risk exposure. The more autonomy AI has, the more important it becomes to connect value measurement with control measurement. This is the shift that many organizations still have not made. They measure what is easy to count rather than what is important to know. A serious executive value system must move from visible activity to operational performance, then to business effect, then to economic consequence, with explicit risk and time horizon assumptions throughout. How weak
How Weak Metrics Break Decisions
SPEAKER_00measurement distorts decisions. Weak measurement does more than blur the ROI story. It distorts prioritization, governance, and capital discipline. When organizations rely too heavily on usage, they often reward visibility rather than value. A team with strong engagement can appear more successful than a team delivering quieter but more material workflow gains. That makes it harder to compare use cases properly, harder to stop weak initiatives, and harder to move resources into domains where AI is genuinely improving performance. There is also a risk of false value. A team may report time saved, but if that time is not redeployed productively, the economic benefit may be limited. A workflow may become faster, but if quality falls, rework rises, or risk increases, the apparent gain may disappear elsewhere in the system. This is why AI value needs to be measured across the workflow, not only at the point of tool use. Weak measurement also makes governance harder. If a use case cannot clearly define output quality, workflow outcomes, and expected economic logic, then oversight becomes reactive. Leaders are left judging momentum by anecdote. Boards see activity but not always business effect. Investors see spend but not always return logic. NACD's guidance reflects this shift directly. It says directors should ask what returns are expected from generative AI, how financial plans change over three to five years, and how ROI and KPIs are being measured. This is why measurement belongs inside strategic control. It is not a reporting exercise after the fact. It is one of the mechanisms by which leaders decide what to back, what to challenge, and what to scale. What
Building A Real Executive Scorecard
SPEAKER_00better AI measurement looks like? Better measurement starts by accepting that one number is not enough. Executives need a measurement stack, not a single ROI figure. Leading indicators matter because they show whether a system is being adopted, whether quality is stable, and whether workflows are changing in the intended direction. Lagging indicators matter because they show whether those changes are producing real business and financial outcomes. NIST's guidance is useful here because it pushes organizations to define fit-for-purpose metrics, acceptable limits, pre- and post-deployment comparisons, and incident or error logic rather than relying on generic adoption signals. A practical executive scorecard therefore needs at least five linked classes of measures activity, output quality, workflow performance, business outcomes, and economic impact. In more mature settings, it should also include risk-adjusted value, such as incident cost, compliance cost, override rates, downside scenarios, and error severity. That is especially important in AI where speed gains can conceal quality degradation or delayed operational cost. The point is not to make measurement more complicated than it needs to be, it is to make it credible enough to support real decisions. If the scorecard cannot explain whether AI is improving work, improving the business, and doing so with acceptable risk, then it is not yet an executive measurement system. Why
Value Shows Up In Workflows
SPEAKER_00value appears at workflow level? One of the most important lessons in this series is that AI value usually appears at workflow and system level, not at tool level. That matters just as much in measurement as it does in work design. A common mistake is to measure AI at the point of interaction rather than at the point of outcome. A user may complete a task faster, but that does not necessarily improve the wider workflow. If quality checks, handoffs, escalation, or downstream coordination remain weak, local efficiency gains may never translate into meaningful business performance. This is why user metrics are necessary but insufficient. A user can work faster while the wider process remains weak. Value appears when the surrounding workflow improves, not simply when one step becomes quicker. DBS remains a useful anchor example because it shows what more mature value measurement can look like. The bank reports that its data analytics and AI initiatives delivered approximately $1 billion Singapore dollars of economic value in 2025, while also reducing code deployment time and shortening model deployment cycles. The important point is not the headline number alone. It is the link between platform capability, workflow performance, deployment speed, and measurable economic impact. That is the pattern leaders should focus on. AI value becomes more credible when organizations move beyond activity metrics and start measuring how workflows, business performance, and economic outcomes improve together.
How To Measure AI Starting Now
SPEAKER_00How leaders should measure AI now. The first shift is conceptual. Leaders should stop asking only how much AI is being used and start asking where value is being created, how it is being evidenced, and at what level of the system that evidence sits. Activity can show movement. It cannot, on its own, show whether the business is improving. The second shift is structural. The unit of measurement should usually be the workflow or domain, not the individual tool. Pilot stage metrics should establish baselines and prove fit for purpose. System stage metric should prove process change and outcome movement. Scale stage metric should prove economic impact, capital efficiency, and risk adjusted durability. The third shift is managerial. Every important AI initiative should have an explicit value logic before it scales, with an owner, a baseline, a target, a review cadence, and a stop or scale threshold. This is also why the value layer matters so much in the wider human AI operating system. If measurements stay shallow, all the other layers become harder to govern. Leaders cannot prioritize clearly, boards cannot scrutinize credibly, and capital cannot be allocated with conviction. Value is what makes AI scale defensible. AI does not scale because activity looks impressive. It scales when organizations can distinguish visible movement from real value creation. That means measuring more than usage, time saved, and pilot enthusiasm. It means linking AI to output quality, workflow performance, business outcomes, economic impact, and risk-adjusted durability. Until that happens, AI will remain easier to talk about than to govern, easier to deploy than to prioritize, and easier to use than to defend. The core eight-part series ends here, but one more piece remains. In the bonus article, I will bring the argument together and examine the bigger question behind the whole series. What organizational advantage looks like when AI becomes part of how work, decisions, capability, governance, and value operate. This concludes the article. You can also read this article on my LinkedIn page where I share regular insights on AI, strategy, and emerging technologies.