Open letter
Dear leaders of the software tools industry,
I am writing this as an old friend of software performance measurement.
For more than a decade, I have worked in developer analytics, probabilistic forecasting, flow metrics, and software delivery performance. Some people know me for Monte Carlo forecasting. Some know me for the Six Dimensions of software performance - my attempt to move teams away from one-dimensional productivity measures and toward a more balanced view of delivery.
So I say this with affection, not cynicism:
I think we are about to measure AI adoption badly.
Worse, I think we are at risk of repeating some of the oldest mistakes in software measurement, just with newer dashboards, better branding, and much more expensive models behind the scenes.
Section
We have been here before
The software industry has a long history of mistaking visible activity for meaningful performance.
- Lines of code written
- Stories completed
- Tickets closed
- Pull requests merged
- Cycle time reduced
- Deployments increased
None of these measures are useless. Many are helpful in context. I have used them, taught them, and built forecasting models around them.
But they become dangerous when they are treated as the goal.
The point of software delivery was never to produce more artifacts. It was to create more customer and business value, safely, sustainably, and predictably.
AI does not change that.
In fact, AI makes that distinction more important.
Section
Developer productivity is the wrong center of gravity
Much of the current conversation around AI engineering performance still seems to orbit developer productivity.
- How many developers are using AI?
- How many lines of code were generated?
- How many pull requests were created?
- How much faster was the review?
- How much coding time was saved?
These are understandable first measures. They are easy to instrument. They give executives something concrete to look at. They make AI adoption feel real.
But they are not enough.
And if we stop there, they will mislead us.
The future of AI engineering performance is not developer performance. It is value delivery performance.
The question is not simply whether developers can write code faster.
The question is whether the organization can move more valuable work safely through an AI-enabled delivery system - with fewer regrets.
That phrase matters: without regrets.
- More code generated with more defects is not progress.
- More pull requests merged into fragile systems is not progress.
- More features started without clearer strategy is not progress.
- More automation applied to unsafe areas of the codebase is not progress.
- More output that never becomes customer value is not progress.
AI adoption should not be measured by how much AI is used.
It should be measured by how much better the organization becomes at delivering valuable outcomes through AI-assisted work.
Section
The real constraints are moving
Historically, portfolio planning was shaped by the scarcity of expensive delivery capacity.
Teams were expensive. Coordination was expensive. Engineering time was expensive. Starting work was expensive. Because of that, choosing carefully mattered.
That world is changing.
As AI engineering matures, more work will become cheap enough to start. Many features, ideas, prototypes, fixes, and experiments that once took months may take days or hours. By 2027 and beyond, this shift will be impossible to ignore.
But when it becomes easier to start almost anything, the question changes.
The question is no longer only: Can we build this?
The better question is: Should we start this?
And if we start it, what needs to be true for it to become value?
That includes questions like:
- Is the intent clear enough?
- Is the work strategically aligned?
- Is the code area safe for AI acceleration?
- Are there security, compliance, or privacy concerns?
- Is the blast radius understood?
- Are there dependencies that must happen before or after this work?
- Is the system too fragile or debt-heavy for safe automation?
- Will this work actually reach customers, sales, support, marketing, enablement, and adoption?
This is where many current tools are weak.
They are good at seeing the mechanics of software delivery. They are much less good at seeing whether the organization is learning to move the right work through the system safely.
Section
AI adoption is not a usage metric
One of the biggest traps in this next wave will be treating AI adoption as a usage metric.
- Number of AI users
- Number of prompts
- Number of completions
- Number of generated lines
- Percentage of code touched by AI
- Number of AI-assisted pull requests
Again, these are not irrelevant.
But they are not adoption.
Usage tells us that people are touching the tool. It does not tell us whether the organization is improving.
Real AI adoption means the organization is learning where AI helps, where it harms, where it is blocked, where it creates risk, and where it changes the economics of delivery.
Real AI adoption means lessons from one AI-assisted delivery effort become guidance for the next similar effort.
It means the organization can say:
- We tried using AI in this kind of work.
- Here is where it accelerated us.
- Here is where review slowed down.
- Here is where sensitive code constrained us.
- Here is where unclear intent caused rework.
- Here is where tech debt made automation unsafe.
- Here is where downstream go-to-market work, not coding, became the bottleneck.
- Here is what we should do differently next time.
That is adoption.
Not usage. Learning.
Section
The front and back of delivery matter more now
For years, software organizations have over-focused on the middle of delivery: coding, review, testing, merging, deployment.
Those things matter. They always will.
But AI shifts attention to the fuzzy front and back ends of delivery.
The front end is where we decide what is worth doing, whether the work is clear, whether the assumptions are valid, and whether the organization understands the risk.
The back end is where code becomes value: documentation, pricing, packaging, marketing, sales enablement, support readiness, customer adoption, training, measurement, and feedback.
Many organizations were never especially good at either end.
AI will expose that.
If coding gets faster but the front end remains confused, we will start the wrong work faster.
If coding gets faster but the back end remains weak, we will produce more software that never turns into value.
If coding gets faster but risk management does not improve, we will scale regret.
That is why AI engineering performance has to be bigger than developer productivity.
Section
What should the next generation of measurement include?
I believe AI engineering performance measurement needs to move beyond activity and into a more balanced view.
At minimum, it should help answer questions across several dimensions:
- Value
- Intent clarity
- Strategic alignment
- Safety and risk
- Readiness for AI acceleration
- Learning
- Value realization
Are we accelerating work that matters to customers and the business, or merely increasing output?
Is the work clear enough for humans and AI systems to execute safely?
Does this work support the current priorities of the organization?
Is the work touching sensitive code, regulated areas, security boundaries, or high-blast-radius systems?
Is this the kind of work AI can safely help with, or are unclear requirements, tech debt, architecture, or dependencies likely to slow or distort the result?
Are we capturing what happened and applying those lessons to future similar work?
Did the work make it through launch, adoption, enablement, and customer impact - or did it stop at code complete?
This is not anti-DORA. This is not anti-flow metrics. This is not anti-developer productivity.
It is a recognition that the unit of analysis is changing.
The team is no longer the only interesting unit. The pull request is not the only interesting artifact. The developer is not the only actor. The delivery system is becoming part human, part AI, part product strategy, part governance, part learning loop.
Our metrics have to grow up accordingly.
Section
The danger of another productivity theater
There is a real danger that AI engineering measurement becomes the next productivity theater.
Executives will ask whether AI is working. Tools will respond with visible activity. Dashboards will show more code, more completions, more PRs, more usage, more developer time saved.
And for a while, that may look convincing.
But eventually the harder questions will arrive.
- Why are we shipping more but not winning more?
- Why did AI increase review burden?
- Why did we accelerate low-value work?
- Why did we create more rework?
- Why did we automate in parts of the system where we should have been more careful?
- Why are teams using AI heavily but still blocked by dependencies, unclear intent, and go-to-market delays?
- Why did the dashboard look good while the outcomes did not?
The industry has seen this movie before.
We should not be impressed by a new version of lines-of-code accounting just because the lines are now generated by a model.
Section
A request to the software tool industry
So this is my open request to the leaders building the next generation of software engineering tools:
Please do not reduce AI engineering performance to developer activity.
Please do not make more code generated the new productivity proxy.
Please do not frame AI adoption as usage alone.
Please do not stop at pull requests, reviews, cycle time, and coding acceleration.
Those are useful signals, but they are not the destination.
The opportunity is much bigger.
You have the chance to help organizations understand how AI changes the economics, risks, dependencies, and learning loops of software delivery.
You have the chance to connect engineering work to value delivery more honestly than we have in the past.
You have the chance to show leaders not just whether teams are using AI, but whether they are becoming better at choosing, shaping, accelerating, governing, and realizing value from AI-assisted work.
That is the measurement world I hope we build.
Section
Closing thought
I am enthusiastic about AI engineering. I am not skeptical of the technology. I am skeptical of shallow measurement.
The winning organization will not be the one that uses the most AI.
It will be the one that learns fastest how to apply AI to the right work, in the right way, with the right safeguards, and with the fewest regrets.
That is the performance conversation worth having.
And I hope the software tools industry leads us there.

