Beyond Story Points and Timesheets: How Should Delivered Work Be Measured?
For a long time, we have measured work using things that were easy to count.
Hours worked. People assigned. Tickets closed. Story points completed. Utilization. Capacity. Velocity.
None of these are useless. They helped organizations create structure around work that was otherwise difficult to see.
But they were always proxies.
An hour is not an outcome. A story point is not customer value. A full team is not proof that something important happened.
We tolerated the gap because, for most of modern business history, human effort and productive capacity were closely connected. If five people spent six weeks working on something, that effort usually told us something about the scale of the work.
AI is beginning to break that connection.
A task that took an analyst three days may now take an hour. A testing activity that required several people can run continuously through automation. A developer working with coding agents may produce what previously required a small team.
The work has not necessarily become less valuable.
It has simply become less connected to time.
That leaves us with an uncomfortable question.
If time spent is no longer a reliable representation of work delivered, what should we measure instead?
We have confused effort with value for a very long time
Consider a simple example.
Two teams are asked to solve the same problem. Team A takes four weeks. Team B takes two.
If we bill by hours, Team A produces twice as much billable work.
But if both outcomes are equally good, Team B created the same value in half the time.
In fact, if speed mattered to the customer, Team B created more value.
Yet many commercial and management systems reward the opposite.
The longer something takes, the more effort appears to have been invested. More effort can look like more seriousness. More people can look like more capacity. More meetings can make a program appear important.
This is partly because effort is visible.
Value is harder.
A timesheet gives us a number. An outcome requires judgment.
That difference shaped entire industries.
Consulting firms sold people by the day. Outsourcing companies built rate cards around roles. Internal finance teams planned headcount. Engineering organizations used capacity estimates. Professional services firms tracked utilization.
None of this was irrational.
Time was the common denominator because people were the primary productive resource.
But the common denominator is becoming less common.
AI makes the weakness impossible to ignore
Imagine that an experienced engineer needs six hours to diagnose a difficult production issue.
Now imagine the same engineer has an AI agent that reviews logs, compares recent deployments, checks known failure patterns, and prepares a likely root-cause analysis in twelve minutes.
The engineer still needs judgment. They may reject the agent's first conclusion. They may notice something the model missed. They are still accountable for the fix.
But the economics of the work have changed.
If we measure value through hours, the better system appears to create less work.
That is absurd.
The same problem will appear everywhere.
A lawyer using AI may review a document in one hour instead of four. A designer may explore twenty directions before lunch. A finance team may automate reconciliation that previously consumed several days every month.
We cannot keep telling ourselves that time spent is the value when technology is specifically designed to reduce the time spent.
The issue is larger than billing.
It affects how companies understand productivity itself.
Story points solved a different problem
Software teams tried to escape some of these limitations through story points.
The idea was useful. Do not pretend that every task can be estimated precisely in hours. Measure relative complexity instead.
A five-point story is bigger than a two-point story. Teams learn their own velocity. Planning becomes more realistic.
The problem begins when story points escape their original purpose.
They become performance measures.
Management starts comparing teams.
People begin worrying about how many points they completed rather than what changed for the customer.
Then the inevitable gaming begins.
A task that used to be three points mysteriously becomes five.
Velocity improves.
Nothing meaningful changed.
The metric becomes the work.
This happens with almost every measurement system eventually.
When the number becomes the target, people optimize the number.
AI makes story points even stranger.
If an agent can perform much of the implementation, did the complexity disappear? Or did the human effort disappear while the underlying complexity remained?
Those are not the same thing.
Complexity and effort are different
This distinction matters.
A system migration may be extremely complex even if strong tooling allows most of the execution to happen quickly.
A regulatory interpretation may require only two hours of expert time but carry enormous consequence.
A repetitive data-cleaning task may consume hundreds of hours while requiring little judgment.
Effort tells us how much resource was consumed.
Complexity tells us something about the nature of the problem.
Risk tells us something else.
Value tells us something else again.
We have spent years trying to squeeze all of these into one convenient measure.
They do not fit.
A better model has to accept that work has several dimensions.
Start with the simplest question: what became true?
This is where I think measurement should begin.
Not with hours.
Not with points.
Not with people.
Ask what changed.
A customer was not live. Now the customer is live.
A dataset was inconsistent. Now it is reconciled and trusted.
A vulnerability existed. Now it has been removed and verified.
A legacy workflow required manual intervention. Now it runs automatically under agreed controls.
A product capability did not exist. Now customers can use it.
These are changes in state.
They are closer to actual delivery.
This connects directly to the thinking behind Execution Is the New Scarcity. Execution is not activity. It is the movement from intention to reality.
If we want to measure delivery, we need to get closer to that movement.
“Done” should mean something outside the team
There is a strange thing that happens in organizations.
A team can complete its work while the outcome is still unfinished.
Engineering says development is done.
Security says review is pending.
Operations says deployment has not happened.
The customer says they cannot use it yet.
Finance says the business case assumed adoption that has not occurred.
Everybody may be technically correct inside their own boundary.
The organization is still not done.
This is one of the reasons execution fails even when talented people are involved, something we explored in Why Execution Fails Despite Smart People.
Completion needs an external reference point.
Something should be observably different after the work.
The stronger that definition is, the more meaningful our measurement becomes.
Verification is what turns activity into delivery
Suppose a team says a migration is complete.
What does complete mean?
It may mean the scripts ran successfully.
That is useful information, but it is not enough.
Were all records migrated? Were they reconciled? Were critical relationships preserved? Were errors within the accepted threshold? Did the customer validate the result?
The answer to these questions creates evidence.
Without evidence, “done” is an opinion.
This is why verification sits so centrally inside The Virtual Delivery Center Protocol.
A delivered unit of work should have a finish line that can be demonstrated.
Not every finish line will be mathematically precise. Business is messier than that.
But there should be enough evidence for the people involved to agree that the state changed.
This is where the idea of a Delivery Unit becomes useful
A Delivery Unit should not be thought of as a new version of an hour.
That would miss the point.
It is closer to a bounded piece of delivered work whose completion can be understood and verified.
For example, “twenty hours of integration development” is effort.
“Customer identity system integrated, tested, and accepted against the agreed authentication scenarios” is a delivered unit.
The second description tells us much more.
It tells us why the work exists.
It gives us a boundary.
It creates a basis for acceptance.
It can still have complexity, effort, risk, and cost attached to it, but those things surround the delivered unit rather than replacing it.
That is an important shift.
A Delivery Unit should not pretend all work is equal
This is where any new measurement model can become silly very quickly.
We cannot say that one delivered unit is universally equivalent to another.
Implementing a payment integration is not the same as producing a competitive analysis.
Resolving a critical security issue is not the same as updating a webpage.
Even two technically similar modules may differ dramatically because one touches a regulated system and the other does not.
So the goal should not be to create a universal currency of work.
The goal is to create a better representation of completed work.
A Delivery Unit can carry additional attributes such as complexity, capability requirements, risk, urgency, verification burden, and dependencies.
That makes it useful for planning and economics without pretending that all outcomes are identical.
Think of it as packaging the meaning of work
This is perhaps the simplest way to explain it.
Today, work is often packaged as time.
Forty hours of an engineer.
Three months of a team.
Two days of consulting.
A Delivery Unit packages work around meaning.
What is being delivered?
What conditions define completion?
What capabilities are needed?
What dependencies exist?
How do we know it worked?
Once those things are clearer, time becomes one input rather than the product itself.
Time still matters
There is an easy mistake to make here.
If we become enthusiastic about outcomes, we may start pretending time no longer matters.
Of course it matters.
A company has budgets. People have finite lives. Delay has cost. Urgency changes value.
The point is not to stop measuring time.
The point is to stop confusing time with delivery.
If a module took forty hours when we expected twenty, that tells us something useful. Maybe the problem was harder than expected. Maybe the process was inefficient. Maybe we learned something.
But the forty hours are diagnostic information.
They are not proof that the module deserves to cost twice as much.
That distinction changes incentives.
A good measurement system should reward getting better
This is one of the tests I would use.
Suppose a team develops automation that cuts the delivery time of a module from ten days to two.
Does the measurement system reward that improvement?
Or does it reduce the team's revenue and make the innovation economically unattractive?
If efficiency hurts the provider, something is broken.
If a customer pays exactly the same regardless of whether a provider becomes ten times more efficient, something may also be broken.
The productivity gain should create value for both sides.
The customer should benefit from faster or cheaper delivery.
The people or platform creating the improvement should be rewarded for becoming better.
A good economic model encourages learning.
It should not reward slowness.
AI makes this urgent
This tension has existed for decades, but AI makes it harder to hide.
If a consulting firm develops an internal agent that reduces a week's work to one afternoon, should the client receive one afternoon's value?
Probably not.
If the firm continues billing a week as if nothing changed, should the client accept that forever?
Probably not.
The real answer lies somewhere between the two.
The customer is buying a result.
The provider has invested in capability that made the result faster.
Both contributed to the productivity gain.
Outcome-based economics gives us a way to discuss that honestly.
The old model does not.
Some work cannot be defined cleanly in advance
This is where I want to be careful.
There is a tendency in business writing to find one elegant model and force everything into it.
Real work refuses to cooperate.
Research is a good example.
You may spend three weeks investigating a market and conclude that the opportunity should not be pursued.
Was the work unsuccessful?
Not necessarily.
You may have saved the company millions.
Discovery work behaves similarly.
You may begin believing that a system requires modernization and discover that the better answer is retirement.
Creative work can also resist rigid definitions.
You sometimes do not know what the best output looks like until you see it.
In these situations, trying to define a fixed outcome too early can create worse decisions.
Discovery should be measured differently
For uncertain work, the output may be reduced uncertainty.
That is still an outcome.
A discovery module might produce a validated problem definition, architecture options, risk assessment, and recommendation.
It does not need to promise that the eventual product will succeed.
It needs to produce enough evidence for the next decision.
That is a more honest definition.
The delivered state becomes:
“We know enough to decide.”
Sometimes that is enormously valuable.
Operations are different again
A service desk does not “finish.”
Neither does infrastructure monitoring.
Neither does fraud detection.
These are continuous functions.
Trying to convert them into artificial project outcomes would create unnecessary complexity.
Here, service levels and operational health make more sense.
Availability.
Resolution time.
Error rates.
Customer satisfaction.
Exception rates.
Cost per transaction.
The broader lesson is simple.
Measurement should follow the nature of the work.
Not every piece of work needs the same ruler.
What we need is a measurement hierarchy
At the top is the business outcome.
Did the thing we ultimately care about improve?
Below that sit delivered modules.
What meaningful pieces of work were completed and verified?
Below that sit operational measures.
How well did the execution system perform?
Then we can look at resource measures.
How much human time, machine capacity, specialist effort, and money were consumed?
This hierarchy prevents the bottom from swallowing the top.
Hours remain visible.
They simply stop pretending to be value.
AI introduces machine effort, which makes hours even stranger
Imagine a module completed through a combination of one engineer, three agents, automated tests, and a specialist who joins for thirty minutes.
How many hours did the work take?
The engineer may have worked six hours.
The agents may have performed thousands of operations in minutes.
The specialist contributed half an hour.
The automated test system ran for forty minutes.
Which number describes the work?
None of them really does.
They describe resource consumption.
The module itself describes what was delivered.
That is why human-agent execution, discussed in How AI Agents Fit Inside a Virtual Delivery Center, pushes us toward a new unit of measurement almost automatically.
Once the team stops being purely human, labor time stops being a sufficient description of production.
The customer should not have to care how the sausage was made
There is a point where transparency can become unnecessary detail.
If a customer asks for a secure integration and receives a secure integration that passes every agreed verification step, does it matter whether it required three engineers or one engineer using agents?
Sometimes yes.
If security, data, or contractual obligations require disclosure, it matters.
But commercially, the customer should not have to finance inefficiency simply because the provider chose a labor-heavy method.
The customer should care about the result, the quality, the risk, the timing, and the agreed economics.
How the delivery system achieves that result should increasingly become the responsibility of the execution provider.
This is part of the transition described in What Comes After Consulting and Staffing?.
The value moves away from supplying labor and toward supplying reliable execution capacity.
Internal teams need this shift too
It is easy to make this conversation about vendors.
The bigger problem may be inside companies.
Internal teams often measure themselves using activity proxies.
Tickets.
Sprints.
Releases.
Lines of code.
Projects completed.
Budgets consumed.
People utilized.
These numbers make management easier, but they can create false comfort.
A product team can ship twelve releases that customers barely notice.
Another team can make one change that reduces customer churn materially.
Which team delivered more?
The answer depends on what the business needed.
Measurement should reconnect teams to that question.
The danger is measuring only business outcomes
There is another extreme.
Some leaders say, “Fine. Then everyone should be measured purely on revenue, profit, or customer growth.”
That can also be unfair and destructive.
Most people do not control the whole outcome.
An engineer can build an excellent capability that fails commercially because the pricing was wrong.
A marketer can generate strong demand that fails to convert because the product is weak.
A delivery team can implement perfectly while the customer refuses to change its internal process.
Outcome measurement needs boundaries.
People and teams should be accountable for what they can reasonably influence.
This is why decomposing outcomes into meaningful modules matters.
It creates a fairer connection between contribution and consequence.
Verification protects both sides
One of the understated benefits of better work measurement is reduced argument.
When work is billed by time, customers may ask whether the time was really necessary.
Providers may feel pressured to defend every hour.
When work is vaguely “fixed price,” providers worry about endless scope expansion.
A verified delivery unit can create a clearer conversation.
This is what we agreed would be delivered.
These were the conditions.
Here is the evidence.
It is complete.
That does not eliminate disagreement.
Nothing does.
But it moves the disagreement closer to the work itself.
The future commercial conversation becomes more interesting
Instead of asking, “How many people do you need?”
A customer may ask, “What needs to be delivered?”
Instead of asking, “What is your blended hourly rate?”
They may ask, “What is the cost of this verified module?”
Instead of comparing provider headcount, they may compare execution architecture.
Who can deliver faster?
Who uses AI intelligently?
Who verifies independently?
Who preserves context?
Who takes responsibility for the whole outcome?
That changes competition.
It rewards the provider with the better system rather than simply the cheaper labor pool.
This is also where VDCs become economically different
A Virtual Delivery Center should not merely reproduce an outsourcing contract inside a new container.
If the VDC is organized around persistent execution capability, its economics should gradually reflect that.
There may still be a subscription for maintaining the environment.
Some work may still be time and material, especially early discovery.
Some specialist involvement may still be priced by the hour because that is the sensible unit.
But more delivery can move toward bounded modules with clear verification.
Over time, the organization gains historical information.
It knows how similar work behaved.
It knows which capabilities were required.
It knows where risks emerged.
It knows what agents reduced effort.
It knows which specialists performed well.
That history can improve pricing.
The system learns.
Delivery Units can become a memory of execution
This may eventually be more valuable than the billing model itself.
Imagine an enterprise that can look back across hundreds of completed modules.
It can see which types of work repeatedly take longer.
Which dependencies create delay.
Which capabilities are scarce.
Where AI consistently reduces effort.
Where verification catches failures.
Which work should be internal.
Which work can be composed externally.
Now the organization is not just measuring delivery.
It is learning how it executes.
That knowledge compounds.
A Delivery Unit becomes not only a commercial object, but a piece of execution history.
The measurement should become invisible where possible
There is one thing I would strongly resist.
We should not replace timesheet bureaucracy with Delivery Unit bureaucracy.
People should not spend their evenings categorizing work into elaborate new boxes.
The system should infer as much as possible from the normal flow of execution.
Work already happens in customer tools.
Code is merged.
Tests run.
Tickets move.
Customers accept.
Systems generate logs.
Agents produce records.
Financial systems record payment.
The measurement layer should observe reality rather than constantly asking people to describe reality.
If the cost of measuring the work becomes a meaningful part of doing the work, we have failed.
Human judgment will still matter
There is no formula that can perfectly measure the value of every piece of work.
I do not think we should try to create one.
There will always be judgment.
Was the architecture elegant?
Did the customer feel understood?
Was the decision courageous?
Did the team handle an unexpected situation well?
Did a specialist notice something no metric would have captured?
Human beings are not going to disappear from evaluation simply because we become better at measuring outcomes.
The goal is not perfect quantification.
The goal is better honesty.
Better measurement changes behavior
This is why the topic matters more than it first appears.
People optimize what organizations measure.
Measure hours, and people protect hours.
Measure utilization, and managers try to keep everyone busy.
Measure story points, and teams learn to produce points.
Measure headcount, and executives treat workforce size as capacity.
Measure verified outcomes, and something begins to shift.
Teams ask what completion actually means.
Leaders clarify ownership earlier.
Providers have reasons to automate.
Customers focus more on acceptance.
AI becomes useful when it improves execution, not merely when it produces impressive demonstrations.
Measurement quietly redesigns the organization.
The future unit of work is a verified change in state
I do not think story points will disappear.
Timesheets will not disappear either.
Some industries will continue using hourly billing for a very long time, and sometimes that will still be the right choice.
But they should move down the hierarchy.
They are useful details about how work happened.
They should not remain our primary definition of what the work was.
The more AI participates, the more obvious this becomes.
A future team may consist of two humans, six agents, one specialist, and several systems. Another team may consist of twelve employees. Both may deliver the same module.
What matters is not how crowded the room was.
What matters is what changed.
Was it real?
Was it accepted?
Was it verified?
Did it create value?
Did we learn enough to do it better next time?
Those are harder questions than asking how many hours were spent.
That is probably why we avoided them for so long.
But the tools are changing now. The workforce is changing. The economics are changing.
Our measurement system has to catch up.
For the last century, we measured work largely by the human effort consumed in producing it.
The next era will need to measure something closer to the truth.
Not how long we worked.
Not how busy we looked.
Not how many people we assigned.
What did we actually make true?