In short
Estimates are systematically optimistic because people estimate the work itself and then deliver the work plus waiting, rework, interruption and coordination. The fix is not to pad, which gets negotiated away, but to estimate from evidence: find three similar past pieces of work and use what they actually took. Then quote a range rather than a number, and track the ratio between your estimates and your outcomes until it stabilises. Most teams discover a consistent multiplier, and applying it openly is more accurate than any refinement of the estimating method.
Key takeaways
- Estimates fail in one direction, which means the error is structural and therefore correctable.
- You estimate working time and deliver elapsed time. The gap is waiting, rework and interruption.
- Past actuals beat expert judgement. Three comparable jobs is enough to start.
- Padding privately gets negotiated away. A stated multiplier survives scrutiny because it has evidence behind it.
- Quote ranges. A single number communicates a confidence you do not have and cannot deliver.
Wrong in one direction is a solvable problem
If estimates were wrong randomly, half of all projects would finish early. They do not. Overruns dominate, in every organisation, on every kind of work, which tells you something useful: this is not a skill problem distributed across individuals. It is a structural bias, and structural biases can be corrected without anyone becoming better at anything.
The bias has a simple mechanism. Asked how long something will take, a person mentally simulates doing it. That simulation contains the steps of the work. It does not contain the client who takes four days to send the file, the review that comes back with changes, the morning lost to an unrelated emergency, or the fact that the person doing it also has three other things running. All of those happen every time, and none of them appears in the imagined version.
You are asked how long the work takes. You answer accurately. Then you deliver the work plus everything that happens around the work, which nobody asked about.
What estimates leave out, reliably
The omissions are consistent enough to list, which is the first step to including them.
| Omitted | Typical share of elapsed time |
|---|---|
| Waiting for inputs, feedback or decisions | Frequently the largest single component |
| Rework after review | Significant, and almost never estimated separately |
| Coordination: briefing, handoffs, status, questions | Grows with the number of people involved |
| Interruption and context switching | Invisible per instance, substantial in aggregate |
| The two or three surprises every project has | Individually unpredictable, collectively certain |
That last row is the one people resist, because each surprise genuinely is unforeseeable. But the existence of surprises is entirely foreseeable. Every project has some. An estimate built on the assumption that this one will not is an estimate of the best case being presented as a plan.
The waiting row is usually the largest and the most fixable. If a fifth of your elapsed time is working time, as a first current-state map often shows, then estimating the working time and quoting it as a delivery date is guaranteed to fail by a factor with no relationship to how well the work was estimated.
Estimate from evidence, not from imagination
The most effective single change is to stop asking “how long will this take?” and start asking “how long did the last three like it take?”
This works because past actuals include everything the imagined version omits, without anyone having to remember to include it. The waiting, the rework and the surprises are already in the number.
- Pick three comparable pieces of completed work. Same type, roughly similar size. Three is enough to start; more is better but three beats none by a wide margin.
- Use elapsed time from start to delivery, not logged hours. The client experiences elapsed time and the deadline is set in it.
- Note what was different about each, so you can adjust rather than average blindly.
- Adjust for known differences only. Resist adjusting downwards because this one will run more smoothly. It will not.
Comparable past projects is usually enough to beat expert judgement on a new one. The information is already in your email, your project tool and your invoices; almost nobody looks at it before quoting.
The common objection is that every project is different. They are, and they are different in ways that mostly cancel out across three examples, while the systematic omissions do not cancel out at all. A rough number that includes the overhead beats a precise number that excludes it.
Quote ranges, and say what the range means
A single-number estimate communicates a confidence nobody has. It also invites the recipient to treat the number as a commitment, which is how a best case becomes a deadline.
A range communicates the real state of knowledge, provided both ends are meaningful:
- The low end is what happens if inputs arrive promptly and nothing unusual occurs.
- The high end is what happens with normal delays and one surprise.
- Say what would move it. “Three to five weeks; the difference is mostly how quickly we get the brand assets and feedback on the first draft.”
That third element matters more than the range itself. It converts the estimate into a shared responsibility with a visible lever, and it tends to make clients faster at the things that were going to delay you, because now they can see the connection.
Plan internally against the high end and communicate the range. Planning against the low end and hoping is how a business ends up committing capacity it does not have, which is a capacity planning failure that presents as an estimating failure.
Why padding fails and multipliers do not
Most experienced people already pad. They estimate four days, say six, and are still late. Padding fails for two reasons.
First, it gets negotiated away. A padded number is a number the estimator cannot justify, so under commercial pressure it is compressed back towards the unpadded figure, which was the honest estimate of the best case.
Second, padding is applied to the estimate rather than derived from outcomes. It is a feeling about uncertainty, not a measurement of it, so it is applied inconsistently and usually insufficiently.
A stated multiplier behaves differently. Track, for ten completed pieces of work, the ratio of what it actually took to what was estimated. Most teams find a stable number, commonly somewhere between 1.3 and 1.8. Then apply it openly.
- Estimation ratio
- The observed relationship between what work was estimated to take and what it actually took, measured across a set of completed jobs. Applied openly it is evidence, which survives negotiation, unlike a private buffer, which does not.
“Our estimates have run about 1.5 times over the last ten projects, so I am quoting six weeks for four weeks of work” is a defensible sentence. “Six weeks, to be safe” is not, and will become four in the next conversation.
Improving estimates without a measurement programme
None of this requires time tracking or a new tool. Two habits are sufficient.
| Habit | Effort | What it produces |
|---|---|---|
| Record the estimate where the work lives, before starting | Seconds | Without this there is nothing to compare against, which is why most businesses cannot calculate their ratio |
| At delivery, note actual elapsed time and the main cause of any gap | A minute | The ratio, plus the recurring causes, which are usually two or three things |
After ten jobs you will have both a multiplier and a short list of what consistently goes wrong. That second output is frequently more valuable, because the causes are typically fixable: inputs arriving late, a review step with no deadline, rework from an unstated standard. Each of those is a specific operational problem rather than an estimating one, and fixing it improves delivery rather than merely improving the prediction of poor delivery.
Keep the review blameless. If noting the gap between estimate and actual becomes performance monitoring, the estimates will drift upwards defensively and the data becomes useless. The purpose is a better multiplier and a shorter list of causes, which is the same discipline that makes rework measurement work. The Mayim Ops assessment measures flow efficiency, rework and manual handling directly, which between them account for most of the distance between what a team estimates and what it delivers.
Frequently asked questions
Why are project estimates always too low?
Because people estimate the work and then deliver the work plus everything around it: waiting for inputs, revisions, interruptions, coordination and the two or three surprises that occur on any real project. The error is one-directional, which means it is structural rather than random and can be corrected with a multiplier drawn from your own history.
What is the planning fallacy?
The well-documented tendency to estimate a task by imagining it going as intended, which excludes the delays and complications that have accompanied every similar task. It persists even among people who know about it, which is why the practical remedy is comparison with past actuals rather than trying harder to think of everything.
How do you estimate a project accurately?
Find three past pieces of work of a similar type, use what they actually took from start to finish including waiting, and adjust for known differences. This outperforms bottom-up task estimation in most small businesses because it captures the overhead that task lists systematically omit.
Should you add a buffer to estimates?
A visible multiplier works better than a hidden buffer. Padding added privately gets negotiated away by whoever is under pressure, because nobody can defend a number they cannot explain. A stated ratio drawn from your own history survives that conversation because it is evidence rather than caution.
How do you estimate work you have never done before?
Estimate the closest thing you have done, state explicitly what is different, and widen the range rather than the midpoint. For genuinely novel work, timebox a small piece first and estimate the rest from what that reveals, which converts an unbounded guess into a measurement.