Delivery Forecasting In 2026: Why Software Estimates Fail And How Serious Teams Fix Them

Team planning session with whiteboard used for software delivery forecasting

A logistics company I worked with in 2025 approved a 16-week dashboard build. It shipped in week 31. The budget was fixed, the scope was fixed, and the date turned out to be a wish. By month five the weekly call had turned into an interrogation: “Where are we, really?” Nobody lied. Nobody coasted. The estimate was a guess dressed up as a plan, and the whole organisation had agreed to treat it as a promise.

That pattern repeats across nearly every software project I audit. It is not a discipline problem. It is a measurement problem. Teams estimate with one number, report progress with percentages, and then act surprised when reality shows up. Fixing it does not require a statistics degree. It requires better inputs and a willingness to report ranges instead of false precision.

Why Single-Number Estimates Keep Failing

An estimate is a snapshot of what a team believes at one moment. A forecast is a distribution of outcomes based on what similar work actually took. Most organisations blend the two and call the result a commitment. Then they manage the project against a number that was never designed to survive contact with reality.

Four root causes explain most of the damage.

1. The Optimism Tax

When someone asks “how long will this take?”, they are usually asking for reassurance as much as information. Engineers know the pressure. They round down. A task that has a 50% chance of finishing in three days gets reported as three days, even though the honest answer is “three days if nothing goes sideways, and one thing always goes sideways.”

Across the projects I have reviewed, the gap between the most likely estimate and the actual outcome averages about 60% on anything involving an unfamiliar integration. It is not random. It is systematic, which means it can be corrected once you stop pretending it does not exist.

2. Unknown-Unknowns Are Not Edge Cases

Every project of real size contains work nobody identified. A payment provider’s sandbox behaves differently from production. A legacy database has a column that stores dates as strings. An API rate limit only bites at scale.

None of this appears in a backlog written before the first line of code. Teams that plan for zero unknowns are planning for a world that does not exist. The practical response is not to list every possible surprise. It is to hold explicit capacity every sprint for the work you cannot name yet.

3. Handoff Costs Get Ignored

A feature is not finished when the code compiles. It is finished when it is tested, deployed, monitored, documented, and understood by whoever supports it at 2 AM. Those steps are real work, and they are routinely left out of estimates because they are uncomfortable to price.

In one backend team I studied, deployment, QA coordination, and review overhead accounted for 34% of total calendar time. The estimate had allocated 8%. The team did not miss the deadline because they were slow. They missed it because the estimate described a different project.

4. Scope Moves Quietly

Scope rarely changes in a dramatic announcement. It changes in twenty-three small Slack messages. “Can we add export?” “The client wants SSO now.” “Legal needs an audit trail.” Each addition looks trivial in isolation and costs a day. Together they can add weeks.

If you do not log and price these requests, your forecast silently becomes fiction. Track them. Even a simple running list with an hour figure per item changes behaviour, because people start to see the meter running.

Estimates, Forecasts, And Commitments Are Three Different Things

The vocabulary matters more than most teams expect, because the words carry different obligations.

  • Estimate: a point-in-time opinion about effort, usually produced by the people who will do the work. It is a tool for deciding, not a contract.
  • Forecast: a probability statement about what will be delivered by when, based on observed data from comparable work. It changes as new data arrives, and that is a feature.
  • Commitment: a business promise made with awareness of the risk, and typically backed by trade-offs in scope or staffing.

Most project disasters come from collapsing all three into a single sentence delivered in a kickoff meeting. Once the word “deadline” is attached to an estimate, everyone starts protecting their position instead of the outcome. Separate them in your documents, and the conversations get dramatically less defensive.

The Numbers Most Teams Never Track

You cannot forecast what you do not measure. The good news is that the required data is small. You need three things: how long items take, how many you finish per period, and how often you are interrupted.

Metric What it answers Healthy signal
Cycle time How long does one item take from start to done? Use the median, not the average. Stable or falling over 8 weeks
Throughput How many items complete per sprint or week? Predictable count, regardless of size mix
Work in progress How much is open at once? Low enough that items actually finish
Percent of items with blocked time How often is the team waiting on something outside its control? Under 15%
Escaped defects per release How much of the finished work comes back to bite you? Flat or trending down
Replan rate How often does the forecast change materially? Occasional, not weekly

None of these require new tooling. Most ticketing systems can produce them in an afternoon. The teams that struggle are not missing dashboards. They are missing the discipline of writing down what happened, including the unflattering parts.

There is a temptation to over-engineer this step. I have seen teams spend six weeks building a metrics pipeline before producing a single forecast. Do not do that. A CSV export, a few minutes of cleaning, and a chart in a spreadsheet will tell you more in an hour than a custom dashboard will tell you in a quarter. Instrumentation can come later, once you know which questions you actually need answered.

One caution: collect the data at the point where work truly finishes, not where the responsible engineer marks it complete. If the last 10% of a task is deployment and acceptance, those hours belong to the item. Moving the finish line to make numbers look better is the fastest way to make the whole exercise pointless.

How To Forecast Without A Statistics Degree

Forget regression models for a moment. Three practical techniques cover the vast majority of real delivery work.

Sample From Your Own History

Take the last 20 to 30 completed items of similar type. Write down how long each took, in days. You now have an empirical distribution. If 80% of comparable items took 6 days or fewer, then any new item of that type has a reasonable chance of landing in that band.

This is not guessing. It is sampling from something your team already proved it can do. It also has a pleasant side effect: people argue less about estimates when the evidence comes from their own work.

Run A Simple Monte Carlo Simulation In A Spreadsheet

You do not need Python or a dedicated tool. Put your weekly throughput numbers in a column. Then, in a separate sheet, simulate: pick a random week’s throughput, subtract it from your remaining item count, repeat until the count hits zero. Record the number of weeks. Repeat that loop 500 times.

You end up with a distribution: “there is a 50% chance of finishing in 12 weeks, 85% in 15 weeks, and 95% in 18 weeks.” That last number is the one to build commitments on, and it is almost always 40% to 60% higher than the original single estimate. Clients rarely object to that gap once they can see how it is produced.

Use A Reference Class

Before estimating a new initiative, ask a blunt question: when we did something comparable, what actually happened? If your last three integrations averaged 19 days, planning this one at 9 days is not ambition. It is denial with a Gantt chart.

Reference classes are especially valuable with outsourced partners, because they strip out the optimism that comes from wanting to win the deal.

Making Forecasting Work When Part Of The Team Is External

Outsourced delivery adds distance, time zones, and contractual pressure. It also adds a useful forcing function, because vague process becomes visible very quickly when two organisations have to cooperate on it.

Write Risk Into The Contract, Not Just The Date

A contract that promises a fixed date for an unmeasured scope is a dispute waiting to happen. A better structure fixes the team, the rates, and the cadence, then manages scope against a forecast with an explicit confidence level.

If a hard date is genuinely required, price the risk openly. Ask the partner what confidence level they can commit to, then pay for the difference in scope reduction or added capacity. The conversation is uncomfortable for exactly one meeting. Afterwards it prevents six months of blame.

Modern tooling has made this easier than it was a few years ago. Small teams now run their own internal workflow and invoicing systems rather than buying heavy enterprise suites, and some of that tooling gets built in-house. It is worth seeing how product-minded studios structure their own operations, for example at pagii.co, where the same forecasting habits show up in how features ship.

Insist On Demonstrable Increments

A status report is a claim. A working increment is evidence. Require something runnable every two weeks, even if small. When the demo becomes the unit of progress, percentage-complete reporting disappears naturally, and the forecast gains a real anchor.

One client I worked with cut its reporting overhead by roughly 40% simply by replacing written status decks with a 20-minute demo and a three-line written update. The project did not get faster overnight. The visibility did, and problems surfaced weeks earlier.

Share One Source Of Truth

Two boards, one for the vendor and one for the client, guarantee two versions of reality. Pick a single tool, give both sides access, and agree on the definition of “done” in writing. It sounds bureaucratic. In practice it removes the single largest source of end-of-project arguments.

If procurement blocks external access to internal systems, a shared spreadsheet with a weekly export is an acceptable compromise. Perfect integration is not required. Consistency is. Teams that keep their own private trackers usually pay for it later, and always at the worst moment.

Review The Forecast Monthly, Not The Blame Weekly

Forecasts should be reviewed on a schedule everyone accepts, with the assumption that changes are information rather than failure. The moment a slipped forecast triggers a hostile call, people start hiding variance. Once variance is hidden, the forecast is worthless.

A Worked Example: Replatforming A B2B Portal

Here is how this plays out in practice, based on a mid-size engagement I reviewed in early 2026.

A B2B distributor wanted to replace a 9-year-old portal handling 640 active accounts and roughly 4,300 orders per month. The original plan: 20 weeks, two backend engineers, one frontend engineer, one part-time designer. The single-number estimate came from a bidding process, which is the weakest possible source of data.

Instead of arguing about the number, the team spent two weeks measuring. They found that over the previous quarter they had completed 27 comparable items, with a median cycle time of 4.5 days and a weekly throughput of 6 to 9 items. The backlog of real, specified work was 118 items. Using the simulation approach described above, the forecast read: 14 weeks at 50% confidence, 18 weeks at 85%, 21 weeks at 95%.

They committed to the 85% figure with a fixed scope and a monthly review clause. Delivery landed in week 19, one week beyond the commitment and well inside the 95% band. Two things drove the difference: a third-party shipping API that behaved differently under load, and a late compliance requirement that consumed roughly 90 hours.

Compare that with the original 20-week guess. The real answer was 19 weeks, close in hindsight, but only because optimism and luck cancelled out. Had the shipping API been worse, or the compliance request larger, the same process would have produced a 26-week project with a 20-week promise attached.

The financial difference is easy to underestimate. Extending an engagement by six weeks on a team of four costs roughly 480 engineer-hours, plus the internal coordination time on the client side, plus the revenue impact of a delayed launch. On this project, the reforecast conversation took about three hours in total across two meetings. Those three hours protected nearly 500 hours of downstream work.

There was also an unexpected benefit. Because the team had measured their history, they could push back on the shipping API timeline with evidence rather than instinct, and the integration was ultimately scheduled behind a feature flag. That decision alone removed a dependency from the critical path. Good forecasting does not just predict the future. It gives you the information to rearrange it.

The lesson is not that 20-week estimates are always wrong. It is that a single number gives you no way to know whether you are lucky, accurate, or quietly doomed.

Seven Forecasting Mistakes That Cost Real Money

  1. Treating velocity as a productivity score. The moment velocity becomes a target, it stops measuring anything except the team’s ability to game the metric.
  2. Comparing different teams’ velocity. Story points are not a currency. They are local units, like a house’s own light switch.
  3. Estimating in hours for anything longer than a week. Precision is fake at that distance. Use ranges and confidence levels instead.
  4. Ignoring queue time. Work often waits longer than it takes. In many organisations, queue time is 3 to 5 times the actual work time.
  5. Letting support interrupts stay invisible. If 30% of capacity goes to unplanned work, say so in the forecast rather than absorbing it silently.
  6. Re-baselining without recording the change. A forecast that is rewritten each month with no history teaches you nothing about your own bias.
  7. Reporting percentages instead of outcomes. “72% complete” is an opinion. “Four of seven modules in production” is a fact.

What A Healthy Forecast Actually Looks Like

Forecasts fail loudly, so it helps to define what the opposite looks like. When a delivery process is genuinely under control, a handful of signposts show up repeatedly.

  • Ranges are normal language. Nobody says “three weeks” without also saying how confident they are. Stakeholders ask “at what confidence?” as a reflex.
  • Bad news travels early. The first sign of a slip appears in the weekly note, not in the end-of-quarter surprise. Teams that surface risk early usually finish closer to plan.
  • Scope changes have a price tag. Every new request arrives with an estimate of what it displaces. Trade-offs become conversations instead of arguments.
  • The forecast is boring. If the document is dramatic every month, the process is probably measuring politics rather than work.
  • Historical accuracy is visible. Teams can show what they predicted six months ago and what happened. That record is the single best input for improving future forecasts.

Notice that none of these signposts depend on a particular methodology, framework, or tool. They depend on measurement and honesty. That is why forecasting tends to spread through an organisation by reputation: the team that stops over-promising quietly becomes the team that gets trusted with the important work.

There is also a commercial angle. Buyers increasingly ask vendors for evidence of delivery history, not just references. A partner who can show cycle time trends, replan rates, and past forecast accuracy is signalling something that a slide deck cannot. It tells you the team measures itself even when nobody is asking.

A 30-Day Plan To Start

You do not need a transformation programme. You need four small commitments.

Week 1: Capture The Data You Already Have

Export the last 60 days of completed work. For each item, record start date, end date, and type. Do not clean it excessively. Real, messy data beats a beautiful empty dataset.

Week 2: Publish A First Forecast

Take your remaining backlog and produce a range with three confidence levels. Share it with stakeholders alongside the assumptions. Expect one hard question. Answer it with your own numbers.

Week 3: Add A Risk Buffer You Can Defend

Reserve a fixed slice of each iteration for unknowns, commonly 15% to 20% of capacity. Track what actually consumes it. Within two months you will have a defensible number instead of a hunch.

Week 4: Change The Reporting Format

Replace percentage-complete updates with three lines: what shipped, what is at risk, what changed in the forecast. Keep it short enough that people actually read it. This single change often does more for stakeholder trust than any dashboard.

Frequently Asked Questions

Does forecasting only work for agile teams?

No. Forecasting works wherever work items are recorded and completed over time. Waterfall projects can use the same throughput logic at phase level. What matters is having a consistent unit of work and honest completion data.

How much historical data do I need before forecasting is useful?

Around 20 to 30 completed items of comparable type is enough to start. Fewer than that and your distribution is noisy, but it is still more informative than a guess. You can borrow data from similar past projects, as long as you note the difference in context.

What confidence level should I commit to with a client?

Most organisations commit at 80% to 85%. That leaves room for the work you know exists but cannot yet name. Committing at 95% is safer but usually means a shorter scope, and many clients prefer more functionality with a realistic range over less with an artificial certainty.

Is a fixed-price, fixed-date contract ever reasonable?

Yes, when scope is genuinely stable and small, typically under three months of effort, or when the partner has detailed historical data from very similar work. Beyond that, fix the team and cadence instead, and manage scope against a forecast.

How do I handle a stakeholder who insists on the original date?

Give them the trade-offs rather than an argument. Show what fits inside the date at the current confidence level, and what falls outside. People rarely insist on an impossible date once they can see the price of holding it.

Does AI change any of this in 2026?

Coding assistants have shifted where time goes, not eliminated variance. Teams ship more code per engineer in some areas, but review, integration, and deployment remain the dominant costs. Forecasting still works, and the data should be updated as those patterns change.

How often should the forecast be updated?

Monthly for most projects, or after any material scope change. Continuously rewriting it daily creates noise and erodes trust. Updating it monthly, and recording what changed, builds a record of your organisation’s actual bias over time.

How do we forecast when requirements keep changing?

That is precisely when forecasting earns its keep, because the question stops being “when will everything be done” and becomes “what can we finish by the next review date at 85% confidence.” Freeze a horizon, usually two to four weeks, and treat everything beyond it as an update rather than a commitment.

Is this only relevant for large projects?

Small projects benefit too, but the overhead should scale down. A two-person team can work from a simple list of completion dates and a five-minute weekly review. The technique matters less than the habit of comparing what you predicted with what happened.

What if the team is split across two vendors?

Agree on one measurement system and one definition of done before work begins. Joint retrospectives every month help, even when they feel ceremonial. Most multi-vendor failures I have seen were integration problems in reporting, not engineering problems.

Conclusion

Software delivery is uncertain work. Pretending otherwise is expensive, and the bill arrives as missed launches, damaged trust, and teams that quietly stop believing their own plans. Forecasting replaces that theatre with something simpler: measure what happened, sample from it, and communicate in ranges.

The logistics team from the start of this article shipped, eventually. Their next programme used a probabilistic forecast, an explicit risk buffer, and demos every two weeks. It landed within one week of the plan, and nobody spent month five in an interrogation call. Same engineers. Same client. Different measurement.

If you want a place to start this week, pick one metric from the table above and report it honestly for a month. The rest of the system tends to build itself once the data is real. For teams that want to see how disciplined delivery looks in practice, the way products are shipped at pagii.co is a useful reference point.

Leave a Reply