Back to Blog

The Sprint Plan That Acts Like a Contract: How to Treat Estimates as Informed Commitments

EngineeringLeadership

Last updated: August 8, 2026

By Kris Drouet, Engineering Executive, in partnership with KORE1

An estimate becomes an informed commitment when it names its assumptions, carries a range instead of one date, and comes with a standing right to renegotiate scope the moment an assumption breaks. An aspirational guess carries none of that. It carries a date and a hope. It collapses the first week reality disagrees with it.

A contract binds two parties. That is the entire idea of one.

Most sprint plans bind one party. The team signs. The team gets measured. Then somewhere around Wednesday of the first week an executive adds a request, or a customer escalation pulls two engineers for three days, or the requirement quietly changes in a conversation nobody on the team attended, and none of that counts as a breach of anything. The estimate stayed a commitment right up until the other side needed it not to be.

That asymmetry is the whole problem. Not the estimation technique. Not the tooling. Not whether you size in points or days, which is the argument everyone would rather be having because it is the one with a clean answer.

The signatures matter.

Twenty-five years of my career have gone into regulated software, most of it mortgage lending, and I have watched more delivery plans fail on the second signature than on the first. The thing I keep coming back to is what I wrote about operational discipline and why it is not bureaucracy. A team needs a consistent operating rhythm. Stand-ups that stay focused, retrospectives that produce actionable outcomes, and planning sessions where estimates are treated as informed commitments rather than aspirational guesses.

That last clause is the one people ask me to explain. So here is the long version.

Three engineers discussing a sprint estimate as an informed commitment in a modern office

The Advice on the Internet Is Right and Useless

Search for anything about sprint estimation and you will find the same answer, repeated with slight variation across a decade of blog posts. An estimate is not a commitment. Keep the two separate. Educate your stakeholders.

Every word of that is technically correct.

It also loses every argument it enters, because it hands leadership nothing. A CFO planning a fiscal year, a board tracking a product launch, a customer with a contractual go-live date, none of them can plan against “we prefer not to commit.” So the number gets taken anyway. It gets taken out of a Jira field, or out of a Slack message, or out of something an engineer said in passing four weeks earlier that has since hardened into a date on a slide. The only thing the advice actually accomplishes is that the team stops being present for the conversation where its own number turns into a promise. So they lose twice.

Then I reread the source document, first time in maybe four years. The standard everybody cites already contains the contract.

The 2020 Scrum Guide attaches a commitment to each artifact. For the Sprint Backlog, that commitment is the Sprint Goal. Not the list of tickets. The goal. Harder work than anyone expected gets its own clause, too. Scope “may be clarified and renegotiated with the Product Owner as more is learned,” and the developers “collaborate with the Product Owner to negotiate the scope of the Sprint Backlog within the Sprint without affecting the Sprint Goal.”

Read that twice. The commitment is to an outcome. The item list is the current plan for reaching it, and the plan is explicitly renegotiable inside the sprint. That is a contract with a change-order clause written directly into it.

Almost every organization I walk into has inverted it. The inversion is the bug. They commit to the eleven tickets and treat the goal as decoration on the top of the board. Then they are shocked when the eleven tickets turn out to be a worse predictor of value delivered than the goal nobody read.

Aspirational Guess or Informed Commitment

The two things look identical in a planning tool. Same field, same number, same little colored badge. They are not the same object, and you can tell them apart in about ninety seconds by asking what each one carries.

QuestionAspirational guessInformed commitment
What does it name?A dateA goal, a range, and the assumptions holding both up
What is the unit?One numberA band, widest on the work the team understands least
Who produced it?Whoever was in the room when it was asked forThe people who will do the work
What happens when it slips?Someone explains themselvesScope gets renegotiated against the goal, in writing
Who may change it mid-sprint?Anyone senior enoughThe team and the product owner, together

Run your last planning session through that right-hand column. Count honestly. Most teams score two out of five, and the two they score are the easy ones.

The assumption line is where I focus, because it is the cheapest to add and the one almost nobody writes down. An engineer who says “eight days, assuming the vendor sandbox is available by Tuesday and the schema does not change” has given me something I can act on as a leader. I can go make the sandbox appear. I can freeze the schema. If I do neither, and it comes in at fourteen days, that is a result I purchased. Nobody has to relitigate it in a retro three weeks later, and the engineer does not have to defend a number that was never the thing that broke.

An engineer who says “eight days” has given me a wish with a decimal point. That is all.

The Tail Is the Risk, Not the Average

Leaders who resist ranges usually resist them for a reasonable-sounding reason. A range feels like hedging. A single number feels like accountability. I understand the instinct and I think it is backwards, and the data on this is not particularly close.

Bent Flyvbjerg and Alexander Budzier studied 1,471 IT projects and published the result in Harvard Business Review in 2011. The average cost overrun was 27%. That number is the one everybody quotes, and it is the least interesting number in the study, because a 27% average is survivable and it makes the whole thing sound like a rounding problem. It is not.

The finding that matters is the shape of the distribution. One in six of those projects was what the authors call a black swan, with an average cost overrun of 200% and a schedule overrun of nearly 70%. One in six. The risk is not that your estimates are a little optimistic on average. The risk is that roughly every sixth significant thing you commit to has the potential to run three times its budget, and the average buries that completely.

A single-point estimate is a claim about the average. A range is a claim about the distribution. Only one of those is honest about a body of work that behaves like this, and the one that feels like hedging is the one telling the truth.

Which is why the range is not the deliverable either. The deliverable is knowing where in the plan the tail is hiding. In practice it is nearly always the same three places. Work that crosses a system boundary the team does not own. Work that depends on a vendor’s timeline. Work that has never been done before by anyone currently on the payroll.

Everything else, you can estimate. Those three, you can only bound and monitor. Nothing else works. Say so out loud during planning and put the widest band on them, because the alternative is that they absorb the entire buffer that the rest of the sprint was quietly relying on, and then the team gets to explain why “simple CRUD work” ran late.

Engineering leader and product partner renegotiating sprint scope in a hallway conversation

The Half of the Contract Leadership Breaks

Now the uncomfortable side, and the reason I write about this at all.

Every engineering leader I know can list the ways teams break estimates. Optimism. Not reading the ticket. Forgetting about code review, or the deploy, or the two days of production support that happen every sprint whether or not anyone plans for them. That list is real, and the industry has spent twenty years building process around it.

Almost nobody keeps the equivalent list for their own side. So here is mine, the four I have personally committed at least once. All four.

Adding work after the sprint starts without naming what comes out. This is the common one and it is the most expensive, because it does not just cost the hours of the new work. It costs the team’s belief that the plan means anything, which is a much slower thing to rebuild than a sprint.

Leaving an assumption unresolved after the team flagged it. They told you the vendor sandbox was the risk. You nodded. Nobody chased it. Two weeks later the miss gets attributed to the estimate.

Reassigning people mid-sprint and treating capacity as unchanged. Pulling one engineer off a five-person team for an escalation does not cost twenty percent. It costs twenty percent plus the context, plus whatever the remaining four now have to pick up cold. Call it thirty-five.

Then there is the quiet one. Asking for a number under time pressure and then treating the answer as considered. “Ballpark it, I just need something for the deck.” Everyone in that exchange knows what is happening, and everyone participates anyway. Including me.

The research here is blunt about what this costs. Google’s DORA program, in the Accelerate State of DevOps Report 2024, found that unstable organizational priorities “cause meaningful decreases in productivity and substantial increases in burnout.” Then comes the line I keep quoting back to executives. That negative impact is “highly resistant to mitigation and persists even in environments with strong leaders and high-quality documentation.”

Sit with that for a second. Your good manager does not cancel it out. Your excellent documentation does not cancel it out. If priorities keep moving, the productivity and the people both degrade, and the usual remedies do not reach it. Show me the data on any other single leadership behavior that is that hard to compensate for. I have looked.

Priority stability is not a nice-to-have that supports the estimation process. It is the second signature on the contract. Without it there is no contract, and everything the team does to estimate better is theater performed for an audience that has already left.

How I Run the Planning Session

None of this requires a new framework. It requires about twenty extra minutes and a willingness to say some things out loud that most rooms prefer to leave implicit.

I start with the goal, not the backlog. One sentence, written where everyone can see it, describing what will be true at the end of the sprint that is not true now. If the room cannot produce that sentence in five minutes, the sprint is not ready to plan and the rest of the meeting is decoration. That failure is worth catching. It usually means priorities were never actually settled, which is a problem that lives upstream of the planning session, in whether priorities are written down at all.

Then the work gets sized by the people doing it, in a range, out loud. Not in a form. Out loud, so that the third engineer can hear the second one say “four to six days” and respond with “only if the migration is already run, and it is not.”

Assumptions get captured as they surface, in the ticket, in plain language. This is the twenty minutes. It is also the only part of the meeting that produces an artifact anyone will use later.

Next comes the part I had to learn the hard way, which is asking who is allowed to change this once it starts. Not rhetorically. By name. If a director can add a ticket without talking to the product owner, then the plan is advisory and everyone should know that going in. That question belongs to a broader problem I laid out in the decision-rights map that stops hallway decisions, and estimation is where its absence gets expensive fastest.

Finally I say the trade out loud, in the room, in front of the people who will be affected by it. The team commits to the goal and to keeping the range honest. I commit to protecting the sprint boundary and to resolving the assumptions they just flagged. If I am not willing to say my half in front of them, I have no business holding them to theirs.

Twenty minutes. That is the whole intervention. I have never had a team push back on it, and I have had plenty of executives push back on it, which tells you where the friction actually lives.

When the Sprint Misses Anyway

It will. Not every sprint, but often enough that how you handle the miss determines whether any of this survives past the first quarter.

There are only two useful questions, and neither one is “why are we late.”

The first is which assumption broke. Sometimes the answer is that none of them did and the work was simply larger than it looked, which is a real answer and a much rarer one than teams claim. Usually a named assumption failed, and it failed in a way that was visible on day three to somebody who did not think it was their job to raise it. Day three. Not day nine.

The second question is whether the range was honest at the time it was given. Note the phrasing. Not whether it was right. Ranges are not right or wrong. They are calibrated or not, and you can only judge calibration across ten sprints rather than one. A team whose eight-to-twelve-day items land at fourteen every single time does not have an honesty problem. It has a systematic bias, which is fixable arithmetic. A team that lands inside the range four times out of five is doing this correctly, and the fifth one is not a failure. It is the distribution.

What I refuse to do is treat a miss as a performance event. The moment a missed estimate has a personal cost attached, every future estimate quietly inflates to protect the person giving it, and you have traded a measurable problem for an invisible one. You will never see it happen. You will just notice, two quarters later, that everything takes twice as long as it used to and nobody can explain why.

The honest version of this is that most engineering orgs are not slow. They are carrying a tax they installed themselves, and the tax compounds. I wrote a whole diagnostic about telling a genuinely slow org apart from a tired one, and inflated estimation is one of the tells that shows up in both.

Engineering team in a retrospective reviewing which assumption broke on a missed sprint estimate

Sometimes the Estimate Is Fine and the Bench Is Not

I want to be honest about the limit of the process argument, because I have watched leaders install every piece of this and still miss quarter after quarter.

An estimate is a statement about what a specific group of people can do. Change the people and you change the estimate, and no amount of ceremony fixes a team that is three senior engineers short in a stack where the tail risk lives. If the same category of work keeps blowing its range, and the assumptions were resolved, and the sprint boundary held, then the problem is not the planning session. You are looking at a capability gap that is being described in the language of estimation because that is where it shows up on a chart.

That was the mistake I made longest. For years. I would run better retros at a problem that needed a hire.

The tell is specific. Estimation gets unreliable in exactly one area of the system, the one where a single person is the only one who really understands what is under the hood. Everything they touch is accurate. Everything anybody else touches in that area runs long by a factor of two. That is not an estimation problem. That is a bus-factor problem wearing an estimation costume, and the fix is another person who can hold that domain.

KORE1 has been doing this since 2005. Their footprint runs past 30 U.S. metros, first qualified submittals land around day 17, and 92 percent of their placements clear the twelve-month mark. The recruiters there average fifteen years inside engineering staffing, which matters mostly because the person screening for your platform gap can tell the difference between someone who has run a migration and someone who has read about one. Most of these gaps close through direct hire, since the knowledge you are buying needs to stay. Some of them are genuinely temporary, and software engineering staffing on a contract basis covers those without a permanent headcount fight.

What Gets Argued With in the Room

Doesn’t calling it a commitment just teach the team to pad?

Padding is a response to punishment, not to commitment. Teams inflate when a miss costs them something and an accurate estimate earns them nothing, and no estimation training reaches that incentive.

If your team pads, the useful question is what happened the last three times somebody was honest about a range and the honest range was inconvenient. Somebody usually remembers. In one organization I joined, engineers had learned that any estimate above two weeks triggered a meeting with a VP who would “help them find a faster path,” which everyone correctly understood as pressure. So estimates capped at two weeks. Every single one. The number was worthless and the calendar was a work of fiction, and it was a completely rational adaptation to the environment they were in.

We have no historical data. What do the first estimates anchor on?

Reconstruct the last six things you shipped and how long each one actually took, start to finish. Your ticket history holds enough to do this in an afternoon, and six data points beat zero by more than you would expect.

Measure from when work genuinely started to when it reached production, not from ticket creation, and include the review and deploy time that estimates always forget. You are not building a model. You are building a reference set, so that the next time someone says “that is about like the payments reconciliation work,” there is a real number attached to that comparison instead of a feeling. Do that for a quarter and you will have something better than most teams running formal estimation practices. Start there.

The estimate was honest and the sprint missed anyway. Now what?

Then the contract worked. An honest miss with a named broken assumption is a functioning system doing its job, and the only thing that should change is the range on the next item that looks like it.

This is the answer people find hardest to accept, and I understand why, because it sounds like an excuse for missing. It is not. The purpose of estimating is not to be right. It is to make the organization’s plans progressively less wrong, and a system that produces honest misses does that, while a system that produces comfortable hits by inflating everything does not. One of them gets more accurate over ten sprints. The other one gets slower and calls it predictability.

We run Kanban with no sprints. Does any of this transfer?

All of it except the calendar. Swap the sprint boundary for a service-level expectation on cycle time, and both obligations survive intact.

The team’s half becomes a commitment to a cycle-time distribution, usually stated as something like “eighty-five percent of items of this class finish within nine days.” Leadership’s half becomes work-in-progress limits that are actually respected and a queue that does not get reordered every morning. Honestly the contract is cleaner in a flow system, because the distribution is explicit from the start and nobody has to pretend a single number was ever the point.

Product keeps injecting work mid-sprint. How do I stop that without becoming the fun police?

Do not block the request. Price it. Every mid-sprint addition arrives with a written note of what it displaces, and the person asking picks which item drops.

You will lose this argument if you frame it as protecting your team from interruption, because that sounds like preference and theirs sounds like urgency. Framed as a trade, it stops being about who wins. Half of the requests evaporate at that point, which tells you they were never worth a sprint’s worth of disruption. Half. Consistently. The ones that survive are usually right, and now they are in the plan properly instead of sitting on top of it. Keep the notes. Six of them in a quarter is a conversation about roadmap discipline that you will be very glad to have with receipts.

Sign Your Half First

The reason I care about this is not that estimation is interesting. It is not. It is arithmetic and honesty, and every genuinely useful idea in it was published before 2010. Mostly earlier.

I care because the estimate is where a team finds out whether the plan is real. Everything a leader says about trust and ownership and autonomy gets tested at the exact moment somebody has to decide whether to give a number they actually believe or the number they think will make the meeting end faster. If the environment punishes the first one, you will get the second one forever, and you will never be told that is what happened.

So sign your half first. Say out loud what you are committing to before you ask the team what they are committing to, and then hold that line for one full sprint. It is a small thing that takes about ninety seconds in a planning meeting. It changes what the number means.

Wrestling with a delivery plan nobody believes, or trying to explain to a board why the dates keep moving? I have sat on both ends of that meeting, and neither end is a good seat. Message me. Kris Drouet on LinkedIn.

And if the honest diagnosis is that your ranges are wide because the team is missing a person rather than a process, that is a search, not a retro. Talk to a KORE1 recruiter about what the gap actually is.

Leave a Comment