The Code Took a Morning. The Story Took a Week. How AI Coding Agents Broke Software Estimation

Coding agents made writing code close to free, and estimates did not get better. They got worse. Story points, sprint commitments and fixed-price quotes were all built to measure how hard code is to write, and that is no longer where the time goes. Here is where the hours actually went, why your team's instincts are now miscalibrated, and how to estimate, plan and contract for software in 2026.

EngineeringThe Code Took a Morning. The Story Took a Week. How AI Coding Agents Broke Software Estimation

The Story That Took a Morning and Then Four More Days

Every engineering team that adopted coding agents has a version of this meeting. In a widely shared post, Sujit Vishnu Kamthe of Sahaj Software describes a planning session in which the team argued over whether a story was a 5 or an 8, based on how tricky the algorithm looked. Once the story was picked up, Claude produced a working implementation before lunch. The story then took four more days to finish, spent on review, testing across services and security sign-off. The estimate was not wrong because the team was careless. It was measuring the one part of the work that had stopped costing anything.

The same pattern appears in aggregate data. The Faros AI Productivity Paradox report, built on telemetry from more than 10,000 developers across 1,255 teams, found that developers on high-adoption teams completed 21% more tasks and merged 98% more pull requests. Average PR size rose 154% and PR review time rose 91%. Faros found no significant correlation between AI adoption and improvement at company level. Individual output went up, but the organisation did not ship noticeably faster.

The explanation is mostly arithmetic. In its 2025 State of Developer Experience research, Atlassian reported that coding makes up roughly 16% of a developer's working week. Over two-thirds of developers said AI saved them at least 10 hours a week, and half said they still lost more than 10 hours a week to organisational friction. If you speed up the 16% dramatically and leave the other 84% alone, the total barely moves. An estimation system that mostly prices the 16% will be wrong more often than before, not less.

This is now one of the most argued-over topics in engineering management, on Hacker News, on LinkedIn and in sprint retrospectives. Some teams have dropped story points altogether. Others still play planning poker and quietly ignore the result. Neither approach answers the question clients and CFOs keep asking: when will it be done, and what will it cost?

What an Estimate Was Actually For

Before deciding what to replace, it helps to be honest about what estimates did. Derek Jones, who has spent years analysing estimation datasets, points out that with roughly two-thirds of estimates landing within a factor of four of the actual effort, accuracy was never the main thing estimation delivered. Its real value was as a planning tool. Estimating forced a team to break a large piece of work into small, well-defined chunks, and to discover what it did not understand before it started.

That function has not gone away. It has moved. Jones's argument is that coding agents make implementation cheap but follow the specification they are given, and large, complicated programs need large, complicated specifications that take considerable human effort to write. Estimation moves upstream, to the work of deciding precisely what to build. He also notes that agents separate testing from implementation as visible, distinct work, where it used to be folded into the same estimate.

This matters for anyone buying software. If you treat estimation as a way to predict typing time, AI makes it look obsolete. If you treat it as the discipline of decomposing work and exposing unknowns, it is more important than ever, because an agent will build whatever ambiguity you hand it at high speed, and someone will pay to review, test and unwind it afterwards. We made a similar argument about spec-driven development in outsourced teams. The specification is now the most important cost driver in the project.

Where the Hours Went: Review, Testing and Waiting

The newest data shows the cost moving downstream in detail. The Faros AI Engineering Report 2026, covering 22,000 developers across 4,000 teams, found that under high AI adoption median time in PR review rose 441.5%, average PR size rose 51.3%, bugs per PR rose 54%, and 31.3% more PRs were merged with no review at all. Tasks with code completed rose 210%, but average time in progress also rose 225.2%. More work is being started and coded. It is not getting finished any faster.

Google's research programme sees the same tension from a different angle. The 2025 DORA report found that 90% of technology professionals now use AI at work and that AI adoption is now associated with higher delivery throughput, a reversal of the previous year's result. It is also associated with higher delivery instability. Teams ship more changes, and more of those changes fail. Without strong automated testing and fast feedback loops, extra change volume turns into rework.

For estimation, this means the effort profile of a story has flipped. A ticket used to cost roughly in proportion to how hard the code was to write. Now it costs in proportion to how much a human has to read, how many services it touches, how much testing it needs and how long it waits in someone else's queue. Kamthe's supersonic jet analogy captures it well: replace the plane on a one-hour flight with something twice as fast and the door-to-door journey still takes most of the day, because the flight was never the bottleneck.

We looked at the capacity side of this in our piece on the verification gap. The estimation consequence is simpler. If your sizing conversation does not talk about review load, test burden and integration risk, it is sizing the wrong thing.

Key Takeaways

  • Coding is now the cheapest part of most stories; review, testing and integration are the expensive part
  • Larger AI-generated pull requests wait longer in review, and more of them merge without one
  • More work gets started and coded, but time in progress grows, so delivery dates do not move forward
  • Higher throughput comes with higher instability unless testing and feedback loops keep up

Your Team's Gut Feel Is Now Miscalibrated

Most estimation techniques, from planning poker to T-shirt sizing, rely on experienced people's instincts. That worked because instincts were trained on years of feedback about how long code took to write. AI has broken that feedback loop, and the evidence that it has is uncomfortable.

In METR's randomised controlled trial of 16 experienced open-source developers working on 246 real issues in repositories they knew well, developers took 19% longer when allowed to use AI tools. Beforehand they expected a 24% speedup. Afterwards, having just been slowed down, they still believed AI had made them about 20% faster. METR's 2026 follow-up with 57 developers produced murkier numbers and METR itself describes them as an unreliable signal, partly because many developers now refuse to work without AI at all. The point for estimation is not whether today's tools speed people up. They probably do, more than in early 2025. The point is that people cannot reliably tell by how much.

Agents misjudge their own work as well. In a 2026 study of token consumption in agentic coding tasks, models asked to predict their own token usage before starting reached correlations of at most 0.39 with what they actually used, and consistently underestimated. Expert-rated task difficulty correlated with actual token cost at only 0.32. What looks hard to a senior engineer and what turns out to be expensive for an agent are not the same thing.

Put those findings together and the risk is clear. A team that estimates by consensus will tend to assume AI has made everything faster, will undercount verification, and will be surprised again every sprint. The solution is not better instincts. It is measuring what actually happened and estimating from that.

A New Variable Cost in Every Ticket

There is now a second budget hidden in every story: the compute the agent burns while doing the work. The token consumption study cited above found that agentic tasks consume around 1,200 times more tokens than a multi-turn chat, and that on identical tasks the most expensive run cost roughly twice as much as the cheapest. Accuracy often peaked at intermediate cost and dropped at the highest cost levels. The expensive runs were frequently the ones where the agent was wandering, not the ones where it was working hardest.

Specification quality drives that cost directly. In a Stanford study by Jakub Smékal covering 2,700 agent runs, cutting a full task specification down to a bare user story raised token spend by 29.7% across every task tested, while run-to-run variance stayed about the same whatever the prompt. The same paper showed that one cheap probe run, costing about $0.11, could predict the token cost of other specification variants within 36%, compared with 161% error without a probe. That is a useful practical finding. A short, instrumented trial run tells you far more about cost than a planning meeting does.

Researchers have started building models around this. ACEM, a cost estimation model for agentic software engineering proposed by Mohammad El-Ramly, replaces the COCOMO-era assumptions with three cost dimensions: token consumption, human-in-the-loop oversight effort and infrastructure. It adds factors for retries and context growth, and maps existing story points and function points onto expected token use so teams can reuse old scoping data. The constants are not yet empirically calibrated, so treat it as a framework rather than a calculator, but its structure is right. Human oversight is now a separately priced component of every feature.

For most teams, token spend is still small compared with salaries. It is volatile, though, and it scales with vague requirements. We covered the budgeting side in our piece on usage-based billing for AI coding tools. For estimation, the lesson is that a poorly specified ticket now costs more twice: once in tokens, and again in human review of whatever the agent guessed.

What to Estimate Instead

Teams that have made estimation work again in 2026 have not found a magic replacement for story points. They have changed what they ask during estimation. Kamthe's framework, which he calls residual attention estimation, is a good template because it keeps the existing tooling and changes only the substance.

First, a readiness gate. A story gets no estimate until it has testable, agreed acceptance criteria and a single agreed implementation approach. Ambiguous stories go back to refinement or become time-boxed spikes. This is where the upstream specification work gets done, and it is the cheapest point at which to catch scope problems.

Second, size by human attention rather than code complexity. The drivers are review load, the system context a reviewer must hold in their head, blast radius across repositories and teams, test burden, coordination with other teams or release windows, and how hard the change is to reverse. Kamthe collapses the Fibonacci scale into three sizes, small, medium and large, and refuses to estimate anything larger. Large work gets sliced.

Third, change the capacity math and the forecast. Velocity should count stories that are verified done, not stories with code written, and teams should cap the number of items awaiting review, because review is now the constraint. Forecasts should be ranges with confidence levels, such as 50% confidence by one date and 85% by a later one, rather than a single date that everyone knows is a guess.

AI features need one more step. As Tian Pan argues, LLM features break estimation for structural reasons: outputs are non-deterministic, the number of experiments needed is unknown, and model updates change behaviour underneath you. His recommendation is to split work into a time-boxed experimentation phase with a defined outcome and a production phase that can be estimated normally, and to require an evaluation specification before any AI ticket is sized. That connects directly to the eval gap we wrote about earlier.

How the drivers of a software estimate have shifted with coding agents
Estimation questionTraditional estimateEstimate in an agent-assisted team
What is being sizedEffort to write the codeHuman attention needed to specify, review, test and release the change
Main uncertaintyAlgorithmic difficultySpecification quality, blast radius and review queue length
Entry conditionStory is roughly understoodTestable acceptance criteria and one agreed approach
Velocity countsStories with code completeStories verified done in production
Forecast formatA single target dateA date range with stated confidence levels
AI featuresEstimated like any other featureTime-boxed experiments first, with an eval specification before sizing

Key Takeaways

  • Gate estimation on testable acceptance criteria and a single agreed approach
  • Size work by review load, blast radius, test burden, coordination and reversibility, not by code complexity
  • Count only verified-done work in velocity and cap the review queue
  • Forecast with ranges and confidence levels, and time-box experimentation for AI features

What This Means for Quotes and Contracts

Clients expect AI to make software cheaper, and in one sense it has. Implementation hours are falling. But a vendor quote built on the old estimation model is now wrong in a predictable way. It overprices typing, underprices verification and integration, and does not mention specification work at all. Whether the quote is fixed price or time and materials, that mismatch eventually lands on someone.

On fixed-price work, the risk is a quote that looks competitive because it assumes AI-speed implementation and then runs into review, testing and integration costs it never priced. The vendor recovers through change requests or by cutting verification, and you find out which one later. On time-and-materials work, the risk is paying for agent-generated code that sits in review, or that gets merged without review and comes back as defects. We set out the general trade-offs in our guide to fixed price versus time and materials. AI does not change that choice, but it does change what a credible quote has to contain.

Ask any vendor to break their estimate into specification, implementation, review and testing, and integration and release. If implementation dominates, the quote reflects how projects used to run. Ask how they measure velocity: code written or work verified in production. Ask what happens to the estimate when a requirement turns out to be ambiguous, and whether a readiness gate exists. And if you are tempted to demand a flat AI discount, read our analysis of pricing outsourcing in the AI era first. The savings are real, but they are not where most people look.

The most reliable way to get an honest number is still to buy a small amount of real delivery before committing to a large one. A discovery phase produces the specification that estimation now depends on, and a paid pilot gives you a measured delivery rate from the actual team rather than a sales estimate.

Key Takeaways

  • A credible quote separates specification, implementation, review and testing, and integration
  • Quotes dominated by implementation hours reflect a delivery model that no longer applies
  • Ask whether velocity counts code written or work verified in production
  • Buy discovery and a short pilot to get a measured delivery rate before signing a large contract

Why a Stable Dedicated Team Estimates Better

Every method described above depends on one thing: a trailing record of how a specific team actually delivers verified work in a specific codebase. That record cannot exist if the team changes every quarter, if the people writing code are not the people reviewing it, or if the vendor reports activity rather than outcomes. Estimation in the agent era is an empirical exercise, and empirical exercises need stable conditions.

That is the core of how Stepto works. A dedicated development team from our engineering base in Serbia stays with your product, builds up context in your codebase and your agent setup, and develops a delivery history you can forecast from. Our teams do the specification work before implementation, treat review and testing capacity as planned work rather than slack, and report velocity as work verified in production. When a story is ambiguous, it goes back to refinement instead of into an agent.

Working hours matter here too. Specification and review are the most conversation-heavy parts of the delivery chain, and they are now the parts that decide the schedule. Central European time overlaps fully with European clients and gives a solid daily window with US East Coast teams, so the questions that block a story get answered the same day instead of the next morning. That shortens time in review, which is now the main thing driving delivery dates.

If your sprint commitments have stopped meaning anything, or a vendor's quote no longer matches the invoices, the cause is usually not the people. It is an estimation model built for a different way of working. We are happy to look at your current planning process and delivery data and show where the time is actually going.

Estimate the Attention, Not the Typing

Coding agents did not make software estimation obsolete. They exposed what it was always for. When writing code was the slow part, sizing by code complexity was a reasonable shortcut. Now that the code arrives in minutes, the schedule is set by everything around it: how clearly the work was specified, how much a person has to read and test, how many systems it touches and how long it waits for review. Teams that keep estimating the old way will keep missing commitments in both directions and keep blaming the tools. Teams that gate work on a real specification, size it by human attention, measure verified delivery and forecast in ranges will get dates and budgets they can rely on again, and so will their clients. If you want a team that already works that way, talk to Stepto about a dedicated development team.

Building a team in Eastern Europe?

StepTo helps European and US companies build senior-led nearshore engineering teams in Serbia. Let's talk about what your next engagement could look like.

Start a conversation
I

Written by

Igor Gazivoda

Founder & CEO · StepTo

Igor has 15+ years in software engineering and business development. He specializes in scaling engineering teams, nearshore strategy, and AI-driven product development. He holds a Master's in Computer Science from the University of Belgrade.

LinkedIn →
Performance-led engineering

Want senior engineers who move work forward, not just tickets?

Work with accountable, English-fluent professionals who communicate clearly, protect quality, and deliver with a steady operating rhythm. Cost efficiency matters, but performance is why clients stay with us.

Delivery signals · senior engineering team
Senior ownership
Lead-level
Delivery rhythm
Weekly
Timezone overlap
CET
1 teamaccountable for outcomes, communication, and execution