Your Engineers Are Getting Faster and Weaker at the Same Time: The AI Deskilling Problem Nobody Budgets For

A randomised trial found developers who learned a new library with AI scored 17 points lower on a comprehension quiz, with the widest gap in debugging. Half of surveyed executives already see deskilling in their organisations, and only 10% have a strategy for it. Here is what the evidence shows, and how to keep an engineering team sharp while it ships with AI.

LeadershipYour Engineers Are Getting Faster and Weaker at the Same Time: The AI Deskilling Problem Nobody Budgets For

The Productivity Story Has a Second Column

For two years, the business case for AI coding tools has been written in one column: output. More pull requests, faster prototypes, shorter cycle times. Adoption followed. In the 2025 Stack Overflow Developer Survey, 84% of developers said they use or plan to use AI tools, while more developers actively distrusted the accuracy of those tools (46%) than trusted it (33%).

That gap between use and trust is the starting point for a problem most engineering budgets do not model: the people using these tools every day may be losing the very skills they need to catch the tools' mistakes. The idea used to sit in opinion pieces and conference talks. In 2026 it moved into controlled experiments, executive surveys and, outside software, peer-reviewed medical research.

The conversation reached practitioners this week too. At Craft Conference, software engineer Chelsea Troy argued that learning and execution are different modes of work, and that real skill development means you "have to sit with something that we feel bad at". AI tools are designed to remove exactly that discomfort. That is what makes them useful, and it is also what makes them risky for a team that still has to understand its own systems.

This article is not an argument against AI-assisted development. Stepto's engineers use these tools every day. It is an argument that a second column belongs next to output on every engineering scorecard: is the team's ability to reason about its own software growing, stable or shrinking?

What the Anthropic Trial Actually Found

The most-cited piece of evidence comes from an AI vendor, which makes it harder to dismiss. In a randomised controlled trial, Anthropic asked 52 mostly junior engineers to learn Trio, a Python library for asynchronous programming that none of them had used before. Half had an AI assistant available; half coded by hand. Afterwards, everyone took a quiz on the concepts they had just used.

The AI group averaged 50% on the quiz, against 67% for the hand-coding group, a 17-point gap that the researchers described as nearly two letter grades. The AI group did finish about two minutes faster, but that difference was not statistically significant. In other words, participants gave up a large amount of understanding for a speed gain the study could not reliably detect.

The detail that matters most for engineering leaders is where the gap was widest. According to the same study, the largest difference between the groups was on debugging questions. Debugging is the skill you need when generated code is almost right, when a production incident does not match any pattern, or when an agent confidently fixes the wrong thing. It is also the skill a team practises least when the assistant is doing the fixing.

Anthropic's own recommendation was direct: managers should consider systems or intentional design choices that ensure engineers continue to learn, because cognitive effort, and even getting painfully stuck, is likely important for building mastery.

Key Takeaways

  • AI-assisted group scored 50% on the comprehension quiz versus 67% for developers who coded by hand
  • The speed advantage of about two minutes was not statistically significant
  • The widest gap was on debugging, the skill most needed when AI output is wrong
  • The study was run by an AI vendor, which recommends designing work so engineers keep learning

It Is Not How Much AI You Use, It Is How You Use It

The most useful part of the Anthropic study is not the headline number. It is the breakdown of how individual participants worked. The researchers grouped people by interaction pattern, and the patterns separated cleanly into high and low scorers.

Low scorers fell into three patterns: wholesale delegation, where the AI wrote the code; progressive reliance, where people started by asking questions and ended up handing over all the writing; and iterative AI debugging, where participants relied on the assistant to debug or verify their code. High scorers also used AI, but differently. Some generated code and then asked follow-up questions to understand it. Some asked for code together with explanations. The largest high-scoring group asked only conceptual questions and resolved errors themselves.

This is the practical takeaway. Two engineers with the same tool licence and the same throughput can be on opposite trajectories. One is using AI as a tutor and accelerator and getting more capable each sprint. The other is using it as a substitute and is slowly losing the ability to work without it. Velocity dashboards cannot tell them apart.

The pattern matches what Microsoft Research and Carnegie Mellon found in a survey of 319 knowledge workers: higher confidence in generative AI was associated with less critical thinking, while higher self-confidence in the task was associated with more. Critical thinking did not disappear, but it shifted toward verification, integration and stewardship of AI output. That shift is only safe if people can still tell good output from bad.

The Clearest Warning Came From Medicine

Software teams rarely get clean before-and-after data on human skill. Medicine sometimes does, and one study has become the reference case for AI deskilling. In August 2025, The Lancet Gastroenterology & Hepatology published research from four colonoscopy centres in Poland that tracked 19 experienced endoscopists before and after AI-assisted polyp detection was introduced.

When those doctors performed colonoscopies without AI after months of using it, their adenoma detection rate fell from 28.4% to 22.4%, a 20% relative reduction. With the AI switched on, the rate was 25.3%. The tool was propping up performance while the underlying human skill drained away. As one of the authors, Dr Marcin Romańczyk, put it, this was the first study to suggest a negative impact of regular AI use on professionals' ability to complete a patient-relevant task.

Aviation learned the same lesson a decade earlier. In 2013 the US Federal Aviation Administration issued Safety Alert for Operators 13002 after identifying an increase in manual handling errors, warning that continuous use of autoflight systems does not reinforce manual flying skills and could degrade a pilot's ability to recover an aircraft from an undesired state. The FAA's response was not to remove the autopilot. It was to require deliberate opportunities for manual practice.

The parallel for software is uncomfortable but exact. The coding agent is the autopilot. It works well most of the time. The moment it does not, in an outage, a security incident or a subtle data-corruption bug, you need engineers who can still fly the plane by hand.

Executives See It Coming. Almost Nobody Has a Plan.

In June 2026, BCG published a survey of 70 C-suite leaders and senior executives on what it calls distributed de-skilling: a collective erosion of human skills that undermines organisational intelligence and resilience over time. Half said they are already observing de-skilling in their organisations, and more than 60% expect it to pose a material threat within three to five years.

The symptoms they reported read like a description of an engineering organisation under AI pressure. In the same BCG survey, 53% of leaders reported slower development of junior talent, 49% saw less diversity of thinking, 43% cited fewer constructive debates and 33% noted a decline in mentoring and knowledge sharing. Almost 90% pointed to overreliance on AI outputs without stress testing as the starting point.

The most telling number is the response. BCG found that only 10% of companies have an organisation-wide strategy for de-skilling. The skills leaders rated most at risk, judgement and decision-making, problem framing, causal reasoning and solution evaluation, are the same skills that separate a senior engineer from a fast typist. Nine in ten companies are leaving them to chance.

BCG frames this as a system design problem rather than a talent problem, and that framing is right. Individual engineers cannot be expected to resist a tool that their sprint commitments are built around. If the workflow rewards delegation and never tests understanding, delegation is what you will get.

Key Takeaways

  • 50% of surveyed executives already observe deskilling; more than 60% expect it to become a material threat
  • 53% report slower development of junior talent and 33% see less mentoring
  • Almost 90% trace the problem to unchallenged reliance on AI output
  • Only 10% of companies have an organisation-wide strategy to address it

Where Deskilling Shows Up in a Software Organisation

Deskilling rarely appears as a single failure. It shows up as a set of small signals that are easy to explain away. Incident response takes longer because fewer people can reason from symptoms to cause without asking an assistant first. Code reviews get shallower because reviewers approve what looks plausible rather than what they have traced. Architecture discussions converge quickly on whatever the model suggested, because nobody wants to argue from first principles against a confident answer.

Juniors are the most exposed. The Anthropic trial was run mostly on junior engineers, and the BCG survey found slower junior development to be the most common symptom. A junior who has never debugged a race condition without help will not become the senior who can. That is a pipeline problem with a delay of several years, which is exactly why it is easy to ignore in quarterly planning. We have written before about the junior developer pipeline and about comprehension debt in AI-generated codebases. Deskilling is the human side of the same ledger: the codebase gets harder to understand while the team gets less practised at understanding it.

Seniors are not immune. Senior engineers often use AI most effectively, but they also delegate the routine work that used to keep their knowledge of a codebase current. Over a year, the person who once knew every corner of the payment service becomes the person who knows how to prompt about it. That is fine until the day the prompt stops working.

There is also a vendor dimension. If you work with an outsourcing partner whose business model is maximum output per billed hour, deskilling is not their risk. It is yours. Their engineers can generate code that nobody on either side truly understands, and the bill for that arrives later, during an incident, a migration or a handover.

How to Keep a Team Sharp While It Ships With AI

The answer is not to ban AI tools. That would give up real productivity and push usage into the shadows. The answer is to treat understanding as a deliverable, the way the FAA treats manual flying proficiency, and to build it into normal work rather than into an annual training budget.

Start with review. Require that the author of every pull request, whether a human wrote it or an agent did, can explain the change without the assistant open: why this approach, what it breaks if it is wrong, how it was tested. A short walkthrough in review takes minutes and is the cheapest deskilling control available. It also shifts AI use toward the high-scoring patterns from the Anthropic study, generation followed by comprehension.

Next, protect deliberate practice. Rotate engineers through incident drills and debugging sessions where assistants are switched off for the diagnosis phase. Assign juniors tasks where the learning is the point and the deadline is soft. Run architecture decisions as written arguments with alternatives, not as accepted model suggestions. None of this is expensive. All of it disappears unless someone owns it.

Finally, measure it. Add a small number of capability signals to your engineering metrics alongside throughput: time to diagnose incidents, how many engineers can independently own each critical service, how often reviews catch substantive issues. If output rises while those signals fall, you are borrowing against the team's future.

AI usage patterns and what they do to engineering capability
Usage patternShort-term effectLong-term effect on skillsWhat to do
Full delegation: AI writes and fixes the codeFast outputWeakest understanding and debugging abilityRequire authors to explain every change in review
AI debugging loop: paste error, accept fixProblems close quicklyDiagnostic skill is not practisedRun assistant-off diagnosis in incident drills
Generate, then interrogate the outputSlightly slowerUnderstanding keeps pace with outputMake this the team default
Conceptual questions, human writes and fixesSlower on unfamiliar workStrongest skill growthUse for onboarding and junior development

Key Takeaways

  • Make explaining the change, without the assistant, a condition of every merge
  • Schedule assistant-off debugging and incident drills, the software equivalent of manual flying practice
  • Give juniors learning-first tasks with soft deadlines
  • Track capability signals such as diagnosis time and service ownership alongside velocity

What to Ask of a Development Partner

If part of your engineering capacity comes from an external team, deskilling becomes a procurement question. A partner can deliver impressive velocity for six months while building a codebase that neither its engineers nor yours can reason about. The warning signs are vague answers in technical reviews, reliance on the assistant during live debugging, and high rotation of engineers who never stay long enough to own anything.

Ask direct questions. How do your engineers use AI tools, and how do you check that they understand what they ship? Who on your team can explain the architecture of our system without notes? What happens to knowledge when an engineer rotates off? How are your juniors developed, and who reviews their work? A serious partner will have specific answers. A partner selling only output will talk about speed.

This is where the dedicated development team model has a structural advantage. Stable, named engineers who stay on your product build deep knowledge of it over time, and that knowledge compounds instead of resetting with every new ticket. At Stepto, our engineers in Serbia use AI tools as accelerators, but they own their code: they review it, explain it and debug it, and our seniors mentor juniors on real client work rather than leaving them to the assistant. Working hours that overlap with European and US teams mean those walkthroughs, pairing sessions and incident drills happen live, not through overnight tickets.

Nearshore development also helps with the pipeline problem. If your in-house team is too small or too stretched to grow juniors properly, a nearshore partner with a real seniority structure gives you experienced engineers now and a mentoring culture that keeps producing them, without asking your own leads to run a training programme on top of delivery.

Treat Understanding as Part of the Delivery

AI coding tools are not going away, and they should not. But the evidence from 2026 is consistent: when people hand over the thinking along with the typing, the skills they need most when the tools fail start to fade, and the decline is invisible on a velocity chart. Medicine and aviation have already seen what happens when that goes unmanaged. Software teams can learn from them cheaply now or expensively later. Make understanding a condition of merging, keep debugging and design skills in regular practice, measure capability as well as output, and choose partners whose engineers own what they ship. If you want a team that uses AI to move faster without losing the judgement your systems depend on, talk to Stepto about a dedicated development team.

Building a team in Eastern Europe?

StepTo helps European and US companies build senior-led nearshore engineering teams in Serbia. Let's talk about what your next engagement could look like.

Start a conversation
I

Written by

Igor Gazivoda

Founder & CEO · StepTo

Igor has 15+ years in software engineering and business development. He specializes in scaling engineering teams, nearshore strategy, and AI-driven product development. He holds a Master's in Computer Science from the University of Belgrade.

LinkedIn →
Performance-led engineering

Want senior engineers who move work forward, not just tickets?

Work with accountable, English-fluent professionals who communicate clearly, protect quality, and deliver with a steady operating rhythm. Cost efficiency matters, but performance is why clients stay with us.

Delivery signals · senior engineering team
Senior ownership
Lead-level
Delivery rhythm
Weekly
Timezone overlap
CET
1 teamaccountable for outcomes, communication, and execution