Nobody Can Tell You How Many AI Agents Your Company Is Running. That Is the Whole Problem.

Employees are creating agents faster than IT can list them, and fewer than one in five organisations has a complete inventory. The fix is boring: a registry, an owner per agent, and a retirement rule.

AI StrategyNobody Can Tell You How Many AI Agents Your Company Is Running. That Is the Whole Problem.

What Does Agent Sprawl Actually Look Like Inside a Company?

It does not look like a security incident. It looks like a spreadsheet that will not reconcile.

In May 2026 the Wall Street Journal reported on what internal teams at several large employers had started calling agent sprawl, and the numbers in that reporting are the clearest picture anyone has published. Employees at the dialysis provider DaVita had created more than 10,000 AI agents. At FICO, roughly 3,500 employees were creating dozens of new agents every day at every tier of the organisation. At Lyft, teams built so many overlapping agents that the company ended up building an internal platform whose entire purpose was to track them (Belle Lin's reporting for the WSJ CIO Journal, summarised for non-subscribers by Futurism).

The quote from that reporting that engineering leaders should sit with belongs to Michael Friedlander, CIO for the Americas at the Magnum Ice Cream Company: "Because everybody can do it, we're probably going to end up with a lot of people having the same types of agents." That is the mechanism in one sentence. Agent creation stopped being an engineering activity somewhere in late 2025. It is now a thing a finance analyst does on a Tuesday afternoon, in a chat interface, with no ticket, no repository and no approval step, and the artefact it produces is a long-lived autonomous process with access to real systems.

The trajectory is what turns an annoyance into a governance problem. Gartner projected in April 2026 that the average global Fortune 500 enterprise will have over 150,000 agents in use by 2028, up from fewer than 15 in 2025. Treat the precise figure as a directional claim rather than a forecast you can plan capacity against, and it still says something unavoidable: the thing you are being asked to govern grows by four orders of magnitude inside three years, and the governance model you apply has to work at the end of that curve, not the beginning.

The reason this reads as a new problem rather than a familiar one is the ownership pattern. Shadow IT put unapproved software on the network. Agent sprawl puts unapproved workers on the payroll, each with credentials, a budget line and the ability to take actions that have consequences outside your systems.

Why Can Nobody Count Their Own Agents?

Because the surveys that ask are getting a consistent and uncomfortable answer: deployment is effectively universal, and the management layer underneath it is not there.

The most rigorous data point comes from the IBM Institute for Business Value, which surveyed 2,000 senior executives responsible for IT, technology or AI decisions across 33 geographies and 19 industries between January and April 2026, in cooperation with Oxford Economics. Just 18% of organisations maintain a current and complete inventory of their AI agents. The companion findings are worse in a specific way: 70% of respondents report teams deploying technology faster than IT can track it, 77% say AI adoption is outpacing their governance capabilities, and only 11% feel fully prepared for the scale of agent deployment they expect. Two thirds of the CIOs and CTOs surveyed described being held accountable for systems they do not fully control.

A second survey of nearly 1,900 global IT leaders, conducted for OutSystems between December 2025 and January 2026, lands in the same place from a different angle: 96% of organisations are already using AI agents in some capacity while only 12% have implemented a centralised platform to manage them. 94% said sprawl is compounding complexity, technical debt and security risk, and 38% reported mixing custom-built and pre-built agents, which is the architectural detail that makes a single inventory hard: half your fleet lives in a vendor's console and half lives in your own repositories, and neither knows about the other.

The vendor surveys agree even when their interests do not. An August 2026 SAP LeanIX survey found 98% of companies have deployed AI agents or plan to, while fewer than half have visibility into an inventory of them. And the coordination cost of an uncounted fleet shows up directly: Salesforce's 2026 Connectivity Benchmark reports organisations using 12 or more AI agents on average, with 50% of those agents operating in silos rather than as part of a coordinated system.

Put plainly, the gap is not between companies that govern agents well and companies that govern them badly. It is between the rate at which agents appear and the rate at which any process can register them, and no amount of policy writing closes a gap that is structurally a throughput problem.

Key Takeaways

  • Only 18% of organisations maintain a current and complete inventory of their AI agents, per a survey of 2,000 executives across 33 geographies
  • 96% are already using agents; 12% have a centralised platform to manage them
  • 94% of IT leaders say sprawl is compounding complexity, technical debt and security risk
  • 38% mix custom-built and pre-built agents, which splits the inventory across systems that do not talk to each other

What Does the Uncounted Fleet Cost You?

Three bills arrive, and only one of them is obvious.

The obvious one is tokens. Duplicate agents doing near-identical work consume inference capacity twice, and because they were created outside any provisioning flow, the spend rarely maps to a cost centre. This is the same failure we described in our analysis of AI API costs as an engineering line item, with an extra twist: you cannot attribute what you have not enumerated. The Magnum CIO's framing in the WSJ reporting was precisely this, that the agents are easy and the tokens, and therefore the cost, come later. Note that IBM's survey expects AI spend to move from roughly 15% of IT budgets in 2025 to roughly 25% by 2027. A quarter of the IT budget is not a line you leave unattributed.

The second bill is incidents, and here the data is specific enough to plan against. Surveyed organisations averaged 54 AI agent incidents over the preceding year, meaning an unintended or harmful occurrence that required human correction. 17% were high severity, taking more than four hours to contain; 37% resulted in data exposure or a security breach; 33% caused cascading system failures; and 17% triggered compliance issues. Fifty-four incidents a year is roughly one a week. That is not an emerging risk category, that is an operational cadence, and most organisations are absorbing it without a runbook because the agent that caused it was not on any list.

The third bill is the one nobody invoices: conflicting work. Two agents built by two departments against the same records, each confident, each partly right, produce a data quality problem that surfaces months later as a reconciliation failure or a customer complaint. An agent registry does not prevent that on its own, but it is the only artefact that makes the collision visible before it happens, because the duplicate shows up as two entries with the same data scope and two different owners.

The counterargument to all of this is that governance slows deployment down, and the survey data says the opposite. Organisations that embed controls into the platform rather than relying on manual governance deploy 16 times more agents, report 25% fewer incidents, and show 18% higher operating margins, while those with strong financial discipline deploy 2.4 times more agents on the same budget. Read that as a throughput argument rather than a compliance one. The registry is not a brake. It is the thing that lets you stop being afraid of the next thousand agents.

Key Takeaways

  • Surveyed organisations averaged 54 agent incidents in a year, with 17% high severity and 37% involving data exposure
  • AI spend is expected to rise from roughly 15% to roughly 25% of IT budgets between 2025 and 2027
  • Teams that embed controls in the platform deploy 16x more agents with 25% fewer incidents
  • Duplicate agents produce conflicting outputs that surface as data quality failures long after the fact

Every Agent Is an Identity You Did Not Provision

The count problem and the identity problem are the same problem, approached from two directions, and the identity side has been measured for longer.

CyberArk's machine identity research puts the ratio of machine to human identities at more than 80 to 1, and the Cloud Security Alliance's work on non-human identity governance describes an enterprise average nearer 45 to 1, rising to 144 to 1 in cloud-native environments. Those identities are service accounts, API keys, OAuth tokens, machine certificates and the credentials an agent carries. Each agent an employee creates adds at least one, usually several, and each one was issued by a flow that was designed for a different kind of principal.

The practical consequence is that your joiner-mover-leaver process, which is the single control that keeps human access from accumulating forever, has no equivalent on the agent side. A person changes teams and their access is reviewed. An agent's creator changes teams and the agent keeps running with the permissions it was granted, against the systems it was pointed at, indefinitely. We have written about the access side of this in the agent overprivilege piece and about the credential side in our work on secrets sprawl. Sprawl is what happens when neither of those controls has an inventory to operate on.

There is a third layer that turns a linear problem into a compounding one, which is agents that create agents. Once a fraction of your fleet can spawn sub-agents to decompose a task, the count stops being a function of how many employees are building things and starts being a function of how much work the fleet is doing. That is also why containment matters alongside counting, a point we made in detail in our piece on agent execution boundaries.

The registry is what makes identity governance tractable here, because it supplies the missing join key. Without it, your IAM system sees a population of non-human principals with no business context attached. With it, every credential maps to an agent, every agent maps to an owner and a purpose, and an access review becomes a question a human can answer.

Key Takeaways

  • Machine identities outnumber human ones by ratios measured between 45 to 1 and more than 80 to 1, reaching 144 to 1 in cloud-native estates
  • There is no joiner-mover-leaver equivalent for agents, so permissions granted at creation persist indefinitely
  • Agents that spawn sub-agents make the population a function of workload rather than headcount
  • A registry supplies the business context that makes non-human access reviews answerable

Why Blocking Agent Creation Makes the Problem Worse

The first instinct of a governance function confronted with 10,000 employee-built agents is to require approval for the next one. It is the wrong move, and the analyst who named the problem said so in the same breath as naming it.

Max Goss, senior director analyst at Gartner, framed it as a balance: organisations need to govern agents and manage sprawl while also safely empowering employees to innovate with these tools. The failure mode of the restrictive approach is well documented and has a name you already know. Block the sanctioned path and the work moves to personal accounts, personal API keys and consumer tools, which is exactly the dynamic we covered in our analysis of shadow AI inside the enterprise. You have not reduced the number of agents. You have reduced the number of agents you can see, which is the only variable that actually matters.

There is also a value argument that survives contact with the finance team. Most of those 10,000 agents are genuinely useful to the person who built them. They automate a reconciliation, summarise a queue, draft a response. The aggregate productivity is real, and an organisation that suppresses it to regain tidiness has made a trade it cannot defend at the next board meeting. The point of the registry is not to reduce the fleet. It is to make the fleet legible, so the small fraction that is dangerous or duplicated can be found.

The design principle that follows is that registration must be easier than evasion. If registering an agent means filing a ticket and waiting three days, your inventory will be a permanent undercount. If it means the agent simply does not get credentials until it exists in the registry, and the registry entry is created automatically at deployment, the inventory is complete by construction. That is a platform decision, not a policy one, and it is the single highest-leverage thing an engineering organisation can build in this space.

In Europe, the Inventory Is Not Optional

For any organisation operating in the EU, the registry has stopped being a good practice and started being an evidentiary requirement, and the timing is immediate rather than distant.

The obligations on deployers of high-risk AI systems under the EU AI Act apply from August 2026, and they are the kind of obligations that presuppose a list. Article 26 requires deployers to use systems in accordance with instructions, assign human oversight to competent people, monitor operation, and retain the logs under their control for at least six months unless other law requires longer. Every one of those duties is per-system. If you cannot enumerate the systems, you cannot demonstrate any of it, and the demonstration is the obligation. We walked through the engineering implications of the August deadline in our piece on the AI Act as an engineering problem.

The awkward part for a sprawling fleet is classification. The Act's obligations depend on what a system does, so an employee-built agent that touches recruitment screening, creditworthiness assessment or access to essential services can carry high-risk duties that nobody assessed, because nobody knew the agent existed. This is the failure mode that makes sprawl a legal exposure rather than an operational one: not a deliberate decision to skip a conformity step, but an inability to know which of 10,000 artefacts the step applied to.

The management-system standards are the practical bridge here, because they supply the auditable structure that the regulation assumes. An AI management system built to the ISO standard gives you the record-keeping, risk management and documentation backbone that a regulator or a customer's procurement team will ask to see, which is why it has quietly become a procurement gate in enterprise AI vendor assurance. Either way the first artefact is the same one: a complete, current list of the systems in scope, with an owner against each.

What Belongs in an Agent Registry

The good news is that this is a well-bounded engineering deliverable rather than a research problem. The registry is a small service, a schema and a set of enforcement hooks, and it can ship in a quarter.

Gartner's April 2026 guidance sets out six steps, and they are worth reading in order because the sequence matters: establish governance and policies covering who may build agents and which connectors are permitted; build a centralised inventory spanning sanctioned tools and shadow deployments; define agent identity, permissions and lifecycle, including continual review and retirement of redundant agents; create information governance controlling what data each agent can reach; monitor and remediate agent behaviour in production; and invest in training so employees build responsibly rather than around you. Steps two and three are the engineering core. The rest depends on them.

Concretely, each registry entry needs a named human owner and a named backup, because unowned agents are the ones that survive reorganisations. It needs a stated purpose in plain language, which is what makes duplicate detection possible at all. It needs the model and version in use, so that a provider deprecation becomes a query rather than an archaeology project. It needs the explicit tool and connector scopes, the data classifications it can reach, and the identity it authenticates as, so that the entry joins cleanly to your IAM system. It needs a cost centre, so the tokens land on someone's budget. And it needs an evaluation reference and an expiry date, which are the two fields almost every implementation leaves out.

Enforcement is what separates a registry from a wiki page. The rule that makes it real is that credentials are issued by the registry: an agent with no entry gets no token, no connector and no service account, and the entry is created at deployment time by the platform rather than by a human remembering to file it. Dynamic inventories that auto-update on deployment outperform manual documentation, which is an unsurprising finding stated politely. Manual inventories of fast-moving populations are always wrong, and a governance artefact everyone knows is wrong stops being consulted.

One more field earns its place in regulated environments: the classification decision, recorded with a date and a rationale. When someone asks which of your agents fall under a high-risk category, the answer should be a filter on the registry rather than a three-week discovery exercise.

Key Takeaways

  • Minimum viable schema: owner and backup, plain-language purpose, model and version, tool and data scopes, identity, cost centre, eval reference, expiry date
  • Credentials should be issued by the registry, so an unregistered agent simply cannot authenticate
  • Entries must be created automatically at deployment; manual inventories of fast-moving populations are wrong by default
  • Record the regulatory classification with a date and rationale so scope questions become a query

Retirement Is the Step Everyone Skips

Every organisation that builds a registry builds the creation path first and the deletion path never. That asymmetry is why sprawl is a ratchet.

The mechanism that works is an expiry date on every entry, set at creation, defaulting to something short like 90 days, with renewal requiring the owner to confirm the agent is still in use. This is not a novel control; it is how well-run organisations already handle service accounts and feature flags. Applied to agents, it converts the default outcome from accumulation to decay, which is the correct default when creation is cheap and the creator has no incentive to clean up. Lifecycle systems that trigger regular reviews and automatically surface dormant agents prevent exactly the resource drain that undecommissioned systems create.

Dormancy detection is the companion control, and it is cheap because the telemetry already exists. An agent that has not been invoked in 60 days, or whose last 200 invocations produced no downstream write, is a candidate for retirement regardless of what its owner believes. Surface those in a monthly report to owners and the fleet shrinks without anyone having to run a political consolidation programme.

Deduplication is the higher-value and harder half. Two agents with overlapping purpose statements, the same connector scopes and the same data classification are almost certainly the same agent built twice, and the registry is the only place that comparison can be made mechanically. The organisational work after detection is genuine, someone has to decide which one survives and tell the other owner, but the detection itself is a query, and companies that have done this consolidated on a shared library of instruction sets rather than a shared agent, which is a lighter change to ask for.

Finally, retirement has to be complete rather than cosmetic. Deleting the registry entry without revoking the identity leaves a credential in the wild that nothing now accounts for, which is strictly worse than the agent you started with. Revoke first, delete second, and keep the audit trail for the retention period your regulator expects.

Who Builds and Runs the Control Plane

The reason this work keeps not getting done is not disagreement. Almost every engineering leader who reads the numbers above agrees with the conclusion. The problem is that an agent registry belongs to nobody.

It is not a product feature, so product teams will not prioritise it. It is not an application security review, so the security function can only escalate it. It is internal platform infrastructure with a compliance requirement attached and a permanent operational burden, which is precisely the category of work that falls between org charts and stays there until an incident assigns it. It is also, unhelpfully, a project whose success looks like nothing happening.

It is, on the other hand, an unusually well-specified mandate. A small senior team can deliver a registry service with an enforced schema, wire credential issuance through it, integrate discovery against the two or three agent platforms your organisation actually uses, implement dormancy and duplicate detection, and hand back a working expiry and revocation flow inside a quarter. That is a deliverable with an acceptance test, not an open-ended programme, and it survives being owned by a team that is not your product team. The skills it needs are platform engineering, DevOps and security engineering, in that order.

That is the shape of engagement StepTo is built around. We are a senior-led nearshore partner based in Serbia, working with European and US clients in overlapping hours, which matters here because governance work generates a steady stream of small decisions that are cheap to resolve in a same-day conversation and expensive to resolve across a twelve-hour gap. For an internal platform mandate we default to a dedicated development team rather than fixed-scope delivery, because the registry is a system you operate rather than a project you finish, and the team that built it is the cheapest team to run it. Where the harder question is which agents should exist at all rather than how to catalogue them, that is AI strategy work and it is worth separating from the build.

For EU-based clients there is a jurisdictional convenience worth naming: a nearshore development team inside the EU removes an entire category of cross-border argument from the compliance review, since the registry holds metadata about systems that may themselves be in regulatory scope. Our rates are published on the pricing page if you want to size the work before speaking to anyone.

Count Them First, Then Decide What to Keep

Agent sprawl is unusual among the problems on a 2026 engineering agenda in that the remedy is not contested. Nobody thinks a registry with an owner, a scope and an expiry date per agent is the wrong answer; the surveys, the analysts and the regulation all point at the same artefact. What is missing is an owner for the work itself, and that absence has a compounding cost, because every month the fleet grows and every new agent inherits whatever controls existed when it launched. Start with enumeration, because it is cheap and it makes every subsequent decision possible: discover what exists across the platforms your employees actually use, assign a human to each entry, and accept that the first count will be embarrassing. Then make registration the only path to credentials, so the inventory stays complete by construction rather than by diligence. Then add the expiry date, because a population that only grows is a population you will eventually stop governing. If the honest answer to "how many agents do we run" is a shrug, that is the finding, and it is a solvable one, whether you staff it internally or hand it to a <a href="/nearshore-development" class="underline decoration-dotted">nearshore team</a> that can treat it as a quarter-long mandate with an acceptance test at the end.

Building a team in Eastern Europe?

StepTo helps European and US companies build senior-led nearshore engineering teams in Serbia. Let's talk about what your next engagement could look like.

Start a conversation
I

Written by

Igor Gazivoda

Founder & CEO · StepTo

Igor has 15+ years in software engineering and business development. He specializes in scaling engineering teams, nearshore strategy, and AI-driven product development. He holds a Master's in Computer Science from the University of Belgrade.

LinkedIn →
Performance-led engineering

Want senior engineers who move work forward, not just tickets?

Work with accountable, English-fluent professionals who communicate clearly, protect quality, and deliver with a steady operating rhythm. Cost efficiency matters, but performance is why clients stay with us.

Delivery signals · senior engineering team
Senior ownership
Lead-level
Delivery rhythm
Weekly
Timezone overlap
CET
1 teamaccountable for outcomes, communication, and execution