Articles

Beyond the Adoption Metric: Why AI Agent Value Requires Human-Capital-Style Management

The Measurement Gap 

Botanu came out of stealth in June 2026 with a number that should have made more noise than it did: the average enterprise now spends $186 million a year on AI. A KPMG figure released around the same time put the payoff in perspective. Only 8% of enterprises have gotten meaningful business returns from that spending, despite adoption that’s basically universal at this point. The startup’s founders framed it plainly: companies aren’t failing because the technology doesn’t work, they’re failing because they can’t tell where their agents are earning their keep. 

That framing is right, and it’s worth sitting with for a second, because the obvious response, “just measure harder,” misses what’s actually broken. For two years the enterprise AI story has been a story about adoption. How many tools did you pilot. How fast did you roll them out. Usage rates, query volume, tokens burned, all trending up and to the right on some dashboard in a board deck. Easy numbers to produce, flattering numbers to show. None of them tell you whether the thing paid for itself. 

I’d call this a management failure before I’d call it a technology failure. No company would judge a new hire purely on hours logged or emails sent. You give someone a goal, you check what they produced, you compare it to something. That’s not a controversial idea; it’s just basic management. And yet that basic idea is nowhere in how most companies evaluate their AI agents. They track how much the agents get used. Almost nobody tracks what the usage bought them. 

So, here’s the argument. Adoption numbers and token counts can’t tell you whether agents are creating real business value, because they’re activity metrics dressed up as outcome metrics, and no amount of squinting at them changes that. The fix isn’t a smarter dashboard. It’s borrowing the accountability structures companies already use on their human employees, goal setting, output attribution, baseline comparison, and pointing them at the agents instead of building some agent-specific measurement system from a blank page. 

We’ve Made This Mistake Before 

None of this is new, which is honestly the most annoying part. Industrial and knowledge-work management spent most of the twentieth century making the exact same error with human labor, and it took decades to correct. 

Early factory management measured people the way companies now measure AI, by watching them work. Time-and-motion studies counted hours on the floor, units touched, tasks completed, on the theory that more activity meant more value. That theory held up fine on an assembly line, where the connection between motion and output is basically mechanical. It fell apart almost everywhere else. Put the same logic on a sales floor or in an office and you get a worker who looks busy all day and produces almost nothing, right next to one who barely seems to move and quietly generates most of the department’s revenue. 

The correction didn’t happen all at once. Drucker pushed management by objectives in the 1950s, arguing employees should be judged against goals they’d agreed to, not against how busy they looked. Andy Grove tightened that into OKRs at Intel, tying measurable, time-bound targets to actual business results. What we now call performance management, goal setting, output attribution, quarterly review against a baseline, grew directly out of that correction. It exists because someone eventually asked the only question that matters: did this labor produce value worth what it cost. Activity metrics couldn’t answer that question. So, companies stopped relying on them, at least for anything that mattered. 

That history is useful here for one specific reason. It’s proof that the fix works, not just a nice parallel. Nobody invented performance management in a vacuum. It got built and rebuilt over decades to solve exactly the problem enterprises are now stumbling into with AI. When Botanu’s founders say companies can’t locate where their agents are creating value, they’re describing the same condition performance management was built to stamp out in human teams: labor with no defined goal, no baseline, and nobody on the hook for the result. 

I don’t think the parallel is just a nice rhetorical trick, either. Both problems trace back to the same root cause, delegated work with no attribution attached to it. A salesperson dialing hundreds of numbers a day with nothing to show for it in revenue, and an agent burning through millions of tokens with no attributable outcome, are failing in identical ways. Same blind spot, different labor. Whatever fixed it for one is a reasonable place to start for the other, not because the agent is secretly a person, but because the shape of the failure is the same, and the fix was built to address that shape. 

The Objection, and Why It Doesn’t Hold 

Somebody’s going to push back here, and they should. AI agents aren’t employees. No wage, no motivation to manage, no engagement survey, and you can clone one for the cost of a compute cycle. A framework built to manage people sounds, on its face, like a strange thing to point at software. If the analogy snaps here, the whole argument snaps with it. 

But I think that objection is aimed at the wrong target. Goal setting and output attribution were never about managing anyone’s feelings. They’re mechanisms for handling labor that’s been handed a goal, produces a result nobody can fully predict in advance, and needs to be checked against that goal after the fact. Drucker’s questions were never “is this person fulfilled.” They were narrower and more useful: what was this labor supposed to accomplish, what did it produce, how does that stack up against a target. None of that requires a pulse. It requires delegation, attributable output, and a baseline. Agents check all three boxes, arguably better than most human employees do. 

Look closer at the differences and they cut the other way. Agents don’t draw a salary, sure, but they burn compute and API spend, and that’s a cost basis every bit as real as payroll, one that slots neatly into the cost-per-output math companies already run on contractors and outsourced teams. Human performance data has always been messy because self-reported effort is easy to inflate and painful to audit. Agent activity has no such problem. Every token, every tool call, gets logged automatically. If anything, agents are easier to hold accountable than people ever were, because the paper trail comes free. 

Then there’s scale, which people usually raise as a point against the analogy, and I think is the strongest point for it. One underperforming employee is a contained problem. One underperforming agent pattern, copied across every workflow that touches it, is not contained at all, it multiplies. That’s what Gartner is describing when it says agentic projects without a baseline or an accountable owner get cancelled at much higher rates. The scale that makes agents so appealing is the same scale that makes skipping accountability so expensive. Which means the goal-output-baseline mechanism isn’t just something you could port over to agents. It’s more urgent there than it ever was for people, because mistakes compound faster. 

To be clear, I’m not arguing agents are human capital in some literal sense. That would be a much weirder and harder claim to defend. I’m arguing something narrower: agents do delegated, goal-directed work that looks structurally like what performance management was built to govern, and reusing a framework that already works beats building a new one from scratch. The real differences between agents and employees, how you set the baseline, how you calculate cost, how often you check in, are implementation details. They don’t touch the core mechanism that makes any of this work in the first place. 

Why Reuse Beats Reinvention 

Fine, say the analogy holds. A skeptic could still ask why bother retrofitting an HR framework at all, when you could just design measurement tools built for software from day one. If a purpose-built AI metric system would be faster or more accurate than adapting something meant for people, then portability is a nice-to-have, not a requirement, and the argument gets a lot weaker. 

I don’t think that objection survives contact with how companies are actually behaving. Three things push toward reuse: cost, speed, and how mature the organization already is at doing this kind of work. 

Cost first. Building AI-native measurement from the ground up means inventing new goal-setting conventions, new attribution logic, new baselines, new review cycles, essentially rebuilding performance management’s entire apparatus and slapping a new label on it. But the pattern behind failed AI deployments isn’t a lack of tooling. It’s a lack of ownership and baseline discipline, and that discipline already exists, fully formed, inside every company’s HR function. Rebuilding it from zero for agents specifically is just paying twice for the same thing. 

Speed, too. Gartner expects over 40% of agentic AI projects to get cancelled by 2027 over unclear ROI, and a good chunk of that is a speed problem, projects getting killed before any measurement framework has time to mature. The ones that survive share a specific set of traits: infrastructure built before deployment, governance documented up front, baselines captured before the pilot even starts, and one person accountable afterward. Every one of those is standard performance-management practice. Nobody invented anything new to get there. They just pointed an existing HR playbook at a new kind of worker. That’s reuse, and it’s faster than building parallel infrastructure from scratch, full stop. 

And then there’s maturity, the part people underrate. Performance management isn’t just a process sitting on a shelf somewhere, it’s wired into how budgets get set, how reviews happen, how underperformance gets escalated. Managers already know how to set a quarterly target and flag a gap to leadership. Asking them to point that same instinct at an agent costs the organization nothing new to learn. Asking them to instead run a second, parallel management culture, with its own vocabulary and its own escalation path, just for agents, is a much heavier lift, and for labor that doesn’t require it. 

None of this means token cost or latency or error rate are useless numbers. They’re not. They’re necessary, just not sufficient. The point is where they sit in the system. They should feed into an outcome-attribution framework borrowed from performance management, not stand in for one. Token cost is the agent version of tracking someone’s hours: good operational data, and a bad answer to whether the work was worth what it cost. 

Where This Leaves Us 

The enterprise AI conversation has spent two years chasing the wrong number. “How much are we using it” was never the real question, it just felt like one because it was easy to measure. The distance between record spending and rare, verified returns isn’t proof the technology doesn’t work. It’s proof that measurement never caught up to deployment, that companies scaled a whole new category of labor without scaling the accountability that’s supposed to come with it. 

We’ve closed a gap like this before. Industrial and knowledge-work management spent decades learning that activity, hours logged, calls made, units touched, can’t tell you whether labor earned its keep. The fix wasn’t a better activity metric. It was a different question altogether: what was this labor supposed to produce, what did it actually produce, and how does that measure up against a baseline. Agents aren’t people, but they’re stuck in the same blind spot people used to be stuck in, and the framework that got people out of it transfers here with less friction and less cost than anything built from scratch would. 

None of this makes token counts or adoption rates worthless, they’re still useful operational signals, the same way hours or uptime are useful. What they can’t do on their own is answer the only question that matters to a company spending nine figures a year on this stuff: was the work worth what it cost. That question already has a tested answer sitting inside every company that’s ever managed a team. The job now isn’t inventing a new discipline for measuring AI value. It’s admitting one’s already been built and using it. 

Notes 

  • The $186 million average annual enterprise AI spend and the framing of AI’s “measurement problem” come from Botanu’s June 11, 2026, stealth launch announcement (GlobeNewswire), where co-founders Alina Vrsaljko and Deborah Jacob introduced the company as “the COO for AI agents.” 
  • The finding that only 8% of enterprises have achieved meaningful business returns from AI is drawn from a 2026 KPMG report, cited in coverage of the Botanu launch. 
  • The forecast that over 40% of agentic AI projects will be cancelled by 2027 due to unclear ROI and weak risk controls is drawn from Gartner’s 2025 research on enterprise AI adoption. 
  • The traits shared by successful AI deployments (pre-deployment infrastructure investment, governance documentation, baseline metrics captured before launch, and dedicated business ownership) are drawn from industry analysis of enterprise AI agent adoption patterns published in 2026. 
  • The account of time-and-motion studies and early industrial management practice reflects the broader history of scientific management associated with Frederick Taylor in the early twentieth century. 
  • Peter Drucker’s management-by-objectives framework is outlined in The Practice of Management (1954). 
  • Andy Grove’s OKR framework at Intel is described in High Output Management (1983). 

These sources were accessed via web search in July 2026; figures attributed to Gartner and KPMG reflect secondary reporting on those firms’ research rather than the original reports themselves and should be verified against primary sources before publication.