The Capability Overhang Is Real. Most Companies Are Pointing It at the Wrong Problem.
August 2026
There's a term circulating through 2026 AI strategy discourse worth taking seriously: the capability overhang — the gap between what AI systems are capable of doing and what organizations actually use them for. OpenAI itself has argued its models already outperform most current deployments, and the usage data backs it up: API reasoning-token consumption per enterprise grew roughly 320x in a year, while the typical worker only reclaims 40–60 minutes a day from AI assistance. The frontier keeps moving. Deployed value mostly doesn't.
The standard explanation is adoption friction — data fragmentation, change management, governance debt. All real. But there's a second explanation getting less attention, and it's more uncomfortable because it's not a maturity problem you grow out of. It's a targeting problem: most enterprises are pointing frontier-capability models at capacity creation, when the actual value in the overhang sits almost entirely in capability increase. These are different economic bets, and confusing them is why so many AI programs generate an impressive demo and a disappointing income statement.
Two different jobs, and only one of them plays to the technology's strength
Capacity creation means using AI to do more of a thing you already do, faster or cheaper — deflect more support tickets, draft more first-pass contracts, process more claims per analyst-hour. The value proposition is legible: headcount avoided, cycle time cut, cost-per-unit down. It's the version of AI ROI that fits cleanly into a spreadsheet, which is exactly why it's the default pitch in almost every enterprise AI business case.
Capability increase means using AI to do something that wasn't economically viable to do at all before — not faster, but at a fidelity, scope, or personalization level that no headcount budget could previously justify. A diligence team that used to sample 5% of contracts for risk review because reading all of them wasn't affordable now reviews all of them. A single strategist who used to build one scenario model per quarter because each one took two analysts a week now runs twenty variations before a decision gets made. The comparison isn't "faster than a person." It's "wasn't happening before, at any price."
Enterprises reach for capacity creation first because it's measurable against an existing baseline. Capability increase is harder to build a business case for precisely because there's no baseline — you're not replacing a line item, you're proposing to do something the organization has never priced. That's a real barrier. It's also exactly backwards from where the technology's comparative advantage sits.
Why capacity creation keeps breaking in production
Capacity-creation deployments are where the overhang collapses hardest, and not by accident. When you point a general-purpose reasoning system at a narrow, high-volume, cost-sensitive task, you're asking it to behave like a much cheaper, much more specialized system — and at production volume, the long tail of edge cases the model wasn't optimized for starts showing up constantly. Klarna's widely reported reversal — cutting customer service roles to AI, then rehiring humans for the same work — is the canonical example of this failure mode playing out at real scale. The model didn't get worse between the pilot and the rollout. The task's tolerance for variance just went from "acceptable in a demo" to "unacceptable across millions of interactions."
This is a structural mismatch, not an execution failure that better prompting fixes. Throughput tasks reward narrow, predictable, boring reliability — the profile of a purpose-built system tuned tightly to one job, not a general reasoning model asked to approximate one. You can eventually get there, but you're spending frontier-model capability to simulate a system that didn't need frontier capability in the first place, and you're paying for the mismatch in reliability, not just in compute cost.
Where the value actually is
Capability increase doesn't have this problem, because there's no narrow baseline to underperform against — you're comparing "happened" to "didn't happen," not "AI's version" to "the human's version at scale." That's a much easier bar to clear, and it's where a general-purpose reasoning system's actual strength — breadth, synthesis, handling genuinely novel inputs — gets used for what it's good at instead of fighting against it.
Three patterns are worth looking for when scoping where to point AI capability inside a business.
Work that was rationed by expert scarcity, not by demand — Legal review, actuarial judgment, senior technical diagnosis, M&A diligence. Work where the constraint was never interest or budget, it was that only a handful of people could do it well, and their time was the bottleneck. AI doesn't need to replace the expert here; it needs to extend how much of that judgment gets applied per unit of expert time.
Work that used to be all-or-nothing due to fixed cost — Sampling instead of full review. One persona instead of segment-specific messaging. One scenario instead of a distribution of scenarios. Anywhere the organization made a coverage trade-off because doing the thorough version wasn't affordable at any previous headcount level.
Work nobody was doing at all, because it never cleared the bar of "worth assigning a person to." Continuous competitive monitoring at a granularity no analyst would maintain by hand. Personalized synthesis for every customer instead of every segment. This is the least legible category and the hardest to build a business case for in advance — and also usually the highest-ceiling one, once someone tries it.
The diagnostic question
Before greenlighting an AI initiative, ask one question: would we have hired a person to do exactly this, at any price?
If yes, you're in capacity-creation territory — proceed, but budget for the reliability engineering a narrow, high-volume task actually demands, and don't expect frontier-model economics to save you from doing that work. If no — if the honest answer is "we never would have staffed this, it wasn't worth it" — you're looking at capability increase, and that's where a general-purpose reasoning system stops competing with a cheaper narrow tool and starts doing something genuinely new.
The overhang isn't going to close by getting better at automating what you already pay people to do. It closes by using the model for the part of the value curve that was never staffed in the first place — which, not coincidentally, is also the part of the value curve nobody's competitor has gotten to yet either.
Sources — OpenAI enterprise usage data via AI CERTs (Jan 2026); capability overhang framing via Duperrin (Jan 2026) and Applied Compute; Klarna customer service reversal as widely reported in 2024–2025 coverage.