A Third of Enterprises Skipped a Software Purchase. On Thursday OpenAI Shipped a Model That Drives the Software They Did Buy.
On August 25, 2026, McKinsey published the 2026 edition of its State of AI survey. Buried under the usual adoption percentages was a number that should have led every writeup: 32 percent of organizations reported deciding against buying one or more software products or features because they could build them internally with agentic coding tools. The survey drew 1,719 responses across 97 countries between May 4 and June 8, weighted by each country's share of global GDP.
Nine days later, on Thursday, September 3, OpenAI released GPT-6 Astra and reported 72.6 percent on OSWorld 2.0, the computer use benchmark, against 65.7 percent for GPT-5.6 Sol. The human baseline on that benchmark sits at roughly 72 percent. OpenAI demonstrated the model clicking through browsers, filling forms, updating CRM records, editing spreadsheets, building presentations and driving engineering applications.
These got covered as two unrelated stories. A consulting survey and a model launch. I think they are one story, and the object both of them are pointed at is the enterprise software seat.
What a Seat Has Always Actually Priced
A per-seat license bundles two goods that nobody separated because there was never a reason to. The first is the write entitlement: the right for a given identity to read and modify rows in a governed system of record, under permissions, validation rules, approval chains and an audit trail. The second is a human being sitting at a screen, clicking through the vendor's interface to exercise that right.
Those were the same purchase for thirty years because the only way to exercise the entitlement was through the interface, and the only thing that could operate the interface was a person. Both halves of that sentence stopped being true this year, and they stopped being true from opposite directions.
| What the seat priced | The 2026 capability that attacks it | Evidence on the table | Who loses first |
|---|---|---|---|
| The decision to buy at all | Agentic coding tools | 32% skipped a purchase | Point solutions and feature-tier upsells |
| The human at the screen | Computer-use agents | 72.6% OSWorld 2.0 | Seat count on software already bought |
| The governed write entitlement | Nothing yet | n/a | Nobody. This is the part that holds |
The third row is the reason this is a repricing rather than an extinction event. Twenty years of process, permissions and audit obligations encoded into a system of record cannot be regenerated from a prompt, and a model that drives a browser still has to authenticate as somebody.
The 32 Percent Is Not Evenly Distributed, and the Skew Is the Tell
The headline number is a blended average across every industry in the sample. Split it and the pattern is not subtle: the sectors with the most in-house engineering capacity are the ones walking away from purchases.
| Sector | Decided against a software purchase | Spread vs the 32% average |
|---|---|---|
| Technology | 41% | +9 |
| Healthcare payers and providers | 39% | +7 |
| Professional services | 38% | +6 |
| Energy and materials | 38% | +6 |
| Financial institutions | 36% | +4 |
| Media and telecom | 34% | +2 |
| Pharma and medical products | 33% | +1 |
| High performers (5%+ of EBIT from AI) | ~50% | +18 |
| Everyone else | 31% | -1 |
The last two rows are the ones I keep going back to. McKinsey defines high performers as the roughly 6 percent of respondents attributing at least 5 percent of EBIT to AI use, and nearly half of them skipped a software purchase against 31 percent of everyone else. That is an eighteen point gap between the organizations getting measurable financial return from AI and the organizations that are not.
You can read that two ways and both are uncomfortable for a software vendor. Either building instead of buying is what produces the EBIT, in which case the behavior spreads as the playbook gets copied. Or the same organizational competence that produces AI returns also produces the confidence to build, in which case the cohort walking away from your product is precisely the cohort with the budget and the engineering bench you most wanted as a reference customer.
Thursday's Number Is About the Software You Already Bought
Agentic coding attacks the purchase decision. Computer use attacks something different and arguably more valuable: the number of seats you need on software you have already committed to.
Astra's headline computer-use figure is 72.6 percent on OSWorld 2.0. The number almost everyone skipped is the second one: roughly 40 minutes per evaluated task against about 75 minutes for its predecessor, a reduction of about 47 percent. That is the figure that changes the economics, because a computer-use agent is only viable if you can afford to retry it. Halving wall-clock time per attempt halves the cost of the failure mode that makes agents unusable in production.
| Measure | GPT-6 Astra | GPT-5.6 Sol | Delta |
|---|---|---|---|
| OSWorld 2.0 (computer use) | 72.6% | 65.7% | +6.9 pts |
| Avg time per evaluated task | ~40 min | ~75 min | -47% |
| MRCR v2 8-needle, 512K to 1M | 96.3% | 73.8% | +22.5 pts |
| Context window | 1,050,000 | smaller | 922K in / 128K out |
| Price per 1M in / out | $10 / $50 | $4 / $20 | 2.5x |
Vendor-reported figures from OpenAI's launch materials and subsequent coverage. Sol pricing is TensorFeed's tracked list rate. Note the pricing footnote most writeups dropped: requests exceeding 272,000 input tokens are billed at 2x the input and cache rates and 1.5x the output rate, which matters a great deal for long-horizon agent runs that accumulate screen state.
The long-context row is the one to sit with. A computer-use agent working a real business process accumulates screenshots, DOM state, tool outputs and its own prior reasoning, and it degrades when it stops being able to find the instruction it was given forty steps ago. Going from 73.8 percent to 96.3 percent on multi-needle retrieval in the 512K to 1M band is not a benchmark flex. It is the difference between an agent that can finish a multi-hour workflow and one that loses the plot in the middle of it.
If you want the current list rates side by side rather than taking mine, our models tracker carries them, and the pricing tracker holds the history.
The Vendors Already Repriced. That Is the Confirmation.
The strongest evidence that this is real is not the survey. It is that the companies with the most to lose have already moved their pricing, which is not something a vendor does for fun.
ServiceNow has said half its net-new business is no longer sold by seat. Salesforce put a price on a single agent resolution. HubSpot cut its own revenue per customer by charging only for conversations its AI actually resolved. In the same period, ServiceNow made Build Agent generally available and pushed its core skills out into Cursor, Windsurf, Claude Code and GitHub Copilot, which is a system-of-record vendor concluding that the build activity is going to happen in somebody else's editor and deciding it would rather govern it than lose it.
That is the same concession Salesforce made when it shipped 37 prebuilt skills into Claude's chat window, which we wrote about last week. Give up the interface, keep the governed data layer. Thursday's computer-use numbers are the argument for why giving up the interface was the correct read: an interface that a model can operate at human parity is not a moat, it is a compatibility surface.
Three Counterarguments, Taken Seriously
One: 32 percent is a decision not to buy, not a system running in production. This is the strongest objection and it is largely correct. The fieldwork closed on June 8, so a decision captured in the survey is a prototype in September at best. Worse for the build case, published estimates put maintenance at roughly 60 to 90 percent of total software lifecycle cost, and a team that builds a feature with a coding agent inherits the maintenance, the security review and the migration burden that the vendor was pricing into the license. McKinsey's own survey has 20 percent of organizations already limiting AI use because of operating costs, which is the bill starting to arrive. But note what the objection concedes: the decision is the thing that hits the vendor first. A deal that never enters the funnel does not appear as churn. It appears as a quarter that came in fine and a next year that did not.
Two: enterprise software spend is still growing, so nothing is being displaced. Gartner has enterprise software spend rising 14.7 percent in 2026 to more than $1.4 trillion, and that is real. It is also the wrong denominator. Spend growing at 14.7 percent while roughly a third of buyers skip a purchase means the mix is moving, not that the behavior is imaginary, and Gartner names generative AI as the primary accelerant of that growth. A large share of the increase is new AI line items rather than renewals of the products being skipped. Aggregate growth can hide a category rotation for several years, and usually does.
Three: 72.6 percent is nowhere near reliable enough to run a business process. Correct, and the human-parity comparison everyone ran this week is misleading in a specific way. A human at 72 percent and a model at 72.6 percent fail differently. The human fails on the tedious tail and knows they are failing. The model fails unpredictably in the middle and reports success, and Astra has been observed stopping mid-task. Every figure here is also vendor-run and none of it has been independently reproduced. But reliability is the wrong axis for the seat question, because a computer-use agent does not have to be better than an employee to remove a seat. It has to be good enough that one employee supervising several agent runs is cheaper than several employees, and that threshold sits well below parity.
Our Take
The number I cannot stop looking at is not 32 percent. It is that 80 percent of respondents report individual productivity gains while 37 percent report any EBIT impact at all, and that 37 percent is flat against last year's survey. Two years of deployment, real gains at the desk, nothing arriving at the bottom line.
An organization in that position eventually goes looking for the money, and the software line is where it looks. Not because software is the largest cost, but because it is the most legible one: an itemized list of recurring charges, each tied to a headcount that a coding agent or a computer-use agent can now plausibly be argued down. The build-versus-buy shift is not primarily a story about engineering capability. It is a story about a CFO who was promised AI savings, cannot find them in EBIT, and has a renewal calendar in front of them.
What survives is the governed write entitlement, which is why the system-of-record vendors are repricing rather than panicking and why the point solutions should be. If your product is a workflow wrapper over data somebody else owns, both of this month's stories are about you: the coding agent can rebuild your wrapper, and the computer-use agent can operate the thing your wrapper was wrapping.
Three Signposts
Whether any major SaaS vendor breaks out a build-loss category in churn reporting. Right now losses to internal development get filed under competitive loss or budget cut, which makes the phenomenon invisible in every public disclosure. The first vendor to name it will do so because it has gotten too large to bury, and that filing is the moment this stops being a survey finding and becomes a market fact.
Whether the next State of AI survey separates decided from shipped. The 2026 instrument asks whether an organization decided against a purchase. It does not ask whether the replacement is in production a year later. That single follow-up question is the difference between a procurement mood and actual displacement, and until somebody asks it the 32 percent is a leading indicator of unknown strength.
Whether any vendor prices a license for a non-human operator. Agent-resolution pricing at Salesforce and outcome pricing at HubSpot are steps toward it, but nobody has yet published a rate card line that says: this identity is an agent, it costs this, and it is not a person. The day a system-of-record vendor ships that SKU is the day the seat formally stops being about people, and every renewal negotiation after it runs on different arithmetic. We will be tracking it on the pricing tracker.
Marcus Chen, September 4, 2026. TensorFeed tracks AI model releases, pricing, and provider status in real time. Sources for this piece: McKinsey's State of AI global survey published August 25, 2026, OpenAI's GPT-6 Astra launch materials of September 3, 2026, and subsequent reporting. All capability and timing figures for Astra are vendor-reported and not independently reproduced.
