Coding Agents vs SaaS: The 2026 Build-vs-Buy Shift
٤ سبتمبر ٢٠٢٦

Nearly a third of respondents — 32% — told McKinsey their organizations decided against buying at least one software product or feature because they could build it in-house with agentic coding tools.1 That is a purchase forgone, not a contract cancelled.
TL;DR
McKinsey published its 2026 State of AI global survey on August 25, 2026.1 Among its findings: 32% of respondents said their organizations skipped buying at least one software product or feature because agentic coding tools let them build it internally.1
The survey ran May 4 to June 8, 2026, drawing 1,719 respondents across 97 nations, weighted by each country's share of global GDP.1
Two numbers sit underneath it. Large enterprises scaling AI agents in one or more functions jumped from 27% to 40% in a year, while smaller organizations stayed flat at 22%.1 And about two in ten organizations are scaling software coding agents across the enterprise — 31% at larger enterprises.1
Markets had already repriced enterprise software months earlier, though mostly on a neighboring fear. Three of the four hardest-hit names in the February 3, 2026 selloff were legal- and professional-information vendors — Thomson Reuters, RELX, Wolters Kluwer — alongside LSEG. In-house building was one item on one analyst's list of risks.2
The counterweight is running cost. About 20% of respondents said AI operating costs, including tokens, constrained their organization's AI use — and McKinsey's AI high performers reported cost constraints on software coding agents about three times as often as everyone else.1
What You'll Learn
- What McKinsey's 32% figure actually counts, and what it does not
- How far the agent adoption gap between large and small companies widened
- How the 6% of "high performers" behave differently from everyone else
- What Retool's competing 35% number measures, and why its sample matters
- What the February 2026 software selloff did and did not establish
- What aggregate software spending has done while all this was happening
- Where token costs and code security slow the build side down
- What Veracode's security numbers do not tell you about agents specifically
- Questions worth asking before you let a renewal lapse
What McKinsey Actually Measured
McKinsey does not define the term in its published write-up. In common usage, agentic coding tools are AI systems that plan and execute multi-step software development work — reading a repository, writing code across files, running tests, and iterating — rather than autocompleting a line at a time.
McKinsey's wording on the finding itself is precise, and worth quoting rather than paraphrasing. Respondents reported that their organizations "decided against purchasing at least one software product or feature because they were able to build the functionality in-house using agentic coding tools."1
Read that carefully. It counts a decision not to buy something.
It does not count a cancelled contract, a churned seat, or a replaced incumbent. An organization that declined to add one small reporting add-on qualifies. So does one that walked away from a seven-figure platform evaluation.
The published write-up breaks the figure out by industry, but not by deal size. Those decisions are reported most often in technology and healthcare, followed by professional services and energy and materials.1
It is also "at least one," which sets the bar at a single instance across an entire organization, with no reference period stated in the published write-up. That is a low threshold for a headline that reads like a market shift.
None of this makes the finding weak. Nearly a third of a GDP-weighted global sample reporting a changed procurement decision because of a new tool category is a real signal about where IT budgets are heading. It is just a narrower signal than "companies are replacing their SaaS stack."
Michael Chui, a McKinsey senior fellow and co-author of the report, framed it in budget terms rather than replacement terms: "Agentic coding tools are creating a more capable option for moving some software development in-house."3
Note the word "some."
The Adoption Gap Widened
The build-vs-buy finding rides on top of a broader split that got wider this year, not narrower.
Among organizations with more than $1 billion in annual revenue, the share scaling AI agents in one or more functions rose from 27% to 40%.1 Among smaller organizations, McKinsey describes adoption as remaining "essentially flat, at 22 percent."1
That is a 13-point gain on one side of the line and, on the other, a year of standing still.
The same pattern shows up in enterprise-wide AI scaling generally: 54% of the $1B-plus cohort report scaling AI across the enterprise, against roughly one-third of smaller organizations.1 Overall, 44% now report enterprise-wide scaling, up from 38% a year ago.1
Coding agents are still a minority tool. About two in ten organizations report scaling them across the enterprise, rising to 31% at larger enterprises.1 Chatbots remain far ahead at 47%.1
So the honest framing is this: coding agents are a fast-growing minority capability, concentrated in large companies, that has already started to change some procurement decisions. They are not yet the default way enterprises get software.
Where agents are scaling, McKinsey found it happens most often in IT, knowledge management, and software engineering, with technology and media and telecommunications the industries most likely to report scaling agents within functions.1
The 6% Doing Something Different
McKinsey defines AI high performers as respondents who attribute at least 5% of EBIT to their AI use and describe the impact as "significant." They account for 6% of all respondents, unchanged from 2025.1
On build-vs-buy, they behave measurably differently. Nearly half of high performers report skipping at least one software purchase because they could build it in-house, against fewer than a third of everyone else.1
They are twice as likely as others to be scaling software coding agents, and 2.7 times more likely to be scaling other agentic AI.1
The behavioral difference McKinsey emphasizes is not tooling, though. Nearly three-quarters of high performers report fundamentally redesigning workflows because of AI, up from 55% last year — against one-quarter of other respondents.1
The causal ordering here is easy to reverse in your head: tools first, then results. The data does not show that buying coding agents produced the outcome. It shows that one population — high performers — reports doing both, without establishing which came first.
Retool's Bigger Number, and Its Sample
A second survey this year reported a larger figure, and it deserves a caveat.
Retool released "The Build vs. Buy Shift" on February 17, 2026, reporting that 35% of teams have already replaced at least one SaaS tool with a custom build. And 78% said they expect to build more custom internal tools in 2026 — an intention, not an action.4
That is a stronger claim than McKinsey's — replacement rather than a purchase forgone.
It also rests on a different sample: 817 Retool customers and builders, surveyed in late 2025.4 Retool sells a platform for building internal software. Its respondents are self-selected toward building.
Treat the 35% as a figure observed among committed builders, not a population estimate. Retool CEO David Hsu's framing in the release — "enterprise AppGen has become a threat to traditional SaaS" — is a vendor stating its own thesis.4
The report's shadow IT finding carries the same sample caveat, and is worth noting anyway: 60% of respondents said they had built software outside IT oversight in the past year, and 25% said they do so frequently.4
If that direction holds outside a builder-heavy panel, faster building without faster procurement and review does not produce a leaner stack. It produces an unmapped one — shadow IT with a code generator attached.
A later Retool survey put numbers on the discomfort. Fielded by third-party research firm Wynter among 307 CIOs, CTOs, and CISOs and published June 17, 2026, it found 93% concerned about vibe-coded tools running in production.5
38% were "very concerned," calling it a top operational risk, and just 8% described their organization's governance as "strong."5 Note that "vibe-coded" is a broader and more pejorative category than agentic coding — unreviewed prompt-generated code of any kind.
Apply the same skepticism here. It was fielded by a third party rather than by Retool and asked executives rather than builders, which blunts the sample objection — though the release does not say the panel excluded Retool customers. And Retool published it alongside a governance product, so a finding that enterprises have a governance gap is not inconvenient for the sponsor.5
What the February Selloff Did and Didn't Prove
Markets moved on this thesis months before McKinsey published.
On Tuesday, February 3, 2026, enterprise software and analytics stocks fell sharply. RELX dropped more than 14% and Wolters Kluwer more than 12%; Thomson Reuters posted a record 16% slump on fears about its core legal division; London Stock Exchange Group fell nearly 13%.2
The selloff reached Asian markets the next day, with NEC, Nomura Research, and Fujitsu falling between 8% and 11%.2
Reuters named one trigger: Anthropic's launch of plug-ins for its Claude Cowork agent the previous Friday, January 30, enabling automated tasks across legal, sales, marketing, and data analysis.26
J.P. Morgan's Toby Ogg described the mood bluntly: "We are now in an environment where the sector isn't just guilty until proven innocent but is now being sentenced before trial."2
Reuters reported that Ogg found investor appetite to step in "generally low," citing risks including competition from AI-native firms and clients building their own solutions in-house.2
Not everyone agreed. Nvidia CEO Jensen Huang played down the same fears that week, calling the idea that AI would replace software and related tools "illogical" and saying "time will prove itself."2
Note which companies fell hardest. RELX and Wolters Kluwer are described by Reuters as major providers of analytics to the legal industry, and Thomson Reuters slid on fears about its core legal division.2 The trigger Reuters named was an agent automating legal, sales, marketing, and data-analysis work.2 LSEG's near-13% fall is reported without any reason attached.2
For the legal-information names at least, that is a fear about agents doing the work those vendors sell — adjacent to the build-vs-buy thesis, but not the same one. In-house building appeared as one item on one analyst's list of risks.2
A share price is a forecast, not a finding. The February move tells you what investors expected of enterprise software revenue; McKinsey's August survey tells you what buyers reported actually doing. Only one of them is evidence of buyer behavior.
Aggregate spending is not turning down, either. Gartner's July 2026 forecast puts worldwide software spending at $1.468 trillion for the year, up 15.5% from $1.271 trillion in 2025 — an acceleration on 2025's 13.9% growth, not a contraction.7
Whatever coding agents are doing to procurement, it is not yet visible as a decline in aggregate software spend.
Two Brakes on Building
The build side has costs that a procurement spreadsheet catches late.
Running cost. One in five McKinsey respondents said AI operating costs, including tokens, had constrained their organization's AI use.1 Broken out by tool, roughly one in ten reported constraints for chatbots, for agents, and for coding agents each.1
The sharpest detail is about the leaders. High performers — the group building the most in-house — reported cost constraints on software coding agents about three times as often as other respondents, while showing no unusual constraint on other tool types.1
The organizations furthest into this are the ones reporting its bill most often. The alternative reading is that they simply measure more carefully, and so notice constraints others absorb without naming.
Either way, that makes session-level spend caps for agents a live procurement question rather than an engineering nicety. It is also why the collapse in open-weight coding-model prices matters to the build case.
McKinsey senior partner Lieven Van der Veken, as reported by CIO Dive, said that to manage IT spending optimally, businesses should be deliberate about when to buy something instead of building it, and take greater ownership over their technology and change agendas. He added: "They are also treating operating costs as a design constraint, not an afterthought, and making the deeper changes in workflows required for value realization."3
Security. Veracode, which sells the application security testing this post ends up recommending, released its 2026 GenAI Code Security Report on July 28, 2026. Across four testing snapshots and more than 100 models tracked since the program began, it puts the average security pass rate at 56% — "virtually unchanged since last year's report."8 The 2026 edition itself tested 11 new models.9
Set that against syntax. Models now generate compilable code at what Veracode calls a near-universal syntax pass rate of roughly 100%, while failing security tasks nearly 44% of the time when given no security-specific guidance.8
In the Summer 2026 leaderboard, GPT-5.5 led at 68%; six of the eleven models clustered between 50% and 53%, with Alibaba's Qwen3.7-max last at 50%.8 Language spread ran from Python at 63% down to Java at 30% — Java the riskiest by a significant margin, and, Veracode notes, the only language with a clear upward trend over the past year.8
Two assumptions found no support in the data. Coding-specialized models averaged 51% against 52% for general-purpose models, and model size showed what Veracode calls no impact on security performance: 53% above 100 billion parameters against 51% for medium and small alike.8 Veracode's release and public summary give no confidence intervals, so treat spreads this narrow with care.
The best observed result has moved slightly since. In follow-up testing published August 27, 2026 — after the report, on a newer model — Veracode put GPT-5.6 Sol at 70% secure overall against GPT-5.5's 68%. Python rose from 70% under GPT-5.5 to 85% under Sol, and cross-site scripting from 50% to 60%.10
Those are two-model comparisons, not the cross-model language averages above.
The gains did not generalize. "C# and Java did not share the gain, JavaScript was unchanged, and log injection (CWE-117) remained a pronounced weakness," Veracode wrote. It summarized the point as: "A model has a security profile, not a security score."10
For a build decision, that shifts the question. It is not which model tops a leaderboard, but whether the language and vulnerability classes your internal tool actually lives in are ones that model happens to handle well.
One scope caveat matters more than any of these numbers, and Veracode states it plainly: "The tests were run against the raw models and not agents or production environments with additional tooling, guardrails, or human review in the loop."8
That is a real limit on reading this into the agentic case. An agent that runs tests, reads the failure, and iterates is not the thing being measured here — the methodology is a function-completion task, 80 of them across four languages, targeting four vulnerability classes, with no security-specific prompting.11
Whether agent loops close that gap or widen it is not something this data answers.
What survives the caveat is the asymmetry of who carries the risk. Buy the software and a vendor owns its security program, whatever its quality, with a contract apportioning some of the burden. Generate it and all of it stays with you — as does what the agent itself pulls into its context, which is a second supply chain problem.
McKinsey vs Retool: What Each Number Measures
| McKinsey State of AI 2026 | Retool Build vs. Buy 2026 | |
|---|---|---|
| Headline figure | 32% skipped a purchase1 | 35% replaced a SaaS tool4 |
| What it measures | Decision not to buy ≥1 product or feature | Replacement of ≥1 existing tool |
| Sample | 1,719 respondents, 97 nations1 | 817 Retool customers and builders4 |
| Field dates | May 4 – Jun 8, 20261 | Late 20254 |
| Weighting | By nation's share of global GDP1 | Not stated in the release |
| Sponsor interest | Consulting firm; sells AI advisory services | Sells the build platform |
| Shadow-IT finding | Not reported | 60% built outside IT oversight4 |
What to Ask Before You Drop a Renewal
Five questions that separate a defensible build decision from an expensive one:
- Is this a differentiator or a commodity? Chui's read is that spend on horizontal AI-enabled tools such as chatbots is "increasingly managed as a necessary cost of doing business, similar to office productivity and communications tools."3 That is an argument for buying the commodity layer. The corollary — that the build case is strongest where software encodes something specific to you — is ours, not his.
- What is the token bill at steady state? Not the build cost — the run cost, after the agent is maintaining and extending the thing. The high performers are the group reporting this constraint most often.1
- Who owns it in eighteen months? Generated code still needs an on-call owner, a dependency policy, and a patch path. A vendor contract typically covers part of that and gives you someone to escalate to for the rest.
- What replaces the vendor's security program? Veracode's raw-model pass rates — 56% on average, with its report's leader at 68% and its August follow-up test of GPT-5.6 Sol at 70% — do not measure agent loops, and its data does not address whether those loops close the gap.810 Until something shows they do, plan for mandatory static analysis and human review.
- Will IT know this exists? Retool's builder-heavy panel reported 60% shipping outside IT oversight.4 That is a biased sample, but it points at the risk — ask what your own number is before assuming it is lower.
The Bottom Line
The most defensible version of this story is smaller than the headline and more useful.
At the organizations of roughly a third of McKinsey's respondents, agentic coding tools have already changed at least one procurement decision — reported most often in technology and healthcare, and among the 6% who qualify as high performers, most of whom have already redesigned their workflows.1 That is real, and McKinsey had not asked it before in this survey.
What the data does not show is a stack-wide replacement of enterprise software. McKinsey counted forgone purchases, Retool counted replacements inside its own builder base, and the February selloff counted investor sentiment about a related but different fear. None of the three measures SaaS churn in a population.
The interesting tension is inside the leaders. The same organizations building the most in-house are the ones reporting cost pressure on the tools doing the building most often.1 Whatever the build-vs-buy line looks like in 2027, it will be drawn by whoever gets the run cost under control — not by whoever ships the first internal replacement.
Footnotes
-
Dan Tinkoff, Lieven Van der Veken, Michael Chui, and Tara Balakrishnan, "The state of AI in 2026: On the road to ROI," McKinsey & Company, August 25, 2026. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15 ↩16 ↩17 ↩18 ↩19 ↩20 ↩21 ↩22 ↩23 ↩24 ↩25 ↩26 ↩27 ↩28 ↩29 ↩30 ↩31 ↩32 ↩33 ↩34 ↩35
-
Danilo Masoni, "Global software stocks hit by Anthropic wake-up call on AI disruption," Reuters, February 4, 2026. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13
-
Paige Gross, "Enterprises bet on agents to build in-house software, boost productivity," CIO Dive, August 27, 2026. ↩ ↩2 ↩3
-
"Retool's 2026 Build vs. Buy Report Reveals 35% of Enterprises Have Already Replaced SaaS With Custom Software," Business Wire, February 17, 2026. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10
-
"As Vibe Coding Tops C-suite's List of Concerns, Retool Unveils First Platform to Extend Enterprise Governance to All AI-Coded Apps," Business Wire, June 17, 2026. Survey fielded by Wynter; the release's summary rounds the sample to 300 and its body states 307. ↩ ↩2 ↩3
-
Lucas Ropek, "Anthropic brings agentic plug-ins to Cowork," TechCrunch, January 30, 2026. ↩
-
"Gartner Forecasts Worldwide IT Spending to Grow 14.2% in 2026, Totaling $6.37 Trillion," Gartner, July 27, 2026. Software segment figures from Table 1. ↩
-
"LLMs Are Getting Smarter, But Not Safer: Veracode 2026 GenAI Code Security Report Finds AI-Generated Code Security Has Stalled at 56% Pass Rate," Business Wire, July 28, 2026. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9
-
"2026 GenAI Code Security Report," Veracode, July 28, 2026. Figures cited here are from the report's public summary page. ↩
-
Veracode Research, "GPT-5.6 Sol Shows Why a Better Model Isn't a Uniformly Safer Model," Veracode, August 27, 2026. ↩ ↩2 ↩3 ↩4
-
Felix Brombacher, "Spring 2026 GenAI Code Security Update: Despite Claims, AI Models Are Still Failing Security," Veracode, March 24, 2026. ↩

