Insights

The Commoditization of AI, Revisited

In February 2025, a few days after DeepSeek rattled the market, I wrote that large language models were rapidly commoditizing. Open models were closing in on the frontier labs, prices had fallen 92% in a year, and I argued the value was moving to integration and application.

This week gave me a reason to check that argument. On September 22, Anthropic shipped Claude Opus 5.5, and OpenAI answered the same day with GPT-6 Sol and Luna. Anthropic led with capability. Opus 5.5 took the top spot on the Artificial Analysis Intelligence Index at $4 per million input tokens and $20 per million output. OpenAI led with price. Sol and Luna cut API costs by half, and the pitch was cost per task rather than a new frontier.

Two companies with nearly equivalent technology made opposite bets on the same day. That sent me back to the 2025 piece. The short version is that I was right about the models and underestimated what the market would do about it.

The Models Did Commoditize

The original piece made two claims about the models themselves: that competition at the frontier was growing, and that prices were collapsing. Both held up, and the data is clearer now than it was then.

We pulled Arena’s text leaderboard (formerly LMSYS Chatbot Arena) month by month from October 2023 to September 2026 and measured two gaps. The first is the spread between the number one and number ten models. In October 2023 it was 172 points, back when Claude 1 sat at the top and a 30-billion-parameter MPT model rounded out the top ten. By September 2026 it was 13. It wasn’t a straight line. The spread widened again in the spring of 2025, when Gemini 2.5 Pro pulled ahead of everyone for several months, but the direction over three years is unmistakable.

Line chart of two gaps on Arena’s text leaderboard from October 2023 to September 2026: the spread between the first and tenth ranked models falls from 172 points to 13, and the gap between the best closed and best open model falls from 122 points to 21.
The spread across Arena’s top ten models and the gap between the best proprietary and best open-weight model, sampled monthly. No archived snapshots exist for October 2025 to February 2026. Source: Eskridge analysis of Arena leaderboard archives

The second gap is between the best proprietary model and the best open-weight one, and it fell from 122 points to about 21 over the same period. Twenty-one points is barely more than the spread across the entire top ten. Epoch AI’s analysis puts open models about four months behind the frontier on its capability index, which matches what Arena shows: close, and consistently a step behind.

Where switching is easy, buyers already treat models as interchangeable. On OpenRouter, a marketplace where developers can change models by editing one line of code, US models’ share of token volume fell from about 70% to about 30% in a year as cheaper open models, led by DeepSeek, took the volume. OpenRouter skews toward price-sensitive developers, so I wouldn’t read that as market share. I’d read it as what buyers do when nothing holds them in place.

The cost of a given level of capability keeps falling too. Epoch finds that the price of reaching a fixed level of performance has dropped about 13x per year since 2023, faster than any technology they compared it against.

If the story ended there, the market should by now look like any other commodity market, with dozens of interchangeable suppliers, thin margins, and no pricing power. Nothing about the actual market looks like that.

Commodity Models, Concentrated Market

Consumer usage is a long tail with a few very large heads. In May 2026, ChatGPT held 52.7% of AI chatbot web traffic, Gemini 27.3%, and Claude 8.9%, according to Similarweb. That’s about 89% for three products, with DeepSeek, Grok, Copilot, and Perplexity splitting most of the rest in low single digits. Similarweb counts web traffic only, which leaves out mobile apps, the API, and AI embedded in tools like Microsoft 365, so the exact percentages should be read loosely.

Horizontal bar chart of AI chatbot web traffic share in May 2026: ChatGPT 52.7%, Gemini 27.3%, Claude 8.9%, DeepSeek 4.0%, Grok 2.8%, Copilot 2.0%, Perplexity 1.3%.
Share of AI chatbot web traffic, May 2026. Three products account for about 89%. Web traffic only; excludes mobile apps and API usage. Source: Similarweb, via PPC Land

Enterprise spend is concentrated the same way. Menlo Ventures’ year-end data has Anthropic, OpenAI, and Google holding roughly 88% of enterprise LLM API spend at the end of 2025, at 40%, 27%, and 21%. Google, second in consumer traffic and third in enterprise spend, competes mainly on distribution, with Gemini bundled into Workspace and Android. Same goes for Microsoft and Copilot. The leaders also have pricing power, which is the one thing a commodity supplier never has. Anthropic reportedly earns API gross margins above 80% and reported its first profitable quarter this year on $11.5 billion of booked revenue.

So the capability commoditized, yet the value didn’t. Yes, cheap models are winning the tokens wherever switching is trivial, but a handful of companies are still collecting most of the revenue. That gap between a commodity product and a concentrated market is what the 2025 piece missed, and it raises an obvious question: how do a few companies hold their position when the thing they make is becoming interchangeable? The same way market leaders always have, with product and brand.

Turns Out the Value Was in the Application

In 2025 I wrote that “the real value increasingly lies in integration and application.” At the time I meant other companies building on top of the models: Cursor wrapping Claude, startups building vertical tools. That happened. What I didn’t expect was that the labs would reach the same conclusion and build the applications themselves.

Anthropic is the clearest case. In early 2025, Claude was mostly the model inside someone else’s coding tool. Anthropic’s own tool, Claude Code, now holds an estimated 54% of the enterprise coding market. Projects, the feature I singled out in the original piece, has grown into Cowork and, this month, into “one Claude”, with Docs, Slides, and Design available inside the conversation. OpenAI is doing the same thing from the consumer side, merging ChatGPT, Codex, and its Atlas browser into a single desktop app.

Why does that matter in a commodity market? Because the switching cost moved. Swapping one model for another is a configuration change. One of our clients built on OpenAI, then moved to Google mid-build when their parent company standardized on it, and the swap was a minor engineering task. Leaving a product where a team’s documents, workflows, and automations live is a migration, with a budget and a project owner, and most companies never start one. We’ve taken to calling that product the harness: the thing your team actually works in, whichever model happens to be running underneath. When the model is a commodity, the harness is what keeps you.

Product is one answer. The other, which I didn’t see coming, is brand.

The Leaders Are Choosing Different Customers

Put this week’s launches on Artificial Analysis’s chart of intelligence against cost per task and the two labs are plainly aiming at different targets. OpenAI owns the fast, cheap end of the curve. GPT-6 is a three-tier lineup, Luna, Sol, and Astra, built so a workflow can route each step to the cheapest adequate model, and Luna, at seven cents per task on the index, is available to free users. Anthropic owns the top. Opus 5.5 holds the top three positions on the index at three different effort levels, and it spends more tokens per task to get there.

Scatter plot of Artificial Analysis Intelligence Index score against cost per task on a log scale. OpenAI’s GPT-6 Luna, Sol, and Astra sit along the cheap end of the Pareto line; Claude Opus 5.5 variants occupy the top of the index at higher cost per task.
Intelligence Index score against cost per task for the models released in September 2026. OpenAI’s GPT-6 lineup runs along the low-cost end of the frontier; Anthropic’s Opus 5.5 holds the top at a higher cost per task. Source: Artificial Analysis

The business models match the positioning. OpenAI is playing for breadth. It began testing ads in ChatGPT’s Free and Go tiers in February, and the app is approaching a billion users. Anthropic is playing for depth. About 80% of its revenue comes from the API and enterprise customers, and its share of enterprise API spend went from 12% in 2023 to 40% at the end of 2025. OpenAI is pushing into enterprise as well, with more than five million business users on ChatGPT Enterprise, so think of these as more centers of gravity rather than hard lines of delineation.

The brands say it out loud. Anthropic’s Super Bowl debut in February was a set of ads mocking AI that runs ads, aimed squarely at OpenAI’s new revenue model, and it pushed Claude into the App Store’s top ten. Sam Altman called the campaign “clearly dishonest.”

This is what market leaders do when their product commoditizes. When buyers can’t tell the products apart on specs and cheap alternatives exist, a price war is a race every incumbent loses, so the leaders build ecosystems and choose an identity. The closest (albeit imperfect) analogy is smartphones. OpenAI’s approach, free at enormous scale and supported by ads, with a model at every price point, is the Android playbook. Anthropic’s approach, premium pricing, no ads, and an integrated product for people who will pay for quality, is the iPhone. iOS and Android kept their customers through their ecosystems long after the hardware converged, and that is the job the harness does now. The smartphone market never collapsed into one winner either. It settled into two dominant platforms with a long tail of smaller players behind them, which is what the AI market is starting to look like.

If the frontier AI companies are focusing on specific customers, the practical question is which customer you are. We’ve had to answer that for ourselves.

Our Answer Is Capability-First

The obvious objection is that if the models are commoditizing, paying for the top one is a waste. At the top of the curve, though, a few points still buy something. On the Artificial Analysis index, Opus 5.5 scores 58, GPT-6 Astra 53, and the best-value open models sit in the mid-40s. For high-volume routine work that difference is invisible. For work where the output is the product, a client memo, a financial model, a campaign concept, it shows up in what comes back from the AI and in the fewer revisions you have to make.

Token costs are rarely the constraint anyway. A creative agency we work with moved its copywriting tool from Gemini to Claude because the writers preferred the output, and at full scale the tool costs about $50 a month in tokens. Across our client work, the far greater budget scrutiny is on whether the tool produced a result someone can point to, and almost never on the model bill.

The premium also doesn’t last, which makes it easier for us to justify. Epoch finds that costs fall fastest right after a capability level debuts as state of the art, about 75x per year at the frontier, and open models trail by only a few months. What you pay a premium for today will be mid-market pricing within a year, so the question is whether a year’s head start is worth it for your work. For ours, it is. Quick disclaimer: Eskridge is a certified Anthropic partner. That said, we chose to invest in that partnership because their enterprise orientation aligns with ours, and they have a track record in innovation. Plus we use Claude heavily ourselves. So while we’re biased, we still have and will continue to deploy other models when it makes sense.

Which Segment Are You?

Nobody argues about whether an iPhone or an Android is better anymore. They ask which one fits their needs better, and the same question now applies to AI. In our work we see three kinds of buyers.

Capability-first. Your output quality is your revenue: professional services, agencies, finance, strategy. Pay for the frontier, commit to one platform, and go deep. The lock-in is the price of moving fast, and for you it’s worth paying.

Cost-first. Your AI workloads are high-volume and routine: support, extraction, classification, back-office processing. This is where commoditization works in your favor. If you’re interested in close to frontier capabilities wrapped in a user-friendly application, OpenAI is an excellent choice. That said, open and low-cost models are now close enough for most of this work, and a model-agnostic setup is worth building because the workloads are narrow and stable enough to easily switch between.

Requirements-first. You have regulatory or data-residency obligations, or a strong preference for owning your own infrastructure. Open-weight models are now good enough to make that realistic. Before assuming you have to build, though, check what the labs already offer: business associate agreements, zero data retention, and availability through cloud platforms you’ve already approved. The model is the commoditized part. The ongoing cost of everything built around it is what to budget for.

Keep reading

All insights