BusinessIssue #225

Why AI Prices Stay Flat While Inference Margins Hit 80%

Compute costs are plunging and provider margins have surged by 50 percentage points, yet our subscription fees remain frozen.

Why AI Prices Stay Flat While Inference Margins Hit 80%

Opening

Reader, every time an AI company brings in $100, $35–$40 immediately flows straight into another company’s account. Specifically, to AWS, Microsoft Azure, and Google Cloud. That is the raw compute cost of running inference1. Out of that sum, the Big 3 cloud providers pocket $10–$20 in operating profit—an operating margin of roughly 35–45%.

These figures come from Barclays’ AI industry unit economics report released on August 28. Having reviewed the industry summary charts myself, I should note that all numbers reflect Barclays estimates, and figures from 2026 onward are forward-looking projections rather than realized results.

The same report points out another striking trend: paid inference margins for AI labs jumped from the mid-teens in 2025 to over 50–65% in 2026, with adjusted gross margins climbing 30–50 percentage points in a single year. For direct APIs, margins crossed 80%.

This sounds contradictory. If $40 out of every $100 drains away in computing costs, how could profit margins possibly exceed 80%? Here is the bottom line: the two figures measure entirely different denominators, and the gap between them holds the real story. It explains why, even as provider margins skyrocketed, what we pay every month hasn’t budged at all.

Why the Two Numbers Seem at Odds

First, let’s untangle the denominators. Here is the industry summary from Barclays’ estimates:

Unit: $B202420252026E2027E2028E
AI Lab Revenue726137376690
Inference Cost
(% of revenue)
4 (57%)17 (66%)58 (42%)157 (42%)292 (42%)
Training Cost
(% of revenue)
7 (96%)19 (70%)66 (48%)132 (35%)210 (30%)
Hyperscaler AI Revenue1136124289502
As % of AI Lab Revenue153%136%90%77%73%

These are Barclays Research estimates, with figures from 2026 onward being forecasts.

Looking at the 2026 column, inference costs make up 42% of revenue, and training costs 48%. The sum of these two—$124 billion—becomes, in full, the hyperscalers’ AI revenue. That is 90% of AI lab revenue.

The “$35 to $40 out of every $100” figure isolates only the inference portion of this total. Training is excluded.

An “80% inference margin” narrows the scope by another layer. It measures only API revenue from running an existing, trained model minus the direct compute cost of processing those specific requests. Training expenses, research headcount, and data acquisition costs do not enter the equation.

In other words, the former figure represents the company-wide income statement, while the latter represents the contribution margin of a single production line. Both are accurate, but placing them on the same axis is a category error. For context, the actual gross margins for AI products tracked by ICONIQ sit at roughly 41% in 2024, 45% in 2025, and 52% in 2026. While improving, they remain far from 80%.

This dynamic becomes even clearer at the far left of the table. In 2024, combined compute costs for inference and training reached 153% of revenue. Compute alone exceeded total incoming revenue. It was an era when hyperscalers took home more than the AI labs themselves.

This table is also easy to misread. One frequently sees claims that “90% of hyperscaler AI revenue comes from AI labs,” but that reverses the direction. The denominator is AI lab revenue. It means that for every $1 an AI lab earns, hyperscalers generate 90 cents—not that 90% of hyperscaler revenue originates from AI labs. The table tells us nothing about the latter. It is precisely this denominator problem that this issue keeps returning to.

Subscriptions Are the Least Profitable Product

Here is where it gets interesting.

Barclays broke down inference margins by product line: direct APIs exceed 80%, indirect APIs sit in between, and subscription products hover around 70%. Flat-rate subscription products, like Claude Code or Codex, yield the lowest margins among the three lines.

This defies conventional wisdom, where bundled offerings typically carry better margins.

The reason is simple: AI labs are deliberately subsidizing the token costs of subscription users. These are the users labs must hold on to. While API customers have already built products on top of a model—making migration difficult—subscribers can simply cancel their plans next month.

The recent frequent resets of usage limits on subscription tiers can be interpreted through the same lens. Improved model efficiency certainly helps, but pressure to prevent churn appears to be driving these decisions just as strongly. This is Barclays’ observation, not a confirmed fact.

In short: AI price tags do not reflect raw costs; they reflect who must be retained. Capital is flowing not to where costs are heaviest, but to the users most likely to walk away.

Different Business Mixes Make the Books Look Completely Different

Barclays modeled two hypothetical frontier labs for comparison. These are analytical models, not specific companies.

Lab A derives 70% of its revenue from APIs and 30% from subscriptions. Lab B flips that mix: 80% comes from subscriptions and 20% from APIs. Operating in the same industry with comparable products, their adjusted gross margins come out to roughly 55% for Lab A and about 38% for Lab B—a 17%p gap.

Half of this divergence stems from the underlying margin structure we just examined, and the other half comes from accounting. Lab A recognizes indirect API revenue on a gross basis2, whereas Lab B uses net reporting or does not recognize the portion operated by strategic partners as revenue at all. Barclays likened this to the dynamic between Uber and Lyft: identical underlying businesses where differing revenue recognition standards produce vastly different reported figures.

From the cloud provider’s perspective, the picture flips entirely. For every $100 in revenue generated by Lab A, the cloud provider pulls in $35 in revenue, generates $11.8 in operating profit, and achieves a 34% margin. For Lab B, partner revenue sharing boosts those numbers to $41 in cloud revenue, $19.1 in profit, and a 47% margin.

Barclays adds an important caveat here: revenue sharing merely inflates the headline margin percentage, while the actual profit pocketed per token remains identical. Furthermore, they expect this revenue-sharing arrangement to phase out after 2028.

There is one more dynamic worth highlighting: agentic subscription products create additional value for cloud providers. Because an agent is a persistent, stateful runtime rather than a single ephemeral chat exchange, it continuously calls upon higher-tier resources like databases and storage. That means each $1 in AI revenue pulls along significantly more downstream consumption. In some arrangements, cloud providers and AI labs even share the proceeds from this attached spend.

This dynamic applies directly to enterprise buyers as well. The more agents you deploy, the more you pay not just in model inference fees, but in the accompanying infrastructure consumption trailing behind them. If your cost estimate only factors in model pricing per token, you are looking at only half the bill.

Once AI labs start publishing GAAP financial statements, these discrepancies will land squarely in investors’ laps. We have covered before how headline metrics become difficult to take at face value. Back then, the issue lay in how ARR was calculated; this time, it is the denominator of the margin equation itself.

But Costs Aren’t Moving in Just One Direction

The discussion so far has assumed that costs are falling. And inference cost per token is indeed declining. Completing the same task now requires fewer tokens, techniques like quantization and speculative decoding have matured, and new-generation compute hardware has arrived.

Yet the price of purchasing that compute hardware is moving in the exact opposite direction. This trend is intensifying as high-end components like CCL[^3] become mainstream.

CategoryIncrease RateNotes
Memory+45% q-qQ2 2026
CCL+12~22%M8/M9 high-end
MLB+15%For AI accelerators
Package Substrate+15%For memory
MLCC+15~35%High-capacity for AI servers
AI Servers+15%+

These figures were compiled by Samsung Securities from media reports. The memory price increase is lower than TrendForce’s Q2 forecast of 58~63% for commodity DRAM, likely reflecting realized blended figures across mixed product lines.

Two distinct cost curves are moving in opposite directions: the cost of generating a single token is falling, while the cost of buying the machines that produce those tokens is rising. And this price inflation has not yet fully hit AI labs’ income statements, because once servers are purchased, their expense flows in through depreciation over several years.

The 80% gross margin we see today is therefore not a cushion of luxury. It is closer to a defensive buffer built to absorb the oncoming wave of procurement cost increases.

Why This Matters

Three practical takeaways follow from this structure.

First, never bake price cuts into your baseline budget assumptions. When planning AI adoption, many organizations build 3-year roadmaps assuming “token prices will inevitably keep falling.” Yet over the past 1 year, while gross margins improved by 30–50 percentage points, the price drops passed on to buyers were nowhere close. The cost reductions stayed inside the AI companies rather than flowing down to customers. In some tiers, nominal API list prices actually went up.

Barclays’ projections point in the same direction. Inference costs as a share of revenue are projected to plunge from 66% in 2025 to 42% in 2026, but then hold flat at 42% through 2027 and 2028. The view is that the era of massive cost collapses is already behind us. What will improve going forward is the share of training expenses, not per-token inference unit costs.

Second, looking at a vendor’s revenue mix reveals where your bargaining leverage lies. Vendors with a high API revenue share have plenty of margin cushion, but also strong incentives to defend price floors. Vendors with a high subscription share operate on thinner margins, making them far more sensitive to churn. For the latter, cancellation options carry more bargaining power than volume commitments. You will get very different responses to the exact same request depending on which bucket the vendor falls into.

Third, the revenue hyperscalers collect today is not a permanent rent. The ratio of hyperscaler AI revenue to AI lab revenue is projected to drop from 153% in 2024 to 90% in 2026, 77% in 2027, and 73% in 2028. Training spend as a share of total compute outlays will likewise tumble from 96% in 2024 to 35% in 2027 and 30% in 2028. To be clear, the absolute dollar amounts will keep swelling: hyperscaler AI revenue is projected to expand from $124 billion in 2026 to $502 billion in 2028. It is their relative share that is shrinking, not their absolute scale.

make moneyStill, the trajectory is unmistakable. The tenants are growing faster than the landlords, and from 2028 onward, their own committed in-house infrastructure projects will come online. Any cloud infrastructure procurement or investment decision needs to incorporate this timeline.

Total industry revenue is projected to rise from $7 billion in 2024 to $137 billion in 2026 and $690 billion in 2028. These are all forecasts, of course; if market expansion fails to hit these figures, the ratios above will unravel along with them.

Closing

I want to leave you with 3 things.

First, when looking at AI margin figures, check the denominator first. An inference margin of 80% and a gross margin of 52% are both correct numbers. The same goes for 90%. You just need to verify what is being divided by what.

Second, today’s price sheet is not a cost sheet—it is a policy sheet. Prices are set by who can walk away most easily.

Third, costs do not move in only one direction. While token prices drop, hardware costs are climbing, and that bill has not arrived yet.

Over the past 1 year, have you experienced an actual drop in the unit price of the AI tools you use? Or did the usage limits simply expand while the final bill stayed the same? I would love to know your take.

References & Further Reading

Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

Glossary

Footnotes

  1. Inference: The process of running a trained model to generate answers. Unlike training (which builds the model), inference occurs continuously as long as the service remains operational.

  2. Gross vs. Net revenue accounting: The distinction between recognizing the full transaction value as revenue (gross) versus recognizing only the commission or take-rate (net) in intermediary transactions. The same underlying business will show dramatically different revenue scales and margin profiles depending on which method is used.