Loading...


Updated 1 Sep 2026 • 4 mins read

AI inference is a real, variable cost of goods sold, and it is pulling SaaS gross margins down from 80 percent toward 50 to 60 percent. This guide explains how AI spend hits SaaS profitability, the three-layer AI P&L, the heavy-user trap, and the levers that protect gross margin.
For twenty years, software had a magic number: an 80 percent gross margin. Build the product once, serve the next customer for almost nothing, and watch the margin hold as you scale. It was the whole reason software commanded the valuations it did. In 2026 that number is quietly breaking, and AI is the reason. ICONIQ's 2026 data puts the average gross margin for AI products at about 52 percent, against the 70 to 80 percent that defined mature SaaS. The magic is not gone, but it is no longer free.
The cause is not waste or mismanagement. It is structural: AI inference is a real, variable cost that scales with usage, and it sits squarely in cost of goods sold. This piece explains exactly how AI spend hits SaaS profitability, what the new P&L looks like, the trap that makes a healthy-looking average dangerous, and the levers that protect gross margin as you add AI.
The short version Traditional SaaS runs at 70 to 80 percent gross margin because serving another customer costs almost nothing. AI changes that: every query spends real inference, so AI-native products land around 50 to 60 percent (ICONIQ's 2026 average is 52 percent), and traditional SaaS adding AI features typically loses 10 to 17 points of gross margin. The fix is to treat AI inference and infrastructure as explicit COGS layers, measure AI-adjusted gross margin, and pull the levers that control it, model routing, caching, and usage-aware pricing, before a flat price quietly goes negative on your heaviest users.
Classic SaaS economics rested on one fact: near-zero marginal cost. Once the software existed, the cost to serve one more customer was a rounding error, some database queries, a little hosting. Revenue scaled far faster than cost of goods sold, so gross margin stayed high and stable as the business grew. That is what let software outrun every other business model on profitability, and it is the assumption baked into how investors value SaaS.
AI cannot copy that trick. Traditional software is built once and copied for free; an AI feature spends real compute on every single call. Each inference has a cost that scales directly with usage, which makes it a variable cost of goods sold, not a fixed R&D expense you amortize away. As we argued in AI costs are cloud costs now, AI spend behaves like cloud spend, only more variable, and that variability lands in the one line investors watch most: gross margin. The moment your product makes a model call to serve a customer, that call is COGS.
The compression is well documented across 2026 benchmarks
| Metric | Traditional SaaS | AI product |
|---|---|---|
| Typical gross margin | 70–80%+ | 50–60% |
| 2026 benchmark average | ~80% | ~52% (ICONIQ) |
| Cost structure | Near-zero marginal cost | Inference COGS scales with use |
| Margin hit from adding AI | Baseline | 10–17 points lower |
| Inference as share of revenue | Negligible | ~4–9% (public), up to ~23% at scale |
Bessemer and a16z frame the AI band at 50 to 60 percent versus 80 to 90 percent for software, and public SaaS companies are now disclosing gross margins 10 to 17 points below their pre-AI baselines, with inference broken out as its own line in financial filings. The trend is not temporary: ICONIQ found inference rising as a share of spend as products mature and usage grows, so this is a structural reset, not a launch-phase blip. Understanding it starts with the hidden cost multipliers that make real inference cost higher than the rate card suggests.
The cleanest way to see the damage, and manage it, is to stop hiding AI cost inside a generic infrastructure line and stack the P&L explicitly:
| P&L line | What it contains |
|---|---|
| Revenue | Total revenue (ideally split into AI and non-AI) |
| less Traditional COGS | Hosting, support, customer success, services |
| less AI Inference Layer | Model API fees, per-call inference, self-hosted compute |
| less AI Infrastructure Layer | Vector databases, model monitoring, amortized fine-tuning |
| = AI-Adjusted Gross Profit | The number that actually matters |
AI-Adjusted Gross Margin is that gross profit divided by revenue. Pulling inference and AI infrastructure out as their own layers is what turns a vague sense that AI is expensive into a number you can track, defend, and improve. If you can, compare AI revenue directly against AI COGS to see whether the AI part of your product actually pays for itself.
Here is the failure that catches teams by surprise. A flat monthly price can look perfectly healthy on a blended average while quietly running negative on your heaviest 10 percent of users, because an agentic workflow can consume orders of magnitude more tokens than a single chat turn. The average hides the loss; the power users create it. GitHub watched exactly this happen to Copilot and moved every plan to usage-based billing in June 2026 to stop the bleed. The defense is per-customer visibility: knowing cost per customer and per feature, not just in aggregate, which is what cost allocation and cloud unit economics make possible.
Margin compression is a structural headwind, but it is manageable. The levers, in rough order of impact:
The tailwind helps too: model prices keep falling, so teams with cost discipline benefit on the input side as vendors pass through cheaper inference. The mechanics of these levers are worked through in our LLM cost optimization and FinOps for AI guides.
Gross margin is not just an accounting line; it drives valuation. A SaaS business at 25 percent growth and 80 percent gross margin scores far higher on the Rule of 40 than the same business at 67 percent gross margin, even with identical cash flow, so margin compression mechanically lowers how the market values you. Sophisticated investors have already moved to a post-AI-COGS view, comparing companies on apples-to-apples gross-margin bands and reading eval-engineering hires as structural COGS. In that world, being able to show a defensible, well-instrumented AI gross margin is not just good hygiene, it is a competitive advantage in the room where your company gets priced.
AI did not make software a worse business. It made it a different one. The near-zero marginal cost that defined SaaS for two decades is gone for anything with a model in the loop, and pretending otherwise, by burying inference in an infrastructure line and quoting a blended margin, is how good companies get surprised by their own P&L.
The teams that win the next few years will not be the ones with the cheapest AI or the most features. They will be the ones who treat inference as the cost of goods it is, know their gross margin per customer to the cent, and price and route accordingly. The 80 percent era is over. The era of AI-adjusted gross margin, measured, defended, and managed on purpose, is the one that replaces it, and it rewards the companies that started counting first.
AI gross margin is gross profit as a percentage of revenue after treating AI inference and AI infrastructure as explicit costs of goods sold. Because every AI query spends real compute, it is lower than traditional SaaS gross margin, typically 50 to 60 percent versus 70 to 80 percent.
Because AI inference is a variable cost that scales with usage and sits in COGS. Traditional software copies for free, supporting 80 percent margins, but an AI feature spends compute on every call, so margins compress. ICONIQ's 2026 benchmark puts the average AI product margin at about 52 percent.
Public SaaS companies are reporting gross margins 10 to 17 points below their pre-AI baselines after adding AI features, with inference costs often disclosed as a separate line running roughly 4 to 9 percent of revenue, and higher at scale.
As their own layers within COGS: an AI Inference Layer (model API fees, per-call inference, self-hosted compute) and an AI Infrastructure Layer (vector databases, model monitoring, amortized fine-tuning). Subtracting both from revenue, after traditional COGS, gives AI-adjusted gross profit.
A flat monthly price can look healthy on a blended average while running negative on the heaviest users, because agentic workflows consume far more tokens than light use. The average hides the loss. GitHub moved Copilot to usage-based billing in 2026 for exactly this reason.
Route work to the cheapest capable model, cache repeated context, adopt usage-based or metered pricing so revenue tracks cost, measure cost per customer and feature to find negative-margin cohorts, and treat inference and AI infrastructure as explicit COGS you track monthly.