We plan to bring eChai across 100 cities in India. It starts with eChai Startup Demo Day on 26 September, all in person. 11 cities confirmed, 313 founders registered. Any city that reaches 20 interested founders is on too. See your city
How the best do it

My AI features cost real money every time someone uses them. How do I stop gross margin from collapsing?

Accept the new baseline first: AI native products are running somewhere near 50 to 65 percent gross margin, not the 80 to 90 percent that defined the last decade of SaaS, and inference cost tends to rise as a share of spend rather than fall as usage grows. Then get instrumented. You cannot manage a cost you cannot see per customer, so meter inference per account and watch the tail, because a small minority of users typically drive the bulk of consumption. After that the levers are ordinary: route cheap work to cheap models, cache aggressively, cap or credit the heaviest actions, and put a consumption component in your price so the customers costing you the most also pay the most. If you earn in rupees and pay for compute in dollars, add currency to that list, because a weak rupee quietly eats a margin point at a time.

Go deeper

6 resources, 2 India-specific, 6 link-checked.

📰 Newsletter
✓ Link checked Freemium Advanced

The margin defence most companies are actually reaching for. Poyar's number is the one to remember: roughly 70 to 80 percent of token consumption comes from about 10 percent of users, which is why a flat seat price on an AI product bleeds quietly.

Why everyone's switching to AI credits

From Growth Unhinged by Kyle Poyar 14 min read

  • 70 to 80 percent of token consumption comes from about 10 percent of users, which is what breaks flat pricing.
  • Microsoft, Salesforce, Cursor and OpenAI all moved to credit based models within months of each other.
  • Model costs are not falling as promised: GPT-5 output is around 10 dollars per million tokens, close to GPT-4o in early 2024.
  • Credits exist as a guardrail because AI gross margins are already thin, not because customers asked for them.
Open growthunhinged.com
📄 Article
✓ Link checked Free Advanced

Gives you the current benchmarks to set expectations with your board (roughly 45 to 53 percent depending on how much of the stack you own) and four concrete moves: pick a margin posture deliberately, meter cost per customer, route to cheaper models by default, price with usage.

AI Gross Margins: Why the 90% SaaS Benchmark Is Gone

From Jeff Brokaw by Jeff Brokaw 12 min read

  • ICONIQ's survey of about 300 software executives puts AI product gross margins at 45 to 53 percent in 2026.
  • The 90 percent SaaS margin was always mythology, 80 percent was the real industry assumption.
  • Every AI request carries real variable cost: inference, retrieval infrastructure and human review.
  • Cost curves diverge: a fixed capability gets 5 to 10x cheaper a year while frontier model pricing rises 3 to 18x.
Open jeffbrokaw.com
📄 Article
✓ Link checked Free Advanced

One metric you can start tracking this month: AI product revenue divided by inference cost. It moves before your gross margin does, which makes it the early warning you want rather than the postmortem you get from the P&L.

How to Calculate the Inference Efficiency Ratio

From The SaaS CFO by Ben Murray 8 min read

  • Inference efficiency ratio is AI product revenue divided by inference cost for the same period.
  • For AI infused SaaS, 8:1 is the floor and 10:1 or better is healthy, below 5:1 is a warning.
  • For AI native products the healthy zone is 5:1 or better, since inference is structurally about 20 percent of revenue.
  • Unlike gross margin, which lags, IER is a leading signal: model routing, prompt caching and tiered pricing moved one example from 4.4:1 to 8.0:1 at the same revenue.
Open thesaascfo.com
📄 Article
✓ Link checked India Free Intermediate

Names the version of this problem that is specific to Indian companies: revenue in rupees, inference billed in dollars, and a currency move you did not budget for eating margin. If you sell in INR and run on foreign models, this is your risk in one page.

India's AI Inference Problem

From Inc42 by Team Inc42 8 min read

  • India is projected to need around 7 GW of AI compute capacity by 2030.
  • Inference is becoming the largest operating cost of AI native products because it scales with usage, not customer count.
  • Indian companies earn in rupees and pay for inference in dollars, so FX moves hit margins directly.
  • Yotta, CtrlS and Reliance Jio are building out GPU clusters to cut that foreign dependency.
Open inc42.com

Browse all 796 resources →

The same ground, at another level

How pricing and packaging reads from a different seat.

Terms in this answer

People also ask

Also in Starting Up

The same ground, over in Money, pricing & model, our Starting Up track.

Also in D2C

The same ground, over in Money, pricing & unit economics, our D2C track.

eChai Partner Brands