A strategic view of vendor pricing games and how buyers should respond.
Inside the AI Arms Race: LLMs, Economics, and Strategy
CXOTalk with Nate B. Jones 2025
Open cxotalk.com →You pay per token (roughly per word) in and out, so cost scales with usage; a small feature can run on tens of dollars a month while a chat-heavy product can quietly eat 30-60% of revenue. The good news is prices for the same capability keep falling about 10x a year, and three habits (prompt caching, routing easy requests to cheap models, and batching non-urgent work) routinely cut bills 50-80%. Instrument cost per user per feature from day one so pricing decisions are based on data, not fear.
19 resources.
A strategic view of vendor pricing games and how buyers should respond.
CXOTalk with Nate B. Jones 2025
Open cxotalk.com →Willison narrates the price collapse and capability gains in one listenable year-in-review.
Latent Space with Simon Willison Jan 2025
Open latent.space →The definitive chart: equivalent-quality inference gets 10x cheaper per year, which changes what you can afford to build.
Guido Appenzeller (a16z) Nov 2024
Open a16z.com →The five-minute version of the cost-decline argument, useful for convincing a worried co-founder.
a16z Nov 2024
Open x.com →Cached reads cost 10% of normal input price; the single highest-leverage cost lever, from the primary source.
Anthropic 2025
Open platform.claude.com →A flat 50% discount for anything that can wait up to 24 hours, like digests, enrichment, and backfills.
OpenAI 2025
Open developers.openai.com →Grounded numbers on what AI COGS actually look like across real startups, not vibes.
Tanay Jaipuria (Wing VC) 2025
Open tanayj.com →Fresh survey data showing margins improving from 41% to 45% and what the improvers did differently.
Upstarts Media 2025
Open upstartsmedia.com →Routing, compaction, prompt trims, caching, batching: each lever quantified so you know what to do first.
Morph 2025
Open morphllm.com →Engineering-level detail on token budgets and semantic caching from an infra company that measures it.
Redis 2025
Open redis.io →Diagram-first explanations your whole team can absorb in ten minutes.
Divy Yadav (Towards AI) Apr 2026
Open pub.towardsai.net →A gateway gives you routing, fallbacks, and spend visibility in one move; this compares the options.
Helicone 2025
Open helicone.ai →A sober side-by-side of per-token prices across the big three, including budget tiers.
IntuitionLabs 2025
Open intuitionlabs.ai →A builder's receipts: one technique, a 90% real bill reduction, with the code that did it.
Du'An Lightfoot 2025
Open medium.com →Actual production numbers for agent workloads, the costliest and least predictable AI feature type.
Nikhil Sood 2025
Open medium.com →Explains why output tokens cost more than input and how billing categories really work.
Introl 2025
Open introl.com →Understand your supplier's economics to predict where API prices go next.
Azeem Azhar (Exponential View) 2025
Open exponentialview.co →Practical recipes for attributing AI spend per user and per feature, which is how you find the leaks.
Helicone 2025
Open docs.helicone.ai →Live price, speed, and latency comparison across 500+ endpoints; check before you commit spend.
Artificial Analysis 2026
Open artificialanalysis.ai →The same ground, over in Build the product, our Starting Up track.