Skip to content
All posts

Model Pricing Isn't Just a Cost Line — It Changes What You Build

As per-token costs drop, the interesting shift isn't cheaper bills. It's that entire product ideas that didn't pencil out a year ago suddenly do.

Nathan Levine

2 min read

Model Pricing Isn't Just a Cost Line — It Changes What You Build

Every time a new model generation ships with a lower price per token, the conversation defaults to "great, cheaper bills." That's true but misses the more interesting effect: pricing isn't just a cost line on an existing product, it's a constraint that determines which products are viable to build at all.

Price determines the design, not just the budget

When a capability is expensive per call, it gets used sparingly — a single summarization pass, a one-shot classification, gated behind a feature that only fires when it clearly justifies the cost. When the same capability gets meaningfully cheaper, the design space opens up: you can afford to call a model multiple times per request, run a verification pass on its own output, or use it for something that used to be "too expensive to justify" like real-time suggestions on every keystroke instead of on submit.

This is the same dynamic that played out with cloud compute and with storage — the interesting change wasn't that existing systems got cheaper to run, it's that systems that didn't make sense at the old price point became normal at the new one.

What actually changes as prices drop

  • Multi-call patterns become affordable. Self-critique loops, multiple draft-and-refine passes, or running the same task through a cheap model as a first pass and a stronger model as a check — patterns that were cost-prohibitive at scale become the default architecture.
  • Background and speculative work becomes viable. Pre-computing likely next actions, or running lightweight analysis continuously instead of on-demand, only makes sense once the per-call cost is low enough that being wrong most of the time is cheap.
  • The smallest capable model becomes a real engineering decision, not an afterthought. With a wider spread between the cheapest and most capable models, picking the right one per task — not just defaulting to the biggest — becomes a meaningful lever on both cost and latency.

The mistake to avoid

Treating a price drop as purely a cost optimization on your current architecture is underusing it. The more valuable question when a new pricing tier lands isn't "how much do we save" — it's "what can we now afford to build that we couldn't before." That's usually where the actual product opportunity is, and it's easy to miss if you're only looking at the invoice.

Thanks for reading. If this was useful, the newsletter below is the best way to catch the next one.

Keep reading

More essays