Skip to content
All posts

A New Model Shipped and Your Prompts Broke. Now What?

Upgrading to a newer model is rarely a free win. If your prompts and evals aren't versioned, you won't notice what quietly changed until a user does.

Nathan Levine

3 min read

A New Model Shipped and Your Prompts Broke. Now What?

The instinct when a new, better-benchmarking model ships is to swap it in immediately — better model, same prompt, free upgrade. Sometimes that's exactly what happens. Often, something quieter breaks: a prompt that was tuned around a previous model's specific quirks stops producing the same output shape, and nobody notices until a downstream parser chokes on a response that looks slightly different than before.

Why prompts are more model-specific than they feel

A prompt isn't just instructions — it's instructions plus an implicit model of how the model on the other end will interpret them. Models differ in how literally they follow formatting instructions, how verbose they default to being, how they handle contradictory instructions, and how they weight instructions placed early versus late in a prompt. A prompt tuned against one model's tendencies can produce a subtly different output shape on another, even when both models are, in the abstract, "better."

This shows up most in the places people notice least: JSON that used to come back clean now has a stray explanatory sentence before it, a summary that used to be three bullet points is now a paragraph, a classification task that used to return one of three fixed labels occasionally returns a close paraphrase instead.

Treat model upgrades like dependency upgrades

The teams that don't get burned by this treat a model swap the same way they'd treat bumping a major dependency version — not something to do silently in production:

  • Version your prompts alongside your code. If a prompt lives in a string buried in a function with no history, you have no way to know what changed when output quality shifts.
  • Keep a small eval set for anything that matters. A handful of representative inputs with expected output shapes, run against any model swap before it ships, catches most of this class of regression before a user does.
  • Roll out gradually, not all at once. Shadow-testing a new model against real traffic, or a staged rollout, surfaces the edge cases that a handful of manual spot-checks won't.
  • Re-tune, don't just re-point. A prompt written for one model's tendencies often needs adjustment for another's, even between versions from the same lab. Budget the time instead of assuming zero-cost migration.

The upgrade is still usually worth it

None of this is an argument against upgrading — a newer model is usually a real improvement, and staying on an old one indefinitely out of caution has its own cost. It's an argument against treating the swap as risk-free just because the benchmark chart went up. The chart doesn't know about your specific prompt, your specific parser, or your specific users. Your eval set does.

Thanks for reading. If this was useful, the newsletter below is the best way to catch the next one.

Keep reading

More essays