How I Actually Pick Which Model to Use
With a new model shipping from someone every few weeks, picking one for a task has become its own small skill. Here's the checklist I actually use.
Nathan Levine
3 min read

With a new model landing from one lab or another every few weeks, "which model should I use" has quietly become its own recurring decision, separate from the actual task at hand. It's easy to default to whichever one is newest, or whichever one is loudest in your feed that week. Neither is a great method. Here's the actual checklist I run through.
Start with the failure mode, not the capability
The first question isn't "which model is smartest" — it's "what does a wrong answer cost me here." A model drafting a first-pass blog outline can be wrong cheaply; I'll just edit it. A model writing a database migration that runs against production data needs a much higher bar, and probably needs a slower, more deliberate model plus a human review step, regardless of which one is fastest.
Matching the model's reliability to the cost of being wrong is a bigger lever than matching it to a benchmark score.
Then check latency against how you'll actually use it
A model that's slightly more capable but meaningfully slower is often the wrong trade in an interactive loop — if I'm iterating rapidly on a small edit, a fast model that's "good enough" beats a slower one that's marginally better, because the iteration speed itself improves the final result more than the single-shot quality does. For anything I'm kicking off and walking away from — a long refactor, a batch job — the latency trade flips, and I'll take the slower, more capable model.
Test it on my own task, not a demo task
Before committing to a model for a recurring workflow, I run it on a real example from that workflow — not a toy prompt, the actual messy input I'll be feeding it in production. Generic capability doesn't reliably predict performance on a specific, idiosyncratic task, and the only way to know is to actually try it there.
Weigh cost against volume, not against the task in isolation
A cost difference that's negligible for one-off use becomes a real budget line at scale. I ask how often this call actually runs before deciding whether the cheaper or the more capable model is the right default — a model that's 3x more expensive per call is a non-issue at ten calls a day and a real cost at ten thousand.
Leave room to be wrong
None of this is a one-time decision. Model landscape shifts fast enough that a choice that was right two months ago might not be right now. I treat model selection as a setting to revisit periodically, not a decision to make once and forget — the same way I'd revisit a dependency choice, not agonize over it once and never look back.


