Skip to content
All posts

The Context Window Arms Race Isn't the Point

Every release brags about a bigger context window. A bigger window doesn't fix the actual problem most people have with long conversations.

Nathan Levine

2 min read

The Context Window Arms Race Isn't the Point

Context window size has become a headline number the same way megapixels used to be for cameras — easy to put on a slide, not actually the thing that determines the output quality. Every model cycle, the number gets bigger, and every cycle, people assume that alone solves the problem they were having with long tasks. It doesn't, and it's worth being clear about why.

Fitting in the window isn't the same as using it well

A model that can technically hold a million tokens of context doesn't necessarily reason well across all of it. Recall degrades unevenly — information in the middle of a long context is well documented to get less attention than information at the start or end. A bigger window raises the ceiling on what you can hand the model; it doesn't guarantee the model will weigh all of it correctly once it's there.

This matters concretely for engineering work: dumping an entire large repo into context and asking a broad question tends to produce worse results than giving a smaller, curated set of relevant files and a specific question. More context is not free — it's a trade between completeness and focus, and past a certain point, focus wins.

What actually helps on long tasks

  • Curated context beats maximal context. Handing over the five files that matter, not the fifty that might, consistently gets better results than relying on window size to compensate for a vague ask.
  • Explicit structure helps more than raw size. A model does better with a long conversation that's been summarized and re-anchored periodically than with the same conversation left to sprawl untouched, even if both technically fit in the window.
  • Task decomposition still matters. A bigger window makes it tempting to hand over an entire multi-step task at once. Breaking it into stages, with checkpoints, still produces more reliable output than trusting the model to hold the whole plan coherently start to finish.

The number that actually matters

Window size is a capacity limit, not a quality guarantee. The more useful question when a new model ships isn't "how big is the window" — it's "how well does it use the context I actually give it, especially in the middle." That's harder to put on a marketing chart, which is exactly why it doesn't show up on one as often as it should.

Thanks for reading. If this was useful, the newsletter below is the best way to catch the next one.

Keep reading

More essays