Agentic AI: Separating What Works Today From the Hype
'Agent' gets attached to every product announcement now. Most of what's actually reliable today is narrower than the pitch decks suggest.
Nathan Levine
3 min read

"Agentic" has become one of those words that shows up on nearly every AI product announcement, regardless of what the product actually does. Somewhere between "chatbot with a plugin" and "fully autonomous employee," the term has stretched to cover almost anything, which makes it worth being precise about what's actually reliable right now versus what's still mostly a pitch.
What's actually working today
A few categories of agentic behavior are genuinely reliable in production, not just in demos:
- Bounded, tool-using loops with human review at the boundary. An agent that reads code, proposes a change, runs tests, and iterates — with a human reviewing before it merges or deploys — is a solid, proven pattern. The autonomy is real, but it's scoped, and the risk is bounded by the review step.
- Narrow, well-defined workflows with clear success criteria. Triaging support tickets, extracting structured data from documents, running a fixed multi-step research task — these work well because "done" is checkable, so failures are visible instead of silent.
- Long-horizon coding tasks with test feedback. Where the agent has a concrete, verifiable signal (tests passing, a build succeeding), it can iterate productively over many steps, because it's not just guessing whether it succeeded.
What's still mostly aspirational
- Fully autonomous, open-ended goals with no checkpoints. "Go improve our product" or "manage this process end to end with no review" still tends to drift, compound small misunderstandings, or optimize for the wrong proxy, because there's no verifiable signal keeping it anchored to what was actually intended.
- Long-running autonomy across ambiguous, shifting priorities. Agents are good at holding a fixed goal across many steps. They're much less reliable at correctly re-prioritizing when the situation changes mid-task in a way that wasn't specified upfront.
- High-stakes actions with no human in the loop. Anything genuinely hard to reverse — financial transactions, production deployments without a review gate, customer-facing communication — is still a place where full autonomy is a liability, not a feature, regardless of how capable the underlying model is.
The useful filter
When evaluating a product or a workflow described as "agentic," the question worth asking isn't whether it uses the word correctly. It's: what's the verifiable success signal, and where's the human checkpoint before something hard-to-reverse happens. Products with clear answers to both are usually the real thing. Products without either answer are usually a chatbot with better marketing.

