Stop Chasing the Model

A premium futuristic illustration showing a central AI model surrounded by connected systems for memory, context, workflows, guardrails, evaluation, and human oversight, representing how AI value now comes from the system around the model.

The model war is slowing down. Something more interesting is taking its place.

There’s a particular kind of exhaustion setting in among people who work closely with AI. Not exhaustion with the technology itself — but with the news cycle around it. Another model launch. Another benchmark. Another claim of “unprecedented capability.”

And yet, underneath all that noise, something genuinely interesting is happening. It’s just not where most people are looking.

The model race — that breathless sprint to out-benchmark everyone else — is quietly plateauing. Not because progress has stopped, but because the easy wins are mostly gone. The jumps that used to feel like leaps now feel like shuffles. Labs are spending more to move less. And the practitioners who actually build things with these models? Most stopped caring about raw benchmarks months ago.

What they care about now is something harder to headline but far more consequential: what you wrap around the model.


The Bottleneck Has Moved

For a couple of years, the dominant assumption was simple: better model, better outcome. So everyone chased the model. Bigger context windows. Lower hallucination rates. Faster inference. All of that mattered — and still does.

But frontier models are now genuinely good. Good enough that the bottleneck for most real-world applications has quietly shifted from the model itself to everything around it: the workflow design, the context quality, the guardrails, the memory architecture, the integration with systems people already use.

Studies on what’s being called “context rot” show that even the most capable frontier models degrade reliably as input length grows. The drop isn’t dramatic at first. But it’s consistent, and it’s there in every model tested. A model’s advertised context window is not the same thing as its useful context window — and teams building on these systems are learning that the hard way.

The implication is subtle but important: how you feed a model information is now as important as which model you choose. The teams figuring this out early — who think carefully about what goes into context, in what order, with what structure — are getting meaningfully better results than teams chasing the next model release.


Memory Is Becoming a Moat

Here’s one of the most underappreciated shifts happening right now: AI companies are quietly turning memory into a lock-in strategy.

On the surface, it looks like a feature. Anthropic extended memory to free users. Google lets you import conversation history from other AI apps into Gemini. These feel like small quality-of-life improvements. They are not.

What’s actually happening is that memory is being repositioned as a portable asset — and a source of competitive advantage. Once an AI system knows your working style, your preferences, your past decisions, your team’s context — switching to a different tool stops being a model comparison. It becomes a question of how much you’re willing to leave behind.

For individuals, this might just feel like convenience. For enterprises evaluating AI tools over a three-year horizon, it’s a different calculation entirely. The model you choose today may matter less than the memory you accumulate in it over time.

Nobody in procurement is thinking about this yet. They probably should be.


The Agentic Dream Is Running Into Reality

2025 was confidently declared “the year of the agent.” It wasn’t, quite. But 2026 is genuinely different in one important way: agentic AI is no longer a demo category. It’s a deployment category. Teams are shipping agents into real workflows, and the lessons coming back are humbling.

The honest picture, from people building in production: agents are brittle in ways that are hard to predict. They handle the use cases you designed for reasonably well. They handle edge cases with the confidence of someone who doesn’t know they’re at the edge. And because they operate across multiple steps, a small error early in a chain compounds in ways a single-prompt mistake never would.

The teams making the most progress with agents are the ones who’ve stopped thinking about them as autonomous and started thinking about them as supervised. Constrained to narrow domains. With humans in the loop at the moments that matter. That’s not a failure of the technology. It’s a more honest and more useful framing of what agents can actually do right now.

Give an agent enough autonomy and eventually something breaks. The skill is in designing exactly how much rope to give it.


The Real Competitive Edge Is Getting Boring

Here is the part that nobody wants to write a think piece about, because it doesn’t make for a good headline: the organizations pulling ahead with AI right now are mostly winning on fundamentals.

Not the most sophisticated model stack. Not the flashiest agentic pipeline. They’re winning because they invested in clean data. Because their teams understand the business context well enough to know what the AI output actually means. Because they built evaluation processes so they can tell when the system drifts. Because they’re patient enough to let a narrow, well-scoped use case compound over time rather than launching ten pilots that go nowhere.

This is the part of the AI story that doesn’t travel well on LinkedIn. But it’s the part that separates organizations genuinely building capability from the ones accumulating press releases.

The AI advantage in 2026 won’t belong to whoever deploys the most. It will belong to whoever deploys thoughtfully.


What This Means in Practice

If you’re building with AI right now, the most useful question isn’t which model to use. It’s: how good is the context we’re feeding it? How well does the system understand the actual work? Where are the humans, and what are they actually checking?

If you’re evaluating AI tools for your team or your organisation, the memory question deserves serious attention. Not just “does it have memory?” but “what does memory mean for our data governance, our vendor dependency, our switching costs in three years?”

And if you’re watching this space from a distance, wondering when it will settle down enough to engage with — this is probably the closest it’s going to get to settled for a while. The noise is still loud. But the signal is getting clearer.

The models are good enough. The question now is whether the people and systems around them are.


If this was useful, forward it to one person who’d appreciate it.

Read this next

One essay a week. No hype.