FRUIT TREE · AI, SYSTEMS, BUILDING
AI Switching Costs: When Your Model Learns From You, Who Actually Gets Smarter?
I use Cursor, Codex and Claude interchangeably today at zero switching cost. Continual learning will end that — the question is whether the model improves for you alone, or for everyone.
I use three AI tools in the same week — often in the same day. Cursor is my main tool: it drafts my marketing briefs, walked me step by step through installing OpenClaw on a server over SSH, and runs the file-based knowledge system I use to take in something new every day. Codex handles the work that needs a browser — deep research on target domains — because its browsing flow is currently the smoothest of the three. Claude is where I go to understand an industry or a hard problem systematically, or to get unstuck with a different angle; lately also for design, because it works from templates and confirms the mockup with me before building. Moving between them costs me almost nothing. My project context lives in Markdown files inside my repositories: rules files, project briefs, notes the agents maintain for me. Any of the three tools can read them. None of the three tools owns them.
After reading Dwarkesh Patel’s 8 Predictions for the Era of Continual Learning, I started asking why switching is this cheap — and how long it will stay that way. His answer to the first question is convincing. But the piece leaves a second question mostly unexamined, and I think that second question decides who actually wins: when a model improves through your usage, does it improve for you specifically, or does it improve for everyone?
Why switching between AI tools is free today
The reason I can hop between Cursor, Codex, and Claude is that none of them retains anything about me in the model itself. Today’s “memory” features — Cursor’s memories, CLAUDE.md files, custom instructions, retrieval over past chats — are all context injection. The model’s weights are frozen. Everything it “knows” about my projects is text that gets pasted into the prompt at inference time.
Text is portable. If I switch tools tomorrow, I copy my rules files into the new tool’s format and lose close to nothing. The accumulated context is mine, stored in my repository, readable by any competitor.
To be precise, my switching cost isn’t exactly zero — it’s just not about learning. What actually keeps me anchored to Cursor is interface: the left sidebar holds my conversations, outputs, and project folders in one view, and that sense of control is something Codex’s floating file panel doesn’t replicate. That’s an ordinary software switching cost, the kind every product category has, and it’s small enough to break with a week of habit change. Worth naming, because the switching cost this essay is about — a model that has learned things about my work that I cannot export — would be a categorically larger one stacked on top of it.
Dwarkesh’s core argument is that this architecture has a ceiling. Session notes passed between stateless model instances cannot substitute for experience accumulated in weights, the same way notes passed between people cannot substitute for practice. If that’s right, the current regime — frozen weights plus portable text — is temporary, and the tools I switch between freely today will eventually learn from usage at the weight level.
I think the ceiling is real. But I also think it sits higher than his framing suggests — and that changes how long the current regime lasts.
The frozen-weight ceiling is higher than the notes framing suggests
Dwarkesh’s argument treats what carries over between sessions as notes: descriptions of experience that each fresh model instance must read and re-derive correct behavior from. For notes in that strict sense, his point stands.
But that isn’t all a stateless system can accumulate. Lilian Weng’s essay on harness engineering surveys a research line where agents improve by rewriting their own harness — the code around the model that handles planning, tool calls, memory, and evaluation — while the weights stay frozen. The headline result is the Darwin Gödel Machine: a coding agent on a frozen Claude 3.5 Sonnet backbone that iteratively rewrote its own harness code, keeping only versions that scored higher, and pushed SWE-bench Verified from 20% to 50% without touching a single weight.
Thirty percentage points of improvement with frozen weights is hard to square with a low ceiling. The resolution is that what accumulates in a harness includes executable artifacts, not just notes. A debugging procedure becomes a script. A past failure becomes a regression test. A recurring subtask becomes a tool function. These artifacts run at inference time and produce correct behavior directly — the model doesn’t need to have internalized anything, because the competence lives in the artifact.
So the real boundary isn’t “text vs. weights.” It’s externalizable vs. non-externalizable competence:
- Competence that can be compiled into an artifact — procedures, checks, tools, playbooks — accumulates fine under frozen weights. And artifacts, like notes, are portable: they’re files in my repository that any tool can execute.
- Competence that can’t be pre-compiled — judgment on genuinely novel inputs, recognizing an unfamiliar failure pattern, taste — is where Dwarkesh’s argument holds and weight-level learning is irreplaceable.
This matters for the switching-cost story. The territory where continual learning beats portable artifacts is the second category only — narrower than the notes framing implies. Weight-level learning will still arrive, but the zero-switching-cost regime has more room to run than the headline argument suggests, and everything accumulated in the first category stays portable even after it does.
When weight-level learning does arrive, switching stops being free. But how it stops being free depends on where the learning goes.
The question the essay skips: learning for whom?
Dwarkesh touches this in a single clause — there’s a difference between “updating one user’s set of weights” and “pooling all these different weight forks back into the main model” — and moves on. I want to slow down here, because the two mechanisms have completely different consequences, and a third layer sits between them that I haven’t seen discussed anywhere.
Layer 1: your fork — the model improves for you alone
In this mechanism, your usage updates a set of weights (a full fork, or an adapter) that serves only you or your organization. The model that has worked with your codebase for six months is measurably better at your codebase than a fresh instance of the same base model.
This is the layer that creates switching costs, and it is the layer Dwarkesh’s lock-in prediction depends on. Leaving the provider means abandoning weights you cannot export. His comparison is firing an employee with months of organizational context and re-training a replacement from zero.
Note what this layer does not do: it does not make the base model more capable. Your fork’s improvements are, by design, yours. Nobody else benefits, and — the flip side — your fork captures 100% of the learning from your own sessions.
Layer 2: the pool — the model improves for everyone
In this mechanism, learning from millions of deployed sessions flows back into the shared base model. Everyone benefits, including users who contributed nothing.
Unlike Layer 1, this already exists — just in slow motion. Usage data becomes preference signals becomes the next model version, on a cycle measured in months rather than the continuous updates Dwarkesh describes. It’s the reason labs want permission to train on your sessions today.
Here’s the asymmetry that matters: your individual contribution to the pool is statistically negligible — one session among millions. So Layer 2 creates no switching cost for you. The base model got better from everyone’s usage, including yours, but you can walk away tomorrow and lose nothing that was yours specifically. Layer 2 builds a moat between labs (whoever has the most deployment learns fastest), not a moat around users.
Layer 3: the ecosystem — the model improves at your kind of work
This is the layer I find most interesting, and it falls out of one of Dwarkesh’s other predictions — that AI minds will diversify as they learn from different experience — combined with the pooling mechanism.
If models learn from deployment, then a model’s capability profile is shaped by the composition of its deployment. A model whose usage is dominated by coding sessions improves fastest at coding. A model deployed heavily into legal work improves fastest at legal work.
Which means my choice of tool is also a vote. Every session I run in Cursor marginally pushes that model’s learning distribution toward my kind of work. Multiply by millions of users and the products with the most usage in a domain compound their advantage in that domain — not because any individual is locked in, but because the collective usage pattern steers where the shared model gets better. Call it ecosystem lock-in: no single user is captive, but the community’s aggregated choice determines which model becomes the best tool for that community’s work.
Where the line gets drawn — and who captures the value
So the answer to my original question is: both, through different mechanisms, and the split determines the business model.
| Layer | Who benefits | Who is locked in | What the moat is |
|---|---|---|---|
| Your fork | You alone | You — weights aren’t exportable | Per-user switching cost → pricing power |
| The pool | Every user | Nobody individually | Deployment scale → best base model |
| The ecosystem | Users in your domain | The domain community, collectively | Domain-specific compounding |
Before the two-tier bargain, there’s a constraint that decides who can afford Layer 1 at all: inference economics. Dwarkesh’s back-of-the-envelope math puts the optimal inference batch size for a sparse model like DeepSeek V3 above 2,400 concurrent sequences — a set of weights is only served efficiently when thousands of requests decode against it at once. A large company whose employees and agents generate that much traffic can serve its own continually updated fork efficiently; an individual serving a personal fork at batch size 1 could pay a 100x+ compute penalty. So if personalization requires full weight updates rather than low-rank adapters, full-weight forks are economically viable mainly at the organization level, where thousands of employees and agents pool their traffic against one shared fork. Note the unit: the fork belongs to the company, not to any employee. Individual personalization — whether you’re inside an enterprise or on your own — would ride on top as adapters or context: cheaper to serve (adapters from many users can share the base weights within one batch), but shallower. Which means the strongest form of Layer 1 lock-in lands on enterprises first, and individual users like me keep low switching costs for longer than the headline prediction suggests.
The realistic outcome is a two-tier system, and the interesting fight is over where the boundary sits. Labs will want generalizable skills — how to debug a race condition, how to structure a migration — pooled into the base model, because that’s their capability flywheel. Enterprises will want proprietary knowledge — their codebase, their business logic, their data — fenced inside their own fork, because that’s their confidentiality requirement and, increasingly, their negotiating chip.
Dwarkesh predicts labs will use both incentives and restrictions to get training access: discounts for users who allow it, best-model access withheld from those who refuse. I’d sharpen that: the negotiation won’t be binary. It will be about which layer your sessions feed. “Train on the generalizable parts, fence the proprietary parts” is a stable bargain in a way that all-or-nothing isn’t — and drawing that line in a verifiable way is, as far as I can tell, an unsolved technical problem.
What I’m doing about it in the meantime
Until weight-level personalization ships, the practical strategy for an individual user is straightforward: keep your accumulated context in portable form, outside any single product.
Concretely, what I do today:
- Project context lives in Markdown files inside the repository, organized into three folders: Rules (conventions and standards the agents must follow), Skills (reusable skill definitions), and per-project context files that record goals, decisions, and current progress — written explicitly so that any agent, in any tool, can pick up where the last one left off. None of it lives in any one product’s memory feature.
- Tool-specific configuration is thin: a rules file per tool that mostly points at the shared context.
- When I run Codex or Claude against the same work, I point them at these folders. The only real friction is project understanding — a tool that hasn’t been living in the repository needs orientation. My fix is a system-level “read this first” file in the folder, which every agent reads at the start of a session and updates at the end.
This keeps my switching cost at zero for as long as the current regime lasts. The moment a provider offers weight-level personalization that visibly compounds — where month three with the tool is measurably better than month one in a way I can’t replicate by pasting context — the calculus flips, and picking one provider becomes the rational move. That moment is the thing to watch for. Everything before it is positioning.
FAQ
Do AI models learn from my conversations today? Not in real time. Your sessions may be used as training data for future model versions (Layer 2, on a months-long cycle, subject to your data settings), but the model you’re talking to has frozen weights. Its apparent “memory” of you is retrieved text, not learned weights.
What are AI switching costs right now? Close to zero for individuals. Accumulated context (rules files, memory exports, project docs) and executable artifacts (scripts, tests, skill definitions) are files that can be moved between products. Real switching costs arrive only when providers update weights per user or per organization.
Will continual learning make one AI provider win everything? It concentrates advantage — deployment scale feeds the pool, and the pool improves the product — but the ecosystem layer suggests domain-level winners rather than one general winner: the model with the most coding deployment becomes the best at coding, which may differ from the model that dominates other domains.
References: Dwarkesh Patel, 8 Predictions for the Era of Continual Learning, August 2026. Lilian Weng, Harness Engineering for Self-Improvement, July 2026.