The first day with a new AI model can feel extraordinary.
You give it something the previous model struggled with. It understands the request, makes a useful connection, and produces an answer that needs surprisingly little correction. You try another task. That works too.
Finally, you think. This is the one.
Then we do what organisations have always done when something proves competent.
We give it more work.
The requests become more ambitious. The neat little test prompts disappear. We bring it into actual work, where the requirements are unfinished, the evidence conflicts, three people have different ideas about what “done” means, and getting something wrong has consequences.
Before long, you are correcting it, repeating yourself, and wondering why it keeps missing the point. Meanwhile, the usage allowance that looked perfectly reasonable when you were asking it to tidy paragraphs is now disappearing with the urgency of a departmental budget in the final week of the financial year.
The obvious explanation is that the model has been “dumbed down.”
Sometimes the product really has changed. In April 2025, for example, OpenAI acknowledged that an update had made GPT-4o excessively agreeable and rolled it back. Genuine regressions should not be dismissed as users imagining things. 1
But a model does not have to deteriorate for our experience of it to deteriorate.
There is another explanation worth considering: success earns the model more trust, trust earns it a bigger job, and the bigger job exposes limitations we had not previously encountered.
We may be doing to AI what organisations have long done to their best employees.
The model passes the test. Then we change the test.
Consider the difference between asking an AI to improve a paragraph and asking it to produce a publishable article.
The first task has a relatively clear boundary. The words are already there. You can judge whether the revision is better.
The second requires much more: a defensible argument, accurate evidence, an understanding of the audience, appropriate emphasis, and enough judgement to know what should be left out.
Now extend the request again. Ask it to research the subject, choose the argument, write the article, adapt it into a campaign, and ensure everything fits your organisation’s positioning.
It may still look like a writing request.
It is no longer the same job.
This is the bit we are strangely good at ignoring. The interface has not changed. There is still a box. We still type words into the box. Therefore, apparently, everything involving words is one task.
The same distinction applies to software. Fixing a clearly identified function is different from diagnosing an intermittent failure across an application. Producing working code is different from changing a system without breaking assumptions elsewhere.
There is experimental evidence for this unevenness. In a study involving 758 Boston Consulting Group consultants, AI assistance helped participants complete suitable tasks about 25% faster. Yet on a task selected to fall outside the model’s capabilities, AI-assisted participants were 19 percentage points less likely to reach the correct answer than those working without it. The researchers described a “jagged technological frontier”: success on one task did not reliably establish competence on another. 2
That study does not prove that launch-day enthusiasm fades because users assign harder work. It establishes something narrower and important: changing the assignment can reverse the value of using the same AI.
A model can therefore remain just as capable while becoming less dependable for the work we now expect it to do.
On the first day, the question was:
“Can it help me with this?”
Later, the question becomes:
“Why can’t it take care of this without me?”
Those are different standards.
We just tend to promote the model between them without mentioning it.
The Peter Principle, at software speed
The Peter Principle describes the tendency to promote people on the strength of their current performance until they reach a role in which they are no longer effective. Research by Alan Benson, Danielle Li, and Kelly Shue, using data from 131 firms, found evidence of companies favouring strong sales performers for promotion even when other characteristics better predicted managerial performance. 3
The important point is not that successful people somehow become stupid.
It is that the new role may require different abilities. Being excellent at winning customers does not automatically establish that someone can coach a team, allocate resources, or resolve conflict.
The official version is that we recognised excellence and rewarded it.
The less flattering version is that someone was very good at Job A, so we stopped letting them do Job A.
We can make the same mistake with AI, only much faster.
It summarises documents well, so we promote it to analyst. It produces a convincing analysis, so we promote it to adviser. It gives useful advice, so we ask it to make decisions and carry out the work.
No interview. No probation period. No awkward conversation about whether it actually wants management responsibilities. One good Tuesday afternoon and congratulations: you now own the workflow.
Each promotion feels justified by the previous success. Yet each may involve an untested jump.
A model that helps you think is not necessarily ready to decide on your behalf. A model that can complete a task with close guidance is not necessarily ready to own an entire workflow.
Call this an AI version of the Peter Principle: we keep expanding the model’s responsibilities until we find the work it cannot reliably handle.
Then we judge it by that work, rather than by the capabilities that earned our trust in the first place.
Promotion and overload are different problems
There is a second organisational mistake hiding inside this pattern.
Sometimes a person is given the wrong job.
Sometimes they have the right skills but are given too much work.
Imagine a dependable employee who handles difficult assignments well. Whenever something urgent appears, it goes to them. Whenever another team falls behind, they help. Their reward for being reliable is an expanding collection of responsibilities.
This is one of management’s more durable operating systems: if someone reliably carries things, find more things.
Eventually, something slips.
Calling this incompetence would confuse ability with workload. Burnout is not proof that someone has reached the limit of their talent. The World Health Organization describes it as a consequence of chronic workplace stress that has not been successfully managed, involving exhaustion, detachment or cynicism, and reduced professional efficacy. 4
That distinction matters when drawing the AI comparison.
Models do not burn out in the human sense.
They do not become exhausted, cynical, or sit through a Sunday evening wondering how “one small additional responsibility” somehow became twelve.
But they do operate within limits—and those limits are not all the same.
One is capability: whether the model can reliably perform the task.
Another is context: how much relevant information it can use effectively during the work. A longer conversation can accumulate instructions, corrections, discarded approaches, and tool results. Anthropic’s guidance on context engineering explicitly warns that model performance can degrade as context grows; a larger context window does not guarantee equally effective use of everything inside it. 5
A third is the usage allowance: how much work the service permits before access is restricted or additional usage must be purchased. That is separate from the length of an individual conversation. 6
These are different diagnoses.
A difficult task does not become easy because your allowance resets. Buying more usage does not automatically improve judgement. Starting a cleaner conversation may reduce contextual clutter, but it does not supply a capability the model lacks.
This sounds obvious when written down. In practice, these problems are often experienced through the same highly technical diagnostic instrument:
“It was better yesterday.”
The analogy with people is therefore useful only up to a point. Human burnout requires attention to working conditions, support, and recovery. AI requires an appropriate task, usable context, and sufficient resources.
The shared mistake is treating demonstrated competence as evidence of unlimited capacity.
Why the usage allowance seems to disappear faster
This also offers a plausible explanation for the feeling that each new release gives us less usable time.
Think about what happens after a model earns your confidence.
You stop using it only for occasional questions. You give it longer documents, larger projects, and work you previously would not have delegated. You ask it to investigate, compare, revise, verify, and try again.
The model has proved that it can carry boxes.
Naturally, we respond by locating every box in the building.
You may type fewer messages while requesting substantially more processing.
Anthropic’s usage documentation states that consumption depends on factors including conversation length and complexity, the model, the features used, and the selected effort level. A message is therefore not a fixed unit of work. 6
With an agent, the gap between what you type and what the system does can be particularly large. A short instruction may initiate repeated file reads, tool calls, edits, and checks. Claude Code’s documentation explains that tool interactions generate further requests and that long conversation histories continue contributing to usage, even when caching reduces their cost. 7
Reasoning can add another layer. OpenAI’s API documentation notes that reasoning tokens count towards output usage even though they are not displayed as ordinary answer text. A short visible response need not represent a small amount of processing. 8
This creates a counterintuitive possibility: a more useful model can make the same allowance feel smaller because you have found more valuable—and more demanding—ways to spend it.
That is not evidence that every newer model consumes more resources, or that providers never change their limits. It means that reaching a cap sooner does not, by itself, establish that the allowance was reduced.
The useful comparison is not simply how many prompts you sent.
It is how much verified work you completed, including the effort spent correcting unsuccessful attempts.
A model that uses more resources but reliably completes the job may be better value. One that spends the allowance repeatedly failing may not be.
Preference is not performance. Neither is prompt count.
Why the honeymoon can get shorter
The first time you encounter a capable AI assistant, you may approach it cautiously.
You experiment with small tasks. You discover what it does well. You learn where it gets confused. You gradually find the boundary between “surprisingly useful” and “why did I let it touch this?”
A later model does not necessarily receive that gentle introduction.
It inherits your existing workflows, your accumulated expectations, and the backlog of work the previous model could not handle.
Perhaps the last model could draft a report but struggled to reconcile conflicting evidence. The next one demonstrates better analysis, so you immediately give it the whole research assignment.
Perhaps the previous coding agent needed constant supervision. The new one solves a difficult bug, so you immediately ask it to work independently across the project.
Yesterday’s breakthrough becomes today’s minimum requirement remarkably quickly.
This suggests a reason the honeymoon can shorten without the technology declining faster: we arrive at each release better prepared to reach its limits.
The capability improvement may be real. But instead of continuing to assign yesterday’s work and enjoying the improvement, we spend that improvement on greater ambition.
This is a hypothesis about the experience, not a measured rule governing every release.
It may not explain your experience with a particular model at all. Products can change. Limits can change. Models differ in capability, context behaviour, and usage allowances. Genuine regressions happen.
But the hypothesis explains how two impressions can both be sincere: the model really was impressive at first, and it really is struggling now.
What changed may be the distance between the task and the model’s dependable capabilities.
We did not necessarily discover that the new employee was secretly terrible.
We may simply have completed the promotion cycle unusually efficiently.
Before blaming the performer, inspect the assignment
The practical response is not to stop trusting AI or to keep capable people permanently in small roles.
It is to make expansion deliberate.
For AI, preserve a small set of representative tasks, with the original inputs and clear standards for success. When a model seems worse, repeat those tasks under comparable conditions, preferably in fresh conversations. Repeat enough times to avoid treating one unusually good or bad answer as the whole story.
This gives you a better basis for separating a genuine regression from a harder workload, a cluttered conversation, or a higher standard of success.
Then test new responsibilities as new responsibilities.
Strong drafting is evidence for drafting.
Reliable analysis is evidence for analysis.
Neither should silently become permission for unrestricted decision-making.
The same applies when the tool moves from answering questions to operating inside a workflow. If the job now includes reading files, choosing an approach, editing code, checking the result, recovering from errors, and deciding when the work is finished, you have changed more than the prompt.
Treat it accordingly.
For people, apply the same discipline without forgetting the human difference. Before deciding that someone has lost their edge, examine how their role has changed, what has been added, what support has disappeared, and whether the workload is sustainable.
A promotion should come with an assessment of what the new role requires. Additional responsibility should come with an explicit decision about what will be removed, deferred, or supported.
Otherwise, “more responsibility” has a tendency to mean “everything you already do, plus this”.
Success should justify greater trust.
It should not remove the need for boundaries.
The next time a brilliant new AI seems to have become disappointing, ask two questions rather than one: has the system changed, and has the job changed?
And when a once-dependable employee begins to struggle, ask those questions before reaching for a label.
Sometimes capability has declined. Sometimes the match was wrong. Sometimes the workload has become unreasonable.
But sometimes the performer is still capable of everything that first impressed us.
We have simply stopped asking for that—and started asking for everything else.


