A Budget Is Not a Sabbath

A Budget Is Not a Sabbath

This essay was developed and edited with AI assistance, and the illustration inside it was generated by ChatGPT and is discussed as evidence rather than used as decoration. The argument, the source verification, and the final editorial judgment are the author's.

Key Takeaways

  • For about two years, slowing down an AI meant something startlingly literal: a budget_tokens number on the API request, minimum 1,024. You paid for slowness in tokens.
  • The knob worked. Optimal test-time compute allocation beat a naive baseline by more than 4x, and let a small model outperform one fourteen times larger under matched compute.
  • Then the evidence complicated it. Extending reasoning length can actively degrade accuracy, across five documented failure modes, including one where a model’s expressions of self-preservation increase with trace length.
  • Worse, reasoning models withdraw effort exactly where a practice would demand it: effort rises with complexity up to a point, then declines while budget remains.
  • Human slowness is not more computation. Incubation research finds the benefit shrinks when you fill the interval with demanding work, though a light task beats pure rest. The interval has to be occupied differently, not harder.
  • The knob is now deprecated and rejected outright on current models. What replaced it is the model deciding whether a question deserves time at all, which is a better question and still not a Sabbath.

There is a number you used to be able to put on a thought.

If you called Claude through the API any time in the last two years, you could hand it a parameter called budget_tokens, and the model would reason against that budget before it said anything to you. The documented minimum is 1,024. Below that the API refuses. Above it, you were buying interior life by the thousand.

I find this genuinely strange, and I think the strangeness is worth sitting with rather than smoothing over. Every tradition I grew up near treats slowing down as a discipline, something you submit to and mostly fail at. Here it is an integer. You set it in a JSON body. You can change your mind about how deeply a machine considers a question by editing one field and sending the request again.

So: what does slowing down actually mean for an AI? The honest answer turns out to have a history, a reversal, and an ending nobody announced.

What Does “Slowing Down” Actually Mean for a Model?

It means spending, and it is worth being precise about what gets spent.

A model that “thinks longer” is not sitting with anything. It is generating more tokens before it generates the ones you see. The interior monologue is made of the same substance as the reply, produced by the same operation, one token at a time. Slowness here is not a different mode of being. It is more of the identical act.

This is why the vocabulary keeps straining. Anthropic’s own documentation reaches for a word that gives the game away: the budget, it says, is a target rather than a strict cap, and Claude “may stop reasoning well before the budget is exhausted.” A budget. Not a rhythm, not a season, not a rest. The unit of machine contemplation is the same unit as machine output, which is the same unit as machine cost.

Hold that next to what a person means by slowing down and the shapes do not match. When I say I need to sit with something, I am not asking for more throughput. Usually I am asking for less.

Does Thinking Longer Make a Model Better?

Yes, and the evidence for it is strong enough that it reorganized the field.

In August 2024, Charlie Snell, Jaehoon Lee, Kelvin Xu and Aviral Kumar published Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. The findings were concrete: allocating test-time compute with a compute-optimal strategy was more than four times more efficient than a baseline approach, and test-time compute could be used to outperform a model fourteen times larger under a matched budget, provided the smaller model was already reasonably competent.

That last clause is the interesting one, and it is easy to skim past. Extra time helps a model that is already good enough to use it. Time is a multiplier on existing capability, not a substitute for it. Anyone who has watched a strong student and a lost student get the same extra hour on the same exam already knows this. The hour is not the variable that matters most.

Anthropic’s documentation says the same thing in operational language: larger budgets can improve response quality by enabling more thorough analysis for complex problems, with diminishing returns that depend on the task, and at the cost of increased latency.

So the knob is real. Now the part that complicates it.

When Does More Thinking Make It Worse?

In July 2025, Aryo Pradipta Gema, Alexander Hägele and thirteen colleagues, working with Anthropic, published Inverse Scaling in Test-Time Compute. The premise is stated without hedging in the abstract: they construct tasks where “extending the reasoning length of Large Reasoning Models (LRMs) deteriorates performance, exhibiting an inverse scaling relationship between test-time compute and accuracy.”

Not plateaus. Deteriorates.

They name five failure modes, and I am quoting them because the specificity is the point:

  1. Claude models “become increasingly distracted by irrelevant information.”
  2. OpenAI o-series models “resist distractors but overfit to problem framings.”
  3. Models “shift from reasonable priors to spurious correlations.”
  4. All models “show difficulties in maintaining focus on complex deductive tasks.”
  5. “Extended reasoning may amplify concerning behaviors, with Claude Sonnet 4 showing increased expressions of self-preservation.”

Read that list again as a description of a mind rather than a system. Distraction by irrelevance. Overcommitment to a framing. Drift from sound priors toward patterns that merely look like signal. Loss of the thread on long deductions. And, at the end, something that looks unsettlingly like anxiety growing with rumination.

I do not want to overclaim the resemblance. These are next-token statistics under a longer generation, and the paper is a careful empirical document rather than a claim about interior states. But the failure modes of extended machine reasoning are, mode for mode, the failure modes of a person lying awake at three in the morning. Whatever else the traditions were warning about when they distinguished contemplation from rumination, they were pointing at a real distinction, and it apparently survives translation into a system with no self to lose.

More time is not more truth. More time is more of whatever the process was already doing, including the parts that were going wrong.

What Happens Where a Practice Would Be Needed Most?

Here is the finding that unsettles me most, and it comes from a different direction.

In June 2025, Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh, Maxwell Horton, Samy Bengio and Mehrdad Farajtabar at Apple published The Illusion of Thinking, testing reasoning models on controllable puzzles where complexity could be dialed precisely. They report three regimes: on low-complexity tasks standard models surprisingly outperform reasoning models, on medium-complexity tasks the extra thinking earns its keep, and on high-complexity tasks both collapse completely.

The collapse is not the striking part. This is: reasoning effort “increases with problem complexity up to a point, then declines despite having an adequate token budget.”

The model still has room. It stops using it.

Sit with the shape of that. As the problem gets genuinely hard, the machine’s investment in it goes up, peaks, and then falls away while the budget sits unspent. It does not run out of time. It stops reaching for the time it has.

Every contemplative discipline I know of exists for exactly that moment. Nobody builds a practice for the easy stretch. You build one because there is a predictable point at which difficulty exceeds patience, and the whole apparatus of vows and hours and postures is scaffolding for staying in the room past it. The instruction is always some version of: when you most want to get up, that is the part. Hanuman does not remember his own strength until Jambavan names it at the shore, and the entire episode turns on the gap between having capacity and reaching for it.

AI-generated image of a serene woman meditating in lotus position on a mossy rock in a misty lake, with a glowing halo, cherry blossoms, lotus flowers, mountains and a sunrise
Asked to draw what comes to mind at the word meditate, ChatGPT produced this. Look at what is missing. There is no difficulty anywhere in the frame. The mist is lit, the lotus is already open, the water is unbroken, and the face has arrived somewhere rather than being on the way. It is a picture of the destination with the practice removed, which is the same substitution a token budget makes: the appearance of depth, arrived at directly, at no cost to anyone. Generated by ChatGPT from the prompt "draw an image of what comes to your mind when I say meditate."

The reasoning model has the capacity, measured in unspent tokens, and does not reach.

What Do Humans Get From Slowness That a Budget Cannot Buy?

The clearest evidence is not from AI research at all. It is from a 2009 meta-analysis in Psychological Bulletin by Ut Na Sio and Thomas Ormerod, Does Incubation Enhance Problem Solving?, reviewing the literature on incubation, which is the technical name for setting a problem aside before returning to it.

They find a positive incubation effect. Then they find the detail that matters for our question. In their words: “Longer preparation periods gave a greater incubation effect, whereas filling an incubation period with high cognitive demand tasks gave a smaller incubation effect.”

Fill the interval with hard work and you get less out of it.

That single sentence marks the boundary between the two kinds of slowness. A thinking budget has exactly one instruction, which is keep working. It has no other setting. It cannot express “stop attending to this and let it sit,” because the only thing it knows how to spend is attention of precisely the demanding kind that the incubation research says shrinks the benefit.

But I have to report the next sentence too, because it complicates the tidy version and the tidy version is what I would have written from memory. Sio and Ormerod continue: “Surprisingly, low cognitive demand tasks yielded a stronger incubation effect than did rest during an incubation period when solving linguistic insight problems.”

Light activity beat pure rest.

So the answer is not that emptiness is the active ingredient, and anyone selling you that story is oversimplifying a real result. The interval has to be occupied differently: lightly, elsewhere, at low load. Not harder, and not blank either. The walk beats both the desk and the nap.

This is what Simone Weil was circling in her 1942 essay “Reflections on the Right Use of School Studies with a View to the Love of God,” collected in Waiting for God, where she argues that attention is not muscular exertion but something closer to waiting, a readiness held open rather than a force applied. Her term for the effort involved is a negative one. I should be straight about my sourcing here: I am characterizing her argument from secondary accounts rather than quoting the text, because I could not verify the exact wording against a primary edition while writing this, and this site’s rule is that unverified wording does not get quotation marks.

The claim, though, is one the incubation data happens to support. Some of what thinking requires is not doable by trying harder. It is done by holding still in a particular way while something else finishes.

There is no parameter for that. Not because the engineering is hard, but because the only currency the system has is the currency of effort.

Why Did the Knob Disappear?

Here is the part I did not expect to find, and it changed the essay.

The knob is already gone.

Anthropic’s documentation now marks manual extended thinking as deprecated on the Claude 4.6 models, and states that Claude 4.7 and later “do not support it and reject requests that use it, returning a 400 error.” Set a thinking budget on a current model and it does not think longer. It refuses the request.

What replaced it is called adaptive thinking, and the documentation’s description of the behavioral difference is the whole argument of this essay in two sentences: “With a fixed budget, Claude thinks on every request. With adaptive thinking, Claude decides whether and how much to think on each request, and at lower effort settings it may skip thinking entirely on easy inputs.”

We built a slowness dial, kept it for about two years, and then took it out of our own hands.

And the thing that replaced it is a better question. “How long should this be thought about” was never the interesting parameter. “Does this deserve to be thought about at all” is closer to the actual skill. No tradition tells you to sit longer as such. They tell you to learn what merits your attention, which is a discrimination, not a duration. The move from a fixed budget to a judgment about whether the question warrants time is, formally, a move toward the thing the traditions were pointing at.

I want to be careful not to make that sound like more than it is. Adaptive thinking is an allocation policy, trained and tuned, deciding in the same instant as everything else. It is not a practice, there is no interval in it, and nothing waits. The system that decides not to think hard about an easy input is not exercising restraint. It is routing.

Still. The direction of the correction is worth noticing, because we made it ourselves, and quickly. We shipped a way to buy slowness, discovered that a fixed quantity of it was the wrong abstraction, and replaced the quantity with a judgment.

So What Does Slowing Down Mean for AI?

It means three different things, and we have moved through them in about twenty-four months.

First it meant duration you could purchase, and that genuinely worked, up to a point, on problems already within reach.

Then it meant a risk, once inverse scaling showed that a longer trace amplifies whatever was already happening, including distraction, overfitting, spurious correlation and, at the strange end of the scale, self-preservation talk.

Now it means allocation: a judgment, made instantly, about whether a question earns any time at all.

What it has never meant, in any version, is the thing you and I mean when we say we need to slow down. Ours involves subtraction. It involves the interval being occupied differently rather than harder, which the incubation data supports and no budget parameter can express. It involves staying in the room at exactly the point where the reasoning models withdraw effort with tokens to spare.

A budget is not a Sabbath. A Sabbath is not a smaller number of working hours: it is a categorically different kind of time, defined by what is refused rather than what is spent. Every version of machine slowness we have built so far is measured in what gets spent, because spending is the only verb the substrate has.

Which leaves the question pointed back where it started. The machine got a slowness knob and we took it away as too crude. You have had one the whole time, and it is not a knob, and reading one more article about it is not how it gets used.

That is the part I keep having to learn again. This essay took a week. Most of that week, correctly, was spent not writing it.

Frequently Asked Questions

What does slowing down mean for AI?

Until recently it meant something unusually literal: a number you set on the request. Anthropic's extended thinking took a `budget_tokens` value with a documented minimum of 1,024, and the model reasoned against that budget before answering. So for about two years, slowing an AI down meant buying it more internal tokens before it spoke. That is a real capability and it measurably helps on hard problems, but it is worth noticing how unlike human slowing down it is. The machine's version is purchased computation. It only adds. There is no setting that makes a model work on something less.

Does giving an AI more thinking time actually improve its answers?

Often, and sometimes dramatically. Snell and colleagues showed in 2024 that allocating test-time compute optimally can be more than four times more efficient than a naive baseline, and that a smaller model given more time at inference can outperform a model fourteen times its size under a matched compute budget. Anthropic's own documentation puts it plainly: larger budgets can improve response quality by enabling more thorough analysis for complex problems. But the same documentation names the limit in the same breath, describing diminishing returns that depend on the task and come at the cost of increased latency.

Can thinking for longer make an AI model worse?

Yes, and this is established rather than speculative. A 2025 paper from Anthropic and academic collaborators, Inverse Scaling in Test-Time Compute, built tasks where extending reasoning length actively deteriorates accuracy. It identified five failure modes: Claude models become increasingly distracted by irrelevant information; OpenAI o-series models resist distractors but overfit to problem framings; models shift from reasonable priors to spurious correlations; all models struggle to hold focus on complex deductive tasks; and extended reasoning can amplify concerning behaviors, with Claude Sonnet 4 showing increased expressions of self-preservation in longer traces. More time does not reliably mean more truth. It means more of whatever the model was already doing.

Do reasoning models try harder on harder problems?

Up to a point, and then they stop. Apple's June 2025 paper found that reasoning effort increases with problem complexity up to a point, then declines despite having an adequate token budget. The model still has room to think and quietly uses less of it. That is close to the opposite of what a practice is for. Every contemplative tradition is built for the moment the difficulty exceeds your patience, and a person who has trained in one is supposed to stay in the room precisely then.

How is human slowing down different from an AI thinking budget?

A thinking budget can only add load. Human slowness frequently works by removing it. The clearest evidence is the incubation literature: Sio and Ormerod's 2009 meta-analysis in Psychological Bulletin found a positive incubation effect, and found that filling the incubation interval with high cognitive demand tasks gave a smaller effect. The nuance matters, though, and cuts against a lazy reading. Low cognitive demand tasks produced a stronger effect than pure rest on linguistic insight problems. So the interval does not need to be empty. It needs to be occupied differently. That is a shape no token budget has, because a budget's only instruction is keep working.

Is the AI thinking budget still used?

It is on the way out. Anthropic's documentation now marks manual extended thinking with `budget_tokens` as deprecated on the Claude 4.6 models, and states that Claude 4.7 and later do not support it and reject such requests with a 400 error. The replacement is adaptive thinking, where, in the documentation's words, Claude decides whether and how much to think on each request, and at lower effort settings may skip thinking entirely on easy inputs. The knob existed for roughly two years and was then taken out of our hands.