How Much Does It Cost to Ask Claude an Interesting Question?
spoiler: about a nickel
Imagine you had an interesting question1: something juicy enough that you’d want an AI to really chew on it. You open Claude, but before you even type anything, you’re already making decisions: Opus or Sonnet? Extended thinking on or off? Each choice implies a tradeoff between intelligence and cost. Intelligence is hard to quantify, but cost is straightforward.
So how much are you actually saving by switching to Sonnet? And what does that “extended thinking” toggle cost you in practice?
I ran a small experiment to find out.
Method
Claude users get a rolling usage allowance measured in five-hour sessions. The settings page at claude.ai/settings/usage displays a “current session” tracking bar as a percentage. This works as a basic instrument for measuring relative inference cost across configurations.
The procedure was straightforward. Before each query, I recorded the current session percentage:
Then I launched the query, waited for the full response, and recorded the new percentage:
The delta between the two readings is the cost of that single exchange, expressed as a fraction of one session’s total budget.
All readings were noted contemporaneously in a Notion page. To check repeatability, I ran queries in both a standard Chat window and a Cowork instance.
The prompt held constant across all six trials was the following:
I’m 33. What’s the percentage likelihood that I live to see the 22nd century? Quantify your uncertainty and consider exotic scenarios such as aging research and AI-driven post-scarcity.
This gives the model room to flex: probability estimation, calibration language, speculative reasoning, hedging, tool use.
Results
The range spans from 2% (Sonnet without Extended Thinking) up to 7% (Opus with Extended Thinking). Assuming one full session’s budget corresponds to roughly a dollar of inference cost2, this means a single interesting question costs anywhere from about two cents on the frugal end to seven cents on the expensive end. Split the difference and call it five cents on average.
Extended thinking is the primary cost driver, and its effect is more pronounced with Opus than Sonnet. Opus with extended thinking consumed 3x to 3.5x more budget than the same model without it. Sonnet’s extended thinking overhead was more modest, roughly 1x to 1.5x above its baseline. Without extended thinking, Opus and Sonnet were indistinguishable at this level of measurement precision.
Conclusion
Compare the following two answers:
The seven cent result (Opus Extended Thinking)
They’re…both fine?
Here is the paradox of these numbers: the seven-cent response is decent, and the two-cent response is also decent. Both gave reasonable, well-structured answers to the longevity question. The expensive answer included more nuance and more thorough scenario analysis, but it wasn’t 3.5 times better than the cheap answer.
And yet, the incentive structure of Claude Pro’s pricing pushes in the opposite direction. Subscribers pay a flat monthly fee for a rolling usage budget. Any usage below the session cap is functionally free at the margin. There is zero incremental cost to choosing Opus with Extended Thinking over plain Sonnet, as long as you stay under the ceiling of roughly twenty queries per five hour period. The rational behavior for a cost-conscious subscriber under that limit is therefore to always use the most expensive configuration available. You are leaving quality on the table if you use Sonnet, and the money you “save” by doing so evaporates unused at the end of your session window.
This is a strange equilibrium. For many the system incentivizes maximum consumption per query, which is the exact opposite of what you’d want if you were optimizing inference infrastructure costs at scale.
Limitations and Next Steps
This experiment has obvious constraints. The session percentage bar is a blunt instrument; direct token-cost measurement via the API would yield far more precise data. The sample size is six queries with one prompt across two models, which is enough to sketch the landscape but nowhere near enough for statistical confidence. The Cowork and Chat deltas for the same configuration didn’t always agree, suggesting either meaningful variance in response length or some overhead difference between interfaces that deserves investigation.
More interesting open questions sit downstream. This was all done on Claude Pro, which is the default tier for a large share of users. For power users willing to go further: how cheap of a model can you actually tolerate before the quality gap becomes noticeable? Is the degradation curve smooth or is there a cliff? And for anyone considering self-hosting open-weight models, what does the cost picture look like when the only expense is energy and hardware amortization? Like $5 Ubers, the two-cent query might turn out to be a VC-fueled luxury that we look back on wistfully in the 22nd century.3
Not strictly required.
I asked the models to estimate the cost of their token usage; their estimates were consistent at roughly a few cents each, so this seems directionally correct.
I give it a 15% chance.





