◉CLAUDE LABJP
●2.1.292 — Claude Code 2.1.292 (Oct 6) fixes cloud sessions dropping answers to permission prompts and the last messages of a session being lost on quit●MODELS API — The Models API now reports in capabilities whether a model accepts thinking turned off (Oct)●11/30 — 54 days until Sonnet 4.5 retires on the Claude API. Move to Sonnet 5.5, and check your thinking settings and tool_choice first●MODS — A Zenn post walks through installing Claude Code Mods safely. The real question is what to check before you install, and how to back out●NEW — When Pro looks unpaid, the screens to check depend on whether you bought on the web or in the store. I sorted that out step by step●THINKING — Sonnet 5.5 thinking blocks only work in the account that produced them. Sent from another account, they are silently dropped●2.1.292 — Claude Code 2.1.292 (Oct 6) fixes cloud sessions dropping answers to permission prompts and the last messages of a session being lost on quit●MODELS API — The Models API now reports in capabilities whether a model accepts thinking turned off (Oct)●11/30 — 54 days until Sonnet 4.5 retires on the Claude API. Move to Sonnet 5.5, and check your thinking settings and tool_choice first●MODS — A Zenn post walks through installing Claude Code Mods safely. The real question is what to check before you install, and how to back out●NEW — When Pro looks unpaid, the screens to check depend on whether you bought on the web or in the store. I sorted that out step by step●THINKING — Sonnet 5.5 thinking blocks only work in the account that produced them. Sent from another account, they are silently dropped
Articles/API & SDK
⬡ API & SDK/2026-07-09Advanced

The 200K-Token Cliff That Doubled My Nightly Bill — A count_tokens Preflight to Avoid Long-Context Pricing

Sonnet 5's native 1M context switches the entire request to long-context pricing the moment input crosses 200K tokens. Here's how I caught that silent cost cliff before sending, using the billing-exempt count_tokens API, with working Python and TypeScript code.

claude-api83token-counting3cost5sonnet-51m-context

✦ Premium Article

One morning I froze while reviewing the cost of the previous night's article-generation batch. The workload was nearly identical to the day before, yet the bill had roughly doubled.

It wasn't a failed job or the wrong model. That particular night, accumulated conversation history plus the files I'd loaded pushed the input past 200K tokens, and the entire request flipped to long-context pricing. The rate changed discontinuously across a single token. A genuine cliff.

As an indie developer running Dolice's four sites through nightly automation, I couldn't leave that cliff alone. Failures leave logs. This one succeeds quietly and only shows up on the invoice. The only clue was the bill itself, and that scared me.

This article shares the mechanism I now run to detect that cliff before sending, with the actual code.

Why the bill jumps at 200K tokens

Claude Sonnet 5 handles a native 1M-token context. But the price isn't flat. A request whose input exceeds 200K tokens is billed under a separate long-context tier.

The crucial part: this is not a marginal surcharge on the tokens above 200K. As I understand it, the instant you cross the boundary the entire request — all input and all output, not just token 200,001 — is priced at the long-context rate. That's exactly why the total jumps discontinuously rather than sloping up smoothly.

The rate relationship looks like this. The base tier is published, but long-context multipliers can change, so confirm the current figure on the pricing page and feed it into the config value below rather than hardcoding it.

Input tokensInput price (per 1M)Output price (per 1M)Notes
≤ 200KBase tier (Sonnet 5 intro $2)Base tier ($10)Intro pricing through 2026-08-31
> 200KLong-context tier (roughly 2× base)Long-context tier (verify)Crossing applies to the whole request

I design around an assumption of roughly 2× on the input side, but I always verify the exact multiplier before going live and keep it in config rather than in code. Prices change; the structure — the whole request re-pricing at the boundary — should hold for the foreseeable future. That structure is what's worth defending.

Estimate before sending — count_tokens is billing-exempt

The key to avoiding the cliff is to learn the input token count before sending, rather than discovering it on the invoice. This is where the Token Counting API (count_tokens) earns its place.

It has two properties that suit unattended operation. First, count_tokens itself is not billed. Second, it runs on a separate track from ordinary message sends and does not consume your rate limit. So you can safely call it before every request, without spending real budget or rate headroom on an estimate.

For general token accounting I've written How to estimate token counts up front and optimize cost; here I narrow it to a single decision — will this request cross the boundary or not.

✦

Thank you for reading this far.

Continue Reading

What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.

WHAT YOU'LL LEARN
✦If your unattended jobs sometimes cost twice as much for no visible reason, you'll be able to trace it to one specific line: the 200K-token boundary
✦You'll get a preflight function in both Python and TypeScript that estimates input tokens with count_tokens and stops the request before it falls off the cliff
✦You'll come away with a clear rule for handling an over-budget request: compact by summarizing, split into smaller calls, or defer the run
Secure payment via Stripe · Cancel anytime
✦

Unlock This Article

Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.

or
Unlock all articles with Membership →
Share

Thank You for Reading

Claude Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • ✦Copy-paste ready implementation code
  • ✦New advanced guides published daily
  • ✦$5/mo or $15 for lifetime access
View Membership →

Related Articles

⬡ API & SDK2026-09-22
Character counts won't predict PDF input tokens — bracket the cost per page before you send
How to estimate the input tokens of a PDF before sending it to the Claude API, by bracketing from page count rather than character count. Includes a measured 6.6x gap on a nine-page A4 document, the two separate increases that come with current models, and the count_tokens limitation that bites when you use file_id.
⬡ API & SDK2026-09-08
Writing 25 as your session budget caps it at twenty-five cents
Managed Agents session budgets are expressed in minor units as a string. Here is how I mixed up the unit, hit budget_reached in minutes, and what raising versus clearing a budget actually does.
⬡ API & SDK2026-09-07
A 75% cheaper cache read moved one bill by 29% and another by 2.8%
Fable 5.1 and Mythos 5.1 changed the cache read multiplier from 0.1 to 0.025. Here is how to estimate your own saving from your token mix, and how to fix code that hard-codes the multiplier.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links