◉CLAUDE LABJP
●2.1.288 — Claude Code 2.1.288 (Oct 2) adds a re-authentication prompt for MCP and `/code-review --max-findings`. No newer release has appeared yet●10/07 — 3 days left until the old spellings of Claude Desktop / Cowork managed-config keys stop being accepted. After the cutoff they fail closed●NEW — My Usage Limit Ran Out Before Noon, and the Reason Was Conversation Length, Not Message Count●LIMIT — Reports of Max-plan usage limits running out early have passed 870 comments on the official issue. Many readers are looking for a way to isolate the cause●SWITCH — A request to switch accounts quickly in Desktop (#18435) has 200 comments. A good case for thinking about how to separate work and personal use●KEY — When an API key is set it overrides the subscription and can trigger "Organization has been disabled" (#8327). The cause can be checked in a few short steps●2.1.288 — Claude Code 2.1.288 (Oct 2) adds a re-authentication prompt for MCP and `/code-review --max-findings`. No newer release has appeared yet●10/07 — 3 days left until the old spellings of Claude Desktop / Cowork managed-config keys stop being accepted. After the cutoff they fail closed●NEW — My Usage Limit Ran Out Before Noon, and the Reason Was Conversation Length, Not Message Count●LIMIT — Reports of Max-plan usage limits running out early have passed 870 comments on the official issue. Many readers are looking for a way to isolate the cause●SWITCH — A request to switch accounts quickly in Desktop (#18435) has 200 comments. A good case for thinking about how to separate work and personal use●KEY — When an API key is set it overrides the subscription and can trigger "Organization has been disabled" (#8327). The cause can be checked in a few short steps
Articles/Claude Code
⟐ Claude Code/2026-07-05Advanced

When to Use Claude Code's Native 1M Context — and When Not To: A Cost-Based Rule

With Sonnet 5 as the default, Claude Code now handles a native 1M-token context. A big window is convenient, but every token you park in it is billed again each turn. Should you load the whole repo, or feed slices? Here is an estimable token model and a decision rule that gives a concrete answer per situation, with working code and the traps to avoid.

Claude Code258Sonnet 58context10cost optimization131M

✦ Premium Article

I handed a large repository to Claude Code, felt reassured that it would "read all of it," and half an hour later opened the usage screen. My hand stopped.

That single session had consumed several times my usual tokens. The work itself finished correctly. But that job did not need that window size. Feed it only the slices it needed, and the same result would have cost far less.

On June 30, 2026, Claude Sonnet 5 became the default across all plans, and Claude Code gained a native one-million-token context. The old 1M was a beta limited to specific models. Now it is within reach by default. Which is exactly why the instinct to "load everything because the window is wide" quietly melts money.

This article turns "when to use a big window and when not to" into something you decide by arithmetic rather than by feel, with code you can run against your own price sheet and repo.

A big window changes cost, not speed

Let me clear up one misconception first. Widening the context does not make the model smarter, nor necessarily faster. What changes is what you can show it at once and what you pay every turn to do so.

As a conversation progresses, Claude Code repeatedly resends the prior exchange as input. Whatever sits in the window is billed as input tokens on every response. That is the crux. Keep a 10,000-token file resident across 20 round trips, and (outside of any cache) those 10,000 tokens can be billed roughly 20 times.

So the cost of a big window scales as "amount loaded × times you touch it." For a one-shot read it is noise; for long exploration or iterative refactoring, that multiplication is what bites.

The estimate: "everything resident" vs "sliced"

Let us put the decision in symbols.

SymbolMeaning
T_ctxTokens kept resident in the window (e.g. the whole repo)
T_qTokens per instruction / question
T_outTokens per response
NRound trips in the session
p_in / p_outInput / output unit price (per million tokens)
s_in / s_outLong-context surcharge multipliers (e.g. input ×2)
cPrompt cache hit rate (0 to 1)

When you keep everything resident, the input cost is dominated by resending T_ctx every turn. The cached fraction c is billed at the cheaper cache-read rate, so the effective input cost is roughly:

input_cost(resident) ≈ N × T_ctx × p_in × s_in × (1 - c + c × r_cache)

Here r_cache is the ratio of cache-read price to normal input price (around 0.1 in many setups, i.e. about one tenth). The formula makes it visible: the higher c is, the cheaper resident becomes.

When you slice and feed only what each turn needs, you do not resend T_ctx; you send only the file fragment T_slice(i):

input_cost(sliced) ≈ Σ_i ( T_slice(i) × p_in × s_in' )

s_in' can be the no-surcharge multiplier (1.0) if the total you load stays under the long-context tier. That is where slicing earns its keep. Pricing commonly changes the surcharge based on whether you cross the 200K-token tier, so slicing under the tier lowers the multiplier itself.

Written out it looks obvious, but the practically important point is a single one: resident cost spikes precisely when N is large, c is low, and T_ctx straddles the tier.

✦

Thank you for reading this far.

Continue Reading

What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.

WHAT YOU'LL LEARN
✦How to estimate, with an explicit formula and Python, whether widening the window or slicing is cheaper — with the long-context surcharge left as a variable so you can plug in your own price sheet
✦A decision function that mechanically decides whether to use the 1M window from repo size, revisit count, and cache hit rate — with the reasoning behind each threshold
✦The typical ways a large window fails to help while quietly inflating cost, and how to avoid each
Secure payment via Stripe · Cancel anytime
✦

Unlock This Article

Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.

or
Unlock all articles with Membership →
Share

Thank You for Reading

Claude Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • ✦Copy-paste ready implementation code
  • ✦New advanced guides published daily
  • ✦$5/mo or $15 for lifetime access
View Membership →

Related Articles

⟐ Claude Code2026-08-14
One generated file outweighed all 70 hand-written source files, so I redrew what Claude Code may read
I profiled a repository I actually run to see which areas consume the most context. Here is what I found, the deny rules I settled on, how search selectivity changes the math, and the options I considered but rejected.
⟐ Claude Code2026-05-04
Build a Pipeline Where Docs Update Automatically Every Time Your Code Changes
Build a CI/CD pipeline that auto-generates README, CHANGELOG, and API docs whenever code changes. Use Claude Haiku 4.5 for cost-efficient classification and Sonnet 4.6 for quality output — cutting API costs by up to 70% while keeping documentation accurate.
⟐ Claude Code2026-09-29
Two Default-Model Changes Haven't Reached Claude Code's Stable Channel Yet
This morning, Claude Code's stable channel sat 6 releases and 10 days behind latest. A small standard-library Python script that counts the gap in real, published releases instead of version arithmetic, and lists what hasn't arrived yet with a headline for each.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links