◉CLAUDE LABJP
●2.1.288 — Claude Code 2.1.288 (Oct 2) adds a re-authentication prompt for MCP and `/code-review --max-findings`. No newer release has appeared yet●10/07 — 3 days left until the old spellings of Claude Desktop / Cowork managed-config keys stop being accepted. After the cutoff they fail closed●NEW — My Usage Limit Ran Out Before Noon, and the Reason Was Conversation Length, Not Message Count●LIMIT — Reports of Max-plan usage limits running out early have passed 870 comments on the official issue. Many readers are looking for a way to isolate the cause●SWITCH — A request to switch accounts quickly in Desktop (#18435) has 200 comments. A good case for thinking about how to separate work and personal use●KEY — When an API key is set it overrides the subscription and can trigger "Organization has been disabled" (#8327). The cause can be checked in a few short steps●2.1.288 — Claude Code 2.1.288 (Oct 2) adds a re-authentication prompt for MCP and `/code-review --max-findings`. No newer release has appeared yet●10/07 — 3 days left until the old spellings of Claude Desktop / Cowork managed-config keys stop being accepted. After the cutoff they fail closed●NEW — My Usage Limit Ran Out Before Noon, and the Reason Was Conversation Length, Not Message Count●LIMIT — Reports of Max-plan usage limits running out early have passed 870 comments on the official issue. Many readers are looking for a way to isolate the cause●SWITCH — A request to switch accounts quickly in Desktop (#18435) has 200 comments. A good case for thinking about how to separate work and personal use●KEY — When an API key is set it overrides the subscription and can trigger "Organization has been disabled" (#8327). The cause can be checked in a few short steps
Articles/Claude.ai
◉ Claude.ai/2026-10-04Intermediate

claude-sonnet-4-5 stops on November 30 — I checked the bill before calling its replacement cheaper

claude-sonnet-4-5-20250929 retires on November 30, 2026. The recommended claude-sonnet-5-5 has lower list prices, but a new tokenizer and a 400 on temperature change the math. After reading, you can audit your own calls and plan the switch in order.

Claude API124Model deprecationSonnet 5.5CostMigration5

On the evening of September 30, a notice from Anthropic was waiting in my inbox. The subject line announced the retirement of Claude Sonnet 4.5; the body named November 30 and pointed to claude-sonnet-5-5 as the recommended replacement.

As an indie developer I run a wallpaper app, and a small script that polishes store descriptions and release notes into several languages has been pinned to the dated ID claude-sonnet-4-5-20250929 for a long time. The pricing page says the replacement drops input from $3 to $2 and output from $15 to $10 per million tokens. If it gets cheaper, there's no rush — I had almost closed the email on that thought when my hand stopped.

I had decided it was cheaper from one column of a price table. I hadn't checked how the bill itself would change.

Three dates in the notice, and the one I misread first

The official deprecations page lists two dates per model: the day it becomes deprecated and the day it retires. claude-sonnet-4-5-20250929 became deprecated on September 30, 2026 and retires on November 30, 2026. While deprecated, it still answers as before. Once retired, requests fail.

What I misread was the meaning of "still works." As long as you pin a dated ID, you never see an error. Nothing changes until the morning of November 30, when everything stops at once. An alias would roll you onto the next generation automatically, but the places where you pinned a date for reproducibility are exactly the places you forget to update.

There are two doors into the inventory. One is the Usage page in Claude Console: press Export and you get a CSV broken down by API key and model. The other is grep on your code. The first tells you what is actually being called; the second tells you where it is written. You need both, and you reconcile them against each other.

# What this solves: read the usage CSV exported from Console and list
# which API keys are still calling a model that is about to retire.
import csv
import sys
from collections import defaultdict
 
RETIRING = {"claude-sonnet-4-5-20250929": "2026-11-30"}
 
 
def summarize(path: str) -> None:
    hits: dict[tuple[str, str], int] = defaultdict(int)
    with open(path, newline="", encoding="utf-8") as f:
        for row in csv.DictReader(f):
            model = row.get("model", "")
            if model in RETIRING:
                key = row.get("api_key_name") or row.get("api_key", "unknown")
                tokens = int(row.get("input_tokens", 0) or 0) + int(row.get("output_tokens", 0) or 0)
                hits[(key, model)] += tokens
    if not hits:
        print("No retiring models in this export")
        return
    for (key, model), tokens in sorted(hits.items(), key=lambda x: -x[1]):
        print(f"{key}\t{model}\tretires {RETIRING[model]}\t{tokens:,} tokens")
 
 
if __name__ == "__main__":
    summarize(sys.argv[1] if len(sys.argv) > 1 else "usage.csv")

Column names in the export have shifted over time, so I read them with row.get rather than trusting a fixed header. On the code side, a single grep -rn "claude-sonnet-4-5" --include="*.py" --include="*.ts" --include="*.env*" . is enough. In my script the ID turned up in two places: the config file, and a test fixture I had forgotten about.

A third off on paper — except the tokens are counted differently

Read naively, the price table says input goes from $3 to $2 and output from $15 to $10, a one-third cut on both. Further down the same page, though, there's a sentence that's easy to skim past: Claude 4.7 and later models use a newer tokenizer that produces roughly 30% more tokens for the same text. Sonnet 4.6 and earlier keep the old tokenizer, so moving from 4.5 to 5.5 crosses exactly that boundary.

So the unit price becomes two-thirds, and the number of tokens becomes about 1.3 times. Multiply them and the bill settles at roughly a 13% reduction. Rebudget as if you'd saved a third, and you'll be puzzling over the invoice at the end of the month.

ItemSonnet 4.5Sonnet 5.5 (list)Sonnet 5.5 (at 1.3x tokens)
Input / MTok$3.00$2.00≈ $2.60 effective
Output / MTok$15.00$10.00≈ $13.00 effective
Cache write (5 min)$3.75$2.50≈ $3.25 effective
Cache read$0.30$0.20≈ $0.26 effective

That 30% is an official rule of thumb, and it moves with the kind of text. Japanese descriptions and English ones, one-line headings and long paragraphs — each inflates differently. So before you decide on a replacement, I'd ask you to count once with your own material. The Messages API has count_tokens, which returns only the token count without generating anything; send the same prompt under both model names and you have your ratio on the spot.

# What this solves: compare input token counts for the old and new model on your own prompts,
# so the pricing page's "about 30%" becomes a number for your text.
import os
import anthropic
 
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY", "YOUR_API_KEY"))
 
OLD = "claude-sonnet-4-5-20250929"
NEW = "claude-sonnet-5-5"
 
 
def count(model: str, system: str, user: str) -> int:
    res = client.messages.count_tokens(
        model=model,
        system=system,
        messages=[{"role": "user", "content": user}],
    )
    return res.input_tokens
 
 
if __name__ == "__main__":
    system = open("prompts/store_description_system.txt", encoding="utf-8").read()
    user = open("samples/ja_release_note.txt", encoding="utf-8").read()
    old_n = count(OLD, system, user)
    new_n = count(NEW, system, user)
    ratio = new_n / old_n if old_n else float("nan")
    print(f"{OLD}: {old_n} tokens")
    print(f"{NEW}: {new_n} tokens")
    print(f"ratio: {ratio:.2f}x -> effective input price ${2.0 * ratio:.2f} / MTok (was $3.00)")

Count the real production system prompt together with a representative user input. If you measure only a short sample, the fixed system prompt dominates and you'll misread the ratio. I prepared a few store-description drafts per language and used the pair with the largest increase as the basis for the budget.

What stopped me first wasn't the price — it was temperature

With the counting done, I switched the model name on a single call and sent it. What came back wasn't a response but a 400. The body pointed at temperature.

The lower part of the deprecations page covers request parameters as well as models. temperature, top_p and top_k are deprecated from Claude 4.7 onward, and a non-default value returns a 400. On top of that, the Python SDK removed those arguments in v1.0, so once you upgrade the SDK the failure moves earlier and becomes a TypeError before anything is sent.

My script had carried temperature=0.2 for years to keep translations from wandering. A line that sailed through on 4.5 is turned away at the door on 5.5 — and if I'd only looked at the price table, I would have met that 400 for the first time on retirement morning.

The official replacement is to drop the parameter and steer behavior with the prompt. In my case, adding one paragraph to the system prompt — follow the attached glossary, don't add paraphrases — brought the variance back to what I'd been getting with 0.2, at least as far as I could tell by reading the output.

# What this solves: automatically drops temperature for models that no longer accept it,
# so one function can serve both the old and the new model ID during the migration window.
import os
import anthropic
 
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY", "YOUR_API_KEY"))
 
# Only the old-tokenizer generation (4.6 and earlier) accepts sampling parameters
LEGACY_SAMPLING_MODELS = {"claude-sonnet-4-5-20250929", "claude-sonnet-4-6"}
 
 
def translate(model: str, system: str, text: str, temperature: float | None = 0.2) -> str:
    kwargs = {
        "model": model,
        "max_tokens": 1024,
        "system": system,
        "messages": [{"role": "user", "content": text}],
    }
    if temperature is not None and model in LEGACY_SAMPLING_MODELS:
        kwargs["temperature"] = temperature
    try:
        res = client.messages.create(**kwargs)
    except anthropic.BadRequestError as e:
        # You land here if temperature reaches a 4.7+ model. Fail loudly and say where to fix it.
        raise RuntimeError(f"{model} does not accept temperature; steer it from the system prompt instead: {e}") from e
    return "".join(block.text for block in res.content if getattr(block, "type", "") == "text")
 
 
if __name__ == "__main__":
    system = "You translate for app stores. Follow the glossary and do not add paraphrases."
    print(translate("claude-sonnet-5-5", system, "Added 30 new wallpapers."))

Why write it this way? Because during the migration window, old and new IDs live in the same codebase. Branching the arguments on the model name means a one-line change in the config switches over, and the same one line switches back. When you move the SDK to v1.0 or later, delete LEGACY_SAMPLING_MODELS and the temperature argument together — I left a comment so I don't do those two steps in the wrong order.

The order I followed, and the line I drew

Here is the sequence I ended up with after the notice arrived.

  1. Export Usage from Console and confirm which keys still call the retiring ID.
  2. Grep the code and count dated IDs in config, fixtures and docs.
  3. Count production-like prompts with count_tokens and estimate the bill from the ratio.
  4. Remove temperature and friends; add the equivalent guidance to the system prompt.
  5. Switch a single call, run it for a few days, and watch both output quality and spend.
  6. Switch the rest and delete the old ID from the code before retirement day.

I did not, however, rewrite everything at once, and there's a reason. Whether to point at the alias claude-sonnet-5-5 or pin its dated ID again depends on how much reproducibility you need. I kept production on a dated ID and moved only the verification script to the alias, so that when the next generation lands I see the differences early. If you have many places to update, the habit of checking the model catalog your local CLI carries before a bulk replace is something I wrote up in Before bulk-replacing model IDs, look up the catalog your local CLI carries.

If I compress what this whole exercise taught me into one line, it's this:

Check the price on the table; check the bill with your own text.

Good news about a price cut is welcome, but it isn't "cheaper" until the tokenizer and the parameter changes are in the calculation too. One 400 was enough to teach me that.

For today, I'd suggest just one thing: export Usage from Console and look at whether claude-sonnet-4-5 still appears in the model column. If it does, the days left until retirement are your deadline. That one row is where I started as well.

Share

Thank You for Reading

Claude Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • ✦Copy-paste ready implementation code
  • ✦New advanced guides published daily
  • ✦$5/mo or $15 for lifetime access
View Membership →

If you found this article helpful, a small tip ($1.50) would mean a lot to us. Your support helps keep this site ad-free and covers server and hosting costs.

Related Articles

⬡ API & SDK2026-08-14
Move the Prompt Tools API Into Your Own Scripts Before Workbench Closes on August 17
The legacy Workbench and three experimental prompt endpoints shut down on August 17, 2026. Here is how to count what actually depends on them, plus working local replacements for templatize_prompt and improve_prompt with real output.
◉ Claude.ai2026-10-03
My Usage Limit Ran Out Before Noon, and the Reason Was Conversation Length, Not Message Count
When the usage-limit banner shows up on Pro or Max sooner than you expected, here is the order I check: the two bars in Settings > Usage, the length of the conversation and its attachments, then the model and effort level. The morning I got cut off before lunch, and the lines I drew afterwards.
◉ Claude.ai2026-09-19
When a proofreading request comes back as a rewrite: keeping the original wording
Ask for proofreading and you often get a polished replacement instead. Fixing the output shape first, splitting edits into three layers, and showing one keep-and-change example pair will hold on to the author's sentence endings and word order.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links