On the evening of September 30, a notice from Anthropic was waiting in my inbox. The subject line announced the retirement of Claude Sonnet 4.5; the body named November 30 and pointed to claude-sonnet-5-5 as the recommended replacement.
As an indie developer I run a wallpaper app, and a small script that polishes store descriptions and release notes into several languages has been pinned to the dated ID claude-sonnet-4-5-20250929 for a long time. The pricing page says the replacement drops input from $3 to $2 and output from $15 to $10 per million tokens. If it gets cheaper, there's no rush — I had almost closed the email on that thought when my hand stopped.
I had decided it was cheaper from one column of a price table. I hadn't checked how the bill itself would change.
Three dates in the notice, and the one I misread first
The official deprecations page lists two dates per model: the day it becomes deprecated and the day it retires. claude-sonnet-4-5-20250929 became deprecated on September 30, 2026 and retires on November 30, 2026. While deprecated, it still answers as before. Once retired, requests fail.
What I misread was the meaning of "still works." As long as you pin a dated ID, you never see an error. Nothing changes until the morning of November 30, when everything stops at once. An alias would roll you onto the next generation automatically, but the places where you pinned a date for reproducibility are exactly the places you forget to update.
There are two doors into the inventory. One is the Usage page in Claude Console: press Export and you get a CSV broken down by API key and model. The other is grep on your code. The first tells you what is actually being called; the second tells you where it is written. You need both, and you reconcile them against each other.
# What this solves: read the usage CSV exported from Console and list
# which API keys are still calling a model that is about to retire.
import csv
import sys
from collections import defaultdict
RETIRING = {"claude-sonnet-4-5-20250929": "2026-11-30"}
def summarize(path: str) -> None:
hits: dict[tuple[str, str], int] = defaultdict(int)
with open(path, newline="", encoding="utf-8") as f:
for row in csv.DictReader(f):
model = row.get("model", "")
if model in RETIRING:
key = row.get("api_key_name") or row.get("api_key", "unknown")
tokens = int(row.get("input_tokens", 0) or 0) + int(row.get("output_tokens", 0) or 0)
hits[(key, model)] += tokens
if not hits:
print("No retiring models in this export")
return
for (key, model), tokens in sorted(hits.items(), key=lambda x: -x[1]):
print(f"{key}\t{model}\tretires {RETIRING[model]}\t{tokens:,} tokens")
if __name__ == "__main__":
summarize(sys.argv[1] if len(sys.argv) > 1 else "usage.csv")Column names in the export have shifted over time, so I read them with row.get rather than trusting a fixed header. On the code side, a single grep -rn "claude-sonnet-4-5" --include="*.py" --include="*.ts" --include="*.env*" . is enough. In my script the ID turned up in two places: the config file, and a test fixture I had forgotten about.
A third off on paper — except the tokens are counted differently
Read naively, the price table says input goes from $3 to $2 and output from $15 to $10, a one-third cut on both. Further down the same page, though, there's a sentence that's easy to skim past: Claude 4.7 and later models use a newer tokenizer that produces roughly 30% more tokens for the same text. Sonnet 4.6 and earlier keep the old tokenizer, so moving from 4.5 to 5.5 crosses exactly that boundary.
So the unit price becomes two-thirds, and the number of tokens becomes about 1.3 times. Multiply them and the bill settles at roughly a 13% reduction. Rebudget as if you'd saved a third, and you'll be puzzling over the invoice at the end of the month.
| Item | Sonnet 4.5 | Sonnet 5.5 (list) | Sonnet 5.5 (at 1.3x tokens) |
|---|---|---|---|
| Input / MTok | $3.00 | $2.00 | ≈ $2.60 effective |
| Output / MTok | $15.00 | $10.00 | ≈ $13.00 effective |
| Cache write (5 min) | $3.75 | $2.50 | ≈ $3.25 effective |
| Cache read | $0.30 | $0.20 | ≈ $0.26 effective |
That 30% is an official rule of thumb, and it moves with the kind of text. Japanese descriptions and English ones, one-line headings and long paragraphs — each inflates differently. So before you decide on a replacement, I'd ask you to count once with your own material. The Messages API has count_tokens, which returns only the token count without generating anything; send the same prompt under both model names and you have your ratio on the spot.
# What this solves: compare input token counts for the old and new model on your own prompts,
# so the pricing page's "about 30%" becomes a number for your text.
import os
import anthropic
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY", "YOUR_API_KEY"))
OLD = "claude-sonnet-4-5-20250929"
NEW = "claude-sonnet-5-5"
def count(model: str, system: str, user: str) -> int:
res = client.messages.count_tokens(
model=model,
system=system,
messages=[{"role": "user", "content": user}],
)
return res.input_tokens
if __name__ == "__main__":
system = open("prompts/store_description_system.txt", encoding="utf-8").read()
user = open("samples/ja_release_note.txt", encoding="utf-8").read()
old_n = count(OLD, system, user)
new_n = count(NEW, system, user)
ratio = new_n / old_n if old_n else float("nan")
print(f"{OLD}: {old_n} tokens")
print(f"{NEW}: {new_n} tokens")
print(f"ratio: {ratio:.2f}x -> effective input price ${2.0 * ratio:.2f} / MTok (was $3.00)")Count the real production system prompt together with a representative user input. If you measure only a short sample, the fixed system prompt dominates and you'll misread the ratio. I prepared a few store-description drafts per language and used the pair with the largest increase as the basis for the budget.
What stopped me first wasn't the price — it was temperature
With the counting done, I switched the model name on a single call and sent it. What came back wasn't a response but a 400. The body pointed at temperature.
The lower part of the deprecations page covers request parameters as well as models. temperature, top_p and top_k are deprecated from Claude 4.7 onward, and a non-default value returns a 400. On top of that, the Python SDK removed those arguments in v1.0, so once you upgrade the SDK the failure moves earlier and becomes a TypeError before anything is sent.
My script had carried temperature=0.2 for years to keep translations from wandering. A line that sailed through on 4.5 is turned away at the door on 5.5 — and if I'd only looked at the price table, I would have met that 400 for the first time on retirement morning.
The official replacement is to drop the parameter and steer behavior with the prompt. In my case, adding one paragraph to the system prompt — follow the attached glossary, don't add paraphrases — brought the variance back to what I'd been getting with 0.2, at least as far as I could tell by reading the output.
# What this solves: automatically drops temperature for models that no longer accept it,
# so one function can serve both the old and the new model ID during the migration window.
import os
import anthropic
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY", "YOUR_API_KEY"))
# Only the old-tokenizer generation (4.6 and earlier) accepts sampling parameters
LEGACY_SAMPLING_MODELS = {"claude-sonnet-4-5-20250929", "claude-sonnet-4-6"}
def translate(model: str, system: str, text: str, temperature: float | None = 0.2) -> str:
kwargs = {
"model": model,
"max_tokens": 1024,
"system": system,
"messages": [{"role": "user", "content": text}],
}
if temperature is not None and model in LEGACY_SAMPLING_MODELS:
kwargs["temperature"] = temperature
try:
res = client.messages.create(**kwargs)
except anthropic.BadRequestError as e:
# You land here if temperature reaches a 4.7+ model. Fail loudly and say where to fix it.
raise RuntimeError(f"{model} does not accept temperature; steer it from the system prompt instead: {e}") from e
return "".join(block.text for block in res.content if getattr(block, "type", "") == "text")
if __name__ == "__main__":
system = "You translate for app stores. Follow the glossary and do not add paraphrases."
print(translate("claude-sonnet-5-5", system, "Added 30 new wallpapers."))Why write it this way? Because during the migration window, old and new IDs live in the same codebase. Branching the arguments on the model name means a one-line change in the config switches over, and the same one line switches back. When you move the SDK to v1.0 or later, delete LEGACY_SAMPLING_MODELS and the temperature argument together — I left a comment so I don't do those two steps in the wrong order.
The order I followed, and the line I drew
Here is the sequence I ended up with after the notice arrived.
- Export Usage from Console and confirm which keys still call the retiring ID.
- Grep the code and count dated IDs in config, fixtures and docs.
- Count production-like prompts with
count_tokensand estimate the bill from the ratio. - Remove
temperatureand friends; add the equivalent guidance to the system prompt. - Switch a single call, run it for a few days, and watch both output quality and spend.
- Switch the rest and delete the old ID from the code before retirement day.
I did not, however, rewrite everything at once, and there's a reason. Whether to point at the alias claude-sonnet-5-5 or pin its dated ID again depends on how much reproducibility you need. I kept production on a dated ID and moved only the verification script to the alias, so that when the next generation lands I see the differences early. If you have many places to update, the habit of checking the model catalog your local CLI carries before a bulk replace is something I wrote up in Before bulk-replacing model IDs, look up the catalog your local CLI carries.
If I compress what this whole exercise taught me into one line, it's this:
Check the price on the table; check the bill with your own text.
Good news about a price cut is welcome, but it isn't "cheaper" until the tokenizer and the parameter changes are in the calculation too. One 400 was enough to teach me that.
For today, I'd suggest just one thing: export Usage from Console and look at whether claude-sonnet-4-5 still appears in the model column. If it does, the days left until retirement are your deadline. That one row is where I started as well.