When the retirement notice arrived, I assumed the job was a one-line change: swap the model ID for claude-sonnet-5-5 and move on.
I read the official migration guide only afterwards. Thinking settings, sampling arguments, prefill, tool_choice, beta headers — a long list of things outside the model ID that turn the same request into a 400. As an indie developer who calls the API from small scripts, I found that the real work was the shape of the request, not the string.
The short version: count what will break before you replace anything. Below is a scanner for that, and the output from running it on a small sample.
The dates, and the replacement
On the official deprecations page (checked October 8, 2026), claude-sonnet-4-5-20250929 was deprecated on September 30, 2026, with a tentative retirement date of November 30, 2026. The recommended replacement is claude-sonnet-5-5. After retirement, requests fail.
The page says these dates apply to the Claude API, Claude Platform on AWS, and Microsoft Foundry. Amazon Bedrock and Google Cloud set their own schedules, so check your own route separately.
Pricing and tokenizer differences are covered in another article. This one is only about whether the request goes through.
What the migration guide says will break
These are the items from the Sonnet 5.5 migration guide that matter when moving from 4.5, ordered by how easy they are to spot in code.
| Item | How 4.5 code writes it | On 5.5 |
|---|---|---|
| Sampling | temperature / top_p / top_k | Non-default values return 400. Remove them |
| Thinking budget | {"type": "enabled", "budget_tokens": N} | Error. Use adaptive with output_config.effort |
| Turning thinking off | {"type": "disabled"} | 400. Use between_tools |
| Prefill | Trailing assistant message | Error. End the conversation with a user message |
| tool_choice | "tool" / "any" | Error. Use auto with strict: true |
| Structured output | output_format | Deprecated. Use output_config.format |
| Beta headers | interleaved-thinking-2025-05-14 and others | Remove, or replace with eager_input_streaming |
| Reading the reply | content[0].text | A thinking block can come first. Find the block by type |
The last row is the dangerous one. The request succeeds, but content[0] is a thinking block and .text is gone — a failure with no error message, which tends to surface late in production.
Also, a request with no thinking field now runs with adaptive thinking by default, which 4.5 did not. Since max_tokens covers thinking too, it is worth revisiting.
The table, turned into regular expressions
Each row became one pattern. This is not a parser; it is a tool that produces candidates for a human to confirm.
#!/usr/bin/env python3
"""Simple scanner for Sonnet 4.5 -> 5.5 patterns that can return a 400.
Usage: python3 sonnet55_scan.py <dir> exits 1 on any hit"""
import re, sys, pathlib
RULES = [
("model-id", r"claude-sonnet-4-5(?:-\d{8})?", "change the model ID to claude-sonnet-5-5"),
("sampling", r"\b(temperature|top_p|top_k)\s*=", "non-default value returns 400; remove"),
("budget", r"budget_tokens", "switch to adaptive + output_config.effort"),
("think-off", r"""["']type["']\s*:\s*["']disabled["']""", "disabled returns 400; use between_tools"),
("prefill", r"""["']role["']\s*:\s*["']assistant["']""", "prefill if it ends the conversation; check"),
("tool-choice", r"""["']type["']\s*:\s*["'](?:tool|any)["']""", "tool/any returns 400; use auto + strict"),
("output-fmt", r"\boutput_format\b", "move to output_config.format"),
("beta-hdr", r"interleaved-thinking-2025-05-14|fine-grained-tool-streaming-2025-05-14", "remove or replace the header"),
("content0", r"content\[0\]\.text", "a thinking block can come first; find by type"),
]
EXTS = {".py", ".ts", ".js", ".mjs", ".json", ".yaml", ".yml", ".env", ".md"}
def scan(root):
hits = []
for p in sorted(pathlib.Path(root).rglob("*")):
if p.suffix not in EXTS or not p.is_file() or "node_modules" in p.parts:
continue
for n, line in enumerate(p.read_text(errors="ignore").splitlines(), 1):
for tag, pat, hint in RULES:
if re.search(pat, line):
hits.append((str(p), n, tag, hint, line.strip()))
return hits
if __name__ == "__main__":
hits = scan(sys.argv[1] if len(sys.argv) > 1 else ".")
for f, n, tag, hint, src in hits:
print(f"{f}:{n} [{tag}] {hint}\n {src}")
files = len({h[0] for h in hits})
print(f"\n{len(hits)} hits / {files} files")
sys.exit(1 if hits else 0)The prefill rule matches every "role": "assistant" line, including legitimate conversation history. Treat it as "check whether the conversation ends on this line."
What it printed
I wrote two files that deliberately use the 4.5 style and ran the scanner on them. The first has temperature, budget_tokens, a prefill, and content[0].text; the second has a beta header, tool_choice, and output_format.
app/extract.py:4 [model-id] change the model ID to claude-sonnet-5-5
model="claude-sonnet-4-5",
app/extract.py:6 [beta-hdr] remove or replace the header
betas=["interleaved-thinking-2025-05-14"],
app/extract.py:8 [tool-choice] tool/any returns 400; use auto + strict
tool_choice={"type": "tool", "name": "pick"},
app/extract.py:9 [output-fmt] move to output_config.format
output_format={"type": "json_schema", "schema": {}},
app/summarize.py:6 [model-id] change the model ID to claude-sonnet-5-5
model="claude-sonnet-4-5-20250929",
app/summarize.py:8 [sampling] non-default value returns 400; remove
temperature=0.3,
app/summarize.py:9 [budget] switch to adaptive + output_config.effort
thinking={"type": "enabled", "budget_tokens": 2000},
app/summarize.py:12 [prefill] prefill if it ends the conversation; check
{"role": "assistant", "content": "{"},
app/summarize.py:14 [content0] a thinking block can come first; find by type
).content[0].text
9 hits / 2 filesThe exit code was 1. Only two of the nine hits are model-ID lines; the other seven would survive a find-and-replace. A plain replace would have left seven places returning a 400 or misbehaving silently.
This confirms the scanner works. I did not send the fixed code to the live API for this article, so please run your own version once. The fixes follow the migration guide.
The order I would fix things in
- Remove sampling arguments,
budget_tokens, anddisabledfirst. They fail loudly with a 400. - Remove prefill. If it was forcing JSON to start with
{, the guide points to structured outputs (output_config.format). - Set
tool_choicetoautoand addstrict: trueto tools. Strict tools needadditionalProperties: falseon every object. - Last, change
content[0].textto a lookup by block type. It raises no error, so it is the one most likely to slip through.
A fixed summarizer might look like this:
resp = client.messages.create(
model="claude-sonnet-5-5", # no date suffix
max_tokens=8000, # revisit: it now includes thinking
thinking={"type": "adaptive"},
output_config={"effort": "medium"}, # set it explicitly
messages=[{"role": "user", "content": text}],
)
# safe even if a thinking block comes first
answer = next(b.text for b in resp.content if b.type == "text")effort did not exist on 4.5. Sonnet 5.5 has five levels and the API default is high. Setting it explicitly gives you something to compare against if behavior shifts later.
One thing to do today
Run the scanner once on your repository. Zero hits means you have plenty of room before November 30. Any other number is your estimate of the work.
When a deadline is fixed, count what remains before you replace anything. I intend to keep to that at the next retirement as well.
Sources: Model deprecations, Claude Sonnet 5.5 migration guide