One night I changed a single model ID string and ran my script.
A 400 came back. The message made it clear the model wasn't at fault. One argument I had carried over from the old setup no longer matched what the new model accepts. The fix took a few minutes. What made my stomach sink was that this was the first request production would send.
Since then I replay my own call shapes against a candidate model before switching. It isn't an evaluation platform. It's a handful of tiny requests.
What breaks on a switch is the shape of the call, not the model's quality
When we swap models, we tend to worry about output quality and cost. Rightly so.
But before quality comes a door: is the request accepted at all? If you trip there, comparing quality is beside the point. The causes usually fall into three groups.
| Cause | How it looks in your code | When you notice |
|---|---|---|
| A parameter no longer fits | A thinking setting or an argument you've always passed | 400 on the first request |
| A combination no longer fits | Forcing a tool with tool_choice while using another feature | 400 only on one code path |
| A capability differs | A feature the old model supported isn't there | When a screen or batch job reaches it |
The second and third are the nasty ones. They live on paths you rarely hit, so a single smoke test won't find them.
The migration guide is a map; your calls are the current location
Migration guides list what changed per model, and I'm grateful for them. But they can't tell you which of those changes your code actually steps on.
So I split the work:
- Write down the shapes of the calls my code really makes.
- Send those shapes, as they are, to the candidate and see what's accepted.
- Read the guide only for the shapes that were rejected.
That's far shorter than reading every item and cross-checking, because the checking is limited to my own calls.
Put your call shapes in one file
Start with an inventory: one entry per distinct combination of arguments your code sends.
# call_shapes.py
# The shapes of calls the production code actually makes, collected in one place.
# No model ID here (the candidate is injected at replay time).
# max_tokens is tiny: we only want an accept/reject verdict.
DUMMY_TOOL = {
"name": "ping",
"description": "Dummy tool for connectivity checks",
"input_schema": {"type": "object", "properties": {}},
}
SHAPES = {
"plain": {
"max_tokens": 16,
"messages": [{"role": "user", "content": "Reply with ok"}],
},
# Keep this only if your own code sets thinking explicitly
"thinking_disabled": {
"max_tokens": 16,
"thinking": {"type": "disabled"},
"messages": [{"role": "user", "content": "Reply with ok"}],
},
"tool_choice_any": {
"max_tokens": 16,
"tools": [DUMMY_TOOL],
"tool_choice": {"type": "any"},
"messages": [{"role": "user", "content": "Call ping"}],
},
"system_prompt": {
"max_tokens": 16,
"system": "You are an assistant for connectivity checks.",
"messages": [{"role": "user", "content": "Reply with ok"}],
},
}The rule: don't list shapes your code never sends. Chasing every possible argument bloats the file and dulls the exercise. Do include the call that only a monthly batch job makes. Those rarely-hit paths are exactly what a replay should catch.
Replay against the candidate: accepted, rejected, unknown
I used plain HTTP rather than the SDK to keep dependencies down.
# replay.py
import json
import os
import sys
import httpx
from call_shapes import SHAPES
BASE = "https://api.anthropic.com/v1"
HEADERS = {
"x-api-key": os.environ["ANTHROPIC_API_KEY"],
"anthropic-version": "2023-06-01",
"content-type": "application/json",
}
def replay(model: str) -> dict:
"""Send each shape with max_tokens 16 and report whether it was accepted."""
results = {}
with httpx.Client(timeout=30) as client:
for name, body in SHAPES.items():
r = client.post(f"{BASE}/messages", headers=HEADERS,
json={"model": model, **body})
if r.status_code == 200:
results[name] = {"verdict": "accepted"}
elif r.status_code == 400:
msg = r.json().get("error", {}).get("message", "")
results[name] = {"verdict": "rejected", "reason": msg[:300]}
else:
# 429 / 529 / 5xx say nothing about shape -> "unknown"
results[name] = {"verdict": "unknown", "status": r.status_code}
return results
if __name__ == "__main__":
out = replay(sys.argv[1])
print(json.dumps(out, indent=2))
sys.exit(1 if any(v["verdict"] == "rejected" for v in out.values()) else 0)I separate "accepted / rejected / unknown" on purpose. If a 529 overload or a 429 counted as a rejection, a busy day would block your switch for no reason. Only a 400 is a shape problem. Deciding that up front keeps the verdict stable.
python replay.py <candidate-model-id>
echo $? # 0: every shape accepted, 1: at least one rejectedWhen something is rejected, read the reason. Only then open the migration guide, and fix just that part.
Use the Models API's capability info as a shortcut, not as the verdict
Sending real requests is certain, if a bit blunt. As I read the release notes, the Models API response gained capability information on 7 October, including whether a model accepts thinking being disabled.
I plan to use it as a quick lookup before the replay. I don't read it with a fixed schema, since the exact shape may change between versions.
# capabilities.py
import httpx
from replay import BASE, HEADERS
def thinking_disabled_supported(model: str):
"""Return True / False / None (no information)."""
r = httpx.get(f"{BASE}/models/{model}", headers=HEADERS, timeout=30)
if r.status_code != 200:
return None
node = r.json().get("capabilities", {})
for key in ("thinking", "types", "disabled"):
if not isinstance(node, dict) or key not in node:
return None # not found -> "unknown"; let the replay decide
node = node[key]
if isinstance(node, bool):
return node
if isinstance(node, dict) and "supported" in node:
return bool(node["supported"])
return NoneThe None path is the point. When the capability info isn't there, I read it as "unknown", not "unsupported", and hand the decision to the replay. Capability info is the shortcut you check first; the replay is the final confirmation. With two layers, either one going stale doesn't turn into an incident.
The switch itself takes four steps
- Query capabilities with
capabilities.py; if you getFalse, fix the code paths using that shape first. - Run
replay.pyagainst the candidate. It takes seconds and costs next to nothing. - Only when the exit code is 0, change the production model ID.
- Right after the change, run
replay.pyagain with the production configuration.
The fourth step is unglamorous but earns its keep. Typos in the change and environment mismatches never show up in a pre-check.
What the replay cannot catch
To be fair about limits: this only confirms that the request is accepted.
Output quality, cost changes, and differences in behavior need their own checks. Accepted doesn't mean safe. The replay keeps you from tripping at the door; it doesn't certify the migration.
Still, tripping less often leaves more energy for the quality checks. I'd like to spend fewer sleepless nights staring at production logs, and that's the reason this small script sits in my toolbox.
Switch your own calls first, then the model. Since I fixed that order, I haven't seen a 400 as the first thing production says.
As a next step, try writing down just three calls your code really makes in SHAPES today. Three is enough for the replay to start running.