◉CLAUDE LABJP
●2.1.292 — Claude Code 2.1.292 (Oct 6) fixes cloud sessions dropping answers to permission prompts and the last messages of a session being lost on quit●MODELS API — The Models API now reports in capabilities whether a model accepts thinking turned off (Oct)●11/30 — 54 days until Sonnet 4.5 retires on the Claude API. Move to Sonnet 5.5, and check your thinking settings and tool_choice first●MODS — A Zenn post walks through installing Claude Code Mods safely. The real question is what to check before you install, and how to back out●NEW — When Pro looks unpaid, the screens to check depend on whether you bought on the web or in the store. I sorted that out step by step●THINKING — Sonnet 5.5 thinking blocks only work in the account that produced them. Sent from another account, they are silently dropped●2.1.292 — Claude Code 2.1.292 (Oct 6) fixes cloud sessions dropping answers to permission prompts and the last messages of a session being lost on quit●MODELS API — The Models API now reports in capabilities whether a model accepts thinking turned off (Oct)●11/30 — 54 days until Sonnet 4.5 retires on the Claude API. Move to Sonnet 5.5, and check your thinking settings and tool_choice first●MODS — A Zenn post walks through installing Claude Code Mods safely. The real question is what to check before you install, and how to back out●NEW — When Pro looks unpaid, the screens to check depend on whether you bought on the web or in the store. I sorted that out step by step●THINKING — Sonnet 5.5 thinking blocks only work in the account that produced them. Sent from another account, they are silently dropped
Articles/API & SDK
⬡ API & SDK/2026-10-07Intermediate

Before You Switch Models, Replay Your Own Call Shapes Against the Candidate for a Few Tokens

Change a model ID and the first production request is where a 400 shows up. Here is a small replay script that checks your own call shapes against a candidate model with max_tokens 16 before you switch.

Claude API125model migration3Models APIpreflight4design5

One night I changed a single model ID string and ran my script.

A 400 came back. The message made it clear the model wasn't at fault. One argument I had carried over from the old setup no longer matched what the new model accepts. The fix took a few minutes. What made my stomach sink was that this was the first request production would send.

Since then I replay my own call shapes against a candidate model before switching. It isn't an evaluation platform. It's a handful of tiny requests.

What breaks on a switch is the shape of the call, not the model's quality

When we swap models, we tend to worry about output quality and cost. Rightly so.

But before quality comes a door: is the request accepted at all? If you trip there, comparing quality is beside the point. The causes usually fall into three groups.

CauseHow it looks in your codeWhen you notice
A parameter no longer fitsA thinking setting or an argument you've always passed400 on the first request
A combination no longer fitsForcing a tool with tool_choice while using another feature400 only on one code path
A capability differsA feature the old model supported isn't thereWhen a screen or batch job reaches it

The second and third are the nasty ones. They live on paths you rarely hit, so a single smoke test won't find them.

The migration guide is a map; your calls are the current location

Migration guides list what changed per model, and I'm grateful for them. But they can't tell you which of those changes your code actually steps on.

So I split the work:

  1. Write down the shapes of the calls my code really makes.
  2. Send those shapes, as they are, to the candidate and see what's accepted.
  3. Read the guide only for the shapes that were rejected.

That's far shorter than reading every item and cross-checking, because the checking is limited to my own calls.

Put your call shapes in one file

Start with an inventory: one entry per distinct combination of arguments your code sends.

# call_shapes.py
# The shapes of calls the production code actually makes, collected in one place.
# No model ID here (the candidate is injected at replay time).
# max_tokens is tiny: we only want an accept/reject verdict.
 
DUMMY_TOOL = {
    "name": "ping",
    "description": "Dummy tool for connectivity checks",
    "input_schema": {"type": "object", "properties": {}},
}
 
SHAPES = {
    "plain": {
        "max_tokens": 16,
        "messages": [{"role": "user", "content": "Reply with ok"}],
    },
    # Keep this only if your own code sets thinking explicitly
    "thinking_disabled": {
        "max_tokens": 16,
        "thinking": {"type": "disabled"},
        "messages": [{"role": "user", "content": "Reply with ok"}],
    },
    "tool_choice_any": {
        "max_tokens": 16,
        "tools": [DUMMY_TOOL],
        "tool_choice": {"type": "any"},
        "messages": [{"role": "user", "content": "Call ping"}],
    },
    "system_prompt": {
        "max_tokens": 16,
        "system": "You are an assistant for connectivity checks.",
        "messages": [{"role": "user", "content": "Reply with ok"}],
    },
}

The rule: don't list shapes your code never sends. Chasing every possible argument bloats the file and dulls the exercise. Do include the call that only a monthly batch job makes. Those rarely-hit paths are exactly what a replay should catch.

Replay against the candidate: accepted, rejected, unknown

I used plain HTTP rather than the SDK to keep dependencies down.

# replay.py
import json
import os
import sys
 
import httpx
 
from call_shapes import SHAPES
 
BASE = "https://api.anthropic.com/v1"
HEADERS = {
    "x-api-key": os.environ["ANTHROPIC_API_KEY"],
    "anthropic-version": "2023-06-01",
    "content-type": "application/json",
}
 
 
def replay(model: str) -> dict:
    """Send each shape with max_tokens 16 and report whether it was accepted."""
    results = {}
    with httpx.Client(timeout=30) as client:
        for name, body in SHAPES.items():
            r = client.post(f"{BASE}/messages", headers=HEADERS,
                            json={"model": model, **body})
            if r.status_code == 200:
                results[name] = {"verdict": "accepted"}
            elif r.status_code == 400:
                msg = r.json().get("error", {}).get("message", "")
                results[name] = {"verdict": "rejected", "reason": msg[:300]}
            else:
                # 429 / 529 / 5xx say nothing about shape -> "unknown"
                results[name] = {"verdict": "unknown", "status": r.status_code}
    return results
 
 
if __name__ == "__main__":
    out = replay(sys.argv[1])
    print(json.dumps(out, indent=2))
    sys.exit(1 if any(v["verdict"] == "rejected" for v in out.values()) else 0)

I separate "accepted / rejected / unknown" on purpose. If a 529 overload or a 429 counted as a rejection, a busy day would block your switch for no reason. Only a 400 is a shape problem. Deciding that up front keeps the verdict stable.

python replay.py <candidate-model-id>
echo $?   # 0: every shape accepted, 1: at least one rejected

When something is rejected, read the reason. Only then open the migration guide, and fix just that part.

Use the Models API's capability info as a shortcut, not as the verdict

Sending real requests is certain, if a bit blunt. As I read the release notes, the Models API response gained capability information on 7 October, including whether a model accepts thinking being disabled.

I plan to use it as a quick lookup before the replay. I don't read it with a fixed schema, since the exact shape may change between versions.

# capabilities.py
import httpx
 
from replay import BASE, HEADERS
 
 
def thinking_disabled_supported(model: str):
    """Return True / False / None (no information)."""
    r = httpx.get(f"{BASE}/models/{model}", headers=HEADERS, timeout=30)
    if r.status_code != 200:
        return None
    node = r.json().get("capabilities", {})
    for key in ("thinking", "types", "disabled"):
        if not isinstance(node, dict) or key not in node:
            return None          # not found -> "unknown"; let the replay decide
        node = node[key]
    if isinstance(node, bool):
        return node
    if isinstance(node, dict) and "supported" in node:
        return bool(node["supported"])
    return None

The None path is the point. When the capability info isn't there, I read it as "unknown", not "unsupported", and hand the decision to the replay. Capability info is the shortcut you check first; the replay is the final confirmation. With two layers, either one going stale doesn't turn into an incident.

The switch itself takes four steps

  1. Query capabilities with capabilities.py; if you get False, fix the code paths using that shape first.
  2. Run replay.py against the candidate. It takes seconds and costs next to nothing.
  3. Only when the exit code is 0, change the production model ID.
  4. Right after the change, run replay.py again with the production configuration.

The fourth step is unglamorous but earns its keep. Typos in the change and environment mismatches never show up in a pre-check.

What the replay cannot catch

To be fair about limits: this only confirms that the request is accepted.

Output quality, cost changes, and differences in behavior need their own checks. Accepted doesn't mean safe. The replay keeps you from tripping at the door; it doesn't certify the migration.

Still, tripping less often leaves more energy for the quality checks. I'd like to spend fewer sleepless nights staring at production logs, and that's the reason this small script sits in my toolbox.

Switch your own calls first, then the model. Since I fixed that order, I haven't seen a 400 as the first thing production says.

As a next step, try writing down just three calls your code really makes in SHAPES today. Three is enough for the replay to start running.

Share

Thank You for Reading

Claude Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • ✦Copy-paste ready implementation code
  • ✦New advanced guides published daily
  • ✦$5/mo or $15 for lifetime access
View Membership →

If you found this article helpful, a small tip ($1.50) would mean a lot to us. Your support helps keep this site ad-free and covers server and hosting costs.

Related Articles

⬡ API & SDK2026-09-28
The Night Claude Returned 529, My Fallback Failed Too — Strip cache_control at the Provider Boundary
My Claude API fallback threw a TypeError because content blocks still carried cache_control when they reached another vendor. Here is the stripping function, the pre-send guard, and the weekly forced-fallback routine that finally made the fallback a real safety net.
⬡ API & SDK2026-09-17
I Send Images Twenty at a Time — The Day the 21st Image Changed the Rules for the Other 61
Sending 62 images in one request failed with invalid_request_error. The cause was the image count, not the payload size. Here is how I recounted visual tokens as 28px patches and built batches that respect count, dimensions, and payload at once.
⬡ API & SDK2026-09-06
The translation read perfectly and still crashed at runtime
When you translate app strings with the Claude API, the meaning can be right while the format specifiers quietly break. Here is a severity-aware acceptance check and a repair loop that only re-translates the broken lines.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links