For a while I had a nightly batch that checked store listings for my wallpaper apps in several languages. One night the Claude API kept returning 529 for a stretch of time. Overload errors were something I had planned for, and there was a fallback to another vendor's model waiting behind it. When I opened the log the next morning, the fallback path had thrown a TypeError, and not a single result had been saved.
What made my stomach sink was not the overload on Anthropic's side. It was realizing that the insurance I had written myself had torn first. That fallback, wired up about half a year earlier, had never fired once until that night.
The cause was the cache_control marker I had attached for prompt caching. It was still sitting on the content blocks when they were handed to the other vendor. The same shape of failure is reported in langchain-ai/langchain issue #33709, where with_fallbacks from Anthropic to another model breaks the moment it switches. My code did not use LangChain, but I had treated the boundary the same careless way.
cache_control is a decoration you attach just before Claude
In the Anthropic Messages API, adding cache_control: {"type": "ephemeral"} to a content block in system or messages caches the prefix up to that point, and you can place up to four such breakpoints per request. With a glossary pinned in the system prompt, cache reads on Opus 5.5 come down to $0.20 per million tokens, which matters when a batch sends the same glossary dozens of times a night. How much your bill actually drops depends on the token mix, and I wrote about that in Cache reads got 75% cheaper, yet the bill can drop 10x differently depending on your mix.
The key means nothing outside Anthropic's API, though. When a content block reaches another vendor's chat format with anything besides type and text on it, either the client library's type check raises a TypeError or the server rejects the request as a 400. Either way, nothing arrives.
I will admit that my first suspect was the other vendor's client library. The traceback pointed at its message-conversion function, so I assumed an outdated version and upgraded it, and it failed on exactly the same line. What stopped my hands was printing the payload right before the send. There, on the glossary block, sat the familiar "cache_control": {"type": "ephemeral"}, exactly as I had written it for Claude.
That is the uncomfortable part of a bug like this: the object that breaks the fallback is the one you are proudest of. The cache marker had been saving me money every night, and I had no reason to look at it as a liability. It only became one at the boundary, and I had never looked at the boundary.
What I had missed was structural. I kept a single messages array as "the conversation record," decorated it with cache markers for Claude, and then passed that very same array to the fallback. The record and the wire format were never separated.
Attach the cache marker just before Claude, and never carry it across the boundary. Once I wrote that sentence down, the code became simple.
A stripping function and a guard right before sending
Stripping comes first. Two rules keep it short: never mutate the original array, and run system through the same cleaner.
ANTHROPIC_ONLY_KEYS = {"cache_control"}
def strip_anthropic_extensions(messages, system):
"""Strip Claude-specific keys right before handing messages to another vendor. Does not mutate the input."""
def clean(content):
if isinstance(content, str):
return content
return [
{k: v for k, v in block.items() if k not in ANTHROPIC_ONLY_KEYS}
for block in content
]
return (
[{**m, "content": clean(m["content"])} for m in messages],
clean(system),
)Stripping alone would let me repeat the same accident the day I add another Claude-only key. So I placed a guard immediately before the send, and it refuses to send if anything is left over.
def assert_portable(obj, path="$"):
"""Guard right before crossing the boundary. Stops the send if a Claude-only key survived."""
if isinstance(obj, dict):
for k, v in obj.items():
if k in ANTHROPIC_ONLY_KEYS:
raise ValueError(f"{path}.{k} is still present")
assert_portable(v, f"{path}.{k}")
elif isinstance(obj, list):
for i, v in enumerate(obj):
assert_portable(v, f"{path}[{i}]")There is one more difference in shape. Anthropic takes system as a top-level argument, while most chat formats expect it as the first message with the system role. Flattening block arrays into plain strings happens in the same place.
def to_chat_format(messages, system):
"""Convert Anthropic's shape (separate system, block-array content) to a generic chat format."""
def flatten(content):
if isinstance(content, str):
return content
return "\n".join(b["text"] for b in content if b.get("type") == "text")
out = [{"role": "system", "content": flatten(system)}]
out += [{"role": m["role"], "content": flatten(m["content"])} for m in messages]
return outI ran these three functions locally against arrays that carried cache_control, and confirmed three things: the guard stops the decorated payload, the stripped payload passes, and the original arrays are untouched. Blocks other than text (tool_use, tool_result, images) cannot be flattened, so I drew a line there: conversations that use tools are excluded from the fallback entirely. I would rather wait for Claude to recover and rerun than receive results in a broken shape.
Switch on 529 and connection errors, never on 429
The calling side ended up like this. The Claude function wraps only the failures that deserve a switch into ClaudeUnavailable.
import os
import time
import anthropic
client = anthropic.Anthropic() # ANTHROPIC_API_KEY=YOUR_API_KEY via environment
MODEL = "claude-sonnet-5"
class ClaudeUnavailable(Exception):
pass
def call_claude(messages, system):
if os.environ.get("FORCE_FALLBACK") == "1":
raise ClaudeUnavailable("forced by FORCE_FALLBACK=1")
try:
r = client.messages.create(
model=MODEL, max_tokens=1024, system=system, messages=messages
)
except (anthropic.InternalServerError, anthropic.APIConnectionError) as e:
# Only 5xx (including 529 overloaded) and connection failures trigger a switch
raise ClaudeUnavailable(str(e)) from e
return r.content[0].text
def call_other(chat_messages):
# Replace with your own client; it receives an OpenAI-compatible chat payload
...
def ask(messages, system):
for attempt in range(2):
try:
return call_claude(messages, system), MODEL
except ClaudeUnavailable:
if attempt == 0:
time.sleep(30) # a short congestion often clears within this window
stripped_messages, stripped_system = strip_anthropic_extensions(messages, system)
payload = to_chat_format(stripped_messages, stripped_system)
assert_portable(payload) # stops here if anything survived
return call_other(payload), "fallback"Leaving anthropic.RateLimitError out of that except is deliberate. A 429 means I am sending too fast; it is my problem, not the vendor's congestion. Escaping to another vendor only postpones the same wall to tomorrow, and waiting a little clears it. The Python SDK already retries twice by default before raising, so a 5xx or connection failure that still gets through is the only thing I read as "Claude is unavailable right now."
I also decided how long to wait before switching. The SDK's retries use exponential backoff and finish within a few seconds, so after that I add a 30-second pause of my own, and only a 5xx that survives the pause sends the request to the other vendor. The batch just has to finish by morning, so tens of seconds cost nothing, and a short congestion often clears within that window, which keeps the request on Claude where the glossary cache still applies.
One more habit came out of this: the fallback vendor's output is stored in the same table, but every row now carries a model column. It sounds trivial, but on a morning when three languages look slightly off, being able to filter for model=fallback in a second turns a vague suspicion into a five-minute check.
Returning which model answered is another lesson from that night. The fallback model has different habits with terminology, and I wanted the morning review to tell me which rows to doubt. Bouncing between Sonnet and Haiku inside Claude is a different, gentler story that I covered in High availability with the Claude API — making Sonnet / Haiku / Opus multi-model fallback work in production; a path that leaves for another vendor gets one extra notch of caution.
A fallback that never fires has never been tested
The biggest oversight was calling a path "insurance" when it had not been walked once in six months. Now I run the batch once a week with FORCE_FALLBACK=1 set, pushing only the first item through the fallback route. It costs one request, and it leaves exactly one model=fallback line in the log. If that line is missing in the morning, I know the insurance itself is what broke.
Put simply, the stripping function, the pre-send guard, and the weekly forced trigger — it took all three before the fallback deserved the name.
If you have a fallback of your own, I'd suggest forcing it to fire once on a day when Claude is perfectly healthy. If it is going to break, that day is a far better time than an overloaded night. As an indie developer running this alone at Dolice, I now put a guard on every piece of data that crosses a vendor boundary, and I may be overcautious, but that habit has held up so far.