◉CLAUDE LABJP
●2.1.292 — Claude Code 2.1.292 (Oct 6) fixes cloud sessions dropping answers to permission prompts and the last messages of a session being lost on quit●MODELS API — The Models API now reports in capabilities whether a model accepts thinking turned off (Oct)●11/30 — 54 days until Sonnet 4.5 retires on the Claude API. Move to Sonnet 5.5, and check your thinking settings and tool_choice first●MODS — A Zenn post walks through installing Claude Code Mods safely. The real question is what to check before you install, and how to back out●NEW — When Pro looks unpaid, the screens to check depend on whether you bought on the web or in the store. I sorted that out step by step●THINKING — Sonnet 5.5 thinking blocks only work in the account that produced them. Sent from another account, they are silently dropped●2.1.292 — Claude Code 2.1.292 (Oct 6) fixes cloud sessions dropping answers to permission prompts and the last messages of a session being lost on quit●MODELS API — The Models API now reports in capabilities whether a model accepts thinking turned off (Oct)●11/30 — 54 days until Sonnet 4.5 retires on the Claude API. Move to Sonnet 5.5, and check your thinking settings and tool_choice first●MODS — A Zenn post walks through installing Claude Code Mods safely. The real question is what to check before you install, and how to back out●NEW — When Pro looks unpaid, the screens to check depend on whether you bought on the web or in the store. I sorted that out step by step●THINKING — Sonnet 5.5 thinking blocks only work in the account that produced them. Sent from another account, they are silently dropped
Articles/API & SDK
⬡ API & SDK/2026-06-22Advanced

Putting a Ceiling on the pause_turn Loop: Running Long Server Tools Safely Unattended

A production design for continuing pause_turn safely in unattended runs, where long server tools like web_search and code execution are involved. Covers branching all four stop_reason values in one loop, capping continuations and wall-clock time, and accumulating usage across paused segments.

Claude49API28pause_turntool-use24production111

✦ Premium Article

As an indie developer, I batch-generate articles for several sites overnight, and one morning a few of them were missing the fresh information they were supposed to have searched for. Nothing in the error log. When I opened a saved response, stop_reason was pause_turn — and my generation loop had happily stopped there. web_search hadn't finished in a single round trip; it had returned a "pause," and my loop read that pause as completion.

You don't see pause_turn often, because short prompts never produce it. But the moment you involve a long server tool — web_search, web_fetch, or code execution — it can show up. And it arrives as a normal, successful response, not an exception, so as long as you swallow it you'll never notice. In unattended runs, that's exactly where silent truncation hides.

pause_turn Is a Third State, Neither Error Nor Done

If you treat every stop_reason as "the reason the response ended," pause_turn will trip you up every time. The starting point is to split the values into "finished" and "still going."

stop_reasonStateWhat you must do next
end_turnDone normallyNothing. The output is final
max_tokensCut off mid-outputDecide: continue, or record as incomplete
tool_useContinue (client side)Append a tool_result and re-request
pause_turnContinue (server side)Append the response as-is and re-request
refusalSafety refusalDon't retry; handle it by design

tool_use and pause_turn look similar, but what you append differs. With tool_use you run the tool yourself and add a new user message containing the tool_result. With pause_turn, the partial output from the server-side tool is already inside the assistant response, so you append the response blocks as-is, with no extra input, and the turn keeps going. The basic branching itself is laid out in implementation patterns for not dropping stop_reason, so if you aren't even checking max_tokens yet, start there. This piece focuses on what comes after: how to design for pause_turn when you're running long tools unattended.

Reproduce It First — Which Tools Produce pause_turn

Before defending against it blindly, it helps to make your own code emit a pause_turn once. Enable a server-side tool and ask something that likely needs several searches.

import anthropic
 
client = anthropic.Anthropic(api_key="YOUR_API_KEY")
 
resp = client.messages.create(
    model="claude-opus-4-8",
    max_tokens=4096,
    tools=[{"type": "web_search_20260318", "name": "web_search"}],
    messages=[{
        "role": "user",
        "content": "Summarize three June 2026 Claude Developer Platform updates with sources",
    }],
)
print(resp.stop_reason)        # may be pause_turn
print([b.type for b in resp.content])
# e.g. ['text', 'server_tool_use', 'web_search_tool_result', 'text']

The thing to internalize: even on pause_turn, content already holds partial blocks — text, server_tool_use, web_search_tool_result. It is not an empty response. That's precisely why swallowing it leaves you with a half-finished body presented as final. My own first mistake was exactly this: mistaking the in-progress text for the finished product.

✦

Thank you for reading this far.

Continue Reading

What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.

WHAT YOU'LL LEARN
✦Branch pause_turn, tool_use, end_turn, and max_tokens in a single continuation loop so server tools never get silently truncated
✦Add a continuation cap, a wall-clock budget, and cross-segment usage accumulation so a runaway turn can't quietly rack up cost in an overnight batch
✦Take away paste-ready code for streaming event ordering and an observability log that stops you from misreading pause_turn as end_turn
Secure payment via Stripe · Cancel anytime
✦

Unlock This Article

Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.

or
Unlock all articles with Membership →
Share

Thank You for Reading

Claude Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • ✦Copy-paste ready implementation code
  • ✦New advanced guides published daily
  • ✦$5/mo or $15 for lifetime access
View Membership →

Related Articles

⬡ API & SDK2026-07-24
When Memory Store Listings Returned Half the Rows: Migrating to agent-memory-2026-07-22
agent-memory-2026-07-22 changed memories.list: fixed ordering, stricter depth, segment-based path_prefix matching. How an audit quietly halved, an access layer that survives it, finding every call site with ast, and measured depth=1 traversal cost.
⬡ API & SDK2026-05-26
Stabilizing Claude API Structured Responses in Production — Notes on tool_use, JSON Schema, and Layered Validation
Getting Claude to return JSON takes a few lines. Keeping that JSON usable in production is a different problem. Here is the layered design I landed on after running a wallpaper classification pipeline through Claude API, built around tool_use, JSON Schema, and domain validation.
⬡ API & SDK2026-05-22
Why tool_result could not be submitted Keeps Coming Back, and How to Build a Recovery Handler That Actually Holds
Run a Claude agent long enough and one day it starts: 'tool_result could not be submitted', back to back, and retries change nothing. The error message hides four completely different root causes. Here is what I learned debugging it on the always-on agent jobs I run as an indie developer, with the TypeScript recovery handler I now ship in production.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links