On the night I finished tidying a set of runbooks, I wrote one line at the end of the change log: "about 150.9 KB (about 45,700 tokens) down to about 101 KB (about 30,600 tokens), a 33% cut." The moment I saved it, I stopped. A 33% share of what, exactly?
What I had counted was only the size of the documents a scheduled task loads when it starts, the bytes read once. How many times those bytes get read again during a run appeared nowhere in the number.
Here is the recount first. Trimming 12,960 bytes from the preload came to about 5.9% of cumulative re-reads, roughly 1.3 average tool calls' worth. Folding three check-only calls into other calls, with almost no extra reading, removed about 14.7%, roughly 3.2 calls' worth. Cutting calls beat cutting documents.
I'm an indie developer who leaves the upkeep of four Lab sites to scheduled tasks, and this number changed how I spend the evenings I set aside for tidying runbooks. Here is how to build the ruler.
The premise: context is re-read on every call
Each time the agent calls a tool, the conversation so far and every document already loaded go back in as input. Read a 20 KB document on the first call and those 20 KB ride along on every call after it.
I can't see from the outside how a plan's usage limit is counted internally. API billing discounts cached input, but how much that discount helps depends heavily on the mix, as I covered in A 75% cheaper cache read moved one bill by 29% and another by 2.8%.
So the ledger below does not predict a bill. It is a relative ruler for asking which change probably helps more than which other change. Two values are assumptions, not measurements:
- Base tokens (system prompt, tool definitions, auto-loaded settings files) are set to 30,000. The real value differs per environment, so I recompute at 20,000 and 40,000 to confirm the conclusion holds.
- Mixed Japanese-and-English text is set to 3.3 bytes per token. That is the figure I used in my own estimate, not the tokenizer's real ratio.
A ledger that counts re-reads
This small script reads a TSV of call number, label, and output bytes, and sums the context each call re-read.
#!/usr/bin/env python3
"""Count cumulative re-reads for one scheduled run.
Every tool call re-reads the context built up so far.
Input TSV: call_no<TAB>label<TAB>output_bytes (same call_no is summed)
Usage: reread_ledger.py run.tsv [base_tokens=30000] [bytes_per_token=3.3]
"""
import csv, sys
from collections import OrderedDict
def load(path):
calls = OrderedDict()
with open(path, newline="", encoding="utf-8") as f:
for r in csv.reader(f, delimiter="\t"):
if len(r) < 3:
continue
no = int(r[0])
label, nbytes = calls.get(no, ("", 0))
calls[no] = ((label + "+" if label else "") + r[1], nbytes + int(r[2]))
return [(no, lab, b) for no, (lab, b) in sorted(calls.items())]
def ledger(rows, base_tokens=30000, bytes_per_token=3.3):
ctx, total, out = float(base_tokens), 0.0, []
for no, label, nbytes in rows:
total += ctx # context this call re-reads
out.append((no, label, nbytes, round(ctx), round(total)))
ctx += nbytes / bytes_per_token # its output joins the context afterwards
return total, out
def as_calls(rows, edited, base_tokens=30000, bpt=3.3):
"""How many average calls' worth of re-reads the edit removed"""
t0, _ = ledger(rows, base_tokens, bpt)
t1, _ = ledger(edited, base_tokens, bpt)
return (t0 - t1), (t0 - t1) / (t0 / len(rows))
if __name__ == "__main__":
rows = load(sys.argv[1])
base = float(sys.argv[2]) if len(sys.argv) > 2 else 30000
bpt = float(sys.argv[3]) if len(sys.argv) > 3 else 3.3
total, out = ledger(rows, base, bpt)
for no, label, nbytes, ctx, cum in out:
print(f"{no:>3} {label:<24} +{nbytes:>6}B ctx {ctx:>7,} cum {cum:>10,}")
print(f"cumulative re-reads: {round(total):,} tokens ({len(rows)} calls)")The order of the two lines matters. Call i re-reads the context before its own output lands, so add to the total first and grow the context second. Swap them and every call is charged for its own output too.
Rows sharing a call number are merged because one bash call often cats several documents. What joins the context is a call, not a file.