# Context Budget

**Most tokens are I/O, not thinking.** Keep raw input and mechanical output
out of your main context. Keep reasoning in it. Spotify measured 82 to 94
percent fewer main-context tokens this way.

**Read the slice, not the file.** Search first, then read the exact range.
Read a whole file only when it is short (about 350 lines) or the task needs
all of it. Example: `grep -n "def save"`, then read lines 120 to 180.

**Delegate bulk reads to a worker.** A worker is a subagent or a cheaper
model with its own context. Delegate when a file is long, when a question
spans three or more files, or when a diff is large. Ask for findings, not
the corpus: symbol, file, fact, uncertainty, what needs an exact read.

**Never delegate reasoning.** Debugging, architecture, security,
concurrency and ambiguous contracts stay with you. A cheap worker misses
subtle faults; Spotify's missed a thread-safety bug.

**Re-read the original before every edit.** A summary is not a source.
Worker line numbers are guesses; locate with a tool, then read the exact
lines you change.

**Write predictable output straight to disk.** When a reference file makes
the result about 80 percent predictable (tests after a pattern, config,
stubs, adapters), give the worker the reference and the target path. Take
back only path, line count and lint and test status. Read the artifact
only where a check fails.

**Demand a strict output contract.** Tell every worker the exact fields to
return. No greeting, no restated task, no summary, no code fences. Unknown
values say `unspecified`; they are never guessed.

**Do not bypass the gate.** Do not use `cat`, `head` or `tail` in a shell
to read a file you would not read whole. Small reads stay direct:
delegating 100 lines costs more time than it saves.

**Put stable text first when you build prompts.** Tools, rules, schemas
and reference data go before the changing task, byte for byte identical,
so provider prompt caches hit. Send the same content once; refer to it by
path or hash afterwards.

**Compress prose, never code you will edit.** Summaries and learned
compression suit logs, transcripts and old history. They do not suit
source before an edit, migrations, exact config or security logic. Test a
compact data format against the real model before you adopt it.

**Measure before you claim a saving.** Count main-context tokens, worker
tokens and cost apart. Never cut tokens without a quality check: a
smaller context that loses the fact is a regression.

**Limits.** In Claude Code a hook refuses a whole read of a long text
file in the main session. Read a range or delegate. The owner can raise
the threshold or switch it off with `BESTAICONFIG_READ_GATE_LINES`
(`0` means off).
