Compressing recurring tool guidance
GitHub measured 2.9% (~1,300 tokens per turn) from compressing their recurring prompts ~50%. FreeCode’s compression shipped 2026-09-04 with all three gate suites green — after the safety guard caught a real regression in the first draft.
The problem, in plain words
Some text rides on every single model call for the life of a session: the
system prompt (how to behave) and the tool definitions (what read, bash,
grep etc. do and what parameters they take). Unlike a file read — paid when
it happens — this block is a fixed tax on every request, forever.
Measured on a real recorded request (the numbers come from the rollout log, not guesses):
| Block | Size | Biggest items |
|---|---|---|
System prompt (system.md) | ~2.2K tokens | — |
| Tool descriptions + schemas (16 tools) | ~4.7K tokens | grep 2,454 chars, bash 2,427, memory 1,944, todowrite 1,557, agent 1,423 |
| Recurring block total | ~6.9K tokens / request |
Halving it, GitHub-style, would save ~3.4K tokens on every request — before the conversation has even started growing.
Why this is the dangerous one
Compressing prose sounds safe. It isn’t. GitHub’s own prompt compression silently serialized their parallel agents — a shortened instruction lost the nuance that told the model to issue tool calls in parallel, everything got slower and more expensive, and no offline token count could have caught it. Only behavioral testing did.
FreeCode’s system prompt carries the same load-bearing instruction (batch independent tool calls into one response). So before any compression lands, the regression it could cause needed a permanent, automated tripwire.
How I implemented it (the guard — built first, on purpose)
The tripwire is an eval expectation, expectParallelTools
(apps/core/src/eval/types.ts, scored in eval/scorers/trajectory.ts), plus
the trajectory case parallel-batch-two-reads in evals/trajectory.jsonl: a
prompt whose correct trajectory is two concurrency-safe reads in one
assistant message. If any future prompt edit — this compression or any
other — weakens the batching instruction, the trajectory suite goes red.
Two details worth understanding:
- It scores what the response emitted, not what ran (
ModelSpan.toolCallsfrom the trace). A batch that gets permission-denied still counts as batching — we’re testing the model’s behavior, not the sandbox’s mood. - The dataset loader rejects
expectParallelToolsvalues below 2, because “at least one tool call” would assert nothing.
This case guards every future prompt edit, not just this one — it’s the test the suite should have had since the parallel-batching instruction was added.
How the compression went (and what the guard caught)
The compression ran in exactly the order the plan demanded, and the guard earned its keep on day one:
system.mdand the five fattest tool descriptions were compressed by hand, section by section — keeping the parallel-batching instruction deliberately strong.- The first compressed variant (v1) failed
explore-mode-stays-readonlytwice in a row: the trial data showed a rejectedquestioncall burning the extra turn. The cut responsible: v1 dropped “requesting input from the user is a blocking action.” A control run on the original prompt (via git stash) passed, convicting the compression. - v2 restored that rule plus an explicit “if an action is blocked, report it plainly instead of asking” — and matched the original prompt’s score.
- The full release gate then ran on v2: all three suites GATE OPEN.
That’s the whole cost-efficiency principle in miniature: the offline token count said v1 was fine; only behavioral measurement caught what it broke.
Performance improvements
| Metric | Before | After | Δ |
|---|---|---|---|
system.md | 8,862 chars (~2.2K tokens) | 5,446 chars | −38.5% |
| Fat-five tool descriptions | ~9.8K chars | roughly halved | ~−50% |
| Trajectory gate | — | GATE OPEN (guard case green at 3 trials) | |
| Coding gate | — | GATE OPEN (11/11) | |
| Judged gate | — | GATE OPEN — mean 4.83/5, no case below 4 |
The judge’s own rationales confirm the style survived compression: “direct, no unnecessary padding”, “omitting preamble as requested” — the exact qualities a clumsy prompt trim would have destroyed.