Skip to Content
InternalsCompaction

Compaction

Every request sends the whole conversation. The model has no memory of the last turn, so turn 40 re-sends turns 1–39 in full. Two things break as that grows:

  1. The window. Eventually the request doesn’t fit and the provider rejects it.
  2. The bill. Long before it stops fitting, you are paying to re-send the same history on every single turn.

Compaction is the answer to both: summarize the old part of the conversation, drop the messages it replaces, and carry on. In the memory layer vocabulary this is the working memory layer — the only one scoped to a single session.

Name collision, worth knowing up front. The class that implements this is called MemoryService (compaction/service.ts:36) and its state lives in ~/.freecode/memory/<sessionId>/memory.json. That is not the persistent memory system, which is MemoryStore under ~/.freecode/projects/<project>/memory/. Same word, different layer.

Two ways the context gets smaller

Compaction is one of two mechanisms, and confusing them makes both look broken:

CompactionTool-result pruning
What it doessummarizes old turns, deletes themreplaces oversized tool results with a short marker
Scoperewrites history wholesaleonly the view sent to the provider
Prompt cacheinvalidates it — expected, and recordeddesigned to never invalidate it
Lives incompaction/agent/prune-state.ts + loop.ts:422

Pruning is the gentle one and runs constantly: every decision is recorded and re-applied verbatim on later turns, and any result already sent at full size is frozen forever, so the bytes at a given position never change and the provider’s cached prefix keeps growing instead of being invalidated two turns back on every turn. Compaction is the blunt one, and it rebuilds the prefix — so it is used rarely and deliberately.

When it fires

Measured beats estimated

The service keeps its own running token count, and it is an estimate over the memory transcript — not a measurement of the request. A turn that calls tools used to be recorded as the four-token stub [Executed N tools], so tool arguments and results — the bulk of a coding session — were invisible to it.

The consequence was not subtle: it under-reported a real 196K context as 14.5K, and auto-compaction never fired before the window overflowed. Two things fixed it. addToolTurn now writes a real transcript of the turn (tool-transcript.ts), and shouldCompact prefers the provider’s own count of the last request’s input (service.ts), falling back to the estimate only when no measurement exists.

The two numbers stay far apart even so, and that is worth knowing when reading a trace: on a measured run of the compaction-boundary eval case the request was 16K tokens while the memory transcript the summarizer works over was 872. The trigger is judged on the former; compact.occurred records the latter.

Cost is the real trigger, not fit

Compaction used to fire only when the next request wouldn’t fit. That is a hard constraint, not the objective. On a 1M-window model it puts the trigger near 968K, so a session peaking at 270K never compacts — and every one of its requests carries the full history. One recorded 7-message session cost 48.1M input tokens that way (tokens.ts:31).

So the threshold is whichever binds first:

const limit = Math.min(windowLimit, getCompactTarget()); // tokens.ts:131 return tokenCount >= Math.max(0, limit - bufferTokens);

DEFAULT_COMPACT_TARGET_TOKENS = 120_000 is the knee: below it a request is cheap at cache-read rates, above it the per-turn cost dominates. A 200K model and a 1M model therefore compact at the same point, because a 120K request costs the same either way.

SettingDefaultPurpose
autoCompactBufferTokens13,000headroom below the limit
preserveRecentTurns2user turns kept verbatim
maxPreserveRecentTokens8,000cap on what “recent” may cost
maxToolOutputChars2,000truncation for tool text in the record
FREECODE_COMPACT_TARGET_TOKENS120,000the cost ceiling; raise for fit-only behaviour
FREECODE_AUTO_COMPACT_TOKENS—force the trigger low, for testing
FREECODE_MAX_TURN_TOKENS—opt-in per-run spend circuit breaker

The env overrides are read per call, not at module load, so a long-running daemon and a test both see a change without restarting. A malformed value logs a warning rather than being ignored silently — a typo there is indistinguishable from compaction being broken, which is usually the thing being tested (tokens.ts:64).

The context-window limit itself comes from models.dev when available; FALLBACK_CONTEXT_LIMIT = 100_000 is the offline floor. Being wrong there only makes the agent compact slightly early, never lose data, so one safe constant beats a per-model table that drifts.

What gets kept

selectForCompaction() (selector.ts:16) splits the history in two:

  1. Walk backwards to the second-to-last user message — that boundary and everything after it is preserved.
  2. If that block still exceeds maxPreserveRecentTokens, drop from its front until it fits.
  3. Everything before the boundary is summarized.

If the session has fewer user turns than preserveRecentTurns, nothing is summarized and compaction is a no-op that reports success. The boundary is always a user message, in both the in-memory selection and the on-disk trim (keepLastNUserTurns in session/compact-apply.ts), because a history starting with an assistant reply to a question the model can no longer see is not a valid conversation — and some providers reject it outright.

Writing the summary

Two summarizers, and the second is the fallback for the first.

The LLM summarizer (llm-summarizer.ts) asks the current provider for a handoff document with seven fixed sections — Goal, Done, In Progress, Blocked, Decisions, Relevant Files, Next Steps — at temperature: 0 and 1,024 max tokens. The prompt is explicit about what compaction must not do: don’t invent progress that isn’t in the transcript, preserve exact file paths, identifiers, commands, and unresolved errors, and under Blocked, list only real, still-open blockers. A summary that hallucinates completed work is worse than no summary, because the model will believe it.

The heuristic summarizer (summarizer.ts:113) needs no provider at all. It builds the same section layout by keyword: file paths by regex, status buckets by word-boundary matching (with negation handling, so “no errors” is not a blocker), the first user message as the goal, the last three assistant messages as progress. It runs first, and the LLM result replaces it only if the call succeeds — so a provider failure downgrades the summary instead of failing the compaction (service.ts:168).

Summaries chain. The previous summary is passed into the next one with “carry forward and update, don’t repeat verbatim”, so a session compacted five times still has a thread back to the beginning rather than five disconnected fragments.

Applying it

applyCompaction() (session/compact-apply.ts) is what makes compaction stick:

Without the middle step the next turn’s loadHistory() would reload the full log and silently undo the work.

Around that, runCompaction() (loop.ts:1078) handles the consequences:

  • compaction_start / compaction_complete on the bus, so every frontend can show it.
  • recordInvalidation(...) + bumpCacheGeneration(...) — compaction rebuilds the prefix, so the next request must be a cache miss. Recording it means cache analysis reports an expected miss instead of a harness bug.
  • compact.occurred in the rollout log with before/after tokens, so the erased turns are still accounted for in replay.
  • lastMeasuredContextTokens = 0 — the measurement describes the pre-trim request and would immediately re-trigger compaction if left in place.

Hooks can veto it

PreCompact runs inside compact() and can block (service.ts:139). A block is not a crash: the result comes back { success: false, blocked: true } with a reason, and history is left intact.

The interesting part is what happens next. Without a guard, the next message would re-cross the threshold and re-run the blocked hook — forever. So the service records blockedAtTokenCount and refuses to retry until the conversation has grown another 5,000 tokens (service.ts:34). The count remembered is the same kind of count the decision was made on: mixing a measured 196K with an estimated 14.5K would make the guard compare across scales and never hold (service.ts:101).

PostCompact runs after state is saved, receives success, and can inject context into the freshly-shortened conversation.

When the provider says no anyway

Sometimes the request is rejected as too long despite all of the above — a bad estimate, a sudden huge tool result. compactAndRetry() (loop.ts:996) catches only context-overflow errors, compacts, and sends once more.

Three guards, each earned:

  • Anything that isn’t an overflow is rethrown untouched, so ordinary failures still surface as themselves.
  • If compaction freed nothing — already minimal, or a hook blocked it — the original error is thrown rather than re-sending an identical request.
  • MAX_OVERFLOW_COMPACTIONS = 3, and the retry is deliberately not wrapped in the same handler again. A second overflow in one turn means compaction isn’t converging, and looping burns quota: claude-code recorded 1,279 sessions hitting 50+ consecutive compaction failures — up to 3,272 — for roughly 250K wasted calls a day before adding the same cap.

Where the state lives

~/.freecode/memory/<sessionId>/memory.json

Messages, summaries, token count, compaction count. Written atomically (temp file, then rename) so a crash mid-write can’t corrupt it, and a file that is unreadable is treated as “no memory” and starts fresh rather than bricking the session (storage.ts:39).

A fresh MemoryService is constructed per turn and loads this file in its constructor, which is why the manual /compact path (session.compact, a separate instance) and the loop’s own instance stay consistent across turns.

Known gaps

  1. compact.occurred under-reports what compaction did. It records MemoryService’s estimated transcript sizes — 872 → 748 tokens on a measured run — while the actual trim is keepLastNUserTurns over the session store, against a request that measured 16K. So freecode trace and the eval harness’s Trace.compactedTokens understate the saving by an order of magnitude. ApplyCompactionResult already carries messagesBefore / messagesAfter; neither is recorded.
  2. renderPromptMemoryContext() is dead code. Exported from selector.ts:64, referenced only by comments in loop.ts explaining why it must not be used (it renders recentMessages, which are already in history verbatim and would rewrite the prefix every turn).
  3. getContextLimit(model) ignores its argument and returns the constant floor (tokens.ts:20). Harmless, but the name promises a per-model table that no longer exists.
  4. The estimator is chars ÷ 4 for every model and provider (tokens.ts:13) — no tokenizer, no per-provider adjustment. Acceptable only because the measured count is preferred; it is the fallback that under-reports.
  5. context.overflow is never recorded. The loop detects overflow and retries (loop.ts:1003) but never calls recorder.recordContextOverflow(), so the one event designed to make this visible in replay is absent. It is one of seven rollout event types with no emitter — see Sessions, store & rollout.
  6. The blocked-retry threshold is a fixed 5,000 tokens (service.ts:34) and not part of CompactionConfig, so a hook that wants tighter control can’t have it.

Where to look

You wantFile
The service, hooks, blocked-retry guardapps/core/src/compaction/service.ts
Thresholds, estimator, env overridesapps/core/src/compaction/tokens.ts
What is kept vs summarizedapps/core/src/compaction/selector.ts
The heuristic summaryapps/core/src/compaction/summarizer.ts
The provider-backed summaryapps/core/src/compaction/llm-summarizer.ts
Persisted stateapps/core/src/compaction/storage.ts
Trimming the session logapps/core/src/session/compact-apply.ts
Trigger, retry, cache bookkeepingapps/core/src/agent/loop.ts (maybeCompact, runCompaction, compactAndRetry)
Cache-stable tool-result pruningapps/core/src/agent/prune-state.ts

Related: Sessions, store & rollout for the log this trims, Lifecycle hooks for PreCompact/PostCompact, and Memory for the layer that survives the session entirely.