Skip to Content
InternalsAgent loop

Agent loop

A language model cannot read a file, run a test, or save an edit. It can only read text and produce text. Everything else — the reading, the running, the saving — is done by the program around it, and that program is the agent loop.

The deal is simple and it repeats:

Here is the project, here is the conversation so far, and here is a list of tools you may ask me to run. Answer, or ask for a tool.

If the model answers with plain text, the loop is done. If it asks for tools, the loop runs them, appends the results to the conversation, and asks again. That is the whole engine (agent/loop.ts). Everything else on this page is detail about how each of those steps is made cheap, safe, and hard to get stuck in.

Vocabulary

Four words get used interchangeably in most write-ups about agents. Here they mean different things, and the difference matters when reading the code.

TermWhat it isWhere it lives
SessionOne conversation, persisted, resumable. Survives restarts.session/, store/
RunOne run(UserInput) call — everything that happens because of one user message.loop.ts:543
Turn / iterationOne provider call plus the tools it asked for. A run has many.loop.ts:1281 (executeTurn)
BatchA group of tool calls from one turn that may execute concurrently.tools/batching.ts

turnCount and iterationCount are both incremented in the same place and are always equal; the loop keeps both because SessionState predates the merge.

A fresh AgentLoop is constructed for every user message (server.ts:199) and reloads history from the session store. So “per-run state” and “per-instance state” are the same thing on the interactive path — which is load-bearing in one place, and a bug in another (see Known gaps).

The shape of one run

Before the first request

run() starts by throwing away everything that belonged to the previous prompt (loop.ts:543). Loop-health counters in particular are per-run: carrying them forward would let one prompt’s history stop the next prompt before it ever reached the provider.

Then, in order:

  1. Watch the project — ensureWatching(projectPath), idempotent per project.
  2. Load permission rules — project and user .freecode/settings.json, watched for changes (permission/settings.ts).
  3. Collect context — name, path, file tree, git HEAD, and an hour-rounded clock, via getFrozenSessionContext (context/session-context.ts:39). Snapshotted on the session’s first turn and reused verbatim afterwards.
  4. SessionStart hook, then session.created on the bus.
  5. Load history from the session store and rebuild Message[].
  6. Append the user message. Attached images become image parts only if modelSupportsImages(provider, model) — otherwise the user gets an explicit notice on the stream rather than watching the model claim it saw nothing.

Why the context is frozen

The file tree looks like the thing you’d want fresh every turn. It is the thing you least want fresh.

The tree is rendered into the conversation’s first message, and prompt caching matches on an exact byte prefix. Every byte of the request depends on position 0. Re-render the tree because a build dropped a file into dist/, or because a 5-minute cache TTL expired, and the entire conversation prefix is invalidated — every message re-billed as a cache write instead of a read.

So the prompt gets a snapshot, taken once per session, and the clock is rounded to the hour for the same reason.

One turn, end to end

1. Build the request

The system prompt is built as blocks, not one string, and the split is the whole point (context/compiler.ts:137):

BlockcacheContentsChanges
Statictruebase prompt, model identity, mode prompt, CLAUDE.md/AGENTS.md, skill list, memory guidancenever, within a session
Sessionfalsecompaction summary, retrieved memories, todo list, transient remindersmost turns

Session blocks sit at the tail of the array, after the cache breakpoint, so rewriting them costs nothing.

The dynamic context — tree, git HEAD, clock — is not a system block at all. It is inlined as the conversation’s first user message with a fixed id and timestamp: 0 (loop.ts:1672). Claude Code splits at the same boundary for the same reason.

Two things reach the model here that are easy to miss:

  • Retrieved memories (memGraph.prepareMemories, loop.ts:1281) are asked for on every turn, not just when the user speaks. Warm turns return the previous set instantly; a cold-start retrieval that lands just after the loop stopped waiting still gets injected on the next inner turn instead of being lost for the rest of the request.
  • Reminders (agent/reminders.ts) are <system-reminder> blocks queued by the gates below, drained into exactly one turn, and never persisted to history.

Finally the UserPromptSubmit hook sees the joined system text and may rewrite it. A rewrite is re-split back into static + session blocks by applySystemPromptHookRewrite — collapsing them into one cached blob would push todos and memories under the cache breakpoint — and is recorded as a deliberate cache invalidation so later analysis reports it as explained rather than as a bug.

2. Prune, then send

Before the messages go out, pruneHistoryToolResults (loop.ts:422) replaces the largest tool results with a short marker until the model-visible total is under TOOL_RESULT_BUDGET_CHARS (200,000 by default).

The subtle part is not what gets replaced but that the decision is recorded (agent/prune-state.ts). Every result is one of three things:

  • must re-apply — already replaced once; the stored marker is re-sent byte for byte.
  • frozen — already sent at full size; off-limits, because it is in the cached prefix and shrinking it now costs more than it saves.
  • fresh — never sent; eligible for replacement this turn, largest first.

The previous implementation used a sliding “keep the last N turns whole” window, which mutated the prefix two turns back on every single turn — saving ~250 tokens and paying a partial cache invalidation for it. Results in the final assistant message are protected (the model hasn’t reasoned over them yet) but still recorded as seen, so they freeze rather than becoming eligible again.

When a read result is replaced, markReadPruned forgets that the file was ever shown — otherwise read-deduplication would answer a re-read with “it’s already above”, pointing at a marker.

3. The provider call

callProviderOnce writes model.request to the rollout log before the call (loop.ts:1720). That asymmetry is the diagnostic: a request with no matching response or error is what a hang looks like in the log. Timeouts themselves live at the fetch layer (providers/fetch-timeout.ts), never around the chunk iterator — bounding silence between normalized chunks kills long tool-argument streams, which arrive as tool-input-delta parts the normalizer drops.

Streaming is preferred when the provider supports it. Each chunk is both accumulated and re-emitted on the bus (text_delta, thinking_delta), so the TUI renders while the loop is still reading.

Recovery (agent/recovery/manager.ts) classifies before it retries:

ClassDetected byResponse
Rate limit (429)status5 attempts, exponential from 5s, honouring retry-after
Transient (5xx, ECONNRESET, fetch timeout)status / code / message3 attempts, exponential from 1s
Quota exhausted402, or 429 + billing wordingno retry — waiting cannot help
Context overflow400/413 + provider-specific wordingcompact and re-send once
AbortAbortErrorstop the whole chain immediately

The AI SDK wraps failures in AI_RetryError, which exposes no status of its own, so errors are unwrapped to the innermost real error before classification — without that, every nested 429 reads as statusless and therefore fatal. Backoff sleeps are abortable, so Ctrl+C is never stuck behind a 5-second delay. When everything is exhausted the chain fails with a message that says what to do (top up, or configure recovery.fallbackProviders) and deliberately drops the SDK error object, which carries the entire conversation on requestBodyValues.

Context overflow is handled one level up, in compactAndRetry (loop.ts:996): compact, send once more, and do not re-wrap the retry. A second overflow in one turn means compaction isn’t converging, and the cap is MAX_OVERFLOW_COMPACTIONS = 3. See Compaction.

4. Read the answer

Tool calls come back natively from the provider. If there are none, the text is scanned for a [TOOL_CALLS]…[/TOOL_CALLS] block (loop.ts:1959) — a fallback for providers and mocks without native tool calling.

An assistant Message is assembled from the text plus one tool part per call and pushed to history. Usage is attached to the first persisted message only and then cleared: one provider response becomes several stored messages, and attaching usage to each would make any per-session sum a multiple of the truth.

If there are no tool calls, the turn returns here — the run is about to finish.

5. Execute the tools

planToolBatches (tools/batching.ts:22) walks the calls in order and groups runs of concurrency-safe tools into one parallel batch. Anything else — and anything whose behavior is unknown — gets a batch of its own, so ordering guarantees around writes, edits, and shell commands hold. Post-processing is always in the original order, so results are deterministic regardless of who finished first.

Every single call goes through the same gauntlet (loop.ts:2045):

Results carry two payloads: displayOutput (full, for the UI) and modelOutput (head+tail truncated by adaptiveTruncate, for the provider). The full text is also stashed in the per-session output store so the output tool can page through it later. Only modelOutput is written into history and persisted — re-sending untruncated stdout every turn is what used to overflow context windows.

Two side effects are recorded per call:

  • Mutation tracking. A tool with behavior.isDestructive === true that returned without an error marks the run as having changed files and adds its path to mutatedFiles — which is what the verification gates key off.
  • Loop health (below).

Images are a special case. A tool result is a string on the wire, so a tool that produced an image signals it via metadata.image, and the loop re-emits those images as a user message after the results — putting them in the tool result would break the tool_use/tool_result pairing that providers require.

Deciding to stop

“No tool calls” means the model wants to stop. It is not, by itself, enough. Three gates run in order, and any one of them can force another turn (loop.ts:831).

GateFires whenCapWhat it injects
Todothe todo list still has pending / in_progress items3 per runa reminder listing the unfinished items
Verifythis run mutated files and the project has a typecheck/build script2 per runthe failing output
Verifierthis run touched ≥ 3 distinct files2 per runthe verifier’s findings, on FAIL

Verify (agent/verify.ts) resolves a command from package.json scripts, in priority order typecheck, type-check, check-types, check, build. Tests are deliberately excluded — slow and flaky — and stay the model’s job via the prompt. No script, no gate; there is no guessing. The run is capped at 120 seconds and 4,000 chars of captured output.

Verifier (agent/subagent.ts:164) spawns an independent read-only subagent in explore mode that must end its final message with VERDICT: PASS | FAIL | PARTIAL. A missing or garbled verdict parses as PARTIAL, never a false PASS. Its prompt spends most of its length preventing the failure mode this kind of check actually has — inventing new objections: missing edge cases, extra validation, and absent tests are explicitly not grounds for FAIL unless they were the request, and on a re-check the bar does not rise between rounds.

Once the gates are satisfied, the loop kicks memory extraction without awaiting it (loop.ts:936) and returns. The user’s answer arrives now; mining the transcript for durable memories finishes behind it, and can never turn into a task failure. See Memory.

Loop health

The loop watches for three stuck patterns — two while tools execute (updateLoopHealth, loop.ts:2440), one at the end of each turn (advanceStagnation, loop.ts:2498) — and checks the verdict at the top of the next turn. The verdict itself comes from one shared evaluator, createLoopHealthEvaluator (effect/loop-health.ts); the loop holds no second copy of the policy.

A. Repeated identical calls. The last 10 tool + JSON.stringify(args) signatures are kept; repeatedTools is how many of them match the current call.

B. Stagnation. Counted per turn: a turn in which no mutating tool succeeded increments stagnantTurns, and one that changed a file resets it to 0. The counter used to advance per tool call, which meant five consecutive reads — what reading a codebase looks like — registered as “no progress”.

C. Oscillation. Counting edits per file flags real work as a loop — a feature file legitimately gets edited twenty times. What actually signals a stuck agent is an edit that undoes an earlier one: X→Y followed by Y→X. Both sides of every edit are hashed and kept in a 30-entry window (agent/oscillation.ts), and only an inverse scores. A no-op edit (from === to) never counts, and a failed edit changed nothing so it cannot be part of a cycle. oscillationScore is a count of the reverts still inside that window (countReverts), not a running total, so it falls again as a recovering run pushes the scored pair out — one edit/revert pair no longer leaves the counter armed for the rest of the session.

Braking is two-tier: the first breach only warns, and a hard stop is reserved for twice the threshold.

HeuristicDefault thresholdWarn atStop at
repeatedIdenticalThreshold336
stagnantTurnsThreshold55never
oscillationScoreThreshold448
totalIterationLimitInfinity—never

A stop ends the run with Loop stopped: <reason>, runs the Stop hook, and preserves everything — history is already persisted turn by turn.

Trajectory redirection

A warn on its own goes to logger.debug and reaches nobody — the loop keeps funding the circle until it is twice as bad, then kills the run. Redirection is what makes the warn tier useful: on a warning, fold the rollout log into a bounded evidence packet, make one small model call asking for up to three materially different next directions, and push them into the next turn as a <system-reminder> — the same mechanism the todo nudge already uses.

It is off by default. Turn it on per project or per user:

// .freecode/settings.json { "redirect": { "enabled": true, "maxPerRun": 2 } }

FREECODE_DISABLE_REDIRECT=1 overrides the setting, matching the FREECODE_DISABLE_MEMORY_* convention.

What can trigger it. Only a warn, and only for repeated_identical_tool, oscillation_detected or no_progress. Never a stop: at that point the run is over and the honest answer is to hand back to the user, not to spend more money re-planning a corpse. max_iterations_reached is a budget rather than a pathology, and wrapUpReminder() already covers it.

CapValueWhy
Per run2Mirrors MAX_VERIFY_ATTEMPTS. A third means the advice is not what is wrong
Per reason1Re-advising on an unchanged reason produces the same advice
Debounce3 turnsThe model needs turns to act before being judged again
SubagentsoffAlready turn-capped and disposable; the parent re-plans

After a redirection fires, the counter that triggered it is reset — along with the window it is derived from, since zeroing the number alone would let the next call re-derive the old value.

The supervisor is advisory and nothing else (agent/redirect/supervisor.ts). One non-streaming call, 400 tokens, 15-second timeout, the run’s own provider. No tools, no permissions, no verifier access; it cannot end a run, alter the agent mode, or extend a budget. Its tokens are added to the run totals and to the daily usage file, so the spend circuit breaker sees them — a supervisor whose cost is invisible to FREECODE_MAX_TURN_TOKENS would reintroduce the hole that breaker was built to close.

It fails closed, always. Provider error, timeout, empty response, nothing parseable, or an evidence packet with nothing in it → no redirection, a recorded redirect.skipped, and the loop continues exactly as it would have. The call is synchronous rather than one-turn-behind on purpose: the turn being delayed is by definition an unproductive one, and advice that arrives a turn late can arrive after the hard-stop tier has already killed the run.

What is recorded. redirect.triggered carries the reason, the count and size of the directions, latency, tokens, and evidenceEventIds — but not the advice text. buildEvidence() is pure, so the ids are enough to reconstruct the exact packet from the log, while model-authored prose (which can quote code) stays out of a log that OTLP exports. The text is already durable in the transcript. freecode trace shows the counts.

Ending badly

Iteration cap. maxIterations defaults to Infinity, matching Claude Code and opencode: interactive sessions are ended by the health checks and the gates, not by a turn count. Subagents pass one explicitly (50 for the agent tool, 20 for executeSubagent, 15 for the verifier). One turn before the cap trips, a wrap-up reminder tells the model to stop calling tools and summarize — so a capped run hands back a real summary rather than being truncated mid-task, and the returned text is the model’s own last response with a note appended.

Spend. FREECODE_MAX_TURN_TOKENS is an opt-in circuit breaker on input + output tokens across the run (loop.ts:812). Loop-health only warns on a stuck pattern; nothing else caps actual money.

Interrupt. interrupt() flips status to stopped and aborts the AbortController threaded into both the provider request and every tool context, so in-flight work stops immediately rather than at the next loop check. A turn that fails while abort.signal.aborted is reported as a clean "Interrupted" completion, not a failure.

Failure. Anything else returns fail(), which emits session.error on the bus and returns success: false with the state marked error.

Every exit path reports accumulated usage. An abnormal stop still spent whatever it spent, and dropping the totals would render the run as free.

Where the numbers come from

Usage accounting has one rule that is easy to get wrong in both directions:

  • inputTokens from the provider mapper is inclusive — cache reads and cache writes are already folded in. Adding cache writes back on top was a real double-count bug against Anthropic.
  • outputTokens is likewise inclusive of reasoningTokens, which is carried separately only so the TUI can show reasoning cost.
  • Context occupancy is the last call’s input, not a sum. Every call re-sends the whole conversation, so summing would multiply it. That single number (lastMeasuredContextTokens) is what drives auto-compaction — it replaced MemoryService’s own estimate, which never saw tool activity and under-reported a 196K context as 14.5K.

Per-turn totals are also written to the daily usage heatmap (~/.freecode/usage.json) and streamed to the frontend as usage_totals. Cache read/write counts are surfaced as cache_status, and an unexplained cache miss — one the invalidation journal has no recorded cause for — is escalated to a warning, because it means something mutated an already-sent message.

Knobs

SettingDefaultEffect
maxIterationsInfinityhard turn cap for this loop instance
repeatedIdenticalThreshold3identical-call warn threshold (stop at 2×)
stagnantTurnsThreshold5no-progress warn threshold, in turns
oscillationScoreThreshold4revert-cycle warn threshold (stop at 2×)
TODO_NUDGE_TURNS / TODO_NUDGE_GAP3 / 5when to nudge the model to keep a plan
TODO_GATE_MAX_FORCES3forced continues for unfinished todos
MAX_VERIFY_ATTEMPTS2typecheck/build gate retries
VERIFIER_MIN_FILES / MAX_VERIFIER_ATTEMPTS3 / 2when the adversarial verifier runs, and how often
MAX_OVERFLOW_COMPACTIONS3compact-and-retry attempts per run
FREECODE_TOOL_RESULT_BUDGET_CHARS200,000history-wide budget for tool-result text
FREECODE_MAX_TURN_TOKENSunsetper-run spend circuit breaker
recovery.fallbackProviders[]providers to try after the primary is exhausted

FREECODE_TOOL_RESULT_BUDGET_CHARS is read at module load, so changing it needs a restart — unlike the compaction variables, which are read per call.

Known gaps

  1. A loop-health warn reaches nobody unless redirection is switched on. Trajectory redirection is what acts on a warning, and it is off by default until the measurement in §9 of its spec says otherwise. With it off, a warn is still just logger.debug (loop.ts:737) and nothing happens until the pattern doubles into a stop. The trigger itself is recorded either way (redirect.skipped, reason disabled, once per run), which is how the flip decision gets its evidence.
  2. The spend circuit breaker is off by default (loop.ts:812, compaction/tokens.ts:105). FREECODE_MAX_TURN_TOKENS is unset unless the user sets it, so nothing caps actual spend. freecode run has --max-turns for a turn cap, but nothing caps tokens by default, and loop-health only warns on the stuck patterns most likely to burn quota (stagnation never stops at all).

Previously listed here and since closed: the Stop hook now fires on every run() exit (complete()/fail()), not only on an abnormal stop(); SessionStart/session.created now fire once per session (gated on empty history) instead of once per message; the compiler’s second, independently-keyed file-tree cache — the thing that made the tree watcher’s fix defeat itself for up to five minutes — was deleted rather than fixed, since it was formatting a string and buying nothing by being cached; and heuristic D (repeated reasoning similarity) is implemented via word-set Jaccard similarity between consecutive turns’ reasoning text, feeding warn/stop the same way A–C do (it is not, however, wired into trajectory redirection’s trigger list — see gap 1).

Where to look

You wantFile
The loop itselfapps/core/src/agent/loop.ts
Heuristic defaults, SessionState, LoopResultapps/core/src/agent/types.ts
Revert detection and the oscillation windowapps/core/src/agent/oscillation.ts
The loop-health policy (the only copy)apps/core/src/effect/loop-health.ts
Redirection: caps, evidence, prompt, supervisorapps/core/src/agent/redirect/
Todo nudge / gate / wrap-up textapps/core/src/agent/reminders.ts
The typecheck/build gateapps/core/src/agent/verify.ts
The adversarial verifier + subagentsapps/core/src/agent/subagent.ts
Cache-stable tool-result pruningapps/core/src/agent/prune-state.ts
Retry, backoff, provider fallback, error classificationapps/core/src/agent/recovery/manager.ts
System-block compilationapps/core/src/context/compiler.ts
The frozen project snapshotapps/core/src/context/session-context.ts
Parallel batchingapps/core/src/tools/batching.ts
Validation, coercion, truncation, tool retryapps/core/src/tools/orchestrator.ts
Turn construction and the follow-up queueapps/core/src/server.ts (runSessionTurn)

Related: Tool system for what the loop executes, Permission engine for the gauntlet each call runs, Lifecycle hooks for every interception point, Compaction for what happens when history stops fitting, and Sessions, store & rollout for where each turn is written down.