Skip to Content
InternalsMemory

Memory

A language model does not remember you. Every request is answered from whatever text is in front of it; close the session and the model is exactly as ignorant as it was the first time. Memory is the small system that makes something survive that reset: a folder of markdown notes that FreeCode writes during a session and reads back into the prompt in later ones.

That is the entire trick, and it is worth saying plainly before the detail starts, because “AI memory” usually sounds like a mysterious model capability:

Nothing is stored inside the model. Memory is files on your disk, plus rules about when to write one and which ones to paste back into the prompt.

Everything below is those two rules — the write path and the read path — and the reasons each of them is shaped the way it is.

Which memory is this?

“Memory” in agent systems is four different things wearing one word. The standard split is by what kind of thing is stored, and it is worth learning once because every memory paper, library, and blog post assumes it:

LayerStoresLifespanExample
Workingcurrent task statethe session“editing loop.ts right now”
Episodicspecific past eventsuntil pruned or consolidated“on Aug 12 the deploy failed with OOM”
Semanticgeneral facts and rulesdurable, revisable“this host OOMs above two instances”
Proceduralreusable know-howdurable, revisable“steps to diagnose a crash loop”

All four exist in FreeCode. Only one of them is this page:

LayerIn FreeCodeWhereLearned automatically?
Workingin-session history, summary, compactionCompaction (compaction/service.ts:36)n/a — session-scoped
Episodicdated episode memories, one sentence each — plus rollout event sourcing underneath, which logs every turn verbatimthis page, and Sessions, store & rollout for the raw logyes — written by consolidation
Semanticthe markdown memory store + its graphthis page, plus Knowledge graphyes — tool + extraction + consolidation
Proceduralskills and custom commandsSkills, Custom commandsno — hand-authored

A name collision worth knowing before you read the code. The class called MemoryService (compaction/service.ts:36) is working memory — the message history the loop compacts. The system on this page is MemoryStore + MemoryGraphService. Two different layers, similar names, no relationship.

The healthy flow in that model is working → episodic → semantic + procedural: something happens, it gets logged as an episode, and only after episodes repeat does consolidation distill a durable fact or a reusable method.

FreeCode now does both arrows. A finished session is mined for durable facts (extraction), and roughly once a day a separate pass reads what has accumulated, merges near-duplicates, records what happened as a dated episode, and promotes anything the period made clear (consolidation). What is still missing is narrower than it used to be: consolidation reads recent memory, not the rollout archive, so hundreds of historical sessions have never been distilled and nothing goes back for them. The procedural arrow is also still manual — nothing turns a successful tool sequence into a skill. Both are revisited under Known gaps.

The shape of the whole thing

┌──────────────────────────────┐ WRITE │ markdown files on disk │ READ │ (the source of truth) │ memory tool ──────────▶│ │──────▶ MemoryGraphService (the model decides) │ memory/<type>/<name>.md │ .prepareMemories() │ memory/MEMORY.md │ │ extract.ts ───────────▶│ │ ▼ (turn-end mining) │ │ retrieval judge │ │ │ consolidate.ts ───────▶│ │ ▼ (daily merge/prune) └──────────────┬───────────────┘ "# Relevant memories" │ system block (≤2 KB) onMemoryChange │ (fire-and-forget) │ ▼ │ ┌──────────────────────────────┐ │ │ .graph/ — derived sidecar │◀── citations ──┘ │ embeddings.bin · graph.json │ (what got used) │ usage.json · consolidation │ │ DELETE IT ANY TIME │ └──────────────────────────────┘

The invariant that makes the rest safe: the markdown files are the only source of truth. MemoryStore.list() reads the filesystem directly (mem-store.ts:146), so anything that writes a valid file — including you, in your editor — is picked up on the next sync. The .graph/ sidecar that makes retrieval smart is derived: corruption, a schema change, or a swapped embedding model are all fixed by deleting it and letting it rebuild. That half is documented separately in Knowledge graph; this page is the files and the paths in and out of them.

What a memory is

One fact per file. Small, human-readable, diffable:

--- name: prefers-tables description: User wants comparisons rendered as tables type: feedback tags: style, formatting supersedes: old-formatting-note --- Render comparisons as markdown tables. **Why:** easier to scan than prose. **How to apply:** any time you compare 2+ options. See [[response-style]].
FieldRequiredDoes
nameyeskebab-case identity. Saving the same name updates that memory
descriptionyesone line; this is what relevance is judged against later
typeyesone of five, below
tagsnotopic tags → HasTag edges in the graph
supersedesnonames this memory replaces → Supersedes edges
happened_atnoISO date; episodes only — what the entry is about, not when it was written
bodyyesthe fact itself; [[wikilinks]] become RelatesTo edges for free

tags, supersedes, and happened_at were all added after the first version and are parsed back-compatibly — absent means none (mem-types.ts:100), so every file written before they existed still parses unchanged. Frontmatter parsing is a deliberate ~50 lines of string handling rather than a YAML dependency, and the list parser accepts both a, b and [a, b] because different providers emit different shapes.

The five types

The type is not decoration; it is the directory the file lives in, and it is how the model is taught what deserves saving at all. The first four are semantic in the sense above — timeless claims, not dated events — though feedback leans procedural, since a How to apply: line is a small piece of know-how rather than a fact. The fifth is the episodic layer.

TypeHoldsExample
userwho they are, how they like to work“Prefers TypeScript strict mode everywhere”
feedbackguidance they gave you, with Why: and How to apply:“Don’t add comments unless asked”
projectdecisions, constraints, deadlines not derivable from the code“We chose SQLite over Postgres because the CLI must run offline”
referencepointers to external systems“Dashboards live at grafana.internal/freecode”
episodewhat happened, when — one sentence, with a happened_at date“2026-08-23 — settled on a 180s SSE stall timeout at the fetch layer”

Episodes are machine-written. The memory tool refuses type: "episode" on save and points at the four durable types; only consolidation creates them. The model can read its own history and delete from it, but not author it — a model narrating its sessions into memory is the noise failure mode the whole consolidation design exists to prevent.

Episodes also decay, and the other four do not. An episode’s retrieval score is multiplied by max(floor, 0.5 ^ (ageDays / 30)), where use raises the floor: min(0.9, 0.25 + 0.15·ln(uses + 1)). So an old episode nobody needed sinks, while one the user keeps hitting stays reachable. “User prefers tables” is never decayed — demoting a durable fact for being old is how a system forgets a standing instruction.

The negative rule matters more than the positive one, and it is stated in the system prompt (mem-prompt.ts:116): do not save what the repo already records. Code structure, how a bug was fixed, git history, CLAUDE.md content — all of that is re-derivable by reading the project, so storing it buys nothing and costs prompt space forever. Nor should anything be saved that stops mattering when the current task ends.

Where it lives

~/.freecode/projects/<project-name>/memory/ ├── user/*.md ├── feedback/*.md ├── project/*.md ├── reference/*.md ├── episode/*.md written by consolidation, not by the model ├── MEMORY.md generated index — never hand-edit ├── .git/ baseline for consolidation's "what changed" diff └── .graph/ derived, rebuildable (gitignored)

The project key is the full reversible path (formatSessionDirName, the same scheme session dirs use), so two projects both called api in different parent directories keep separate stores. A store written under the old basename key is renamed to the new key on first access (migrateLegacyDir, mem-store.ts).

MEMORY.md is a regenerated index of names, descriptions, and links, rewritten on every save and delete by updateIndex() (mem-store.ts:182), and truncated at 200 lines with a visible warning rather than growing without bound.

The write path

Three things can create a memory, and they exist for different reasons: one is the model choosing, one is a safety net for when it doesn’t, and one is housekeeping over what the other two left behind.

1. The memory tool — the model’s deliberate choice

memory(action: "save", type, name, description, content, tags?, supersedes?) memory(action: "delete", type, name) memory(action: "list", type?)

The model could, in principle, just use the write tool — memories are only markdown files. It doesn’t, and the reason is the difference between a file existing and a file being part of the system:

via MemoryStore.save()via a raw write
frontmatter serialized from a typed entrymodel hand-writes YAML; a typo degrades it silently
updateIndex() refreshes MEMORY.mdthe index goes stale
emitMemoryChange() → incremental embed + new edgesno event; only self-heals on a later full sync
containsSecret() refuses credentialsa token lands on disk in plaintext

Behaviours worth knowing:

  • Idempotent on (type, name). Saving an existing name updates it, and the result tells the model it clobbered something, including the previous description (tools/memory.ts:219). No accidental duplicates named prefers-tables-2.
  • list returns names and descriptions only, never bodies (tools/memory.ts:153). It is a dedup check before saving, not a recall path — recall is the graph’s job, and dumping every body would defeat the point of selective retrieval.
  • Comma-separated tags/supersedes are coerced, not rejected (tools/memory.ts:32), because providers with strict schema decoding send lists that way.
  • Blocked in read-only modes. memory declares file.write permissions and is deliberately not in READONLY_TOOLS, so plan, review, and explore modes block it like any other writing tool. Fail-closed was the deliberate choice; the cost is that a preference stated while planning isn’t captured.

2. Turn-end extraction — the safety net

Users state preferences in passing and never think “save that”. So when a run finishes naturally — the model answered with no further tool calls — the finished transcript is mined for anything durable (loop.ts:936 → extract.ts:117).

completion → shouldExtract() gates → one provider call → JSON array of proposals → validate → cap at 3 → MemoryStore.save()

Four properties keep this from becoming a liability:

  1. Fire-and-forget. kickMemoryExtraction is never awaited (loop.ts:1212). Your result returns first; extraction finishes behind it.
  2. Never throws. Malformed JSON, a fenced code block, a dead provider, an unknown type — every failure degrades to “saved nothing” (extract.ts:170). A memory failure must never surface as a task failure.
  3. Capped at three saves per run (MAX_SAVES_PER_RUN, extract.ts:17). There is no consolidation pass yet, so an unbounded extractor would fill the store faster than anything cleans it.
  4. Sub-agents never extract. executeSubagent passes memoryExtraction: false (subagent.ts:70). Their transcript is delegated machine work, and without this every verifier and explorer would fire its own extraction, turning one user turn into several billed calls.

The transcript is text parts only, clipped to 12,000 characters — tool arguments and results are the how; what’s wanted is what was said and concluded. The system prompt tells the extractor that returning [] is the common, correct answer (extract.ts:36), because a weak memory is worse than none.

Why a one-shot provider call rather than a sub-agent? SubagentType is a closed union with no per-agent tool allowlist, so “allow only the memory tool” isn’t expressible. Parsing proposals ourselves is also what makes the cap deterministic rather than a request the model may ignore.

The gates: when extraction is allowed to run

Extraction is a fresh, full-price provider call. Running one after every turn would roughly double the calls in a chatty session for very little gain, so shouldExtract() (extract-policy.ts:137) checks six things, cheapest first:

#GateBehaviour
1FREECODE_DISABLE_MEMORY_EXTRACTIONenv kill switch
2memory.autoExtract: falsesettings kill switch
3The model already saved this runskip and reset the counter — paying a second model to second-guess it is the least valuable call available
4Too short (< 200 chars or < 2 turns)“fix the typo” never held a memory
5Topic changed (similarity < 0.12)extract now, before the old thread is buried
6Interval (extractEveryNRuns, default 8)otherwise throttle

Throttling loses nothing. buildTranscript() (loop.ts:1176) rebuilds from the session’s whole history, not just the current run, so a skipped run is still covered by the next extraction rather than dropped.

A session that ends before the interval still extracts. Ending a session — switching away, archiving, or quitting — runs one final pass with force, which bypasses gate 6 and nothing else: both kill switches, “the model already saved”, and the too-short gate all still apply, because a kill switch a code path can bypass is not a kill switch. Deleting a session deliberately does not flush; you discarded it, so mining it is the one case where a write is clearly unwanted.

Topic detection is free. lexicalSimilarity() — Jaccard overlap on tokens — and the 0.12 threshold are already computed every turn for retrieval, so the policy reuses the same function and constant and the two can never disagree. Measured over a simulated 20-run session, the gates turn 20 provider calls into 3.

// .freecode/settings.json (project) beats ~/.freecode/settings.json (user) { "memory": { "autoExtract": true, "extractEveryNRuns": 8, "retrievalJudge": true, "autoConsolidate": true, "consolidateMinHours": 24, "consolidateMinSessions": 5 } }

An unparseable settings file falls through to defaults rather than disabling memory (extract-policy.ts:92) — failing closed there would silently kill the feature and nobody would know why.

3. Consolidation — the pass that shrinks the store

At most once per project per day, and only after five sessions have finished since the last run, one cheap-model call is handed the memory index, a git diff of everything written since last time, and the ~20 memories most likely to need attention. It returns merges, an optional episode, and promotions.

The memory directory is a git repository (.graph/ is gitignored, being derived state). Each successful run commits and moves a baseline tag, so “what changed since we last looked” costs one git diff and comes with content and deletions included. A failed run does not commit, so the next diff spans both windows — nothing written during a failed window is missed. A bad merge is one git revert in a directory whose entire history is memory edits.

There is no delete verb. A memory can only be removed as the supersedes list of a merge, and only if the model was actually shown that name. An unconstrained delete on a cheap model running unattended against your memory is the one failure a retry cannot undo. This also makes consolidation the first writer in the system’s history to emit supersedes: — the Supersedes edges the graph has always supported finally have something to describe.

Off with memory.autoConsolidate: false or FREECODE_DISABLE_MEMORY_CONSOLIDATION=1. It is mutually exclusive with extraction, so there is at most one memory-related provider call per completion, ever.

Citations — how the store learns what earned its place

The loop knows exactly which memories it surfaced, so the only missing half was whether any of them helped. The injected block now ends with an instruction to mark what was used:

<memory-used>project/sse-stall-timeout</memory-used>

The loop strips the tag before anything renders or persists, and credits only ids that were actually injected — a model can name a memory it was never shown. Counters land in .graph/usage.json: useCount, lastUsedAt, and injectedCount. That third field is the point of the design, because useCount alone cannot distinguish never useful from never shown; the ratio is a per-memory precision estimate. freecode memory graph usage prints it.

Counters live in the sidecar rather than the memory files because writing one into a file would bump its updatedAt, dirty the git baseline every turn, and change the content hash that gates re-embedding. Delete .graph/ and retention falls back to age.

Secrets are refused at write time

containsSecret() (graph/secret-filter.ts) matches private-key headers, sk-ant-*, sk-*, AKIA*/ASIA*, ghp_*, github_pat_*, xox[baprs]-*, AIza*, glpat-*, and key=value assignments with secret-looking names. It is enforced in both writers — the tool (tools/memory.ts:190) and the extractor (extract.ts:137) — and again before embedding.

Checking only at embed time was a real hole: a secret-bearing memory was never vectorized but was still written to disk in plaintext and still reachable through the keyword fallback. When the tool refuses, the error tells the model what to do instead — record where the credential lives, not the credential.

The read path

The write path fills a folder. The read path answers the harder question: with 500 memories on disk, which ones belong in this prompt?

Two blocks, two places, one reason

BlockWhere it goesCached?Changes when
buildMemoryGuidanceBlock() — how to use memorystatic system prefix (compiler.ts:164)yesnever
renderRetrievedMemories() — the actual memoriesper-session block, rebuilt each turn (loop.ts:1285)noevery turn

The guidance block takes no arguments and returns constant text on purpose (mem-prompt.ts:104). It sits in the prompt-cache prefix, so if it depended on the store, every single save would rewrite the prefix and bust the whole session’s cache. There is a test that locks this: the block must be byte-identical regardless of what the store contains.

For the same reason, MEMORY.md is deliberately never injected. Loading the whole index every turn would make the cached prefix depend on the store, and would grow linearly with how much you’ve told the agent. Semantic retrieval already surfaces what’s relevant, and memory(action: "list") answers “what exists?” on demand. Net effect: a user with 0 memories and a user with 500 pay the same ~150 cached tokens.

Choosing what to inject

Every turn, the loop asks the graph service for the memories relevant to the last user message (loop.ts:1281):

seed: embed(query) → cosine top-10, threshold 0.4 ┐ BM25 over the same entries → top-10 ┘ fused by rank (RRF) cascade: walk edges from the seeds, 2 hops, score × edge weight × 0.7 per hop judge: a cheap model keeps only what is actually relevant (or nothing) render: ≤ 2 KB → "# Relevant memories"

Vector and lexical retrieval are peers, not a primary and a fallback. Their ranks are combined with reciprocal rank fusion — ranks, because a cosine similarity and a BM25 score have no common scale, and any weighted sum of them would be a hidden calibration constant nobody ever revisits.

The scoring, edge weights, and vector storage are the knowledge graph’s business. Four properties belong here, because they are what keeps retrieval out of your way:

  • One turn behind (graph/index.ts:455). A warm turn returns the previous set instantly and refreshes in the background; the loop never blocks on retrieval. A cold turn — the session’s first message, or right after a topic change cleared the set — waits COLD_BUDGET_MS = 60, then gives up and lets the background fill land on the next turn.
  • Per session, not per project. Vectors and the graph are shared across a project, but the surfaced set is keyed by sessionId and LRU-bounded at 64, so two open sessions never clobber each other’s context.
  • Topic changes clear the stash. If the new message shares almost nothing with the last one, the old set is dropped rather than injected — stale memories about a finished subject are worse than no memories.
  • The block has a hard 2 KB ceiling. Entries render in relevance order; once the budget is spent the rest degrade to a one-line name — description, and past that they drop. A count cap could not fix the failure that matters — one memory with a 4 KB body, injected uncached on every turn.

Why there is a model call in the read path

Retrieval scores cannot tell relevant from merely nearby, and this was measured rather than assumed. On the benchmark corpus, the top cosine for on-topic queries spans 0.674–0.932 and for irrelevant ones 0.588–0.719 — they overlap, so no threshold separates them. A within-query z-score overlaps too: “write a haiku about the sea” outscores 13 of 22 genuinely on-topic queries. The cause is structural, not a bad constant — bi-encoder similarity between short texts has a high, corpus-dependent floor.

So a cheap model decides, and it is affordable because of where it runs: on the background prefetch the loop already never waits for, behind a cadence carry that fires it on a topic change rather than on every message. It fails closed — on a dead provider or an unparseable verdict the candidates are dropped rather than surfaced, because a missed memory costs one turn while an injected irrelevant one biases the answer invisibly. Every outcome is a named variant, so a silent degradation is countable rather than invisible.

Off with memory.retrievalJudge: false or FREECODE_DISABLE_MEMORY_JUDGE=1.

When the embedder isn’t there

Embeddings are optional (a native ONNX dependency that some builds don’t ship). If it can’t load, available() flips to false permanently and retrieval falls back to lexical retrieval alone — BM25 over name, description, and body, still expanded by the tag and wikilink graph walk.

It used to be much dumber: the original scorer awarded points for any substring overlap, with no IDF and no length normalization, so a term in every memory counted as much as a rare one and a long memory beat a precise one. BM25 is ~40 lines, has no dependencies, and cannot fail. Memory never throws into the agent loop; the worst case is that retrieval gets less precise.

What the user sees

Nothing is ever recorded about you silently. The two write paths surface differently on purpose:

The model saves itExtraction saves it
A normal tool call: ● Memory — Saved memory feedback/prefers-tables.A system notice: “Remembered 2 things for next time: … — /graph to view, memory.autoExtract: false to stop.”
Rendered inline in the turnNames each memory and points at the off switch (apps/tui/src/index.ts:872)

The notice travels on the bus, not the turn stream: extraction runs after the turn’s done, so the stream is already closed. The bus speaker wire is subscribed at startup and writes regardless, so out-of-band notices still arrive (memory.saved on the bus → memory_saved as a StreamEvent). Injection is announced too — memory_injected fires once per user message that gets a hit (loop.ts:1296), not on every inner-loop turn.

Surfaces

SurfaceWhat
Toolmemory — save / delete / list
IPCmemory.list · memory.get · memory.save · memory.delete · memory.query
CLIfreecode memory graph stats · graph usage · graph rebuild · ui-install
Benchpnpm bench:recall — recall@k / MRR / nDCG / abstention over a fixed corpus
TUI/graph — the optional explorer addon

Known gaps

Knowing where a system is thin is more useful than a tour of its strengths. Read in the layer vocabulary from the top of the page, these are specific rather than vague. Consolidation used to head this list; now that it exists, what remains clusters around revision (contradiction, valid time) and evidence — most of what follows is “we built it and have not yet measured it in the wild”.

  1. Consolidation is unproven in the field. It merges near-duplicates, writes episodes, and emits the first supersedes: this system has ever produced — but at most once per project per day, so real stores accumulate evidence slowly. Its value claim (better recall at constant token cost) is testable with pnpm bench:recall and has not been tested against a real store yet.
  2. Revision is one-directional and untimed. Entries carry createdAt and updatedAt — transaction time, when we learned something — and supersedes: is the entire contradiction story. There is no valid time, so “the deploy host ran Apache until March” is not expressible; the old memory can only be replaced, not dated. Two contradictory memories that never reference each other both stay retrievable: consolidation emits Supersedes (the writer-knows case), but nothing detects that two independently-written memories disagree.
  3. Procedural memory is authored, not learned. Skills and custom commands are real procedural memory, but a human writes them. Nothing turns “I did A, B, C and it worked” into a reusable procedure, which is a harder distillation problem than fact extraction because order and preconditions matter. codex’s consolidation writes skills/; ours does not.
  4. The rollout archive is still never mined. The end-of-session flush closes the “short session loses a preference” case, but hundreds of historical session directories under ~/.freecode/rollout/sessions/ have never been read and nothing goes back for them. codex’s answer is a bounded, leased, parallel backfill at startup.
  5. Retention is enforced only by consolidation. If it never fires, episodes accumulate — demoted by decay, but never deleted.
  6. Citation is self-reported. Usage counts come from the model saying which memories it used. A model may credit one it ignored or use one silently, so useCount is a biased but directional signal. It ranks and retains; nothing deletes on it alone.
  7. Tuning values are guesses. Cap 3, interval 8, 200-character minimum, minHours 24, minSessions 5, MAX_EPISODES 50. The retrieval-side constants are no longer guesses — RRF_K, the episode half-life, and the BM25 parameters can be swept by the benchmark — but the consolidation-side ones remain unmeasured.
  8. The benchmark corpus is self-written. It catches regressions well and proves little about absolute quality, because the people who wrote the retriever wrote the queries. LongMemEval-S is the intended external corpus and is not wired up yet.

Debugging

SymptomCheck
No memories retrievedfreecode memory graph stats — nodes: 0 means the store is empty, not that retrieval is broken
embedder: falsethe native embedder failed to load → keyword fallback. Expected in some binary builds
Retrieval feels stalefreecode memory graph rebuild, or delete .graph/
Nothing is ever saved automaticallycheck the gates in order: env, memory.autoExtract, run count, topic change
Extraction seems to never fireFREECODE_DEBUG=1; the line [MemoryExtract] skipped: <reason> names the exact gate
Consolidation never runsFREECODE_DEBUG=1 → [MemoryConsolidate] skipped: <reason>. It needs 5 finished sessions and 24h; a failed run also backs off for 10 minutes
Memories are surfaced but usage shows 0/nthe model is not emitting <memory-used>. Retention still works, on age alone
A scripted freecode run never consolidatesexpected — the process exits before fire-and-forget background work lands

Where to look

You wantFile
The entry shape and frontmatterapps/core/src/memory/mem-types.ts
Files, the MEMORY.md index, change eventsapps/core/src/memory/mem-store.ts
Lexical retrieval (BM25)apps/core/src/memory/bm25.ts, mem-query.ts
Rank fusionapps/core/src/memory/graph/fusion.ts
The retrieval judgeapps/core/src/memory/judge.ts
Both prompt blocksapps/core/src/memory/mem-prompt.ts
Turn-end miningapps/core/src/memory/extract.ts
The gates and settingsapps/core/src/memory/extract-policy.ts
Consolidationapps/core/src/memory/consolidate.ts (the call), consolidate-run.ts (the schedule)
The git baselineapps/core/src/memory/git-baseline.ts
Usage counters and citationsapps/core/src/memory/usage-store.ts, citations.ts
The recall benchmarkapps/core/src/memory/bench/ (start with its README.md)
The toolapps/core/src/tools/memory.ts
Call sites in the loopapps/core/src/agent/loop.ts (kickMemoryExtraction, executeTurn)

Related: Knowledge graph for how relevance is actually computed, Context engine for what else goes into the prompt, and Compaction for the other kind of memory — surviving a full context window inside one session.