Skip to Content
ExtrasBatching background completions inline

Batching background completions inline

GitHub measured 2.3% saved by delivering background results inside already-happening requests instead of dedicated polling turns. FreeCode gets this for free today — this page records why, and the rule that keeps it that way.

The problem, in plain words

Imagine the agent kicks off two long-running commands in the background, then later asks “is task 1 done?” (one model call), gets the result (another), asks “is task 2 done?” (a third), gets that result (a fourth). Four model calls to collect two results — and remember, every call re-sends the entire conversation. That polling pattern is pure waste: the results should ride along inside model calls that were going to happen anyway.

GitHub fixed this in Copilot by attaching finished background completions to the next already-scheduled request. Their saving: 2.3% of total cost.

Why FreeCode doesn’t have this problem

When this page was written, FreeCode had no background bash and no async subagents. When the model emits several independent tool calls in one response, the batching layer (apps/core/src/tools/batching.ts + tools/orchestrator.ts) runs them concurrently and returns all their results in the single follow-up model call. There was no “check status later” turn to eliminate.

Both now exist, so the rule below is live rather than theoretical:

  • Background sub-agents (agent(run_in_background: true), 2026-09-27) follow it. A finished one is pushed, never polled. If the parent is mid-turn, the <task-notification> rides the steering queue into the next model call the turn was already going to make, which is this rule exactly (as a user message rather than a tool result). If the parent is idle, core starts one turn, and a burst of completions within 250 ms shares it. That turn is the only call that happens, not an extra one on top of a schedule, so it is not the polling anti-pattern. See Sub-agents.
  • Background shells (bash(run_in_background: true), 2026-09-27) follow it the same way: an exit is pushed as a <task-notification>, and none is sent when the model already read the finished shell with bashoutput — which is the one case that would otherwise buy a redundant turn. See long-running commands.

How I implemented it (as a recorded invariant)

The risk isn’t today; it’s the roadmap. FreeCode’s autonomous/ subsystem (Phase 0: types, budgets, storage) will eventually grow detached execution, and async subagents are a natural future feature. The day either lands, the polling anti-pattern becomes possible.

So the spec (2026-09-04-harness-cost-efficiency.md, D5) records this as a hard constraint on all future work:

A background completion is delivered as an ordinary tool-result in the next already-happening model call — never via a dedicated retrieval turn.

Four calls to process two results is the named anti-pattern. Any future background-execution design gets reviewed against this line before it builds.

Performance improvements

MetricValue
Dedicated polling turns for sub-agents and shells0 — results are pushed
Cost of enforcing the invariant0 tokens, 0 code — it’s a design rule
Measured savingnot measured yet; GitHub’s 2.3% is the reference point

The cheapest technique of the four: it costs nothing because it was adopted before the feature that would violate it.