Context compaction
Compaction replaces old messages with a structured checkpoint generated by the current provider and keeps recent turns intact. The replaced messages stay in the session database marked as compacted (for auditing), but they are no longer sent to the model.
Checkpoint sections
The summarizer is asked for JSON validated with zod. The checkpoint has these sections:
| Section | Field | Content |
|---|---|---|
| Goal | goal | What the session is trying to achieve |
| User instructions/constraints | instructions | Instructions and constraints given by the user |
| Discoveries | discoveries | Facts learned along the way |
| Accomplished | accomplished | Work already done |
| Current state | currentState | Where things stand now |
| Next steps | nextSteps | What remains |
| Relevant files | relevantFiles | Files that matter |
If the summarizer does not return valid JSON, its text is used as-is (a text-only checkpoint).
Guarantees
- A tool call is never separated from its results, and tool call IDs are never altered.
keepTurnsrecent turns are kept; if there are fewer, at least the last one is kept. Inside a single long turn, the cut happens at the last safe boundary between calls.- A session with uncertain tool results is not compacted until it is recovered.
- Persistence is transactional.
Manual and automatic
- Manual:
/compact [focus]in the TUI, with optional focus instructions. - Automatic: before a model call, when the used context reaches
thresholdof a known context window (provider.contextWindow, the active/connectprofile'svalues.contextWindow, orGET /models). When the window is unknown — or a declared window is absurdly large (beyond2_000_000tokens, sowindow × thresholdcannot hide real pressure) — Alisio falls back tolimits.maxContextChars: est. tokens (≈ characters / 4) reachingmaxContextChars / 4also compacts. Independently of that, the character budget is also a trigger for every model: once the conversation (instructions + transcript) reachesthreshold(85 %) oflimits.maxContextChars, Alisio compacts, even when a large window (for example 1M tokens) says there is plenty of room. The window rule and the character rule are alternatives: either one fires compaction. Both window and fallback triggers measure the full request (instructions, transcript and the fixed tool catalog), because they protect what the model really sees.maxContextCharsadditionally stays as the post-compaction hard limit — but there it counts only reducible content (instructions + transcript), never the tool catalog (see below). If compressing cannot get the conversation under it, the retained tail is reduced (see below), and only an irreducible session fails with an actionable error instead of sending an oversized request.auto: falsedisables automatic compaction entirely. - Last chance: if the conversation is still over
maxContextCharsafter the clipping below, Alisio runs one automatic compaction (reason"budget") and clips again before failing. It is attempted at most once per turn.
The fallback default is 800000 characters (≈ 200000 tokens): an assumption for unknown windows, the same ~200k-token budget OpenCode assumes for custom providers, so local servers that do not report a window (for example llama.cpp) get about a 200k-token budget instead of a 40k one. A known window (up to the 2M-token trust boundary) always overrides it, and /settings → Context char budget can lower the fallback at any time.
Configuration
{
"compaction": { "auto": true, "threshold": 0.85, "keepTurns": 2, "maxOutputTokens": 16000 }
}| Field | Default | Description |
|---|---|---|
auto | true | Automatic compaction |
threshold | 0.85 | Fraction of a known context window (0.1–0.99) |
keepTurns | 2 | Recent turns kept verbatim (0–20) |
maxOutputTokens | 16000 | Output token budget for the summarizer call; independent of limits.maxOutputTokens |
The context window comes from provider.contextWindow, the active /connect profile's values.contextWindow, or GET /models. Compaction token usage is not counted against limits.maxTokens, and before/after estimates are approximate (about 4 characters per token).
Truncated summaries
The summarizer has its own output budget (compaction.maxOutputTokens, default 16000 — larger than the agent loop's limits.maxOutputTokens on purpose, since a summary must fit the whole transcript). If a summary is cut by this budget:
- A usable partial summary (structured checkpoint JSON or plain text) is kept: compaction completes and
compaction_completedincludes"partial": true. The TUI notice marks the checkpoint as partial and suggests raisingcompaction.maxOutputTokens. - If the cut produced nothing usable, compaction fails with an actionable message pointing at
compaction.maxOutputTokens.
A partial checkpoint preserves everything the model produced before the cut, but may omit later context; treat it as a degraded fallback, not a full summary.
Reducing an oversized kept tail
Compaction keeps keepTurns recent turns verbatim. If those turns hold huge tool outputs (for example grep over a large repository), even a perfect checkpoint cannot bring the session's reducible content under limits.maxContextChars, and the session used to die on every prompt. Instead, after a compaction the runner now checks the hard limit and, when still over, reduces the retained messages in place before sending anything:
- The reduction targets a total character budget for the transcript:
target = max(4 000, limits.maxContextChars − instructions), so it also handles sessions with many MEDIUM tool results (for example MCP outputs of a few thousand characters each) that individually stay under the per-message caps but together exceed the limit. - Content is clipped iteratively, largest first, in descending cap rounds: tool results at 8 000 → 4 096 → 2 048 → 1 024 → 512 characters, user/assistant texts at 16 000 → 8 192 → 4 096 → 2 048 → 1 024. Each round cuts the most expensive messages first (ties by position), so the smallest number of changes reaches the target; the walk stops as soon as the transcript fits, or at the minimum caps.
- Every cut carries the explicit
… [truncated by context budget]marker. Only message content changes: roles, call IDs, order and message boundaries stay identical, so the transcript remains valid and replayable and a tool call is never separated from its results. The reduced transcript is persisted (originals stay in the database marked as compacted, like compaction itself). - The
context_reducedevent reports how many messages were cut. - Only a pathological session —
instructionsalone (plus the 4 000-character floor) still exceeding the limit, so even the reduction floor cannot fit — fails with an actionable error naming the approximate conversation size and suggesting/compact, trimming large tool outputs, starting a new session, or disabling unneeded MCP servers with/plugins. The message says what compaction did:Context budget exceeded (approximately N characters of conversation; limit L). Automatic compaction ran but could not reduce it enough.(orwas skipped because there is no safe boundary to summarize,failed (<error>), oris disabled). A single huge turn with no safe boundary can still fail, now with that accurate message. The reduction is still persisted, so the session keeps working for later prompts.
Checkpoint summaries (summary: true) are bounded by design and are never cut by this step.
The tool catalog is not counted
The hard limit above deliberately measures reducible content only — instructions plus the transcript. The serialized tool catalog (toolsText: every registered tool's name, schema and description) is a fixed deployment reality: it is part of every request regardless of history, so a server exposing dozens of tools (for example a large MCP catalog) can be worth 100k+ characters on its own, and a small transcript can no longer fit under the limit once the catalog is counted. Before this distinction, such sessions hit the fatal "Context budget exceeded" error even with an almost-empty transcript — the reported failure for subagents and explores in repos with several MCP servers.
- Catalog size is a
/pluginsdecision, not session growth: disable unneeded MCP servers there (/mcpsshows each server's tool count), instead of the runner corrupting a healthy transcript to "fit" a catalog it cannot shrink. - Auto-compaction still measures the full request (window thresholds protect everything the model sees, tools included); only the post-compaction hard cap and the transcript reduction exclude the catalog, so the fatal error never fires because of
toolsTextalone.
Events
| Event | When |
|---|---|
compaction_started | Compaction begins |
compaction_completed | Done; includes before/after estimates, the size of the summarized span and of the checkpoint, plugin reports, and partial: true when the summary was cut but kept |
compaction_skipped | Not enough history to compact |
context_reduced | After compaction, retained messages still exceeded maxContextChars and were clipped to a total character target (messages = how many) |
compaction_failed | The summarizer call failed (including a cut summary with nothing usable) |
plugin_hook_failed | A plugin hook failed or timed out; compaction continues without it |
Events are emitted as versioned JSONL with --json.
Responses mode
In Responses mode, opaque continuation items (for example encrypted reasoning) of the summarized span are discarded together with those messages. Kept messages retain theirs unchanged.
Plugin hooks
Plugins can extend compaction with compaction.register({ beforeCompact, afterCompact }):
beforeCompactadds instructions and extra JSON fields (outputFields) to the same summarizer call.afterCompactreceives the checkpoint and the plugin's own extracted fields, and can returninjectContext(text added after the checkpoint) and areport(itssummaryis shown in the TUI).
Hooks run with the host timeout pluginHooks.timeoutMs. The built-in memory plugin uses these hooks. See Writing plugins.
Uncertain tool results
If a crash leaves a tool with an uncertain result, Alisio does not repeat it automatically. Inspect its effects, then run:
alisio sessions recover <session> --acknowledgeRecovery records the uncertainty as the result; it does not ensure that an external effect completed, and it does not undo changes.
