We Read 20 Agent Harnesses. 15 Compact Your Context Automatically — 10 Let You Set When.
Last Updated on September 25, 2026 by Editorial Team
Author(s): Decoding AI by Nueravi
Originally published on Towards AI.
We Read 20 Agent Harnesses. 15 Compact Your Context Automatically — 10 Let You Set When.

Key takeaways
- We read the agent context compaction settings in 20 open-source agent harnesses at HEAD on 22 September 2026. 15 compact your conversation automatically. Only 10 let you set when.
- 4 give you a switch and nothing else —
opencodeandcrushexpose on/off with no threshold,continuehard-codes the ratio at 0.8, andcrewAI's summariser takes no threshold argument at all. - 5 have no automatic conversation compaction we could find:
aider,letta,agno,aichatandbrowser-use. One of those five,agno, compresses tool results — not conversation history. They are different things and it would have been easy to miss. - The two richest controls are
zedandopenai/codex.zedtakes a percentage, an absolute token count, or a negative number meaning "tokens remaining".codexexposes three separate knobs including a post-turn threshold. clinewe could not classify. Auto-compaction exists and the interface displays a threshold, but we found no setting that writes it. It is counted for neither side.- 20 repositories cloned at HEAD, 0 clone failures, every classification read by hand and then re-checked with a second, wider vocabulary. Five classifications moved on that re-check — all in the same direction.
TL;DR
A 176-setting ablation published on 17 September reported that a single context-management setting moved one model from 6.40% to 58.40% on SWE-Bench Verified — a 52-point swing, against a 43-point spread between the best and worst models in the same study. That result is not ours and we have not reproduced it. But it raises a question anyone can check, and nobody had: if the harness setting matters that much, can you actually reach it? In half of the harnesses we read, you cannot.
Can I control when my agent compacts context?
In 10 of the 20 harnesses we read, yes — there is a named setting that sets the threshold or the budget. In 4 more, compaction happens and you can only turn it off. In 5, there is no automatic conversation compaction to control. And in 1, cline, the interface shows you a threshold we could not find a way to set. The median agent user is running something that decides when to forget for them.
Status as at 22 September 2026, 19:30 UTC. 20 harnesses at HEAD; 15 compact automatically, 10 expose a threshold. Re-check with:
git clone --depth 1 <repo> h && grep -rniE 'auto[_-]?compact|auto[_-]?condense|compressionThreshold|auto[_-]?summariz' h— then read the hits, because the grep alone is wrong.
What did we actually measure?
For each harness: does it automatically reduce the conversation history when context fills, and is the trigger point settable by the user? We counted a control only where a named configuration key, preference, or environment parameter sets the threshold or budget. A value the interface merely displays does not count. A setting that compresses tool output does not count, because that is a different mechanism on a different part of the prompt.
We cloned 20 repositories at --depth 1 on 22 September 2026, with 0 clone failures, screened with a keyword pass, and then opened the configuration surface of every one of them by hand.

Which harnesses let you set the threshold?
Ten, and the controls are not equivalent:
zed—agent.auto_compact.enabledand.threshold, default"90%". It accepts a percentage, an absolute token count, or a negative integer meaning "compact once this many tokens remain". The richest control in the set.openai/codex— three keys:model_auto_compact_token_limit, a..._scopecontrolling whether the limit counts the whole context or only tokens after the carried prefix, andmodel_post_turn_compact_threshold_percent.RooCodeInc/Roo-Code—autoCondenseContextPercent, in the global settings schema.microsoft/vscode-copilot-chat—chat.advanced.summarizeAgentConversationHistoryThreshold.google-gemini/gemini-cli—compressionThreshold, "the fraction of context usage at which to trigger context compression".block/goose—GOOSE_AUTO_COMPACT_THRESHOLD, read as a runtime parameter, and also a named preference key.All-Hands-AI/OpenHands— a dedicated condenser settings screen bound toagent_settings.condenser.plandex-ai/plandex—maxConvoTokens, above which a dedicated summariser model role runs.Significant-Gravitas/AutoGPT—ActionHistoryConfiguration, withmax_tokens,full_message_countandenable_compression.danny-avila/LibreChat—maxContextTokensper agent, which drives the effective budget.
The controls are not even the same kind of thing, and that matters more than the count. Roo-Code, gemini-cli and goose express the trigger as a proportion of the context window. plandex, AutoGPT and LibreChat set an absolute token budget instead. zed and codex accept either. A proportion follows the model when you move to one with a larger window; a fixed token budget does not, and will keep compacting at the old point on a model with five times the room. We did not audit which shape each of the ten defaults to, so treat this as a difference in kind, not a tally.

And the ones that decide for you?
Four, in two different ways. opencode and crush give you a switch: OPENCODE_DISABLE_AUTOCOMPACT and disable_auto_summarize respectively. Both are on/off with no threshold, so your only choice is between the vendor's timing and none at all. continue hard-codes AUTO_COMPACT_BUFFER_RATIO = 0.8 as an exported constant and derives the trigger from your context limit. crewAI calls summarize_messages() with no threshold argument in the signature.
A switch is not a setting. If the timing is what moves the result, “off” is not the other end of the dial — it is a different failure.
If you are using one of those four right now: do you know what its threshold is?
What about the five with nothing?
aider, letta, aichat and browser-use had no automatic conversation compaction we could find under either vocabulary we tried. browser-use has a model.turn.context_overflow event, which is a way of noticing the problem rather than handling it.
agno is the interesting one, and it nearly became a wrong number. It does compress — models/base.py compresses existing tool results before the API call, explicitly to avoid context overflow. That is real context management, and it is not conversation compaction. Counting it would have made our "15 compact automatically" read 16 on a mechanism that does something else to a different part of the prompt.
The one we could not classify
cline ships auto-compaction: there is an auto_compaction event kind, and ContextWindowSummary.tsx renders autoCompactThreshold to the user as a percentage. What we could not find is the setting that writes it. A displayed value implies a source, and we could not locate one in the settings schema.
So it is counted for neither side. The honest position is “not established”, not “no” — and both our 10 and our 4 would look better if we had guessed.
What this does not show
It does not show that a settable threshold is better. The 176-setting ablation is somebody else’s result, on a different question, and we have not reproduced it. It is entirely possible that a well-chosen fixed threshold beats a badly-chosen user one, and the harnesses with no control may simply have picked well.
It does not measure defaults. We recorded that zed ships "90%" and continue ships 0.8, but we did not compare defaults across the set or test any of them, and a default is what almost every user runs.
It is not a sample of agent harnesses. These 20 were chosen for being open-source, locally runnable and widely installed. There is no reliable enumeration of the population, so this is a named list, not a census.
And it says nothing about compaction quality. Whether a harness summarises well — what it keeps, what it throws away — is invisible to this method and is probably the more important question. That needs a harness and a task suite, not a grep and a careful read.
The re-check, and why every flip went the same way
Our first vocabulary was the obvious one: compact, condense, auto-summarize. Running only that, five of these harnesses looked like they had nothing.
The second pass added the words the other half of the ecosystem uses — recursive summarization, context overflow, memory pressure, trim history, prune messages, rolling summary. Five classifications moved, and every single one moved from "no control" toward "has one": plandex has a summariser model role gated on max-convo-tokens; AutoGPT has a whole ActionHistoryConfiguration; LibreChat has a per-agent token budget; crewAI summarises without a threshold; agno compresses, but the wrong thing.
That is the second time in one day that a first-pass filter on this account was wrong in the direction that made the story louder. A narrow filter produces a scarcity headline — “only five harnesses let you do this” — and scarcity headlines travel. The correct response is to assume the first number is the interesting one because it is wrong.
FAQ
What is context compaction, exactly? When a conversation approaches the model’s context limit, the harness replaces older messages with a generated summary so the session can continue. It is lossy by design. The threshold decides how much history is still present when that loss happens, and different thresholds produce measurably different agent behaviour on long tasks.
Why does the threshold matter more than the model? We did not measure that and will not claim it. What the 17 September ablation reported is that within its setup, one context-management setting moved a single model further than the gap between the best and worst models it tested. Treat that as a reason to look at your configuration, not as a settled fact.
Is “off” a safe default? Not obviously. With compaction off, a long session eventually hits the context limit and the harness has to do something worse — drop messages, error, or restart. The four on/off harnesses are not offering you “no compaction”, they are offering you “no automatic compaction”.
How do I find my own threshold? Search your harness’s settings for compact, condense, compress or summariz. If nothing matches, widen to max_context, maxConvoTokens or contextLimit — several of these ship the control as a token budget rather than a percentage, which is why our own first pass missed them.
Does this apply to hosted agents too? The mechanism does; the visibility does not. Every one of these numbers comes from reading source we could clone. A hosted agent’s compaction policy is not observable this way, which is a real limit on what any measurement like this can cover.
One question for you
Open your agent’s settings and search for compact, condense or summariz. If you found a threshold — had you ever changed it, and do you know what it shipped as?
Decoding AI publishes one original measurement a week, with the method and the raw counts attached. Next: what the ten settable thresholds are actually set to out of the box, since a default is what almost everyone runs and we counted the knob without reading the dial.
Related measurements: what twelve agent runtimes do when two skills share a name, what nine MCP servers spend on tool definitions before you type anything, and every measurement we have published, in one place.
Sources & References
- SWE-bench — the benchmark the 17 September ablation reports against; the ablation’s figures are cited here and are not ours.
- Harness source, read at the HEAD commits recorded in our measurement log: zed, codex, Roo-Code, vscode-copilot-chat, gemini-cli, goose, OpenHands, plandex, AutoGPT, LibreChat, opencode, crush, continue, crewAI, cline, aider, letta, agno, aichat and browser-use.
Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.
Published via Towards AI
Towards AI Academy
We Build Enterprise-Grade AI. We'll Teach You to Master It Too.
15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.
Start free — no commitment:
→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day
→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages
Our courses:
→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.
→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.
→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.
Note: Article content contains the views of the contributing authors and not Towards AI.