danielhuber.dev@proton.me Saturday, October 3, 2026
Context Layer

Context Language Models: Giving the Model Write Access to Its Own Context

An examination of Context Language Models, a University of Washington and Meta proposal that exposes an agent's live context as an editable file, and what it changes about compaction, caching, and training.


Almost every production agent harness treats the context window as an append-only log. The model emits tokens, tool results get appended, and when the log approaches a length threshold the harness steps in with a fixed policy, most often a summarisation pass of the kind Codex and Cursor run. Even the more adaptive approaches published this year, such as Self-Compact, AutoCompact, and ACM, give the model a small menu of predefined actions (compact now, offload this, retrieve that) rather than control over the context itself.

Context Language Models (CLMs), from the University of Washington and Meta Superintelligence Labs, remove the menu. The model’s live context is mirrored to a file, the path is put in the system prompt, and the model can edit that file with ordinary Bash and Python. Whatever the file contains becomes the next turn’s context. The code is released at facebookresearch/context-language-models.

From append-only to model-controlled transitions

The paper frames the difference as a change to the context transition. A standard LM appends its output: c[t+1] = c[t] ⊕ f(c[t]). A CLM produces the next context directly: c[t+1] = f(c[t]), where f can be any edit the model chooses to write. Prior context-management strategies are special cases of this function; the CLM simply stops fixing which cases are allowed.

ApproachWho decides whenWho decides howExamples
Harness-scheduledHarness (length threshold or every turn)Harness (fixed summarisation prompt)Codex, Cursor, MEM1
Constrained actionsModelHarness (predefined tool semantics)Self-Compact, AutoCompact, ACM, Context-as-a-Tool
Context as a REPL variableModelModel, but read-only over the inputRecursive Language Models
Context Language ModelModelModel, read-write over the full live contextCLM

The distinction from Recursive Language Models is worth noting: RLMs let the model decide what to read from a large input held outside the context, but retrieved material still accumulates in an append-only history. CLMs make the history itself editable. The authors describe the two as complementary.

How context-as-a-file works

The implementation is deliberately thin. Edits to the context file are synchronised to the model server and used for the next generation. If the model does not touch the file on a given turn, its tokens are appended as usual, which keeps the common case cache-friendly.

Context as a file
turn t:   live context  ──mirrored──▶  /workspace/ctx_0.txt
                                            │
          model writes Bash / Python  ──────┤  (regex replace, truncate,
                                            │   rewrite a span, add a role)
                                            ▼
turn t+1: edited file  ──synchronised──▶  live context sent to server

no edit this turn  ⇒  generated tokens are appended (default)
multi-agent        ⇒  one file per agent; spawning a subagent = creating a file

Because the interface is general-purpose code rather than a fixed tool, the qualitative behaviours the authors report are ones no harness designer specified. On a circle-packing run, a model kept an orchestrator scoreboard for 21 subagents up to date through 163 in-place edits while holding its own context at 6–8K tokens. On BrowseComp-Plus, a model defined a compact_turns helper at step 109 and called it 37 times; another wrote a loop that replaced irrelevant search results with a one-line “No relevant results” marker; another introduced a new notes role alongside the chat template’s existing roles. All of this was zero-shot, with no training for the behaviour.

Where fixed strategies break

To isolate context management from reasoning, the authors built ContextBench, four synthetic tasks run under a 32K limit at up to 24× context pressure (input volume divided by context limit). Each targets a capability a fixed strategy lacks: verbatim retention of selected “needles” (summaries paraphrase or lose them), surgical in-place updates to a Sudoku board as moves stream in (append-only methods must regenerate the full board), and exact key-value and log recall (coding tools can offload data to disk but cannot evict it from the live context afterwards). With GPT-5.4, every baseline degrades as pressure rises; the CLM stays near perfect on all four. The repository lists ContextBench as “coming soon”, so it cannot yet be reproduced.

Results on long-horizon tasks

The zero-shot evaluations use existing models with a shared minimal harness, mostly at a 32K context budget.

TaskCLMStrongest baselineCompute
BrowseComp-Plus (Qwen3.6-27B)59.4%Codex-style summary, 11.4% lower (relative)21.5% fewer prefix-reuse FLOPs
TerminalBench 2.1Matches summaryCodex-style summary~70% of baseline FLOPs
TBLite73.7%67.0% (summary)91% of baseline FLOPs
EdgeBench-10, 12 h (Qwen3.6-27B)44.642.3 (summary)179 vs 437 PFLOPs per trial
EdgeBench-10, 12 h (Claude 4.6 Sonnet)51.042.3 (summary)—
Software World, 24 h six-repo swarm (GPT-5.6-Sol)1.044× speedup1.026× (summary swarm)Same spend

On four AlphaEvolve-style mathematical optimisation problems, a general CLM agent given the evolutionary algorithm only as in-context guidance beat the specialised OpenEvolve workflow on all four with Claude 4.6 Sonnet. Running five subagents added little on single-repository EdgeBench, which suggests the gains come from context control rather than parallelism.

Note

A 32K budget is far below what current models support, and it is chosen to force context management to matter. The comparison is fair across methods, but the size of the gap at 200K+ budgets on typical production workloads is not something this paper measures, apart from the Software World swarm run at 272K.

The cache cost of editing the middle

Editing context has a cost that token counts hide. Prompt caching reuses computed attention states only up to the first changed token; an edit in the middle of a 30K context forces everything after it to be prefilled again. A strategy that keeps contexts short but edits early spans can end up more expensive than one that lets the context grow.

The authors account for this with prefix-reuse FLOPs: prefill cost for the unmatched suffix plus decode cost for new tokens. All headline efficiency numbers use this metric under standard serving, so the reported savings already include the re-prefill penalty.

They then reduce that penalty with Suffix Cache Reuse (SCR), implemented in SGLang: after an edit, cached states for the unchanged suffix are kept rather than recomputed, and only the inserted tokens are prefilled. Those suffix states were computed against the old prefix, so they are technically stale, but on BrowseComp-Plus accuracy was identical (60.2% under both) while server-side compute fell by 35%. SCR also applies outside CLMs: serving stacks that strip prior reasoning tokens between turns hit the same re-prefill problem.

Steering, evolving, and training the strategy

Because context management now happens through the model’s own outputs, it responds to the same levers as any other behaviour.

  • Instruction. A single sentence in the task prompt changes the policy. “Compact to 4k tokens once you reach Y tokens” moved the first compaction to track Y; asking the model to back up to disk before compacting raised the rate of full backups from 0 to 0.68.
  • Skill evolution. A prompt-evolution loop (rollouts, a proposer drafts candidate skills, selection on a dev split) improved held-out ContextBench accuracy by up to 35.9 points while also lowering compute. It worked both with a stronger proposer model and with the agent proposing its own skills.
  • Reinforcement learning. Rewarding context edits directly invites reward hacking, since the model can delete useful information to look efficient. The authors use stepwise GRPO with a success-gated efficiency advantage: only successful trajectories are re-ranked by compute, and failures receive no efficiency bonus. Qwen3.5-9B went from 28.8% to 42.5% on BrowseComp-Plus, starting six points behind a summary harness trained the same way and finishing 0.4 points ahead of it (42.1%), at 1.34 versus 2.19 PFLOPs per question.

The authors suggest a further step: distilling existing harnesses into CLM behaviour, treating a harness as procedural memory that a model can later absorb. That is the same pressure described in the harness-model training loop, applied to one specific part of the harness.

Risks and limits

The paper’s discussion section names the main risk directly. Editable context is a persistence channel: a prompt injection that reaches the model can now write itself into the context that every subsequent turn reads, and so can a model’s own mistaken instruction. OpenAI’s alignment team has already documented self-generated prompt injections in compaction summaries, and arbitrary edits give that failure more room. A second, quieter issue is auditability: once the model rewrites history, the trace a reviewer sees is no longer the input the model acted on unless every context version is logged.

Other limits are practical. The code is released under CC BY-NC 4.0, which rules out direct commercial use. ContextBench is not yet public. The RL results cover one 9B model on one task family.

Practical takeaways

  • The zero-shot result is the cheap one to test. It needs no training: expose the context as a file, keep append-by-default, and compare against your current compaction policy on your own long-horizon tasks.
  • Keep an immutable log outside the editable file. Store every context version (or every diff) so debugging, evaluation, and incident review still see what the model actually saw.
  • Measure cost with caching in the model. Prefix-reuse FLOPs, or cached versus uncached input tokens on a hosted API, are the relevant numbers. Shorter contexts are not automatically cheaper.
  • State the policy in words before building it in code. The steering results suggest many compaction rules that harnesses hard-code today could be a sentence in the system prompt.
  • Treat context edits as writes to a trusted store. Diff each edit and flag inserted text that reads like instructions, the same way you would guard agent memory against poisoning.

For background on the problems CLMs address, see Context Engineering and Context Bloat & Context Rot.

Tags: researchcontext engineeringcompactionlong-horizon agentsprompt cachingreinforcement learning

This article is an AI-generated summary. Read the original paper: Context Language Models .