Papers
arxiv:2609.37725

Context Language Models

Published on Sep 29
· Submitted by
Rulin Shao
on Sep 30
Authors:
,
,
,
,
,
,
,
,
,
,

Abstract

We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.

Community

Paper submitter

The Bitter Lesson for context management: Giving LMs unrestricted access to their own context beats human-designed SOTA!

Introducing 🩵 Context Language Models (CLMs) 🩵

  • Natively manage their own context
  • Treat context as a file
  • Learn policies in CLM weights, no harness

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Interesting result. I've only read the abstract, so forgive me if the paper covers this. If the model can rewrite its own context file without restriction, how does someone tell afterward what it dropped and why? I'd want each edit kept as a diff with a reason, so a person can audit what the model decided to forget. Does the file history survive the run?

·

Thanks for the question. Yes, the full history survives the run, and every edit can be audited after the fact. In the released harness, each trial records:

  • the exact context sent to the model at every call (context_snapshots/), so diffing consecutive snapshots shows precisely what an edit removed or rewrote;
  • the edit itself, which is a command the model runs on its context file, together with the reasoning the model wrote before running it (trajectory.json);
  • the intermediate context segments in a structured trajectory format (trajectory.ctx.json, our ATIF-CTX extension of Harbor's ATIF), plus per-run counts of edits (usage.json).

The system prompt and the initial task are protected from edits. A stored diff with an explicit, model-written reason per edit would make this audit more direct, and we think it is a good addition. See the Outputs section of the harness README: https://github.com/facebookresearch/context-language-models/tree/main/clm/clm_harness

My main concern with this paper is whether the reported results justify the broader engineering claims about production agents. Treating context as an editable file is straightforward to implement, but its practical value depends on how it compares with simpler alternatives under realistic constraints.

  1. End-to-end latency and actual cost
    The paper accounts for editing-related computation and re-prefilling, and the 12-hour EdgeBench experiment provides evidence under a real time budget. However, for interactive agents, I would also want task completion time, time to the next useful response, p50/p95 latency, and actual serving or API cost alongside task quality.
    Context editing saves future inference work while adding work to the current execution loop. The net benefit depends on edit frequency, cache behavior, tool overhead, and workload. Lower FLOPs do not by themselves establish better responsiveness or lower deployment costs.
  2. Stronger low-cost baselines
    A practically important comparison would include a cheaper dedicated compaction model, deterministic filtering of repetitive tool output, and asynchronous or hybrid context management. These alternatives have weaknesses too, so I would not assume they win. But without well-tuned comparisons, it is difficult to establish when using the main reasoning model for context maintenance is worth the cost.
  3. Policy control and reliability
    Different applications need different information preserved. Learned context-management behavior and domain-specific instructions can coexist. The harder question is whether explicit retention requirements remain reliable under domain shifts, conflicting instructions, and long trajectories.
    What happens when an edit removes a critical user constraint, a source reference, or evidence needed for an audit? How is that detected and recovered? These failure modes need explicit evaluation before unrestricted self-editing can be treated as a dependable production mechanism.
  4. Scope of the conclusions
    The experiments support self-editing as a potentially useful option for the evaluated workloads. I am less convinced that they establish a general case for replacing externally controlled context management. Where does this approach regress? Could a simpler hybrid preserve most of the gains while keeping critical information and recovery under deterministic control?
    The adoption criterion should be a measurable improvement over simpler alternatives at comparable quality, latency, cost, and operational risk. That is the gap I would like the paper to address more directly.
·

Thanks for the careful read. Some clarifications, and where we agree:

1. Latency and cost. We report two kinds of cost on purpose. For the method comparison we use prefix-reuse FLOPs, a token-level count that includes the cost of edits and of re-prefilling after an edit, so a method that edits too often pays for it. We chose an analytic measure here because wall-clock time depends on each method's implementation and on how many evaluations share a server, which would bias the comparison across methods. On BCP, CLM uses 21.5% fewer prefix-reuse FLOPs than Codex-style summarization while scoring higher (Sec. 5.1). For serving, we measure on the server itself. The main serving cost an edit adds is re-prefilling the cache after the edit point, which prefix caching cannot reuse. Suffix Cache Reuse (Sec. 3.3 and 5.3, released in the repo) targets this, and SGLang's own counters show it serves BCP with 35% less compute than standard SGLang at matched accuracy. In a vLLM port with Qwen3.6-27B on one H200 and 8 concurrent conversations, it raises throughput from 1.14 to 1.78 requests/s and lowers median time to first token from 1.80 s to 1.09 s. The EdgeBench runs use a real 12-hour wall-clock budget. We agree that interactive p50/p95 latency and serving cost under production load are worth reporting alongside quality.

2. Baselines. All methods share one backbone and are evaluated out of the box: Codex-style summarization, ACM, Self-Compact, MEM1, RLM, and OpenEvolve for the optimization tasks. A cheaper dedicated compaction model and deterministic filtering of repetitive tool output are not among them, and they are good comparisons to add. They also combine naturally with CLM: since the context is a file the model edits, such a filter or a small summarizer can be offered to the model as a tool or skill.

3. Retention and auditability. These are three separate questions.

  • Prevention: the released harness protects the system prompt and the initial task from edits, so the original task constraints always stay in context. Constraints introduced later in a session are not protected this way.
  • Auditing: the context sent at every call is recorded (context_snapshots/, ATIF-CTX), so any edit can be inspected after the fact.
  • Detection and recovery during a run: not addressed in this paper.

ContextBench measures retention under context pressure: what an agent keeps, updates, and offloads. A natural extension is to mark policy-critical content as non-editable and let the model manage the rest of its working context. Sec. 6 discusses editable context as a channel through which injected or self-generated instructions can persist.

4. Scope. Our main claim is that context management can be a capability of the model itself, learnable in context and in weights, rather than only a fixed harness operation. The evaluations show it is competitive and often cheaper on these workloads; on TB2.1, for example, CLM matches the strongest baseline at 70% of its FLOPs. We do not argue against external control: Sec. 6 frames existing harnesses as strategies that can be distilled into CLMs, and skills let domain-specific retention rules coexist with learned behavior. We agree that adoption in production should be decided on the full quality, latency, cost, and reliability trade-off.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.37725
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 1

Spaces citing this paper 1

Collections including this paper 4