Context compression

Most of what an agent sends, it has already sent

A model keeps nothing between calls. Every turn re-sends the whole conversation - files it already read, tool output it already used, its own reasoning. Compression is the work of sending less of that without losing what the next turn needs.

What it is

A language model is stateless. The context window is not memory: it is an argument, passed again in full on every call. A coding agent twenty turns into a task is paying for the first nineteen every time it speaks.

Context compression reduces what goes into that argument while keeping the meaning the model needs. It happens between the client and the model, so nothing about the client changes and nothing about the model changes.

What SOMA compresses

SOMA compresses the conversation and the chain-of-thought at the gateway, on the way to the upstream model. Two things are deliberately left alone: the system and developer instructions, and the user's first turn. Those set the task, and a task description that arrives paraphrased is a different task.

If the compressor fails, the request continues uncompressed. Failing open is the only acceptable behaviour for something that sits in front of every call - a compressor having a bad minute must never become an outage.

The families of methods

Extractive methods keep a subset of what is already there - whole messages, sentences or spans - and drop the rest. They are cheap and they never invent anything, because every token in the output was in the input.

Abstractive methods rewrite. A summary of nineteen turns is shorter than any subset of them, and it can carry the shape of a long exchange in a paragraph - at the cost of a second model's judgement about what mattered.

Learned methods train a compressor for the model that will read the result, sometimes replacing tokens with vectors the model consumes directly. Highest ceiling, tightest coupling: what works for one model rarely transfers to another.

In practice, mixtures win. Keeping the recent turns verbatim, summarising the old ones and dropping tool output that has already been used is three families in one pipeline.

Where compression does not pay

Prompt caching and compression solve the same problem from opposite ends, and they interfere. When a cache is hitting on 87 to 93 percent of a prefix, the tokens compression would remove are the cheap ones - and rewriting them changes the prefix, which costs the cache hit as well.

Compression earns its place where context is long, changes between turns, and is paid for at full price. Where the same prefix is sent over and over unchanged, caching is simply the better instrument.

How the numbers are reported

Every saving figure carries its basis. A measured figure comes from a real run; an estimate is what the system expects before it has measured, and it is labelled as one. The API refuses to present the second as the first, and this site follows that rule.

The published estimate of about fifteen percent is an estimate. The measured figure on Copilot CLI traffic is far lower, and a single benchmark run showed around ten percent with the author's own caveat that it was not a clean baseline. Compression is worth what it saves on your traffic, which is why the dashboard reports per-task token counts rather than one headline number.

Where you can use it

The SOMA app is the gateway: an OpenAI-compatible endpoint that compresses on the way through, so any client that speaks that protocol works unchanged.

Somarizer is the same idea for a single document - paste text or drop a PDF and get it back shorter. The OpenClaw compressor is open source, for people who would rather run it themselves.