The log is the truth

September 2026

On March 12 I tweeted: "i don't like the context summarization feature in claude code because it deletes the past context. sometimes i want to go back and resume from the beginning of the conversation or refer to a message before the summarization, but i can't..."

Three weeks earlier someone had opened issue #27242 on the Claude Code repo with the diagnosis: "The full data is always preserved in transcript.jsonl, but the TUI provides zero functional paths to access it." Claude Code's session docs describe what compaction does. It "sends one summarization request over the full history, then replaces the history with the summary, your most recent exchanges, and up to five recently read files." And: "whatever the summary leaves out is no longer in Claude's context."

So nothing was deleted. It was on disk, one JSON object per line. What got deleted was every path back to it, for me and for the model.

In December 2013 Jay Kreps, who built Kafka at LinkedIn, published The Log. It opens with how databases already work: "The log is the record of what happened, and each table or index is a projection of this history into some useful data structure or index." Databases solved this decades ago. Nobody truncates the redo log because an index got too big.

My claim is that a context window is a table, not the database. A harness should persist everything that happens, in order, and separately decide what the model reads on its next call. The window is a derived view: a list of references into the log, rebuilt from the log. Compaction, pruning and summarization edit that view. They never touch the log. Most harnesses merge the two, so their memory policy and their prompt policy are the same policy, and every time they shrink the prompt they make the agent forget. The industry now calls this work "context engineering," and most of it curates the window as if it were the agent's state. The window should be a query.

I've been building this into a personal assistant harness I call truth. This post is the argument, with code from truth and the places where I think it breaks.

Two things that got merged

Martin Fowler wrote the pattern down in 2005 in Event Sourcing: "Event Sourcing ensures that all changes to application state are stored as a sequence of events." His framing is the one I keep coming back to: "we are persisting two different things an application state and an event log." Keep the events and you get two things for free. You can do a complete rebuild of the state, and you can do a temporal query, including "multiple time-lines (analogous to branching in a version control system)." That second one is my tweet. I wanted to go back to the start of a conversation and branch from there.

Most agent harnesses have one thing where they need two. The messages array is both the event log and the application state. The harness appends to it, sends it to the model, and when it gets too big, rewrites it. Claude Code at least keeps a JSONL copy on the side, which is more than most. But the running session, and the model, live on the rewritten array.

I wrote the fix in my notes on September 18. My notes are in Portuguese, so every quote from them here is my translation: "the logs of everything that happened are different from what will be ingested into the model's context window." Everything goes in the log: user messages, tool calls with their arguments, full tool outputs, background runs, heartbeats, and the harness's own decisions about what to show. The view is derived from it. Ideally the view is only references to log entries, so it's cheap to store and exact to rebuild.

The idea isn't new in AI either. MemGPT framed it in 2023 as an operating system problem: "virtual context management, a technique drawing inspiration from hierarchical memory systems in traditional operating systems which provide the illusion of an extended virtual memory via paging between physical memory and disk." What's changed is that there are now frameworks built around the log. truth runs on Tardigrade, a TypeScript framework whose README states the whole design in one line, { view, transitions } = f(event log), and whose homepage says: "The durable log is also the trace." truth's own rule, from the top of its TODO file: "The log is the memory. A decision that changes what the model sees, or changes the world, is an event."

Once the log is the source, you get more than one view out of it. The model reads one projection. truth's console shows the conversation, the model's current view and the raw thread events side by side, which is a second projection. An audit trail is a third, and an eval replay forked from any checkpoint is a fourth. Only the log is authoritative.

The view has to stay small

If you keep everything, why not send everything? Because models don't read everything equally.

Position matters. Lost in the Middle (Liu et al., 2023) put the one document that answers a question at different positions among 20 retrieved documents. GPT-3.5-Turbo scored 75.8% when the answer was first, 53.8% when it was tenth, and 63.2% when it was last. With no documents at all, answering from memory, it scored 56.1%. In the middle of the context, the document with the answer made the model worse than having nothing.

Accuracy by position of the answer among 20 documents, from Lost in the Middle

The same paper found that more retrieval barely helps. Going from 20 to 50 retrieved documents improved GPT-3.5-Turbo by about 1.5% and Claude-1.3 by about 1%, "while significantly increasing the input context length (and thus latency and cost)."

Length alone hurts. Those are 2023 models. Chroma's Context Rot report (July 2025) tested 18 newer ones, including GPT-4.1, Claude 4, Gemini 2.5 and Qwen3, on tasks kept deliberately simple: "models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows." The test closest to an assistant's life was LongMemEval, conversational memory questions. Chroma compared focused prompts of about 300 tokens, holding only what the question needs, with the full versions of about 113,000 tokens: "Across all models, we see significantly higher performance on focused prompts compared to full prompts." They also found that "even a single distractor reduces performance relative to the baseline."

The builders say the same thing. Anthropic's Effective context engineering for AI agents names the effect: "context rot: as the number of tokens in the context window increases, the model's ability to accurately recall information from that context decreases." It also rules out the easy fix: "Waiting for larger context windows might seem like an obvious tactic. But it's likely that for the foreseeable future, context windows of all sizes will be subject to context pollution and information relevance concerns." Manus, in Context Engineering for AI Agents: "Model performance tends to degrade beyond a certain context length, even if the window technically supports it."

My rule is to keep the view under 40% of the window. The note from September 18 says: "keep the context window always small, <40% (optimal Pareto frontier of fast and intelligent)." I haven't measured the knee of the curve for each model. 40% is a budget I picked from experience and from curves like the ones above, and it's one constant in truth (VIEW_RATIO = 0.4) so the traces can move it later. Whatever the number, the budget applies to the view. The log has no budget.

Compaction edits the view

Anthropic's post defines compaction as "taking a conversation nearing the context window limit, summarizing its contents, and reinitiating a new context window with the summary." Then it warns that "overly aggressive compaction can result in the loss of subtle but critical context whose importance only becomes apparent later." Manus says the same thing more sharply: "you can't reliably predict which observation might become critical ten steps later. From a logical standpoint, any irreversible compression carries risk." Their answer is that "our compression strategies are always designed to be restorable." A web page's content can leave the context as long as the URL stays.

The log/view split turns that trick into a property. If the view is references into a log that never loses anything, no compression is irreversible. A summary becomes one more event in the log, and the events it summarized are still there.

In truth, the thing that can leave the view is called a unit: one model response's batch of tool calls and their results, or a whole closed turn. Hiding a unit appends a ContextHidden event that lists unit ids. Tardigrade's stock compaction works the same way: it appends a CompactionCompleted event, and the projection shows the model a "Summary of earlier work" in place of the older turns while the turns stay in the log. Rebuilding the view is a fold over events. From my harness truth:

export const visible = (events: ReadonlyArray<Event>, hidden: ReadonlySet<string>): ReadonlyArray<Event> =>
  hidden.size === 0 ? events : owned(events).flatMap((entry) => (hides(entry, hidden) ? [] : [entry.event]))

// viewOf rebuilds, from a thread's log, what the model reads on its next call and why.
export const viewOf = (events: ReadonlyArray<Event>) => {
  const hidden = new Set(events.flatMap((event) => (isContextHidden(event) ? event.hide : [])))
  const trajectory = visible(projectedOutput(events), hidden)
  const routed = events.findLast((event) => event.type === "EffortRouted")
  return { tier: routed?.tier, tokens: estimateTokens(trajectory), hidden: [...hidden], messages: renderMessages(trajectory) }
}

viewOf takes only the log. It returns exactly what the model will read next, how many tokens that is, which units are hidden, and which reasoning effort the turn was given. There's no other state to consult, so a replay produces the same view. Even the reasoning effort is read from the log at request time, so "a replay reads the same event."

The second half is the read path. truth's memory package describes itself to the model as "Your notes under memory/ and the complete history of this conversation, including what no longer fits in view." The system prompt says: "Search notes and the full history of this conversation with memory.search before saying you do not know something." That's the piece Claude Code is missing. It has the JSONL. The model has no tool that reads it.

A System One model prunes the view

Which units should go? Anthropic points at the obvious candidates: "once a tool has been called deep in the message history, why would the agent need to see the raw result again? One of the safest lightest touch forms of compaction is tool result clearing." Clearing every old result is blunt, and summarizing is slow and lossy. What I want is a cheap yes-or-no per unit.

That became practical this month. On September 15 TypeSafe released Jev, which they call a System One model: "unstructured state in, typed probabilistic decisions out." It doesn't generate text. It returns probabilities for typed questions, priced at $0.042 per million input tokens with no output charge, and TypeSafe reports end-to-end responses of 70 to 500 ms. Vercel's guide says: "Jev is a probabilistic decision model, not a chat model." Tamara Tran then published fast-jev-compaction, a "Claude Code plugin that replaces the compaction summary with Jev decisions: every tool call and result is scored in one fast request, stale ones are dropped or truncated, everything kept stays verbatim." I took that idea into truth with one change: nothing gets dropped from the log.

Before the first model call of every turn, truth runs what it calls the prelude. Two Jev questions go out in parallel.

The first is asked about each unit from earlier closed turns: "Must units.{{id}} remain visible for the assistant to continue correctly?" The answer "yes" is defined as: "The current work still depends on this unit. It holds a constraint, an error, an identifier, a decision, or evidence that later visible work does not already record." The answer "no" is: "Later visible work already records the outcome, or this unit is an old detour the current work does not use." A unit whose keep probability is under 0.2 leaves the view. In my notes the rule was to hide anything more than 80% likely to be useless. If the view is still over 40% after that, the least-needed batches go next, then the oldest turns, until it fits. My note from September 22 spells out the fallback: cut the oldest part "not in the log but in the effective window" until it's under 40%, and "that should be the algorithm."

The second question routes effort. Jev scores how much work the request needs, "the work, not the length of the reply," as the router prompt puts it, and the tier maps to the model's reasoning effort: low, xhigh or max. Low confidence takes at least the middle tier. Three tiers is a starting guess. I haven't measured whether finer ones pay for themselves.

Both decisions land in the log as events (ContextHidden and EffortRouted) before inference is allowed to start. Both are purely additive. If Jev is down or the key is missing, the judge returns nothing and the router returns the middle tier. The view still fits the budget by dropping the oldest units. The harness works without a System One model. It just works worse.

The rubric also answers an objection from Manus, whose post argues for keeping failures visible: "Erasing failure removes evidence. And without evidence, the model can't adapt." Errors are on the keep list by name. Jev is asked whether anything still visible already records what the unit taught, not whether the unit is interesting.

Cache geometry

There's a second reason to keep the view small and stable, and it's money.

Time to first token with and without a cached prefix, from Anthropic

Anthropic's prompt caching announcement lists what early customers saw. Chatting with a book behind a 100,000-token cached prompt went from 11.5 s to 2.4 s to first token (-79%) and cost 90% less. A 10-turn conversation with a long system prompt went from about 10 s to about 2.5 s and cost 53% less. On Claude 3.5 Sonnet, base input was $3 per million tokens, cache writes $3.75, cache reads $0.30.

Manus goes further: "the KV-cache hit rate is the single most important metric for a production-stage AI agent." Their agent's average input-to-output ratio is "around 100:1," so an agent is mostly prefill, and prefill is what the cache saves. Their rules follow from how the cache works. "Even a single-token difference can invalidate the cache from that token onward," so a timestamp at the top of the system prompt "kills your cache hit rate." Serialization must be deterministic. And: "Make your context append-only."

My own rule is to keep the system prompt and the tool list as small as the agent's work allows. Whatever changes goes at the very end, so the prefix before it can be reused. The static part goes first and stays byte-identical: a short system prompt with no clock in it, then a stable tool list. truth's system prompt is 1,521 characters, and nothing before the conversation changes between calls. Heartbeats and other runtime events arrive as messages at the tail, and notes are searched on demand instead of injected.

Pruning the view looks like it breaks this. Hiding a unit rewrites the middle of the prompt, and Manus says never to do that. The way out is timing. truth only hides units in the prelude, once per turn, and only from closed turns. Inside a turn the view is append-only, and that's where the tool loop runs. Manus says a typical task takes "around 50 tool calls," so that's where the 100:1 prefill compounds. The cost is at most one partial cache miss per turn, starting at the earliest newly hidden unit, and a view held under 40% makes that miss cheaper. I haven't measured truth's hit rate yet.

Memory is prose you can grep

The log answers "what happened." Notes answer "what did I learn." They're different objects, and I want both.

On August 14 I wrote that we shouldn't impose "structural despotism" on the .md files: agents should be free to write what they want, be complete, and put down tacit knowledge. Two weeks earlier I'd written that an agent's work should get better over time, since it would have more rhizomes to search with BM25. I call this rhizomatic memory, a word I took from Deleuze and Guattari: plain Markdown notes with no schema and no root, linked to each other, searched on demand instead of preloaded.

In truth, notes are a git repository on the host. Every write a tool call makes is one commit, with a Truth-Call trailer naming the call that wrote it, so the notes have a log of their own. Links are written [label](other.md): why, and memory.graph reads them back. memory.search runs BM25 (SQLite FTS5) and embeddings over the notes and the full thread history, then merges the two ranked lists with reciprocal rank fusion. BM25 goes first because what an assistant needs to find is usually an exact token: a name, an order id, an error string, the word the person used. Embeddings catch the paraphrase BM25 misses.

A friend found the same shape in the wild. On September 20 he dropped Dhravya Shah's reverse-engineering of Instinct's memory into our group chat. According to the article, Instinct's memories "are stored as git-tracked markdown files," and there's "no vector indexing, or BM25 search." He pasted two passages. The first: "Instinct quite literally just uses keyword matching / grep-style queries to look things up in the file system. This is why every file has aliases associated, so that every file has a good chance of showing up when the agent is looking for it." The second was about the split between writing and reading: "Memory is read only, atleast for the agent," with Dhravya adding, "A background process does the work of combining things, not the main agent." He estimated consolidation runs about once a day. He's also clear that it's inference from the outside: "because i haven't seen their code, some things may be wrong."

Two things from that thread changed how I think. Aliases are a cheap fix for grep's biggest weakness, which is not knowing the word the note used. And a read-only agent with a background writer is the log/view split applied to memory: raw conversation in, derived notes out, and the derived side can always be rebuilt. truth still lets the main agent write notes directly. I haven't settled which is better.

Claude Code navigates code the same way, with grep instead of an index. Anthropic describes it as a hybrid: "CLAUDE.md files are naively dropped into context up front, while primitives like glob and grep allow it to navigate its environment and retrieve files just-in-time."

Your trajectories are memory

The largest log most of us own is our history with agents. On August 25 I wrote that my end plan was a skill with RAG over it, so that during trajectories the model could check what I'd said before and what my opinions were. I built it as a CLI over every session in every harness I use, untruncated.

This essay came out of that log. The quotes from my notes are from a digest of 4,985 messages I sent to AI agents between March and September. If each harness had compacted those sessions in place, there'd have been nothing to search. In Intelligence is a trajectory I argue the trajectory is the unit of work. It's also the unit of memory.

A team that remembers

Michael Polanyi opened The Tacit Dimension (1966) with: "I shall reconsider human knowledge by starting from the fact that we can know more than we can tell."

Most of what a team knows never gets written down: which service lies about its latency, or why a migration was reverted. All day it passes through agent sessions, and it dies when those sessions are compacted or closed. I think every team's agents should write to a shared memory: the same free-form, linked Markdown notes, but one graph for everyone, so what one person's agent learned is there when someone else's agent searches. On August 14 I wrote that a memory like that should hold "all the content that might be important that passes trough the sessions, all the tacit knowledge, insights."

Polanyi's point cuts two ways here. Nobody will write a note about what they can't tell. But a trajectory records what was done, not only what was said, and an agent that watches the work can write down some of what the person never would. Writing is nearly free for an agent. How swarms share memory is a separate essay, Agent civilizations.

Speculation: the panopticon

This part is speculation, and I'll mark where it stops being evidence.

On August 12 I wrote that I want a panopticon program I can later connect to an AI companion, "where my AI companion will have infinite context about me." What I type, click, say and hear, recorded all day. I chose the ugly word on purpose.

The log/view split makes the idea coherent. An infinite log and a finite view don't contradict each other. The companion would read a small, fresh projection and search the rest. What I don't know, and nobody has measured, is whether a model with a perfect record of your life is more useful than one with good notes. The products that tried to find out mostly died, which is where the counterarguments start.

Where the log leaks

Lifelogging has a bad record. Microsoft Research's MyLifeBits started in 2001 as "a lifetime store of everything," with Gordon Bell as its subject. By 2016, per Wikipedia, Bell had stopped using the project's wearable camera and described the smartphone as largely fulfilling the Memex vision. Humane sold its assets to HP for $116 million. Its notice to customers said pins would stop working on February 28, 2025, when "all remaining consumer data will be permanently deleted." Rewind became Limitless and was bought by Meta in December 2025. Limitless's FAQ says the Rewind app "disables all screen and audio capture starting December 19, 2025," and service ended after that date in Brazil, China, the European Union, Israel, South Korea, Turkey and the UK.

My reading is that storage was never their problem. The hardware disappointed, and none of them had a good way to use what they recorded. But the fair version of the objection is stronger: maybe people don't want to be recorded, and no read path changes that. A log held by a vendor dies with the vendor. Humane gave its users ten days' notice. That's why truth keeps memory and thread logs on a host the person controls, and treats the sandbox as disposable.

Long context might win. Chroma notes that frontier models "achieve near-perfect scores on widely adopted benchmarks like Needle in a Haystack." If context rot gets solved, the 40% budget can loosen. Cost doesn't go away, though. Manus: "Long inputs are expensive, even with prefix caching. You're still paying to transmit and prefill every token." And the split survives either way. A bigger view is still a view.

Retrieval is sometimes wrong. A wrongly retrieved note is a distractor, and Chroma found a single distractor hurts. BM25 misses synonyms, which is why Instinct adds aliases and truth fuses in vectors. Dhravya rated Instinct's memory "Weak" on multi-hop questions across sessions and gave it a fail on implicit personalization. Search only helps when the model knows to search.

The model doesn't know what it can't see. This is the weakness I worry about most. truth hides units without leaving a trace in the view. The system prompt tells the model to search before saying it doesn't know, but a model rarely searches for something it doesn't suspect exists. Jev can also misjudge, and it judges each unit from at most 2,000 characters of it. A one-line stub per hidden unit ("3 tool calls to the calendar API, hidden, id u7") would cost a few tokens and might close the gap. I haven't tried it yet. Anthropic's warning about context "whose importance only becomes apparent later" applies to my pruning too. The difference is that my mistakes are recoverable, if the model thinks to look.

Privacy. A log of everything is also a liability. An append-only log and a right to be forgotten pull in opposite directions, and a shared team memory will collect secrets and customer data unless something stops it. truth handles one part: credentials go through a vault proxy on the sandbox's egress, and the model only ever sees key names, so secrets don't enter the log. For deletion I don't have a clean answer. The notes repository even keeps its reflog forever, which is the opposite of forgetting. The best I have is that the log lives on hardware its owner controls, and that deletion should itself be an event.

A log/view checklist

  1. Append everything. User messages, tool calls with arguments, full tool results, background runs, heartbeats, and your own routing and hiding decisions. Keep one durable store. A second engine gives the same fact a second home that can disagree with the log.
  2. Make the view a pure function of the log. If you can't rebuild what the model saw on call N from the log alone, you can't replay that call or fork from it.
  3. Record every view decision as an event. Hiding, compaction and effort routing are appends with ids. Nothing is mutated and nothing is deleted.
  4. Hide by reference, in whole units. A tool call and its result leave together, so the view never shows an orphaned result.
  5. Budget the view, not the log. My default is 40% of the window. Make it a constant you can move once you have traces.
  6. Prune once per turn, on closed turns only. Keep the tool loop append-only so the cache holds where the prefill is.
  7. Put a cheap judge in front, and make it additive. If the judge fails, fall back to dropping the oldest units. Never block a turn on it.
  8. Tell the judge what to keep. Constraints, errors, identifiers, decisions, and evidence nothing later records.
  9. Give the model a read path into everything hidden. Search over the full log and the notes, and a system prompt that tells it to search before saying it doesn't know.
  10. Keep the prefix static. A short system prompt with no clock, a stable tool list, deterministic serialization. Time, memory and runtime events go at the tail.
  11. Write memory as free-form Markdown in git. Links with a reason, aliases for the words people will search, BM25 first and vectors fused in.
  12. Host the log where its owner can reach it. Keep secrets out of it with a credential proxy, and plan for deletion before someone asks.

A projection function

The whole idea fits in two functions. project is the view: references into the log, rebuilt from the log and the hide decisions. fit proposes the next decision, and the harness appends it. Nothing is ever removed.

type Entry = { id: string; unit: string; tokens: number; closed: boolean }
type Hidden = { type: "ContextHidden"; hide: string[] }

const hiddenBy = (decisions: Hidden[]) => new Set(decisions.flatMap((d) => d.hide))

// The view is a list of log ids, rebuilt from the log and the decisions alone.
export const project = (log: Entry[], decisions: Hidden[]): string[] => {
  const hidden = hiddenBy(decisions)
  return log.filter((e) => !hidden.has(e.unit)).map((e) => e.id)
}

// keep maps a unit to the judge's probability that the model still needs it; missing means unjudged.
export const fit = (log: Entry[], decisions: Hidden[], window: number, keep: Map<string, number>, ratio = 0.4): Hidden => {
  const hidden = hiddenBy(decisions)
  const shown = log.filter((e) => !hidden.has(e.unit))
  let over = shown.reduce((n, e) => n + e.tokens, 0) - Math.floor(window * ratio)
  const units = new Map<string, number>()
  for (const e of shown) if (e.closed) units.set(e.unit, (units.get(e.unit) ?? 0) + e.tokens)
  const score = (unit: string) => keep.get(unit) ?? 0.5
  const hide: string[] = []
  for (const [unit, tokens] of [...units].toSorted(([a], [b]) => score(a) - score(b))) {
    if (over <= 0 && score(unit) >= 0.2) break
    hide.push(unit)
    over -= tokens
  }
  return { type: "ContextHidden", hide }
}

The sort is stable, so when the judge is absent every unit scores 0.5 and fit drops the oldest closed units first until the view fits. When the judge is present, anything under 0.2 goes regardless of budget, then the least-needed units until the view fits. Render the result as [static prefix, ...project(log, decisions), dynamic tail] and send it. Keep the log.

Sources

← Optimize the harness, not the model
Compute is not the bottleneck →

Markdown version: /blog/the-log-is-the-truth.md. Every essay: /agents.