ANALYSIS 10 min read

Long-Running AI Agents Need Three Memories: Working Context, Durable State, and an Audit Record

Context Language Models, Pi Durable, and OpenAI's long-horizon Codex guidance converge on one lesson: persistence is not one feature, and reliability requires separate memory layers.

By EgoistAI ·
Long-Running AI Agents Need Three Memories: Working Context, Durable State, and an Audit Record

Three current sources describe long-running AI work from different layers. The Context Language Models paper lets a model rewrite the material inside its active context. Pi Durable persists conversations, tasks, and application state across crashes. OpenAI’s long-horizon Codex case study emphasizes specification files, milestone plans, verification, and a continuously updated audit log.

Taken together, they suggest that “agent memory” is too vague to be a useful architecture. A production agent needs at least three distinct memories, each with different mutability and trust.

Layer one: editable working context

The working context is what the model can see now. It must be compact enough to fit the model and relevant enough to guide the next action. Context Language Models make that layer explicitly editable: the agent can preserve exact facts, compress old turns, delete noise, or maintain a current scorecard.

This layer benefits from flexibility because relevance changes. A test failure that matters during diagnosis may become noise after repair. A user constraint and a definition of done should survive much longer. The CLM results indicate that model-directed editing can improve task performance and reduce measured compute under the authors’ benchmarks.

The same mutability creates risk. Working context can drift, lose provenance, or retain an injected instruction. It should therefore be treated as a derived view, not the authoritative record.

Layer two: durable operational state

Operational state answers different questions: which task was running, which tool call started, whether it completed, what should retry, and which external side effect already happened. Pi Durable addresses that layer with checkpoints, task records, typed application documents, request identifiers, and replay policies.

A summary paragraph cannot replace these facts. “The deployment probably finished” is not a safe recovery state. The system needs the deployment identifier, destination, version, response, and retry contract. Durable state must be structured and committed around the action it represents.

This is where ordinary distributed-systems concepts return. Idempotency, atomic updates, ownership, cancellation, and versioning matter more as agents receive wider tool access. A smarter model does not make a duplicated payment acceptable.

Layer three: immutable evidence and human-readable status

OpenAI’s long-horizon case study describes a roughly 25-hour Codex run that used about 13 million tokens and produced around 30,000 lines of code. The article labels it an experiment rather than a production result. Its most transferable practice was not a model setting; it was a set of durable project files.

A specification fixed the target and non-goals. A plan divided the work into milestones with validation commands. An implementation runbook defined behavior after failures. A documentation file tracked status, decisions, known issues, and the demo path.

This third layer serves humans and later audits. It should not be freely rewritten to make a run look successful. Tool logs, diffs, approvals, source links, and verification outputs should remain available even when the model works from a smaller summary.

Why the separation matters

Combining the layers creates predictable failures. If the transcript is the database, recovery becomes guesswork. If the database is copied wholesale into the prompt, cost and distraction grow. If the model can rewrite the only audit record, reviewers cannot distinguish a genuine event from a later summary.

A better flow is one-directional at the trust boundaries. Immutable events feed structured state. Structured state and selected evidence produce an editable working view. The model proposes actions against that view. Consequential actions create new events, often after policy checks or human approval.

The model may help summarize the event history, but the summary points back to the evidence. It may propose a state transition, but the harness validates and commits it. It may rewrite working memory, but it cannot erase the original user constraint or authorization record.

A practical architecture

For a coding agent, the working context can contain the active milestone, relevant files, recent test failures, and a compact decision history. Durable state can hold job status, worktree, commit identifiers, pending approvals, and tool-call records. The audit layer can retain the original specification, full logs, diffs, review comments, and verification results.

For a research agent, working context can hold unresolved questions and current source notes. Durable state can track fetch jobs, deduplication, and publication workflow. The audit record can preserve URLs, retrieval timestamps, excerpts, and the distinction between primary evidence and community reaction.

Each layer needs independent tests. Working context should be evaluated for retention of critical constraints and resistance to injection. Durable state should be tested by killing processes at every boundary. The audit layer should be checked for completeness, provenance, and access control.

What teams should measure

Success rate remains essential, but it is not enough. Measure recovery after interruption, duplicate external actions, missing approvals, context-edit diffs, loss of cited evidence, time to human diagnosis, and the fraction of claims that can be traced to an artifact.

Cost should include model inference, tool execution, storage, re-prefill after context edits, and human review. A method that saves tokens but increases duplicate work may be worse. A durable system that resumes perfectly but preserves a corrupted instruction may also be worse.

METR’s time-horizon work provides a useful external framing: longer task duration is becoming a meaningful capability dimension. But a task completed once in a controlled environment is not the same as a service operating safely across failures and users.

Limitations

The three sources do not form a common benchmark. CLM reports research results on selected tasks. Pi Durable presents a framework and demonstrations. OpenAI describes one ambitious case study. Their claims should not be merged into a synthetic performance number.

The architecture also does not solve authorization, model error, or adversarial input. Separating memory layers limits the blast radius and improves diagnosis; it does not guarantee correct decisions. Human review remains necessary where errors affect money, privacy, production systems, or people.

The emerging lesson is nevertheless concrete: long-running agents are becoming stateful software. Reliability will depend on what they are allowed to forget, what they must remember exactly, and what nobody—including the model—may quietly rewrite.

Share this article

> Want more like this?

Get the best AI insights delivered weekly.

By subscribing, you agree to our Privacy Policy. You can unsubscribe at any time.

> Related Articles

Tags

AI agentslong-horizon tasksagent architecturedurable stateanalysis

> Stay in the loop

Weekly AI tools & insights.