Pi Durable Checkpoints Long-Running Agents and Replays Only Tools Marked Safe
Earendil's experimental harness persists conversations, tasks, and application state so an agent can resume after a crash. Its replay contract is the feature to study closely.
Earendil released Pi 1.0 and an experimental companion called Pi Durable on October 2 in Asia. Pi 1.0 remains a terminal coding agent; Pi Durable is a framework for applications whose conversations and tool calls must survive process failure, support multiple clients, and continue over a much longer period.
The public response is already substantial. At 1:08 p.m. Malaysia time, the Pi 1.0 Hacker News submission showed 957 points and 311 comments, while Pi Durable had 291 points and 36 comments. GitHub’s public API showed 111,348 stars and 14,138 forks for the Pi repository. Those numbers measure developer attention, not correctness, but they make this more than an unnoticed prototype.
What happened
Pi Durable stores conversation history, task state, queued input, and application documents behind a harness. A process opening the same storage can find unfinished work and resume it. The project currently offers memory, SQLite, and JSONL backends, with an interface for other stores.
Every model request and tool execution is represented as a task. A checkpoint is persisted before the harness advances. If a model request was interrupted, the request can run again while the partial answer remains marked as aborted. If a tool call was interrupted, automatic replay depends on the tool’s declared policy. A read-only search can be marked safe; a deployment or payment action should not be replayed blindly.
The framework also supports multiple concurrent conversations, forks at a particular transcript point, conversation-specific models and tools, queued steering messages, and background tasks. Its application state uses typed JSON documents stored with the conversation history so state changes and the actions that produced them can commit together.
Pi Durable is MIT-licensed and lives in the Pi monorepo. Earendil describes it as experimental and says APIs may change.
Why it matters
An agent that runs for ten minutes can rely on process memory and a person watching the terminal. An agent that runs for hours or days needs the same disciplines as any distributed job system: durable state, idempotency, cancellation, retries, versioning, and visibility.
The replay declaration is particularly important. Many agent demos treat every failed tool call as something to retry. That is safe for a deterministic read, but dangerous for actions that charge a card, send a message, merge code, create a cloud resource, or update a production record. Pi Durable makes the difference explicit at the tool boundary.
The framework’s exactly-once request identifier addresses a related edge case. If a client loses its connection after submitting work, it can repeat the submission with the same identifier and receive the original operation instead of creating a duplicate. Exactly-once behavior still depends on the complete application path; an external service without idempotency guarantees can remain a source of duplication.
Evidence
The announcement includes implementation examples rather than benchmark claims. Earendil says the source excluding tests is about 15,000 lines, and documents a small vacation-planning demonstration in which several searches run in parallel. The process is stopped before one search completes; after restart, only the work declared safe to replay is run again.
The design also keeps active memory bounded through background compaction. Older messages are summarized as the model approaches its context limit, while original records remain in storage. A conversation can reset into a fresh context with a handoff note and later retrieve earlier material through a separate search path.
Hot-swappable extensions provide another operational feature. Conversations store extension and tool names rather than serialized code. An in-flight call finishes under the implementation it started with, while the next call can use a replacement. That is useful for fixing a long-running system without discarding its work, but it increases the need to record which version produced each outcome.
Practical takeaway
Developers evaluating Pi Durable should begin with a deliberately failure-prone sandbox. Kill the process during model calls, safe reads, file writes, and externally visible actions. Confirm what resumes, what is marked interrupted, and what is never repeated automatically.
Define replay semantics conservatively. A tool is not safe merely because it usually succeeds. Reads that consume a queue, generate a signed URL, or advance a cursor can still mutate state. Writes may be safe only when they carry an idempotency key enforced by the destination.
Keep approval boundaries outside the model’s discretion. Durability should preserve a pending approval, not convert it into permission after restart. Store actor, timestamp, tool version, arguments, output, and external request identifiers for every consequential call.
Limitations
Pi Durable is an early framework, not a managed reliability guarantee. It uses a single process as the owner of a storage backend, so deployment topology and leader failover remain application concerns. Durable records also create privacy and retention obligations: a system that remembers everything must decide who can retrieve it and when it should be deleted.
Compaction can preserve the transcript while still changing the information visible to the model. Multi-client steering can create conflicting instructions. Hot-swapping code can make identical transcript states behave differently across time. None of these are fatal flaws, but each needs explicit product policy.
The release is useful because it frames agent reliability as systems engineering. The model is only one component. Recovery semantics, permissions, and state integrity determine whether a long-running agent is dependable.
Sources
> Want more like this?
Get the best AI insights delivered weekly.
By subscribing, you agree to our Privacy Policy. You can unsubscribe at any time.
> Related Articles
Context Language Models Let Agents Rewrite Their Own Working Memory—With New Failure Modes
A Meta–UW research project treats an agent's context as an editable file. The reported efficiency gains are notable, but the design makes context integrity a first-class security problem.
Long-Running AI Agents Need Three Memories: Working Context, Durable State, and an Audit Record
Context Language Models, Pi Durable, and OpenAI's long-horizon Codex guidance converge on one lesson: persistence is not one feature, and reliability requires separate memory layers.
Gemini 4 Argon Starts With Cyber Defenders, a 1M-Token Output Limit, and Phased Access
Google's newest frontier model is being released first to trusted cybersecurity teams. The notable story is not only benchmark scores, but the operational burden created by million-token trajectories.
Tags
> Stay in the loop
Weekly AI tools & insights.