Skip to main content
Durable by default, on every runtime. The in-process (local) runtime is durable via durable-go — no config needed, no external server. Temporal and Restate add horizontal scaling and a client/worker split on top of the same durability guarantees, plus token-level (not step-level) reconnect fidelity. See Runtime support for the one fidelity difference on local.
The problem: Agent runs invoke LLM calls, tools, and memory operations that take time — and any of these processes can crash mid-run. Workers or endpoints can restart, your subscriber process can disconnect, and prompts get lost. Re-running the entire agent from scratch wastes tokens and time, and loses partially completed work. The solution: The durable runtime records every step of the run as durable history — on every runtime, including local. If the executing process crashes, completed steps are not re-run — the run resumes exactly where it left off. If your client process crashes mid-run, call GetAgentRun (non-stream) or GetAgentStream (stream, optionally with WithOffset on Temporal/Restate) to reconnect without restarting the agent. The core reconnect API is the same everywhere; local just can’t seek to an arbitrary offset (see Runtime support). These guarantees are independent:

Keep the run alive across disconnects

The durable runtime already records history. Your call-site context decides whether the live run is still there when you reconnect:
  • Pass a long-lived ctx to Run / Stream (context.Background(), process-lifetime, or context.WithoutCancel(req.Context())) so disconnect or process shutdown does not cancel the durable run.
  • Cancel only Get / Events when the HTTP client leaves or the API is shutting down — that drops the subscriber, not the run.
  • Bound run length with WithTimeout (or a long-lived deadline on the Stream/Run ctx).
  • Avoid passing a request/shutdown ctx into Run / Stream if it will be cancelled when the handler exits — that stops the run before reconnect.
Full matrix: Timeouts & Modes — What each context does.

Server-side durability

The durable runtime records every step of the agent run — LLM calls, tool executions, approvals, memory operations — in durable history (a local durable-go journal, Temporal workflow history, or Restate journaled steps). When the executing process crashes or restarts:
  • In-flight steps are automatically retried on a recovered or different executor
  • Completed steps are not re-executed — the run fast-forwards through recorded history
  • The final result is identical regardless of how many interruptions occurred
This is automatic on every runtime — durable by default on local, no SDK changes needed. Temporal/Restate are additional opt-in choices, not what makes a run durable. The local journal is plaintext JSON by default (user prompt, LLM content, tool args and results). Do not put secrets or raw PII in prompts, tool I/O, or errors. For at-rest encryption, a journal MAC, or approval tokens that survive a restart, pass a caller-owned durable.Engine on local.LocalConfig.Engine — same pattern as Temporal Cloud / TLS via WithTemporalClient. See durable-go data privacy.

Step retry and streaming

When an LLM step retries (after a transient error or process crash), the LLM call restarts on the new attempt. The SDK detects this and signals any connected subscribers to discard token events from the failed attempt. Subscribers transparently receive clean token delivery from the new attempt — no duplicate or garbled output. On local, this shows up differently on reconnect: instead of clean live token delivery for an already-completed step, you get one coalesced step_replayed custom event per completed step (see Client-side stream recovery below).

Experiencing it

The Durable Agent (Local) example is the fastest way to feel this — zero infrastructure, kill and restart the same process, reconnect via GetAgentStream (no WithOffset — see Runtime support). The Durable Agent (Temporal) example is a hands-on Temporal lab for this behavior (split NewAgentWorker topology). Scenarios 3, 4, and 6 deliberately crash or stop workers mid-run and show:
  • Runs completing after the worker that started them is killed
  • Graceful and crash restarts with no data loss
  • Multiple workers sharing a task queue
Restate provides the same durability guarantees through journaled steps and an embedded endpoint — see Durable Agent (Restate) and Restate runtime. Restart the agent process (or another registered deployment) and Restate resumes from the journal.

Client-side stream recovery

If YOUR process — the one consuming the event stream — crashes or disconnects, the run continues on the server side but your subscriber loses the live connection. Use GetAgentStream to resubscribe and resume from exactly where you left off.
Persist runID from the handle immediately. Both Stream and Run return a handle whose ID() is available before events or Get complete. Persist that ID (and stream offsets) before consuming the channel so a mid-run crash can reconnect.

The five-step protocol

1. Persist runID before consuming events.
2. Track the offset of each received event. Events from a durable stream carry a monotonic offset. Persist the last seen offset before processing each event — this is your resume point.
Offset() returns (offset int64, ok bool). When ok is false the event was emitted client-side by the SDK (e.g. RUN_STARTED) and has no stream position — do not use it as a resume point. 3. On restart, call GetAgentStream then Events with the saved offset.
4. Skip already-processed events, then resume normally. The stream may redeliver from savedOffset. Skip events at or below that point before resuming normal handling.
5. Clear saved state on RUN_FINISHED or RUN_ERROR.

Reconnect requires a live run

GetAgentStream only works while the durable run is still executing. Once it completes, fails, or times out, the stream log is no longer available for replay. Calling GetAgentStream on a finished run returns ErrRunAlreadyCompleted — clear your saved state and continue from conversation/memory or start a new run.
Important: ErrRunAlreadyCompleted does not mean the agent failed. The durable runtime completed the run — the LLM responded, tools ran, everything finished. What is lost is only the streaming view of those events. If you have WithConversation configured, the final response is already stored in conversation history and a follow-up turn will have full context. Practical timing: reconnect only works while the run is still live. For short queries that complete in seconds, it may finish before you reconnect. After GetAgentRun / GetAgentStream, WithTimeout (if set) starts fresh — not the remaining time from the original Run / Stream. See Timeouts & Modes.

Approval events on reconnect

If an approval was already resolved while your subscriber was disconnected, the SDK filters it out automatically on reconnect — you will not be re-prompted for an approval that was already actioned.

Client-side run recovery

If YOUR process started a non-stream Run and then crashed or exited before Get returned, the durable run continues on the server. Persist runID from AgentRun.ID() immediately after Run, then call GetAgentRun on restart to wait for the result:
Same rule as stream reconnect: GetAgentRun only works while the durable run is still executing. Once it completes, fails, or times out, you get ErrRunAlreadyCompleted — clear saved state and continue from conversation/memory. Cancelling the GetAgentRun / Get context does not cancel the run — use AgentRun.Cancel. After reconnect, WithTimeout starts fresh. See Run and Timeouts & Modes.

API reference

ErrRunAlreadyCompleted — returned when the durable run is no longer live. Clear saved state and start a new turn/run. ErrRunNotFound / ErrStreamNotFound — returned when runID is unknown, or on the local runtime when durability is off (local.DurabilityOff()) so there is no journal to reconnect to. ErrStreamOffsetNotSupported — returned by the local runtime for any fromOffset > 0 (no per-token durable log — only step-level replay from the start).

Runtime support

Examples

Durable Agent (Local)

Zero-infrastructure lab: durable by default, kill and restart one process

Durable Agent (Temporal)

Temporal lab: crash workers, kill processes, observe durability

Durable Agent (Restate)

Restate lab: single process, kill mid-stream, reconnect

Reconnect

Temporal demo of the shared reconnect API (same calls work on Restate)

Distributed Execution

Split agent client and Temporal worker into separate processes

Run

GetAgentRun — reconnect a non-stream Run

Approvals

Human-in-the-loop tool approval and reconnect interaction

Streaming

RunID, event types, and the streaming API

Temporal Runtime

Temporal durability architecture and event delivery