WithErrorControl when you want the agent to keep going after a failure instead of stopping the run:
- Swap to a cheaper or larger-window model after the primary LLM fails
- Give the same conversation a few more rounds when
MaxIterationsis hit - Stop calling a tool that is stuck in a failure loop
ExecutionConfig retries. They are not lifecycle WithHooks — those inspect or mutate a successful call.
Zero config keeps today’s defaults: LLM failure aborts the run, max iterations does one last no-tools LLM call. An unknown tool name continues the run (it is never executed).
Configure
Register extra models withWithNamedLLMClients, then point error control at a name:
FallbackLLMClient is a name in WithNamedLLMClients, not a second client on the config. NewAgent rejects an unknown name. If the hook asks for fallback and the name is empty or missing at run time, the run aborts.
A nil hook field uses the default. An invalid action logs a warning and uses the default.
Tool execute failures already skip-and-continue. Use the breaker when the same tool keeps failing.
Fallback after LLM failure
Built-in providers wrap Generate / stream errors asinterfaces.LLMError. Branch with errors.As — do not parse err.Error(). Custom clients that do not wrap stay unknown.
Retries from
WithLLMExecutionConfig run first. Then OnLLMFailure can abort or switch models for the rest of that run. The shared primary client is not mutated. Fallback uses the same WithLLMSampling and WithResponseFormat as the agent.
On in-process and Restate, the fallback client is retried with the same LLM execution policy as the primary. On Temporal, fallback is one Generate/stream on the last LLM activity attempt. If that call fails, the activity fails and Temporal may retry from the primary (OnLLMFailure can run again).
context_exceeded can FallbackModel (a model with a larger window) or Abort.
Extra iterations
When the last iteration still returns tool calls, the default is one LLM call with tools skipped andFinishReasonMaxIterations.
Return ExtendIterations with ExtraIterations to keep the same messages and continue. The SDK:
- Grants this once per run
- Requires
ExtraIterations≥ 1 and caps it at the originalMaxIterations(unset max uses 10) - Ignores a second grant and does the final no-tools call
Circuit breaker
Optional config — no callback. Tracks tool name (not MCP server). Only tools that count toward tool telemetry participate (native and MCP tools; not unknown, sub-agent, A2A, or retriever tools).
On trip the SDK skips authorize/execute, appends a tool-role warning, and logs. In-process and Restate increment
agent.tool.call.circuit_open. Temporal logs the skip and does not increment that counter. The run continues.
A later success clears the failure streak; it does not reopen a tripped tool. Both detectors off means the tracker is not created.
State is kept across Temporal continue-as-new and Restate / in-process durable journals, so a restart does not reset the trip.
Unknown tool names
If the model names a tool that is not registered, the run continues. The SDK appends a synthetic tool-role message and logs a warning. That name is never authorized or executed. In-process and Restate incrementagent.tool.call.unknown. Temporal logs and continues; it does not increment that counter.
Temporal client and worker
When the starter and worker are separate processes, pass the sameWithNamedLLMClients and WithErrorControl on both. The fingerprint includes:
- Whether each error-control hook is set (not the function body)
FallbackLLMClientname- Each named client’s model and provider
- Breaker thresholds
Example
Error Control
Fallback on LLM error, extra iteration, breaker skip
Related
Hooks
Lifecycle middleware — inspect, mutate, or abort a call
Execution Config
Timeouts and retries run before error-control hooks
Budget
Token and cost caps — separate from max iterations
Distributed execution
Match named LLMs and error-control on client and worker