Restate Durable Agents: Duplicate Tool Calls Tested (2026)
October 2, 2026

TL;DR: A Restate agent that crashes between a tool's side effect and the moment ctx.run records its result can repeat that side effect. In my tests, killing the worker at exactly that point made a mock payment API receive two charges in 3 of 3 trials.1
Kills before the effect, or after ctx.run had returned, gave one charge every time. Passing a key from ctx.rand.uuidv4() to the payment call made the second request a no-op in 3 of 3 trials, as long as the receiver honors the key.
Restate's AI docs say "Tool side effects are not duplicated (no double bookings, no duplicate emails)."2 That holds once a step is journaled. This post shows where the gap is, how to measure it with a SIGKILL, and what to add to your tools.
Restate durable agents: the short answer
Restate is a durable execution engine: your agent runs as a handler, and each ctx.run step is recorded in a journal so a restarted handler can skip finished work.23 It protects completed steps. It cannot protect a step whose effect happened but whose result was never recorded.
For tools that spend money or send messages, the fix is an idempotency key. ctx.rand.uuidv4() returns the same value on a retry, so the downstream API can recognize the repeat.3
What you'll learn
- What Restate's docs promise for agent tool calls, and the open issue that challenges it
- How to build a crash-testing rig with a mock LLM, a mock payment API and SIGKILL
- The result of nine scenarios, from a plain loop to a keyed Restate agent
- Why an effect outside
ctx.runand an unjournaled LLM call also repeat - A short checklist for making agent tools safe to retry
- What these tests did not cover
Why Restate durable agents are in the news
Restate announced a $20 million Series A on September 30, 2026, led by Singular, with Redpoint Ventures and Capital One Ventures also investing.4 Trade press covered it on October 1.4
The pitch is crash recovery for agents. Restate's announcement says a single agent can perform hundreds of model calls, tool calls, waits and retries, so a crash halfway through is expensive if the run has to start over.4
Restate's docs describe the mechanism. The Durable Agents page says: "Completed steps are replayed from the journal (no re-execution)."2
It then lists two consequences:2
- "LLM calls are not repeated (saving cost and time)"
- "Tool side effects are not duplicated (no double bookings, no duplicate emails)"
For the hand-written TypeScript loop, the same page says: "Side effects are executed exactly once. On recovery, the result is replayed."2 That page was last changed on June 23, 2026.2
The open issue: a window between effect and journal
On September 24, 2026, a user opened issue #410 on Restate's docs repository, titled: Durable agents: "Tool side effects are not duplicated", but a ctx.run action re-executes when the process dies after the action and before its result is journaled.5
The reporter used a Pydantic AI agent in Python and a receiver that cannot deduplicate. They killed the worker with SIGKILL after the receiver applied the request, and report that it was applied twice in 30 of 30 trials.5
They write that "In that interval, the only record of the action is the effect it had."5 As of October 2, 2026, the issue is open with no comments.5
Restate's own architecture page is consistent with this. It says the moment a step's journal entry "is replicated to quorum defines 'the step happened.' From then on, the step will be recovered on retries and won't be re-executed."6
A comment in a database guide on the same site, about an update that is not wrapped in ctx.run, describes "a very small window time where the query gets re-executed after success."6
So the guarantee is real, and it starts at the journal write. I searched the AI page for "window", "at-least-once" and "idempotency" and found none of them. I wanted to see it for myself, in TypeScript, with the current release.
The test rig: mock LLM, mock payments, SIGKILL
I used @restatedev/restate-server 1.7.13 (published October 1, 2026) and @restatedev/restate-sdk 1.17.2 (published September 21, 2026).1 There is no real model in the loop. A scripted mock "LLM" asks for a charge, then a send_email, then answers.
One small server plays both the mock model and the mock payment and email APIs. It logs every request to events.jsonl, and it deduplicates only when a request carries an Idempotency-Key header.
mkdir restate-agent-test && cd restate-agent-test
npm init -y
npm i @restatedev/restate-server@1.7.13 @restatedev/restate-sdk@1.17.2 tsx typescript @types/node
Save this as mocks.ts:
// Mock LLM + mock "payments/email" receiver on one port. Every hit is appended to events.jsonl.
import http from "node:http";
import fs from "node:fs";
const LOG = process.env.EVENTS_LOG!;
const seen = new Map<string, unknown>();
const log = (e: object) => fs.appendFileSync(LOG, JSON.stringify({ ...e, at: Date.now() }) + "\n");
http
.createServer((req, res) => {
let body = "";
req.on("data", (c) => (body += c));
req.on("end", () => {
const input = body ? JSON.parse(body) : {};
const send = (o: object) => {
res.setHeader("content-type", "application/json");
res.end(JSON.stringify(o));
};
if (req.url === "/llm") {
// scripted "model": 0 tool results -> charge, 1 -> send_email, 2 -> final answer
const n = input.toolResults as number;
log({ t: "llm", toolResults: n });
if (n === 0) return send({ type: "tool", tool: { name: "charge", args: { amount: 49 } } });
if (n === 1) return send({ type: "tool", tool: { name: "send_email", args: { to: "a@example.com" } } });
return send({ type: "final", text: "Charged $49 and emailed the receipt." });
}
if (req.url === "/charge" || req.url === "/send_email") {
const kind = req.url.slice(1);
const key = req.headers["idempotency-key"] as string | undefined;
if (key && seen.has(key)) {
log({ t: kind, key, applied: false });
return send(seen.get(key) as object);
}
const out = { id: `${kind}_${Math.random().toString(36).slice(2, 8)}` };
if (key) seen.set(key, out);
log({ t: kind, key: key ?? null, applied: true });
return send(out);
}
res.statusCode = 404;
res.end();
});
})
.listen(7001, "127.0.0.1");
Now the agent, agent.ts. It is the same shape as the TypeScript loop in Restate's docs: one ctx.run per model call and one per tool call.2 Two environment variables control the experiment: MODE picks how tools are wrapped, and CRASH picks where the process kills itself, once.
import * as restate from "@restatedev/restate-sdk";
import fs from "node:fs";
const MODE = process.env.MODE ?? "journaled"; // journaled | key | bare
const CRASH = process.env.CRASH ?? "none"; // none | before_effect | after_effect | after_tool_returned | llm_after_call
const MARKER = process.env.CRASH_MARKER!;
const BASE = "http://127.0.0.1:7001";
function crashOnce(point: string) {
if (CRASH !== point || fs.existsSync(MARKER)) return;
fs.writeFileSync(MARKER, point);
process.kill(process.pid, "SIGKILL"); // no cleanup, no goodbye
}
async function post(path: string, body: object, key?: string) {
const r = await fetch(BASE + path, {
method: "POST",
headers: { "content-type": "application/json", ...(key ? { "idempotency-key": key } : {}) },
body: JSON.stringify(body),
});
return (await r.json()) as any;
}
async function callTool(name: string, args: object, key?: string) {
if (name === "charge") crashOnce("before_effect");
const out = await post("/" + name, args, key);
if (name === "charge") crashOnce("after_effect");
return out;
}
const agent = restate.service({
name: "agent",
handlers: {
run: async (ctx: restate.Context, input: { task: string }) => {
let toolResults = 0;
for (let turn = 0; turn < 6; turn++) {
const decision = await ctx.run(`llm turn ${turn}`, async () => {
const d = await post("/llm", { task: input.task, toolResults });
if (turn === 0) crashOnce("llm_after_call");
return d;
});
if (decision.type === "final") return decision.text as string;
const { name, args } = decision.tool;
if (MODE === "bare") {
await callTool(name, args); // NOT wrapped in ctx.run: re-executes on every replay
} else {
const key = MODE === "key" ? ctx.rand.uuidv4() : undefined;
await ctx.run(`tool ${name}`, () => callTool(name, args, key));
}
if (name === "charge") crashOnce("after_tool_returned");
toolResults++;
}
return "gave up";
},
},
});
restate.serve({ services: [agent], port: 9080 });
The marker file makes each crash fire once, so the restarted worker finishes the run. For the baseline with no durable runtime, naive.ts is the same loop as plain code:
// The same loop with no durable runtime: a crash means "start over".
import fs from "node:fs";
const CRASH = process.env.CRASH ?? "none";
const MARKER = process.env.CRASH_MARKER!;
const BASE = "http://127.0.0.1:7001";
const post = async (p: string, b: object) =>
(await (await fetch(BASE + p, { method: "POST", headers: { "content-type": "application/json" }, body: JSON.stringify(b) })).json()) as any;
async function loop(task: string) {
let toolResults = 0;
for (let turn = 0; turn < 6; turn++) {
const d = await post("/llm", { task, toolResults });
if (d.type === "final") return d.text;
await post("/" + d.tool.name, d.tool.args);
if (d.tool.name === "charge" && CRASH === "after_tool_returned" && !fs.existsSync(MARKER)) {
fs.writeFileSync(MARKER, "x");
process.kill(process.pid, "SIGKILL");
}
toolResults++;
}
}
loop("refund order 1042").then((t) => { console.log("RESULT " + t); });
How to run one crash scenario
Use four terminals in the project folder. In the first, start the server. My sandbox had no IPv6, so I bound everything to the loopback address; on a normal machine plain npx restate-server works.
RESTATE_BIND_ADDRESS=127.0.0.1:5122 RESTATE_INGRESS__BIND_ADDRESS=127.0.0.1:8080 \
RESTATE_ADMIN__BIND_ADDRESS=127.0.0.1:9070 npx restate-server --no-logo
In the second, start the mock receiver. In the third, run the agent under a loop that restarts it, the way an orchestrator would after a crash:
# terminal 2
EVENTS_LOG=events.jsonl npx tsx mocks.ts
# terminal 3
export MODE=journaled CRASH=after_effect CRASH_MARKER=/tmp/crashed
rm -f "$CRASH_MARKER"
while true; do npx tsx agent.ts; echo "worker exited, restarting"; sleep 0.4; done
In the fourth, register the service, call the agent, and count what the receiver saw:
curl -s 127.0.0.1:9070/deployments -H 'content-type: application/json' -d '{"uri":"http://127.0.0.1:9080"}' > /dev/null
curl -s 127.0.0.1:8080/agent/run -H 'content-type: application/json' -d '{"task":"refund order 1042"}'
jq -s '{llm: map(select(.t=="llm"))|length, charges_received: map(select(.t=="charge"))|length, charges_applied: map(select(.t=="charge" and .applied))|length}' events.jsonl
Restate's own quickstart uses the same ports: the admin UI on 9070, the service on 9080 and the ingress on 8080.2 To try another row of the table below, stop terminal 3, change MODE and CRASH, delete events.jsonl and the marker, and start again.
Results: nine scenarios, three trials each
I ran the whole matrix three times, with a fresh Restate data folder and fresh mock servers for every scenario.1 Every row gave the same counts in all three trials.

Real output of the test harness on October 2, 2026, with the result and restart columns dropped. "Charges received" counts requests that reached the mock payment API. Red box: duplicate charge. Green: the key fix. Orange: other repeats.1
| Scenario | MODE / CRASH | LLM calls | Charges received | Charges applied |
|---|---|---|---|---|
| S0 Plain loop, no crash | n/a | 3 | 1 | 1 |
| S1 Plain loop, kill after the charge | after_tool_returned | 4 | 2 | 2 |
| S2 Restate, no crash | journaled / none | 3 | 1 | 1 |
| S3 Restate, kill before the charge request | journaled / before_effect | 3 | 1 | 1 |
| S4 Restate, kill after the charge, before the journal | journaled / after_effect | 3 | 2 | 2 |
S5 Restate, kill after ctx.run returned | journaled / after_tool_returned | 3 | 1 | 1 |
| S6 Same kill as S4, with an idempotency key | key / after_effect | 3 | 2 | 1 |
S7 Restate, charge outside ctx.run, kill after it | bare / after_tool_returned | 3 | 2 | 2 |
| S8 Restate, kill after LLM call 0 answered | journaled / llm_after_call | 4 | 1 | 1 |
Each Restate row also applied exactly one email, and every run returned the final answer. In every Restate crash row the invocation completed on its own once my restart loop brought the worker back.1
What the results show
The plain loop starts over. After a kill following the charge, the restarted loop asked the model again, charged again and finished with four model calls and two charges. This is the problem durable execution exists to solve.
Restate resumed from the journal. With the kill before the charge request, or after ctx.run had returned, the restarted worker did not repeat any finished step. Model calls stayed at three and the charge was applied once.
The gap is between the effect and the journal. In the kill-after-effect row, the charge reached the receiver, the worker died before ctx.run could record the result, and the restarted handler ran the action again. The payment API received two charges and applied both.
That matches issue #410, found there with a Python agent and 30 trials.5 My run reproduces it with the TypeScript SDK and a hand-written loop, in 3 of 3 trials, plus 5 of 5 earlier trials while I was building the rig.1
The three crash windows
Here is one tool call, with the kill points marked.

The three kill points in the matrix. Only window B duplicates the effect.
Window A is before the request leaves the worker. Nothing happened, so the retry is the first attempt. Window C is after await ctx.run(...) has returned. The step is journaled, so the replay skips it.
Window B is the one in between. The only evidence of the charge is the charge itself, which is the issue reporter's point.5 How wide it is in real life depends on your network, your host and how the worker dies. The issue reports a median of about 16 ms on its setup.5
Restate's architecture page puts the line in the same place: a step has "happened" once its journal entry is replicated to quorum, and not before.6 My tests did not measure the window's width, only that it exists.
The fix: an idempotency key from ctx.rand.uuidv4()
Restate's docs say its random helpers are "seeded by the invocation ID" and return "the same result on retries." They show ctx.rand.uuidv4() for "stable UUIDs for things like idempotency keys."3
That is exactly what a retried tool call needs. In the keyed row, the agent generates the key outside ctx.run, then sends it as an Idempotency-Key header:
const key = MODE === "key" ? ctx.rand.uuidv4() : undefined;
await ctx.run(`tool ${name}`, () => callTool(name, args, key));
With the same kill as before, the receiver saw two charge requests with the same key. It applied the first and answered the second from its record, so one charge was applied in 3 of 3 trials.1
The restarted handler reused the key without any storage on my side. That is the property you want: the replay computes the same key, so the downstream service can tell the repeat from a new charge.
There is one catch. The key only helps if the service on the other end honors it. My mock receiver does; a real payment, email or ticketing API may or may not, so check its documentation before you rely on this.
Two more traps: bare effects and unjournaled model calls
Effects outside ctx.run repeat on every replay. In the bare row, the charge was a plain fetch in the handler. After the kill, Restate replayed the handler from the top and the charge ran again. Journaled steps are skipped, but code outside them runs again.
An unjournaled LLM call is paid twice. In the last row I killed the worker after the first model call returned but before ctx.run recorded it. The restarted handler called the model again, so the receiver logged four model calls instead of three.
That is the same window as the charge, applied to the model. The docs say "LLM calls are not repeated"; for calls that were recorded, that is what I saw.21
My mock model is scripted, so it gave the same answer the second time. A real model may answer differently on the repeated call. I did not test that, and I did not use a real model.
If you are counting the cost of this, my post on AI agent cost control and session caps covers the spend side. One repeated model call is small, but it is not zero.
A short checklist for retry-safe agent tools
- Wrap every side effect in
ctx.run. Anything outside it can run again on replay. - Give every non-idempotent tool an idempotency key from
ctx.rand.uuidv4(), generated outside thectx.runclosure. - Check that the downstream API honors the key. If it does not, add a lookup of your own, such as a record keyed by the same value, before you repeat the action.
- Budget for a repeated model call after a crash. In my tests a crash cost one.
- Test with a real SIGKILL at each point. The rig above is 123 lines of TypeScript.
- Watch the retry settings. Restate retries failed
ctx.runsteps and can limit them withmaxRetryAttemptsand related options.3
For retries, timeouts and routing in a different framework, I ran a similar set of tests on Google ADK 2.0 Workflow. For the structural side of agent reliability, see AI agent reliability and verification loops.
Limits of these tests
These are single-node tests on one Linux sandbox with a scripted model, a mock payment API and SIGKILL as the only fault. I tested Restate server 1.7.13 and TypeScript SDK 1.17.2 only.
I did not test the Python SDK, the Vercel AI SDK integration, a multi-node cluster, a frozen worker, or a crash of the Restate server itself. The issue reporter describes a frozen-worker case; I did not run it.5
I also did not test other durable execution engines, so I cannot say how they behave in the same window.
Bottom line
Restate does what it says for finished steps: in my tests, nothing that had been journaled ran twice. The gap is a step that has acted but has not been recorded yet.
If your agent charges cards, sends email or files tickets, add an idempotency key to those tools and kill your own worker to prove it works. For read-only tools and model calls, the gap costs time and tokens, not money.
Footnotes
-
Author's measurement, October 2, 2026: Node 22.22.0 on Linux x86_64;
@restatedev/restate-server1.7.13 (npm publish date 2026-10-01) and@restatedev/restate-sdk1.17.2 (npm publish date 2026-09-21), dates fromnpm view <package> time. A fresh Restate server data folder and fresh mock servers per scenario; the full nine-scenario matrix run three times, with identical counts in every trial. The two key rows (kill after the effect, with and without a key) were also run five times each earlier while building the rig, with the same counts. The code in this post was extracted from the finished post and re-run in a fresh folder. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 -
Restate docs, "Durable Agents" page (docs.restate.dev/ai/patterns/durable-agents); quotations checked against its source,
docs/ai/patterns/durable-agents.mdxon the restatedev/docs-restate GitHub main branch, read 2026-10-02. The last commit to that file was dated June 23, 2026. Port numbers: Agent Quickstart, fetched 2026-10-02. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 -
Restate docs, "Durable Steps" (TypeScript), quotations checked against
docs/develop/ts/durable-steps.mdxon GitHub main, read 2026-10-02. ↩ ↩2 ↩3 ↩4 ↩5 -
Restate raises $20M Series A — Restate blog, dated September 30, 2026 as displayed, fetched 2026-10-02; investor list and date also checked against The AI Insider's October 1, 2026 report, fetched 2026-10-02. ↩ ↩2 ↩3
-
Issue #410, restatedev/docs-restate, opened 2026-09-24 by keshav9926; state, comment count and text read through the GitHub API on 2026-10-02 (open, 0 comments). The 30-of-30 result and the roughly 16 ms median are the reporter's own measurements on Restate server 1.7.10 and Python SDK 1.0.5; I did not reproduce the Python setup. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10
-
Restate docs, "Architecture" (step "Durable step commit (ctx.run)") and the database guide's code comment in
docs/guides/databases.mdx, both read on the docs-restate GitHub main branch on 2026-10-02. The quoted database-guide sentence spans two comment lines in the source; I removed the comment markers. ↩ ↩2 ↩3 ↩4



