Google ADK 2.0 Workflow: Retries and Timeouts Tested (2026)
September 29, 2026

TL;DR: I ran the graph Workflow engine in google-adk 2.10.0 with plain Python function nodes (no LLM calls) on Python 3.10, 3.11, 3.12 and 3.13, and three results are worth knowing. First, a bare RetryConfig() makes up to 5 attempts, but its default jitter is a full 100%, so each wait is randomized: in one batch of 20 runs on Python 3.11.15, the time from the start of the first attempt to the start of the fifth ranged from 6.6 s to 25.4 s, not a fixed 15 s. Second, a node's timeout cannot interrupt blocking code: with timeout=0.3, every node that blocked for 1.0 s ran the full second, and whether the run then completed or failed with NodeTimeoutError depended on what the node did next and on the Python version (a node that returned a value straight after blocking completed on 3.10 and 3.11 and failed on 3.12 and 3.13). Third, a route value that matches none of a router's routed edges raises no error: when all of the router's edges were routed, the branch ended with one warning in ADK's log, and when it also had an unrouted edge, that edge's node ran and nothing was logged.
Google ADK 2.0 Workflow: the short answer
A Workflow in Google's Agent Development Kit (ADK) 2.0 is a graph of nodes joined by an edges list. Nodes can be Python functions, agents or nested workflows, and the engine follows the graph you define instead of leaving the sequencing to a prompt. You can attach a RetryConfig and a timeout to a node, choose a branch with Event(route=...), fan out and join, and pause for a human with RequestInput. The checks below were run on the Python package, google-adk 2.10.0. Dates and quotations come from the linked sources, and statements about ADK's internals are marked as reads of the package source. I did not test the Go or TypeScript SDKs.
What you'll learn
- What ADK 2.0's graph
Workflowis and which release line ships it - What the default retry policy does, and how its default jitter changes the timing
- How
RetryConfig(exceptions=...)scopes retries, and what happens with noretry_config - Why a node
timeoutcan't interrupt blocking code, and how the result changes with what the node does next and with the Python version - What happens when a route matches nothing, and how
DEFAULT_ROUTEcatches it - Why fan-out only runs in parallel when workers yield to the event loop
- How a human-approval pause and resume looks in the event stream, and which replies fail validation
- A production checklist built from these results, and a script that reproduces every measurement
What is the ADK 2.0 graph Workflow?
ADK Python 2.0 reached general availability on May 19, 2026,1 ADK Go 2.0 followed on June 30, 2026, and ADK TypeScript 2.0 on August 21, 2026.2 The documentation lists three main features: graph-based workflows, dynamic workflows, and collaborative workflows built from coordinator agents and multiple sub-agents.2 The Python package requires Python 3.10 or later, according to its PyPI metadata and the Python quickstart.34
Google's "Why we built ADK 2.0" post, published July 1, 2026, states the motivation directly: "Large language models are frequently tasked with execution orchestration—handling tasks like routing, scheduling, and error handling that traditional code already excels at. While they can get the job done, they are slow, expensive, and exhibit variance compared to a workflow or deterministic code."5 In other words, the Workflow engine moves routing and error handling out of the prompt and into code, which is what the rest of this post tests.
The version I tested, google-adk 2.10.0, is listed on PyPI as released on September 25, 2026.3 The ADK 2.0 documentation also warns about a migration trap that affects retries and human pauses. If you migrate a tool and leave a broad except Exception: block inside it, "this code masks the failure from the framework, permanently disabling the new 2.0 automatic retry mechanisms for that step," and catching BaseException "inadvertently traps NodeInterruptedError, which breaks the framework's ability to pause the workflow for Human-in-the-Loop (HITL) input."2
How I tested it
Every check runs a real Workflow through ADK's InMemoryRunner with plain Python function nodes (plus one JoinNode), so nothing depends on model latency or API keys. I installed google-adk 2.10.0 (with google-genai 2.25.0) into four virtual environments, Python 3.10.20, 3.11.15, 3.12.3 and 3.13.13, and ran the same script in each. Durations use time.monotonic(), one warm-up workflow runs first so one-time import costs stay out of the timings, and ADK's log records at WARNING and above are captured, because that is where one important behavior turned out to live. The full script is in the "Reproduce it" section below. This is its harness:
log_lines = []
class Capture(logging.Handler):
def emit(self, record):
if record.levelno >= logging.WARNING:
log_lines.append(record.getMessage())
logging.getLogger("google_adk").addHandler(Capture())
def user_msg(text):
return types.Content(role="user", parts=[types.Part(text=text)])
async def run(wf, message=None, runner=None, session=None):
"""Run a workflow once. Returns (runner, session, events, exception_or_None)."""
runner = runner or InMemoryRunner(agent=wf, app_name="demo")
session = session or await runner.session_service.create_session(app_name="demo", user_id="u")
events, error = [], None
try:
async for ev in runner.run_async(user_id="u", session_id=session.id,
new_message=message or user_msg("go")):
events.append(ev)
except Exception as exc:
error = exc
return runner, session, events, error
def status(error):
return type(error).__name__ if error else "completed"
This covers function nodes and the runtime around them. I did not test LLM-agent nodes, tool nodes, dynamic workflows, persistent session backends, deployed runners, or resuming a workflow after a crash, after a restart or mid-retry; the only resume I tested is the in-process reply to a human-input pause described below.
What is the default retry behavior of an ADK node?
A node without a retry_config is not retried on its own, although it can run again when a workflow that contains it, nested or top-level, is retried (see the next section). A node with a bare RetryConfig() makes up to 5 attempts, waits 1, 2, 4 and 8 seconds on average between them, and re-raises the last exception. The waits are random, because the default jitter is 1.0.
The defaults come from the package source. RetryConfig documents max_attempts as defaulting to 5 including the original request, initial_delay to 1.0 second, max_delay to 60.0 seconds, backoff_factor to 2.0 and jitter to 1.0.6 The first nominal wait is initial_delay, and each later one is multiplied by backoff_factor. When jitter is above zero, the delay function first caps the nominal wait at max_delay / (1 + jitter), then adds a random offset drawn uniformly between minus and plus jitter × wait, so a wait never exceeds max_delay.6 With the defaults and 5 attempts, every wait lands between 0 and twice its nominal 1, 2, 4 or 8 seconds, so the four waits can add up to 30 s.
My check uses a node that always raises ConnectionError:
stamps = []
def always_fail(node_input: str):
stamps.append(time.monotonic())
raise ConnectionError("x")
node = FunctionNode(func=always_fail, name="f", retry_config=RetryConfig(jitter=0.0))
With jitter=0.0, the gaps between the starts of the five attempts were 1.0, 2.01, 4.01 and 8.05 s on Python 3.11.15 and within 0.05 s of the nominal 1, 2, 4 and 8 seconds on every interpreter (the script rounds gaps to 0.01 s), and the fifth failure reached the caller as a ConnectionError. With the default RetryConfig(), I ran 20 copies of the same node at once. All 20 made 5 attempts, and the time from the start of the first attempt to the start of the fifth ranged from 6.6 s to 25.4 s (mean 14.3 s) on Python 3.11.15, not a fixed 15 s. The gap before the second attempt varied from 0.18 s to 2.00 s and the gap before the fifth from 1.32 s to 15.85 s. Each gap is the wait plus a little overhead for the failed attempt and ADK's bookkeeping, which came to 0.05 s or less per gap in the jitter-free runs. A second batch of 20 on the same interpreter gave 6.3 s to 26.2 s (mean 13.3 s), so treat any single batch as one draw. Across the four interpreters, the first batches ranged from 4.4 s to 26.1 s, under the 30 s ceiling.

Three practical consequences follow. First, budget for the ceiling, not the mean: under the defaults, a node that fails every time can spend up to 30 s waiting between attempts, on top of the time the five attempts themselves take. Second, each failed attempt of a function node adds an error event to the run (5 events for 5 attempts in my check), because the node runner enqueues one before deciding whether to retry.6 That holds even when the retry eventually succeeds: a node that failed twice and then succeeded made 3 attempts, completed, and still left 2 error events in the run. Third, every local retry logs a WARNING containing "retry count is not persisted across resuming" (4 warnings for the 4 retries above; the call is in _node_runner.py).6 I did not test resuming a workflow mid-retry, so I can't say what a resumed node does with its attempt counter. Treat max_attempts as a budget for one run of the node, not a durable one: when a workflow around the node has its own retry_config, each workflow attempt runs the failed node again with a fresh count, so the two budgets multiply.6
How do you scope retries to specific exceptions?
Pass exceptions=[...] to RetryConfig with exception classes or class names. ADK converts classes to their names and retries an exception if its class, or any class it inherits from other than object, has one of the listed names.6 That covers subclasses, so exceptions=[ConnectionError] also covers ConnectionResetError. It also covers unrelated classes that happen to share a name: requests.exceptions.ConnectionError, which inherits from OSError rather than from the built-in ConnectionError, was retried too.6 I checked each case:
| What the node raised | retry_config | Calls made | Outcome |
|---|---|---|---|
ValueError | exceptions=[ConnectionError], max_attempts=4 | 1 | ValueError raised |
ConnectionResetError | exceptions=[ConnectionError], max_attempts=4 | 4 | ConnectionResetError raised |
ConnectionResetError | exceptions=["ConnectionError"] (a string), max_attempts=4 | 4 | ConnectionResetError raised |
requests.exceptions.ConnectionError | exceptions=[ConnectionError], max_attempts=4 | 4 | requests.exceptions.ConnectionError raised |
KeyError | exceptions unset, max_attempts=4 | 4 | KeyError raised |
ValueError | none | 1 | ValueError raised |
Retries are opt-in: with no retry_config on the node or on any workflow that contains it, including the top-level Workflow you pass to the runner, a failure is not retried. I found no workflow-level default for a workflow's children, since the Workflow source contains no retry logic and retry_config is a per-node field. A Workflow is itself a node, though, and according to the retry_config docstring, a retry_config on it retries that workflow as a whole: children that already produced an output or a state change are replayed rather than run again, and a child that produced neither (typically the one that failed) runs again, even without a retry_config of its own.6 Leaving exceptions unset retried every exception type I raised, including timeouts and a programming error (KeyError), so list the transient exception types you actually expect. One side effect of listing them: a timeout raises NodeTimeoutError, which inherits directly from Exception and not from Python's built-in TimeoutError, so a list that names neither NodeTimeoutError nor one of its base classes (Exception or BaseException) switches off retries for timeouts, as the next section shows.6
Why can't a node timeout interrupt blocking code?
Because ADK enforces timeout with asyncio.wait_for, and blocking code never hands control back to the event loop, nothing can cancel it midway. In _node_runner.py, a node with a timeout runs inside asyncio.wait_for(...), and the resulting asyncio.TimeoutError becomes NodeTimeoutError.6 In _function_node.py, a plain (non-generator) def function is called inline (result = self._func(**kwargs)), so time.sleep(1.0) holds the event loop for the full second, and so does a blocking call inside an async def.6 The package's own timeout docstring says a node that does not finish in time "is cancelled and treated as a failure," and in my tests that happened on time only for nodes that were waiting in an await when the deadline passed.6
I gave a node a body that takes 1.0 s and a timeout of 0.3 s, in seven styles, on four interpreters (time.sleep stands in for the blocking call, and the rows differ in what the node does around it):
| Node body (takes 1.0 s) | Python 3.10 | Python 3.11 | Python 3.12 | Python 3.13 |
|---|---|---|---|---|
def + time.sleep(1.0) (sync_block) | completed, 1.01 s | completed, 1.01 s | NodeTimeoutError, 1.01 s | NodeTimeoutError, 1.01 s |
async def + await asyncio.sleep(1.0) (async_wait) | NodeTimeoutError, 0.31 s | NodeTimeoutError, 0.31 s | NodeTimeoutError, 0.31 s | NodeTimeoutError, 0.31 s |
async def + await asyncio.to_thread(time.sleep, 1.0) (async_thread) | NodeTimeoutError, 0.31 s | NodeTimeoutError, 0.31 s | NodeTimeoutError, 0.31 s | NodeTimeoutError, 0.31 s |
async def + blocking time.sleep(1.0) (async_block) | completed, 1.01 s | completed, 1.01 s | NodeTimeoutError, 1.01 s | NodeTimeoutError, 1.01 s |
async def + blocking time.sleep(1.0), then await asyncio.sleep(0) (block_then_await) | NodeTimeoutError, 1.01 s | NodeTimeoutError, 1.01 s | NodeTimeoutError, 1.01 s | NodeTimeoutError, 1.01 s |
async def + blocking time.sleep(1.0), then await a coroutine that returns without suspending (block_then_noop) | completed, 1.04 s | completed, 1.01 s | NodeTimeoutError, 1.01 s | NodeTimeoutError, 1.01 s |
def + time.sleep(1.0), returns nothing (block_no_output) | completed, 1.01 s | completed, 1.01 s | completed, 1.01 s | completed, 1.01 s |

Nothing interrupted the blocking bodies: all five blocking variants ran for the full second on every interpreter (1.01 to 1.04 s across these runs), and in none of these runs did the timeout fire at 0.3 s for a blocking body. What happened once the call returned depended on what the node did next and on the interpreter. A body that returned a value straight after blocking (sync_block, async_block) completed on Python 3.10 and 3.11 and failed with NodeTimeoutError on 3.12 and 3.13, about 1.0 s in. A body that then yielded to the event loop (block_then_await, which awaits asyncio.sleep(0)) failed on all four. Awaiting is not the same as yielding: block_then_noop, which awaits a coroutine that returns without suspending, had the same outcome as async_block on every interpreter. A body that returned nothing and wrote no session state (block_no_output) completed on all four.
That pattern follows from when the event loop gets control back, which is only when the node's code suspends or finishes. When a node emits an output event, ADK hands the event to the runner and then waits for the runner to process it (await processed.wait() in invocation_context.py); that wait is the first point after the blocking call where the loop, and with it the overdue timer, can run. A body that suspends on its own after blocking reaches such a point sooner. A body that returns nothing and writes no state emits no event at all, so it never suspends; a None return that does write ctx.state still emits a state event.6 What the timer does then differs by interpreter, and the results are consistent with each version's asyncio.wait_for.7 In 3.12 and 3.13, wait_for uses asyncio.timeout(), which cancels the task once the loop runs the overdue timer, so every blocking body that suspended afterwards failed and the one that never suspended completed. In 3.10 and 3.11, wait_for runs the node as a separate task and, when it wakes, returns the result if that task has already finished (if fut.done(): return fut.result()), so the outcome depends on the order in which the loop runs the callbacks queued behind the blocking call. I did not trace that ordering step by step, so treat this as the likely explanation and the table as the measured result.
The fix is to keep the event loop free: use an async client for network calls, or wrap a blocking call in asyncio.to_thread, which timed out at 0.31 s on every interpreter. But to_thread only abandons the wait; it does not stop the thread. With timeout=0.2 and RetryConfig(max_attempts=3, initial_delay=0.05, jitter=0.0) around a one-second blocking call, all 3 attempts had started and none had finished when run_async raised NodeTimeoutError, and all 3 had finished 1.5 s later (identical on all four Python versions). So each retry of a to_thread node can leave another copy of the blocking call running. If that call has side effects, make it idempotent.
Timeouts and retries compose, but only if the retry policy covers the timeout. With exceptions=[ConnectionError] or exceptions=[TimeoutError], a node that timed out made 1 attempt and was not retried, while with exceptions=['NodeTimeoutError'] it made 3, on all four interpreters. With timeout=0.3 and RetryConfig(max_attempts=3, initial_delay=0.1, jitter=0.0) and no exceptions list, an async node that never finished also made 3 attempts, starting at 0.0 s, 0.4 s and 0.9 s (the third at 0.91 s on Python 3.10 and 3.11), and then raised NodeTimeoutError. So for a node that yields, worst-case latency is roughly max_attempts × timeout plus the backoff waits: here 3 × 0.3 + 0.1 + 0.2 = 1.2 s with jitter=0.0, and with the default jitter each wait can reach twice its nominal value. A blocking body is not cut off, so each attempt can run for its full body time. With the same settings and a 1.0 s blocking body, Python 3.12 and 3.13 made three attempts and took 3.31 to 3.32 s in total, in line with 3 × 1.0 s plus the 0.1 s and 0.2 s waits, while Python 3.10 and 3.11 completed on the first attempt after 1.01 s.
What happens when a route matches nothing?
No exception is raised. When no routed edge matches and there is no DEFAULT_ROUTE, none of the routed successors run, and the branch ends there unless the router also has an unrouted edge, which is still followed. A routing node returns Event(output=..., route="..."), and the edge map picks the successor. In my check the node emitted the route "nope", and I ran it against three edge lists:
def classify(node_input: str):
return Event(output=node_input, route="nope")
def yes(node_input: str): return "YES"
def fallback(node_input: str): return "FALLBACK"
def always(node_input: str): return "ALWAYS"
cases = {
"no default": [("START", classify, {"yes": yes})],
"DEFAULT_ROUTE": [("START", classify, {"yes": yes, DEFAULT_ROUTE: fallback})],
"no default, plus an unrouted edge": [("START", classify, {"yes": yes}), (classify, always)],
}
With only routed edges and no default, the workflow completed normally with the single output 'go', nothing downstream ran, and ADK logged one warning: Node 'classify' has conditional/DEFAULT edges but none were matched by the emitted route(s): nope. The branch will end. With DEFAULT_ROUTE: fallback in the map, the fallback node ran, the outputs were ['go', 'FALLBACK'], and there was no warning. With an extra unrouted edge from the router to always, the unmatched route raised no error and logged nothing: always ran and the outputs were ['go', 'ALWAYS']. In _graph.py, edges without a route are always followed, and the warning is logged only when a node that has routed edges ends up scheduling nothing.6
The routes page describes DEFAULT_ROUTE in its TypeScript section as a setting that "matches when no other route on the same source node matches," and its Go section documents the equivalent workflow.Default. I found no Python example of it on that page and no statement about the no-match case without a default.8 What happens there comes from running it and from _graph.py.6 Give every router a DEFAULT_ROUTE. If the route value comes from something you do not control, such as a classifier's label, point the default at a node that records the problem and escalates. Alerting on the warning text helps too, but it only catches routers whose outgoing edges are all routed.
How do fan-out and JoinNode behave?
Fan-out works, and JoinNode is defined as a node that "waits for all specified predecessors" before outputting.6 In every run the join output was a dict keyed by node name, {'w1': 'done', 'w2': 'done'} (key order varied between runs). What varied was the elapsed time, which depended on whether the workers yielded to the event loop. Each of the two workers waits 0.3 s:
| Worker body (two workers, 0.3 s each) | Elapsed across the four interpreters |
|---|---|
def + time.sleep(0.3) | 0.61 to 0.62 s |
async def + await asyncio.sleep(0.3) | 0.31 s |
async def + await asyncio.to_thread(time.sleep, 0.3) | 0.31 to 0.33 s |
async def + blocking time.sleep(0.3) | 0.61 s |
The rule is the same as for timeouts: workers that block the event loop take turns, and workers that yield while waiting overlap. Marking a function async def is not enough. The fourth row is an async def worker that blocks, and it gets no speed-up.
How does human-in-the-loop work in a Workflow?
A node returns or yields a RequestInput, the run pauses, and the client resumes it with a function response that carries the interrupt ID. In my checks, both a node that returned a RequestInput and one that yielded it produced an adk_request_input function call. The documentation lists three options for RequestInput (message, payload and response_schema), and its Python section says that once the system receives input from the user, "that input is passed to the next node."9 The TypeScript section adds that with the leaf default rerunOnResume: false, "the reply is routed to the node's successor as input, bypassing the interrupted node."9 In the Python package, FunctionNode also defaults rerun_on_resume to false.6
My main check chains draft, approve and execute, where approve returns RequestInput(message=..., payload=..., response_schema=str). The first run stopped at approve. An event carried a function call named adk_request_input with interruptId, payload, message and a response_schema of {'type': 'string'}, and the counters read draft 1, approve 1, execute 0. I resumed with a FunctionResponse whose response was {"result": "yes approved"}:
def reply(call, value):
return types.Content(role="user", parts=[types.Part(function_response=types.FunctionResponse(
id=call.id, name="adk_request_input", response=value))])
execute received the string from the result key and returned EXECUTED after human said: 'yes approved'. The counters then read draft 1, approve 1, execute 1, so neither earlier node ran again.
The response schema is enforced. Replying with the number 42 ({"result": 42}) to the response_schema=str request made run_async raise WorkflowDataError, and execute ran 0 times. The same reply was accepted when the RequestInput had no response_schema, and execute then received the integer 42. ADK also runs the text inside a {"result": "..."} reply through Python's json.loads before validating it (_rehydration_utils.py), so any text that parses to something other than a string fails a str schema.6 Replies typed as the text 42 and as true both raised WorkflowDataError, with execute running 0 times, while 042, which json.loads rejects because of the leading zero, reached execute unchanged as the string '042'. Python's parser is also more lenient than JSON: it turns NaN, Infinity and -Infinity into floats, so those replies fail a str schema as well. The human-input page also says that ADK Go does not automatically parse or validate the structure of the reply, so don't assume the Python behavior carries across SDKs.9 In Python, define a response_schema for every pause, be ready to handle the error where you deliver replies, and expect a person who types true or 42 into a string field to trigger it.
One caution from the resume documentation, which covers resuming a stopped agent workflow by invocation ID: the Resume feature "ensures that the Tools in an agent are run at least once, and may run more than once when resuming a workflow."10 That page's version banner lists Python v1.16.0 and Kotlin v0.1.0, it covers ResumabilityConfig with sequential, loop, parallel and custom agents, and it does not mention graph workflows or RequestInput, so I can't say how it applies to graph nodes. In my run, no earlier node re-ran after a human reply. As a precaution, keep any node with side effects that could be replayed after an interruption idempotent.
Production checklist for ADK Workflows
| Concern | What I found | What to do |
|---|---|---|
| Retry budget | A bare RetryConfig() means 5 attempts; with default jitter, first-to-last-attempt times of 4.4 to 26.1 s in the first batches on four interpreters, and the four waits can total up to 30 s | Set max_attempts, initial_delay and jitter on purpose; budget for the ceiling |
| Retry scope | No retry_config means the node is not retried on its own; exceptions unset retried every exception I raised, timeouts included; listed names match the exception's class or any base class except object, by name | List the transient exception types, and include NodeTimeoutError if timeouts should be retried |
| Timeouts | Cannot interrupt blocking code; afterwards the node completes or fails late, depending on what it does next and on the Python version | Use non-blocking async code, or asyncio.to_thread for blocking calls |
Timeout with to_thread | Threads kept running after the timeout: 3 of 3 unfinished when run_async raised | Make the blocking call idempotent before combining it with retries |
| Timeout plus retry | Yielding node: attempts at about 0.0, 0.4 and 0.9 s with timeout=0.3; blocking body: 3 attempts took 3.31 to 3.32 s on 3.12 and 3.13 | Budget max_attempts × timeout for yielding nodes and max_attempts × body time for blocking ones, plus the backoff waits in both cases |
| Routing | An unmatched route raises no error; with only routed edges the branch ends with one warning, and with an unrouted edge as well nothing is logged | Add DEFAULT_ROUTE to every router; alert on the warning, but don't rely on it alone |
| Fan-out | Blocking workers ran in 0.61 to 0.62 s, yielding workers in 0.31 to 0.33 s | Make parallel workers yield |
| Human pause | With the default rerun_on_resume=False, the reply goes to the successor; a reply that breaks response_schema raises WorkflowDataError, and for a str schema that includes text that json.loads turns into a non-string value | Set response_schema and handle the error |
| Crash recovery | Every retry logs that the retry count is not persisted across resuming (resuming mid-retry or after a crash not tested) | Do not treat max_attempts as a durable budget |
For related reading on NerdLevelTech, Managed Agent Runtimes: AWS, Google, Alibaba in 2026 covers Google's Gemini Enterprise Agent Platform alongside AWS and Alibaba, AI Agents in Go: Microsoft Joins Google's Bet in 2026 covers Google's ADK for Go and Microsoft's Agent Framework for Go, and if one of your nodes waits on a slow MCP tool call, MCP Tasks Extension: A Working Python Server (2026) shows how to hand back a task handle instead of blocking.
Reproduce it
Install google-adk 2.10.0 in a fresh virtual environment (pip install google-adk==2.10.0 google-genai==2.25.0; this also installs requests, a google-adk dependency that the script imports), save the script below as adk_workflow_checks.py and run python adk_workflow_checks.py. A full run takes about a minute on my machine. Pass one or more of retry, scope, timeout, route, fanout or hitl to run only those checks. The default-jitter figures are random and every duration depends on your machine, so expect the same pattern rather than the same decimals.
"""Checks for google-adk 2.10.0 Workflow behavior. Run: python adk_workflow_checks.py [retry|scope|timeout|route|fanout|hitl]"""
import asyncio, logging, statistics, sys, time
import requests # installed with google-adk, which depends on it
from google.adk import Event
from google.adk.events import RequestInput
from google.adk.runners import InMemoryRunner
from google.adk.workflow import DEFAULT_ROUTE, FunctionNode, JoinNode, RetryConfig, Workflow
from google.genai import types
log_lines = []
class Capture(logging.Handler):
def emit(self, record):
if record.levelno >= logging.WARNING:
log_lines.append(record.getMessage())
logging.getLogger("google_adk").addHandler(Capture())
def user_msg(text):
return types.Content(role="user", parts=[types.Part(text=text)])
async def run(wf, message=None, runner=None, session=None):
"""Run a workflow once. Returns (runner, session, events, exception_or_None)."""
runner = runner or InMemoryRunner(agent=wf, app_name="demo")
session = session or await runner.session_service.create_session(app_name="demo", user_id="u")
events, error = [], None
try:
async for ev in runner.run_async(user_id="u", session_id=session.id,
new_message=message or user_msg("go")):
events.append(ev)
except Exception as exc:
error = exc
return runner, session, events, error
def status(error):
return type(error).__name__ if error else "completed"
async def retry():
log_lines.clear()
stamps = []
def always_fail(node_input: str):
stamps.append(time.monotonic())
raise ConnectionError("x")
node = FunctionNode(func=always_fail, name="f", retry_config=RetryConfig(jitter=0.0))
_, _, events, error = await run(Workflow(name="retry", edges=[("START", node)]))
gaps = [round(b - a, 2) for a, b in zip(stamps, stamps[1:])]
errors = sum(1 for e in events if e.error_message)
print(f"retry jitter=0: {len(stamps)} attempts, gaps between attempt starts {gaps}s, "
f"{errors} error events, outcome {status(error)}")
persisted = [m for m in log_lines if "not persisted across resuming" in m]
print(f" {len(persisted)} warnings say the retry count is not persisted across resuming")
tries = []
def flaky(node_input: str):
tries.append(1)
if len(tries) < 3:
raise ConnectionError("x")
return "ok"
node = FunctionNode(func=flaky, name="f", retry_config=RetryConfig(jitter=0.0, initial_delay=0.05))
_, _, events, error = await run(Workflow(name="flaky", edges=[("START", node)]))
errors = sum(1 for e in events if e.error_message)
print(f"retry fails twice, then succeeds: {len(tries)} attempts, {errors} error events, outcome {status(error)}")
async def one_run(): # default RetryConfig(): jitter left at its default
marks = []
def fail(node_input: str):
marks.append(time.monotonic())
raise ConnectionError("x")
wf = Workflow(name="j", edges=[("START", FunctionNode(func=fail, name="f", retry_config=RetryConfig()))])
await run(wf)
return [b - a for a, b in zip(marks, marks[1:])]
runs = await asyncio.gather(*[one_run() for _ in range(20)])
totals = [sum(r) for r in runs]
print(f"retry default jitter, 20 runs: attempts per run {sorted({len(r) + 1 for r in runs})}, "
f"first to last attempt min {min(totals):.1f}s mean {statistics.mean(totals):.1f}s max {max(totals):.1f}s")
for i in range(4):
col = [r[i] for r in runs]
print(f" gap before attempt {i + 2} (nominal wait {2 ** i}s): {min(col):.2f}s to {max(col):.2f}s")
async def scope():
async def calls(label, raises, retry_config):
n = {"calls": 0}
def flaky(node_input: str):
n["calls"] += 1
raise raises("x")
node = FunctionNode(func=flaky, name="f", retry_config=retry_config)
_, _, _, error = await run(Workflow(name="scope", edges=[("START", node)]))
print(f"scope {label}: {n['calls']} call(s), outcome {status(error)}")
only_conn = RetryConfig(max_attempts=4, initial_delay=0.05, jitter=0.0, exceptions=[ConnectionError])
by_name = RetryConfig(max_attempts=4, initial_delay=0.05, jitter=0.0, exceptions=["ConnectionError"])
anything = RetryConfig(max_attempts=4, initial_delay=0.05, jitter=0.0)
await calls("exceptions=[ConnectionError], ValueError raised", ValueError, only_conn)
await calls("exceptions=[ConnectionError], ConnectionResetError raised", ConnectionResetError, only_conn)
await calls("exceptions=['ConnectionError'] (a string), ConnectionResetError raised", ConnectionResetError, by_name)
await calls("exceptions=[ConnectionError], requests.exceptions.ConnectionError raised",
requests.exceptions.ConnectionError, only_conn)
await calls("exceptions unset, KeyError raised", KeyError, anything)
await calls("no retry_config, ValueError raised", ValueError, None)
async def timeout():
def sync_block(node_input: str):
time.sleep(1.0); return "late"
async def async_wait(node_input: str):
await asyncio.sleep(1.0); return "late"
async def async_thread(node_input: str):
await asyncio.to_thread(time.sleep, 1.0); return "late"
async def async_block(node_input: str):
time.sleep(1.0); return "late"
async def block_then_await(node_input: str):
time.sleep(1.0); await asyncio.sleep(0); return "late"
async def returns_at_once():
return None
async def block_then_noop(node_input: str):
time.sleep(1.0); await returns_at_once(); return "late"
def block_no_output(node_input: str):
time.sleep(1.0)
print(f"python {sys.version.split()[0]}, FunctionNode(timeout=0.3), body takes 1.0s")
for fn in (sync_block, async_wait, async_thread, async_block, block_then_await, block_then_noop,
block_no_output):
node = FunctionNode(func=fn, name=fn.__name__, timeout=0.3)
start = time.monotonic()
_, _, _, error = await run(Workflow(name="t", edges=[("START", node)]))
print(f" {fn.__name__:16} -> {status(error):16} after {time.monotonic() - start:.2f}s")
# timeout + retry: attempt start offsets
starts = []
async def slow(node_input: str):
starts.append(time.monotonic()); await asyncio.sleep(2)
node = FunctionNode(func=slow, name="slow", timeout=0.3,
retry_config=RetryConfig(max_attempts=3, initial_delay=0.1, jitter=0.0))
_, _, _, error = await run(Workflow(name="tr", edges=[("START", node)]))
print(f" timeout+retry(3): attempts at {[round(s - starts[0], 2) for s in starts]}, outcome {status(error)}")
# does a timeout retry when exceptions is scoped?
for label, excs in (("exceptions=[ConnectionError]", [ConnectionError]),
("exceptions=[TimeoutError]", [TimeoutError]),
("exceptions=['NodeTimeoutError']", ["NodeTimeoutError"])):
seen = []
async def slow_scoped(node_input: str):
seen.append(1); await asyncio.sleep(2)
node = FunctionNode(func=slow_scoped, name="slow_scoped", timeout=0.3,
retry_config=RetryConfig(max_attempts=3, initial_delay=0.05, jitter=0.0, exceptions=excs))
_, _, _, error = await run(Workflow(name="ts", edges=[("START", node)]))
print(f" timeout + retry {label}: {len(seen)} attempt(s), outcome {status(error)}")
# a blocking body with timeout and retry
begins = []
async def blocking_retry(node_input: str):
begins.append(time.monotonic()); time.sleep(1.0); return "late"
node = FunctionNode(func=blocking_retry, name="blocking_retry", timeout=0.3,
retry_config=RetryConfig(max_attempts=3, initial_delay=0.1, jitter=0.0))
start = time.monotonic()
_, _, _, error = await run(Workflow(name="br", edges=[("START", node)]))
print(f" blocking body + timeout + retry(3): attempts at {[round(b - begins[0], 2) for b in begins]}, "
f"{time.monotonic() - start:.2f}s in total, outcome {status(error)}")
# threads survive a timeout
live = {"started": 0, "finished": 0}
async def threaded(node_input: str):
def blocking():
live["started"] += 1; time.sleep(1.0); live["finished"] += 1
await asyncio.to_thread(blocking)
node = FunctionNode(func=threaded, name="threaded", timeout=0.2,
retry_config=RetryConfig(max_attempts=3, initial_delay=0.05, jitter=0.0))
_, _, _, error = await run(Workflow(name="th", edges=[("START", node)]))
at_return = dict(live)
await asyncio.sleep(1.5)
print(f" to_thread+timeout+retry(3): outcome {status(error)}; at return {at_return}; 1.5s later {live}")
async def route():
def classify(node_input: str):
return Event(output=node_input, route="nope")
def yes(node_input: str): return "YES"
def fallback(node_input: str): return "FALLBACK"
def always(node_input: str): return "ALWAYS"
cases = {
"no default": [("START", classify, {"yes": yes})],
"DEFAULT_ROUTE": [("START", classify, {"yes": yes, DEFAULT_ROUTE: fallback})],
"no default, plus an unrouted edge": [("START", classify, {"yes": yes}), (classify, always)],
}
for label, edges in cases.items():
log_lines.clear()
_, _, events, error = await run(Workflow(name="r", edges=edges))
print(f"route {label}: {status(error)}, outputs {[e.output for e in events]}, warnings {log_lines}")
async def fanout():
async def split(node_input: str): return node_input
def sync_w(node_input: str):
time.sleep(0.3); return "done"
async def async_w(node_input: str):
await asyncio.sleep(0.3); return "done"
async def thread_w(node_input: str):
await asyncio.to_thread(time.sleep, 0.3); return "done"
async def block_w(node_input: str):
time.sleep(0.3); return "done"
for fn in (sync_w, async_w, thread_w, block_w):
w1, w2 = FunctionNode(func=fn, name="w1"), FunctionNode(func=fn, name="w2")
wf = Workflow(name="f", edges=[("START", split, (w1, w2), JoinNode(name="join"))])
start = time.monotonic()
_, _, events, error = await run(wf)
joined = [e.output for e in events if isinstance(e.output, dict)]
print(f"fanout {fn.__name__:8} -> {time.monotonic() - start:.2f}s, join output {joined}, {status(error)}")
async def hitl():
counts = {"draft": 0, "approve": 0, "execute": 0}
def draft(node_input: str):
counts["draft"] += 1; return f"REFUND $500 for {node_input}"
def make_approve(schema):
def approve(node_input: str):
counts["approve"] += 1
return RequestInput(message=f"Approve? {node_input}", payload=node_input, response_schema=schema)
return approve
def execute(node_input):
counts["execute"] += 1; return f"EXECUTED after human said: {node_input!r}"
def reply(call, value):
return types.Content(role="user", parts=[types.Part(function_response=types.FunctionResponse(
id=call.id, name="adk_request_input", response=value))])
async def pause(schema, order):
for key in counts: counts[key] = 0
wf = Workflow(name="h", edges=[("START", draft, make_approve(schema), execute)])
runner, session, events, _ = await run(wf, user_msg(order))
return wf, runner, session, [fc for e in events for fc in e.get_function_calls()][0]
wf, runner, session, call = await pause(str, "order-17")
print(f"hitl pause: function call {call.name!r}, args {call.args}, counts {counts}")
_, _, events, error = await run(wf, reply(call, {"result": "yes approved"}), runner, session)
print(f"hitl resume: {status(error)}, outputs {[e.output for e in events]}, counts {counts}")
def approve_yield(node_input: str):
yield RequestInput(message=f"Approve? {node_input}", response_schema=str)
wf_yield = Workflow(name="hy", edges=[("START", draft, approve_yield, execute)])
_, _, events, _ = await run(wf_yield, user_msg("order-20"))
print(f"hitl with yield instead of return: function calls {[fc.name for e in events for fc in e.get_function_calls()]}")
wf, runner, session, call = await pause(str, "order-18")
_, _, events, error = await run(wf, reply(call, {"result": 42}), runner, session)
print(f"hitl reply 42 sent as a number, response_schema=str: {status(error)}, execute ran {counts['execute']} time(s)")
for text in ("42", "true", "042"):
wf, runner, session, call = await pause(str, f"order-{text}")
_, _, events, error = await run(wf, reply(call, {"result": text}), runner, session)
print(f"hitl reply typed as the text {text}, response_schema=str: {status(error)}, "
f"execute ran {counts['execute']} time(s), outputs {[e.output for e in events]}")
wf, runner, session, call = await pause(None, "order-19")
_, _, events, error = await run(wf, reply(call, {"result": 42}), runner, session)
print(f"hitl reply 42 sent as a number, no response_schema: {status(error)}, outputs {[e.output for e in events]}")
async def warm_up(): # first run pays one-time import costs; keep them out of the timings
await run(Workflow(name="warm", edges=[("START", FunctionNode(func=lambda node_input: "ok", name="ok"))]))
CHECKS = {"retry": retry, "scope": scope, "timeout": timeout, "route": route, "fanout": fanout, "hitl": hitl}
async def main():
await warm_up()
for name in sys.argv[1:] or CHECKS:
await CHECKS[name]()
asyncio.run(main())
Limits of these tests
These checks cover function nodes and the runtime around them in the Python package, on one sandbox machine, with wall-clock durations and time.sleep standing in for blocking work. Most timing cells come from a single run per interpreter, and the jitter figures come from 20-run batches. I did not test LLM-agent nodes, tool nodes, dynamic workflows, persistent session services, deployed runners, the Go or TypeScript SDKs, Python 3.14, or resuming after a crash, after a restart or mid-retry. A TODO in _node_runner.py says retrying dynamic-node failures is to be considered later: in 2.10.0, when a node's dynamic child (started with ctx.run_node) fails and the resulting DynamicNodeFailError propagates out of the calling node, the node runner does not retry the calling node, even with exceptions unset.6 The timeout result depends on what the node does after the blocking call and on the interpreter, so test your real node on the version you deploy. The 30 s ceiling comes from reading the delay formula, not from observation.
Bottom line
In these tests, routing, fan-out, retries and human pauses all worked, and the surprises came from how they interact with the event loop and with defaults: jittered retry waits, timeouts that cannot interrupt blocking code and then behave differently depending on what the node does next and on the Python version, and unmatched routes that raise no error and at most leave a log line. Write nodes that yield to the event loop, set the retry policy explicitly, give every router a DEFAULT_ROUTE, define a response_schema for every human pause, and run the script above on your own graph and your own Python version.
Footnotes
-
Release v2.0.0 — google/adk-python on GitHub, fetched 2026-09-29 (release v2.0.0, published May 19). ↩ ↩2
-
Welcome to ADK 2.0 — Agent Development Kit documentation, fetched 2026-09-29 (Python 2.0 GA May 19, 2026; Go 2.0 GA June 30, 2026; TypeScript 2.0 GA August 21, 2026; the three main features; the exception-handling migration notes). ↩ ↩2 ↩3 ↩4
-
google-adk 2.10.0 — PyPI, fetched 2026-09-29 ("Released: Sep 25, 2026"; "Requires: Python >=3.10"), and Release v2.10.0 — google/adk-python on GitHub (published September 25; the changelog heading is dated 2026-09-24). ↩ ↩2 ↩3
-
Python quickstart — ADK documentation, fetched 2026-09-29 ("Python 3.10 or later"). ↩ ↩2
-
Why we built ADK 2.0 — Google Developers Blog, published July 1, 2026, fetched 2026-09-29. ↩
-
google-adk2.10.0 package source, installed from PyPI and read on 2026-09-29:google/adk/workflow/_retry_config.py(documented defaults),workflow/utils/_retry_utils.py(attempt limit, name-based exception matching, delay and jitter),workflow/_errors.py(NodeTimeoutErrorbase class),workflow/_node_runner.py(asyncio.wait_fortimeout, error events, retry warning, per-run attempt count, dynamic-node TODO),workflow/_function_node.py(sync functions called inline,Nonereturns and state events,rerun_on_resumedefault),workflow/_base_node.py(timeoutandretry_configdocstrings),workflow/_graph.py(edge matching and the unmatched-route warning),workflow/_join_node.py(JoinNodedocstring),workflow/_workflow.py(no retry logic),agents/invocation_context.py(event hand-off to the runner),agents/context.py(run_node),workflow/_node_runner_utils.py(the runner drives the root node through aNodeRunner, so a top-levelWorkflow'sretry_configapplies),workflow/_dynamic_node_scheduler.py(DynamicNodeFailErrorraised when arun_nodechild fails) andworkflow/utils/_rehydration_utils.py(JSON-parsing of{"result": ...}text). Therequests.exceptions.ConnectionErrorclass hierarchy was checked inrequests2.34.2, installed as a google-adk dependency. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15 ↩16 ↩17 ↩18 ↩19 ↩20 ↩21 -
Python standard library
asyncio.wait_forsource, inspected withinspect.getsourcein Python 3.10.20, 3.11.15, 3.12.3 and 3.13.13 on 2026-09-29. ↩ -
Graph routes — ADK documentation, fetched 2026-09-29. ↩ ↩2
-
Human input — ADK documentation, fetched 2026-09-29. ↩ ↩2 ↩3 ↩4
-
Resume Agents — ADK documentation, fetched 2026-09-29. ↩



