Skip to content

span 01 · note · 2026-10-04 · 9 min read

Agent reliability is a queueing problem

A tool-using agent is a distributed system with an expensive, non-deterministic node in the middle. Retries with jitter, idempotency keys, deadlines, backpressure, checkpoints and dead letters — the backend patterns that keep agents alive in production.

The demo runs once. Production runs forever.

An agent demo is a single run on a good day: one prompt, a few tool calls, a clean answer. Production is the same loop running thousands of times while the world misbehaves around it. The model API returns 529 overloaded_error when traffic spikes. A 429 arrives with a retry-after header. A stream that started with 200 OK fails halfway through — Anthropic's own error docs note that errors can occur "after the API returns a 200 response." A tool times out after it already sent the email.

None of these are AI problems. They are distributed-systems problems, and they have well-known answers.

Here is the framing I keep coming back to: a tool-using agent is a distributed system with an expensive, slow, non-deterministic node in the middle. Every step is a remote call that can fail, stall, or succeed twice. That is exactly the shape of the queue-backed systems I've built — RabbitMQ pipelines that decouple producers from consumers, and webhook ingestion for Shopify, WooCommerce and Global Payments where the same event can arrive more than once. The patterns that kept those systems honest map almost one-to-one onto agent runs.

Why multi-step makes it worse

Anthropic's "Building effective agents" is refreshingly direct about the cost of autonomy: "The autonomous nature of agents means higher costs, and the potential for compounding errors." It recommends "stopping conditions (such as a maximum number of iterations) to maintain control," and "finding the simplest solution possible, and only increasing complexity when needed."

Compounding is easy to underestimate. As illustrative arithmetic — not a measured figure — if each of 20 independent steps succeeds 95% of the time, the whole run succeeds about 36% of the time (0.95²⁰). Real steps aren't independent, but the direction holds, and benchmarks see it too: in τ-bench (Yao et al., 2024), state-of-the-art function-calling agents succeeded on under 50% of tasks and were "quite inconsistent", with pass^8 — all eight repeated trials succeeding — under 25% in the retail domain.

You can't prompt your way out of that. You engineer around it, the same way you engineer around flaky networks.

1. Retries: at one layer, with jitter, and only for the right errors

The good news is that you start with sensible defaults. Both official Anthropic SDKs retry connection errors, 408, 409, 429 and 5xx responses "2 times by default, with a short exponential backoff," honouring retry-after.

The bad news is that it's easy to add retries on top. Marc Brooker's Builders' Library article on timeouts, retries and backoff explains why that hurts: retries are "selfish," and in a five-deep stack where each layer retries three times, "the load on the database will increase 243x." AWS's answer is to "retry at a single point in the stack." For an agent, that means deciding who retries a model call — the SDK, your step runner, or your queue — and switching the others off.

The second rule is jitter. Brooker's analysis of backoff strategies found that plain exponential backoff still produces "clusters of calls," that the no-jitter approach "is the clear loser," and that full jitter "should be considered a standard approach for remote clients."

The third rule is the one people miss: not every failure deserves a retry. Anthropic's rate-limit docs note that a 429 caused by hitting your spend limit has no retry-after header, and retrying — "including the SDK's automatic retries" — "fails until access resumes." That run should stop and land somewhere a human will see it.

type Failure = { status?: number; retryAfterMs?: number };

/** Decide what to do with a failed step: retry (when), or dead-letter it. */
function nextAction(f: Failure, attempt: number, maxAttempts = 4): { retryInMs: number } | "dead-letter" {
  const retryable = f.status === undefined || [408, 409, 429, 500, 502, 503, 504, 529].includes(f.status);
  const spendLimited = f.status === 429 && f.retryAfterMs === undefined;
  if (!retryable || spendLimited || attempt >= maxAttempts) return "dead-letter";
  if (f.retryAfterMs !== undefined) return { retryInMs: f.retryAfterMs }; // the server knows best
  const cap = 30_000;
  const base = 500;
  return { retryInMs: Math.random() * Math.min(cap, base * 2 ** attempt) }; // full jitter
}

2. Idempotency: every side effect needs a key

Queues deliver at least once in practice. RabbitMQ's documentation on acknowledgements says any unacknowledged delivery "is automatically requeued when the channel (or connection) on which the delivery happened is closed," and that "consumers must be prepared to handle redeliveries and otherwise be implemented with idempotence in mind."

An agent step is no different. If the worker crashes after the tool ran but before the result was recorded, the step will run again. For read-only tools that's harmless. For send_invoice, refund_order or create_ticket, it's an incident.

Stripe solved this for payments years ago. Their API saves "the resulting status code and body of the first request made for any given idempotency key, regardless of whether it succeeds or fails," and replays it for later requests with the same key. Inngest's retry guide states the agent version of the same idea plainly: a retry can call an external API again "when the first call succeeded but its response was lost. Use the same idempotency key for every attempt."

For a tool call, the key derives from things that don't change between attempts: the run, the step, and the arguments.

import { createHash } from "node:crypto";

async function runToolOnce<T>(runId: string, step: number, tool: string, args: unknown, exec: () => Promise<T>) {
  const key = createHash("sha256").update(JSON.stringify([runId, step, tool, args])).digest("hex");

  // A unique constraint on `key` makes the claim atomic: exactly one attempt wins.
  const claimed = await db.query(
    "INSERT INTO tool_calls (key, status) VALUES ($1, 'running') ON CONFLICT (key) DO NOTHING RETURNING key",
    [key],
  );
  if (claimed.rowCount === 0) return (await waitForResult(key)) as T; // another attempt already ran it

  const result = await exec();
  await db.query("UPDATE tool_calls SET status = 'done', result = $2 WHERE key = $1", [key, JSON.stringify(result)]);
  return result;
}

The model's output isn't deterministic, so the decision to call a tool may differ between runs. That's fine. The key protects the side effect once a decision has been recorded, which is the moment it matters.

3. Deadlines: an agent step should know when to give up

Defaults are generous. The Anthropic SDKs default to a 10-minute request timeout, retried twice. That's the right default for a library, and almost never the right budget for a step inside a user-facing run.

The Builders' Library recommends choosing an acceptable false-timeout rate — say 0.1% — and setting the timeout at the matching latency percentile of the dependency, p99.9. For agents I'd add one thing: give every run a deadline, and give each step whatever is left of it. A run that has already spent 50 seconds of a 60-second budget should not start a 10-minute model call.

Queues force the same discipline. RabbitMQ's consumer timeout defaults to 30 minutes; a consumer that holds a delivery longer without acknowledging it has "its channel … closed with a PRECONDITION_FAILED channel exception." A slow model call holding an unacknowledged message is exactly how you trip it. Checkpoint the step, acknowledge it, and enqueue the next one — don't hold a delivery across a long generation.

4. Backpressure: size concurrency to the token budget, not the CPU

RabbitMQ's prefetch is backpressure in one number: it "defines the max number of unacknowledged deliveries that are permitted on a channel." Agent workers need the same limit, but the scarce resource isn't CPU. It's the model provider's rate limits.

Anthropic's API uses "the token bucket algorithm," and its acceleration limits ask you to "ramp up your traffic gradually." A worker pool that pulls 200 runs off a queue at once doesn't go faster — it converts a backlog into a wall of 429s, each of which costs a retry. Cap concurrent model calls per worker, size the cap from your tokens-per-minute budget, and let the queue absorb the burst. That's what it's for.

5. Durable checkpoints: never redo a step that finished

The most expensive failure in an agent run is losing ten good steps because the eleventh crashed the process. Durable-execution systems exist to prevent exactly that, and their docs read like a checklist for agents:

  • DBOS: "Every workflow input and step output is durably stored," "steps should be idempotent," and "once a step completes and is checkpointed, it is never re-executed."
  • Inngest: it "saves the result of each completed step," so a failing step can retry "without re-running steps that already succeeded."
  • Temporal's guidance on agents puts model calls and tools in Activities, "where the actual work happens," and the Workflow "replays your agent's progress using the recorded LLM decisions from the Event History." The non-deterministic decision is recorded once and replayed, never re-asked.

You don't need a framework to adopt the idea. A steps table keyed by run and step index — status, input, output, attempts — gives you resumption, an audit log and a debugger in one. When a framework earns its place later, the mental model is already there.

6. Dead letters: failure needs somewhere to go

RabbitMQ republishes a message to a dead-letter exchange when it's rejected without requeue, expires, overflows the queue, or is redelivered past a delivery limit, and the x-death header records why. Agent runs need the same exit: the runs that exhausted retries, hit the iteration cap, or were refused by policy go to a dead-letter queue with their full trace attached, so a person can see exactly which step failed and replay it after a fix.

A run that silently stops is the worst outcome. A run in a dead-letter queue is a bug report.

7. Tracing: a run is a trace, a step is a span

You can't operate what you can't see. OpenTelemetry's GenAI semantic conventions — still marked status "Development" — already give agents a shared vocabulary:

  • A model call is a CLIENT span named {gen_ai.operation.name} {gen_ai.request.model}.
  • A tool call is an INTERNAL span named execute_tool {gen_ai.tool.name}.
  • An agent invocation is invoke_agent {gen_ai.agent.name}.
  • Attributes cover what you need to debug and to bill: gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reasons, gen_ai.tool.call.id, gen_ai.conversation.id.

With every retry, timeout and dead-letter decision recorded as a span event, "why did this run cost four times the average?" becomes a query, not an investigation.

The mapping

Queue-backed backendAgent run
MessageOne step: a model call or a tool call
At-least-once deliveryA step may run more than once
Idempotency keyhash(run, step, tool, args) on every side effect
Retry with backoff + jitterRetry model and tool calls at exactly one layer
Prefetch / QoSConcurrency capped by the token budget
Consumer timeoutPer-run deadline, per-step remaining budget
Dead-letter exchangeRuns that exhaust retries or hit the iteration cap, with their trace
Durable offsets / checkpointsA steps table; never redo a completed step
Distributed traceRun = trace, step = span (gen_ai.*)

The agent on this site is a small application of the same ideas: a hard cap on tool iterations, validation and rate limits before any paid call, no retries once the visitor has disconnected, and a labelled fallback instead of a silent failure. You can read how it's built, or ask it a question with ⌘K.

The model is the newest part of an agent. The reliability work around it is the oldest part of backend engineering — and that's good news, because we already know how to do it.

Sources