Skip to content

span 01 · note · 2026-10-04 · 6 min read

What broke while shipping the agent on this site

A closed tab that cost money, an endpoint any website could spend, an answer stuck on “thinking…”, a focus trap broken by the Enter key and a WebGL hero that scored 50 on Lighthouse — and the tests that now pin each fix.

The demo worked on the first day

The agent behind the ⌘K palette on this site answered its first question correctly almost immediately. A Claude tool-use loop, four tools over my portfolio content, events streamed to the browser over server-sent events, each tool call drawn as a span in a live trace. On a fast laptop with a good connection, it looked done.

It wasn't. Everything below broke — or would have broken in front of a visitor — and none of it showed up by looking at the demo. Each problem was caught by a test, a measurement or a review, and each fix now has a test that failed before the fix and passes after it. The model call was the easy part. These are the parts that weren't.

1. Closing the tab could cost money

The route streams events with a ReadableStream. When a visitor closes the tab mid-answer, the stream is cancelled, and the next controller.enqueue throws a TypeError.

My agent loop had a deliberately narrow retry: if the SDK failed to parse a streamed tool input, re-issue the turn, at most twice. But the catch block treated any non-API error as that case. So a closed tab looked like a parse failure, and the loop re-issued the whole request — two more paid model calls for an answer nobody would read. The model also kept generating after the disconnect, because no abort signal reached the SDK.

The fix was three small changes: pass the request's abort signal into the SDK call, turn emit into a no-op once the stream is cancelled, and never retry after an abort.

} catch (err) {
  // Never retry once the visitor has gone: every retry is a paid model call.
  if (signal?.aborted || err instanceof Anthropic.APIError || parseRetries >= MAX_PARSE_RETRIES) throw err;
  parseRetries += 1;
}

Lesson: in an agent, a retry isn't free. Every retry branch needs to ask "is anyone still waiting for this?"

2. Any website could spend my budget

The endpoint had per-IP rate limiting, input-length limits and a cap of four tool rounds per question. All reasonable — and all bypassable by a third-party page that calls the endpoint from its own visitors' browsers. The response is unreadable cross-origin, but the request still runs, the model is still billed, and every visitor is a fresh IP.

Browsers send Sec-Fetch-Site on fetch(), so the fix is two lines at the top of the route:

const site = req.headers.get("sec-fetch-site");
if (site && site !== "same-origin") return Response.json({ error: "Forbidden" }, { status: 403 });

It isn't a complete defence against a determined attacker with a script, which is why a provider-side spend limit is still part of the deployment checklist. But it closes the cheapest abuse path entirely.

Lesson: for an LLM endpoint, "who can make this request?" is a cost question as much as a security one.

3. An answer could hang on "thinking…" forever

The model runs with adaptive thinking, and thinking tokens count toward max_tokens. A hard question can spend the whole budget thinking and return no text and no tool call. My loop saw no tool calls, assumed the turn was finished, and emitted done. The palette faithfully kept showing "thinking…" — forever.

The fix: track whether any text was emitted, and if a run ends without any, say so.

if (!emittedText) emit({ type: "error", message: EMPTY_TEXT, offline: false });
emit({ type: "done" });

Lesson: "the stream ended" is not the same as "the user got an answer". Make the empty case an explicit state.

4. One hiccup switched the agent off for the whole visit

When the API is unreachable, the palette switches to clearly labelled scripted answers instead of failing. The first version made that switch permanent for the session after any failure. One transient 500 or a cold start, and a recruiter would see "offline mode" for every later question — on the feature the whole site is built to show off.

Now only two failures are sticky: 503 (the agent isn't configured) and 429 (this visitor is rate-limited). Anything else falls back for that one answer and tries live again on the next question.

5. Asking a question with the keyboard broke the focus trap

The palette is a modal dialog with a focus trap, and the end-to-end tests covered opening it, tabbing around inside it and closing it with Escape. What they didn't cover was asking a question with the keyboard.

Pressing Enter disabled the input while the answer streamed. Disabling a focused element sends focus to <body>, so the next Tab landed on the "Skip to content" link behind the backdrop, and focus never came back.

The fix is one attribute — readOnly instead of disabled while streaming — plus a test that does exactly what a keyboard user does: type, press Enter, wait for the answer, press Tab, assert focus is still inside the dialog. A related race came from moving focus into the dialog one animation frame after opening it; a fast Tab could escape in that gap. Moving focus in a layout effect closed it.

Lesson: accessibility bugs live in the transitions between states, not in the states. Test the transitions.

6. The WebGL hero scored 50 on Lighthouse

The hero is a particle field that condenses into an agent graph as you scroll. The first Lighthouse run on mobile scored 50 for performance: 3.3 seconds of total blocking time, almost all of it from mounting the three.js scene about three seconds after load.

The scene is decoration. It shouldn't compete with the first paint. Now a static SVG version of the graph renders on the server, and the WebGL scene mounts on the visitor's first interaction — a pointer move, scroll, touch or key press. Performance went to 91–92, with total blocking time around 60 ms and zero layout shift.

That fix created its own bug. The hero section grew from one viewport to 180% of a viewport when the scene mounted, because the pinned scroll range only existed for the WebGL version. On a desktop the document grew by 640 px after a single mouse move; on a phone, where the first interaction is the scroll itself, the page grew by most of a screen under the visitor's finger. The section's height now depends only on whether the device can render the scene, never on whether it has.

Lesson: deferring work is free only if deferring it doesn't change the layout.

7. Three bugs that were mine alone

A few problems had nothing to do with agents, and I'm including them because the way they were found matters as much as the fix.

  • The particles started fully formed. Scroll progress came from measuring the hero section. The measurement was taken before the section grew, so the scene believed the user had already scrolled to the end. Probing a production build showed the graph's node labels at full opacity at scrollY = 0. Deriving progress from window scroll fixed it.
  • Every scroll-revealed heading was invisible. The reveal watched the heading line with an IntersectionObserver, but the line starts translated out of its overflow: hidden wrapper. An element that's fully clipped never intersects, so the reveal never fired. The observer now watches the wrapper.
  • 800 × 0.56 is not 448. In floating point it's 448.00000000000006, so a progress value meant to reach exactly 1 never did. A unit test asserting toBe(1) caught it. The fix was to rearrange the arithmetic.

What I'd tell myself on day one

  • Instrument before you polish. Lighthouse, a 360 px overflow test and a reduced-motion test found more real problems than any amount of looking at the animation.
  • Every paid call needs an owner. Retries, aborts and cross-origin requests are all cost decisions.
  • Make the empty and failure states explicit. "Offline", "no answer" and "cut off" are designed states with their own copy, not accidents.
  • Write the test that does what the user does. The focus bug survived four keyboard tests because none of them asked a question.

You can read the architecture of the agent, or try it yourself with ⌘K. It's the same code, with every fix above in place.