Best Way to Debug a Multi-Step AI Agent in Production
Here is the debug loop we use on multi-step agents in production — not a framework tour. Runtime notes from Aiden where useful; the steps are shape-agnostic.
TL;DR
- One golden prompt you can re-run cold.
- One debug zip / session export (one zip, one conversation).
- Grade per step: what the model saw, which tools paid rent, which were waste.
- Fix host contracts before prose.
- Rematch the golden prompt and publish absolute wall/tokens/correctness.
Explain like I’m five
When a Rube Goldberg machine fails, you do not rewrite the instruction manual first. You watch which gear stuck, replace that gear, then run the same marble again.
Step 1 — Freeze the errand
Dual-part jobs are ideal: alert triage + log anomaly (our combo). If the agent can fake Part A and skip Part B, your gate is soft.
Record: model, orchestration shape (single-agent ReAct vs hierarchical ReAcTree — primer · PDF), host flags, tool endpoint, session id.
Step 2 — Export once
Prefer a single artifact: transcript + tool results + stage timeline. If you only have logs, at least pull:
- wall seconds
- token in/out
- tool names + error strings
- final Theory / Unknowns / Do-this-now
For hierarchical runs, merge parent and child spans from session traces. Parent-only logs lie.
Step 3 — Grade the loop, not the vibes
Ask four questions:
| Question | Fail looks like |
|---|---|
| Did measurement tools run? | Only search_tools / skills |
| Did failures look like failures? | HTTP 200 empty treated as success |
| Did observations change? | Same failed LogQL five times |
| Did the close name receipts? | Undetermined with no blocked query |
Bring-up multi-stage workflows like hardware: green one stage at a time.
Step 4 — Change the host
Order of operations that saved us time:
- Typed failure / no-progress / fan-out (stop retrying)
- Tool path contracts (cursor paging for truncated tool output)
- Prefer parent measurement before children (single vs multi)
- Prompt / skill text last
Mid-run steer belongs here too (steer agents mid-run) — operators need a cancel path when the loop is clearly stuck.
Step 5 — Rematch and refuse relative theater
If wall doubled, say wall doubled (relative efficiency lies). If correctness jumped on one seat, show the Theory text.
That is the whole craft: same marble, one gear, honest stopwatch.
FAQ
What is the best way to debug a multi-step AI agent?
Freeze one golden prompt, export one execution debug bundle, grade what the model saw per step, separate tool good-vs-waste, then change one host gate or tool contract — not the system prompt first.
Should I start by rewriting the system prompt?
No. Prompt changes hide harness bugs. Fix typed failures, fan-out, and completion gates first; use prompts for role and output shape only.
What artifacts do I need?
Session id, wall clock, token totals, tool call list with outcomes, final Theory text, and ideally a single zip that holds the conversation plus tool payloads.
Stay in the loop — production notes on AI agents, workflows, and SRE.
Low volume — new posts and curated reading lists. Unsubscribe anytime.