Multi-Agent Handoff Testing: Catch Context Loss Before Production
In one run, a worker sent its entire transcript back to the parent agent. The parent exceeded its input limit before it could finish the user’s task. The failure was not a missing fact; it was a handoff too large to use. A shorter handoff has the opposite risk: dropping an unresolved question and letting the parent claim the work is complete.
A handoff is the information passed from a worker agent to the parent that assigned it a subtask. A brief is the bounded summary the parent can act on. Discussions by TestMu and QASkills likewise focus on the transfer boundary: checking each agent separately can miss information lost or dumped between them.
The parent should receive both findings and what remains unknown.
Specify what crosses the boundary
| Participant | Responsibility |
|---|---|
| Worker (provider) | Return the goal, settled facts with evidence pointers, open gaps, and why it stopped, within a payload limit |
| Parent (consumer) | Use those facts without re-asking unnecessarily, keep unresolved gaps visible, stop assigning work when they are closed, and write the user-facing deliverable |
This is a contract test: it checks what the provider sends and what the consumer requires. A pointer to a note can stand in for its full body in the brief, provided the evidence remains retrievable. The payload and parent-rollup caps depend on the model window and product budget; there is no universal byte count here. Stop Spawning Duplicate Workers covers the related stop condition.
If shortening a brief would remove an open gap, mark the parent’s summary unfinished rather than silently treating the gap as closed. A brief should save context space without converting uncertainty into certainty.
Input limits and production settings matter
When a prompt exceeds the serving model’s input window, shrink tool output once according to your compaction policy, then refuse if it still does not fit. Repeatedly feeding the worker’s transcript back into the parent makes overflow likely. Test both the acceptable brief and the refusal path.
Our trials slowed when we turned summaries off to match production. Increasing the trial time budget was preferable to grading a faster, summarized configuration we did not ship. This is a tradeoff: realistic settings cost more time, but different settings answer a different question.
Test the transfer, one change at a time
- Supply a brief with goal, settled facts, open gaps, and stop reason. Check that each field reaches the parent.
- Remove one open gap during trimming and assert that the parent cannot mark the task closed.
- Give the receiver an already settled fact and check it does not delegate simply to rediscover it.
- Exercise the hop budget—the limit on parent-to-worker transfers—and check that another spawn is refused when the budget is spent or the gaps are closed. Log when the cap fires.
- Confirm the parent writes its deliverable at the path the grader reads; keep worker files in isolated folders so they cannot overwrite it.
- Supply an oversize prompt and verify one compaction attempt followed by refusal if necessary.
- Run a live trial with production summary settings after the fast contract tests pass.
These checks have different costs. A unit test can verify fields and bounds quickly; a live trial can show whether the parent actually uses the brief. The testing pyramid helps keep the repeated structural checks out of the slow layer.
The desired outcome is a brief the parent can use, not merely a short one. Size, preserved uncertainty, and correct downstream behavior all matter.
Previous: How to Test AI Agent Loops. Next: AI Agent Evals in CI/CD.
FAQ
How do you test context loss between agents?
Check that the handoff preserves required fields and unresolved questions, stays within its size bound, and reaches the receiver. Test whether the receiver avoids re-asking settled facts and stops delegating when gaps are closed.
What is multi-agent handoff testing?
Testing the transfer from one agent to another: payload completeness, provenance, permissions where relevant, hop budget, and a clear terminal state. A good final answer alone may not reveal a faulty transfer.
Should a worker return its full transcript?
Generally return a bounded brief with settled facts, evidence pointers, and open questions rather than replaying the full transcript. Preserve access to the original evidence when a later audit needs it.
How does this relate to contract tests?
The worker provides a brief and the parent consumes it. Specify and test what that boundary requires, including how the parent behaves if a field or piece of evidence is missing.
Stay in the loop — production notes on AI agents, workflows, and SRE.
Low volume — new posts and curated reading lists. Unsubscribe anytime.