Stop Spawning Duplicate Workers: What a Handoff Gate Changed in Agent Evals
Rule for this benchmark run: once a delegated worker returns successfully, later create_agent (launch-worker) attempts return that result instead of launching duplicate workers. This avoids repeating the same AppWorld work. It does not remove the planner turns spent attempting those launches, and a production system may need a way to start a genuinely different follow-up job.
We fixed fair agent evals and cut agent orchestration tax on a five-task AppWorld slice — the wins were in measurement and cost, not in declaring a benchmark champion. Strict judge pass stayed 0/5 on this slice; the gate did not “win AppWorld,” it stopped invalid comparisons.
Parts one, two, and three documented unfair handoffs, ~1.6× token overhead, and failure classes. This sequel covers what changed after we shipped the handoff gate, re-enabled note tooling, soft-dropped missing infra names on spawn, and capped worker nodes to one.
Dataset unchanged. Judge unchanged (AppWorld /evaluate, strict success bit). Observability: Langfuse aggregates only.

What changed in the five-task rerun
With the worker tools registered, optional missing names handled, and duplicate launches blocked, 5/5 pairs passed tool-access checks versus 2/5 before. The planner-to-single-agent average token ratio moved from about 1.34× to 0.86×; the planner averaged 27.8 turns instead of 38.4. Multiple changes happened together, and the single-agent token average changed too, so this is not an isolated estimate of the gate’s effect. Neither mode passed strict AppWorld task-goal completion on the rerun. The practical check is to count actual worker sessions as well as attempted launch calls.
Sequel: How to evaluate AI agents: clarify + zero-tool failures — workers that report success without app calls.
What we fixed (behavior, not blueprint)
Caption: The handoff gate avoids duplicate workers; spawn-attempt telemetry and external task success remain separate measures.
After failure-mode runs on the same plan_fit_5 tasks, we addressed harness and runtime issues that made planner mode look worse than it was:
| Fix | What it does |
|---|---|
| Note tools registered | Infra note/read tools available so spawn payloads do not hard-fail on names the registry lacked |
| Soft-drop on spawn | Optional infra names dropped instead of aborting the whole create_agent call |
| Successful handoff gate | First completed worker handoff is stored; repeat spawn attempts return that result instead of launching another AppWorld runner |
| Worker node cap = 1 | Planner tree cannot grow a second worker branch on these eval configs |
| Unattended mode | No clarify / human-wait tools in benchmark runs |
Same five delegation-fit tasks (phone → notes → SMS, inbox + contacts + Splitwise, workout note → Spotify, batch Venmo, trip ledger → Splitwise). Same model family and MCP tool pack.
Before vs after (aggregate)
| Metric | Pre-gate cohort | Post-gate (plan-fit5-gate) |
|---|---|---|
| Fairness pairs OK | 2 / 5 | 5 / 5 |
| Avg tokens (simple / plan) | 136,322 / 198,691 | 175,138 / 157,623 |
| Planner / single token ratio | ~1.34× | ~0.86× |
| Avg iterations (simple / plan) | 14.6 / 38.4 | 14.0 / 27.8 |
Strict AppWorld TGC (judge_pass) |
not cleared (5/5) | not cleared (5/5) |
Avg judge pass_percentage |
(pre-gate noisy) | 41.7% / 56.0% |
| Head-to-head (strict) | ties 5/5 | ties 5/5 |
Caption: plan_fit_5 cohort, n=5 task pairs · fairness and token ratio improved post-gate.
Interpretation: the combined changes improved tool-access parity and lowered measured cost on this slice. The rerun cannot attribute that difference to the gate alone. Strict AppWorld TGC still failed; budgets, tool discovery and redaction merit investigation, but task correctness may also be at issue. These numbers do not evaluate production SRE agents.
Per-task scorecard (post-gate)
| Task | Simple pass% / tokens | Plan pass% / tokens | Strict outcome | Notes |
|---|---|---|---|---|
29caf6f_1 |
50.0% / 110,581 | 50.0% / 124,037 | both fail_judge |
tie; simple cheaper |
3aa1a22_3 |
28.6% / 275,342 | 50.0% / 193,229 | both fail (plan budget_no_eval) |
plan higher pass%, lower tokens |
b0a8eae_3 |
50.0% / 185,992 | 50.0% / 264,356 | both fail_judge |
tie; simple cheaper |
afc0fce_2 |
50.0% / 105,322 | 100.0% ⚠️ / 0 tokens | simple fail_judge; plan budget_no_eval |
telemetry bug — 100% pass% with zero Langfuse tokens is not a victory |
32616b5_1 |
30.0% / 198,455 | 30.0% / 206,495 | both fail_judge |
tie; similar pass% |
Scoreboard on pass% alone: plan higher on 2/5 tasks, though one of those rows has zero-token telemetry and no confirmed judge run; treat that row as an unresolved diagnostic rather than a measured quality gain. Scoreboard on strict TGC: neither mode cleared the bar on this slice; compare modes on fairness and tax first.
Two metrics, two stories
| Metric | Use it for | Do not use it for |
|---|---|---|
pass_percentage |
Diagnosing partial progress, comparing runs after harness fixes | Declaring a routing winner |
judge.success (strict) |
Shippable / benchmark pass gate | Explaining away budget or telemetry gaps |
We saw planner prose claim evaluate PASS while the harness labeled budget_no_eval. We saw 100% pass_percentage with zero Langfuse tokens on one plan row — a telemetry hole, not a victory. Same lesson as evidence-based verification and the agent-done checklist: systems of record vote; narration does not.
This is also why from vibes to contracts separates correctness, consistency, and reliability — partial test passes are not pass^k.
Metric blind spot: spawn count vs real workers
Post-gate logs verified one real worker spawn per task when the handoff gate returned the prior result on repeat create_agent calls (stop_after_successful_handoff behavior).
SSE (server-sent event) aggregates still showed ~3.0 create_agent events per plan run. Those are blocked retries — the planner burning orchestrator iterations on spawn attempts the gate rejected.
Copy for your harness:
| Log | What it tells you |
|---|---|
create_agent_count (SSE) |
Tool invocations, including blocked duplicates |
| Worker session / trace roots | How many subcontractors actually ran |
| Handoff gate return events | When duplicate spawns were suppressed |
Without all three, you will over-count workers and under-estimate wasted planner loops.
Residual blockers (honest ceiling)
Even with fairness 5/5 and lower token tax:
wrong_method_422— wrong documented API names after docs searchbudget_no_eval— iteration cap before evaluate on both modespii_poison— redacted placeholders copied into tool calls (plan, 1/5 tasks)discovery_hit = 0— harness signal still flat despite search tools being called
Strict judge.success did not flip on this five-task rerun. Next increments: discovery quality, eval budgets, and redaction policy — not another spawn-policy tweak alone.
Lessons learned
- Fix fairness before you re-rank modes. 2/5 → 5/5 pairs OK changed the cost story entirely.
- A handoff gate fixes duplicate work, not every benchmark blocker. Fairness and cost improved; strict TGC is still tuned separately.
- Average token ratio can flip sign once duplicate workers stop — do not immortalize one ratio from an unfair run.
pass_percentagewithoutsuccessis a diagnostic, not a product gate.- Count real worker traces, not spawn tool calls — blocked retries still tax the planner.
Next: Running AppWorld Locally for Aiden Agent Evals — reproduce the three-server stack before your next cohort. Then multi-agent vs single-agent quality and token tax and simple vs plan: when to use which.
Related reading
On this site
- Running AppWorld Locally — ops prerequisite
- Fair Agent Evals
- Agent Orchestration Tax
- AI Agent Eval Failure Modes
- How to evaluate AI agents (clarify + zero-tool)
- Multi-agent vs single-agent
- Is the Task Actually Done?
Elsewhere
Acknowledgments. Built with the StackGen Aiden team — the engineers behind the agent runtime and platform this series describes.
🚀 We’re building AI-powered SRE at StackGen. If you’re tired of 3 AM pages and want AI agents that triage incidents, run diagnostics, and draft RCA reports — check out ai.stackgen.com and try our new SRE offering.
FAQ
What is an agent handoff gate in planner-worker evals?
A per-run guard that records the first successful worker handoff and returns that result when the planner tries to spawn again. It stops duplicate AppWorld runs after the subcontractor already finished — without it, orchestration tax and fairness metrics lie.
Why did judge pass_percentage rise but strict success stay at zero?
pass_percentage counts partial test passes inside AppWorld's evaluate harness. success is the strict TGC gate. We saw plan at 56% average pass% vs simple at 41.7% after fixes — a partial diagnostic, while strict success still failed on this five-task slice; a missing judge result and a telemetry gap limit interpretation.
Did the planner become cheaper than single-agent after the gate fix?
On average, yes on this five-task slice: planner tokens were about 0.86× single-agent (down from ~1.34× pre-gate). Per-task it still varies; three tasks were cheaper on single-agent. Do not treat average ratio as a universal routing rule.
Why does create_agent count still show ~3 when only one worker ran?
SSE metrics count blocked spawn tool calls the planner attempted after the handoff gate returned the prior worker result. Log verification showed one real worker per task; the extra counts are wasted orchestrator iterations, not extra subcontractors.
How does this relate to fair agent evals?
Fairness (tool-access parity) must pass before you compare cost or pass%. This sequel fixed fairness from 2/5 to 5/5 pairs OK, then re-measured. See the trilogy starting with fair agent evals before performance.
Stay in the loop — production notes on AI agents, workflows, and SRE.
Low volume — new posts and curated reading lists. Unsubscribe anytime.