Rule for this benchmark run: once a delegated worker returns successfully, later create_agent (launch-worker) attempts return that result instead of launching duplicate workers. This avoids repeating the same AppWorld work. It does not remove the planner turns spent attempting those launches, and a production system may need a way to start a genuinely different follow-up job.

We fixed fair agent evals and cut agent orchestration tax on a five-task AppWorld slice — the wins were in measurement and cost, not in declaring a benchmark champion. Strict judge pass stayed 0/5 on this slice; the gate did not “win AppWorld,” it stopped invalid comparisons.

Parts one, two, and three documented unfair handoffs, ~1.6× token overhead, and failure classes. This sequel covers what changed after we shipped the handoff gate, re-enabled note tooling, soft-dropped missing infra names on spawn, and capped worker nodes to one.

Dataset unchanged. Judge unchanged (AppWorld /evaluate, strict success bit). Observability: Langfuse aggregates only.

Handoff gate blocking duplicate worker spawns after successful subcontractor return


What changed in the five-task rerun

With the worker tools registered, optional missing names handled, and duplicate launches blocked, 5/5 pairs passed tool-access checks versus 2/5 before. The planner-to-single-agent average token ratio moved from about 1.34× to 0.86×; the planner averaged 27.8 turns instead of 38.4. Multiple changes happened together, and the single-agent token average changed too, so this is not an isolated estimate of the gate’s effect. Neither mode passed strict AppWorld task-goal completion on the rerun. The practical check is to count actual worker sessions as well as attempted launch calls.

Sequel: How to evaluate AI agents: clarify + zero-tool failures — workers that report success without app calls.


What we fixed (behavior, not blueprint)

Coordinator sends one worker, records a successful return, rejects a duplicate spawn, then separates spawn attempts from actual workers and judge outcome

Caption: The handoff gate avoids duplicate workers; spawn-attempt telemetry and external task success remain separate measures.

After failure-mode runs on the same plan_fit_5 tasks, we addressed harness and runtime issues that made planner mode look worse than it was:

Fix What it does
Note tools registered Infra note/read tools available so spawn payloads do not hard-fail on names the registry lacked
Soft-drop on spawn Optional infra names dropped instead of aborting the whole create_agent call
Successful handoff gate First completed worker handoff is stored; repeat spawn attempts return that result instead of launching another AppWorld runner
Worker node cap = 1 Planner tree cannot grow a second worker branch on these eval configs
Unattended mode No clarify / human-wait tools in benchmark runs

Same five delegation-fit tasks (phone → notes → SMS, inbox + contacts + Splitwise, workout note → Spotify, batch Venmo, trip ledger → Splitwise). Same model family and MCP tool pack.

Handoff gate sequence: one real worker, blocked duplicate spawn retries


Before vs after (aggregate)

Metric Pre-gate cohort Post-gate (plan-fit5-gate)
Fairness pairs OK 2 / 5 5 / 5
Avg tokens (simple / plan) 136,322 / 198,691 175,138 / 157,623
Planner / single token ratio ~1.34× ~0.86×
Avg iterations (simple / plan) 14.6 / 38.4 14.0 / 27.8
Strict AppWorld TGC (judge_pass) not cleared (5/5) not cleared (5/5)
Avg judge pass_percentage (pre-gate noisy) 41.7% / 56.0%
Head-to-head (strict) ties 5/5 ties 5/5

Fairness pairs OK and planner token ratio before vs after handoff gate

Caption: plan_fit_5 cohort, n=5 task pairs · fairness and token ratio improved post-gate.

Interpretation: the combined changes improved tool-access parity and lowered measured cost on this slice. The rerun cannot attribute that difference to the gate alone. Strict AppWorld TGC still failed; budgets, tool discovery and redaction merit investigation, but task correctness may also be at issue. These numbers do not evaluate production SRE agents.


Per-task scorecard (post-gate)

Task Simple pass% / tokens Plan pass% / tokens Strict outcome Notes
29caf6f_1 50.0% / 110,581 50.0% / 124,037 both fail_judge tie; simple cheaper
3aa1a22_3 28.6% / 275,342 50.0% / 193,229 both fail (plan budget_no_eval) plan higher pass%, lower tokens
b0a8eae_3 50.0% / 185,992 50.0% / 264,356 both fail_judge tie; simple cheaper
afc0fce_2 50.0% / 105,322 100.0% ⚠️ / 0 tokens simple fail_judge; plan budget_no_eval telemetry bug — 100% pass% with zero Langfuse tokens is not a victory
32616b5_1 30.0% / 198,455 30.0% / 206,495 both fail_judge tie; similar pass%

Scoreboard on pass% alone: plan higher on 2/5 tasks, though one of those rows has zero-token telemetry and no confirmed judge run; treat that row as an unresolved diagnostic rather than a measured quality gain. Scoreboard on strict TGC: neither mode cleared the bar on this slice; compare modes on fairness and tax first.


Two metrics, two stories

Metric Use it for Do not use it for
pass_percentage Diagnosing partial progress, comparing runs after harness fixes Declaring a routing winner
judge.success (strict) Shippable / benchmark pass gate Explaining away budget or telemetry gaps

We saw planner prose claim evaluate PASS while the harness labeled budget_no_eval. We saw 100% pass_percentage with zero Langfuse tokens on one plan row — a telemetry hole, not a victory. Same lesson as evidence-based verification and the agent-done checklist: systems of record vote; narration does not.

This is also why from vibes to contracts separates correctness, consistency, and reliability — partial test passes are not pass^k.


Metric blind spot: spawn count vs real workers

Post-gate logs verified one real worker spawn per task when the handoff gate returned the prior result on repeat create_agent calls (stop_after_successful_handoff behavior).

SSE (server-sent event) aggregates still showed ~3.0 create_agent events per plan run. Those are blocked retries — the planner burning orchestrator iterations on spawn attempts the gate rejected.

Copy for your harness:

Log What it tells you
create_agent_count (SSE) Tool invocations, including blocked duplicates
Worker session / trace roots How many subcontractors actually ran
Handoff gate return events When duplicate spawns were suppressed

Without all three, you will over-count workers and under-estimate wasted planner loops.


Residual blockers (honest ceiling)

Even with fairness 5/5 and lower token tax:

  • wrong_method_422 — wrong documented API names after docs search
  • budget_no_eval — iteration cap before evaluate on both modes
  • pii_poison — redacted placeholders copied into tool calls (plan, 1/5 tasks)
  • discovery_hit = 0 — harness signal still flat despite search tools being called

Strict judge.success did not flip on this five-task rerun. Next increments: discovery quality, eval budgets, and redaction policy — not another spawn-policy tweak alone.


Lessons learned

  1. Fix fairness before you re-rank modes. 2/5 → 5/5 pairs OK changed the cost story entirely.
  2. A handoff gate fixes duplicate work, not every benchmark blocker. Fairness and cost improved; strict TGC is still tuned separately.
  3. Average token ratio can flip sign once duplicate workers stop — do not immortalize one ratio from an unfair run.
  4. pass_percentage without success is a diagnostic, not a product gate.
  5. Count real worker traces, not spawn tool calls — blocked retries still tax the planner.

Next: Running AppWorld Locally for Aiden Agent Evals — reproduce the three-server stack before your next cohort. Then multi-agent vs single-agent quality and token tax and simple vs plan: when to use which.

On this site

Elsewhere


Acknowledgments. Built with the StackGen Aiden team — the engineers behind the agent runtime and platform this series describes.

🚀 We’re building AI-powered SRE at StackGen. If you’re tired of 3 AM pages and want AI agents that triage incidents, run diagnostics, and draft RCA reports — check out ai.stackgen.com and try our new SRE offering.

FAQ

What is an agent handoff gate in planner-worker evals?

A per-run guard that records the first successful worker handoff and returns that result when the planner tries to spawn again. It stops duplicate AppWorld runs after the subcontractor already finished — without it, orchestration tax and fairness metrics lie.

Why did judge pass_percentage rise but strict success stay at zero?

pass_percentage counts partial test passes inside AppWorld's evaluate harness. success is the strict TGC gate. We saw plan at 56% average pass% vs simple at 41.7% after fixes — a partial diagnostic, while strict success still failed on this five-task slice; a missing judge result and a telemetry gap limit interpretation.

Did the planner become cheaper than single-agent after the gate fix?

On average, yes on this five-task slice: planner tokens were about 0.86× single-agent (down from ~1.34× pre-gate). Per-task it still varies; three tasks were cheaper on single-agent. Do not treat average ratio as a universal routing rule.

Why does create_agent count still show ~3 when only one worker ran?

SSE metrics count blocked spawn tool calls the planner attempted after the handoff gate returned the prior worker result. Log verification showed one real worker per task; the extra counts are wasted orchestrator iterations, not extra subcontractors.

How does this relate to fair agent evals?

Fairness (tool-access parity) must pass before you compare cost or pass%. This sequel fixed fairness from 2/5 to 5/5 pairs OK, then re-measured. See the trilogy starting with fair agent evals before performance.