If you are trying to evaluate AI agents on a tool-using benchmark, the first question is not “which mode wins?” It is whether both modes could use the same tools on the worker that does the work.

We ran paired evals on AppWorld — a controllable multi-app benchmark (paper) — through our agent runtime in Aiden, with domain APIs exposed via Model Context Protocol (MCP). The early cost comparison was misleading. The planner path looked cheaper until we found that its delegated workers often made zero calls to the app APIs, while the single-agent path actually attempted the task.

This post is part one of the AppWorld agent eval series: fix tool parity, then measure orchestration tax, then classify failure modes. Later posts cover handoff gates, local setup, unattended harness rules, MCP tool tax + pass@k, and simple vs plan routing. Part two: orchestration tax. Part three: failure modes.

AppWorld agent eval series: fair evals, orchestration tax, failure modes, handoff gate

Fair agent evals: single-agent with tools vs planner worker without domain MCP access


What this series measures (and what it does not)

This is a methodology series about how to run fair planner-vs-single-agent evals — not a claim that AppWorld scores predict production outcomes.

Lens What we measured What Aiden ships in production
Benchmark AppWorld — multi-app API mutation tasks (phone, payments, notes, …) SRE triage, evidence-gated RCA, operator workflows
Goal of the run Calibrate fairness gates, orchestration tax, failure classes Incident response, diagnostics, audit trails
Sample size 5–10 tasks per cohort (directional) Customer environments with different SLOs

On these small AppWorld slices, strict TGC/SGC success (judge.success) did not clear — both modes often landed at 30–50% partial pass_percentage with harness blockers (budget, API discovery, PII redaction in eval mode). That tells us where the eval harness still needs work. It does not mean the runtime fails at jobs it was built for. When we say “judge did not pass,” read it as “this benchmark bar on this slice” — the same way a unit test red does not mean the whole product is broken.


The comparison that failed first

In an early three-task check, a single agent averaged about 14 calls to the benchmark apps while the planner’s workers made none. The planner could still produce a convincing account of completion, but the app state had not changed. After tool-access wiring was fixed, all 10/10 pairs in a mixed ten-task cohort passed the access check; workers averaged 20.5 app calls versus 13.8 for single-agent. That is evidence that both paths could attempt the task, not that either passed it: neither cleared strict AppWorld task-goal completion on this slice. Before comparing cost, inspect the actual worker’s tool list and the judge result.


What we were actually comparing

Two execution shapes on the same AppWorld tasks:

Shape Who calls domain APIs Typical orchestration
Single-agent loop Root agent holds the full MCP tool pack One plan → tool → context cycle
Planner + worker Meta root delegates; worker should hold the MCP pack Spawn subcontractor with explicit tool_names

The benchmark judge is AppWorld’s own evaluate step (TGC/SGC mean task-goal completion and scenario-goal completion: checks of the resulting app state). We treat that as ground truth, not assistant prose. Telemetry came from Langfuse generation counts (aggregates only in this write-up).

This is the eval sequel to single-agent vs multi-agent orchestration — same “match the axes” discipline, applied to a public tool benchmark instead of a single-plane triage dig.


The unfair run (why the planner looked cheap)

On an early three-task smoke, averages looked like this:

Mode Avg domain API calls Avg worker spawns What actually happened
Single-agent ~14 0 Linear load → login → call → evaluate
Planner 0 ~1.3 Root spawned workers; workers often had no domain tools

The planner path burned tokens on create_agent / search_tools while never calling the apps under test. In at least one run the transcript claimed export success and evaluation pass with no tool receipts — classic self-report without a judge call.

Interpretation: this does not show that delegation saves resources. The two paths were not doing comparable work.

Eval pipeline: fairness gate before comparing tokens or modes

Bar chart: domain API calls per run before handoff fix (single-agent ~14, planner ~0) vs after fix on n=10 cohort (13.8 vs 20.5)

Caption: After the handoff fix, both paths call domain APIs. Strict TGC on this slice: see benchmark context above.


The fairness gate (what “fair” means here)

We defined tool-access fairness separately from task success:

  1. Single-agent: domain MCP tools present; no worker spawn tool on the root.
  2. Planner: at least one worker spawn; worker registry includes the full domain MCP pack; no hard-fail because optional infra names were missing from the registry.
  3. Unattended: no clarify / permission prompts (eval mode must not wait for a human).
  4. Judge: evaluate result parsed; fallback DB scoring labeled when the agent stopped on budget.

On a two-task fair canary, 2/2 pairs passed fairness. Domain calls averaged 12 (single-agent) vs 31.5 (planner workers) — both sides touched the APIs.

On a ten-task mixed cohort (simple-band and advanced-band tasks), 10/10 pairs passed fairness. Averages: 13.8 vs 20.5 domain calls, 14.6 vs 43.8 model iterations, ~1.6× tokens. Neither mode cleared strict AppWorld TGC on this harness slice — expected while budgets, discovery, and eval redaction policy are still being tuned.

Fairness does not mean matched compute budget (planner runs allowed slightly higher caps and worker node limits in this harness). It means both sides could do the work.


Copy-paste fairness scorecard

Before you compare planner vs single-agent on any tool benchmark, record:

# Check Pass?
1 Worker (or root) registry lists every domain tool the task needs  
2 Planner root is not required to call domain APIs if design is meta-only  
3 Spawn payload does not hard-fail on benign extra tool names  
4 Parent telemetry does not count subagent tools as root tools (or you split roles in analysis)  
5 Unattended mode — no HITL / clarify stalls  
6 External judge invoked and parsed; prose PASS ≠ judge PASS  

If row 1 fails, stop. Fix wiring, then rerun. Tokens and “wins” before that are noise.


Lessons learned

  1. Cheaper can mean “did not run the benchmark.” Zero domain calls is an abort signal, not a routing win.
  2. Fairness is a gate, not a score. Passing tool-access checks is prerequisite to comparing modes — it does not by itself clear a hard multi-app benchmark bar.
  3. Handoff bugs look like mode failures. “Planner has no APIs” was registry wiring, not a law of trees.
  4. Self-report is not evaluate. Fluent PASS lines without judge receipts are a failure class, not a tie-breaker.
  5. The method generalizes. A planner-worker runtime needs an access check before you trust a comparison; the measured ratios here do not automatically transfer to another runtime.

On this site

Elsewhere


Acknowledgments. Built with the StackGen Aiden team — the engineers behind the agent runtime and platform this series describes.


🚀 We’re building AI-powered SRE at StackGen. If you’re tired of 3 AM pages and want AI agents that triage incidents, run diagnostics, and draft RCA reports — check out ai.stackgen.com and try our new SRE offering.

FAQ

What makes an AI agent evaluation fair when comparing planner vs single-agent?

Both paths must reach the same domain tools on the worker that actually mutates state. The planner root may stay meta-only, but the subcontractor must get the full MCP pack. Also match unattended mode, judge parsing, budgets, and record whether tools were hard-denied vs soft-dropped.

Why did our planner look cheaper before we fixed the handoff?

The planner path spent budget on spawn and repair while workers called zero domain APIs. That is a wiring bug, not evidence that orchestration is efficient. Cheap runs that never touch the benchmark APIs are invalid comparisons.

How do you evaluate AI agents that use MCP tools?

Log tool-access parity first: which names were on the root vs worker registry, whether create-worker calls hard-failed on missing names, and whether parent telemetry double-counts subagent calls. Only then compare tokens, iterations, and external judge pass rate.

Does AppWorld judge pass mean the agent succeeded?

AppWorld publishes TGC/SGC through its own evaluate harness — not your runtime's self-report. Strict success is a high bar on multi-app mutation tasks. This series uses AppWorld to calibrate eval fairness and failure taxonomy; it is not a scorecard for Aiden's production SRE workflows.

Is this only about one agent runtime?

No. Any planner-worker shape — LangGraph, CrewAI, AutoGen, OpenAI Agents, or an in-house tree — needs the same fairness gate before you compare modes.