Reasoning effort is a model setting that allocates more or less internal computation to a response. It can change latency and behavior, but higher effort does not guarantee a better outcome for an AI agent that uses tools.

We compared settings on a site reliability engineering (SRE) triage job: same alert class, same observability tools, same wall-clock ceiling. The lesson was not “high is smarter.” It was effort is a scarce budget that competes with tool turns, and where you spend it matters more than the global knob.

This post compares three live configurations — blanket high, blanket low, adaptive low→high — plus what broke when we removed the known alert schema and forced discovery.


What changed in the test

  • On a tool-heavy API-gateway error-rate dig, blanket low finished; blanket high hit the wall clock without a usable Summary or completion gate.
  • Letting the parent set effort per child worked: collectors got low, synthesis got high — and dug deeper than blanket low.
  • Adaptive still timed out: spawn fan-out and a missing gate tool on the synthesizer beat the dial.
  • Newer models often refuse tools + thinking on Chat Completions; the runtime must call the Responses API (or try Completions, then Responses after an error), or you pay for failed requests. On hierarchical plan, still split models: thinking model for the planner, normal chat model for workers (hybrid plan · Completions vs Responses).
  • In unknown waters, discovery and negative evidence worked — but correlated ≠ drain-the-cluster, and child completion receipts must count for the parent.

The time ceiling applies to the whole run. Spending more time on each collector turn leaves less time for tool calls and final synthesis. Reserve costly reasoning for decisions where it changes the conclusion, then verify that the run can still complete.


The job

We used a fixed API gateway high error-rate card against a live Grafana/Loki plane: collect locus and shape, falsify tempting dependencies, then close with Theory / Unknowns / Do-this-now behind a completion gate.

That is deliberately tool-heavy: Prometheus Query Language (PromQL) and Loki query language (LogQL) calls, plus query repairs, consume much of the wall clock. If reasoning effort only made prose prettier, we would have wasted the A/B. We wanted to know whether the dial helped finish the method.


Blanket high vs blanket low

Same model family that still accepts tools plus effort on Chat Completions. Same temperature rules that family requires. Same wall-clock ceiling. Only the global reasoning dial changed.

Shape Finished under the ceiling? Completion gate Rough feel
High everywhere No — hit the wall Never called More exploration turns, slower LLM hops, empty close
Low everywhere Yes Called with a correlated theory Cleaner Collect → close; usable operator Summary

What low got right

  • Named a concrete cluster / host / namespace locus
  • Described a step-change error-rate cliff that matched the alert timing
  • Stayed honest that mechanism was correlated, not proven causal
  • Proposed next checks instead of inventing a root cause

What high did instead

  • Burned the budget on broader search and command loops
  • Never reached the gate or operator Summary
  • Looked busy; delivered almost nothing an on-call could act on

Finding 1. In this bounded triage comparison, higher effort did not improve the delivered result. It taxes every model turn. If your product is latency-bounded assist, that tax can erase the dig.

Finding 2. Tool strategy shifted with effort. Low preferred focused observability queries and closed. High preferred broader exploration and never synthesized. The dial changed how the agent spent time, not just how carefully it wrote.

Related: single-agent vs multi-agent — architecture taxes show up the same way. Effort is another tax.


Adaptive: low collectors, high synthesis

Fixed wall-clock budget moves from low-effort collection to high-effort synthesis and a completion gate, with failure branches for fan-out and missing tools

Spend expensive reasoning on synthesis, while reserving time and tools to finish.

Global knobs are blunt. The better product question is: can the parent choose effort per delegated dig?

We taught the orchestrator a portable effort field on child spawn: cheap for bounded evidence collection, expensive for competing-hypothesis synthesis. Omission keeps the model’s configured default. Bad values should fail soft (warn and fall back), not abort the whole investigation because the LLM invented "turbo".

On the same alert class, the parent actually toggled:

Child role Effort chosen
Rule / window / locus / shape / differentials / recovery collectors low
Mechanism synthesis high

What improved vs blanket low

  • Stronger validation that the ratio cliff was a real 5xx increase, not a denominator artifact
  • Concrete gateway status codes in logs (not just a red metric)
  • A telemetry confounder: unfiltered request totals disagreed with status-labelled totals — the kind of thing that steals an afternoon if you miss it

What still failed

  • The parent spawned too many collectors in parallel, then a recovery agent, then high synthesis — effort selection did not cap fan-out
  • The synthesis child sometimes lacked the completion gate tool, so there was no receipt to close on
  • Redaction placeholders on trusted alert ids made rule-scoped digs thrash even while metric/log digs continued

Finding 3. Per-child effort changed what the run inspected, but this adaptive run still timed out. Effort selection is not a substitute for spawn budgets, tool allowlists, and completion contracts (is the task actually done?).


Model routing is part of the dial

While probing newer models for the adaptive config, we hit a sharp platform edge:

  • Some current models reject tools + thinking on /v1/chat/completions
  • They want the Responses API for that combination
  • Completions-only runtimes (ours included, for a while) paid for HTTP 400 errors, or a failed Completions call plus a Responses retry, before the dig even started

So the practical mix became two layers. First, which API: Responses (or Completions then Responses on refuse) when tools and thinking are both required. Second, which model: who is allowed to think hard.

Role Safe shape
Plan / synthesis Reasoning model on Responses, medium thinking unless you measured otherwise
Tool workers / collectors Normal chat model (or thinking off / low) — do not put every worker on high-thinking Responses
Parent / summarizer Normal chat model, or thinking off, when the job is mostly coordination prose

Finding 4. “Turn reasoning up on the newest model” can still be a 400 on Completions, or a slow dig on Responses if every tool turn thinks. Treat model + effort + API as one routing problem. Supporting Responses does not make high effort free (Completions vs Responses · prompt caching).


Unknown waters: discovery without a card

We then removed the supplied alert unique identifier (UID), metric name, label schema, and named failing dependency. The agent had to learn the telemetry vocabulary and change strategy after no-data.

What went well

  • Connectivity checks before inventing an alert identity
  • Label discovery: one aggregation dimension worked; another returned empty and was abandoned
  • Widening search after no-data instead of cosmetic query thrash
  • Keeping negative evidence (healthy auth / API / DB paths) instead of clinging to a favorite villain
  • Closing with a correlated theory and terminal avenues — uncertainty survived synthesis

What still hurt

  • The parent tried to delegate a parent-only tool; repair loops burned retries
  • One high-effort synthesis hop went through the wrong model lane and paid for a Chat Completions rejection before fallback
  • A child accepted the gate, but the parent’s completion contract did not fully credit that receipt — orchestration kept chewing wall clock after useful close text existed
  • First-draft remediation overreached: a newly appearing cluster target looked like a rollout. It can also be telemetry onboarding. Correlated is not permission to drain production traffic

We tightened the prompt afterward: treat new targets as both hypotheses, require a read-only segmented check before change talk, and mark traffic moves as human authority until a causal split appears (curiosity before confidence, be creative — do not invent).

Finding 5. Adaptive effort did not remove the need for disciplined discovery and authority boundaries. Correlation alone cannot authorize a production traffic change.


The principle

Spend thinking budget where it changes the decision — and meter everything else.

Put low / none here Put high here Never confuse with effort
Bounded metric/log collection Competing mechanisms, confounders, synthesis Spawn count, wall clock, Max tool/LLM calls
Parent coordination on tool-hostile models One bounded synthesis pass after evidence exists Completion gate ownership
Retries and repair Causal claims that need careful language Permission to change production

Effort is a scheduling and cost choice, not an independent measure of answer quality.


Builder checklist

  • Global high effort is not the default for tool-heavy on-call assist
  • Parents can set effort per child; invalid values fail soft
  • Collectors are cheap; one synthesis pass is expensive
  • Spawn fan-out and per-child timeouts are capped independently of effort
  • Synthesis owns the completion gate tool — and child receipts count for the parent
  • Parent-only tools cannot be delegated
  • Model routing knows which APIs accept tools + thinking (Chat Completions vs Responses)
  • Workers are not silently mapped to high-thinking Responses just because that API works
  • Correlated findings cannot recommend production traffic changes without human authority
  • New entities are both “change” and “telemetry artifact” until split

Shareable cousins: evidence-gated RCA checklist · agent done checklist.


Where to go next

These runs favor testing effort by role under the actual time ceiling, while measuring both investigation depth and whether an operator-facing answer was delivered.


Debating where to spend reasoning budget in on-call agents? Find me on GitHub or LinkedIn.


StackGen develops AI tools for site reliability engineering (SRE), including incident triage and diagnostic workflows. Product details are at ai.stackgen.com.

FAQ

Does higher reasoning effort make SRE agents better?

Not automatically. On tool-heavy triage, higher effort can burn the time budget on slower model turns and extra exploration, so the agent never finishes with a usable summary.

What is adaptive reasoning effort for multi-agent triage?

Give collectors a cheap thinking budget and reserve expensive reasoning for one synthesis pass after evidence exists — instead of one global setting for every delegated dig.

Why do some new models reject tools plus reasoning on Chat Completions?

Several current OpenAI models only allow tools and thinking together on the Responses API. Chat Completions returns an error unless you turn thinking off or switch APIs. See Completions vs Responses for how to tell which API a run used.

Is the Responses API slower than Chat Completions?

Usually no. Same model without extra thinking takes about the same time on either API. Runs feel slow when the model spends hidden thinking tokens, when every worker thinks hard, or when the client fails on Completions and retries on Responses.

What still breaks after effort dials work?

Too many parallel workers, missing completion tools on the synthesis agent, parent-only tools delegated to children, and child completion receipts that the parent never counts.