Observability Tools Agents Actually Call on Triage
If you are wiring an agent runtime for SRE, the temptation is a fat tool catalog. Our six-combo bench was blunt: the closes that scored full correctness leaned on a short observability menu.
We ran both a single-agent ReAct loop and a hierarchical ReAcTree planner (what is ReAcTree? · PDF). Tool names are easiest to audit on the single-agent stream.
TL;DR
- Live signal tools that mattered: alert rule fetch, firing instances, PromQL/LogQL query, occasional execute_command.
- Context tools that appeared but are not evidence:
search_tools,discover_skills,load_skill, notes. - Common dead ends: empty
search_filepath, absolute paths into truncated tool dumps, LogQL transport timeouts. - Count measurement separately from catalog. Mixing them fools first-turn “progress” gates.
Explain like I’m five
A good detective kit is a flashlight, a notepad, and a door key. Dumping the whole hardware store in their backpack just means they trip on wrenches while the house burns.
The menu that showed up in good closes
Across single-agent seats (where tool names are easiest to audit in the stream):
| Tool family | Role |
|---|---|
*_get_alert_rule |
Confirm expr, for, labels |
*_list_firing_instances |
Anchor fired_at / current series |
*_observability_query |
PromQL slopes + LogQL volume/errors |
*_execute_command |
Denser follow-ups when query wrapper is awkward |
search_tools / discover_skills |
Orientation only |
note / budget checks |
Bookkeeping |
Hierarchical ReAcTree streams sometimes collapse tool names in the parent seat when the parent measures first and children dig later. Do not read “0 grafana_* in parent log” as “no measurement happened.” Check child or merged tool spans in your session traces.
Dead ends worth instrumenting
From the same runs:
search_filewith empty path — instant reject. Preview reasoning seats hit this early.search_contenton absolute/tmp/...dump paths — path policy correctly refuses; agent retries waste turns unless the host returns a relative dump id.- LogQL 504 / transport errors — look like “try again with different args” to a high-thinking model. Treat typed
failed/ empty envelopes as upstream failure, not soft success.
Design rules for the catalog
- Pin Collect tools by exact name for evals (fair agent evals).
- Separate retrieval / catalog classes from live-signal classes in your loop detectors.
- Prefer one good query primitive over five overlapping wrappers.
- Truncated-output tools must accept the ids your summarizer emits, not absolute host paths.
Related: What is ReAcTree?, single-agent vs multi-agent, loop detection, empty query ≠ absent signal.
FAQ
Which Grafana tools mattered on the dual-part alert+logs job?
get_alert_rule, list_firing_instances, observability_query (metrics and LogQL), and execute_command for denser probes. Catalog tools (search_tools, discover_skills) showed up but did not count as measurement.
Do agents need dozens of observability tools exposed?
Not for this job. Successful closes clustered on a handful of query primitives. Extra catalog and filesystem tools mostly produced invalid-path and empty-path failures.
What failed repeatedly?
Absolute paths into truncated tool-output stores in search_content, empty path on search_file, and transport timeouts on heavy LogQL. Hosts should treat those as typed failures, not successes.
Stay in the loop — production notes on AI agents, workflows, and SRE.
Low volume — new posts and curated reading lists. Unsubscribe anytime.