Stop Retrying the Same Failed Observability Query
Prompt poetry does not stop a high-thinking model from poking a dead PromQL with slightly different labels.
We already had loop detection and habits that prefer parent measurement before spawning children (single vs multi). They were not enough. The missing piece is observation identity: if the world did not change, stop paying for another thought.
Shapes in the rematch: single-agent ReAct vs hierarchical ReAcTree (what is ReAcTree? · PDF).
TL;DR
- Typed query failures (
n/outcome: failed, all-failed arrays) must count as failures, not transport success. - Hash terminal failed/empty observations; identical hashes → steer once → halt with partial findings.
- Cap same-tool fan-out in one wave — even for retrieval tools.
- Context-only tools (
search_tools,load_skill, notes) must not satisfy “live evidence” gates. - On our rematch, the big product win was quality on a reasoning-preview hierarchical seat (scorecard), not a free latency cut.
Explain like I’m five
If the smoke detector keeps saying “battery dead” the same way, write it down once, try another room, then stop pressing the same button.
What we saw without the gates
On a dual-part alert+logs job, a Responses hierarchical seat finished “fast” with:
- corr 0.69
- both parts Undetermined
- almost no measured Part B numbers
Single-agent on the same model scored 1.0 and took longer. The planner was not “more careful.” It was done pretending.
Meanwhile generate and xAI seats already closed full Theories. The failure mode is concentrated where thinking budget is large and host feedback is polite.
Gates that are not prompts
| Gate | Behavior |
|---|---|
| Typed failure class | All-failed query envelope → upstream failure |
| No-progress ledger | Same failed/empty observation hash twice → steer, then stop |
| Fan-out cap | Nth parallel copy of the same tool → hard error |
| Tool classes | Catalog/skill/notes ≠ live measurement |
Do not reuse “required completion tool names” for gain-exhausted tools. That field means something else; mixing it drops required gates when a plane is exhausted.
What still lies to you
- Tools that return HTTP 200 without
n/outcomestill look healthy. Fix the envelope at the integration. - Contended tool servers make everyone slower. Gates do not delete queueing.
- Valid credentials + broken upstream image can keep stale tools until refresh.
Monday checklist
- Log observation hashes for failed/empty query tools.
- Alert when steer-then-halt fires more than N times per session.
- Re-run one golden dual-part prompt after every gate change.
- Read the Theory — if it says Undetermined with no blocked query named, fail the build.
Related: empty query is data, deliver findings at the budget cap, how models write Theories.
FAQ
Why do reasoning models retry the same failed PromQL or LogQL?
High-thinking seats treat HTTP 200 typed failures as soft success and tweak args forever. Prompt text alone does not stop that. The host must classify failed/empty observations and halt after one steer.
What host gates help?
Treat all-failed query envelopes as upstream failure, hash terminal failed/empty observations so identical failures match, steer once then stop with partial findings, and cap parallel copies of the same tool in one wave.
Did these gates make our bench faster?
Not on absolute wall for a contended six-way Grafana rematch. They did coincide with a hierarchical ReAcTree Responses seat jumping from 0.69 to 1.0 correctness on the same prompt.
Stay in the loop — production notes on AI agents, workflows, and SRE.
Low volume — new posts and curated reading lists. Unsubscribe anytime.