Tag: sre
Posts tagged sre:
-
Same Incident, Five Models: Rank Investigators Without Mixing the Exam — Random CI reruns mix model skill with incident difficulty. Freeze one checkout outage, swap only the bound model, and rank on Correct first, then wall time and new input tokens.
-
Session Efficiency Should Not Beat Accuracy — We labeled a score 0-100 and watched a perfect, cheap, fast agent run print 125. Cap budget bonuses at 1.0 so efficiency never grades above how often you were Correct.
-
Hybrid Plan Mode: Smart Planner, Generate Digs — Plan roots already route to planning; digs default to tool calling. Putting a high-thinking reasoning model on every task makes workers slow. Use a thinking model for the planner and a normal chat model for digs.
-
AI Agent Runtime: What to Measure Before You Buy — Buyer checklist for an AI agent runtime: loop, tools, gates, and an eval receipt — with numbers from a real multi-model triage bench.
-
How to Evaluate an AI Agent for Root Cause Analysis — A checklist-first RCA eval: Theory, Unknowns, next action, measured signal, honesty about gaps — plus wall, tokens, and cost. Built from a live alert+logs bench.
-
Reasoning vs Generate Models for Tool-Heavy Agents — Merits and demerits of reasoning vs normal chat models on a live Grafana triage job — elapsed time, tokens, Theory style, and when each earns the bill.
-
Hierarchical vs Single-Agent on a Dual-Part Observability Job — When a hierarchical ReAcTree planner earns its tax on alert+logs triage — and when a single-agent ReAct loop is the better default. Real ratios from a six-combo bench.
-
Observability Tools Agents Actually Call on Triage — From live combo benches: the short Grafana/Loki tool set that showed up in successful closes — and the dead ends that burned turns.
-
What Reasoning Models Write When You Ask Them to Triage — Same alert+logs prompt, three model classes. How their Theories diverge: leak vs measurement artifact, and what that means for operators.
-
Six Model×Orchestration Combos on One Alert+Logs Job — Same dual-part observability prompt across three model classes and single-agent ReAct vs hierarchical ReAcTree. Wall, tokens, cost, correctness, and an efficiency score you can steal.
-
What Is ReAcTree? Hierarchical Agent Trees for Long-Horizon Work — Plain English: ReAct is one reason→act→observe loop. ReAcTree is a tree of agent nodes with sequence, parallel, and fallback control flow — for SRE-style tool-heavy jobs.
-
Cursor paging for spilled agent tool output — When observability tool results spill to disk, agents should page with opaque cursors, not grep the preview and declare every pod Unavailable.
-
Weekly Reflection: Context Engineering Beat the Language Wars — Weekly reflection: context engineering over language wars, plus public libraries and servers for write/select/compress/isolate with multi-language or MCP clients.
-
Observation Masking Cut Token Cost 43%—Then Lost the RCA — Context engineering A/B: observation masking (tool-result clearing) made an AI SRE agent ~43% cheaper—then failed equal-evidence RCA. Scorecard + knobs.
-
Debug AI Agent Failures: One Zip Per Conversation — Support said the AI investigation went wrong. Give one conversation debug zip—not per-execution scavenger hunts—then grade batches into product gates.
-
Why AI SRE Feels Stuck Before the First Tool Call — On-call sees ‘the bot is stuck’ while vault re-checks and MCP catalog re-index burn cold start. Measure time to the first useful tool call.
-
From Vibes to Contracts: How We Rebuilt Agent Evals Around an Industry Standard — From vibes to contracts: how we rebuilt agent evals around eval sets, rubrics vs criteria, a grader stack, and pass^k reliability.
-
Canary First: Lessons from Black-Box Consistency Evals for Live SRE Investigate — Canary-first evals for live SRE investigate: check the canary before burning judge tokens, and never treat a draft RCA as done.
-
AI SRE Agent Benchmarks: Wall Time, Tool Calls, Tokens, and ReAcTree Tax — AI SRE agent benchmarks: wall time, tool calls, tokens, and ReAcTree tax — fair A/B numbers so you know when orchestration is worth the cost.
-
Reasoning Effort Is Not a Free Upgrade for Tool-Heavy Agents — Reasoning effort is not a free upgrade for tool-heavy AI agents. Live SRE A/Bs: blanket high timed out, adaptive low→high dug deeper.
-
Single-Agent vs Multi-Agent Orchestration: How to Choose — Single-agent vs multi-agent for SRE triage: a fair A/B, what each shape wins at, and a decision framework so you stop defaulting to either.
-
What Are SRE AI Agents? — What are SRE AI agents? AI for incident triage, diagnostics, and RCA with bounded autonomy — not a chatbot and not open-ended remediation demos.
-
Don’t Invent PromQL: Measure the Alert Rule Query First — Alert title said latency; the rule was ClickHouse. AI SRE agents that invent PromQL from titles ship wrong RCA. Measure the stored query first.
-
SRE for Agentic Systems: Why Uptime Isn’t Enough Anymore — SRE for agentic systems: uptime isn’t enough. Judgment SLOs and how to measure agentic drift when the agent can be ‘up’ and still wrong.
-
AI Agent Loop Detection — Don’t Throw Away the Answer — AI agent loop detection can erase a good answer. Preserve the best evidence-backed result when a stalled run ends — don’t throw the work away.
-
Empty PromQL ≠ Missing Data: Fix AI SRE Scope Blindness — A Grafana no_data result applies to one query and time range; check labels, time, data source, and tool failures before reporting missing evidence.
-
AI Agent Root Cause Analysis — Evidence Discarded After the Lead — An incident agent can collect a useful metric lead and omit it from the final report; check time windows, unrelated alerts, and transcript-to-summary consistency.
-
Be Creative. Don’t Invent. — When an incident agent reaches an empty result, broaden searches using recorded identifiers instead of guessing rule IDs or measurements.
-
Ungrounded Synthesis Must Read as Hypothesis — If a check finds unsupported details in an AI-written incident analysis, correct the primary message rather than hiding the warning in a note.
-
AI Agent Root Cause Analysis — Curiosity Before Confidence — For AI-assisted incident analysis, record required checks, return missing fields together, and limit confidence to what the evidence supports.
-
Is the Task Actually Done? — Completion Loops for Production Agents — Check an AI agent’s deliverable and tool outcomes before reporting a goal as complete; return unverified work explicitly when checking stops.
-
When Your AI Agent Scorecard Lies — An AI agent scorecard can grade the wrong work; check trace identity, coverage, and evaluations before using its reliability or quality scores.
-
AI Agent Hit Max Turns? Deliver Partial RCA, Not Apology — When an incident agent reaches a model-call limit, preserve observed findings and label a partial analysis incomplete.
-
The Hypothesis Ladder — Ruling Things Out Before You Narrate — A structured way to investigate service incidents: identify the affected system and onset, test competing explanations, and report uncertainty.
-
From Demo to Deploy — Failure Modes with Receipts — Questions for evaluating AI incident-response pilots: recorded evidence, staged checks, approval policies, and test results beyond a fluent demo.
-
Beyond Confluence Runbooks: Why GitOps Triage Steps Matter in the AI Era — Beyond Confluence runbooks: why GitOps triage steps matter for AI agents — version-controlled procedures that change with your stack.
-
Your RCA Agent Doesn’t Need Another Runbook — It Needs a Map — Your RCA agent doesn’t need another runbook — it needs a map. Topology, gates, and verify-first navigation beat a forty-page notebook.
-
AI Agents Call Truncated Grafana ‘No Data’—It’s Spill — Large tool results get preview-truncated; AI SRE agents invent ‘Unavailable.’ Spill recovery and COMPLETE/PARTIAL/FAILED fix dishonest RCA.
-
How We Debug Multi-Stage AI Agent Workflows — How to debug multi-step and multi-stage AI agent workflows and execution logs — green one stage at a time, score tool effects not transcripts.
-
Evidence-Based Verification — Don’t Trust Self-Report, Check the System — Evidence-based verification for AI agents: don’t trust self-report — pull proof from ArgoCD, Datadog, and systems of record, then let Go own pass/fail.
-
Evidence-Gated RCA — Prove, Then Narrate — Evidence-gated RCA for AI SRE agents: prove with receipts, then narrate. Fixed stages, structural evals, and token-aware tool loops.
-
AI Incident Triage for SREs — What Actually Helps On-Call — AI incident triage for SREs — what actually helps on-call versus demo theater. Canonical copy now on StackGen.
-
When the Operator Asks to Correlate, Make It a Gate — Natural-language correlation goals need server-side gates — not hope the LLM remembers to search prior incidents.
-
Same Alert, Different Verdict: Entry Path Is Context — Don’t paste the alert in the UI and wonder why Slack gave a different impact score — entry path carries investigation context.
-
Slack Is a Triage Board, Not a Log Dump — Make incident replies in Slack scannable: show counts, findings and uncertainty, with links to evidence and searchable Activity.
-
Stop Re-Investigating the Same Alert — Reuse recent incident findings for follow-ups on the same alert, with a clear way to investigate again when evidence changes.
-
Service Rendered Efficiently: SRE AI Is Not an Engineering Credibility Project — Service Rendered Efficiently: judge AI investigation by whether it helps on-call reuse results, understand uncertainty, and hand off work.