Tag: sre
Posts tagged sre:
-
From Vibes to Contracts: How We Rebuilt Agent Evals Around an Industry Standard — From vibes to contracts: how we rebuilt agent evals around eval sets, rubrics vs criteria, a grader stack, and pass^k reliability.
-
Canary First: Lessons from Black-Box Consistency Evals for Live SRE Investigate — Canary-first evals for live SRE investigate: check the canary before burning judge tokens, and never treat a draft RCA as done.
-
AI SRE Agent Benchmarks: Wall Time, Tool Calls, Tokens, and ReAcTree Tax — AI SRE agent benchmarks: wall time, tool calls, tokens, and ReAcTree tax — fair A/B numbers so you know when orchestration is worth the cost.
-
Reasoning Effort Is Not a Free Upgrade for Tool-Heavy Agents — Reasoning effort is not a free upgrade for tool-heavy AI agents. Live SRE A/Bs: blanket high timed out, adaptive low→high dug deeper.
-
Single-Agent vs Multi-Agent Orchestration: How to Choose — Single-agent vs multi-agent for SRE triage: a fair A/B, what each shape wins at, and a decision framework so you stop defaulting to either.
-
What Are SRE AI Agents? — What are SRE AI agents? AI for incident triage, diagnostics, and RCA with bounded autonomy — not a chatbot and not open-ended remediation demos.
-
SRE for Agentic Systems: Why Uptime Isn’t Enough Anymore — SRE for agentic systems: uptime isn’t enough. Judgment SLOs and how to measure agentic drift when the agent can be ‘up’ and still wrong.
-
AI Agent Loop Detection — Don’t Throw Away the Answer — AI agent loop detection can erase a good answer. Preserve the best evidence-backed result when a stalled run ends — don’t throw the work away.
-
AI Agent Root Cause Analysis — Evidence Discarded After the Lead — AI agent root cause analysis fails when a dig finds a lead and discards it — peer noise, wrong fire-time windows, and transcript gates for AI SRE.
-
Be Creative. Don’t Invent. — When an AI SRE agent hits a dead end, be creative — don’t invent. Search harder instead of hallucinating rule IDs, metrics, or a tidy RCA.
-
AI Agent Root Cause Analysis — Curiosity Before Confidence — AI agent root cause analysis for SRE: curiosity before confidence. Soft prompts don’t stop bad RCAs — checklists, hard gates, and batched validation do.
-
Is the Task Actually Done? — Completion Loops for Production Agents — Is the AI agent task actually done? Why production agents need an independent completion check — not a self-graded ‘I’m finished.’
-
When Your AI Agent Scorecard Lies — When your AI agent scorecard lies: measure telemetry quality before you trust reliability, correctness, cost, or latency scores.
-
The Hypothesis Ladder — Ruling Things Out Before You Narrate — Hypothesis-driven AI SRE root cause analysis: climb identity and onset before deploy theories, keep parallel branches, prove first and narrate last.
-
From Demo to Deploy — Failure Modes with Receipts — From demo to deploy: production-ready AI agents need receipts, not fluent demos — evidence gates, HITL tiers, and eval checklists for enterprise pilots.
-
Beyond Confluence Runbooks: Why GitOps Triage Steps Matter in the AI Era — Beyond Confluence runbooks: why GitOps triage steps matter for AI agents — version-controlled procedures that change with your stack.
-
Your RCA Agent Doesn’t Need Another Runbook — It Needs a Map — Your RCA agent doesn’t need another runbook — it needs a map. Topology, gates, and verify-first navigation beat a forty-page notebook.
-
Evidence-Based Verification — Don’t Trust Self-Report, Check the System — Evidence-based verification for AI agents: don’t trust self-report — pull proof from ArgoCD, Datadog, and systems of record, then let Go own pass/fail.
-
Evidence-Gated RCA — Prove, Then Narrate — Evidence-gated RCA for AI SRE agents: prove with receipts, then narrate. Fixed stages, structural evals, and token-aware tool loops.
-
AI Incident Triage for SREs — What Actually Helps On-Call — AI incident triage for SREs — what actually helps on-call versus demo theater. Canonical copy now on StackGen.