A short path for on-call and platform SREs evaluating AI agents. No blueprint — problem → lesson → when not to trust the demo.

The five posts

  1. What Are SRE AI Agents? — triage vs RCA vs remediation without theater
  2. AI Incident Triage for SREs — what actually helps on-call (on StackGen)
  3. Evidence-Gated Multi-Plane RCA — prove before you narrate
  4. You Can’t Debug What You Can’t See — why APM misses agent failures (CNCF reprint)
  5. Is the Task Actually Done? — completion checks that don’t self-grade

Pocket checklist

Downloadable principles (no proprietary schemas): Evidence-gated RCA checklist · “Done” checklist

Next

FAQ

What should an SRE read first about AI agents?

Start with what SRE AI agents are, then incident triage that helps on-call, then evidence-gated RCA and observability. Skip autonomous remediation demos until triage and receipts are solid.

How long is this reading pack?

Five posts. Most readers finish in one sitting if they skim TL;DRs; deeper reads take an evening.