SRE on-call starter pack
A short path for on-call and platform SREs evaluating AI agents. No blueprint — problem → lesson → when not to trust the demo.
The five posts
- What Are SRE AI Agents? — triage vs RCA vs remediation without theater
- AI Incident Triage for SREs — what actually helps on-call (on StackGen)
- Evidence-Gated Multi-Plane RCA — prove before you narrate
- You Can’t Debug What You Can’t See — why APM misses agent failures (CNCF reprint)
- Is the Task Actually Done? — completion checks that don’t self-grade
Pocket checklist
Downloadable principles (no proprietary schemas): Evidence-gated RCA checklist · “Done” checklist
Next
- Hub: AI agents for SRE · AI incident triage
- Full series: Building an Enterprise AI Agent Platform in Go
- Runtime basics: Go / runtime starter pack
Stay in the loop — production notes on AI agents, workflows, and SRE.
Low volume — new posts and curated reading lists. Unsubscribe anytime.
FAQ
What should an SRE read first about AI agents?
Start with what SRE AI agents are, then incident triage that helps on-call, then evidence-gated RCA and observability. Skip autonomous remediation demos until triage and receipts are solid.
How long is this reading pack?
Five posts. Most readers finish in one sitting if they skim TL;DRs; deeper reads take an evening.