Building an Enterprise AI Agent Platform in Go
This series follows our work on an AI agent runtime—the software that manages a language model’s tool calls, state, and retries—and Aiden, StackGen’s platform for running agents across teams. Many examples involve site reliability engineering (SRE): investigating incidents and keeping services available. The posts mix implementation details, experiments, and lessons from things that went wrong. Results from our setup may not transfer directly to yours; the useful part is often the test or tradeoff, not our choice.
Start with a pack
| Pack | For |
|---|---|
| Go agent runtime | Runtime definition, Go vs Python, platform split |
| SRE on-call | Triage, RCA, observability |
| SRE as service | Service Rendered Efficiently culture pack |
| Evidence-gated RCA checklist | Operator review of agent write-ups |
| “Done” checklist | When not to trust agent completion |
Topic hubs
Dive by theme:
- AI agent workflows — multi-stage pipelines, bring-up, evidence-gated RCA
- AI agents for SRE — incident triage, observability, tokenomics
- Service Rendered Efficiently — AI investigation as a service product (sibling series)
- Go AI agents — language choice, platform architecture, IaC config
- AI agent runtime — loop vs platform
Suggested for you
-
Start here
Python vs Go for AI Agents — Why We Chose Go
Go vs Python for AI agents: why we chose Go for a production agent runtime — concurrency, single-binary ops...
-
Then read
TOML Over YAML and PKL — How We Stopped Fighting Config and Started Shipping
Why we chose TOML for flat, typed agent configuration after considering YAML, PKL and CUE, and when those a...
-
Then read
Go Platform Architecture at Speed — Without Drowning
Go platform architecture for a production AI agent codebase — patterns that keep rapid development sustaina...
Every post, month by month
76 posts across 5 months.
October 2026 6 posts
-
Keep live agent trials deliberate, report missing tasks and failed attempts, and match trial IDs to traces before trusting a CI result.
-
Test the brief passed from a worker to its parent: settled facts, open gaps, size limits, and a clear stop condition.
-
Detect stalled agent loops from outcomes and evidence, not repeated tool names; test the stop rule with stubbed calls.
-
Check verifiable facts with code and use a calibrated model judge for questions that require reading and interpretation.
-
Use fast tests for agent bounds, contract checks for facts, and a small number of live trials for whole jobs.
-
Separate the agent from its grader, check facts with code, and run the installed build in the trial environment.
September 2026 16 posts
-
Random CI reruns mix model skill with incident difficulty. Freeze one checkout outage, swap only the bound model, and rank on Correct first, then wall time and new input tokens.
-
We labeled a score 0-100 and watched a perfect, cheap, fast agent run print 125. Cap budget bonuses at 1.0 so efficiency never grades above how often you were Correct.
-
OpenAI Chat Completions and the Responses API are two ways to call a model. Reasoning models are a different choice. How to tell them apart in a debug export, and why Responses is not automatically slower.
-
Plan roots already route to planning; digs default to tool calling. Putting a high-thinking reasoning model on every task makes workers slow. Use a thinking model for the planner and a normal chat model for digs.
-
Buyer checklist for an AI agent runtime: loop, tools, gates, and an eval receipt — with numbers from a real multi-model triage bench.
-
A checklist-first RCA eval: Theory, Unknowns, next action, measured signal, honesty about gaps — plus wall, tokens, and cost. Built from a live alert+logs bench.
-
Merits and demerits of reasoning vs normal chat models on a live Grafana triage job — elapsed time, tokens, Theory style, and when each earns the bill.
-
A practical debug loop for multi-step agents: one zip, one golden prompt, tool families, observation hashes, and Theory receipts — not more prompts.
-
Efficiency normalized to a cohort median can rise while every seat gets slower. Publish absolute wall, tokens, and correctness first.
-
Host-side gates for typed query failures, unchanged observations, and same-tool fan-out — why prompt-only retries fail on high-thinking models.
-
When a hierarchical ReAcTree planner earns its tax on alert+logs triage — and when a single-agent ReAct loop is the better default. Real ratios from a six-combo bench.
-
From live combo benches: the short Grafana/Loki tool set that showed up in successful closes — and the dead ends that burned turns.
-
Same alert+logs prompt, three model classes. How their Theories diverge: leak vs measurement artifact, and what that means for operators.
-
Same dual-part observability prompt across three model classes and single-agent ReAct vs hierarchical ReAcTree. Wall, tokens, cost, correctness, and an efficiency score you can steal.
-
Plain English: ReAct is one reason→act→observe loop. ReAcTree is a tree of agent nodes with sequence, parallel, and fallback control flow — for SRE-style tool-heavy jobs.
-
When observability tool results spill to disk, agents should page with opaque cursors, not grep the preview and declare every pod Unavailable.
August 2026 25 posts
-
Weekly reflection: context engineering over language wars, plus public libraries and servers for write/select/compress/isolate with multi-language or MCP clients.
-
Context engineering A/B: observation masking (tool-result clearing) made an AI SRE agent ~43% cheaper—then failed equal-evidence RCA. Scorecard + knobs.
-
Simple vs plan on AppWorld skills smoke: plan scores higher on some tasks, ties on another, and costs more in this three-task sample. Choose the mode that fits the job — both belong in the toolkit.
-
Single-agent MCP loaded 453 tools (~17k peak). A planner peaked lower, spent ~2× total tokens, and sometimes won quality. Report pass@1 and pass@3 — that is not pass^k.
-
Unattended AI agent evaluation fails when the model asks the user or reports done with zero MCP calls. Put those rules in the harness, not the framework.
-
Run AppWorld locally with the Aiden agent runtime: docker-compose, MCP HTTP wiring, health gates, and lessons from replacing custom shims with the official three-server stack.
-
Stop duplicate agent workers: combined handoff and tool-registration fixes improved fairness to 5/5 and reduced measured planner tokens on one AppWorld slice — partial pass% rose while strict TGC remains a separate harness goal.
-
AI agent eval failure modes on AppWorld: budget stops before evaluate, PII placeholders poison tool calls, 422 wrong methods, and prose PASS vs judge FAIL.
-
Agent orchestration tax after a fair eval: on AppWorld tasks, planner paths cost ~1.6× tokens and ~3× iterations — measure coordination cost separately from benchmark success.
-
Fair agent evals for planner vs single-agent: match domain tool access on workers before you compare tokens, pass rate, or declare a routing winner.
-
From vibes to contracts: how we rebuilt agent evals around eval sets, rubrics vs criteria, a grader stack, and pass^k reliability.
-
Canary-first evals for live SRE investigate: check the canary before burning judge tokens, and never treat a draft RCA as done.
-
AI SRE agent benchmarks: wall time, tool calls, tokens, and ReAcTree tax — fair A/B numbers so you know when orchestration is worth the cost.
-
Reasoning effort is not a free upgrade for tool-heavy AI agents. Live SRE A/Bs: blanket high timed out, adaptive low→high dug deeper.
-
Claim-aware evidence packing: don't accuse AI agents of inventing what you truncated. Why hallucination guards fail when they drop the receipts.
-
Single-agent vs multi-agent for SRE triage: a fair A/B, what each shape wins at, and a decision framework so you stop defaulting to either.
-
Aiden the easy way: one OpenTofu module from a vague GitHub issue to a review PR — Specify, Research, Plan, without wiring every resource by hand.
-
Aiden the hard way: turn a vague GitHub issue into a review PR — GitHub integration, agent, workflow, webhook, and status poll wired by hand.
-
What are SRE AI agents? AI for incident triage, diagnostics, and RCA with bounded autonomy — not a chatbot and not open-ended remediation demos.
-
What is an AI agent runtime? The production loop that plans, calls tools, and manages context — not a GenAI platform, chatbot, or notebook demo.
-
SRE for agentic systems: uptime isn't enough. Judgment SLOs and how to measure agentic drift when the agent can be 'up' and still wrong.
-
PII redaction for AI agents needs two views: a protected model history and authorized operator visibility when debugging tool calls.
-
Can you interrupt or redirect an AI agent mid-response? Yes. How to send steer or cancel signals mid-stream and change the active task without restarting.
-
Prompt caching for AI agents is an architecture problem: stable prefixes, early compaction, and references beat copying a turbulent payload.
-
AI agent loop detection can erase a good answer. Preserve the best evidence-backed result when a stalled run ends — don't throw the work away.
July 2026 17 posts
-
An incident agent can collect a useful metric lead and omit it from the final report; check time windows, unrelated alerts, and transcript-to-summary consistency.
-
When an incident agent reaches an empty result, broaden searches using recorded identifiers instead of guessing rule IDs or measurements.
-
For AI-assisted incident analysis, record required checks, return missing fields together, and limit confidence to what the evidence supports.
-
Check an AI agent’s deliverable and tool outcomes before reporting a goal as complete; return unverified work explicitly when checking stops.
-
An AI agent scorecard can grade the wrong work; check trace identity, coverage, and evaluations before using its reliability or quality scores.
-
A structured way to investigate service incidents: identify the affected system and onset, test competing explanations, and report uncertainty.
-
Questions for evaluating AI incident-response pilots: recorded evidence, staged checks, approval policies, and test results beyond a fluent demo.
-
AI agent learning loop: daily digests become human-approved workflow and policy changes — not a bigger vector store.
-
Beyond Confluence runbooks: why GitOps triage steps matter for AI agents — version-controlled procedures that change with your stack.
-
Your RCA agent doesn't need another runbook — it needs a map. Topology, gates, and verify-first navigation beat a forty-page notebook.
-
How to debug multi-step and multi-stage AI agent workflows and execution logs — green one stage at a time, score tool effects not transcripts.
-
LLM tokenomics for production agents: token budget strategies, context compression, and FinOps loops that keep sessions finishing.
-
Evidence-based verification for AI agents: don't trust self-report — pull proof from ArgoCD, Datadog, and systems of record, then let Go own pass/fail.
-
Evidence-gated RCA for AI SRE agents: prove with receipts, then narrate. Fixed stages, structural evals, and token-aware tool loops.
-
AI incident triage for SREs — what actually helps on-call versus demo theater. Canonical copy now on StackGen.
-
Production agent tool calls need layered JSON repair — why one pass fails on malformed LLM output and what we learned in Go middleware.
-
Contributing back while building a commercial product: how we shipped a proprietary platform and still merged PRs into the agent framework we depend on.
June 2026 12 posts
-
AI agent runtime vs platform: why we split the Go runtime from Aiden, a multi-tenant orchestration layer for enterprise GenAI agents.
-
Terraform for AI agent configuration — why we use infrastructure as code, not YAML dashboards, to govern production agents.
-
Observability for AI agents: session traces, tool attribution, token budgets, and audit trails — the signals traditional APM misses in production.
-
Defense-in-depth for production AI agents: layered policy, HITL, and tool governance when the agent has root — prompts are not security.
-
HITL approvals can make AI agents worse. How to find the human-in-the-loop balance so review gates protect production without stalling the agent.
-
Agent skill distillation without fine-tuning: teach reusable skills from production traces instead of GPU-trained distilled models.
-
AI agent memory that forgets on purpose — Pensieve manages four memory types with decay and self-pruning so RAG stops stuffing stale context.
-
Six runtime bugs we found implementing ReAcTree: delegated approvals, graph wiring, timeouts, session isolation, recursion and memory.
-
Go platform architecture for a production AI agent codebase — patterns that keep rapid development sustainable without drowning in process.
-
Why we chose TOML for flat, typed agent configuration after considering YAML, PKL and CUE, and when those alternatives may fit better.
-
Go vs Python for AI agents: why we chose Go for a production agent runtime — concurrency, single-binary ops, deployment, and when Python still wins.
-
LLM and agent latency: first-token wait, output rate, prompt processing, completion time, and validation failures, using web metrics as a debugging analogy.
No posts match that search. Try a broader term, or .
Posts outside the numbered series (e.g. cloud entitlements, web→LLM metrics) live on the homepage archive.
Stay in the loop — production notes on AI agents, workflows, and SRE.
Low volume — new posts and curated reading lists. Unsubscribe anytime.
FAQ
Why build an enterprise AI agent platform in Go?
We chose Go for its types, deployment model, and concurrency support. The series explains that choice alongside its costs, including access to Python-first AI libraries; it is not a general recommendation to switch languages.
Where should I start reading?
Use a starter pack: Go agent runtime (definition, why Go, platform split) or SRE on-call (triage, RCA, observability). Then open the searchable series archive and follow series_order. Each post is self-contained but builds on prior lessons.
Who is this series for?
Engineers interested in how agents run and fail in practice. The starter packs introduce the terms before the more detailed implementation posts.