Tag: ai-agents
Posts tagged ai-agents:
-
AI Agent Evals in CI/CD: Flakes, Retries, and Missing Results — Keep live agent trials deliberate, report missing tasks and failed attempts, and match trial IDs to traces before trusting a CI result.
-
Multi-Agent Handoff Testing: Catch Context Loss Before Production — Test the brief passed from a worker to its parent: settled facts, open gaps, size limits, and a clear stop condition.
-
How to Test AI Agent Loops Without Overfitting — Detect stalled agent loops from outcomes and evidence, not repeated tool names; test the stop rule with stubbed calls.
-
Deterministic Checks vs LLM-as-a-Judge for Agent Evals — Check verifiable facts with code and use a calibrated model judge for questions that require reading and interpretation.
-
The AI Agent Testing Pyramid: What Belongs in CI? — Use fast tests for agent bounds, contract checks for facts, and a small number of live trials for whole jobs.
-
AI Agent Evaluation: Why the Grader Must Be Separate — Separate the agent from its grader, check facts with code, and run the installed build in the trial environment.
-
Same Incident, Five Models: Rank Investigators Without Mixing the Exam — Random CI reruns mix model skill with incident difficulty. Freeze one checkout outage, swap only the bound model, and rank on Correct first, then wall time and new input tokens.
-
Session Efficiency Should Not Beat Accuracy — We labeled a score 0-100 and watched a perfect, cheap, fast agent run print 125. Cap budget bonuses at 1.0 so efficiency never grades above how often you were Correct.
-
Chat Completions vs Responses API (and Why Agents Feel Slow) — OpenAI Chat Completions and the Responses API are two ways to call a model. Reasoning models are a different choice. How to tell them apart in a debug export, and why Responses is not automatically slower.
-
Hybrid Plan Mode: Smart Planner, Generate Digs — Plan roots already route to planning; digs default to tool calling. Putting a high-thinking reasoning model on every task makes workers slow. Use a thinking model for the planner and a normal chat model for digs.
-
AI Agent Runtime: What to Measure Before You Buy — Buyer checklist for an AI agent runtime: loop, tools, gates, and an eval receipt — with numbers from a real multi-model triage bench.
-
How to Evaluate an AI Agent for Root Cause Analysis — A checklist-first RCA eval: Theory, Unknowns, next action, measured signal, honesty about gaps — plus wall, tokens, and cost. Built from a live alert+logs bench.
-
Reasoning vs Generate Models for Tool-Heavy Agents — Merits and demerits of reasoning vs normal chat models on a live Grafana triage job — elapsed time, tokens, Theory style, and when each earns the bill.
-
Best Way to Debug a Multi-Step AI Agent in Production — A practical debug loop for multi-step agents: one zip, one golden prompt, tool families, observation hashes, and Theory receipts — not more prompts.
-
Relative Efficiency Scores Lie When Absolute Wall Doubles — Efficiency normalized to a cohort median can rise while every seat gets slower. Publish absolute wall, tokens, and correctness first.
-
Stop Retrying the Same Failed Observability Query — Host-side gates for typed query failures, unchanged observations, and same-tool fan-out — why prompt-only retries fail on high-thinking models.
-
Hierarchical vs Single-Agent on a Dual-Part Observability Job — When a hierarchical ReAcTree planner earns its tax on alert+logs triage — and when a single-agent ReAct loop is the better default. Real ratios from a six-combo bench.
-
Observability Tools Agents Actually Call on Triage — From live combo benches: the short Grafana/Loki tool set that showed up in successful closes — and the dead ends that burned turns.
-
What Reasoning Models Write When You Ask Them to Triage — Same alert+logs prompt, three model classes. How their Theories diverge: leak vs measurement artifact, and what that means for operators.
-
Six Model×Orchestration Combos on One Alert+Logs Job — Same dual-part observability prompt across three model classes and single-agent ReAct vs hierarchical ReAcTree. Wall, tokens, cost, correctness, and an efficiency score you can steal.
-
What Is ReAcTree? Hierarchical Agent Trees for Long-Horizon Work — Plain English: ReAct is one reason→act→observe loop. ReAcTree is a tree of agent nodes with sequence, parallel, and fallback control flow — for SRE-style tool-heavy jobs.
-
Cursor paging for spilled agent tool output — When observability tool results spill to disk, agents should page with opaque cursors, not grep the preview and declare every pod Unavailable.
-
Weekly Reflection: Context Engineering Beat the Language Wars — Weekly reflection: context engineering over language wars, plus public libraries and servers for write/select/compress/isolate with multi-language or MCP clients.
-
Observation Masking Cut Token Cost 43%—Then Lost the RCA — Context engineering A/B: observation masking (tool-result clearing) made an AI SRE agent ~43% cheaper—then failed equal-evidence RCA. Scorecard + knobs.
-
Simple vs Plan: When to Use Which (AppWorld smoke cohort) — Simple vs plan on AppWorld skills smoke: plan scores higher on some tasks, ties on another, and costs more in this three-task sample. Choose the mode that fits the job — both belong in the toolkit.
-
Multi-Agent vs Single-Agent: When Planning Beats Reacting (MCP Tool Tax + pass@k) — Single-agent MCP loaded 453 tools (~17k peak). A planner peaked lower, spent ~2× total tokens, and sometimes won quality. Report pass@1 and pass@3 — that is not pass^k.
-
How to Evaluate AI Agents: Ban Clarifying Questions and Zero-Tool “Success” — Unattended AI agent evaluation fails when the model asks the user or reports done with zero MCP calls. Put those rules in the harness, not the framework.
-
Debug AI Agent Failures: One Zip Per Conversation — Support said the AI investigation went wrong. Give one conversation debug zip—not per-execution scavenger hunts—then grade batches into product gates.
-
Running AppWorld Locally for Aiden Agent Evals: Docker Compose, MCP, and Things We Wish We Knew Upfront — Run AppWorld locally with the Aiden agent runtime: docker-compose, MCP HTTP wiring, health gates, and lessons from replacing custom shims with the official three-server stack.
-
Stop Spawning Duplicate Workers: What a Handoff Gate Changed in Agent Evals — Stop duplicate agent workers: combined handoff and tool-registration fixes improved fairness to 5/5 and reduced measured planner tokens on one AppWorld slice — partial pass% rose while strict TGC remains a separate harness goal.
-
AI Agent Eval Failure Modes: Budget, PII Placeholders, and Self-Reported Passes — AI agent eval failure modes on AppWorld: budget stops before evaluate, PII placeholders poison tool calls, 422 wrong methods, and prose PASS vs judge FAIL.
-
Agent Orchestration Tax: Tokens, Iterations, and Tool Calls After a Fair Eval — Agent orchestration tax after a fair eval: on AppWorld tasks, planner paths cost ~1.6× tokens and ~3× iterations — measure coordination cost separately from benchmark success.
-
Fair Agent Evals: Don’t Compare Planner vs Single-Agent Until Tools Match — Fair agent evals for planner vs single-agent: match domain tool access on workers before you compare tokens, pass rate, or declare a routing winner.
-
Why AI SRE Feels Stuck Before the First Tool Call — On-call sees ‘the bot is stuck’ while vault re-checks and MCP catalog re-index burn cold start. Measure time to the first useful tool call.
-
From Vibes to Contracts: How We Rebuilt Agent Evals Around an Industry Standard — From vibes to contracts: how we rebuilt agent evals around eval sets, rubrics vs criteria, a grader stack, and pass^k reliability.
-
Canary First: Lessons from Black-Box Consistency Evals for Live SRE Investigate — Canary-first evals for live SRE investigate: check the canary before burning judge tokens, and never treat a draft RCA as done.
-
AI SRE Agent Benchmarks: Wall Time, Tool Calls, Tokens, and ReAcTree Tax — AI SRE agent benchmarks: wall time, tool calls, tokens, and ReAcTree tax — fair A/B numbers so you know when orchestration is worth the cost.
-
Reasoning Effort Is Not a Free Upgrade for Tool-Heavy Agents — Reasoning effort is not a free upgrade for tool-heavy AI agents. Live SRE A/Bs: blanket high timed out, adaptive low→high dug deeper.
-
Claim-Aware Evidence Packing — Don’t Accuse Agents of Inventing What You Truncated — Claim-aware evidence packing: don’t accuse AI agents of inventing what you truncated. Why hallucination guards fail when they drop the receipts.
-
Single-Agent vs Multi-Agent Orchestration: How to Choose — Single-agent vs multi-agent for SRE triage: a fair A/B, what each shape wins at, and a decision framework so you stop defaulting to either.
-
Aiden the Easy Way: One Module from Vague Issue to Review PR — Aiden the easy way: one OpenTofu module from a vague GitHub issue to a review PR — Specify, Research, Plan, without wiring every resource by hand.
-
Aiden the Hard Way: Vague GitHub Issues to Review PRs — Aiden the hard way: turn a vague GitHub issue into a review PR — GitHub integration, agent, workflow, webhook, and status poll wired by hand.
-
What Are SRE AI Agents? — What are SRE AI agents? AI for incident triage, diagnostics, and RCA with bounded autonomy — not a chatbot and not open-ended remediation demos.
-
Don’t Invent PromQL: Measure the Alert Rule Query First — Alert title said latency; the rule was ClickHouse. AI SRE agents that invent PromQL from titles ship wrong RCA. Measure the stored query first.
-
What Is an AI Agent Runtime? — What is an AI agent runtime? The production loop that plans, calls tools, and manages context — not a GenAI platform, chatbot, or notebook demo.
-
SRE for Agentic Systems: Why Uptime Isn’t Enough Anymore — SRE for agentic systems: uptime isn’t enough. Judgment SLOs and how to measure agentic drift when the agent can be ‘up’ and still wrong.
-
PII Redaction for AI Agents — Two Views, One Trace — PII redaction for AI agents needs two views: a protected model history and authorized operator visibility when debugging tool calls.
-
How to Steer an AI Agent Mid-Run Without Starting Over — Can you interrupt or redirect an AI agent mid-response? Yes. How to send steer or cancel signals mid-stream and change the active task without restarting.
-
Prompt Caching for AI Agents Is an Architecture Problem — Prompt caching for AI agents is an architecture problem: stable prefixes, early compaction, and references beat copying a turbulent payload.
-
AI Agent Loop Detection — Don’t Throw Away the Answer — AI agent loop detection can erase a good answer. Preserve the best evidence-backed result when a stalled run ends — don’t throw the work away.
-
Empty PromQL ≠ Missing Data: Fix AI SRE Scope Blindness — A Grafana no_data result applies to one query and time range; check labels, time, data source, and tool failures before reporting missing evidence.
-
AI Agent Root Cause Analysis — Evidence Discarded After the Lead — An incident agent can collect a useful metric lead and omit it from the final report; check time windows, unrelated alerts, and transcript-to-summary consistency.
-
Be Creative. Don’t Invent. — When an incident agent reaches an empty result, broaden searches using recorded identifiers instead of guessing rule IDs or measurements.
-
Ungrounded Synthesis Must Read as Hypothesis — If a check finds unsupported details in an AI-written incident analysis, correct the primary message rather than hiding the warning in a note.
-
AI Agent Root Cause Analysis — Curiosity Before Confidence — For AI-assisted incident analysis, record required checks, return missing fields together, and limit confidence to what the evidence supports.
-
Is the Task Actually Done? — Completion Loops for Production Agents — Check an AI agent’s deliverable and tool outcomes before reporting a goal as complete; return unverified work explicitly when checking stops.
-
When Your AI Agent Scorecard Lies — An AI agent scorecard can grade the wrong work; check trace identity, coverage, and evaluations before using its reliability or quality scores.
-
AI Agent Hit Max Turns? Deliver Partial RCA, Not Apology — When an incident agent reaches a model-call limit, preserve observed findings and label a partial analysis incomplete.
-
The Hypothesis Ladder — Ruling Things Out Before You Narrate — A structured way to investigate service incidents: identify the affected system and onset, test competing explanations, and report uncertainty.
-
From Demo to Deploy — Failure Modes with Receipts — Questions for evaluating AI incident-response pilots: recorded evidence, staged checks, approval policies, and test results beyond a fluent demo.
-
The Diary Learning Loop — From Daily Agent Digests to Human-Approved Policy — AI agent learning loop: daily digests become human-approved workflow and policy changes — not a bigger vector store.
-
Beyond Confluence Runbooks: Why GitOps Triage Steps Matter in the AI Era — Beyond Confluence runbooks: why GitOps triage steps matter for AI agents — version-controlled procedures that change with your stack.
-
Your RCA Agent Doesn’t Need Another Runbook — It Needs a Map — Your RCA agent doesn’t need another runbook — it needs a map. Topology, gates, and verify-first navigation beat a forty-page notebook.
-
AI Agents Call Truncated Grafana ‘No Data’—It’s Spill — Large tool results get preview-truncated; AI SRE agents invent ‘Unavailable.’ Spill recovery and COMPLETE/PARTIAL/FAILED fix dishonest RCA.
-
How We Debug Multi-Stage AI Agent Workflows — How to debug multi-step and multi-stage AI agent workflows and execution logs — green one stage at a time, score tool effects not transcripts.
-
LLM Tokenomics for Production Agents — Context Budgets as an Operating Model — LLM tokenomics for production agents: token budget strategies, context compression, and FinOps loops that keep sessions finishing.
-
Evidence-Based Verification — Don’t Trust Self-Report, Check the System — Evidence-based verification for AI agents: don’t trust self-report — pull proof from ArgoCD, Datadog, and systems of record, then let Go own pass/fail.
-
Evidence-Gated RCA — Prove, Then Narrate — Evidence-gated RCA for AI SRE agents: prove with receipts, then narrate. Fixed stages, structural evals, and token-aware tool loops.
-
AI Incident Triage for SREs — What Actually Helps On-Call — AI incident triage for SREs — what actually helps on-call versus demo theater. Canonical copy now on StackGen.
-
Why One JSON Repair Pass Isn’t Enough for Production Agent Tool Calls — Production agent tool calls need layered JSON repair — why one pass fails on malformed LLM output and what we learned in Go middleware.
-
When the Operator Asks to Correlate, Make It a Gate — Natural-language correlation goals need server-side gates — not hope the LLM remembers to search prior incidents.
-
Contributing Back While Building a Commercial Product — Contributing back while building a commercial product: how we shipped a proprietary platform and still merged PRs into the agent framework we depend on.
-
AI Agent Runtime vs Platform — Why We Split Them — AI agent runtime vs platform: why we split the Go runtime from Aiden, a multi-tenant orchestration layer for enterprise GenAI agents.
-
Terraform for Agent Configuration — Infrastructure as Code Meets AI Governance — Terraform for AI agent configuration — why we use infrastructure as code, not YAML dashboards, to govern production agents.
-
You Can’t Debug What You Can’t See — Observability for AI Agents — Observability for AI agents: session traces, tool attribution, token budgets, and audit trails — the signals traditional APM misses in production.
-
Your Agent Has Root — Defense-in-Depth for AI Agents That Wield Real Tools — Defense-in-depth for production AI agents: layered policy, HITL, and tool governance when the agent has root — prompts are not security.
-
Same Alert, Different Verdict: Entry Path Is Context — Don’t paste the alert in the UI and wonder why Slack gave a different impact score — entry path carries investigation context.
-
The HITL Paradox — When Human Approval Makes Agents Worse — HITL approvals can make AI agents worse. How to find the human-in-the-loop balance so review gates protect production without stalling the agent.
-
Agent Skill Distillation Without Fine-Tuning — Agent skill distillation without fine-tuning: teach reusable skills from production traces instead of GPU-trained distilled models.
-
Pensieve — Memory Management for AI Agents That Actually Forget — AI agent memory that forgets on purpose — Pensieve manages four memory types with decay and self-pruning so RAG stops stuffing stale context.
-
Implementing ReAcTree — 6 Production Bugs the Paper Didn’t Warn You About — Six runtime bugs we found implementing ReAcTree: delegated approvals, graph wiring, timeouts, session isolation, recursion and memory.
-
Go Platform Architecture at Speed — Without Drowning — Go platform architecture for a production AI agent codebase — patterns that keep rapid development sustainable without drowning in process.
-
TOML Over YAML and PKL — How We Stopped Fighting Config and Started Shipping — Why we chose TOML for flat, typed agent configuration after considering YAML, PKL and CUE, and when those alternatives may fit better.
-
Python vs Go for AI Agents — Why We Chose Go — Go vs Python for AI agents: why we chose Go for a production agent runtime — concurrency, single-binary ops, deployment, and when Python still wins.
-
Slack Is a Triage Board, Not a Log Dump — Make incident replies in Slack scannable: show counts, findings and uncertainty, with links to evidence and searchable Activity.
-
Stop Re-Investigating the Same Alert — Reuse recent incident findings for follow-ups on the same alert, with a clear way to investigate again when evidence changes.
-
Service Rendered Efficiently: SRE AI Is Not an Engineering Credibility Project — Service Rendered Efficiently: judge AI investigation by whether it helps on-call reuse results, understand uncertainty, and hand off work.
-
LLM Performance Metrics — From Lighthouse to the Token Era — LLM and agent latency: first-token wait, output rate, prompt processing, completion time, and validation failures, using web metrics as a debugging analogy.