Python vs Go for AI Agents — Why We Chose Go
We chose Go for the part of our agent system that runs model calls, tools and policy checks. An agent runtime is the software that coordinates those calls and tool executions. We needed to embed it in an existing Go platform and distribute it in environments where managing a Python installation would be costly. That is a local architectural choice, not a general rule against Python.
Many agent frameworks and examples use Python, including LangChain, LangGraph, AutoGen, and CrewAI. When we started building an enterprise AI agent platform in Go at StackGen, the obvious choice was Python.
We chose Go instead. Several months of production development later, here’s why — and the trade-offs we had to accept. Related: Go agent runtime starter pack, What Is an AI Agent Runtime?.
Update (July 2026): The trade-off in §5 — Dynamic LLM JSON parsing is softer than when this post first published. We now route shared LLM JSON decode paths through
github.com/kaptinlin/jsonrepair— a Go port of Jos de Jong’sjsonrepair, tuned for malformed model output. See the revised section below.
The Problem We Were Solving
We weren’t building a chatbot. We were building an agent runtime — software that lets AI models invoke shell commands, call APIs, manage infrastructure, and delegate work to sub-agents. In production. On your servers. With your credentials.
The requirements:
- Run many agents concurrently per deployment, each with its own tools, memory, and policies
- Embed as a library inside a larger platform (Aiden), not run as a separate service
- Deploy as a single binary to air-gapped environments, edge nodes, and developer laptops
- Low overhead per tool call — agents call many tools per task; middleware cost compounds
- Handle untrusted input — agents process user prompts that might contain injection attacks
Let’s walk through how Go handles each of these.
1. Concurrency: goroutines fit our runtime
An AI agent in production does many things at once:
- Calls a large language model (LLM) API and waits for streaming tokens
- Executes a shell command on a remote server
- Queries a vector database (a similarity-search index) for relevant context
- Listens for human approval on a pending tool call
- Runs scheduled work that triggers another agent
Python’s asyncio runs cooperative tasks on an event loop. Blocking calls in that loop need separate handling; threads are another option, and the global interpreter lock (GIL) has historically limited CPU-parallel Python bytecode in standard builds. Free-threaded Python 3.13+ changes that tradeoff when supported by the deployment and its libraries. Go’s goroutines are lightweight concurrent functions; a context can carry cancellation through model and tool calls. Neither language makes cancellation or shared state correct automatically.
Real example: When an agent delegates to several sub-agents in parallel, each runs with its own model session, tool set, and memory context. The parent waits for all of them; if one fails, the context cancels and the others wind down cleanly.
What about Python’s asyncio?
It is a good fit when dependencies support it. Blocking database or HTTP calls must not run on the event loop without accommodation. Go avoids propagating async function signatures, but we still need to bound concurrent work, pass cancellation, and test races.
2. Single Binary: Ship One File, Not a Dependency Tree
Our agent runs on developer laptops, Kubernetes clusters, air-gapped enterprise servers, and CI/CD pipelines as a CLI tool.
For our targets, Go lets us cross-compile and package a single binary, avoiding a separate Python interpreter and package installation. Static linking and portability depend on build options and native dependencies; a binary still needs compatible target architecture and operating-system support.
Containers, PyInstaller and uv can make Python deployments manageable, particularly when the host already has a container platform. For our regulated and air-gapped targets, reducing the shipped dependency set simplified installation and review. Image size alone would not justify a language switch.
3. Embed as a Library, Not a Service
This boundary comparison explains why embedding the runtime in the existing Go process avoided a separate deployment and cross-process API.
This is the decision that sealed it.
Our agent runtime needs to run inside our orchestration platform (Aiden), which uses Temporal for durable workflow orchestration. Temporal is itself built in Go, with a first-class Go SDK. The agent runtime runs as part of the platform process — not a sidecar, not a subprocess, but a library import in the same binary.
Same process, same memory space, shared types. Embedding avoids a network boundary between runtime and platform. A Python runtime alongside this Go process would need an interprocess boundary, such as gRPC (remote procedure calls), or a separate service; that boundary would require deployment, serialization, and failure handling.
Python packages embed naturally in Python applications. Our host was Go, so the cross-language boundary was the issue. Go modules also require ordinary dependency pinning and review; they do not remove supply-chain risk.
4. Compile-Time Safety: Catch Breaking Changes Before Production
We have a large codebase organized around small interfaces — every method takes a context and a request struct, returns typed results. When we change an interface, the compiler immediately tells us every place that needs updating.
“But we use Pydantic and mypy!” Yes, Python has type-checking now, and Pydantic is excellent. But Pydantic validates data at runtime; mypy can check annotated Python code before deployment when enforced in CI. Go checks declared types at compilation. This catches interface mismatches early, but does not prove incoming model output is valid or that a tool invocation is safe.
We generate test doubles from interfaces automatically. These fakes are type-safe — if the interface changes, the fake won’t compile, and every test using it fails with a clear compiler error, not a runtime surprise.
5. Performance: When Every Tool Call Counts
Agents are chatty. A single task might involve many tool calls. Each call passes through a middleware stack — logging, audit, loop detection, approval gates, timeouts, rate limits, and more.
In Go, middleware is function closures — lightweight per call. For a single agent, that overhead is noise compared to LLM latency. The real performance win is memory footprint at scale.
Goroutines start with a tiny stack that grows as needed. Goroutines can be inexpensive compared with operating-system threads, but memory use depends on the libraries, session state, and workload. We valued the option to run many sessions on constrained hosts; without matched benchmarks, this is not a general Go-versus-Python memory claim.
The Trade-offs We Accepted
It’s not all sunshine. Here’s what we gave up:
1. Smaller AI/ML ecosystem
Python has HuggingFace, PyTorch, scikit-learn, and thousands of AI libraries. Go doesn’t.
How we handled it: We don’t run ML models. We call LLM APIs via HTTP. The heavy ML work happens on the provider’s infrastructure. Our Go code handles orchestration, governance, and tool execution.
It’s worth noting that while Go lacks ML modeling libraries, Some AI infrastructure is Go-based: Ollama, Kubernetes, Docker, and Temporal are all written in Go. Models run in Python/C++/Rust; Go was familiar in our infrastructure layer.
2. Fewer agent framework examples
Every “build an agent” tutorial is in Python. Our team had to translate concepts, not copy code.
How we handled it: This was actually a feature. Translating forced us to understand the algorithms rather than copy an example. When we implemented ReAcTree, integration tests exposed production bugs in graph wiring and governance; those were not language-specific bugs (more in a dedicated post).
3. Prototyping speed
Python is faster for throwaway experiments. Go requires more upfront structure.
How we handled it: We accepted slower early iterations in exchange for faster late-stage development. By month three, typed interfaces and generated test doubles helped us change the runtime with confidence. That experience does not establish a general difference in defect rates between the languages.
4. Hiring
More ML engineers know Python than Go.
How we handled it: We’re building infrastructure, not ML models. Systems engineers who know Go are exactly the profile we need. A team without Go experience should include training and maintenance costs in this decision.
5. Dynamic LLM JSON parsing
This used to be Go’s genuine pain point for AI work. LLMs return subtly malformed JSON — fences, single quotes, trailing commas, truncated objects. Python’s json.loads() and Go’s encoding/json both reject malformed JSON. Go also supports dynamic values; we chose strict structs at tool boundaries because they make expected fields explicit.
How we handled it (2026 update): We still use strict structs for typed boundaries — that’s a feature, not a bug. But for syntax problems models create, we centralized repair in github.com/kaptinlin/jsonrepair. Shared decode helpers try strict parsing first; when that fails, they repair and retry.
We still layer other tactics where semantics matter: deferring argument parsing until the tool handler knows its schema, path-based extraction when we need specific values without full structs, and re-prompting the model when JSON is syntactically valid but semantically wrong (repair can’t fix wrong shapes).
It’s more ceremony than json.loads() in a REPL. The upside is unchanged: when JSON parses into a struct, its basic types match the declared shape; additional semantic and policy validation is still needed. See the JSON repair post for why one repair pass isn’t enough in production.
When You Should NOT Use Go for Agents
Be honest about when Python is the right choice:
- You’re prototyping and need to test ideas in hours, not days
- You’re running local models with HuggingFace/PyTorch and need direct GPU access
- Your team is ML-first and everyone thinks in NumPy
- You’re building on LangChain/LangGraph and the ecosystem matters more than performance
- You’re a solo developer and can’t invest in the upfront structure Go requires
Go suited our long-running runtime and its Go host. A Python team can also build governed, multi-tenant infrastructure. If deployment integration and team expertise point to Python, the existence of goroutines is not a reason to switch.
The Scorecard
| Criterion | Go | Python |
|---|---|---|
| Concurrency model | Goroutines; explicit cancellation and synchronization | asyncio (opt-in, colored functions) |
| Deployment | Often one binary; target-specific builds | Interpreter/packages or packaged container |
| Library embedding | Module import, same process | Natural in Python host; cross-language boundary in Go host |
| Type safety | Compile-time checks for declared types | Runtime validation plus optional static checking |
| Memory footprint at scale | Workload-dependent | Workload-dependent |
| Middleware overhead | Typically small versus model latency | Typically small versus model latency |
| Dynamic JSON parsing | Strict structs; repair libraries help | json.loads() rejects malformed syntax; dynamic structures available |
| AI/ML ecosystem | Minimal (strong AI infra) | Dominant |
| Prototyping speed | Slower start | Rapid iteration |
| Hiring pool (ML) | Smaller | Larger |
What’s Next
In the next post, I’ll cover why we chose TOML over YAML and PKL for agent configuration — and why the config format you pick matters more than you think.
Related reading
- What Is an AI Agent Runtime? — the production loop this language bet is for
- Go Platform Architecture at Speed — Without Drowning — growing the Go codebase after the language bet
- AI Agent Runtime vs Platform — Why We Split Them — what we built on top of the runtime
- More on Go AI agents · full series
If you’re building agents in Go, I’d love to hear about your experience. Find me on GitHub or LinkedIn.
I work on AI-assisted incident investigation at StackGen; project information is at ai.stackgen.com.
FAQ
Why choose Go over Python for AI agents?
Go fit our existing Go host, concurrent sessions, typed tool interfaces, and deployment targets. Python remains a strong choice for research, ML libraries, and Python-based services.
When should you still use Python for GenAI agents?
When your team lacks Go depth, when you need tight notebook or ML-library integration, or when prompt iteration speed matters more than runtime discipline and deployment shape.
What problem was the Go agent runtime solving?
A production agent runtime that runs many concurrent agents with tools, memory, and policies — embeddable as a library, deployable as a single binary, with tool-call policy checks and untrusted-input handling.
Stay in the loop — production notes on AI agents, workflows, and SRE.
Low volume — new posts and curated reading lists. Unsubscribe anytime.