A Kubernetes-native framework for building production AI agents. Tools, Skills, and Agents are declared as custom resources and launched as one-shot Jobs by a dedicated controller — so operators manage the catalog declaratively and the orchestrator stays focused on reasoning, not infrastructure.
Hard-coding tool definitions and Job-launching logic inside an orchestrator
couples infrastructure concerns to application code: rotating a secret, tuning
resource limits, or adding a new tool all require an image rebuild and
redeployment. It also means the orchestrator's ServiceAccount needs broad
batch/jobs create permissions, every configuration change bypasses version
control, and there is no Kubernetes-native way to inspect what the agent has
been doing.
Modelling tools and agents as custom resources flips this:
| Concern | Baked into orchestrator | With core-controller |
|---|---|---|
| Tool definition | Config files / env vars | Tool CR — live-editable, kubectl-discoverable |
| Secret injection | Orchestrator builds the full Job spec | Declared once on the CR; controller injects at run time |
| Job RBAC | Orchestrator SA needs batch/jobs create |
Controller's SA creates Jobs; orchestrator only needs toolruns create |
| Lifecycle visibility | Orchestrator polls Job status | ToolRun.status.phase + conditions — any cluster tenant can read |
| Skill / Agent catalog | Static code or restart-requiring config | Skill and Agent CRs — kubectl apply, no rebuild |
| Audit trail | Application logs only | Kubernetes events + per-invocation CR status |
Operational benefits:
- Change a tool's description, roles, or limits with
kubectl apply— no image rebuild. - Operators control the catalog; developers control the orchestrator. Neither needs the other's access.
- Tools and full sub-agent loops share one execution architecture. Adding a new workload kind is a new CRD, not new Job-launching code.
kubectl get toolruns,kubectl describe agentrun, and standard controller metrics work out of the box.
graph TD
User([User / Chat client])
Orchestrator[agent-orchestrator\nDeployment]
Controller[core-controller\nDeployment]
subgraph CRDs["Custom Resources (core.controller-agent.dev/v1alpha1)"]
ToolCR[Tool CR\nimage · SA · secretEnv · allowedRoles]
SkillCR[Skill CR\nmarkdown · toolRefs]
AgentCR[Agent CR\nimage · SA · skillRefs · toolRefs]
ToolRunCR[ToolRun CR\ntoolRef · args · callback]
AgentRunCR[AgentRun CR\nagentRef · goal · callback]
end
subgraph Jobs["Kubernetes Jobs (one-shot)"]
ToolJob[Tool Job pod]
AgentJob[Agent Job pod]
end
CallbackSrv[CallbackReceiver\nport 8080]
Qdrant[(Qdrant\nvector store)]
User -->|POST /v1/chat/completions| Orchestrator
Orchestrator -->|RAG skill + tool selection| Qdrant
Orchestrator -->|reads CRs at startup| ToolCR & SkillCR
Orchestrator -->|kubectl create| ToolRunCR & AgentRunCR
Controller -->|watches + reconciles| ToolRunCR & AgentRunCR & ToolCR & SkillCR & AgentCR
Controller -->|creates Job with secret injection| ToolJob & AgentJob
ToolJob & AgentJob -->|HMAC POST /callback/:id| CallbackSrv
CallbackSrv -->|resolves pending promise| Orchestrator
Orchestrator -->|final SSE chunk| User
This is an npm workspace monorepo: packages/ holds shared libraries,
tools/ holds on-demand tool containers, apps/ holds the long-lived
orchestrator service, and controllers/ holds the Go controller.
.
├── README.md # this file — general overview & conventions
├── package.json # npm workspaces root (packages/* + tools/* + apps/*)
├── docs/ # shared standards every tool follows
│ ├── messaging.md # event protocol & transports
│ ├── security.md # threat model & mitigations
│ ├── orchestrator.md # orchestrator architecture
│ ├── integrations-gateway.md # event integrations proposal (GitHub Issues implemented)
│ └── adr/ # Architecture Decision Records
├── packages/
│ ├── messaging/ # @controller-agent/messaging — shared event protocol
│ └── github-app-auth/ # @controller-agent/github-app-auth — GitHub App JWT/token auth
├── tools/ # on-demand tool containers (example implementations)
│ ├── recipe-scraper/ # URL → recipe Markdown
│ ├── recipe-publisher/ # recipe Markdown → Mealie instance
│ └── github/ # gh CLI command → GitHub, as the calling user's own identity
├── apps/
│ ├── agent-orchestrator/ # RAG skill selection + ToolRun/AgentRun creator
│ └── integration-gateway/ # GitHub Issues → agent-orchestrator webhook adapter
├── controllers/
│ └── core-controller/ # Go controller — watches CRDs, launches Jobs
│ ├── api/v1alpha1/ # Tool, Skill, Agent, ToolRun, AgentRun types
│ └── internal/controller/ # reconciliation logic
└── charts/ # Helm charts
├── agent-controller/ # system chart: orchestrator + core-controller (+ CRDs) + optional Redis/Qdrant/NATS/Open WebUI/integration-gateway
└── community-components/ # catalog chart: Tool/Skill/Agent custom resources
General, cross-cutting documentation lives at the repo root (README.md and
docs/). Anything specific to a single tool — its inputs, configuration,
build/run steps, and troubleshooting — lives in that tool's own README.md.
Code shared by more than one tool belongs in packages/, not copied between
tools.
| Component | Language | Docs |
|---|---|---|
| core-controller | Go (kubebuilder) | controllers/core-controller/README.md |
The controller watches Tool, Skill, Agent, ToolRun, and AgentRun CRs
(API group core.controller-agent.dev/v1alpha1) and manages all Job creation,
secret injection, and lifecycle tracking.
| App | Docs |
|---|---|
| agent-orchestrator | apps/agent-orchestrator/README.md |
A long-lived LangGraph.js service that handles the agent loop: resolves caller
identity, selects a Skill via RAG, plans an action, and creates a ToolRun or
AgentRun CR for the controller to execute.
| Tool | Input | Output | Docs |
|---|---|---|---|
| recipe-scraper | any recipe URL (web page, video, or image) | recipe Markdown | tools/recipe-scraper/README.md |
| recipe-publisher | recipe Markdown | published/updated recipe in a Mealie instance | tools/recipe-publisher/README.md |
| github | a single gh CLI command line |
gh's own output, authenticated as the calling user's own linked GitHub identity |
tools/github/README.md |
| glyph | a JSON note/task command (create/read/update/search) | a Markdown summary of the result, authenticated as the calling user's own linked Glyph identity | tools/glyph/README.md |
Every tool container is expected to conform to these repo-wide standards:
- Message passing — the event protocol
(
accepted → progress* / warning* → succeeded | failed), the transports (stdout / file / HTTP callback), and the correlation/idempotency rules. Implemented once as the @controller-agent/messaging package. - Security model — SSRF defense, prompt-injection containment, the hardened container run contract, and secret handling.
recipe-scraper is the reference implementation of both standards.
Two independent Helm charts cover the full system:
| Chart | What it installs |
|---|---|
| charts/agent-controller | The system: CRDs + core-controller operator Deployment/RBAC, the agent-orchestrator Deployment/Services, and optional Redis/Qdrant/NATS/Open WebUI |
| charts/community-components | The catalog: Tool/Skill/Agent custom resources (recipe-scraper, recipe-publisher, recipe-refining skill, opencode-swe-agent) |
Install agent-controller first (it owns the CRDs), then
community-components on top of it. See each chart's README for
prerequisites and values. Tools are never deployed as long-running pods — the
controller launches them as one-shot Jobs via ToolRun/AgentRun CRs.
From a local checkout:
helm install agent-controller charts/agent-controller -n controller-agent --create-namespace
helm install community-components charts/community-components -n controller-agentBoth charts are also published as OCI artifacts to GitHub Container Registry
on every merge to main that touches charts/** (see
.github/workflows/release.yml),
so you can install without cloning the repo:
helm install agent-controller oci://ghcr.io/imaustink/charts/agent-controller --version 0.1.0 \
-n controller-agent --create-namespace
helm install community-components oci://ghcr.io/imaustink/charts/community-components --version 0.1.0 \
-n controller-agentEvery image is published publicly to GitHub Container Registry by
.github/workflows/release.yml on each merge to
main that touches its sources. No login is needed to pull. Each is tagged
latest (tracks main) and with the full commit SHA (immutable, for pinning),
e.g. docker pull ghcr.io/imaustink/agent-controller/agent-orchestrator:latest.
The chart defaults are bare <name>:latest, so point image values at these
paths (as values-production.yaml does).
| Image | Built from |
|---|---|
ghcr.io/imaustink/agent-controller/agent-orchestrator |
apps/agent-orchestrator/Dockerfile |
ghcr.io/imaustink/agent-controller/opencode-swe-agent |
apps/opencode-swe-agent/Dockerfile |
ghcr.io/imaustink/agent-controller/claude-code-swe-agent |
apps/claude-code-swe-agent/Dockerfile |
ghcr.io/imaustink/agent-controller/integration-gateway |
apps/integration-gateway/Dockerfile |
ghcr.io/imaustink/agent-controller/recipe-scraper |
tools/recipe-scraper/Dockerfile |
ghcr.io/imaustink/agent-controller/recipe-publisher |
tools/recipe-publisher/Dockerfile |
ghcr.io/imaustink/agent-controller/web-search |
tools/web-search/Dockerfile |
ghcr.io/imaustink/agent-controller/web-fetch |
tools/web-fetch/Dockerfile |
ghcr.io/imaustink/agent-controller/image-gen |
tools/image-gen/Dockerfile |
ghcr.io/imaustink/agent-controller/kubectl-readonly |
tools/kubectl-readonly/Dockerfile |
ghcr.io/imaustink/agent-controller/signoz-query |
tools/signoz-query/Dockerfile |
ghcr.io/imaustink/agent-controller/github |
tools/github/Dockerfile |
ghcr.io/imaustink/agent-controller/glyph |
tools/glyph/Dockerfile |
ghcr.io/imaustink/agent-controller/ssh |
tools/ssh/Dockerfile |
ghcr.io/imaustink/agent-controller/core-controller |
controllers/core-controller/Dockerfile |
ghcr.io/imaustink/agent-controller/temporal-engine-worker |
engines/temporal/Dockerfile.worker |
ghcr.io/imaustink/agent-controller/temporal-engine-gateway |
engines/temporal/Dockerfile.gateway |
ghcr.io/imaustink/agent-controller/temporal-engine-catalog-sync |
engines/temporal/Dockerfile.catalog-sync |
ghcr.io/imaustink/agent-controller/localtool-executor-node |
sidecars/localtool-executor/Dockerfile |
ghcr.io/imaustink/agent-controller/localtool-executor-python |
sidecars/localtool-executor/Dockerfile |
ghcr.io/imaustink/agent-controller/localtool-executor-go |
sidecars/localtool-executor/Dockerfile |
ghcr.io/imaustink/agent-controller/localtool-executor-shell |
sidecars/localtool-executor/Dockerfile |
The list comes from .github/release-images.json; adding an entry there publishes a new image.
scripts/dev-up.sh automates the full local deploy (start minikube, build images, install/upgrade both Helm releases, apply CRs). The only step it can't do for you is creating secrets — do this once per fresh cluster, in your own terminal (never paste real secrets into chat or files):
# Generate a random callback HMAC secret
CALLBACK_SECRET=$(openssl rand -hex 32)
# ...and the secret the gateway signs a webhook's sender login with, which the
# orchestrator verifies (docs/adr/0030 §6). ONE value, set on both sides: it
# decides which human a webhook turn acts as, and hence whose stored Claude
# credentials the run gets. Left unset, that login is trusted unsigned.
SENDER_ASSERTION_SECRET=$(openssl rand -hex 32)
kubectl create namespace controller-agent
# OpenAI key + callback HMAC secret (used by the orchestrator)
kubectl -n controller-agent create secret generic agent-orchestrator-secrets \
--from-literal=OPENAI_API_KEY=<your-openai-api-key> \
--from-literal=AGENT_CALLBACK_SECRET="$CALLBACK_SECRET" \
--from-literal=AGENT_SENDER_ASSERTION_SECRET="$SENDER_ASSERTION_SECRET"
# Mealie long-lived API token (create at /user/profile/api-tokens in your Mealie instance)
kubectl -n controller-agent create secret generic recipe-publisher-secrets \
--from-literal=MEALIE_API_TOKEN=<your-mealie-api-token>
# Google OAuth client secret (from Google Cloud Console)
kubectl -n controller-agent create secret generic agent-orchestrator-openwebui-google-oauth \
--from-literal=client-secret=<your-google-oauth-client-secret>Then run:
./scripts/dev-up.shThese secrets survive normal minikube stop/start cycles. If the minikube
container is ever deleted (e.g. after minikube delete), all secrets are lost
and must be re-created before running the script again.
- Create
tools/<tool-name>/with its ownDockerfile, source, andREADME.md. - Add it to the root
package.jsonworkspaces (covered by thetools/*glob) and depend on @controller-agent/messaging for the event protocol — seetools/recipe-scraper/src/messaging/index.tsfor the wiring pattern. - Follow the security model: treat all input as untrusted,
guard outbound requests (SSRF), constrain any LLM output, ship a hardened
run.sh. - If the Dockerfile depends on a
packages/*library, build from the repo root:docker build -f tools/<tool-name>/Dockerfile -t <name>:latest . - Create a
ToolCR intools/<tool-name>/tool.yamlreferencing the image and any requiredsecretEnventries, thenkubectl apply -fit — the controller picks it up immediately, and the orchestrator indexes it on next restart. No manifest file, no orchestrator rebuild required (ADR 0010). - Document tool-specific inputs, config, and build/run steps in the tool's
README.md.
- Language: TypeScript (Node, ESM). Prefer it unless a tool's core dependencies dictate otherwise.
- Isolation: each tool is self-contained (its own dependencies, image, and run contract); tools do not import from one another.
- Untrusted by default: a tool never trusts its input or the content it fetches, and never needs more secrets than the single credential its job requires.