Action Agent
Service: action-agent-srv · Port: 8094
The action agent is the part of the reactive pipeline that changes your cluster. It carries out the analysis agent's decisions — immediately for decisions approved to run automatically, and after approval for the rest — and it exposes the same actions to people and AI agents through its API.
Every action is safe by design: it respects PodDisruptionBudgets, verifies its own effect, and is recorded with its reason and outcome.
Actions
| Action | What it does | Triggered by |
|---|---|---|
| Scale deployment | Changes a Deployment's replica count. | Decisions, API, agents |
| Restart pod | Deletes a pod so its controller recreates it. | Decisions, API, agents |
| Cordon node | Stops new pods being scheduled on a node. | Decisions, API, agents |
| Uncordon node | Allows scheduling on a node again. | API, agents |
| Drain node | Cordons a node and safely evicts its pods. | API, agents |
| Roll back deployment | Returns a Deployment to its previous revision. | API, agents |
Scale deployment
Sets the Deployment's replicas to a target — or, without a target, to the current count plus one.
{ "spec": { "replicas": 5 } }
Scale-downs check every PodDisruptionBudget in the namespace first and are refused if they would break one.
Scale-ups are verified. After scaling up, the agent waits (60 seconds by default), then compares the cluster's health score with its value before the action. If health dropped by more than the allowed amount (10 points by default), the scale-up is reverted automatically and the outcome is recorded as failed.
Restart pod
Deletes the pod; its Deployment, StatefulSet or DaemonSet immediately creates a fresh one. This clears stuck state without touching the rest of the workload.
Cordon and uncordon node
Cordon marks a node unschedulable (spec.unschedulable: true): no new pods are placed on it, and existing pods keep running. Uncordon reverses this.
Cordoning does not move existing pods. Use drain when you need the node empty.
Drain node
Prepares a node for maintenance or removal:
- Cordons the node.
- Lists every pod on it except DaemonSet pods (which belong on every node).
- Evicts them one at a time through the Kubernetes Eviction API, with a 30-second grace period.
Eviction respects PodDisruptionBudgets. When an eviction is blocked, the agent retries with exponential backoff (5 seconds, doubling to a 60-second cap, up to 5 attempts), so a node with tightly budgeted workloads may take a few minutes to drain.
POST /api/v1/actions/drain { "node_name": "ip-10-0-1-45.ec2.internal" }
Roll back deployment
Rolls a Deployment back to its previous revision — the ReplicaSet it ran before the latest rollout — as a zero-downtime rolling update. Pass a revision to return to a specific earlier revision.
POST /api/v1/actions/rollback { "namespace": "production", "deployment_name": "payments-api" }
How decisions become actions
For each decision consumed from analysis.decisions:
- If
AutoApproveis true, the agent executes the action. If it's false, the action is added to the approval queue instead. - An action record is written with status
executing. - When the action completes, the record is updated to
successorfailed. - An
ActionOutcomeis published toaction.outcomesfor the feedback agent.
Approving actions
Actions waiting for approval appear on /agents/actions with the decision that produced them. Anyone with the approve actions permission can approve or reject them — and so can AI agents that have been granted that permission (for example, an Incident Responder working an incident).
| Method | Path | Description |
|---|---|---|
GET | /api/v1/actions/pending | Actions waiting for approval. |
POST | /api/v1/actions/{id}/approve | Approve and execute. |
POST | /api/v1/actions/{id}/reject | Reject; nothing changes in the cluster. |
Both approval and rejection are recorded with who made them.
Permissions
The agent runs with a ServiceAccount limited to exactly what its actions need:
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["list", "get", "delete"]
- apiGroups: [""]
resources: ["nodes"]
verbs: ["list", "get", "patch"]
- apiGroups: ["apps"]
resources: ["deployments", "replicasets"]
verbs: ["list", "get", "patch", "update"]
- apiGroups: ["policy"]
resources: ["pods/eviction", "poddisruptionbudgets"]
verbs: ["create", "list"]
Inside Kubernetes it uses its in-cluster service account; for local development it uses the kubeconfig in KUBECONFIG (or ~/.kube/config).
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/actions | Action history (?cluster_id=&limit=). |
GET | /api/v1/actions/{id} | Action detail. |
GET | /api/v1/actions/pending | Actions waiting for approval. |
POST | /api/v1/actions/{id}/approve | Approve a pending action. |
POST | /api/v1/actions/{id}/reject | Reject a pending action. |
POST | /api/v1/actions/scale | Scale a deployment ({ "namespace", "deployment_name", "replicas" }). |
POST | /api/v1/actions/restart | Restart a pod ({ "namespace", "pod_name" }). |
POST | /api/v1/actions/cordon | Cordon a node ({ "node_name" }). |
POST | /api/v1/actions/uncordon | Uncordon a node ({ "node_name" }). |
POST | /api/v1/actions/drain | Drain a node ({ "node_name" }). |
POST | /api/v1/actions/rollback | Roll back a deployment ({ "namespace", "deployment_name", "revision"? }). |
GET | /healthz | Health check. |
Configuration
| Variable | Default | Description |
|---|---|---|
CANARY_ENABLED | true | Verify scale-ups and revert them if health drops. |
CANARY_SETTLE_SECONDS | 60 | How long to wait before comparing health. |
CANARY_ROLLBACK_HEALTH_DROP | 10.0 | Health-score drop that triggers a revert. |
K8S_MONITOR_URL | http://k8s-monitor:8085 | Source of health scores for verification. |
KUBECONFIG | ~/.kube/config | Kubeconfig for local development. |
AUTH_JWT_ACCESS_SECRET | — | Validates bearer tokens on the REST API. |
DATABASE_URL | — | PostgreSQL connection. |
RABBITMQ_URL | — | RabbitMQ connection. |
PORT | 8094 | HTTP port. |