Skip to main content
Version: 2.0

Action Agent

Service: action-agent-srv · Port: 8094

The action agent is the part of the reactive pipeline that changes your cluster. It carries out the analysis agent's decisions — immediately for decisions approved to run automatically, and after approval for the rest — and it exposes the same actions to people and AI agents through its API.

Every action is safe by design: it respects PodDisruptionBudgets, verifies its own effect, and is recorded with its reason and outcome.

Actions​

ActionWhat it doesTriggered by
Scale deploymentChanges a Deployment's replica count.Decisions, API, agents
Restart podDeletes a pod so its controller recreates it.Decisions, API, agents
Cordon nodeStops new pods being scheduled on a node.Decisions, API, agents
Uncordon nodeAllows scheduling on a node again.API, agents
Drain nodeCordons a node and safely evicts its pods.API, agents
Roll back deploymentReturns a Deployment to its previous revision.API, agents

Scale deployment​

Sets the Deployment's replicas to a target — or, without a target, to the current count plus one.

{ "spec": { "replicas": 5 } }

Scale-downs check every PodDisruptionBudget in the namespace first and are refused if they would break one.

Scale-ups are verified. After scaling up, the agent waits (60 seconds by default), then compares the cluster's health score with its value before the action. If health dropped by more than the allowed amount (10 points by default), the scale-up is reverted automatically and the outcome is recorded as failed.

Restart pod​

Deletes the pod; its Deployment, StatefulSet or DaemonSet immediately creates a fresh one. This clears stuck state without touching the rest of the workload.

Cordon and uncordon node​

Cordon marks a node unschedulable (spec.unschedulable: true): no new pods are placed on it, and existing pods keep running. Uncordon reverses this.

note

Cordoning does not move existing pods. Use drain when you need the node empty.

Drain node​

Prepares a node for maintenance or removal:

  1. Cordons the node.
  2. Lists every pod on it except DaemonSet pods (which belong on every node).
  3. Evicts them one at a time through the Kubernetes Eviction API, with a 30-second grace period.

Eviction respects PodDisruptionBudgets. When an eviction is blocked, the agent retries with exponential backoff (5 seconds, doubling to a 60-second cap, up to 5 attempts), so a node with tightly budgeted workloads may take a few minutes to drain.

POST /api/v1/actions/drain   { "node_name": "ip-10-0-1-45.ec2.internal" }

Roll back deployment​

Rolls a Deployment back to its previous revision — the ReplicaSet it ran before the latest rollout — as a zero-downtime rolling update. Pass a revision to return to a specific earlier revision.

POST /api/v1/actions/rollback   { "namespace": "production", "deployment_name": "payments-api" }

How decisions become actions​

For each decision consumed from analysis.decisions:

  1. If AutoApprove is true, the agent executes the action. If it's false, the action is added to the approval queue instead.
  2. An action record is written with status executing.
  3. When the action completes, the record is updated to success or failed.
  4. An ActionOutcome is published to action.outcomes for the feedback agent.

Approving actions​

Actions waiting for approval appear on /agents/actions with the decision that produced them. Anyone with the approve actions permission can approve or reject them — and so can AI agents that have been granted that permission (for example, an Incident Responder working an incident).

MethodPathDescription
GET/api/v1/actions/pendingActions waiting for approval.
POST/api/v1/actions/{id}/approveApprove and execute.
POST/api/v1/actions/{id}/rejectReject; nothing changes in the cluster.

Both approval and rejection are recorded with who made them.

Permissions​

The agent runs with a ServiceAccount limited to exactly what its actions need:

rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["list", "get", "delete"]
- apiGroups: [""]
resources: ["nodes"]
verbs: ["list", "get", "patch"]
- apiGroups: ["apps"]
resources: ["deployments", "replicasets"]
verbs: ["list", "get", "patch", "update"]
- apiGroups: ["policy"]
resources: ["pods/eviction", "poddisruptionbudgets"]
verbs: ["create", "list"]

Inside Kubernetes it uses its in-cluster service account; for local development it uses the kubeconfig in KUBECONFIG (or ~/.kube/config).

REST API​

MethodPathDescription
GET/api/v1/actionsAction history (?cluster_id=&limit=).
GET/api/v1/actions/{id}Action detail.
GET/api/v1/actions/pendingActions waiting for approval.
POST/api/v1/actions/{id}/approveApprove a pending action.
POST/api/v1/actions/{id}/rejectReject a pending action.
POST/api/v1/actions/scaleScale a deployment ({ "namespace", "deployment_name", "replicas" }).
POST/api/v1/actions/restartRestart a pod ({ "namespace", "pod_name" }).
POST/api/v1/actions/cordonCordon a node ({ "node_name" }).
POST/api/v1/actions/uncordonUncordon a node ({ "node_name" }).
POST/api/v1/actions/drainDrain a node ({ "node_name" }).
POST/api/v1/actions/rollbackRoll back a deployment ({ "namespace", "deployment_name", "revision"? }).
GET/healthzHealth check.

Configuration​

VariableDefaultDescription
CANARY_ENABLEDtrueVerify scale-ups and revert them if health drops.
CANARY_SETTLE_SECONDS60How long to wait before comparing health.
CANARY_ROLLBACK_HEALTH_DROP10.0Health-score drop that triggers a revert.
K8S_MONITOR_URLhttp://k8s-monitor:8085Source of health scores for verification.
KUBECONFIG~/.kube/configKubeconfig for local development.
AUTH_JWT_ACCESS_SECRET—Validates bearer tokens on the REST API.
DATABASE_URL—PostgreSQL connection.
RABBITMQ_URL—RabbitMQ connection.
PORT8094HTTP port.