Nodes Manager
Service: nodes-manager · Port: 8115 · Database schema: nodes
nodes-manager looks after the nodes your workloads run on. Working with Karpenter, it keeps the right amount and the right kind of capacity: it places workloads where they fit best, takes advantage of cheap spot capacity safely, consolidates underused nodes, and repairs node states that would otherwise need a person.
What it does continuously
Every scan cycle, nodes-manager:
- Syncs node state — every node's status, capacity, pods and pricing, for the dashboard and API.
- Scores placement — classifies each workload and checks it's on a suitable node (see below).
- Plans consolidation — finds underused nodes that can be emptied safely and, when enabled, consolidates them.
- Repairs stale cordons — uncordons nodes left cordoned by spot rebalance warnings that never became interruptions.
- Publishes events — every scaling, consolidation and repair action goes to
nodes.events.
Workload placement
Each pod is classified into a workload class — latency_sensitive, gpu, batch, stateful or stateless — from its QoS class, GPU requests, owning controller and volumes. Candidate nodes are then scored on five weighted dimensions:
| Dimension | Weight | Favors |
|---|---|---|
| Resource fit | 35% | Nodes where the pod fits without waste. |
| Topology spread | 25% | Spreading replicas across zones. |
| Cost efficiency | 20% | Spot over on-demand where appropriate. |
| Interruption risk | 15% | Stable capacity for sensitive workloads. |
| Data locality | 5% | Nodes near the pod's volumes. |
Each decision comes with a plain-language rationale, so you can see why a workload belongs where it does. nodes-manager feeds these preferences to Karpenter as node pool requirements and pod affinities.
Node pools
nodes-manager manages Karpenter NodePool and NodeClaim resources through the Kubernetes API. From the dashboard's Node Pools page (or the API) you can create pools for different workload classes — for example a spot pool for batch work and an on-demand pool for latency-sensitive services — and set their limits and disruption policies.
Spot capacity
The spot market advisor reads current spot prices and interruption frequency per instance type and zone (cached for five minutes per region), so pools use the cheapest capacity that's reliable enough for their workloads.
When AWS issues a spot rebalance recommendation, the node termination handler cordons the node as a precaution. Often no interruption follows. nodes-manager watches for nodes that are still Ready, carry only the plain unschedulable taint and have stayed that way past a grace period (15 minutes by default — well beyond a real interruption's two-minute notice), and uncordons them. Manual maintenance cordons and Karpenter's own disruption taints are never touched.
Consolidation
nodes-manager continuously looks for underused nodes whose pods can move elsewhere without breaking any PodDisruptionBudget, estimates the savings, and — when consolidation is enabled — drains and removes them through Karpenter. Review the current plan on the Node Pools page or with GET /api/v1/optimize/plan.
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/nodes | Nodes with state, resources and pod counts. |
GET | /api/v1/nodes/{id} | Node detail with its pods. |
POST | /api/v1/nodes/{id}/drain | Drain a node safely. |
GET | /api/v1/decisions | Placement decisions with scores and rationale. |
GET | /api/v1/workloads/placement | Every workload's node, class, score and rationale. |
GET | /api/v1/optimize/plan | Consolidation plan with savings and PDB impact. |
POST | /api/v1/optimize/trigger | Run an optimization cycle now. |
POST | /api/v1/simulate | Simulate scheduling N pods and see the nodes needed. |
GET · POST | /api/v1/nodepools | List or create node pools. |
GET · PATCH · DELETE | /api/v1/nodepools/{name} | Node pool detail, limits and disruption policy. |
GET | /api/v1/nodepools/{name}/nodeclaims | A pool's node claims. |
POST | /api/v1/nodepools/{name}/consolidate | Consolidate a pool now. |
GET | /api/v1/spot/market | Spot prices and interruption data (?region=). |
GET | /healthz | Health check. |
Configuration
| Variable | Default | Description |
|---|---|---|
SCAN_INTERVAL_SECONDS | 60 | Time between scan cycles. |
CONSOLIDATION_ENABLED | true | Consolidate underused nodes automatically. |
AUTO_UNCORDON_ENABLED | true | Repair stale spot-rebalance cordons. |
AUTO_UNCORDON_GRACE_PERIOD_MINUTES | 15 | How long a node must stay cordoned before it's repaired. |
AWS_REGION | — | Region for spot pricing. |
DATABASE_URL | — | PostgreSQL connection. |
RABBITMQ_URL | — | RabbitMQ connection. |
AUTH_JWT_ACCESS_SECRET | — | Validates bearer tokens. |
APP_PORT | 8115 | HTTP port. |
AWS access uses the service's workload identity (IRSA) — no static keys.