Skip to main content
Version: 2.0

Nodes Manager

Service: nodes-manager · Port: 8115 · Database schema: nodes

nodes-manager looks after the nodes your workloads run on. Working with Karpenter, it keeps the right amount and the right kind of capacity: it places workloads where they fit best, takes advantage of cheap spot capacity safely, consolidates underused nodes, and repairs node states that would otherwise need a person.

What it does continuously​

Every scan cycle, nodes-manager:

  1. Syncs node state — every node's status, capacity, pods and pricing, for the dashboard and API.
  2. Scores placement — classifies each workload and checks it's on a suitable node (see below).
  3. Plans consolidation — finds underused nodes that can be emptied safely and, when enabled, consolidates them.
  4. Repairs stale cordons — uncordons nodes left cordoned by spot rebalance warnings that never became interruptions.
  5. Publishes events — every scaling, consolidation and repair action goes to nodes.events.

Workload placement​

Each pod is classified into a workload class — latency_sensitive, gpu, batch, stateful or stateless — from its QoS class, GPU requests, owning controller and volumes. Candidate nodes are then scored on five weighted dimensions:

DimensionWeightFavors
Resource fit35%Nodes where the pod fits without waste.
Topology spread25%Spreading replicas across zones.
Cost efficiency20%Spot over on-demand where appropriate.
Interruption risk15%Stable capacity for sensitive workloads.
Data locality5%Nodes near the pod's volumes.

Each decision comes with a plain-language rationale, so you can see why a workload belongs where it does. nodes-manager feeds these preferences to Karpenter as node pool requirements and pod affinities.

Node pools​

nodes-manager manages Karpenter NodePool and NodeClaim resources through the Kubernetes API. From the dashboard's Node Pools page (or the API) you can create pools for different workload classes — for example a spot pool for batch work and an on-demand pool for latency-sensitive services — and set their limits and disruption policies.

Spot capacity​

The spot market advisor reads current spot prices and interruption frequency per instance type and zone (cached for five minutes per region), so pools use the cheapest capacity that's reliable enough for their workloads.

When AWS issues a spot rebalance recommendation, the node termination handler cordons the node as a precaution. Often no interruption follows. nodes-manager watches for nodes that are still Ready, carry only the plain unschedulable taint and have stayed that way past a grace period (15 minutes by default — well beyond a real interruption's two-minute notice), and uncordons them. Manual maintenance cordons and Karpenter's own disruption taints are never touched.

Consolidation​

nodes-manager continuously looks for underused nodes whose pods can move elsewhere without breaking any PodDisruptionBudget, estimates the savings, and — when consolidation is enabled — drains and removes them through Karpenter. Review the current plan on the Node Pools page or with GET /api/v1/optimize/plan.

REST API​

MethodPathDescription
GET/api/v1/nodesNodes with state, resources and pod counts.
GET/api/v1/nodes/{id}Node detail with its pods.
POST/api/v1/nodes/{id}/drainDrain a node safely.
GET/api/v1/decisionsPlacement decisions with scores and rationale.
GET/api/v1/workloads/placementEvery workload's node, class, score and rationale.
GET/api/v1/optimize/planConsolidation plan with savings and PDB impact.
POST/api/v1/optimize/triggerRun an optimization cycle now.
POST/api/v1/simulateSimulate scheduling N pods and see the nodes needed.
GET · POST/api/v1/nodepoolsList or create node pools.
GET · PATCH · DELETE/api/v1/nodepools/{name}Node pool detail, limits and disruption policy.
GET/api/v1/nodepools/{name}/nodeclaimsA pool's node claims.
POST/api/v1/nodepools/{name}/consolidateConsolidate a pool now.
GET/api/v1/spot/marketSpot prices and interruption data (?region=).
GET/healthzHealth check.

Configuration​

VariableDefaultDescription
SCAN_INTERVAL_SECONDS60Time between scan cycles.
CONSOLIDATION_ENABLEDtrueConsolidate underused nodes automatically.
AUTO_UNCORDON_ENABLEDtrueRepair stale spot-rebalance cordons.
AUTO_UNCORDON_GRACE_PERIOD_MINUTES15How long a node must stay cordoned before it's repaired.
AWS_REGION—Region for spot pricing.
DATABASE_URL—PostgreSQL connection.
RABBITMQ_URL—RabbitMQ connection.
AUTH_JWT_ACCESS_SECRET—Validates bearer tokens.
APP_PORT8115HTTP port.

AWS access uses the service's workload identity (IRSA) — no static keys.