Skip to main content
Version: 1.0

Nodes Manager

Repo: o-apps/nodes-manager · Port: 8115 · DB Schema: nodes

nodes-manager owns node-lifecycle reconciliation for the cluster. It is worth being precise about what it actually does today versus what it's designed to eventually do, because the two are further apart here than for most services in this platform: a substantial amount of intelligent scheduling and Karpenter integration code exists and is exposed over its REST API, but is not yet wired into the service's own background loop, and Karpenter itself is not yet installed in any environment. What genuinely runs continuously today is node state synchronization and an auto-uncordon reconciler described below.

What actually runs continuously​

Every scan interval (SCAN_INTERVAL_SECONDS, no fixed default — set per environment), nodes-manager does two things:

  1. Syncs every Kubernetes node's state into Postgres (nodes.nodes, upserted), so the REST API below has something real to read.
  2. Runs the auto-uncordon reconciler. AWS EC2 issues a Spot rebalance recommendation when a Spot instance is at elevated risk of reclamation. aws-node-termination-handler, running as a DaemonSet, reacts to that recommendation by cordoning the affected node — but a rebalance recommendation is only a risk signal, not a guaranteed interruption, and most of the time no actual reclaim follows. Left alone, that node would sit Ready but permanently unschedulable, since nothing else in the cluster ever reverses the cordon. nodes-manager watches for exactly this pattern — a node cordoned by nothing but the plain node.kubernetes.io/unschedulable taint (never a manual maintenance cordon or, once Karpenter is installed, one of its own disruption taints), continuously Ready throughout, and past a configurable grace period (AUTO_UNCORDON_GRACE_PERIOD_MINUTES, default 15 minutes — comfortably longer than a real Spot interruption's roughly two-minute final notice) — and uncordons it automatically, recording the action as a scaling_events row and, when RabbitMQ is configured, a nodes.events message.

A heartbeat is published to the nodes.events RabbitMQ topic exchange alongside these events when RABBITMQ_URL is set.

Built, exposed over the API, but not yet wired into the running loop​

The remaining capabilities below are real, tested code — reachable through the REST endpoints in the table further down — but the background scheduling logic they represent (ScaleUpUseCase, ScaleDownUseCase, OptimizeUseCase) is constructed at startup and explicitly not invoked from the management loop yet. Calling their endpoints directly still works; nothing currently triggers them on a schedule.

  • AI-scored workload placement. Each pod is classified into a WorkloadClass (latency_sensitive, gpu, batch, stateful, or stateless, by QoS class, GPU requests, owning controller, and volume mounts) and scored across five weighted dimensions — resource fit (35%), topology spread across zones (25%), cost efficiency of spot versus on-demand (20%), spot interruption risk (15%), and data locality to existing volumes (5%) — to recommend a placement, with a plain-language rationale attached to each decision.
  • Karpenter NodePool management. NodePool and NodeClaim are accessed through the Kubernetes dynamic client, so no Karpenter Go library is required. If the CRDs aren't installed — the case in every environment as of this writing — every Karpenter-related endpoint degrades gracefully to an empty list with a log warning rather than an error, rather than failing.
  • Spot market advisor. Queries EC2's DescribeSpotPriceHistory (5-minute cache per region) for current spot pricing and a rough interruption-frequency estimate per instance type and zone.
  • Consolidation planning. Identifies underutilized nodes that would be safe to drain given PodDisruptionBudget constraints, and estimates the resulting savings.
  • On-demand AWS provisioning. A thin wrapper over RunInstances/DescribeInstanceTypes, for provisioning outside of Karpenter.

REST API​

MethodPathDescription
GET/api/v1/nodesList nodes with state, resources, and pod counts
GET/api/v1/nodes/{id}Node detail with running pods
POST/api/v1/nodes/{id}/drainGracefully drain a node
GET/api/v1/decisionsAI scheduling decision log (populated once the scheduler is wired in)
GET/api/v1/optimize/planConsolidation plan with savings and PDB impact
POST/api/v1/optimize/triggerManually trigger an optimisation cycle
POST/api/v1/simulateSimulate scheduling N pods → projected node additions
GET/api/v1/nodepoolsList Karpenter NodePools (empty until Karpenter is installed)
GET/api/v1/nodepools/{name}NodePool detail
POST/api/v1/nodepoolsCreate NodePool
PATCH/api/v1/nodepools/{name}Update NodePool limits or disruption policy
DELETE/api/v1/nodepools/{name}Delete NodePool
GET/api/v1/nodepools/{name}/nodeclaimsList NodeClaims for a pool
POST/api/v1/nodepools/{name}/consolidateTrigger Karpenter consolidation
GET/api/v1/spot/marketSpot pricing and interruption data by zone
GET/api/v1/workloads/placementAll workloads with node, class, score, and rationale
GET/healthzHealth check

Environment Variables​

VariableDefaultDescription
DATABASE_URL—PostgreSQL connection
AUTH_JWT_ACCESS_SECRET—Required — shared JWT signing secret for the HTTP API
RABBITMQ_URL—Optional — enables publishing to nodes.events
AUTO_UNCORDON_ENABLEDtrueEnables the reconciler described above
AUTO_UNCORDON_GRACE_PERIOD_MINUTES15How long a node must have been cordoned, with nothing but the bare unschedulable taint, before it's reversed
KUBECONFIGin-clusterPath to kubeconfig (local development only)
AWS_REGION—Region for spot pricing queries
AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY—AWS credentials for the on-demand provisioner
APP_PORT8115HTTP port

A note on the service's history​

This service has, for most of its deployment, been silently non-functional: its Docker image never actually shipped its own database migrations, so every startup ran goose against an empty directory, logged a warning, and continued — meaning nodes.nodes and the other tables never existed, and every read or write against them failed. A second, related bug (missing type:jsonb annotations on the ORM's JSON-holding columns) was masked entirely by the first, and only surfaced once the migration itself was fixed. Both are now fixed, and the auto-uncordon reconciler described above has been confirmed live: during its own rollout verification, it found several real nodes that had sat cordoned for 33 minutes to over an hour from an actual Spot rebalance event, and correctly uncordoned all of them with no manual intervention.