Skip to main content
Version: 2.0

Agent Runtime

Service: agent-runtime · Port: 8111 · Database schema: agents

agent-runtime is KubeOpera's reasoning engine. It runs Claude-powered agents that investigate your clusters the way an experienced engineer would: form a hypothesis, gather evidence with live tools, refine, and conclude — then recommend (or, where allowed, take) action. Every step is streamed live and saved.

How an agent works​

Agents follow the ReAct loop:

┌──────────┐    ┌──────────┐    ┌──────────┐
│ Reason │ → │ Act │ → │ Observe │ ─┐
│ (think) │ │(call tool)│ │ (result) │ │
└──────────┘ └──────────┘ └──────────┘ │
↑ │
└──────────── until it can answer ───────┘
  1. The agent receives a goal (your prompt) and context (the cluster, and for automatic runs the triggering insight).
  2. It thinks about what it needs to know.
  3. It calls a tool — for example, get_cluster_health or get_incidents.
  4. It reads the result and decides the next step.
  5. It repeats until it has an answer, or until it reaches its iteration limit.
  6. It writes its findings: a situation summary, root cause, and recommended actions.

Agents with interleaved thinking reason between every tool call, which produces better tool choices on complex investigations.

Agent types​

AgentModelTool scopeMax iterationsInterleaved thinking
sre_orchestratorClaude Sonnet 4.6All tools20Yes
incident_responderClaude Sonnet 4.6Incidents, diagnostics, runbooks, write actions15Yes
node_opsClaude Sonnet 4.6Node management, diagnostics, drain15Yes
app_advisorClaude Sonnet 4.6App Advisor, diagnostics12No
load_test_analystClaude Sonnet 4.6Performance, diagnostics10No
security_auditorClaude Haiku 4.5Security10No
cost_optimizerClaude Haiku 4.5Cost, scaling decisions10No

What each agent focuses on​

  • SRE Orchestrator — starts broad, correlates signals across every service, and reports: Situation summary · Key findings · Root cause · Recommended actions · Monitoring checkpoints.
  • Incident Responder — works an incident with the OODA loop (observe, orient, decide, act): assesses blast radius, finds the root cause, runs runbooks, and writes the post-incident review.
  • NodeOps — reviews node pools and claims, scheduling, spot savings and consolidation, respecting PodDisruptionBudgets; drains nodes when it's safe.
  • App Advisor — looks at one application in its business context and gives specific, prioritized advice.
  • Load Test Analyst — compares load-test results with the Optimizer's baseline and live metrics, finds saturation points, and produces a prioritized performance report.
  • Security Auditor — separates immediate threats from technical debt across vulnerabilities, RBAC, network policy and compliance.
  • Cost Optimizer — finds quick wins and strategic savings: right-sizing, waste, scaling efficiency and multi-cluster placement; reviews pending scaling decisions.

Tools​

Agents only have the tools for their domain; the SRE Orchestrator has them all. Write tools change your cluster or your records, so they are only available to agents — and users — with the matching permission, and every call is audited.

Cluster and cost​

ToolSourceParameters
get_cluster_healthk8s-monitor—
get_cluster_costk8s-monitor—
get_optimization_reportk8s-monitornamespace?, view?
get_pod_metricsk8s-monitornamespace?
get_node_metricsk8s-monitor—
get_cluster_listkubeopera-api—
get_multi_cluster_overviewkubeopera-api, k8s-monitor—

Security and pipelines​

ToolSourceParameters
get_security_posturesecurity-apicluster_id?
get_pipeline_statuscicd-gatewaycluster_id?, limit?

Anomalies and scaling​

ToolSourceParameters
get_anomaly_eventsanomaly-detectorcluster_id?, severity?, limit?
get_alert_rulesanomaly-detectorcluster_id?
get_scaling_forecastspredictive-scalercluster_id?, status?
approve_scaling_decision ✎predictive-scalerdecision_id
reject_scaling_decision ✎predictive-scalerdecision_id, reason?

Incidents​

ToolSourceParameters
get_incidentsincident-managercluster_id?, status?
execute_runbook ✎incident-managerincident_id, runbook_id
get_runbook_executionincident-managerexecution_id
generate_post_incident_report ✎incident-managerincident_id

Nodes​

ToolSourceParameters
get_node_poolsnodes-manager—
get_node_claimsnodes-managerpool_name
get_scheduling_decisionsnodes-managercluster_id?, limit?
get_spot_marketnodes-managerregion?
get_consolidation_plannodes-manager—
get_workload_placementnodes-manager—

Actions​

ToolSourceParameters
drain_node_action ✎action-agentnode_name
rollback_deployment ✎action-agentnamespace, deployment_name, revision?
approve_action ✎action-agentaction_id
reject_action ✎action-agentaction_id, reason?

Reactive pipeline​

ToolSourceParameters
get_agent_telemetryobservability-agentcluster_id?, limit?
get_analysis_resultsanalysis-agentcluster_id?, limit?
get_action_logaction-agentcluster_id?, limit?
get_feedback_outcomesfeedback-agentcluster_id?, limit?
get_recommendationsrecommendation-agentcluster_id?, limit?

Optimizer​

ToolSourceParameters
get_optimization_statuskubeopera-ai—
get_load_test_analysiskubeopera-airesults_path?

App Advisor​

ToolSourceParameters
get_app_profileapp-advisor-srvapp_id
get_app_adviceapp-advisor-srvapp_id, status?, limit?
get_app_live_metricsapp-advisor-srvapp_id

✎ = write tool.

Running agents​

Start a run​

POST /api/v1/runs
Content-Type: application/json

{
"agent_type": "incident_responder",
"cluster_id": "prod-us-east",
"prompt": "Incident INC-142: payments-api 5xx rate above 2%. Find the cause and mitigate."
}

The response includes the run_id and a stream_url. The run executes in the background; you can close the browser and come back.

Watch it live​

curl -N /api/v1/runs/{run_id}/stream

The stream delivers thinking, tool_call, tool_result, text, done and error events (see AI Agents UI for the format). Any number of clients can watch the same run at once without affecting it, and clients that reconnect receive the events they missed.

Automatic runs​

When the analysis agent publishes an insight with a risk score above the auto-trigger threshold (70 by default), agent-runtime starts an SRE Orchestrator run with trigger: "auto":

Cluster {clusterID} has elevated risk score {score}/100.
Key anomalies: {N} detected. Summary: {summary}.
Investigate and recommend remediation.

To avoid duplicate investigations, only one automatic run is started per cluster while a previous one is still running.

How streaming works internally​

POST /api/v1/runs
→ create AgentRun (status: pending) → start executeRun() → return run_id

executeRun():
→ RunManager.NewSink(runID)
→ StreamingRunner.Execute(ctx, run, sink)
for each turn and each event (thinking, text, tool_call, tool_result):
→ instrumentedTool records the call in agents.tool_calls
→ sink → RunManager.Broadcast() → every subscriber
→ AgentRun.status = completed | failed

GET /api/v1/runs/{id}/stream
→ replay stored events, then subscribe to live ones
→ write "data: {json}\n\n" until "done" or "error"

AI credentials​

Each run uses the AI credential of the tenant it runs for — the tenant's own provider key if they've configured one, otherwise the platform key within the tenant's quota — resolved from auth-service for every run. Usage is attributed to the tenant.

REST API​

MethodPathDescription
POST/api/v1/runsCreate and start a run.
GET/api/v1/runsList runs (?agent_type=&status=&cluster_id=).
GET/api/v1/runs/{id}Run detail and findings.
GET/api/v1/runs/{id}/streamServer-Sent Events stream.
GET/api/v1/runs/{id}/tool-callsStored tool-call history.
POST/api/v1/runs/{id}/cancelStop a running agent.
GET/healthzHealth check.

Configuration​

VariableDefaultDescription
PORT8111HTTP port.
DATABASE_URL—PostgreSQL connection.
RABBITMQ_URL—Enables automatic runs from analysis insights.
AUTO_TRIGGER_RISK_SCORE70Risk score above which an SRE investigation starts automatically.
AUTH_SERVICE_BASE_URL—auth-service, for AI credential resolution.
AI_CREDENTIAL_INTERNAL_API_KEY—Authenticates credential resolution calls.
K8S_MONITOR_BASE_URLhttp://k8s-monitor:8085
SECURITY_API_BASE_URLhttp://security-api:8086
CICD_GATEWAY_BASE_URLhttp://cicd-gateway:8087
KUBEOPERA_API_BASE_URLhttp://kubeopera-api:8090
ANOMALY_DETECTOR_BASE_URLhttp://anomaly-detector:8088
PREDICTIVE_SCALER_BASE_URLhttp://predictive-scaler:8089
INCIDENT_MANAGER_BASE_URLhttp://incident-manager:8090
NODES_MANAGER_BASE_URLhttp://nodes-manager:8115
OBSERVABILITY_AGENT_SRV_BASE_URLhttp://observability-agent-srv:8092
ANALYSIS_AGENT_SRV_BASE_URLhttp://analysis-agent-srv:8093
ACTION_AGENT_SRV_BASE_URLhttp://action-agent-srv:8094
FEEDBACK_AGENT_SRV_BASE_URLhttp://feedback-agent-srv:8095
RECOMMENDATION_AGENT_SRV_BASE_URLhttp://recommendation-agent-srv:8096
KUBEOPERA_AI_BASE_URLhttp://kubeopera-ai:8113
APP_ADVISOR_BASE_URLhttp://app-advisor-srv:8105