Skip to main content
Version: 1.0

Agent Runtime

Repo: o-apps/agent-runtime · Port: 8111 (deployed; the code's own fallback default is a stale 8097, left over from before this service split off from k8s-optimizer, which really does own port 8097 today) · DB Schema: agents

agent-runtime is the reasoning engine. It runs Claude LLM agents that can investigate clusters, correlate signals across services, and produce prioritised recommendations using the Anthropic BetaToolRunnerStreaming API.

Agent Types​

AgentModelToolsMax iterationsThinking
sre_orchestratorClaude Sonnet 4.6All 3420Yes (interleaved)
security_auditorClaude Haiku 4.55 (security)10No
cost_optimizerClaude Haiku 4.55 (cost)10No
incident_responderClaude Sonnet 4.613 (incident + write)15Yes (interleaved)
app_advisorClaude Sonnet 4.67 (app advisor)12No
load_test_analystClaude Sonnet 4.67 (performance)10No
node_opsClaude Sonnet 4.610 (node management)15Yes (interleaved)

Interleaved thinking (interleaved-thinking-2025-05-14 beta) enables SRE Orchestrator, Incident Responder, and NodeOps to think before each tool call — producing more accurate tool selection and better reasoning chains. Thinking blocks are streamed to the UI as collapsible cards.

Tool Subsets​

SRE Orchestrator — all 34 tools (Tools: nil in the agent registry means "no restriction"). Full visibility across every service, including App Advisor's own tools.

Security Auditor — get_security_posture, get_cluster_list, get_anomaly_events, get_incidents, get_agent_telemetry

Cost Optimizer — get_cluster_cost, get_optimization_report, get_scaling_forecasts, get_cluster_list, get_multi_cluster_overview

Incident Responder — get_incidents, get_anomaly_events, get_cluster_health, get_pod_metrics, get_node_metrics, get_action_log, get_recommendations, get_agent_telemetry, execute_runbook, get_runbook_execution, generate_post_incident_report, drain_node_action, rollback_deployment

App Advisor — get_app_profile, get_app_advice, get_app_live_metrics, get_cluster_health, get_pod_metrics, get_anomaly_events, get_recommendations

NodeOps — get_node_pools, get_node_claims, get_scheduling_decisions, get_spot_market, get_consolidation_plan, get_workload_placement, get_cluster_health, get_incidents, get_anomaly_events, drain_node_action

Load Test Analyst — get_load_test_analysis, get_optimization_status, get_cluster_health, get_pod_metrics, get_node_metrics, get_scaling_forecasts, get_anomaly_events

All 34 Tools​

Cluster & Cost Observability​

ToolSource serviceParameters
get_cluster_healthk8s-monitor /api/healthnone
get_cluster_costk8s-monitor /api/costnone
get_optimization_reportk8s-monitor /api/optimizernamespace?, view?
get_pod_metricsk8s-monitor /api/metrics/podsnamespace?
get_node_metricsk8s-monitor /api/metrics/nodesnone
get_cluster_listkubeopera-api /api/v1/clustersnone
get_multi_cluster_overviewkubeopera-api + k8s-monitornone

Security & Pipelines​

ToolSource serviceParameters
get_security_posturesecurity-api /api/v1/posture/summarycluster_id?
get_pipeline_statuscicd-gateway /api/v1/pipelinescluster_id?, limit?

Anomaly Detection & Alerts​

ToolSource serviceParameters
get_anomaly_eventsanomaly-detector /api/v1/anomaliescluster_id?, severity?, limit?
get_alert_rulesanomaly-detector /api/v1/alert-rulescluster_id?
get_scaling_forecastspredictive-scaler /api/v1/scaling-decisionscluster_id?, status?

Incident Management​

ToolSource serviceParameters
get_incidentsincident-manager /api/v1/incidentscluster_id?, status?
execute_runbookincident-manager /api/v1/incidents/{id}/runbooks/{rbId}/executeincident_id, runbook_id, started_by?
get_runbook_executionincident-manager /api/v1/executions/{execId}execution_id
generate_post_incident_reportincident-manager /api/v1/incidents/{id}/pirincident_id

Node Management​

ToolSource serviceParameters
get_node_poolsnodes-manager /api/v1/nodepoolsnone
get_node_claimsnodes-manager /api/v1/nodepools/{name}/nodeclaimspool_name
get_scheduling_decisionsnodes-manager /api/v1/decisionscluster_id?, limit?
get_spot_marketnodes-manager /api/v1/spot/marketregion?
get_consolidation_plannodes-manager /api/v1/optimize/plannone
get_workload_placementnodes-manager /api/v1/workloads/placementnone

Write Actions​

ToolSource serviceParameters
drain_node_actionaction-agent-srv /api/v1/actions/drainnode_name
rollback_deploymentaction-agent-srv /api/v1/actions/rollbacknamespace, deployment_name

Agentic Pipeline​

ToolSource serviceParameters
get_agent_telemetryobservability-agent /api/v1/telemetrycluster_id?, limit?
get_analysis_resultsanalysis-agent /api/v1/resultscluster_id?, limit?
get_action_logaction-agent /api/v1/actionscluster_id?, limit?
get_feedback_outcomesfeedback-agent /api/v1/feedbackcluster_id?, limit?
get_recommendationsrecommendation-agent /api/v1/recommendationscluster_id?, limit?

Continuous Optimisation (kubeopera-ai)​

ToolSource serviceParameters
get_optimization_statuskubeopera-ai /api/v1/statusnone
get_load_test_analysiskubeopera-ai /api/v1/analyze/load-testresults_path?

App Advisor​

ToolSource serviceParameters
get_app_profileapp-advisor-srv /api/v1/apps/{id}/profileapp_id
get_app_adviceapp-advisor-srv /api/v1/apps/{id}/adviceapp_id, status?, limit?
get_app_live_metricsapp-advisor-srv /api/v1/apps/{id}/liveapp_id

A same-named get_app_live_metrics also appears in the frontend AI chat's own tool list (app/api/ai/chat/route.ts) — that's a separate, unrelated definition specific to the browser-facing chat sidebar, and as of this writing it's a dead reference there (listed in the customer tool allowlist but never actually added to the chat's own TOOLS array). The version documented here, used by agent-runtime's App Advisor and SRE Orchestrator agent types, is fully implemented and calls a real endpoint.

Streaming Architecture​

POST /api/v1/runs
→ creates AgentRun (status: pending)
→ starts goroutine: executeRun()
→ returns { run_id, status, stream_url }

executeRun() goroutine:
→ RunManager.NewSink(runID) → broadcast channel
→ StreamingRunner.Execute(ctx, run, sink)
→ BetaToolRunnerStreaming.AllStreaming(ctx)
for each turn:
for each event (text_delta, thinking_delta, tool_call, tool_result):
→ instrumentedTool.Execute() → emits SSE events to sink
→ sink → RunManager.Broadcast() → all subscribers
→ on complete: update AgentRun.status = completed

GET /api/v1/runs/{id}/stream
→ RunManager.Subscribe(runID) → receive channel
→ HTTP Flusher loop: write "data: {json}\n\n" for each SSE event
→ closes on "done" or "error" event

The RunManager fan-out pattern means multiple browser tabs or SSE clients can subscribe to the same run simultaneously without affecting agent execution.

Auto-Trigger​

When analysis-agent publishes an insight with RiskScore > 70.0 to the analysis.insights exchange, agent-runtime automatically creates an SRE Orchestrator run with trigger: "auto" and a generated prompt:

Cluster {clusterID} has elevated risk score {score}/100.
Key anomalies: {N} detected. Summary: {summary}.
Investigate and recommend remediation.

REST API​

MethodPathDescription
POST/api/v1/runsCreate and start an agent run
GET/api/v1/runsList runs (?agent_type=&status=&cluster_id=)
GET/api/v1/runs/{id}Run detail
GET/api/v1/runs/{id}/streamSSE stream of events
GET/api/v1/runs/{id}/tool-callsPersisted tool call history

System Prompts​

SRE Orchestrator — ReAct framework. Starts with broad situational awareness, cross-correlates signals, outputs: Situation Summary / Key Findings / Root Cause / Recommended Actions / Monitoring Checkpoints.

Security Auditor — focuses on vulnerabilities, RBAC misconfigurations, network policy gaps, and compliance violations. Distinguishes immediate threats from technical debt.

Cost Optimizer — FinOps framework: right-sizing, waste elimination, scaling efficiency, multi-cluster arbitrage. Outputs quick wins (this week) and strategic optimizations.

Incident Responder — OODA loop (Observe / Orient / Decide / Act). Assesses blast radius, identifies root cause from correlated signals, provides kubectl commands plus runbook steps. Can execute runbooks and generate post-incident reports directly.

App Advisor — focuses on one application's own domain profile and health, using business-context-aware advice rather than generic cluster-wide recommendations. A separate, independent implementation reaches the same app-advisor-srv data from the frontend's customer-facing AI chat sidebar (get_app_profile/get_app_advice) — that's a direct tool-use loop in the frontend itself, not this agent type, though the two are functionally similar in intent.

NodeOps — node pool specialist. Reviews NodeClaim phases, scheduling scores, spot market savings, consolidation feasibility, and PDB constraints. Can drain nodes when safe.

Load Test Analyst — performance engineering expert. Fetches the latest kubeopera-ai optimisation status as a baseline, triggers load test analysis for p95/p99 assessment, cross-references with live pod/node metrics to identify saturation points, and produces a structured performance report with a numbered remediation action list ordered by priority.

Environment Variables​

VariableDescription
ANTHROPIC_API_KEYRequired. Anthropic API key
DATABASE_URLPostgreSQL connection
RABBITMQ_URLOptional. Enables auto-trigger from analysis insights
K8S_MONITOR_BASE_URLk8s-monitor URL (default: http://localhost:8085)
SECURITY_API_BASE_URLsecurity-api URL (default: http://localhost:8086)
CICD_GATEWAY_BASE_URLcicd-gateway URL (default: http://localhost:8087)
KUBEOPERA_API_BASE_URLkubeopera-api URL (default: http://localhost:8080)
ANOMALY_DETECTOR_BASE_URLanomaly-detector URL (default: http://localhost:8088)
PREDICTIVE_SCALER_BASE_URLpredictive-scaler URL (default: http://localhost:8089)
INCIDENT_MANAGER_BASE_URLincident-manager URL (default: http://localhost:8090)
NODES_MANAGER_BASE_URLnodes-manager URL (code default: http://localhost:8098 — a local-dev value; the deployed service actually listens on 8115)
OBSERVABILITY_AGENT_SRV_BASE_URLobservability-agent URL (default: http://localhost:8092)
ANALYSIS_AGENT_SRV_BASE_URLanalysis-agent URL (default: http://localhost:8093)
ACTION_AGENT_SRV_BASE_URLaction-agent URL (default: http://localhost:8094)
FEEDBACK_AGENT_SRV_BASE_URLfeedback-agent URL (default: http://localhost:8095)
RECOMMENDATION_AGENT_SRV_BASE_URLrecommendation-agent URL (default: http://localhost:8096)
KUBEOPERA_AI_BASE_URLkubeopera-ai URL (code default: http://localhost:8104 — a local-dev value; the deployed service actually listens on 8113)
APP_ADVISOR_BASE_URLapp-advisor-srv URL, needed for the App Advisor agent type's tools (default: http://localhost:8105)
PORTHTTP port (deployed value: 8111; the code's own fallback default is a stale 8097)