Agent Runtime
Repo: o-apps/agent-runtime · Port: 8111 (deployed; the code's own fallback default is a stale 8097, left over from before this service split off from k8s-optimizer, which really does own port 8097 today) · DB Schema: agents
agent-runtime is the reasoning engine. It runs Claude LLM agents that can investigate clusters, correlate signals across services, and produce prioritised recommendations using the Anthropic BetaToolRunnerStreaming API.
Agent Types
| Agent | Model | Tools | Max iterations | Thinking |
|---|---|---|---|---|
sre_orchestrator | Claude Sonnet 4.6 | All 34 | 20 | Yes (interleaved) |
security_auditor | Claude Haiku 4.5 | 5 (security) | 10 | No |
cost_optimizer | Claude Haiku 4.5 | 5 (cost) | 10 | No |
incident_responder | Claude Sonnet 4.6 | 13 (incident + write) | 15 | Yes (interleaved) |
app_advisor | Claude Sonnet 4.6 | 7 (app advisor) | 12 | No |
load_test_analyst | Claude Sonnet 4.6 | 7 (performance) | 10 | No |
node_ops | Claude Sonnet 4.6 | 10 (node management) | 15 | Yes (interleaved) |
Interleaved thinking (interleaved-thinking-2025-05-14 beta) enables SRE Orchestrator, Incident Responder, and NodeOps to think before each tool call — producing more accurate tool selection and better reasoning chains. Thinking blocks are streamed to the UI as collapsible cards.
Tool Subsets
SRE Orchestrator — all 34 tools (Tools: nil in the agent registry means "no restriction"). Full visibility across every service, including App Advisor's own tools.
Security Auditor — get_security_posture, get_cluster_list, get_anomaly_events, get_incidents, get_agent_telemetry
Cost Optimizer — get_cluster_cost, get_optimization_report, get_scaling_forecasts, get_cluster_list, get_multi_cluster_overview
Incident Responder — get_incidents, get_anomaly_events, get_cluster_health, get_pod_metrics, get_node_metrics, get_action_log, get_recommendations, get_agent_telemetry, execute_runbook, get_runbook_execution, generate_post_incident_report, drain_node_action, rollback_deployment
App Advisor — get_app_profile, get_app_advice, get_app_live_metrics, get_cluster_health, get_pod_metrics, get_anomaly_events, get_recommendations
NodeOps — get_node_pools, get_node_claims, get_scheduling_decisions, get_spot_market, get_consolidation_plan, get_workload_placement, get_cluster_health, get_incidents, get_anomaly_events, drain_node_action
Load Test Analyst — get_load_test_analysis, get_optimization_status, get_cluster_health, get_pod_metrics, get_node_metrics, get_scaling_forecasts, get_anomaly_events
All 34 Tools
Cluster & Cost Observability
| Tool | Source service | Parameters |
|---|---|---|
get_cluster_health | k8s-monitor /api/health | none |
get_cluster_cost | k8s-monitor /api/cost | none |
get_optimization_report | k8s-monitor /api/optimizer | namespace?, view? |
get_pod_metrics | k8s-monitor /api/metrics/pods | namespace? |
get_node_metrics | k8s-monitor /api/metrics/nodes | none |
get_cluster_list | kubeopera-api /api/v1/clusters | none |
get_multi_cluster_overview | kubeopera-api + k8s-monitor | none |
Security & Pipelines
| Tool | Source service | Parameters |
|---|---|---|
get_security_posture | security-api /api/v1/posture/summary | cluster_id? |
get_pipeline_status | cicd-gateway /api/v1/pipelines | cluster_id?, limit? |
Anomaly Detection & Alerts
| Tool | Source service | Parameters |
|---|---|---|
get_anomaly_events | anomaly-detector /api/v1/anomalies | cluster_id?, severity?, limit? |
get_alert_rules | anomaly-detector /api/v1/alert-rules | cluster_id? |
get_scaling_forecasts | predictive-scaler /api/v1/scaling-decisions | cluster_id?, status? |
Incident Management
| Tool | Source service | Parameters |
|---|---|---|
get_incidents | incident-manager /api/v1/incidents | cluster_id?, status? |
execute_runbook | incident-manager /api/v1/incidents/{id}/runbooks/{rbId}/execute | incident_id, runbook_id, started_by? |
get_runbook_execution | incident-manager /api/v1/executions/{execId} | execution_id |
generate_post_incident_report | incident-manager /api/v1/incidents/{id}/pir | incident_id |
Node Management
| Tool | Source service | Parameters |
|---|---|---|
get_node_pools | nodes-manager /api/v1/nodepools | none |
get_node_claims | nodes-manager /api/v1/nodepools/{name}/nodeclaims | pool_name |
get_scheduling_decisions | nodes-manager /api/v1/decisions | cluster_id?, limit? |
get_spot_market | nodes-manager /api/v1/spot/market | region? |
get_consolidation_plan | nodes-manager /api/v1/optimize/plan | none |
get_workload_placement | nodes-manager /api/v1/workloads/placement | none |
Write Actions
| Tool | Source service | Parameters |
|---|---|---|
drain_node_action | action-agent-srv /api/v1/actions/drain | node_name |
rollback_deployment | action-agent-srv /api/v1/actions/rollback | namespace, deployment_name |
Agentic Pipeline
| Tool | Source service | Parameters |
|---|---|---|
get_agent_telemetry | observability-agent /api/v1/telemetry | cluster_id?, limit? |
get_analysis_results | analysis-agent /api/v1/results | cluster_id?, limit? |
get_action_log | action-agent /api/v1/actions | cluster_id?, limit? |
get_feedback_outcomes | feedback-agent /api/v1/feedback | cluster_id?, limit? |
get_recommendations | recommendation-agent /api/v1/recommendations | cluster_id?, limit? |
Continuous Optimisation (kubeopera-ai)
| Tool | Source service | Parameters |
|---|---|---|
get_optimization_status | kubeopera-ai /api/v1/status | none |
get_load_test_analysis | kubeopera-ai /api/v1/analyze/load-test | results_path? |
App Advisor
| Tool | Source service | Parameters |
|---|---|---|
get_app_profile | app-advisor-srv /api/v1/apps/{id}/profile | app_id |
get_app_advice | app-advisor-srv /api/v1/apps/{id}/advice | app_id, status?, limit? |
get_app_live_metrics | app-advisor-srv /api/v1/apps/{id}/live | app_id |
A same-named get_app_live_metrics also appears in the frontend AI chat's own tool list (app/api/ai/chat/route.ts) — that's a separate, unrelated definition specific to the browser-facing chat sidebar, and as of this writing it's a dead reference there (listed in the customer tool allowlist but never actually added to the chat's own TOOLS array). The version documented here, used by agent-runtime's App Advisor and SRE Orchestrator agent types, is fully implemented and calls a real endpoint.
Streaming Architecture
POST /api/v1/runs
→ creates AgentRun (status: pending)
→ starts goroutine: executeRun()
→ returns { run_id, status, stream_url }
executeRun() goroutine:
→ RunManager.NewSink(runID) → broadcast channel
→ StreamingRunner.Execute(ctx, run, sink)
→ BetaToolRunnerStreaming.AllStreaming(ctx)
for each turn:
for each event (text_delta, thinking_delta, tool_call, tool_result):
→ instrumentedTool.Execute() → emits SSE events to sink
→ sink → RunManager.Broadcast() → all subscribers
→ on complete: update AgentRun.status = completed
GET /api/v1/runs/{id}/stream
→ RunManager.Subscribe(runID) → receive channel
→ HTTP Flusher loop: write "data: {json}\n\n" for each SSE event
→ closes on "done" or "error" event
The RunManager fan-out pattern means multiple browser tabs or SSE clients can subscribe to the same run simultaneously without affecting agent execution.
Auto-Trigger
When analysis-agent publishes an insight with RiskScore > 70.0 to the analysis.insights exchange, agent-runtime automatically creates an SRE Orchestrator run with trigger: "auto" and a generated prompt:
Cluster {clusterID} has elevated risk score {score}/100.
Key anomalies: {N} detected. Summary: {summary}.
Investigate and recommend remediation.
REST API
| Method | Path | Description |
|---|---|---|
POST | /api/v1/runs | Create and start an agent run |
GET | /api/v1/runs | List runs (?agent_type=&status=&cluster_id=) |
GET | /api/v1/runs/{id} | Run detail |
GET | /api/v1/runs/{id}/stream | SSE stream of events |
GET | /api/v1/runs/{id}/tool-calls | Persisted tool call history |
System Prompts
SRE Orchestrator — ReAct framework. Starts with broad situational awareness, cross-correlates signals, outputs: Situation Summary / Key Findings / Root Cause / Recommended Actions / Monitoring Checkpoints.
Security Auditor — focuses on vulnerabilities, RBAC misconfigurations, network policy gaps, and compliance violations. Distinguishes immediate threats from technical debt.
Cost Optimizer — FinOps framework: right-sizing, waste elimination, scaling efficiency, multi-cluster arbitrage. Outputs quick wins (this week) and strategic optimizations.
Incident Responder — OODA loop (Observe / Orient / Decide / Act). Assesses blast radius, identifies root cause from correlated signals, provides kubectl commands plus runbook steps. Can execute runbooks and generate post-incident reports directly.
App Advisor — focuses on one application's own domain profile and health, using business-context-aware advice rather than generic cluster-wide recommendations. A separate, independent implementation reaches the same app-advisor-srv data from the frontend's customer-facing AI chat sidebar (get_app_profile/get_app_advice) — that's a direct tool-use loop in the frontend itself, not this agent type, though the two are functionally similar in intent.
NodeOps — node pool specialist. Reviews NodeClaim phases, scheduling scores, spot market savings, consolidation feasibility, and PDB constraints. Can drain nodes when safe.
Load Test Analyst — performance engineering expert. Fetches the latest kubeopera-ai optimisation status as a baseline, triggers load test analysis for p95/p99 assessment, cross-references with live pod/node metrics to identify saturation points, and produces a structured performance report with a numbered remediation action list ordered by priority.
Environment Variables
| Variable | Description |
|---|---|
ANTHROPIC_API_KEY | Required. Anthropic API key |
DATABASE_URL | PostgreSQL connection |
RABBITMQ_URL | Optional. Enables auto-trigger from analysis insights |
K8S_MONITOR_BASE_URL | k8s-monitor URL (default: http://localhost:8085) |
SECURITY_API_BASE_URL | security-api URL (default: http://localhost:8086) |
CICD_GATEWAY_BASE_URL | cicd-gateway URL (default: http://localhost:8087) |
KUBEOPERA_API_BASE_URL | kubeopera-api URL (default: http://localhost:8080) |
ANOMALY_DETECTOR_BASE_URL | anomaly-detector URL (default: http://localhost:8088) |
PREDICTIVE_SCALER_BASE_URL | predictive-scaler URL (default: http://localhost:8089) |
INCIDENT_MANAGER_BASE_URL | incident-manager URL (default: http://localhost:8090) |
NODES_MANAGER_BASE_URL | nodes-manager URL (code default: http://localhost:8098 — a local-dev value; the deployed service actually listens on 8115) |
OBSERVABILITY_AGENT_SRV_BASE_URL | observability-agent URL (default: http://localhost:8092) |
ANALYSIS_AGENT_SRV_BASE_URL | analysis-agent URL (default: http://localhost:8093) |
ACTION_AGENT_SRV_BASE_URL | action-agent URL (default: http://localhost:8094) |
FEEDBACK_AGENT_SRV_BASE_URL | feedback-agent URL (default: http://localhost:8095) |
RECOMMENDATION_AGENT_SRV_BASE_URL | recommendation-agent URL (default: http://localhost:8096) |
KUBEOPERA_AI_BASE_URL | kubeopera-ai URL (code default: http://localhost:8104 — a local-dev value; the deployed service actually listens on 8113) |
APP_ADVISOR_BASE_URL | app-advisor-srv URL, needed for the App Advisor agent type's tools (default: http://localhost:8105) |
PORT | HTTP port (deployed value: 8111; the code's own fallback default is a stale 8097) |