Service: agent-runtime · Port: 8111 · Database schema: agents
agent-runtime is KubeOpera's reasoning engine. It runs Claude-powered agents that investigate your clusters the way an experienced engineer would: form a hypothesis, gather evidence with live tools, refine, and conclude — then recommend (or, where allowed, take) action. Every step is streamed live and saved.
How an agent works
Agents follow the ReAct loop:
┌──────────┐ ┌──────────┐ ┌──────────┐
│ Reason │ → │ Act │ → │ Observe │ ─┐
│ (think) │ │(call tool)│ │ (result) │ │
└──────────┘ └──────────┘ └──────────┘ │
↑ │
└──────────── until it can answer ───────┘
- The agent receives a goal (your prompt) and context (the cluster, and for automatic runs the triggering insight).
- It thinks about what it needs to know.
- It calls a tool — for example,
get_cluster_health or get_incidents.
- It reads the result and decides the next step.
- It repeats until it has an answer, or until it reaches its iteration limit.
- It writes its findings: a situation summary, root cause, and recommended actions.
Agents with interleaved thinking reason between every tool call, which produces better tool choices on complex investigations.
Agent types
| Agent | Model | Tool scope | Max iterations | Interleaved thinking |
|---|
sre_orchestrator | Claude Sonnet 4.6 | All tools | 20 | Yes |
incident_responder | Claude Sonnet 4.6 | Incidents, diagnostics, runbooks, write actions | 15 | Yes |
node_ops | Claude Sonnet 4.6 | Node management, diagnostics, drain | 15 | Yes |
app_advisor | Claude Sonnet 4.6 | App Advisor, diagnostics | 12 | No |
load_test_analyst | Claude Sonnet 4.6 | Performance, diagnostics | 10 | No |
security_auditor | Claude Haiku 4.5 | Security | 10 | No |
cost_optimizer | Claude Haiku 4.5 | Cost, scaling decisions | 10 | No |
What each agent focuses on
- SRE Orchestrator — starts broad, correlates signals across every service, and reports: Situation summary · Key findings · Root cause · Recommended actions · Monitoring checkpoints.
- Incident Responder — works an incident with the OODA loop (observe, orient, decide, act): assesses blast radius, finds the root cause, runs runbooks, and writes the post-incident review.
- NodeOps — reviews node pools and claims, scheduling, spot savings and consolidation, respecting PodDisruptionBudgets; drains nodes when it's safe.
- App Advisor — looks at one application in its business context and gives specific, prioritized advice.
- Load Test Analyst — compares load-test results with the Optimizer's baseline and live metrics, finds saturation points, and produces a prioritized performance report.
- Security Auditor — separates immediate threats from technical debt across vulnerabilities, RBAC, network policy and compliance.
- Cost Optimizer — finds quick wins and strategic savings: right-sizing, waste, scaling efficiency and multi-cluster placement; reviews pending scaling decisions.
Agents only have the tools for their domain; the SRE Orchestrator has them all. Write tools change your cluster or your records, so they are only available to agents — and users — with the matching permission, and every call is audited.
Cluster and cost
| Tool | Source | Parameters |
|---|
get_cluster_health | k8s-monitor | — |
get_cluster_cost | k8s-monitor | — |
get_optimization_report | k8s-monitor | namespace?, view? |
get_pod_metrics | k8s-monitor | namespace? |
get_node_metrics | k8s-monitor | — |
get_cluster_list | kubeopera-api | — |
get_multi_cluster_overview | kubeopera-api, k8s-monitor | — |
Security and pipelines
| Tool | Source | Parameters |
|---|
get_security_posture | security-api | cluster_id? |
get_pipeline_status | cicd-gateway | cluster_id?, limit? |
Anomalies and scaling
| Tool | Source | Parameters |
|---|
get_anomaly_events | anomaly-detector | cluster_id?, severity?, limit? |
get_alert_rules | anomaly-detector | cluster_id? |
get_scaling_forecasts | predictive-scaler | cluster_id?, status? |
approve_scaling_decision ✎ | predictive-scaler | decision_id |
reject_scaling_decision ✎ | predictive-scaler | decision_id, reason? |
Incidents
| Tool | Source | Parameters |
|---|
get_incidents | incident-manager | cluster_id?, status? |
execute_runbook ✎ | incident-manager | incident_id, runbook_id |
get_runbook_execution | incident-manager | execution_id |
generate_post_incident_report ✎ | incident-manager | incident_id |
Nodes
| Tool | Source | Parameters |
|---|
get_node_pools | nodes-manager | — |
get_node_claims | nodes-manager | pool_name |
get_scheduling_decisions | nodes-manager | cluster_id?, limit? |
get_spot_market | nodes-manager | region? |
get_consolidation_plan | nodes-manager | — |
get_workload_placement | nodes-manager | — |
Actions
| Tool | Source | Parameters |
|---|
drain_node_action ✎ | action-agent | node_name |
rollback_deployment ✎ | action-agent | namespace, deployment_name, revision? |
approve_action ✎ | action-agent | action_id |
reject_action ✎ | action-agent | action_id, reason? |
Reactive pipeline
| Tool | Source | Parameters |
|---|
get_agent_telemetry | observability-agent | cluster_id?, limit? |
get_analysis_results | analysis-agent | cluster_id?, limit? |
get_action_log | action-agent | cluster_id?, limit? |
get_feedback_outcomes | feedback-agent | cluster_id?, limit? |
get_recommendations | recommendation-agent | cluster_id?, limit? |
Optimizer
| Tool | Source | Parameters |
|---|
get_optimization_status | kubeopera-ai | — |
get_load_test_analysis | kubeopera-ai | results_path? |
App Advisor
| Tool | Source | Parameters |
|---|
get_app_profile | app-advisor-srv | app_id |
get_app_advice | app-advisor-srv | app_id, status?, limit? |
get_app_live_metrics | app-advisor-srv | app_id |
✎ = write tool.
Running agents
Start a run
POST /api/v1/runs
Content-Type: application/json
{
"agent_type": "incident_responder",
"cluster_id": "prod-us-east",
"prompt": "Incident INC-142: payments-api 5xx rate above 2%. Find the cause and mitigate."
}
The response includes the run_id and a stream_url. The run executes in the background; you can close the browser and come back.
Watch it live
curl -N /api/v1/runs/{run_id}/stream
The stream delivers thinking, tool_call, tool_result, text, done and error events (see AI Agents UI for the format). Any number of clients can watch the same run at once without affecting it, and clients that reconnect receive the events they missed.
Automatic runs
When the analysis agent publishes an insight with a risk score above the auto-trigger threshold (70 by default), agent-runtime starts an SRE Orchestrator run with trigger: "auto":
Cluster {clusterID} has elevated risk score {score}/100.
Key anomalies: {N} detected. Summary: {summary}.
Investigate and recommend remediation.
To avoid duplicate investigations, only one automatic run is started per cluster while a previous one is still running.
How streaming works internally
POST /api/v1/runs
→ create AgentRun (status: pending) → start executeRun() → return run_id
executeRun():
→ RunManager.NewSink(runID)
→ StreamingRunner.Execute(ctx, run, sink)
for each turn and each event (thinking, text, tool_call, tool_result):
→ instrumentedTool records the call in agents.tool_calls
→ sink → RunManager.Broadcast() → every subscriber
→ AgentRun.status = completed | failed
GET /api/v1/runs/{id}/stream
→ replay stored events, then subscribe to live ones
→ write "data: {json}\n\n" until "done" or "error"
AI credentials
Each run uses the AI credential of the tenant it runs for — the tenant's own provider key if they've configured one, otherwise the platform key within the tenant's quota — resolved from auth-service for every run. Usage is attributed to the tenant.
REST API
| Method | Path | Description |
|---|
POST | /api/v1/runs | Create and start a run. |
GET | /api/v1/runs | List runs (?agent_type=&status=&cluster_id=). |
GET | /api/v1/runs/{id} | Run detail and findings. |
GET | /api/v1/runs/{id}/stream | Server-Sent Events stream. |
GET | /api/v1/runs/{id}/tool-calls | Stored tool-call history. |
POST | /api/v1/runs/{id}/cancel | Stop a running agent. |
GET | /healthz | Health check. |
Configuration
| Variable | Default | Description |
|---|
PORT | 8111 | HTTP port. |
DATABASE_URL | — | PostgreSQL connection. |
RABBITMQ_URL | — | Enables automatic runs from analysis insights. |
AUTO_TRIGGER_RISK_SCORE | 70 | Risk score above which an SRE investigation starts automatically. |
AUTH_SERVICE_BASE_URL | — | auth-service, for AI credential resolution. |
AI_CREDENTIAL_INTERNAL_API_KEY | — | Authenticates credential resolution calls. |
K8S_MONITOR_BASE_URL | http://k8s-monitor:8085 | |
SECURITY_API_BASE_URL | http://security-api:8086 | |
CICD_GATEWAY_BASE_URL | http://cicd-gateway:8087 | |
KUBEOPERA_API_BASE_URL | http://kubeopera-api:8090 | |
ANOMALY_DETECTOR_BASE_URL | http://anomaly-detector:8088 | |
PREDICTIVE_SCALER_BASE_URL | http://predictive-scaler:8089 | |
INCIDENT_MANAGER_BASE_URL | http://incident-manager:8090 | |
NODES_MANAGER_BASE_URL | http://nodes-manager:8115 | |
OBSERVABILITY_AGENT_SRV_BASE_URL | http://observability-agent-srv:8092 | |
ANALYSIS_AGENT_SRV_BASE_URL | http://analysis-agent-srv:8093 | |
ACTION_AGENT_SRV_BASE_URL | http://action-agent-srv:8094 | |
FEEDBACK_AGENT_SRV_BASE_URL | http://feedback-agent-srv:8095 | |
RECOMMENDATION_AGENT_SRV_BASE_URL | http://recommendation-agent-srv:8096 | |
KUBEOPERA_AI_BASE_URL | http://kubeopera-ai:8113 | |
APP_ADVISOR_BASE_URL | http://app-advisor-srv:8105 | |