AI-Powered Investigation
This tutorial walks through using KubeOpera's agentic AI to investigate a cluster incident — from triggering an agent run to interpreting the live reasoning stream.
Prerequisites:
- KubeOpera platform running (all services)
agent-runtimeservice running (port 8111 as deployed) — with its Anthropic key resolvable, either from a tenant's own configured key or the platform's shared one, per Auth Service's credential resolution rather than a staticANTHROPIC_API_KEY- At least one cluster registered in kubeopera-api
Time: ~10 minutes
1. Open the Agent Runs Page
Navigate to AI Agents → Agent Runs in the sidebar.
You will see the runs list, which shows all previous and active agent runs with status badges:
- Gray clock — pending
- Blue spinner — running
- Green checkmark — completed
- Red X — failed
2. Start a New Agent Run
Click New Run to open the agent launcher modal.
Fill in the form:
Agent Type — select SRE Orchestrator. This agent uses Claude Sonnet 4.6 with interleaved thinking and has access to all 17 platform tools.
Cluster ID — enter the ID of the cluster to investigate (e.g. prod-us-east).
Prompt — describe what you want the agent to investigate. Be specific:
The payments-api deployment has been showing elevated CPU usage for the last 30 minutes.
Investigate the root cause, check for anomalies and incidents, and recommend immediate action.
Click Run to submit. The modal closes and the new run appears at the top of the list with status running.
3. Open the Run Detail
Click the run row to open the streaming detail view.
The page has two panels:
Left panel — Reasoning Trace
This shows everything the agent thinks and does in real time:
- Purple blocks — thinking (the agent's internal reasoning before acting)
- Blue blocks — tool calls (which tool the agent is calling and with what parameters)
- Green blocks — tool results (the data returned by each tool, with execution time)
- Red blocks — tool errors (if a backend service is unavailable)
Right panel — Text Output
The agent's final written response streams here as the agent completes its analysis. This is the human-readable summary and recommendation.
4. Watch the Investigation Unfold
A typical SRE investigation follows this pattern:
Step 1 — Thinking (initial planning)
I need to investigate elevated CPU usage on payments-api. I'll start by checking
overall cluster health, then look at pod-level metrics, recent anomalies, and
any open incidents before making recommendations.
Step 2 — Tool calls (data gathering)
The agent calls tools in sequence, using the result of each to decide what to call next:
Tool call: get_cluster_health
→ health_score: 72, 4 pods in failed state, API latency 340ms
Tool call: get_pod_metrics (namespace: production)
→ payments-api pods: CPU 92%, 96%, 88% (3 replicas all saturated)
Tool call: get_anomaly_events (cluster_id: prod-us-east, severity: high)
→ 2 critical anomalies: cpu_usage_pct z-score 3.8, crash_loop_count z-score 2.6
Tool call: get_incidents (cluster_id: prod-us-east, status: open)
→ 1 open incident: INC-4421 "High CPU on payments-api" (opened 28 min ago)
Tool call: get_scaling_forecasts (cluster_id: prod-us-east)
→ payments-api forecast: 97% CPU in 15 minutes, scaling decision pending
Tool call: get_recommendations
→ 1 recommendation: scale payments-api from 3 to 5 replicas immediately
Step 3 — Thinking (synthesis)
The data is consistent: all three payments-api replicas are saturated.
The anomaly detector flagged this 28 minutes ago and an incident is already open.
The predictive scaler has a pending recommendation to scale to 5 replicas.
The HPA maxReplicas is currently 4, which would be insufficient. I should
recommend both approving the scaling decision and raising the HPA limit.
Step 4 — Text output (recommendation)
## Cluster Health Investigation — prod-us-east
### Root Cause
All three `payments-api` replicas (production namespace) are CPU-saturated at
88–96%. The HPA has not triggered because maxReplicas is set to 4 and only 3
replicas are currently running — there is headroom for one more replica, but
this will not resolve the saturation given forecasted load.
### Current State
- Cluster health score: 72/100 (degraded)
- Open incident: INC-4421 (28 minutes old, not yet mitigated)
- Anomaly Z-score: 3.8 (critical threshold: 3.5)
- 4 pods in failed state (likely cascading from overload)
### Immediate Actions Required
1. **Scale payments-api to 5 replicas** — approve the pending scaling decision
in the Predictive Scaling page, or run:
kubectl scale deployment payments-api -n production --replicas=5
2. **Raise HPA maxReplicas to 8** — the current limit of 4 is insufficient for
this traffic pattern. Update the HPA to allow the cluster to auto-scale during
future spikes without manual intervention.
3. **Investigate the 4 failed pods** — run:
kubectl get pods -n production --field-selector=status.phase=Failed
kubectl logs <failed-pod> -n production --previous
### Forecast
CPU load is projected to reach 97% across all replicas within 15 minutes if no
action is taken. Scale immediately to prevent a complete service outage.
5. Approve the Scaling Recommendation (Optional)
After the agent completes, navigate to Analytics → Scaling Decisions to approve the scaling recommendation identified during the investigation.
Click Approve on the pending decision for payments-api. The predictive-scaler will patch the deployment replicas immediately.
6. Triggering Investigations via MCP
If you have the KubeOpera MCP Server configured in Claude Desktop or Claude Code, you can trigger the same investigation with a natural language prompt:
Run a full SRE investigation on cluster prod-us-east.
The payments-api has had high CPU for 30 minutes.
Claude will call run_sre_agent automatically, receive the run_id, and can then fetch the completed summary via get_recommendations or direct API.
7. Auto-Triggered Investigations
You do not need to manually start every investigation. The agent-runtime service listens to the analysis.insights RabbitMQ exchange and automatically starts an SRE Orchestrator run when the analysis agent publishes an insight with risk_score > 70.
To see auto-triggered runs, filter the Agent Runs list by Trigger: auto.
Auto-triggered runs use the following prompt template:
Cluster {cluster_id} has elevated risk score {score}.
Key anomalies: {anomaly_list}.
Investigate and recommend remediation.
Interpreting the SSE Stream Directly
For programmatic access, the agent reasoning stream is a standard Server-Sent Events stream. Each event has a type field:
| Event type | Meaning |
|---|---|
thinking | Internal reasoning text (Claude Sonnet only) |
tool_call | Agent is calling a tool — includes tool name and input |
tool_result | Tool returned successfully — includes output and duration_ms |
tool_error | Tool call failed — includes error message |
text | Delta of the agent's written response |
done | Agent run completed |
error | Agent run failed |
curl -N http://localhost:8111/api/v1/runs/{run_id}/stream
data: {"type":"thinking","run_id":"run-abc","payload":"I need to check cluster health first..."}
data: {"type":"tool_call","run_id":"run-abc","payload":{"tool":"get_cluster_health","input":{}}}
data: {"type":"tool_result","run_id":"run-abc","payload":{"tool":"get_cluster_health","output":"{...}","duration_ms":42}}
data: {"type":"text","run_id":"run-abc","payload":"## Cluster Health Investigation\n\n"}
data: {"type":"done","run_id":"run-abc","payload":null}
Next Steps
- Agent Runtime Reference — understand the four agent types and their tool access
- MCP Server Overview — use KubeOpera tools from Claude Desktop
- Analytics Page — review anomalies and approve scaling decisions from the UI