Your First AI Investigation
In this tutorial you'll ask KubeOpera's SRE Orchestrator agent to investigate a performance problem, watch it reason in real time, and act on what it finds.
Time: about 10 minutes · You'll need: a KubeOpera account with a registered cluster and at least one running app. (Tenant users can follow along with the App Advisor agent.)
The scenario
Your payments-api has been running hot for half an hour. Users say checkout is slow. You want to know why, and what to do — fast.
Step 1 — Open Agent Runs
In the sidebar, open AI Agents → Agent Runs. Each run in the list shows a status:
| Icon | Status |
|---|---|
| Gray clock | Pending |
| Blue spinner | Running |
| Green check | Completed |
| Red cross | Failed |
Step 2 — Start a run
Select New Run and fill in:
-
Agent type:
SRE Orchestrator— the generalist, with access to every tool and extended thinking. -
Cluster: the cluster your app runs in, for example
prod-us-east. -
Prompt: be specific about the symptom and the time frame:
The payments-api deployment has had elevated CPU for the last 30 minutes
and checkout is slow. Find the root cause, check anomalies and incidents,
and recommend immediate action.
Select Start Run. The run opens straight away.
Say what you're seeing, where (namespace, app), and since when. "Checkout slow since 14:00 in production" beats "something is wrong".
Step 3 — Watch the agent think
The run view has two panels.
Left — the reasoning trace:
- Purple — thinking. The agent planning its next move.
- Blue — tool calls. What it's asking for, and with which parameters.
- Green — tool results. What came back, and how long it took.
- Red — tool errors. A call that failed, and how the agent works around it.
Right — the output. The agent's findings, written as it goes.
A typical investigation unfolds like this:
1. Plan.
I need to investigate elevated CPU on payments-api. I'll check overall cluster
health, then pod-level metrics, recent anomalies and open incidents, then look
at forecasts before recommending anything.
2. Gather evidence. Each result shapes the next call:
get_cluster_health
→ health_score 72, 4 pods failed, API latency 340ms
get_pod_metrics (namespace: production)
→ payments-api pods at 92%, 96%, 88% CPU — all three saturated
get_anomaly_events (cluster_id: prod-us-east, severity: high)
→ cpu_usage_pct z-score 3.8 (critical); crash_loop_count z-score 2.6
get_incidents (status: open)
→ INC-4421 "High CPU on payments-api", opened 28 minutes ago
get_scaling_forecasts
→ 97% CPU forecast in 15 minutes; scaling decision to 5 replicas pending
get_logs (app: payments-api, since: 30m)
→ retry storm against the fraud-check service starting 14:02
3. Synthesize.
All three replicas are saturated, and the logs show a retry storm against
fraud-check starting just before the CPU climb. Scaling will relieve pressure
now; the retry storm is the underlying cause. The HPA's maxReplicas of 4 is
too low for the forecast.
4. Report.
## payments-api investigation — prod-us-east
### Root cause
A retry storm against fraud-check (from 14:02) is multiplying request volume;
all three payments-api replicas are CPU-saturated (88–96%).
### Immediate actions
1. Approve the pending scaling decision: payments-api 3 → 5 replicas.
2. Raise the HPA's maxReplicas from 4 to 8.
3. Add backoff to payments-api's fraud-check client; check fraud-check health.
### Watch
CPU is forecast to reach 97% within 15 minutes without action.
Step 4 — Act on the findings
From the run's output, select Approve scaling decision — or open Analytics → Scaling Decisions and approve it there. KubeOpera scales payments-api to 5 replicas.
If you've given the agent permission to act, you can instead reply in the run: "Go ahead and approve the scaling decision." The agent calls approve_scaling_decision, and the action is recorded in the run's history.
Then watch the Dashboard: CPU drops, and the anomaly clears on the next analysis cycle.
Step 5 — Link it to the incident
Open Incidents → INC-4421. The agent run is linked on the incident's timeline, and once you resolve it you can Generate review to draft a post-incident review that draws on the investigation.
Automatic investigations
You don't have to start every investigation yourself. When the analysis agent's risk score for a cluster passes 70, KubeOpera starts an SRE Orchestrator run automatically. Filter Agent Runs by Trigger: auto to see them — often, the investigation is finished by the time you open the page.
Stream a run from the API
Every run is a standard Server-Sent Events stream you can consume from your own tools:
curl -N -H "Authorization: Bearer $KUBEOPERA_TOKEN" \
https://kubeopera.example.com/api/agents/runtime/api/v1/runs/{run_id}/stream
data: {"type":"thinking","run_id":"run-abc","payload":"I need to check cluster health first..."}
data: {"type":"tool_call","run_id":"run-abc","payload":{"tool":"get_cluster_health","input":{}}}
data: {"type":"tool_result","run_id":"run-abc","payload":{"tool":"get_cluster_health","duration_ms":42}}
data: {"type":"text","run_id":"run-abc","payload":"## payments-api investigation\n\n"}
data: {"type":"done","run_id":"run-abc","payload":null}
| Event | Meaning |
|---|---|
thinking | The agent's reasoning between steps. |
tool_call | A tool call, with its input. |
tool_result | A tool's result, with its duration. |
tool_error | A failed tool call. |
text | Part of the agent's written output. |
done / error | The run finished or failed. |
What you learned
- How to launch an agent with a focused prompt.
- How to read the reasoning trace: plan → evidence → synthesis → report.
- How to act on findings — yourself or through the agent — with everything audited.
- That high-risk situations trigger investigations automatically.
Next steps
- Agent runtime — every agent type and tool.
- Reactive AI Pipeline — what happens before an investigation starts.
- Incidents — manage the incident end to end.