Skip to main content
Version: 2.0

Your First AI Investigation

In this tutorial you'll ask KubeOpera's SRE Orchestrator agent to investigate a performance problem, watch it reason in real time, and act on what it finds.

Time: about 10 minutes · You'll need: a KubeOpera account with a registered cluster and at least one running app. (Tenant users can follow along with the App Advisor agent.)

The scenario​

Your payments-api has been running hot for half an hour. Users say checkout is slow. You want to know why, and what to do — fast.

Step 1 — Open Agent Runs​

In the sidebar, open AI Agents → Agent Runs. Each run in the list shows a status:

IconStatus
Gray clockPending
Blue spinnerRunning
Green checkCompleted
Red crossFailed

Step 2 — Start a run​

Select New Run and fill in:

  • Agent type: SRE Orchestrator — the generalist, with access to every tool and extended thinking.

  • Cluster: the cluster your app runs in, for example prod-us-east.

  • Prompt: be specific about the symptom and the time frame:

    The payments-api deployment has had elevated CPU for the last 30 minutes
    and checkout is slow. Find the root cause, check anomalies and incidents,
    and recommend immediate action.

Select Start Run. The run opens straight away.

Writing good prompts

Say what you're seeing, where (namespace, app), and since when. "Checkout slow since 14:00 in production" beats "something is wrong".

Step 3 — Watch the agent think​

The run view has two panels.

Left — the reasoning trace:

  • Purple — thinking. The agent planning its next move.
  • Blue — tool calls. What it's asking for, and with which parameters.
  • Green — tool results. What came back, and how long it took.
  • Red — tool errors. A call that failed, and how the agent works around it.

Right — the output. The agent's findings, written as it goes.

A typical investigation unfolds like this:

1. Plan.

I need to investigate elevated CPU on payments-api. I'll check overall cluster
health, then pod-level metrics, recent anomalies and open incidents, then look
at forecasts before recommending anything.

2. Gather evidence. Each result shapes the next call:

get_cluster_health
→ health_score 72, 4 pods failed, API latency 340ms
get_pod_metrics (namespace: production)
→ payments-api pods at 92%, 96%, 88% CPU — all three saturated
get_anomaly_events (cluster_id: prod-us-east, severity: high)
→ cpu_usage_pct z-score 3.8 (critical); crash_loop_count z-score 2.6
get_incidents (status: open)
→ INC-4421 "High CPU on payments-api", opened 28 minutes ago
get_scaling_forecasts
→ 97% CPU forecast in 15 minutes; scaling decision to 5 replicas pending
get_logs (app: payments-api, since: 30m)
→ retry storm against the fraud-check service starting 14:02

3. Synthesize.

All three replicas are saturated, and the logs show a retry storm against
fraud-check starting just before the CPU climb. Scaling will relieve pressure
now; the retry storm is the underlying cause. The HPA's maxReplicas of 4 is
too low for the forecast.

4. Report.

## payments-api investigation — prod-us-east

### Root cause
A retry storm against fraud-check (from 14:02) is multiplying request volume;
all three payments-api replicas are CPU-saturated (88–96%).

### Immediate actions
1. Approve the pending scaling decision: payments-api 3 → 5 replicas.
2. Raise the HPA's maxReplicas from 4 to 8.
3. Add backoff to payments-api's fraud-check client; check fraud-check health.

### Watch
CPU is forecast to reach 97% within 15 minutes without action.

Step 4 — Act on the findings​

From the run's output, select Approve scaling decision — or open Analytics → Scaling Decisions and approve it there. KubeOpera scales payments-api to 5 replicas.

If you've given the agent permission to act, you can instead reply in the run: "Go ahead and approve the scaling decision." The agent calls approve_scaling_decision, and the action is recorded in the run's history.

Then watch the Dashboard: CPU drops, and the anomaly clears on the next analysis cycle.

Open Incidents → INC-4421. The agent run is linked on the incident's timeline, and once you resolve it you can Generate review to draft a post-incident review that draws on the investigation.

Automatic investigations​

You don't have to start every investigation yourself. When the analysis agent's risk score for a cluster passes 70, KubeOpera starts an SRE Orchestrator run automatically. Filter Agent Runs by Trigger: auto to see them — often, the investigation is finished by the time you open the page.

Stream a run from the API​

Every run is a standard Server-Sent Events stream you can consume from your own tools:

curl -N -H "Authorization: Bearer $KUBEOPERA_TOKEN" \
https://kubeopera.example.com/api/agents/runtime/api/v1/runs/{run_id}/stream
data: {"type":"thinking","run_id":"run-abc","payload":"I need to check cluster health first..."}
data: {"type":"tool_call","run_id":"run-abc","payload":{"tool":"get_cluster_health","input":{}}}
data: {"type":"tool_result","run_id":"run-abc","payload":{"tool":"get_cluster_health","duration_ms":42}}
data: {"type":"text","run_id":"run-abc","payload":"## payments-api investigation\n\n"}
data: {"type":"done","run_id":"run-abc","payload":null}
EventMeaning
thinkingThe agent's reasoning between steps.
tool_callA tool call, with its input.
tool_resultA tool's result, with its duration.
tool_errorA failed tool call.
textPart of the agent's written output.
done / errorThe run finished or failed.

What you learned​

  • How to launch an agent with a focused prompt.
  • How to read the reasoning trace: plan → evidence → synthesis → report.
  • How to act on findings — yourself or through the agent — with everything audited.
  • That high-risk situations trigger investigations automatically.

Next steps​