Skip to main content
Version: 2.0

AI Agents UI

The AI Agents section is where you see what KubeOpera's AI is doing — the always-on reactive pipeline, and the reasoning agents you (or the platform) launch to investigate problems. To learn how the agents themselves work, see AI Agents.

Agent overview (/agents)​

A status card for each service in the reactive pipeline, updated continuously from its health endpoint:

AgentWhat it does
Observability agentCollects telemetry snapshots.
Analysis agentDetects anomalies and makes decisions.
Action agentCarries out approved remediations.
Feedback agentMeasures whether actions helped.
Recommendation agentTurns insights into recommendations.

Below the cards, a flow diagram shows how events move between them — the same closed loop described in Reactive AI Pipeline.

Each agent's page​

Every pipeline agent has its own page with its latest output:

PageShows
/agents/observabilityRecent telemetry snapshots per cluster.
/agents/analysisAnalysis results: anomalies, risk scores and decisions.
/agents/actionsActions taken, their reasons and outcomes.
/agents/feedbackFeedback signals and how thresholds have adapted per cluster.
/agents/recommendationsRecommendations, by priority and category.

Agent runs (/agents/runs)​

An agent run is one session of a reasoning agent working on a goal. The runs list shows each run's agent type, cluster, prompt, status (with a spinner while running), duration, trigger (manual or automatic) and start time.

Start a run​

  1. Select New Run.

  2. Choose an agent type:

    AgentUse it for
    SRE OrchestratorOpen-ended investigation across everything.
    Security AuditorSecurity posture, findings and RBAC review.
    Cost OptimizerSpend, waste and right-sizing.
    Incident ResponderTriage and resolve an active incident.
    App AdvisorReview one application against best practices.
    Load Test AnalystInterpret performance and load-test results.
    NodeOpsNode pools, capacity and consolidation.
  3. Choose the cluster.

  4. Describe what you want in the prompt — be specific about the symptom and time frame.

  5. Select Start Run. The run opens immediately so you can watch it work.

The same thing through the API:

POST /api/agents/runtime/api/v1/runs
Content-Type: application/json

{
"agent_type": "sre_orchestrator",
"cluster_id": "prod-us-east",
"prompt": "Investigate why memory usage has been climbing for the last 2 hours"
}

Runs also start automatically: when the analysis agent's risk score for a cluster exceeds 70, an SRE Orchestrator run is launched with the analysis as context.

Watching a run (/agents/runs/{id})​

The run view streams the agent's work live.

Left — reasoning trace:

  • Thinking (purple, collapsible) — the agent's reasoning between steps.
  • Tool calls (blue) — each tool the agent calls, with its input.
  • Tool results (green) — the result and how long it took.
  • Tool errors (red) — calls that failed, and how the agent recovered.

Right — output: the agent's findings and recommendations, streamed as they're written.

If your connection drops, the view reconnects automatically and catches up — the run itself continues on the server regardless, and its full history is saved.

How streaming works​

The page subscribes to a Server-Sent Events stream:

GET /api/agents/runtime/api/v1/runs/{id}/stream
Content-Type: text/event-stream

Each event is a JSON object:

{ "type": "thinking", "run_id": "abc-123", "payload": "Let me start by checking..." }
{ "type": "tool_call", "run_id": "abc-123", "payload": { "tool": "get_cluster_health", "input": "{}" } }
{ "type": "tool_result", "run_id": "abc-123", "payload": { "tool": "get_cluster_health", "duration_ms": 143 } }
{ "type": "text", "run_id": "abc-123", "payload": "The cluster health score is 72/100..." }
{ "type": "done", "run_id": "abc-123", "payload": null }

The stream ends with done (or error). Use the same endpoint to build your own integrations — for example, posting run results to a chat channel.

Next steps​