Recommendation Agent
Service: recommendation-agent-srv · Port: 8096
The recommendation agent turns what the analysis agent notices into advice a person can act on. For each significant insight, it asks Claude to explain the problem, its impact and the concrete steps to fix it, and stores the answer as structured recommendations you'll find on /agents/recommendations, in the AI Chat, and through the API.
How it works
When the analysis agent publishes an insight to analysis.insights (only for insights that pass its significance threshold):
- The agent reads the insight: the cluster or tenant, risk score, summary and anomalies.
- It resolves the AI credential for that tenant (see below).
- It builds a prompt describing the current state and asks Claude for recommendations as JSON.
- It validates the response against the recommendation schema.
- It stores each recommendation, linked to the insight that produced it.
If a step fails temporarily — for example, the AI provider is rate-limiting — the insight is retried with backoff and, after repeated failures, moved to a dead-letter queue for inspection. No insight is silently lost.
AI credentials
Each tenant's recommendations are generated with that tenant's AI credential. The agent resolves it for every call from auth-service (GET /internal/ai-credentials/resolve), using the tenant carried on the insight:
- If the tenant has configured their own AI provider key (Settings → AI Provider Key), it is used, and usage is attributed to them.
- Otherwise the platform key is used, within the tenant's quota.
When a tenant has reached its quota, new insights are queued until the quota resets, and the dashboard shows why recommendations are paused.
Recommendation schema
type Recommendation struct {
ID string
ClusterID string
TenantID string
Category string // security | cost | performance | reliability
Priority string // low | medium | high | critical
Title string
Description string
Impact string // what happens if it isn't addressed
Action string // concrete steps to resolve it
RiskScore float64 // from the triggering insight
InsightID string
GeneratedAt time.Time
}
The prompt
The agent asks Claude to respond with JSON only:
You are a Kubernetes SRE expert. Respond only with a valid JSON array
of recommendation objects. Each object must have:
category (security|cost|performance|reliability)
priority (low|medium|high|critical)
title
description
impact
action
No markdown, no extra text.
Example
[
{
"category": "performance",
"priority": "high",
"title": "Scale payments-api deployment",
"description": "CPU utilization has exceeded 85% for the last 4 cycles, with a Z-score of 3.1.",
"impact": "Request latency will keep increasing, with a risk of OOM kills under current load.",
"action": "Increase replicas from 3 to 5 now. Review the HPA's maxReplicas, currently 4."
}
]
Working with recommendations
On /agents/recommendations you can filter by category and priority, and for each recommendation:
- Apply — when it maps to an action KubeOpera can take (such as scaling), apply it directly; it goes through the normal approval workflow.
- Investigate — launch an agent run with the recommendation as context.
- Dismiss — hide it, with an optional reason that improves future recommendations.
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/recommendations | List recommendations (?cluster_id=&category=&priority=&limit=). |
GET | /api/v1/recommendations/{id} | Recommendation detail. |
POST | /api/v1/recommendations/{id}/dismiss | Dismiss a recommendation. |
GET | /healthz | Health check. |
Tenant users only see recommendations for their own tenant.
Configuration
| Variable | Default | Description |
|---|---|---|
AI_MODEL | claude-sonnet-4-5 | The Claude model used for recommendations. |
AUTH_SERVICE_BASE_URL | — | auth-service, for AI credential resolution. |
AI_CREDENTIAL_INTERNAL_API_KEY | — | Authenticates calls to auth-service's credential endpoint. |
MAX_RETRIES | 5 | Retries before an insight goes to the dead-letter queue. |
DATABASE_URL | — | PostgreSQL connection. |
RABBITMQ_URL | — | RabbitMQ connection. |
PORT | 8096 | HTTP port. |