Skip to main content
Version: 1.0

Recommendation Agent

The recommendation agent consumes analysis insights and calls the Anthropic API (Claude Sonnet 4.5) to generate structured, actionable recommendations for improving cluster health, security, cost, and reliability.

How It Works​

When the analysis agent publishes an InsightPublishMessage to the analysis.insights exchange (only for insights that clear its own significance gate — see Analysis Agent):

  1. recommendation-agent-srv deserialises the message
  2. Resolves which Anthropic API key to use for this call (see below) — if none is available or the platform's usage quota is exceeded, the message is dropped rather than retried
  3. Constructs a prompt describing the cluster state, risk score, and detected anomalies
  4. Calls client.Messages.New() with Claude Sonnet 4.5 and a strict JSON-only system prompt
  5. Parses the response into a []Recommendation struct — Claude occasionally wraps its JSON array in a ```json fence despite being told not to, so the parser trims both leading and trailing junk around the outermost [/] before unmarshalling
  6. Persists each recommendation to PostgreSQL

If message handling fails at any of these steps, the RabbitMQ delivery is dropped (not requeued) — a lost insight just means no recommendation for that one anomaly cycle, not a lost delivery worth retrying indefinitely.

AI credential resolution​

This service has no real per-tenant identity to resolve a key against today — InsightMessage only carries a ClusterID string, which is sometimes a tenant vCluster namespace and sometimes the platform-wide "default", never a UUID tenant ID. So every call resolves against the platform/host Anthropic key via auth-service's GET /internal/ai-credentials/resolve — this is not a loss of isolation this service never had; it's the same platform-wide key it always used, just resolved dynamically per call instead of read once from a static ANTHROPIC_API_KEY env var at startup. Real per-tenant scoping across the observability → analysis → recommendation pipeline is a larger, separate effort spanning several services' domain models, not something this service alone can add.

Recommendation Schema​

type Recommendation struct {
ID string
ClusterID string
Category string // security | cost | performance | reliability
Priority string // low | medium | high | critical
Title string
Description string
Impact string // what happens if not addressed
Action string // concrete steps to resolve
RiskScore float64 // inherited from the triggering insight
GeneratedAt time.Time
}

System Prompt​

You are a Kubernetes SRE expert. Respond only with a valid JSON array
of recommendation objects. Each object must have:
category (security|cost|performance|reliability)
priority (low|medium|high|critical)
title
description
impact
action
No markdown, no extra text.

Example Output​

[
{
"category": "performance",
"priority": "high",
"title": "Scale payments-api deployment",
"description": "CPU utilisation has exceeded 85% for the last 4 cycles, with a Z-score of 3.1.",
"impact": "Request latency will continue to increase. Risk of pod OOMKill under current load.",
"action": "Increase deployment replicas from 3 to 5 immediately. Review HPA maxReplicas limit — currently set to 4."
}
]

REST API​

MethodPathDescription
GET/api/v1/recommendationsList recommendations (?cluster_id=&limit=)
GET/api/v1/recommendations/{id}Recommendation detail
GET/healthzHealth check

Environment Variables​

VariableDefaultDescription
DATABASE_URL—PostgreSQL connection
RABBITMQ_URL—RabbitMQ connection
AUTH_SERVICE_BASE_URL—Required for AI credential resolution
AI_CREDENTIAL_INTERNAL_API_KEY—Required. Authenticates this service's calls to auth-service's internal credential-resolve endpoint
PORT8096HTTP port
tip

There's no static ANTHROPIC_API_KEY for this service anymore — the actual Anthropic key is resolved per-call from auth-service (see above). If no platform host key is configured there, or the platform's usage quota is exceeded, the service logs a warning and drops that insight without generating a recommendation. Other services in the pipeline are not affected.