Skip to main content
Version: 1.0

Feedback Agent

The feedback agent evaluates automated remediation actions after the fact and publishes an adjustment signal intended for the analysis agent. This service's own half of the work is real and fully functional — it genuinely evaluates every action outcome and publishes real FeedbackSignal messages. The loop it's named for isn't actually closed today, though: analysis-agent-srv never starts a consumer for the exchange these signals go to, so nothing currently reads them. See Analysis Agent for the other half of this story.

How It Works​

The feedback agent consumes ActionOutcome messages from the action.outcomes exchange. The evaluation here is simpler than the metric-comparison story the service's name might suggest: it doesn't wait for a follow-up telemetry snapshot or compare a metric's before/after value. Instead, for each outcome:

  1. Persists an OutcomeRecord (for the trend history the REST endpoint below serves)
  2. Looks up which metric this action type is associated with via a fixed map (see below) — if the action type isn't in that map, no signal is published for it at all
  3. Reads the action's own execution status: did the Kubernetes API call the action agent made actually succeed?
  4. Publishes a FeedbackSignal — reinforcement if the action succeeded, correction if it failed or errored

Metric Mapping​

Each action type maps to the metric that its (never-currently-consumed) threshold adjustment would apply to:

ActionAssociated metric
scale_deploymentcpu_usage_pct
restart_podcrash_loop_count
cordon_nodememory_usage_pct
notifyhealth_score

Signal Types​

Reinforcement — published when the action's own status was success. Analysis agent's threshold logic — real, tested, but not currently invoked — would decrease the Z-score threshold by 0.05 on this signal (making detection more sensitive).

Correction — published when the action's status was anything other than success (failed or errored). Analysis agent's threshold logic would increase the threshold by 0.10 on this signal (reducing false positives) — again, real code, just not wired to a consumer today.

Note what this does and doesn't measure: it's evaluating whether the action itself executed without error, not whether cluster health actually improved afterward. A scale_deployment call that succeeds at the Kubernetes API level counts as a reinforcement here even if CPU usage never actually came down — this service has no telemetry-comparison step to catch that difference.

FeedbackSignal Schema​

type FeedbackSignal struct {
ClusterID string `json:"cluster_id"`
Metric string `json:"metric"`
Signal string `json:"signal"` // "reinforcement" | "correction"
Magnitude float64 `json:"magnitude"` // always 1.0 today
}

REST API​

MethodPathDescription
GET/api/v1/outcomesOutcome record history (?cluster_id=&limit=)
GET/healthzHealth check

Environment Variables​

VariableDefaultDescription
DATABASE_URL—PostgreSQL connection
RABBITMQ_URL—RabbitMQ connection
PORT8095HTTP port