Skip to main content
Version: 2.0

Feedback Agent

Service: feedback-agent-srv · Port: 8095

The feedback agent closes the reactive pipeline's loop. After every automated action it asks one question — did it actually help? — by measuring the cluster before and after. It sends the answer to the analysis agent as a feedback signal, which tunes detection for that cluster and metric.

Without it, an automated system can repeat the same unhelpful action forever. With it, KubeOpera learns which actions work on which clusters.

How it works​

For each ActionOutcome consumed from action.outcomes, the agent:

  1. Records the outcome, for the trend history on /agents/feedback.
  2. Finds the metric the action is meant to affect (see metric mapping). Actions without a mapped metric are recorded but produce no signal.
  3. Takes the baseline — the metric's value in the snapshot just before the action.
  4. Waits for the next snapshot after the action has settled, and reads the metric again.
  5. Compares the two. If the action succeeded and the metric moved in the right direction, the result is a reinforcement; if the action failed, or the metric didn't improve, it's a correction.
  6. Publishes a FeedbackSignal to feedback.signals, with a magnitude reflecting how much the metric changed.

Metric mapping​

Each action is judged against the metric it is meant to improve:

ActionMetric"Improved" means
scale_deploymentcpu_usage_pctCPU utilization went down.
restart_podcrash_loop_countFewer pods are crash-looping.
cordon_nodememory_usage_pctMemory pressure went down.
notifyhealth_scoreThe health score went up.

Signal types​

Reinforcement — the action succeeded and its metric improved against the baseline. The analysis agent lowers the threshold for that metric slightly (−0.05 × magnitude), so it detects similar problems sooner.

Correction — the action failed, or it succeeded but the metric didn't improve. The analysis agent raises the threshold (+0.10 × magnitude), so it's more cautious about repeating that action.

Because the judgement is based on measured telemetry, not just whether the Kubernetes API call succeeded, a scale-up that executes cleanly but never brings CPU down is treated as a correction.

FeedbackSignal schema​

type FeedbackSignal struct {
ClusterID string `json:"cluster_id"`
Metric string `json:"metric"`
Signal string `json:"signal"` // "reinforcement" | "correction"
Magnitude float64 `json:"magnitude"` // 0–1: how strongly the metric responded
ActionID string `json:"action_id"` // the action being judged
}

REST API​

MethodPathDescription
GET/api/v1/outcomesOutcome history with before/after values and the resulting signal (?cluster_id=&limit=).
GET/api/v1/signalsPublished feedback signals (?cluster_id=&metric=).
GET/healthzHealth check.

Configuration​

VariableDefaultDescription
SETTLE_SNAPSHOTS1How many snapshots to wait after an action before measuring.
IMPROVEMENT_MIN_PCT5The minimum relative change that counts as an improvement.
DATABASE_URL—PostgreSQL connection.
RABBITMQ_URL—RabbitMQ connection.
PORT8095HTTP port.