Feedback Agent
Service: feedback-agent-srv · Port: 8095
The feedback agent closes the reactive pipeline's loop. After every automated action it asks one question — did it actually help? — by measuring the cluster before and after. It sends the answer to the analysis agent as a feedback signal, which tunes detection for that cluster and metric.
Without it, an automated system can repeat the same unhelpful action forever. With it, KubeOpera learns which actions work on which clusters.
How it works
For each ActionOutcome consumed from action.outcomes, the agent:
- Records the outcome, for the trend history on
/agents/feedback. - Finds the metric the action is meant to affect (see metric mapping). Actions without a mapped metric are recorded but produce no signal.
- Takes the baseline — the metric's value in the snapshot just before the action.
- Waits for the next snapshot after the action has settled, and reads the metric again.
- Compares the two. If the action succeeded and the metric moved in the right direction, the result is a reinforcement; if the action failed, or the metric didn't improve, it's a correction.
- Publishes a
FeedbackSignaltofeedback.signals, with amagnitudereflecting how much the metric changed.
Metric mapping
Each action is judged against the metric it is meant to improve:
| Action | Metric | "Improved" means |
|---|---|---|
scale_deployment | cpu_usage_pct | CPU utilization went down. |
restart_pod | crash_loop_count | Fewer pods are crash-looping. |
cordon_node | memory_usage_pct | Memory pressure went down. |
notify | health_score | The health score went up. |
Signal types
Reinforcement — the action succeeded and its metric improved against the baseline. The analysis agent lowers the threshold for that metric slightly (−0.05 × magnitude), so it detects similar problems sooner.
Correction — the action failed, or it succeeded but the metric didn't improve. The analysis agent raises the threshold (+0.10 × magnitude), so it's more cautious about repeating that action.
Because the judgement is based on measured telemetry, not just whether the Kubernetes API call succeeded, a scale-up that executes cleanly but never brings CPU down is treated as a correction.
FeedbackSignal schema
type FeedbackSignal struct {
ClusterID string `json:"cluster_id"`
Metric string `json:"metric"`
Signal string `json:"signal"` // "reinforcement" | "correction"
Magnitude float64 `json:"magnitude"` // 0–1: how strongly the metric responded
ActionID string `json:"action_id"` // the action being judged
}
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/outcomes | Outcome history with before/after values and the resulting signal (?cluster_id=&limit=). |
GET | /api/v1/signals | Published feedback signals (?cluster_id=&metric=). |
GET | /healthz | Health check. |
Configuration
| Variable | Default | Description |
|---|---|---|
SETTLE_SNAPSHOTS | 1 | How many snapshots to wait after an action before measuring. |
IMPROVEMENT_MIN_PCT | 5 | The minimum relative change that counts as an improvement. |
DATABASE_URL | — | PostgreSQL connection. |
RABBITMQ_URL | — | RabbitMQ connection. |
PORT | 8095 | HTTP port. |