Feedback Agent
The feedback agent evaluates automated remediation actions after the fact and publishes an adjustment signal intended for the analysis agent. This service's own half of the work is real and fully functional — it genuinely evaluates every action outcome and publishes real FeedbackSignal messages. The loop it's named for isn't actually closed today, though: analysis-agent-srv never starts a consumer for the exchange these signals go to, so nothing currently reads them. See Analysis Agent for the other half of this story.
How It Works
The feedback agent consumes ActionOutcome messages from the action.outcomes exchange. The evaluation here is simpler than the metric-comparison story the service's name might suggest: it doesn't wait for a follow-up telemetry snapshot or compare a metric's before/after value. Instead, for each outcome:
- Persists an
OutcomeRecord(for the trend history the REST endpoint below serves) - Looks up which metric this action type is associated with via a fixed map (see below) — if the action type isn't in that map, no signal is published for it at all
- Reads the action's own execution status: did the Kubernetes API call the action agent made actually succeed?
- Publishes a
FeedbackSignal—reinforcementif the action succeeded,correctionif it failed or errored
Metric Mapping
Each action type maps to the metric that its (never-currently-consumed) threshold adjustment would apply to:
| Action | Associated metric |
|---|---|
scale_deployment | cpu_usage_pct |
restart_pod | crash_loop_count |
cordon_node | memory_usage_pct |
notify | health_score |
Signal Types
Reinforcement — published when the action's own status was success. Analysis agent's threshold logic — real, tested, but not currently invoked — would decrease the Z-score threshold by 0.05 on this signal (making detection more sensitive).
Correction — published when the action's status was anything other than success (failed or errored). Analysis agent's threshold logic would increase the threshold by 0.10 on this signal (reducing false positives) — again, real code, just not wired to a consumer today.
Note what this does and doesn't measure: it's evaluating whether the action itself executed without error, not whether cluster health actually improved afterward. A scale_deployment call that succeeds at the Kubernetes API level counts as a reinforcement here even if CPU usage never actually came down — this service has no telemetry-comparison step to catch that difference.
FeedbackSignal Schema
type FeedbackSignal struct {
ClusterID string `json:"cluster_id"`
Metric string `json:"metric"`
Signal string `json:"signal"` // "reinforcement" | "correction"
Magnitude float64 `json:"magnitude"` // always 1.0 today
}
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/outcomes | Outcome record history (?cluster_id=&limit=) |
GET | /healthz | Health check |
Environment Variables
| Variable | Default | Description |
|---|---|---|
DATABASE_URL | — | PostgreSQL connection |
RABBITMQ_URL | — | RabbitMQ connection |
PORT | 8095 | HTTP port |