Skip to main content
Version: 1.0

Tool Reference

Complete reference for all 49 tools exposed by the KubeOpera MCP Server.

Read-Only Data Tools​

These tools query live data from KubeOpera services. All are marked ReadOnlyHint: true and OpenWorld: true.


get_cluster_health​

Fetches the current cluster health score, node status summary, pod counts, and API server latency.

Source service: k8s-monitor (port 8085) Endpoint: GET /api/health

Parameters: None

Example response:

{
"health_score": 87,
"status": "healthy",
"nodes": {
"total": 6,
"ready": 6,
"not_ready": 0
},
"pods": {
"total": 142,
"running": 138,
"pending": 2,
"failed": 2,
"succeeded": 0
},
"api_server_latency_ms": 12,
"checked_at": "2026-04-18T10:00:00Z"
}

get_cluster_cost​

Fetches current cloud cost breakdown by namespace, workload, and resource type.

Source service: k8s-monitor (port 8085) Endpoint: GET /api/cost

Parameters: None

Example response:

{
"total_monthly_usd": 4820.50,
"by_namespace": [
{ "namespace": "production", "monthly_usd": 3200.00, "cpu_pct": 66, "memory_pct": 70 },
{ "namespace": "staging", "monthly_usd": 980.00, "cpu_pct": 20, "memory_pct": 18 }
],
"by_resource": {
"compute_usd": 3100.00,
"memory_usd": 1200.00,
"storage_usd": 520.50
}
}

get_optimization_report​

Returns over-provisioned and idle workload recommendations with estimated monthly savings.

Source service: k8s-monitor (port 8085) Endpoint: GET /api/optimizer

Parameters:

NameTypeRequiredDescription
namespacestringNoFilter recommendations to a specific namespace
viewstringNosummary (default) or detailed

Example response:

{
"total_savings_usd": 640.00,
"recommendations": [
{
"workload": "payments-api",
"namespace": "production",
"type": "over_provisioned",
"current_cpu_request": "2000m",
"recommended_cpu_request": "500m",
"savings_usd": 220.00
}
]
}

get_pod_metrics​

Returns per-pod CPU and memory usage for all pods or a specific namespace.

Source service: k8s-monitor (port 8085) Endpoint: GET /api/metrics/pods

Parameters:

NameTypeRequiredDescription
namespacestringNoFilter to a specific namespace

Example response:

{
"pods": [
{
"name": "payments-api-6d8f9b-xkj2p",
"namespace": "production",
"cpu_usage_millicores": 420,
"memory_usage_mb": 312,
"cpu_limit_millicores": 2000,
"memory_limit_mb": 512
}
],
"collected_at": "2026-04-18T10:00:00Z"
}

get_node_metrics​

Returns per-node CPU, memory, and disk metrics for all cluster nodes.

Source service: k8s-monitor (port 8085) Endpoint: GET /api/metrics/nodes

Parameters: None

Example response:

{
"nodes": [
{
"name": "ip-10-0-1-45.ec2.internal",
"cpu_usage_pct": 62,
"memory_usage_pct": 74,
"disk_usage_pct": 38,
"ready": true,
"roles": ["worker"]
}
]
}

get_security_posture​

Returns compliance score, vulnerability counts, and RBAC findings.

Source service: security-api (port 8086) Endpoint: GET /api/v1/posture/summary

Parameters:

NameTypeRequiredDescription
cluster_idstringNoFilter to a specific cluster

Example response:

{
"compliance_score": 78,
"vulnerabilities": {
"critical": 2,
"high": 8,
"medium": 24,
"low": 51
},
"rbac_findings": 5,
"privileged_pods": 1,
"last_scan_at": "2026-04-18T09:30:00Z"
}

get_pipeline_status​

Returns CI/CD run history, success rate, and failed stages.

Source service: cicd-gateway (port 8087) Endpoint: GET /api/v1/clusters/{cluster_id}/runs/summary

Parameters:

NameTypeRequiredDescription
cluster_idstringNoFilter to a specific cluster
limitintegerNoNumber of recent runs (default: 20)

Example response:

{
"success_rate_pct": 92,
"avg_duration_secs": 187,
"total_runs_24h": 38,
"failed_runs": [
{
"id": "run-abc123",
"pipeline": "payments-api",
"branch": "main",
"failed_stage": "integration-tests",
"started_at": "2026-04-18T08:15:00Z"
}
]
}

get_cluster_list​

Returns all registered clusters with health score, node count, and resource data.

Source service: kubeopera-api (port 8080) Endpoint: GET /api/v1/clusters

Parameters: None

Example response:

{
"clusters": [
{
"id": "prod-us-east",
"name": "Production US East",
"provider": "aws",
"region": "us-east-1",
"node_count": 6,
"health_score": 87,
"status": "running"
}
],
"total": 3
}

get_anomaly_events​

Returns anomaly events with Z-score, severity, and auto-remediation status.

Source service: anomaly-detector (port 8088) Endpoint: GET /api/v1/anomalies

Parameters:

NameTypeRequiredDescription
cluster_idstringNoFilter by cluster
severitystringNolow, medium, high, or critical
limitintegerNoNumber of results (default: 20)

Example response:

{
"anomalies": [
{
"id": "evt-7f3d",
"cluster_id": "prod-us-east",
"metric": "cpu_usage_pct",
"value": 94.2,
"baseline": 45.1,
"z_score": 3.8,
"severity": "critical",
"namespace": "production",
"resource": "payments-api",
"healing_id": "heal-abc",
"acknowledged": false,
"detected_at": "2026-04-18T09:55:00Z"
}
]
}

get_scaling_forecasts​

Returns Holt-Winters load forecasts and proactive scaling recommendations.

Source service: predictive-scaler (port 8089) Endpoint: GET /api/v1/scaling-decisions and GET /api/v1/forecasts

Parameters:

NameTypeRequiredDescription
cluster_idstringNoFilter by cluster
statusstringNopending, approved, applied, or rejected

Example response:

{
"forecasts": [
{
"workload": "payments-api",
"namespace": "production",
"metric": "cpu_usage_pct",
"horizon_mins": 60,
"predicted_peak": 87.4,
"confidence": 0.91
}
],
"scaling_decisions": [
{
"id": "dec-9a2f",
"workload": "payments-api",
"current_replicas": 3,
"recommended_replicas": 5,
"reason": "CPU forecast exceeds 80% in 45 minutes",
"status": "pending"
}
]
}

get_incidents​

Returns active and recent incidents with MTTR and runbook status.

Source service: incident-manager (port 8090) Endpoint: GET /api/v1/incidents and GET /api/v1/incidents/metrics

Parameters:

NameTypeRequiredDescription
cluster_idstringNoFilter by cluster
statusstringNoopen, investigating, mitigating, or resolved

Example response:

{
"incidents": [
{
"id": "inc-f4a1",
"title": "High CPU usage on payments-api",
"severity": "high",
"status": "investigating",
"cluster_id": "prod-us-east",
"namespace": "production",
"mttr_secs": null,
"runbook_id": "rb-cpu-scale",
"opened_at": "2026-04-18T09:56:00Z"
}
],
"metrics": {
"open_count": 2,
"avg_mttr_secs": 840,
"sla_breach_pct": 0
}
}

get_multi_cluster_overview​

Returns aggregated health, cost, and incident count across all registered clusters.

Source service: kubeopera-api (port 8080) + k8s-monitor (port 8085) — parallel fetch Endpoint: Multiple

Parameters: None

Example response:

{
"total_clusters": 3,
"total_nodes": 18,
"avg_health_score": 83,
"total_cost_monthly_usd": 14200.00,
"open_incidents": 2,
"clusters": [
{
"id": "prod-us-east",
"health_score": 87,
"node_count": 6,
"cost_monthly_usd": 4820.50,
"open_incidents": 1
}
]
}

get_agent_telemetry​

Returns recent telemetry snapshots from the observability agent.

Source service: observability-agent-srv (port 8092) Endpoint: GET /api/v1/telemetry

Parameters:

NameTypeRequiredDescription
cluster_idstringNoFilter by cluster
limitintegerNoNumber of snapshots (default: 10)

get_analysis_results​

Returns statistical analysis results with risk scores and anomaly decisions.

Source service: analysis-agent-srv (port 8093) Endpoint: GET /api/v1/results

Parameters:

NameTypeRequiredDescription
cluster_idstringNoFilter by cluster
limitintegerNoNumber of results (default: 10)

get_action_log​

Returns the automated Kubernetes action log with outcomes.

Source service: action-agent-srv (port 8094) Endpoint: GET /api/v1/actions

Parameters:

NameTypeRequiredDescription
cluster_idstringNoFilter by cluster
limitintegerNoNumber of actions (default: 20)

get_feedback_outcomes​

Returns feedback signals showing whether automated actions improved cluster state.

Source service: feedback-agent-srv (port 8095) Endpoint: GET /api/v1/feedback

Parameters:

NameTypeRequiredDescription
cluster_idstringNoFilter by cluster
limitintegerNoNumber of signals (default: 20)

get_recommendations​

Returns AI-generated recommendations from the recommendation agent.

Source service: recommendation-agent-srv (port 8096) Endpoint: GET /api/v1/recommendations

Parameters:

NameTypeRequiredDescription
cluster_idstringNoFilter by cluster
limitintegerNoNumber of recommendations (default: 10)

Example response:

{
"recommendations": [
{
"id": "rec-b3c2",
"cluster_id": "prod-us-east",
"category": "performance",
"priority": "high",
"title": "Scale payments-api deployment",
"description": "CPU utilisation exceeded 85% for 4 consecutive cycles (Z-score 3.1).",
"impact": "Request latency will increase. Risk of OOMKill under current load.",
"action": "Increase replicas from 3 to 5. Review HPA maxReplicas (currently 4).",
"risk_score": 72.5,
"generated_at": "2026-04-18T09:58:00Z"
}
]
}


get_alert_rules​

Returns anomaly alert rules with metric, threshold, severity, and configured auto-heal action.

Source service: anomaly-detector (port 8088) Endpoint: GET /api/v1/alert-rules

Parameters:

NameTypeRequiredDescription
cluster_idstringNoFilter by cluster

execute_runbook​

Triggers execution of an automated remediation runbook against an incident. Execution is asynchronous — use get_runbook_execution to track progress.

Source service: incident-manager (port 8090) Endpoint: POST /api/v1/incidents/{id}/runbooks/{rbId}/execute

Parameters:

NameTypeRequiredDescription
incident_idstringYesIncident to execute against
runbook_idstringYesRunbook to run
started_bystringNoOperator identifier (default: mcp-agent)

Returns: { id, status: "pending", incident_id, runbook_id, started_by, started_at }


get_runbook_execution​

Returns the current status and per-step results of a runbook execution.

Source service: incident-manager (port 8090) Endpoint: GET /api/v1/executions/{execId}

Parameters:

NameTypeRequiredDescription
execution_idstringYesExecution ID from execute_runbook

generate_post_incident_report​

Generates an AI-powered post-incident report for a resolved incident using Claude Haiku. Returns structured summary, root_cause, impact, and action_items.

Source service: incident-manager (port 8090) Endpoint: POST /api/v1/incidents/{id}/pir

Parameters:

NameTypeRequiredDescription
incident_idstringYesIncident to generate report for

get_node_pools​

Lists Karpenter NodePools with resource limits, consolidation policy, and node/claim counts.

Source service: nodes-manager (port 8098) Endpoint: GET /api/v1/nodepools

Parameters: None


get_node_claims​

Lists NodeClaims for a Karpenter NodePool showing phase (Pending/Bound/Disrupted), instance type, and zone.

Source service: nodes-manager (port 8098) Endpoint: GET /api/v1/nodepools/{name}/nodeclaims

Parameters:

NameTypeRequiredDescription
pool_namestringYesNodePool name

get_scheduling_decisions​

Returns the AI scheduling decision log with workload class (latency_sensitive/batch/gpu/stateful/stateless), 5-dimension score breakdown, and placement rationale.

Source service: nodes-manager (port 8098) Endpoint: GET /api/v1/decisions

Parameters:

NameTypeRequiredDescription
cluster_idstringNoFilter by cluster
limitintegerNoMax results (default: 20)

get_spot_market​

Returns spot instance pricing, savings vs on-demand, and interruption frequency by zone and instance family.

Source service: nodes-manager (port 8098) Endpoint: GET /api/v1/spot/market

Parameters:

NameTypeRequiredDescription
regionstringNoAWS region (e.g. us-east-1)

Example response:

[
{
"instance_type": "m5.2xlarge",
"zone": "us-east-1a",
"current_price": 0.1212,
"on_demand_price": 0.384,
"savings_pct": 68.4,
"interruption_freq": "low",
"available": true
}
]

get_consolidation_plan​

Returns the node consolidation plan showing which nodes can be safely drained, projected cost savings, and PDB impact analysis.

Source service: nodes-manager (port 8098) Endpoint: GET /api/v1/optimize/plan

Parameters: None


get_workload_placement​

Returns all workloads with their scheduled node, AI-assigned workload class, placement score (0–1 across 5 dimensions), and rationale.

Source service: nodes-manager (port 8098) Endpoint: GET /api/v1/workloads/placement

Parameters: None


Write Action Tools​

These tools trigger Kubernetes mutations via action-agent-srv. All are marked ReadOnlyHint: false.


drain_node_action​

Gracefully drains a Kubernetes node: cordons it then evicts all non-DaemonSet pods with a 30-second grace period.

Source service: action-agent-srv (port 8094) Endpoint: POST /api/v1/actions/drain

Parameters:

NameTypeRequiredDescription
node_namestringYesKubernetes node name to drain

rollback_deployment​

Triggers a rolling restart of a Kubernetes Deployment by patching the restartedAt annotation.

Source service: action-agent-srv (port 8094) Endpoint: POST /api/v1/actions/rollback

Parameters:

NameTypeRequiredDescription
namespacestringYesKubernetes namespace
deployment_namestringYesDeployment to roll back

Continuous Optimisation Tools (kubeopera-ai)​

These tools call kubeopera-ai (port 8104) to access the continuous monitoring agent's data and trigger AI analysis.


get_optimization_status​

Returns the latest AI-generated cluster health assessment from the kubeopera-ai monitoring agent: health score, status, severity-ranked issues with fix commands, last resource optimisation recommendations, and service uptime.

Source service: kubeopera-ai (port 8104) Endpoint: GET /api/v1/status

Parameters: None

Example response:

{
"last_health": {
"health_score": 68,
"status": "warning",
"summary": "Two pods in CrashLoopBackOff, memory pressure on worker-3",
"issues": [
{
"severity": "critical",
"component": "payments-api",
"description": "3 pods restarting every 2 minutes",
"recommendation": "Increase memory limits or fix OOM leak",
"command": "kubectl describe pod -n production -l app=payments-api"
}
],
"generated_at": "2026-04-25T10:00:00Z"
},
"namespace": "kubeopera-prod",
"uptime": "4h32m"
}

get_load_test_analysis​

Triggers a Claude-powered load test analysis via kubeopera-ai. Reads test output and resource snapshots (before and after) from the load-generator pod, then produces a structured Markdown report with p95/p99 latency assessment, bottleneck identification, and a numbered remediation action list.

Source service: kubeopera-ai (port 8104) Endpoint: POST /api/v1/analyze/load-test

Parameters:

NameTypeRequiredDescription
results_pathstringNoPath to results directory (uses latest if omitted)

Thresholds evaluated: p95 < 500ms, p99 < 1000ms, error rate < 1%, resource usage 60–80% of limits at peak.


Agent Launcher Tools​

These tools trigger a reasoning agent run in agent-runtime (port 8111) and return a run_id. The agent then autonomously calls read-only tools to investigate, reason, and produce a structured response.

The returned run_id can be used to stream agent reasoning via the KubeOpera UI at /agents/runs/{run_id}.


run_sre_agent​

Triggers a full cluster investigation by the SRE Orchestrator agent (Claude Sonnet 4.6 with interleaved thinking). Has access to all 17 read-only tools.

Parameters:

NameTypeRequiredDescription
cluster_idstringYesCluster to investigate
promptstringYesInvestigation goal or question

Example:

run_sre_agent(
cluster_id: "prod-us-east",
prompt: "Investigate the high CPU alert on payments-api and recommend immediate action"
)

Returns:

{
"run_id": "run-abc123",
"status": "running",
"stream_url": "http://localhost:8111/api/v1/runs/run-abc123/stream"
}

run_security_agent​

Triggers a security posture analysis by the Security Auditor agent (Claude Haiku 4.5). Uses 5 security-focused tools: get_security_posture, get_cluster_list, get_anomaly_events, get_incidents, get_agent_telemetry.

Parameters:

NameTypeRequiredDescription
cluster_idstringYesCluster to audit
promptstringYesSecurity question or audit scope

run_cost_agent​

Triggers a cost analysis by the Cost Optimizer agent (Claude Haiku 4.5). Uses 5 cost-focused tools: get_cluster_cost, get_optimization_report, get_scaling_forecasts, get_cluster_list, get_multi_cluster_overview.

Parameters:

NameTypeRequiredDescription
cluster_idstringYesCluster to optimize
promptstringYesCost analysis question

run_incident_agent​

Triggers an incident triage by the Incident Responder agent (Claude Sonnet 4.6 with interleaved thinking). Can investigate, execute runbooks, and generate post-incident reports autonomously.

Parameters:

NameTypeRequiredDescription
cluster_idstringYesCluster with the incident
promptstringYesIncident description or triage question

run_node_ops_agent​

Triggers a node pool analysis by the NodeOps agent (Claude Sonnet 4.6 with interleaved thinking). Analyses node pool efficiency, scheduling decisions, spot market savings, and consolidation opportunities.

Parameters:

NameTypeRequiredDescription
cluster_idstringYesCluster to analyse
promptstringYesInvestigation goal (e.g. "Identify consolidation opportunities and spot savings")

run_load_test_agent​

Triggers a load test performance analysis by the Load Test Analyst agent (Claude Sonnet 4.6). Fetches kubeopera-ai optimisation status as a baseline, triggers full load test analysis, and cross-references with live cluster metrics to produce a prioritised performance report.

Parameters:

NameTypeRequiredDescription
cluster_idstringYesCluster to analyse
promptstringYesAnalysis goal (e.g. "Analyse the latest load test and identify scaling bottlenecks")

App Advisor Tools​

These are tenant-scoped rather than cluster-wide — each call is bound to one specific app, and the service enforces that a caller can only ever reach an app belonging to their own tenant.


get_app_profile​

Fetches a discovered domain profile for one app — what kind of workload it is, its dependencies, and its observed behavior.

Source service: app-advisor-srv (port 8105) Endpoint: GET /api/v1/apps/{app_id}/profile

Parameters:

NameTypeRequiredDescription
app_idstringYesThe app to profile

get_app_advice​

Fetches SRE-style advice for one app — configuration and operational recommendations grounded in its discovered profile.

Source service: app-advisor-srv (port 8105) Endpoint: GET /api/v1/apps/{app_id}/advice

Parameters:

NameTypeRequiredDescription
app_idstringYesThe app to advise on

run_app_advisor_agent​

Triggers a full App Advisor investigation for one app, combining its profile and advice into a single reasoning pass.

Parameters:

NameTypeRequiredDescription
app_idstringYesThe app to investigate
promptstringYesInvestigation goal

Observability Gap-Fill Tools​

These fill in signal types the original read-only tool set didn't cover — logs, SLOs, APM metrics, traces, and root-cause analysis each have their own dedicated backend service rather than being folded into get_cluster_health or get_pod_metrics.


get_logs​

Fetches recent logs for a pod or deployment.

Source service: log-gateway (port 3100)

Parameters:

NameTypeRequiredDescription
namespacestringYesNamespace
pod_name | deployment_namestringOne of the twoWhat to fetch logs for

get_slo_status​

Fetches SLO compliance and error-budget burn for a service.

Source service: slo-manager (port 9090)

Parameters:

NameTypeRequiredDescription
servicestringYesThe service to check

get_apm_metrics​

Fetches application performance metrics (latency, throughput, error rate) for a service.

Source service: apm-gateway (port 9090)

Parameters:

NameTypeRequiredDescription
servicestringYesThe service to query

get_traces​

Fetches distributed traces for a service.

Source service: tracing-gateway (port 16686)

Parameters:

NameTypeRequiredDescription
servicestringYesThe service to fetch traces for

get_rca_analysis​

Fetches root-cause analysis results for an incident or anomaly.

Source service: rca-engine (port 8116)

Parameters:

NameTypeRequiredDescription
incident_idstringYesThe incident to analyse

GitOps Tools​

These talk directly to Argo CD and a Git-hosting API rather than to another KubeOpera backend service — see Configuration for the separate env vars (ARGOCD_API_URL, ARGOCD_TOKEN, FLUX_WEBHOOK_URL, FLUX_WEBHOOK_TOKEN, GITOPS_REPO_API_URL, GITOPS_REPO_TOKEN) they read.


sync_argocd_app​

Triggers an Argo CD application sync.

Parameters:

NameTypeRequiredDescription
app_namestringYesThe Argo CD application to sync

get_argocd_app_status​

Fetches an Argo CD application's current sync and health status.

Parameters:

NameTypeRequiredDescription
app_namestringYesThe Argo CD application to check

push_gitops_patch​

Commits a patch to a GitOps repository.

Parameters:

NameTypeRequiredDescription
repositorystringYesTarget repository
pathstringYesFile path within the repository
patchstringYesThe patch content
commit_messagestringYesCommit message

trigger_flux_reconcile​

Forces a Flux Kustomization to reconcile immediately, rather than waiting for its next scheduled interval.

Parameters:

NameTypeRequiredDescription
kustomization_namestringYesThe Kustomization to reconcile
namespacestringYesIts namespace

Tool Availability by Agent​

The App Advisor, observability gap-fill, and GitOps tools above were added after this matrix was last updated, and their exact per-agent availability hasn't been re-confirmed since — treat their absence below as "not yet documented," not as "unavailable to every agent."

ToolSRESecurityCostIncidentNodeOpsLoad Test
get_cluster_healthYes——YesYesYes
get_cluster_costYes—Yes———
get_optimization_reportYes—Yes———
get_pod_metricsYes——Yes—Yes
get_node_metricsYes——Yes—Yes
get_security_postureYesYes————
get_pipeline_statusYes—————
get_cluster_listYesYesYes———
get_anomaly_eventsYesYes—YesYesYes
get_alert_rulesYes—————
get_scaling_forecastsYes—Yes——Yes
get_incidentsYesYes—YesYes—
execute_runbookYes——Yes——
get_runbook_executionYes——Yes——
generate_post_incident_reportYes——Yes——
get_multi_cluster_overviewYes—Yes———
get_node_poolsYes———Yes—
get_node_claimsYes———Yes—
get_scheduling_decisionsYes———Yes—
get_spot_marketYes———Yes—
get_consolidation_planYes———Yes—
get_workload_placementYes———Yes—
drain_node_actionYes——YesYes—
rollback_deploymentYes——Yes——
get_agent_telemetryYesYes—Yes——
get_analysis_resultsYes—————
get_action_logYes——Yes——
get_feedback_outcomesYes—————
get_recommendationsYes——Yes——
get_optimization_statusYes————Yes
get_load_test_analysisYes————Yes