Auth Service
Repo: o-apps/auth-service · Port: 8081 · DB Schema: auth
auth-service is the identity provider for the entire platform: it issues and validates the JWTs every other backend service trusts, owns the user/role/permission model, and brokers OAuth2 logins. There is no NextAuth.js or any other frontend-side auth library — the Next.js frontend proxies auth calls straight through to this service.
Key Capabilities
- Email/password auth — registration, login, email verification, password reset, MFA (TOTP-style enable/disable/verify)
- JWT issuance and rotation — HMAC-signed access + refresh tokens; accepts a trusted-list of secrets so a rotation can accept both the old and new signing key during cutover
- Token introspection —
GET /api/v1/auth/introspectlets any other service verify a bearer token's validity, claims, and expiry without sharing the signing secret directly - OAuth2 login — Google via a standard authorize/callback flow.
GitHubandKeycloakexist asOAuthProviderenum values in the domain model but have no working implementation inOAuthService— only"google"is handled; every other provider returns "unsupported OAuth provider" - RBAC — roles, permissions, and role-to-user assignment, enforced via
RequireRolemiddleware on admin routes - API keys — users can mint/list/revoke long-lived API keys as an alternative to session JWTs
- CloudSpace and app management — provisions per-tenant CloudSpaces and app registrations, with self-service onboarding routes alongside the admin-only management routes
- Rate limiting — login/register attempts are rate-limited via cache-service (not a direct Redis dependency — see below)
Shared JWT Trust Model
Every other backend service that requires auth (12 as of this writing, via pkg/authmiddleware) validates tokens against the same shared secret auth-service signs with (AUTH_JWT_ACCESS_SECRET, distributed as the shared-jwt-access-secret Kubernetes secret). None of them call back into auth-service to validate a token per request — they verify the signature locally. GET /api/v1/auth/introspect exists for callers that specifically need auth-service's own opinion on a token (e.g. the frontend, which doesn't hold the signing secret).
Rate Limiting
auth-service does not talk to Redis directly. It calls cache-service (a thin Redis-or-in-memory wrapper other services also use) over HTTP for its IncrementWithExpiry sliding-window rate limiter — login and registration attempts are capped per identifier within a rolling window. If CACHE_SERVICE_BASE_URL is unset, it falls back to an in-memory cache: rate limits then don't survive a restart or apply across replicas, which is fine for local dev but not for a multi-replica production deployment.
AI provider credential resolution
auth-service is also where every AI-calling service in the platform (agent-runtime, recommendation-agent-srv, Optimizer, app-advisor-srv, and the frontend's own AI chat) resolves its Anthropic API key — there is no single platform-wide key set as an environment variable anywhere anymore. The resolution policy is deliberately simple and the same for every caller: if the calling tenant has configured their own key, use it, unmetered; otherwise, if a platform/host key is configured and sharing is enabled, use that one, metered against the tenant's configured quota; otherwise, tell the caller plainly that no key is configured rather than failing with a raw upstream error. This one policy covers two different deployment shapes without needing separate code paths for each — an enterprise running KubeOpera for its own internal teams typically configures one shared host key with per-team quotas and lets individual teams override it, while an independent SaaS customer's deployment simply never enables host-key sharing, so every tenant is required to bring their own.
The tenant-facing self-service endpoint lets a tenant set, view the status of (never the plaintext again, only the last four characters and whether one is configured), or remove their own key. A separate super_admin-only endpoint manages the host key, its sharing toggle, and its quota. Every AI-calling service resolves its key through one internal, API-key-authenticated endpoint, with a short-lived cache so a resolution doesn't require a network round trip on every single AI call.
| Method | Path | Description |
|---|---|---|
GET | PUT | DELETE | /api/v1/ai-credentials | Tenant self-service — set, check, or remove their own key |
GET | PUT | /api/v1/admin/ai-credentials/host | super_admin only — host key, sharing toggle, quota |
GET | /internal/ai-credentials/resolve | Internal — the resolution policy above, called by every AI-calling service |
POST | /internal/ai-credentials/usage | Internal — records usage against a tenant's quota when they drew on the shared host key |
Registry token issuance
auth-service also issues short-lived, narrowly-scoped access tokens for KubeOpera's self-hosted container registry, implementing the standard Docker Registry v2 token-auth protocol so the registry itself needs no knowledge of tenants at all — it only ever sees a signed token and trusts whatever scope that token carries. The scoping is enforced by the issuer, not by the credential's secrecy alone: a request authenticated with one tenant's push or pull credential is only ever granted access to that tenant's own path in the registry (tenant-<id>/*), regardless of what scope the request itself asks for. This is what makes it safe for a Kaniko build running on shared, host-cluster infrastructure to push an image without ever being able to read or overwrite another tenant's images.
| Method | Path | Description |
|---|---|---|
GET | /v2/registry-token | Docker Registry v2 token-auth endpoint — validated against the requesting credential's own tenant scope, not the scope the request asks for |
POST | /internal/registry-credentials/{tenantID} | Internal — idempotently provisions a tenant's push/pull registry credentials, returning the plaintext once |
Cloud provider credential storage
The newest of the three runtime credential stores, backing Cluster Management's Create New path — where cluster-provisioner gets the real AWS credentials it needs to actually run Terraform. Structurally the simplest of the three: super_admin-only, with no tenant dimension at all (provisioning a host cluster is inherently a platform-level operation, never something scoped to one customer), and no resolution policy to speak of — a caller always asks for one specific stored credential by ID rather than auth-service picking one on their behalf. Encrypted at rest via the same reversible CryptoService (AES-256-GCM) the AI-credential store above uses, keyed off MASTER_KEY; once stored, the plaintext is only ever returned again to the one internal caller resolving it immediately before a provisioning run, never back to the browser — the UI only ever sees a label and the stored value's last four characters.
A stored credential can be access_key (an AWS access key ID + secret) or assume_role (an IAM role ARN, optionally with an external ID). Only access_key is wired all the way through to a working provisioning run today — an assume_role credential can be stored, but nothing yet performs the actual sts:AssumeRole call a real Terraform apply or an EKS kubeconfig's exec plugin would need to use one.
| Method | Path | Description |
|---|---|---|
GET | POST | /api/v1/admin/cloud-credentials | super_admin only — list, or store a new credential |
DELETE | /api/v1/admin/cloud-credentials/{id} | super_admin only — remove a stored credential |
GET | /internal/cloud-credentials/{id}/resolve | Internal — returns one credential's decrypted payload, called by cluster-provisioner exactly once per provisioning run |
Super-admin bootstrap
auth-service's migrations seed the identity-free Super Admin role only — never an account. BootstrapSuperAdmin creates or updates the account itself, upserting by email: an existing account's username/password/active/verified state is updated and the Super Admin role (re-)assigned; a new email creates a platform-wide account with no tenant attached. It's a real API-driven action rather than a one-time migration step, so it's also the correct way to reset a lost super-admin password — safe to call again with the same email at any time.
Cluster Management calls this automatically once a new environment's apps have reconciled, when an admin email/username were supplied at cluster-creation time; the standalone Helm chart's post-install hook Job calls it the same way on first install; the manual GitOps path calls it by hand, once, after authapi first comes up.
| Method | Path | Description |
|---|---|---|
POST | /internal/bootstrap/super-admin | Internal — upserts the super-admin account by email; body {email, username, password} |
REST API
| Method | Path | Description |
|---|---|---|
POST | /api/v1/auth/register | Create an account |
POST | /api/v1/auth/login | Email/password login |
POST | /api/v1/auth/refresh | Exchange a refresh token for a new access token |
POST | /api/v1/auth/logout | Invalidate the current session |
POST | /api/v1/auth/send-verification-email / /verify-email | Email verification flow |
POST | /api/v1/auth/forgot-password / /reset-password | Password reset flow |
POST | /api/v1/auth/mfa/enable / /disable / /verify | TOTP MFA management |
GET | /api/v1/auth/introspect | Validate a bearer token, return its claims |
POST | /api/v1/apikeys | Create an API key |
GET | /api/v1/apikeys | List the caller's API keys |
DELETE | /api/v1/apikeys/{id} | Revoke an API key |
GET | POST | PUT | DELETE | /api/v1/users, /api/v1/roles | Admin-only user and role management (RequireRole("admin")) |
POST | GET | PUT | DELETE | /api/v1/cloudspaces, /api/v1/cloudspaces/{id}/apps | CloudSpace and app management (create is self-service; management is admin-gated) |
GET | /oauth/auth-url | Get the provider authorize URL |
GET | /oauth/{provider} | Redirect to provider login |
GET | /oauth/{provider}/callback | OAuth2 callback |
GET | /health | Health check |
Environment Variables
| Variable | Description |
|---|---|
DATABASE_URL | PostgreSQL connection (schema: auth) |
JWT_ACCESS_SECRET / JWT_REFRESH_SECRET | HMAC signing secrets — required in production, no insecure default is accepted outside an explicit local/dev profile |
MASTER_KEY | Encryption key for sensitive stored fields |
CACHE_SERVICE_BASE_URL | cache-service endpoint for rate limiting; falls back to in-memory if unset |
CACHE_SERVICE_API_KEY | Shared secret cache-service's API-key middleware validates against |
PORT | HTTP port (default: 8080 in code; 8081 in the live deployment) |