Skip to content

Observability

Everything is under Observability in the site, per controller and time range.

  • Reconciler: reconciles and errors per second, reconcile time (95th percentile) and work queue depth, from controller-runtime’s metrics.
  • Pod resources: CPU, memory and network of the controller.

ctrlplane scrapes your controller at :8080/metrics every 15 seconds and keeps 15 days. If a controller shows no metrics, check it serves them there over HTTP; see Controllers.

Your controllers’ output (stdout and stderr), searchable, with a live tail. Search uses LogsQL, for example error or "failed to sync". Logs on a controller shows its current run, and the run before its last restart.

What your controllers (and your API server) record about objects, newest first, from your namespace and from default (events about cluster-scoped objects). Warnings also show on the overview under Needs attention. Events expire after an hour.

Your API server serves this data, pinned to your control plane, to anything with a token:

Data source URL
Prometheus https://<id>.api.ctrlplane.run/platform/metrics
VictoriaLogs https://<id>.api.ctrlplane.run/platform/logs
  1. Create a service account and a long-lived token for it:

    Terminal window
    kubectl create serviceaccount grafana
    kubectl create token grafana --duration=8760h
  2. In each data source, add the header Authorization: Bearer <token>.

  3. Under TLS, add your control plane’s CA certificate: the certificate-authority-data from your kubeconfig, base64-decoded:

    Terminal window
    kubectl config view --raw --minify \
    -o jsonpath='{.clusters[0].cluster.certificate-authority-data}' | base64 -d

That service account may read your control plane’s metrics and logs and nothing else, so the token is only as powerful as a dashboard needs.