User Tools

Site Tools


services:loki

Central Logging (Loki)

Summary

All lab logs land in one queryable store: a LokiStack in the openshift-logging namespace on the OpenShift cluster. See https://github.com/ceph/sepia-openshift/tree/main/openshift-logging for the full deployment runbook.

Two tenants:

  • infrastructure - journald + auditd from ~205 baremetal lab hosts (pushed by Grafana Alloy, deployed via the alloy role in ceph-cm-ansible) plus the OCP node journals (collected in-cluster). 30d retention.
  • audit - the full OCP cluster audit trail including request bodies for writes. Answers “who changed what.” 90d retention.

OCP infra-compute/infra-storage nodes and testnodes are excluded from the Alloy play. The OCP nodes are already covered by the in-cluster collector.

Reading logs

Grafana

https://ci-metrics.ceph.com has a Loki Infrastructure datasource. Any dashboard built on it must go in the Private folder (spec.folderRef: private) or it's world-readable.

OpenShift console

Observe -> Logs in the console. This is the easiest way to query the audit tenant.

Useful queries

Logs from a baremetal host

{host="soko01.front.sepia.ceph.com"} |= "nginx"
{job="auditd", host="soko03.front.sepia.ceph.com"}

Who changed something in OpenShift

{log_type="audit"} | json | objectRef_namespace="paddles" |= "requestObject"

Who shelled into a pod

{log_type="audit"} |= "pods/exec"

Troubleshooting

One host stopped shipping logs

Check the agent on the host. Permission-denied on /var/log/audit/audit.log means the file-permission tasks in the role didn't take.

journalctl -u alloy

ALL hosts stopped shipping at once

The cluster's ingress cert rotated and the agents no longer trust the gateway. Re-extract the CA and re-run the play.

oc get cm -n openshift-config-managed default-ingress-cert \
  -o jsonpath='{.data.ca-bundle\.crt}'
# -> ceph-sepia-secrets/ansible/files/lokistack-ingress-ca.crt, then
ansible-playbook alloy.yml

Audit tenant is empty

The collect-audit-logs ClusterRoleBinding is missing. The CLF status stays green with no error, so this is easy to miss.

oc apply -f rbac-collector.yaml

Grafana/agent tokens stopped working

Don't use oc create token - it issues a 1-hour token. Use the long-lived secrets created by the rbac yamls:

oc -n openshift-logging get secret sepia-hosts-token -o jsonpath='{.data.token}' | base64 -d
oc -n openshift-logging get secret grafana-reader-token -o jsonpath='{.data.token}' | base64 -d
services/loki.txt · Last modified: by djgalloway