This is an old revision of the document!
Table of Contents
Central Logging (Loki)
Summary
All lab logs land in one queryable store: a LokiStack in the openshift-logging namespace on the OpenShift cluster. See https://github.com/ceph/sepia-openshift/tree/main/openshift-logging for the full deployment runbook.
Two tenants:
- infrastructure - journald + auditd from ~205 baremetal lab hosts (pushed by Grafana Alloy, deployed via the
alloyrole in ceph-cm-ansible) plus the OCP node journals (collected in-cluster). 30d retention. - audit - the full OCP cluster audit trail including request bodies for writes. Answers “who changed what.” 90d retention.
OCP infra-compute/infra-storage nodes and testnodes are excluded from the Alloy play. The OCP nodes are already covered by the in-cluster collector.
Reading logs
Grafana
https://ci-metrics.ceph.com has a Loki Infrastructure datasource. Any dashboard built on it must go in the Private folder (spec.folderRef: private) or it's world-readable.
OpenShift console
Observe → Logs in the console. This is the easiest way to query the audit tenant.
Useful queries
Logs from a baremetal host
{host="soko01.front.sepia.ceph.com"} |= "nginx"
{job="auditd", host="soko03.front.sepia.ceph.com"}
Who changed something in OpenShift
{log_type="audit"} | json | objectRef_namespace="paddles" |= "requestObject"
Who shelled into a pod
{log_type="audit"} |= "pods/exec"
Troubleshooting
One host stopped shipping
Check the agent on the host. Permission-denied on /var/log/audit/audit.log means the file-permission tasks in the role didn't take.
journalctl -u alloy
ALL hosts stopped shipping at once
The cluster's ingress cert rotated and the agents no longer trust the gateway. Re-extract the CA and re-run the play.
oc get cm -n openshift-config-managed default-ingress-cert \
-o jsonpath='{.data.ca-bundle\.crt}'
# -> ceph-sepia-secrets/ansible/files/lokistack-ingress-ca.crt, then
ansible-playbook alloy.yml
Audit tenant is empty
The collect-audit-logs ClusterRoleBinding is missing. The CLF status stays green with no error, so this is easy to miss.
oc apply -f rbac-collector.yaml
Grafana/agent tokens stopped working
Don't use oc create token - it issues a 1-hour token. Use the long-lived secrets created by the rbac yamls:
oc -n openshift-logging get secret sepia-hosts-token -o jsonpath='{.data.token}' | base64 -d
oc -n openshift-logging get secret grafana-reader-token -o jsonpath='{.data.token}' | base64 -d
