This is an old revision of the document!
Table of Contents
paddles
Summary
paddles is a very simple JSON-based API used to report teuthology test results. See https://github.com/ceph/paddles
It runs in OpenShift in the paddles namespace. The database is a Crunchy PGO postgres cluster in the paddles-db namespace; the app connects through the paddles-db-pgbouncer service. See https://github.com/ceph/sepia-openshift/tree/main/paddles
paddles is not publicly exposed. paddles.front.sepia.ceph.com is VPN-only. External read-only consumers get https://pulpito.ceph.com/_paddles/runs/... (GET-only, IP-allowlisted nginx proxy on soko01).
Backups
The 'paddles' db is backed up daily by the backup.sh script on gitbuilder-archive.
Backups are located in gitbuilder-archive:/home/backup/paddles.front.sepia.ceph.com-psql/paddles
Admin Tasks
Updating/Fixing Zombie Jobs
For jobs that indicate they're running but aren't, expire_jobs can be used.
The following example would expire any queued jobs 14 days old or older and any running jobs that haven't been updated in 30 minutes.
oc exec -n paddles $(oc get pods -n paddles -o name | grep -v build | head -1) -- pecan expire_jobs config.py -q 14 -r 30
Adding testnodes to the inventory/DB
From your workstation or anywhere with a valid teuthology.yaml,
cd ~/src/teuthology source ./virtualenv/bin/activate # Edit docs/_static/create_nodes.py # (paddles_url, machine_type, lab_domain, and machine_index_range) # These can all be found in teuthology.yaml on a teuthology host python docs/_static/create_nodes.py
Upgrade Paddles
The paddles API / Web UI services live in Openshift in the paddles namespace. There is a cronjob that checks quay.io/ceph/paddles for a new image every 5 minutes. The ImageStream will automatically deploy any new images. See https://github.com/ceph/sepia-openshift/tree/main/paddles
Troubleshooting
Check the logs
oc -n paddles logs deploy/paddles
# or who changed something in the namespace (console Logs view, audit tenant):
{log_type="audit"} |= "paddles" | json | objectRef_namespace="paddles"
See loki.
Pods getting killed every few minutes (exit 137)
Almost always the /jobs/?description= query. It does a LIKE '%X%' over the 8.4M-row jobs table and postgres walks ix_jobs_posted instead of the trigram index (3-5s per call). Enough of them saturates the gunicorn workers, liveness fails, and the pod gets killed. Find the caller (usually a scraper hitting pulpito-ng's job history pages) and block it at soko01.
Config changes don't take effect
The image's container_start.sh always uses the sample config baked into the image. The paddles-config configmap mounted at /etc/paddles/config.py is ignored unless the pod command: is overridden. Don't burn time editing the configmap and wondering why nothing changed.
Pods Ready but API dead
The readiness/startup probes only test DB connectivity, not HTTP. Check the route from a VPN host before trusting pod status.
Dumping the DB
Run pg_dump against the primary, not a replica - replicas cancel long dumps. Budget ~4m to dump, ~18m to restore.
