====== paddles ====== ===== Summary ===== paddles is a very simple JSON-based API used to report teuthology test results. See https://github.com/ceph/paddles It runs in OpenShift in the ''paddles'' namespace. The database is a Crunchy PGO postgres cluster in the ''paddles-db'' namespace; the app connects through the ''paddles-db-pgbouncer'' service. See https://github.com/ceph/sepia-openshift/tree/main/paddles paddles is **not** publicly exposed. ''paddles.front.sepia.ceph.com'' is VPN-only. External read-only consumers get ''%%https://pulpito.ceph.com/_paddles/runs/...%%'' (GET-only, IP-allowlisted nginx proxy on soko01). ===== Backups ===== The 'paddles' db is backed up daily by the [[services:backups#backupsh|backup.sh]] script on gitbuilder-archive. Backups are located in ''gitbuilder-archive:/home/backup/paddles.front.sepia.ceph.com-psql/paddles'' ===== Admin Tasks ===== ==== Updating/Fixing Zombie Jobs ==== For jobs that indicate they're running but aren't, ''expire_jobs'' can be used. The following example would expire any **queued** jobs 14 days old or older and any **running** jobs that haven't been updated in 30 minutes. oc -n paddles exec deploy/paddles -- pecan expire_jobs config.py -q 14 -r 30 ==== Adding testnodes to the inventory/DB ==== From your workstation or anywhere with a valid ''teuthology.yaml'', cd ~/src/teuthology source ./virtualenv/bin/activate # Edit docs/_static/create_nodes.py # (paddles_url, machine_type, lab_domain, and machine_index_range) # These can all be found in teuthology.yaml on a teuthology host python docs/_static/create_nodes.py ==== Upgrading Paddles ==== Merges to paddles ''main'' deploy themselves: GitHub Actions pushes ''quay.io/ceph-infra/paddles:main'' (~10 min), the ImageStream polls Quay every 5 minutes, and the trigger annotation on the Deployment rolls out the new digest. Alembic migrations are applied by the new pod at startup - see the paddles-db README for the migration runbook. To check what's deployed or force a roll: oc -n paddles get istag paddles:main oc -n paddles rollout restart deploy/paddles ==== Restarting ==== oc -n paddles rollout restart deploy/paddles ==== Getting a psql shell ==== Replicas are read-only, so find the leader first: oc -n paddles-db exec -it $(oc -n paddles-db get pod -l postgres-operator.crunchydata.com/role=master -o name) -c database -- psql paddles # or: patronictl list from any inst pod ===== Troubleshooting ===== ==== Check the logs ==== oc -n paddles logs deploy/paddles # or who changed something in the namespace (console Logs view, audit tenant): {log_type="audit"} |= "paddles" | json | objectRef_namespace="paddles" See [[services:loki]]. ==== Pods getting killed every few minutes (exit 137) ==== Almost always the ''/jobs/?description='' query. It does a ''LIKE %%'%X%'%%'' over the 8.4M-row jobs table and postgres walks ''ix_jobs_posted'' instead of the trigram index (3-5s per call). Enough of them saturates the gunicorn workers, liveness fails, and the pod gets killed. Find the caller (usually a scraper hitting pulpito-ng's job history pages) and block it at soko01. ==== Config changes don't take effect ==== The image's ''container_start.sh'' always uses the **sample config baked into the image**. The ''paddles-config'' configmap mounted at ''/etc/paddles/config.py'' is ignored unless the pod ''command:'' is overridden. Don't burn time editing the configmap and wondering why nothing changed. ==== Pods Ready but API dead ==== The readiness/startup probes only test DB connectivity, not HTTP. Check the route from a VPN host before trusting pod status. ==== Dumping the DB ==== Run pg_dump against the **primary**, not a replica - replicas cancel long dumps. Budget ~4m to dump, ~18m to restore.