User Tools

Site Tools


services:paddles

This is an old revision of the document!


paddles

Summary

paddles is a very simple JSON-based API used to report teuthology test results. See https://github.com/ceph/paddles

It runs in OpenShift in the paddles namespace. The database is a Crunchy PGO postgres cluster in the paddles-db namespace; the app connects through the paddles-db-pgbouncer service. See https://github.com/ceph/sepia-openshift/tree/main/paddles

paddles is not publicly exposed. paddles.front.sepia.ceph.com is VPN-only. External read-only consumers get https://pulpito.ceph.com/_paddles/runs/... (GET-only, IP-allowlisted nginx proxy on soko01).

Backups

The 'paddles' db is backed up daily by the backup.sh script on gitbuilder-archive.

Backups are located in gitbuilder-archive:/home/backup/paddles.front.sepia.ceph.com-psql/paddles

Admin Tasks

Updating/Fixing Zombie Jobs

For jobs that indicate they're running but aren't, expire_jobs can be used.

The following example would expire any queued jobs 14 days old or older and any running jobs that haven't been updated in 30 minutes.

oc exec -n paddles $(oc get pods -n paddles -o name | grep -v build |  head -1) -- pecan expire_jobs config.py -q 14 -r 30

Adding testnodes to the inventory/DB

From your workstation or anywhere with a valid teuthology.yaml,

cd ~/src/teuthology
source ./virtualenv/bin/activate

# Edit docs/_static/create_nodes.py
# (paddles_url, machine_type, lab_domain, and machine_index_range)
# These can all be found in teuthology.yaml on a teuthology host

python docs/_static/create_nodes.py

Upgrade Paddles

The paddles API / Web UI services live in Openshift in the paddles namespace. There is a cronjob that checks quay.io/ceph/paddles for a new image every 5 minutes. The ImageStream will automatically deploy any new images. See https://github.com/ceph/sepia-openshift/tree/main/paddles

Troubleshooting

Check the logs

oc -n paddles logs deploy/paddles
# or who changed something in the namespace (console Logs view, audit tenant):
{log_type="audit"} |= "paddles" | json | objectRef_namespace="paddles"

See loki.

Pods getting killed every few minutes (exit 137)

Almost always the /jobs/?description= query. It does a LIKE '%X%' over the 8.4M-row jobs table and postgres walks ix_jobs_posted instead of the trigram index (3-5s per call). Enough of them saturates the gunicorn workers, liveness fails, and the pod gets killed. Find the caller (usually a scraper hitting pulpito-ng's job history pages) and block it at soko01.

Config changes don't take effect

The image's container_start.sh always uses the sample config baked into the image. The paddles-config configmap mounted at /etc/paddles/config.py is ignored unless the pod command: is overridden. Don't burn time editing the configmap and wondering why nothing changed.

Pods Ready but API dead

The readiness/startup probes only test DB connectivity, not HTTP. Check the route from a VPN host before trusting pod status.

Dumping the DB

Run pg_dump against the primary, not a replica - replicas cancel long dumps. Budget ~4m to dump, ~18m to restore.

services/paddles.1790966860.txt.gz · Last modified: by djgalloway